High resolution virtual staining of label-free tissue using diffusion models

The diffusion model-based approach effectively transforms lower-resolution auto-fluorescence microscopy images into high-resolution, digitally stained images, overcoming the limitations of traditional chemical staining methods, enhancing spatial resolution and fidelity, and reducing variance in biomedical imaging applications.

WO2026090345A1PCT designated stage Publication Date: 2026-04-30RGT UNIV OF CALIFORNIA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RGT UNIV OF CALIFORNIA
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing diffusion models generate high variance in image outputs for biomedical applications, particularly in clinical diagnosis, and traditional chemical staining methods are time-consuming and costly.

Method used

A diffusion model-based approach that uses a Brownian bridge process and attention-based U-Net for image-conditional inference, combined with mean and skip sampling strategies, to transform low-resolution images into high-resolution, virtually stained images, enhancing spatial resolution and reducing variance.

Benefits of technology

This method achieves high-resolution, accurate virtual staining without the need for chemical reagents, reducing processing time and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052148_30042026_PF_FP_ABST
    Figure US2025052148_30042026_PF_FP_ABST
Patent Text Reader

Abstract

A diffusion model-based high resolution virtual staining method is disclosed that uses a Brownian bridge process to enhance both the spatial resolution and fidelity of label-free virtual tissue staining. The method integrates novel sampling techniques into a diffusion model-based image inference process to significantly reduce the variance in the generated virtually stained images, resulting in more stable and accurate outputs. Blindly applied to lower-resolution auto-fluorescence images of label-free human lung tissue samples, the diffusion-based super-resolution virtual staining model consistently outperformed conventional approaches in resolution, structural similarity and perceptual accuracy. Diffusion-based super-resolved virtual tissue staining improves resolution and image quality as well as enhances the reliability of virtual staining without traditional chemical staining.
Need to check novelty before this filing date? Find Prior Art

Description

HIGH RESOLUTION VIRTUAL STAINING OF LABEL-FREE TISSUE USING DIFFUSION MODELSRelated Applications

[0001] This Application claims priority to U. S. Provisional Patent Application No. 63 / 712,348 filed on October 25, 2024 and U. S. Provisional Patent Application No.63 / 717,882 filed on November 7, 2024, which are hereby incorporated by reference in its entirety. Priority is claimed pursuant to 35 U. S. C. § 119 and any other applicable statute.Technical Field

[0002] The technical field generally relates to methods and systems used to image unstained (i.e., label-free) tissue. In particular, the technical field relates to microscopy methods and systems that utilize a diffusion model-based image inference process for digitally or virtually staining of images of unstained or unlabeled tissue. In particular, the technical field relates to a diffusion-based model that takes, in one embodiment, low-resolution auto-fluorescence images of label-free tissue and generates high resolution virtual stained images of the tissue that are substantially equivalent to histochemically-stained samples. Other embodiments use diffusion models to enhance the spatial resolution and digitally introduce cellular morphological contrast into mass spectrometry7images of label-free human tissue.Statement Regarding Federally SponsoredResearch and Development

[0003] This invention was made with government support under EB032840 awarded by the National Institutes of Health. The government has certain rights in the invention.Background

[0004] Generative Al models have achieved considerable advances over the last decade and wide applications in various fields. These models fostered the emergence of computational pathology7, showcasing unprecedented performance in image transformation, segmentation, and reconstruction. As one of the state-of-the-art generative techniques, diffusion models have shown a strong ability to approximate multi-modal distributions with the versatility7to be conditioned through multiple forms of guidance, including text andimage. There have been synergetic studies on applying diffusion models in computational pathology. For example, diffusion models have been demonstrated to generate photorealistic histopathology images, given guidance of segmentation masks, domain knowledge and tissue genomics. Researchers have also studied image translation between multiple histopathological image domains, a process termed stain transformation. This includes, for example, transforming a histopathological image stained with Hematoxylin and Eosin (H& E) into an immunohistochemistry (IHC)-stained image of the same tissue slice. In addition, diffusion models have been explored for histological image enhancement and segmentation to facilitate downstream analysis and diagnosis.

[0005] To better adapt to conditional image generation tasks, researchers have also exploited stochastic bridges connecting two image domains and demonstrated various conditional diffusion models. Among them, the Brownian bridge is one of the well-known and widely utilized stochastic processes that stems from the standard Brownian (diffusion) process and is conditioned on both the start and end states. Instead of using the standard Brownian motion in the common forward diffusion process that converges to white noise, the Brownian bridge diffusion model (BBDM) learns the mapping from the target image domain to the input (conditional) image domain via a Brownian bridge. BBDM has been reported to outperform standard diffusion models in various image restoration and translation applications. Nevertheless, all diffusion models inherently generate outputs with relatively high variance (from run to run) compared to some of the existing generative models. including e.g., conditional Generative Adversarial Networks (cGANs); such stochastic image variations for the same specimen raise concerns regarding their impact on biomedical image synthesis or reconstruction tasks, especially for potential uses in clinical diagnosis.Summary

[0006] Here, a diffusion model-based high resolution virtual staining (VS) model is disclosed that, in one embodiment, transforms lower-resolution auto-fluorescence (AF) microscopy images of label-free tissue samples into high resolution images, digitally matching the histochemically stained higher-resolution images of the same tissue samples without the need for traditional chemical staining. In one embodiment, the high-resolution images are super-resolved brightfield images as disclosed herein. This diffusion-based superresolved VS model significantly outperforms traditional VS methods that process the same lower-resolution AF images of label-free tissue samples, and it drastically reduces inference variance from the diffusion process, converging to stable and accurate image inference that matches the histochemically stained higher-resolution brightfield images of the same tissue samples. The approach is built on an image-conditional diffusion model leveraging the Brownian bridge process to effectively integrate the lower-resolution conditional image and the noise estimation from an attention-based U-Net that incorporates the time step information to reconstruct a higher-resolution histological image - performing two tasks at the same time: (z) spatial resolution enhancement and (zz) virtual staining of label-free tissue. The term “super-resolution” as used herein should not be confused with nanoscopy techniques that beat the diffraction limit of light. Herein, super-resolution or super-resolved images refer to the capability of the VS model in synthesizing brightfield equivalent stained images with higher spatial resolution compared to the input label-free images, hence increasing the space-bandwidth product and the effective number of useful pixels in the VS images - all within the diffraction limit of light. Earlier uses of the term super-resolution are also aligned with this terminology. It should be appreciated that the methods and systems disclosed herein may generate high resolution, virtually stained images of a label-free test sample without the output images being “super-resolved.” That is to say, the resolution of the autofluorescence images of label-free test samples is lower than the high-resolution microscopic images of the same label-free test samples that have been chemically stained. High resolution microscopic images, in this context, refers to microscopic images of a sample-like tissue with sub-micron level resolution sufficient for pathological examination by an expert diagnostician.

[0007] In comparison with other VS models, the conditional diffusion model generates better V S images with higher resolution and image fidelity' matching the ground truth histochemically stained brightfield images. Besides, to mitigate the inherent high variance of diffusion models for VS applications in pathology, novel sampling process engineering techniques are introduced, i.e., the mean and skip sampling strategies as illustrated in FIGS. IB, 21, 2J. Based on the analysis of the posterior sampling variance over time steps (f) as shown in FIG. 2K, exit points are selected and the additive random noise in the following sampling steps is removed or skip to the estimated value at Z=0, which significantly enhances the VS fidelity and reduces output image variance. A post-sampling averaging strategy is also disclosed, which can be combined with the aforementioned sampling process engineering techniques to further reduce the output variance and improve the utility of diffusion-based V S techniques in pathology.

[0008] Virtual tissue staining using Al is critically important because it eliminates the need for chemical reagents, reduces tissue processing time and costs, and enables nondestructive, high-resolution analysis of tissue samples, paving the way for faster, more scalable diagnostics and unlocking new possibilities for digital pathology and precision medicine. A powerful generative model for super-resolution virtual tissue staining tasks is disclosed that surpasses traditional deep learning-based VS models, but also introduces sampling process engineering techniques that provide enhanced control over diffusion model image outputs during the testing without the need for retraining or fine-tuning of the model, offering significant benefits in biomedical imaging and related applications including digital pathology'.

[0009] In one embodiment, a method of generating one or more high resolution, virtually stained microscopic images of a label-free test sample is provided. The method includes providing a trained diffusion-based image inference model that is executed by one or more processors of a computing device, wherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s); obtaining one or more autofluorescence images of the label-free test sample using a fluorescence microscope; inputting the one or more autofluorescence images of the label-free test sample to the trained diffusion-based image inference model; and the trained diffusion-based image inference model outputting one or more virtually stained microscopic images of the label-free test sample with high resolution that are substantially equivalent to corresponding high resolution microscopic image(s) of the same label-free test sample that has been chemically stained, wherein the trained diffusion-based image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

[0010] In another embodiment, a method of generating one or more virtually stained microscopic images of a label-free test sample includes providing a trained diffusion-based image inference model that is executed by one or more processors of a computing device, wherein the trained diffusion-based image inference model is trained with a plurality' training images comprising matched chemically stained images or image patches and their corresponding imaging mass spectrometry (IMS) images or image patches of the sametraining sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s); obtaining one or more IMS images of the label-free test sample using an imaging mass spectrometer; inputting the one or more IMS images of the label-free test sample to the trained diffusionbased image inference model; and the trained diffusion-based image inference model outputting one or more virtually stained microscopic images of the label-free test sample with enhanced spatial resolution that are substantially equivalent to corresponding high resolution microscopic image(s) of the same label-free test sample that has been chemically stained, wherein the trained diffusion-based image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality' of intermediate denoising steps as part of the diffusion model inference.

[0011] In one embodiment, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0.

[0012] In another embodiment, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to te and wherein the variance δ̃tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0.

[0013] In another embodiment, the resolution of the autofluorescence images or the IMS images of label-free test samples is lower (i.e., worse resolution) than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

[0014] In another embodiment, the plurality training images and / or the one or more autofluorescence images or IMS images of the label-free training sample(s) and / or label-free test sample(s) obtained using the fluorescence microscope or the image mass spectrometer are input through a shallow neural network for processing of information (e.g., dimension matching).

[0015] In another embodiment, the trained diffusion-based image inference model comprises a convolutional neural network architecture.

[0016] In another embodiment, the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.

[0017] In another embodiment, a system for generating one or more high resolution, virtually stained microscopic images of a label-free test sample includes a computing device having one or more processors configured to execute a trained diffusion-based image inference model that receives one or more autofluorescence images of the label-free test sample and outputs one or more virtually stained microscopic images of the label-free test sample with high resolution that are substantially equivalent to corresponding high resolution microscope image(s) of the same label-free test sample that has been chemically stained; and wherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s) and wherein the trained diffusionbased image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

[0018] In another embodiment, a system for generating one or more virtually stained microscopic images of a label-free test sample includes a computing device having one or more processors configured to execute a trained diffusion-based image inference model that receives one or more imaging mass spectrometry (IMS) images of the label-free test sample and outputs one or more virtually stained microscopic images of the label-free test sample with enhanced resolution that are substantially equivalent to corresponding high resolution microscope image(s) of the same label-free test sample that has been chemically stained; and wherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding IMS images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model leams virtual staining of the microscopic images of the label-free training sample(s) and wherein the trained diffusion-based image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

[0019] In one embodiment, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0.

[0020] In another embodiment, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0.

[0021] In another embodiment, the system further includes a fluorescence microscope or an image mass spectrometer.

[0022] In another embodiment, the one or more processors comprise one or more graphics processing units (GPUs).

[0023] In another embodiment, the resolution of the autofluorescence images or IMS images of label-free test samples is lower (i.e., worse resolution) than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

[0024] In another embodiment, the system further includes a shallow neural network, wherein the plurality training images and / or the one or more autofluorescence images or the IMS images of the label-free training sample(s) and / or label-free test sample(s) obtained using the fluorescence microscope or an image mass spectrometer are input through the shallow neural network for processing of information.

[0025] In another embodiment, the trained diffusion-based image inference model comprises a convolutional neural network architecture.

[0026] In another embodiment, the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.Brief Description of the Drawings

[0027] FIG. 1 A schematically illustrates a system for the virtual staining of unlabeled sample (e.g., tissue sections) that use a diffusion model-based model for virtual staining.

[0028] FIG. IB schematically illustrates the reverse sampling process including the plurality of intermediate denoising steps that span T to step 0. Illustrated are the mean embodiment and the skip embodiment.

[0029] FIGS. 2A-2J illustrate diffusion model-based super-resolution virtual staining of unlabeled tissue sections. FIG. 2A illustrates the diffusion model-based virtual tissue staining pipeline for fluorescence images while FIG. 2B illustrates the pipeline or workflow for IMS images. The image-conditional VS diffusion model is designed based on the Brownian bridge process for both the forward and reverse processes. FIG. 2C is a schematic diagram of the forward process of the Brownian bridge diffusion model. FIGS. 2D-2F illustrates three different reverse sampling processes: vanilla, mean, and skip sampling strategies. FIG. 2G illustrates the detailed workflow for training the diffusion-based VS model. FIGS. 2H-2J illustrate the respective workflows for a single vanilla, mean, and skip sampling step, used individually or in combination in the reverse sampling process in FIGS. 2C-2F.

[0030] FIG. 2K is a plot of the posterior variance added during the reverse sampling process, plotted against the reverse sampling step t, where t — 0 marks the end of the reverse sampling process.

[0031] FIGS. 2A-2C show a comparison of super-resolution virtual staining performances of diffusion-based VS models and cGAN-based VS methods. FIG. 2A illustrates visual comparisons of virtually stained H& E images generated from cGAN-based VS models and the diffusion-based VS models using the mean diffusion sampling strategy. Both models were trained and tested using lower-resolution AF images of unlabeled lung tissue sections, with super-resolution factors ranging from 1 x to 5 x in each lateral direction. Arrowed regions show failures of the cGAN-based VS model. FIG. 2B shows bar plots displaying the SSIM, LPIPS, and PSNR metrics averaged across testing virtually stained images for diffusionbased and cGAN-based VS models. These metrics were calculated on virtually stained and histochemically stained H& E images from 15 blind testing lung samples. The error bars represent the standard error of the mean. Dx, Gxdenote diffusion-based and cGAN-based VS models, respectively, where x represents the super-resolution factor. FIG. 2C shows bar plots of Escores calculated between the diffusion-based VS models and their cGAN-based counterparts with the same super-resolution factor. The dark areas show the statisticallysignificantly improved virtual staining performance of the diffusion virtual staining model over the cGAN-based VS model.

[0032] FIGS. 3A-3C illustrate a comparison of super-resolution virtual staining performances of diffusion-based VS models and cGAN-based VS methods. FIG. 3 A is a visual comparison of virtually stained H& E images generated from cGAN-based VS models and the diffusion-based VS models using the mean diffusion sampling strategy. Both models were trained and tested using lower-resolution AF images of unlabeled lung tissue sections, with pixel super-resolution factors ranging from lx to 5 x in each lateral direction. Arrowed regions show7failures of the cGAN-based VS model. FIG. 3 show's bar plots displaying the SSIM, LPIPS, and PSNR metrics averaged across testing virtually stained images for diffusion-based and cGAN-based VS models. These metrics were calculated on virtually stained and histochemically stained H& E images from 15 blind testing lung samples. The error bars represent the standard error of the mean. Dx, Gx denote diffusion-based and cGAN-based VS models, respectively, where x represents the pixel super-resolution factor. FIG. 3C shows bar plots of t-scores calculated between the diffusion-based VS models and their cGAN-based counterparts with the same super-resolution factor. The shaded areas show the statistically significantly improved virtual staining performance of the diffusion virtual staining model over the cGAN-based VS model.

[0033] FIGS. 4A-4B illustrate spatial frequency spectrum analysis of virtually stained images generated by diffusion-based VS models. FIG. 4A show s input autofluorescence DAPI images, the virtually stained images generated by the diffusion models for different pixel super-resolution factors. The corresponding histochemically stained image is also displayed as ground truth. FIG. 4B illustrates the radially averaged power spectrum crosssections corresponding to the images in FIG. 4A.

[0034] FIGS. 5A-5D illustrates a comparison of performance for different diffusion sampling strategies using the x super-resolution diffusion-based VS model. FIG. 5 A illustrates visual comparisons of virtually stained H& E images generated using different sampling strategies with the diffusion-based VS model trained for 5x super-resolution factor. The virtually stained images produced by the cGAN-based VS model (for the 1 x case, without super-resolution) and the histochemically stained image of the same FOV are also presented for comparison. An assessment conducted by a certified pathologist (N. P.) revealed strong structural similarity across all image subsegments (e.g., alveoli, blood vessels, and scattered lymphocytes). FIG. 5B illustrates bar plots showing the averaged quantitativemetrics, including SSIM and LPIPS, comparing the virtually stained images generated from different diffusion sampling strategies shown in FIG. 5A against their corresponding histochemically stained ground truth images. The cGAN results are also displayed for comparison. FIG. 5C shows bar plots of t-scores calculated between the inference results obtained using the mean diffusion sampling strategy and those from other sampling strategies. The dark areas show the statistically significant superiority of the mean diffusion sampling strategy. FIG. 5D illustrates comparisons of VS image inference time per ~1 mm2 of label-free tissue between the cGAN-based VS model and the diffusion-based VS model using three different sampling strategies. The error bars in FIG. 5B and 5D represent the standard error of the mean.

[0035] FIGS. 6A-6C illustrate a comparison of the coefficient of variation (CV) for diffusion-based virtually stained images generated in different sampling runs using different diffusion sampling engineering approaches. FIG. 6A is a visualization of the CV maps for the YCbCr channels of the generated virtually stained tissue images, obtained using three different approaches: mean sampling, mean sampling with 5-times averaging, and vanilla sampling with 5-times averaging. The generated virtually stained images of these approaches are also presented in the last column. FIG. 6B is a plot of the mean CV for mean / skip sampling strategies with different averaging times. The mean CV was calculated across all color channels and pixels of all test image FOVs. FIG. 6C illustrates a histochemically stained image of the same FOV in FIG. 6 A.

[0036] FIGS. 7A and 7B illustrates super-resolution virtual staining performances on human heart tissue samples using transfer learning. FIG. 7 A illustrates virtually stained H& E images of label-free heart tissue samples, generated by transfer-learned diffusion-based VS models employing the mean sampling strategy. All transfer-learned models (lx to 5 x) were trained using five heart tissue sections. FIG. 7B illustrates bar plots illustrating the SSIM, LPIPS, and PSNR metrics, averaged across virtually stained testing images for the transfer learned diffusion VS models. These metrics were calculated by comparing 178 virtually stained images to their corresponding histochemically stained H& E images from 25 unseen, unlabeled human heart tissue samples. Error bars indicate the standard error of the mean. Dxrepresents the diffusion-based virtual heart H& E model transfer learned from the Dx(virtual lung H& E model), where the x represents the super-resolution factor.

[0037] FIG. 8 illustrates optimization of the sampling exit point tefor both the mean and skip diffusion sampling strategies. Bar plots display the average values of the SSIM andLPIPS metrics for image inference results obtained using the mean and skip diffusion sampling strategies configured with different sampling exit points te. This exit point optimization was performed using the diffusion-based VS model for the 1 x case.

[0038] FIGS 9A-9C illustrate a comparison of super-resolution virtual staining performances of dedicated and universal diffusion-based VS models. FIG. 9A illustrates visual comparisons of virtually stained H& E images generated by separate dedicated models (top) and the universal model (bottom), utilizing the mean diffusion sampling strategy. The dedicated diffusion models were trained and tested individually for a particular spatial downsampling factor, whereas the universal model was simultaneously trained and tested across all pixel super-resolution factors (from lx to 5x). All models were evaluated using autofluorescence images on n = 180 unique FOVs from 15 unlabeled lung samples. FIG. 9B shows bar plots presenting the SSIM, LPIPS, and PSNR metrics, averaged across testing virtually stained images for the dedicated and the universal models. Error bars represent the standard error of the mean. Dxrepresents the dedicated diffusion-based VS model for a particular super-resolution factor x, while Dudenotes the universal diffusion-based VS model. FIG. 9C illustrates bar plots of / -scores comparing the performances of the dedicated models and the universal model for the same super-resolution factors. Dark regions highlight statistically significant improvements in the virtual staining performance achieved by the dedicated models over the universal model.

[0039] FIGS. 10A-10B illustrate evaluation of super-resolution virtual staining performances of diffusion-based VS models using a reduced number of input autofluorescence channels. FIG. 10A illustrates visual comparisons of virtually stained H& E images generated by diffusion-based VS models trained using 2 (DAPI and TxRed), 3 (DAPI, TxRed, Cy5), and 4 (DAPI, TxRed, FITC, Cy5) autofluorescence channels, with a 2x super-resolution factor. For reference, the corresponding virtually stained images generated by the cGAN model and the histochemically stained ground truth are also shown. FIG. 10B illustrates bar plots illustrating the SSIM, LPIPS, and PSNR metrics averaged across testing virtually stained images for diffusion-based and cGAN-based VS models. These metrics were computed by comparing n = 180 sampled virtually stained images to histochemically stained H& E images from 15 blind testing lung samples. Error bars represent the standard error of the mean. The -values resulting from paired / -tests comparing the diffusion VS models to the cGAN model are also provided.

[0040] FIGS. 11 A-l IB illustrate the network architecture for the diffusion-based superresolution virtual staining model for AF images (FIG. 11 A) and IMS images (FIG. 1 IB). The pipeline or workflow of the forward and reverse sampling processes are illustrated respectively. The detailed architecture of the shallow convolutional neural network used for dimension matching is also illustrated. The detailed architecture of the denoising network used at each step of both the forward and reverse sampling processes is also illustrated.

[0041] FIGS. 12A-12C illustrate comparative evaluation of pixel super-resolution virtual staining performance of the diffusion VS models against DDPM-based models. FIG. 12A illustrates visual comparisons of virtually stained H& E images generated by the diffusion VS models (top row) and the DDPM-based diffusion VS models (bottom row). Each model was independently trained and evaluated for specific super-resolution factors ranging from 1 x to 5x. Evaluations were conducted using autofluorescence images from 180 distinct FOVs obtained from 15 unlabeled lung samples. FIG. 12B shows bar plots illustrating quantitative comparisons of SSIM, LPIPS, and PSNR metrics, averaged across n = 180 test images virtually stained by the diffusion VS models and the DDPM-based VS models. Error bars indicate the standard error of the mean. Labels Dxand Pxrepresent the VS models and DDPM-based VS models, respectively, for each super-resolution factor x. FIG. 12C shows bar plots depicting t-scores comparing the performance differences between the VS models and the DDPM-based VS models at identical super-resolution factors. Dark shaded areas highlight statistically significant improvements in virtual staining performance achieved by the VS models relative to the DDPM-based VS models.

[0042] FIGS. 13A-13C illustrate visual comparisons between the virtually stained PAS images generated from label-free IMS data of different patients and their histochemically stained counterparts. FIG. 13 A illustrates imaging mass spectrometry data (images) of label-free tissue, consisting of 1,453 ion (m / z) channels with a pixel size of 10 pm. FIG. 13B illustrates virtually stained images digitally generated from the IMS data using the diffusion-based VS model. FIG. 13D illustrates histochemically stained ground truth images. Both the virtually stained and the histochemically stained images have a pixel size of 1 pm. The concordant localization of glomeruli (G), proximal convoluted tubules (P), and distal convoluted tubules (D), as annotated by a board-certified pathologist, can be visualized on both the VS and HS images.

[0043] FIGS. 14A-14E illustrate quantitative comparisons between virtually stained PAS images generated from label-free IMS data and their histochemically stained counterparts.FIG. 14A shows the color histogram comparisons of virtually stained PAS images and their histochemically stained counterparts in Y. Cb. Cr channels, calculated separately. FIG. 14B illustrates the Hellinger distances between the color histograms of VS and HS image pairs in Y, Cb, Cr channels. FIG. 14C illustrates the box plots of image contrast of virtually stained PAS images and their histochemically stained counterparts across a test dataset comprising 201 distinct FOVs. A statistically significant improvement of the image contrast values of the VS images over the HS images was determined by a two-tailed paired / -test (p=2.60><10’15). FIG. 14D illustrates the low-resolution single MS channel image, the corresponding high-resolution grey-scale virtually and histochemically stained images (top left), and their respective spatial frequency spectra (amplitude, displayed in bottom left). The radially averaged power spectrum cross-section for each case is also presented in FIG. 14E.

[0044] FIGS. 15A-15B illustrate spatial frequency analysis comparing paired and unpaired virtually stained (VS) and histochemically stained (HS) images. FIG. 15A illustrates a visual comparison of virtually stained images, the paired histochemically stained reference images from the same FOV as shown in FIG. 14D, and unpaired histochemically stained images from separate tissue sections. FIG. 15B shows radially averaged power spectrum cross-sections corresponding to the images shown in FIG. 15 A.

[0045] FIG. 16 illustrates a comparison of the Hellinger distances for paired and unpaired VS and HS images across the Y. Cb, and Cr channels. The Hellinger distances between the color histograms of paired VS / HS images and unpaired VS / HS images were computed separately for the Y, Cb, and Cr channels. For each channel, a statistically significant increase in the Hellinger distance for unpaired V S / HS image pairs compared to paired V S / HS image pairs was determined using a two-tailed paired t-test.

[0046] FIGS. 17A-17B illustrates a comparison of virtual staining performance for diffusion models trained with different numbers of IMS channels and channel selection strategies. FIG. 17A illustrates a visual comparison of virtually stained PAS images generated from diffusion-based virtual staining models trained using all the IMS channels and three reduced channel sets obtained through three different selection strategies (resulting in nine configurations in total). FIG. 17B shows bar plots displaying the quantitative PSNR, LPIPS and PCC metrics averaged across testing virtually stained images for all the configurations. Higher PSNR, PCC, and lower LPIPS values are desired.

[0047] FIGS. 18A-18C illustrates a comparison of virtual staining performance of mass spectrometry images using different sampling strategies. FIG. 18A shows a visualcomparison of the virtually stained PAS images generated using different sampling strategies. FIG. 18B illustrates bar plots showing the average LPIPS values, comparing the virtually stained images generated using different diffusion sampling strategies shown in FIG. 18A against their corresponding histochemically stained ground truth images. FIG. 18C illustrates bar plots showing the averaged pixel -wise coefficient of variation (CV) across all YCbCr channels for the virtually stained images generated in different sampling runs using different diffusion sampling approaches.

[0048] FIGS. 19A-19B illustrate optimization of the sampling exit point tefor both the mean and skip diffusion sampling strategies. FIG. 19A shows visual comparison of the virtually stained PAS images generated using mean / skip sampling strategies configured with different exit points, te. FIG. 19B shows bar plots showing the average LPIPS values, FID values, and NIQE values, comparing the virtually stained images generated using different diffusion sampling configurations show n in FIG. 19A against the corresponding histochemically stained ground truth images. Lower values across these three performance metrics are desired, indicating a higher similarity to the ground truth images.

[0049] FIGS. 20A-20B illustrate a comparison of the coefficient of variation (CV) for different noise sampling engineering approaches. FIG. 20A shows a visualization of the CV maps for the YCbCr channels of virtually stained images obtained using three different approaches: vanilla, mean, and skip sampling approaches. The CV was calculated using five repetitions of virtual staining inference on the same FOV. FIG. 20B illustrates plots of the mean CV calculated across all pixels of all test image FOVs for the three sampling approaches.

[0050] FIGS. 21A and 21B illustrate multimodal data acquisition and image registration. FIG. 21 A illustrates the multimodal data acquisition pipeline. An unstained tissue section was first imaged using autofluorescence microscopy (prior-IMS AF imaging) and subsequently subjected to IMS data acquisition. After the IMS data acquisition, the tissue underwent AF microscopy again to produce post-IMS AF images. Finally, the same tissue section was stained histochemically and imaged using bright-held microscopy. FIG. 21B illustrates the multimodal registration pipeline. The registration process started with bright-field microscopy images of histochemically stained tissue (source modality), and these are computationally aligned them with post-IMS AF images (target modality) using the Elastix registration algorithm. The post-IMS AF images were aligned with the IMS pixel map by using laser ablation marks as reference points.Detailed Description of Illustrated Embodiments

[0051] FIG. 1 A schematically illustrates one embodiment of a system 2 for outputting digitally or virtually stained images 40 from an input microscope image 20 of a sample 22. The input microscope image 20 may include an autofluorescence (AF) microscope image in one embodiment obtained from a fluorescence microscope 110. In another embodiment, the input microscope image 20 is an IMS image obtained from imaging mass spectrometer 112. As explained herein, the input image 20 is an autofluorescence image or an IMS image of a sample 22 (such as tissue in one embodiment) that is not stained or labeled with a fluorescent stain or other label. The input image 20 may be an autofluorescence image 20 of the sample 22 in which the fluorescent light that is emitted by the sample 22 is the result of one or more endogenous fluorophores or other endogenous emitters of frequency-shifted light contained therein. Frequency-shifted light is light that is emitted at a different frequency (or wavelength) that differs from the incident frequency (or wavelength). Endogenous fluorophores or endogenous emitters of frequency-shifted light may include molecules, compounds, complexes, molecular species, biomolecules, pigments, tissues, and the like. In some embodiments, the input image 20 (e.g., the raw autofluorescence image) is subject to one or more linear or non-linear pre-processing operations selected from contrast enhancement, contrast reversal, image filtering. For IMS images 20, the image data may include mass spectrometry images containing a plurality of channels obtained during a pixelwise raster scan of the sample 22.

[0052] The system 2 includes a computing device 100 that contains one or more processors 102 therein and software 104 that executes the trained diffusion-based image inference model 10. The computing device 100 may include, as explained herein, a personal computer, laptop, mobile computing device, remote server, or the like, although other computing devices may be used (e.g., devices that incorporate one or more graphic processing units (GPUs)) or other application specific integrated circuits (ASICs). GPUs or ASICs can be used to accelerate training as well as final image output. The computing device 100 may be associated with or connected to a monitor or display 106 that is used to display the digitally or virtually stained images 40 as seen in FIG. 1 A. The display 106 may be used to display a Graphical User Interface (GUI) that is used by the user to display and view the digitally or virtually stained images 40. In one embodiment, the user may be able to trigger or toggle manually between multiple different digital / virtual stains for a particularsample 22 using, for example, the GUI. Alternatively, the triggering or toggling between different stains may be done automatically by the computing device 100.

[0053] Network training of the diffusion-based image inference model 10 may be performed the same or different computing device 100 that executes the inference model 10. For example, in one embodiment a personal computer 100 may be used to train the diffusionbased image inference model 10 although such training may take a considerable amount of time. To accelerate this training process, one or more dedicated GPUs may be used for training. As explained herein, such training and testing was performed on one or more GPUs obtained from a commercially available graphics card. Once the diffusion-based image inference model 10 has been trained, the diffusion-based image inference model 10 may be used or executed on a different computing device 100 which may include one with less computational resources used for the training process (although GPUs may also be integrated into execution of the diffusion -based image inference model 10).

[0054] The software 104 can be implemented using Python and TensorFlow although other software packages and platforms may be used. The diffusion-based image inference model 10 is not limited to a particular software platform or programming language and the diffusion-based image inference model 10 may be executed using any number of commercially available software languages or platforms. The software 104 that incorporates or executes the diffusion-based image inference model 10 may be run in a local environment or a remote cloud-type environment. In some embodiments, some functionality of the software 104, which may include image processing software of functionality, may run in one particular language or platform (e.g., image normalization) while the diffusion-based image inference model 10 may run in another particular language or platform.

[0055] In one embodiment, the trained diffusion-based image inference model 10 receives one or more autofluorescence images 20 (FIG. 2A) or ISM images 20 (FIG. 2B) of an unlabeled sample 22. In some embodiments, for example, where multiple excitation channels are used in the fluorescence microscope 110, there may be multiple fluorescence input images 20 of the unlabeled sample 22 that are input to the trained diffusion-based image inference model 10. The trained diffusion-based image inference model 10 is trained specifically on the type of images to be input during use. For example, the trained diffusionbased image inference model 10 that is used with input microscopy images 20 obtained from a fluorescence microscope 110 will be different from the trained diffusion-based imageinference model 10 that is used with input microscopy images 20 obtained from an imaging mass spectrometer 112.

[0056] The autofluorescence input images 20 may, in some implementations, include a wide-field fluorescence image 20 of an unlabeled tissue sample 22. Wide-field is meant to indicate that a wide field-of-view (FOV) is obtained by scanning of a smaller FOV, with the wide FOV being in the size range of 10-2,000 mm2. For example, smaller FOVs may be obtained by a scanning fluorescence microscope 110 that uses software 104 to digitally stitch the smaller FOVs together to create a wider FOV. Wide FOVs, for example, can be used to obtain w hole slide images (WSI) of the sample 22. The autofluorescence image is obtained using a fluorescence microscope 110. The fluorescence microscope 110 includes an excitation light source that illuminates the sample 22 as well as one or more image sensor(s) (e.g., CMOS image sensors) for capturing fluorescent light that is emitted by fluorophores or other endogenous emitters of frequency -shifted light contained in the sample 22. The fluorescence microscope 110 may, in some embodiments, include the ability to illuminate the sample 22 with excitation light at multiple different wavelengths or wavelength ranges / bands. This may be accomplished using multiple different light sources and / or different filter sets (e g., standard UV or near-UV excitation / emission filter sets). In addition, the fluorescence microscope 110 may include, in some embodiments, multiple filter sets that can filter different emission bands. For example, in some embodiments, multiple autofluorescence input images 20 may be captured, each captured at a different emission band using a different filter set.

[0057] The sample 22 may include, in some embodiments, a portion of tissue that is disposed on or in a substrate 23 (FIG. 1 A). The substrate 23 may include an optically transparent substrate in some embodiments (e g., a glass or plastic slide or the like). The sample 22 may include a tissue section that is cut into thin sections using a microtome device or the like. Thin sections of tissue 22 can be considered a weakly scattering phase object, having limited amplitude contrast modulation under brightfield illumination. The sample 22 may be imaged with or without a cover glass / cover slip. The sample may involve frozen sections or paraffin (wax) sections. The tissue sample 22 may be fixed (e.g., using formalin) or unfixed. The tissue sample 22 may include mammalian (e.g., human or animal) tissue or plant tissue. The tissue sample 22 may include healthy tissue or diseased tissue (e.g., cancerous tissue). The sample 22 may also include other biological samples, environmental samples, and the like. Examples include particles, cells, cell organelles, pathogens, or othermicro-scale objects of interest (those with micrometer-sized dimensions or smaller). The sample 22 may include smears of biological fluids or tissue. These include, for instance, blood smears, Papanicolaou or Pap smears. As explained herein, for the autofluorescencebased embodiments, the sample 22 includes one or more naturally occurring or endogenous fluorophores that fluoresce and are captured by the fluorescence microscope device 110. Most plant and animal tissues show some autofluorescence when excited with ultraviolet or near ultra-violet light. Endogenous fluorophores may include by way of illustration proteins such as collagen, elastin, fatty acids, vitamins, flavins, porphyrins, lipofuscins, co-enzymes (e.g., NAD(P)H). For the IMS-based embodiments, kidney tissue was specifically used but it should be understood that other tissues types may also be used in additional to other types of samples 22 as outlined above.

[0058] The trained diffusion-based image inference model 10 in response to the input image 20 outputs or generates a digitally or virtually stained output image 40 (FIGS. 1 A, 2 A, 2B). The digitally / virtually stained, high resolution output image 40 (which may be optionally super-resolved in some embodiments) has "‘staining'' that has been digitally integrated into the output image 40 using the trained diffusion-based image inference model 10. In some embodiments, such as those involved tissue sections as the sample 22, the trained diffusion-based image inference model 10 appears to a skilled observer (e.g., a trained histopathologist) to be substantially equivalent to a corresponding high resolution microscopic image (e.g.. brightfield) of the same tissue section sample 22 that has been chemically stained. This digital or virtual staining of the tissue section sample 22 appears just like the tissue section sample 22 had undergone histochemical staining even though no such staining operation was conducted. In addition to the output image 40 being digitally or virtually stained, in some embodiments, it has a higher resolution (e.g., high resolution or super-resolved image).

[0059] In one embodiment, the trained diffusion model 10 is a diffusion model-based high resolution virtual staining (V S) model that transforms input microscope images 20 that are lower-resolution auto-fluorescence (AF) microscopy images of label-free tissue samples 22 into high resolution brightfield output images 40 that digitally match the histochemically-stained higher-resolution images of the same tissue samples without the need for traditional chemical staining. That is to say, in this embodiment, the virtual or digitally stained higher-resolution image(s) 40 is / are substantially equivalent to a corresponding brightfield image(s) of the same label-free sample 22 that has been chemically stained. The diffusion model 10described herein is based on an image-conditional diffusion model that is based on the Brownian bridge process as seen in FIGS. 2C-2K. The diffusion model 10 integrates lower-resolution conditional images and noise estimation from an attention-based convolutional neural network architecture that incorporates the time step information to reconstruct a higher-resolution virtual histological image 40; thus performing two tasks at the same time: (1) spatial resolution enhancement and (2) virtual staining of label-free tissue.

[0060] The conditional diffusion model 10 generates better VS images 40 with higher resolution and image fidelity that, in some embodiments, substantially matching the ground truth histochemically stained brightfield images. In addition, to mitigate the inherent high variance of diffusion models for VS applications in pathology, novel sampling process engineering techniques are introduced, i.e., the mean and skip sampling strategies described herein (FIG. IB and FIGS, 21, 2J). After training (FIG. 2G), in the reverse process of the diffusion model 10 (e.g., going from low resolution autofluorescence images 20 to higher resolution, virtual stained images 40 or going from IMS images 20 to virtual stained images 40), exit points (te) are selected and remove the additive random noise in the following sampling steps or skip to the estimated value at t=0, which significantly enhances the fidelity’ of the output image 40 and reduces variance of the output image 40. In another embodiment, a post-sampling averaging strategy7is used, which can be combined with the aforementioned sampling process engineering techniques to further reduce the output variance and improve the utility- of diffusion-based VS techniques in pathology.

[0061] In one embodiment, a method of generating one or more high resolution, virtually stained microscopic images 40 of a label-free test sample 22 includes providing a trained diffusion-based image inference model 10 that is executed by one or more processors 102 of a computing device 100, wherein the trained diffusion-based image inference model 10 is trained with a plurality training images that include matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model 10 learns virtual staining of the microscopic images of the label-free training sample(s) 22. Once training is complete and the trained diffusion model 10 is in place, one or more autofluorescence input images 20 of the label-free test sample 22 are obtained using a fluorescence microscope 110. The one or more autofluorescence input images 20 of the label-free test sample 22 are input to the trained diffusion-based image inference model 10. The trained diffusion-based image inference model 10 then outputs one or more virtuallystained microscopic images 40 of the label-free test sample 22 with high resolution that, in one embodiment, are substantially equivalent to a corresponding high resolution microscopic image(s) of the same label-free test sample 22 that has been chemically stained.

[0062] Here, the trained diffusion-based image inference model 10 generates the one or more virtually stained images 40 using a reverse sampling process (FIGS. 2H-2J) that includes a plurality of intermediate denoising steps t as part of the diffusion model inference 10. In one embodiment, the plurality of intermediate denoising spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 8tat the intermediate denoising steps from T to te and wherein the variance 8tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0. See FIGS. IB and 21. Here, T is denoted as the starting state / point of the reverse diffusion process. This embodiment is referred to as "mean sampling.” In this configuration, the variance 8tis omitted from a portion of the plurality of intermediate denoising steps leaving just mean sampling steps from the exit point teto the step 0. The mean sampling step is illustrated in FIGS. IB and 21.

[0063] In another embodiment, the trained diffusion-based image inference model 10 generates the one or more virtually stained images 40 using a reverse sampling process that includes a plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0. This embodiment is referred to as “skip sampling.” In this configuration, the variance 8tis omitted from a portion of the plurality of intermediate denoising steps from the exit point teand skips directly to step 0. The skip sampling step is illustrated in FIGS. IB and 2J.

[0064] In another embodiment, a system 2 for generating one or more high resolution, virtually stained microscopic images 40 of a label-free test sample 22 is provided. A computing device 100 having one or more processors 102 executes a trained diffusion-based image inference model 10 that receives one or more one or more autofluorescence input images 20 of the label-free test sample 22 and outputs one or more virtually stained microscopic images 40 of the label-free test sample 22 with high resolution that, in one embodiment, are substantially equivalent to a corresponding high resolution microscopicimage(s) (e.g., brightfield image(s) in one embodiment) of the same label-free test sample 22 that has been chemically stained.

[0065] The trained diffusion-based image inference model 10 is trained with a plurality training images including matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s). The trained diffusion-based image inference model 10 leams virtual staining of the microscopic images of the label-free training sample(s) 22. The trained diffusion-based image inference model 10 generates the one or more virtually stained images 40 using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference. In one embodiment (mean sampling), the plurality of intermediate denoising steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0. In another embodiment (skip sampling), the plurality’ of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to te and wherein the variance δ̃tis omitted at the exit point te and the trained diffusion model inference skips directly to step 0. The system 2, in some embodiments, may include a fluorescence microscope 110 that is associated with the computing device 100.

[0066] In another embodiment, a method of generating one or more virtually stained microscopic images 40 of a label-free test sample 22 is disclosed. The method includes providing a trained diffusion-based image inference model 10 that is executed by one or more processors 102 of a computing device 100, wherein the trained diffusion-based image inference model 10 is trained with a plurality training images that include matched chemically stained images or image patches and their corresponding IMS images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model 10 leams virtual staining of the microscopic images of the label-free training sample(s) 22. Once training is complete and the trained diffusion model 10 is in place, one or more IMS input images 20 of the label-free test sample 22 are obtained using an image mass spectrometer 112. The one or more IMS input images 20 of the label-free test sample 22 are input to the trained diffusion-based image inference model 10. The trained diffusion-basedimage inference model 10 then outputs one or more virtually stained microscopic images 40 of the label-free test sample 22 that, in one embodiment, are substantially equivalent to a corresponding high resolution microscopic image(s) of the same label-free test sample 22 that has been chemically stained.

[0067] Here, the trained diffusion-based image inference model 10 generates the one or more virtually stained images 40 using a reverse sampling process that includes a plurality of intermediate denoising steps t as part of the diffusion model inference 10. This includes the mean sampling and skip sampling processes described previously.

[0068] In another embodiment, a system 2 for generating one or more virtually stained microscopic images 40 of a label-free test sample 22 is provided. A computing device 100 having one or more processors 102 executes atained diffusion-based image inference model 10 that receives one or more one or more IMS input images 20 of the label-free test sample 22 and outputs one or more virtually stained microscopic images 40 of the label-free test sample 22. In one embodiment, the one or more virtually stained microscopic images 40 are substantially equivalent to a corresponding high resolution microscopic image(s) (e.g., bnghtfield image(s) in one embodiment) of the same label-free test sample 22 that has been chemically stained.

[0069] The trained diffusion-based image inference model 10 is trained with a plurality' training images including matched chemically stained images or image patches and their corresponding IMS images or image patches of the same training sample(s). The trained diffusion-based image inference model 10 learns virtual staining of the microscopic images of the label-free training sample(s) 22. The trained diffusion-based image inference model 10 generates the one or more virtually stained images 40 using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference. In one embodiment (mean sampling - FIG. 21), the plurality of intermediate denoising steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 6tat the intermediate denoising steps from T to teand wherein the variance 6tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point te to step 0. In another embodiment (skip sampling - FIG. 2J), the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 8tat the intermediate denoising steps from Tto te and wherein the variance 8tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0. The system 2, in some embodiments, may include an image mass spectrometer 112 that is associated with the computing device 100 as seen in FIG. 1A and 2B.

[0070] Results

[0071] Super-resolved virtual staining of unlabeled tissue sections using a diffusion model

[0072] The workflow of diffusion model-based virtual staining is illustrated in FIGS. 2A and 2B. Workflow begins by capturing label-free autofluorescence images of unstained human lung tissue samples 22 which are used for training of the inference model 10. Then, these slides were sent for histopathological H& E staining (for generation of ground truth stained samples), followed by digital imaging with a brightfield optical microscope to obtain the corresponding histochemical images, which serve as the ground truth images for the training and testing phases. The workflow for IMS images proceeds similarly and is described below. Each one of the label-free autofluorescence images was paired and registered with respect to the corresponding labeled histochemical images to create the training / testing dataset for virtual staining. The data capturing and preprocessing steps are detailed in the Methods section.

[0073] Diffusion model training and sampling represent two directions of propagation of the same stochastic process. In this approach, the forward process is modeled by a Brownian bridge, as illustrated in FIG. 2C, which is a Gaussian process with a linearly scheduled mean from x0to xTand a quadratically scheduled variance with respect to the time step. Here, the set of xQrepresents the target image domain, corresponding to histochemically stained images, while the set of xTdenotes the input image domain, comprising lower resolution autofluorescence images (or IMS images) that went through a shallow convolutional neural network 50 (used for dimension matching), as illustrated in FIGS. 2A-2B. In contrast, the reverse process (FIG. 2D) is a step-wise denoising process agnostic of the ground truth image x0. However, the exact distribution of xtconditioned on xT, Wt < T is intractable and therefore a neural network is employed to estimate the posterior mean of xtconditioned on xT(see the detailed derivations in the Methods section and below). As seen in FIG. 2G, a U-Net-based denoising network 52 is trained to estimate the difference between xtand0, where t is an arbitrary, random time step between 0 and T. After the training, the denoising network 52 predicts the difference between the current state xtand0, as shown in FIG. 2H.Notice that an additional posterior noise with variance 6tis introduced to match the variance schedule of the Brownian bridge. The posterior variance 8tis illustrated in FIG. 2K, where a large posterior noise is added to diversify the distribution; this choice, however, also introduces excessive randomness in the inference stage, which presents a challenge to achieving high consistency among multiple inferences - an important feature to have for VS applications since one needs to virtually stain the tissue sample of “the patient”. To mitigate this challenge of diffusion-based VS models 10, engineered sampling strategies, the mean and skip diffusion sampling methods, illustrated in FIGS. 21 and 2J, respectively are utilized. The mean sampling strategy eliminates the posterior noise addition in the final few sampling steps from an empirically defined exit point te, and the skip sampling strategy estimates0|tdirectly. Detailed sampling algorithms for these two inference strategies are elucidated in the Methods section.

[0074] The super-resolution virtual staining performance of the diffusion-based virtual staining model 10 was validated on human lung tissue histomorphology and compared its performance to that of a cGAN-based state-of-the-art virtual staining method. For a direct comparison between the two approaches, multiple super-resolution VS models 10 (covering different super-resolution factors) separately trained using the diffusion-based VS approach as well as the traditional cGAN approach (see the training details in the Methods section). The trained diffusion-based model 10 was termed as Dxand the trained cGAN model as Gx. where x denotes the pixel super-resolution factor. To span different levels of resolution loss, the input autofluorescence images 20 underwent pixel binning for spatial undersampling, covering super-resolution factors of I x to 5 x in each lateral direction; for example, a superresolution factor of 5x indicates a space-bandwidth reduction of 25-fold at the input AF microscopy images 20 of label-free tissue.

[0075] During the blind testing phase, both approaches (diffusion vs. cGAN) were applied to 180 autofluorescence image sets (with each autofluorescence channel having 960x960 pixels) from 15 unlabeled lung tissue sections that were never used during training to generate virtual stained images 40 at various super-resolution factors, as shown in FIG. 3A. The blind testing results revealed that, across all the super-resolution factors (2x to 5x), the diffusion-based VS models 10 using the mean sampling strategy consistently outperformed cGAN-based models; this performance advantage can be visually confirmed in FIG. 3A and was quantitatively demonstrated in FIG. 3B by the higher structural similarity index measure (SSIM) values, lower learned perceptual image patch similarity (LPIPS) values, and higherpeak signal-to-noise ratio (PSNR) values calculated with respect to the high-resolution brightfield images of the same histochemically stained tissue samples (ground truth).Specifically, as indicated in the arrowed region of FIG. 3 A, due to spatial resolution loss of the AF microscopy images 20 of label-free tissue samples 22, cGAN-based models failed to reconstruct stained regions of anthracotic pigment, which are important in lung pathology for disease diagnosis. In contrast, the diffusion-based super-resolution VS models 10 consistently stained these black pigments, presenting a good match to the histochemically stained images

[0076] To further quantify the advantages of diffusion-based super-resolution VS models 10, the statistical significance of potential improvements observed in the SSIM and LPIPS values of the two methods were tested (see the Methods section for details). The / -scores illustrated in FIG. 3C highlight the statistically significant performance improvements of the diffusion-based super-resolution VS models 10 over traditional cGAN-based models for super-resolution factors of 2* to 5 x; for the 1 x case without super-resolution, both of these approaches perform similar in virtual staining image quality, without a statistically significant difference between them.

[0077] To highlight the super-resolution capabilities of the diffusion-based VS model 10, a spatial frequency spectrum analysis was performed on the input autofluorescence DAPI images, the virtually stained images generated by the diffusion model 10, and their corresponding histochemically stained ground truth images for all super-resolution factors. The results are presented in FIG. 4A; also refer to the Methods section for details. The crosssections of the radially averaged power spectra, as shown in FIG. 4B, reveal a significant enhancement in the spatial frequency spectra of the VS images 40 compared to the lower-resolution autofluorescence inputs 20. Notably, the spectra of the VS images 40 have a good alignment with those of the histochemically stained ground truth, underscoring the diffusionbased VS model's ability to enhance spatial resolution.

[0078] Diffusion sampling engineering techniques for super-resolved virtual staining of label-free tissue

[0079] The performance of three different sampling strategies, named vanilla (FIG. 2H), mean (FIG. 21), and skip (FIG. 2J) methods was compared using the VS diffusion model 10 trained for a 5x super-resolution factor, as shown in FIG. 5A. Since the diffusion reverse process introduces additional noise with a variance of 8t, different inferences using the same sampling strategy can produce output images 40 with inherent variance. To further mitigate this variability, in addition to these sampling methods, an averaging strategy was employed torefine the virtual staining quality and suppress the stochasticity at the output super-resolved VS images 10. Specifically, repeated image inferences were performed using the vanilla diffusion sampling strategy to generate multiple super-resolved VS images 40 for the same field of view (FOV). These images were then averaged on a pixel-wise basis, using 2, 3, and 5 repeats and the resulting averaged virtually stained images are presented in FIG. 5A.Quantitative metrics were used to compare these super-resolved VS images generated using different strategies against their corresponding histochemically stained ground truth images, as shown in FIG. 5B. Specifically, these quantitative metrics were calculated using the virtually stained images from a subset of the testing dataset, consisting of 12 paired image patches (each measuring 960 x 960 pixels) derived from a single patient. This comparative analysis revealed that the mean diffusion sampling strategy achieved the second-highest SSIM scores and the second-lowest LPIPS scores, outperforming both the vanilla and skip diffusion sampling strategies. It is worth noting that both the vanilla sampling and the skip sampling approaches may not surpass the VS performance of cGAN, as illustrated in FIG. 5B. However, by employing the mean sampling and averaging strategy, the diffusion-based virtual staining model can achieve image quality that outperforms that of cGAN. To further highlight the advantages of the mean diffusion sampling strategy, a paired / -test was conducted between it and the other strategies, see the Methods section and FIG. 5C. The / -test results demonstrated that the mean diffusion sampling strategy is statistically superior to the other approaches, except for the vanilla sampling strategy with 5-times averaging. However, if the same number of averaging were to be used for the mean sampling-based diffusion strategy, it would perform superior compared to the vanilla sampling strategy.

[0080] FIG. 5D compares the VS image inference time per ~1 mm2of label-free tissue section for each strategy. Compared to the cGAN-based models, the diffusion-based VS models 10 have a much longer inference time because of the repeated runs of the denoising network during the reverse sampling process - this is a general weakness of the diffusionbased VS models 10. Since the skip sampling strategy does not require the denoising network inference after the exit point te, it achieves a decrease in the VS image inference time compared to the vanilla and mean sampling strategies, which is an advantage. Moreover, the inference time of the vanilla sampling strategy with post-sampling averaging can be estimated by multiplying the inference time of the vanilla sampling strategy7by the number of averages. Although the five-time averaging strategy in vanilla sampling offers better performance, its significantly longer inference times may limit its applicability in time-sensitive scenarios; in such cases, the mean diffusion sampling strategy should be the preferred method for competitive image inference.

[0081] Evaluation of diffusion-based virtual staining image variance

[0082] In clinical practice, consistent tissue staining quality is crucial for pathologists to make reliable diagnoses. Therefore, the inherent variance in virtually stained images 10 of the same tissue FOV, resulting from different trajectories of the diffusion process, may hinder its clinical applicability. To demonstrate that the diffusion sampling engineering techniques can significantly reduce this variance, the coefficient of variation (CV) was analyzed for virtually stained images 10 generated in different runs of the diffusion reverse process. For this analysis, the trained diffusion model 10 was used for lx virtual staining (i.e., without downsampling of the input AF microscopy images 20), and tested the CV performance of three diffusion sampling engineering techniques: mean sampling, mean sampling with 5-times averaging, and vanilla sampling with 5-times averaging. Specifically, each diffusion sampling technique was performed five times, i.e. five independent runs of the mean sampling strategy’ and 25 runs of both the mean and vanilla sampling strategies. This resulted in five virtually stained images 40 for the same tissue FOV for each approach. The CV map of each method was then calculated by taking the pixel-wise ratio of the standard deviation to the mean for the YCbCr channels of the virtually stained images, as illustrated in FIG. 6A. The chroma components (Cb and Cr) reflect the blue and red color information of the virtually stained images 40, respectively, evaluating the staining quality of H& E. where hematoxylin stains nuclei purplish-blue and eosin stains the extracellular matrix and cytoplasm pink. As shown in FIG. 6A, both of the averaging methods drastically reduced the variance of the diffusion-based image inference across all three channels. Furthermore, w ith the 5-times averaging approach, the mean diffusion sampling strategy’ outperformed the vanilla sampling strategy, achieving a mean CV of less than 0.5% for both the Cb and Cr channels. A similar analysis is reported in FIG. 6B with the averaging factors ranging from 2 to 5. The results further confirmed that the mean diffusion sampling strategy outperformed the others at all the averaging factors, revealing its superiority over the vanilla sampling strategy by yielding consistent results with smaller variance at its VS images. Note that the reduction in the output variance achieved by averaging with the mean sampling strategy exhibits diminishing returns beyond 5 -time averaging; therefore, 5 -time averaging was selected as the optimal strategy to showcase the virtual staining performance. The results demonstrate that a more deterministic diffusion image inference with negligible variationscan be achieved in virtual tissue staining using the diffusion sampling process engineering approaches, which is important for practical applications of diffusion-based VS models 10 in digital pathology.

[0083] Generalization to human heart tissue samples

[0084] To further demonstrate the robustness and the generalization ability7of the diffusion-based VS models 10 on a new type of organ tissue, transfer learning was employed on lung H& E diffusion models (l tos) using a small human heart tissue dataset. This dataset included autofluorescence and histochemically stained image pairs from five heart samples. The resulting heart-specific H& E diffusion models 10, denoted as(where x represents the super-resolution factor, e.g., D transfer-learned from D ). were subsequently tested on 178 autofluorescence image sets (with each autofluorescence channel having 960 x 960 pixels) from 25 unlabeled heart tissue sections not included in the transfer learning process. The blind testing results reveal that the virtually stained heart H& E images 40 align closely with the histochemically stained ground truth, regardless of the super-resolution factor, as shown in FIG. 7A. This agreement is consistently observed across various FOVs of the heart tissue obtained from different patients. Furthermore, the quantitative metrics presented in FIG. 7B confirm the extended success and consistency of the virtual staining models 10. Additionally, paired Etests were performed to compare the performance of D2Hwith D. The / ^-values shown in FIG. 7B indicate that these models deliver statistically comparable virtual staining performance for heart samples, even though D2Hand were tested on AF images 20 with lower spatial resolution. These findings highlight the robustness and super-resolution capability of the diffusion-based VS models 10, reinforcing their potential for accurate and adaptable staining across various tissue and organ types.

[0085] Discussion

[0086] The presented success of the Brownian bridge diffusion model 10 for superresolution virtual staining of label-free tissue, combined with the sampling process engineering, can be readily extended to various existing image reconstruction or enhancement tasks in biomedical imaging. As demonstrated in the Results section, the diffusion-based VS models 10 outperformed cGAN-based alternatives in super-resolution virtual staining of label-free tissue images 20, while the two approaches performed statistically similar for the VS without super-resolution. Diffusion models in general exhibit stable training dynamics and are well-suited to addressing some of the most challenging image reconstruction problems.

[0087] One of the major achievements of this invention is the development of diffusion sampling process engineering to suppress performance variations in virtually stained images 40 of label-free tissue 22. A vanilla diffusion model 10 was derived to synthesize images matching the distribution of x0. where the additional step-wise variance introduced during the reverse process not only stands for an essential component of the diffusion model to match posterior distributions p^x^ \xt,y but also enables diversity in the generated images. On the other hand, the consistency of the virtual staining results is crucial for pathologists to make reliable diagnoses. Considering that the profde of 3tshows a drastic increase at the end stages of the reverse sampling process (as illustrated in FIG. 2K), using a mean or skip diffusion sampling step that avoids 6tin the final sampling steps is important to suppress the stochastic variations in the output of the diffusion model for the same label-free tissue FOV. It is worth noting that this sampling process engineering does not require modifications to the training process or finetuning of a pre-trained model, since the network 10 consistently predicts the error at xt.

[0088] The diffusion sampling process engineering approaches presented herein for superresolution virtual tissue staining can be further refined and optimized. Both the mean and skip sampling strategies are constricted by the famous bias-variance tradeoff. In other words, the reduction in the variance of a stochastic estimator would inevitably cause an increase of the error bias between the mean of the estimator and the ground truth value. To better understand this trade-off for super-resolution virtual staining of label-free tissue, the exit point for both the mean and skip diffusion sampling strategies was optimized. In a diffusion process consisting of 1000 total steps, the performance of the mean and skip sampling strategies was evaluated across nine different exit points, ranging from 10 to 500. Quantitative image metrics were calculated by comparing the generated VS images to their ground truth counterparts, as shown in FIG. 8. The superior quantitative metrics observed across all exit points tefurther validate the advantages of the mean sampling strategy over the skip sampling strategy. Specifically, for the mean sampling strategy, an increment in LP1PS scores with larger exit points was observed, confirming the necessity of the posterior noise and the bias-variance trade-off. However, the effect of the exit point may differ under different evaluation metrics. The SSIM scores improved at larger exit points, indicating that the addition of the variance term results in a trade-off. Therefore, picking a comprehensive evaluation metric and an appropriate exit point for an optimal result is very important for uses of diffusion-based image inference models in biomedical applications. Furthermore, beyondoptimizing the exit points, combining different sampling strategies may also enhance the image inference performance. As demonstrated in FIGS. 6A-6B, integrating the mean diffusion sampling strategy with subsequent averaging yields superior CV performance compared to other competing strategies. The design of these combined strategies can be more sophisticated; for example, one can apply the averaging strategy' to the results inferred from the mean diffusion sampling strategies implemented with different exit points te. One can also simultaneously apply accelerated sampling strategies, e.g. Denoising Diffusion Implicit Models (DDIM) and Pseudo Linear Multi-Step method (PLMS), with the variance-reduction strategies introduced herein to achieve faster and better results. These various combinations can be further explored and tailored for different image reconstruction and synthesis applications, beyond the virtual staining of label-free tissue sections.

[0089] The diffusion-based VS models 10 (Dx) presented herein are specifically trained for each integer super-resolution factor, making them dedicated models for a particular spatial downsampling factor. To demonstrate the versatility and robustness of this framework, a universal model (Dy) was explored that was simultaneously trained and tested across all super-resolution factors (from 1 x to 5x). See FIGS. 9A-9C. The quantitative evaluation of the blind test results for this universal model is illustrated in FIGS. 9B-9C. As expected, the universal model (Dy) exhibits a performance trade-off and cannot achieve statistically equivalent performance compared to the dedicated diffusion models (£>x), likely due to accommodating diverse super-resolution tasks. Nevertheless, Duconsistently produces high-quality virtually stained images that closely match the corresponding histochemically stained ground truth images across all super-resolution factors. Notably, it successfully reconstructs anthracotic pigment features that the cGAN model failed to capture, as shown in FIG. 3A.

[0090] The application of the diffusion-based VS models 10 was also investigated for faster sample scanning by evaluating their performance with a reduced number of input AF channels. Specifically, diffusion VS models 10 were trained and tested using two AF channels (DAPI and TxRed) and three AF channels (DAPI, TxRed. and Cy5) with a 2x super-resolution factor. The visual and quantitative results of these models, termedchand Dch, are presented in FIGS. 10A-10B. These quantitative results demonstrate that Dch, with one AF channel removed, still achieves statistically significant improvements in virtual staining performance compared to the cGAN model (ti2) that used all four AF channels. Furthermore,c / l. with tw o AF channels removed, maintains statistically equivalentperformance to the cGAN model G2. These results highlight the robustness of the diffusion-based framework and its potential for faster sample scanning by reducing the input image AF channel requirements without compromising performance.

[0091] Denoising Diffusion Probabilistic Models (DDPM) are among the most widely used and powerful diffusion frameworks for image generation, editing, and enhancement. To further demonstrate the effectiveness of the diffusion-based VS models 10 (Dx), their performance against DDPM-based VS models (Px) was compared, as shown in FIGS. 12A-12C. To provide a fair comparison, the DDPM-based models employed the original DDPM sampling process and used the same denoising U-Net architecture as the diffusion models 10. Quantitative evaluations on blind test data, presented in FIGS. 12B and 12C, clearly show that the diffusion models 10 outperform the DDPM-based virtual staining counterparts, underscoring the effectiveness of the disclosed approach.

[0092] A diffusion-based model 10 for super-resolution virtual staining of label-free, lower resolution autofluorescence microscopy images is disclosed, demonstrating its superiority over the state-of-the-art virtual tissue staining approaches. Furthermore, several diffusion sampling engineering techniques were developed, which not only improved the image quality of virtually stained images 40 but also resulted in more consistent image inference with a lower statistical variation, which is crucial for the wide-spread uses of these Al-based image reconstruction / synthesis approaches in biomedical settings.

[0093] Methods

[0094] Sample preparation, image acquisition, and histochemical H& E staining

[0095] Lung and heart tissue samples 22 for this research were sourced from existing, deidentified tissue blocks housed at the UCLA Translational Pathology Core Laboratory (TPCL), with authorization under IRB approval # 18-001029. The study involved lung specimens from 33 individual patients and heart samples from 30 individual patients. For each patient, a tissue section approximately 4 pm thick was sliced from the unlabeled tissue blocks, deparaffmized, and mounted on glass slides. Autofluorescence images 20 of the lung tissue sections were acquired using a Leica DMI8 microscope 110 equipped with a 40x / 0.95 NA objective lens (Leica HC PL APO 40* / 0.95 DRY), controlled by the Leica LAS X software for automated microscopy. The autofluorescence images 20 of label-free tissue sections 22 were captured under four distinct fluorescence filter cubes: DAPI (Semrock OSFI3-DAPI-5060C, EX 377 / 50 nm, EM 447 / 60 nm), TxRed (Semrock OSFI3-TXRED-4040C. EX 562 / 40 nm, EM 624 / 40 nm), F1TC (Semrock F1TC-2024B-OFX. EX 485 / 20 nm,EM 522 / 24 nm), and Cy5 (Semrock CY5-4040C-OFX, EX 628 / 40 nm, EM 692 / 40 nm). Images 20 were captured using a scientific complementary metal-oxide-semiconductor (sCMOS) sensor (Leica DFC 9000 GTC), with an exposure time of around 300 ms for all four autofluorescence channels. After performing autofluorescence imaging, the same unlabeled tissue sections 22 were sent to UCLA TPCL for standard histochemical H& E staining (used as ground truth). The stained tissue slides were then scanned and digitized using a brightfield slide scanner (Leica Biosystems Aperio AT2).

[0096] Dataset division and preparation

[0097] The training and testing dataset comprised paired autofluorescence images and their corresponding bright-field histochemically stained H& E images. For the lung experiments, 1,051 paired autofluorescence-H& E microscopic image patches (each with 2048x2048 pixels) from 18 patients were used for training, and 180 paired image patches (each with 960x960 pixels) were reserved for blind testing, obtained from 15 de-identified patients not included in the training set. As for the heart experiments, 163 paired autofluorescence-H& E images (each with 2048x2048 pixels) from 5 patients were used for transfer learning, and 178 image pairs (each with 960x960 pixels) from 25 patients were used for blind testing. During each training epoch, the paired image FOVs were subdivided into smaller 192xl92-pixel patches, normalized for zero mean and unit variance, and further augmented through random flipping and rotation to ensure robust model training.

[0098] In the image registration workflow, a rigid registration approach was initially applied at the whole slide image (WSI) level. The maximum cross-correlation coefficient was calculated for each WSI pair, allowing for the estimation of the rotation angles and shifts. This facilitated the spatial alignment of the histochemically stained WSIs to their autofluorescence counterparts. Following this, the slides underwent a finer registration at the image patch level. The WSIs were segmented into smaller FOV pairs of 3248x3248 pixels (-528x528 pm2), followed by a multi-modal affine image registration algorithm to adjust for shifts, sizing differences, and rotations between the histology and autofluorescence image FOVs. Before cropping WSIs into smaller local FOVs, an intensity normalization of the autofluorescence channels was performed. In the final phase, small local paired FOVs were cropped to 2048x2048 pixels (-333x333 pm2) to reduce edge artifacts and underwent an iterative elastic pyramid cross-correlation registration to achieve pixel-level alignment. During this elastic registration process, an initial virtual staining netw ork w as trained to match the style of the autofluorescence images to the style of the brightfield H& E images.The resulting roughly-stained images and their histochemically stained counterparts were fed into the elastic pyramid registration algorithm to obtain transformation maps which were then applied to correct the local discrepancies in the ground truth images to better align with their corresponding autofluorescence images. These training and registration cycles were repeated until precise pixel-level registration was achieved. As a final step, the manual data cleaning was employed to remove images with obvious artifacts like tissue tearing or image blurring in out-of-focus areas.

[0099] Baseline VS models using cGAN

[0100] All baseline VS models, which were compared with the diffusion-based VS models 10, were built on a state-of-the-art structurally-conditioned GAN architecture. The generator network employs a five-level U-Net structure, while the discriminator network is a convolutional neural network-based classifier. The training loss for the generator includes adversarial loss and pixel-based structural loss terms, such as mean absolute error (MAE) loss and total variation (TV) loss. Meanwhile, the training loss for the discriminator utilizes a least squares loss function. The generator and the discriminator networks were updated at a frequency ratio of 3: 1. The learning rate for optimizing the generator network was set at 1 X 10“4, while for the discriminator network, it was set at 1 X 10-5. The Adam optimizer was employed for network training, and the batch size was set as 8 for all the model training.

[0101] Brownian bridge diffusion process

[0102] The Brownian bridge diffusion process was used to model the conditional diffusion and apply it to transform the lower resolution autofluorescence images 20 of label -free tissue y0(where N is the super-resolution factor) into the target histochemically stained higher resolution brightfield images xQG [j^wxivxsdiffusion sampling process requires the conditional and target images to have identical dimensions, the autofluorescence images 20 of label-free tissue y0were first processed to match the dimensions of the target image domain through a shallow convolutional neural network (CNN) 50, as depicted in FIG. 11 A. This shallow network 50 is composed of two convolutional layers and a pixel-shuffle layer for dimension adaptation, denoted as:

[0103] y = fcy0) (1)

[0104] where y G 1RHX M / x3has a dimension matched with the histochemically stained ground truth images x0. The forward Brownian bridge with an initial state x0and terminal state y is defined as:

[0106] where T is the total sampling steps, and t is the intermediate time step. T was set to 1000 during both the forward and reverse processes. By denoting mt= - and 8t= — — —, one can reparametrize the distribution of xtas:

[0108] Eq. (3) shows that the arbitrary intermediate step xtcan be directly sampled using x0and y during the forward diffusion process, represented as "direct sampling" in FIG. 2C. The shallow network 50 output y is concatenated with the noisy image x£and fed into a denoising network 52 ee. This denoising network 52 is trained to estimate x0from xtand t. In other words, it is trained to estimate mt(y — x0) +This denoising network 52 egis designed based on a U-Net structure and consists of one down-sampling path and one up-sampling path with skip connections in-between the two paths. FIG. 11B illustrates the architecture of the U-Net-based denoising network 52, where each path has four consecutive levels and each level contains a residual block and an attention block. The adjacent levels in the down-sampling and up-sampling path are connected with a 2 X 2 average pooling and 2 x 2 nearest interpolation, respectively. The middle block of the denoising network 52 is the concatenation of two residual blocks and one attention block. The attention block adopts a multi-head attention mechanism. The timestep t is embedded through a linear layer with SiLU pre-layer activation and added to the input features of the residual block.

[0109] The loss function of eeis defined as:

[0111] wheretis the weight for each t. During the training, a uniform sampling schedule was utilized, settingt= 1 for all the sampling steps, t.

[0112] The vanilla reverse process can be shown to be a Gaussian process with mean HtXt, y)arqd variance 6tI

[0116] The expressions of cxt, cyt, cet, and Sqt-i and the details of Eqs. (4) and (6) are detailed further herein.

[0117] During the blind testing, a test autofluorescence image y0is first processed by fcto generate the image xT= y. From this, the image at time T — 1, xr-15can be estimated using

[0118] This step is repeated for T iterations to estimate the target virtual stained image x0, as denoted with "sample step by step" in FIGS. 2D-2F. This process is the vanilla sampling strategy illustrated in FIG. 2D, 2H. This strategy can be represented as a Markov chain with learned Gaussian transitions, starting with p(xr) = J\f (xT; y, 0):

[0120] The mean sampling strategy’, as depicted in FIGS. 2E and 21. go through the same vanilla sampling steps when te< t < T, where teis defined as the exit point in reverse sampling process. For t < te, this strategy7estimates x^ 'rom x, using a mean sampling step as follows:

[0121] Pm^Xt-^y) = J^' xt-1iit,xt,y), 0) (9)

[0122] where the x^ is directly estimated from n't(xt, y), without adding additional variance pattern 8tin the sampling step, i.e., xt-1= cxtxt+ cyty — c£teext, t) for t < te. Therefore, the mean sampling strategy can be formulated as:

[0124] Similarly, the skip sampling strategy follows the same vanilla sampling steps until the exit point t = te:

[0126] However, it diverges by directly estimating the final virtual stained image from the state xte, as illustrated in FIGS. 2F, 2J, i.e.:

[0128] During the training, the shallow convolution neural network and the denoising network were all optimized using the AdamW optimizer, starting with a learning rate of 1 x 104. A batch size of 16 was maintained throughout the training phase, and the network converged after approximately 72 hours of training. During the blind testing, te= 50 was empirically set.

[0129] Quantitative image evaluation metrics

[0130] To quantitatively evaluate the performance of H& E virtual staining, standard metrics of SSIM and LPIPS were utilized. The SSIM is defined as:[100131

[0132] where iaand i-ibare the mean values of a and b. which represent the two images being compared. < Jaand < Jbare the standard deviations of a and b. ua bis the crosscovariance of a and b. Crand C2are constants that are used to avoid division by zero.

[0133] For LPIPS metrics, pretrained VGG model was used to evaluate the learned perceptual similarity between the generated virtually stained images m and their corresponding histochemically stained images m0. These image pairs were fed into the pretrained VGG network and their feature stack from L layers were extracted as ft1, nQlG ]^w!xwixC( £oriayeri TeLPIPS score was calculated as:

[0136] where A represents the histochemically stained brightfield H& E images and max (d) is the maximum pixel value of the image A. The mean squared error (MSE) is defined as:

[0137] where B represents the virtually stained H& E images, m, n are the pixel indices, and MN denotes the total number of pixels in each image.

[0138] To perform the spatial frequency spectrum analysis, the raw autofluorescence DAPI image was first bilinearly upsampled from its original size ofpixels (x is the spatial undersampling factor) to 960x960 pixels, matching the dimensions of the grayscale virtually stained and histochemically stained images. A two-dimensional Fourier Transform was then applied to the upsampled autofluorescence DAPI image, as well as the virtually stained and the histochemically stained image. For consistency, both the virtual images 40 and histochemical H& E images were processed in grayscale for this analysis. The radially averaged power spectrum was calculated following the method reported in Wang et al., Deeplearning enables cross-modality super-resolution in fluorescence microscopy. Nat. Methods 16, 103-110 (2019), which is incorporated by reference.

[0139] Statistical significance analysis

[0140] In FIGS. 3A, 3B, paired two-sided / -tests were used to assess whether the diffusion-based VS models 10 showed statistically significant difference from the virtual staining performance of cGAN-based VS models. These tests were performed across 180 unique FOVs, using the SSIM and LPIPS metrics calculated for both the diffusion-based VS model 10 and the corresponding cGAN-based VS counterpart for the same super-resolution factor. The null hypothesis for the paired two-sided / -test assumes that the two models, given the same super-resolution factor, should have identical means for the same FOV. A statistical significance level of 0.05 was used to reject the null hypothesis in favor of the alternative, i.e., the two-sided t-test assumed 2.5% probability7level to each side (superiority and inferiority). Note that a lower (higher) score for LPIPS (SSIM) metrics is desired for improved VS performance; therefore, the results in FIGS. 3A-3B revealed a statistically significant improvement of the diffusion-based VS models 10 over their cGAN-based counterparts for LPIPS when t < 0, p < 0.05 and for SSIM when t > 0, p < 0.05.Similarly, as illustrated in FIGS. 4A-4B, paired two-sided / -tests were performed using the SSIM and LPIPS metrics to compare the virtual staining performance of the mean sampling VS strategy’ against other diffusion-based VS strategies. Same as before, for both LPIPS (when t < 0) and SSIM (when t > 0), a.p value of < 0.05 from the two-sided t test revealed the statistical superiority of the mean sampling-based diffusion strategy7over the other VS approaches, showing inferiority only to the vanilla sampling strategy with 5 -times averaging.

[0141] Other implementation details

[0142] All image preprocessing and registration were performed using MATLAB software 104 version R2022b. All network training and testing tasks were conducted on a desktop computer 100 equipped with an Intel Core i9-13900K CPU, 64 GB of memory, and an NVIDIA GeForce RTX 4090 GPU. The code for training the diffusion models was developed in Python 3.9.19 using PyTorch 2.2.1. Example testing images and network models are available together with the code at: https: / / github.com / Yijie-Zhang / Super-resolved-virtual-staining which is incorporated by7reference herein.

[0143] Brownian bridge diffusion process

[0144] Given a target image x0G j^xivxc andaconditional image y G ^HxWxCthe forward process of Brownian Bridge diffusion model with a total sampling step T can be defined as:

[0148] Given the transition probability q xt|x0, xT= y), the intermediate states can be computed as:

[0150] where et, et-1~ N(0, 1). By removing x0in Eq. (16) and Eq. (17), the transition probability q(xt|xt-1,y) can be derived as:

[0154] In the reverse process, the diffusion process starts from the conditional image y Gseting the input xT= y. Given the input xT. the xt^ can be predicted by:

[0156] where 14 (xt, y) is the predicted mean value of the noise, and 8tis the variance of noise at each step.

[0157] The training objective for the Brownian Bridge diffusion process is based on optimizing an Evidence Lower Bound (ELBO), which can be denoted as:

[0159] where q(xt-1\xt, x0, y) in the second term can be derived through Bayes’ theorem and the Markov chain property:

[0161] By comparing with Eq. (15) and Eq. (18), Eq. (22) can be derived as:

[0163] where the mean term Tt(xt, x0, y) and the variance term 6tare given by:

[0166] By combining Eq. (15) and Eq. (24), / lt(xt, x0, y) can be reformulated as:

[0168] where cxt, cytic£tare defined as:

[0172] During training, the neural networkis trained to estimate the noise term

[0174] Thus, the training objective Eq. (21) can be simplified to optimizing the lossbetween the sampled noise in the forward process and the estimated noise by the neuralnetwork:

[0176] where ytis the weight for each t.

[0177] Image Mass Spectrometry7(IMS) Embodiment

[0178] Imaging mass spectrometry (IMS) enables untargeted, highly multiplexed mapping of molecular species in biological tissue with unparalleled chemical specificity andsensitivity. However, most IMS platforms lack microscopy-level spatial resolution andcellular morphological contrast, necessitating subsequent histochemical staining, microscopic imaging and advanced image registration steps to correlate / link molecular distributions with specific tissue features and cell types. In this embodiment, a diffusion model-based virtual histological staining approach is used that enhances spatial resolution and digitally introducescellular morphological contrast into mass spectrometry images of label-free human tissue. Blind testing on human kidney tissue demonstrated that the virtually stained images of label-free samples closely match their histochemically stained counterparts (with Periodic Acid-Schiff staining), showing high concordance in identifying key renal pathology structures despite utilizing IMS data with 10-fold larger pixel size. Additionally, this approach employs optimized noise sampling during the diffusion model's inference to achieve reliable and repeatable virtual staining.

[0179] Here, a diffusion model-based virtual histological staining technique is disclosed that transforms IMS-measured ion images 20 reporting molecular species’ distributions in label-free tissue samples into super-resolved brightfield microscopy images 40, closely matching their histochemically stained (HS) counterparts (as illustrated in FIG. 2B). The approach utilizes an image-conditional diffusion model 10, underpinned by the Brow nian bridge process, which integrates the low-resolution conditional input with noise estimation through an attention-based U-Net 52 (FIGS. 2G, 2H, 21, 2J) that incorporates the time-step information to reconstruct the high-resolution histological image 40 of the label-free sample. Following a one-time training effort, the diffusion-based virtual staining (VS) model 10 was able to generate brightfield microscopy equivalent Periodic Acid-Schiff (PAS)-stained images 40 from IMS-measured ion images 20 of label-free human kidney tissue samples 22 (never seen before) despite the fact that the input IMS data had a 10-fold larger pixel size. Quantitative evaluations confirmed that the VS approach effectively overcomes key limitations of IMS for histological interpretation by digitally creating high-resolution histological stain images using low-resolution IMS data without the need for chemical staining and microscopic imaging of stained tissue, also eliminating image registration steps since the virtually stained images are automatically registered with the IMS data used as the label-free input of the VS model.

[0180] A board-certified pathologist was able to identify key pathological structures directly from the virtually stained PAS images 40, demonstrating a high degree of concordance with the features observed in the corresponding histochemically stained images and emphasizing the human interpretability of this approach. Furthermore, to mitigate the inherent high variance associated with diffusion models 10, an optimized noise sampling strategy was employed, which eliminates additive random noise in the final stages of the reverse diffusion process as explained herein. This approach not only quantitatively reduces output variance but also ensures that the pathological features in the diffusion model outputsfrom different test runs remain consistent and histologically equivalent. In summary, the diffusion-based virtual staining approach overcomes some of the limitations of traditional IMS, including its relatively low spatial resolution and the absence of cellular morphological contrast, also eliminating the need for time-intensive histological staining process and complex image co-registration after the IMS data acquisition.

[0181] Results - Virtual Staining of IMS images

[0182] Virtual histological staining of IMS-measured ion images of label-free tissue using a diffusion model

[0183] The dataset used herein comprises IMS-measured ion images 20 of label -free human kidney tissues 22 and high-resolution brightfield images of the histochemically stained versions of the same tissue samples 22. As shown in FIG. 2B. the IMS data of each tissue sample 22 were acquired using pixel-wise raster scanning with 10 pm lateral spacing. Subsequent data processing selected the most representative channels, resulting in individual ion images 20 containing 1,453 mass-over-charge (m / z) channels, each with a pixel size of 10 pm. After the IMS data acquisition, the label-free tissue slides were subjected to histopathological PAS staining and digitally imaged using a benchtop brightfield optical microscope. The brightfield images of the histochemically stained tissue samples were then registered to their corresponding ion images 20, which formed the label-free input (IMS) and ground truth (brightfield and labeled) image pairs used for training the inference model 10. It is worth noting that autofluorescence microscopy imaging was performed both prior to and following the IMS data acquisition to facilitate the registration between the IMS data and the brightfield microscopy images of histochemically stained tissue. Details of the dataset collection and preprocessing are provided in the Methods section below.

[0184] The training and sampling process of the Brownian Bridge Diffusion Model (BBDM) reflect the two-directional propagation of a Brownian bridge process, as depicted in FIGS. 2C-2F. The forward process starts from x0, representing the target image domain of the histochemically stained brightfield images (ground truth, GT), and progresses toward xT. the input image domain, corresponding to lower-resolution IMS images 20 (with 1,453 m / z channels), processed through a shallow convolutional neural network 50 for image dimension alignment / match, as shown in FIG. 2B. The mean of the forw ard process is linearly scheduled from x0to xT, while the variance evolves quadratically over time steps (t). In contrast, the reverse process aims to denoise the input IMS images 20 of label-free tissue step-by-step, gradually refining the data without directly using the ground truth image x0. AU-Net-based denoising neural network 52 is trained to estimate the posterior mean of xtbased on xT. As shown in FIGS. 2G, 2H, 21, 2J. the denoising network 52 is employed to consistently estimate the difference between the current state xtand the target image x0at arbitrary time steps between 0 and T (see the Methods section). Importantly, this Brownian bridge framework departs fundamentally from the classic denoising diffusion probabilistic model, DDPM. In unconditional DDPM, the forward process degrades an initial sample x0by adding Gaussian noise at each timestep according to a fixed variance schedule, ultimately producing a pure-noise state xT. A denoising neural network 52 is then trained to invert this process: beginning from xT~ N(0,I), it iteratively predicts the noise component and partially removes it, gradually reconstructing x0. When adapted for conditional image translation, the conditioning signal — whether a text embedding or an image — is concatenated with the noisy image at every forward and reverse step, ensuring that the model retains direct access to the context throughout the diffusion trajectory. In contrast, the Brownian bridge diffusion model 10 reframes the forward process so that the terminal state is the conditional image xTitself (the label-free IMS data). Noise is incrementally injected into x0(histochemically stained, ground truth image) until it converges to xT, yielding intermediate states xtthat stochastically interpolate between x0and xTalong a Brownian bridge.Consequently, in the learning stage, the denoising neural network learns to denoise these blended states. Crucially, for the Brownian bridge reverse process, the initial input is solely the conditional image (label-free IMS data). These forward and reverse processes of the BBDM framework result in more stable image translation outputs.

[0185] To improve the consistency / repeatability of the virtual staining process and minimize stochasticity in the diffusion process-generated images, a deterministic noise sampling strategy was implemented alongside the standard noise sampling method.Specifically, a mean sampling strategy7w as used (FIGS. 2E, 21) that eliminates, after a certain time point is reached, the posterior noise introduced during the vanilla sampling process (FIGS. 2D, 2H). Detailed sampling algorithms for this mean sampling strategy are provided in the Methods section. The performance advantages of this deterministic sampling are further evaluated in the subsection “Deterministic diffusion model inference via noise sampling engineering techniques.”

[0186] Following the training phase, the BBDM-based VS model 10 was tested on human kidney tissue excluded from both the training and validation datasets. FIGS. 13A-13C showcases the VS images 40 generated by the BBDM-based model 10 using IMS images 20from different patients. The virtually stained outputs 40, as presented in FIG. 13B, exhibit a good resemblance to the GT brightfield images, despite being generated from 10-fold larger pixel size IMS images 20, as illustrated in FIG. 13B. Furthermore, a board-certified pathologist annotated key renal structures — glomeruli (denoted as G), as well as proximal (P) and distal (D) convoluted tubules — on both the VS images 40 and HS images. These structures are of clinical importance since they are involved in most renal pathological conditions. As depicted in FIGS. 13B-13C, there was very good concordance in the identification of these structures between the VS and HS ground truth PAS images. This alignment was consistently demonstrated across multiple fields of view (FOVs) of tissue, underscoring the robustness and generalizability of the framework for the virtual staining of low-resolution IMS data.

[0187] To quantitatively evaluate the fidelity of the virtually generated high-resolution PAS images 40 in replicating their HS counterparts, a comparative analysis was conducted on a test dataset comprising 201 distinct tissue FOVs, each with 640x640 pixels, from 10 patients. This evaluation focused on three critical aspects: (1) image color distribution; (2) image contrast; and (3) spatial frequency spectrum (detailed in the Methods section). These quantitative metrics reported in FIGS. 14A-14E were selected to evaluate whether the VS images 40 generated from ion images 20 can meaningfully contribute to the interpretation of IMS-measured tissue samples without physically performing the chemical staining and microscopic imaging of stained tissue, and, thus, effectively safeguarding the tissue for other assay types and subsequent analysis. While the approach requires multimodal image registration between IMS data and microscopy images during the training phase, which is a one-time effort, this eliminates cumbersome image registration steps during the inference phase. During image inference, the V S images 40 are inherently aligned with their input IMS data, removing the need for additional registration.

[0188] As depicted in FIG. 14A, the color distribution histograms in the YCbCr color space across all the test FOVs demonstrate strong color agreement between the VS images and the HS brightfield images. This agreement is further supported by the small Hellinger distances calculated between the color histograms of each VS and HS image pair (FIG. 14B). Moreover, as illustrated in FIG. 14C, the generated VS images 40 achieve image contrast comparable to that of the brightfield HS images.

[0189] Furthermore, to showcase the super-resolution capabilities of the VS staining method, a spatial frequency spectrum analysis was conducted on the raw single-channel MSimages 20 (picked from 1,453 m / z channels based on the contrast of glomeruli), network-generated VS images 40, and their corresponding HS ground truth images; see the Methods section for details. This analysis, illustrated in FIG. 14D, includes cross-sections of the radially averaged power spectra, which demonstrate an excellent match between the spatial frequency spectra of the VS and HS image pairs, as desired. The relative error between the radially averaged power spectra (log scale, normalized) of the V S and the corresponding HS ground truth images is 6.70% ± 4.84% (mean ± standard deviation) across all the frequencies (FIG. 14E). These results further confirm the effectiveness of the VS output images 40, which successfully align with the spatial frequency spectra of the high-resolution HS images. Additionally, the relative increase of the radially averaged power spectra (log scale, normalized) of the VS images 40 compared to their lower resolution MS input is 144.23% ± 81.19% (mean ± standard deviation) over all the frequencies (FIG. 14E). These results demonstrate a marked improvement over the spatial frequency spectra of low-resolution MS images 20, underscoring the capacity' of the diffusion-based virtual staining inference model 10 to enhance spatial resolution significantly.

[0190] To further illustrate the effectiveness of the spatial frequency analysis in quantifying the agreement between the virtually stained images 40 generated by the V S inference model 10 and their corresponding paired histochemically stained images, an additional comparison was conducted involving two unpaired histochemically stained image FOVs with the paired VS / HS images presented in FIG. 14D. With reference to FIGS. 15A-15B, there is a clear discrepancy in the radially averaged power spectra between the unpaired histochemically stained images and the paired VS / HS images. This result further underscores the effectiveness and specificity of the virtual staining framework in accurately enhancing spatial frequency features to closely match the corresponding histochemically stained ground truth image. Similarly, to validate the significance of the color distribution agreement quantified using the Hellinger distance, an additional comparative analysis was performed involving unpaired histochemically stained images and their virtually stained counterparts. FIG. 16 presents the calculated Hellinger distances between the color histograms of the paired VS / HS images and those of unpaired VS / HS image pairs across the Y, Cb, and Cr channels. As illustrated in FIG. 16, there is a statistically significant increase in the Hellinger distances for the unpaired VS / HS images compared to the paired VS / HS images, reinforcing the specificity of the virtual staining method in reproducing the color characteristics inherent in histochemical staining.

[0191] Reduction analysis on mass spectrometry image channels

[0192] The high-resolution image translation capability of virtual staining for IMS data relies on the rich molecular information captured in ion images. To further shed light on this, a series of diffusion-based VS models 10 were trained using different numbers of mass spectrometry (m / z) channels and evaluated their performance on the same test dataset. For the channel selection, three distinct strategies were evaluated. In the primary approach, the IMS channel indices were selected from the top-ranking channels of the sorted list of 1,453 channels, prioritized based on signal-to-noise ratio (SNR) values (in descending order). The SNR definition used in this study entails that for each IMS channel, the mean is divided by the standard deviation across all pixels within that channel. In addition to this SNR-based ranking and channel selection, two alternative channel selection strategies were explored: frequency -based selection and uniform selection. For the frequency-based selection method, a 2D Fourier transform was applied to the ion image of each channel, ranking the channels by the ratio of the average power in the angular spectrum of the high-frequency components to that of the low-frequency components. The half-maximum frequency was set as the threshold to differentiate between high and low frequencies. For the uniform selection strategy, the channels were sampled at fixed intervals (every 4th, 16th, or 64th channel for reductions of 4-fold, 16-fold, and 64-fold, respectively), starting from the first channel. Using these three distinct selection approaches, channels were selected from the full set of 1,453 IMS channels, progressively reducing the total number of channels utilized per input image from 363 down to 23, corresponding to 4-fold, 16-fold, and 64-fold reductions. These selected IMS channels remained consistent throughout the training and testing of each VS model 10 for a given channel selection strategy'.

[0193] As shown in FIG. 17A, reducing the number of MS channels from 1,453 to 23 resulted in a gradual degradation of the virtual staining performance for all three channel selection strategies, with noticeable losses in critical features such as nuclear morphology. The relationship between the number of IMS channels and V S quality' w as further quantified using the peak signal-to-noise ratio (PSNR) and learned perceptual image patch similarity (LPIPS) and Pearson correlation coefficient (PCC). FIG. 17B presents the PSNR, LPIPS and PCCs metrics for the ten VS models 10, each trained with a combination of a different number of IMS channels and channel selection approaches, evaluated across a test dataset comprising 201 distinct tissue FOVs from ten patients, each with 640x640 pixels. These results confirmed that as the number of used MS channels increased, the diffusion-based VSmodels 10 achieved a statistically significantly higher V S fidelity, underscoring the rich molecular information present in MS data and its strong utility for label-free histological staining. Among the three distinct channel selection strategies, the SNR-based selection consistently delivered the best performance compared to the other two channel sampling approaches, further validating its effectiveness in preserving the critical molecular information for high-quality virtual staining with a reduced number of IMS channels.

[0194] Repeatable diffusion model inference via noise sampling engineering

[0195] Tissue heterogeneity and inherent histochemical staining variability often lead to minor differences between adjacent tissue sections. These spatial changes usually do not influence the overall slide-level diagnosis and are well tolerated by human pathologists in their clinical workflow. However, computational models applied for tissue analysis are frequently misled by these slide-to-slide variations, resulting in lower performance. For the virtual staining model 10, the primary source of variation in the output V S results comes from the stochastic nature of the noise sampling process during the backward diffusion. To address this, noise sampling process engineering techniques were applied to improve VS consistency and reduce output variance without the need for fine-tuning or transfer learning on the trained model 10. In addition to the mean sampling strategy illustrated in FIGS. 2E and 21, an alternative “skip sampling'’ strategy7was used for comparison (FIGS. 2F and 2J), which directly estimates0from the denoising network's output. The skip sampling strategy was previously described herein.

[0196] The rationale behind the mean and skip sampling strategies arises from the significant increase in the additional variance 8tduring the final stages of the reverse diffusion process (see FIG. 2K). Both of these sampling strategies — mean and skip — avoid this increasing noise variance in the reverse path after an engineered exit point te(see the Methods section); this effectively reduces stochastic variations (observed from run to run) in the output of the diffusion model for the same label-free tissue FOV, which is highly desired for VS applications. Detailed descriptions of these engineered noise sampling techniques are provided in the Methods section and FIGS. 18A-18C and 19A-19B.

[0197] To quantitatively assess the effectiveness of these noise sampling strategies, the trained BBDM was tested using the vanilla, mean, and skip sampling strategies, repeating each method five times and calculating the pixel-wise coefficient of variance (CV) across these repetitions of the diffusion-based VS process. FIG. 20A shows the CV maps for the three distinct strategies in the YCbCr color channels, demonstrating that both the mean andskip sampling strategies effectively reduce the variance in the sampled VS images compared to the vanilla method. Additionally, the average CV values were computed across all the pixels in the test image FOVs and plotted them for each YCbCr channel, as shown in FIG. 20B. The results further corroborate that the mean and skip sampling strategies are effective in achieving lower output variances, indicating the repeatability of the diffusion-based VS process using these engineered noise sampling techniques.

[0198] This comparative analysis in FIGS. 20A-20B further reveals that the skip sampling strategy yields a lower average CV value compared to the mean sampling method. However, the mean sampling strategy7produces results that exhibit better perceptual similarity7to the ground truth histochemically stained PAS images. Additional visual and quantitative comparisons between different noise sampling strategies in diffusion-based VS models 10 are presented in FIGS. 18A-18C. These results confirm that the mean sampling strategy is superior to both the vanilla and skip sampling strategies, achieving a lower average LPIPS as desired. Additionally, the performance of an averaging strategy7was evaluated, which involves averaging independent test runs of different inferences for the same tissue FOV. While FIG.18C highlights the advantage of the averaging strategy in reducing pixel-level CV, the resulting images exhibit relatively low er contrast with a significantly w orse LPIPS score, falling short of the performance achieved by the mean sampling strategy7alone. Therefore, this averaging strategy would only be preferable in scenarios where high consistency is prioritized over image contrast.

[0199] The noise sampling strategies demonstrated herein can be further optimized. For instance, the performance of the mean and skip sampling strategies were evaluated across eight different exit points (te), ranging from 0 (equivalent to the vanilla sampling strategy) to 100, as illustrated in FIGS. 19A-19B. The findings indicate that both the mean and skip sampling strategies consistently outperformed the vanilla strategy7across a range of exit points. Specifically, both strategies achieved optimal LPIPS performance at an exit time point te~10. Consequently, selecting a well-suited evaluation metric and determining the appropriate exit point are critical for optimizing the performance of diffusion-based VS models 10, particularly for clinical applications.

[0200] Discussion - IMS

[0201] The effectiveness of the presented label-free virtual staining of IMS-measured ion images 20, despite the 10-fold larger pixel size of the input IMS data, can be attributed to the capabilities of diffusion models 10 in effectively capturing and modeling complex datadistributions. Historically, Generative Adversarial Network (GAN)-based approaches have been a predominant choice in image restoration for biomedical applications, particularly in super-resolution image reconstruction tasks. However, recent advancements in the field have revealed that GANs might struggle with highly challenging image reconstruction problems, especially at extreme super-resolution factors. In contrast, diffusion models have emerged as a superior alternative, consistently producing more realistic and accurate spatial features even at these high super-resolution factors. Moreover, the complexity of multiplexed inference tasks that require simultaneous resolution enhancement and cross-domain image translation has further underscored the limitations of GAN-based techniques. It has been demonstrated that diffusion models 10 outperform GANs in such multiplexed tasks, owing to their robustness in generating high-fidelity images. Additionally, the issue of mode collapse, which is a well-documented limitation of GANs when confronted with sparse or low-quality data, is well mitigated in diffusion models. Diffusion models, in general, exhibit more stable training dynamics, even when faced with significant discrepancies between the input data and ground truth images, making them more resilient in handling challenging datasets. For the same super-resolution VS task presented herein, the cGAN framework was evaluated, however, the model collapsed at an early stage and failed to converge effectively, resulting in images containing unacceptable staining artifacts. Taking these factors into account, one can conclude that the success of the VS models 10 in generating high-resolution virtual stains from low-resolution IMS data / images 20 of label-free tissue samples 22 is a testament to the diffusion models’ inherent strengths.

[0202] The repeatability and consistency of the IMS-generated virtually stained images 40 are crucial for digital pathology interpretation and were achieved using the mean and skip sampling strategies employed during the reverse process of the diffusion model. Another notable achievement of the disclosed VS method is its ability to enable medical experts to directly identify diagnostically relevant renal pathology structures such as glomeruli, proximal, and distal convoluted tubules using virtually stained images generated from lower-resolution IMS images 20. Historically, recognizing these structures directly from IMS images 20 was not feasible. Conventional histology-directed IMS analysis requires both IMS data and brightfield microscopy images of the histochemically stained tissue, followed by a complex and time-consuming registration process to link the IMS data with pathology-annotated regions in optical microscopy images of stained tissue. Furthermore, histochemical staining of IMS slides prevents their utilization for genomics / epigenetics analyses andhinders IMS-molecular comparison studies. The diffusion-based VS model 10, however, allows these limitations to be bypassed by offering a means of histological interpretation of IMS data for regions of interest within the kidney tissue, immediately after an IMS scan, without the need for histochemical staining. Moreover, once the diffusion model is trained, this technique can be seamlessly integrated into the post-IMS data processing pipeline without requiring any modifications to the existing hardware / setup.

[0203] It is also important to note that the quality of IMS-based virtual staining can be further enhanced. Here, the virtual histological staining of ion images 20 with a pixel size of -10 pm w as demonstrated, as this is a common spatial resolution for MALDI IMS systems. However, with recent advances in IMS technology, achieving higher spatial resolution is now possible. For instance, transition mode ion sources allow IMS to achieve 1 pm pixel sizes. Unfortunately, these specialized workflows often inflict significant tissue damage, as the tissue is ablated during the sampling process, making the post-IMS multimodal analysis impossible. By applying this VS framew ork to these high-resolution IMS images 20, the post-IMS multimodal analysis challenge of tissue damage is obviated. In addition, this enhances the fidelity of the virtually stained images, further elevating the precision and quality of the diffusion-model results. Furthermore, although the efficacy of the technique was demonstrated using PAS staining on human kidney samples, this label-free approach can be extended to other types of histochemical stains and various organs. This adaptability is supported by the prior success of various virtual staining techniques, making it versatile across different staining protocols and sample 22 (e g., tissue) types.

[0204] The VS method described herein can augment existing IMS datasets by generating virtually stained images 40, offering advantages to the broader biomedical research community. In addition, the reduction analyses on IMS channels further confirm the rich information encoded in IMS data and can potentially reveal the relationship of IMS channels with histochemical staining, which might prospectively facilitate a better understanding of both the virtual and histochemical staining processes. Importantly, the virtual staining method remains effective even when using a reduced number of IMS channels, as selected via a signal-to-noise ratio (SNR)-based reduction strategy. This provides several practical advantages. First, computational load and hardware requirements for both model training and inference scale with the number of input channels; thus, reducing the channel count enhances the method’s accessibility and efficiency. Second, the results suggest that this SNR-based feature reduction makes the method more compatible with high-spatial-resolution IMSexperiments, such as those performed using MALDI IMS. One of the key limitations to achieving high spatial resolution in MALDI IMS is the trade-off between spatial resolution and molecular sensitivity. As the laser beam diameter is reduced (e.g., via decreased fluence, improved focus, or oversampling), the sampled tissue area — and hence the amount of ablated material — also decreases. This results in lower signal intensities and the potential loss of low-SNR molecules, as they may fall below the detection threshold. In such scenarios, only high-intensity, high-SNR molecules remain detectable. These findings suggest that, even under these conditions, the virtual staining method could remain viable and effective, further underscoring its applicability.

[0205] Methods - IMS

[0206] Sample Preparation

[0207] Human kidney tissue 22 was surgically removed during a full nephrectomy, and remnant tissue was processed for research purposes by the Cooperative Human Tissue Network at Vanderbilt University7Medical Center. Remnant biospecimens were collected in compliance with the Cooperative Human Tissue Network standard protocols and the National Cancer Institute’s Best Practices for the procurement of remnant surgical research material. The study received ethical approval from Vanderbilt University7’s Institutional Review Board (IRB #210190). The study involved kidney specimens from 5 individual patients. Kidney tissue was flash-frozen over an isopentane-dry ice slurry, embedded in carboxymethylcellulose (CMC) and stored at -80 °C. The tissue was cryosectioned into 10-pm-thick sections using a CM3050 S cryostat (Leica Biosystems, Wetzlar, Germany). The sections were then thaw-mounted onto indium tin oxide (ITO) coated glass slides (Delta Technologies, Loveland, CO) for IMS analysis or regular glass slides for histological staining. Slides were stored at -80 °C and returned to ~20 °C within a vacuum desiccator prior to further processing. To remove endogenous salt for IMS, the section was washed three times with chilled (4 °C) 150 mM ammonium formate 3 times for 45 seconds each. It was then dried with nitrogen gas to remove excess moisture.

[0208] Autofluorescence microscopy and histochemical staining

[0209] To enable co-registration of IMS data with histochemically stained images, autofluorescence (AF) images were used as an intermediate modality, to which both IMS and histochemically stained images can be registered, as shown in FIG. 21 A (also see details in the Multimodal image registration section). AF microscopy images 60 were acquired on each tissue prior to IMS analysis using DAPI, eGFP and DsRed filters on a Zeiss AxioScan. Zlslide scanner (Carl Zeiss Microscopy GmbH, Oberkochen, Germany). Additionally, AF was also collected after IMS data acquisition prior to matrix removal, enabling visualization of ablation marks created by the MALDI laser. The resulting images have a pixel size of 0.65 pm.

[0210] It is worth noting that acquiring AF microscopy images 60 after the IMS data acquisition but before the chemical matrix is removed offers a clear visualization of the ablation craters for each IMS pixel in the microscopy coordinate space, thereby facilitating precise IMS-microscopy registration. During the MALDI process, the tissue surface including its chemical matrix layer is irradiated with a laser that causes the ablation and desorption of material and the ionization of endogenous molecules for mass spectrometry detection. Since the laser is focused on a diameter slightly smaller than that of an IMS pixel to avoid oversampling, this results in an array of distinct ablation marks on the matrix-coated tissue surface. Importantly, in most experiments, only the chemical matrix layer is ablated during MALDI, while the underlying tissue remains intact.

[0211] After the acquisition of the post-IMS autofluorescence images 60, the same unlabeled tissue was stained using a standard PAS staining protocol. The PAS staining process takes 30-60 minutes for each batch of processed slides. The stained tissue slides w ere scanned and digitized using a brightfield slide scanner (Leica Biosystems Aperio AT2). The resulting PAS-stained brightfield microscopy images have a pixel size of 0.22 pm.

[0212] Imaging mass spectrometry

[0213] Samples for IMS analysis w ere coated with 20 mg / mL solution of DAN dissolved in THF using a TM Sprayer M3 (HTX Technologies, LLC, Chapel Hill, NC, USA), yielding a 1.67 mg / cm2coating (0.05 mL / hr, 4 passes, 40 °C spray nozzle). MALDI IMS was performed on a prototype timsTOF fleX mass spectrometer (Bruker Daltonik, Bremen, Germany). The ion images were collected in negative ionization mode at 10 pm pixel size with a beam scan set to 6 pm, using 150 laser shots per pixel and 18.6% laser power (30% global attenuator and 62% local laser power) at 10 Hz. Data w ere acquired in negative ionization qTOF mode, covering an m / z range from 150 to 2000. The acquisition speed of IMS is ~5 hours / slide.

[0214] Multimodal image registration

[0215] The multimodal registration pipeline is illustrated in FIG. 21B. Due to the 10-fold difference in the pixel size and inherent differences in the imaging modalities, direct registration of the histochemically stained images with IMS data is not feasible. To addressthis challenge, prior-IMS and post-IMS AF images 60 are used to facilitate the registration process. The pipeline consists of three sequential image registration steps: (1) the first step aligns the histochemically stained images with the prior-IMS AF images 60; (2) next, the prior-IMS AF images 60 are registered to the post-IMS AF images 60; and (3) finally, the IMS data are registered to the post-IMS AF images 60. Notably, the third step leverages the laser ablation craters left on the chemical matrix after the IMS acquisition to guide image registration. However, this ablation process causes localized feature loss in the post-IMS AF images 60, making them unsuitable for direct alignment with the histochemically stained images. Therefore, the prior-IMS AF images sen e as an intermediary', as they share the same modality as the post-IMS AF images 60 and retain structural features needed for robust image registration. The microscopy modalities were co-registered using the elastix framework (https: / / github.com / SuperElastix / elastix which is incorporated herein by reference) integrated into wsireg software (https: / / github.com / NHPatterson / wsireg which is incorporated herein by reference). The post-IMS AF was selected as the target modality7to enable integration with IMS, with the PAS-stained brightfield microscopy images becoming the source modalities. The rigid and affine transformations were used since the microscopy modalities were collected on the same tissue section. All registered whole-slide images were stored in the vendor-neutral pyramidal OME-TIFF format at a common pixel size (i.e., the PAS image was resampled to the resolution of the prior-IMS AF and post-IMS AF). Furthermore, the MALDI IMS datasets were manually registered to the post-IMS AF images using the laser ablation marks and IMS pixels 62. The manual registration was performed using IMS Microlink software, where 8-12 fiducial markers were selected in both modalities to estimate the affine transformation.

[0216] Data division and preparation

[0217] The collected IMS data were exported from the Bruker timsTOF file format (.d) to a custom binary format for ease of access and improved performance. Each pixel / frame contains between 104and 105centroid peaks covering the entire acquisition range, which can be reconstructed into a pseudoprofile mass spectrum using Bruker’s SDK (v2.21). The dataset was m / z-aligned using six internally identified peaks (appearing in at least 50% of the pixels) through the msalign library7(v0.2.0). This step corrects spectral misalignment (drift along the m / z axis), resulting in increased overlap between spectral features (peaks) across the experiment. Subsequently, the mass axis of the data set was calibrated using the theoretical masses of the six peaks, achieving a precision of approximately ±1 ppm.Following the preprocessing steps, normalization correction factors were computed, and a total ion current (TIC) approach was used for mass spectral and ion image normalization. Subsequently, an average mass spectrum based on all pixels was calculated for each dataset. Since samples from multiple donors were used in this study, an average mass spectrum of all samples was generated and peak-picked, identifying 1,453 that were used for further analysis. The resulting average spectrum had a resolving power of -40,000 at m / z 885.55.

[0218] To better match the dimensions of the IMS data and facilitate the transformation from IMS to histochemically stained images, the histochemically stained images were downsampled to achieve a pixel size of 1 pm, which is ten times smaller than the pixel size of IMS. The whole slide images (WSIs) of IMS and histochemically stained images, acquired from four patients, were then segmented into smaller FOV pairs of approximately 1400x1400 pm2(each IMS image 20 with 140x140 pixels and each histochemically stained image with 1400x1400 pixels), with 10% overlap between neighboring regions. To augment the data for robust model training, the paired WSIs were spatially transformed using a combination of rotation and flipping. The transformed WSIs were segmented as described above. This process generated a training dataset containing 712 paired IMS-PAS microscopic image patches obtained from 4 de-identified patients. Additionally, 201 paired IMS-PAS microscopic image patches without data augmentation (each FOV with 64x64 pixels for IMS and 640x640 pixels for the histochemically stained image) were reserved for blind testing, obtained from ten de-identified patient not included in the training set. During each training epoch, the paired image FOVs were further subdivided.

[0219] Brownian bridge diffusion process

[0220] The processed dataset provided input-output image pairs, where x0G nkHxM / x3H Wcorresponds to histochemically stained image, and y0G HRioxiox 3represents the label-free IMS-measured ion images of the same tissue FOV. Prior to training the diffusion model, the IMS dataset y0was preprocessed by a shallow neural network 50 or gs, as shown in FIG. 2B. This shallow^ network consisted of two convolutional layers followed by a pixel-shuffle layer, which upscaled the image dimensions to match the target histochemically stained image, i.e., the corresponding ground truth image. Specifically, the processed output y G ntfyxivxsthe same dimension as the histochemically stained target image. This shallow network 50 can be denoted as:y = 9 s (To) (32)

[0221] Subsequently, the processed label-free image y and the corresponding histochemically stained image x0were used as the starting and ending points of the Brownian bridge, respectively, thereby forming an image-conditional diffusion process. In this Brownian bridge diffusion process, y acts as the image condition input, while x0represents the target image. The forward diffusion process of the Brownian bridge with an initial state x0and terminal state y is defined as:

[0222] where atand the variance 6tcan be calculated as:

[0223] Here, t refers to the intermediate time step, and T is the total sampling steps. T was set to 1000 during both the forward and reverse processes. Factor s here is used to control the sampling diversity and is set to 1 in the forw ard process. During the forward sampling process, the intermediate state xtcan be derived from Eq. (33) as:

[0224] As shown in Eq. (4), the target image x0can be recovered from any intermediate state xtby removing the add-on term at(y — x0) + 8^et. Therefore, a denoising network 52 or edwas also trained as part of the BBDM to estimate this add-on term to recover the target image x0. Since the term ut(y — x0) +is a function of time step t, the denoising network takes both the noisy image xtand the time step t as input. This network 52 is built upon a U-Net architecture, as shown in FIG. 11B. The U-Net structure features two connected parts: a downsampling path and an upsampling path, with a central layer and skip connections linking them. The downsampling path comprises four levels, each consisting of a residual block followed by an attention block, and adjacent levels are connected via 2x2 average pooling. In parallel, the upsampling path mirrors this structure but with 2x2 nearest-neighbor interpolation connecting the levels. At the core of the U-Net, the middle block incorporates two residual blocks and one attention block, utilizing multi-head attention for enhanced feature processing. Additionally, the time-step t is embedded through a linear layer with a SiLU activation function and integrated into the residual blocks as an added feature.

[0225] The loss function of the denoising network egis defined as:

[0226] where ytis the weight for each t. During the training, ytwas empirically set as 1 for all the sampling steps, t.

[0227] During the blind testing, the reverse sampling process of BBDM starts from the conditional label-free input, xT= y = gs(yo)- Based on the main idea of the denoising diffusion method, the reverse process aims to estimate xt-1from xt. Hence the vanilla reverse process can be defined as:

[0228] where Mt(xt, y) is the estimated mean relying on the trained denoising network, while the 8tis the additional variance added during the sampling process. The expressions of cxt,cyt>cet-, and Sfit-! can be formulated as:_ < i (1 — nt) 8t t-xt 6tr (il - a nt-i A) d" 6 s't(1 ^t-i) (40)

[0229] During the blind testing phase, the reverse process of the diffusion model predicts the target image x0starting from xT= y. The vanilla sampling strategy operates as a Markov chain with learned Gaussian transitions, initiating from p(xr) = J\f (xT; y, 0). This process can be expressed as:

[0230] Each transition probability p(xt\xt, y) estimates the intermediate statext-1from the previous state xtand can be formulated as:

[0231] By iterating Eq. (45), each intermediate state is sequentially estimated, starting from xTand continuing until the final target image x0is obtained, revealing the virtuallystained PAS image of the label-free test tissue. This iterative process is referred to as the vanilla sampling strategy’ and is depicted in FIGS. 2D, 2H.

[0232] For the mean sampling strategy, shown in FIGS. 2E, 21, the procedure initially mirrors the vanilla sampling approach up until the time step te< T, where teis defined as the exit point in the reverse sampling process. Once the process reaches t < te. this strategy estimatest-1fromxtusing a mean sampling step as follows:Pm(xt—ilXt, y) = xt-1nt' (ixt,y'), 0') (46)

[0233] where the estimation of xt-1relies solely on nt' xt, y). without incorporating additional variance 6t, i.e.. xt-= cxtxt+ cyty — ceteext, t) for t < te. Thus, the mean sampling strategy can be formulated as:

[0234] The skip sampling strategy follows the same vanilla sampling steps up to an exit point t = te, i.e.

[0235] However, instead of continuing the iterative inference process, this skip sampling method estimates the final virtual stained image directly’ from the state at xte

[0236] During the training phase, both the shallow convolutional network and the denoising network were trained jointly. The parameters of these networks were optimized using the AdamW optimizer with an initial learning rate of 1 X 10-4. The training process used a batch size of 16 and took approximately 72 hours to converge. For blind testing, the exit point, tewas empirically set to 10, where both the mean and skip sampling strategies achieved better performance, as illustrated in FIGS. 19A-19B.

[0237] Quantitative performance evaluation metrics

[0238] To quantitatively evaluate the performance of PAS virtual staining results for the study reported in FIGS. 14A-14E and 17A-17B. 201 FOVs of virtually stained PAS images were used together with their corresponding histochemically stained images for paired image comparisons. In the quantitative evaluation illustrated in FIGS. 14A-14E, a comprehensive comparative analysis was conducted using several features: image contrast, histogram distributions in the YCbCr color space, Hellinger distance calculated between the colorhistograms of the VS and HS image pairs, and spatial frequency spectrum analysis. The image contrast is defined as:Contrast

[0239] where figov, andrepresents the 90thpercentile and 10thpercentile of the intensity values in image A, respectively. In this case, image A refers to the grey-scale version of the VS or HS image. For the color analysis, the paired VS and HS images were converted from RGB to YCbCr color space to facilitate a detailed comparison of the distributions / histograms in the Y, Cb, and Cr channels, performed separately. The Hellinger distances were calculated between the YCbCr color histograms for each VS and GT image pair as follows:Hellinger distance

[0240] where p(x) and q(%) are the probability distributions of the YCbCr color (approximated using the extracted color distribution points). The Hellinger distance ranges between 0 (identical) and 1 (completely different). As for the spatial frequency spectrum analysis, the single channel raw MS image was bilinearly upsampled by a factor of 10, from 64x64 pixels to 640x640 pixels, matching the dimensions of the grey-scale VS and HS images. The frequency spectrum of each image was obtained by performing a two-dimensional (2D) Fourier Transform on the 10x bilinearly upsampled single-channel IMS image 20 (selected based on the contrast of glomeruli), the VS output and the corresponding HS ground truth image. For this analysis, both the VS and HS images were processed in greyscale. The radially averaged power spectrum was calculated according to Wang et al.Moreover, the relative error between the radially averaged power spectra (log scale, normalized) of the VS image and their HS ground truth (FIG. 14E) is defined as:relative err

[0241] where Pvsand PHSdenote the radially averaged power spectra (log scale, normalized) of the VS image and its HS counterpart. Similarly, the relative gain of the radially averaged power spectra (log scale, normalized) of the V S images compared to their lower resolution MS input (FIG. 14E) is defined as:relative gain

[0242] where Pvsand PMSrepresent the radially averaged power spectra (log scale, normalized) of the VS image and the corresponding single-channel MS input.

[0243] For the evaluations reported in FIGS. 17A-17B, the metrics of PSNR, PCC and LPIPS w ere used. The PSNR is defined based on mean squared error (MSE):

[0244] where A and B present the histochemically and virtually stained brightfield PAS images, respectively, m, n are the pixel indices, and MN denotes the total number of pixels in each image. PSNR can be denoted as:

[0245] where max ( / I) is the maximum pixel value of the ground truth histochemically stained PAS image.

[0246] The Pearson correlation coefficient (PCC) was used to quantify the linear correlation between the histochemically stained PAS images and the corresponding VS images. PCC is calculated as:

[0247] where A and B represent the mean pixel value of A and B.

[0248] The calculation of the LPIPS metrics utilized a pre-trained VGG network to evaluate the learned perceptual similarity between the generated VS images n and their corresponding HS images n0. These compared image pairs were fed into the pre-trained VGG network and their feature stack from L layers can be extracted as in1, inQlG IEHi*wtxCi for layer I. The LPIPS score can be denoted as:The Frechet inception distance (FID) values, reported in FIGS. 19A-19B, are calculated as follows:

[0249] where N (fi, Z) represents the multivariate normal distribution estimated from the Inception v3 features of the ground truth PAS-stained images and V( / zw, Sw) represents the distribution estimated from the Inception v3 features of the generated virtually stainedimages. The 36 VS-HS image pairs (each with 640 x 640 pixels) were used as input to the Inception v3 network to extract 2048-dimensional feature vectors. From these extracted features, the mean vectors Qi, / / ^) and covariance matriceswere estimated.

[0250] The Naturalness Image Quality Evaluator (NIQE) values reported in FIGS.19A019B, are calculated as:

[0251] where Viand 2^ represent the mean vector and covariance matrix of the customized reference Multivariate Gaussian (MVG) model, which was fitted using 800 ground truth PAS image patches (each with 640 x 640 pixels) sampled from both the training and testing datasets. Meanwhile, v2and S2arethe mean vector and covariance matrix for the MVG model of a generated virtually stained image. The MATLAB functions fitniqe and niqe were used to fit the reference MVG model and compute the final NIQE values for the virtually stained images.

[0252] Statistical analysis

[0253] In FIGS. 14A-14E, a two-tailed paired / -test was conducted to assess whether the image contrast between the virtually stained output and its histochemically stained PAS counterpart was statistically equivalent, with a statistical significance level of 0.05. This analysis was performed on 201 paired VS and HS images. A p-value greater than 0.05 indicates no statistically significant difference in the image contrast betw een the VS images and their histochemically stained counterparts.

[0254] Other implementation details

[0255] All netw ork training and testing tasks were conducted on a desktop computer 100 equipped with an Intel Core i9-13900K CPU processor 102, 64 GB of memory, and an NVIDIA GeForce RTX 4090 GPU processor 102. The code for training the diffusion models was developed in Python 3.9.19 using PyTorch 2.2.1.

[0256] While embodiments of the present invention have been shown and described, various modifications may be made without departing from the scope of the present invention. The invention, therefore, should not be limited, except to the following claims, and their equivalents.

Claims

What is claimed is:

1. A method of generating one or more high resolution, virtually stained microscopic images of a label-free test sample comprising:providing a trained diffusion-based image inference model that is executed by one or more processors of a computing device, wherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s);obtaining one or more autofluorescence images of the label-free test sample using a fluorescence microscope;inputting the one or more autofluorescence images of the label-free test sample to the trained diffusion-based image inference model; andthe trained diffusion-based image inference model outputting one or more virtually stained microscopic images of the label-free test sample with high resolution that are substantially equivalent to corresponding high resolution microscopic image(s) of the same label-free test sample that has been chemically stained, wherein the trained diffusion-based image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

2. The method of claim 1, wherein the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 8tat the intermediate denoising steps from T to te and wherein the variance 8tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0.

3. The method of claim 1, wherein the plurality7of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with avariance of 6tat the intermediate denoising steps from T to teand wherein the variance 6tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0.

4. The method of claim 1, wherein the resolution of the autofluorescence images of label-free test samples is lower than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

5. The method of claim 1, wherein the plurality training images and / or the one or more autofluorescence images of the label-free training sample(s) and / or label-free test sample(s) obtained using the fluorescence microscope are input through a shallow neural network for processing of information.

6. The method of claim 1, wherein the trained diffusion-based image inference model comprises a convolutional neural network architecture.

7. The method of claim 1, wherein the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.

8. A system for generating one or more high resolution, virtually stained microscopic images of a label-free test sample comprising:a computing device having one or more processors configured to execute a trained diffusion-based image inference model that receives one or more autofluorescence images of the label-free test sample and outputs one or more virtually stained microscopic images of the label-free test sample with high resolution that are substantially equivalent to corresponding high resolution microscope image(s) of the same label-free test sample that has been chemically stained; andwherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding autofluorescence images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s) and wherein the trained diffusionbased image inference model generates the one or more virtually stained images using areverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

9. The system of claim 8, wherein the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0.

10. The method of claim 8, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of δ̃tat the intermediate denoising steps from T to teand wherein the variance δ̃tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0.

11. The system of claim 8, further comprising a fluorescence microscope.

12. The system of claim 8, wherein the one or more processors comprise one or more graphics processing units (GPUs).

13. The system of claim 8, wherein the resolution of the autofluorescence images of label-free test samples is lower than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

14. The system of claim 8, further comprising a shallow neural network, wherein the plurality training images and / or the one or more autofluorescence images of the label-free training sample(s) and / or label-free test sample(s) obtained using the fluorescence microscope are input through the shallow neural network for processing of information.

15. The system of claim 8, wherein the trained diffusion-based image inference model comprises a convolutional neural network architecture.

16. The system of claim 8, wherein the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.

17. A method of generating one or more virtually stained microscopic images of a label-free test sample comprising:providing a trained diffusion-based image inference model that is executed by one or more processors of a computing device, wherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding imaging mass spectrometry (IMS) images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s);obtaining one or more IMS images of the label-free test sample using an imaging mass spectrometer;inputting the one or more IMS images of the label-free test sample to the trained diffusion-based image inference model; andthe trained diffusion-based image inference model outputting one or more virtually stained microscopic images of the label-free test sample with enhanced spatial resolution that are substantially equivalent to corresponding high resolution microscopic image(s) of the same label-free test sample that has been chemically stained, wherein the trained diffusionbased image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

18. The method of claim 17, wherein the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 6tat the intermediate denoising steps from T to teand wherein the variance 8tisomitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point te to step 0.

19. The method of claim 17, wherein the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point te within the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 6tat the intermediate denoising steps from T to teand wherein the variance 8tis omitted at the exit point teand the trained diffusion model inference skips directly to step 0.

20. The method of claim 17, wherein the resolution of the IMS images of label-free test samples is lower than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

21. The method of claim 17, wherein the plurality training images and / or the one or more IMS images of the label-free training sample(s) and / or label-free test sample(s) obtained using the imaging mass spectrometer are input through a shallow neural network for processing of information.

22. The method of claim 17, wherein the trained diffusion-based image inference model comprises a convolutional neural network architecture.

23. The method of claim 17, wherein the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.

24. A system for generating one or more virtually stained microscopic images of a label-free test sample comprising:a computing device having one or more processors configured to execute a trained diffusion-based image inference model that receives one or more imaging mass spectrometry (IMS) images of the label-free test sample and outputs one or more virtually stained microscopic images of the label-free test sample with enhanced resolution that are substantially equivalent to corresponding high resolution microscope image(s) of the same label-free test sample that has been chemically stained; andwherein the trained diffusion-based image inference model is trained with a plurality training images comprising matched chemically stained images or image patches and their corresponding IMS images or image patches of the same training sample(s) and wherein the trained diffusion-based image inference model learns virtual staining of the microscopic images of the label-free training sample(s) and wherein the trained diffusion-based image inference model generates the one or more virtually stained images using a reverse sampling process that includes a plurality of intermediate denoising steps as part of the diffusion model inference.

25. The system of claim 24, wherein the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point tewithin the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 8tat the intermediate denoising steps from T to te and wherein the variance 8tis omitted from a portion of the plurality of intermediate denoising steps within the range spanning the exit point teto step 0.

26. The method of claim 24, the plurality of intermediate denoising steps spans steps within a range of T to step 0 and includes an exit point te within the range of T to step 0 and wherein during the reverse sampling process additional noise is introduced with a variance of 8tat the intermediate denoising steps from T to teand wherein the variance 8tis omitted at the exit point te and the trained diffusion model inference skips directly to step 0.

27. The system of claim 24, further comprising an imaging mass spectrometer.

28. The system of claim 24, wherein the one or more processors comprise one or more graphics processing units (GPUs).

29. The system of claim 24, wherein the resolution of the IMS images of label-free test samples is lower than the high-resolution microscopic images of the same label-free test samples that have been chemically stained.

30. The system of claim 24, further comprising a shallow neural network, wherein the plurality training images and / or the one or more IMS images of the label-free training sample(s) and / or label-free test sample(s) obtained using an imaging mass spectrometer are input through the shallow neural network for processing of information.

31. The system of claim 24, wherein the trained diffusion-based image inference model comprises a convolutional neural network architecture.

32. The system of claim 24, wherein the trained diffusion-based image inference model outputs one or more virtually stained microscopic images of the label-free test sample with reduced variance.

Citation Information

Patent Citations

  • Method and system for digital staining of microscopy images using deep learning

    US20230030424A1

  • Video generation with latent diffusion probabilistic models

    US20240087179A1

  • Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models

    WO2024187031A2