Method for virtual stain and computational pathology based on label-free tissue

HK40135096APending Publication Date: 2026-07-17THE HONG KONG UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
THE HONG KONG UNIV OF SCI & TECH
Filing Date
2026-04-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Current techniques require physical staining of tissue samples, which is then interpreted by pathologists. This process is time-consuming, labor-intensive, and destructive to the tissues. Furthermore, the limited number of pathologists makes it impossible to process large volumes of digital WSIs.

Method used

By acquiring the AF WSI of tissue samples, using a virtual staining ML model to generate virtual H&E stained WSI, and combining it with a MIL network for tumor diagnosis, automated tumor diagnosis can be achieved without physical staining and interpretation by pathologists.

Benefits of technology

It enables rapid and accurate tumor diagnosis of tissue samples without the need for actual staining processes and pathologist interpretation, reducing time and labor costs and improving the efficiency of processing large-scale WSIs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A machine learning (ML) workflow for virtual staining and computational pathology based on label-free tissue samples is provided to avoid the cumbersome and time-consuming processes that need to be consumed by pathologists in interpreting formalin fixation and paraffin embedding histology. A virtual staining ML model converts a label-free autofluorescence (AF) full-slide image (WSI) of the tissue sample into virtual hematoxylin and eosin (Hamp; e), the virtual hematoxylin and eosin (Hamp; e) further converting the dyed WSI into a WSI which is virtually dyed with a special dye by means of a dye conversion ML model. Each of the ML models may be implemented as a U-Frame model. A multi-instance learning (MIL) network with attention-based pooling may be applied to the virtual Hamp; e-stained WSI or AF WSI to provide rapid and accurate tumor diagnosis. Further, the attention-based MIL network may be configured to cluster constrained attention multi-instance learning (CLAM) models for providing multiple classes of tumor subtypes or even multiple classes of disease types in the diagnosis of the tissue sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 510,366, filed June 27, 2023, the disclosure of which is incorporated herein by reference in its entirety.

[0003] List of abbreviations

[0004] AF autofluorescence

[0005] AI (Artificial Intelligence)

[0006] CLAM clustering constrained attention multi-instance learning

[0007] CNN (Convolutional Neural Network)

[0008] DCIS ductal carcinoma in situ

[0009] DSMIL Dual Stream MIL

[0010] EVG elastic fiber dyeing

[0011] FFPE formalin fixation and paraffin embedding

[0012] Generative Adversarial Networks (GANs)

[0013] H&E Hematoxylin and Eosin

[0014] LUAD (Lung Adenocarcinoma)

[0015] MIL Multi-Instance Learning

[0016] MIL-RNN with RNN

[0017] ML Machine Learning

[0018] MT Masson Tricolor

[0019] NA (Numerical Aperture)

[0020] NSCLC (Non-small Cell Lung Cancer)

[0021] PAM photoacoustic microscope

[0022] QPI quantitative phase imaging

[0023] RNN (Recurrent Neural Network)

[0024] SRH stimulated Raman histology

[0025] UV ultraviolet rays

[0026] WSI all-slide image Technical Field

[0027] This disclosure generally relates to the histological examination of tissue samples. Specifically, this disclosure relates to performing virtual staining on tissue samples without requiring physical staining, and to performing tumor diagnosis on tissue samples without first requiring a real histochemical staining process to stain the tissue sample, and secondly without requiring a pathologist to interpret the WSI of the histochemical staining of the tissue sample. Background Technology

[0028] The routine workflow for histopathological examination requires interpretation of histochemically stained slide images (more commonly H&E stained images) by an experienced pathologist. However, the staining process is time-consuming, laborious, and destructive to tissue. In recent years, various label-free image-based virtual staining methods have been developed to replace the actual histochemical staining process, which significantly saves reagent and time costs and preserves the tissue for further analysis. FFPE histology remains the gold standard for postoperative diagnosis. This process requires high-quality slide preparation but is also lengthy and laborious, typically taking 3 to 5 days [Reference 1]. For rapid intraoperative assessments sometimes required during surgery, frozen section histology is widely used within 30 minutes. However, image quality with freezing artifacts is unsatisfactory and can affect diagnosis [Reference 2]. Various advanced slide-free and label-free imaging techniques have been developed to meet the needs of rapid histology, including AF microscopy [Reference 3], QPI [Reference 4], PAM [Reference 5], etc.

[0029] Various GAN-based models [Reference 6] have been designed for virtual H&E staining on unlabeled thin tissue sections and unprocessed thick tissues, including: supervised GAN models [Reference 3], [Reference 4], [Reference 7], weakly supervised GAN models [Reference 8], and unsupervised GAN models [Reference 9]-[Reference 11]. In addition to standard H&E staining, various specialized staining methods are widely used to assess a wide range of diseases, though these require additional time and cost. For example, EVG staining can highlight elastic fibers in connective tissue for diagnosing vascular diseases, and MT staining can visualize collagen fibers for diagnosing non-neoplastic diseases. Supervised GAN models are widely used for virtual specialized staining on thin tissue sections. Virtual specialized staining can be derived from unlabeled tissue sections [Reference 4], [Reference 7], or H&E-stained tissue sections [Reference 12].

[0030] Despite the success of various virtual staining methods, the interpretation of histological images, primarily H&E-stained images, still requires expertise that can currently be performed by pathologists. Due to the limited number of trained pathologists globally, pathologists face a significant burden in processing the vast amount of digital WSIs generated annually. Advances in deep learning-based computational pathology have enabled automated tumor diagnosis with pathologist-level accuracy and interpretability. Fully supervised models trained using patch-level annotations of WSIs have achieved promising results on various types of tumor datasets, including lymph node metastases of breast cancer [Reference 13], non-small cell lung cancer subtypes [Reference 14], and brain tumor subtypes [Reference 15]. Common supervised models such as VGG16 [Reference 16] and Inception v3 [Reference 17] require millions of patches labeled by experienced pathologists. The labeling task is labor-intensive, costly, and time-consuming. To address this issue, a weakly supervised model using MIL with only slide-level annotations has been employed for WSI classification. Deep neural networks combined with different MIL pooling strategies have been applied to tumor diagnosis, including attention-based deep MIL [Reference 18], MIL-RNN [Reference 19], and DSMIL [Reference 20]. The probability or attention score of the patch can be visualized as a heatmap to locate the tumor tissue region of the WSI.

[0031] In light of the foregoing observations, there is a need in the field to perform tumor diagnosis on tissue samples without performing a real histochemical staining process and without requiring a pathologist to interpret the histochemical staining of the tissue sample via WSI. Summary of the Invention

[0032] A first aspect of this disclosure is to provide a computer-implemented first method for virtually staining tissue samples.

[0033] The first method includes: obtaining an AF WSI of the tissue sample; and generating a virtual H&E stained WSI from the AF WSI using a virtual staining ML model, such that the tissue sample is virtually H&E stained to form a virtual H&E stained WSI. Advantageously, the virtual H&E stained WSI is obtained without performing a real histochemical staining process of staining the tissue sample with H&E dye.

[0034] In some embodiments, the first method further includes using a staining transformation ML model to convert a virtual H&E-stained WSI into a virtual special-stained WSI, such that the tissue sample is virtually stained with a special stain other than H&E staining to form a virtual special-stained WSI. Advantageously, the virtual special-stained WSI is obtained without physically staining the tissue sample with a special stain.

[0035] Virtual staining ML models can be implemented as U-Frame models. Staining transformation ML models can also be implemented as U-Frame models. Special staining can be MT staining.

[0036] In some embodiments, the first method further includes training the virtual staining ML model before using the virtual staining ML model to generate a WSI of virtual H&E staining.

[0037] In some embodiments, a virtual staining ML model is trained using a first training dataset. The first training dataset is prepared depending on whether thin tissue slices or thick tissue are used as tissue samples.

[0038] In some embodiments, the first method further includes training the staining transformation ML model before using the staining transformation ML model to generate a virtual special staining WSI.

[0039] In some embodiments, a second training dataset is used to train the staining transformation ML model. The second training dataset is prepared based on whether thin tissue slices or thick tissue are used as tissue samples.

[0040] A second aspect of this disclosure is to provide a second method for virtually staining tissue samples.

[0041] The second method includes: generating an AF WSI of a tissue sample using a fluorescence microscopy system; and performing a process of virtually staining the tissue sample according to any embodiment of the first method by one or more computers.

[0042] A third aspect of this disclosure is to provide a computer-implemented third method for performing tumor diagnosis on tissue samples.

[0043] The third method includes: acquiring an AF WSI of a tissue sample; selecting and acquiring a working WSI for tumor diagnosis, wherein the working WSI is selected from the AF WSI of the tissue sample and a virtually stained WSI, and wherein the virtually stained WSI is generated from the AF WSI; trimming the working WSI to retain one or more tissue regions of the working WSI; dividing the one or more tissue regions into multiple tissue patches; and using a tumor diagnosis ML model to process the corresponding tissue patches among the multiple tissue patches to diagnose any tumor in the working WSI, such that tumor diagnosis of the tissue sample is performed without first performing a real histochemical staining process on the tissue sample and without requiring a pathologist to interpret the histochemically stained WSI of the tissue sample.

[0044] In some embodiments, the tumor diagnostic ML model is implemented as a MIL network. The MIL network includes: a backbone for extracting multiple patch-level features from each tissue patch in a plurality of tissue patches; and a MIL pooling module implemented using MIL pooling methods to aggregate the corresponding multiple patch-level features extracted from the multiple tissue patches and predict a score indicating the probability of tumor presence observed on the working WSI.

[0045] In some embodiments, attention-based pooling is used as the MIL pooling method, making the MIL network an attention-based MIL network.

[0046] In some embodiments, the attention-based MIL network is configured as a CLAM model. The CLAM model can be configured to provide multi-tumor subtyping in tumor diagnosis of tissue samples. The CLAM model can also be configured to provide multi-disease subtyping in the diagnosis of tissue samples.

[0047] In some embodiments, max pooling or average pooling is used as the MIL pooling method.

[0048] In some embodiments, the backbone is implemented as a CNN. The backbone can also be implemented as a ResNet50 model.

[0049] In some embodiments, the backbone is implemented as a visual converter.

[0050] In some embodiments, the backbone is implemented as a diffusion-based model.

[0051] In some embodiments, the tumor diagnosis ML model is implemented as a weakly supervised ML model.

[0052] In some embodiments, the tumor diagnostic ML model is implemented as a supervised ML model.

[0053] In some embodiments, the third method further includes, after diagnosing one or more tumors in the working WSI, using a tumor diagnostic ML model to perform prognostic prediction of patient outcome probabilities. Patient outcome probabilities may include recurrence probability, patient survival probability, and other outcomes following surgery and / or drug treatment.

[0054] In some embodiments, the third method further includes, after diagnosing one or more tumors in the working WSI, using a tumor diagnostic ML model to perform mutation prediction of the patient's genes.

[0055] In some embodiments, the third method further includes training the MIL network before the tumor diagnostic ML model is used to process the corresponding tissue patches. In the case of training the tumor diagnostic ML model with a third training dataset, the third training dataset can be prepared depending on whether thin tissue slices or thick tissue are used as tissue samples.

[0056] In some embodiments, the virtually stained WSI is selected as the virtual H&E stained WSI of the tissue sample, wherein the virtual H&E stained WSI is obtained by virtually stained the tissue sample with H&E staining. Obtaining a working WSI includes obtaining a virtually stained WSI, wherein obtaining a virtually stained WSI includes generating a virtual H&E stained WSI from an AF WSI using a virtual staining ML model. In some embodiments, the virtual staining ML model is implemented as a U-Frame model.

[0057] In some embodiments, the virtually stained WSI is selected as the virtual special staining WSI of the tissue sample, wherein the virtual special staining WSI is obtained by virtually staining the tissue sample with a special stain other than H&E staining. Obtaining the working WSI includes obtaining the virtually stained WSI. Obtaining the virtually stained WSI includes: generating a virtual H&E staining WSI from the AF WSI using a virtual staining ML model, wherein the virtual H&E staining WSI is obtained by virtually staining the tissue sample with H&E staining; and converting the virtual H&E staining WSI into a virtual special staining WSI using a staining conversion ML model. In some embodiments, the special staining is MT staining. In some embodiments, each of the virtual staining ML model and the staining conversion ML model is implemented as a U-Frame model.

[0058] A fourth aspect of this disclosure is to provide a fourth method for performing tumor diagnosis on tissue samples.

[0059] The fourth method includes: generating an AF WSI of a tissue sample using a fluorescence microscopy system; and performing a tumor diagnosis on the tissue sample based on the generated AF WSI by one or more computers according to any embodiment of the third method.

[0060] Other aspects of this disclosure are disclosed as illustrated in the following examples. Attached Figure Description

[0061] Figure 1 A deep learning-based workflow for virtual staining and computational pathology is described.

[0062] Figure 2An exemplary computational pathology workflow using an attention-based MIL model is described, in which... Figure 2 In the middle: Subfigure a depicts a tumor diagnostic model based on AF images; Subfigure b depicts a tumor diagnostic model based on H&E staining images; Subfigure c depicts attention heatmaps predicted directly on AF images by a well-trained MIL model or on virtual H&E staining images generated by a U-Frame model.

[0063] Figure 3 Histological images of human lung cancer tissue sections are depicted, in which... Figure 3 In the middle: Subfigure a depicts the AF image of a section of resected human lung tissue; Subfigure b depicts the baseline truth MT staining image of the same section; Subfigure c depicts the virtual H&E staining image generated from the AF image of subfigure a by a virtual staining model; Subfigure d depicts the virtual MT staining image generated from the virtual H&E staining image of subfigure c by a staining conversion model; Subfigure e depicts the baseline truth H&E staining image of the same section; Subfigure f depicts the virtual MT staining image generated from the baseline truth H&E staining image of subfigure e by a staining conversion model.

[0064] Figure 4 Histological images of lung adenocarcinoma tissue sections are depicted, in which... Figure 4 In the middle: Subfigure a depicts an AF image of a resected human lung tissue; Subfigure b depicts a virtual H&E stained image generated from the AF image of subfigure a using a U-Frame model; Subfigure c depicts a baseline truth H&E stained image of the same slice; Subfigures d to f depict magnified images of the first rectangular region 410 marked in subfigures a to c, respectively; and Subfigures g to i depict magnified images of the second rectangular region 420 marked in subfigures d to f, respectively.

[0065] Figure 5 The illustration shows the localization and visualization of tumor regions on a slice of human lung, where... Figure 5 In the middle: Subfigure a illustrates the annotation of the tumor region on the baseline truth H&E stained image by a pathologist; Subfigure b illustrates the full slide attention heatmap generated from the virtual H&E stained image by the CLAM model; Subfigure c illustrates the full slide attention heatmap generated from the baseline truth H&E stained image by the CLAM model; Subfigures d to f depict magnified images of the first rectangular region 511 in subfigures a to c, respectively; Subfigures g to i depict magnified images of the second rectangular region 512 in subfigures a to c, respectively.

[0066] Figure 6 Histological images of breast cancer biopsy sections are depicted, in which... Figure 6In the middle: Subfigure a depicts an AF image of human breast biopsy tissue; Subfigure b depicts a virtual H&E stained image generated from the tissue in subfigure a using a U-Frame model; Subfigure c depicts a baseline truth H&E stained image of the same slice; Subfigures d to f depict magnified images of the first rectangular region 610 in subfigures a to c, respectively; Subfigures g to i depict magnified images of the second rectangular region 620 in subfigures d to f, respectively.

[0067] Figure 7 The illustration shows the localization and visualization of tumor regions on a human breast tissue slice, where... Figure 7 In the middle: Subfigure a depicts the tumor region annotation 730 on the baseline truth H&E stained image by a pathologist; Subfigure b depicts the full slide attention heatmap generated from the virtual H&E stained image by the CLAM model; Subfigure c depicts the full slide attention heatmap generated from the baseline truth H&E stained image by the CLAM model; Subfigures d to f depict magnified images of the first rectangular region 711 in subfigures a to c, respectively; Subfigures g to i depict magnified images of the second rectangular region 712 in subfigures a to c, respectively.

[0068] Figure 8 The illustration shows the localization and visualization of tumor regions on a thick human lung sample, where... Figure 8 Subfigure a depicts an AF image of a thick human lung resection tissue; subfigure b depicts a full-slide attention heatmap generated from the thick AF image using a CLAM model; subfigure c depicts a reference H&E stained image of an adjacent thin section with pathologist annotations; subfigures d through f depict magnified images of the first rectangular region 811 in subfigures a through c, respectively; subfigures g through i depict magnified images of the second rectangular region 812 in subfigures a through c, respectively; subfigure j depicts AF images of adjacent thin human lung sections; subfigure k depicts a full-slide attention heatmap generated from the thin AF image using a CLAM model; subfigure l depicts a reference truth H&E stained image of a thin section with pathologist annotations; subfigures m through o depict magnified images of the third rectangular region 813 in subfigures j through l, respectively; and subfigures p through r depict magnified images of the fourth rectangular region 814 in subfigures j through l, respectively.

[0069] Figure 9 It provides a comparison of heatmaps generated by different workflows, where... Figure 9 In the middle: Subfigure a depicts a pathologist's annotation of a tumor on a baseline truth H&E stained image; Subfigure b depicts an attention heatmap generated by an AF-based diagnostic workflow; Subfigure c depicts an attention heatmap generated by a virtual H&E-based diagnostic workflow.

[0070] Figure 10A first exemplary workflow according to a first method, implemented by a computer and disclosed herein, is described for virtually staining a tissue sample with a special stain to produce a virtual special staining of the tissue sample (WSI).

[0071] Figure 11 A second exemplary workflow is described, which incorporates a first exemplary workflow, according to a second method for virtually staining tissue samples disclosed herein.

[0072] Figure 12 A third exemplary workflow for performing tumor diagnosis on tissue samples according to a third method is described, which is implemented by a computer and disclosed herein.

[0073] Figure 13 An exemplary structure of a MIL network used as a tumor diagnostic ML model in a third exemplary workflow is depicted for processing working WSIs to perform tumor diagnosis on tissue samples.

[0074] Figure 14 A fourth exemplary workflow is described, which incorporates a third exemplary workflow, according to a fourth method disclosed herein for performing tumor diagnosis on tissue samples.

[0075] Those skilled in the art will understand that the elements in the accompanying drawings are illustrated for simplicity and clarity and are not necessarily depicted to scale. Detailed Implementation

[0076] This disclosure develops an automated tumor diagnosis method for tissue samples, which does not perform a real histochemical staining process on the tissue samples and does not require a pathologist to interpret the WSI of the histochemical staining of the tissue samples. In particular, the development of automated tumor diagnosis technology is simplified by unifying different virtual staining and computational pathology.

[0077] To unify different virtual staining and computational pathology approaches, this disclosure proposes a novel deep learning-based method using virtual H&E images of label-free tissue. Label-free AF ​​images, exhibiting negative nuclear contrast under deep UV excitation, can be further transformed into virtual H&E stained images using a virtual staining network. Virtual H&E staining can be further transformed into other specialized staining methods using a staining transformation network. The virtual staining and staining transformation networks can have the same model architecture but different staining data for model training. The proposed virtual staining workflow was validated on human lung cancer tissue using both virtual H&E and MT staining. For tumor classification and localization on virtual H&E stained images or potential AF images, workflows based on virtual H&E diagnosis and AF diagnosis have been developed using an attention-based MIL network trained on corresponding large-scale WSI datasets with the same image types. The versatility of the proposed computational pathology workflow was experimentally demonstrated on thin sections of human lung resection and human breast biopsy tissue.

[0078] This general virtual staining and computational pathology approach has been demonstrated using deep learning models on thin tissue sections, and it can potentially be extended to thick, unprocessed tissue samples. Virtual H&E staining images of unprocessed thick tissue using an advanced unsupervised virtual staining model [Reference 11] can potentially be applied to downstream tasks such as specialized staining and tumor diagnosis. Similar ideas have been demonstrated in existing work, such as CNN-based diagnosis of SRH images of unprocessed surgical specimens for intraoperative diagnosis [Reference 15]. Furthermore, AF images of unprocessed thick tissue can be directly used for tumor diagnosis, as demonstrated on a small dataset of thick lung cancer samples. Therefore, it is believed that the workflow proposed in this application has the potential to be generalized for rapid and accurate intraoperative and postoperative pathological examination.

[0079] A. Methods and Materials

[0080] A.1. General Work Process

[0081] Figure 1A general deep learning-based workflow 100 for virtual staining and computational pathology of label-free tissues is described. For virtual staining of label-free AF ​​WSI, this application employs a weakly supervised GAN-based model, namely the U-Frame model [Reference 8], which does not require precise image registration and exhibits better performance compared to fully supervised methods. This application also employs the U-Frame model to achieve the conversion from virtual H&E to other virtual special stains. This application uses an attention-based MIL network as the preferred ML model to directly achieve tumor classification and localization on virtual H&E-stained WSI or AF WSI. Besides the weakly supervised attention-based MIL network, other weakly supervised or even supervised ML models can also be used for tumor classification and localization. Nevertheless, as mentioned above, using a weakly supervised machine learning ML model has practical advantages over a supervised machine learning model.

[0082] A.2. Computational Pathology

[0083] Computational pathology workflows for tumor diagnosis can have two variations: workflows based on virtual H&E diagnosis and workflows based on AF diagnosis. Figure 2 (Figure c) They can achieve the same results as standard H&E histology interpreted by pathologists ( Figure 2 (d) Comparable tumor classification and localization.

[0084] Given a large number of open-source H&E staining images, this application uses open-source H&E staining images and inferences from virtual H&E staining images of AF images to train an H&E diagnostic model 124. Figure 2 Sub-images b and c). This virtual H&E-based diagnostic workflow consists of two deep learning models for virtual staining and tumor diagnosis. Virtual staining model 115 converts AF image 110 into a virtual H&E staining image 120 equivalent to FFPE H&E histology, which can be further predicted by a well-trained H&E diagnostic model 124. To evaluate the performance of the diagnostic model 124, slide-level predictions can be compared with diagnoses from pathologists, and heatmaps with suspicious tumor areas generated by the H&E diagnostic model 124 can be compared. Figure 2 Subfigure c) and the baseline truth H&E staining image already interpreted by pathologists ( Figure 2 Comparing it with sub-figure d). To demonstrate the idea of ​​performing artificial intelligence diagnosis on label-free tissue sections, this application also designs a workflow based on AF diagnosis ( Figure 2 Subfigure c), which uses the open-source H&E WSI and combines it with 13 other AF WSIs and corresponding slide-level markings ( Figure 2Subgraph a) was used to train the AF diagnostic model 125. To demonstrate the feasibility of the AF-based diagnostic workflow for thick tissues, this application also trained two AF diagnostic models 125 purely on a small-scale thick AF WSI dataset and a neighboring thin AF WSI dataset. These two AF diagnostic models showed comparable performance in tumor localization on attention heatmaps. Figure 8 ).

[0085] For tumor diagnosis tasks, the standard MIL method can be used for binary classification of tumors and normal tissue. Each WSI can be considered as a package, and patches cut from a WSI can be considered as instances. Consider each package X for n = 1, ..., N and m = 1, ..., M. n ={x n,1 ,x n,2 ,…,x n,m Each of the following has M instances, with a corresponding package tag Y. n It can be positive or negative, Y n ∈{0,1}, this can be achieved by labeling {y} with an unknown instance of the aggregation. n,1 ,y n,2 ,…,y n,m},Y n,m Defined as ∈{0,1}. Specifically,

[0086]

[0087] For a positive tumor slide, there is at least one tumor patch. For a negative normal slide, all patches on the slide are negative. Deep learning models within the MIL framework use a pre-trained CNN backbone to extract feature representations of the patches. Besides implementing the backbone with CNNs, other types of ML models, such as visual converters and diffusion-based models, can also be used as the backbone. The MIL pooling method aggregates these patch-level features from the same slide and predicts the final score or probability of the entire slide, comparing this predicted final score or probability of the entire slide with a benchmark truth slide-level label using a cross-entropy loss function. Therefore, the choice of different MIL pooling methods is crucial to model performance. Assume the deep learning model outputs patch-level embedding h. n,m =f(x) n,m ),in and x n,m Let m represent the m-th patch used for the n-th slice, where n = 1, ..., N and m = 1, ..., M.

[0088] For max pooling, the slide-level probability p n The positive patch is represented by the one with the highest probability among the M patches for each WSI. This assumption can be relaxed by using the top K patches with the highest probabilities as the positive patches. Therefore, p nGiven by the following formula

[0089]

[0090] For average pooling, the slide-level probability p n It is represented by the average probability value of all patches for each WSI, and is given by the following formula.

[0091]

[0092] For attention-based pooling, the slide-level probability p n The weighted average of patch-level embeddings is represented [Reference 18]. Therefore...

[0093]

[0094] During model training, the weights a can be updated using a neural network. n,m ,in and These are parameters in the fully connected layer. Attention weights reflect the importance of different patches within the slice. In some embodiments, a n,m Calculated by the following formula

[0095]

[0096] In this disclosure, the CLAM model is applied to the proposed workflow for WSI classification [Reference 21]. Tissue regions of WSI are detected using contours, and patches are created at 20x magnification using a 512×512 pixel size. Figure 2 (Subgraphs a and b). These tissue patches are embedded into low-dimensional feature vectors using a ResNet50 model pre-trained on ImageNet. Then, a gated attention module can be used to aggregate patch features from the same slice for final slide-level prediction. A clustering method with smoothed support vector machine loss based on attention scores is implemented to enhance supervision of different categories. Thus, the combination of slide-level classification loss and instance-level clustering loss is minimized during model training, where attention scores can be learned directly from the data-driven model. Attention heatmaps can be used to locate tumor and normal regions. Compared to the conventional MIL algorithm using max pooling, average pooling, or RNN aggregation, the CLAM model is more data-efficient and can be scaled to multi-tumor subtype classification or even multi-disease classification.

[0097] A.3.WSI Dataset

[0098] For the public human breast cancer lymph node metastasis dataset, this application collected 899 annotated WSIs from CAMELYON16 and CAMELYON17 Grand Challenges [Reference 22]. For the public human lung cancer dataset, this application collected a total of 1983 lung resection WSIs, including 1059 normal and LUAD WSIs from the CPTAC-LUAD project, 383 normal WSIs from the CPTAC-LSCC project, and 541 LUAD WSIs from the TCGA-LUAD project. These public datasets were used for training tumor diagnostic models.

[0099] For the human breast and lung cancer thin-section dataset, tissue samples were collected from the applicant's collaborating hospitals, and AF images and corresponding H&E staining images were obtained in the applicant's laboratory. Human breast cancer tissue was extracted by biopsy, and human lung tissue was surgically removed. These tissues underwent the FFPE process. After sectioning with a microtome, optically thin tissue sections with a thickness of 4 μm were placed on quartz slides and dewaxed for AF imaging. Label-free AF ​​images were acquired using an AF imaging system. The same slides were stained with H&E and digitized to WSI using a full-slide scanner (NanoZoomer-SQ, Hamamatsu Photonics KK) equipped with a 20× / 0.75NA objective lens. The AF images and H&E staining images were used for training a virtual staining model. The generated virtual H&E staining images were used for testing the tumor diagnostic model, and details are shown in Table 1.

[0100] Table 1. Virtual H&E WSI dataset.

[0101] Organization type Number of patients Number of virtual H&E WSIs tumor normal Lung cancer resection 9 13 12 1 Breast cancer biopsy 3 5 5 0

[0102] For a specific stained thin-section dataset, this application obtained AF, H&E, and MT stained images of human lung cancer sections. To obtain MT stained images of the same section, H&E stained samples were destained by soaking in acidic alcohol overnight, followed by a second staining process using standard MT staining. H&E and MT stained images were used for training the staining transformation network.

[0103] For the thick human lung cancer dataset, this application collected a total of 33 AF WSIs from thick human lung tissue of 29 patients, including 31 tumor WSIs and 2 normal WSIs. Standard H&E staining images of adjacent layers were obtained as references. For the adjacent thin AF dataset, this application collected a total of 43 AF WSIs from 26 patients, including 37 tumor WSIs and 6 normal WSIs.

[0104] A.4. Training Details

[0105] For virtual H&E staining of AF images and staining conversion from H&E staining to MT staining, this application follows the same data preprocessing steps and model implementation methods as in the original U-Frame paper [Reference 8]. This application uses an internal WSI dataset with AF images, H&E-stained images, and MT-stained images for model training and testing. For tumor classification and localization on H&E-stained images, this application uses the default tumor vs. normal binary classification mode only for human breast and lung cancer samples in the CLAM model. This model is trained purely on public datasets and tested on the applicant's internal dataset. For tumor diagnosis purely on thick and thin AF WSI images, normal data is augmented for model training to balance image categories. These models are trained on a workstation using PyTorch 1.8.1 and Python 3.8.8, equipped with an Ubuntu 20.04 operating system, an Intel Core i9-10980XE CPU, and an Nvidia GeForce RTX 3090 GPU.

[0106] B. Result

[0107] B.1. Special staining on virtual H&E on human lung tissue sections

[0108] To demonstrate the feasibility of the virtual staining workflow, this application validated virtual H&E and MT staining on human lung cancer tissue sections with abundant fibrosis by using a well-trained virtual staining and staining conversion model, where excessive accumulation of collagen fibers could be highlighted in blue by MT staining.

[0109] Figure 3 Label-free grayscale AF images of human lung cancer tissue sections are shown. Figure 3 Sub-image a) and baseline truth MT staining image ( Figure 3 Sub-image b) and baseline truth H&E staining image ( Figure 3 Subgraph e). From the AF image (by a virtual staining model). Figure 3 The virtual H&E staining image generated by subgraph a) Figure 3 Sub-image c) has a similarity to the baseline truth H&E staining image ( Figure 3 The quality is comparable to that of sub-image e). Virtual MT staining image ( Figure 3 Subgraphs d and f are transformed from virtual H&E using the same coloring transformation model. Figure 3 Subgraph c) and real H&E ( Figure 3Sub-image e) was generated. These two virtual MT-stained images showed similar blue collagen information compared to the baseline truth MT-stained image. The results indicate that the virtual special staining workflow can provide important diagnostic features on label-free images, just like conventional special staining methods.

[0110] B.2. Tumor diagnosis on virtual H&E of human lung resection sections

[0111] To demonstrate the feasibility of the computational pathology workflow, this application first validated virtual staining and tumor classification on human lung resection tissue using well-trained U-Frame and CLAM models, respectively. Figure 4 Label-free grayscale AF images of human lung adenocarcinoma tissue are shown. Figure 4 Sub-image a) was transformed into a baseline truth H&E stained image ( Figure 4 Subgraph c) Comparable virtual H&E staining images ( Figure 4 The result of subgraph b).

[0112] In the magnified first rectangular region 410 ( Figure 4 In subgraphs d to f, in the AF image ( Figure 4 Sub-image d), and the corresponding virtual H&E staining image ( Figure 4 Sub-image e) and baseline truth H&E staining image ( Figure 4 In sub-figure f), different features such as large tumor cells and blood can be well identified. In another magnified second rectangular region 420 ( Figure 4 In sub-images g to i), the size and shape of tumor cells in the AF image ( Figure 4 Sub-image g), and the corresponding virtual H&E staining image ( Figure 4 Sub-image H) and baseline truth H&E stained image ( Figure 4 The results in sub-figure i) are similar. Virtual staining results show that the U-Frame model has good performance for style transfer of histological images of lung cancer tissues with complex cellular and tissue structures.

[0113] For the downstream tumor classification and localization task, all 13 samples were correctly classified. Table 2 shows the confusion matrix of the classification results on the test virtual H&E dataset. Figure 5 Tumor localization results for a lung cancer sample with correct slide-level classification are shown, including the pathologist's interpretation of the baseline truth H&E staining image. Figure 5 Subfigure a), CLAM model prediction of virtual H&E staining images ( Figure 5 Sub-image b), and baseline truth H&E staining image ( Figure 5 Subgraph c).

[0114] Table 2. Tumor classification results in human lung resection.

[0115]

[0116] exist Figure 5 In subgraph a, the baseline truth tumor region is annotated by multiple polygons 530. For the attention heatmap generated by the CLAM model ( Figure 5 In sub-figures b and c), most of the darker areas with high attention scores are very close to the baseline truth annotation, while the brighter areas with low attention scores are more likely to be normal tissue areas and background areas. In the magnified first rectangular area 511 (… Figure 5 Subgraphs d to f) and the second rectangular region 512 ( Figure 5 In subgraphs g to i), respectively, from the virtual H&E staining image ( Figure 5 Sub-images e and h) and baseline truth H&E staining images ( Figure 5 The heatmap generated by subgraphs f and I) uses high attention to locate tumor regions, which are compared with the baseline truth annotation ( Figure 5 The CLAM model shows a high correlation with sub-images d and g. This encouraging result indicates that tumor cell features in virtual H&E stained images can be well classified by the CLAM model, and its performance is comparable to that of the benchmark true H&E stained images.

[0117] B.3. Tumor diagnosis on virtual H&E of human breast biopsy sections

[0118] To demonstrate the universality of computational pathology workflows for different tissues treated with different clinical protocols, this application validates the workflow on human breast cancer biopsy tissue sections. Figure 6 Histological images of human breast biopsy sections with DCIS are shown. AF images ( Figure 6 Sub-image a) is further transformed into a virtual H&E staining image using a U-Frame model with a similar cell distribution. Figure 6 Subgraph b), which is equivalent to the H&E staining image of the baseline truth ( Figure 6 Subgraph c).

[0119] In the magnified first rectangular area 610 ( Figure 6 In subgraphs d to f, the AF image ( Figure 6 Sub-image d), and the corresponding virtual H&E staining image ( Figure 6 Sub-image e) and baseline truth H&E staining image ( Figure 6 Abnormal ducts with neoplastic proliferation of ductal epithelial cells were observed in subfigure f). In another magnified second rectangular region 620 (… Figure 6 In sub-images g to i), the density and fibrous structure of tumor cells in AF images ( Figure 6 Sub-image g), and the corresponding virtual H&E staining image ( Figure 6 Sub-image h) and baseline truth H&E staining image ( Figure 6 The results in subgraph i) show a high correlation. Virtual staining results show that the U-Frame model can be used for various human tissue samples.

[0120] For downstream tumor classification and localization tasks, although the CLAM model was trained on real H&E stained images of breast excision tissue, it performed well on virtual H&E stained images of breast biopsy tissue. Table 3 shows the confusion matrix of classification results on the test virtual H&E WSI, where 4 samples were correctly classified as tumors, while 1 false negative sample was predicted by the diagnostic model.

[0121] Table 3. Tumor classification results from human breast biopsy.

[0122]

[0123] Figure 7 This illustrates the interpretation of the tumor region in a breast cancer sample with correct slide-level prediction using the CLAM model. (From a virtual H&E staining image) Figure 7 Sub-image b) and baseline truth H&E staining image ( Figure 7 The attention heatmap generated in sub-image c) is almost identical, which demonstrates that the virtual H&E stained images have image quality comparable to standard H&E histology. Figure 7 Compared to the annotation region in subgraph a), the attention heatmap ( Figure 7 Subgraphs b and c) can accurately locate most DCIS areas.

[0124] In the magnified first rectangular area 711 ( Figure 7 Subgraphs d to f) and the second rectangular region 712 ( Figure 7 In subgraphs g to i), from the virtual H&E staining image ( Figure 7 Subgraphs e and h) and baseline truth H&E staining images ( Figure 7 In the heatmaps generated from sub-figures f and i), the DCIS region was correctly located with high attention. These images were also annotated by pathologists in reference truth 731. Figure 7 The results are marked in subgraphs d and g). The results show that the workflow proposed in this application is universal for different human tissue samples.

[0125] B.4. Tumor diagnosis on AF images of thick human lung samples

[0126] To demonstrate the feasibility of the AF-based diagnostic workflow for thick tissues, this application validated the approach on thick human lung samples by referencing adjacent thin human lung slices. Figure 8 Thick human lung cancer tissues are shown separately. Figure 8 Subgraph a) and adjacent thin slices ( Figure 8 The AF image of sub-image j), and the corresponding attention heatmaps generated by the thick AF diagnostic model and the thin AF diagnostic model. Figure 8 Sub-images b and k), and reference H&E staining images with polygonal annotations of 831 and 832 (sub-images b and k), and reference H&E staining images. Figure 8 Subgraph c and subgraph l).

[0127] Thermograph of thick AF image ( Figure 8 The tumor regions with high attention scores shown in subfigure b) are very close to those in the reference H&E ( Figure 8 Reference truth annotation 831 in subgraph c). Heatmap of adjacent thin AF images ( Figure 8 Subgraph k) is also consistent with the baseline truth H&E ( Figure 8 Annotation 832 in subfigure l matches well. This is achieved by referencing the H&E staining image ( Figure 8 Subgraph f and subgraph o Figure 8 Subgraphs i and r), thick human lung samples and adjacent thin slices in the tumor region ( Figure 8 Subplots d and m) and normal region ( Figure 8 Sub-images g and p have similar diagnostic features. Thick AF images ( Figure 8 The tumor tissue region in subfigure a) Figure 8 Sub-figure d) and normal tissue area ( Figure 8 Subgraph g) can be used in heatmaps through thick AF diagnostic models. Figure 8 Subgraph b) has high attention ( Figure 8 Subgraph e) and low attention ( Figure 8 Sub-image (h) is well classified. Thin AF image ( Figure 8 Adjacent tumor tissue features of subgraph j) Figure 8 Sub-figure m) and normal tissue characteristics ( Figure 8 Subgraph p) can be used to diagnose heatmaps using a thin AF diagnostic model. Figure 8 High attention is present in subgraph k. Figure 8 Subgraph n) and low attention ( Figure 8 The results showed that the AF-based diagnostic workflow performed well for thick tissues compared to the workflow used for thin tissue sections and the pathologist's interpretation of standard H&E.

[0128] C. Discussion

[0129] It has been shown that the virtual H&E-based diagnostic workflow proposed in this application enables rapid and interpretable tumor diagnosis from virtual staining of label-free AF ​​images of various thin human tissue sections. Attention heatmaps of tumor tissue samples with correct slide-level predictions have demonstrated good performance for tumor region localization comparable to the baseline truth annotations of pathologists. Figure 5 Subgraph a to subgraph c; Figure 7 Subplots a to c). However, some false positive areas with moderate attention scores still exist in the prediction heatmap of human lung sections, such as areas of large hemorrhage (…). Figure 5 Subgraphs a to c) likely contribute to this heterogeneity in human lung tissue samples between open-source training data and internal test data. Small false-negative regions with low attention scores appear in the prediction heatmap of human breast slices. Figure 7 (Subfigures a to c) because the model may not be able to distinguish poorly differentiated tumor regions in biopsy sections. Furthermore, false-negative slide-level predictions still exist in the test dataset (Table 3). To improve the generalization ability and localization accuracy of the H&E diagnostic model, tissue samples with different structures and staining hue variations should be considered for H&E training data augmentation.

[0130] This application also demonstrates the feasibility of an AF-based diagnostic workflow that directly uses label-free tissue sections. Figure 2 Subgraph c and Figure 9 It should be noted that since the AF diagnostic model is trained on a mixture of limited AF data and a large amount of H&E data, the attention heatmap generated by the AF diagnostic model is expected to be inaccurate; therefore, this figure is merely a simple demonstration of the idea. Although this application uses only limited AF data for model training, it is comparable to benchmark truth annotations (…). Figure 9 Sub-image a) and from the virtual H&E staining image ( Figure 9 Compared to the attention heatmap generated from the AF image (sub-image c) (marked with three blue arrows at different locations 920), the attention heatmap generated from the AF image (sub-image c) is a different image. Figure 9 Subfigure b) can still roughly indicate the correct tumor area. Furthermore, the baseline truth notes from the pathologist, marked with brown arrow 910, are also present. Figure 9 Compared to sub-figure a), in the virtual H&E diagnostic heatmap ( Figure 9 The false positive areas in subfigure c) are correctly classified as in the AF diagnostic heatmap ( Figure 9 The true negative region in subfigure b). However, the overall AF diagnostic heatmap ( Figure 9 Sub-image b) is not as good as the virtual H&E diagnostic heatmap ( Figure 9Sub-image c) is accurate, but incomplete tumor regions and large false-positive regions marked with three green arrows (930) are present. This application suggests that the performance of the AF diagnostic model can be improved by collecting sufficient label-free AF ​​WSI and training the model purely using label-free image data.

[0131] Although this application validated the computational pathology workflow only on various thin human tissue sections, the proposed workflow holds promise for application to thick and unprocessed tissues. For workflows based on virtual H&E diagnosis, virtual staining of label-free AF ​​images of thick tissue samples is very challenging because tissue information between AF images and H&E staining images cannot be aligned for supervised model training. The applicant believes that this problem will be solved with the development of unsupervised virtual staining models [Reference 11]. Using virtual staining as a bridge, it is easier to interpret diagnostic features identified by tumor diagnostic models. For workflows based on AF diagnosis, this application demonstrates the usefulness of using label-free AF ​​images without slides for model training and testing on thick human samples. Figure 8 This can further accelerate the process, as a virtual staining step is not required. AF images and corresponding H&E stained images of adjacent thin sections can be obtained as diagnostic references. Pathologists can interpret thick and thin AF images by referring to adjacent H&E stained images. Assuming that training only requires slide markers obtainable from adjacent thin section references, the thick AF diagnostic model can be trained in the same way as the thin AF diagnostic model. Although the results on small-scale training data are encouraging, many false positive areas still exist in some test samples. Therefore, a large amount of AF data collection is still needed to improve model performance and generalization ability.

[0132] D. Details of the embodiments of this application

[0133] In the specification and appended claims, the terms "virtual staining network" and "virtual staining ML model" are used interchangeably to refer to an ML model used to generate virtual H&E staining images from AF images of biological samples. The virtual H&E staining image may be a virtual H&E staining WSI. The AF image may be an AF WSI.

[0134] In the specification and appended claims, the terms "stain conversion network" and "stain conversion ML model" are used interchangeably to refer to an ML model used to generate a virtual special staining image from an H&E staining image or a virtual H&E staining image of a biological sample. The (virtual) H&E staining image can be a (virtual) H&E staining WSI. The virtual special staining image can be a virtual special staining WSI.

[0135] The terms "thin tissue section" and "thick tissue" are used to describe biological tissues. As used herein, a "thin tissue section" refers to a tissue layer with a thickness of up to 10 micrometers, while a "thick tissue" refers to a piece of tissue with a thickness greater than 10 micrometers. Generally, thin tissue sections are on the μm scale, while thick tissues are on the mm scale or even larger. For example, a tissue with a thickness of 4 μm is considered a thin tissue section, while another tissue measured to be at least 2 mm thick is considered thick tissue. Thin tissue sections are most commonly obtained by sectioning thick tissue so that they are optically thin enough for microscopic examination.

[0136] Based on the details, examples, applications, etc. of the workflows mainly disclosed in Chapter AC above, and possibly in combination with generalizations and extensions, embodiments of this disclosure are developed as follows.

[0137] A first aspect of this disclosure is to provide a computer-implemented first method for virtually staining tissue samples. The tissue samples are label-free.

[0138] The first method relies on Figure 10 and Figure 1 The explanation is as follows. Figure 10 A first exemplary workflow 1000 for generating virtual special stains by virtually staining tissue samples with special staining is described. Figure 1 A general deep learning-based workflow for virtual staining and computational pathology is described.

[0139] In some embodiments intended to virtually stain tissue samples using H&E staining, workflow 1000 includes steps 1010 and 1030. In step 1010, an AF WSI 110 of the tissue sample is obtained. It should be noted that the AF WSI 110 is unlabeled. In step 1030, a virtual staining ML model 115 is used to generate a virtual H&E-stained WSI 120 from the AF WSI 110, such that the tissue sample is virtually stained with H&E staining to form the virtual H&E-stained WSI 120. Advantageously, the virtual H&E-stained WSI 120 is obtained without performing a real histochemical staining process to stain the tissue sample with H&E staining.

[0140] In some embodiments designed to virtually stain tissue samples using special staining other than H&E staining, workflow 1000 further includes step 1050. After generating a virtual H&E staining WSI 120 in step 1030, in step 1050, the virtual H&E staining WSI 120 is converted into a virtual special staining WSI 130 using a staining conversion ML model 122, such that the tissue sample is virtually stained by special staining to form the virtual special staining WSI 130. Similarly, it is advantageous to obtain the virtual special staining WSI 130 without physically staining the tissue sample using special staining.

[0141] In some embodiments, the special stain is MT staining.

[0142] Preferably, the virtual coloring ML model 115 is implemented as a U-Frame model. More preferably, the coloring transformation ML model 122 is implemented as a U-Frame model. Details of the U-Frame model can be found in [Reference 8], the disclosure of which is incorporated herein by reference. It should be noted that the virtual coloring ML model 115 and the coloring transformation ML model 122 have the same network structure (e.g., derived from the U-Frame model), but the operating parameters of the two models 115 and 122 are different and are determined during model training.

[0143] It should be noted that in steps 1030 and 1050, the virtual staining ML model 115 and the staining transformation ML model 122 are operated after the two ML models 115 and 122 have been trained.

[0144] In some embodiments, the virtual staining ML model 115 used in step 1030 is a pre-trained model. A pre-trained virtual staining ML model is obtained by loading pre-computed operational parameters into an untrained virtual staining MT model. This parameter loading process can also be used to obtain a pre-trained staining transformation ML model.

[0145] Alternatively, in some embodiments, workflow 1000 further includes one or both of steps 1020 and 1050. Step 1020, which precedes step 1030 in execution, is used to train the virtual staining ML model 115. Step 1040, which precedes step 1050 in execution, is used to train the staining transformation ML model 122.

[0146] As disclosed above, by using appropriate training datasets when training the two ML models 115 and 122, the virtual staining performed by workflow 1000 can be applied to both thin tissue sections and thick tissue. Consider training the virtual staining ML model 115 using a first training dataset and the staining conversion ML model 122 using a second training dataset. Preferably, each of the first and second training datasets is prepared according to whether a thin tissue section or a thick tissue section is used as the tissue sample.

[0147] A second aspect of this disclosure is to extend workflow 1000 to provide a second method that virtually stains a tissue sample with a special stain other than H&E staining to produce a virtual special staining WSI of the tissue sample.

[0148] Figure 11 A second exemplary workflow 1100 for virtually staining a tissue sample according to a second method is described, incorporating a first exemplary workflow 1000. In step 1110 of the second exemplary workflow 1100, a fluorescence microscopy system is used to examine the tissue sample under illumination with some pre-selected excitation beams to excite the tissue sample to generate an AF, thereby generating an AF WSI 110 of the tissue sample. After generating the AF WSI 110 in step 1110, the first exemplary workflow 1000, implemented according to any embodiment of the first method, is performed using one or more computers.

[0149] A third aspect of this disclosure is to provide a computer-implemented third method for performing tumor diagnosis on tissue samples. The tissue samples are label-free.

[0150] The third method relies on Figure 12 and Figure 1 The explanation is as follows. Figure 12 A third exemplary workflow 1200 for performing tumor diagnosis on tissue samples is described. Exemplarily, workflow 1200 includes steps 1210, 1220, 1230, and 1250.

[0151] In step 1210, AF WSI 110 of the tissue sample is obtained. It should be noted that AF WSI 110 is unlabeled.

[0152] In step 1220, a working WSI for the tissue sample is selected and obtained, enabling tumor diagnosis 140 to be performed based on the working WSI. Specifically, the working WSI is selected from the group consisting of AF WSI 110 and virtually stained WSIs of the tissue sample. The virtually stained WSIs are generated from AF WSI 110. On the one hand, performing tumor diagnosis 140 directly on AF WSI 110 instead of on virtually stained WSIs offers advantages in terms of cost and processing time savings. On the other hand, using virtual staining as a bridge between AF WSI acquisition and tumor diagnosis makes it easier to interpret diagnostic features identified by the tumor diagnosis model. The choice between using AF WSI 110 or virtually stained WSI 140 for tumor diagnosis can be determined by those skilled in the art based on the specific circumstances considered.

[0153] After identifying and obtaining the working WSI in step 1220, the working WSI is preprocessed in step 1230 before performing tumor diagnosis 140. Step 1230 includes cropping the working WSI to retain one or more tissue regions of the working WSI. That is, one or more tissue regions are initially located on the working WSI. Existing image segmentation techniques can be used to identify and delineate tissue regions on medical images. Step 1230 also includes dividing the one or more tissue regions on the working WSI into multiple tissue patches. In most cases, the corresponding tissue patches in the multiple tissue patches have the same size and dimensions, for example, each tissue patch has a size of 512 × 512 image pixels. As an illustrative example, Figure 2 Subfigure b depicts a set of cut patches 230 obtained from the tissue region 220 identified on WSI 210.

[0154] After obtaining multiple tissue patches in step 1230, in step 1250, a tumor diagnostic ML model (124 or 125) is used to process the corresponding tissue patches among the multiple tissue patches to diagnose any tumors in the working WSI.

[0155] In some embodiments, the tumor diagnostic ML model is implemented as a MIL network. It should be noted that if the working WSI is a virtual H&E staining WSI 120, then the tumor diagnostic ML model is an H&E diagnostic model 124; and if the working WSI is an AF WSI 110, then the tumor diagnostic ML model is an AF diagnostic model 125. Therefore, tumor diagnosis 140 of the tissue sample is advantageously performed without first requiring the execution of a real histochemical staining process on the tissue sample, and secondly without requiring a pathologist to interpret the histochemical staining WSI of the tissue sample.

[0156] Details of the MIL network are given in Section A.2. Its background references can be found in publicly available literature, for example, in [References 18] through [References 21], whose publications are incorporated herein by reference. Based on the publications in Section A.2, Figure 13 An exemplary MIL network 1300 for implementing a tumor diagnostic ML model (124 or 125) is depicted. The MIL network 1300 includes a backbone 1310 and a MIL pooling module 1320. The backbone 1310 is used to extract multiple patch-level features from various tissue patches across multiple tissue patches. The MIL pooling module 1320 is implemented using MIL pooling methods to aggregate the corresponding multiple patch-level features extracted for the multiple tissue patches and predict a score indicating the probability of tumor presence observed on a working WSI. This score can be simply the probability of tumor presence, or it can be a number indicating the likelihood of any tumor present in the tissue sample.

[0157] As disclosed in Section A.2, max pooling, average pooling, and attention-based pooling can be used as MIL pooling methods. For some advantages, it is preferred to use attention-based pooling as the MIL pooling method, so that the MIL network 1300 is an attention-based MIL network. However, max pooling or average pooling can also be used as MIL pooling methods.

[0158] In some embodiments, the attention-based MIL network is configured as a CLAM model. Background details of the CLAM model can be found in publicly available literature, for example, in [Reference 21], the contents of which are incorporated herein by reference. As described in Section A.2, using the CLAM model instead of max pooling, average pooling, or RNN aggregation offers the advantage of greater data efficiency.

[0159] In some embodiments, the CLAM model used to configure the attention-based MIL network is configured to provide multi-class tumor subtyping in tumor diagnosis of tissue samples.

[0160] In some embodiments, the CLAM model is configured to provide multi-disease subtyping in the diagnosis of tissue samples. It should be noted that special staining can be used to assist in the identification of non-tumor diseases, vascular diseases, etc. Utilizing virtual special staining to perform artificial intelligence diagnosis can lead to applications beyond simply subtyping tumor subtypes.

[0161] Besides implementing tumor diagnosis ML models as MIL networks, they can also generally be implemented as weakly supervised ML models. Alternatively, they can be implemented as supervised ML models. It should be noted that, as mentioned above, using weakly supervised ML models has practical advantages over supervised models, such as reducing the cost of preparing training datasets.

[0162] In some embodiments, a tumor diagnostic ML model (124 or 125) is further used in step 1250 to perform prognostic prediction of patient outcome probabilities and / or mutation prediction of patient genes after one or more tumors have been diagnosed in the working WSI. In prognostic prediction, one or more pieces of information extracted from the working WSI are used to predict the probabilities of patient outcomes. The probabilities of patient outcomes may include recurrence probability, patient survival probability, other outcomes following surgery and / or drug treatment, etc. One or more items may include tumor grade, tissue composition (e.g., stromal content), mutation information, etc. Mutation prediction results can be used as pre-screening prior to immunohistochemistry or next-generation sequencing to improve cost-effectiveness. Examples of prognostic prediction and mutation prediction based on H&E staining images can be found in [Reference 23] and [Reference 24], respectively.

[0163] The other implementation details of the third method are described below.

[0164] In some embodiments, the backbone 1310 of the MIL network 1300 is implemented as a CNN. Alternatively, the backbone 1310 can be implemented as a ResNet50 model.

[0165] In some embodiments, the backbone 1310 is implemented as a visual converter.

[0166] In some embodiments, the backbone 1310 is implemented as a diffusion-based model. Diffusion-based models used in the field of machine learning are generally also called diffusion models, which are probabilistic generative models that gradually corrupt data by injecting noise and then learning to reverse this process to generate samples.

[0167] It should be noted that in step 1250, the tumor diagnosis ML model (124 or 125) is operated to perform tumor diagnosis only after the tumor diagnosis ML model (124 or 125) has been trained.

[0168] In some embodiments, the tumor diagnostic ML model (124 or 125) used in step 1250 is a pre-trained model. The pre-trained tumor diagnostic ML model is obtained by loading pre-computed operational parameters into an untrained tumor diagnostic MT model.

[0169] Alternatively, in some embodiments, workflow 1200 further includes step 1240. Step 1240, which precedes step 1250 in execution, is used to train the tumor diagnostic ML model (124 or 125), i.e., the MIL network 1300.

[0170] It should also be noted that by using an appropriate training dataset when training the MIL network 1300, the tumor diagnosis performed by workflow 1200 can be applied to both thin tissue sections and thick tissue. Consider that the MIL network 1300 is trained using a third training dataset. Preferably, the third training dataset is prepared based on whether thin tissue sections or thick tissue are used as tissue samples.

[0171] The virtually stained WSI can be selected as the virtual H&E staining WSI 120 of the tissue sample. It should be noted that the virtual H&E staining WSI 120 is obtained by virtually staining the tissue sample using H&E staining. In step 1220, a working WSI is selected and obtained. Obtaining the working WSI may include obtaining the virtually stained WSI. Obtaining the virtually stained WSI may include generating the virtual H&E staining WSI 120 from the AF WSI 110 using the virtual staining ML model 115 as disclosed above. As described above, the virtual staining ML model 115 can be implemented as a U-Frame model.

[0172] Alternatively, the virtually stained WSI can be selected as the virtual special staining WSI 130 of the tissue sample, wherein the virtual special staining WSI 130 is obtained by virtually stained the tissue sample using a special stain other than H&E staining. An example of a special stain is MT staining. In step 1220, similarly, obtaining the working WSI may include obtaining the virtually stained WSI. Obtaining the virtually stained WSI may include: (a) generating a virtual H&E staining WSI 120 from the AF WSI 110 using the virtual staining ML model 115; and (b) converting the virtual H&E staining WSI 120 into a virtual special staining WSI 130 using the staining conversion ML model 122. As described above, each of the virtual staining ML model 115 and the staining conversion ML model 122 may be implemented as a U-Frame model.

[0173] The fourth aspect of this disclosure is to extend workflow 1200 to provide a fourth method for performing tumor diagnosis on tissue samples.

[0174] Figure 14A fourth exemplary workflow 1400 for performing tumor diagnosis on a tissue sample according to a fourth method is described, incorporating a third exemplary workflow 1200. In step 1410 of the fourth exemplary workflow 1400, a fluorescence microscopy system is used to examine the tissue sample under illumination with some pre-selected excitation beams to excite the tissue sample to generate an AF, thereby generating an AF WSI 110 of the tissue sample. After generating the AF WSI 110 in step 1410, one or more computers are used to execute the third exemplary workflow 1200 implemented according to any embodiment of the third method.

[0175] The embodiments of the various computer-implemented methods disclosed above can be implemented by one or more computers. The personal computer may be a general-purpose computer, a portable computer, a mobile computing device such as a smartphone, a computing server, a cloud server, or any computing device that a person skilled in the art deems suitable.

[0176] The invention may be implemented in other specific forms without departing from the spirit or essential characteristics thereof. Therefore, these embodiments are to be considered illustrative rather than restrictive in all respects. The scope of the invention is indicated by the appended claims rather than by the foregoing description, and thus all variations within the equivalent meaning and scope of the claims are encompassed therein.

[0177] Reference List

[0178] The following is a list of references occasionally cited in this specification. The publication details of each of these references are incorporated herein by reference in their entirety.

[0179] [1] XHGao et al., “Comparison of Fresh Frozen Tissue With Formalin-Fixed Paraffin-Embedded Tissue for Mutation Analysis Using a Multi-Gene Panel in Patients With Colorectal Cancer”, Front. Oncol. Vol. 10, March, pp. 1-8, 2020, doi:10.3389 / fonc.2020.00310.

[0180] [2] Y. Rivenson, K. de Haan, WD Wallace and A. Ozcan, “Emerging Advances to Transform Histopathology Using VirtualStaining”, BME Front, Vol. 2020, pp. 1-11, Aug. 2020, doi:10.34133 / 2020 / 9647163.

[0181] [3] Y. Rivenson et al., “Virtual histological staining of unlabeled tissue-autofluorescence images via deep learning”, Nature Biomedical Engineering, Vol. 3, No. 6, pp. 466-477, June 2019, doi:10.1038 / s41551-019-0362-y.

[0182] [4] Y. Rivenson, T. Liu, Z. Wei, Y. Zhang, K. de Haan, and A. Ozcan, “PhaseStain: the digital staining of label-free quantitative phase microscopy images using deep learning”, Light Sci. Appl., Vol. 8, No. 1, p. 23, December 2019, doi:10.1038 / s41377-019-0129-y.

[0183] [5] TTWWong et al., "Fast label-free multilayered histology-like imaging of human breast cancer by photoacoustic microscopy", Sci. Adv., Vol. 3, No. 5, p.e1602168, May 2017, doi:10.1126 / sciadv.1602168.

[0184] [6] IJ Goodfellow et al., “Generative Adversarial Networks”, Neural Information Processing Systems, 2014, pp. 2672-2680.

[0185] [7] Y. Zhang, K. de Haan, Y. Rivenson, J. Li, A. Delis, and A. Ozcan, “Digital synthesis of histological stains using micro-structured and multiplexed virtual staining of label-free tissue”, Light Sci. Appl., Vol. 9, No. 1, p. 78, December 2020, doi:10.1038 / s41377-020-0315-y.

[0186] [8] W. Dai, IHMWong, and TTWWong, “Exceeding the limit for microscopic image translation with a deep learning-based unified framework”, PNAS Nexus, March 2024, p. 133.

[0187] [9] L.Kang, X.Li, Y.Zhang, and TTWWong, “Deep learning enables ultraviolet photoacoustic microscopy based histological imaging with near real-time virtual staining,” Photoacoustics, Vol. 25, p. 100308, 2022, doi:10.1016 / j.pacs.2021.100308.

[0188]

[10] Y. Zhang et al., “High-Throughput, Label-Free and Slide-Free Histological Imaging by Computational Microscopy and Unsupervised Learning”, Adv. Sci., Vol. 9, No. 2, p. 2102358, January 2022, doi:10.1002 / advs.202102358.

[0189]

[11] L. Shi, IHMWong, CTKLo, and TTWWong, “One-side Virtual Histological Staining Model for Complex Human Samples”, Proceedings of the BHI-BSN2022-IEEE-EMBS International Conference on Biomedicine, Healing and Informatics, IEEE-EMBS International Conference on Wearable Implants, Body Sensing Networks, 2022, doi:10.1109 / BHI56158.2022.9926959.

[0190]

[12] K. de Haan et al., “Deep learning-based transformation of H&E stained tissues into special stains”, Nature Communications, Vol. 12, No. 1, pp. 1-13, 2021, doi:10.1038 / s41467-021-25221-2.

[0191]

[13] BEBejnordi et al., “Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer”, JAMA, Vol. 318, No. 22, pp. 2199-2210, 2017, doi:10.1001 / jama.2017.14585.

[0192]

[14] N. Coudray et al., “Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning”, Nature Medicine, Vol. 24, No. 10, pp. 1559-1567, October 2018, doi:10.1038 / s41591-018-0177-5.

[0193]

[15] TCHollon et al., “Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks”, Nature Medicine, Vol. 26, No. 1, pp. 52-58, January 2020, doi:10.1038 / s41591-019-0715-9.

[0194]

[16] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition”, Learn.Represent., September 2015, [online], see: http: / / arxiv.org / abs / 1409.1556.

[0195]

[17] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision”, Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., Vol. 2016-Decem, pp. 2818-2826, 2016, doi:10.1109 / CVPR.2016.308.

[0196]

[18] M. Ilse, J. Mtomczak and M. Welling, “Attention-based Deep Multiple Instance Learning”, 2018.

[0197]

[19] G. Campanella et al., “Clinical-grade computational pathology using weakly supervised deep learning on whole slide images”, Nature Medicine, Vol. 25, No. 8, pp. 1301-1309, August 2019, doi:10.1038 / s41591-019-0508-1.

[0198]

[20] B.Li, Y.Li, and K.W.E.L.C.I., “Dual-stream Multiple Instance Learning Network for Whole Slide Image Classification with Self-supervised Contrastive Learning”, IEEE Computer Society Conference Proceedings, Computer Vision and Pattern Recognition, pp. 14313-14323, 2021, doi:10.1109 / CVPR46437.2021.01409.

[0199]

[21] MY Lu, DFK Williamson, TY Chen, RJ Chen, M. Barbieri, and F. Mahmood, “Data-efficient and weakly supervised computational pathology on whole-slide images”, Nature Biomedical Engineering, Vol. 5, No. 6, pp. 555-570, 2021, doi:10.1038 / s41551-020-00682-w.

[0200]

[22] G. Litjens et al., “1399 H&E-stained sentinel lymph node sections of breast cancer patients: The CAMELYON dataset”, Gigascience, Vol. 7, No. 6, pp. 1-8, 2018, doi:10.1093 / gigascience / giy065.

[0201]

[23] N. Wahab et al., “AI-enabled routine H&E image based prognostic marker for early-stage luminal breast cancer”, npj Precision Oncology, Vol. 7, No. 122 (2023).

[0202]

[24] M. Chen et al., “Classification and mutation prediction based on histopathology H&E images in liver cancer using deep learning”, npj Precision Oncology, Vol. 4, No. 14 (2020).

Claims

1. A computer-implemented method for virtually staining tissue samples, the computer-implemented method comprising: Acquire autofluorescence (AF) whole-slide images (WSI) of the tissue sample; as well as A virtual staining machine learning (ML) model is used to generate virtual hematoxylin and eosin (H&E) stained WSIs from the AF WSIs, such that the tissue sample is virtually stained with H&E to form the virtual H&E stained WSIs.

2. The computer-implemented method according to claim 1, further comprising: The WSI of the virtual H&E staining is converted into the WSI of the virtual special staining using a staining transformation ML model, such that the tissue sample is virtually stained with special stains other than the H&E staining to form the WSI of the virtual special staining.

3. The computer-implemented method according to claim 1, wherein the virtual coloring ML model is implemented as a U-Frame model.

4. The computer-implemented method according to claim 2, wherein the color transformation ML model is implemented as the U-Frame model.

5. The computer-implemented method according to claim 2, wherein the special staining is Masson's trichrome (MT) staining.

6. The computer-implemented method according to claim 1, further comprising: The virtual staining ML model is trained before it is used to generate the WSI of the virtual H&E staining.

7. The computer-implemented method of claim 6, wherein the virtual staining ML model is trained using a first training dataset, the first training dataset being prepared based on whether thin tissue slices or thick tissue are used as tissue samples.

8. The computer-implemented method according to claim 2, further comprising: The staining transformation ML model is trained before it is used to generate the WSI of the virtual special stain.

9. The computer-implemented method of claim 8, wherein the staining conversion ML model is trained with a second training dataset, the second training dataset being prepared based on whether thin tissue slices or thick tissue are used as tissue samples.

10. A method for virtually staining a tissue sample, the method comprising: The autofluorescence (AF) WSI of the tissue sample was generated using a fluorescence microscopy system; and A process of virtually staining the tissue sample by performing a computer-implemented method according to any one of claims 1 to 9 by one or more computers.

11. A computer-implemented method for performing tumor diagnosis on a tissue sample, the computer-implemented method comprising: Acquire autofluorescence (AF) whole-slide images (WSI) of the tissue sample; Select and obtain a working WSI for tumor diagnosis, wherein the working WSI is selected from an AF WSI or a virtually stained WSI of the tissue sample, and wherein the virtually stained WSI is generated from the AF WSI; Trim the working WSI to retain one or more tissue regions of the working WSI; Divide the one or more tissue regions into multiple tissue patches; and A tumor diagnostic ML model is used to process corresponding tissue patches among the plurality of tissue patches to diagnose any tumor in the working WSI, such that a tumor diagnosis can be performed on the tissue sample without first performing a real histochemical staining process to stain the tissue sample, and without requiring a pathologist to interpret the histochemical staining WSI of the tissue sample.

12. The computer-implemented method of claim 11, wherein the tumor diagnosis ML model is implemented as a multi-instance learning (MIL) network, the MIL network comprising: The backbone is used to extract multiple patch-level features of each tissue patch from the plurality of tissue patches; as well as The MIL pooling module, implemented using the MIL pooling method, is used to aggregate corresponding patch-level features extracted for the plurality of tissue patches and to predict scores indicating the probability of tumor presence observed on the working WSI.

13. The computer-implemented method of claim 12, wherein attention-based pooling is used as the MIL pooling method, such that the MIL network is an attention-based MIL network.

14. The computer-implemented method of claim 13, wherein the attention-based MIL network is configured as a clustering-constrained attention multiple instance learning (CLAM) model.

15. The computer-implemented method of claim 14, wherein the CLAM model is configured to provide multi-tumor subtyping in tumor diagnosis of the tissue sample.

16. The computer-implemented method of claim 14, wherein the CLAM model is configured to provide multi-disease typing in the diagnosis of the tissue sample.

17. The computer-implemented method of claim 12, wherein max pooling or average pooling is used as the MIL pooling method.

18. The computer-implemented method of claim 12, wherein the backbone is implemented as a convolutional neural network (CNN).

19. The computer-implemented method of claim 18, wherein the backbone is implemented as a ResNet50 model.

20. The computer-implemented method of claim 12, wherein the backbone is implemented as a vision converter.

21. The computer-implemented method of claim 12, wherein the backbone is implemented as a diffusion-based model.

22. The computer-implemented method of claim 11, wherein the tumor diagnosis ML model is implemented as a weakly supervised ML model.

23. The computer-implemented method of claim 11, wherein the tumor diagnosis ML model is implemented as a supervised ML model.

24. The computer-implemented method according to claim 11, further comprising: After one or more tumors are diagnosed in the working WSI, the tumor diagnostic ML model is further used to perform prognostic predictions on the likelihood of patient outcomes.

25. The computer-implemented method according to claim 11, further comprising: After one or more tumors are diagnosed in the working WSI, the tumor diagnostic ML model is further used to perform mutation prediction of the patient's genes.

26. The computer-implemented method according to claim 11, further comprising: The tumor diagnostic ML model is trained before it is used to process the corresponding tissue patch.

27. The computer-implemented method of claim 26, wherein the tumor diagnostic ML model is trained with a third training dataset, the third training dataset being prepared based on whether thin tissue slices or thick tissue are used as the tissue samples.

28. The computer-implemented method according to claim 11, wherein: The virtual ground staining WSI is the virtual hematoxylin and eosin (H&E) staining WSI of the tissue sample, wherein the virtual H&E staining WSI is obtained by virtually staining the tissue sample with H&E staining; and The acquisition of the working WSI includes acquiring the virtual ground staining WSI, wherein acquiring the virtual ground staining WSI includes generating the virtual H&E staining WSI from the AF WSI using a virtual staining ML model.

29. The computer-implemented method of claim 28, wherein the virtual coloring ML model is implemented as a U-Frame model.

30. The computer-implemented method according to claim 11, wherein: The virtually stained WSI is the virtual special staining WSI of the tissue sample, wherein the virtual special staining WSI is obtained by virtually staining the tissue sample with a special stain other than hematoxylin-eosin (H&E) staining; and The obtaining working WSI includes obtaining a virtual ground staining WSI, wherein the obtaining virtual ground staining WSI includes: A virtual H&E staining WSI is generated from the AF WSI using a virtual staining ML model, wherein the virtual H&E staining WSI is obtained by virtually staining the tissue sample with the H&E stain; and The WSI of the virtual H&E staining is converted into the WSI of the virtual special staining using a staining transformation ML model.

31. The computer-implemented method of claim 30, wherein the special staining is Masson's trichrome (MT) staining.

32. The computer-implemented method of claim 30, wherein each of the virtual staining ML model and the staining transformation ML model is respectively implemented as a U-Frame model.

33. A method for performing tumor diagnosis on a tissue sample, the method comprising: The autofluorescence (AF) WSI of the tissue sample was generated using a fluorescence microscopy system; and The computer-implemented method according to any one of claims 11 to 32, wherein, based on the generated AF WSI, one or more computers perform a process for tumor diagnosis of the tissue sample.