Digital analysis of pre-analytical factors in tissues used for histological staining
Patent Information
- Application Number
- JP2024529716
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-23
- Filing Date
- 2022-11-23
- Publication Date
- 2025-11-17
AI Technical Summary
Existing methods for assessing preanalytical factors in tissue staining are subjective, manual, and not reproducible, leading to inconsistencies in staining quality and potential misdiagnosis due to variations in preanalytical factors such as fixation time and fixation quality.
A computer-implemented method using machine learning models to analyze images of histopathology slides, trained on datasets with preanalytical factors, to objectively determine the quality and consistency of tissue staining by predicting factors like fixation time and other preanalytical variables.
The method provides an objective, reproducible assessment of staining quality, reducing the risk of misdiagnosis by identifying abnormal preanalytical factors and enabling standardized tissue processing protocols, thereby improving diagnostic accuracy.
Smart Images

Figure 00000038_0000 
Figure 00000038_0001 
Figure 00000039_0000
Abstract
Description
[Technical field]
[0001] Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 282,249, filed November 23, 2021, the contents of which are incorporated herein by reference.
[0002] The present invention, in some embodiments thereof, relates to pre-analytical factors, and more particularly, but not exclusively, to systems and methods for the estimation of pre-analytical factors in tissue used in tissue staining. [Background technology]
[0003] Preanalytical factors (also called preanalytical variables) include fixation and processing variables that can affect the processing of tissue formalin fixation and paraffin embedding for tissue preservation and tissue staining. Summary of the Invention
[0004] According to a first aspect, a computer-implemented method for training a pre-analytical factor machine learning model includes creating a pre-analytical training dataset of a plurality of records, the pre-analytical records including images of subject pathology tissue slides processed with at least one pre-analytical factor and ground truth labels representing the at least one pre-analytical factor, and training a pre-analytical machine learning model on the pre-analytical training dataset to generate results of at least one target pre-analytical factor used to process tissue represented in the target image in response to an input of a target image.
[0005] According to a second aspect, a computer-implemented method for obtaining at least one pre-analysis factor of a target image of a slide of pathology tissue of a subject includes the steps of: feeding the target image to a pre-analysis machine learning model, the pre-analysis machine learning model being trained on a pre-analysis training dataset of a plurality of records, the pre-analysis records including images of the slide of pathology tissue of the subject processed with the at least one pre-analysis factor and ground truth labels representing the at least one pre-analysis factor; and obtaining results of the at least one target pre-analysis factor used to process the pathology tissue represented in the target image.
[0006] According to a third aspect, a device for training a pre-analytical factor machine learning model comprises at least one hardware processor that executes code for: creating a pre-analytical training dataset of a plurality of records, the pre-analytical records including images of subject pathology tissue slides processed with at least one pre-analytical factor and ground truth labels representing the at least one pre-analytical factor; and training a pre-analytical machine learning model on the pre-analytical training dataset to generate results of at least one target pre-analytical factor used to process tissue represented in the target image in response to an input of a target image.
[0007] According to a fourth aspect, a device for obtaining at least one pre-analysis factor of a target image of a slide of pathology tissue of a subject comprises at least one hardware processor for executing code for: feeding the target image to a pre-analysis machine learning model, the pre-analysis machine learning model being trained on a pre-analysis training dataset of a plurality of records, the pre-analysis records including images of the slide of pathology tissue of the subject processed with the at least one pre-analysis factor and ground truth labels representing the at least one pre-analysis factor; and obtaining results of the at least one target pre-analysis factor used to process the pathology represented in the target image.
[0008] In further embodiments of the first, second, third and fourth aspects, the method further includes creating a secondary training dataset of a plurality of records, the secondary records including images of slides of pathology tissue from subjects processed with at least one pre-analysis factor, the at least one pre-analysis factor and ground truth labels representing a secondary indication, and training a secondary machine learning model on the secondary training dataset to generate results of a target secondary indication in response to an input of a target image and the at least one target pre-analysis factor used to process the tissue represented in the target image.
[0009] In further embodiments of the first, second, third, and fourth aspects, the secondary training dataset comprises a clinical indication training dataset, the secondary indication comprises a clinical indication, and the secondary machine learning model comprises a clinical machine learning model.
[0010] In a further embodiment of the first, second, third and fourth aspects, the clinical indication is selected from the group comprising a clinical score, a medical condition and a pathology report.
[0011] In further embodiments of the first, second, third and fourth aspects, the method further comprises treating the subject with a therapy effective for the medical condition according to the clinical score and / or according to the pathology report.
[0012] In further embodiments of the first, second, third and fourth aspects, the ground truth labels are selected from the group consisting of tags, metadata, images and segmentation results of a segmentation model fed with the images.
[0013] In further embodiments of the first, second, third and fourth aspects, the input of the at least one pre-analysis factor provided to the secondary machine learning model is obtained as a result of the pre-analysis machine learning model being provided with the target image.
[0014] In further embodiments of the first, second, third, and fourth aspects, the pre-analysis machine learning model and the secondary machine learning model are jointly trained using at least common images and common labels of the pre-analysis factors.
[0015] In further embodiments of the first, second, third and fourth aspects, the at least one pre-analytical factor of the secondary record includes at least one feature map extracted from a hidden layer of the pre-analytical machine learning model fed with an image of the subject's pathology slide processed with the at least one pre-analytical factor, and the secondary machine learning model generates a target secondary instruction result in response to an input of a target image and the target feature map extracted from the hidden layer of the pre-analytical machine learning model fed with the target image.
[0016] In further embodiments of the first, second, third and fourth aspects, the method includes creating an image translation training dataset including two or more sets of image translation records, where source image translation records of a source set of the image translation records include source images of the subject's histopathology slides processed with at least one pre-analysis factor and ground truth representing source labels, and destination image translation records of a destination set of the image translation records include destination images of the subject's histopathology slides processed with at least one pre-analysis factor and ground truth representing destination labels; and training an image translation machine learning model on the image translation training dataset to translate target source images of the histopathology slides of the source set of the image translation records to result destination images of the histopathology slides of the destination set of the image translation records.
[0017] In further embodiments of the first, second, third and fourth aspects, the source label represents pathological tissue that has been abnormally processed with at least one pre-analytical factor and the destination label represents pathological tissue that has been normally processed with at least one pre-analytical factor.
[0018] In further embodiments of the first, second, third and fourth aspects, the target source image includes the input image and additional metadata representing the abnormally processed source pre-analysis factors, and metadata representing normally processed destination pre-analysis factors.
[0019] In further embodiments of the first, second, third and fourth aspects, the target source image includes the input image, further comprising providing a reference image from a destination set that is used to infer a destination for the input image.
[0020] In further embodiments of the first, second, third and fourth aspects, the source set is selected according to an input of at least one pre-analysis factor obtained as a result of a pre-analysis machine learning model fed with the target image.
[0021] In further embodiments of the first, second, third, and fourth aspects, the method further includes creating an image correction training dataset of a plurality of records, the image correction records including images of a subject's tissue slide processed with at least one pre-analysis factor, where the at least one pre-analysis factor is classified as abnormal, and the images of the slides including images representing the abnormally processed tissue, the at least one pre-analysis factor, and ground truth labels representing normal images of the tissue slides processed with the at least one pre-analysis factor classified as normal; and training an image correction machine learning model on the image correction training dataset to generate synthetic corrected image results of the tissue slides simulating what the target images of the slides would look like when processed with at least one pre-analysis factor classified as normal in response to a target image of the slide processed with the at least one target pre-analysis factor classified as abnormal.
[0022] In further embodiments of the first, second, third and fourth aspects, the input of the at least one pre-analysis factor supplied to the image correction machine learning model is obtained as a result of the pre-analysis machine learning model being supplied with the target image.
[0023] In further embodiments of the first, second, third, and fourth aspects, the image correction machine learning model and the pre-analysis machine learning model are jointly trained using common images and common ground truth labels of the pre-analysis factors.
[0024] In further embodiments of the first, second, third, and fourth aspects, the method further comprises training the baseline model using a self-supervised and / or unsupervised technique on an unlabeled training dataset of a plurality of unlabeled images of the pathological tissue of the subject processed with the at least one pre-analysis factor, wherein the training comprises further training the baseline model on the pre-analysis training dataset to create a pre-analysis machine learning model.
[0025] In further embodiments of the first, second, third, and fourth aspects, the ground truth labels representing the at least one pre-analysis factor include ground truth labels representing correctly applied pre-analysis factors or anomalous application of the pre-analysis factors, and the training step includes training an implementation of the pre-analysis machine learning model to learn a distribution of inlier images labeled as correctly applied pre-analysis factors to detect images as outliers representing incorrectly applied pre-analysis factors.
[0026] In further embodiments of the first, second, third and fourth aspects, the method further comprises extracting features from the image using a pre-trained feature extractor, wherein the pre-analysis record comprises the extracted features, and wherein the pre-trained feature extractor is applied to the target image to obtain extracted target features that are fed to the pre-analysis machine learning model.
[0027] In further embodiments of the first, second, third and fourth aspects, the pre-trained feature extractor is implemented as a neural network, and the extracted features are obtained from at least one feature map prior to a classification layer of the neural network when the neural network is fed with a target image.
[0028] In further embodiments of the first, second, third, and fourth aspects, the neural network is an image classifier that is trained on an image training dataset of non-tissue images labeled with ground truth classification categories.
[0029] In further embodiments of the first, second, third and fourth aspects, the neural network is a nuclei segmentation network that is trained on a segmentation training dataset of images of pathology slides that have been labeled with a ground truth segmentation of nuclei.
[0030] In further embodiments of the first, second, third and fourth aspects, the method further comprises extracting a plurality of patches from the image, and wherein extracting features comprises extracting features from the plurality of patches.
[0031] In further embodiments of the first, second, third and fourth aspects, the method further comprises, for each patch, reducing the extracted features extracted from the patch to a feature vector using a global max pooling layer and / or a global average pooling layer, the pre-analysis record comprising the feature vector, and the pre-analysis machine learning generates results of the at least one target pre-analysis factor in response to an input of the feature vector calculated for the features extracted from the patch of the target image.
[0032] In further embodiments of the first, second, third and fourth aspects, the method further comprises, for each pre-analysis record, providing the image to a nucleus segmentation machine learning model to obtain a result of a segmentation of the nuclei in the image, creating a mask that covers pixels outside the nucleus segmentation based on the result of the segmentation, and applying the mask to the image to create a covered image, wherein the images of the pre-analysis record include the covered image, and the target covered image created from the target image is provided to the pre-analysis machine learning model trained on the pre-analysis training dataset.
[0033] In further embodiments of the first, second, third and fourth aspects, the method further comprises, for each pre-analysis record, feeding the image to a nuclei segmentation machine learning model to obtain a segmentation result of nuclei in the image, and cropping a boundary around each segmentation to create single nuclei patches, wherein the image of the pre-analysis record includes a plurality of single nuclei patches, and the target segmentations of nuclei created from the target image are fed to a pre-analysis machine learning model trained on a pre-analysis training dataset.
[0034] In further embodiments of the first, second, third and fourth aspects, the method further comprises, for each pre-analysis record, converting a color version of the image to a grayscale version of the image, and the target grayscale version of the target image is fed to a pre-analysis machine learning model trained on the pre-analysis training dataset.
[0035] In further embodiments of the first, second, third and fourth aspects, the method further comprises, for each pre-analysis record, providing the image to a (RBC) segmentation machine learning model to obtain results of segmentation of red blood cells (RBCs) and / or patches representing RBCs in the image, wherein the images of the pre-analysis record include segmentation of RBCs and / or patches representing RBCs, and the target segmentation of RBCs and / or patches representing RBCs from the target image are provided to a pre-analysis machine learning model trained on a pre-analysis training dataset.
[0036] In further embodiments of the first, second, third, and fourth aspects, the pre-analysis machine learning model is pre-trained on a separate image training dataset including a plurality of images each labeled with a respective ground truth indication of a particular classification category, and the pre-trained pre-analysis training dataset is further trained on the pre-analysis training dataset.
[0037] In further embodiments of the first, second, third, and fourth aspects, the pre-analysis record further includes metadata representing at least one known pre-analysis factor, and the ground truth label is for the at least one unknown pre-analysis factor, and the at least one known pre-analysis factor associated with the target image is further provided to a pre-analysis machine learning model that is trained on the pre-analysis training dataset.
[0038] In further embodiments of the first, second, third and fourth aspects, the method further comprises training an interpretability machine learning model to generate an interpretability map representing the relative importance of pixels of the target image to obtain at least one target pre-analysis factor, the target image being low resolution, and further comprises sampling a plurality of high resolution patches of the target image; and feeding the plurality of high resolution patches to the pre-analysis machine learning model to obtain the at least one target pre-analysis factor.
[0039] In a further embodiment of the first, second, third and fourth aspects, the at least one pre-analytical factor comprises a fixation time.
[0040] In a further embodiment of the first, second, third and fourth aspects, the at least one pre-analytical factor comprises tissue thickness obtained by sectioning the FFPE block.
[0041] In further embodiments of the first, second, third, and fourth aspects, the at least one pre-analysis factor is selected from the group consisting of fixative type, warm ischemia time, cold ischemia time, duration and delay of temperature during pre-fixation, fixative formulation, fixative concentration, fixative pH, duration of fixation of reagents, source of fixative preparation, tissue to fixative volume ratio, fixation method, primary and secondary fixation conditions, post-fixation wash conditions and duration, post-fixation storage reagents and duration, processor type, frequency of maintenance and reagent changes, tissue to reagent volume ratio, number of co-processed sample positions, dehydration and clearing reagents, dehydration and clearing temperature, number of dehydration and clearing changes, dehydration and clearing duration, bake time, and temperature.
[0042] In a further embodiment of the first, second, third and fourth aspects, the at least one pre-analytical factor is an indication of the quality of staining of the histopathology of the slide.
[0043] In further embodiments of the first, second, third, and fourth aspects, the staining is selected from the group consisting of immunohistochemistry (IHC) staining, in situ hybridization (ISH) staining, fluorescent ISH (FISH), chromogenic ISH (CISH), silver ISH (SISH), hematoxylin and eosin (H&E), hematoxylin, acridine orange, bismarck brown, carmine, Coomassie blue, cresyl violet, crystal violet, 4',6-diamidino-2-phenylindole ("DAPI"), eosin, endothelial cell suspension ... The marker is selected from the group consisting of tetrazine intercalating, acid fuchsin, Hoechst stain, iodine, malachite green, methyl green, methylene blue, neutral red, Nile blue, Nile red, osmium tetroxide, propidium iodide, rhodamine, safranine, antibody-based stains, or label-free imaging markers obtained using imaging techniques including Raman spectroscopy, near-infrared ("NIR") spectroscopy, autofluorescence imaging, and phase imaging that highlight features of interest without the use of external dyes.
[0044] In a further embodiment of the first, second, third and fourth aspects, the slide comprises formalin-fixed paraffin-embedded (FFPE) tissue.
[0045] In further embodiments of the first, second, third and fourth aspects, the method further includes the steps of supplying the target image and at least one target pre-analysis factor to a secondary machine learning model, the secondary machine learning model being trained on a secondary instruction training dataset of a plurality of records, the secondary instruction records including images of subject's pathology tissue slides processed with the at least one pre-analysis factor, the at least one pre-analysis factor, and ground truth labels representing the secondary instruction, and obtaining results of the target secondary instruction.
[0046] In further embodiments of the first, second, third and fourth aspects, the method further comprises, in response to classifying the at least one target pre-analysis factor as abnormal, providing the target image and the at least one target pre-analysis factor to an image correction machine learning model, the image correction machine learning model being trained on a corrected image training dataset of a plurality of records, the image correction records including images of a slide of pathology tissue of the subject processed with the at least one pre-analysis factor, where the at least one pre-analysis factor is classified as abnormal, the images of the slide representing the abnormally processed pathology tissue, the at least one pre-analysis factor, and a ground truth label representing a normal image of the slide of pathology tissue processed with the at least one pre-analysis factor classified as normal; and obtaining a corrected image result simulating what the target image of the slide would look like when processed with the at least one pre-analysis factor classified as normal.
[0047] In further embodiments of the first, second, third, and fourth aspects, the method further includes, in response to classifying the at least one target pre-analysis factor as abnormal, providing the target image and the at least one target pre-analysis factor to an image translation machine learning model, the image translation machine learning model being trained on an image translation training dataset including two or more sets of image translation records, where the source image translation records of the source set of image translation records include source images of the subject's histopathology slides processed with the at least one pre-analysis factor and ground truth representing source labels, and the destination image translation records of the destination set of image translation records include destination images of the subject's histopathology slides processed with the at least one pre-analysis factor and ground truth representing destination labels, and obtaining a resultant destination image of the histopathology slide of the destination set of image translation records, which is a transformation of the abnormally processed target image to a normally processed image.
[0048] Unless otherwise specified, all technical and / or scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. Although methods and materials similar or equivalent to those described herein may be used in carrying out or testing embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including the provisions, shall control. In addition, the materials, methods and examples are only illustrative and are not necessarily intended to be limiting.
[0049] Some embodiments of the present invention are described herein by way of example with reference to the accompanying drawings. With particular reference now to the drawings in detail, the details shown are emphasized by way of example and for purposes of illustrative discussion of embodiments of the present invention. In this regard, the description made with the drawings below will make apparent to those skilled in the art how embodiments of the present invention may be practiced. [Brief description of the drawings]
[0050] [Figure 1] FIG. 1 is a block diagram of components of a system for training ML model(s) for generating instructions for pre-analysis factor(s) used to process tissue represented in a target image in response to input of a target image(s) depicting a tissue sample(s) according to some embodiments of the present invention. [Diagram 2] 1 is a flowchart of a process for training an ML model(s) to generate an indication of pre-analysis factor(s) to be used to process tissue represented in a target image in response to an input of the target image, according to some embodiments of the present invention. [Diagram 3]1 is a flowchart of a process for using ML model(s) to obtain an indication of pre-analysis factor(s) in response to input of target image(s) representing tissue sample(s) according to some embodiments of the present invention. [Figure 4] 11A-11C are example images representing slides of tissue samples having different fixation times, according to some embodiments of the present invention. [Diagram 5] 13 is another example depicting slides of tissue samples having different fixation times, according to some embodiments of the present invention. [Figure 6] FIG. 1 is a schematic diagram illustrating a process of training an ML model using extracted features according to some embodiments of the present invention. [Figure 7] FIG. 1 depicts an image of tissue processed with one or more pre-analysis factors, having segmented nuclei segmented by a nuclei segmentation ML model, according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0051] The present invention, in some embodiments thereof, relates to pre-analytical factors, and more particularly, but not exclusively, to systems and methods for the estimation of pre-analytical factors in tissues used in histological staining.
[0052] An aspect of some embodiments of the present invention relates to a system, method, computing device, and / or code instructions (stored in memory and executable by one or more hardware processors) for training a pre-analysis factor machine learning model. A pre-analysis training dataset of a plurality of records is created. The pre-analysis records include images of slides of pathology tissues of subjects processed with pre-analysis factor(s) and ground truth labels representing the pre-analysis factor(s). A pre-analysis machine learning model is trained on the pre-analysis training dataset to generate results of target pre-analysis factor(s) used to process tissues represented in the target image in response to an input of the target image.
[0053] An aspect of some embodiments of the present invention relates to a system, method, computing device, and / or code instructions (stored in memory and executable by one or more hardware processors) for obtaining pre-analysis factor(s) for a target image of a slide of a pathology tissue of a subject. The target image is provided to a pre-analysis machine learning model. Results of the target pre-analysis factor(s) used to process the target image are obtained from the pre-analysis machine learning model.
[0054] Optionally, the target pre-analysis factor(s) obtained as a result from the pre-analysis machine learning model are fed to the secondary machine learning model in combination with the target image. The target secondary instruction result is obtained as a result of the secondary machine learning model. The secondary machine learning model may be trained on a secondary instruction training dataset including a plurality of records. The secondary instruction record includes an image of a slide of a subject's pathology tissue processed with the pre-analysis factor(s), an instruction of the pre-analysis factor(s), and a ground truth label representing the secondary instruction, e.g., tags, metadata, images, and a segmentation result of the segmentation model fed with the image. The secondary training dataset may be implemented as a clinical instruction training dataset, the secondary instruction may be implemented as a clinical instruction, and the secondary machine learning model may be implemented as a clinical machine learning model.
[0055] Optionally, if the target pre-analysis factor(s) obtained as a result from the pre-analysis machine learning model are determined to be abnormal, e.g., outside a range and / or threshold representing a correct value for the target pre-analysis factor(s), the target image and the target pre-analysis factor(s) are provided to an image correction machine learning model. A corrected image result simulating how the target image of the slide would look when processed with the pre-analysis factor(s) classified as normal is obtained as a result of the image correction machine learning model. The image correction machine learning model is trained on an image correction training dataset of a plurality of records. The image correction record includes an image of a slide of pathology tissue of the subject processed with the pre-analysis factor(s) where the pre-analysis factor(s) are classified as abnormal and the image of the slide represents abnormally processed pathology tissue, an indication of the pre-analysis factor(s), and a ground truth label representing a normal image of the slide of pathology tissue processed with the pre-analysis factor(s) classified as normal.
[0056] Alternatively or additionally, if the target pre-analysis factor(s) obtained as a result from the pre-analysis machine learning model are determined to be anomalous, a heat map (e.g., as described herein) and / or a score (e.g., probability of being anomalous) may be presented on the display. A user may view the heat map and / or score to help determine how to interpret the image and / or whether the image should be discarded.
[0057] At least some embodiments of the systems, methods, apparatus (e.g., computing devices), and / or code instructions (e.g., stored on a data storage device and executable by one or more hardware processors) described herein address the technical problem of determining pre-analysis factors of tissues to be processed represented within an image, e.g., a whole slide image of a pathology tissue. At least some embodiments of the systems, methods, apparatus, and / or code instructions described herein improve the technical and / or medical fields of analysis of tissue specimens by determining pre-analysis factors used to process the tissues from images representing those tissue specimens. At least some embodiments of the systems, methods, apparatus, and / or code instructions described herein improve the technical field of machine learning by providing machine learning model(s) that generate results of pre-analysis factor(s) in response to an input of an image of a tissue specimen.
[0058] Processed tissues, such as formalin-fixed paraffin-embedded (FFPE) tissue samples stained using immunohistochemistry (IHC) techniques, are routinely analyzed by pathologists in clinical and research laboratories worldwide. However, the quality of the final IHC staining depends on multiple pre-analysis factors, such as tissue fixation, operating variables, assay validity, and others described herein. Staining quality can refer to the staining intensity for both primary and counterstains and / or the appearance of tissue structure within the tissue specimen. Staining quality is primarily influenced by the number of finite tissue antigens preserved within the tissue specimen throughout the preanalytical workflow, as described, for example, in KB Engel and HM Moore, “Effects of preanalytical variables on the detection of proteins by immunohistochemistry in formalin-fixed, paraffin-embedded tissue”, Arch Pathol Lab Med. 2011; 135(5); 537-43 (hereinafter “Engel”), and / or DRBauer, M. Otter and DR Chafin, “A New Paradigm for Tissue Diagnostics: Tools and Techniques to Standardize Tissue Collection, Transport, and Fixation”, Current pathobiology reports, 2018; Vol. 6; 135-143 (hereinafter “Bauer”), which are incorporated herein by reference in their entireties.
[0059] The preanalysis phase begins as soon as the tissue pieces are removed from the blood supply as tissue denaturation caused by autolysis within cells. Fixatives are therefore used to preserve as much of the tissue structure and antigens as possible. The staining quality depends mainly on the over / underfixation of the tissue specimen, for example, underfixed tissue specimens show weak staining signals in IHC staining, as described, for example, by Bauer. Tissue denaturation can be further accelerated by temperature increases in the surrounding environment, making the time to fixation a key parameter for obtaining good staining quality. Poor fixation can also cause morphological tissue changes that can remove important tissue information that can be used for manual or automated cancer diagnosis. There is no standard preanalysis workflow, so it is not known how each parameter can affect the final staining quality, as described, for example, by Bauer. Lack of standardization leads to significant variations in staining protocols both within and between institutions, as described, for example, in Engel and / or Lanng, M. et al., “Quality assessment of Ki67 staining using cell line proliferation index and stain intensity features”, 2019, Cytometry. Part A: the journal of the International Society for Analytical Cytology. 95, 4, s. 381-388 (hereinafter “Lanng”), the entire contents of which are incorporated herein by reference.
[0060] The ML models described herein provide an objective and / or reproducible approach to determine pre-analytical factors used in processing tissues represented in an image, which may be used to standardize pre-analytical tissue collection workflows, to evaluate the staining quality of newly developed staining protocols, and / or to improve disease (e.g., cancer) diagnosis and / or treatment. For example, following a protocol, images of tissues may be fed into the ML model(s) to determine whether the pre-analytical factors used to process the tissue are within the correct range (or above / below a threshold) or abnormal (e.g., outside the correct range and / or above / below a threshold). Improved clinical workflows based on evaluation of histologically stained tissue specimens may be obtained by analyzing the extent to which pre-analytical factors affect the quality of the prepared stained tissue specimens by using the approaches described herein, for example, images of the final stained tissue specimen and / or staining responses in new tissue specimens may be used to evaluate staining quality. Human pathologists and / or automated cancer (or other disease) diagnosis processes (e.g., applications running on a computer) may take into account the predicted staining quality when making a diagnosis. For example, poor staining quality may result in no or indeterminate diagnosis, whereas good staining quality may result in a diagnosis with high certainty. The staining quality assessment tool may serve as a gold standard staining quality assessment for developing more robust staining protocols and / or testing products.
[0061] The effect of pre-analytical factors on staining quality is also antigen-dependent, making some IHC stains more sensitive to pre-analytical variations than others. HER2 is an example of a sensitive epitope that is routinely used in breast cancer diagnosis to determine optimal treatment, as described, for example, with reference to Bauer and / or EC Colley & R.H. Stead, "Fixation and Other Pre-Analytical Factors", in Immunohistochemical staining methods, IHC Guidebook, chapter 2, 6th edition, Dako Denmark A / S, An Agilent Technologies Company (hereinafter "Colley"). However, insufficient staining can result in insufficient staining response in HER2-positive tissue structures, which will therefore go undetected by pathologists. Staining variations caused by pre-analytical factors can then directly affect the diagnostic process, and thereby the treatment and outcome for patients, as described, for example, with reference to Engel.
[0062] A technical challenge is that pathologists have difficulty finding tissue changes resulting from inadequate pre-analysis processing. Extensive tissue deterioration from warm ischemia, for example, produces significant morphological changes, but requires expertise from the pathologist to be able to recognize the subtle differences resulting from over- and under-fixation. One parameter that pathologists can find is the geometry of red blood cells, although it is not common to assess this metric and the changes are so subtle that they are rarely found. Another criterion that can be assessed is the clarity of mitotic events, since over-fixation appears to slightly blur mitotic nuclear changes.
[0063] At least some embodiments described herein improve upon standard approaches for assessing the quality of tissue specimens. Conventional approaches for quality control measurements of tissue samples are manual-based and rely on pathologists who are trained enough to recognize issues with pre-analysis parameters and make decisions about tissue specimens based on this prior knowledge. Moreover, such manual approaches are subjective and not necessarily reproducible. In contrast, at least some embodiments described herein use machine learning models to determine pre-analysis factors used during tissue processing by providing automatic, objective, reproducible, and / or accurate analysis of tissue specimens. The pre-analysis factors may represent the quality of the tissue specimen, such as an indication of whether fixation time is acceptable or not. The embodiments based on the machine learning model(s) described herein may significantly improve the workflow of pathologists assessing tissue specimens by making the analysis less subjective, improving decision making.
[0064] Among many pre-analytical variables, tissue fixation time (i.e., a specific pre-analytical factor) probably has the most significant impact on the quality of IHC and in situ hybridization (ISH) staining, because it affects many other variables, such as antigen retrieval and epitope binding. At least some embodiments described herein make it possible to provide an automatic, objective, reproducible, and / or accurate technique for predicting fixation time depending on an image of a tissue specimen, for example, in a HER2-stained tissue specimen. Staining quality may be determined according to the predicted fixation time, for example, good staining quality when the predicted fixation time is within the correct range, and poor staining quality when the predicted fixation time is outside the correct range, representing an abnormal (e.g., incorrect) fixation time.
[0065] Considering the consequences of changing staining quality based on changing pre-analysis conditions, the ability to help end users better interpret IHC staining by informing them about potential biases in the staining would help reduce the risk of incorrect patient treatment due to false positive / false negative interpretation of the staining. In some cases, it is known that for overfixation cases, increasing the pre-processing time can effectively overcome the problem of overfixation. In that case, it would be possible to inform the pathologist that a given sample was under / overfixed using the embodiments described herein and that a modified diagnostic staining protocol for that sample is necessary to give accurate results, without introducing additional hardware other than a brightfield slide scanner. Such a tool may solve the technically challenging development of having limited knowledge of the fixation condition of the incoming biological tissues used during the development of new diagnostic tests. Although official guidelines are commonly used by laboratories performing fixation, these have a wide range and tissue density, size and geometry greatly affect the degree of fixation of the tissue sample. The ability to measure the relative fixation of tissues using the embodiments described herein will, in the short term, make it possible to have an objective approach to tissue selection for test development, thus enabling the development of more robust staining protocols and diagnostic products.
[0066] Before describing at least one embodiment of the invention in detail, it is to be understood that the invention is not necessarily limited in its application to the details of construction and arrangement of components and / or methods set forth in the following description and / or illustrated in the drawings and / or in the examples. The invention is capable of other embodiments or of being practiced or carried out in various ways.
[0067] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium(s) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.
[0068] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. Computer-readable storage media may include, but are not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), static random access memories (SRAMs), portable compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), memory sticks, floppy disks, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed as being a transitory signal itself, such as an electric wave or another freely propagating electromagnetic wave, an electromagnetic wave propagating through a wave guide or another transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electric signal transmitted through an electric wire.
[0069] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0070] The computer readable program instructions for carrying out the operations of the present invention are either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, for carrying out aspects of the present invention.
[0071] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0072] These computer readable program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or another programmable data processing apparatus to create a machine, such that the instructions, which execute via the processor of the computer or another programmable data processing apparatus, create means for performing the function / operation specified in the block or blocks of the flowcharts and / or block diagrams. These computer readable program instructions may also be stored in a computer readable storage medium that may cause a computer, programmable data processing apparatus, and / or another device to function in a particular manner, such that the computer readable storage medium having instructions stored thereon comprises an article of manufacture including instructions that implement an aspect of the function / operation specified in the block or blocks of the flowcharts and / or block diagrams.
[0073] The computer readable program instructions may also be loaded onto a computer, another programmable data processing apparatus, or another device to cause a series of operational steps to be performed on the computer, another programmable apparatus, or another device to generate a computer-implemented process, such that the instructions executing on the computer, another programmable apparatus, or another device perform the function / operation specified in the block or blocks of the flowcharts and / or block diagrams.
[0074] The flowcharts and block diagrams in the drawings represent the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing a specified logical function(s). In some alternative embodiments, the functions shown in the blocks may occur in an order different from the order shown in the drawings. For example, two blocks shown in succession may in fact be executed substantially in parallel, or the blocks may sometimes be executed in reverse order depending on the functionality involved. Also, each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or executes a combination of dedicated hardware and computer instructions.
[0075] Reference is now made to FIG. 1, which is a block diagram of components of a system 100 for training ML model(s) to generate instructions for pre-analysis factor(s) used to process tissue represented in a target image in response to an input of a target image and / or to use the ML model(s) to obtain instructions for pre-analysis factor(s) in response to an input of a target image representing tissue specimen(s), according to some embodiments of the present invention. Reference is also made to FIG. 2, which is a flow chart of a process for training ML model(s) to generate instructions for pre-analysis factor(s) used to process tissue represented in a target image in response to an input of a target image, according to some embodiments of the present invention. Reference is also made to FIG. 3, which is a flow chart of a process for using ML model(s) to obtain instructions for pre-analysis factor(s) in response to an input of a target image representing tissue specimen(s), according to some embodiments of the present invention. Reference is also made to FIG. 4, which is an example of images representing slides of tissue specimens having different fixation times, according to some embodiments of the present invention. Reference is also made to Figure 5, which is another example depicting slides of tissue specimens with different fixation times, according to some embodiments of the present invention. Reference is also made to Figure 6, which is a schematic diagram depicting a process 600 of training an ML model using extracted features, according to some embodiments of the present invention. Reference is now made to Figure 7, which depicts an image 702 of tissue processed with one or more pre-analysis factors, having segmented nuclei 704 (one nucleus shown for clarity) segmented by a nucleus segmentation ML model, according to some embodiments of the present invention.
[0076] Referring again to FIG. 4, image 402 represents a specimen of normally fixed red blood cells that underwent a normal fixation time of 26 hours. In contrast, image 404 represents another specimen of over-fixed red blood cells that underwent an over-fixation time of 143 hours. Even an expert pathologist would have difficulty visually distinguishing between the cells in image 404 and image 402, especially since the images are of different cell specimens. Thus, it would be difficult to determine that the cells in image 404 are over-fixed while the cells in image 402 are normally fixed. As discussed herein, in some cases, the geometry of the red blood cell criteria may be calculated and used to try and determine the fixation time, but evaluating this criteria is not common since the changes are minor and they are found infrequently. In at least some embodiments described herein, the trained ML model generates a result in response to the input of images 402 and / or 404 that represents the fixation time and / or whether the fixation time is normal or abnormal.
[0077] Referring again to FIG. 5, image 502 represents a normally fixed tissue specimen that received a normal fixation time of 26 hours. In contrast, image 504 represents another specimen of overfixed tissue that received an overfixation time of 143 hours. Even an expert pathologist would have difficulty visually distinguishing between cells in image 504 and image 502, especially since the images are of different cytological specimens. Thus, it would be difficult to determine that the cells in image 504 are overfixed while the cells in image 502 are normally fixed. As discussed herein, in some cases, the geometry of the red blood cell criteria may be calculated and used to try and determine the fixation time, but evaluating this criterion is not common as the changes are minor and they are rarely found. Although overfixation such as in 504 slightly blurs the mitotic nuclear changes compared to normal fixation in 502, the changes are difficult to find. In at least some embodiments described herein, the trained ML model generates results indicative of fixation times and / or indicative of whether fixation times are normal or abnormal in response to input of images 502 and / or 504.
[0078] Referring again now to FIG. 1, the system 100 may optionally perform the acts of the methods described with reference to FIGS. 2-7 by a hardware processor(s) 102 of a computing device 104 executing code instructions 106A and / or 106B stored in a memory 106.
[0079] Computing device 104 may be implemented as, for example, a client terminal, a server, a virtual server, a laboratory workstation (e.g., a pathology workstation), a procedure (e.g., operating) room computer and / or server, a virtual machine, a computing cloud, a mobile device, a desktop computer, a thin client, a smartphone, a tablet computer, a laptop computer, a wearable computer, a glasses computer, and a wristwatch computer. Computing 104 may sometimes include an advanced visualization workstation implemented as an add-on to a laboratory workstation and / or another device to present images of tissue specimens to a user (e.g., a pathologist).
[0080] Different architectures of the system 100 based on the computing device 104 may be implemented, for example, in a central server-based embodiment and / or a locally based embodiment.
[0081] In an example of a central server-based embodiment, the computing device 104 may include locally stored software performing one or more of the operations described with reference to FIGS. 2-7 and / or may operate as one or more servers (e.g., network servers, web servers, computing clouds, virtual servers) providing services (e.g., one or more of the operations described with reference to FIGS. 2-7) over the network 110 to one or more client terminals 108 (e.g., remotely located laboratory workstations, remote picture archiving and communication system (PACS) servers, remote electronic medical record (EMR) servers, remote image archive servers, remotely located pathology computing devices, user client terminals such as desktop computers), such as providing software as a service (SaaS) to the client terminal(s) 108, providing applications for local download to the client terminal(s) 108 as an addition to a web browser and / or tissue specimen imaging viewer application, and / or providing functionality using a remote access session to the client terminal 108, such as via a web browser. In one embodiment, multiple client terminals 108 each acquire an image of the tissue specimen from a different imaging device(s) 112. Each of the multiple client terminals 108 provides the image to the computing device 104. The computing device may feed the received image(s) to one or more machine learning model(s) 122A to obtain a result indicative of a pre-analysis factor (e.g., an estimated fixation time, and / or whether the fixation time was normal or abnormal, e.g., it was too little (i.e., overfixation) or too much (i.e., underfixation), and / or another as described herein), and / or another result of a different ML model, such as a secondary indication and / or a corrected image, as described herein.Results obtained from the computing device 104 may be provided to each client terminal 108, respectively, for example, for presentation on a display and / or storage in local storage and / or for feeding to another process, such as a diagnostic application. Training of the machine learning model(s) 122A may be performed centrally by the computing device 104 based on images of tissue specimens and / or annotations of data obtained from one or more client terminal(s) 108, optionally multiple different client terminals 108, and / or may be performed by another device (e.g., server(s) 118) and provided to the computing device 104 for use.
[0082] In a locally based embodiment, each computing device 104 is used by a particular user, e.g., a particular pathologist, and / or a group of users within an institution, such as a hospital and / or a pathology laboratory. The computing device 104 receives specimen images from the imaging device 112, for example, directly and / or via an image repository 114 (e.g., a PACS server, cloud storage, hard disk). Annotations may be received from a user (e.g., manually entered via an interface) and / or extracted from another source, e.g., from metadata output by the tissue processing device(s) 150 representing pre-analysis factors used during tissue processing. The images may be locally provided to one or more machine learning model(s) 122A to obtain one or more result(s) as described herein. The result(s) may be, for example, presented on the display 126, may be locally stored in the data storage device 122 of the computing device 104, and / or may be provided to another application that may be locally stored on the data storage device 122. The training of the machine learning model(s) 122A may be performed locally by each respective computing device 104 based on annotations of images and / or data of specimens acquired from each imaging device 112, for example, different users may each train their own set of machine learning models 122A using specimens used by the user that have been processed using a particular processing protocol and / or using a particular tissue processing device 150, and / or different pathology laboratories may each train their own set of machine learning models using their own images that have been processed using their own particular tissue processing protocol and / or using their own particular tissue processing device 150. For example, a pathologist specializing in the analysis of bone marrow biopsies trains an ML model on images of bone marrow biopsy specimens that have been processed using pre-analysis factors suitable for bone marrow.Another laboratory that specializes in kidney biopsies trains an ML model on images representing kidney tissue obtained via biopsy that have been processed using pre-analysis factors suitable for kidney tissue. In another example, the trained machine learning model(s) 122A are obtained from another device, such as a central server.
[0083] The computing device 104 receives images of the tissue specimen captured by one or more imaging device(s) 112. Exemplary imaging device(s) 112 include scanners scanning in standard color channels (e.g., red, green, blue), multispectral imagers that acquire images in four or more channels, confocal microscopes, black and white imagers, and imaging sensors.
[0084] Optionally, one or more tissue processing devices 150 process the tissue using analytical agent(s), which may be known and / or unknown, as determined as described herein, for example, to fix and / or apply a stain to the tissue specimen that is then imaged by imaging device 112.
[0085] The imaging device(s) 112 may create a two-dimensional (2D) image of the specimen, optionally an entire slide image.
[0086] Images captured by the imaging machine 112 may be stored in an image repository 114, for example, a storage server (eg, PACS, EHR server), a computing cloud, virtual memory, and a hard disk.
[0087] Training data set(s) 122B may be created based on the captured images, as described herein.
[0088] The machine learning model(s) 122A may be trained on training dataset(s) 122B as described herein.
[0089] Exemplary ML model(s) 122A include one or more of a pre-analysis ML model, a secondary ML model (e.g., a clinical ML model), an image correction ML model, and another ML model used in an optional pre-processing step, such as a nucleus segmentation ML model, an RBC segmentation ML model, and / or an interpretability ML model (e.g., as described with reference to 206 in FIG. 2).
[0090] Exemplary architectures of the machine learning models described herein include, for example, statistical classifiers and / or other statistical models, neural networks of various architectures (e.g., convolutional, fully connected, one or more convolutional layers with one or more subsequent connection layers, deep, encoder-decoder, recurrent, graph), support vector machines (SVMs), logistic regression, k-nearest neighbors, decision trees, boosting, random forests, regressors, and / or any other commercial or open source package that enables regression, classification, dimensionality reduction, supervised, unsupervised, semi-supervised, or reinforcement learning. Machine learning models may be trained using supervised and / or unsupervised techniques.
[0091] The machine learning models described herein can be fine-tuned and / or updated. An existing trained ML model trained on a particular type of tissue, such as bone marrow biopsy, can be used as a basis for training another ML model using a transfer learning approach for another type of tissue, such as blood smear. The transfer learning approach of using an existing ML model can improve the accuracy of the newly trained ML model and / or reduce the size of the training dataset for training the new ML model, and / or reduce the time and / or computational resources for training the new ML model compared to the standard approach of training the new ML model "from scratch".
[0092] The computing device 104 may receive images for analysis from the imaging device 112 and / or image repository 114 using one or more imaging interfaces 120, which may include, for example, a wired connection (e.g., a physical port), a wireless connection (e.g., an antenna), a local bus, a port for connecting a data storage device, a network interface card, another physical interface implementation, and / or a virtual interface (e.g., a software interface, a virtual private network (VPN) connection, an application programming interface (API), or a software development kit (SDK)). Alternatively or additionally, the computing device 104 may receive images from the client terminal(s) 108 and / or the server(s) 118.
[0093] The hardware processor(s) 102 may be implemented, for example, as a central processing unit(s) (CPU), a graphics processing unit(s) (GPU), a field programmable gate array(s) (FPGA), a digital signal processor(s) (DSP), an application specific integrated circuit(s) (ASIC). The processor(s) 102 may include one or more processors (homogeneous or heterogeneous), which may be arranged for parallel processing as a cluster and / or as one or more multi-core processing units.
[0094] The memory 106 (also referred to herein as a program store and / or data storage device) stores code instructions for execution by the hardware processor(s) 102, e.g., random access memory (RAM), read-only memory (ROM), and / or storage devices, e.g., non-volatile memory, magnetic media, semiconductor memory devices, hard drives, removable storage devices, and optical media (e.g., DVD, CD-ROM). The memory 106 stores code 106A and / or training code 106B that implements one or more acts and / or features of the methods described with reference to FIGS. 3-7.
[0095] The computing device 104 may include a data storage device 122 for storing data, e.g., machine learning model(s) 122A as described herein, and / or a training dataset 122B for training the machine learning model(s) 122A as described herein. The data storage device 122 may be implemented, for example, as a memory, a local hard drive, a removable storage device, an optical disk, a storage device, and / or as a remote server and / or a computing cloud (e.g., accessed via the network 110). It should be noted that executable code portions of the data stored in the data storage device 122 may be loaded into the memory 106 for execution by the processor(s) 102.
[0096] The computing device 104 may include one or more of a data interface 124, optionally a network interface, e.g., a network interface card, a wireless interface for connecting to a wireless network, a physical interface for connecting to a cable for network connection, a virtual interface implemented in software, network communication software providing a higher layer of network connectivity, and / or another embodiment, to connect to the network 110. The computing device 104 may access one or more remote servers 118 using the network 110 to download, for example, updated versions of the machine learning model(s) 122A, code 106A, training code 106B, and / or training dataset(s) 122B.
[0097] The computing device 104 may use a network 110 (or another communications channel, such as via a direct link (e.g., cable, wireless) and / or an indirect link (e.g., through a computing device such as a server and / or through a storage device)) to communicate with the *Client terminal(s) 108, for example when the computing device 104 acts as a server providing image analysis services (e.g., SaaS) to remote lab terminals as described herein; * A server 118 implemented in conjunction with a PACS and / or electronic medical records, for example, that may store images of specimens from different individuals (e.g., patients) for processing, as described herein; * an image repository 114 that stores images of the specimen captured by the imaging device 112; may communicate with one of
[0098] It should be noted that the imaging interface 120 and the data interface 124 may exist as two independent interfaces (e.g., two network ports), as two virtual interfaces on a common physical interface (e.g., a virtual network on a common network port), and / or may be combined into a single interface (e.g., a network interface).
[0099] The computing device 104 includes or communicates with a user interface 126 that includes mechanisms designed for a user to input data (e.g., manual input of pre-analysis factors for annotation of an image) and / or view data (e.g., pre-analysis factors predicted by an ML model(s)). Exemplary user interfaces 126 include, for example, one or more of a touch screen, a display, a keyboard, a mouse, and voice-activated software using a speaker and microphone.
[0100] Referring again now to FIG. 2, at 200, one or more images (e.g., of slides) of tissue, optionally pathological tissue of one or more subjects, that have been processed with at least one pre-analysis factor, are obtained and / or accessed.
[0101] Multiple images of multiple slides may be acquired, each representing a tissue specimen obtained from a different subject. The multiple images may be from different slides of the same tissue. Alternatively or additionally, multiple images are acquired from different slides of different tissues of the same subject. The images may be of the same type of tissue specimen obtained from different subjects, for example, blood smears, bone marrow biopsies, surgically removed tumors, and polyps extracted from biopsies. The provided and / or trained ML model(s) may correspond to one or each tissue type, or alternatively, images representing different tissue types from different patients.
[0102] The tissue on the slide may include formalin-fixed paraffin-embedded (FFPE) tissue.
[0103] The images may be obtained, for example, from an image sensor that captures the images, from a scanner that captures the images, or from a server that stores the images (e.g., a PACS server, an EMR server, a pathology server). For example, tissue images are sent for analysis automatically after capture by an imaging modality and / or as soon as the images are stored after being scanned by an imaging modality.
[0104] As used herein, the term "image" may refer to an entire slide image (WSI), and / or a patch extracted from the WSI, and / or a portion of a specimen. For example, a phrase describing an image being fed to an ML model may refer to a patch extracted from the WSI that is fed to the ML model.
[0105] The images may be of a specimen taken at high magnification, for example, with an objective lens of about 20X to 40X or another value. Such high magnification imaging may produce very large images, for example, on the order of gigapixel size. Each large image may be divided into patches of smaller size that are then analyzed. Alternatively, the large image is analyzed as a whole. The images may be scanned along different xy planes at different axial (i.e., z-axis) depths.
[0106] Tissue may be obtained, for example, during a biopsy procedure, a fine needle aspiration (FNA) procedure, a core biopsy procedure, a liquid biopsy procedure, a colonoscopy for removal of a colon polyp, surgery for removal of an unknown mass, surgery for removal of a benign cancer, surgery for removal of a malignant cancer, and / or surgery for treatment of a medical condition. Tissue may be obtained from fluids, for example, urine, synovial fluid, blood, and cerebrospinal fluid. Tissue may be in the form of a connected group of cells, for example, a histological slide. Tissue may be in the form of individual cells or clumps of cells suspended in a fluid, for example, a cytological specimen.
[0107] At 202, an indication of the pre-analysis factor(s) to be used during processing of the tissue represented in each image within the range is obtained and / or accessed, e.g., automatically extracted (e.g., from a record associated with the slide as output by a side preparation device) and / or manually entered by a user. The indication may be stored, for example, as metadata, tags, and / or field values.
[0108] Exemplary pre-analysis factors include fixation time, tissue thickness obtained by sectioning of FFPE blocks, fixative type, warm ischemia time, cold ischemia time, temperature duration and delay during pre-fixation, fixative formulation, fixative concentration, fixative pH, fixation duration of reagents, source of fixative preparation, tissue to fixative volume ratio, fixation method, primary and secondary fixation conditions, post-fixation wash conditions and duration, post-fixation storage reagents and duration, processor type, frequency of maintenance and reagent changes, tissue to reagent volume ratio, number of co-processed sample positions, dehydration and clearing reagents, dehydration and clearing temperature, number of dehydration and clearing changes, dehydration and clearing duration, bake time, and temperature.
[0109] The pre-analysis factor(s) may include an indication of the staining quality of the slide. Exemplary stains include IHC stains, in situ hybridization (ISH) stains, alternative methods for ISH such as fluorescent ISH (FISH), chromogenic ISH (CISH), silver ISH (SISH), etc., hematoxylin and eosin (H&E), hematoxylin, acridine orange, bismarck brown, carmine, Coomassie blue, cresyl violet, crystal violet, 4',6-diamidino-2-phenylindole ("DAPI"), eosin, ethidium bromide intercalate, acid fuchsin, Hoechst stain, iodine ... The contrast may be derived from the use of imaging techniques including, but not limited to, uran, malachite green, methyl green, methylene blue, neutral red, Nile blue, Nile red, osmium tetroxide, propidium iodide, rhodamine, safranine, antibody-based stains, or label-free imaging markers (which may result from the use of imaging techniques including, but not limited to, Raman spectroscopy, near infrared "NIR" spectroscopy, autofluorescence imaging, or phase imaging, and / or which may be used to highlight features of interest without the use of external dyes, etc.) and / or others. In some cases, contrast when using label-free imaging techniques may be generated without the use of additional markers such as fluorescent or chromogenic dyes.
[0110] At 204, one or more additional data items may be obtained and / or accessed, such as for each subject, automatically (e.g., extracted from records such as each subject's electronic health record) and / or manually provided by a user. The additional data items may be stored, for example, as metadata, tags, and / or field values.
[0111] The additional data items may serve as ground truth in the record(s) of the training dataset(s) for training one or more ML models, as described herein, and / or may be used as inputs to the ML models.
[0112] Optionally, the additional data items may include a secondary indication for each subject. Examples of secondary indications include tags, metadata, images, and segmentation results of a segmentation model fed with the images. The secondary indication may be a clinical indication (e.g., for a clinical indication record of a clinical indication training dataset for training a clinical indication ML model), such as a clinical score (e.g., ratio of specific immune cells to total immune cells, a score of cancer invasiveness into tissue), a clinical diagnosis of a medical condition (e.g., malignant tumor, benign tumor, adenoma, lung cancer), and a pathology report.
[0113] Alternatively or additionally, the additional data item may be an indication that each pre-analysis factor(s) is / are classified as either normal (e.g., normally applied) or abnormal (e.g., incorrectly applied, incorrect operating value, abnormal application). The quality of the slide is determined depending on whether the pre-analysis factor(s) are normal or abnormal, for example, whether the pre-analysis factor(s) are within a range defined as a correct operating range suitable for obtaining good quality slides, or whether the pre-analysis factor(s) are outside the correct operating range (i.e., erroneous), thereby reducing the quality of the slide. The indication of whether the pre-analysis factor(s) are normal or abnormal may be used to select images representing normal pre-analysis factors to serve as ground truth and other images representing abnormal pre-analysis factors for inclusion in an image correction training dataset, as described herein.
[0114] Alternatively or additionally, the additional data item may be metadata representing unknown pre-analysis factor(s). For each image of each slide, some pre-analysis factor(s) may be known and some pre-analysis factor(s) may be unknown.
[0115] At 206, one or more (e.g., each respective) images may be pre-processed to, for example, perform patch extraction, feature extraction, nuclear segmentation, color transformation, RBC segmentation, and computation of an interpretability map.
[0116] Optionally, features are extracted from each image. The features may be extracted using a pre-trained feature extractor. The extracted features may serve as ground truth in record(s) of training dataset(s) for training one or more ML models, as described herein, and / or may be used as input to the ML models.
[0117] The pre-trained feature extractor may be implemented as a neural network (e.g., a deep neural network) and / or another ML model architecture and / or another feature extraction architecture that may be non-ML based (e.g., Scale Invariant Feature Transform (SIFT) and / or Speed Up Robust Features (SURF)). The extracted features are obtained from at least one feature map before the classification layer of the neural network when the neural network is fed with the target image. For example, from the layer immediately preceding the classification layer and / or from one or more deeper layer(s), for example, using a projection head on top of the learned representation. The neural network may be, for example, an image classifier trained on an image training dataset of non-tissue images labeled with ground truth classification categories. Alternatively or additionally, the neural network is a nucleus segmentation network trained on a segmentation training dataset of images of pathology slides labeled with ground truth segmentation of nuclei and / or nucleoli. The bottleneck layer may be extracted from the nucleus segmentation network. In such an embodiment, the extracted features are the segmentation of the nuclei and / or the mask of the nuclei segmentation output by the neural network. Alternatively or additionally, other features may be extracted, e.g., hand-created features and / or features automatically identified by a feature discovery process (e.g., SIFT, SURF).
[0118] Alternatively or additionally, patches are extracted from the image. Patches, rather than entire slides, may be used to increase the computational efficiency of a computing device during training and / or inference, i.e., patches are smaller than entire slide images, and therefore less computational resources are required to process the patches across the entire slide image. In some cases, the same pre-analysis factor(s) may be applied to the entire tissue specimen shown in an image (e.g., on a slide). In such cases, determining the pre-analysis factor(s) for the patch infers the pre-analysis factor(s) for the entire image. In other cases, the pre-analysis factor may vary locally for different regions of an image (e.g., on a slide), e.g., tissue thickness may vary, which may affect local pre-analysis factors, fixation times may vary locally, autolysis may vary locally. In such cases, different patches of the same image may have varying values for the pre-analysis factors.
[0119] Features may be extracted from the patches, for example, using techniques described herein for features extracted from images. The patches may be taken from a region of interest (ROI), which may be rectangular with a preset size (e.g., number of pixels in length and / or width), optionally at a preset scale factor. The ROI may be a region of the WSI. The patches may be extracted within a grid covering the ROI. The patches may be overlapping (e.g., with a preset amount of overlap) and / or non-overlapping. The features extracted from the patches may be stitched together to create an enhanced feature map and / or may be used as individual features.
[0120] For each image and / or each patch, the features extracted from each patch and / or image may be reduced to a feature vector. The reduction may be done, for example, using a global max pooling layer and / or a global average pooling layer. The pre-analysis record (used to train the pre-analysis ML model) may include the feature vector. Optionally, during training of a neural network embodiment of the pre-analysis ML model (e.g., convolutional neural networks (CNNs), fully connected networks, and attention-based (transformer) networks), the convolutional layer(s) may operate directly on the input feature patches. Optionally, non-neural network embodiments of the ML model (e.g., tree-based methods such as gradient boosting trees (GBTs) and random forests, etc.) may operate on features extracted by another method (e.g., SIFT, SURF). The pre-analysis machine learning generates the results of the target pre-analysis factor in response to the input of a feature vector extracted from a patch of the target image and / or calculated for features to be extracted from the target image.
[0121] Referring again now to Fig. 6, the features described in Fig. 5 may be implemented as, combined with and / or substituted for the features described with reference to Fig. 6. At 602, an image of a tissue specimen, optionally an entire slide image, processed with one or more pre-analysis factors is obtained, for example as described with reference to 200 of Fig. 2. A ground truth representing the pre-analysis factor(s) used to process the tissue represented in the image is obtained, for example as described with reference to 202 of Fig. 2. At 604, a patch is extracted from the image of the tissue, optionally from an ROI. At 606, the extracted features are applied to the patch to extract features, for example as described with reference to 206 of Fig. 2. At 608, a feature map may be extracted, for example as described with reference to 206 of Fig. 2. At 610, a training dataset includes, for example, the feature map records and / or the extracted features that are labeled with the ground truth, as described with reference to 208A of Fig. 2. At 612, the ML model is trained using a loss function, e.g., as described with reference to 208B of Figure 2. Alternatively, features 606 and / or 608 are omitted, in which case the patches of 604 are labeled with their respective ground truth indications of the pre-analysis factor(s) and included in the training dataset records of 610.
[0122] Referring back to 206 of FIG. 2, alternatively or additionally, the image is fed to a nuclei segmentation machine learning model to obtain a result of a segmentation of nuclei in the image. A mask may be created based on the result of the segmentation to cover pixels outside the nuclei segmentation. The mask is applied to the image to generate a cover image. The cover image may be used in the recording (e.g., pre-analysis recording) instead of and / or in addition to the image itself to train the ML model(s) (e.g., pre-analysis machine learning model). During inference, a target cover image created from the target image is fed to a trained (e.g., pre-analysis) machine learning model to obtain, for example, target pre-analysis factor(s).
[0123] Alternatively or additionally, once an image is fed to a nucleus segmentation machine learning model to obtain the results of segmentation of nuclei in the image, a boundary (e.g., a minimum bounding rectangle or another context to allow inference from the periphery of the nuclei) may be created around each segmentation to create a single nucleus patch. The single nucleus patch may be used in the recording (e.g., pre-analysis recording) instead of and / or in addition to the image itself to train the ML model(s) (e.g., pre-analysis machine learning model). During inference, the target segmentation of nuclei generated from the target image is fed, for example, to a trained (e.g., pre-analysis) machine learning model to obtain target pre-analysis factors.
[0124] Referring again now to FIG. 7, an image 702 of tissue processed with one or more pre-analysis factors is shown, which includes segmented nuclei 704 (one nucleus is shown for clarity) that have been segmented by a nuclei segmentation ML model. The nuclei segmentation ML model may, for example, be trained on a training dataset of images of cells labeled with a ground truth segmentation of nuclei. The nuclei segmentation ML model may also compute the segmentation using, for example, another approach that analyzes the color distribution of cells to identify segmented nuclei.
[0125] Referring back to 206 in FIG. 2, alternatively or additionally, a color version of the image is converted to a grayscale version of the image. The grayscale image may be used for recording (e.g., pre-analysis recording) instead of and / or in addition to a color image for training the ML model(s) (e.g., pre-analysis machine learning model). During inference, the target grayscale version of the target image is fed, for example, to a trained (e.g., pre-analysis) machine learning model to obtain the target pre-analysis factor(s). The use of a grayscale image instead of and / or in addition to a color image may prevent the ML model from learning irrelevant color variations resulting, for example, from different stains, different imaging sensors, etc.
[0126] Alternatively or additionally, the image is fed to a red blood cell (RBC) segmentation machine learning model to obtain the results of segmentation of RBCs in the image and / or patches representing RBCs. The segmentation of RBCs and / or patches representing RBCs may be used in the recording (e.g., pre-analysis recording) instead of and / or in addition to the image itself to train the ML model(s) (e.g., pre-analysis machine learning model). During inference, the target segmentation of RBCs from the target image and / or patches representing RBCs are fed, for example, to a trained (e.g., pre-analysis) machine learning model to obtain the target pre-analysis factor(s). RBCs are more sensitive to the fixation process and so may be a good indication of whether the pre-analysis factor is normal or abnormal, e.g., representing over-fixation and / or under-fixation.
[0127] Alternatively or additionally, an interpretability machine learning model is trained to generate an interpretability map that represents the relative importance of pixels of a target image, thereby obtaining the target pre-analysis factors. The interpretability map may be implemented, for example, as an attention map, a probability map, and / or a class activation map. The target image used to obtain the interpretability map may be of low resolution. High-resolution patches of the target image may then be sampled according to the interpretability map calculated from the low-resolution target image. The high-resolution patches may be selected, for example, as K sampled patches, where K indicates a hyperparameter of the ML model based on the relevance of the patches and / or another consideration, such as selecting the K most relevant ones and / or trying to select the most relevant ones without selecting all patches from the sample area of the sample. In another example, the high-resolution patches may be selected as those having a relative significance above a threshold. The high-resolution patches may be used in the recording (e.g., pre-analysis recording) instead of and / or in addition to the image itself to train the ML model(s) (e.g., pre-analysis machine learning model). During inference, high-resolution patches extracted from the target image are fed into a trained (e.g., pre-analysis) machine learning model to obtain the target pre-analysis factor(s).
[0128] 2, the features described with reference to 208A-B, 210A-B, and 212A-B represent different ML models that may be trained using the data obtained in features 200-206. Training may be performed using a loss function, for example a standard cross-entropy loss function.
[0129] At 208A, a pre-analysis training dataset of a plurality of records is created. The pre-analysis records include images of each subject's (e.g., pre-analysis factor) tissue slides processed with the pre-analysis factor(s), ground truth labels representing the pre-analysis factors, and optionally other data as described with reference to 204 and / or 206. The other data may be appended to the images and / or may be an embodiment of the images, such as patches extracted from the images. The other data may include one or more of patches extracted from the images, features extracted from the images, segmented nuclei, color transformed images (e.g., black and white images), RBC segmentation, and interpretability map(s).
[0130] The pre-analysis record may further include metadata representing two types of pre-analysis factors: (i) known pre-analysis factor(s) and (ii) pre-analysis factor(s) predicted to be unknown during inference (but known during training). The known pre-analysis factors may be associated with the pre-analysis factor(s) that are unknown at inference time. During inference, the values of the known pre-analysis factors are fed to the ML model and used to help determine the values of the unknown pre-analysis factor(s). For example, the pre-analysis factor FISH is highly sensitive to over-fixation. During inference, the known pre-analysis factor FISH may be fed to the ML model and used to help the ML model infer information about the degree of fixation and / or the degree of tissue autolysis within the tissue block where such pre-analysis factor(s) are unknown. To train such a model, ground truth labels are for the pre-analysis factor(s) predicted to be unknown during inference (but known during training).
[0131] At 208B, the pre-analysis machine learning model is trained on the pre-analysis training dataset to generate results of pre-analysis factor(s) used to process tissue represented in the target image in response to an input of the target image.
[0132] Optionally, the ground truth labels representing the pre-analysis factors include ground truth labels representing whether the applied pre-analysis factors were applied correctly or not, or whether the application of the pre-analysis factors is anomalous or not. In such a case, an embodiment of a machine learning model may be trained to learn the distribution of inlier images labeled as correctly applied pre-analysis factors to detect images as outliers representing incorrectly applied pre-analysis factors. An embodiment of a ML model may be, for example, an autoencoder, a variational autoencoder (VAE), a generative adversarial network (GAN), and the like.
[0133] Optionally, the pre-analysis machine learning model is pre-trained on another image training dataset, each of which includes images labeled with a respective ground truth indication of a particular classification category. The pre-trained pre-analysis training dataset is further trained on the pre-analysis training dataset.
[0134] At 210A, a secondary instruction training data set of records is created, the secondary instruction records including each image of each subject's pathology slide processed with the pre-analysis factor(s), an indication of the pre-analysis factor(s), and a ground truth label representing the secondary instruction, and optionally other data as described with reference to 204 and / or 206 (e.g., the example described at 208A).
[0135] Optionally, the pre-analysis factor(s) of the secondary instruction record includes at least one feature map extracted from a hidden layer(s) of the pre-analysis machine learning model fed with an image of the subject's pathology slide processed with the pre-analysis factor(s). The hidden layer(s) may include one or more layers, which may be the last layer before the classification layer or a separate layer. During inference, the secondary machine learning model generates a target secondary instruction outcome in response to an input of a target image and the target feature map extracted from the hidden layer of the pre-analysis machine learning model fed with the target image.
[0136] At 210B, a secondary machine learning model is trained on the secondary instruction training dataset to generate a target secondary instruction outcome in response to input of the target image and the target pre-analysis factor(s) used to process the tissue represented in the target image. The target pre-analysis factor(s) may be obtained as a result of the pre-analysis machine learning model fed with the target image.
[0137] At 212A, an image correction training dataset of a plurality of records is created. The image correction record includes images of slides of pathology tissue of the subject processed with the pre-analysis factor(s). The record includes images of slides representing pathology tissue processed abnormally. The record also includes an indication that the pre-analysis factor(s) are classified as abnormal. Images for which the pre-analysis factor(s) are classified as normal are excluded. The record further includes an indication of the pre-analysis factor(s). The record further includes a ground truth label of a normal image of the slide (e.g., another image that may be the same tissue as the abnormal slide, or similar tissue to the slide labeled as abnormal), and optionally the pathology tissue processed with the pre-analysis factor(s) classified as normal.
[0138] Alternatively or additionally, an image translation training dataset of two or more sets of image translation records is created, each set including a source set of source image translation records and a destination set of destination image translation records. The sets may be divided by classification of the pre-analysis factors. The source image translation records of the source set of image translation records may include source images of subject histopathology slides processed with the pre-analysis factors and ground truth representing source labels. The source labels may represent pathology processed abnormally with the pre-analysis factors. The destination image translation records of the destination set of image translation records may include destination images of subject histopathology slides processed with the pre-analysis factors and ground truth representing destination labels. The destination labels may represent pathology processed normally with the pre-analysis factors.
[0139] At 212B, an image correction machine learning model is trained on an image correction training dataset to generate a result of a synthetic corrected image of a pathology tissue slide that simulates what the target image of the slide would look like when processed with a pre-analysis factor(s) classified as normal, in response to an input of a target image of the slide processed with a target pre-analysis factor classified as abnormal.
[0140] Alternatively or additionally, an image translation machine learning model is trained on the image translation training dataset for translating target source images of histopathology slides in a source set of image translation records to result destination images of histopathology slides in a destination set of image translation records.
[0141] Exemplary architectures for implementing the image correction ML model and / or the image translation ML model include unsupervised image translation, self-supervised image translation, CycleGAN, StarGAN, unsupervised image-to-image translation (UNIT), and multimodal unsupervised image-to-image translation (MUNIT).
[0142] At 214, the pre-analysis machine learning model and the secondary machine learning model may be trained together (e.g., end-to-end) using at least common images and common labels of the pre-analysis factors. For example, some of the images and / or labels are common and some of the images and / or labels are unique to one or both of the pre-analysis and secondary ML models. For example, common images and / or labels may be used for joint (e.g., end-to-end) training while unique images and / or labels are used, and no secondary results are present but pre-analysis factor(s) are present to enable joint training.
[0143] At 216, the image correction machine learning model and the pre-analysis machine learning model may be jointly trained using common images and common ground truth labels of the pre-analysis factors.
[0144] At 218, a baseline model may be trained using self-supervised and / or unsupervised techniques on an unlabeled training dataset of unlabeled images of subject(s) tissue or, optionally, pathological tissue processed with pre-analysis factor(s). The unlabeled images may be of similar tissue and / or different tissue as used in the recording described herein. The unlabeled images may be of similar pre-analysis factor(s) and / or different pre-analysis factor as used in the recording described herein. The baseline model is then trained on the pre-analysis training dataset to create a pre-analysis machine learning model. It should be noted that the baseline model may be trained on a secondary instruction training dataset to create a secondary ML model and / or on an image correction training dataset to create an image correction ML model.
[0145] The baseline model may be used as an alternative to and / or in addition to using a feature extractor. Feature extraction may be used for rapid training under cross-validation schemes. Using a fine-tuning procedure in which a baseline model (e.g., a pre-trained network) is used as the initial state and some or all of the network layers are trained using the training dataset may allow the network to learn lower level, more relevant features.
[0146] Referring again now to Figure 3, at 302, ML model(s) are trained and / or provided, for example, as described with reference to Figure 2. The ML model(s) include one or more of a pre-analysis ML model, a secondary ML model, an image correction ML model, and another ML model used in an optional pre-processing step, such as a nuclei segmentation ML model, an RBC segmentation ML model, and / or an interpretability ML model (e.g., as described with reference to 206 in Figure 2).
[0147] At 304, a target image of a subject's tissue, optionally a pathology tissue specimen, is acquired and / or accessed, for example as described with reference to 200 in FIG.
[0148] At 306, the target image may be pre-processed, e.g., by one or more of patch extraction, feature extraction, nuclear segmentation, color transformation, RBC segmentation, and computation of an interpretability map, as described with reference to 206 of Figure 2. The pre-processing corresponds to the pre-processing performed at 206 of Figure 2 to obtain data for each training dataset used to train each ML model, as described with reference to Figure 2.
[0149] At 308, the target image (optionally pre-processed) is fed to a pre-analysis machine learning model. Alternatively or additionally, one or more of the following obtained as described with reference to 306 are fed to the pre-analysis ML model: feature extraction, patches, segmented nuclei, transformed color image, RBC segmentation, interpretability map, and / or other data obtained from the target image:
[0150] At 310, the results of the target pre-analysis factor(s) used to process the target image are obtained from the pre-analysis machine learning model.
[0151] At 312, the target pre-analysis factor(s) may be, for example, presented on a display, stored in a data storage device (e.g., as an image tag), and / or forwarded to another process for input and / or further processing.
[0152] Alternatively or additionally, in 314A, the target image, the pre-analysis factor(s), and optionally one or more additional data obtained as described with reference to 306 are provided to a secondary machine learning model.
[0153] The pre-analysis factor(s) inputs supplied to the secondary machine learning model may be obtained as a result of the pre-analysis machine learning model being supplied with at least the target image, as described with reference to 310.
[0154] At 314B, results of the target secondary instruction are obtained from the secondary machine learning model.
[0155] In 314C, the subject may be treated with a therapy effective for the medical condition according to the target secondary instructions, for example, if the secondary score is above a threshold, the subject may be treated with chemotherapy.
[0156] At 316A, in response to the target pre-analysis factor being classified as anomalous, the target image and the target pre-analysis factor(s) are provided to an image correction machine learning model and / or an image translation ML model.
[0157] It should be noted that the pre-analytical factor classification is not necessarily binary, e.g., normal-abnormal. In some cases, binary classification is not necessarily possible, e.g., when the pre-analytical factor is applied to the whole tissue, to a block, when it is not reversible or incremental, and / or when there is no specific "right" or "wrong" but rather a range of possibilities. There may be multiple categories, e.g., three or more classifications, which may depend on the particular pre-analytical factor. For example, if the pre-analytical factor is time, there may be five categories, e.g., 0 to 9 hours, 9 to 20 hours, 20 to 60 hours, 60 to 120 hours, and more than 120 hours.
[0158] For an image translation ML model, the target source image may include an input image and additional metadata representing source pre-analysis factors that represent the state of the input image. The source pre-analysis factors may be an acquired indication such as normal, abnormal, or another classification result acquired as in 310. For example, the source pre-analysis factors may represent abnormal processing. Another optional metadata represents destination pre-factors for the desired result image to be generated, for example, to generate an image where processing is performed for a selected classification category such as 20 to 60 hours to generate an image that is normally processed. For example, the target source image has a pre-analysis factor of 9 to 20 hours, and an image representing 20 to 60 hours is desired. The metadata may be explicit, for example, automatically generated and / or selected by a user. The metadata may be implicit as a default, for example, the desired pre-analysis factors for the result image may be those that are normal, or optimal or otherwise "best" pre-analysis factors. Alternatively or additionally, if no explicit metadata is provided, the target source image may include an input image that does not have explicit metadata. Optionally, a reference image from the destination set is used to infer a destination for the input image.
[0159] There may be multiple image translation ML models and / or different image correction ML models trained on different source sets and / or different training sets, e.g., different training sets representing different pre-analysis factors. For example, the image translation ML model and / or the image correction ML model may be selected and / or the source set may be selected according to the input of the pre-analysis factors obtained as a result of the pre-analysis machine learning model being fed with the target image.
[0160] The target pre-analysis factors may be classified as normal or abnormal, for example, by applying a set of rules to the target pre-analysis factors obtained as a result of the pre-analysis ML model. In another example, applying a range and / or a threshold defines a correct value for the target pre-analysis factor. When the target pre-analysis factor is within the range or below (or above) the threshold, the target pre-analysis factor is classified as normal, and when the target pre-analysis factor is outside the range or above (or below) the threshold, the target pre-analysis factor is classified as abnormal. In another example, the result of the pre-analysis ML model may include a classification label indicating whether the target pre-analysis factor is classified as normal or abnormal. To obtain such a result, the records of the pre-analysis training dataset may include a ground truth indication of normality or abnormality for each pre-analysis factor of each record.
[0161] The pre-analysis factor(s) inputs provided to the image correction machine learning model and / or the image translation ML model may be obtained as a result of the pre-analysis machine learning model being provided with the target image, as described with reference to 310.
[0162] In 316B, a corrected image result is obtained as a result of the image correction machine learning model, simulating what the target image of the slide would look like when processed with the pre-analysis factor(s) classified as normal.
[0163] Alternatively or additionally, resultant destination images of pathology slides of the destination set of image translation records that are translations of abnormally processed target images to normally processed images are obtained from the image translation ML model.
[0164] Various embodiments, embodiments, and aspects of the present invention as outlined above and as claimed in the claims section below find experimental and / or computational support in the following examples. EXAMPLES
[0165] Reference is now made to the following examples which, together with the above description, illustrate certain embodiments and / or embodiments of the invention in a non-limiting fashion.
[0166] The inventors have conducted experiments to investigate at least some embodiments of a machine learning model that is trained to generate results representative of fixation time in response to images of fixed specimens of tissue and / or features extracted from the images, as described herein.
[0167] material Access to tissue from freshly prepared porcine tissue was obtained by the University of Copenhagen. As described herein, we considered fixation time to be the main effector on staining quality results. Therefore, we prepared a training dataset where we had full control over ischemia time, with the only variable being fixation time in neutral buffered formalin. In total, 144 blocks were created, representing 6 different fixation times across 8 different organ systems performed in trial number 3. For feasibility, section specimens were cut from blocks from the liver tissue organ system and stained with eosin and hematoxylin (H&E) using a standardized protocol on a Dako Coverstainer instrument. The specimens were scanned on a Phillips UltraFast slide scanner to create a training dataset of whole slide images, which we tested using various machine learning computational approaches described herein below. We successfully trained several networks that were able to distinguish between fixation times.
[0168] method We evaluated a first feature extraction approach in which features are extracted from patches taken from a whole slide image (WSI) tuned to be within the material cross-section. Features were extracted using a pre-trained feature extractor, e.g., a deep neural net, or some other feature extraction mechanism, as described herein. These features were then used to train a pre-analysis machine learning model, such as a classification and / or regression model, to infer fixation time, as described herein. We evaluated two pre-trained networks for feature extraction: ResNet18 and UNet.
[0169] ResNet18 is a publicly available image classifier trained on the ImageNet dataset, from which we extracted the final feature map before the classification layer. The patches extracted using ResNet18 were of size 224x224x3. The features extracted from ResNet18 have a vector dimension of 512.
[0170] UNet is a customized pre-trained nuclear segmentation network from which we extracted the bottleneck layer. The patches extracted using the customized UNet network were of size 256x256x3. The features extracted from UNet have a vector dimension of 2048.
[0171] From each whole slide image, a rectangular region of interest (ROI) of 10,000 to 20,000 pixels horizontally (magnification X40) was selected for extraction. Patches were extracted within a grid covering the ROI. We attempted to extract both partially overlapping and non-overlapping patches. The extracted features of each patch were either stitched together to create an extracted feature map or saved as individual features.
[0172] The extracted features or feature maps were split into training / validation datasets in a 5-fold cross-validation (CV) fashion. For each CV split, all features extracted from the same WSI were selected together, either all for training or all for validation. If feature maps were extracted rather than individual features, the feature maps were split during training into a grid of non-overlapping or partially overlapping feature patches. A variety of feature patch grids were used, ranging in spatial dimensions from 1x1 to 20x20.
[0173] We trained neural networks using different architectures to classify the extracted dataset according to the fixed time. The architectures we explored were Convolutional Neural Networks (CNNs) and Fully Connected Neural Networks (FCNNs). CNNs consisted of one or more convolutional layers with one or more subsequent fully connected layers.
[0174] When training the FCNN, each feature patch was spatially reduced to a feature vector using either a global max pooling layer or a global average pooling layer. When training the CNN, convolutional layers operated directly on the input feature patches. The network was trained using a standard cross-entropy loss. Model performance was evaluated by measuring the F1 score of each validation fold, with the best average F1 score across folds achieved being ~0.7.
[0175] We also evaluate an alternative pipeline in which WSI patches were fed directly into a customized CNN without prior feature extraction. The final layer was a classification layer with outputs for each of the various fixation times. From each WSI, a rectangular region of interest (ROI) of 10,000 to 20,000 pixels horizontally (magnification X40) was selected for extraction, and patches were extracted within a grid covering the ROI 256x256 RGB patches. Patches were split into training and validation sets based on either the WSI slides or a random distribution of patches.
[0176] The loss function was the standard cross-entropy loss. The accuracy score for random patch selection was high (>95%), whereas it was significantly lower (<60%) when the validation / training split was performed at the level of the WSI, corresponding to the results obtained using feature extraction.
[0177] The description of various embodiments of the present invention has been presented for illustrative purposes, but they are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used in this specification are selected to best explain the principles of the embodiments, practical applications, or technical improvements to the art found in the market, or to enable another skilled in the art to understand the embodiments disclosed herein.
[0178] It is expected that during the life of the patent resulting from this application, many relevant ML models will be developed, and the scope of the term ML model is intended to include a priori all such new technologies.
[0179] As used herein, the term "about" refers to ±10%.
[0180] The terms "comprise," "comprising," "including," "including," "having," and their cognates mean "including but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of."
[0181] The phrase "consisting essentially of" means that a composition or method may include additional elements and / or steps, provided that the additional elements and / or steps do not materially alter the basic and novel characteristics of the composition or method recited in the claim.
[0182] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" may include a plurality of compounds, including mixtures thereof.
[0183] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments and / or as excluding the incorporation of features from other embodiments.
[0184] The word "optionally" is used herein to mean "provided in some embodiments and not provided in other embodiments." Any particular embodiment of the invention may include more than one "optional" feature, unless such features conflict.
[0185] Throughout this application, various embodiments of the present invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity, and should not be construed as an inflexible limitation on the scope of the present invention. Thus, the description of a range should be considered to have specifically disclosed all possible subranges as well as individual numerical values within that range. For example, a description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0186] Whenever a numerical range is given herein, it is meant to include any recited numbers (fractional or integer) within the given range. The phrases "ranging between" a first designated number and a second designated number, and "ranging from" a first designated number to a second designated number, are used interchangeably herein and are meant to include the first and second designated numbers and all fractional and integer numbers therebetween.
[0187] It should be understood that certain features of the invention that are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention that are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as appropriate in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperable without those elements.
[0188] While the present invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
[0189] It is the intention of the applicant(s) that all publications, patents, and patent applications mentioned herein shall be incorporated herein in their entirety by reference as if each individual publication, patent, or patent application was specifically and individually set forth herein by reference. In addition, citation or identification of any reference in this application should not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document(s) of this application are incorporated herein in their entirety by reference.
Claims
1. 1. A computer-implemented method for training a pre-analysis factor machine learning model, comprising: creating a pre-analysis training dataset of a plurality of records, the pre-analysis records comprising: an image of a slide of the subject's pathology tissue processed with at least one pre-analytical factor; and a ground truth label representing the at least one pre-analysis factor; and training the pre-analysis machine learning model on the pre-analysis training dataset to generate, in response to input of a target image, a result of at least one target pre-analysis factor used to process tissue represented in the target image; A computer-implemented method comprising:
2. creating a secondary training data set of a plurality of records, the secondary records comprising: the image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; the at least one pre-analytical factor; and The ground truth labels representing the secondary instructions are including the steps: training a secondary machine learning model on the secondary training data set to generate a target secondary indication result in response to an input of a target image and at least one target pre-analysis factor used to process tissue represented in the target image; The computer-implemented method of claim 1 further comprising:
3. 3. The computer-implemented method of claim 2, wherein the secondary training dataset comprises a clinical indication training dataset, the secondary indication comprises a clinical indication, and the secondary machine learning model comprises a clinical machine learning model.
4. The computer-implemented method of claim 3 , wherein the clinical indication is selected from the group comprising a clinical score, a medical condition, and a pathology report.
5. 5. The computer-implemented method of claim 4, further comprising treating the subject with a therapy effective for the medical condition according to the clinical score and / or according to the pathology report.
6. The computer-implemented method of claim 2 , wherein the ground truth labels are selected from the group consisting of tags, metadata, images, and segmentation results of a segmentation model fed with the images.
7. 3. The computer-implemented method of claim 2, wherein the input of the at least one pre-analysis factor supplied to the secondary machine learning model is obtained as the result of the pre-analysis machine learning model being supplied with the target image.
8. The computer-implemented method of claim 2 , wherein the pre-analysis machine learning model and the secondary machine learning model are jointly trained using at least common images and common labels of pre-analysis factors.
9. 3. The computer-implemented method of claim 2, wherein the at least one pre-analysis factor of the secondary record includes at least one feature map extracted from a hidden layer of the pre-analysis machine learning model fed with the image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor, and the secondary machine learning model generates the result of the target secondary instruction in response to an input of the target image and a target feature map extracted from a hidden layer of the pre-analysis machine learning model fed with the target image.
10. creating an image translation training dataset comprising two or more sets of image translation records, The source image translation records in the source set of image translation records are: a source image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; and Ground truth representing source labels Including, The destination image translation record in the destination set of the image translation record is a destination image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; and Ground truth representing destination labels and training an image translation machine learning model on the image translation training dataset to transform target source images of histopathology slides in the source set of image translation records to result destination images of histopathology slides in the destination set of image translation records; The computer-implemented method of claim 1 further comprising:
11. 11. The computer-implemented method of claim 10, wherein the source label represents pathological tissue that has been abnormally processed with the at least one pre-analysis factor and the destination label represents pathological tissue that has been normally processed with the at least one pre-analysis factor.
12. 12. The computer-implemented method of claim 11, wherein the target source image includes additional metadata representing the input image and abnormally processed source pre-analysis factors, and metadata representing normally processed destination pre-analysis factors.
13. 11. The computer-implemented method of claim 10, wherein the target source images include input images, further comprising providing a reference image from the destination set used to infer the destination for the input image.
14. 11. The computer-implemented method of claim 10, wherein the source set is selected according to input of the at least one pre-analysis factor obtained as a result of the pre-analysis machine learning model supplied with the target image.
15. creating an image correction training dataset of a plurality of records, the image correction records comprising: the image of the slide of the subject's pathological tissue processed with the at least one pre-analytical factor, wherein the at least one pre-analytical factor is classified as abnormal, and the image of the slide represents abnormally processed pathological tissue; the at least one pre-analytical factor; and Ground truth labels representing normal images of histopathology slides processed with at least one pre-analysis factor classified as normal. and training an image correction machine learning model on the image correction training dataset to generate a resultant synthetic corrected image of the pathology slide that simulates what the target image of the slide would look like when processed with the at least one target pre-analysis factor classified as normal, in response to the target image of the slide processed with at least one target pre-analysis factor classified as abnormal; The computer-implemented method of claim 1 further comprising:
16. 16. The computer-implemented method of claim 15, wherein the input of the at least one pre-analysis factor provided to the image correction machine learning model is obtained as the result of the pre-analysis machine learning model being provided with the target image.
17. 16. The computer-implemented method of claim 15, wherein the image correction machine learning model and the pre-analysis machine learning model are jointly trained using common images and common ground truth labels of pre-analysis factors.
18. 10. The computer-implemented method of claim 1, further comprising training a baseline model using a self-supervised and / or unsupervised technique on an unlabeled training dataset of a plurality of unlabeled images of pathological tissue from a subject processed with at least one pre-analysis factor, wherein training comprises further training the baseline model on the pre-analysis training dataset to create the pre-analysis machine learning model.
19. 2. The computer-implemented method of claim 1, wherein the ground truth labels representing the at least one pre-analysis factor include ground truth labels representing correctly applied pre-analysis factors or anomalous application of pre-analysis factors, and wherein training includes training an implementation of the pre-analysis machine learning model to learn a distribution of inlier images labeled as correctly applied pre-analysis factors to detect images as outliers representing incorrectly applied pre-analysis factors.
20. 2. The computer-implemented method of claim 1, further comprising extracting features from the image using a pre-trained feature extractor, wherein the pre-analysis record includes the extracted features, and wherein the pre-trained feature extractor is applied to the target image to obtain extracted target features that are fed to the pre-analysis machine learning model.
21. 21. The computer-implemented method of claim 20, wherein the pre-trained feature extractor is implemented as a neural network, and the extracted features are obtained from at least one feature map prior to a classification layer of the neural network when the neural network is fed the target image.
22. 22. The computer-implemented method of claim 21, wherein the neural network is an image classifier that is trained on an image training dataset of non-tissue images labeled with ground truth classification categories.
23. 23. The computer-implemented method of claim 22, wherein the neural network is a nuclei segmentation network trained on a segmentation training dataset of images of pathology slides labeled with a ground truth segmentation of nuclei.
24. 21. The computer-implemented method of claim 20, further comprising extracting a plurality of patches from the image, wherein extracting features comprises extracting features from the plurality of patches.
25. 25. The computer-implemented method of claim 24, further comprising: for each patch, using a global max pooling layer and / or a global average pooling layer to reduce the extracted features extracted from the patch to a feature vector, wherein the pre-analysis record includes the feature vector, and wherein the pre-analysis machine learning generates the result of at least one target pre-analysis factor in response to the input of the feature vector calculated for features extracted from the patch of the target image.
26. 2. The computer-implemented method of claim 1, further comprising the steps of: for each pre-analysis record, feeding the image to a nucleus segmentation machine learning model to obtain a segmentation result of nuclei in the image; creating a mask that covers pixels outside the segmentation of the nuclei based on the result of the segmentation; and applying the mask to the image to create a cover image, wherein the images of the pre-analysis record include the cover image, and the target cover image created from the target image is fed to the pre-analysis machine learning model trained on the pre-analysis training dataset.
27. 2. The computer-implemented method of claim 1, further comprising: for each pre-analysis record, feeding the image to a nuclei segmentation machine learning model to obtain segmentations of nuclei in the image; and cropping a boundary around each segmentation to create single nuclei patches, wherein the images of the pre-analysis record include a plurality of single nuclei patches, and the target segmentations of nuclei created from the target image are fed to the pre-analysis machine learning model trained on the pre-analysis training dataset.
28. 2. The computer-implemented method of claim 1, further comprising, for each pre-analysis record, converting a color version of the image to a grayscale version of the image, wherein the target grayscale version of the target image is fed to the pre-analysis machine learning model trained on the pre-analysis training dataset.
29. 2. The computer-implemented method of claim 1, further comprising: for each pre-analysis record, providing the image to an (RBC) segmentation machine learning model to obtain segmentation results of red blood cells (RBCs) and / or patches representing RBCs in the image, wherein the image of the pre-analysis record includes the segmentation of RBCs and / or patches representing RBCs, and the target segmentation of RBCs and / or patches representing RBCs from the target image are provided to the pre-analysis machine learning model trained on the pre-analysis training dataset.
30. 2. The computer-implemented method of claim 1, wherein the pre-analysis machine learning model is pre-trained on a separate image training dataset including a plurality of images each labeled with a respective ground truth indication of a particular classification category, and the pre-trained pre-analysis training dataset is further trained on the pre-analysis training dataset.
31. 2. The computer-implemented method of claim 1, wherein the pre-analysis record further includes metadata representing at least one known pre-analysis factor, the ground truth label is for at least one unknown pre-analysis factor, and the at least one known pre-analysis factor associated with the target image is further supplied to the pre-analysis machine learning model trained on the pre-analysis training dataset.
32. 2. The computer-implemented method of claim 1, further comprising: training an interpretability machine learning model to generate an interpretability map representing the relative importance of pixels of the target image to obtain the at least one target pre-analysis factor, wherein the target image is low resolution; and sampling a plurality of high-resolution patches of the target image; and feeding the plurality of high-resolution patches to the pre-analysis machine learning model to obtain the at least one target pre-analysis factor.
33. The computer-implemented method of claim 1 , wherein the at least one pre-analysis factor comprises a fixed time.
34. The computer-implemented method of claim 1 , wherein the at least one pre-analysis factor comprises a tissue thickness obtained by cutting the FFPE block.
35. 2. The computer-implemented method of claim 1, wherein the at least one pre-analysis factor is selected from the group consisting of fixative type, warm ischemia time, cold ischemia time, duration and delay of temperature during pre-fixation, fixative formulation, fixative concentration, fixative pH, duration of fixation reagent, source of fixative preparation, tissue to fixative volume ratio, fixation method, primary and secondary fixation conditions, post-fixation wash conditions and duration, post-fixation storage reagent and duration, processor type, frequency of maintenance and reagent changes, tissue to reagent volume ratio, number of co-processed sample locations, dehydration and clearing reagent, dehydration and clearing temperature, number of dehydration and clearing changes, dehydration and clearing duration, bake time, and temperature.
36. The computer-implemented method of claim 1 , wherein the at least one pre-analysis factor is an indication of the quality of staining of the histopathology of the slide.
37. The stains include immunohistochemistry (IHC) stains, in situ hybridization (ISH) stains, fluorescent ISH (FISH), chromogenic ISH (CISH), silver ISH (SISH), hematoxylin and eosin (H&E), hematoxylin, acridine orange, Bismarck brown, carmine, Coomassie blue, cresyl violet, crystal violet, 4',6-diamidino-2-phenylindole ("DAPI"), eosin, ethidium bromide intercalate, acid fuchsin, Hoechst stain, iodine , malachite green, methyl green, methylene blue, neutral red, Nile blue, Nile red, osmium tetroxide, propidium iodide, rhodamine, safranine, antibody-based stains, or label-free imaging markers obtained using imaging techniques including Raman spectroscopy, near-infrared ("NIR") spectroscopy, autofluorescence imaging, and phase imaging that highlight features of interest without the use of external dyes.
38. 10. The computer-implemented method of claim 1, wherein the slide comprises formalin-fixed, paraffin-embedded (FFPE) tissue.
39. 1. A computer-implemented method for obtaining at least one pre-analysis factor of a target image of a histopathology slide of a subject, the method comprising: providing a target image to a pre-analysis machine learning model, The pre-analysis machine learning model is trained on a pre-analysis training dataset of a plurality of records, the pre-analysis records comprising: an image of a slide of the subject's pathology tissue processed with at least one pre-analytical factor; and a ground truth label representing the at least one pre-analysis factor; and obtaining results of at least one target pre-analysis factor used to process the pathological tissue represented in the target image; A computer-implemented method comprising:
40. providing the target image and the at least one target pre-analysis factor to a secondary machine learning model; The secondary machine learning model is trained on a secondary instruction training dataset of a plurality of records, the secondary instruction records comprising: the image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; the at least one pre-analytical factor; and Ground truth labels representing the secondary instructions and obtaining a result of the target secondary instruction; 40. The computer-implemented method of claim 39, further comprising:
41. In response to classifying the at least one target pre-analysis factor as an anomaly, providing the target image and the at least one target pre-analysis factor to an image correction machine learning model, The image correction machine learning model is trained on a corrected image training dataset of a plurality of records, the image correction records comprising: the image of the slide of the subject's pathological tissue processed with the at least one pre-analytical factor, wherein the at least one pre-analytical factor is classified as abnormal, and the image of the slide represents abnormally processed pathological tissue; the at least one pre-analytical factor; and Ground truth labels representing normal images of histopathology slides processed with at least one pre-analysis factor classified as normal. and obtaining a corrected image result that simulates what the target image of the slide would look like when processed with the at least one pre-analysis factor classified as normal; 40. The computer-implemented method of claim 39, further comprising:
42. In response to classifying the at least one target pre-analysis factor as an anomaly, providing the target image and the at least one target pre-analysis factor to an image translation machine learning model, the image translation machine learning model is trained on an image translation training dataset comprising two or more sets of image translation records; The source image translation records in the source set of image translation records are a source image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; and Ground truth representing source labels Including, The destination image translation record in the destination set of the image translation record is a destination image of the slide of the subject's pathology tissue processed with the at least one pre-analysis factor; and Ground truth representing destination labels and obtaining a resulting destination image of the pathology slide of the destination set of image translation records, which is a translation of the abnormally processed target image to a normally processed image; 40. The computer-implemented method of claim 39, further comprising:
43. 1. A device for training a pre-analysis factor machine learning model, comprising: at least one hardware processor executing code, said code comprising: creating a pre-analysis training dataset of a plurality of records, the pre-analysis records comprising: an image of a slide of the subject's pathology tissue processed with at least one pre-analytical factor; and a ground truth label representing the at least one pre-analysis factor; and training the pre-analysis machine learning model on the pre-analysis training dataset to generate results of at least one target pre-analysis factor used to process tissue represented in the target image in response to the input of the target image; The device is for.
44. 1. A device for acquiring at least one pre-analysis parameter of a target image of a histopathology slide of a subject, comprising: at least one hardware processor executing code, said code comprising: providing the target image to a pre-analysis machine learning model, The pre-analysis machine learning model is trained on a pre-analysis training dataset of a plurality of records, the pre-analysis records comprising: an image of a slide of the subject's pathology tissue processed with at least one pre-analytical factor; and a ground truth label representing the at least one pre-analysis factor; and obtaining results of at least one target pre-analysis factor used to process the pathological tissue represented in the target image; The device is for.