Label-free assessment of biomarker expression using vibrational spectroscopy

By combining vibrational spectroscopy imaging and machine learning algorithms, the problem of reliance on manual examination in pathological sample analysis has been solved, enabling label-free and automated estimation of biomarker expression and improving the accuracy and efficiency of sample evaluation.

CN114270174BActive Publication Date: 2026-03-10VENTANA MEDICAL SYSTEMS INC
View PDF 35 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing methods for analyzing pathological samples rely on manual examination, which is affected by the operator's skills and experimental conditions, leading to inaccurate diagnoses, especially in borderline and similar cases of diseases with unpredictable outcomes. Furthermore, the methods of sample preparation and preservation also affect the results.

Method used

By employing vibrational spectral imaging technology and training a biomarker expression estimation engine, biomarker expression can be estimated label-free based on vibrational spectral data. Combined with machine learning algorithms such as dimensionality reduction and neural networks, automated biomarker expression prediction can be achieved, adapting to different preparation conditions and fixation states.

Benefits of technology

It enables accurate and reproducible assessment of biomarker expression in unstained samples, reduces human error, improves the reliability and efficiency of diagnosis, and is adaptable to samples with different preparation and fixation states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270174B_ABST
    Figure CN114270174B_ABST
Patent Text Reader

Abstract

This disclosure relates to automated systems and methods for predicting the expression of one or more biomarkers in a biological sample. In some embodiments, the sample is a sample having an unknown fixed state or a sample subjected to a fixed duration of unknown duration. In some embodiments, the predicted expression is a quantitative estimate of the positive percentage of one or more biomarkers. In other embodiments, the predicted expression is a quantitative estimate of the staining intensity of one or more biomarkers. In some embodiments, the system and method utilize a trained biomarker expression estimation engine that has been trained with multiple training samples, wherein the trained biomarker expression estimation engine is adapted to derive biomarker expression features from the sample. In some embodiments, the trained biomarker expression estimation engine includes a machine learning algorithm based on a projection to a latent structural regression model. In some embodiments, the trained biomarker expression estimation engine includes a neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims the benefit of U.S. Patent Application No. 62 / 892,680, filed on August 28, 2019, the disclosure of which is incorporated herein by reference in its entirety. Background Technology

[0003] Over the past few years, disease diagnosis based on the interpretation of tissue or cell samples collected from diseased organisms has seen tremendous development. Now, in addition to traditional tissue staining techniques and immunohistochemical (IHC) assays, in situ techniques such as in situ hybridization (ISH) and in situ polymerase chain reaction (PCR) are also used to aid in the diagnosis of human disease states and to elucidate gene expression sites in tissue sites. Therefore, a variety of techniques can assess not only cell morphology but also the presence of specific molecules (e.g., DNA, RNA, and proteins) within cells and tissues. Each of these techniques requires sample cell or tissue preparation procedures, which may include fixing the sample with chemicals such as aldehydes (e.g., formaldehyde, glutaraldehyde), formalin substitutes, alcohols (e.g., ethanol, methanol, isopropanol); or embedding the sample in an inert material such as paraffin, collodion, agar, polymers, resins, cryogenic media, or various plastic embedding media (e.g., epoxy and acrylic resins). Other sample tissue or cell preparation requires physical manipulation, such as freezing (frozen tissue sections) or fine-needle aspiration (FNA).

[0004] Subsequently, the sample cells or tissue are embedded in a solid medium (usually paraffin) to obtain one or more well-preserved two-dimensional sections. These sections are typically 3–7 μm thick and are placed on a glass slide for a microscope. The slide is then cleaned and stained using specific methods, ready for microscopic observation or pre-imaging. A trained pathologist then analyzes the stained sample to determine tissue morphology and changes, such as those caused by disease, the expression of one or more biomarkers, etc.

[0005] Pathologists are increasingly using molecular techniques to aid in tissue characterization and disease diagnosis. Immunohistochemical (IHC) staining is used to identify proteins in cells within tissue sections and is therefore widely used to study different cell types, such as cancer cells and immune cells in biological tissues. Thus, IHC staining can be used to investigate the distribution and localization of differentially expressed biomarkers of immune cells (such as T cells or B cells) in cancerous tissues for immune response studies. For example, tumors often contain infiltrations of immune cells, which may inhibit tumor development or promote tumor growth.

[0006] In situ hybridization (ISH) can be used to determine the presence of genetic abnormalities or, for example, the specific amplification of oncogenes in cells that appear morphologically malignant under a microscope. ISH uses labeled DNA or RNA probe molecules that are antisense to the target gene sequence or transcript to detect or locate targeted nucleic acid genes within cell or tissue samples. ISH is performed by exposing a cell or tissue sample immobilized on a glass slide to labeled nucleic acid probes that specifically hybridize with a given target gene in the cell or tissue sample. Multiple target genes can be analyzed simultaneously by exposing a cell or tissue sample to multiple nucleic acid probes labeled with multiple different nucleic acid tags. Using labels with different emission wavelengths, simultaneous multicolor analysis of a single target cell or tissue sample can be performed in a single step.

[0007] Analyzing histological and cytological samples to identify disease is a manual process that requires spatial morphology recognition. For example, pathologists must identify morphology and assess cellular details in any histopathological or cytological sample. Through these visual cues, pathologists determine diagnostic information from the sample, such as assessing evidence of cancer and / or characterizing its severity. Many problems in pathology are believed to stem from the manual nature of stained samples. Furthermore, sample quality and preparation are also believed to affect a pathologist's ability to accurately assess a sample. Similarly, IHC and ISH staining rely on operator skill and experimental conditions and methods for accurate diagnosis. Worse still, unpredictable borderline and similar cases can further exacerbate potential problems when evaluating samples. Regardless of the tissue or cellular sample, or its preparation or preservation methods, the goal of both technicians and pathologists is to obtain accurate, readable, and reproducible results for accurate interpretation of the data. Summary of the Invention

[0008] A robust, automated method for detecting diseases and their spatial morphology is highly needed. As mentioned above, clinicopathological techniques employ histological or cytological staining to reveal morphological patterns in biomedical samples. Typically, obtaining individual tissue sections for each target biomarker is expensive and time-consuming. On the other hand, vibrational spectral imaging is believed to provide information about multiple biomarkers from a single tissue section.

[0009] This disclosure describes systems and methods for estimating the expression of one or more biomarkers (e.g., positive percentage, staining intensity) in biological sample samples. In some embodiments, this disclosure provides systems and methods that allow for completely label-free molecular analysis of biomarkers in biological samples. In some embodiments, the estimation of the expression of one or more biomarkers in a sample is based on the identification of biomarker expression characteristics present in vibrational spectral data acquired from the biological sample. In some embodiments, the biomarker expression characteristics present in vibrational spectral data acquired from the biological sample are identified using a trained biomarker expression estimation engine; and the estimated expression of one or more biomarkers (e.g., positive percentage; staining intensity) can be calculated based on the expression characteristics of those identified biomarkers. Therefore, the systems and methods of this disclosure can enable “label-free” diagnostics (e.g., predicting the expression of one or more biomarkers in a biological sample without staining in IHC or ISH assays). It goes without saying that while the currently available systems and methods can be used alone to provide “label-free” diagnostics, they can also be used in combination with one or more IHC and / or ISH assays, for example, to provide further analysis of samples on the same or sequential sections of formalin-fixed, paraffin-embedded tissue (FFPET) samples.

[0010] In some embodiments, the biological samples are unstained. In these embodiments, the systems and methods of this disclosure enable the estimation of biomarker expression in unstained samples, for example, for samples with an unknown fixed duration or an unknown demasking state. In other embodiments, the biological samples are stained for the presence of one or more biomarkers, such as one biomarker, two biomarkers, three biomarkers, four or more biomarkers.

[0011] This disclosure also describes systems and methods for training a biomarker expression estimation engine capable of label-free quantitative estimation of the expression of one or more biomarkers in biological samples based on real-world data, such as training vibrational spectral data containing one or more class labels. In some embodiments, the training vibrational spectral data includes differentially prepared biological samples, such as biological samples that have been differentially fixed and / or differentially demasked. In this way, a biomarker expression estimation engine can be trained to estimate the different degrees of expression of one or more biomarkers in prepared (e.g., fixed and / or demasked) biological samples (e.g., variable fixed samples; variable demasked samples). As described herein, sample preparation can affect biomarker expression, and the systems and methods described herein for estimating biomarker expression take into account this variability. These and other embodiments are described in more detail herein, for example.

[0012] One aspect of this disclosure is a system for predicting the expression of one or more biomarkers in a test biological sample, the system comprising: (i) one or more processors, and (ii) one or more memories coupled to the one or more processors, the one or more memories storing computer-executable instructions, which, when executed by the one or more processors, cause the system to perform operations including: obtaining test spectral data from the test biological sample, wherein the obtained test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine; and predicting the expression of the one or more biomarkers in the test biological sample based on the derived biomarker expression features. In some embodiments, the test biological sample is unstained. In some embodiments, the test biological sample is stained for the presence of one or more biomarkers.

[0013] In some embodiments, predicted biomarker expression includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, predicted biomarker expression includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, the fixation state (e.g., fixation quality, fixation duration) of the test biological sample is unknown. In some embodiments, the demasking state (e.g., demasking quality) is unknown.

[0014] In some embodiments, the biomarker expression estimation engine is trained using one or more training spectral datasets, each training spectral dataset comprising multiple training vibrational spectra derived from multiple training tissue samples, each training tissue sample being stained for the presence of one or more biomarkers, and each training vibrational spectrum comprising one or more class labels. In some embodiments, the one or more class labels contain known biomarker expression levels of one or more biomarkers. In some embodiments, the known biomarker expression level comprises at least one of a known positive percentage of one or more biomarkers and a known staining intensity of one or more biomarkers. In some embodiments, the system further comprises one or more additional class labels selected from the group consisting of a known demasking duration, a known demasking temperature, a qualitative assessment of the demasking state, a known fixed duration, and a qualitative assessment of the fixed state.

[0015] In some embodiments, the training spectral dataset is derived by: (i) obtaining training biological samples; (ii) dividing the obtained training biological samples into multiple training tissue samples; (iii) staining each of the multiple training tissue samples for the presence of one or more biomarkers; and (iv) quantitatively assessing the expression of one or more biomarkers. In some embodiments, each training tissue sample is differentially prepared prior to staining. In some embodiments, each of the multiple training tissue samples is differentially demasked, differentially fixed, or differentially demasked and differentially fixed. In some embodiments, the quantitative assessment of one or more biomarkers in the training tissue samples includes determining the staining intensity of one or more biomarkers. In some embodiments, the quantitative assessment of one or more biomarkers in the training tissue samples includes determining the positive percentage of one or more biomarkers. In some embodiments, the quantitative assessment is performed by a pathologist. In some embodiments, the quantitative assessment is performed using one or more image analysis algorithms. In some embodiments, the multiple training tissue samples are stained in immunohistochemical assays. In some embodiments, the multiple training tissue samples are stained in in situ hybridization assays. In some embodiments, the multiple training tissue samples are stained in multiplex assays.

[0016] In some embodiments, the test spectral data includes an average vibrational spectrum derived from a plurality of normalized and corrected vibrational spectra. In some embodiments, the plurality of normalized and corrected vibrational spectra are obtained by: (i) identifying a plurality of spatial regions within the test biological sample; (ii) acquiring vibrational spectra from each individual region of the plurality of identified regions; (iii) correcting the vibrational spectra acquired from each individual region to provide a corrected vibrational spectrum for each individual region; and (iv) normalizing the amplitude of the corrected vibrational spectra from each individual region to a predetermined global maximum value to provide an amplitude-normalized vibrational spectrum for each region. In some embodiments, the vibrational spectra acquired from each individual region are corrected by: (i) compensating each acquired vibrational spectrum for atmospheric effects to provide an atmospherically corrected vibrational spectrum; and (ii) compensating the atmospherically corrected vibrational spectrum for scattering.

[0017] In some embodiments, the trained biomarker expression estimation engine includes a dimensionality reduction-based machine learning algorithm. In some embodiments, dimensionality reduction includes projection onto a latent structural regression model. In some embodiments, dimensionality reduction includes principal component analysis plus discriminant analysis. In some embodiments, the trained biomarker expression estimation engine includes a neural network.

[0018] In some embodiments, the system further includes operations for correcting the predicted expression of one or more biomarkers for poor demasking and / or poor fixation of test biological samples. For example, the predicted expression of one or more biomarkers in the test biological sample obtained by using a trained biomarker expression estimation engine can be corrected by: (i) obtaining a biomarker fixation sensitivity curve; (ii) estimating the actual fixation time of the test biological sample; and (iii) using the obtained fixation sensitivity curve to correct the predicted biomarker expression level of the obtained test biological sample to a fixed compensating expression level.

[0019] In some embodiments, the system further includes operations for comparing the actual expression of a biomarker in a test biological sample with the predicted expression of one or more biomarkers in the test biological sample. In some embodiments, the obtained test spectral data includes vibrational spectral information of at least one amide I band. In some embodiments, the obtained test spectral data includes wavelengths in the range of about 3200 to about 3400 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from about 2800 to about 2900 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from about 1020 to about 1100 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from approximately 1520 to approximately 1580 cm⁻¹. -1 Vibrational spectral information between them.

[0020] A second aspect of this disclosure is a non-transitory computer-readable medium storing instructions for predicting the expression of one or more biomarkers in a processed test biological sample, comprising: obtaining test spectral data from the test biological sample, wherein the test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine, wherein the biomarker expression estimation engine is trained using a training spectral dataset acquired from multiple differentially prepared training biological samples, wherein the training spectral dataset includes class labels for known biomarker expressions of one or more biomarkers; and predicting the expression of another biomarker in the test biological sample based on the derived biomarker expression features. In some embodiments, the test biological sample has an unknown fixed state and / or an unknown demasking state. In some embodiments, the predicted expression of one or more biomarkers includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, the predicted expression of one or more biomarkers includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, the predicted expression of one or more biomarkers is quantitative. In some embodiments, the test biological sample is unstained. In some embodiments, the test biological sample is stained in response to the presence of one or more biomarkers.

[0021] In some embodiments, each training spectral dataset is derived by: (i) obtaining training biological samples; (ii) dividing the obtained training biological samples into multiple training tissue samples; and (iii) preparing each of the multiple training tissue samples under different preparation conditions; (iv) staining each of the multiple training tissue samples for the presence of one or more biomarkers; and (v) quantitatively evaluating the expression of one or more biomarkers. In some embodiments, the different preparation conditions include different demasking conditions. In some embodiments, the different preparation conditions include different fixed durations. In some embodiments, the training biological samples include the same tissue type as the test biological samples. In some embodiments, the training biological samples include a different tissue type than the test biological samples.

[0022] In some embodiments, the obtained test spectral data includes vibrational spectral information of at least one amide I band. In some embodiments, the obtained test spectral data includes wavelengths in the range of about 3200 to about 3400 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from approximately 2800 to approximately 2900 cm⁻¹. -1Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from about 1020 to about 1100 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from approximately 1520 to approximately 1580 cm⁻¹. -1 Vibrational spectral information between them.

[0023] A third aspect of this disclosure is a method for predicting the expression of one or more biomarkers in a test biological sample, comprising: obtaining test spectral data from the test biological sample, wherein the test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine, wherein the biomarker expression estimation engine is trained using a training spectral dataset collected from multiple differentially prepared training biological samples, and wherein the training spectral dataset includes class labels of known biomarker expressions of one or more biomarkers; and predicting the expression of one or more biomarkers in the test biological sample based on the derived biomarker expression features.

[0024] In some embodiments, predicted biomarker expression includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, predicted biomarker expression includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, one or more biomarkers include at least one cancer biomarker. In some embodiments, the test biological sample has an unknown fixed state and / or an unknown demasking state. In some embodiments, the test biological sample is unstained. In some embodiments, the test biological sample is stained in response to the presence of one or more biomarkers.

[0025] In some embodiments, each training spectral dataset is derived by: (i) obtaining training biological samples; (ii) dividing the obtained training biological samples into multiple training tissue samples; and (iii) preparing each of the multiple training tissue samples under different preparation conditions. In some embodiments, the method further includes staining each of the multiple training tissue samples for the presence of one or more biomarkers; and quantitatively assessing the known positive percentage and / or known staining intensity of one or more biomarkers.

[0026] In some embodiments, the trained biomarker expression estimation engine includes a dimensionality reduction-based machine learning algorithm. In some embodiments, dimensionality reduction includes projection onto a latent structural regression model. In some embodiments, the trained biomarker expression estimation engine includes a neural network. In some embodiments, the method further includes compensating for poor demasking and / or poor fixation of one or more biomarkers in the test biological sample. For example, the predicted expression of one or more biomarkers in the test biological sample obtained by using the trained biomarker expression estimation engine can be corrected by: (i) obtaining a biomarker fixation sensitivity curve; (ii) estimating the actual fixation time of the test biological sample; and (iii) correcting the predicted biomarker expression level of the obtained test biological sample to a fixed compensated expression level using the obtained fixation sensitivity curve.

[0027] In some embodiments, the obtained test spectral data includes vibrational spectral information of at least one amide I band. In some embodiments, the obtained test spectral data includes wavelengths in the range of about 3200 to about 3400 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from approximately 2800 to approximately 2900 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from about 1020 to about 1100 cm⁻¹. -1 Vibrational spectral information between [wavelength ranges]. In some embodiments, the obtained test spectral data includes wavelengths ranging from approximately 1520 to approximately 1580 cm⁻¹. -1 Vibrational spectral information between them. Attached Figure Description

[0028] For a general understanding of the features of this disclosure, please refer to the accompanying drawings. In the drawings, the same reference numerals are used throughout to identify the same elements.

[0029] Figure 1 illustrates a representative digital pathology system according to an embodiment of the present disclosure, the system including an image acquisition device and a computer system.

[0030] Figure 2 illustrates various modules that can be used in a system or in a digital pathology workflow, according to an embodiment of this disclosure, to quantitatively or qualitatively predict the demasking state of test biological samples.

[0031] Figure 3 illustrates a flowchart of the steps of using a trained biomarker expression estimation engine according to an embodiment of the present disclosure to estimate the expression of one or more biomarkers in an unstained test biological sample.

[0032] Figure 4A illustrates a process for obtaining multiple training tissue samples according to one embodiment of the present disclosure. For example, training samples 1, 2, 3, 4, 5, and 6 for differential preparation (e.g., for differential fixation and / or differential demasking) come from two different training biological samples. In some embodiments, training tissue samples 1, 2, and 3 belong to a first set of training tissue samples from which a first training spectral dataset can be collected; while training tissue samples 4, 5, and 6 belong to a second set of training tissue samples from which a second training dataset can be collected.

[0033] Figure 4B illustrates the differential preparation of multiple training tissue samples obtained from two different training biological samples according to an embodiment of the present disclosure, and further illustrates the preparation of two different training spectral datasets.

[0034] Figure 5A illustrates the preparation of multiple training tissue samples according to an embodiment of the present disclosure.

[0035] Figure 5B illustrates the preparation of multiple training tissue samples according to an embodiment of the present disclosure.

[0036] Figure 5C illustrates the preparation of multiple training tissue samples according to an embodiment of the present disclosure.

[0037] Figure 5D illustrates the preparation of multiple training tissue samples according to an embodiment of the present disclosure.

[0038] Figure 5E illustrates the preparation of multiple training tissue samples according to an embodiment of the present disclosure.

[0039] Figure 6 shows a flowchart illustrating the steps of acquiring vibrational spectra of training biological samples according to an embodiment of the present disclosure.

[0040] Figure 7 shows a flowchart illustrating the steps of acquiring the average vibrational spectrum of a test biological sample according to an embodiment of the present disclosure.

[0041] Figure 8 shows a flowchart illustrating various steps for correcting, normalizing, and averaging acquired spectra derived from biological samples (including test biological samples and training biological samples) according to an embodiment of the present disclosure.

[0042] Figures 9A, 9B, and 9C show the quantitative analysis of IHC expression (positive percentage) of BCL2 (Figure 9A), ki-67 (Figure 9B), and FOXP3 (Figure 9C).

[0043] Figure 9D shows a comparison of IHC expression for all three biomarkers with time at a fixed point, where mean expression levels are plotted on a normalized scale to allow observation of the relative change of each biomarker with time at a fixed point. Bars represent significance levels determined by a two-way rank-sum test (p < 0.05).

[0044] Figure 10 provides an example of tonsil tissue labeled with antiserum generated against Ki-67. Image analysis was performed only on the tonsil tissue (circled in the left image). Connective tissue that sometimes showed high background but was not present in other sections was excluded.

[0045] Figure 11 provides a visualization of an example tissue slice with multiple identified regions. The figure further provides examples of collected, averaged, processed, and normalized vibrational spectra from the indicated regions in the visualization.

[0046] Figure 12A provides the mid-IR absorption spectrum, specifically showing the protein bands within the acquired mid-IR spectrum.

[0047] Figure 12B shows the first derivative of the amide I band and the peak position of the FWHM of this band, which shows that the unrepaired tissue has a significantly different spectrum from other repaired tissues.

[0048] Figure 13 illustrates an example of training a biomarker expression estimation engine (specifically, the PLSR machine learning algorithm). Initially, the model is trained using input vibrational spectra with known classifications. A model is developed to assign weights to each wavelength, roughly corresponding to the degree of correlation (or inverse correlation) between that wavelength and the response (e.g., demasking time). Finally, the model is applied to the vibrational spectral data used during training to evaluate its accuracy in predicting demasking time.

[0049] Figure 14 shows typical FR-IR and Raman spectra of collagen.

[0050] Figure 15 illustrates a biomarker expression estimation engine based on the PLSR model, where the trained biomarker expression estimation engine (trained using acquired mid-IR spectra) can predict C4d staining. The prediction accuracy for C4d positive cells in the blinded spectrum is 0.4%.

[0051] Figure 16 illustrates a biomarker expression estimation engine based on the PLSR model, where the trained biomarker expression estimation engine (trained using acquired mid-IR spectra) can predict Ki-67 staining. The prediction accuracy for Ki-67 positive cells in the blinded spectrum is 0.8%.

[0052] Figure 17 shows photographs of four tissues imaged using mid-IR over a time-temperature process. The biomarker expression estimation engine was trained on tissues within the circled regions, which included three tissue samples (right and bottom of the figure); and the predictive power of the biomarker expression estimation engine was evaluated using tissues within a “smaller” circled region that included only one tissue sample (left of the figure).

[0053] Figure 18 shows the prediction accuracy of the trained biomarker expression estimation engine at all times and temperatures in the tonsillar blind sample. At all test times and temperatures, the trained biomarker expression estimation engine predicted functional C4d staining intensity with an accuracy greater than approximately 10%. The values ​​at the time-temperature crossover represent the percentage error between the predicted and actual C4d staining intensity.

[0054] Figure 19 provides a table listing the infrared and Raman characteristic frequencies of biological samples.

[0055] Figure 20 shows the quantitative analysis of IHC expression (staining intensity) of BCL2.

[0056] Figure 21 shows the quantitative analysis of FOXP3 IHC expression (staining intensity).

[0057] Figure 22 shows the quantitative analysis of IHC expression (staining intensity) of ki-67.

[0058] Figure 23A shows a comparison of estimated and predicted DAB staining for the BCL2 biomarker in fixation experiments. Specifically, Figure 23A provides a box-and-whisker plot of BCL2 concentrations in tissue samples fixed at room temperature NBF for different time periods (only in BCL2-positive cells), ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / appropriate fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis algorithm. The predicted concentrations represent the estimated BCL2 concentrations predicted using a trained biomarker expression estimation engine trained on the PLSR algorithm. The left box (“training”) represents BCL2 predictions made based on a MID-IR spectral training set; the right box (“Holdout”) represents BCL2 predictions made on blind spectra (e.g., validation spectra) that the model has never “seen” before. The results demonstrate that the PLSR prediction model can accurately predict BCL2 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0059] Figure 23B plots the cumulative distribution function of DAB staining for the estimation and prediction of the BLC2 biomarker shown in Figure 23A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from the analysis of bright-field images and the predicted protein concentrations calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error on the training set (“training”) is similar to the prediction error on the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0060] Figure 24A provides a box-and-whisker plot of FOXP3 concentrations in tissue samples fixed at different times in room temperature NBF (only in FOXP3-positive cells), ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / proper fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis program. The predicted concentrations represent estimated FOXP3 concentrations predicted using a trained biomarker expression estimation engine trained on the PLSR algorithm. The boxes on the left (“dashed boxes”) represent FOXP3 predictions made based on the training set MID-IR spectra, and the boxes on the right (“diagonally bordered boxes”) represent FOXP3 predictions made on blind spectra (e.g., validation spectra) that the model has never seen before. The results demonstrate that the PLSR prediction model can accurately predict FOXP3 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0061] Figure 24B plots the cumulative distribution function of DAB staining for the estimation and prediction of FOXP3 biomarkers shown in Figure 24A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from analyzing bright-field images and the predicted protein concentrations calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error for the training set (solid line) is similar to the prediction error for the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0062] Figure 25A provides a box-and-whisker plot of Ki-67 concentrations in tissue samples fixed at different times in room temperature NBF (only in Ki-67-positive cells), ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / proper fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis program. The predicted concentrations represent Ki-67 estimates predicted using a trained biomarker expression estimation engine trained on the PLSR algorithm. The boxes on the left (“dashed boxes”) represent Ki-67 predictions made based on the training set MID-IR spectra, and the boxes on the right (“diagonally bordered boxes”) represent Ki-67 predictions made on blind spectra (e.g., validation spectra) that the model has never seen before. The results demonstrate that the PLSR prediction model can accurately predict Ki-67 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0063] Figure 25B plots the cumulative distribution function of DAB staining for the estimated and predicted Ki-67 biomarkers shown in Figure 25A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from analyzing bright-field images and the predicted MID-IR protein concentration calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error for the training set (solid line) is similar to that for the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0064] Figure 26A provides a box plot of FOXP3-positive tissues fixed in room temperature NBF for different time periods, ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / proper fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis program. The predicted concentrations represent estimated FOXP3 concentrations predicted using a trained biomarker expression estimation engine trained on the PLSR algorithm. The boxes on the left (“dashed boxes”) represent FOXP3 predictions made based on the training set MID-IR spectra, and the boxes on the right (“diagonally bordered boxes”) represent FOXP3 predictions made on blind spectra (e.g., validation spectra) that the model has never seen before. The results demonstrate that the PLSR prediction model can accurately predict FOXP3 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0065] Figure 26B plots the cumulative distribution function of the estimated and predicted percentages of FOXP3 biomarker-positive tissues shown in Figure 26A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from analyzing bright-field images and the predicted MID-IR protein concentration calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error for the training set (solid line) is similar to the prediction error for the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0066] Figure 27A provides a box plot of BCL2-positive tissues fixed in room temperature NBF for different time periods, ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / appropriate fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis program. The predicted concentrations represent estimated BCL2 concentrations predicted using a trained biomarker expression estimation engine trained on the PLSR algorithm. The boxes on the left (“dashed boxes”) represent BCL2 predictions made based on the training set MID-IR spectra, and the boxes on the right (“diagonally bordered boxes”) represent BCL2 predictions made on blind spectra (e.g., validation spectra) that the model has never seen before. The results demonstrate that the PLSR prediction model can accurately predict BCL2 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0067] Figure 27B plots the cumulative distribution function of the estimated and predicted percentages of BCL2 biomarker-positive tissues shown in Figure 27A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from analyzing bright-field images and the predicted MID-IR protein concentration calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error for the training set (solid line) is similar to the prediction error for the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0068] Figure 28A is a box-and-whisker plot of the percentage of Ki-67 positive tissue in tissue samples fixed at room temperature NBF for different time periods, ranging from 0 hours (e.g., inadequate / poor fixation) to 24 hours (e.g., complete / appropriate fixation). Experimental protein concentrations were determined by analyzing bright-field images using an image analysis program. The predicted concentrations represent estimated Ki-67 concentrations predicted using a trained prediction engine based on the PLSR algorithm. The boxes on the left (“dashed boxes”) represent Ki-67 predictions made based on the training set MID-IR spectra, and the boxes on the right (“diagonally bordered boxes”) represent Ki-67 predictions made on blind spectra (e.g., validation spectra) that the model has never seen before. The results demonstrate that the PLSR prediction model can accurately predict Ki-67 concentrations in differentially fixed tissues (from incomplete to complete fixation).

[0069] Figure 28B plots the cumulative distribution function of the estimated and predicted percentage of positive tissues for the Ki-67 biomarker, as shown in Figure 25A. The horizontal axis represents the absolute value of the model error, defined as the difference between the actual protein concentration obtained from analyzing bright-field images and the predicted MID-IR protein concentration calculated using MID-IR spectra from tissues and the PLSR prediction engine. The model prediction error for the training set (solid line) is similar to that for the prediction / validation data, indicating that a well-trained model does not overfit to the noise in the MID-IR spectra.

[0070] Figure 29A shows the C4d staining results of tissue samples repaired for 30 minutes at temperatures of 9.6℃, 110℃, 120℃, 130℃, or 140℃. The left panel shows that, using a PLSR-based trained biomarker expression estimation engine, blind training facilitates the prediction of the C4d positivity percentage in all tissues, regardless of the antigen retrieval temperature, and even though the inflection point is at 120℃. The right panel shows that both staining intensity (top, curve, diamond) and positivity percentage (bottom, curve, square) increase with retrieval temperature until 130℃, at which point the detected C4d amount decreases (from the DAB image analysis algorithm).

[0071] Figure 29B shows the Ki-67 staining results for tissue samples repaired for 60 minutes at temperatures of 25°C, 70°C, 80°C, 90°C, 100°C, 105°C, or 110°C. The left panel shows that both staining intensity (diamonds) and positive percentage (squares) increase with repair temperature, but saturate near 100°C according to data from the DAB image analysis algorithm. The right panel shows that, using a PCDA-based trained biomarker expression estimation engine, MID-IR spectroscopy can be used to determine the Ki-67 positive percentage staining in all tissues, regardless of the antigen repair temperature and despite saturation at higher repair temperatures.

[0072] Figure 30A shows a flowchart illustrating the steps of correcting the predicted biomarker expression level according to an embodiment of the present disclosure.

[0073] Figure 30B shows a flowchart illustrating the steps of correcting the predicted biomarker expression level according to an embodiment of the present disclosure. Detailed Implementation

[0074] It should also be understood that, unless the opposite is specified, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are described.

[0075] References to "an embodiment," "a particular embodiment," "an illustrative embodiment," etc., in the specification indicate that the described embodiment may include a specific feature, structure, or characteristic, but each embodiment may or may not include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a particular feature, structure, or characteristic is described in connection with an embodiment, whether explicitly described or not, it is believed to affect that feature, structure, or characteristic in relation to other embodiments to the extent that those skilled in the art possesses.

[0076] As used herein, unless the context clearly indicates otherwise, the singular forms “a / an” and “the / that” include plural referents. Similarly, unless the context clearly indicates otherwise, the word “or” is intended to include “and”. The term “including” is defined as inclusive, such as “including A or B” meaning including A, B, or A and B.

[0077] As used herein in the specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” should be interpreted as inclusive, such as several elements or at least one element in a list of elements, but also including more than one element, and optionally including additional unlisted items. Only terms indicating the opposite, such as “only one of them” or “exactly one of them,” or “consisting of…” as used in the claims, will refer to several elements or exactly one element in a list of elements. In general, the term “or” as used herein should only be interpreted as indicating an exclusive alternative (e.g., “one or the other, but not two”) when preceded by exclusive terms such as “or,” “one of them,” “only one,” or “exactly one.” “Constitutes substantially of…” as used in the claims should have the ordinary meaning used in the field of patent law.

[0078] The terms “comprising,” “including,” and “having” are used interchangeably and have the same meaning. Similarly, the terms “comprising,” “including,” and “having” are used interchangeably and have the same meaning. Specifically, the definition of each term is consistent with the definition of “comprising” under ordinary U.S. patent law, and therefore each term can be understood as an open-ended term meaning “at least the following,” and can also be interpreted as not excluding additional features, limitations, aspects, etc. Thus, for example, “an apparatus having components a, b, and c” means that the apparatus includes at least components a, b, and c. Similarly, the phrase “a method relating to steps a, b, and c” means that the method includes at least steps a, b, and c. Furthermore, although the steps and processes may be described herein in a specific order, those skilled in the art will recognize that the order of steps and processes may vary.

[0079] As used herein in the specification and claims, with respect to a list of one or more elements, the phrase "at least one" should be understood as at least one element selected from any one or more elements in the list, but does not necessarily include at least one of each element specifically listed in the list, nor exclude any combination of elements in the list. In addition to the elements specifically identified in the list of elements referred to by the phrase "at least one," this definition also allows other elements to optionally be present, whether or not they are related to the specifically identified elements. Thus, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") in one embodiment may refer to at least one optional inclusion of more than one A, but no B (and optionally include elements other than B); in another embodiment, it refers to at least one optional inclusion of more than one B, but no A (and optionally include elements other than A); in yet another embodiment, it refers to at least one optional inclusion of more than one A, and at least one optional inclusion of more than one B (and optionally include other elements), and so on.

[0080] As used herein, the term “antigen” refers to a substance bound to an antibody, antibody analogue (e.g., aptamer), or antibody fragment. Antigens can be endogenous, meaning they are produced within cells as a result of normal or abnormal cellular metabolism, or due to viral or intracellular bacterial infection. Endogenous antigens include xenogeneic (heterogeneous), autologous, and idiotyped or allogeneic (homologous) antigens. Antigens can also be tumor-specific antigens or presented by tumor cells. In this case, they are called tumor-specific antigens (TSA) and are typically produced by tumor-specific mutations. Antigens can also be tumor-associated antigens (TAA), which are presented by both tumor cells and normal cells. Antigens further include CD antigens, which refer to any of a variety of cell surface markers expressed by leukocytes and can be used to distinguish cell lineages or developmental stages. Such markers can be identified by specific monoclonal antibodies and numbered by their differentiation clusters.

[0081] As used herein, the terms “biological sample,” “sample,” or “tissue sample” refer to any sample obtained from any organism, including viruses, that includes biomolecules such as proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof. Examples of other organisms include mammals (such as humans; mammals such as cats, dogs, horses, cattle, and pigs; and laboratory animals such as mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and needle biopsies of tissues), cell samples (such as cytological smears, such as cervical smears or blood smears, or cell samples obtained through microdissection), or cell fractions, fragments, or organelles (such as those obtained by lysing cells and separating their components by centrifugation or other means). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical or needle biopsy), nipple aspiration, cerumen, breast milk, vaginal secretions, saliva, swabs (e.g., oral swabs), or any material containing biomolecules derived from the first biological sample. In some embodiments, the term "biological sample" as used herein refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or a portion thereof obtained from a subject.

[0082] As used herein, the term "biomarker" or "marker" refers to a measurable indicator of a biological state or condition. Specifically, a biomarker can be a nucleic acid, lipid, carbohydrate, protein, or peptide, such as a surface protein, that can be specifically stained and indicates a cellular biological characteristic, such as cell type or physiological state. Biomarkers can be used to determine the extent of a body's response to treatment for a disease or condition, or whether a subject is susceptible to a disease or condition. Immune cell markers are biomarkers that selectively indicate characteristics associated with the immune response in mammals. In the case of cancer, a biomarker refers to a biological substance that indicates the presence of cancer in the body. Biomarkers can be molecules secreted by a tumor or a specific response of the body to the presence of cancer. Genetic, epigenetic, proteomic, carbohydrate, and imaging biomarkers can be used for the diagnosis, prognosis, and epidemiology of cancer. Such biomarkers can be measured in non-invasively collected biological fluids, such as blood or serum. Several gene- and protein-based biomarkers have been used in patient care, including but not limited to AFP (liver cancer), BCR-ABL (chronic myeloid leukemia), BRCA1 / BRCA2 (breast cancer / ovarian cancer), BRAF V600E (melanoma / colorectal cancer), CA-125 (ovarian cancer), CA19.9 (pancreatic cancer), CEA (colorectal cancer), EGFR (non-small cell lung cancer), HER-2 (breast cancer), KIT (gastrointestinal stromal tumor), PSA (prostate-specific antigen), and S100 (melanoma). Biomarkers can be used for diagnosis (identifying early-stage cancer) and / or prognosis (predicting cancer invasiveness and / or the degree of response to specific treatments and / or the likelihood of cancer recurrence).

[0083] As used herein, the term "cytological sample" refers to a cell sample in which the cells have been partially or completely depolymerized, such that the sample no longer reflects the spatial relationships of the cells (as if the cells were present in the subject from whom the cell sample was obtained). Examples of cytological samples include tissue scrapings (such as cervical scrapings), fine-needle aspirates, and samples obtained by irrigating the subject.

[0084] As used herein, the term "immunohistochemistry" refers to a method for determining the presence or distribution of an antigen in a sample by detecting the interaction of the antigen with a specific binding agent, such as an antibody. The sample is contacted with the antibody under conditions that allow antibody-antigen binding. Antibody-antigen binding can be detected by means of a detectable tag conjugated to the antibody (direct detection) or by means of a detectable tag conjugated to a secondary antibody that specifically binds to a primary antibody (indirect detection). In some instances, indirect detection may include tertiary or higher-level antibodies to further enhance the detectability of the antigen. Examples of detectable tags include enzymes, fluorophores, and haptens, which (in the case of enzymes) may be used in conjunction with chromogenic or fluorescent substrates.

[0085] As used in this article, the term "positive percentage" refers to the number of positively stained cells divided by the sum of the number of positively stained cells and the number of negatively stained cells.

[0086] As used herein, the term "slide" refers to any suitable-sized substrate on which a biological sample can be placed for analysis (e.g., a substrate made wholly or partially of glass, quartz, plastic, silicon, etc.), and more particularly to "microscope slides" such as standard 3 x 1-inch microscope slides or standard 75 mm x 25 mm microscope slides. Examples of biological samples that can be placed on a slide include, but are not limited to, cytological smears, thin tissue sections (e.g., from biopsies), and arrays of biological samples, such as tissue arrays, cell arrays, DNA arrays, RNA arrays, protein arrays, or any combination thereof. Thus, in one embodiment, tissue sections, DNA samples, RNA samples, and / or proteins are placed at specific locations on a slide. In some embodiments, the term "slide" may refer to SELDI and MALDI chips, as well as silicon wafers.

[0087] As used herein, the term "specific binding entity" refers to a member of a specific binding pair. A specific binding pair is a pair of molecules characterized by binding to each other to substantially exclude binding to other molecules (e.g., the binding constant of a specific binding pair can be at least 10 greater than the binding constant of either of the two members of a binding pair of other molecules in a biological sample). 3 M -1 10 4 M -1 Or 10 5 M -1 Specific examples of specific binding moieties include specific binding proteins (e.g., antibodies, lectins, streptavidin, and avidin proteins such as protein A). Specific binding moieties may also include molecules (or portions thereof) specifically bound by such specific binding proteins.

[0088] As used herein, the term "spectral data" includes raw image spectral data acquired from biological samples or any part thereof, for example, using a spectrometer.

[0089] As used herein, the term "spectrum" refers to information (absorption, transmission, reflection) obtained "at" or within a certain wavelength or wavenumber range of electromagnetic radiation. The wavenumber range can be as large as 4000 cm⁻¹. -1 It can also be as small as 0.01 cm. -1 Please note that measurements performed at a so-called “single laser wavelength” typically cover a small spectral range (e.g., laser linewidth), and therefore this spectral range is included whenever the term “spectrum” is used throughout the document. For example, transmission measurements at a fixed wavelength setting of a quantum cascade laser should fall under the term “spectrum” in this application.

[0090] As used herein, the terms “staining,” “performing staining,” or similar terms generally refer to any treatment of a biological sample that detects and / or distinguishes the presence, location, and / or amount (e.g., concentration) of a specific molecule (e.g., lipid, protein, or nucleic acid) or a specific structure (e.g., normal or malignant cell, cytoplasm, nucleus, Golgi apparatus, or cytoskeleton) in the biological sample. For example, staining can compare a specific molecule or cellular structure of a biological sample to surrounding parts, and the intensity of the stain can provide a determination of the amount of the specific molecule in the sample. Staining can be used not only with bright-field microscopy but also with other observational tools (such as phase-contrast microscopy, electron microscopy, and fluorescence microscopy) to aid in the observation of molecules, cellular structures, and organisms. Some systematic staining can make the outlines of cells clearly visible. Other staining by said systems may depend on staining specific cellular components (e.g., molecules or structures) that do not stain other cellular components or stain relatively little of other cellular components. Examples of various staining methods by said systems include, but are not limited to, histochemical methods, immunohistochemical methods, and other methods based on intermolecular reactions (including non-covalent binding interactions) such as hybridization reactions between nucleic acid molecules. Specific staining methods include, but are not limited to, primary staining methods (such as H&E staining, cervical staining, etc.), enzyme-linked immunohistochemistry methods, and in situ RNA and DNA hybridization methods, such as fluorescence in situ hybridization (FISH).

[0091] As used herein, the term "target" refers to any molecule whose presence, location, and / or concentration can be determined. Examples of target molecules include proteins, epitopes, nucleic acid sequences, and haptens, such as haptens that covalently bind to proteins. Typically, target molecules are detected using conjugates of one or more specific binding molecules and a detectable label.

[0092] As used herein, the term “tissue sample” refers to a cell sample that preserves the cross-sectional spatial relationships between cells as if the cells were present in the subject from whom the cell sample was obtained. “Tissue sample” includes original tissue samples (e.g., cells and tissues generated by the subject) and xenografts (e.g., foreign cell samples implanted in the subject).

[0093] As used herein, the term "unmask" or "unmasking" refers to the restoration of an antigen or target in a fixed tissue and / or the improvement of the detection of antigens, amino acids, peptides, proteins, nucleic acids, and / or other targets. For example, antigenic sites that might otherwise go undetected may be revealed, for instance, by disrupting some protein cross-links surrounding the antigen during unmasking. In some embodiments, antigens and / or other targets are unmasked by applying one or more unmasking agents (as defined below), heat, and / or pressure. In some embodiments, only one or more unmasking agents are applied to the sample to achieve unmasking. In other embodiments, only heat is applied to achieve unmasking. In some embodiments, unmasking may occur only in the presence of water and heat. Example of an unmasking operation is described in U.S. Patent Publication No. 2009 / 01700152 (the entire disclosure of which is incorporated herein by reference).

[0094] Overview

[0095] In some embodiments, this disclosure relates to systems and methods capable of enabling “label-free” diagnostics, such as predicting the expression of biomarkers in unstained biological samples, as in IHC and / or ISH assays. In some embodiments, the systems and methods disclosed herein utilize a trained biomarker expression estimation engine to evaluate vibrational spectral data acquired from biological samples, and based on the evaluation of the vibrational spectral data, provide an estimated expression of one or more biomarkers as output.

[0096] In some embodiments, the output of the disclosed systems and methods is a quantitative estimate of the staining intensity of one or more biomarkers, or a quantitative estimate of the positive percentage of one or more biomarkers. In some embodiments, quantitative estimates of the staining intensity and / or positive percentage of one or more biomarkers can be provided for biological samples prepared under unknown conditions, such as when the duration of fixation and / or the demasking state of the biological sample is unknown.

[0097] In summary, the applicant proposes that the disclosed system and method can rapidly and accurately predict the expression of one or more biomarkers in unstained biological samples using machine learning algorithms, ultimately facilitating improved IHC and / or ISH assay results and patient care. It is believed that the system and method can also save time and costs because, in some embodiments, staining assays are not required. Furthermore, also in some embodiments, the assessment of the expression of one or more biomarkers is unaffected by inconsistencies in sample preparation or IHC and / or ISH analysis. These and other embodiments are described in more detail herein.

[0098] system

[0099] At least some embodiments of this disclosure relate to a computer system for analyzing vibrational spectral data acquired from biological samples. In some embodiments, the test biological sample is stained for the presence of one or more biomarkers. In some embodiments, the test biological sample is unstained.

[0100] In some embodiments, the biological sample has an unknown fixed state and / or unmasked state. According to this disclosure, a trained biomarker expression estimation engine can be used to provide quantitative estimates of the expression of one or more biomarkers within a biological sample (e.g., an unstained test biological sample). In some embodiments, the system of this disclosure can receive test vibrational spectral data from a test biological sample (e.g., an unstained test biological sample) as input and can provide quantitative estimates of the expression of one or more biomarkers as output, including positive percentage or staining intensity. In some embodiments, and depending on how the biomarker expression estimation engine is trained, in addition to the estimation of biomarker expression, the trained biomarker expression estimation engine can also provide quantitative or qualitative estimates of one or both of the fixed state and / or unmasked state as output.

[0101] In some embodiments, the output may be in the form of a generated report. In other embodiments, the output may be an overlay map superimposed on an image of the test biological sample. In other embodiments, any output may be stored in a memory associated with the system (e.g., storage system 240) and may be associated with the test biological sample and / or other patient data.

[0102] Figures 1 and 2 illustrate a system 200 for acquiring spectral data (e.g., vibrational spectral data) and biological samples (including test biological samples and training biological samples) for analysis. The system may include a spectral acquisition device 12, such as a spectral acquisition device configured to acquire vibrational spectra (e.g., mid-IR spectra or Raman spectra) of biological samples (or any portion thereof), and a computer 14, whereby the spectral acquisition device 12 and the computer may be communicatively coupled together (e.g., directly or indirectly via network 20). The computer system 14 may include a desktop computer, laptop computer, tablet computer, or the like, digital electronic circuitry, firmware, hardware, memory 201, computer storage medium (240), computer programs or instruction sets (such as programs stored in the memory or storage medium), one or more processors (209) (including programming processors), and any other hardware, software, or firmware modules or combinations thereof (as further described herein). For example, the system 14 shown in Figure 1 may include a computer having a display device 16 and a housing 18. The computer system can store the collected spectral data locally, such as in a memory, on a server, or on another device connected to a network.

[0103] Vibrational spectroscopy involves transitions caused by the absorption and emission of electromagnetic radiation. These transitions are believed to occur between 10² and 10⁴ cm⁻¹. -1 The vibrations originate from the vibrations of the nuclei that make up the molecules in any given sample. It is believed that chemical bonds in a molecule can vibrate in many ways, and each vibration is called a vibrational mode. There are two types of molecular vibrations: stretching and bending. Stretching vibrations are characterized by movement along the bond axis as the distance between atoms increases or decreases, while bending vibrations involve changes in bond angles relative to the rest of the molecule. Two widely used vibrational energy-based spectroscopic techniques are Raman spectroscopy and infrared spectroscopy. Both mid-infrared (MIR) absorption spectroscopy and Raman spectroscopy utilize the inelastic scattering of laser light to probe specific vibrational energy levels of molecules within a target volume. These two techniques are complementary, probing different vibrational modes based on vibrational selection rules and the fact that within any molecule, atoms vibrate at some well-defined frequencies. When a sample is irradiated with an incident beam of radiation, the sample absorbs energy at frequencies characterized by the vibrational frequencies of the chemical bonds in the molecule. This absorption of energy through the vibrations of chemical bonds produces an infrared spectrum.

[0104] Although both IR and Raman spectroscopy can measure the vibrational energy of molecules, these two methods rely on different selection rules, such as absorption processes and scattering effects. While the contrast mechanisms of these two methods differ, and each has its own advantages and disadvantages, the synthetic spectra from each mode are generally correlated (see, for example, Figure 14 and...). Figure 19 ).

[0105] Infrared spectroscopy is based on the absorption of electromagnetic radiation, while Raman spectroscopy relies on the inelastic scattering of electromagnetic radiation. Infrared spectroscopy offers a wide range of analytical tools, from absorption to reflection and dispersive techniques, extending to a broad range of wavenumbers and including the near, mid, and far-infrared regions. The diverse bonds present in sample molecules provide many general and characteristic bands suitable for both qualitative and quantitative purposes. In IR spectroscopy, the sample is illuminated with IR light, and vibrations induced by the electric dipole moment are detected.

[0106] Raman spectroscopy is a scattering phenomenon caused by the difference between the incident and scattered radiation frequencies. It utilizes scattered light to obtain information about molecular vibrations, providing insights into molecular structure, symmetry, electronic environment, and bonding. In Raman spectroscopy, a sample is illuminated with monochromatic visible or near-IR light from a laser source, and its vibrations during changes in polarizability are determined.

[0107] Any vibrational spectral acquisition device can be used in the system disclosed herein. Examples of suitable spectral acquisition devices or components of such devices for acquiring mid-infrared spectra are described in U.S. Patent Publications 2018 / 0109078a and 2016 / 0091704, and U.S. Patents 10,041,832, 8,036,252, 9,046,650, 6,972,409, and 7,280,576 (the entire contents of which are incorporated herein by reference).

[0108] Any suitable method can be used to generate representative mid-infrared spectra for biological samples. Fourier transform infrared spectroscopy and its biomedical applications are discussed in the following literature, for example, in P. Lasch, J. Kneipp (Eds.) "Biomedical Vibrational Spectroscopy" 2008 (John Wiley & Sons). However, recently tunable quantum cascade lasers have enabled rapid spectroscopic and microscopic examination of biomedical samples due to their high spectral power density (see N. Kröger et al., "Biomedical Vibrational Spectroscopy VI: Advances in Research and Industry", edited by A. Mahadevan-Jansen and W. Petrich, Proceedings of the SCO 8939, 89390Z; N. Kröger et al., J. Biomed Opt. 19 (2014) 111607; N. Kröger-Lui et al., Analyst 140 (2015) 2086). The entire contents of each of the aforementioned publications are incorporated herein by reference. This work is believed to represent an advancement in applicability (compared to the aforementioned infrared microscopy setups) due to significantly faster imaging speeds (e.g., 5 minutes instead of 18 hours) at a considerably reduced cost, the elimination of liquid nitrogen cooling, and the provision of more pixels per image. A particular advantage of QCL-based microscopy in the absence of stained tissue quality assessment is the larger field of view (compared to FT-IR imaging), achieved by a microbolometer array detector, for example, 640 x 480 pixels.

[0109] In some embodiments, the spectrum may be obtained over a wide wavelength range, one or more narrow wavelength ranges, or only at a single wavelength or a combination thereof. For example, spectra of amide I and amide II bands may be acquired. As another example, the spectrum may be obtained in the range from about 3200 to about 3400 cm⁻¹. -1 Approximately 2800 to approximately 2900 cm -1 Approximately 1020 to approximately 1100 cm -1 and / or about 1520 to about 1580 cm -1 The spectrum is collected at wavelengths ranging from about 3200 to about 3400 cm⁻¹. In some embodiments, the spectrum can be collected in the range of about 3200 to about 3400 cm⁻¹. -1 Spectra can be acquired at wavelengths ranging from about 2800 to about 2900 cm⁻¹. In some embodiments, the range can be from about 2800 to about 2900 cm⁻¹. -1The spectrum is collected at wavelengths ranging from about 1020 to about 1100 cm⁻¹. In some embodiments, the range can be from about 1020 to about 1100 cm⁻¹. -1 The spectrum is collected at wavelengths ranging from about 1520 to about 1580 cm⁻¹. In some embodiments, the range can be from about 1520 to about 1580 cm⁻¹. -1 The spectrum is acquired at specific wavelengths. It is believed that narrowing the spectral range is generally advantageous in terms of acquisition speed, especially when using quantum cascade lasers. In some embodiments, individual tunable lasers are tuned one after another to their respective wavelengths. Alternatively, a set of fixed-frequency, untunable lasers can be used, allowing wavelength selection to be accomplished by turning on and off any laser required for a specific frequency measurement.

[0110] Spectra can be acquired using measurements (e.g., transmission or reflection). For transmission measurements, barium fluorite, calcium fluoride, silicon, polymer films, or zinc selenide are typically used as substrates. For reflection measurements, gold- or silver-plated substrates and standard microscope slides or slides coated with mid-infrared reflective coatings (e.g., multilayer dielectric coatings or thin silver coatings) are commonly used. Furthermore, surface-enhancing techniques (e.g., SEIRS), such as structured surfaces like nanoantennas, can be employed.

[0111] In some embodiments, other computer devices or systems may be utilized, and the computer systems described herein may be communicatively coupled to additional components, such as microscopes, imaging devices, scanners, other imaging systems, automated slide preparation equipment, etc. Some of these additional components, as well as various available computers, networks, etc., will be further described herein.

[0112] For example, in some embodiments, system 200 may further include an imaging device and any images captured from the imaging device may be stored in binary form, such as locally or on a server. In some embodiments, the captured images may be stored together with biomarker expression estimates and / or any patient data, such as in storage subsystem 240. The captured digital images may also be divided into a pixel matrix. The pixels may include one or more bits of digital value defined by bit depth. Generally, the imaging device (or other image source including pre-scanned images stored in memory) may include, but is not limited to, one or more image capture devices. Image capture devices may include, but are not limited to, cameras (such as analog cameras, digital cameras, etc.), optics (such as one or more lenses, sensor focusing lens groups, microscope objectives, etc.), imaging sensors (such as charge-coupled devices (CCDs), complementary metal-oxide-semiconductor (CMOS) image sensors, etc.), photographic film, etc. In digital embodiments, the image capture device may include multiple lenses that cooperate to demonstrate instantaneous focusing capability. The image sensor, for example, a CCD sensor, may capture digital images of the sample.

[0113] In some embodiments, the imaging apparatus is a bright-field imaging system, a multispectral imaging (MSI) system, or a fluorescence microscopy system. The digitized tissue data can be generated, for example, by an image scanning system such as the VENTANA MEDICALSYSTEMS, Inc. (Tucson, Arizona) VENTANA DP200 scanner, or other suitable imaging devices. Other imaging devices and systems will be described further herein. In some embodiments, the digital color image acquired by the imaging apparatus is typically composed of basic color pixels. Each color pixel can be encoded on three digital components, each containing the same number of bits, and each component corresponds to a primary color, typically red, green, or blue, also referred to by the term "RGB" components.

[0114] Figure 2 provides an overview of the system 200 of this disclosure and the various modules used within the system. In some embodiments, the system 200 employs a computer device or computer-implemented method having one or more processors 209 and one or more memories 201, the one or more memories 201 storing non-transitory computer-readable instructions for execution by the one or more processors to cause the one or more processors to perform specific instructions as described herein.

[0115] In some embodiments, and as described above, the system includes a spectral acquisition module 202 for acquiring vibrational spectra, such as mid-IR spectra or Raman spectra, or, for example, step 310 of FIG. 3, of a biological sample (see, for example, step 320 of FIG. 3). In some embodiments, the system 200 further includes a spectral processing module 212 adapted to process the acquired vibrational spectral data. In some embodiments, the spectral processing module 212 is configured to preprocess the spectral data. In some embodiments, the spectral processing module 212 corrects and / or normalizes the acquired vibrational spectra, or converts the acquired transmission spectra into absorption spectra. In other embodiments, the spectral processing module 212 is configured to average multiple acquired vibrational spectra from a single biological sample. In other embodiments, the spectral processing module 212 is configured to further process any acquired vibrational spectra, such as calculating the first derivative, second derivative, etc., of the acquired vibrational spectra.

[0116] In some embodiments, the system 200 further includes a training module 211 adapted to receive training vibrational spectral data and to use the received training vibrational spectral data to train a biomarker expression estimation engine 210.

[0117] In some embodiments, system 200 includes a biomarker expression estimation engine 210 trained to detect biomarker expression features within test vibrational spectral data (see, for example, step 340 of FIG. 3) and provide an estimate of biomarker expression in a biological sample (e.g., staining intensity or positive percentage) based on the detected biomarker expression features (see, for example, step 350 of FIG. 3). In some embodiments, biomarker expression estimation engine 210 includes one or more machine learning algorithms. In some embodiments, the one or more machine learning algorithms are based on dimensionality reduction as further described herein. In some embodiments, dimensionality reduction utilizes principal component analysis, such as principal component analysis with discriminant analysis. In other embodiments, dimensionality reduction is projected onto latent structure regression. In some embodiments, biomarker expression estimation engine 210 includes a neural network. In other embodiments, biomarker expression estimation engine 210 includes a classifier, such as a support vector machine.

[0118] In some embodiments, additional modules may be incorporated into the workflow or system 200. In some embodiments, an image acquisition module is run to acquire digital images of biological samples or any portion thereof. In other embodiments, automated image analysis algorithms may be run to detect, classify, and / or score cells (see, for example, U.S. Patent Publication No. 2017 / 0372117, the entire contents of which are incorporated herein by reference). Other suitable image analysis algorithms are described in PCT Publications WO / 2019 / 121564, WO / 2019 / 110583, WO / 2019 / 110567, WO / 2019 / 110561, WO / 2019 / 025533, WO / 2019 / 025515, and WO / 2018 / 122056 (the entire contents of which are incorporated herein by reference).

[0119] Spectral acquisition module and acquired spectral data

[0120] Referring to Figure 2, in some embodiments, system 200 operates spectral acquisition module 202 to acquire vibrational spectra from at least a portion of a biological sample (e.g., a test biological sample or a training biological sample) using spectral imaging device 12, such as any of those described above. In other embodiments, the test biological sample (further described herein) is unstained; for example, the sample does not include any staining indicating the presence of one or more biomarkers. In some embodiments, and for training biological samples (further described herein), the biological sample is stained for the presence of one or more biomarkers. Once vibrational spectra have been acquired using spectral acquisition module 202, the acquired vibrational spectra can be stored in storage module 240 (e.g., a local storage module or a network storage module).

[0121] In some embodiments, vibrational spectra can be acquired from a portion of a biological sample (and this is independent of whether the sample is a training or test biological sample, as further described herein). In this case, the spectral acquisition module 202 can be programmed to acquire vibrational spectra from a predetermined portion of the sample, for example, by random sampling or by sampling at regular intervals across a grid covering the entire sample. The spectral acquisition module is also useful when only a specific region of the sample is relevant to the analysis.

[0122] For example, compared to another target region, a target region may include a specific type of tissue or a relatively high number of a specific type of cells. For instance, a target region might selectively include tonsillar tissue but exclude connective tissue. In this case, the spectral acquisition module 202 can be programmed to acquire vibrational spectra from predetermined portions of the target region, for example, by randomly sampling the target region or by sampling at regular intervals across a grid covering the entire target region. In embodiments where the sample includes one or more stains, vibrational spectra can be obtained from those target regions that do not contain any stains or include less staining than other regions.

[0123] In some embodiments, at least two regions of a biological sample are sampled, and vibrational spectra are acquired for each of the at least two regions (again, this is independent of whether the sample is a training or test biological sample). In other embodiments, at least 10 regions of a biological sample are sampled, and vibrational spectra are acquired for each of the at least 10 regions. In other embodiments, at least 30 regions of a biological sample are sampled, and vibrational spectra are acquired for each of the at least 30 regions. In a further embodiment, at least 60 regions of a biological sample are sampled, and vibrational spectra are acquired for each of the at least 60 regions. In a further embodiment, at least 90 regions of a biological sample are sampled, and vibrational spectra are acquired for each of the at least 90 regions. In a further embodiment, even approximately 30 to approximately 150 regions of a biological sample are sampled, and vibrational spectra are acquired for each region.

[0124] In some embodiments, a single vibrational spectrum is acquired from each region of the biological sample. In other embodiments, at least two vibrational spectra are acquired from each region of the biological sample. In still other embodiments, at least three vibrational spectra are acquired from each region of the biological sample.

[0125] In some embodiments, the acquired vibrational spectra or acquired vibrational spectral data (which may be used interchangeably herein) stored in storage module 240 include “training spectral data”. In some embodiments, the training spectral data is derived from training biological samples, wherein the training biological samples may be histological samples, cytological samples, or any combination thereof.

[0126] In some embodiments, training spectral data is used to train the biomarker expression engine 210, for example, by using training module 211 as described herein. In some embodiments, the training spectral data includes category labels such as biomarker expression levels (e.g., positive percentage, staining intensity), demasking status (e.g., demasking time, demasking duration, relative demasking quality information, such as "not repaired," "fully repaired," and "partially repaired"), fixation status (e.g., fixation duration, relative fixation quality, such as "partially fixed," "fully fixed," "fully fixed," and "not sufficiently fixed"), etc. In some embodiments, the training spectral data includes multiple category labels. In some embodiments, category labels include identification of tissue type, specific binding agent used in any staining assay, tissue preparation information, patient information, etc.

[0127] In some embodiments, multiple training vibrational spectral datasets are used to train the biomarker expression estimation engine. In some embodiments, each training spectral dataset can be derived from a single training biological sample, which is divided into multiple parts (see Figure 4A), such as multiple training tissue samples (e.g., a first training tissue sample, a second training tissue sample, and an nth training tissue sample), and each training tissue sample can be differentially prepared. For example, and as further described below, each training tissue sample can be differentially prepared, e.g., differential staining, differential fixation, and / or differential demasking (see Figure 4B). In this respect, a single training biological sample can generate multiple differentially prepared samples representing a continuum of different conditions and / or tissue preparation states. Of course, each different training vibrational spectral dataset may come from different subjects or patients, may come from different tissue types (e.g., a comparison of tonsil tissue and breast tissue), and / or may be treated with different specific binding entities (e.g., a comparison of specific binding entities recognizing CD8 biomarkers and specific binding entities recognizing CD3 biomarkers; a comparison of specific binding entities recognizing CD8 from a first manufacturer and specific binding entities recognizing CD8 from a second manufacturer).

[0128] In some embodiments, training biological samples and each training tissue sample derived therefrom are stained for the presence of one or more biomarkers, enabling the assessment of biomarker expression (e.g., positive percentage and / or staining intensity) in each training sample (e.g., by a trained pathologist or using one or more image analysis algorithms). For example, each individual training sample may be stained with one or more of BCL2, C4d, Ki-67, FOXP3, etc. Other biomarkers suitable for detection and classification are described herein.

[0129] In some embodiments, each training tissue sample is stained for the presence of a single biomarker, and then an imaging device is used to capture and analyze images of the training tissue samples (so that the staining intensity and / or positive percentage of the biomarker in each individual training tissue sample can be determined). In other embodiments, each training tissue sample is stained for the presence of two or more biomarkers, and then an imaging device is used to capture and analyze images of the training tissue samples (again, so that the staining intensity and / or positive percentage of each of the two or more biomarkers can be analyzed independently). For the training tissue samples stained for the presence of two or more biomarkers, the captured images of these training tissue samples may first be unmixed, and then each unmixed image channel image can be evaluated, so that the staining intensity and / or positive percentage can be assessed by the staining signal present in a particular unmixed image channel image. PCT Publication WO / 2019 / 110583 describes a method of unmixing, the disclosure of which is incorporated herein by reference in its entirety.

[0130] In some embodiments, the preparation of any training tissue sample, including sample fixation and demasking steps of intrasample targets (e.g., protein and / or nucleic acid targets), may have an impact on biomarker expression. Example 1 of this paper illustrates the effect of fixation time on the expression of three different biomarkers, namely BLC2, Ki-67, and FOXP3, particularly the effect of fixation time on the measured positive percentage (see also Figures 9A-9D). Similarly, Figures 20-22 show the effect of fixation time on the staining intensity of these three identical biomarkers.

[0131] Example 2 in this paper similarly illustrates the effect of demasking quality on the expression of the Ki-67 or C4d biomarkers. As further described in Example 2, it shows that different biomarkers may exhibit different responses to increased demasking treatment. For example, C4d staining intensity and the number of labeled cells decrease after reaching a certain point. Conversely, even under demasking conditions that would otherwise impair the biological sample, Ki67 continuously increases intensity and positivity throughout the duration of the applied demasking process until saturation is reached (see, for example, the dots and associated tissue images in Figure 15).

[0132] In view of the foregoing, in some embodiments, the training vibrational spectral dataset may include training tissue samples that have been differentially fixed and / or differentially demasked as described below. In this way, a biomarker expression estimation engine can be trained using training spectral data spanning a continuum of differentially fixed and / or demasked states, enabling the biomarker expression estimation engine to determine the expression of one or more biomarkers in an unstained test biological sample, regardless of the actual fixed and / or demasked state of the test biological sample, and / or regardless of whether the fixed and / or demasked state of the test biological sample is known or unknown.

[0133] In some embodiments, training biological samples are differentially fixed. Differential fixing is a process in which each of a plurality of training tissue samples (each from a single training biological sample as described above) undergoes a different fixation process. In some embodiments, any training tissue sample can be fixed for any predetermined amount of time, such as 1 hour, 2 hours, 4 hours, 6 hours, 12 hours, etc. In this respect, the plurality of training tissue samples can each be partially fixed (e.g., not treated with fixative for a duration sufficient to make the sample appear “fully fixed” or “well fixed”), to varying degrees. Furthermore, the group of training tissue samples may include tissue samples that have not yet been fixed (e.g., fixed for 0 hours).

[0134] In some embodiments, training biological samples are differentially demasked. Differential fixation is a process in which each training tissue sample (each from a single training biological sample as described above) undergoes different demasking conditions, such as different demasking reagents, different demasking durations, different demasking temperatures, and / or different demasking pressures. For example, in some embodiments, multiple training samples derived from a single training biological sample are demasked at the same temperature, but for different durations. For example, each training tissue sample from a single training biological sample may be demasked at the same temperature (e.g., 98.6°C), but the duration of demasking may vary (5 minutes, 30 minutes, 60 minutes, etc.).

[0135] As another example, and in other embodiments, multiple training tissue samples derived from a single training biological sample are masked for the same duration but at different temperatures. For example, each training tissue sample may be masked for the same duration (e.g., 10 minutes) but at different temperatures (98.6°C, 110°C, 120°C, 130°C, etc.). In some embodiments, both the masking time and temperature may be varied. As in the embodiments described above, a first group of training tissue samples may be masked at a first temperature but for different durations, thus providing a first group of training tissue samples. A second and third group of training tissue samples may be masked at a second and a third temperature, respectively, and similarly for different durations, thus providing a second and a third group of training tissue samples.

[0136] In some embodiments, a single training biological sample may be divided into multiple training tissue samples, and each individual training tissue sample may be (i) fixed for the same predetermined duration (e.g., 12 hours), but (ii) demasked differently. In some embodiments, each tissue sample may be fixed for a period of time, which will provide “sufficient” or “complete” fixation. Figure 5A illustrates the above.

[0137] As an example, and referring again to Figure 5A, "Pre-fixed 1" can be a fixed duration of 12 hours; "Staining 1" can refer to one or more stainings applied to the training tissue sample; and "Demasking conditions 1, 2, 3, and 4" can each have a duration of 10 minutes, but the demasking temperatures are different, for example, 98.6°C, 110°C, 120°C, and 130°C, respectively. Although Figure 5A shows the preparation and acquisition of a single set of training spectral data, multiple additional training spectral datasets can be prepared and acquired similarly, but any of the fixed duration, demasking conditions, applied staining, tissue type, etc., can vary.

[0138] In other embodiments, a single training biological sample can be divided into two groups of training tissue samples, with each group comprising multiple individual training tissue samples. Following this particular example, each of the first group of training tissue samples can be immobilized for a period of time, providing a sample considered "fully immobilized." Each of the individual training tissue samples in the first group can then be differentially demasked. Similarly, each of the second group of training tissue samples can be immobilized for a period of time, providing a sample considered "insufficiently immobilized." Each of the individual training tissue samples in the second group can then be differentially demasked. Figure 5B illustrates the above.

[0139] In other embodiments, a single training biological sample may be divided into multiple training tissue samples, and each individual training tissue sample may be (i) differentially fixed (e.g., 12 hours), but (ii) demasked under the same demasking conditions. Figure 5C illustrates the above. In some embodiments, the demasking conditions may be those considered to "sufficiently" demask the sample, taking into account the fixed duration and the given tissue type and demasking reagent used.

[0140] In some embodiments, the length of the fixation process can be a determining factor for any conditions used in the demasking process (e.g., a longer demasking time may be required for samples that have been fixed for a longer duration). Therefore, in a further embodiment, a single training biological sample can be divided into multiple training tissue sample groups, and each different training tissue sample group comprises multiple individual training tissue samples, and each different training tissue sample group is fixed for a different duration.

[0141] Within each distinct group of training tissue samples with a fixed predetermined duration, each individual training tissue sample can be differentially demasked, as shown in Figure 5D. In this way, each of these differentially fixed training tissue samples can be demasked for a predetermined amount of time and under predetermined conditions that allow each sample to be "sufficiently" demasked. In other words, each differentially fixed sample can be demasked for a specific amount of time and under set conditions to ensure that the particular training tissue sample is "sufficiently" demasked. Each training tissue sample can then be stained for the presence of one or more biomarkers.

[0142] Figure 5E illustrates a flowchart of the process of obtaining one or more training spectral datasets from training biological samples with a fixed duration of unknown time. Here, the training biological samples are separated, dissimilatory masked, and stained for the presence of one or more biomarkers. The resulting stained training tissue samples are then imaged to detect and / or classify cells, followed by the acquisition of vibrational spectra for each training tissue sample. The resulting dataset (e.g., images, class labels, vibrational spectral data, etc.) can be stored on a server or other storage device for later retrieval. Example 3 further describes the method. The applicant has found that even training biological samples with unknown fixed durations in a biomarker expression estimation engine is valuable. Indeed, as shown in Figures 15 and 16 and as described in Example 3, a biomarker expression estimation engine trained solely on training spectral datasets derived from training biological samples with unknown fixed durations is able to estimate one or more biomarkers in test biological samples with high accuracy.

[0143] Figure 6 illustrates the process of acquiring spectral data from differentially prepared samples stained for the presence of one or more biomarkers. As described above, one or more training biological samples are first acquired (step 410). Then, each of the one or more training biological samples is divided into at least two parts (step 420). In this way, each of the one or more training biological samples provides at least two “training tissue samples.” Each of these training tissue samples can be differentially prepared, for example, each can be differentially fixed and / or differentially demasked (step 430). After differential preparation of the at least two training tissue samples, staining is performed for the presence of one or more biomarkers in each of the at least two training tissue samples, including protein and / or nucleic acid biomarkers (step 435). After staining, multiple regions in each of the at least two differentially prepared and stained training tissue samples are identified (step 440).

[0144] Next, at least one vibrational spectrum is acquired for each of the multiple identification regions (step 450). The average of each acquired vibrational spectrum from each identification region (or a variant thereof, as further described below) is calculated to provide the average vibrational spectrum of the training sample (step 460). Steps 400 to 460 can be repeated for multiple different training biological samples (see dashed line 470). In some embodiments, the average vibrational spectra of all training tissue samples from all training biological samples (referred to as the “training spectral dataset”) are stored (step 480), for example, in storage module 240. In this way, training module 211 can retrieve training spectral data or training spectral datasets from storage module 240 for training the biomarker expression estimation engine 210. In addition to storing the average vibrational spectra from all training samples, storage module 240 is also suitable for storing any category labels associated with the average vibrational spectra (e.g., actual measured expression of one or more biomarkers (e.g., as assessed by a pathologist or determined using one or more image analysis algorithms), demasking state, fixed state, etc.).

[0145] The process described above for preparing training biological samples and acquiring spectral data from these samples (see step 470) can be repeated for multiple different training biological samples, each of which may be of the same tissue type or may be of different tissue types (e.g., tonsil tissue or breast tissue). The Examples section of this document further describes the method for preparing training biological samples and the acquisition of spectral data for training the biomarker expression estimation engine 210.

[0146] In some embodiments, the acquired spectral data stored in storage module 240 includes "test spectral data". In some embodiments, the test spectral data is derived from a test biological sample, such as a sample derived from a subject (e.g., a human patient), wherein the test biological sample can be a histological sample, a cytological sample, or any combination thereof. In some embodiments, the test spectral data is derived from an unstained test sample. In other embodiments, the test spectral data is derived from a biological sample stained for the presence of one or more biomarkers.

[0147] Referring to Figure 7, a test biological sample can be obtained (step 510), and then multiple spatial regions within the test biological sample can be identified (step 520). At least one vibrational spectrum can be acquired for each identified region (step 530). The vibrational spectra acquired from all regions can then be corrected, normalized, and averaged to provide an average vibrational spectrum of the test biological sample (“test spectral data”). As further described herein, the test spectral data can be provided to a trained biomarker expression estimation engine 210, enabling the prediction of the expression of one or more biomarkers within the test biological sample. The predicted expression of one or more biomarkers can then be used in downstream processes or downstream decision-making, such as sample scoring, where scored samples can be used to guide treatment protocols. In some embodiments, the test biological sample has been fixed for unknown quantities over time and / or has been demasked under unknown conditions.

[0148] As described above, regardless of whether the spectral data is collected from training or testing biological samples, multiple vibrational spectra are collected from each biological sample, for example, to interpret the spatial heterogeneity of space samples. In some embodiments, each collected vibrational transmission spectrum is first converted into a vibrational absorption spectrum using the spectral processing module 212. In some embodiments, the transmission and absorption spectra are directly correlated by the equation absorbance = ln(blank transmission / transmission through tissue), thus the collected transmission spectrum can be converted into an absorption spectrum.

[0149] In some embodiments, once all vibrational spectra have been converted from transmission to absorption spectra, the spectral processing module 212 averages all acquired spectra from all different regions, and the averaged vibrational spectra are used for downstream analysis, e.g., for training or predicting biomarker expression. In some embodiments, and referring to FIG8, vibrational spectra acquired from each of the multiple spatial regions are first normalized and / or corrected before they are averaged. In some embodiments, vibrational spectra from each region are individually corrected (step 620) to provide corrected vibrational spectra. For example, correction may include compensating each acquired vibrational spectrum for atmospheric effects (step 630), and then compensating each atmospherically corrected vibrational spectrum for scattering (step 640). Next, each corrected vibrational spectrum is normalized, e.g., to a maximum value of 2 to mitigate differences in sample thickness and tissue density (step 650). Subsequently, the set of amplitude-normalized spectra is averaged (step 660).

[0150] Biomarker expression estimation engine

[0151] The system and method disclosed herein employ machine learning techniques to mine spectral data. When the biomarker expression estimation engine is in training mode, it can learn features from multiple acquired and processed training vibrational spectra (e.g., training vibrational spectra stored in storage module 240) and correlate those learned features with category labels associated with the training spectra (e.g., known biomarker expression of one or more biomarkers, known demasking temperature, known demasking duration, tissue quality, etc.). With the trained biomarker expression estimation engine, it can derive biomarker expression features from unstained test biological samples and, based on the learned dataset, predict the expression of one or more biomarkers within the unstained test biological samples based on the derived biomarker expression features.

[0152] Machine learning can generally be defined as a type of artificial intelligence (AI) that gives computers the ability to learn without explicit programming. Machine learning focuses on developing computer programs that can learn and change by being exposed to new data. In other words, machine learning can be defined as a subfield of computer science that gives computers the ability to learn without explicit programming.

[0153] Machine learning explores the research and construction of algorithms that can learn from data and make predictions. These algorithms overcome the problem of strictly adhering to static procedural instructions by building models from sample inputs to make data-driven predictions or decisions. The machine learning described herein can be further described in “Introduction to Statistical Machine Learning,” by Sugiyama and Morgan Kaufmann, 2016, 534 pages; “Discriminative, Generative, and Imitative Learning,” Jebara, MIT Thesis, 2002, 212 pages; and “Principles of Data Mining (Adaptive Computation and Machine Learning),” Hand et al., MIT Press, 2001, 578 pages; which are incorporated herein by reference as if fully expounded herein. The embodiments described herein can be further configured as described in these references.

[0154] In some embodiments, the biomarker expression estimation engine 210 employs "supervised learning" to predict biomarker expression in test spectra derived from test biological samples. Supervised learning is a machine learning task that learns a function that maps inputs to outputs based on example input-output pairs. It infers a function from labeled training data (here, biomarker expression is the label associated with the training spectral data) consisting of a set of training examples (training spectra). In supervised learning, each example is a pair consisting of an input object (typically a vector) and a desired output value (also called a supervision signal). The supervised learning algorithm analyzes the training data and generates an inference function that can be used to map new examples. The optimal scenario is that the algorithm can correctly determine the class label of instances that have never been seen before.

[0155] The biomarker expression estimation engine 210 may include any type of machine learning algorithm known to those skilled in the art. Suitable machine learning algorithms include regression algorithms, similarity-based algorithms, feature selection algorithms, regularization-based algorithms, decision tree algorithms, Bayesian models, kernel-based algorithms (e.g., support vector machines), clustering-based methods, artificial neural networks, deep learning networks, ensemble methods, and dimensionality reduction methods. Examples of suitable dimensionality reduction methods include principal component analysis (e.g., principal component analysis plus discriminant analysis) and projection to latent structure regression.

[0156] In some embodiments, the biomarker expression estimation engine 210 uses principal component analysis (PCA). The main idea behind PCA is to reduce the dimensionality of a dataset consisting of many interrelated variables while preserving as much variation as possible within the dataset. This is accomplished by transforming the variables into a new set of variables called principal components (or simply PCs), and by orthogonal ordering such that the preservation of variation present in the original variables decreases as they are moved down the hierarchy. In this way, the first principal component preserves the greatest variation present in the original components. Principal components are eigenvectors of the covariance matrix, and therefore they are orthogonal. Principal component analysis and its methods of use are described in U.S. Patent Publication No. 2005 / 0123202 and U.S. Patents 6,894,639 and 8,565,488 (the entire contents of which are incorporated herein by reference). Khan et al. further described PCA and linear discriminant analysis in “Principal Component Analysis - Linear Discriminant Analysis Feature Extractor for Pattern Recognition”, “IJCSI International Journal of Computer Sciences Issues”, Volume 8, Issue 6, No. 2, November 2011, the full text of which is incorporated herein by reference.

[0157] In some embodiments, the biomarker expression estimation engine 210 utilizes Projection to Latent Structure Regression (PLSR). PLSR is a technique that combines and generalizes features of PCA and multiple linear regression. Its goal is to predict a set of dependent variables from a set of independent or predictor variables. This prediction is achieved by extracting a set of orthogonal factors called latent variables from the predictor variables, which have the best predictive power. These latent variables can be used to create a display similar to that of PCA. The quality of predictions obtained from the PLS regression model is evaluated using cross-validation techniques such as bootstrapping and knife-cutting. There are two main variants of PLS ​​regression: the most common one separates the roles of the dependent and independent variables; the second—giving the dependent and independent variables equal roles. Abdi further describes PLSR in “Partial Least Squares Regression and Projection on Latent Structure Regression (PLS Regression),” “WIREs Computational Statistics,” John Wiley & Sons, 2010, the full text of which is incorporated herein by reference. The examples provided in this paper describe a PLSR-based trained biomarker expression estimation engine and demonstrate that the PLSR-based trained biomarker expression estimation engine 210 can at least be used to provide quantitative estimates of biomarker expression levels.

[0158] In some embodiments, the biomarker expression estimation engine 210 utilizes T-distributed random neighborhood embedding (t-SNE). T-SNE is a nonlinear dimensionality reduction technique well-suited for embedding high-dimensional data for visualization in a low-dimensional space of two or three dimensions. Specifically, it models each high-dimensional object using a two- or three-dimensional point, such that similar objects are modeled using nearby points, while dissimilar objects are modeled with a high probability using distant points.

[0159] The t-SNE algorithm comprises two main stages. First, t-SNE constructs a probability distribution over high-dimensional object pairs, such that similar objects are highly likely to be selected, while dissimilar points are extremely unlikely to be selected. Second, t-SNE defines a similar probability distribution over points in the low-dimensional map, and t-SNE minimizes the Kullback-Leibler divergence between the two distributions relative to the positions of points in the map. Note that while the original algorithm uses Euclidean distance between objects as the basis for its similarity measure, this should be modified accordingly. T-SNE is further described in PCT Publication No. WO / 2019 / 084697 and US Patent Publications Nos. 2018 / 0356949 and 2018 / 0340890 (the entire contents of which are incorporated herein by reference).

[0160] In some embodiments, the biomarker expression estimation engine 210 uses reinforcement learning. Reinforcement learning (RL) is a machine learning method in which an agent receives a delayed reward at the next time step to evaluate its previous actions. In other words, RL is a model-free machine learning paradigm that focuses on some concepts of how a software agent should act in an environment to maximize cumulative rewards. Typically, an RL setup consists of two components: an agent and an environment. The environment refers to the object the agent is acting on, while the agent represents the RL algorithm. The environment first sends a state to the agent, and then the agent acts in response to that state based on its knowledge. Afterward, the environment sends back a pair of next states and rewards to the agent. The agent updates its knowledge using the reward returned by the environment to evaluate its last action. The loop continues until the environment sends a terminating state, which ends the episode. Reinforcement learning algorithms are further described in U.S. Patent Nos. 10,279,474 and 7,395,252 (the disclosures of which are incorporated herein by reference in their entirety).

[0161] For example, in some embodiments, the machine learning algorithm is a Support Vector Machine (“SVM”). Typically, an SVM is a classification technique based on statistical learning theory, where a non-linear input dataset is transformed into a high-dimensional linear feature space using a kernel for the non-linear case. The SVM projects a set of training data E representing two distinct classes into a high-dimensional space using a kernel function K. In this transformed data space, the non-linear data is transformed to form a flattened hyperplane to separate the classes, maximizing the separation. The test data is then projected into the high-dimensional space through K, and the test data (e.g., features or metrics listed below) is classified based on its position relative to the hyperplane. The kernel function K defines the method for projecting data into a high-dimensional space.

[0162] In some embodiments, the biomarker expression estimation engine 210 includes a neural network. In some embodiments, the neural network is configured as a deep learning network. Generally speaking, "deep learning" is a branch of machine learning, based on a set of algorithms that attempt to model high-level abstractions in data. Deep learning is part of a broader family of machine learning methods based on learning representations of data. Observations can be represented in various ways, such as an intensity value vector for each pixel, or in a more abstract way as a set of edges, regions of a specific shape, etc. Some representations are superior to others in simplifying learning tasks. One of the prospects of deep learning is to replace handcrafted features with efficient algorithms to achieve unsupervised or semi-supervised feature learning and hierarchical feature extraction.

[0163] In some embodiments, the neural network is a generative network. A “generative” network can generally be defined as a model that is inherently probabilistic. In other words, a “generative” network is not a network that performs forward simulations or rule-based methods. Instead, a generative network can be learned based on a suitable training dataset (e.g., multiple training spectral datasets) because its parameters can be learned. In some embodiments, the neural network is configured as a deep generative network. For example, the network can be configured to have a deep learning architecture, as the network can include multiple layers performing numerous algorithms or transformations.

[0164] In some embodiments, the neural network includes an autoencoder. An autoencoder neural network is an unsupervised learning algorithm that applies backpropagation, setting the target value to be equal to the input value. The purpose of an autoencoder is to learn a representation (encoding) of a set of data by training the network to ignore signal “noise,” often used for dimensionality reduction. Along with the simplification aspect, a reconstruction aspect is learned, where the autoencoder attempts to generate a representation from the simplified encoding that is as close as possible to its original input. Additional information about autoencoders can be found at http: / / ufldl.stanford.edu / tutorial / unsupervised / Autoencoders / , the contents of which are incorporated herein by reference in their entirety.

[0165] In some embodiments, the neural network may be a deep neural network with a set of weights that model the world based on data that has been fed back to train the world. Neural networks typically consist of multiple layers, with signal paths traversing from front to back between the layers. Any neural network can be implemented for this purpose. Suitable neural networks include LeNet, AlexNet, ZFnet, GoogLeNet, VGGNet, VGG16, DenseNet, and ResNet. In some embodiments, fully convolutional neural networks are utilized, such as those described by Long et al., "Fully Convolutional Networks for Semantic Segmentation," Computer Vision and Pattern Recognition (CVPR), 2015 IEEE Conference, June 2015 (INSPEC Registry No.: 15524435), the disclosure of which is incorporated herein by reference.

[0166] In some embodiments, the neural network is configured as AlexNet. For example, a classification network architecture can be AlexNet. The term "classification network" is used herein to refer to a CNN, which includes one or more fully connected layers. Typically, AlexNet contains multiple convolutional layers (e.g., 5), followed by multiple fully connected layers (e.g., 3) configured and trained in combination for classifying data.

[0167] In other embodiments, the neural network is configured as GoogleNet. While the GoogleNet architecture may contain a relatively large number of layers (especially compared to some other neural networks described herein), some of these layers may run in parallel, and groups of layers running in parallel are often referred to as initial modules. Other layers may operate sequentially. Thus, GoogleNet differs from other neural networks described herein in that not all layers are arranged in a sequential structure. An example of a neural network configured as GoogleNet is described in “Going Deeper with Convolutions,” Szegedy et al., CVPR 2015, which is incorporated herein by reference as if fully described herein.

[0168] In other embodiments, the neural network is configured as a VGG network. For example, a classification network structure can be VGG. A VGG network is created by increasing the number of convolutional layers while fixing other parameters of the architecture. Convolutional layers can be added to increase depth by using essentially small convolutional filters in all layers.

[0169] In other embodiments, the neural network is configured as a deep residual network. For example, a classification network architecture can be a deep residual network or ResNet. Like some other networks described herein, a deep residual network can contain convolutional layers followed by fully connected layers, which are configured and trained in a combined manner for detection and / or classification. In a deep residual network, each layer is configured to learn residual functions as reference layer inputs, rather than learning unreferenced functions. In particular, instead of expecting each of the few stacked layers to directly fit the desired base map, it is explicitly allowed that these layers fit the residual map, which is achieved through a feedforward neural network with shortcut connections. Shortcut connections are connections that skip one or more layers.

[0170] Deep residual networks can be created by employing a standard neural network architecture containing convolutional layers and inserting shortcut connections, thereby taking a standard neural network and converting it into a residual learning copy. An example of a deep residual network described in “Deep Residual Learning for Image Recognition” by He et al., NIPS 2015, is incorporated herein by reference as fully illustrated herein. The neural networks described herein can be further configured as described in that reference.

[0171] Training a biomarker expression estimation engine

[0172] In some embodiments, the biomarker expression estimation engine 210 is adapted to operate in a training mode. In some embodiments, the training module 211 may be operable to provide training spectral data to the biomarker expression estimation engine 210 and operate the biomarker expression estimation engine 210 in its training mode according to any suitable training algorithm. In some embodiments, the training module 211 communicates with the biomarker expression estimation engine 210 and is configured to receive training spectral data (or further processed variations of training absorption spectral data, such as the first or second derivative of the training spectral data, the amplitude of a single band within the training spectral data, the integral of a band within the training spectral data, the ratio of the intensities of two or more bands within the training spectral data, the ratio of the second and third derivatives of the training spectral data, etc.) and provide the training spectral data to the biomarker expression estimation engine 210.

[0173] In some embodiments, the training module 211 is also adapted to provide category labels associated with the training spectral data, including actual biomarker expression values ​​(e.g., positive percentage, staining intensity). In some embodiments, the category labels associated with the training spectral data may include actual biomarker expression values ​​(such as those determined by a trained pathologist or calculated using one or more image analysis algorithms) and information related to sample preparation prior to staining (e.g., fixed state, demasked state).

[0174] In some embodiments, the training algorithm utilizes a set of known training vibrational spectral data (as described herein) and a set of corresponding known output class labels (e.g., biomarker expression levels, etc.), and is configured to modify the internal connections within the biomarker expression estimation engine 210 such that the processing of the input training spectral data provides the required corresponding class labels.

[0175] The biomarker expression estimation engine 210 can be trained using any method known to those skilled in the art. For example, any training method is disclosed in U.S. Patent Publications 2018 / 0268255, 2019 / 0102675, 2015 / 0356461, 2016 / 0132786, 2018 / 0240010, and 2019 / 0108344 (the entire contents of which are incorporated herein by reference).

[0176] In some embodiments, a cross-validation method is used to train the biomarker expression estimation engine 210. Cross-validation is a technique that can be used to help with model selection and / or parameter tuning when developing a classifier. Cross-validation uses one or more subsets of cases from a labeled case set as the test set. For example, in k-fold cross-validation, the labeled case set is divided equally into k “folds.” K-fold cross-validation is, for example, a resampling procedure used to evaluate machine learning models. A series of training-then-test loops are performed, iterating through the k folds so that in each loop a different fold is used as the test set, while the remaining folds are used as the training set. Since each fold is used as the test set at some point, non-randomly selected cases in the labeled case set may bias the cross-validation. For example, in the scenario of 5-fold cross-validation (k=5), the dataset is split into 5 folds. In the first iteration, the first fold is used to test the model, and the rest are used to train the model. In the second iteration, the second fold is used as the test set, and the rest are used as the training set. This process is repeated until every one of the 5 folds is used as the test set. U.S. Patent Publications 2014 / 0279734 and 2005 / 0234753 (the entire contents of which are incorporated herein by reference) further describe methods for performing k-fold cross-validation.

[0177] Figure 13 illustrates, within the context of a biomarker expression estimation engine 210 utilizing a PLSR-based machine learning algorithm, that a PLSR model is trained to mine vibrational spectra of biomarker expression features within a training spectrum. In some embodiments, the PLSR model is also trained to identify variations in these features across different tissue types and / or different types of molecules (proteins, nucleic acids). In some embodiments, the PLSR algorithm employs vibrational spectral data (e.g., absorption spectra, first derivatives, second derivatives) and creates a model to determine which features (wavelengths) best predict the response variable (biomarker expression, etc.). In some embodiments, the performance of the generated model can be further evaluated using the same and unknown vibrational spectral data used for performance evaluation and optimization.

[0178] In the context of a biomarker expression estimation engine 210 utilizing a principal component analysis-based machine learning algorithm, PCA is performed on an initial training dataset with a default sample size to generate a PCA transformation matrix. A second PCA is performed on a combined dataset including the initial training dataset and the test dataset. The number of samples in the initial training dataset is then increased to generate an expanded training dataset. PCA is performed on the expanded training dataset to determine if the number of PCA operations on the expanded training dataset is the same as that on the initial training dataset. If so, the error between the initial test dataset and the expanded test dataset is evaluated based on the PCA signals and the PCA transformation matrix to estimate the final solution error. The PCA matrix of the combined dataset is transformed back to the domain of the initial training dataset (e.g., the spectral domain) using the transformation matrix from the first PCA to generate a test dataset estimate. This method iteratively expands the size of the training matrix until the number of PCA operations converges and a final error target is reached. After the error target is reached, the training dataset of the identified size sufficiently represents the training objective function information contained within a specified range of input parameters. The machine learning system (e.g., biomarker expression estimation engine 210) can then be trained using the training matrix of the identified size. Further aspects of training using PCA are disclosed in U.S. Patent Nos. 8,452,718 and 7,734,087 (the entire disclosures of which are incorporated herein by reference).

[0179] In embodiments where the biomarker expression estimation engine 210 includes a neural network, a backpropagation algorithm can be used to train the biomarker expression estimation engine 210. Backpropagation is an iterative process in which the connections between network nodes are given some random initial values, and the network is operated to compute a corresponding output vector for a set of input vectors (training spectral dataset). The output vector is compared with the expected output of the training spectral dataset, and the error between the expected output and the actual output is calculated. The calculated error is propagated back from the output node to the input node to modify the values ​​of the network connection weights to reduce the error. After each such iteration, the training module 211 can compute the total error for the entire training set, and then the training module 211 can repeat the process with another iteration. When the total error reaches a minimum, the training of the biomarker expression estimation engine 210 is complete. If the total error does not reach a minimum after a predetermined number of iterations and if the total error is not constant, the training module 211 can consider the training process to have not converged.

[0180] In the context of training using acquired spectral data derived from multiple differentially prepared, stained training tissue samples as described above, each acquired training spectrum is associated with a known expression level of one or more biomarkers (where the known expression level of one or more biomarkers is used as a category label, as described herein). In some embodiments, and again in the context of training using acquired spectral data derived from multiple differentially prepared, stained training tissue samples, each acquired training spectrum may be associated with (i) a known expression level of one or more biomarkers, and (ii) known sample preparation conditions and / or sample preparation states (e.g., fixed duration, fixed quality, demasking conditions, demasking state). For example, the two training spectral datasets shown in Figure 4B (see the dashed boxes listing groups 1 and 2) may be provided to training module 211 for training biomarker expression estimation engine 210, along with the known expression levels of one or more biomarkers, and any additional category labels.

[0181] When the biomarker expression estimation engine 210 has completed training, the system 200 is ready to run to detect biomarker expression features within test spectral data and, based on the detected biomarker expression features, estimate the expression levels of one or more biomarkers in unstained test biological samples. In some embodiments, the biomarker expression estimation engine 210 can be periodically retrained to adapt to changes in the input data.

[0182] Estimation of biomarker expression

[0183] Once the biomarker expression estimation engine 210 is properly trained, as described above, it can be used to detect biomarker expression features within test vibrational spectral data, such as test spectral data collected from unstained test biological samples, and predict the expression of one or more biomarkers in the unstained test biological samples based on the detected biomarker expression features. In some embodiments, and referring to FIG3, an unstained test biological sample is obtained (step 310) (e.g., from a subject suspected of having a disease or known to have a disease), and then test vibrational spectral data is collected from the unstained test biological sample (step 320) (also see FIG7). In some embodiments, the test vibrational spectral data includes absorption spectra, first and / or second derivatives of absorption spectra, amplitudes of individual bands within training spectral data, integrals of bands within training spectral data, ratios of intensities of two or more bands in training spectral data, ratios of second and third derivatives of training spectral data, etc.

[0184] Once the aforementioned test spectral data and / or its variants have been acquired and processed, biomarker expression features can be derived from the test spectral data using the trained biomarker expression estimation engine 210 (step 340). In some embodiments, the derived biomarker expression features include a mapping of the correlation between each wavenumber and the predicted repair status. Values ​​close to zero are largely meaningless. In some embodiments, detectable biomarker expression features include peak amplitude, peak location, peak ratio, sum of spectral values ​​(e.g., integral over a spectral range), one or more variations in slope (first derivative) or curvature (second derivative), etc. Based on the derived biomarker expression features, an estimate of the expression of one or more biomarkers can be calculated (step 350). In some embodiments, the estimated expression of one or more biomarkers includes a quantitative estimate of the staining intensity of one or more biomarkers and / or a quantitative estimate of the positive percentage of one or more biomarkers, thereby enabling a “label-free” scoring of the expression of one or more biomarkers.

[0185] Figures 23A, 24A, and 25A respectively show BCL2 (Figure 23A), FOXP3 (Figure 24A), and Ki-67 (Figure 25A). Figure 25A A comparison of measured (experimental) staining intensity levels with predicted staining intensity levels in BLC2, FOXP3, and Ki-67 positive cells is shown. In each example, a separate model was trained that was able to predict the staining intensity of each of the three biomarkers using MID-IR spectroscopy (see Example 4). In this example, first-derivative spectroscopy was used, and the spectrum was set at 1750–2800 cm⁻¹. -1 and 3700 - 4000 cm-1 The two regions are set to zero, even though different numbers of components are needed in each model to achieve ideal performance.

[0186] As can be seen from the data in Figures 23A, 24A, and 25A, despite significant variations in expression intensity over fixed time periods, the method of this disclosure is able to predict the biomarker intensity of all three proteins. Figures 23A, 24A, and 25A each illustrate that a biomarker expression estimation engine 210 trained with data relating to the expression levels of one or more biomarkers over various fixed durations (e.g., staining intensity levels, such as the staining intensity of DAB) can be used to quantitatively predict the expression levels of one or more biomarkers with high accuracy. Figures 23B, 24B, and 25B list the cumulative distribution function (CDF) of the estimated and predicted DAB staining for each of the aforementioned biomarkers.

[0187] Figures 26A, 27A, and 28A respectively show FOXP3 (Figure 27A), BCL2 (Figure 27A), and Ki-67 (Figure 28A). Figure 28A Figures 26A, 27A, and 28A show a comparison of the measured (experimental) expression levels of positive cells with the predicted expression levels (positive percentage) of FOXP3, BLC2, and Ki-67 positive cells. Figures 26A, 27A, and 28A each illustrate how a biomarker expression estimation engine 210, trained with data relating to the expression levels of one or more biomarkers at various fixed durations, can be used to quantitatively predict the expression levels of one or more biomarkers with high accuracy. Figures 26B, 27B, and 28B list the cumulative distribution function (CDF) of the estimated and predicted tissue positive percentages for each of the aforementioned biomarkers.

[0188] Figures 15 and 16 illustrate the results obtained using a trained biomarker expression estimation engine 210 to determine the expression of two different biomarkers in tissue samples with an unknown fixed duration. Figures 15 and 16 comparatively show the predicted positive percentage for two different biomarkers (cd4 and life-67) against a fixed unknown duration of difference demasking in test biological samples, using the systems and methods described herein, compared to known (e.g., experimentally derived values, such as those derived after tissue staining and analysis with detection and classification algorithms) positive percentage values. At least as shown in the figures above, the biomarker expression estimation engine 210 is able to accurately predict biomarker expression information across difference-demasked samples (as well as samples with an unknown fixed state).

[0189] Figure 18 further illustrates the predictive power of the system and method of this disclosure. In fact, Figure 18 shows the predictive accuracy of the trained biomarker expression estimation engine at all times and temperatures in a tonsil blind sample of unknown fixed duration. At all test times and temperatures, the trained biomarker expression estimation engine predicted functional C4d staining intensity with an accuracy greater than approximately 10%. The values ​​at the time-temperature intersections represent the percentage error between the predicted and actual C4d staining intensities.

[0190] In this example, three independent PLSR prediction engines were trained. In the first model, tissue was repaired at different temperatures (98.6°C, 110°C, 120°C, 130°C, and 140°C) for durations of approximately 5 minutes each. Several tissues were considered as the training set, meaning they were imaged using a MID-IR microscope, and the PLSR model was trained on this dataset. Blind tissue was then imaged using a MID-IR microscope, and the trained biomarker expression estimation engine was used to calculate the expected C4d staining intensity for that tissue. This was calculated based on digital analysis brightfield DAB images, with the error percentage calculated in a standard manner, comparing the model's predictions to the average staining intensity as 100 * (MID-IR predicted staining - brightfield actual staining) / brightfield actual staining.

[0191] The process was then repeated at the same antigen retrieval temperature, but with retrieval durations of 30 minutes and 60 minutes. Thus, three independent engines were trained and validated in this example. Given the foregoing, in some embodiments, the data can be used to train an overall predictive model capable of determining biomarker staining solely based on the MID-IR spectra acquired from the sample, regardless of the sample's retrieval time and temperature.

[0192] In embodiments where the biomarker expression estimation engine 210 is trained with category labels including biomarker expression levels and sample preparation states (e.g., fixed state and / or unmasked state), the trained biomarker expression estimation engine 210 may further provide, as output, a predicted difference between: (i) the expression levels of one or more biomarkers in a test sample based on the test sample preparation state (e.g., fixed duration), and (ii) the expected expression levels of one or more biomarkers in the same test sample prepared under different conditions (e.g., fixed samples for different time periods). It is believed that this may be useful in instances where test biological samples are not fixed for a sufficiently long time and / or are not properly unmasked, thus fixed duration and / or unmasked state of the target biomarker may be considered “insufficient.” In some embodiments, the predicted difference may be used such that the expression levels of one or more biomarkers increase or decrease based on fixed duration and / or unmasked state, and the increase or decrease in fixed level or unmasked state may be used for downstream scoring.

[0193] Referring to Figure 30A, in some embodiments, the system further includes operations for correcting predicted expression of one or more biomarkers to test for poor demasking and / or poor fixation of biological samples. For example, a biomarker fixation sensitivity curve can be obtained (step 910). Figure 9D shows an example of a suitable biomarker fixation sensitivity curve. There, the graph shows the normalized percentage of positive results for three different biomarkers with fixation time, and more specifically, where the average expression is plotted on a normalized scale so that the relative change of each biomarker with fixation time can be observed, as shown in this example, as a biomarker fixation sensitivity curve used to correct the obtained predicted biomarker expression level.

[0194] Next, a fixed time is obtained for the test biological sample (step 911). Subsequently, the trained biomarker expression estimation engine of this disclosure is used to obtain the predicted biomarker expression level of the test biological sample (912). In some embodiments, the test biological sample is an unstained test biological sample. In step 913, the obtained predicted biomarker expression level of the test biological sample is corrected using the obtained fixed sensitivity curve to provide a fixed compensated expression level. Figure 30B illustrates an alternative method in which the actual biomarker expression level is measured (step 914), and then compensated using the obtained fixed sensitivity curve (step 915).

[0195] In some embodiments, the system disclosed herein may include one or more scoring modules that enable the estimation of one or more expression scores (H-scores, etc.) based on predicted biomarker expression data received as output. Any of the scoring methods disclosed in U.S. Patent Publication No. 2015 / 0347702 (the entire contents of which are incorporated herein by reference) can be used to determine biomarker expression scores, wherein the biomarker expression value is estimated using the trained biomarker expression estimation engine 210 described herein.

[0196] In some embodiments, the information provided as output can be used for further downstream processes and can be used to make decisions about whether a test biological sample should be treated with one or more specific binding entities.

[0197] Example 1

[0198] The expression and fixation times of three different biomarkers (BCL2, FOXP3, and ki67) are presented here. Tissue blocks were stained for each biomarker at each fixation time, and the expression across the entire slide was quantified using image analysis algorithms (e.g., an algorithm suitable for quantitatively determining the expression level of each stain, such as an automated algorithm that first segments the tissue onto a slide, then identifies non-target tissue regions; the algorithm then automatically determines whether the given protein biomarker in the tissue is positive or negative). Figures 9A, 9B, and 9C present the summary results for BCL2, ki-67, and FOXP3 in box-and-whisker plots and fixation times, respectively. It was found that BCL2 and FOXP3 were particularly unstable and susceptible to inappropriate fixation, with their expression levels increasing monotonically and steadily with fixation time.

[0199] On the other hand, it was found that Ki-67 is relatively robust to inappropriate fixation as long as the biological samples are fixed in NBF for at least 1 hour. Finally, these three figures are summarized in Figure 9D, which shows a comparison of the average expression level of each of the three biomarkers with the maximum expression at 24 hours after the fixation time is normalized to the scale.

[0200] Turning to Figures 20, 21, and 22, a digital analysis of biomarker expression levels in stained tissues / cells was performed, and the relative concentration of each biomarker was quantified, as shown below. The results indicate that tissues with longer fixation times tend to stain more intensely / deeper. The relationship between box-and-whisker plots and fixation time is shown again. Similar to those mentioned above, BCL2 and FOXP3 were found to be particularly unstable and susceptible to inappropriate fixation, with their expression levels increasing monotonically and steadily with fixation time. On the other hand, Ki-67 was found to be relatively robust to inappropriate fixation.

[0201] Example 2

[0202] MirrIR microscope slides (Kevley Technologies, Chesterland, OH) used for reflectance infrared studies were employed for mid-IR spectroscopy measurements. Four-micron serial sections of formalin-fixed paraffin-embedded (FFPE) tonsil tissue were placed on pretreated MirrIR slides. The tonsil tissue was manually dewaxed according to OP2100-025. Briefly, following the xylene step, the slides were hydrated with a decreasing gradient of ethanol and then transferred to a Rapid Antigen Retrieval (RAR) stage in VENTANA Cell Regulation 1 (CC1) solution.

[0203] Antigen retrieval was performed in CC1 solution in a RAR chamber, pre-pressurized to 30 psi before turning on the heater. The total heating time for any given experiment included a 90-second rise time and a 2-minute cool-down time. After the antigen retrieval step, the slides were gently washed in deionized water and air-dried at room temperature. Dry slides with intact tonsil tissue were used for mid-IR measurements. A description of individual antigen retrieval experiments can be found in LN #3685 (Bohuslav Dvorak), pages 52–59 and 64–69.

[0204] Collect all sample and processed immunoreactivity data analyzed using mid-IR spectroscopy. Briefly, samples were processed using a hybrid procedure involving manual dewaxing and antigen retrieval. Dewaxing was performed using xylene, followed by rehydration via a series of gradient ethanol steps according to OP2100-025. Samples were then placed in CC1 (catalog number: 950-124). Antigen-retrieval samples were transferred to a BenchMark UTLRA instrument in reaction buffer (catalog number: 950-300) for subsequent processing steps from peroxide inhibitor to counterstaining.

[0205] For the study described here, tonsil samples were labeled with antiserum generated against Ki-67 (30-9) or C4d (SP91). These markers were chosen because they exhibited different responses to increased antigen retrieval treatment. It was found that Ki-67 staining intensity and the number of labeled cells increased to a certain extent, after which the intensity and positivity decreased.

[0206] Conversely, it was found that C4d continued to increase intensity and positivity under antigen retrieval conditions, otherwise it would damage the sample. In addition, C4d was chosen because it performed poorly when treated with current retrieval methods, but showed significant immunoreactivity when treated with high-temperature antigen retrieval (this property is described in detail in Appendix D081973 entitled "Improvement of Rapid Antigen Retrieval Staining Quality").

[0207] Example 3 – Estimating biomarker expression using a trained biomarker expression estimation engine

[0208] Instruction Manual Summary

[0209] This experiment utilizes mid-infrared (mid-IR) spectroscopy to examine the vibrational state of molecules in histological tissue sections. In this work, changes in mid-IR spectra induced by differentially repaired tonsillar tissue were investigated, and these changes were used to train a biomarker expression estimation engine. Identified transitions in the mid-IR spectra correlated with immunohistochemical (IHC) staining for Ki-67 and C4d proteins.

[0210] introduce

[0211] Mid-infrared spectroscopy (mid-IR) is a powerful optical technique capable of detecting the vibrational states of individual molecules in tissues and is highly sensitive to the conformational state of proteins. This extremely high sensitivity makes mid-IR spectroscopy ideal for microscopy applications, as the presence of endogenous and exogenous materials, and even their conformational states, can be revealed through changes in the mid-IR absorption curve of biological samples. Vibrational spectroscopy has even been used in diagnostic applications, such as distinguishing between healthy and cancerous tissues.

[0212] Methods and Materials

[0213] Repair

[0214] MirrIR microscope slides (Kevley Technologies, Chesterland, OH) used for reflectance infrared studies were employed for mid-IR spectroscopy measurements. Four-micron serial sections of formalin-fixed paraffin-embedded (FFPE) tonsil tissue were placed on pretreated MirrIR slides. The tonsil tissue was manually dewaxed according to OP2100-025. Briefly, following the xylene step, the slides were hydrated with a decreasing gradient of ethanol and then transferred to a Rapid Antigen Retrieval (RAR) stage in VENTANA Cell Regulation 1 (CC1) solution.

[0215] The antigen retrieval step is performed in a CC1 solution in a RAR chamber, pre-pressurized to 30 psi before turning on the heater. The total heating time for any given experiment includes a 90-second rise time and a 2-minute cooling time. After the antigen retrieval step, the slides are gently washed in deionized water and air-dried at room temperature. Dry slides with intact tonsil tissue are used for mid-IR measurements. A description of individual antigen retrieval experiments can be found in LN #3685 (Bohuslav Dvorak), pages 52-59 and 64-69.

[0216] IHC staining and quantification

[0217] Collect all samples analyzed using mid-IR spectroscopy and processed immunoreactivity data. These samples were generated using the methods described in detail in "D081973 Rapid Antigen Retrieval Product and Process Feasibility Report". In short, samples were processed using a hybrid procedure in which dewaxing and antigen retrieval were performed manually. Dewaxing was performed using xylene, followed by rehydration via a series of gradient ethanol according to OP2100-025. The samples were then placed in CC1 (catalog number: 950-124).

[0218] In this report, antigen retrieval was performed using a RAR benchtop (part number: 101430300) at the times and temperature settings described. Antigen retrieval samples were transferred to a BenchMark UTLRA instrument in reaction buffer (catalog number: 950-300) for subsequent processing steps from peroxide inhibitor to counterstaining.

[0219] Sample slides were scanned using a Leica Aperio AT2 (Leica Biosystems, Nussloch, Germany) slide scanner, and the intensity of immunoreactivity and the proportion of stained tissue were quantified using the "Positive Pixel Count v9" algorithm provided by the Aperio Imagescope software. For each tissue, a target region of interest (ROI) was selected to include tonsil tissue that was expected to be stained. Figure 10 shows that connective tissue that showed high background in some staining treatments but was absent in others was excluded.

[0220] This quantitative method produces intensity units that are repeatable across samples and can be compared within experiments. However, no attempt has been made to plot or harmonize intensity measurements, or the percentage of positive pixels reported to pathologists for scoring.

[0221] Collection of Mid-IR data

[0222] Mid-IR spectra were collected on a Fourier transform infrared (FTIR) microscope (Bruker Hyperion 3000, Bruker Optics, Billerica MA) with an attached optical interferometer (Vertex 70). Serial sections of tonsil blocks were cut into 4-micrometer-thick slices on mid-IR reflective slides (Kevley Technologies, MirrIR), differentially repaired, and imaged using a mid-IR microscope.

[0223] Tonsil tissue sections repaired under different experimental conditions were placed on an FTIR microscope, and the entire tissue section was imaged with a visible objective using a grating scanning field of view (FOV). Bruker software OPUS was used to randomly select tissue regions, from which mid-IR spectra were collected using a mercury-cadmium-tellurium (MCT) detector. Typically, 20–80 spectra were collected from each tissue sample. Absorption spectra were collected at a resolution of 4 cm⁻¹, and each selected ROI was sampled 64 times. These spectra were then averaged together to produce the final spectrum for a given location. Example tissue images, sampling patterns, and the resulting average spectra for a single ROI are shown in Figure 11 below. All spectra were collected using a 15X IR objective, producing an FOV of approximately 200 μm x 200 μm.

[0224] Preprocessing mid-IR data

[0225] The collected spectra were preprocessed to remove artifacts, normalize the spectral format, and separate the tissue's mid-IR absorption. Mid-IR transmittance was measured directly under a microscope. To convert the transmission spectra to absorption spectra, reference transmission spectra were collected at a spatial location outside the sample and used to delimit the spectra collected through the tissue. This calculation provides the amount of light attenuated by the tissue (absorption + scattering). Next, atmospheric absorption (primarily from water vapor and carbon dioxide) was removed using an algorithm in OPUS software. Then, baseline correction was used to correct for tissue scattering corrected with a concave rubber band (8 iterations, 64 baseline points). The resulting spectra represent the absorption of the sample tissue. Finally, all spectra were normalized to a maximum of 2 to mitigate differences in section thickness and tissue density.

[0226] Experimental Design and Results

[0227] Changes in antigen retrieval time at constant antigen retrieval temperature

[0228] In this experiment, tonsillar tissue was subjected to antigen retrieval at 98.6 °C for 0, 10, 30, 60, or 120 minutes. Each treatment was performed on replicate samples. Mid-IR spectroscopy revealed a significant shift in the major protein band (referred to as the amide I band), which was associated with a loose antigen retrieval process. Figure 12A shows an example of this amide I band shift. Quantification of the peak wavelength and full width at half maximum (FWHM) of the amide I band was able to distinguish between unretrieved and partially, completely, and over-retrieved antigen retrieval processes (Figure 12B).

[0229] Several other metrics were evaluated throughout the project, including principal component analysis, integration of the amide I band, normalization of several bands to correct for scattering, and quantification of methyl and methylene peaks. Unfortunately, none of these other metrics improved the stratification of tissue antigen repair status. Finally, a supervised machine learning model was developed to utilize non-obvious features in mid-IR spectra indicating the expression levels of one or more biomarkers.

[0230] These subtle differences in the spectrum are identified using the Projection to Latent Structure Regression (PLSR) method. The algorithm takes mid-IR signals (e.g., absorption spectra, first derivatives, second derivatives) and creates a model to determine which features (wavelengths) best predict response variables (antigen repair state, target repair state, etc.). The performance of the generated model is then evaluated using both identical and unknown mid-IR data for performance assessment and optimization. Figure 13 illustrates how the PLSR model was trained to mine mid-IR spectra of antigen repair features. In this experiment, the model achieved an accuracy of 3 minutes.

[0231] These studies demonstrate that supervised machine learning models can mine data and develop models that can be used to determine the expression levels of biomarkers in tonsil samples. To further validate that the model can identify true biomarker expression characteristics, a series of spectra not trained on the algorithm were provided to determine its ability to make blind predictions. Furthermore, it has been shown that for samples with unknown time intervals, the PLSR model can correlate differences in mid-IR spectra with the IHC staining intensity of Ki-67 and C4d proteins (see Figures 15 and 16).

[0232] Changes in antigen retrieval time and temperature

[0233] In this study, mid-IR spectroscopy was combined with a machine learning model to determine whether it could be used to estimate the expression of one or more biomarkers (e.g., positive percentage; staining intensity) in samples with unknown fixed-time conditions and variable demasking conditions. Five multi-tissue slides with four independent tonsillar tissues were repaired for 5 minutes at temperatures between 98.6 °C and 140 °C.

[0234] Mid-IR spectra from three tonsillar tissues (Fig. 17, circled portion including all three tissue samples) were used to train the PLSR model. This model was then used to infer antigen retrieval conditions in “unknown” tonsillar tissues (Fig. 17, circled portion including only a single tissue sample). The results in Fig. 11 demonstrate that, at least in tonsillar tissues, the combination of mid-IR spectroscopy and PLSR can accurately quantify the extent of retrieval in unknown samples and the degree of C4d staining in unknown samples across all times and temperatures. This is crucial because time and temperature are the two most important variables affecting antigen retrieval.

[0235] Example 4 – Training a model to predict stained regions or intensity

[0236] Functional staining data can be used to train a PLSR model. In this case, the process of selecting and preparing the input data (spectrum) is similar to training a model to predict at a fixed time. However, the training will be different. In this case, all slides are imaged using a brightfield scanner and fed into the digital pathology algorithm. To obtain meaningful protein expression data, all unstained areas of the tissue (matrix, connective tissue, pores, overlapping tissue / folds) are excluded from the analysis area. Cells identified as positive for the protein are identified, and active tissue areas positive for a given biomarker are digitally quantified. The slides are then characterized by the percentage of positive tissue, which is the percentage of potentially stained tissue that is actually stained. This process is repeated for all tissues. The model can then be trained based on one of the following two processes:

[0237] (a) Average biomarker expression at a given fixed time. Training is performed on all tissues at a given fixed time to produce the average expression of the target protein. Similar to a fixed-time training model, since all tissues at a given fixed time are trained on the same output (fixed time / quality). Advantages and disadvantages: Less noise, model optimized for average performance, and can be trained with less data.

[0238] (b) The model can be trained using the biomarker expression of each tissue individually. For example, if two tissues at the same time point have different biomarker expressions, their spectra can be mined individually to find the spectral features that best explain the differential staining. Benefits: More robust and generalizable models, optimized for individual performance, requiring large training sets.

[0239] An alternative approach to determining functional staining is to quantify the intensity of the biomarker in the currently positively stained cells. This is done by identifying cell / tissue regions that are positive for the biomarker, making DAB expression spectrally unmixable to produce a number proportional to the protein concentration (or alternatively, simply using the raw intensity reading from the detector). This final measurement of intensity can be used to train a model that can predict the staining intensity of a tissue for a given protein. Furthermore, models can be trained based on pathologist readings to predict staining positivity or intensity.

[0240] Examples of biomarkers

[0241] The following identifies non-limiting examples of biomarkers whose expression can be estimated using the systems and methods of this disclosure. Some biomarkers are cell-specific, while others have been identified as being associated with specific diseases or conditions. Examples of known prognostic biomarkers include enzyme biomarkers such as galactosyltransferase II, neuron-specific enolases, proton ATPase-2, and acid phosphatase. Hormone or hormone receptor biomarkers include human chorionic gonadotropin (HCG), adrenocorticotropic hormone (ACTH), carcinoembryonic antigen (CEA), prostate-specific antigen (PSA), estrogen receptor, progesterone receptor, androgen receptor, gC1q-R / p33 complement receptor, IL-2 receptor, p75 neurotrophic receptor, PTH receptor, thyroid hormone receptor, and insulin receptor.

[0242] Lymphocyte markers include α-1-antichymotrypsin, α-1-antitrypsin, B cell markers, bcl-2, bcl-6, B lymphocyte antigen 36kD, BM1 (myeloid marker), BM2 (myeloid marker), galactagogue-3, granzyme B, HLA class I antigen, HLA class II (DP) antigen, HLA class II (DQ) antigen, HLA class II (DR) antigen, human neutrophil defensin, immunoglobulin A, immunoglobulin D, immunoglobulin G, immunoglobulin M, κ light chain, κ light chain, λ light chain, lymphocyte / histocyte antigen, macrophage markers, muramicase (lysozyme), p80 anaplastic lymphoma kinase, plasma cell markers, secretory leukocyte protease inhibitors, T cell antigen receptor (JOVI 1), T cell antigen receptor (JOVI 3), terminal deoxynucleotidyl transferase, and non-clustered B cell markers.

[0243] Tumor markers include alpha-fetoprotein, apolipoprotein D, BAG-1 (RAP46 protein), CA19-9 (sialylated Lewis), CA50 (cancer-associated mucin antigen), CA125 (ovarian cancer antigen), CA242 (tumor-associated mucin antigen), chromogranin A, agglutinin (apolipoprotein J), epithelial membrane antigen, epithelial-associated antigen, epithelial-specific antigen, epidermal growth factor receptor, estrogen receptor (ER), macrocystic lesion fluid protein-15, hepatocyte-specific antigen, HER2, hyalin, human gastric mucin, human milk fat globule, MAGE-1, matrix metalloproteinases, melanin A, melanoma marker (HMB45), mesothelin, metallothionein, microophthalmic transcription factor (MITF), and Muc-1 core glycoprotein. Muc-1 glycoprotein, Muc-2 glycoprotein, Muc-5AC glycoprotein, Muc-6 glycoprotein, myeloperoxidase, Myf-3 (rhabdomyosarcoma marker), Myf-4 (rhabdomyosarcoma marker), MyoD1 (rhabdomyosarcoma marker), myoglobin, nm23 protein, placental alkaline phosphatase, prealbumin, progesterone receptor, prostate-specific antigen, prostate acid phosphatase, prostate inhibitory peptide, PTEN, renal cell carcinoma marker, small intestinal mucus antigen, tetrapeptide, thyroid transcription factor-1, Tissue inhibitor of metalloproteinases 1, Tissue inhibitor of metalloproteinases 2, tyrosinase, tyrosinase-associated protein-1, chorionic villi, von Willebrand factor, CD34, CD34 class II, CD51 Ab-1, CD63, CD69, Chk1, Chk2, C-met (a clasp), COX6C, CREB, cyclin D1, cytokeratin, cytokeratin 8, DAPI, myoderm, DHP (1-6-diphenyl-1,3 ...5-Hextriene), E-cadherin, EEA1, EGFR, EGFRvIII, EMA (epithelial membrane antigen), ER, ERB3, ERCC1, ERK, E-selectin, FAK, fibronectin, FOXP3, γ-H2AX, GB3, GFAP, macroprotein, GM130, Golgi protein 97, GRB2, GRP78BiP, GSK3 β, HER-2, histone 3, histone 3_K14-Ace [anti-acetyl-histone H3 (Lys14)], histone 3_K18-Ace [histone H3-acetyl-Lys 18), histone 3_K27-TriMe, [histone H3 (trimethylK27)], histone 3_K4-diMe [anti-dimethylhistone H3 (Lys 4)], histone 3_K9-Ace [acetyl-histone H3 (Lys 4)] 9)], Histone 3_K9-triMe [Histone 3-trimethylLys 9], Histone 3_S10-Phos [Antiphosphohistone H3 (Ser 10), mitotic marker], Histone 4, Histone H2A.X-5139-Phos [Phosphohistone H2A.X (Ser139) antibody], Histone H2B, Histone H3_dimethylK4, Histone H4_trimethylK20-Chip grad, HSP70, Urokinase, VEGF R1, ICAM-1, IGF-1, IGF-1R, IGF-1 receptor β, IGF-II, IGF-IIR, IKB-α IKKE, IL6, IL8, Integrin αV β3, Integrin αV β6, Integrin αV / CD51, Integrin B5, Integrin B6, Integrin B8, Integrin β1 (CD51) 29) Integrin β3, Integrin β5, Integrin B6, IRS-1, Jagged 1, Antiprotein Kinase C β2, LAMP-1, Light Chain Ab-4 (Cocktail), λ Light Chainκ light chain, M6P, Mach 2, MAPKAPK-2, MEK 1, MEK 1 / 2 (Ps222), MEK 2, MEK1 / 2 (47E6), MEK1 / 2 blocking peptide, MET / HGFR, MGMT, mitochondrial antigen, mitotic tracker green FM, MMP-2, MMP9, E-cadherin, mTOR, ATPase, N-cadherin, nephrotic protein, NFKB, NFKB p105 / p50, NF-KB P65, Notch 1, Notch 2, Notch 3, OxPhos complex IV, p130Cas, p38 MAPK, p44 / 42 MAPK antibody, P504S, P53, P70, P70 S6K, Pan cadherin, peg protein, P-cadherin, PDI, pEGFR, AKT phosphate, CREB phosphate, EGF phosphate Receptor, GSK3 β, H3, HSP-70, MAPKAPK-2, MEK1 / 2, p38 MAP kinase, p44 / 42 MAPK, p53, PKC, S6 ribosomal protein, Src, Akt, Bad, IKB-a, mTOR, NF-κB p65, p38, p44 / 42 MAPK, p70 S6 kinase, Rb, Smad2, PIM1, PIM2, PKC β, podocyte marker protein, PR, PTEN, R1, Rb-4H1, Rb cadherin, ribonucleotide reductase, RRM1, RRM11, SLC7A5, NDRG, HTF9C, CEACAM, p33, S6 Ribosomal proteins, Src, survivin, synaptophysin, multiligand proteoglycan 4, ankle protein, tensin, thymidylate synthase, tuberculin, VCAM-1, VEGF, vimentin, lectin, YES, ZAP-70, and ZEB.

[0244] Cell cycle-related markers include apoptosis protease initiation factor-1, bcl-w, bcl-x, bromodeoxyuridine, CAK (CDK initiation kinase), apoptosis susceptibility protein (CAS), caspase 2, caspase 8, CPP32 (caspase-3), cyclin-dependent protein kinase, cyclin A, cyclin B1, cyclin D1, cyclin D2, cyclin D3, cyclin E, cyclin G, DNA fragmentation factor (N-terminus), Fas (CD95), Fas-associated death domain protein, Fas ligand, Fen-1, IPO-38, Mc1-1, microchromosome retainer, mismatch repair protein (MSH2), poly(ADP-ribose) polymerase, proliferating cell nuclear antigen, p16 protein, p27 protein, p34cdc2, p57 protein (Kip2), and p105. Protein, Stat 1 α, topoisomerase I, topoisomerase II α, topoisomerase III α, topoisomerase II β.

[0245] Neurological tissue and tumor markers include α-B lens protein, α-interconnecting protein, α-synuclein, amyloid precursor protein, β-amyloid protein, calcium-binding protein, choline acetyltransferase, excitatory amino acid transporter 1, GAP43, glial fibrillary acidic protein, glutamate receptor 2, myelin basic protein, nerve growth factor receptor (gp75), neuroblastoma markers, neurofilament 68 kD, neurofilament 160 kD, neurofilament 200 kD, neuron-specific enolase, nicotinic acetylcholine receptor α4, nicotinic acetylcholine receptor β2, peripheral proteins, protein gene product 9, S-100 protein, SNAP-25, synaptophysin I, synaptophysin, τ, tryptophan hydroxylase, tyrosine hydroxylase, and ubiquitin.

[0246] Cluster differentiation markers include CD1a, CD1b, CD1c, CD1d, CD1e, CD2, CD3δ, CD3ε, CD3γ, CD4, CD5, CD6, CD7, CD8α, CD8β, CD9, CD10, CD11a, CD11b, CD11c, CDw12, CD13, CD14, CD15, CD15s, CD16a, CD16b, CDw17, CD18, CD19, CD20, CD21, CD22, CD23, CD24, CD25, CD26, CD27, CD28, CD29, CD30, CD31, CD32, CD33, CD34, CD35, CD36, CD37, CD38, CD39, CD40, CD41, CD42a, CD42b, CD42c, CD42d, CD43, CD44, CD44R, CD45, CD46, CD47, CD48, CD49a, CD49b, CD49c, CD49d, CD49e, CD49f, CD50, CD51, CD52, CD53, CD54, CD55, CD56, CD57, CD58, CD59, CDw60, CD61, CD62E, CD62L, CD62P, CD63, CD64, CD65, CD65s, CD66a, CD66b, CD66c, CD66d, CD66e, CD66f, CD68, CD69, CD70, CD71, CD72, CD73, CD74, CDw75, CDw76, CD77, CD79a, CD79b, CD80, CD81, CD82, CD83, CD84, CD85, CD86, CD87, CD88, CD89, CD90, CD91, CDw92, CDw93, CD94, CD95, CD96, CD97, CD98, CD99, CD100, CD101, CD102, CD103, CD104, CD105, CD106, CD107a, CD107b, CDw108, CD109, CD114, CD115, CD116, CD117, CDw119, CD120a, CD120b, CD121a, CDw121b, CD122, CD123, CD124, CDw125, CD126, CD127, CDw128a, CDw128b, CD130, CDw131, CD132, CD134, CD135, CDw136, CDw137, CD138, CD139, CD140a, CD140b, CD141, CD142, CD143, CD144, CDw145, CD146, CD147, CD148, CDw149, CDw150, CD151, CD152CD153, CD154, CD155, CD156, CD157, CD158a, CD158b, CD161, CD162, CD163, CD164, CD165, CD166 and TCR-ζ. ,

[0247] Other cell markers include centromere protein-F (CENP-F), macroprotein, epidermal protein, lamin A&C [XB 10], LAP-70, mucin, nuclear pore complex protein, p180 laminin, ran, r, cathepsin D, Ps2 protein, Her2-neu, P53, S100, epithelial marker antigen (EMA), TdT, MB2, MB3, PCNA, and Ki67.

[0248] tissue staining

[0249] The training biological samples of this disclosure can be stained using any reagent or biomarker that reacts directly with a specific biomarker or with various types of cells or cell compartments, such as dyes or staining agents, histochemicals, nucleic acid probes, or immunohistochemical materials. Such histochemicals can be chromophores detectable by transmission (or reflection) microscopy or fluorophores detectable by fluorescence microscopy. Typically, the training biological samples of this disclosure can be incubated with a solution containing at least one histochemical that will react directly with or bind to the target chemical groups. Some histochemicals must be co-incubated with a mordant or metal for staining. The training biological samples can be incubated with a mixture of at least one histochemical that stains the target component and another histochemical that acts as a counterstain and binds to the outer regions of the target component. Alternatively, a mixture of multiple probes can be used for staining, providing a method for identifying specific probe sites. The training biological samples of this disclosure can be co-incubated with a suitable substrate of an enzyme that is a component of the target cell and a suitable reagent that produces a colored precipitate at the enzyme's active site.

[0250] Immunohistochemistry is one of the most sensitive and specific histochemical techniques. Any training biological sample disclosed herein can be combined with a labeled binding component containing a specific binding agent. Various labels can be used, such as fluorophores or enzymes that produce products that absorb light or fluoresce. Multiple labels are known to provide strong signals associated with a single binding event. Multiple probes used in staining can be labeled with more than one distinguishable fluorescent label. These color differences provide a method for identifying specific probe locations. Methods for preparing fluorophore and protein (such as antibody) conjugates are extensively described in the literature and need not be illustrated here.

[0251] Examples of suitable immunohistochemical staining for research and, in limited circumstances, for the diagnosis of various diseases include, for example, anti-estrogen receptor antibodies (breast cancer), anti-progesterone receptor antibodies (breast cancer), anti-p53 antibodies (various cancers), anti-Her-2 / neu antibodies (various cancers), anti-EGFR antibodies (epidermal growth factor, various cancers), anti-cathepsin D antibodies (breast cancer and other cancers), anti-Bcl-2 antibodies (apoptotic cells), anti-E-cadherin antibodies, anti-CA125 antibodies (ovarian cancer and other cancers), anti-CA15-3 antibodies (breast cancer), anti-CA19-9 antibodies (colon cancer), anti-c-erbB-2 antibodies, anti-P-glycoprotein antibodies (MDR, multidrug resistance), anti-CEA antibodies (carcinoembryonic antigen), anti-retinoblastoma protein (Rb) antibodies, anti-rasoneoprotein (p21) antibodies, anti-Lewis X (also known as CD15) antibodies, anti-Ki-67 antibodies (cell proliferation), anti-PCNA (various cancers) antibodies, anti-CD3 antibodies (T cells), and anti-CD4 antibodies. Antibodies (helper T cells), anti-CD5 antibody (T cells), anti-CD7 antibody (thymocytes, immature T cells, NK killer cells), anti-CD8 antibody (suppressor T cells), anti-CD9 / p24 antibody (ALL), anti-CD10 (also known as CALLA) antibody (common acute lymphoblastic leukemia), anti-CD11c antibody (monocytes, granulocytes, AML), anti-CD13 antibody (granulocytes and monocytes, AML), anti-CD14 antibody (mature monocytes, granulocytes), anti-CD15 antibody (Hodgkin's disease), anti-CD19 antibody (B cells), anti-CD20 antibody (B cells), anti-CD22 antibody (B cells), anti-CD23 antibody (activated B cells, CLL), anti-CD30 antibody (activated T cells and B cells, Hodgkin's disease), anti-CD31 antibody (angiogenic marker), anti-CD33 antibody (bone marrow cells, AML), anti-CD34 antibody. Antibodies (endothelial stem cells, mesenchymal tumors), anti-CD35 antibody (dendritic cells), anti-CD38 antibody (plasma cells, activated T, B, and bone marrow cells), anti-CD41 antibody (platelets, megakaryocytes), anti-LCA / CD45 antibody (leukocyte common antigen), anti-CD45RO antibody (helper and induced T cells), anti-CD45RA antibody (B cells), anti-CD39, CD100 antibodies, anti-CD95 / Fas antibody (apoptosis), anti-CD99 antibody (Ewing's sarcoma marker, MIC2 gene product), anti-CD106 antibody (VCAM-1;Activated endothelial cells), anti-ubiquitin antibodies (Alzheimer's disease), anti-CD71 (transferrin receptor) antibodies, anti-c-myc (oncoprotein and hapten) antibodies, anti-cytokeratin (transferrin receptor) antibodies, anti-vimentin (endothelial cells) antibodies (B cells and T cells), anti-HPV protein (human papillomavirus) antibodies, anti-κ light chain antibodies (B cells), anti-λ light chain antibodies (B cells), anti-melanosome (HMB45) antibodies (melanoma), anti-prostate-specific antigen (PSA) antibodies (prostate cancer), anti-S-100 antibodies (melanoma, saliva, glial cells), anti-tau antibodies. Antigen antibodies (Alzheimer's disease), anti-fibrin antibodies (epithelial cells), anti-keratin antibodies, anti-cytokeratin antibodies (tumors), anti-α-catenin (cell membranes), anti-Tn-antigen antibodies (colon cancer, adenocarcinoma, and pancreatic cancer); anti-1,8-ANS (1-anilinonaphthyl-8-sulfonic acid) antibody; anti-C4 antibody; anti-2C4CASP grade antibody; anti-2C4CASP antibody; anti-HER-2 antibody; anti-α-B lens protein antibody; anti-α-galactosidase A antibody; anti-α-catenin antibody; anti-human VEGF R1 (Flt-1) antibody; anti-integrin B5 antibody; anti-integrin β6 antibody; anti-SRC phosphate antibody; anti-Bak antibody; anti-BCL-2 antibody; anti-BCL-6 antibody; anti-β-chain protein antibody; anti-β-catenin antibody; anti-integrin αVβ3 antibody; anti-cErbB-2 Ab-12 Antibodies; Anti-calnectin antibody; Anti-calreticulin antibody; Anti-calreticulin antibody; Anti-CAM5.2 (anti-low molecular weight cytokeratin) antibody; Anti-cardiacin (R2G) antibody; Anti-cathepsin D antibody; α-polyclonal antibody against chicken galactosidase; Anti-c-Met antibody; Anti-CREB antibody; Anti-COX6C antibody; Anti-cyclin D1 Ab-4 antibody; Anti-cytokeratin antibody; Anti-decontin antibody; Anti-DHP (1,6-diphenyl-1,3,5-hextriene) antibody; DSB-X biotinylated goat anti-chicken antibody; Anti-E-cadherin antibody; Anti-EEA1 antibody; Anti-EGFR antibody; Anti-EMA (epithelial membrane antigen) antibody; Anti-ER (estrogen receptor) antibody; Anti-ERB3 antibody; Anti-ERCC1 ERK (Pan ERK) antibody; Anti-E-selectin antibody; Anti-FAK antibody; Anti-fibronectin antibody; FITC-goat anti-mouse IgM Antibodies; anti-FOXP3 antibody; anti-GB3 antibody; anti-GFAP (glial fibrillary acidic protein) antibody; anti-macroprotein antibody; anti-GM130 antibody; anti-goat ah Met antibody; anti-Golgin 97 antibody;Anti-GRB2 antibody; anti-GRP78BiP antibody; anti-GSK-3β antibody; anti-hepatocyte antibody; anti-HER-2 antibody; anti-HER-3 antibody; anti-histone 3 antibody; anti-histone 4 antibody; anti-histone H2A X antibody; anti-histone H2B antibody; anti-HSP70 antibody; anti-ICAM-1 antibody; anti-IGF-1 antibody; anti-IGF-1 receptor antibody; anti-IGF-1 receptor β antibody; anti-IGF-II antibody; anti-IKB-α antibody; anti-IL6 antibody; anti-IL8 antibody; anti-integrin 3 antibody; anti-integrin 5 antibody; anti-integrin b8 antibody; anti-serrated 1 antibody; anti-protein kinase C β2 antibody; anti-LAMP-1 antibody; anti-M6P (mannose-6-phosphate receptor) antibody; anti-MAPKAPK-2 antibody; anti-MEK 1 antibody; anti-MEK 2 antibody. Antibodies; Anti-mitochondrial antigen antibody; Anti-mitochondrial label antibody; Anti-mitochondrial green fluorescent probe FM antibody; Anti-MMP-2 antibody; Anti-MMP9 antibody; Anti-Na+ / K ATPase antibody; Anti-Na+ / K ATPase α1 antibody; Anti-Na+ / K ATPase α3 antibody; Anti-N-cadherin antibody; Anti-renin antibody; Anti-NF-KB p50 antibody; Anti-NF-KB p65 antibody; Anti-claw protein 1 antibody; Anti-OxPhos complex IV-Alexa488 conjugated antibody; Anti-p130Cas antibody; Anti-p38 MAPK antibody; Anti-p44 / 42 MAPK antibody; Anti-p504S clone 13H4 antibody; Anti-p53 antibody; Anti-p70 S6K antibody; Anti-p70 phosphokinase blocking peptide antibody; Anti-pan-cadherin antibody; Anti-pylamine antibody; Anti-P-cadherin antibody; Anti-PDI antibody; Anti-phosphorylated AKT Antibodies; Antiphosphorylated CREB antibody; Antiphosphorylated GSK-3-β antibody; Antiphosphorylated GSK-3 β antibody; Antiphosphorylated H3 antibody; Antiphosphorylated MAPKAPK-2 antibody; Antiphosphorylated MEK antibody; Antiphosphorylated p44 / 42 MAPK antibody; Antiphosphorylated p53 antibody; Antiphosphorylated NF-KB p65 antibody; Antiphosphorylated p70 S6 kinase antibody; Antiphosphorylated PKC (Pan) antibody; Antiphosphorylated S6 ribosomal protein antibody; Antiphosphorylated Src antibody; Antiphosphorylated-Bad antibody; Antiphosphorylated-HSP27 antibody; Antiphosphorylated-IKB-a antibody; Antiphosphorylated p44 / 42 MAPK antibody; Antiphosphorylated p70 S6 kinase antibody; Antiphosphorylated-Rb (Ser807 / 811) (retinoblastoma) antibody;Antiphosphorylated HSP-7 antibody; antiphosphorylated p38 antibody; anti-Pim-1 antibody; anti-Pim-2 antibody; anti-PKC β antibody; anti-PKC β11 antibody; anti-podocyte marker protein antibody; anti-PR antibody; anti-PTEN antibody; anti-R1 antibody; anti-Rb 4H1 (retinoblastoma) antibody; anti-R-cadherin antibody; anti-RRM1 antibody; anti-S6 ribosomal protein antibody; anti-S-100 antibody; anti-synaptic protein antibody; anti-synaptic protein antibody; anti-Syndecan 4 antibody; anti-ankle protein antibody; anti-tonin antibody; anti-microtubule protein antibody; anti-urokinase antibody; anti-VCAM-1 antibody; anti-VEGF antibody; anti-vimentin antibody; anti-ZAP-70 antibody; and anti-ZEB.

[0252] Fluorescein conjugable to primary antibodies includes, but is not limited to, fluorescein, rhodamine, Texas Red, Cy2, Cy3, Cy5, VECTOR Red, ELF™ (enzyme-labeled fluorescence), Cy0, Cy0.5, Cy1, Cy1.5, Cy3, Cy3.5, Cy5, Cy7, fluorophore X, calcein, calcein-AM, CRYPTOFLUOR™'S, Orange (42 kDa), Tangerine (35 kDa), Gold (31 kDa), Red (42 kDa), Crimson (40 kDa), BHMP, BHDMAP, Br-Oregon, yellow fluorescein, Alexa dye family, N-[6-(7-nitrobenzene-2-oxa-1,3-benzoxadiazol-4-yl)-amino]hexanoyl (NBD), BODIPY™, dipyrrolemethane boron difluoride, Oregon Green, and MITOTRACKER™. Red, DiOC7 (3), DiIC18, phycoerythrin, phycobiliprotein BPE (240 kDa), RPE (240 kDa), CPC (264 kDa), APC (104 kDa), blue spectral, lake green spectral, green spectral, gold spectral, orange spectral, red spectral, NADH, NADPH, FAD, infrared (IR) dyes, cyclic GDP-ribose (cGDPR), Karl Cofloll fluorescent brightener, lissamine, umbelliferone, tyrosine, and tryptophan. Many other fluorescent probes are available from and / or extensively described in the 8th edition of the "Fluorescent Probes and Research Products Handbook" (2001), and are also available from Molecular Probes, Eugene, Oreg, and many other manufacturers.

[0253] Especially when antibodies originate from different species, further signal amplification can be achieved by using a combination of specific binding agents (such as antibodies and anti-antibodies), where the anti-antibody binds to a conserved region of the target antibody probe. Alternatively, specific binding ligand-receptor pairs (such as biotin-streptavidin) can be used, where the primary antibody is conjugated to one member of the pair, while the other member is labeled with a detectable probe. Thus, a sandwich structure of binding members can be efficiently constructed, where the first binding member binds to the cellular component and serves to provide secondary binding, where the secondary binding member may or may not include a label, and this secondary binding can further provide tertiary binding, where the tertiary binding member will provide the label.

[0254] Secondary antibodies, avidin, streptavidin, or biotin are each independently labeled with a detectable moiety, which can be an enzyme that directs a colorimetric reaction of a substrate having a substantially insoluble colorimetric product, a fluorescent dye (staining agent), a luminescent dye, or a non-fluorescent dye. Examples for each option are listed below.

[0255] In principle, any enzyme (i) can be conjugated to or indirectly bound to a primary antibody (e.g., via a conjugated avidin, streptavidin, biotin, secondary antibody), and (ii) uses a soluble substrate to provide an available insoluble product (precipitate). The enzyme used can be, for example, alkaline phosphatase, horseradish peroxidase, β-galactosidase, and / or glucose oxidase; and the substrate can be an alkaline phosphatase, horseradish peroxidase, β-galactosidase, or glucose oxidase substrate, respectively.

[0256] Alkaline phosphatase (AP) substrates include, but are not limited to, AP-Blue substrate (blue precipitate, Zymed, page 61), AP-Orange substrate (orange, precipitate, Zymed), AP-Red substrate (red, red precipitate, Zymed), 5-bromo,4-chloro,3-indole phosphate (BCIP substrate, greenish-blue precipitate), 5-bromo,4-chloro,3-indole phosphate / nitroblue tetrazolium / iodonitrotetrazolium (BCIP / INT substrate, yellowish-brown precipitate, Biomeda), 5-bromo,4-chloro,3-indole phosphate / nitroblue tetrazolium (BCIP / NBT substrate, blue / purple), 5-bromo,4-chloro, 3-Indole phosphate / Nitroblue tetrazolium / Iodonitrotetrazole (BCIP / NBT / INT, brown precipitate, DAKO, Solid Red (red), Magenta Phosphate (magenta), Naphthol AS-bisphosphate (NABP) / Solid Red TR (red), Naphthol AS-BI-phosphate (NABP) / New Magenta Red (red), Naphthol AS-MX-phosphate (NAMP) / New Magenta Red (red), New Magenta Red AP Substrate (red), p-Nitrophenyl phosphate (PNPP, yellow, water soluble), VECTOR™ Black, VECTOR™ Blue, VECTOR™ Red, Vega Red (raspberry red).

[0257] Horseradish peroxidase (HRP, sometimes abbreviated as PO) substrates include, but are not limited to, 2,2′-azino-di-3-ethylphenyl-thiazoline sulfonate (ABTS, green, water-soluble), aminoethylcarbazole, and 3-amino,9-ethylcarbazole AEC (3A9EC, red). Other substrates include α-naphtholpyranol (red), 4-chloro-1-naphthol (4C1N, blue, blue-black), 3,3'-diaminobenzidine tetrahydrochloride (DAB, brown), o-benzidine (green), o-phenylenediamine (OPD, brown, water-soluble), TACS Blue (blue), TACS Red (red), 3,3',5,5'-tetramethylbenzidine (TMB, green or green / blue), TRUE BLUE™ (blue), VECTOR™ VIP (purple), VECTOR™ SG (smoky blue-gray), and Zymed Blue HRP substrate (bright blue).

[0258] Glucose oxidase (GO) substrates include, but are not limited to, nitroblue tetrazolium (NBT, purple precipitate), tetranitroblue tetrazolium (TNBT, black precipitate), 2-(4-iodophenyl)-5-(4-nitrobenzene)-3-phenyltetrazole chloride (INT, red or orange precipitate), tetrazolium blue (blue), nitrotetrazole violet (purple), and 3-(4,5-dimethylthiazol-2-yl)-2,5-diphenyltetrazole bromide (MTT, purple). All tetrazolium substrates require glucose as a co-substrate. Glucose is oxidized, and the tetrazolium salt is reduced to form an insoluble formazan that forms a colored precipitate.

[0259] β-galactosidase substrates include, but are not limited to, 5-bromo-4-chloro-3-indolylβ-D-galactopyranoside (X-gal, blue precipitate). Each of the listed substrates has a unique detectable spectral signature (composition).

[0260] The enzyme may also have a substantially insoluble reaction product capable of luminescence or directing a second reaction of a second substrate for a luminescent reaction of a catalytic substrate (e.g., but not limited to luciferase and jellyfish luminescent protein), such as, but not limited to, luciferin and ATP or coelenterin and Ca2+ as luminescent products.

[0261] Nucleic acid biomarkers can be detected using in situ hybridization (ISH). Typically, nucleic acid sequence probes are synthetic and labeled with a member of a fluorescent probe or ligand-receptor pair (such as biotin / avidin, which is labeled with a detectable moiety). Exemplary probes and moieties are described in a previous section. The sequence probe is complementary to the target nucleotide sequence in the cell. Each cell or cell compartment containing the target nucleotide sequence can bind to the labeled probe.

[0262] The probes used in the analysis can be DNA or RNA oligonucleotides or polynucleotides, and can contain not only naturally occurring nucleotides but also their analogs, such as dioxin dCTP, biotin dcTP, 7-azaguanosine, zidovudine, inosine, or uridine. Other useful probes include peptide probes and their analogs, branched gene DNA, peptide mimics, peptide nucleic acids, and / or antibodies. The probe should be sufficiently complementary to the target nucleic acid sequence to ensure stable and specific binding between the target nucleic acid sequence and the probe. The degree of homology required for stable hybridization varies with the stringency of the hybridization. Leitch et al. in “In Situ Hybridization: a practical guide,” Oxford BIOS Science Press, “Microscopy Handbooks,” 27th edition (1994), and Sambrook, J., Fritsch, E. F., and Maniatis, T. in “Molecular Cloning: A Laboratory Manual,” Cold Spring Harbor Press (1989), describe routine methodologies for ISH, hybridization, and probe selection.

[0263] Other system components

[0264] The system 200 disclosed herein can be coupled to a sample processing device capable of performing one or more preparation processes on the tissue sample. The preparation processes may include, but are not limited to, sample deparaffinization, sample conditioning (e.g., cell conditioning), sample staining, antigen retrieval, immunohistochemical staining (including labeling) or other reactions, and / or in situ hybridization (e.g., SISH, FISH, etc.) staining (including labeling) or other reactions, as well as other processes for preparing samples for microscopic examination, microanalysis, mass spectrometry, or other analytical methods.

[0265] The processing device can apply a fixative to the sample. The fixative may include cross-linking agents (e.g., aldehydes such as formaldehyde, polyoxymethylene, and glutaraldehyde, as well as non-aldehyde cross-linking agents), oxidizing agents (e.g., metal ions and complexes such as osmium tetroxide and chromic acid), protein denaturing agents (e.g., acetic acid, methanol, and ethanol), fixatives with unknown mechanisms (e.g., mercuric chloride, acetone, and picric acid), combination reagents (e.g., Carnoy fixative, Methacarn, Bouin solution, B5 fixative, Rossman solution, and Gendre solution), microwave fixatives, and other fixatives (e.g., exclusion volume fixation and vapor fixation).

[0266] If the sample is embedded in paraffin, it can be deparaffinized using a suitable deparaffin remover. After deparaffin removal, any number of chemical substances can be continuously applied to the sample. These substances can be used for pretreatment (e.g., reversing protein cross-linking, exposing cells to acids, etc.), denaturation, hybridization, washing (e.g., rigorous washing), detection (e.g., linking display or marker molecules to probes), amplification (e.g., amplifying proteins, genes, etc.), counterstaining, coverslips, etc.

[0267] The sample processing device can apply a variety of different chemicals to the sample. These chemicals include, but are not limited to, staining agents, probes, reagents, rinsing agents, and / or conditioning agents. These chemicals can be fluids (such as gases, liquids, or gas / liquid mixtures) or similar substances. The fluids can be solvents (such as polar solvents, non-polar solvents, etc.), solutions (such as aqueous solutions or other types of solutions), or similar substances. Reagents can include, but are not limited to, staining agents, wetting agents, antibodies (such as monoclonal antibodies, polyclonal antibodies, etc.), antigen recovery solutions (such as water-based or non-water-based antigen retrieval solutions, antigen recovery buffers, etc.), or similar substances. Probes can be isolated cellular acids or isolated synthetic oligonucleotides attached to detectable tags or reporter molecules. Labels can include radioactive isotopes, enzyme substrates, cofactors, ligands, chemiluminescent or fluorescent agents, haptens, and enzymes.

[0268] After sample processing, the user can transport the sample slide to an imaging device. In some embodiments, the imaging device is a brightfield imager slide scanner. One brightfield imager is the iScan Coreo brightfield scanner sold by Ventana Medical Systems, Inc. In automated embodiments, the imaging device is the digital pathology device disclosed in International Patent Application No. PCT / US2010 / 002772 entitled “IMAGING SYSTEM AND TECHNIQUES” (Patent Publication No.: WO / 2011 / 049608) or U.S. Patent Publication No. 61 / 533,114 entitled “IMAGING SYSTEMS, CASSETTES, ANDMETHODS OF USING THE SAME”, filed September 9, 2011. The disclosures of International Patent Application No. PCT / US2010 / 002772 and U.S. Patent Application No. 61 / 533,114 are incorporated herein by reference in their entirety.

[0269] Embodiments of the subject matter and operations described herein may be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the architectures disclosed herein and their equivalents, or in one or more combinations thereof. Embodiments of the subject matter described herein may be implemented as one or more computer programs, such as one or more computer program instruction modules encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Any module described herein may include logic executed by a processor. As used herein, “logic” means any information in the form of instruction signals and / or data that can be applied to influence the operation of a processor. Software is an example of logic.

[0270] Computer storage media can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or may be contained therein. Furthermore, while a computer storage medium is not a propagating signal, it can be a source or destination of computer program instructions encoded as artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices), or may be contained therein. The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0271] The term "programmable processor" encompasses all kinds of devices, apparatuses, and machines used for processing data, including, as examples, programmable microprocessors, computers, systems-on-a-chip, or a combination thereof. Devices may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, devices may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. Devices and execution environments can implement a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.

[0272] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages, declarative or procedural languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordinated files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.

[0273] The processes and logic flows described in this specification can be executed by one or more programmable processors, which execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by dedicated logic circuits, and the device can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0274] As an example, processors suitable for executing computer programs include general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for performing actions according to instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices (such as disks, magneto-optical disks, or optical disks) for storing data, to receive data from, transfer data to, or receive data from and transfer data to. However, a computer does not necessarily have to have such devices. Furthermore, a computer can be embedded in another device, to name just a few, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (such as a universal serial bus USB flash drive). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0275] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device and a keyboard and pointing device (e.g., a mouse or trackball), such as an LCD (liquid crystal display), LED (light-emitting diode) display, or OLED (organic light-emitting diode) display, for displaying information to the user, who can provide input to the computer via the keyboard and pointing device. In some embodiments, a touchscreen can be used to display information and receive input from the user. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback (such as visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form (including sound, speech, or tactile input). Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by responding to a request received from a web browser by sending a web page to a web browser on the user's client device.

[0276] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., data servers), or middleware components (e.g., application servers), or front-end components (e.g., client computers with graphical user interfaces or web browsers through which users can interact with embodiments of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. Components of the system can be interconnected via any form or medium of digital data communication (e.g., communication networks). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks). For example, network 20 of Figure 1 may include one or more local area networks.

[0277] A computing system may include any number of clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The relationship between clients and servers is generated by computer programs running on their respective computers and having a client-server relationship with each other. In some embodiments, the server sends data (e.g., an HTML page) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input therefrom). Data generated at the client device (e.g., the result of user interaction) can be received from the client device at the server.

[0278] Alternative embodiments

[0279] Another aspect of this disclosure is a method for predicting the expression of one or more biomarkers in an unstained test biological sample with an unknown duration of fixed treatment. The method includes obtaining test spectral data from the unstained test biological sample, wherein the test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine; and predicting the expression of one or more biomarkers in the test biological sample based on the biomarker expression features. In some embodiments, the predicted biomarker expression includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, the predicted biomarker expression includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, the fixed state of the unstained test biological sample is unknown.

[0280] In some embodiments, a biomarker expression estimation engine is trained using one or more training spectral datasets, wherein each training spectral dataset includes multiple training vibrational spectra derived from multiple training tissue samples stained for the presence of one or more biomarkers, and wherein each training vibrational spectrum includes one or more class labels. In some embodiments, the one or more class labels contain known biomarker expression levels of one or more biomarkers. In some embodiments, known biomarker expression levels include at least one of a known positive percentage of one or more biomarkers and a known staining intensity of one or more biomarkers. In some embodiments, the system further includes one or more class labels selected from the group consisting of known demasking duration, known demasking temperature, qualitative assessment of demasking state, known fixation duration, and qualitative assessment of fixation state.

[0281] In some embodiments, the training spectral dataset is derived by: (i) obtaining training biological samples; (ii) dividing the obtained training biological samples into multiple training tissue samples; (iii) staining each of the multiple training tissue samples for the presence of one or more biomarkers; and (iv) quantitatively assessing the expression of one or more biomarkers. In some embodiments, each of the multiple training tissue samples is differentially demasked, differentially fixed, or both differentially demasked and differentially fixed. In some embodiments, the quantitative assessment of one or more biomarkers includes determining the staining intensity of one or more biomarkers. In some embodiments, the quantitative assessment of one or more biomarkers includes determining the positive percentage of one or more biomarkers. In some embodiments, the quantitative assessment is performed by a pathologist. In some embodiments, the quantitative assessment is performed using one or more image analysis algorithms. In some embodiments, the multiple training tissue samples are stained in an immunohistochemical assay. In some embodiments, the multiple training tissue samples are stained in an in situ hybridization assay.

[0282] In some embodiments, the test spectral data includes an averaged vibrational spectrum derived from a plurality of normalized and corrected vibrational spectra. In some embodiments, the plurality of normalized and corrected vibrational spectra are obtained by: (i) identifying a plurality of spatial regions within the test biological sample; (ii) acquiring vibrational spectra from each individual region of the plurality of identified regions; (iii) correcting the vibrational spectra acquired from each individual region to provide a corrected vibrational spectrum for each individual region; and (iv) normalizing the amplitude of the corrected vibrational spectra from each individual region to a predetermined global maximum to provide an amplitude-normalized vibrational spectrum for each region. In some embodiments, the vibrational spectra acquired from each individual region are corrected by: (i) compensating each acquired vibrational spectrum for atmospheric effects to provide an atmospherically corrected vibrational spectrum; and (ii) compensating the atmospherically corrected vibrational spectrum for scattering.

[0283] In some embodiments, the trained biomarker expression estimation engine includes a dimensionality reduction-based machine learning algorithm. In some embodiments, dimensionality reduction includes projection onto a latent structural regression model. In some embodiments, dimensionality reduction includes principal component analysis plus discriminant analysis. In some embodiments, the trained biomarker expression estimation engine includes a neural network.

[0284] In some embodiments, the method further includes comparing the actual expression of a biomarker in the test biological sample with the predicted expression of one or more biomarkers in the test biological sample. In some embodiments, the method further includes the predicted expression of one or more biomarkers for poor demasking and / or poor fixation of the test biological sample. In some embodiments, the test spectral data includes vibrational spectral information of at least one amide I band. In some embodiments, the test spectral data includes wavelengths ranging from about 3200 to about 3400 cm⁻¹. -1 Approximately 2800 to approximately 2900cm -1 Approximately 1020 to approximately 1100 cm -1 and / or about 1520 to about 1580 cm -1 Vibrational spectral information.

[0285] Another aspect of this disclosure is a method for obtaining test spectral data from a test biological sample, predicting the expression of one or more biomarkers in the test biological sample subjected to an unknown amount of time of fixed treatment, wherein the test spectral data includes vibrational spectral data from at least a portion of the biological sample; biomarker expression features are derived from the obtained test spectral data using a trained biomarker expression estimation engine; and the expression of one or more biomarkers in the test biological sample is predicted based on the biomarker expression features. In some embodiments, the predicted biomarker expression includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, the predicted biomarker expression includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, the fixed state of the test biological sample is unknown. In some embodiments, the test biological sample is stained for the presence of one or more biomarkers, including any of the biomarkers listed above. In other embodiments, the test biological sample is unstained.

[0286] Another aspect of this disclosure is a system for predicting the expression of one or more biomarkers in an unstained test biological sample, the system comprising: (i) one or more processors, and (ii) one or more memories coupled to the one or more processors, the one or more memories storing computer-executable instructions that, when executed by the one or more processors, cause the system to perform operations including: obtaining test spectral data from the test biological sample, wherein the test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine, wherein the biomarker expression estimation engine is trained using a training spectral dataset collected from multiple differentially prepared training biological samples, and wherein the training spectral dataset includes class labels for known biomarker expressions of one or more biomarkers; and predicting the expression of another biomarker in the unstained biological sample based on the derived biomarker expression features.

[0287] In some embodiments, predicted biomarker expression includes either a predicted positive percentage or a predicted staining intensity. In some embodiments, predicted biomarker expression includes both a predicted positive percentage and a predicted staining intensity. In some embodiments, one or more biomarkers include at least one cancer biomarker.

[0288] In some embodiments, each training spectral dataset is derived by: (i) obtaining training biological samples; (ii) dividing the obtained training biological samples into multiple training tissue samples; and (iii) preparing each of the multiple training tissue samples under different preparation conditions. In some embodiments, the method further includes staining each of the multiple training tissue samples for the presence of one or more biomarkers; and quantitatively assessing the known positive percentage and / or known staining intensity of one or more biomarkers. In some embodiments, the trained biomarker expression estimation engine includes a dimensionality reduction-based machine learning algorithm. In some embodiments, dimensionality reduction includes projecting onto a latent structure regression model. In some embodiments, the trained biomarker expression estimation engine includes a neural network. In some embodiments, the method further includes compensating for poor demasking and / or poor fixation of the predicted expression of one or more biomarkers in the test biological samples.

[0289] Another aspect of this disclosure is a system for predicting the expression of one or more biomarkers in a test biological sample, the system comprising: (i) one or more processors, and (ii) one or more memories coupled to the one or more processors, the one or more memories storing computer-executable instructions, which, when executed by the one or more processors, cause the system to perform operations including: obtaining test spectral data from the test biological sample, wherein the test spectral data includes vibrational spectral data derived from at least a portion of the biological sample; deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine, wherein the biomarker expression estimation engine is trained using a training spectral dataset collected from multiple differentially prepared training biological samples, and wherein the training spectral dataset includes class labels for known biomarker expression of one or more biomarkers; and predicting the expression of another biomarker in the biological sample based on the derived biomarker expression features. In some embodiments, the test biological sample is stained for the presence of one or more biomarkers, including any of the biomarkers listed above. In other embodiments, the test biological sample is unstained.

[0290] All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patents, and non-patent publications mentioned in and / or listed in the application data sheets are incorporated herein by reference in their entirety. Modifications may be made to various aspects of the embodiments as necessary to provide other further embodiments employing the concepts of various patents, applications, and publications.

[0291] Although this disclosure has been described with reference to some illustrative embodiments, it should be understood that those skilled in the art can devise many other modifications and embodiments within the spirit and scope of the principles of this disclosure. More specifically, reasonable variations and modifications can be made to the components and / or arrangements of the subject matter combination within the scope of the foregoing disclosure, the drawings, and the appended claims without departing from the spirit of this disclosure. In addition to variations and modifications in the components and / or arrangements, alternative uses will also be apparent to those skilled in the art.

Claims

1. A system (200) for predicting expression of one or more biomarkers in a test biological sample, the system (200) comprising: (i) one or more processors (209), and (ii) one or more memories (201) coupled with the one or more processors (209) for storing computer-executable instructions that, when executed by the one or more processors (209), cause the system (200) to perform operations comprising: a. obtaining test spectral data from the test biological sample, wherein the obtained test spectral data comprises vibrational spectral data derived from at least a portion of the biological sample; b. deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine (210); and c. predicting expression of the one or more biomarkers in the test biological sample based on the derived biomarker expression features, wherein the predicted expression of the one or more biomarkers comprises a predicted positive percentage, wherein the prediction of expression of the one or more biomarkers in the biological test sample is label-free, wherein the biological sample is unstained.

2. The system of claim 1, wherein the system comprises a training module (211) adapted to receive training vibrational spectral data and to train the biomarker expression estimation engine (210) using the received training vibrational spectral data, wherein the biomarker expression estimation engine is trained using one or more training spectral data sets, wherein each training spectral data set comprises a plurality of training vibrational spectra derived from a plurality of training tissue samples stained for presence of one or more biomarkers, and wherein each training vibrational spectrum comprises one or more class labels.

3. The system of claim 2, wherein the one or more class labels comprise known biomarker expression levels of the one or more biomarkers, wherein the known biomarker expression levels comprise known positive percentages of the one or more biomarkers.

4. The system of claim 3, further comprising one or more class labels selected from the group consisting of known unmasking durations, known unmasking temperatures, qualitative assessment of unmasking status, known fixation durations, and qualitative assessment of fixation status.

5. The system of any of claims 2-4, wherein each training spectral data set is derived by: (i) obtaining a training biological sample; (ii) dividing the obtained training biological sample into a plurality of training tissue samples, wherein each training tissue sample of the plurality of training tissue samples is differentially unmasked, differentially fixed, or both differentially unmasked and differentially fixed; (iii) staining the plurality of training tissue samples for the presence of one or more biomarkers; and (iv) quantitatively assessing expression of the one or more biomarkers in each training tissue sample of the plurality of training tissue samples, wherein the quantitative assessment of the one or more biomarkers comprises determining a percentage of positivity of the one or more biomarkers.

6. The system of claim 5, wherein the quantitative assessment of the one or more biomarkers is performed using one or more image analysis algorithms.

7. The system of any of the preceding claims, wherein the trained biomarker expression estimation engine comprises a machine learning algorithm based on dimension reduction, wherein the dimension reduction comprises projection onto a latent structure regression model, or wherein the dimension reduction comprises principal component analysis plus discriminant analysis.

8. The system of any of the preceding claims, further comprising operations for comparing actual biomarker expression of the test biological sample to predicted expression of the one or more biomarkers of the test biological sample.

9. The system of any of the preceding claims, further comprising operations for compensating predicted expression of the one or more biomarkers for poor unmasking and / or poor fixation of the test biological sample.

10. A non-transitory computer-readable medium storing instructions for predicting expression of one or more biomarkers in a test biological sample being processed, the test biological sample having an unknown fixation status and / or an unknown unmasking status, comprising: (a) obtaining test spectral data from the test biological sample, wherein the obtained test spectral data comprises vibrational spectral data derived from at least a portion of the biological sample; (b) deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine (210), wherein the biomarker expression estimation engine is trained using training spectral data sets collected from a plurality of differentially prepared training biological samples, and wherein the training spectral data sets comprise class labels of known biomarker expression of one or more biomarkers; and (c) predicting expression of another biomarker in the test biological sample based on the derived biomarker expression features, wherein the predicted expression of the one or more biomarkers comprises a predicted percentage of positivity, wherein the prediction of expression of the one or more biomarkers in the biological test sample is label-free, wherein the biological sample is unstained.

11. The non-transitory computer-readable medium of claim 10, wherein each training spectral data set is derived by: (i) obtaining a training biological sample; (ii) dividing the obtained training biological sample into a plurality of training tissue samples; and (iii) preparing each training tissue sample of the plurality of training tissue samples under different preparation conditions; (iv) staining each training tissue sample of the plurality of training tissue samples for the presence of one or more biomarkers; and (v) quantitatively assessing expression of the one or more biomarkers in each training tissue sample of the training tissue samples.

12. A method for predicting expression of one or more biomarkers in a test biological sample fixed for an unknown amount of time, comprising: a. obtaining test spectral data from the test biological sample, wherein the obtained test spectral data comprises vibrational spectral data derived from at least a portion of the biological sample; b. deriving biomarker expression features from the obtained test spectral data using a trained biomarker expression estimation engine, wherein the biomarker expression estimation engine is trained using a plurality of training spectral data sets collected from a plurality of differentially prepared training biological samples, and wherein the training spectral data sets contain class labels of known biomarker expression of one or more biomarkers; and c. predicting expression of another biomarker in the test biological sample based on the derived biomarker expression features, wherein the predicted expression of the one or more biomarkers comprises a predicted positive percentage, wherein the prediction of expression of the one or more biomarkers in the biological test sample is label-free, wherein the biological sample is unstained.

13. The method of claim 12, wherein each training spectral data set is derived by: (i) obtaining a training biological sample; (ii) dividing the obtained training biological sample into a plurality of training tissue samples; and (iii) preparing each training tissue sample of the plurality of training tissue samples under different preparation conditions.

14. The method of claim 13, further comprising staining each of the plurality of training tissue samples for the presence of one or more biomarkers; and quantitatively assessing the known positive percentage of the one or more biomarkers.

15. The method of any one of claims 12-14, wherein the trained biomarker expression estimation engine comprises a machine learning algorithm based on dimension reduction, wherein the dimension reduction comprises projection onto a latent structural regression model.

16. The method of any one of claims 12-15, further comprising compensating for the predicted expression of the one or more biomarkers for poor unmasking and / or poor fixation of the test biological sample.

Citation Information

Patent Citations

  • Mid-infrared super-continuum laser

    US10041832B2

  • Method for improving operation of a robot

    US10279474B2

  • Face recognition apparatus and method using PCA learning per subgroup

    US20050123202A1

  • Predictive model validation

    US20050234753A1

  • Performing Cross-Validation Using Non-Randomly Selected Cases

    US20140279734A1