System for automated in situ hybridization analysis

By using automated systems and methods to identify and register images of protein and nucleic acid biomarkers, and to assess gene aberrations in tumor tissue regions, this approach addresses the issue of high subjectivity in existing technologies, improves assessment efficiency and accuracy, and enhances the development of patient treatment plans.

CN112534439BActive Publication Date: 2026-01-30VENTANA MEDICAL SYSTEMS INC
View PDF 25 Cites 0 Cited by

Patent Information

Application Number
CN201980048897.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-07-27
Filing Date
2019-07-10
Publication Date
2026-01-30
Estimated Expiration
2040-07-18

AI Technical Summary

Technical Problem

Existing technologies suffer from high subjectivity and low efficiency when detecting gene aberrations in biological samples, especially when assessing gene aberrations within tumor tissue regions, which is difficult to automate and perform efficiently.

Method used

A system and method are provided to identify cells that meet predetermined protein biomarker staining criteria, derive tumor tissue regions, and register protein biomarker images with nucleic acid biomarker images using common coordinate system registration technology, identify signal points, calculate ratios and assess gene aberrations, generate overlapping images and histograms, and automatically select cells for assessment to reduce subjectivity.

Benefits of technology

It enables more stable and rapid assessment of gene aberrations, reduces the subjectivity of manual cell selection, allows for the analysis of more cells, improves assessment efficiency and accuracy, and enhances the formulation of patient treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112534439B_ABST
    Figure CN112534439B_ABST
Patent Text Reader

Abstract

This disclosure provides image processing systems and methods for automatically analyzing digital images (311) of biological samples stained for the presence of protein and / or nucleic acid biomarkers and automatically detecting and quantifying signals (314) corresponding to one or more biomarkers. This disclosure also provides systems and methods for clinically interpreting dual-ISH slides in which cells to be scored are automatically selected (e.g., using one or more cell detection and identification algorithms (204)). Subjectivity is believed to be reduced or eliminated through automated detection, identification, and selection of cells for assessment. Compared to manual dot-counting methods, the automated systems and methods also allow for an increase in the number of cells considered for scoring, thereby improving detection sensitivity and ultimately improving patient care and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This disclosure claims the benefit of U.S. Provisional Patent Application No. 62 / 711,049, filed July 27, 2018, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure provides systems and methods for detecting and classifying signals in images of stained biological samples. Background Technology

[0004] Digital pathology involves scanning entire histopathological or cytopathological slides into digital images that can be interpreted on a computer screen. These images are then processed by imaging algorithms or interpreted by a pathologist. To examine tissue sections (which are almost transparent), the sections are prepared with color stains that selectively bind to cellular components. Clinicians or computer-aided diagnostic (CAD) algorithms use the color-enhanced or stained cellular structures to identify morphological markers of disease and guide appropriate treatments. Various processes can be achieved through measurement and observation, including diagnosing diseases, assessing treatment response, and developing new antiviral drugs.

[0005] Immunohistochemical (IHC) staining of slides can be used to identify proteins in cells within tissue sections, and is therefore widely used to study different types of cells, such as cancer cells and immune cells in biological tissues. Thus, IHC staining can be used to study the distribution and localization of differentially expressed biomarkers of immune cells (such as T cells or B cells) in cancerous tissues, for use in immune response studies. For example, tumors often contain infiltrations of immune cells, which may either inhibit tumor development or promote tumor growth.

[0006] In situ hybridization (ISH) can be used to determine the presence of genetic abnormalities or oncogenes that exhibit malignant morphology in cells when observed under a microscope. Unique nucleic acid sequences are precisely located in chromosomes, cells, and tissues, and ISH can determine the presence, deletion, and / or amplification of these sequences without significant breakage. ISH uses labeled DNA or RNA probe molecules that are antisense to the target gene sequence or transcript to detect or locate targeted nucleic acid genes within cell or tissue samples. ISH is performed by exposing a cell or tissue sample immobilized on a glass slide to labeled nucleic acid probes that specifically hybridize to a given target gene in the cell or tissue sample. Multiple target genes can be analyzed simultaneously by exposing a cell or tissue sample to multiple nucleic acid probes labeled with multiple different nucleic acid tags. Using labels with different emission wavelengths, multicolor analysis can be performed simultaneously in a single step on a single target cell or tissue sample. Summary of the Invention

[0007] It is believed that IHC and ISH can target different molecules (such as biomarkers), one of which may be a precursor to another. Similarly, performing ISH and IHC together (e.g., simultaneously or sequentially) can provide complementary information to identify the source of secreted proteins, determine complex tissue structures, identify the regulation of gene expression, and / or assess therapy. In view of the foregoing, in some embodiments, this disclosure provides systems and methods for detecting intracellular genetic aberrations (such as high copy numbers, chromosomal abnormalities, etc.) selected for assessment (e.g., automated selection assessment). In some embodiments, the automatically selected cells for assessment are located within a tumor tissue region comprising cells that meet predetermined protein biomarker staining criteria. It is believed that automated cell selection can reduce and / or eliminate any subjectivity introduced during manual cell selection. Furthermore, the systems and methods of this disclosure allow for the use of a larger number of cells to detect genetic aberrations, thereby facilitating more stable assessments and ultimately providing improved patient care and therapy.

[0008] One aspect of this disclosure is a system for assessing genetic aberrations in images of biological samples (e.g., samples stained for the presence of at least one nucleic acid biomarker and / or protein biomarker), the system comprising: (i) one or more processors, and (ii) one or more memories coupled to the processors, the memories storing computer-executable instructions that, when executed by the processors, cause the system to perform operations including: (a) identifying cells in a first image stained for the presence of at least one protein biomarker (e.g., the HER2 protein biomarker) that meet the protein biomarker staining criteria; and (b) deriving the genetic aberrations in the first image. (c) Covering tumor tissue regions (e.g., epithelial tumor tissue regions) of cells that meet the predetermined protein biomarker staining criteria; (d) registering the first and second images with a common coordinate system such that the derived tumor tissue regions in the first image are mapped to the second image to provide mapped tumor tissue regions, wherein the second image includes signals corresponding to the presence of at least one nucleic acid biomarker (e.g., HER2 and chromosome 17 nucleic acid biomarkers); (e) identifying points within the mapped tumor tissue regions corresponding to the signals from the at least one nucleic acid biomarker; and (f) assessing, based on the identified points, whether tumor cell nuclei in the mapped tumor tissue regions in the second image have genetic aberrations (e.g., gene copy number). In some embodiments, the number of cell nuclei assessed in each mapped tissue region is greater than 20.

[0009] In some embodiments, the genetic aberration of the cell nucleus is assessed by determining whether the total number of recognized points in each cell nucleus corresponding to signals from at least one nucleic acid biomarker meets a predetermined threshold (e.g., analyzing a single nucleic acid biomarker). In other embodiments, the genetic aberration of the cell nucleus is assessed by: (i) calculating the ratio of first recognized points in each cell nucleus corresponding to signals from a first nucleic acid biomarker to second recognized points in each cell nucleus corresponding to a second nucleic acid biomarker; and (ii) comparing the calculated ratio for each cell nucleus to a predetermined threshold. In some embodiments, the first nucleic acid biomarker is HER2, and the second nucleic acid biomarker is chromosome 17; and wherein the at least one protein biomarker is a HER2 protein biomarker. In some embodiments, the first nucleic acid biomarker is EGFR, and the second nucleic acid biomarker is chromosome 7; and wherein the at least one protein biomarker is an EGFR protein biomarker. Those skilled in the art will recognize that other protein biomarkers and nucleic acid biomarkers, including biomarkers that are precursors to each other, can be utilized.

[0010] In some embodiments, the system further includes assigning a first indicator (such as a first color or a first symbol) to those rated cell nuclei with a calculation ratio higher than a predetermined threshold, and assigning a second indicator (such as a second color or a second symbol) to those rated cell nuclei with a calculation ratio equal to or lower than the predetermined threshold. In some embodiments, the system further includes generating an overlapping image based on the assigned first indicator and the assigned second indicator. In some embodiments, the generated overlapping image is overlaid on the entire slice image or a portion thereof. In some embodiments, the overlapping image is a foreground segmentation mask.

[0011] In some embodiments, the system further includes sorting the assessed nuclei according to the calculated ratio of each nucleus (e.g., sorting may be in tabular form, where the calculated ratios are categorized in descending order, and the table may optionally include location information, such as the coordinates of nuclei or cells within the image). In some embodiments, the system further includes generating grouped histograms of the calculated ratios of all assessed nuclei. In some embodiments, a separate grouped histogram is generated for each mapped tissue region. In some embodiments, the system further includes identifying the group of histograms with the highest count. In some embodiments, the system further includes determining a treatment course (e.g., whether to implement targeted therapy; whether to implement combination therapy) based on data from the generated grouped histograms.

[0012] In some embodiments, the predetermined protein biomarker staining standard is a staining intensity threshold. In some embodiments, the staining intensity threshold is a cutoff value for the presence of membrane staining. In some embodiments, the predetermined biomarker staining standard is an expression score calculated for the cell. In some embodiments, the gene aberration refers to an abnormal gene copy number. In some embodiments, the abnormal gene copy number is a copy number greater than the normal gene copy number. In some embodiments, the gene aberration refers to a chromosomal abnormality. In some embodiments, the gene aberration refers to RNA overexpression.

[0013] Another aspect of this disclosure is a method for assessing gene aberrations in images of biological samples stained for the presence of at least one nucleic acid biomarker, the method comprising: detecting cells in a first image stained for the presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; deriving a tumor tissue region in the first image encompassing identified cells that meet the predetermined protein biomarker staining criteria; registering the first image and a second image to a common coordinate system such that the derived tumor tissue region in the first image is mapped to the second image to provide a mapped tissue region, wherein the second image includes a signal corresponding to the presence of at least one nucleic acid biomarker; identifying points within the mapped tissue region corresponding to the signal from the at least one nucleic acid biomarker; and assessing whether tumor cell nuclei in the mapped tissue region of the second image have gene aberrations based on the identified points corresponding to the signal from the at least one nucleic acid biomarker. In some embodiments, the gene aberration refers to RNA overexpression. In some embodiments, the gene aberration refers to an abnormal gene copy number. In some embodiments, the abnormal gene copy number is a copy number greater than the normal copy number of the gene.

[0014] In some embodiments, gene aberrations in tumor cell nuclei are assessed by: (i) for each nucleus, calculating the ratio of identified points corresponding to signals from a first nucleic acid biomarker to identified points corresponding to signals from a second nucleic acid biomarker; and (ii) comparing the calculated ratio for each nucleus to a predetermined threshold. In some embodiments, the method further includes assigning a first indicator to those tumor cell nuclei with calculated ratios higher than the predetermined threshold, and assigning a second indicator to those tumor cell nuclei with calculated ratios equal to or lower than the predetermined threshold. In some embodiments, the method further includes generating an overlay image based on the assigned first and second indicators. In some embodiments, the method further includes a grouped histogram of the calculated ratios of all identified nuclei. In some embodiments, the method further includes sorting the assessed tumor cell nuclei according to the calculated ratio for each nucleus.

[0015] In some embodiments, the biological sample is stained for the presence of HER2 and chromosome 17 nucleic acid biomarkers. In embodiments of staining the biological sample for the presence of HER2 and chromosome 17 nucleic acid biomarkers, the method includes detecting points in the mapped tissue region that satisfy absorbance intensity, black unmixed image channel intensity, red unmixed image channel intensity, and Gaussian threshold difference criteria; and classifying the detected points as belonging to a black nucleic acid biomarker signal corresponding to HER2 or a red nucleic acid biomarker signal corresponding to chromosome 17. In embodiments of staining the biological sample for the presence of HER2 and chromosome 17 nucleic acid biomarkers, the tumor cell nuclei are evaluated by: (i) calculating the ratio of points belonging to those classifications of black nucleic acid biomarker signals to points belonging to those classifications of red nucleic acid biomarker signals; and (ii) comparing the calculated ratio with a predetermined threshold. In embodiments of staining the biological sample for the presence of HER2 and chromosome 17 nucleic acid biomarkers, the at least one protein biomarker is a HER2 protein biomarker. In embodiments where the biological sample is stained for the presence of HER2 and chromosome 17 nucleic acid biomarkers, the method further includes identifying whether a patient is HER2-positive or HER2-negative based on the assessed tumor cell nuclei. In embodiments where the biological sample is stained for the presence of HER2 and chromosome 17 nucleic acid biomarkers, the method further includes scoring the biological sample for the presence of at least one additional protein biomarker. In some embodiments, the at least one additional protein biomarker is selected from the group consisting of EGFR.

[0016] Another aspect of this disclosure is a non-transitory computer-readable medium storing instructions for assessing gene aberrations in a biological sample stained for the presence of at least one nucleic acid biomarker, the instructions comprising: identifying cells in a first image stained for the presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; deriving a tumor tissue region in the first image encompassing the identified cells that meet the predetermined protein biomarker staining criteria; registering the first image and a second image to a common coordinate system such that the tumor tissue region derived in the first image is mapped to the second image to provide a mapped tissue region, wherein the second image includes a signal corresponding to the presence of at least one nucleic acid biomarker; detecting points within the mapped tissue region corresponding to the signal from at least one nucleic acid biomarker; counting all detected points within each tumor cell nucleus in each mapped tissue region; and assessing whether each tumor cell nucleus in each mapped region has a gene aberration based on the total number of counted points in each cell nucleus. In some embodiments, the predetermined protein biomarker staining criteria is a staining intensity threshold. In some embodiments, the gene aberration is selected from the group consisting of abnormal gene copy number and chromosomal abnormalities.

[0017] In some embodiments, the biological sample is stained for the presence of at least two nucleic acid biomarkers, and wherein points corresponding to each of the at least two nucleic acid biomarkers are detected and counted in each cell nucleus. In some embodiments, tumor cell nuclei are assessed by: (i) calculating a ratio of a first counted point to a second counted point; and (ii) comparing the calculated ratio to a clinically relevant threshold.

[0018] In some embodiments, the non-transitory computer-readable medium further includes instructions for generating image overlays, wherein each assessed cell nucleus is assigned a color based on a calculated ratio that is (i) at or below a clinically relevant threshold; or (ii) above a clinically relevant threshold. In some embodiments, the non-transitory computer-readable medium further includes instructions for sorting the assessed cell nuclei in each mapped tissue region according to the calculated ratio. In some embodiments, the non-transitory computer-readable medium further includes instructions for generating a grouped histogram of the calculated ratio. Attached Figure Description

[0019] For a general understanding of the features of this disclosure, please refer to the accompanying drawings. In the drawings, the same reference numerals are used throughout to identify the same elements.

[0020] Figure 1 A representative digital pathology system, including an image acquisition device and a computer system, is shown according to some embodiments.

[0021] Figure 2 Various modules that can be used in digital pathology systems or digital pathology workflows, according to some embodiments, are listed.

[0022] Figure 3 Flowcharts illustrating methods for detecting gene aberrations in biological samples according to some embodiments of the present disclosure are provided.

[0023] Figure 4 Flowcharts illustrating methods for detecting gene aberrations in biological samples according to some embodiments of the present disclosure are provided.

[0024] Figure 5 A flowchart illustrating a method for predicting HER2 status according to some embodiments of the present disclosure is provided.

[0025] Figure 6 A flowchart illustrating the steps of registering one or more images to a common coordinate system according to some embodiments of the present disclosure is provided.

[0026] Figure 7A The first slide stained for the presence of the HER2 protein biomarker is shown.

[0027] Figure 7B The second slide, stained for the presence of HER2 nucleic acid biomarkers and chromosome 17 nucleic acid biomarkers, is shown.

[0028] Figure 8A The slides are shown stained for the presence of HER2 protein biomarkers.

[0029] Figure 8B Image analysis results of membrane staining features are also shown.

[0030] Figure 9 The images show the overall picture of the samples stained for the presence of HER2 protein biomarkers before (bottom) and after (top) HER2 image analysis.

[0031] Figure 10 The images show samples stained for the presence of HER2 nucleic acid biomarkers and chromosome 17 nucleic acid biomarkers (bottom), as well as full views of samples stained for the generation of a foreground segmentation mask (top) after cell detection and classification.

[0032] Figure 11 The results of the HER2 dual ISH assay are shown (top: red dots detected; bottom: black dots detected).

[0033] Figure 12 The workflow for testing HER2 protein biomarkers is shown. Detailed Implementation

[0034] It should also be understood that, unless the opposite is specified, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are described.

[0035] As used herein, unless the context clearly indicates otherwise, the singular forms “a / an” and “the / that” include plural referents. Similarly, unless the context clearly indicates otherwise, the word “or” is intended to include “and”. The term “including” is defined as inclusive, such as “including A or B” meaning including A, B, or A and B.

[0036] As used herein in the specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” should be interpreted as inclusive, meaning it includes several elements or at least one element in the list, but also includes more than one element, and optionally includes additional unlisted items. Only terms indicating the opposite, such as “only one” or “exactly one,” or “consisting of…” as used in the claims, will refer to several elements or exactly one element in the list. Generally, the term “or” as used herein should only be interpreted as indicating an exclusive alternative (i.e., “one or the other, but not two”) when preceded by exclusive terms such as “or,” “one of,” “only one,” or “exactly one.” “Constitutes substantially of…” as used in the claims should have the ordinary meaning used in the field of patent law.

[0037] The terms “comprising,” “including,” and “having” are used interchangeably and have the same meaning. Specifically, the definition of each term is consistent with the definition of “comprising” under ordinary U.S. patent law, and therefore each term can be understood as an open-ended term meaning “at least the following,” and can also be interpreted as not excluding additional features, limitations, aspects, etc. Thus, for example, “an apparatus having components a, b, and c” means that the apparatus includes at least components a, b, and c. Similarly, the phrase “a method involving steps a, b, and c” means that the method includes at least steps a, b, and c. Furthermore, although the steps and processes may be described herein in a specific order, those skilled in the art will recognize that the order of steps and processes may vary.

[0038] As used herein in the specification and claims, with respect to a list of one or more elements, the phrase "at least one" should be understood as at least one element selected from any one or more elements in the list, but does not necessarily include at least one of each element specifically listed in the list, nor exclude any combination of elements in the list. In addition to the elements specifically identified in the list of elements referred to by the phrase "at least one," this definition also allows for the optional presence of other elements, whether or not they are related to the specifically identified elements. Thus, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") in one embodiment may mean at least one element that optionally includes more than one A, but no B (and selectively includes elements other than B); in another embodiment, it means at least one element that selectively includes more than one B, but no A (and selectively includes elements other than A); in yet another embodiment, it means at least one element that selectively includes more than one A, and at least one element that selectively includes more than one B (and selectively includes other elements), and so on.

[0039] As used herein, the terms “biological sample,” “tissue sample,” “specimen,” or similar terms refer to any sample obtained from any organism, including viruses, that includes biomolecules (e.g., proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof). Examples of other organisms include mammals (e.g., humans; mammals such as cats, dogs, horses, cattle, and pigs; and laboratory animals such as mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (e.g., tissue sections and needle biopsies of tissues), cell samples (e.g., cytological smears, such as cervical smears or blood smears, or obtained through microdissection), or cell fractions, fragments, or organelles (e.g., obtained by lysing cells and separating their components by centrifugation or other methods). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical or needle biopsy), nipple aspiration, cerumen, breast milk, vaginal secretions, saliva, swabs (e.g., oral swabs), or any material containing biomolecules derived from the first biological sample. In some embodiments, the term "biological sample" as used herein refers to a sample prepared from a tumor or a portion thereof obtained from a subject (e.g., a homogenized or liquefied sample).

[0040] The term "blob" or "dot" as used in this article refers to regions in a digital image where some properties are constant or nearly constant; in a sense, all pixels in a blob can be considered similar to each other. Depending on the specific application and the in-situ signal to be detected, a "dot" typically comprises 5-60 pixels, thus, for example, one pixel may correspond to a tissue slice of approximately 0.25 micrometers by 0.25 micrometers.

[0041] As used herein, the phrase “double in situ hybridization” or “DISH” refers to an in situ hybridization (ISH) method that uses two probes to detect two different target sequences. Typically, the two probes are labeled differently. In some embodiments, DISH can be a assay that determines the HER2 gene amplification status by contacting a tumor sample with a HER2-specific probe and a chromosome 17 centromere probe, and determining the ratio of HER2 genomic DNA to chromosome 17 centromere DNA (e.g., the ratio of HER2 gene copy number to chromosome 17 centromere copy number). The method involves using different detectable labels and / or detection systems for each of the HER2 genomic DNA and chromosome 17 centromere DNA, allowing each to be detected individually and visibly in a single sample.

[0042] As used in this article, the term "EGFR" refers to the epidermal growth factor receptor, a member of the ErbB receptor family, which comprises four closely related subfamilies of receptor tyrosine kinases: EGFR (ErbB-1), HER2 / neu (ErbB-2), Her 3 (ErbB-3), and Her 4 (ErbB-4).

[0043] As used herein, the term "image data" encompasses raw image data acquired from biological tissue samples, such as through optical sensors or sensor arrays, or preprocessed image data. In particular, the image data may include a pixel matrix.

[0044] As used herein, the terms “image,” “image scan,” or “scanned image” encompass raw image data acquired from biological tissue samples, such as by an optical sensor or sensor array, or preprocessed image data. In particular, the image data may include a pixel matrix.

[0045] As used herein, the term “multichannel image” or “multipath image” encompasses digital images obtained from biological tissue samples in which different biological structures, such as cell nuclei and tissue structures, are simultaneously stained using specific fluorescent dyes, quantum dots, chromosomes, etc., each of which fluoresces or can otherwise be detected in different spectral bands, thus constituting one of the channels in a multichannel image.

[0046] As used in this article, the term "probe" or "oligonucleotide probe" refers to a nucleic acid molecule used to detect complementary nucleic acid target genes.

[0047] As used herein, the term "slide" refers to any suitable size substrate on which a biological specimen can be placed for analysis (e.g., a substrate made wholly or partially of glass, quartz, plastic, silicon, etc.), and more particularly to "microscope slides" such as standard 3 x 1-inch microscope slides or standard 75 mm x 25 mm microscope slides. Examples of biological specimens that can be placed on a slide include, but are not limited to, cytological smears, thin tissue sections (e.g., from biopsies), and biological specimen arrays, such as tissue arrays, cell arrays, DNA arrays, RNA arrays, protein arrays, or any combination thereof. Thus, in one embodiment, tissue sections, DNA samples, RNA samples, and / or proteins are placed at specific locations on a slide. In some embodiments, the term "slide" may refer to SELDI and MALDI chips, as well as silicon wafers.

[0048] As used herein, the term "specific binding entity" refers to a member of a specific binding pair. A specific binding pair is a pair of molecules characterized by binding to each other to substantially exclude binding to other molecules (e.g., the binding constant of a specific binding pair may be at least 10^3 M⁻¹, 10^4 M⁻¹, or 10^5 M⁻¹ greater than the binding constant of either of the two members of a binding pair of other molecules in a biological sample). Specific examples of specific binding moieties include specific binding proteins (e.g., antibodies, lectins, streptavidin, and avidin proteins such as protein A). Specific binding moieties may also include molecules (or portions thereof) specifically bound by such specific binding proteins.

[0049] As used herein, the terms “staining agent,” “staining,” or similar terms generally refer to any treatment of a biological specimen that detects and / or distinguishes the presence, location, and / or amount (e.g., concentration) of a specific molecule (e.g., lipid, protein, or nucleic acid) or a specific structure (e.g., normal or malignant cell, cytoplasm, nucleus, Golgi apparatus, or cytoskeleton) in the biological specimen. For example, staining can compare a specific molecule or cellular structure of a biological specimen to surrounding parts, and the intensity of the stain can determine the amount of the specific molecule in the specimen. Staining can be used not only with bright-field microscopy but also with other observational tools such as phase-contrast microscopy, electron microscopy, and fluorescence microscopy to aid in the observation of molecules, cellular structures, and organisms. Some systematic staining can make the outlines of cells clearly visible. Other staining by said systems may depend on specific cellular components (e.g., molecules or structures) that are stained and do not stain other cellular components or stain relatively little of other cellular components. Examples of various staining methods by said systems include, but are not limited to, histochemical methods, immunohistochemical methods, and other methods based on intermolecular reactions (including non-covalent binding interactions) such as hybridization reactions between nucleic acid molecules. Specific staining methods include, but are not limited to, primary staining methods (such as H&E staining, cervical staining, etc.), enzyme-linked immunohistochemistry, and in situ RNA and DNA hybridization methods, such as fluorescence in situ hybridization (FISH).

[0050] As used herein, the term "target" refers to any molecule whose presence, location, and / or concentration can be determined. Examples of target molecules include proteins, nucleic acid sequences, and haptens, such as haptens that covalently bind to proteins. Typically, target molecules are detected using a conjugate of one or more specific binding molecules and a detectable label.

[0051] Overview

[0052] This disclosure provides systems and methods for detecting intracellular genetic aberrations (e.g., high copy number or chromosomal abnormalities), such as those selected automatically for assessment. In some embodiments, the automatically selected cells for assessment are located within a tumor tissue region, said tumor tissue region comprising cells that meet predetermined protein biomarker staining criteria. Automated cell selection for assessment reduces or eliminates subjectivity. It is believed that the disclosed systems and methods can improve patient treatment outcomes and treatment regimen selection because they utilize more data compared to manual analysis methods.

[0053] While the examples described herein may refer to specific tissues and / or the application of specific stains or detection probes to detect specific biomarkers (and thus diseases), those skilled in the art will recognize that different tissues and different stains / detection probes can be applied to detect different biomarkers and different diseases. For example, although some examples may refer to quantifying signals corresponding to HER2 and Chr17 nucleic acid biomarkers, the systems and methods described herein can be applied to detect and quantify signals from a single nucleic acid probe, any combination of two or more nucleic acid probes, etc. In fact, the systems and methods described herein are applicable to determining gene copy numbers or chromosomal aberrations using any ISH assay or dual ISH assay (including those that utilize chromosomes or fluorophores as markers or any combination thereof).

[0054] In the context of HER2 status in breast and / or gastric cancer, HER2-targeted therapies have achieved good clinical results, highlighting the necessity of accurate HER2 status assessment. Trastuzumab and lapatinib have relatively specific effects on HER2-overexpressing cancer cells, resulting in good patient tolerance and minimal side effects. Therefore, determining the HER2 status of breast or gastric cancer is a crucial step in deciding on the treatment plan.

[0055] The HER2 protein is expressed in the cell membranes of both normal and neoplastic breast tissues in humans. The human HER2 gene, located on chromosome 17, encodes the HER2 protein. Overexpression of the HER2 protein, amplification of the HER2 gene, or both occur in approximately 15% to 25% of breast cancers, and this is considered to be associated with aggressive tumor behavior. Breast cancer cells can have up to 25 to 50 copies of the HER2 gene, and the HER2 protein can increase by up to 40 to 100 times, resulting in the expression of up to 2 million receptors on the surface of tumor cells. Therefore, the difference in HER2 expression between normal tissue and tumor helps to define HER2 as an ideal therapeutic target.

[0056] Traditionally, the pattern and staining intensity of tissue sections are examined, including to determine the integrity of cell membrane staining (see [reference]). Figure 12 Against the background of breast tissue stained for the presence of the HER2 protein, staining that completely surrounds the cell membrane is scored as "2+" or "3+". Partial or incomplete staining of the cell membrane is scored as "1+". The most difficult areas to interpret are those falling on the boundary between intensity levels "1+" and "2+", or those where different expression levels are intertwined. In these cases, alternative tests using ISH (such as HER2 double ISH) can be used for further interpretation.

[0057] Visualization is achieved through dual ISH staining, where HER2 appears as a black discrete signal (SISH), and Chr17 appears as a red signal (Red ISH) in the nuclei of normal cells and cancer cells. This strategy determines the status of the HER2 gene based on its chromosomal position. The copy number of both probes is counted in the tumor cell nucleus, and the HER2 / Chr17 ratio is reported as the count result, thus determining the HER2 amplification status (HER2 / Chr17 ratio ≥ 2.0 indicates amplification, and a ratio < 2.0 indicates non-amplification).

[0058] In manual processing, pathologists visually examine double-ISH tissue sections under a microscope, visually selecting a group of twenty to forty tumor cells, recording the count of dual probes (red and silver probes / dots) in each cell, calculating the ratio of the sum of silver / black dots to the sum of red dots, and comparing the calculated ratio to a threshold (=2.0) to classify the patient's tissue section as double-ISH positive or double-ISH negative. In the algorithmic workflow, double-ISH tissue sections are digitized using a digital microscope or a whole-slide scanner (as described herein), and pathologists review the digital images on software applications that allow viewing the entire slide (e.g., Virtuoso software from Ventana Medical Systems, Inc., Tucson, AZ), manually selecting image regions for analysis. Pathologists digitize twenty or forty cells to calculate a slide score. Image analysis algorithms that automatically detect and output the number of red and black / silver dots in each cell are used to analyze the labeled cells. The same scoring criteria as manual methods are applied to output slide scores and the double ISH positive / negative status of the slides for further review and approval by pathologists. These processes are considered time-consuming. This disclosure provides a faster and more efficient workflow for determining the double ISH positive / negative status in some embodiments. Compared to manual processes, this workflow can analyze more cells and / or other structures, enhance analysis, and improve patient treatment and outcomes.

[0059] In view of the foregoing, in some embodiments, this disclosure provides systems and methods for the clinical interpretation of dual ISH slides, wherein (i) cells to be scored are automatically selected, thereby reducing the subjectivity of manual cell selection; (ii) an increased number of cells can be considered during scoring compared to conventional methods, i.e., more than 20 cells can be analyzed; and (iii) relevant feedback (such as visualization) can be provided to pathologists for more stable (and faster) analysis. Ultimately, the systems and methods of this disclosure can enhance patient care and improve patient treatment outcomes.

[0060] At least some embodiments of this disclosure relate to digital pathology systems and methods for analyzing image data captured from biological samples (including tissue samples), the samples being stained with one or more primary staining agents (such as hematoxylin and eosin (H&E)) and one or more detection probes (such as probes containing specific binding entities that help to label targets within the sample). Figure 1 A digital pathology system 200 for imaging and analyzing specimens according to some embodiments is illustrated. In some embodiments, the digital pathology system includes, for example, a digital data processing device (such as a computer, including an interface for receiving image data from a slide scanner, camera, network, and / or storage medium). In other embodiments, the digital pathology system 200 may include an imaging device 12 (such as a device having a microscope slide assembly for scanning specimens) and a computer 14, wherein the imaging device 12 and the computer may be communicatively coupled together (e.g., directly or indirectly via a network 20). The computer system 14 may include a desktop computer, laptop computer, tablet computer, or the like, digital electronic circuitry, firmware, hardware, memory, computer storage medium, computer programs or instruction sets (such as programs stored in the memory or storage medium), one or more processors (including programming processors), and any other hardware, software, or firmware modules or combinations thereof. For example, the... Figure 1 The computing system 14 shown herein may include a computer having a display device 16 and a casing 18. The computer may store digital images in binary form (stored locally, such as in memory, on a server, or on another network-connected device). The digital images may also be divided into a pixel matrix. A pixel may include a digital value of one or more bits defined by a bit depth. Those skilled in the art will recognize that other computer devices or systems can be utilized, and the computer system described herein may be communicatively coupled to additional components (such as specimen analyzers, microscopes, other imaging systems, automated slide preparation equipment, etc.). This document will further describe some of these additional components, as well as various available computers, networks, etc.

[0061] Generally, the imaging device 12 (or other image sources including pre-scanned images stored in memory or one or more memories) may include, but is not limited to, one or more image capturing devices. Image capturing devices may include, but are not limited to, cameras (such as analog cameras, digital cameras, etc.), optics (such as one or more lenses, sensor focusing lens groups, microscope objectives, etc.), imaging sensors (such as charge-coupled devices (CCDs), complementary metal-oxide-semiconductor (CMOS) image sensors, etc.), photographic film, etc. In digital embodiments, the image capturing device may include multiple lenses that cooperate to demonstrate instantaneous focusing capability. The image sensor, for example, a CCD sensor, may capture digital images of the specimen. In some embodiments, the imaging device 12 is a bright-field imaging system, a multispectral imaging (MSI) system, or a fluorescence microscopy system. The digitized tissue data may be generated, for example, by an image scanning system, such as the VENTANAMEDICAL SYSTEMS, Inc. (Tucson, Arizona) VENTANA iScan HT scanner, or other suitable imaging devices. Other imaging devices and systems will be further described herein. Those skilled in the art will recognize that the digital color images acquired by the imaging device 12 can typically consist of basic color pixels. Each color pixel can be encoded on three digital components, each containing the same number of bits, and each component corresponds to a primary color, typically red, green, or blue, also referred to as "RGB" components.

[0062] Figure 2 Various modules utilized within the digital pathology system 200 of this disclosure are outlined. In some embodiments, the digital pathology system 200 employs a computer device or computer-implemented method having one or more processors 220 and at least one memory 201, the at least one memory 201 storing non-transitory computer-readable instructions for execution by one or more processors to cause one or more processors (220) to execute instructions (or stored data) in one or more modules (such as modules 202 to 210).

[0063] Similarly, refer to Figure 2In some embodiments, the system may include: (a) an imaging module 202 adapted to generate image data of stained biological samples, such as a first image stained for the presence of one or more protein biomarkers and a second image stained for the presence of one or more nucleic acid biomarkers; (b) a demixing module 203 for demixing acquired images with more than one staining into single-channel images; (c) a cell detection and classification module 204 for detecting and classifying stained cells, such as cells stained for nuclear or membrane protein biomarkers; (d) a scoring module 205 for assessing staining intensity and / or deriving expression scores; and (e) tissue region identification. Module 206 is used to identify different tissue regions, such as tumor tissue regions; (f) image registration module 207 is used to map regions in a first image to corresponding regions in a second image; (g) point detection module 208 is used to identify signals corresponding to one or more nucleic acid biomarkers; (h) point classification module 209 is used to classify the identified signals as signals corresponding to specific nucleic acid biomarkers; (i) point counting and classification module 210 is used to return the total count of all detected and classified points within a cell or cell nucleus; and (j) visualization module 211 is used to generate overlay images or certain graphs (such as grouped histograms) based on the counted points and any data derived therefrom. Each of these modules will be described in more detail herein.

[0064] Reference Figures 2 to 5 This disclosure provides a computer-implemented system and method for assessing whether cells within an image of a biological sample stained for the presence of one or more biomarkers exhibit genetic aberrations, such as abnormally high gene copy numbers, certain chromosomal abnormalities, etc. In some embodiments, the system operates multiple modules (such as modules 204 and 205) to identify cells in a first image stained for the presence of one or more biomarkers that meet the expression level of a predetermined protein biomarker, such as a predetermined minimum staining intensity level (step 311). For example, in the context of the HER2 protein biomarker, modules 204 and 205 are used to identify those stained cells that meet the minimum membrane staining intensity level. According to this example, and based on an established minimum threshold, cells meeting the minimum membrane staining intensity level are likely to exhibit an abnormal HER2 gene state.

[0065] Subsequently, tissue regions (e.g., tumor tissue regions) encompassing identified cells meeting predetermined protein biomarker expression levels are identified in the first image (step 312; see also step 412) (e.g., using tissue region identification module 206). Continuing the above example, in the context of HER2, epithelial tumor tissue regions encompassing identified cells meeting a minimum membrane staining intensity threshold are derived. It is believed that cells with genetic abnormalities are likely to be found in these identified tumor tissue regions.

[0066] In some embodiments, the identified tissue regions are mapped from the first image to corresponding regions in the second image (step 313; see also step 413) (e.g., using image registration module 207). For example, if the first image is a first continuous slice of tissue and the second image is a second continuous slice of tissue, the image registration technique allows structures, objects, or regions identified in the first image to be identified in the corresponding second image. In the above example, after image registration is completed, those identified epithelial tumor tissue regions are transferred from the first image to the second image. In this way, genetic aberrations in cells belonging to the identified epithelial tumor tissue regions within the second image can be analyzed.

[0067] Next, multiple modules (such as modules 208, 209, and 210) are used to detect and / or quantify signals corresponding to one or more nucleic acid biomarkers within the mapped tissue region (step 314; see also step 414). In some embodiments, the signals corresponding to one or more nucleic acid biomarkers are dots or spots. The detected and / or quantified signals can then be used to assess whether cell nuclei (e.g., tumor cell nuclei) in the mapped tissue region have genetic aberrations, such as high copy numbers or chromosomal abnormalities (step 315; see also step 415). In the context of a HER2 instance, this includes cell nuclei within an identified and registered epithelial tumor tissue region with, for example, normal gene copy numbers and ploidy status (two signals for HER2 and two signals for chromosome 17); HER2 amplification; chromosome 17 polysomy; and / or chromosome 17 polysomy combined with HER2 amplification.

[0068] Subsequently, the assessment performed can be visualized using visualization module 211 (step 315; see also step 415), or the assessment can be stored in database 240 (e.g., assessments of individual cell nuclei and their locations; assessments of entire mapped tissue regions, etc.). For example, an overlay image can be generated and then superimposed onto the entire slice image or any part thereof. In the context of HER2, and by way of example only, cells with a black-to-red dot ratio greater than 2 can be visualized with one color, while cells with a ratio less than or equal to 2 can be visualized with a second, different color. Similarly, grouped histograms of the calculated ratios can be generated for storage or output.

[0069] Those skilled in the art will recognize that, Figure 2 Additional modules or databases not described herein may be incorporated into the workflow. For example, an image preprocessing module may be run to apply certain filters to the acquired images or to identify certain histological and / or morphological structures within the tissue sample. Furthermore, a target region selection module may be used to select specific portions of the image for analysis.

[0070] Image acquisition module

[0071] In some embodiments, refer to Figure 2 The digital pathology system 200 operates the imaging module 202 to capture images or image data (e.g., from scanning device 12) of biological samples containing one or more staining agents (step 310; see also step 410). In some embodiments, the received or acquired images are RGB images or multispectral images (e.g., multi-channel bright-field and / or dark-field images). In some embodiments, the captured images are stored in memory 201.

[0072] Images or image data (used interchangeably herein) may be acquired using the scanning device 12, for example, in real time. In some embodiments, as described herein, the images are acquired from a microscope or other instrument capable of capturing image data of a microscope slide carrying a specimen. In some embodiments, images are acquired using a two-dimensional scanner such as a scanner capable of scanning image patches, or a line scanner such as a VENTANA DP 200 scanner capable of scanning images line by line. Alternatively, the images may be images that have been previously acquired (e.g., scanned) and stored in one or more memories 201 (or images retrieved from a server via network 20).

[0073] In some embodiments, the received input image is the entire slice image. In other embodiments, the received input image is a portion of the entire slice image. In some embodiments, the entire slice image is decomposed into several parts, such as patches, and each part or patch can be analyzed individually (e.g., using...). Figure 2 The modules listed and at least Figure 3 and Figure 4 (The method shown in the figure). After analyzing the portion or patch individually, the data for each portion or patch can be stored separately and / or reported at the entire substrate level.

[0074] The biological sample can be stained by applying one or more staining agents, and the resulting image or image data includes a signal corresponding to each of the one or more staining agents. In some embodiments, the input image is a simplex image with only a single staining agent (e.g., stained with 3,3'-diaminobenzidine (DAB)). In some embodiments, the biological sample can be stained in a multichannel analysis with two or more staining agents (thus providing multichannel images). In some embodiments, the biological sample is stained for at least two biomarkers. In other embodiments, the biological sample is stained for the presence of at least two biomarkers, and the biological sample is also stained with a primary staining agent (e.g., hematoxylin). In some embodiments, the biological sample is stained for the presence of at least one protein biomarker and at least two nucleic acid biomarkers (e.g., DNA, RNA, microRNA, etc.).

[0075] In some embodiments, the biological sample is stained using immunohistochemical assays to detect the presence of one or more protein biomarkers. For example, the biological sample may be stained to detect the presence of human epidermal growth factor receptor 2 protein (HER2 protein). Currently, there are two FDA-approved methods for HER2 assessment in the United States: HerceptTest. TM (DAKO, Glostrup Demark) and HER2 / neu(4B5) rabbit monoclonal primary antibodies (Ventana, Tucson, Arizona).

[0076] In other embodiments, the biological sample is stained for the presence of estrogen receptor (ER), progesterone receptor (PR), or Ki-67. In other embodiments, the biological sample is stained for the presence of EGFR or HER3. Zamay et al., “Current and Prospective Biomarkers of Long Cancer,” Cancers (Basel), November 2018; 9(11), describes examples of other protein biomarkers, the entire contents of which are incorporated herein by reference. Examples of protein biomarkers described by Zamay include CEACAM, CYFRA21-1, PKLK, VEGF, BRAF, and SCC.

[0077] In other embodiments, biological samples are stained for the presence of one or more nucleic acids, including mRNA, in an in situ hybridization (ISH) assay. U.S. Patent No. 7,087,379 (the disclosure of which is incorporated herein by reference in its entirety) describes a method for staining samples with ISH probes to observe and detect single spots (or dots) representing single gene copies. In some embodiments, multiple target genes are analyzed simultaneously by exposing cell or tissue samples to multiple nucleic acid probes that are tagged with multiple different nucleic acid tags.

[0078] For example, the INFORM HER2 dual ISH DNA probe mixture assay from Ventana Medical Systems, Inc. (Tucson, AZ) aims to determine the status of the HER2 gene by calculating the ratio of the HER2 gene to chromosome 17. The HER2 and chromosome 17 probes are detected using two-color chromosomal in situ hybridization on formalin-fixed paraffin-embedded tissue samples, such as human breast cancer tissue specimens or human gastric cancer tissue specimens. For this HER2 dual ISH assay, signals are presented as a silver signal (“black signal”) and a red signal, corresponding to black and red dots in the input image, respectively. For this HER2 dual ISH assay, cell-based scoring involves counting red and black dots within selected cells, where HER2 gene expression is expressed via black dots and chromosome 17 expression via red dots.

[0079] In some embodiments, the biological sample is stained against at least a HER2 protein biomarker and against HER2 and chromosome 17 nucleic acid biomarkers. In some embodiments, the biological sample is stained against at least a HER2 protein biomarker, against HER2 and chromosome 17 nucleic acid biomarkers, and against at least one additional protein biomarker (such as ER, PR, Ki-67, etc.). For example, a first serial tissue section may be stained against the HER2 protein biomarker (and optionally other protein biomarkers), and a second serial tissue section may be stained using a HER2 dual ISH probe mixture (see [link to relevant documentation]). Figure 7A and Figure 7B In some embodiments, the biological sample is stained against at least EGFR protein biomarkers and against EGFR / CEP nucleic acid biomarkers.

[0080] The staining agent may include hematoxylin, eosin, zesso red, or 3,3'-diaminobenzidine (DAB). In some embodiments, the tissue sample is stained with a primary staining agent (such as hematoxylin). In some embodiments, the tissue sample is stained with a secondary staining agent (such as eosin). In some embodiments, the tissue sample is stained for a specific biomarker in an IHC assay. Of course, those skilled in the art will recognize that any biological sample can also be stained with one or more fluorophores.

[0081] Typical biological samples are processed in an automated staining / analysis platform for staining. Several commercially available products suitable for use as staining / analysis platforms are available, such as Discovery from Ventana Medical Systems, Inc. (Tucson, AZ). TM The product is one example. The camera platform may also include a bright-field microscope, such as the Ventana iScan HT or Ventana DP 200 scanner from Ventana Medical Systems, Inc., or any microscope with one or more objectives and a digital imager. Other techniques can be used to capture images at different wavelengths. Furthermore, camera platforms suitable for imaging stained biological specimens are known in the art and are commercially available from companies such as Zeiss, Canon, and Applied Spectral Imaging, and such platforms are readily adaptable for use in the systems, methods, and apparatus disclosed in this subject matter.

[0082] In some embodiments, the input image is masked such that only tissue regions exist in the image. In some embodiments, a tissue region mask is generated to mask non-tissue regions from the tissue regions. In some embodiments, the tissue region mask can be created by identifying the tissue regions and automatically or semi-automatically (i.e., with minimal user input) excluding background regions (such as the entire slice image region corresponding to a piece of glass without a sample, for example, regions containing only white light from the imaging source). Those skilled in the art will recognize that, in addition to masking non-tissue regions from the tissue regions, the tissue masking module can also mask other target regions as needed, such as those identified as belonging to a certain tissue type or as part of a suspected tumor region. In some embodiments, a tissue region masking image is generated by using segmentation techniques to mask the tissue regions from the non-tissue regions in the input image. Similarly, suitable segmentation techniques are known in the prior art (see Digital Image Processing, 3rd Edition, Rafael C. Gonzalez, Richard E. Woods, Chapter 10, p. 689 and Handbook of Medical Imaging, Processing and Analysis, Isaac N. Bankman Academic Press, 2000, Chapter 2). In some embodiments, image segmentation techniques are used to distinguish digitized tissue data and slides in the image, where the tissue corresponds to the foreground and the slides correspond to the background. In some embodiments, the component calculates an AOI (Area of ​​Interest) across the entire slice image to detect all tissue regions within the AOI while limiting the amount of non-tissue regions analyzed in the background. Various image segmentation techniques (such as HSV color-based image segmentation, lab image segmentation, mean-shifted color image segmentation, region growing, level set methods, fast-progression methods, etc.) can be used to determine, for example, the boundaries between tissue data and non-tissue or background data. Based on at least a partial segmentation technique, the component can also generate a tissue foreground mask that can be used to identify portions of the digitized slide data that correspond to the tissue data. Alternatively, the component can generate a background mask that identifies portions of the digitized slide data that do not correspond to the tissue data.

[0083] This recognition can be achieved through image analysis operations (such as edge detection). Tissue region masks can be used to remove non-tissue background noise from images (e.g., non-tissue regions). In some embodiments, the generation of the tissue region mask includes one or more of the following operations (but is not limited to): calculating the brightness of a low-resolution analytical input image, generating a brightness image, applying a standard deviation filter to the brightness image, generating a filtered brightness image, and applying a threshold to the filtered brightness image to set pixels with brightness above a given threshold to 1 and pixels with brightness below the threshold to 0, and generating the tissue region mask. Further information and examples related to the generation of tissue region masks are disclosed in U.S. Publication No. 2017 / 0154420, entitled “An Image Processing Method and System for Analyzing a Multi-Channel ImageObtained from a Biological Tissue Sample Being Stained by Multiple Stains,” the disclosure of which is incorporated herein by reference in its entirety.

[0084] Demixing module

[0085] In some embodiments, the received input image may be a multi-channel image, i.e., an image of a biological sample stained with more than one staining agent (e.g., an image stained for the presence of HER2 and chromosome 17 probes; an image stained for the presence of protein biomarkers or nucleic acid biomarkers). In these embodiments, the multi-channel image is first demixed into its constituent channels, for example using demixing module 203, before further processing, where each demixed channel corresponds to a specific staining agent or signal.

[0086] In some embodiments, in a sample containing one or more staining agents, a single image can be generated for each channel containing one or more staining agents. Those skilled in the art will recognize that features extracted from these channels can be used to describe different biological structures (such as cell nuclei, membranes, cytoplasm, nucleic acids, etc.) present in any image of a tissue.

[0087] For example, in the context of the HER2 dual ISH probe described herein, demixing will generate a first demixed image channel image with a silver (or black) signal (corresponding to black dots), a second demixed image channel image with a red signal (corresponding to red dots), and a third demixed image channel image with a hematoxylin signal. Each of these demixed black and red dot images can be used as input to the dot detection and classification module described herein. Similarly, as another non-limiting example, an input image stained for the presence of two protein biomarkers (such as HER2 and Ki-67) will be demixed into a first image channel image (such as DAB membrane staining) with a signal corresponding to HER2 and a second image channel image with a signal corresponding to Ki-67. Likewise, the HER2 protein biomarker image and the Ki-67 protein biomarker image can be used as input images for cell detection and classification described herein.

[0088] In some embodiments, the multispectral image provided by the imaging module 202 is a weighted mixture of underlying spectral signals associated with individual biomarkers and noise components. At any given pixel, the mixing weight is proportional to the marker expression of the underlying isolocalized biomarker at that particular location in the tissue and the background noise at that location. Therefore, the mixing weights differ between different pixels. The spectral demixing method disclosed herein decomposes the multichannel pixel value vector at each pixel into a set of constituent biomarker endmembers or components and estimates the proportion of individual constituent staining agents for each biomarker.

[0089] Unmixing refers to the procedure of decomposing the measured spectrum of a mixed pixel into a set of component spectra or endmembers representing the proportion of each endmember in the pixel, and a set of corresponding fractions or abundances. Specifically, the unmixing process can extract staining-specific channels, allowing the determination of the local concentration of individual stains using reference spectra known for standard types of tissue and staining combinations. The unmixing can use reference spectra retrieved from control images or estimated from observed images. Unmixing the component signals of each input pixel allows for the retrieval and analysis of staining-specific channels, such as the hematoxylin and eosin channels in H&E images, or the diaminobenzidine (DAB) and counterstain (e.g., hematoxylin) channels in IHC images. The terms "unmixing" and "color deconvolution" (or "deconvolution") or similar terms (e.g., "deconvolution," "unmixing") are used interchangeably in the art.

[0090] In some embodiments, the multiplexed image and unmixing module 205 unmixes the images in a linear unmixing manner. For example, linear unmixing is described in, for instance, in Zimmermann's "Spectral Imaging and Linear Unmixing in Light Microscopy" (Adv Biochem Engin / Biotechnology (2005) 95:245-265) and in C.L. Lawson and R.J. Hanson's "Solving Least Squares Problems" (Prentice Hall, 1974, Chapter 23, p. 161), the contents of which are incorporated herein by reference in their entirety. In linear staining unmixing, the measured spectrum (S(λ)) at any pixel is considered a linear mixture of the staining spectral components and is equal to the sum of the proportions or weights (A) of the color reference (R(λ)) for each individual staining agent represented at that pixel.

[0091] S(λ)=A1·R1(λ)+A2·R2(λ)+A3·R3(λ).......A i ry(λ)

[0092] More generally, it can be represented in matrix form as

[0093] S(λ)=ΣA i ry(λ) or S=R·A

[0094] For example, if there are M acquired channel images and N individual dyes, then the columns of the M x N matrix R are the optimal colorimetric system derived in this paper, the N x 1 vector A is the unknown of the proportion of the individual dyes, and the M x 1 vector S is the multi-channel spectral vector measured at the pixel. In these equations, the signal in each pixel (S) is measured during the acquisition of multiple images, and the reference spectrum, i.e., the optimal colorimetric system, as described in this paper, is derived. By calculating various dyes (A... i The contribution of each point in the measured spectrum is used to determine their contribution. In some embodiments, an inverse least squares fitting method is used to solve this problem, which minimizes the squared difference between the measured and calculated spectra by solving the following set of equations.

[0095]

[0096] In this equation, j represents the number of detection channels, and i equals the amount of dye. The solution to the linear equation typically allows for constrained solution mixing, forcing the weights (A) to be added together.

[0097] In other embodiments, unmixing is performed using the method described in WO2014 / 195193, filed May 28, 2014, entitled “Image Adaptive Physiologically Plausible Color Separation,” the disclosure of which is incorporated herein by reference in its entirety. Typically, WO2014 / 195193 describes an unmixing method that separates component signals of an input image using iteratively optimized reference vectors. In some embodiments, image data in a measurement is correlated with expected or desired results specific to the measurement to determine a quality metric. In cases of low image quality or poor correlation with the desired result, one or more reference column vectors in matrix R are adjusted, and unmixing is iteratively repeated using the adjusted reference vectors until the correlation shows a high-quality image that meets physiological and anatomical requirements. The anatomical, physiological, and measurement information can be used to define rules for applying image data to the measurement to determine a quality metric. This information includes how tissues are stained, which structures within the tissue are intended or not intended to be stained, and the relationships between structures, staining agents, and markers specific to the measurement being processed. The iterative process generates staining-specific vectors that can produce images that accurately identify target structures and biologically relevant information, without any noise or unwanted spectra, making the process suitable for analysis. The reference vectors are adjusted within a search space. The search space defines the range of values ​​for which the reference vectors can represent the staining agent. The search space is determined by scanning a variety of representative training assays, including known or commonly occurring problems, and identifying a high-quality set of reference vectors for the training assays.

[0098] Protein biomarker detection

[0099] After acquiring one or more images from, for example, sequential tissue sections using the imaging module 202 (and optionally unmixing using the demixing module 203) (step 310; see also step 410), cells meeting predetermined criteria, such as threshold protein biomarker expression levels, are identified in the acquired images (or unmixed image channel images) (step 311; see also step 411). In some embodiments, the cell detection and classification module 204 can be used to detect and classify cells stained for the presence of protein biomarkers. After detection and classification, staining intensity levels or expression levels can be derived, for example, using the scoring module 205. Tissue regions, such as tumor tissue regions, are then identified using the identified cells that meet the predetermined criteria.

[0100] For example, and in the context of a sample stained for the presence of the HER2 protein biomarker, cell membranes expressing the HER2 protein can be detected, classified, and / or scored (see, for example). Figure 8B , Figure 9 A and Figure 9 B). In some embodiments, each detected and classified cell can then be evaluated to determine the staining intensity level. In other embodiments, membrane staining can be evaluated and scored using an automated scoring algorithm (e.g., 0, +1, +2, or +3). For example, those detected cells that meet a minimum threshold staining intensity or score can be identified, and this can be used to determine tumor tissue regions (see Figure 8 and...). Figure 9 ).

[0101] Automated cell detection and classification module

[0102] After image acquisition and / or demixing, the input image or demixed image channel is provided to the cell detection and classification module 204 to automatically detect, identify, and / or classify cells and / or cell nuclei. The program and automatic algorithm described herein are adaptable to identify and classify various types of cells or cell nuclei based on features within the input image, including identifying and classifying tumor cells, non-tumor cells, stromal cells, and lymphocytes.

[0103] Those skilled in the art will recognize that the cell nucleus, cytoplasm, and cell membrane possess different characteristics, and that different stained tissue samples can exhibit different biological features. In fact, those skilled in the art will recognize that certain cell surface receptors can have staining patterns localized to the cell membrane or cytoplasm. Therefore, "cell membrane" staining patterns are analytically different from "cytoplasmic" staining patterns. Similarly, "cytoplasmic" staining patterns are analytically different from "nuclear" staining patterns. Each of these different staining patterns can be used as a characteristic for identifying cells and / or the cell nucleus.

[0104] U.S. Patent No. 7,760,927 describes a method for identifying, classifying, and / or scoring cell nuclei, cell membranes, and cytoplasm in an image of a biological sample having one or more staining agents, the disclosure of which is incorporated herein by reference in its entirety. For example, US 7,760,927 describes an automated method for simultaneously identifying multiple pixels in an input image of biological tissue stained with biomarkers, comprising considering a first color plane of multiple pixels in the foreground of the input image to simultaneously identify cytoplasmic and cell membrane pixels, wherein the input image has been processed to remove its background portions and counterstaining components; determining a threshold level between cytoplasmic and cell membrane pixels in the foreground of the digital image; and using the determined threshold level to simultaneously determine whether a selected pixel in the foreground and its eight neighboring pixels is a cytoplasmic pixel, a cell membrane pixel, or a transition pixel in the digital image.

[0105] U.S. Patent Publication No. 2017 / 0103521 also describes suitable systems and methods for automatically identifying biomarker-positive cells in images of biological samples, the disclosure of which is incorporated herein by reference in its entirety. For example, US2017 / 0103521 describes (i) reading a first digital image and a second digital image into one or more memories, the first and second digital images depicting the same region of a first slide comprising a plurality of tumor cells stained with a first staining agent and a second staining agent; (ii) identifying a plurality of cell nuclei and their location information by analyzing the light intensity in the first digital image; (iii) identifying a cell membrane containing the biomarker by analyzing the light intensity in the second digital image and analyzing the location information of the identified cell nuclei; and (iv) identifying a biomarker-positive tumor cell in the region, wherein a biomarker-positive tumor cell is a combination of an identified cell nucleus and an identified cell membrane surrounding the identified cell nucleus. US2017 / 0103521 also discloses methods for detecting staining using HER2 protein biomarkers or EGFR protein biomarkers.

[0106] In some embodiments, tumor cell nuclei are automatically identified by first identifying candidate cell nuclei and then automatically distinguishing tumor cell nuclei from non-tumor cell nuclei. Various methods for identifying candidate cell nuclei in tissue images are known in the prior art. For example, candidate cell nuclei are automatically detected by applying a radial symmetry-based method, such as detection on the unmixed hematoxylin image channel or biomarker image channel (see Parvin, Bahram et al., “Iterative voting for inference of structural saliency and characterization of subcellular events”, Image Processing, IEEE Transactions on 16.3 (2007): 615-623, the disclosure of which is incorporated herein by reference in its entirety).

[0107] More specifically, in some embodiments, the received image as input is processed, for example, to detect the center of the cell nucleus (seed) and / or segment the cell nucleus. For example, instructions may be provided to detect the center of the cell nucleus based on radial symmetry voting using Parvin's technique (as described above). In some embodiments, radial symmetry is used to detect the center of the cell nucleus, and then the cell nucleus is classified based on the staining intensity around the cell center. In some embodiments, the radial symmetry-based cell nucleus detection operation is performed as described in commonly assigned and commonly pending patent application WO / 2014 / 140085A1, which is incorporated herein by reference in its entirety. For example, the image size can be calculated within the image, and one or more votes at each pixel can be accumulated by summing the sizes within a selected region. Mean-shift clustering can be used to find the local center of said region, which represents the actual location of the cell nucleus. Radial symmetry-based cell nucleus detection can be performed on color image intensity data and explicitly uses prior knowledge that cell nuclei are elliptical spots of varying sizes and eccentricities. To accomplish the above operations, in addition to the color intensity in the input image, image gradient information is used for radial symmetry voting and combined with an adaptive segmentation process to accurately detect and locate cell nuclei. For example, the “gradient” used herein refers to the intensity gradient of a particular pixel calculated considering the intensity value gradients of a set of pixels surrounding that particular pixel. Each gradient can have a specific “direction” relative to a coordinate system whose x and y axes are defined by two orthogonal edges of the digital image. For example, the detection of cell nucleus seeds involves defining the seed as a point assumed to be located within the cell nucleus and serving as the starting point for locating the cell nucleus. The first step is to detect the seed point associated with each cell nucleus using a very stable radial symmetry-based method, thereby detecting elliptical blobular structures resembling cell nuclei. In the radial symmetry method, the gradient image can be processed using a kernel-based voting procedure. Each pixel that accumulates votes through a voting kernel is processed, thereby creating a voting response matrix. The kernel is based on the gradient direction calculated at that particular pixel, the expected range of minimum and maximum cell nucleus sizes, and the voting kernel angle (typically in the range [π / 4, π / 8]). In the resulting voting space, local maxima locations with voting values ​​above a predetermined threshold are saved as seed points. Unrelated seeds are discarded during subsequent segmentation or classification. Other methods are discussed in U.S. Patent Publication No. 2017 / 0140246, the disclosure of which is incorporated herein by reference in its entirety.

[0108] Cell nuclei can be identified using other techniques known to those skilled in the art. For example, the image size can be calculated from a specific image channel of one of the H&E or IHC images, and a number of votes based on the sum of the sizes of the regions surrounding each specified size can be assigned to the pixels. Alternatively, mean-shift clustering can be performed to locate local centers within the voting image that represent the actual location of the cell nucleus. In other embodiments, cell nucleus segmentation can be used to segment the entire cell nucleus based on currently known cell nucleus centers, through morphological operations and local thresholding. In other embodiments, model-based segmentation can be used to detect cell nuclei (i.e., learning a shape model of the cell nucleus from a training dataset and using it as prior knowledge to segment cell nuclei in the test image).

[0109] In some embodiments, the cell nuclei are then segmented using a threshold calculated individually for each nucleus. For example, since pixel intensities in the nucleus region are believed to vary, Otsu's method can be used to perform segmentation operations in regions surrounding identified nuclei. As will be appreciated by those skilled in the art, Otsu's method is used to determine the optimal threshold by minimizing intra-class variance, and this method is known to those skilled in the art. More specifically, Otsu's method is used to automatically perform cluster-based image thresholding, or to restore a grayscale image to a binary image. The algorithm assumes that the image contains two classes of pixels (foreground pixels and background pixels) that follow a bimodal histogram. The optimal threshold separating the two classes of pixels is then calculated, thus achieving minimum or equal combinatorial diffusion (intra-class variance) (because the sum of pairwise squared distances is constant), thereby maximizing their inter-class variance.

[0110] In some embodiments, the system and method further include automatically analyzing the spectral and / or shape features of identified cell nuclei in an image to identify cell nuclei of non-tumor cells. For example, spots can be identified in a first digital image in a first step. As used herein, a “spot” can be, for example, a region of a digital image where some properties (such as intensity or grayscale value) remain constant or vary within a specified numerical range. In a sense, all pixels in a spot can be considered similar to each other. For example, spots can be identified using differential methods based on the derivative of a position function on a digital image and methods based on local extrema. A nuclear spot is a pixel and / or contour shape that suggests it may be generated by a cell nucleus stained with a first staining agent. For example, the radial symmetry of a spot can be evaluated to determine whether the spot should be identified as a nuclear spot or any other structure, such as a staining artifact. For example, if a spot is elongated and lacks radial symmetry, it may not be identified as a nuclear spot but rather as a staining artifact. According to embodiments, spots identified as “nuclear spots” can represent a set of pixels identified as candidate cell nuclei and can be further analyzed to determine whether the nuclear spots represent pixels of a cell nucleus. In some embodiments, any type of nuclear spot is directly used as the “identified nucleus.” In some embodiments, the identified nuclei or nuclear spots are filtered to identify nuclei that do not belong to tumor cells that are positive for the biomarker, and the identified non-tumor nuclei are removed from the list of identified nuclei, or the nuclei are not added to the list of identified nuclei in the first place. For example, additional spectral and / or shape characteristics of the identified nuclear spots can be analyzed to determine whether the nuclei or nuclear spots are tumor cell nuclei. For example, lymphocyte nuclei are larger than the nuclei of other tissue cells (such as lung cells). In cases where the tumor cells are derived from lung tissue, lymphocyte nuclei are identified by recognizing all nuclear spots whose minimum size or diameter is significantly larger than the average size or diameter of normal lung cell nuclei. Identified nuclear spots associated with lymphocyte nuclei can be removed from the set of identified nuclei (i.e., “filtered”). Filtering non-tumor cell nuclei can improve the accuracy of the method. Since non-tumor cells can also express the biomarker to some extent, an intensity signal not originating from tumor cells can be generated in the first digital image. The accuracy of identifying biomarker-positive tumor cells can be improved by identifying and filtering cell nuclei that do not belong to tumor cells from the total number of identified cell nuclei. These and other methods are described in U.S. Patent Publication 2017 / 0103521, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, once a seed is detected, a local adaptive thresholding method can be used to create a blob around the center of the detection.In some embodiments, other methods may also be introduced, such as a marker-based watershed algorithm, to identify nuclear spots around the detected nucleus center. These and other methods are described in PCT Publication WO2016 / 120442, the disclosure of which is incorporated herein by reference in its entirety.

[0111] After the cell nucleus is detected, features (or metrics) are derived from the input image. Deriving metrics from cell nucleus features is well-known in the art, and any known cell nucleus features can be used within the context of this disclosure. Non-limiting examples of computable metrics include:

[0112] (A) Measurements derived from morphological features

[0113] For example, the term "morphological feature" as used herein refers to features that indicate the shape or size of a cell nucleus. Without wishing to be bound by any particular theory, morphological features are believed to provide important information about the size and shape of a cell or its nucleus. For example, morphological features can be calculated by applying various image analysis algorithms to pixels contained in or surrounding a nucleus spot or seed. In some embodiments, the morphological features include area, minor and major axis lengths, perimeter, radius, volume, etc. At the cellular level, such features are used to classify cell nuclei as healthy or diseased cells. At the tissue level, these statistical features are fully utilized at the tissue level to classify tissues as diseased or non-diseased tissues.

[0114] (B) Measurements derived from color

[0115] In some embodiments, the measure derived from color includes the color ratio, R / (R+G+B), or the principal components of the color. In other embodiments, the measure derived from color includes local statistics (mean / median / variance / standard deviation) and / or color intensity correlation for each color within a local image window.

[0116] (C) Measures derived from intensity characteristics

[0117] In histopathological slide images, adjacent cell groups with certain specific attribute values ​​are set between the black and white shading of gray cells. Since the correlation of these color features defines an example of size gradation, the intensity of these colored cells can be used to determine the affected cells from the surrounding clusters of dark cells.

[0118] (D) Measures derived from spatial features

[0119] In some embodiments, spatial features include the local density of cells; the average distance between two adjacent detected cells; and / or the distance from the cell to the segmented region.

[0120] Of course, other features known to those skilled in the art can also be considered and used as the basis for feature calculation.

[0121] In some embodiments, the cell detection and classification module 204 is run more than once. For example, the cell detection and classification module 204 runs for the first time to extract features and classify cells and / or nuclei in a first image; then, it runs a second time to extract features and classify cells and / or nuclei in a series of additional images, wherein the additional images may be other simplex images or unmixed image channel images, or any combination thereof.

[0122] After the features are derived, they can be used alone or in conjunction with training data (e.g., during training, presenting example cells along with a basic ground truth identification provided by an expert observer according to procedures known to those skilled in the art) to classify cell nuclei or cells. In some embodiments, the system may include a classifier trained at least in part on a set of training or reference slides for each biomarker. Those skilled in the art will recognize that different sets of slides can be used to train a classifier for each biomarker. Accordingly, for a single biomarker, a single classifier is obtained after training. Those skilled in the art will also recognize that, due to the variability between image data obtained from different biomarkers, different classifiers can be trained for each different biomarker, thereby ensuring better performance on unseen test data, wherein the biomarker-type test data is known. The classifier trained for the slide description may be selected at least in part based on how best to handle the variability of the training data, such as tissue type, staining protocol, and other target features.

[0123] Rating module

[0124] In some embodiments, the scoring module 205 utilizes data acquired during cell detection and classification. For example, as described herein, the cell detection and classification module 204 may include a series of image analysis algorithms and may be used to determine the presence of one or more cell nuclei, cell walls, tumor cells, or other structures within the identified cell clusters. In some embodiments, derived staining intensity values ​​and the count of specific cell nuclei for each field of view may be used to determine various marker expression scores, such as a positive percentage or an H-score. Suitable scoring methods are described in U.S. Patent Publication No. 2017 / 0103521, the disclosure of which is incorporated herein by reference in its entirety.

[0125] For example, the automated image analysis algorithm in the cell detection and classification module 204 can be used to interpret each IHC slide in the series to detect tumor cell nuclei that are positive and negative for specific biomarkers (e.g., Ki67, ER, PR, HER2, etc.). Based on the detected positive and negative tumor cell nuclei, various slide-level scores, such as the percentage of marker positivity, H-Score, etc., can be calculated using one or more methods.

[0126] In some embodiments, the expression score is an H-Score. In some embodiments, the "H" score is used to assess the percentage of tumor cells with a cell membrane staining grade of "weak," "moderate," or "strong." The grades are summed, with a maximum total score of 300 points, and a cutoff of 100 points distinguishing between "positive" and "negative." For example, the membrane staining intensity (0, 1+, 2+, or 3+) of each cell in a fixed field of view (or, in this case, each cell in a tumor or cell cluster) is determined. The H-score can be simply based on a dominant staining intensity, or more complexly, it can include the sum of individual H-scores for each intensity level seen. By one method, the percentage of cells at each staining intensity level is calculated, and finally, an H-score is assigned using the following formula: [1 x (% cells 1+) + 2 x (% cells 2+) + 3 x (% cells 3+)]. The final score, ranging from 0 to approximately 300, provides more relative weight to higher intensity membrane staining in a given tumor sample. The sample can then be considered positive or negative based on a specific discrimination threshold. U.S. Patent Publication No. 2015 / 0347702 describes an additional method for calculating the H-score, the disclosure of which is incorporated herein by reference in its entirety.

[0127] In some embodiments, the expression score is the Allred score. The Allred score is a scoring system that displays the percentage of cells that are positive for the hormone receptor test, and the degree to which the receptor is presented after staining (referred to as "intensity"). This information is then used to score the sample on a scale of 0 to 8. It is believed that a higher score indicates more receptors and greater visibility in the sample.

[0128] In other embodiments, the expression score is a positive percentage. Similarly, in the context of scoring breast cancer samples stained for PR and Ki-67 biomarkers, for the PR and Ki-67 slides, the positive percentage is calculated in a single slide (e.g., by summing the total number of cell nuclei of positive cells (e.g., malignant cells) in each field of view of the digital image of the slide after staining, and then dividing by the total number of positive and negative stained cell nuclei in each field of view of the digital image), as follows: Positive percentage = Number of positively stained cells / (Number of positively stained cells + Number of negatively stained cells).

[0129] In other embodiments, the expression score is an immunohistochemical composite score, which is a prognostic score based on several IHC markers, wherein the number of said markers is greater than one. These composite scores are described in U.S. Patent Publication No. 2017 / 0082627, the disclosure of which is incorporated herein by reference in its entirety.

[0130] Organizational region identification module

[0131] After identifying cells that meet the threshold protein biomarker expression level (step 311; see also step 411), the tissue region (e.g., tumor tissue region) covering the identified cells can be derived, for example, by using the tissue region identification module 206 (step 312; see also step 412). For example, after identifying a single cell with minimal HER2 cell membrane staining intensity, the tumor tissue region covering the identified cells can be derived.

[0132] Tissue type identification is performed according to the method described in PCT Publication WO2015 / 113895, entitled "Adaptive Classification for Whole SlideTissue Segmentation," filed January 23, 2015, the entire disclosure of which is incorporated herein by reference. Generally, PCT Publication WO2015 / 113895 describes segmenting the tumor region from other regions in an image through operations related to region classification. These operations include identifying grid points in the tissue image, classifying the grid points into one of several tissue types, and generating classified grid points based on a database of known tissue type features; assigning at least one high-confidence score and one low-confidence score to the classified grid points; modifying the database of known tissue type features based on the grid points assigned high-confidence scores; and generating a modified database; reclassifying the grid points assigned low-confidence scores based on the modified database; and then segmenting the tissue (e.g., identifying tissue regions in the image).

[0133] Alternatively, or in addition, automated image analysis operations, such as segmentation, thresholding, edge detection, etc., and FOV automatically generated based on the detection area, can be used to automatically detect tumor areas or other areas.

[0134] In some embodiments, a tissue region mask can be derived. U.S. Patent Application Publication No. 2017 / 0154420 describes a method for generating such tissue regions, the disclosure of which is incorporated herein by reference in its entirety.

[0135] Automated image registration module

[0136] After identifying tissue regions in a first image, such as an image with signals corresponding to one or more protein biomarkers (step 312; see also step 412), the identified tissue regions are mapped to a second image, such as an image with signals corresponding to one or more nucleic acid biomarkers (step 313; see also step 413). This mapping of identified tissue regions is particularly useful when using sequential tissue sections, such as a first sequential section stained for the presence of one or more protein biomarkers and a second sequential section stained for the presence of one or more nucleic acid biomarkers. In this way, despite differences between sequential tissue sections, the mapping process can identify the corresponding structures, cells, and tissues in each sequential section.

[0137] Generally, registration involves selecting an input image or a portion thereof (such as a cell cluster) as a reference image and calculating the transformation of each other input image to the coordinate system of the reference image. Therefore, image registration can align all input images to the same coordinate system (e.g., the reference coordinates could be a section of a tissue block in the case of consecutive tissue slices or a slice with specific markers). Thus, each image can be aligned from its old coordinate system to the new reference coordinate system.

[0138] Image registration is the process of transforming different datasets (here referring to images or clusters of cells within an image) into a coordinate system. More specifically, registration is the process of aligning two or more images, generally involving designating one image as a reference (also called the reference image or fixed image) and performing geometric transformations on the other images to align them with the reference. The geometric transformation maps a position in one image to a new position in another image. The step of determining the correct geometric transformation parameters is crucial to the image registration process. Methods for calculating the transformation from each image to the reference image are well known to those skilled in the art. For example, an image registration algorithm was described at the 11th International Symposium on Biomedical Imaging (ISBI) (2014 IEEE, April 29 – May 2, 2014), the disclosure of which is incorporated herein by reference in its entirety. The following provides a detailed overview of image registration methods.

[0139] Any registration method can be used in the systems and methods disclosed herein. In some embodiments, image registration is performed using the method described in WO / 2015 / 049233, filed September 30, 2014, entitled “Line-Based Image Registration and Cross-Image Annotation Devices, Systems and Methods,” the disclosure of which is incorporated herein by reference in its entirety. WO / 2015 / 049233 describes a registration process that includes a coarse registration process, used alone or in combination with a fine registration process. In some embodiments, the coarse registration process may include selecting digital images for alignment, generating a foreground image mask from each of the selected digital images, and matching tissue structures between the thus generated foreground images. In other embodiments, generating the foreground image mask includes generating a soft-weighted foreground image from an entire slice image of a stained tissue section and OTSU thresholding the soft-weighted foreground image to produce a binary soft-weighted image mask. In a further embodiment, generating the foreground image mask includes generating a binary soft-weighted image mask from an entire slice image of a stained tissue section, generating gradient magnitude image masks from the same entire slice image, OTSU thresholding the gradient image masks to generate a binary gradient magnitude image mask, and merging the binary soft-weighted image mask and the binary gradient magnitude image mask by a binary OR operation to generate the foreground image mask. As used herein, the term "gradient" refers to the intensity gradient of a particular pixel calculated considering the intensity value gradients of a set of pixels surrounding that particular pixel. Each gradient may have a specific "direction" relative to a coordinate system whose x and y axes are defined by two orthogonal edges of the digital image. A "gradient direction feature" may be a data value indicating the gradient direction within the coordinate system. In some embodiments, matching the tissue structure includes calculating line-based features from the boundaries of each thus generated foreground image mask, calculating global transformation parameters between a first set of line features on a first foreground image mask and a second set of line features on a second foreground image mask, and globally aligning the first and second images based on the transformation parameters. In another embodiment, the coarse registration process includes mapping the selected digital image onto a common grid based on global transformation parameters, the grid potentially encompassing the selected digital image. In some embodiments, the fine registration process may include identifying a first sub-region of a first digital image in a set of aligned digital images; identifying a second sub-region on a second digital image in the set of aligned digital images, wherein the second sub-region is larger than the first sub-region and the first sub-region is substantially located within the second sub-region on the common grid; and calculating an optimized position of the first sub-region within the second sub-region.

[0140] In this article Figure 6 These methods are described herein, wherein method 600 begins at starting block 602. At block 604, a set of image data or digital images (e.g., scanned or selected from a database) are acquired for operation. Each set of image data includes image data corresponding to a set of adjacent tissue slices, for example, from a single patient. At block 606, if only a single image pair is selected, the process proceeds directly to block 610. If more than one pair of images is selected, the selected image set is grouped into pairs at block 608 before proceeding to block 610. In some embodiments, image pairs are selected as adjacent pairs. Thus, for example, if the selected image set comprises 10 parallel, adjacent slices (L1....L10), L1 and L2 are grouped into a pair, L3 and L4 into a pair, and so on. On the other hand, if there is no information about which image pairs are most similar to each other, in some embodiments, they are grouped according to the distance between the images (e.g., the distance between edges or between images corresponding to the chamfer distance between various image edge maps), pairing the images that are closest to each other together. In exemplary embodiments of this disclosure, image pairing is performed using the distance between edges / between images. In some embodiments, the distance between images / edges can be calculated using edge-based chamfer distance. If the image pair has previously undergone a coarse registration process, such that the images are roughly aligned and the result has been saved, the process proceeds to block 614. Otherwise, a coarse registration process is performed on the selected image pair at block 612. The coarse registration process will be described in further detail below.

[0141] Proceeding to block 614, the currently selected registered (aligned) images are displayed on a common grid, overlaid in a single image, and displayed as separate images, or both together on one display or distributed across several displays. At block 616, the client user can select one image from a pair as the source image. If the source image has been labeled as required, the process proceeds to block 622. Otherwise, the client user labels the source image as required at block 620. At block 622, which may or may not occur substantially simultaneously with block 620, the labels are mapped to the other image in the pair (the target image), and the labels are reproduced graphically on the target image. In embodiments where the labeling occurs before coarse registration, the labels can be mapped from the source image to the target image substantially simultaneously with the registration (alignment) of the image pair. At block 624, the user can choose whether to perform a fine registration process. If the user chooses not to perform fine registration and directly displays the results, the process proceeds to block 626.

[0142] Otherwise, at block 624, a precise registration process is performed on the selected image pairs, for example, to optimize the position of the mapped annotations and / or image alignment. The precise registration process will be discussed in further detail below. At block 626, the annotated image pairs are displayed together with the results of the precise registration process (or, if precise registration is not performed, the annotated image pairs may be displayed only with the results of the coarse registration process). The method then finally ends at block 628.

[0143] Automated nucleic acid biomarker detection

[0144] After mapping the identified tissue region from the first image to the second image (step 313; see also step 413), the point detection module 208 and the point classification module 209 can be used to identify signals (or points) corresponding to one or more nucleic acid biomarkers in the mapped tissue region in the second image (step 314; see also step 414). Subsequently, the point counting and classification module 210 can be used to interpret the identified signals to identify and assess gene aberrations in the cell nucleus (step 315; see also step 415).

[0145] In some embodiments, modules 208, 209, and 210 are used to detect signals corresponding to at least one nucleic acid biomarker in each mapped tissue region of the second image. Based on the detection of the signals corresponding to the at least one nucleic acid biomarker, it can be assessed whether cells or nuclei in the mapped tissue region have genetic aberrations (e.g., abnormally high copy numbers; chromosomal abnormalities). In some embodiments, genetic aberrations in the cell nuclei are assessed by determining whether the total number of identified points in each cell nucleus corresponding to one or more signals from at least one nucleic acid biomarker meets a predetermined threshold. For example, in the case of a biological sample stained with a single nucleic acid probe, points corresponding to signals from the single nucleic acid probe can be detected, classified, and then counted. The copy number present in each cell nucleus can then be compared with a predetermined threshold.

[0146] In other embodiments, gene aberrations in the cell nucleus are assessed by calculating the ratio of a first recognition point corresponding to a first nucleic acid biomarker to a second recognition point corresponding to a second nucleic acid biomarker; and comparing the calculated ratio for each cell nucleus with a predetermined threshold. In some embodiments, points in each cell nucleus are identified, for example, points corresponding to each different signal type in each cell nucleus are identified. For example, in the case of staining a biological sample with two nucleic acid probes, points corresponding to signals from each probe (e.g., black or red dots in a background of staining for chromosomes HER2 and 17) can be detected and classified, and a ratio can be calculated based on the total number of points corresponding to each probe, ultimately using said ratio to determine whether gene aberrations exist within the cell nucleus (see [link to relevant documentation]). Figure 10 and Figure 11 ).

[0147] As another example, biological samples can be stained using an EGFR / CEP 7 dual probe, and the ratio of the EGFR gene to chromosome 7 can be calculated and compared with clinically relevant thresholds. The following can be observed through comparison: disomy (score = 1), low trisomy (score = 2), high trisomy (score = 3), low polysomy (score = 4), high polysomy (score = 5), and amplification (score = 6) (see Dahle-Smith, “Epidermal Growth Factor (EGFR) copy number aberrations in esophageal and gastro-esophageal junction carcinoma,” MolCytogenet, 2015; 8:78, the entire contents of which are incorporated herein by reference).

[0148] This article describes and Figure 10 Further non-limiting examples of dot detection, dot classification, and dot counting against the background of HER2 and chromosome 17 are illustrated in the text. Similarly, the systems and methods described herein are not limited to using dual ISH assays for HER2 to determine genetic aberrations; rather, such examples are for illustrative purposes only.

[0149] Point detection module

[0150] Generally, point detection module 208 is used to perform point detection to identify all point pixels in the input image (e.g., see [link]). Figure 11 The detected points are then provided to the point classification module 209 for further processing and analysis.

[0151] In some embodiments, point detection can be performed using one or more different derived image features, including but not limited to absorbance, multi-scale difference of Gaussian (DoG), and features from the unmixed image channels obtained after color deconvolution. Point detection according to the method of this disclosure supports accurate and stable point detection by considering multiple features such as those described above. In some embodiments, methods capable of specifically detecting signals corresponding to black and / or red dots are employed, such as those used in HER2 double ISH assays. Those skilled in the art will also recognize that while some examples herein may refer to biomarkers labeled with chromosomes (such as black and red dots in HER2 double ISH assays), point detection can be performed using biomarkers labeled with fluorophores (such as FISH).

[0152] In some embodiments, point detection is performed using the method described in U.S. Patent Publication No. 2014 / 0377753, the disclosure of which is incorporated herein by reference in its entirety. For example, points may first be identified by converting a color image of the cells to a monochrome image. In one embodiment, a monochrome image is first created by converting the color space of the color image of the cells from the RGB color space to the L*a*b color space. In the L*a*b color space, the “L” channel represents the brightness of the pixel, the “a” channel reflects the red and green components of the pixel, and the “b” channel represents the blue and yellow components of the pixel. A new image is then created that emphasizes the red and black in the image obtained at each pixel location through a linear combination of the “L”, “a”, and “b” values. In some embodiments, points in the red and black enhanced image are detected by passing the enhanced image through several filters.

[0153] In some embodiments, the filters are Difference-of-Gaussian (“DoG”) filters, wherein the size of each filter is selected based on the expected size of the points / clusters to be detected. Generally, Difference-of-Gaussian is a feature enhancement algorithm that involves subtracting a blurred version of the original image from a less blurred version. In the simple case of grayscale images, the blurred image is obtained by convolving the original grayscale image with Gaussian kernels having different standard deviations. It is believed that blurring an image using Gaussian kernels only suppresses high-frequency spatial information. Subtracting one image from another can preserve the spatial information between the frequency ranges retained in the two blurred images. Therefore, Difference-of-Gaussian is equivalent to a bandpass filter that removes all spatial frequencies except those few that are retained in the original grayscale image. In some embodiments, the size of the DoG filters ranges from about 0.05 micrometers to about 5 micrometers. In some embodiments, the results of each passing through the DoG filter are combined to create a filtered grayscale image that can be used as a mask to represent the stained nuclear material and some “waste material” within each cell. U.S. Patent Application Publication No. 2017 / 0337695 describes a method for generating such a mask, the disclosure of which is incorporated herein by reference in its entirety. The combined grayscale image can then be binarized using, for example, an adaptive thresholding technique based on the Otsu method to produce a dot mask image, wherein each dot outside the dot has a binary value (e.g., logic 0) and each dot inside the dot has the opposite binary value (e.g., logic 1).

[0154] In other embodiments, the detection of the first and second points of in situ hybridization signals representing different colors includes generating a first color channel image and a second color channel image by performing color deconvolution on the digital image (e.g., using demixing module 203), the first color channel image corresponding to the colored spectral contribution of the first dye and the second color channel image corresponding to the colored spectral contribution of the second dye; calculating at least one DoG image from the digital image by applying a pair of Gaussian filters with different standard deviations to the digital image and by subtracting the two filtered images output by the Gaussian filters from each other. The DoG image is a difference-of-Gaussian image; an absorbance image is calculated from the image of the tissue sample; neighboring pixel sets where the absorbance value of the absorbance image in the digital image exceeds an absorbance threshold and the DoG value in the DoG image exceeds a DoG threshold are detected, and the detected neighboring pixel sets are used as expected points; expected points where the intensity value in the first color channel image exceeds a first color intensity threshold are identified, and the identified expected points are output as the first detected point; and expected points where the intensity value in the second color channel image exceeds a second color intensity threshold are identified, and the identified expected points are output as the second detected point.

[0155] In other embodiments, point detection is performed using the method described in U.S. Patent Publication No. 2017 / 0323148, in the context of staining samples with HER2 dual ISH assay, the disclosure of which is incorporated herein by reference in its entirety. According to the method described in '148, point detection is achieved by calculating pixels based on absorbance image features, difference of Gaussian (DoG) image features, and unmixed color image channel features, wherein the pixels are calculated by evaluating whether certain features within the image satisfy predetermined threshold criteria. In some embodiments, the threshold criteria are derived empirically after extensive experimentation. In this case, pixels are calculated based on features that satisfy the predetermined threshold criteria, such as pixels with sufficiently high difference of Gaussian (DoG) and absorbance; black pixels with sufficiently high unmixed image intensity and DoG; and red pixels with sufficiently high unmixed image intensity and DoG. In some embodiments, multiple subsets of pixels are calculated, each satisfying a different threshold criterion, such that the final subset of pixels is strong across all threshold criteria. As is known to those skilled in the art, after calculating the final subset of pixels, a "hole-filling" operation is performed.

[0156] Those skilled in the art will recognize that different standards can be established for different ISH probes, measurements, and protocols to accommodate different signals (such as different colors) in the image and to optimize point detection for that particular ISH protocol accordingly.

[0157] Point classification module

[0158] After point detection is performed by point detection module 208, the system 200 runs point classification module 209, which can assign a color value (such as black or red color value in the HER2 dual ISH measurement background) to all point pixels.

[0159] In some embodiments, dot classification is performed using the method described in U.S. Patent Publication No. 2014 / 0377753, the disclosure of which is incorporated herein by reference in its entirety. According to the method disclosed in '753, once a dot mask image is created (as described above), it is used in conjunction with a classifier (e.g., the classifier in the cell detection and classification module 204) to remove any dots associated with "garbage" and retain dots representing signals corresponding to the target staining agent (e.g., black and red signals in the background of HER2 dual ISH assay). The '753 disclosure describes that, in some embodiments, a linear binary classifier may be utilized. In a first stage using the classifier, the computer system executes instructions to remove dots with weak DoG responses based on a response histogram determined for each cell analyzed. In a second stage, by analyzing the colors of the RGB image of the tissue, coarse red dots are separated from light red dots, and black dots are separated from dark blue dots at pixel locations corresponding to regions within each dot in the dot mask image. This results in a set of dots containing only red and black dots. The remaining dots are considered "garbage" and removed. Points and blobs are then extracted using connected component analysis. To determine whether a point represents a target signal (such as the HER2 gene or chromosome 17), multiple metrics (including blobs) of the point are measured and analyzed. These metrics include the point's size, color, orientation, shape, response to the multiple difference of a Gaussian filter, the relationship or distance between adjacent points, and other factors that can be measured by a computer. The metrics are then fed into a classifier. In some embodiments, the classifier has been previously trained by, for example, a trained pathologist on a training dataset that has been explicitly identified as representing certain genes in the study (such as the HER2 gene or chromosome 17). In some embodiments, the classifier has been trained on a set of training slices containing a range of point variations, and the linear boundary binary classifier model is used for teaching at each stage. The resulting model, with a discriminative hyperplane as a parameter, divides the feature space into two labeled regions that define a first gene or a second gene (such as the HER2 gene or chromosome 17). Once the classifier is trained, metrics measured for unknown points in the image are applied to the classifier. The classifier can then indicate the type of point it represents (e.g., the HER2 gene or chromosome 17).

[0160] In other embodiments, in the context of staining samples for HER2 dual ISH assays, point detection is performed using the method described in U.S. Patent Publication No. 2017 / 0323148, the disclosure of which is incorporated herein by reference in its entirety. According to the method disclosed in '148, since multiple distinct spectral signals may coexist in a single pixel of ISH, a color deconvolution algorithm is run on the image, wherein each pixel is demixed (e.g., using demixing module 203) into three component channels (e.g., red, black, and dark blue channels) using pixels in the optical density domain. After color deconvolution and training of the classification module / classifier, the point pixels are classified by a computer device or system in a manner known to those skilled in the art. In some embodiments, the classification module is a support vector machine (“SVM”) (as described herein with respect to cell detection and classification module 204). In the context of HER2 dual ISH, the detected points are classified as red, black, and / or blue points. For example, when the contribution of the black channel after color deconvolution is significantly greater than the contributions of the other two channels (red and blue), the pixel is more likely to be classified as black; when the contribution of the red channel after color deconvolution is significantly greater than the contributions of the other two channels (blue and black), the pixel is more likely to be classified as red.

[0161] In some embodiments, following point detection and classification, a series of refinement procedures may be run to enhance and clarify the initial point detection and / or classification operations. These steps are described in U.S. Patent Application Publication No. 2017 / 0323148, the disclosure of which is incorporated herein by reference in its entirety. Those skilled in the art will recognize that some or all of these additional steps or procedures may be applied to further enhance the point detection and classification results provided above. Although steps for dual ISH for HER2 detection are disclosed in '148 disclosure, those skilled in the art will be able to adapt and modify these steps to suit any ISH probe or assay used. Those skilled in the art will also recognize that, in terms of refinement classification, it is not necessary to perform all the operations described in '148 disclosure, and those skilled in the art will be able to select appropriate operations based on the output of the detection and classification module and the images provided to the system.

[0162] Point counting and classification module

[0163] The points are counted after point detection, classification, and optional refinement. In some embodiments, data can be compiled based on the number of points counted. For example, the points can be classified based on the count, and / or a ratio of the number of first points (e.g., black points) to the number of second points (e.g., red points) can be calculated, and this ratio can be used for classification or for genetic determination mapping.

[0164] In some embodiments, point counting is performed using the method described in U.S. Patent Publication No. 2017 / 0323148, the disclosure of which is incorporated herein by reference in its entirety. For example, a connected component labeling process can be applied to all first point pixels (e.g., red point pixels) to identify the first blob (“red blob”). Similarly, connected component labeling is applied to second point pixels (e.g., “black point pixels”) to obtain second blobs (e.g., “black blob”). Generally, connected component labeling scans an image and groups pixels based on their connectivity, i.e., all pixels in a connected component have similar pixel intensity values ​​and are connected to each other in some way.

[0165] In the context of dual ISH for HER2 detection, those skilled in the art will understand that the black dots are generally smaller than the red dots, and any counting rule must take this factor into account. For example, if system 200 determines that the size of the dot is within the nominal size range of a single dot, the dot is recorded by the classifier as a HER2 dot or a chromosome 17 dot. If the dot is larger than the nominal size of a single dot, system 200 determines the area of ​​the cluster and divides it by the area of ​​the nominal dot to determine how many dots can be contained in the cluster. The density of dots classified as HER2 and chromosome 17 is used in the counting algorithm. Of course, the counting rules can be modified by those skilled in the art based on the specific ISH probes, assays, and protocols used.

[0166] In some embodiments, the average size (in pixels) of the black spots is used to assign a certain number of seeds to the black spot clusters. For a small black spot, it should have one or both of the following: (i) a voting intensity (using DoG and radial symmetry with respect to absorbance) greater than a certain threshold (e.g., empirically determined and set to 5); and / or (ii) absorbance greater than a certain absorbance threshold (e.g., empirically determined and set to 0.33). In some embodiments, for a small red spot, it should have one or both of the following: (i) a voting intensity (using DoG and radial symmetry with respect to channels) greater than a certain threshold (e.g., empirically determined and set to 15); and / or (ii) an A channel value (from the LAB color space) (a higher A value is a sign of redness) greater than a certain threshold (e.g., empirically determined and set to 133, where the A channel value ranges from 0 to 255), and the absorbance should be greater than a certain threshold (e.g., empirically determined and set to 0.24). For example, a red dot with a diameter less than 7 pixels can be considered a small red dot. A red dot with a diameter of 7 pixels or more can be considered a large red dot. Black dots can be larger on average, so a small black dot may have a diameter of less than 10 pixels, and a large black dot may have a diameter of 10 pixels or more.

[0167] Once the first and second spots (e.g., black and red spots) are identified, the counts of the first and second spots (e.g., black and red spots) are returned. In some embodiments, the ratio of the first spot to the second spot for each cell nucleus is counted.

[0168] In some embodiments, the expression level may be determined as overexpression, underexpression, etc., based on the ratio (step 315; see also step 415). In some embodiments, the score is compared to a clinically relevant threshold. For example, in the context of HER2 and chromosome 17 staining, the clinically relevant threshold may be an integer 2.

[0169] In some embodiments, the ratios of all cell nuclei can be classified or sorted. In some embodiments, the ratios can be stored alone or in combination with cell / nucleus location information (such as the x and y coordinates of cells / nuclei) in a database or storage module 240.

[0170] Visualization module

[0171] A visualization module can be used to illustrate the data generated during signal acquisition (step 314; see also step 414) and evaluation (step 315; see also step 415) for rapid and stable analysis. In some embodiments, an overlay image can be generated based on the derived data and the evaluation performed. For example, in the context of HER2, a first indicator (e.g., a first color) can be assigned to calculated ratios exceeding a certain threshold (e.g., 2), while a second indicator (e.g., a second color) can be assigned to calculated ratios equal to or below the certain threshold (e.g., 2). In some embodiments, if colors are assigned to the first and second indicators, the colors can be delineated along the cell perimeter. In other embodiments, if colors are assigned to the first and second indicators, the cells can be filled with the colors. The resulting overlay image can then be superimposed on the entire slice image or any portion thereof (e.g., to facilitate communication of the results to a reviewer). In some embodiments, the overlay image may include the calculated ratio of each cell / nucleus with or without other indicators (e.g., other colors).

[0172] In some embodiments, a heatmap can be generated that identifies regions that meet certain threshold restrictions. For example, a heatmap can be generated illustrating regions with a calculated ratio between 2.0 and 2.5 in a first color, regions with a calculated ratio between 2.5 and 3.0 in a second color, and regions with a calculated ratio greater than 3.0 in a third color.

[0173] In other embodiments, histograms can be derived from the data and displayed together with any generated overlapping images. In some embodiments, histograms illustrating the distribution of different calculation ratios can be generated.

[0174] In other embodiments, tables can be generated based on the data, and may be accompanied by any generated overlapping images or histograms. Figure 1 The table is displayed. In some embodiments, the table contains a list of a predetermined number of cell nuclei (e.g., 20) and a calculated ratio of the predetermined number of cell nuclei. In some embodiments, the table may also contain location information, such as the location of the cell / nucleus, or the mapped tissue region in which the cell / nucleus is located. In some embodiments, the table may also contain a total dot count for each corresponding nucleic acid biomarker (e.g., the total count of black dots and the total count of red dots for each cell nucleus). In this way, pathologists can use the table to examine individual cells and manually override any automatically calculated ratios or assessments as necessary.

[0175] Other components that implement the embodiments of this disclosure

[0176] The system 200 of this disclosure can be coupled to a specimen processing device capable of performing one or more preparation processes on the tissue specimen. The preparation processes may include, but are not limited to, specimen deparaffinization, specimen conditioning (e.g., cell conditioning), specimen staining, antigen retrieval, immunohistochemical staining (including labeling) or other reactions, and / or in situ hybridization (e.g., SISH, FISH, etc.) staining (including labeling) or other reactions, as well as other processes for preparing specimens for microscopic examination, microanalysis, mass spectrometry, or other analytical methods.

[0177] The processing equipment can apply a fixative to the specimen. Fixatives may include cross-linking agents (e.g., aldehydes such as formaldehyde, polyoxymethylene, and glutaraldehyde, as well as non-aldehyde cross-linking agents), oxidizing agents (e.g., metal ions and complexes, such as osmium tetroxide and chromic acid), protein denaturing agents (e.g., acetic acid, methanol, and ethanol), fixatives with unknown mechanisms (e.g., mercuric chloride, acetone, and picric acid), combination reagents (e.g., Carnoy fixative, Methacarn, Bouin solution, B5 fixative, Rossman solution, and Gendre solution), microwave fixatives, and other fixatives (e.g., exclusion volume fixation and vapor fixation).

[0178] If the specimen is embedded in paraffin, it can be deparaffinized using a suitable deparaffin remover. After deparaffin removal, any number of chemicals can be continuously applied to the specimen. These chemicals can be used for pretreatment (e.g., reversing protein crosslinks, exposing nucleic acids, etc.), denaturation, hybridization, washing (e.g., rigorous washing), detection (e.g., linking display or label molecules to probes), amplification (e.g., amplifying proteins, genes, etc.), counterstaining, coverslips, etc.

[0179] The specimen processing equipment can apply a variety of different chemicals to the specimen. These chemicals include, but are not limited to, staining agents, probes, reagents, rinsing agents, and / or conditioning agents. These chemicals can be fluids (e.g., gases, liquids, or gas / liquid mixtures) or similar substances. The fluids can be solvents (e.g., polar solvents, non-polar solvents, etc.), solutions (e.g., aqueous solutions or other types of solutions), or similar substances. Reagents can include, but are not limited to, staining agents, wetting agents, antibodies (e.g., monoclonal antibodies, polyclonal antibodies, etc.), antigen recovery solutions (e.g., water-based or non-water-based antigen retrieval solutions, antigen recovery buffers, etc.), or similar substances. Probes can be isolated nucleic acids or isolated synthetic oligonucleotides attached to detectable labels or reporter molecules. Labels can include radioactive isotopes, enzyme substrates, cofactors, ligands, chemiluminescent or fluorescent agents, haptens, and enzymes.

[0180] The specimen processing equipment may be automated, such as the BENCHMARK XT and SYMPHONY instruments sold by Ventana Medical Systems, Inc. Ventana Medical Systems, Inc. is the attorney general for several U.S. patents disclosing systems and methods for performing automated analyses, including U.S. Patents Nos. 5,650,327, 5,654,200, 6,296,809, 6,352,861, 6,827,901, and 6,943,029, and U.S. Publication Patent Applications Nos. 20030211630 and 20040052685, the contents of which are incorporated herein by reference in their entirety. Alternatively, specimens may be processed manually.

[0181] After specimen processing, the user can transport the specimen slide to an imaging device. In some embodiments, the imaging device is a bright-field imager slide scanner. One bright-field imager is the iScan HT and DP200 (Griffin) bright-field scanner sold by Ventana Medical Systems, Inc. In automated embodiments, the imaging device is a digital pathology device disclosed in International Patent Application No. PCT / US2010 / 002772 (Patent Publication No.: WO / 2011 / 049608) entitled "IMAGING SYSTEM AND TECHNIQUES" or in U.S. Patent Publication No. 2014 / 0178169 entitled "MAGING SYSTEMS, CASSETTES, AND METHODS OF USING THE SAME," filed September 9, 2011.

[0182] The imaging system or device may be a multispectral imaging (MSI) system or a fluorescence microscopy system. The imaging system used herein is an MSI. Generally, MSI equips pathological specimens with computerized microscopy-based imaging systems by accessing the spectral distribution of images on a pixel-layer scale. Although various multispectral imaging systems exist, they share a common operational feature: the ability to generate multispectral images. A multispectral image is an image that captures image data at specific wavelengths or spectral bandwidths of the electromagnetic spectrum. These wavelengths can be selected using optical filters or other instruments capable of selecting preset spectral components, including electromagnetic radiation with wavelengths beyond the visible light range, such as infrared (IR).

[0183] The MSI system may include an optical imaging system, a portion of which includes a spectral selection system adjustable to define a predetermined number of N discrete optical bands. The optical system is adapted for imaging tissue samples transmitted via a broadband light source onto an optical detector. In one embodiment, the optical imaging system may include a magnification system, such as a microscope, having a single optical axis generally spatially aligned with a single light output of the optical system. When the spectral selection system is adjusted or tuned (e.g., using a computer processor), the system forms a sequence of images of the tissue, thereby ensuring, for example, that images are acquired in different discrete spectral bands. The device may additionally include a display capable of displaying at least one visually perceptible image of the tissue from the acquired image sequence. The spectral selection system may include optical dispersion elements, such as diffraction gratings, a set of optical filters, such as thin-film interference filters, or any other adapted to select a specific passband from the spectrum of light transmitted from the light source through the sample to the detector in response to user input or pre-programmed processor commands.

[0184] In an alternative implementation, the spectral selection system defines multiple light outputs corresponding to N discrete spectral bands. This type of system introduces the output of transmitted light from an optical system and spatially redirects at least a portion of that light output along N spatially distinct optical paths, thus enabling the sample in an identified spectral band to be imaged onto a detector system along an optical path corresponding to the identified spectral band.

[0185] The embodiments of the subject matter and operation described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and similar structures, or one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., as one or more computer program instruction modules for execution by a data processing device or for controlling the operation of a data processing device, said one or more computer program instruction modules may be encoded on a computer storage medium. Any module described herein may include logic executed by said processor. As used herein, "logic" means information having any form of instruction signals and / or data that can affect the operation of a processor. Software is an example of logic.

[0186] Computer storage media may be or may be contained in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or one or more combinations thereof. Furthermore, although a computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in an artificially generated propagating signal. The computer storage medium may also be or may be contained in one or more separate physical components or media (such as multiple CDs, disks, or other storage devices). The operations described herein can be implemented by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0187] The term "programmable processor" encompasses a wide range of devices, apparatuses, and machines that process data, including, for example, programmable microprocessors, computers, systems-on-a-chip, or a combination thereof. These devices may include special-purpose logic circuitry such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, these devices may also include code that creates an execution environment for the relevant computer program, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or one or more combinations thereof. These devices and execution environments can implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.

[0188] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. Computer programs may, but do not necessarily, correspond to files in a file system. A program may be stored in a partial file that holds other programs or data (such as one or more scripts stored in a markup language file), in a single file dedicated to the program, or in multiple coordinating files (such as a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0189] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform operations by processing input data and producing output results. The processes and logic flows can also be executed by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as a special-purpose logic circuit.

[0190] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in a digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes operations according to instructions and one or more storage devices that store instructions and data. Generally, a computer will also include or be effectively coupled to one or more mass storage devices (disk, magneto-optical, or optical disk) for storing data, to receive data from or to said device, or both. However, a computer does not require such a device. Furthermore, a computer may embed another device, such as a mobile phone, personal digital assistant (PDA), mobile audio or video player, game controller, GPS receiver, or portable storage device (such as a Universal Serial Bus (USB) flash drive). Suitable devices for storing computer program instructions and data include various forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices like EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be added to or incorporated into special purpose logic circuitry.

[0191] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user, such as an LCD (Liquid Crystal Display), an LED (Light Emitting Diode) display, or an OLED (Organic Light Emitting Diode) display, as well as a keyboard and pointing devices, such as a mouse or trackball, through which the user can input information into the computer. In some embodiments, a touchscreen can be used to display information and receive user input. Other types of devices can also be used to facilitate user interaction; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and user input can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending and receiving files from the device used by the user; for example, by sending a webpage to the web browser in response to a request received from the web browser on the user's client device.

[0192] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface or a web browser through which a user can interact with embodiments of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (such as the Internet) and peer-to-peer networks (such as dedicated peer-to-peer networks). For example, Figure 1 The network 20 may include one or more local area networks (LANs).

[0193] The computing system may include any number of clients and servers. Typically, clients and servers are remotely configured and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them. In some embodiments, the server transmits data (such as HTML pages) to a client device (e.g., for displaying data to a user interacting with the client device and receiving the user's input). Data generated on the client device (such as the result of user interaction) can be received from the client device on the server.

[0194] Example

[0195]

[0196]

[0197] The table above lists examples of results returned after image analysis used to detect HER2 protein biomarkers and HER2 and chromosome 17 nucleic acid biomarkers. The table also illustrates the calculated ratio of HER2 to chromosome 17, and the number of relevant cells detected.

[0198] Other embodiments

[0199] A system for assessing genetic aberrations in images of biological samples stained for the presence of at least one nucleic acid biomarker, the system comprising: (i) one or more processors, and (ii) one or more memories coupled to the processors, the memories storing computer-executable instructions which, when executed by the processors, cause the system to perform operations including: running a detection algorithm to automatically detect and identify cells in a first image stained for the presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; deriving (e.g., automatically) a tumor tissue region in the first image encompassing the identified cells that meet the predetermined protein biomarker staining criteria; performing automatic registration of the first image and a second image with a common coordinate system such that the derived tumor tissue region in the first image is mapped to the second image to provide a mapped tumor tissue region, wherein the second image includes a signal corresponding to the presence of at least one nucleic acid biomarker; automatically identifying points within the mapped tumor tissue region corresponding to the signal from the at least one nucleic acid biomarker; and assessing (e.g., automatically) whether the nuclei of tumor cells in the mapped tumor tissue region in the second image have genetic aberrations based on the identified points.

[0200] A method for assessing genetic aberrations in an image of a biological sample stained for the presence of at least one nucleic acid biomarker, the method comprising: automatically detecting cells in a first image stained for the presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; deriving a tumor tissue region in the first image encompassing identified cells that meet the predetermined protein biomarker staining criteria; automatically registering the first image and a second image with a common coordinate system such that the derived tumor tissue region in the first image is mapped to the second image to provide a mapped tissue region, wherein the second image includes a signal corresponding to the presence of at least one nucleic acid biomarker; automatically identifying points within the mapped tissue region corresponding to the signal from the at least one nucleic acid biomarker; and assessing (e.g., automatically) whether tumor cell nuclei in the mapped tissue region in the second image have genetic aberrations based on the identified points corresponding to the at least one nucleic acid biomarker.

[0201] A non-transitory computer-readable medium storing instructions for assessing genetic aberrations in a biological sample stained for the presence of at least one nucleic acid biomarker, the instructions comprising: running a detection algorithm to automatically detect and identify cells in a first image stained for the presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; deriving a tumor tissue region in the first image encompassing the identified cells that meet the predetermined protein biomarker staining criteria; performing automatic registration of the first image and a second image with a common coordinate system such that the derived tumor tissue region in the first image is mapped to the second image to provide a mapped tissue region, wherein the second image includes a signal corresponding to the presence of at least one nucleic acid biomarker; automatically detecting points in the mapped tissue region corresponding to the signal from the at least one nucleic acid biomarker; counting all detected points within each tumor cell nucleus in each mapped tissue region; and assessing (e.g., automatically) whether each tumor cell nucleus in each mapped region has a genetic aberration based on the total number of points counted in each cell nucleus.

[0202] All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patents, and non-patent publications mentioned in and / or listed in the application data sheets are incorporated herein by reference in their entirety. Modifications may be made to various aspects of the embodiments as necessary to provide other further embodiments employing the concepts of various patents, applications, and publications.

[0203] Although this disclosure has been described with reference to some illustrative embodiments, it should be understood that those skilled in the art can devise many other modifications and embodiments within the spirit and scope of the principles of this disclosure. More specifically, reasonable variations and modifications can be made to the components and / or arrangements of the subject matter combination within the scope of the foregoing disclosure, the drawings, and the appended claims without departing from the spirit of this disclosure. In addition to variations and modifications in the components and / or arrangements, alternative uses will also be apparent to those skilled in the art.

Claims

1. A system for assessing gene aberrations in an image of a biological sample stained for the presence of at least one nucleic acid biomarker, the system comprising: (i) one or more processors, and (ii) one or more memories coupled with the one or more processors, the memories to store computer-executable instructions that, when executed by the one or more processors, cause the system to perform operations comprising: (a) running a detection algorithm (204) to automatically detect and identify cells (411) in a first image stained for presence of at least one protein biomarker that meet predetermined protein biomarker staining criteria; (b) deriving a tumor tissue region (412) in the first image encompassing the identified cells that meet the predetermined protein biomarker staining criteria; (c) performing automatic registration of the first image and a second image to a common coordinate system such that the derived tumor tissue region in the first image maps to the second image to provide a mapped tumor tissue region, wherein the second image includes signals corresponding to presence of at least one nucleic acid biomarker (413); (d) automatically identifying points (414) within the mapped tumor tissue region that correspond to the signals from the at least one nucleic acid biomarker; and (e) assessing whether tumor cell nuclei in the mapped tumor tissue region in the second image have genetic aberrations based on the identified points (415), wherein genetic aberration of the cell nuclei is assessed by (i) for each cell nucleus, calculating a ratio of first identified points corresponding to a first nucleic acid biomarker to second identified points corresponding to a second nucleic acid biomarker; and (ii) comparing the calculated ratio for each cell nucleus to a predetermined threshold.

2. The system of claim 1, wherein genetic aberration of the cell nuclei is assessed by automatically determining whether a total number of identified points in each cell nucleus that correspond to signals from the at least one nucleic acid biomarker meet a predetermined threshold.

3. The system of claim 1, wherein the first nucleic acid biomarker is HER2 and the second nucleic acid biomarker is chromosome 17, and wherein the at least one protein biomarker is a HER2 protein biomarker.

4. The system of claim 1, wherein the first nucleic acid biomarker is EGFR and the second nucleic acid biomarker is chromosome 7, and wherein the at least one protein biomarker is an EGFR protein biomarker.

5. The system of claim 1, further comprising associating a total number of points associated with a first stain and a total number of points associated with a second stain with cell nucleus location information.

6. The system of claim 1, further comprising assigning a first indicator to those assessed cell nuclei for which the calculated ratio is above the predetermined threshold, and assigning a second indicator to those assessed cell nuclei for which the calculated ratio is at or below the predetermined threshold.

7. The system of claim 6, further comprising generating an overlay image based on the assigned first and second indicators. ​ 8. The system of claim 1, further comprising generating a binned histogram of the computed ratios of all scored nuclei in each mapped tissue region.

9. The system of claim 8, further comprising operations for identifying the histogram bin with the largest count.

10. The system of claim 8, further comprising operations for determining a course of treatment based on data from the generated binned histogram.

11. The system of claim 1, further comprising ranking the scored nuclei according to the computed ratio of each nucleus.

12. The system of any of the preceding claims, wherein the predetermined protein biomarker staining criterion is a staining intensity threshold.

13. The system of claim 12, wherein the staining intensity threshold is a cutoff value for presence of membrane staining.

14. The system of any of the preceding claims, wherein the genetic aberration comprises an abnormal gene copy number.

15. The system of claim 14, wherein the abnormal gene copy number is a copy number greater than a normal copy number of the gene.

16. The system of any of the preceding claims, wherein the genetic aberration comprises a chromosomal abnormality.

17. The system of any of the preceding claims, wherein the number of scored nuclei in each mapped tissue region is greater than 20.

18. A method of assessing a genetic aberration in an image of a biological sample stained for presence of at least one nucleic acid biomarker, the method comprising: (a) automatically detecting cells (411) in a first image stained for presence of at least one protein biomarker that satisfy a predetermined protein biomarker staining criterion; (b) deriving a tumor tissue region (412) in the first image that encompasses the identified cells that satisfy the predetermined protein biomarker staining criterion; (c) automatically registering the first image and a second image to a common coordinate system such that the derived tumor tissue region in the first image maps to the second image to provide a mapped tissue region, wherein the second image comprises signals corresponding to presence of at least one nucleic acid biomarker (413); (d) automatically identifying points within the mapped tissue region that correspond to signals from the at least one nucleic acid biomarker (414); and (e) based on the identified points corresponding to the at least one nucleic acid biomarker, assessing whether tumor cell nuclei in the mapped tissue region in the second image have a genetic aberration (415), wherein the genetic aberration of the nuclei is assessed by: (i) for each nucleus, computing a ratio of first identified points corresponding to a first nucleic acid biomarker to second identified points corresponding to a second nucleic acid biomarker; and (ii) comparing the computed ratio of each nucleus to a predetermined threshold. ​ 19. The method of claim 18, further comprising assigning a first indicator to those tumor cell nuclei for which the computed ratio is above the predetermined threshold and a second indicator to those tumor cell nuclei for which the computed ratio is at or below the predetermined threshold.

20. The method of claim 19, further comprising generating an overlay image based on the assigned first and second indicators.

21. The method of claim 18, further comprising generating a binned histogram of the computed ratios for all identified cell nuclei.

22. The method of claim 21, further comprising ranking the assessed cell nuclei according to their computed ratios.

23. The method of any one of claims 18-22, wherein the biological sample is stained for the presence of the HER2 and chromosome 17 nucleic acid biomarkers.

24. The method of claim 23, wherein the spots are identified by: (a) automatically detecting spots in the mapped tissue region that satisfy criteria of absorbance intensity, black unmixing image channel intensity, red unmixing image channel intensity, and Gaussian difference threshold; and (b) automatically classifying the detected spots as belonging to either black nucleic acid biomarker signal corresponding to HER2 or red nucleic acid biomarker signal corresponding to chromosome 17.

25. The method of claim 24, wherein the tumor cell nuclei are assessed by: (i) computing a ratio of those classified spots belonging to the black nucleic acid biomarker signal and those classified spots belonging to the red nucleic acid biomarker signal; and (ii) comparing the computed ratio to a predetermined threshold.

26. The method of claim 25, wherein the at least one protein biomarker is a HER2 protein biomarker.

27. The method of claim 26, further comprising identifying whether a patient is positive or negative for HER2 based on the assessed tumor cell nuclei.

28. The method of claim 26, further comprising scoring the biological sample for the presence of at least one additional protein biomarker.

29. The method of claim 28, wherein the at least one additional protein biomarker is EGFR.

30. The method of any one of claims 18-29, wherein the genetic aberration is RNA overexpression.

31. A non-transitory computer readable medium storing instructions for assessing a genetic aberration in a biological sample stained for the presence of at least one nucleic acid biomarker, the instructions comprising: (a) running a detection algorithm (204) to automatically detect and identify cells (411) in a first image stained for the presence of at least one protein biomarker that satisfy predetermined protein biomarker staining criteria; (b) deriving a tumor tissue region (412) in the first image encompassing the identified cells that satisfy the predetermined protein biomarker staining criteria; (c) performing an automated registration of the first and second images to a common coordinate system such that the derived tumor tissue region in the first image is mapped to the second image to provide a mapped tissue region, wherein the second image comprises a signal corresponding to the presence of at least one nucleic acid biomarker (413); (d) automatically detecting points within the mapped tissue region corresponding to a signal from the at least one nucleic acid biomarker (414); (e) counting all detected points within each tumor cell nucleus in each mapped tissue region; and (f) assessing whether each tumor cell nucleus in each mapped region has a genetic aberration based on a total number of counted points in each nucleus (415), wherein the genetic aberration of the nucleus is assessed by (i) calculating, for each nucleus, a ratio of first identified points corresponding to a first nucleic acid biomarker to second identified points corresponding to a second nucleic acid biomarker; and (ii) comparing the calculated ratio of each nucleus to a predetermined threshold.

32. The non-transitory computer readable medium of claim 31, wherein the biological sample is stained for the presence of at least two nucleic acid biomarkers, and wherein points corresponding to each of the at least two nucleic acid biomarkers are detected and counted.

33. The non-transitory computer readable medium of claim 32, wherein the tumor cell nuclei are assessed by (i) calculating a ratio of counted first points to counted second points; and (ii) comparing the calculated ratio to a clinically relevant threshold.

34. The non-transitory computer readable medium of claim 33, further comprising instructions for generating an image overlay, wherein each assessed nucleus is assigned a color based on the calculated ratio (i) being at or below the clinically relevant threshold; or (ii) being above the clinically relevant threshold.

35. The non-transitory computer readable medium of claim 33, further comprising instructions for ranking the assessed nuclei in each mapped tissue region according to the calculated ratio.

36. The non-transitory computer readable medium of claim 33, further comprising instructions for generating a binned histogram of the calculated ratios.

37. The non-transitory computer readable medium of any one of claims 31-36, wherein the predetermined protein biomarker staining criterion is a staining intensity threshold.

38. The non-transitory computer readable medium of any one of claims 31-37, wherein the genetic aberration is selected from the group consisting of an abnormal gene copy number and a chromosomal abnormality.

Citation Information

Patent Citations

  • Automated molecular pathology apparatus having independent slide heaters

    US20030211630A1

  • Automated molecular pathology apparatus having independent slide heaters

    US20040052685A1

  • Imaging systems, cassettes, and methods of using the same

    US20140178169A1

  • System for detecting genes in tissue samples

    US20140377753A1

  • Image Analysis for Breast Cancer Prognosis

    US20150347702A1