Ground Truth for Signal Aggregation and Quantification

The method and system convert stain intensity into predicted biomarker intensity and error metrics using a linear prediction function and confidence function, addressing the challenges of high cell density and aggregation in digital pathology, enabling accurate quantitative assessments.

JP2026516643APending Publication Date: 2026-05-26VENTANA MEDICAL SYSTEMS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
VENTANA MEDICAL SYSTEMS INC
Filing Date
2024-04-05
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing digital pathology techniques struggle to accurately quantify biomarker signals due to high cell density, high expression levels, and spatial aggregation, leading to qualitative rather than quantitative signal evaluation, which hinders diagnosis and treatment recommendations.

Method used

A method and system that utilize a linear biomarker intensity prediction function and a confidence function to convert stain intensity into predicted biomarker intensity and error metrics, establishing an artificial ground truth framework to quantify biomarker levels and error, even in saturated regions.

Benefits of technology

Enables accurate and quantitative assessment of biomarker levels, facilitating more precise diagnoses and treatment recommendations by generating reliable quantitative metrics from digital pathology images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516643000001_ABST
    Figure 2026516643000001_ABST
Patent Text Reader

Abstract

Digital pathology images acquired using bright-field imaging depicting slides with stained sample slices are accessed. Stain intensity corresponding to at least a portion of the digital pathology images is detected. A biomarker intensity prediction function is accessed, linearly relating predicted biomarker intensity levels to the detected stain intensity. A nonlinear confidence function is accessed, relating the confidence of the predicted biomarker intensity to the detected intensity of the stain. Predicted biomarker intensity is generated for at least a portion of the slide using the detected stain intensity corresponding to at least a portion of the slide and the linear biomarker intensity prediction function. A confidence metric for the predicted biomarker intensity is generated based on the confidence function using the detected stain intensity corresponding to at least a portion of the slide. Results based on the predicted biomarker intensity and confidence metric are output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 496,207, titled "Ground Truth of Signal Aggregate Quantification", filed on April 14, 2023, and this provisional patent application is hereby incorporated by reference in its entirety for all purposes.

Background Art

[0002] Digital pathology may involve the interpretation of digitized images in order to accurately diagnose a subject and guide treatment decisions. In digital pathology solutions, an image analysis workflow can be established to automatically detect or classify target biological objects, such as positive and negative tumor cells. An exemplary workflow of a digital pathology solution includes obtaining a tissue slide, scanning the pre - selected region or the entire tissue slide with a digital image scanner (e.g., a whole - slide image (WSI) scanner) to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and potentially detecting and quantifying (e.g., counting or identifying the region specific to that object or the cumulative region) each target object based on the image analysis (e.g., quantitative or semi - quantitative scoring such as positive, negative, medium, weak, etc.).

[0003] A common use of digital pathology is the quantification of signals in tissue samples, such as ribonucleic acid signals, to aid in the analysis of gene expression changes in individual diseased cells (e.g., cancer cells), dysregulated normal cells (e.g., immune cells), and / or healthy cells. Typically, image processing is performed to detect signals corresponding to a given stain in order to assess gene expression. More specifically, tissue biopsies or liquid samples (e.g., blood samples) can be processed to fix the sample and introduce stains to indicate where the desired signal is located within the sample.

[0004] The extent to which a signal can be accurately detected and / or precisely localized may depend on the size, density, and / or aggregation of cells of a given type or expressing a given gene. For example, if a tissue sample contains a high-density population of cells expressing a given gene or a population of cells with high expression of a given gene, the image of the section may contain blobs of stain color. Therefore, signal evaluation is often performed relative or qualitatively, rather than quantitatively. These approaches may hinder techniques for facilitating diagnosis, treatment recommendations, prognosis, etc. [Overview of the project]

[0005] In some embodiments, a computer implementation method is provided, comprising: accessing a digital pathology image depicting a slide having slices of a sample stained with a stain, wherein the digital pathology image was acquired using bright-field imaging; detecting the stain intensity corresponding to at least a portion of the digital pathology image; accessing a linear biomarker intensity prediction function that linearly relates a predicted level of biomarker intensity to the detected intensity of a stain, wherein the biomarker intensity prediction function is generated by evaluating digital pathology images of other slides, the other slides containing samples stained with several other concentrations of stains; accessing a confidence function that relates the confidence of the predicted biomarker intensity to the detected intensity of a stain, wherein the confidence function is nonlinear; generating a predicted biomarker intensity for at least a portion of the slide based on the detected stain intensity corresponding to at least a portion of the slide and based on the linear biomarker intensity prediction function; generating a confidence metric for the predicted biomarker intensity based on the detected stain intensity corresponding to at least a portion of the slide and based on the confidence function; and outputting results based on the predicted biomarker intensity and the confidence metric.

[0006] The confidence function may further include a first part that linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the stain, and the confidence function may further include a second part that non-linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the stain.

[0007] The second part of the confidence function may correspond to the saturation of the detected intensity of the stain.

[0008] The method may further include, where at least a portion of the slide is pixels, generating a different predicted biomarker intensity for each of the other sets of pixels in the digital pathology image based on the detected staining intensity corresponding to the other pixels and based on a linear biomarker intensity prediction function; generating a different confidence metric for each of the other sets of pixels in the digital pathology image based on the detected staining intensity corresponding to the other pixels and based on a confidence function; and generating results based on the predicted biomarker intensity, the other predicted biomarker intensity, the confidence metric and the other confidence metric.

[0009] The method may, alternatively or additionally, include determining whether a stored criterion is met based on a confidence metric, and generating results by integrating the predicted biomarker intensity based on the determination that the stored criterion is met.

[0010] The staining agent may be (for example) an RNA staining agent, a nuclear protein, or a cytoplasmic protein.

[0011] In some embodiments, a system is provided that includes one or more data processors and a non-temporary computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods disclosed herein.

[0012] In some embodiments, a computer program product is provided which includes instructions tangibly embodied in a non-temporary machine-readable storage medium and configured to cause one or more data processors to perform some or all of the methods disclosed herein.

[0013] Some embodiments of the present disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-temporary computer-readable storage medium containing instructions that, when executed by one or more data processors, cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-temporary machine-readable storage medium, which contains instructions configured to cause one or more data processors to execute some or all of the methods disclosed herein and / or some or all of one or more processes.

[0014] The terms and expressions used are for illustrative purposes only, not limitation, and in using such terms and expressions there is no intention to exclude equivalents or parts of the features shown and described, however it is acknowledged that various modifications are possible within the scope of the invention as described in the claims. Accordingly, while the invention as described in the claims is specifically disclosed by embodiments and optional features, modifications and variations of the concepts disclosed herein may be used by those skilled in the art, and it should be understood that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawing]

[0015] The patent or application file must include at least one drawing performed in color. A copy of the published patent or patent application, including the color drawing, will be provided by the Patent Office upon request and payment of the required fees.

[0016] The aspects and features of various embodiments will become even clearer by illustrating examples with reference to the accompanying drawings.

[0017] [Figure 1]Exemplary tissue sections stained using in-situ hybridization (ISH) are shown, displaying KAPPA mRNA detected with black silver (Ag) and LAMBDA mRNA detected with purple tyramide SRB. [Figure 1A] The overall slide image shows the six tonsil regions. [Figure 1B] The image shows the field of view at 40x magnification.

[0018] [Figure 2] Models of several embodiments of the present invention demonstrate how signals detected by digital pathology may relate to true biomarker expression.

[0019] [Figure 3] This document presents an exemplary network for generating digital pathology images and accurately quantifying the staining signals depicted in those images.

[0020] [Figure 4] The image shows 11 breast tissue slides stained with different probe concentrations: 0 pM (or "no probe control (NPC)"), 0.625 pM, 0.125 pM, 0.25 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM, with 15 rectangles located in the tumor area of ​​each slide.

[0021] [Figure 5] Figure 4 shows breast FOV images overlaid with green superpixel segments to indicate the number of isolated spots (red dots) and signal aggregate blobs (blue numbers) for probe concentrations of 0 pM (or "no probe control (NPC)"), 0.625 pM, 0.125 pM, 0.25 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM, respectively.

[0022] [Figure 6]Eleven tonsil tissue slides stained with different probe concentrations of 0 pM (or "no probe control (NPC)"), 0.625 pM, 0.125 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM are shown, with 15 rectangles located in the tumor regions of each slide.

[0023] [Figure 7] Data identifying the total number of spots, isolated spots, and signal aggregation blobs plotted for concentrations of 0 pM to 8 pM for breast, prostate, CRC, and tonsil tissue slides, respectively, are shown.

[0024] [Figure 8] Estimated true biomarker intensity as a function of concentration for multiple tissue types and signals detected by digital pathology is shown.

[0025] [Figure 9] Exemplary nuclear detection results and labels overlaid on the original images are shown.

[0026] [Figure 10] The average number of spots per cell with an ideal biomarker expression line, and the error rate plotted for concentrations of 0 pM to 32 pM in different tissue types are shown.

[0027] [Figure 11] Results of tumor cell classification (red dots) from non-target cells (green dots) using automated image analysis in RNA staining images of prostate cancer at concentrations of 0.5 pM (left), 4 pM (middle), and 32 pM (right) are shown.

[0028] [Figure 12] A process for generating a tumor mask starting from cell-by-cell classification, tumor mask, and polygon generation to group objects in a tumor labeled image is shown. The bottom image is the original image at concentrations of 16 pM and 0 pM (NPC) overlaid with tumor polygons.

[0029] [Figure 13] The graph on the right shows the average spot per cell within the tumor area, compared to the average spot per cell (left graph), and includes ideal biomarker expression lines and error rates plotted for concentrations from 0 pM to 32 pM on prostate tissue slides.

[0030] [Figure 14] The box plots show the intensity characteristics of isolated spots with concentrations of 0.625 pM, 0.125 pM, 0.25 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM.

[0031] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, different components of the same type may be distinguished by following the reference label with a dash and a second label to distinguish similar components. Where only the first reference label is used herein, its description is applicable to any one of the similar components having the same first reference label, regardless of the second reference label. [Modes for carrying out the invention]

[0032] Several embodiments relate to techniques and systems that facilitate the conversion of stained specimen images into meaningful quantitative results. For example, techniques and systems can be used to puncture RNA in situ hybridization (ISH) signals to quantify gene expression while preserving tissue condition and enabling single-cell analysis and workflows. Another example is the use of techniques and systems to convert the intensity of bright-field images into predicted gene expression, protein levels, or secretion levels (and / or error metrics thereof) by using data generated by associating signals with various stain concentrations. In some cases, the data used to associate signals with various stain concentrations corresponds to data showing that the stain is staining molecules that are abundantly present on the slide.

[0033] Many signals detected in the context of digital pathology may have been difficult to quantify due to factors such as the high density of cells of a given type or with a given characteristic (e.g., expression of a given gene), the possibility of very high expression (e.g., a given gene), or the tight spatial aggregation of cells of a given type or with a given characteristic. For example, one approach to quantifying a biomarker signal may include all or part of one or more approaches to quantifying a signal, such as the one disclosed in U.S. Patent Application No. 17 / 586,982 filed on January 28, 2022, which is incorporated herein by reference in whole for all purposes. This approach may additionally or alternatively include: 1) detecting isolated spots in an image (e.g., an unmixed image channel image corresponding to a signal from a biomarker); 2) deriving optical density values ​​for representative isolated spots (e.g., based on the characteristics or properties of the calculated signal from the detected isolated spots); and 3) estimating the number of predicted spots of signal aggregates in each of the subregions based on the derived optical density values ​​of representative isolated spots. This approach may further include estimating the total number of spots in each subregion by combining the number of isolated spots detected with the estimated number of predicted spots for signal aggregates in each subregion and aggregating the sum across the entire tissue slide.

[0034] For illustrative purposes, as shown in Figure 1, tonsil tissue stained for KAPPA mRNA can be detected using a black pigment (silver, Ag), and LAMBDA mRNA can be detected using a violet pigment (tyramide sulforhodamine). The presence of the desired signal appears as small spots (e.g., discrete dots), and these spots can accumulate to form larger regions of agglutination signal (hereinafter referred to as "signal agglutination blobs" or "blobs"), depending on the expression level (copy number) of each target mRNA in B cells. For example, plasma cells have approximately 100,000 mRNA copies per cell, and therefore the signal in those cells can appear as blobs. However, given that the expression of stains can agglutinate and / or saturate, quantifying the level of a biomarker (e.g., gene expression level) expressed in any given blob region can be difficult.

[0035] It would be useful to be able to reliably convert each signal spot or signal blob into a quantifiable signal metric representing the amount of absorbed stain. Furthermore, it would be useful to establish a ground truth metric that characterizes the signal aggregate. These capabilities can facilitate the generation of more accurate diagnoses or prognoses and / or more accurate predictions of the extent to which a given treatment can effectively treat the medical condition of a given subject.

[0036] In some embodiments, techniques and systems are provided for generating highly sensitive quantitative immunohistochemical (qIHC) signals. qIHC techniques are novel, highly sensitive detection methods with the ability to detect several types of biomarkers, such as ribonucleic acid (RNA), proteins, and secreted factors, which are low-density biomarkers that govern or influence tumor growth, immune cell activation and / or dysregulation, and angiogenesis. qIHC techniques can overcome the limited capabilities of on-marker immunohistochemistry (IHC), such as ultraVIEW and OptiVIEW DAB, which can only detect protein biomarkers. For example, existing IHC systems cannot detect secreted factors that leave cells and are not present in sufficient density to generate signals that can be reliably interpreted by pathologists.

[0037] The lack of "ground truth" stems from the inability to precisely and accurately quantify the number of target biomarkers in a tissue sample because any method used to extract molecules involves undefined material loss. Therefore, measurements are not entirely accurate, and all methods used to quantify biomarker abundances involve unknown errors and a lack of precision. Establishing this ground truth can be practically impossible, as predictively identifying tissue samples with biomarkers covering the full range of their presence can be extremely difficult. This difficulty can be combined with the reality that stain absorption can vary across tissue type and potentially other variables (e.g., disease type, demographics of the subject, cell type, etc.), making it even more possible to find a training set covering an applicable range of staining intensities and determine how to translate the staining of slides into results useful for generating a function for transforming staining intensity.

[0038] On the other hand, embodiments of the present invention can identify predicted biomarker levels based on a function generated by relating signal intensity (in the unsaturated region) to staining concentration. Furthermore, embodiments can generate a prediction error (or confidence) metric by relating predicted biomarker levels (in the saturated and / or unsaturated zones) to signal intensity. Thus, an artificial ground truth is created, facilitating the conversion of digital pathology images to metrics representing biomarker levels and used to further identify the error (or confidence) metric of the metric.

[0039] In some embodiments, ground truth is unavailable, but methods and systems are provided for constructing and / or using frameworks that can be used to quantify and validate signals. Signals may include signals from digital pathology images. The framework may be based on, or to ensure, the validation of the accuracy of results and the ability of the quantification method in a whole-slide analysis scheme. This framework can overcome the limitations of conventional signal detection techniques (e.g., the development of RNA or protein biomarker detection assays and / or algorithms) due to the fact that there is no “ground truth” of signal aggregation in these conventional situations.

[0040] Some embodiments of the present invention provide systems, methods, and paradigms to avoid this central and significant obstacle. Abundant biomolecules (e.g., 18s ribosomal RNA) can be targeted, and probes for detecting abundant biomolecules can be used in a linear range of subsaturated concentrations. The combination of high-abundance targets and supersaturated probe concentrations eliminates the need to rely on knowing the absolute abundance of the target biomolecule in the sample being tested. Furthermore, the use of probe concentrations in a known linear range allows for the establishment of a linear regression function, which can then be used to assist in determining or documenting metrics that characterize the capability, reliability, and / or error of signal algorithms or predictive biomolecular metrics. For example, a framework can be defined based on the hypothesis of a linear relationship between the number of dot signals expressed on a tissue slide and the probe concentration in the assay.

[0041] Accordingly, the disclosed image processing techniques can detect signal intensity (e.g., at each pixel or for each region) and convert the signal intensity into a predicted biomolecular level and / or error (or confidence or accuracy) metric. The biomolecular level can be predicted using a linear relationship (which can be established by associating signal intensity with different concentrations of staining agents), and the error can be generated based on this relationship. The error can be defined based on a cutoff that distinguishes a first part of the probe concentration x-axis corresponding to a consistent error metric (e.g., it represents a constant additive error amount) and a second part of the x-axis corresponding to a nonlinear and / or non-constant error metric.

[0042] Therefore, an artificial ground truth framework can be defined for stainers and tissue types by using stainers to target abundant biomolecules in a given type of tissue and evaluating the digital pathology signals detected across various concentrations of stainers. The artificial ground truth framework can be defined to include a linear relationship relating the signals detected using digital pathology imaging to quantitatively predicted biomarker amounts. Furthermore, the artificial ground truth framework can be defined to include a function that estimates the error in predicted biomarker amounts based on saturation signal levels and / or signals detected using digital pathology imaging. Then, stainers can be used to target other biomolecules (not necessarily abundant biomolecules) in a given type of tissue, and the artificial ground truth framework can be applied to convert the captured digital pathology signals into quantitative estimates at other biomolecular levels (e.g., per pixel or region on a slide).

[0043] It will be understood that the artificial ground truth framework does not need to be defined in a way that supports the generation of predictions that accurately identify true biomarker levels in an absolute sense. Rather, the artificial ground truth framework can be used to generate new, potentially arbitrary scales in essence, even though it is useful in supporting quantitative comparisons of biomarker levels (e.g., to support comparisons of levels across parts of a given slide, comparisons of levels across different slides associated with a single sample, comparisons of levels across different slides associated with different subjects, comparisons of levels across different slides associated with different sample selection time points, etc.). Furthermore, the artificial ground truth framework can be used to validate digital pathology algorithms that quantify predictive signals.

[0044] The results (e.g., predicted biomolecular levels and / or prediction error) can be output and / or used to inform diagnostic, prognosis, and / or treatment recommendations. Additional or alternative predicted biomolecular levels and / or error metrics can be used to tune and improve algorithm parameters to produce a more robust and reliable system.

[0045] Various embodiments relate to systems and methods built around the hypothesis that the total number of RNA dots has a linear relationship with the RNA probe concentration. To determine the parameters of the linear relationship, tissues (e.g., breast, prostate, CRC, and tonsils) can be stained with different probe concentrations (e.g., 0 pM (NPC), 0.625 pM, 0.125 pM, 0.25 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM). A digital pathology algorithm can be configured to quantify the total number of RNA dot signals in both the form of isolated dots and aggregate signals and report them with respect to the analysis of the entire slide. The data can be processed to detect a first range of probe concentrations where the RNA dot count increases linearly as the probe concentration increases, and a saturation point where the RNA dot count plateaus. In some cases, the saturation point is assumed to be 8 pM.

[0046] This technique can be used to evaluate dot signals within a single cell (e.g., a single healthy cell, immune cell, or tumor cell) within a larger region (e.g., a tumor region), but averaging RNA dots per cell within a tumor region and averaging RNA dots per cell shows a smaller error rate compared to the total number of RNA dot counts within the entire tumor region. This demonstrates that RNA dot counts fit into an ideal biomarker expression line within a certain linear range, and the error rate increases after reaching saturation concentration. On the other hand, by processing images according to the technique disclosed herein, the total number of RNA dot signals can be accurately estimated for both isolated spots and signal aggregates. This methodology may be a validation method for verifying the algorithmic capability of signals in stained aggregates.

[0047] Furthermore, investigations into the characteristics of isolated spots at different concentrations demonstrated a greater understanding of spot properties. While the blur characteristics and size remained the same throughout the increasing concentration, the intensity and roundness changed with increasing concentration.

[0048] Figure 2 shows proposed underlying relationships used in several embodiments to translate staining intensity into predicted biomarker expression and error. Line 202 illustrates how (according to embodiments of the present invention) it may be assumed that the estimated biomolecular intensity changes linearly with probe concentration. However, as shown by line 204, the detected signal changes linearly based on the probe concentration over the first lower portion 206 of the probe concentration, then saturates and remains constant over the second higher portion 208 of the probe concentration. Thus, the error in the detected signal (represented by the upper and lower error boundaries 210 and 212) is relatively small and constant over the entire first lower portion 206 of the probe concentration, but substantially grows over the entire second higher portion 208 of the probe concentration. Thus, the confidence of the intensity estimate remains relatively high over the entire first portion 206 of the probe concentration, but gradually decreases over the entire second portion 208 of the probe concentration.

[0049] Figure 3 shows an exemplary network for generating digital pathology images and accurately quantifying the staining signals depicted in the digital pathology images. The images are generated by the image generation system 305. The fixation / embedding system 310 fixes and / or embeds tissue samples (e.g., samples containing at least a portion of at least one tumor) using fixatives (e.g., liquid fixatives such as formaldehyde solution) and / or embedding materials (e.g., histological waxes such as paraffin wax, and / or one or more resins such as styrene or polyethylene). Each slice may be fixed by exposing the slice to the fixative for a predetermined period (e.g., at least 3 hours) and then dehydrating the slice (e.g., via exposure to an ethanol solution and / or a clearing intermediate). The embedding material can penetrate the slice when it is in a liquid state (e.g., when heated).

[0050] Next, the tissue slicer 315 slices the fixed and / or embedded tissue sample (e.g., a tumor sample) to obtain a series of sections, each section having a thickness of, for example, 4-5 microns. Such sectioning may be performed by first cooling the sample and then slicing it in a hot water bath. The tissue may be sliced ​​using (e.g.) a vibratome or a compressstorm.

[0051] Since tissue sections and the cells within them are nearly transparent, slide preparation generally involves staining the tissue sections (e.g., automatically) to make the relevant structures more visible. In some cases, staining is performed manually. In other cases, staining is performed semi-automatically or automatically using the staining system 320.

[0052] Staining may involve exposing individual sections of tissue to one or more different stains (e.g., sequentially or simultaneously) to express different characteristics of the tissue. For example, each section may be exposed to a predetermined amount of stain for a predetermined period of time. The stains may include (e.g.) RNA probes, protein probes (e.g., nuclear protein probes or cytoplasmic protein probes), immunohistochemical stains, secretion probes, etc. In some cases, the stain is one that stains KAPPA mRNA or LAMBDA mRNA.

[0053] One typical type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes) to stain tissue structures. Histochemical staining can be used to show general aspects of tissue morphology and / or cellular microanatomy (e.g., distinguishing the cell nucleus from the cytoplasm, showing lipid droplets, etc.). An example of a histochemical stain is hematoxylin and eosin (H&E). Other examples of histochemical stains include trichrome stain (e.g., Masson's trichrome), Schiff periodate (PAS), silver stain, and iron stain. The molecular weight of histochemical staining reagents (e.g., dyes) is generally about 500 kilodaltons (kD) or less, although some histochemical staining reagents (e.g., Alcian blue, phosphomolybdic acid (PMA)) can have molecular weights up to 2000 or 3000 kD. An example of a high molecular weight histochemical staining reagent is α-amylase (approximately 55 kD), which is sometimes used to indicate glycogen.

[0054] Another type of tissue staining is immunohistochemistry (IHC, also called "immunostaining"), which uses a primary antibody that specifically binds to a target antigen of interest (also called a biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or fluorophore). In indirect IHC, the primary antibody first binds to the target antigen, and then a secondary antibody conjugated to a label (e.g., a chromophore or fluorophore) binds to the primary antibody. Because antibodies have a molecular weight of approximately 150 kD or more, the molecular weight of IHC reagents is much larger than that of histochemical staining reagents.

[0055] The sections can then be individually mounted on corresponding slides, and the imaging system 325 can then scan or image them to generate raw digital pathology images 330a-n. Each section may be mounted on a slide, then scanned to create a digital image, which can subsequently be analyzed and / or interpreted by a human pathologist (e.g., using image viewer software). The imaging may also include capturing bright-field images of the sections on the slides.

[0056] In some cases, pathologists can review digital images of slides and manually annotate them (e.g., tumor areas, necrosis, etc.). In other cases, annotation of regions of interest is performed automatically using computer vision technology.

[0057] A portion of the digital pathology images 330a-n can be used as training data by the intensity transducer training system 335 to generate a signal prediction function 340 that associates an estimate of a staining substance with the depicted staining intensity, and / or a confidence function 345 that associates the confidence of the estimate of a staining substance with the staining intensity. The intensity transducer training system 335 may be configured to use one or more of the techniques disclosed herein to learn the signal prediction function 340 and / or the confidence function 345. The intensity transducer training system 335 may be configured to use the training data to detect a cutoff at which the relationship between staining intensity (e.g., the number of dots, the intensity of a given color, or a color-independent intensity) and staining concentration transitions from linear to saturation. The intensity transducer training system 335 can predict that the relationship is within a linear range. The cutoff and / or relationship may be determined separately for each staining substance and each tissue type (e.g.). In some examples, the cutoff and / or relationship may be further determined separately for each of one or more demographic and / or disease groups (e.g., separately for each age / sex group for each of multiple cancer staging groups). The intensity transducer training system 335 can infer that this linear relationship of the signal prediction function 340 applies across the entire range of detected staining intensities. However, the intensity transducer training system 335 may further use a confidence function to predict how the confidence of the predicted signal intensity changes across the detected staining intensities. The confidence may be constant and / or change linearly across the entire first lower portion of the signal intensity. However, within the second higher portion of the signal intensity (after the cutoff), the confidence relationship may change. For example, the confidence may be lower and / or the function used to predict the confidence may be based primarily on a nonlinear function.

[0058] The image processing system 350 can be configured to receive other images 330a-n and process each of the other images using a signal prediction function 340 and a confidence function 345 to generate a corresponding transformed image (of transformed images 355a-k). Each transformed image can show a predicted signal intensity and / or confidence metric (which may optionally or additionally include an error metric) for a given pixel, region, or the image itself. The predicted signal intensity may be based on a linear signal prediction function. The confidence metric may be based on a biomodal function, where the confidence may be constant or linear between the first part of the detected staining intensity, but nonlinear between the second part of the detected staining intensity. Each transformed image 355a-k can identify the predicted signal intensity and confidence for each pixel or region.

[0059] The image processing system 350 can generate one or more of the metrics 360a to k. Each metric 360a to k may correspond to (e.g.) a region of interest, a volume of interest, an image, and / or a subject. The image processing system 350 may output one or more transformed images 355a to k and / or one or more metrics 360a to k to a user device that can be operated (e.g.) by a caregiver or subject. The transformed images 355a to k may be presented via a GUI in which detected staining signals (e.g., detected dots or blobs) are superimposed on the original digital pathology image. The GUI may be configured to receive input for adding or removing one or more dots or blobs.

[0060] The output may be used to generate predicted diagnoses, prognoses, or treatment recommendations. In some cases, the image processing system 350 may use one or more rules or protocols to convert metrics 360a-k into treatment recommendations, potential diagnoses, or potential prognoses. [Examples]

[0061] Example 1 - Artificial ground truth assessment performed at the cell population / tissue area level A set of digital pathology slides of breast cancer tissue was accessed, and 15 field-of-view (FOV) images from tumor regions were randomly selected. The total number of spots (including isolated spots and signal aggregates) from all 15 FOVs was determined for each slide. For each selected slide, the sum was also calculated for all concentrations across other different slides by locating the FOV position to similar regions with very similar tissue morphology. Figure 4 shows all slide images with 15 rectangles located in tumor regions, and Figure 5 shows exemplary results of breast FOV images for all probe concentrations: 0 pM (or "no probe control (NPC)"), 0.625 pM, 0.125 pM, 0.25 pM, 0.5 pM, 1 pM, 2 pM, 4 pM, 8 pM, 16 pM, and 32 pM, respectively, with each breast FOV image overlaid with the number of superpixels (green segments), isolated spots (red dots), and signal aggregate blobs (blue numbers) according to our processing method. Superpixels were used for visualization to estimate aggregated dots by segmenting regions of similar intensity. Details of how superpixels are applied are described in U.S. Provisional Application No. 17 / 586,982, filed January 28, 2022, which is incorporated herein by reference in its entirety for all purposes.

[0062] Similar data were collected using digital pathology slides of prostate tissue and digital pathology slides of colorectal cancer (CRC).

[0063] A set of tonsil tissue slides was accessed. The 15 rectangles shown in Figure 6 identify the tumor region on the tissue slides. The different images shown in Figure 6 show slides of a given tonsil stained with 11 concentrations of probe.

[0064] For each of the four cancerous tissue types (breast, prostate, CRC, and tonsil) and each of the 13 concentrations, the number of isolated dots detected was identified, as well as the number of isolated blobs (or “aggregates”). Figure 7 shows graphs of these amounts against concentration, and lines showing the total number of isolated dots and blobs detected. These graphs extend only to 8 pM, due to embodiments of the present invention that provide a framework for estimating that the digital pathology signal linearly reflects biomarker levels up to a saturation point of 8 pM.

[0065] A linear fit was constructed to the sum of the counts of isolated dots and blobs detected against concentration. The linear fit was observed to be strong in breast, prostate, and CRC tissue. The linear fit was less favorable in tonsil tissue, which may be because tonsil tissue is rich in relatively small immune cells, allowing the stain to aggregate more rapidly.

[0066] In particular, given that the amount of detected staining agent is known to change linearly at different concentrations, the experiments performed produced artificial truth data. However, it was unclear whether the detected signals would exhibit this linearity, given that staining agents cannot be absorbed in a consistent manner and stained aggregates (e.g., blobs) can obscure what the true biomarker level is. Figure 8 shows the “true biomarker intensity” as a function of concentration for each tissue type, which is defined as a linear fit from Figure 7. In particular, the x-axis of the graph in Figure 8 extends beyond the x-axis of the graph in Figure 7 into the concentration region beyond the saturation point.

[0067] Figure 8 also shows the dot counts detected using standard digital pathology techniques. The error bars in the figure indicate the error of the detected dot counts relative to the true biomarker level. For breast, prostate, and CRC tissue data, the dot count error was very small up to the cutoff point, but thereafter became very large at higher concentrations. For tonsil tissue, the error after 8 pM concentration was still relatively large, but the error of the signal detected via digital pathology techniques was relatively high at lower concentrations compared to the errors of other tissue types.

[0068] Example 2 - Artificial ground truth assessment performed at the cellular level This example relates to RNA dot counting in individual cells. Automated nuclear detection was established to detect nuclei in field-of-view images based on a modified radial symmetry method. The two images on the left of Figure 9 show the nuclear detection results of the modified radial symmetry method superimposed on the original image (top left) and the strongly stained image (bottom left) without probe control. The two images on the right show the labels of the nuclear detection results superimposed on the original image (top right) and the strongly stained image (bottom right) without probe control.

[0069] As shown in Figure 9, nuclei were detected and counted in the partial images. As a result, the average RNA dots within a single cell can be reported (RNA dots per cell). This statistic can be generated using RNA dots identified for cell type (within the sample) or only for tumor cells (within the sample).

[0070] Using breast, prostate, CRC, and tonsil tissue, spot signals per cell counted by the DP algorithm at different concentrations were plotted, as shown in Figure 10, along with the ideal biomarker expression line and the error rate in the dot count algorithm. Similar to the characteristics of total spot counts, the true biomarker representation is a linearly fitted line of spot counts between 0 pM and 8 pM, which is a linear relationship concentration range, but with a smaller error rate.

[0071] After investigating the characteristics of RNA dot counts in different tissues, RNA dot counts were investigated only within tumor regions. Automated image analysis was applied to classify tumor cells from stromal cells. Figure 11 shows an example of tumor cell classification (red dots) from non-tumor cells (green dots).

[0072] A framework for segmenting tumor regions from non-target regions was established. The top two rows of images in Figure 12 show the framework for generating tumor masks, starting with cell-level classification, tumor masks, and polygon generation, for grouping objects in tumor-labeled images. The bottom row shows the original image, which is a slide stained at concentrations of 16 pM and 0 pM (NPC) overlaid with tumor polygons.

[0073] In particular, given that the amount of detected staining agent is known to change linearly at different concentrations, the experiments performed produced artificial truth data. However, it was unclear whether the detected signals exhibited this linearity, given that staining agents may not be absorbed in a consistent manner and staining aggregates (e.g., blobs) can obscure what the true biomarker level is. While Example 1 demonstrated linearity when evaluating staining levels expressed across cell populations, this experiment explores how staining agent expression can be interpreted in intracellular contexts.

[0074] Figure 13 shows four plots for prostate tissue, with the top two plots relating to the total number of spots across all cells (left) or tumor cells (right). The bottom two plots relating to the number of spots per cell (left) or per tumor cell (right). "True biomarker intensity" is defined as a straight line, which is a linear fit corresponding to concentrations from 0 to 8 pM. Figure 13 also shows the dot counts detected using standard digital pathology techniques. The error bars in the figure indicate the error of the detected dot counts relative to the true biomarker levels. For all four variables shown in Figure 13, the error in dot counts was very small up to the cutoff point, but became very large thereafter at higher concentrations.

[0075] Therefore, the data demonstrates that signal intensity can be accurately correlated to predicted biomarkers up to the saturation point using a linear function (generated by fitting data that correlates signal data to concentrations across pre-saturated concentration data). The linear function can then still be used to indicate the confidence or error of such biomarker predictions (the error can be quite large).

[0076] Example 3 - Spot characteristics at different concentrations The analysis described above focuses on quantitatively predicting the presence of biomarkers using intensity signals from digital pathology images. When examining individual spots (or potentially even blobs), other features of isolated spots may be useful in relation to underlying biomarker levels. Therefore, for each isolated spot, metrics characterizing the size, blurring, and roundness of the spot's signal were determined (in addition to the spot's intensity) across staining concentrations. In particular, the intensity of the drawn box plots indicates the intensity of isolated dots rather than aggregated dots.

[0077] The metrics characterizing spot blurring and size did not change with staining concentration. However, from concentrations of 0.0625 pM to 1 pM, the roundness feature formed a shape close to a perfect circle (=1). However, upon reaching a concentration of 2 pM, the roundness feature began to distribute closer to an imperfect circle (=0), and became more uniform as the standard deviation increased after a concentration of 8 pM. Therefore, using the techniques disclosed herein, it is possible to generate a linear function relating predicted roundness to concentration levels, identify a saturation point applicable to roundness prediction, and / or generate an error or confidence metric for biomarker prediction based on the spot roundness metric. Exemplary techniques for estimating the characterization of intensity, size, roundness, and / or blurring features are described in U.S. Patent Application No. 17 / 586,982, filed on 28 January 2022, which is incorporated herein by reference in its entirety for all purposes.

[0078] Some embodiments of the present disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-temporary computer-readable storage medium containing instructions that, when executed by one or more data processors, cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-temporary machine-readable storage medium, which contains instructions configured to cause one or more data processors to execute some or all of the methods disclosed herein and / or some or all of one or more processes.

[0079] The terms and expressions used are for illustrative purposes only, not limitation, and in using such terms and expressions there is no intention to exclude equivalents or parts of the features shown and described, however it is acknowledged that various modifications are possible within the scope of the invention described in the claims. Accordingly, although the claimed invention is specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be used by those skilled in the art, and such modifications and variations will be considered to fall within the scope of the invention as defined by the appended claims.

[0080] The description presents only preferred typical embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the description of preferred typical embodiments presents to those skilled in the art a description that enables various embodiments to be realized. It will be understood that the function and arrangement of the elements can be varied without departing from the idea and scope described in the appended claims.

[0081] In the following description, specific details are given to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid unnecessarily obscuring the embodiments with excessive detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

Claims

1. A computer implementation method, Accessing digital pathology images depicting slides having slices of a sample stained with a staining agent, wherein the digital pathology images were acquired using bright-field imaging. To detect staining intensity corresponding to at least a portion of the aforementioned digital pathology image, Accessing a linear biomarker intensity prediction function that linearly correlates the predicted level of biomarker intensity with the detected intensity of the staining agent, wherein the biomarker intensity prediction function is generated by evaluating digital pathology images of other slides, the other slides containing samples stained with multiple other concentrations of the staining agent, Accessing a confidence function that relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, wherein the confidence function is nonlinear. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the linear biomarker intensity prediction function, generate predicted biomarker intensities for at least a portion of the slide. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the confidence function, a confidence metric for the predicted biomarker intensity is generated, and Outputting results based on the predicted biomarker intensity and the confidence metric. Computer implementation methods, including those mentioned above.

2. The computer implementation method according to claim 1, wherein the confidence function includes a first portion that linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, and the confidence function includes a second portion that non-linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent.

3. The computer implementation method according to claim 2, wherein the second part of the confidence function corresponds to the saturation of the detected intensity of the dye.

4. The slide, at least a portion of which is a pixel, and the method, For each of the other sets of pixels in the digital pathology image, generate a different predicted biomarker intensity for the other pixel based on the detected staining intensity corresponding to the other pixel and based on the linear biomarker intensity prediction function. For each of the other sets of pixels in the digital pathology image, generate another confidence metric for the other predicted biomarker intensity based on the detected staining intensity corresponding to the other pixel and based on the confidence function, and To generate the results based on the predicted biomarker intensity, the other predicted biomarker intensity, the confidence metric, and the other confidence metric. The computer implementation method according to claim 1, further comprising:

5. Based on the aforementioned confidence metric, it is determined that the stored criteria are met, and Based on the determination that the stored criteria are met, the result is generated in a form that incorporates the predicted biomarker intensity. The computer implementation method according to claim 1, further comprising:

6. The computer implementation method according to claim 1, wherein the staining agent is an RNA staining agent.

7. The computer implementation method according to claim 1, wherein the staining agent is a staining agent for nucleoproteins.

8. The computer implementation method according to claim 1, wherein the staining agent is a staining agent for cytoplasmic proteins.

9. It is a system, One or more data processors, A non-temporary computer-readable storage medium that includes instructions causing one or more data processors to execute a set of operations when executed by one or more data processors, wherein the set of operations is Accessing digital pathology images depicting slides having slices of a sample stained with a staining agent, wherein the digital pathology images were acquired using bright-field imaging. To detect staining intensity corresponding to at least a portion of the aforementioned digital pathology image, Accessing a linear biomarker intensity prediction function that linearly correlates the predicted level of biomarker intensity with the detected intensity of the staining agent, wherein the biomarker intensity prediction function is generated by evaluating digital pathology images of other slides, the other slides containing samples stained with multiple other concentrations of the staining agent, Accessing a confidence function that relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, wherein the confidence function is nonlinear. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the linear biomarker intensity prediction function, generate predicted biomarker intensities for at least a portion of the slide. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the confidence function, a confidence metric for the predicted biomarker intensity is generated, and Outputting results based on the predicted biomarker intensity and the confidence metric. A system that includes this.

10. The system according to claim 9, wherein the confidence function includes a first portion that linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, and the confidence function includes a second portion that nonlinearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent.

11. The system according to claim 10, wherein the second portion of the confidence function corresponds to the saturation of the detected intensity of the dye.

12. The slide, at least a portion of which is a pixel, and the set of actions is For each of the other sets of pixels in the digital pathology image, generate a different predicted biomarker intensity for the other pixel based on the detected staining intensity corresponding to the other pixel and based on the linear biomarker intensity prediction function. For each of the other sets of pixels in the digital pathology image, generate another confidence metric for the other predicted biomarker intensity based on the detected staining intensity corresponding to the other pixel and based on the confidence function, and To generate the results based on the predicted biomarker intensity, the other predicted biomarker intensity, the confidence metric, and the other confidence metric. The system according to claim 9, further comprising:

13. The aforementioned set of operations is Based on the aforementioned confidence metric, it is determined that the stored criteria are met, and Based on the determination that the stored criteria are met, the result is generated in a form that incorporates the predicted biomarker intensity. The system according to claim 9, further comprising:

14. The system according to claim 9, wherein the staining agent is an RNA staining agent.

15. The system according to claim 9, wherein the staining agent is a staining agent for nucleoproteins.

16. The system according to claim 9, wherein the staining agent is a staining agent for cytoplasmic proteins.

17. A computer program product tangibly embodied in a non-temporary machine-readable storage medium, which includes instructions configured to cause one or more data processors to execute a set of operations, wherein the set of operations is Accessing digital pathology images depicting slides having slices of a sample stained with a staining agent, wherein the digital pathology images were acquired using bright-field imaging. To detect staining intensity corresponding to at least a portion of the aforementioned digital pathology image, Accessing a linear biomarker intensity prediction function that linearly correlates the predicted level of biomarker intensity with the detected intensity of the staining agent, wherein the biomarker intensity prediction function is generated by evaluating digital pathology images of other slides, the other slides containing samples stained with multiple other concentrations of the staining agent, Accessing a confidence function that relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, wherein the confidence function is nonlinear. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the linear biomarker intensity prediction function, generate predicted biomarker intensities for at least a portion of the slide. Based on the detected staining intensity corresponding to at least a portion of the slide, and based on the confidence function, a confidence metric for the predicted biomarker intensity is generated, and Outputting results based on the predicted biomarker intensity and the confidence metric. Computer program products, including [this].

18. The computer program product according to claim 17, wherein the confidence function includes a first part that linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent, and the confidence function includes a second part that non-linearly relates the confidence of the predicted biomarker intensity to the detected intensity of the staining agent.

19. The computer program product according to claim 18, wherein the second portion of the confidence function corresponds to the saturation of the detected intensity of the dye.

20. The slide, at least a portion of which is a pixel, and the set of actions is For each of the other sets of pixels in the digital pathology image, generate a different predicted biomarker intensity for the other pixel based on the detected staining intensity corresponding to the other pixel and based on the linear biomarker intensity prediction function. For each of the other sets of pixels in the digital pathology image, generate another confidence metric for the other predicted biomarker intensity based on the detected staining intensity corresponding to the other pixel and based on the confidence function, and To generate the results based on the predicted biomarker intensity, the other predicted biomarker intensity, the confidence metric, and the other confidence metric. The computer program product according to claim 17, further comprising: