Staining de-mixing of multiple bright field images
By identifying and adjusting color vectors in digital pathological images, and combining GUI and machine learning models, the difficulty of demixing multiple stained images in existing technologies has been solved, achieving high-quality signal separation and accurate demixing under different environments.
Patent Information
- Application Number
- CN202480028156.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-28
- Filing Date
- 2024-04-26
- Publication Date
- 2025-11-28
AI Technical Summary
Existing digital pathology unmixing techniques have suboptimal results in multiple staining image processing, including blurring, missing signals, and noise. Especially in triple or more staining cases, it is difficult to accurately separate signals from different staining agents. Furthermore, unmixing methods depend on experimental results and lack versatility for different environments and imaging systems.
By determining the initial color vector associated with the digital pathology stain and using a graphical user interface (GUI) and machine learning models to fine-tune and adjust the color vector, combined with nonnegative matrix factorization (NMF) technology, synthetic single images are generated, improving unmixing performance.
High-quality staining unmixing was achieved under different imaging systems and environments, improving the accuracy and reliability of signal separation, reducing blurring and noise, and enhancing the precision of unmixing results.
Smart Images

Figure CN121039710A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 499,098, filed April 28, 2023, which is incorporated herein by reference in its entirety for all purposes. Background Technology
[0003] Digital pathology facilitates accurate diagnosis of patients and guides treatment decisions. In digital pathology solutions, image analysis workflows are used to automatically detect or classify target biological components, such as cells possessing one or more specific proteins or antigens. An exemplary digital pathology solution workflow includes: obtaining a tissue slide; scanning a pre-selected region or the entire tissue slide with a digital image scanner (e.g., a whole-slide imaging (WSI) scanner) to obtain a digital image; and performing image analysis on the digital image. Using one or more image analysis algorithms to process the digital image facilitates the detection of cells labeled with one or more target signals and the quantification of such signals using image analysis (e.g., quantitative or semi-quantitative scoring, such as positive, negative, moderate, weak, etc.).
[0004] Digital pathology can utilize singleton or multiplex techniques. Singleton uses a single staining agent for only one biomarker, while simultaneously using a reference staining agent. Multiplex, on the other hand, involves staining two or more biomarkers (in addition to a reference staining agent) in a single slide or tissue sample. Therefore, multiplex techniques support the simultaneous detection of multiple biomarkers and their co-expression at the single-cell level. To process multiplex images, a demixing process can be performed to separate the signal from the different biomarkers. More specifically, color demixing methods can be used to decompose the RGB image into its individual component staining agents / dyes for each biomarker. This facilitates the estimation of the staining level of each of the multiple staining agents (for individual cells).
[0005] One exemplary unmixing technique is color deconvolution, which can be used to unmix signals in RGB images with up to three dyes in a transformed optical density space. (See Ruifrok AC, Johnston DA Quantification of histochemical staining by color deconvolution. Anal QuantCytol Histol. Aug 2001;23(4):291-9. PMID: 11531144, which is incorporated herein by reference in its entirety for all purposes). Another exemplary unmixing technique formulates the color unmixing problem as nonnegative matrix factorization (NMF) and performs color decomposition in a fully automated manner, where reference dye color selection is not required. (See Lee, Daniel and H. Sebastian Seung. Algorithms for non-negative matrix factorization. Advances in neural information processing systems 13 (2000), which is incorporated herein by reference in its entirety for all purposes). However, each of these techniques can produce suboptimal results, including blurring, missing signals, and noise. Unmixing methods become particularly challenging as the number of staining agents used increases. For example, in the triplet case (where three biomarker staining agents and one reference staining agent are present), unmixing methods attempt to convert a three-channel RGB image into a four-channel output. This can lead to inaccurate predictions.
[0006] These suboptimal results may be due to, for example, the colocalization of multiple staining agents in the same type of cellular part (e.g., multiple staining agents may colocalize in the nucleus, or multiple staining agents may colocalize in the cell membrane). For example, in multiplex imaging, a commonly used biomarker is hematoxylin, which is used to stain the cell nucleus, allowing pathologists to visualize tissue structures and determine which cells are negative for all biomarkers. Because cells are stained in our images, they are typically divided into three main regions: the nucleus, the cytoplasm, and the cell membrane. Biomarkers can be designed for any of these regions. In dual imaging, up to three staining colors may be present in a cell, and depending on the specific biomarker, more than one staining color may fall into one of these regions. This is known as colocalization. However, colocalization can cause pixels to appear as a different color associated with a single staining agent, and / or may make it difficult to estimate expression levels.
[0007] Suboptimal results may also be due to the application of staining agents across non-biomarker-specific regions of the cell. For example, if a purple staining agent adheres to the cytoplasm or cell membrane, it will typically add purple to the nucleus even if the purple staining agent is not specific to that part of the cell. Furthermore, current unmixing techniques typically rely on experimental results to identify a color vector for each staining agent. However, the true color vector can vary depending on tissue type, illumination, specific imaging device, etc. Therefore, it would be beneficial to identify a new unmixing method that more reliably delivers higher-quality results. Summary of the Invention
[0008] Some embodiments of this disclosure relate to staining unmixing of digital pathology images by determining an initial color vector associated with a digital pathology stain and adjusting the color vector via a graphical user interface (GUI). The computer-implemented method involves determining a color vector associated with (e.g., at least three) digital pathology stains (e.g., chromogens or fluorophores) from a given digital pathology image acquired using bright-field imaging. To facilitate the determination of the color vector, each pixel of the digital pathology image can be mapped from the RGB space to a location in the optical density (OD) space. Although each fluorophore used to stain the sample can be associated with a predefined color vector, using these color vectors for the unmixing process may lead to suboptimal results. This can be due to differences in imaging systems, lighting, etc., between facilities. For example, even when using a “green” dye configured to have only green (and no red or blue components), the ambient lighting or imaging system associated with a particular facility may result in the image having a certain amount of red and / or blue intensity in the stained areas.
[0009] Therefore, fine-tuning or adjusting the color vectors can help improve staining unmixing performance. Adjustments can be made via an interface and / or automated techniques that include: a visual / picture representation of the determined color vectors in a color space (e.g., in Hue Tone Saturation Density (HSD) space); and, hereinafter, a real multiplex digital pathology image depicting a biopsy slide stained with at least three digital pathology stains associated with the determined color vectors. The interface may also include a synthetic singlex image associated with each stain in the multiplex image. The synthetic singlex image can be generated by filtering the multiplex image using the determined color vectors. For example, a tool could allow a user to adjust the color vector for a “green” stain to include an element of red and / or blue, which better captures the light component of the “green” signal in a given image and can facilitate unmixing of the given image.
[0010] Additionally, the interface can provide one or more color adjustment tools, allowing users to interactively adjust or fine-tune each defined color vector. Input received through user interaction with the interface can be detected, particularly targeting adjustments to defined color vectors associated with a specific dye. In response to the detection of this input, the interface can be automatically updated (e.g., to indicate the orientation / proximity of the changed defined color vector in the representation space, and / or to show exemplary unmixing results produced using the adjusted color vector). This process can support precise and responsive customization of dyes based on multiple images provided via the interface, thereby improving dye unmixing performance.
[0011] In some cases, one or more single-stain (or solid-color) images can be used to determine one or more color vectors. These images, compared to multiple images, can depict the same biopsy slide or different biopsy slides. A single-stain image can be stained with only one of the digital pathology stains used to stain multiple images; however, the single-stain image can be captured in the same environment and using the same imaging system, just like multiple images. The representation of each of the determined color vectors can also include markers overlaid at specific locations within the digital pathology image. The specific locations of the markers can be used to define the color vector for the corresponding stain (or initial color vector, which can then be adjusted based on user input or further processing). In one aspect, the color vectors can be determined using nonnegative matrix factorization (NMF).
[0012] Visualizing the synthesized single images through the interface can help verify the correctness of the adjusted color vectors. The synthesized single images can be regenerated using updates to the interface corresponding to the determined adjustments to the color vectors. These synthesized single images can be further combined to generate synthesized multiimages, which can be compared with real multiimages during training and / or as an indicator of the confidence level of the adjusted color vectors. For example, a graphical user interface can present both the synthesized and real multiimages, and user input adjusting one or more color vectors can trigger dynamic updates to the synthesized multiimages.
[0013] In some embodiments, filtering can be performed by utilizing one or more machine learning models, such as generative models, which can be trained to learn a mapping from a given multi-image to a constituent single-image conditioned on one or more color vectors (defined using one or more techniques disclosed herein).
[0014] Once the color vectors are defined (e.g., using one or more techniques disclosed herein), they can be used to unmix new multiple images to generate a set of synthetic single images (each corresponding to a given biomarker stain or reference stain). The new multiple images can be stained using the same stains as those used in the interface and / or automation techniques for color adjustment. In some cases, the fine-tuned color vectors can be determined based on, for example, a multiple image different from the new multiple image used for further staining and unmixing. The fine-tuned color vectors can then be used to generate one or more synthetic single images from the new multiple images. In some cases, these synthetic single images are generated from the same multiple image used in the interface and / or automation techniques to support the fine-tuning of the color vectors.
[0015] In some embodiments, this disclosure provides a method for determining an initial color vector associated with a specific stain based on a given real multiplex image that may not include that stain. The computer-implemented method includes determining an initial color vector associated with, for example, at least three stains from a corresponding pure stain digital pathology image. The method may further include accessing a real multiplex image stained with one or more stains associated with the initial color vector, but excluding at least one of those stains. For reference purposes, at least one stain not present in the real multiplex image is referred to as a “specific stain.” The initial color vector can be fed to a filter configured to generate a filtered output from the real multiplex image based on the specific stain. Filtering can be performed using one or more machine learning models, such as generative models, trained to learn a mapping from a given multiplex image to its constituent singlex images. Generative models, such as GANs, can be conditional on the color vector such that if the model encounters a color vector not present in the input multiplex image, the model may fail to generate any meaningful output associated with that stain, resulting in a zero-valued or empty image. Conversely, if such a model is given a color vector present in a given multi-image, the model can generate a constituent single-image associated with that color vector.
[0016] To assess the quality of the filtered output, a metric can be calculated that characterizes the degree of variation in staining intensity across all or part of the filtered output. For example, when a biomarker corresponding to a given staining agent (or color vector) is predicted (or known) not to exist in a real multiplex image, this metric (e.g., mean, median, mode, variance, standard deviation, and / or range) can be expected to be relatively low when conditioned on an accurate color vector, compared to when using less accurate color vectors. Once a quantitative metric for the filtered output based on the degree of staining agent presence in the real multiplex image is determined, spatial traversal techniques (e.g., gradient descent, Monte Carlo methods) can be used to discover color vector adjustments associated with a specific staining agent. These adjustments can be incorporated into the color vector of the specific staining agent via an interface.
[0017] Once the color vector associated with a specific dye has been adjusted by minimizing the metric, a new multiple image can be received. This multiple image can be stained with at least one of the dyes associated with the initial color vector. The multiple image may also include one or more specific dyes whose color adjustment is calculated based on metric and spatial visit techniques. By utilizing a staining unmixing process with unmixing techniques such as NMF, a new synthetic single image associated with the specific dye can be generated. Finally, the generated unmixed output can be displayed via an interface.
[0018] In another example, techniques can be provided to discover recommended color vectors for a given multiply image. For example, a dual image stained with two specific stains and a counterstain (e.g., hematoxylin). The aim is to transform the dual image into a triple image by identifying potential additional stains that are distinguishable among the existing stains. A computer-implemented method may include determining an initial color vector associated with, for example, at least two stains, from a corresponding pure stain digital pathology image. The method may further include accessing a true multiply image stained with at least two stains associated with the initial color vector. Additional stains can be selected such that they are uncorrelated with any stains in the multiply image's stains in a multidimensional color space.
[0019] Initial staining agents can be selected or engineered, and associated initial color vectors can be determined using techniques such as NMF (Non-Multiple Facing) to determine them. Using this initial color vector, real multiple images can be filtered using a machine learning model configured to map a given multiple image to one of its constituent synthetic single images based on the provided color vector. To characterize the filtered output, metrics (e.g., mean, mode, or median) can be computed, quantifying the amount of staining agents present in the given multiple images. For example, if the selected color staining agents are not distinguishable, the computed metrics (such as the mean of the synthetic OD single image) can have high values, indicating the presence or similarity to existing staining agents. Metrics can be estimated using spatial ergodic techniques that may include one or more targets in the ergodic process. Corresponding adjustments to the initial color vector can be discovered based on spatial ergodic techniques that minimize the metrics for the filtered output. Finally, recommended color vectors unrelated to staining agents already present in the multiple images can be output via an interface.
[0020] Other aspects of this disclosure include a method for determining a performance prediction score that represents the predicted degree to which at least three digital pathology stains are sufficiently separable in practice to reliably support the generation of a synthetic single-weighted image. The method may include determining an initial color vector associated with, for example, at least three stains, from a corresponding pure-stain digital pathology image. A real multiple image stained with one or more stains associated with the initial color vector, but excluding at least one of these stains (referred to as a “specific stain”), may be accessible. The initial color vector may be fed to a filter configured to generate a filtered output from the real multiple image based on the specific stain. One or more machine learning models, such as generative models, may be trained to learn a mapping from a given multiple image to its constituent single-weighted image for filtering purposes. If such a model is conditioned on a color vector not present in the input multiple image, the model may fail to generate any meaningful output associated with this stain, resulting in zero values or an empty image. The performance prediction score may be generated for the filtered output and / or for other synthetic single-weighted images constituting the real multiple image. It can output a performance prediction score, and the initial color vector can be adjusted based on this performance prediction score via the GUI.
[0021] In some cases, performance prediction scores may include the mean intensity, median intensity, or mode intensity of the corresponding filtered output. For example, when a biomarker corresponding to a given stain is predicted (or is known) to exist in a given depicted sample or multiple images, performance prediction scores (e.g., mean, median, mode, variance, standard deviation, and / or range) may be expected to be relatively high when using accurate color vectors compared to when using less accurate color vectors.
[0022] In another scenario, the performance prediction score can be estimated by grouping or clustering similar stains together based on staining features from one or more single-image datasets. For example, staining features could include optical density values, color histograms, or any other features that can effectively capture staining patterns. In other examples, the performance prediction score can be computed for a synthetic single-image dataset by estimating the correlation between each staining pattern observed in the multiple images.
[0023] In some aspects, staining unmixing is performed using constraint methods that can reduce the complexity of multi-images (e.g., stained with four staining agents), thereby supporting more accurate and / or reliable generation of synthetic single-images from multi-images. Computer-implemented methods include determining initial color vectors associated with, for example, at least four staining agents from corresponding pure-stained digital pathology images. In some cases, each pixel in the digital pathology image can be mapped to a location within a multi-dimensional color space. Among these four staining agents, a specific staining agent can be selected such that it can be attributed to a prominent portion of the color space (e.g., quadrants, portions defined by values greater than / less than a certain y-value and greater than / less than a certain x-value, wedge-shaped regions, cylindrical regions, etc.). The method further includes accessing a real multi-image stained with at least three digital pathology staining agents. Each pixel of the real multi-image can also be mapped to a point in the multi-dimensional space. For each pixel, a pixel-specific vector can be generated that predicts the expression level of each of the at least four staining agents in the portion of the biopsy slide depicted at that pixel.
[0024] The process of generating pixel-specific vectors can further include assigning pixels within a specific portion of the color map to a specific color vector that predicts the expression level of the biomarker corresponding to that portion. For each pixel associated with the specific portion, an optical density can be determined. For each additional biomarker corresponding to the multiple images, pixels outside this portion can be assigned "0" (or other predefined expression level). For pixels outside the first portion, demixing techniques can be used to predict the expression level of each of the other biomarkers, and the predicted expression level "0" (or other predefined number) can be associated with the first biomarker. Finally, one or more synthetic single images can be generated using the pixel-specific color vectors.
[0025] Specific staining agents can be selected based on information about which parts of the cell each of at least four digital pathology staining agents is configured to stain. The color space may include the International Commission on Illumination (CIE) color space. A portion of the color space may include a wedge-shaped region. A portion of the color space may include a part of the space defined based on inequalities about the x-coordinate and inequalities about the y-coordinate. A portion of the color space may include a combination of primitives. Demixing techniques may include the use of nonnegative matrix factorization (NMF). Color vectors can be determined based on one or more user inputs received using one or more color-vector adjustment tools available within the interface.
[0026] In some cases, a computer-implemented method is provided, comprising: determining a color vector representing each of at least three digital pathology stains; making an interface available to a user device, wherein the interface includes: a representation of each of the determined color vectors; a real multiple digital pathology image depicting a biopsy slide stained with two or more of at least three digital pathology stains; at least one synthetic single image, wherein each of the at least one synthetic single image is generated by filtering the real multiple digital pathology image using a single color vector of the determined color vectors; and one or more color-vector adjustment tools, wherein each of the one or more color-vector adjustment tools is configured to receive user input corresponding to an adjustment to a color vector representing a corresponding stain among the at least three digital pathology stains; detecting input received via interaction with the interface, the input corresponding to a specific adjustment to a color vector representing a particular stain among the at least three digital pathology stains; and automatically updating the interface in response to the detection of input.
[0027] Each of the determined color vectors is represented by a position within the optical density space. The updated interface may further include at least one synthetic single-color image. One or more color-vector adjustment tools may include at least three color adjustment tools. Determining the color vectors may include processing one or more single-stain images depicting the same biopsy section or other biopsy section stained with only one of at least three digital pathology stains. One or more single-stain images may include markers overlaid at specific locations within a multidimensional color space, and wherein the color vectors are defined based on those specific locations. The true multiple digital pathology image may depict a biopsy section stained with at least four stains. The determined color vectors may be in a two-dimensional color space, and wherein the method further includes: determining a portion of the color space that is predicted to be attributable to a specific stain among the at least three digital pathology stains; wherein automatic updating of the interface is performed using a demixing technique that selectively focuses on at least three digital pathology stains minus the specific stain. The determination of the color vectors may be performed using nonnegative matrix factorization. The method may further include: receiving a new multiplex image stained with at least one of at least three digital pathology stains; generating a new synthetic singlex image based on the new multiplex image and an adjusted color vector; and outputting the new synthetic singlex image. The real multiplex digital pathology image can be filtered using color vectors and a machine learning model.
[0028] In some embodiments, a computer-implemented method is provided, the method comprising: determining a color vector representing each of at least three digital pathology stains; accessing a real multiplex digital pathology image depicting a biopsy section stained with at least one first stain of at least three stains, wherein the depicted biopsy section is not stained with at least one second stain of at least three stains; generating a filtered output by filtering the real multiplex digital pathology image using a color vector representing the second stain of at least three stains; generating a metric characterizing signal characteristics in the filtered output; identifying adjustments to the color vector representing the second stain using a metric and spatial traversal technique; receiving a new multiplex image stained with at least one of at least three digital pathology stains; generating a new synthetic singlex image based on the new multiplex image and the adjusted color vector representing the second stain; and outputting the new synthetic singlex image.
[0029] For each of at least three digital pathology stains, the color vector can be a vector in optical density space. Spatial traversal techniques can include gradient descent. Spatial traversal techniques can include Monte Carlo techniques. Metrics can include mean intensity, median intensity, or mode intensity. Metrics can characterize all or part of the staining level across the filtered output. The filtered output can be generated using a machine learning model.
[0030] In some embodiments, a computer-implemented method is provided, the method comprising: determining a color vector representing each of at least two digital pathology stains; accessing a real multiplex digital pathology image depicting a biopsy slide stained with at least two digital pathology stains; identifying a recommended color vector representing a potential additional stain by: generating a filtered output by filtering the real multiplex digital pathology image using an initial color vector; generating a metric characterizing signal properties in the filtered output; and identifying the recommended color vector using a metric and spatial traversal technique; and outputting the recommended color vector.
[0031] Spatial traversal techniques can be performed to include minimizing the signal in the filtered output as one or more objectives in the traversal. Minimizing the signal in the filtered output can include minimizing the mean intensity, median intensity, or mode intensity of the corresponding filtered output. The color vector can be determined using nonnegative matrix factorization. For each of at least two digital pathology stains, the color vector can be a vector in optical density space. The filtered output can be generated using a machine learning model. Spatial traversal techniques can include gradient descent techniques.
[0032] In some embodiments, a computer-implemented method is provided, comprising: determining a color vector representing each of at least three digital pathology stains; accessing a real multiplex digital pathology image depicting a biopsy section stained with at least one first stain of at least three digital pathology stains, wherein the depicted biopsy section is not stained with at least one second stain of at least three stains; generating a filtered output by filtering the real multiplex digital pathology image using a color vector representing the second stain of at least one second stain; generating a performance prediction score representing the predicted degree to which the at least three digital pathology stains are sufficiently separable in practice to reliably support the generation of a synthetic single image; and outputting the performance prediction score.
[0033] Performance prediction scores can be generated using the filtered output. The performance prediction score may include the mean intensity, median intensity, or mode intensity of the corresponding filtered output. The performance prediction score includes the correlation coefficient between each pair of synthetic single images associated with the filtered output. For each of at least three digital pathology stains, the color vector can be a vector in optical density space. The filtered output can be generated using a machine learning model. The color vector can be adjusted based on the performance prediction score via a graphical user interface (GUI).
[0034] In some embodiments, a computer-implemented method is provided, the method comprising: determining a color vector representing each of at least four digital pathology stains, wherein the determined color vectors are in a multidimensional color space; selecting a specific stain from the at least four digital pathology stains; determining a portion of the color space attributable to a prominent signal corresponding to the specific stain; accessing a real multiplex digital pathology image depicting a biopsy slide stained with at least three of the at least four digital pathology stains, wherein the real multiplex digital pathology image includes a set of pixels; mapping each pixel in the pixel set of the real multiplex digital pathology image to a point in the multidimensional color space; and generating a pixel-specific color vector for each pixel in the pixel set, the pixel-specific color vector predicting the expression of the stain in a portion of the biopsy slide depicted at that pixel for each of the at least four digital pathology stains. The degree, wherein generating a pixel-specific color vector includes: determining that each pixel in a first subset of the pixel set is mapped to a point within that portion of the color space; determining an optical density for each pixel in the first subset of the pixels, wherein the pixel-specific color vector for the pixel identifies the degree of expression of a specific stain corresponding to the optical density; determining that each pixel in a second subset of the pixel set is mapped to a point outside that portion of the color space; and performing a demixing technique to predict the degree of expression of the stain in the portion of the biopsy slide depicted at that pixel for each pixel in the second subset and for each of at least four digital pathological stains, wherein some of the at least four digital pathological stains do not include the specific stain, and wherein the demixing technique uses a color vector determined to represent each of the at least four digital pathological stains; and using the pixel-specific color vector to generate one or more synthetic single images.
[0035] Specific staining agents can be selected based on information about which parts of the cell each of at least four digital pathology staining agents is configured to stain. The color space may include the International Commission on Illumination (CIE) color space. A portion of the color space may include a wedge-shaped region. A portion of the color space may include a part of the space defined based on inequalities about the x-coordinate and inequalities about the y-coordinate. A portion of the color space may include a combination of primitives. Demixing techniques may include the use of nonnegative matrix factorization (NMF). Color vectors may be or may have been determined based on one or more user inputs received using one or more color-vector adjustment tools available within the interface. In some embodiments, a computer program product tangibly embodied in a non-transitory machine-readable storage medium includes instructions configured to cause one or more data processors to perform part or all of the one or more methods or processes disclosed herein.
[0036] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the methods disclosed herein.
[0037] In some embodiments, a system is provided that includes one or more means for performing some or all of the methods or processes disclosed herein.
[0038] The terms and expressions used are descriptive rather than restrictive, and in using such terms and expressions, no equivalents of the features shown and described or portions thereof are intended to be excluded; however, it should be recognized that various modifications are possible within the scope of the claimed invention. Therefore, it should be understood that although the claimed invention has been specifically disclosed by way of examples and optional features, those skilled in the art can employ modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. Attached Figure Description
[0039] This patent or application document contains at least one color drawing. Upon request and payment of the necessary fees, the Patent Office will provide a copy of this patent or application publication with one or more color drawings. This disclosure is described in conjunction with the accompanying drawings:
[0040] Figure 1 illustrates a workflow for acquiring and processing multiple images according to some embodiments of the present disclosure.
[0041] Figure 2 shows an exemplary network used to generate digital pathology images.
[0042] Figure 3A illustrates an illustrative example of a workflow that facilitates the definition of a color vector associated with a dye and the unmixing of dyes according to embodiments of the present disclosure.
[0043] Figure 3B illustrates exemplary linear demixing techniques according to some embodiments of the present disclosure.
[0044] Figure 3C shows an interface component that facilitates fine-tuning of the color vector.
[0045] Figure 3D illustrates an exemplary architecture for generating one or more synthetic images by leveraging multiple machine learning models.
[0046] Figure 3E illustrates an exemplary architecture for generating one or more synthetic images by leveraging a single machine learning model.
[0047] Figure 4 illustrates a flowchart of an exemplary process according to some embodiments of the present disclosure, which facilitates the definition of one or more color vectors for color unmixing.
[0048] Figure 5 illustrates a system for determining an adjustment to an initial color vector based on a given multiple digital pathology image, according to some embodiments of the present disclosure.
[0049] Figure 6 shows a flowchart of an exemplary process for generating a synthetic single image using finely tuned color vectors.
[0050] Figure 7 shows a flowchart illustrating an exemplary process for identifying recommended color vectors.
[0051] Figure 8 illustrates a flowchart of an exemplary process for generating a performance prediction score, which represents the predicted degree to which the constituent digital pathology stains can be sufficiently separated to reliably support the generation of one or more synthetic single images.
[0052] Figure 9A shows an example of an ER-PR-HER2 triple image, where each pixel of the triple image is mapped to a position within a multidimensional color map.
[0053] Figure 9B is an illustration of staining and unmixing of an exemplary triple ER-PR-HER2 image from Figure 9A according to some embodiments of the present disclosure.
[0054] Figure 9C shows an example of coloring unmixing results using the disclosed constraint technique for ER-PR-HER2 triple images and one or more single images.
[0055] Figure 9D shows an example of staining and remixing results using the disclosed constraint technique for ER-PR-HER2 triple images and one or more single images.
[0056] Figure 9E shows the results of triple ER-PR-HER2 staining remixing according to some embodiments of the present disclosure.
[0057] Figure 10A shows a flowchart of an exemplary process for coloring and unmixing multiple images using the disclosed constraint techniques, according to some embodiments of the present disclosure.
[0058] Figure 10B further illustrates an exemplary flowchart of the components from Figure 10A.
[0059] Figure 11A depicts a comparison of color unmixing of a dual image using an initial color matrix and an adjusted color matrix according to an exemplary implementation.
[0060] Figure 11B depicts a comparison of color unmixing of another dual image and a single image using an initial color matrix and an adjusted color matrix according to an exemplary implementation.
[0061] Figure 12A shows an example of a dual image, which overlays a marker (candidate species) at each nucleus detected by automatic nucleus segmentation.
[0062] Figure 12B illustrates a comparison of cell nucleus segmentation results for hematoxylin images obtained by demixing the dual images using linear deconvolution and NMF techniques.
[0063] Figure 13 illustrates an exemplary graphical user interface (GUI) for generating composite pixels according to an exemplary implementation.
[0064] Figure 14A shows an exemplary GUI for evaluating a range of colors from the mixing of multiple dyes using composite pixels.
[0065] Figure 14B shows a comparison of one or more blended colors synthesized from dyes from different reagent sources according to an exemplary implementation.
[0066] Figure 14C shows an example of producing a range of colors by mixing two or more dyes according to an exemplary implementation.
[0067] Figure 14D shows the evaluation of a range of colors assigned to hematoxylin with wedge-shaped region constraints. Detailed Implementation
[0068] Some embodiments of this disclosure relate to the demixing of digital pathology images labeled with more than three markers (e.g., three or more biomarkers and a reference stain), wherein the digital pathology images have three or fewer channels (e.g., red, green, and blue channels). A color vector can be defined for each of the markers, and these color vectors can then be used for demixing techniques to separate the signals in the digital pathology images (corresponding to more than three markers). These color vectors can be defined using an optical density space (e.g., as an alternative to or supplement to using the RGB space). Each pixel in the input multiimage can then be mapped from the RGB space to a location in the optical density space where initial demixing can be performed.
[0069] In some cases, color vectors can be determined by inputting a solid-color image (e.g., depicting a slice or sample colored with a single marker) into a linear technique such as nonnegative matrix factorization (NMF). However, if used for demixing, color vectors derived from NMF can lead to errors. For example, background noise, faded tissue, or unclear morphology can result in initial color vectors that do not take into account the signals represented in one or more images captured in the real-world environment.
[0070] In some embodiments, fine-tuning of one or more color vectors can be performed using an interactive graphical user interface (GUI) and / or automation techniques. Such fine-tuning can be performed using images obtained in a specific environment (e.g., lighting), allowing the definition of one or more color vectors to account for real-world, environment-specific imaging effects. For example, one or more color vectors can be defined and / or adjusted to account for any effects that the imaging system and / or lighting environment may have on the signal depicted on or in a given marker in a digital pathology image.
[0071] The GUI can present realistic multiplexed images depicting slices stained with multiple dyes. The GUI may include one or more input components configured to adjust (i.e., fine-tune) the definitions of one or more color vectors. For example, one or more input components may be configured to move or adjust the representation of the color vectors in a light density space or RGB space. As another example, one or more input components may be configured to adjust the representation of one or more channels in the color space (e.g., the contribution of one or more of the red, blue, or green channels).
[0072] The GUI may also include one or more synthesized single images and / or synthesized multiple images, wherein each synthesized image is generated (e.g., dynamically generated) based on a color vector defined in the interface. The GUI may include one or more input components configured to receive inputs that adjust the contributions of one or more channels corresponding to a given signal.
[0073] For example, the GUI can be configured to receive definitions or adjustments for one or more color or frequency band channels relative to a given marker. As another example, the GUI can be configured to receive definitions or adjustments for hue angles and / or light density (representing intensity) in a light density space relative to a given marker. The light density space can be configured as a two-dimensional space (e.g., a chromaticity cx-cy plane) where each location is a non-fuzzy identifier of an RGB vector (e.g., such that locations in the light density can be deconvolved to identify locations within the light density space). Within this space, arbitrary scaling factors corresponding to angles can be defined, such that the color space spans a predefined space. Within the light density space, saturation can be represented by distance from the center, and / or hue can be captured by angles in polar coordinates.
[0074] The GUI can be configured to dynamically adjust (e.g., in real-time) one or more displayed single images and / or composite multiimages based on a set of color vectors defined for the underlying channels (via the interface). As an example, if the color vectors are set to be the same when the underlying images are tinted with different markers, the GUI can show that all composite single images will be identical and the composite multiimages will lack signals from the corresponding real multiimages. Therefore, the user can use this information to fine-tune the color vectors.
[0075] Once fine-tuning is complete, the color vectors can be used to generate one or more synthetic single images based on the input multiple images. The input multiple images to be demixed can be different from or the same as the multiple images used to determine the color vectors. Utilizing similar multiple images can reduce the impact of color variations across imaging instances (e.g., due to differences in tissue type, illumination, staining scheme, imaging system, etc.) on the degree to which markers can be accurately detected in a given instance.
[0076] In some cases, unmixing can be performed linearly using NMF techniques that leverage finely tuned color vectors and coefficient matrices from multiple input images (same or different), thereby generating a synthetic single image. As another example, color unmixing can be performed non-linearly by leveraging, for example, machine learning models such as autoencoders or generative adversarial networks (GANs).
[0077] In one aspect of this disclosure, the GUI can be configured to generate color vectors for synthetic dyes by synthetically and interactively blending two or more dye colors at different ratios. The colors of the synthetic dyes can be displayed in a chromaticity plane cx-cy via the GUI. The synthetic dyes can be generated by selecting multiple pre-identified chromatids (or fluorophores) for blending via user interaction. User input can then identify the relative contribution for each of the selected chromatids or fluorophores. The dye colors can be blended by generating a weighted average of the corresponding color vectors of the selected dyes in an OD space, where the weights are defined based on the relative contributions. The weighted average in the OD space can then be converted back to the RGB space (e.g., for display).
[0078] In some examples, synthetic pixels or associated adjusted color vectors obtained using the techniques disclosed above can be used to generate synthetic single and / or synthetic multiple images. For example, a machine learning model can be trained to convert an input counterstain image (such as hematoxylin) or an input multiple image into a synthetic image based on a given adjusted color vector. The synthetic image can be used to verify how well a defined (e.g., user-input based) color vector provides a basis for accurate unmixing and / or accurate mixing. Compared to performing actual staining experiments in the laboratory, the architecture for generating synthetic images can also help create additional training data for machine learning models in a faster and more cost-effective manner. This architecture can also enable pathologists to control different staining conditions, intensities, and appropriate combinations of biomarkers (e.g., for synthetic multiple images). These synthetic multiple images can be customized for specific needs and applications.
[0079] In some embodiments, a technique may be provided to determine adjustments to an initial color vector based on a given real digital pathology image. The real image may use one or more (but not all) stains associated with the initial color vector (where unused stains are referred to as “excluded stains” in the ongoing discussion). One or more generative models (e.g., including one or more autoencoders (AEs), one or more image-to-image translation networks, one or more generative adversarial networks (GANs), etc.) may be used to generate one or more synthetic single-color images using corresponding one or more color vectors (e.g., where at least one of the one or more color vectors is defined based on user input received via the interface described herein). Given the known presence of one or more excluded stains, the target output corresponding to these staining channels will lack any signal. Therefore, if the generated synthetic single-color image corresponding to the excluded stains includes a signal (e.g., a signal that is subjectively or objectively above a threshold), it can be inferred that the one or more color vectors used to generate the synthetic single-color image are suboptimal.
[0080] In some cases, a synthetic single-weighted image can be made available to a user device that receives input from it to define one or more color vectors (e.g., for display). This availability can be provided in real-time or near real-time as the user adjusts one or more color vectors. In some cases, a metric (e.g., cumulative absolute intensity, variation across intensity, maximum intensity) can be computed and used to automatically adjust one or more color vectors (e.g., using a loss function that employs this metric and is associated with one or more machine learning models to generate a synthetic single-weighted image). For example, when it is predicted (or known) that there is no biomarker corresponding to a given staining agent (or color vector) in a real multiplexed image, it can be expected that the mean, median, mode, variance, standard deviation, and / or range may be relatively low (or zero) when using accurate color vectors compared to when using less accurate color vectors. Once a quantitative metric for the synthetic single-weighted output based on the degree of presence of the staining agent in the real multiplexed image is determined, spatial traversal techniques can be used to discover color vector adjustments associated with excluded staining agents. Spatial traversal techniques can systematically explore the space of possible adjustments to color vectors representing excluded staining agents. Examples of such techniques may include, but are not limited to, gradient descent, Monte Carlo methods, genetic algorithms, or other probabilistic optimization techniques that iteratively adjust color vectors to optimize certain criteria, such as minimizing the computed metric. This adjustment is repeated iteratively until convergence or a stopping criterion is met. The goal is to find the optimal color vector that minimizes the metric, producing a synthetic single-weighted image that accurately represents excluded biomarkers.
[0081] Once the color vector associated with the excluded staining agent is adjusted by minimizing the metric, a new multiple image can be received. The multiple image can be colored with staining agents associated with the initial color vector, including one or more excluded staining agents whose color adjustments are calculated based on spatial traversal techniques and metric. By utilizing the aforementioned coloring unmixing process, one or more new synthetic single images associated with the excluded staining agents can be generated.
[0082] In yet another example, the disclosed technique can also be used to identify recommended color vectors for staining agents that complement other staining agents depicted in a given multi-image. This multi-image can be stained with at least two staining agents. For example, a dual image stained with two specific staining agents and a counterstain (e.g., hematoxylin). The objective could be to identify potential additional staining agents that are effectively distinguishable among the existing staining agents. Therefore, if the unmixing result can accurately distinguish signals from different staining agents (e.g., signals from one or more existing staining agents and one or more potential additional staining agents), a high score can be assigned via an objective function.
[0083] The interface can be configured to receive user input that identifies the color vector of the additional staining agent and, when the additional staining agent is used in conjunction with one or more existing staining agents, present one or more predicted unmixed outputs (e.g., one or more synthetic single images). Alternatively or concurrently, the color vector of the additional staining agent can be initially selected automatically (e.g., using a predefined selection of the color vector, a default user selection of the color vector, or an initial result from linear or nonlinear processing). For example, the interface can be configured to receive user input that identifies a specific chromogen or fluorophore, and the color vector associated with that specific chromogen or fluorophore can be initially assigned to the additional staining agent.
[0084] Using one or more color vectors defined according to the techniques disclosed herein, a real multiple image can be transformed into one or more synthetic single images (e.g., using unmixing techniques disclosed herein, such as linear unmixing, nonlinear unmixing, or machine learning models). To characterize the quality of the one or more synthetic single images, one or more metrics can be computed. For example, in the case where the input image depicts a sample slice not stained with a given dye (e.g., but stained with one or more other dyes), the metric can quantify the extent to which the signal associated with the given dye is present in the synthetic single image. For example, the metric can be the mean, median, maximum, or range of intensity in the synthetic single image. In this scenario, an ideal synthetic single image would not include the signal (because it is known that the given dye is not present in the initial slice), so an ideal metric would be zero. The metrics and / or synthetic single images can be presented on an interface so that they can inform the user of fine-tuning of one or more color vectors.
[0085] In another scenario, a metric characterizing a synthetic singlet image can be calculated, which corresponds to the stain actually used to stain the corresponding multiple slices. In this scenario, the signal components in the synthetic singlet image will be expected, and therefore a metric not close to zero can be expected (if the slices are known to have biomarkers corresponding to the stain).
[0086] Performance prediction scores can be generated using one or more metrics, and possibly using one or more target metrics. For example, when it is known that the sample depicted in the corresponding multiplexes does indeed have a signal from the stain associated with the synthetic singlexes, the performance prediction score (or its contributing component) can be defined as positively correlated with a metric for the synthetic singlexes characterizing the presence of the signal (e.g., mean, median, mode, maximum) or signal complexity (e.g., variation or range). Conversely, when it is known that the sample depicted in the corresponding multiplexes does not have a signal from the stain associated with the synthetic singlexes (e.g., because the stain was not applied to the sample), the performance prediction score (or its contributing component) can be defined as negatively correlated with a metric for the synthetic singlexes characterizing the presence of the signal or signal complexity. Thus, performance prediction scores can be generated in such a way that the score represents the degree to which the stain can be accurately detected and / or distinguished in the multiplexes.
[0087] In some cases, performance prediction scores can be further or alternatively estimated by performing clustering analysis based on image features associated with multiple synthetic single images. For each synthetic single image, one or more features can be defined or learned to characterize, for example, the optical density values in the image, the RGB values in the image, etc. For example, features can include statistics (e.g., mean, median, range, maximum, variance, mode, etc.) across one or more axes in the optical density or RGB space. As another example, features can characterize the spatial contrast of intensity (e.g., where contrast is related to the amount and / or degree of intensity across neighboring or nearby pixels). Clustering techniques (e.g., k-means, hierarchical clustering, or density-based spatial clustering of application and noise (DBSCAN)) can be used to cluster the features. For example, k-means clustering can be used when the number of clusters is defined (e.g., equal to the number of dyes applied to the scene or the number of dyes plus one or more other categories, such as blank signal categories). This clustering algorithm partitions the feature space into clusters. Ideally, such clusters are well isolated from each other and compact, and the features of the images associated with each given type of dye can be clustered together. The performance prediction score (or its contributing components) can be based on the degree to which clusters are separated in the feature space, the degree to which synthetic single images corresponding to a given color vector / dye are clustered together, and / or the degree to which images assigned to a given cluster are close together in the feature space. (One or more) This degree can be quantified using, for example, contour scoring, the Davis-Boulding index, or distances (e.g., Euclidean distance, Mahalanobis distance, or Manhattan distance).
[0088] Performance prediction scores can be based, alternatively, on the estimated correlation between one or more synthetic single images and their corresponding multiple images. The correlation can be estimated in RGB space, optical density space, feature space, etc. This method can account for variations in staining schemes, image acquisition settings, and tissue characteristics, thus providing a consistent basis for comparison. The values are inherently in a non-negative to positive range relative to the optical density space, thus aligning well with the physical constraints of staining intensity.
[0089] Regarding unmixing, in one aspect of this disclosure, constraints can be introduced to simplify staining analysis, thereby reducing the complexity involved in staining unmixing. This can achieve higher accuracy, precision, and / or reliability in generating synthetic single images from given multiple images. Each pixel of the multiple images can be mapped to a location within a multidimensional color map. Pixels within specific portions of the color map (e.g., quadrants, portions defined by values greater than / less than y and greater than / less than x, wedge regions, etc.) can be assigned and characterized to depict signals corresponding to only a single specific staining agent. For example, in optical density space, a given angular range can be defined as associated with a specific staining agent. For each pixel associated with a location within the angular range, the expression of the given staining agent can be inferred. Additionally, the staining agent intensity can be estimated based (at least in part) on the distance of the pixel's represented location from the axis. For pixels outside the angular range, the unmixing technique can predict the expression levels of other biomarkers, maintaining predefined expression levels, such as "0" for the first biomarker or other predefined numbers.
[0090] To facilitate the extraction of specific portions from a color space, the GUI can provide a set of tools to interactively define portions of a multidimensional space (e.g., OD space, feature space, RGB space, etc.) to map to corresponding rules regarding the definition of signal components. These tools can be configured to define regions in the multidimensional space corresponding to, for example, wedge-shaped regions, facets, exteriors, cylindrical regions, curved regions, or elliptical regions. Alternatively or additionally, tools can be configured to receive freeform input identifying a portion or all of the boundaries of the region. As some examples, the wedge region tool can be configured to receive input identifying a center point and angle; the exterior tool can be configured to receive input selecting one or more points along the boundary of the region to be defined; the brush tool can be configured to receive input corresponding to directly “painting” to the chroma map to define one or more regions in the multidimensional space; and so on. Tools can also be provided to incorporate thresholding techniques, in which the user can specify thresholds for one or more axes (e.g., one or more polar axes or one or more color channel axes in OD space). Once a portion is defined or selected within the color space, specific processing can be applied to each pixel representation assigned to (or not assigned to) that portion. For example, when a pixel represents a portion of space, a specific algorithm can be used to transform the coordinates into a predicted intensity of a specific dye corresponding to that portion. As another example, when a pixel represents a portion of space, it can be inferred that the pixel does not contain signals from a specific dye associated with that portion (e.g., and demixing can be performed based on this inference).
[0091] Figure 1 illustrates a workflow 100 for acquiring and processing multiple images. The image generation system 105 can be configured to acquire images of one or more stained samples. The stained samples can be stained with, for example, one or more biomarker stains and / or one or more reference stains. The acquired images can include solid color images (where the sample is stained with only one stain), single images 108a to 108n (where the sample is stained with a single biomarker stain and a reference stain), or multiple images 110a to 110n (where the sample is stained with two or more biomarker stains and a reference stain). The acquired images can be transmitted to a computer system 115 via a communication network 120.
[0092] Computer system 115 can process images to generate one or more outputs 135a to 135p. In some cases, computer system 115 receives multiple images depicting a sample stained with multiple biomarker stains (two or more stains or three or more stains) and a reference stain, and computer system 115 generates an output predicting the signal from each of at least one of the stains. For example, if triple images are received, computer system 115 can generate output including one or more synthetic single images corresponding to the biomarker stains used to prepare sample slices for the images and / or the reference stains used to prepare sample slices for the images.
[0093] Outputs 135a to 135p can be generated, for example, using automated techniques and / or using input received via interface 112. For instance, interface 112 can be configured to dynamically display a synthetic single-weighted image and / or associated metrics generated based on current color vectors assigned to multiple stains represented in the input multi-image. Interface 112 can also be configured to receive input that directly or indirectly adjusts the color vectors for each of one or more stains among the multiple stains (e.g., thereby triggering automated updates to interface 112).
[0094] Images that can be processed by computing system 115 may include and / or may be transformed (e.g., via computing system 115) into image data, which may include data characterizing one or more intensities (e.g., where each intensity corresponds to a given color channel or a given frequency band) for each of one or more pixels. For example, biological samples (e.g., tissue sections) have been stained by applying a staining assay that includes one or more chromogenic stains (for bright-field imaging), fluorophores (for fluorescence imaging), quantum dots, or combinations thereof. In the analysis of biological samples (e.g., cancerous tissue), different stains are specified to identify one or more types of biomarkers, such as immune cells.
[0095] Communication network 120 may include the Internet, intranet, wired LAN (local area network), wireless LAN (WiLAN), WAN (wide area network), MAN (metropolitan area network), PSTN (public switched telephone network), and other types of communication networks. Communication network 120 may further include communication devices such as one or more gateways, routers, or bridges. By way of example only, communication network 120 may have one or more servers and one or more websites accessible to users for sending and receiving information usable by one or more computer systems 115. Communication network 120 may be any type of network familiar to those skilled in the art, supporting data communication using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Switching), AppleTalk®, etc.
[0096] The computer system 115 of the exemplary system 100 may include a processing system 125 having one or more high-speed central processing units (CPUs), processors, and one or more memories. The computer system 115 may also include memory for storing processing modules or logical instructions executed by the coupled one or more processors. The computer memory storing data may also be maintained on a computer-readable medium, including disks, optical disks, organic memory, and any other volatile (e.g., random access memory (RAM)) or non-volatile (e.g., read-only memory (ROM), flash memory, etc.) mass storage systems readable by the CPU. The computer-readable medium may include cooperative or interconnected computer-readable media that exist only on the processing system, or may be distributed among multiple interconnected processing systems, possibly local to or remote from the processing system.
[0097] One or more databases 130 may store images acquired by the image generation system 105 and / or one or more image processing results (e.g., synthesized single images and / or synthesized multiple images).
[0098] Computer system 115 may include a client terminal that communicates with one or more servers, personal digital / data assistants (PDAs), laptop computers, mobile computers, internet devices, one-way or two-way pagers, mobile phones, or other similar desktop, mobile, or handheld electronic devices. The client terminal may be configured to transmit and / or receive information to or from one or more client systems. For example, the client terminal may provide an interface through which input is received to partially or fully define one or more color vectors or other components of a demixing protocol. The interface may further or alternatively display a representation (e.g., in optical density space) of one or more received images and / or one or more composite images (e.g., generated using a color vector set that may have been generated at least partially using the input received via the interface).
[0099] Figure 2 illustrates an exemplary network 200 of the digital pathology image generation system 105 from Figure 1. The image generation system 105 may include a fixation / embedding system 205 that fixes and / or embeds tissue samples (e.g., liquid fixatives, such as formaldehyde solution) and / or embedding materials (e.g., histological waxes such as paraffin and / or one or more resins such as styrene or polyethylene). Each section can be fixed by exposing the section to a fixative for a predetermined period of time (e.g., at least 3 hours) and then dehydrating the section (e.g., via exposure to an ethanol solution and / or a clarifying intermediate). When the section is in a liquid state (e.g., when heated), the embedding material can impregnate the sample.
[0100] The image generation system 105 may further include a tissue slicer 210 that slices a fixed and / or embedded tissue sample (e.g., a tumor sample) into a series of sections, each section having a thickness of, for example, 4 to 5 micrometers. This slicing can be performed by first cooling the sample and then slicing it in a warm water bath. Tissue slices can be prepared using, for example, a vibratory microtome or a compression microtome.
[0101] Because tissue sections and the cells within them are essentially transparent, section preparation typically involves staining the tissue sections (e.g., automated staining) to make the relevant structures more visible. In some cases, staining is performed manually. In others, staining is performed semi-automatically or automatically using a staining system 215.
[0102] Staining can involve exposing individual sections of tissue to one or more different staining agents (e.g., sequentially or simultaneously) to express different characteristics of the tissue. For example, each section may be exposed to a predefined volume of staining agent for a predefined time period. Staining agents may include, for example, RNA probes, protein probes (e.g., nuclear-protein probes or cytoplasmic-protein probes), immunohistochemical staining agents, probes for secreted substances, etc. In some cases, the staining agent is a staining agent for KAPPA mRNA or LAMBDA mRNA.
[0103] An exemplary type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes) to stain tissue structures. Histochemical staining can be used to indicate general aspects of tissue morphology and / or cellular histology (e.g., to distinguish between cell nuclei and cytoplasm, to indicate lipid droplets, etc.). An example of a histochemical staining agent is hematoxylin and eosin (H&E). Other examples of histochemical staining agents include trichrome staining agents (e.g., Masson's trichrome stain), periodic acid-Scheffler (PAS), silver staining agents, and iron staining agents. The molecular weight of histochemical staining agents (e.g., dyes) is typically about 500 kilodaltons (kD) or less, although some histochemical staining agents (e.g., Alcian blue, phosphomolybdic acid (PMA)) may have molecular weights as high as two or three thousand kD. An example of a high molecular weight histochemical staining agent is α-amylase (about 55 kD), which can be used to indicate glycogen.
[0104] Another type of tissue staining is immunohistochemistry (IHC, also known as "immunohistochemistry"), which uses a primary antibody that specifically binds to a target antigen (biomarker). IHC can be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a tag (such as a chromophore or fluorophore). In indirect IHC, the primary antibody first binds to the target antigen, and then a secondary antibody conjugated to the tag (such as a chromophore or fluorophore) binds to the primary antibody. The molecular weight of IHC reagents is much higher than that of histochemical staining reagents because antibodies have a molecular weight of approximately 150 kDa or higher.
[0105] The slides can then be individually placed on corresponding slides, and the imaging system 225 can scan those slides to generate raw multiplex and / or singlex digital pathology images (e.g., 110a to 110n, 108a to 108m). Each slide can be placed on a slide, and the slide can be scanned to create a digital image, which can then be evaluated using automated digital pathology image analysis and / or using input from a human pathologist (e.g., using image viewer software). Inputs and / or results from automated analysis can identify, for example, annotations identifying one or more segments corresponding to physiological categories (e.g., tumor regions, necrosis, etc.). Alternatively or additionally, inputs and / or results can identify some or all of color vectors or related variables to facilitate unmixing of the same or different slides.
[0106] Digital histopathological images (e.g., 110 or 108) typically comprise an array of pixels, usually a rectangular matrix. Each “pixel” is an image element and a digital quantity representing a property of the image at a position in the array corresponding to a specific location in the image. If the digital histopathological image is a grayscale image, the pixel values of the digital image typically conform to a specified range. For example, each array element can be one byte (e.g., eight bits) representing a pixel value in the range of 0 to 255. In a grayscale image, “255” can represent absolute white, and zero (“0”) can represent absolute black (and vice versa). Color images can include multiple (e.g., three) color channels, such as red, green, and blue (RGB) channels. For a given pixel, there is typically a value for each of these color channels (e.g., a value representing the red component, a value representing the green component, and a value representing the blue component). By varying the intensity of these three components, all colors in the color spectrum are typically formed. It will be appreciated that in some cases, digital histopathological images include signals corresponding to one or more wavelengths outside the visible spectrum (e.g., in the ultraviolet or infrared spectrum).
[0107] Figure 3A illustrates an exemplary workflow 300-A for defining color vectors associated with stains from a digital pathology image and using these color vectors for stain unmixing. At box 302, one or more color vectors associated with one or more corresponding stains are defined and / or adjusted. One or more single-stain (or solid-color) slides (e.g., slides 305a, 305b, and 305c) are accessed.
[0108] The solid-color stained slide 305 can be an IHC image depicting a slide stained with a single staining agent without counterstaining, the slide being stained by replacing the buffer solution in other major biomarkers in a multiplex IHC staining protocol. The user interface can display one or more slides 305a to 305c and can receive user input identifying one or more target regions in a given depiction of at least a given portion of the slide.
[0109] As an illustrative example, Figure 3A depicts three solid-color stained slides (single yellow stain (dansyl sulfonyl) 305c, single purple (TAMRA) 305b, and single blue (hematoxylin) 305a). These solid-color images 305 may also include one or more markers corresponding to specific locations within the image. This overlay can provide reference points by strategically placing the markers at locations that may provide the solid stain or a representative target area within the image. These locations can be selected based on prior knowledge of the staining process, tissue characteristics, user input, or through empirical observation of image features. Color vectors can be defined based on these specific locations.
[0110] The solid-color stained slide 305 can be processed using linear techniques such as nonnegative matrix factorization (NMF) 310, which can produce two nonnegative matrices. In this technique, the solid-color stained RGB image (e.g., 305a to 305c) can be transformed into the optical density (OD) domain based on the Lambert-Beer law. According to this law, optical density is linearly related to the staining concentration. Following this law, a two-dimensional (2D) matrix (D) can be generated, which can be further decomposed using the NMF 310 technique to find the initial color vector 315 associated with the location of the marker. Mathematically, this can be represented in matrix form as D = WH. For the staining application, D is the optical density matrix, W is the nonnegative basis matrix (also called the "color vector" matrix), and H is the nonnegative coefficient matrix (also called the staining intensity matrix). For an RGB stained image, the columns of the W matrix can correspond to the initial color vector 315 based on a specific location for each staining element.
[0111] A color vector derived from a solid-color stained slide (e.g., having a size of (1×3)) can correspond to a representation of a single color in a three-dimensional color space (such as RGB). For example, for dansyl staining, the extracted color vector could have an RGB composition [0.248, 0.374, 0.894]. A matrix W derived from multiple images (such as dual) can have a size of (3×3) with two stainings and one counterstain, or for triple, W can have a size of (3×4). The initial color reference matrix W obtained from NMF310 may ultimately fail to perform well in unmixing the stainings of a given multiple image 330. After unmixing the multiple image 330 from the initial color vector 315, it may lead to errors (e.g., blank or faded counterstain hematoxylin) or the presence of background noise. It is understood that the initial color vector 315 arranged in columns constitutes the initial reference matrix W. To mitigate the errors of the initial color vector 315, the initial color vector can be calibrated. Calibration can be performed to identify adjusted color vectors that produce high-quality synthetic single and / or synthetic multiple images. For this purpose, an interactive graphical user interface (GUI) 112 and / or automation techniques can be provided to facilitate fine-tuning of one or more initially defined color vectors 315, as shown in Figure 3A. For reference, a real multiple image 330 can also be provided in interface 112 for fine-tuning the initial color vectors. Interface 112 can receive user input defining or adjusting one or more color vectors, enabling dynamic definition or adjustment of the color matrix W. As shown in Figure 3A, while interface 112 can be configured to allow the user to fine-tune the color vectors by adjusting the contribution of a given color channel (e.g., red, green, or blue channel), interface 112 can also display the representation of each of the one or more color vectors in the optical density space.
[0112] Using a color matrix W, one or more synthetic single images 340 are dynamically generated from real multiple images 330 or from different multiple images. The synthetic single image 340 can be displayed on interface 112 and dynamically updated as the color vector 325 is adjusted. Once fine-tuning is complete, the color vector 325 can be locked and used to perform color unmixing 335 on the same or different multiple images.
[0113] In some cases, color unmixing 335 can be performed linearly, such as by using NMF techniques that utilize the updated color matrix W 325 and staining coefficient matrix H of the input (same or different) multiple images 330 to generate a synthetic single image 340. Alternatively, color unmixing can be performed non-linearly (e.g., by utilizing a machine learning model explained below with reference to Figures 3D and 3E).
[0114] Figure 3B illustrates exemplary linear unmixing techniques according to some embodiments of the present disclosure. This technique can be used to extract an initial color vector 315 from a given pure glass slide and / or to perform staining unmixing 335. The specific linear unmixing technique shown in Figure 3B is nonnegative matrix factorization (NMF) 310. NMF 310 can be performed in the optical density (OD) domain, where colors / stains are represented as absorbance values rather than the original RGB pixel values. Therefore, preprocessing can be performed to convert RGB intensities to OD values for each pixel in the image. The OD values are then processed based on the physical properties of light absorption by different stains or fluorophores present in the sample, producing more accurate and meaningful results.
[0115] To transform a real / synthetic RGB image (e.g., an IHC pure dye image, such as 305a to 305c) into the OD domain, it can be assumed that the stained image is light-absorbing and satisfies Lambert-Beer's Law. Lambert-Beer's Law states that the intensity of light absorbed or transmitted through a medium is proportional to the thickness of the medium and the concentration of the transmitting material. Mathematically, Lambert's Law can be formulated for the intensity (I) of light after passing through the medium as: I = I0.e -αcd Where I0 is the initial intensity of the light before it enters the medium, and α Let be the absorption coefficient of the medium, c be the concentration of the absorbing material or the staining dose per unit area, and d be the thickness of the medium. Against the background of a digital IHC image (e.g., 305), each color channel (e.g., red, green, and blue) will have a light intensity (IL). R , I G , I B This includes the sample's absorption coefficient, the concentration of staining in the sample, and the sample's thickness. Therefore, Lambert's law can be applied individually to each color channel, describing how each color component is attenuated differently as light passes through the medium to produce the final color appearance of a multi-image.
[0116] Lambert's law describes an exponential (or nonlinear) relationship between the intensity (I) of light passing through a medium and the products of c, α, and d. Due to this nonlinear relationship, intensity values from RGB (digital) images cannot be directly used to unmix each dye. To simplify data analysis and interpretation, calculations can be performed in the optical density domain, which allows for linear relationships and dynamic range compression over a wide intensity range. Optical density (OD), usually denoted as D, is a measure of the degree to which a material attenuates light. Optical density can be formulated as: The diagram illustrates the direct relationship between optical density and the variables α, c, and d. Higher OD values indicate a greater staining concentration in the sample. For each color channel, an OD vector can be formed such that D = {D...} R D G D B}
[0117] NMF 310 operates under the assumption that the observed colors / stains in an image are linear combinations of the individual components of the colors / stains. This assumption allows for the use of linear transformations to separate mixed stains. NMF utilizes an iterative optimization algorithm to decompose the observed data matrix into nonnegative matrices representing the spectral features (basis matrix) and abundance maps (coefficient matrix) of the components.
[0118] Using NMF can be advantageous (e.g., superior to other linear techniques) because, for example, NMF uses intuitive nonnegativity constraints that align well with the physical constraints of staining intensity in pathological images. Furthermore, given that NMF uses basis vectors representing pure staining agents, the results are interpretable. NMF is also configured to be flexible in accepting constraints or prior knowledge and robust to noise and staining variations.
[0119] In NMF 310, the obtained data matrix is in the optical density domain (e.g., In 311), d represents the dimension of each data point (e.g., 3 for an RGB image), and m is the number of data points. It is assumed that this data matrix is non-negative. In other words, for each pixel, there exists an RGB combination in the OD space. This matrix can be decomposed into a color vector matrix 312 ( ) and staining intensity matrix 313, of which Let D(311) be the expected rank of the matrix representing the number of dyes. Both matrices (i.e., W(312) and H(313)) are similarly subject to nonnegativity constraints. Mathematically, D ≈ W × H, which can be solved by the following optimization problem:
[0120]
[0121] in This represents the Frobenius norm.
[0122] The color vector matrix 312 and coefficient matrix 313 can be initialized to achieve convergence to the optimal solution. Initialization can be performed using various techniques, such as random initialization, singular value decomposition (SVD), sparse initialization, k-means, or guided initialization. These techniques can be used individually or in combination, and the choice of initialization technique may depend on the specific characteristics of the data and the desired decomposition properties. In NMF 310, a multiplicative update rule is used to iteratively optimize the objective function. The updates for the basis matrix and coefficient matrix can be formulated as follows: and To avoid scaling issues and non-unique solutions, NMF 310 can be extended to sparse NMF by adding regularization and sparsity terms.
[0123] For staining unmixing 335, the synthetic single OD image can be reconstructed from the color vector matrix W 312 and the staining intensity matrix H 313. For the i-th... Reconstruction of the staining agent, the i-th of W Column (i.e., W) i ) can be related to the j-th line (i.e., H) j Multiplying these together generates a synthetic single-OD image (e.g., 314a). A color vector matrix W312 can be used (e.g., defined based on user input). If needed, these single-OD images (e.g., 314a and 314b) can be converted to the RGB domain. To transform an OD image to the RGB domain, the synthetic / real OD image (e.g., 314a or 314b) is associated with a single or multiple dyes for the single-OD image, and converted to the corresponding synthetic / real single-OD RGB image by applying Lambert's law, which exponentially scales the OD values. The mathematical formula for the conversion can be written as: I = I0 e -D The transformation can be applied to each pixel to obtain the corresponding intensity value for a synthetic / real RGB single or multiple image.
[0124] Figure 3C illustrates an interface component that facilitates fine-tuning of color vectors. The RGB model may be less useful during fine-tuning because the target information (e.g., the color of a dye (determined by absorption properties)) is mixed with variations in dye dosage. One technique that can be used to extract chromaticity (color) information from RGB data uses the Hue Saturation Intensity (HSI) model. The RGB to HSI transformation decouples intensity information from color information. In the HSI model, the hue of a color is its angle measured on the color wheel within the range of 0 to 360 degrees. For example, pure red is 0°, pure green is 120°, and pure blue is 240°. For convenience, neutral colors (such as white, gray, and black) are set to 0°. The HSI of saturation is defined as a measure of the purity / grayscale of a color, which can be estimated by the ratio of the difference between the maximum and minimum RGB values to the maximum RGB value. In the HSI color model, saturation can be considered as the distance from the center of the color wheel. Purer colors have higher saturation values that are further away from the center, while grayer colors have saturation values that are closer to the center.
[0125] Intensity is the overall luminance or lightness of a color, numerically defined as the average of equivalent RGB values, i.e., I = (R+G+B) / 3. However, most of the intensity variations perceived in transmitted light microscopy are likely due to changes in staining density. Therefore, the Hue Saturation Density (HSD) transformation is defined as the RGB-to-HSI transformation applied to the light density values of individual RGB channels, rather than intensity. For a single pixel, the measure of OD can be defined as...
[0126] The RGB to HSD transformation can be defined as: , It is understandable that, since OD is decoupled, the chromaticity coordinates of the HSD model are not equal to those of the HSI model. For the HSD model, the resulting cx-cy plane has the following properties: a single point and α R α G and α B RGB points with the same ratio correspond to each other. Therefore, all information about the absorption curve is represented in a single plane. Similar to the HSI model, hue and saturation values can be calculated from the chromaticity triangle because the dye mixture exhibits a linear pattern in the cx-cy plane of the HSD model.
[0127] In the chromaticity plane (cx-cy), the RGB cube can be represented by an equilateral triangle 321d, which limits the range of the cx-cy coordinates. The cx-cy plane is a 2D coordinate system represented by the equilateral triangle 321d, where the center of each side represents red 321a, green 321b, and blue 321c. In this plane, each color vector can be represented as a point in the cx-cy plane, where the position of the point corresponds to the relative proportions of the primary colors (i.e., red, green, and blue) in the color vector. For example, if a color vector has a high-intensity green channel, the corresponding point will be closer to the green center point 321b. The staining properties of the dye can be modified by adjusting the position of the color vector in the chromaticity plane 321 via GUI 112. The staining properties can also be modified by adjusting the proportions of R, G, and B from slider 322.
[0128] In one example, GUI 323 can be configured to synthetically and interactively blend two or more dye colors at different ratios to obtain a target dye. The resulting composite pixels can be displayed in the chromaticity plane cx-cy 321 via GUI 323. Such composite color pixels can be generated by user interaction by selecting the dyeing dyes from a given list of color sources. The dye dosage for each color source can then be set (e.g., relative to (one or more) other color sources). The dye colors can be blended by adding the amount of dye / color source to the result of multiplying the corresponding color vector in the OD space, which can then be converted back to the RGB space for display purposes. Since the chromaticity planes cx and cy only represent hue and saturation, a cz value may be needed to determine the transformation back to RGB. This can be done by first finding cz according to cz = 1 – cx – cy and then calculating R = cx.cz / cy, G = cz, B = (1 –cx – cy).cz / cy.
[0129] As an illustrative example, a user can generate a composite pixel 323b in the cx-cy plane by first selecting a set of color sources (e.g., 323c) and then manipulating slider 324 to set the relative staining dose for each color source. In Figure 3C, multiple sliders 324 are shown, where setting "blue-green" to "0.8" and "Tamra" to "0.4," and setting the remaining color sources (e.g., dansyl and hematoxylin (HTX)) to "0" produces a blue composite pixel 1 323b (marked with an asterisk "*" in GUI 323). Similarly, another pixel 2 323d for a different set of color sources 323f can be generated by setting "green" to "0.8" and "Tamra" to "0.4" from slider 324. Composite pixel 2 is marked with an "X" in GUI 323. Pure hematoxylin stain 323a can also be provided as a reference for fine-tuning the synthesized pixels in Figure 3C. It can be observed that synthesized pixel 2 is visually closer to pure hematoxylin stain 323a than synthesized pixel 1. The positions of the mixed stain colors in the cx-cy diagram indicate the similarity in hue and saturation between the two synthesized pixels. Increasing the stain dosage while maintaining the relative ratio of the color sources may not change the pixel positions in the cx-cy diagram, but it may change the appearance of the synthesized pixels as shown in the interface. This is consistent with the design of the cx-cy space, which counts only hue and saturation while keeping density constant.
[0130] Figure 3D illustrates an exemplary architecture 300-D for generating one or more synthetic single-color images and / or synthetic multiple-color images by leveraging multiple machine learning models. To generate such synthetic images, architecture 300-D includes a staining unmixing module 335, a color vector 325, and a remixing module 345. To generate synthetic images, the synthetic chromogen can be controlled by adjusting the color vector, and a trained machine learning model can be used to simulate real images to generate cell / tissue-level biomarker staining patterns. The machine learning model can include a generative model (such as a generative adversarial network (GAN), a diffusion model, or an autoencoder) trained to generate single-color images conditioned on an input color vector. As an illustrative example, the staining unmixing module 335 can include a conditional GAN (cGAN) to generate synthetic single-color images conditioned on the input color vector. For example, individual cGANs (e.g., 338a, 338b, and 338c) can be trained to generate individual single-weighted images (e.g., 340a, 340b, and 340c) corresponding to a specific target staining agent (or synthetic pixel), as shown in Figure 3D. The number of models can depend on the number of constituent staining agents in the multi-weighted images.
[0131] In one aspect, this architecture 300-D can be used to generate synthetic images from synthetic pixels or associated adjusted color vectors obtained using the techniques disclosed above. For example, to generate a synthetic single-color image, the cGAN model (e.g., 338a) can use a complex dye image (such as hematoxylin) as input image 332, conditioned on the color vector (e.g., 325a) of the target synthetic pixel (dye). The generated single-color image 340 can be used to verify the correctness of the synthetic pixels for the target dye. In another example, the synthetic single-color image corresponding to the target synthetic pixel can be used to generate a synthetic multiple image. The generated synthetic multiple image 350 can be displayed simultaneously with the real input multiple images used to generate the synthetic single-color images 340a to 340c. The user can then evaluate how similar the real multiple image and the synthetic multiple image look (e.g., compared to instances where some or all signals from the real multiple image are absent in the synthetic multiple image). This can facilitate quality control and / or additional fine-tuning of one or more color vectors. Additionally, when cGAN models 338a to 338c are approved, they can be used to generate multiplex images, thus producing additional training data for machine learning models in a faster and more cost-effective manner compared to performing actual staining experiments in the laboratory. The models also allow pathologists to control the different staining conditions, intensities, and appropriate combinations of biomarkers in the synthetic multiplex images. These synthetic multiplex images can be customized for specific needs and applications.
[0132] In another scenario, each cGAN may receive a real multiple image (e.g., 330) as input image 332 and color vectors (e.g., 325a, 325b, or 325c) obtained from module 302. In this setting, the real multiple images can be filtered using architecture 300-D based on adjustments to the color vectors. This generative model can be trained to filter a given multiple image to generate an output that includes a predicted signal for a specific staining agent associated with the model (this condition is based on the color vector of the specific staining agent defined for the cGAN). These synthetic single images can be further combined by the staining remixing module 345 to generate a synthetic multiple image 350. The synthetic multiple image 350 can be compared (e.g., computationally, automatically, and / or via user review) with the real multiple images provided as input 332. This comparison can be used during training and / or as an indication of confidence in the quality of the generated synthetic single images 338a through 338c. Image quality indicators can be incorporated as feedback into GUI 112, informing the user of decisions regarding whether to further adjust one or more color vectors. It is understood that the number of generative models shown in Figure 3D is for illustrative purposes only. Depending on the nature of the multiple images, aspects of this disclosure are intended to include or otherwise cover any number of generative models.
[0133] Figure 3E illustrates an example of another method for generating one or more synthetic single images 340 using a single model (e.g., a single cGAN). The single model 339 can be configured to receive a real multiplex / dichromatogram image 332 as input, along with an identifier for a specific biomarker (e.g., by receiving a corresponding color vector, e.g., 325a), indicating the characteristics of the synthetic single image being requested. Additionally, as noted above, the synthetic single images 340 can be combined to produce a synthetic multiple image 350. Comparison of the synthetic multiple image 350 with the real multiple image can be used during training and / or as an indication of confidence in the quality of the synthetic single image.
[0134] The staining and remixing module 345 can combine synthetic single images (e.g., 340) to generate a synthetic multiple image 350. The synthetic multiple image 350 can be generated linearly in the optical density (OD) domain, which involves merging the intensity values of each pixel from the individual single images 340a to 340c to create the synthetic multiple image. Depending on the desired result, this process can be implemented using various mathematical operations such as addition, subtraction, multiplication, or weighted averaging. Mathematically, the process can be formulated as D 多重 = w1 D1+…+w c Dc Where D1,…, D c Denotes the single-dimensional OD matrix (e.g., 314a, 314b), w1,…, w c These are weighting factors assigned to each single-image. These weights control the contribution of each staining agent to the final multi-image 350.
[0135] Alternatively, for staining remixing 345, a generative model (such as a GAN or autoencoder) can be trained to learn the complex mapping between synthetic single images and their corresponding multiple counterparts. By training the generative model on a dataset that includes input-output pairs (e.g., single and multiple images), the model can capture the intricate relationships between staining agents and cellular structures. This process can involve learning to fuse features extracted from individual single images 340 to create consistent and visually realistic multiple images. The adversarial loss used to train the generator G and discriminator D to transform synthetic single images 340a to 340c into synthetic multiple images 350 can be formulated as follows:
[0136]
[0137] Where x is the synthesized single image (x1,…, x…). c The set from 340a to 340c, where n is the number of samples in the training data.
[0138] Figure 4 illustrates a flowchart of an exemplary process 400 according to some embodiments of the present disclosure, which determines one or more color vectors for staining unmixing. Process 400 involves staining unmixing a digital pathology image by finding an initial color vector 315 and adjusting the color vector using a graphical user interface (GUI) 112. Process 400 begins at box 405, where a color vector is determined to be associated with a digital pathology stain or chromogen (e.g., at least three). The color vector can be a default color vector associated with the dye (which can be identified using, for example, lookup tables and / or predefined variables). For example, the color vector for a “green” dye can be defined as [0,1,0] in the RGB space.
[0139] Alternatively, the initial color vector determined at box 405 can be determined using initial processing of one or more images received from the user device (or other devices associated with the user device). For example, nonnegative matrix factorization (NMF) can be performed to transform a given OD matrix into two nonnegative matrices, such as a W color vector matrix and an H abundance or coefficient matrix. The determined color vector 315 in W can accurately represent the true spectral characteristics of the chromatic components, although such characteristics may alternatively be missed due to noise, artifacts, or limitations of the imaging system. Thus, the interface can provide dynamic data that facilitates fine-tuning of one or more color vectors.
[0140] At box 410, the interface is made available to the user device. For example, communication can be transferred from a server (e.g., a web server) to the user device, where the communication includes code with instructions for generating and displaying the interface on the user device. As another example, native code can be executed to generate and display the interface.
[0141] The interface may include: a representation of each of the determined color vectors; a real multiple digital pathology image; at least one synthetic single image; and one or more color-vector adjustment tools. Each of the at least one synthetic single image may be generated using the real multiple digital pathology image and the color vectors determined at box 405. Each of the at least one synthetic single image may be generated by processing the real multiple image using techniques described herein, such as linear unmixing (NMF) or nonlinear unmixing (e.g., machine learning models). One or more color-vector adjustment tools may be configured such that, for a given color vector, they can receive input that adjusts the contribution or weight associated with each of one or more contribution axes. For example, a color-vector adjustment tool may be configured to include a slider or numerical input that defines the weights to be assigned to a given color channel (e.g., red, green, or blue channel), a polar coordinate channel (e.g., in optical density space), or a channel in another space.
[0142] At box 415, input corresponding to a specific adjustment to the color vector represented in the interface is detected. Input may include interaction with at least one of one or more color-vector adjustment tools. Input may include, for example, positioning a slider and / or input indicating the absolute or relative contribution of a channel (e.g., a color channel) represented for a given dye. For example, input may include a number or slider position indicating that a given dye will include 5% of the red channel, instead of 0% of the red channel for a “green” dye (where the percentage is absolute or relative to a cumulative percentage across channels). As another example, input may identify a location within the optical density space that will be used as the definition of the color vector for a given dye.
[0143] At box 420, in response to detected input, interface 112 can be automatically updated. Automatic updating can update the displayed representation of the color vector representing a specific staining agent. Alternatively, the update can use adjusted color vectors to update one or more synthetic images (e.g., one or more synthetic single images and / or synthetic multiple images). One or more metrics (e.g., those characterizing absolute or relative statistics associated with single or multiple images) can also be updated.
[0144] Boxes 415 and 420 can be repeated multiple times (e.g., until no input is received within a threshold time period, the session ends, the user indicates that the color vector has been finalized / defined, the automated quality control conditions are met, etc.).
[0145] Figure 5 illustrates an exemplary architecture of a system 500 for determining adjustments to an initial color vector based on a given multiple digital pathology image, according to some embodiments of the present disclosure. The initial color vector 315 can be determined using initialization techniques (e.g., by utilizing NMF technique 310) and adjustments to a particular initial color vector can be discovered based on the real multiple images. In this setting, the real multiple image 502 can be stained using one or more stains from the initial color vector 315, but at least one of the initial reference stains may be absent (because the corresponding slide was not stained with at least one of the initial reference stains).
[0146] In the ongoing discussion, at least one colorant not represented in the real multiple image 502 is referred to as an "excluded colorant". An initial color vector 315 can be fed into a filter 504, which is configured to generate a filtered output from the real multiple image 502 based on the excluded colorants. Similar to the process in Figure 3D, filtering can be facilitated by employing one or more generative models (e.g., autoencoders (AEs), image-to-image translation networks, generative adversarial networks (GANs)), which can be trained to learn a mapping from a given multiple image to its constituent single images conditioned on the constituent color vectors. This approach is driven by the potential of a well-trained machine learning model to generate empty (zero-value) images when conditioned on color vectors that are not present in the input multiple image 502. This behavior is expected because the model has been trained to understand the relationship between color vectors and corresponding colorants present in the input images. If the model encounters a color vector that is not present in the input multiple image, it may fail to generate any meaningful output associated with this colorant, producing a zero-value or empty image for the excluded colorant associated with this color vector.
[0147] To assess the quality and characteristics of the filtered output generated by the machine learning model (filter 504), metrics can be computed. For example, when it is predicted (or known) that there is no biomarker corresponding to a given staining agent (or color vector) in the real multiplex image 502, the mean, median, mode, variance, standard deviation, and / or range in the synthetic singlex image corresponding to the given staining agent can be expected to be ideally very low (or zero). Therefore, metrics can be generated in such a way that the score is negatively dependent on the statistics (e.g., mean, median, mode, variance, standard deviation, and / or range) in the synthetic singlex image that characterizes the presence of the corresponding staining agent signal in the real multiplex image.
[0148] For this metric, the pixel cumulative statistic (e.g., mean or average) can be calculated by taking the pixel intensity values of the synthesized single image in the OD space (e.g., matrices 314a and 314b in Figure 3B) and then dividing by the total number of pixels. Referring to Figure 3B, if, for example, stain 1 is not present in matrix 314a, the corresponding row in the H matrix will be approximately zero, resulting in an empty image in the OD space. Subsequently, the average value for this synthesized OD space matrix will be low. Alternatively, the median can be selected by sorting all pixel intensities of the given OD matrix (314a in this example) associated with the excluded stain in ascending / descending order and identifying the median value.
[0149] Once the metric is determined to quantify the filtered output based on the degree to which the staining agent is present in the real multi-image 502, a spatial traversal technique 508 can be used to discover color vector adjustments 510 associated with the excluded staining agent. The spatial traversal technique 508 can systematically explore the space of possible adjustments to the color vector representing the excluded staining agent. Examples of such techniques may include, but are not limited to, gradient descent, Monte Carlo methods, genetic algorithms, or other probabilistic optimization techniques (such as simulated annealing) that iteratively adjust the color vector to optimize certain criteria, such as minimizing the computed metric. Gradient descent is an optimization algorithm that is typically used to minimize a function by iteratively moving in the direction of its steepest descent. In this example, the objective to be minimized may be a metric computed based on a synthetic single-image. This algorithm can start with an initial color vector representing the excluded staining agent and compute the gradient of the metric with respect to the color vector. This gradient represents the direction of the steepest ascent of the metric. The color vector can be adjusted (scaled with small steps (learning rate)) in the opposite direction of the gradient to minimize the metric. This adjustment is repeated iteratively until convergence or a stopping criterion is met. The objective can be defined as finding a color vector that minimizes the metric, producing a synthetic single-weighted image that accurately represents excluded biomarkers.
[0150] The Monte Carlo method is a stochastic simulation technique that uses random sampling to estimate numerical results. In this case, Monte Carlo simulation can be used to explore the space of possible adjustments to color vectors representing excluded dyes. According to this technique, color vectors representing excluded dyes within a specified range or distribution are randomly adjusted. A metric is calculated for each randomly adjusted color vector. Based on the metric value and the optimization objective (minimize or maximize), adjustments can be probabilistically accepted or rejected, guiding the search for a better solution. This process can be repeated up to a certain number of iterations to achieve a comprehensive exploration of the adjustment space. By iteratively sampling and evaluating adjustments, the Monte Carlo method can efficiently explore the adjustment space and identify promising regions or solutions that can be incorporated via interface 112.
[0151] Once the color vector associated with the excluded dyes (e.g., 510) is adjusted by minimizing the metric, a new multiple image 512 can be generated and / or made available. This multiple image 512 can be colored with dyes associated with the initial color vector 315, including one or more excluded dyes whose color adjustments are calculated based on spatial traversal techniques and metric. By utilizing the aforementioned color unmixing process 335, one or more new synthetic single images 514 associated with the excluded dyes can be generated.
[0152] Figure 6 illustrates an exemplary flowchart of process 600 for generating a synthetic single image using fine-tuned color vectors. At box 605, a color vector is determined for each of at least three digital pathology stains. The color vectors can be determined using the technique described relative to box 405 of process 400 (or another technique disclosed herein).
[0153] At box 610, digital pathology is accessible, where the image depicts a sample stained with one or more stains associated with an initial color vector 315 but without at least one of these stains. Each of the at least three stains that is absent from the sample but whose color vector is determined is referred to herein as an "excluded stain." Digital pathology can be a multiple (e.g., dual) image or a single image.
[0154] At box 615, the initial color vector 315 can be fed into filter 504, which is configured to generate a filtered output from the digital pathology image based on excluded stains. Filtering can be performed using linear techniques (e.g., NMF) or nonlinear techniques (e.g., machine learning models).
[0155] At box 620, a metric characterizing the signal properties in the filtered output is generated. Since the samples depicted in the known digital pathology images are not stained with a second staining agent, the optimal filtered output will not include the signal and will be blank. The metric can include any measure indicating the presence of the signal. For example, the metric can include statistics related to intensity values, such as mean, median, mode, variance, standard deviation, and / or range.
[0156] At box 625, a metric is used to generate an adjusted color vector for the second staining agent. In some cases, the adjusted color vector is used to automatically generate the adjusted color vector. For example, spatial traversal techniques (e.g., gradient descent, Monte Carlo methods) can be used, where the filtered output and metric are dynamically updated as the space is traversed. As another example, the interface and backend system can be configured such that the filtered output and metric are dynamically updated as the user of the interface adjusts the definition of the color vector for the second staining agent.
[0157] At box 630, a new image depicting the sample stained with a second staining agent is received. The sample may also, but need not, have been stained with one or more other biomarkers and / or a reference staining agent (e.g., one or more other staining agents among at least three staining agents).
[0158] At box 635, the adjusted color vector and the new image are used to generate a composite single image. For example, linear or nonlinear techniques can be used to process the new image to generate the composite image. At box 620, the linear or nonlinear techniques (e.g., and their associated parameters) can be the same as those used to generate the filtered output.
[0159] At box 640, the composite single-weight image is output. For example, the composite single-weight image can be transmitted to and / or displayed on a user device. It will be appreciated that in some cases, multiple composite single-weight images are generated and output at boxes 635 and 640, where each composite single-weight image is generated using a different color vector. In some cases, the other color vector is a color vector modified after the metric is generated. For example, at box 625, the interface can be configured to dynamically generate and dynamically render the metric (e.g., and the composite single-weight image) in response to modifications to the color vector representing a second staining agent and / or modifications to one or more other color vectors representing one or more other staining agents among at least three staining agents. In some cases, the other color vector is a color vector initially determined at box 605.
[0160] Figure 7 illustrates an exemplary flowchart of process 700 for identifying recommended color vectors. At box 705, a color vector is determined for each of at least one dye. The color vector(s) can be determined using the technique described relative to box 405 of process 400 (or another technique disclosed herein).
[0161] At box 710, access is available for a true multiple image stained with at least one staining agent associated with one or more initial color vectors. For example, a dual image stained with two biomarker staining agents and a dichromatin (e.g., hematoxylin) can be accessed. As another example, a single image stained with one biomarker staining agent and a dichromatin can be accessed. The goal may be to identify potential additional staining agents that are effectively distinguishable between one or more existing staining agents, such that a triple image using two existing biomarker staining agents and potential additional staining agents can be reliably and accurately demixed into three synthetic single images.
[0162] At box 715, an initial color vector for an additional potential stain can be identified. This identification can be performed automatically or based on user input. For example, the position of each of at least one digital pathology stain in the optical density space can be determined based on the color vector determined at box 705. Automation techniques can use an objective function to identify another position in the optical density space that prioritizes maximizing (or minimizing) the distance in that space relative to one or more positions associated with at least one digital pathology stain. As another example, the interface can display the position and / or vector of at least one digital pathology stain and can receive user input defining another position and / or vector associated with the initial color vector.
[0163] At box 720, a filtered output is generated by filtering the real multiple image using an initial color vector. Filtering can include linear or non-linear filtering. For example, filtering can use NMF or a machine learning model.
[0164] At box 725, a metric characterizing the signal properties in the filtered output is generated. Since the depicted sample is not stained with an additional staining agent, the objective function can be defined such that the filtered output lacks signal and / or information. This can indicate that the signal detected via the additional staining agent is independent of at least one staining agent.
[0165] Signal characteristics can characterize, for example, the quantity, variation, or complexity of a signal. Signal characteristics may include, for example: the mean, median, mode, variance, standard deviation, and / or range of intensity; spatial contrast; etc. Signal characteristics may additionally or alternatively characterize the degree to which a filtered output corresponding to an initial color vector differs from another filtered output corresponding to another color vector (e.g., at least one vector).
[0166] At box 730, a metric is used to identify a recommended color vector. This color vector can be the same as the initial vector or a different vector. In some cases, the metric is used to determine whether to adjust the recommended color vector. For example, automated algorithms can be used to iteratively evaluate the metric and adjust the color vector for additional potential dyes until predefined conditions are met (e.g., reaching a target metric, iterative improvement on the metric falling below an improvement threshold, a predefined number of iterations occurring, etc.). As another example, the metric and color vector for additional potential dyes can be displayed and dynamically updated in an interface, and user input can be received to adjust the color vector, ultimately accepting a given color vector for the additional potential dye.
[0167] Recommended color vectors can be output (e.g., once determined, once accepted, during iteration, etc.). These recommended color vectors can be used for announcements or to select configurations for additional potential colorants.
[0168] Figure 8 shows an exemplary flowchart of the process 800 for determining a performance prediction score, which represents the predicted degree to which at least three digital pathology stains can be sufficiently separated in practice to reliably support the generation of synthetic single images.
[0169] At box 805, a color vector 315 is determined for each of at least one dye. The color vectors (one or more) can be determined using the technique described relative to box 405 of process 400 (or another technique disclosed herein).
[0170] At box 810, access a real digital pathology image depicting a sample stained with one or more stains associated with an initial color vector 315, but excluding at least one of these stains (referred to as the "excluded stain"). The digital pathology image can be, for example, a dual or single image.
[0171] At box 815, an initial color vector 315 is fed into filter 504, which is configured to generate a filtered output from the real digital pathology image 502 based on excluded stains. One or more machine learning models (e.g., one or more generative models) can be trained to learn a mapping from a given multi-image to its constituent single image for filtering purposes. As an example, a conditional GAN can be used as a filter such that if the model is conditional on a color vector that is not present in the input multi-image, the model may fail to generate any meaningful output associated with this stain, resulting in zero values or empty images. Conversely, if such a model is given a color vector present in a given multi-image, the model can generate a constituent synthetic single image associated with this color vector.
[0172] At box 825, a performance prediction score is generated for the filtered output and / or for other synthetic single images constituting the real image. Finally, at box 830, the performance prediction score is output (e.g., transmitted to a user device and / or displayed at the user device). When it is predicted (or known) that a biomarker corresponding to a given staining agent is present in a given depicted sample or multiple images, it can be expected that the performance prediction score (e.g., mean, median, mode, variance, standard deviation, and / or range) may be relatively high when using an accurate color vector compared to when using a less accurate color vector. When it is predicted (or known) that no biomarker corresponding to a given staining agent is present in a given depicted sample, it can be expected that the mean, median, mode, variance, standard deviation, and / or range may be relatively low when using an accurate color vector compared to when using a less accurate color vector. Therefore, the performance prediction score can be generated in such a way that when the presence of a biomarker for the corresponding staining agent in the depicted sample is known or predicted, the score is positively correlated with the mean, median, mode, variance, standard deviation, range, and / or the degree to which the staining agent can be effectively distinguished in the synthetic single-weighted image.
[0173] In one scenario, performance prediction scores can be estimated by grouping similar staining agents together based on staining features. For example, staining features can include optical density values, color histograms, or any other features that can effectively capture staining patterns. These features can be clustered using clustering techniques such as k-means, hierarchical clustering, or density-based spatial clustering of applications and noise (DBSCAN). For example, k-means clustering can be used when the number of clusters is previously known. This clustering algorithm partitions the feature space into clusters, where each cluster represents a group of stained regions with similar staining patterns. The clustering process aims to minimize intra-cluster distances (distances between points within the same cluster) and maximize inter-cluster distances (distances between points in different clusters). Finally, performance prediction scores that evaluate the quality of the clusters can be estimated using metrics such as contour scores, the Davis-Bolding index, distances (e.g., Euclidean distance, Mahalanobis distance, or Manhattan distance), or visual inspection.
[0174] In another scenario, performance prediction scores can be calculated for synthesized single-weighted images by estimating the correlation between each staining pattern observed in multiple images. For this purpose, a correlation coefficient (ρ) can be calculated for OD single-weighted images derived from RGB, which provides a standardized and quantitative representation of staining intensity by measuring the absorption of light by the stained tissue. This method takes into account variations in staining scheme, image acquisition settings, and tissue characteristics, achieving a consistent basis for comparison. Furthermore, OD values are inherently in a non-negative to positive range, which aligns well with the physical constraints of staining intensity. The correlation coefficient between two single-weighted OD images A and B can be calculated according to… The Pearson correlation was used to calculate, where A i and B i For the columns of the OD matrix, and , The corresponding average values are given. The absolute values of the correlation coefficients range from 0 to 1, where 1 indicates a perfect linear relationship, and values close to 0 indicate that the staining is sufficiently separable. This score represents the degree to which staining is separable in a synthetic single image and can be used as a measure of the suitability of the synthetic image for various applications, such as image analysis, pathology, and medical diagnosis.
[0175] Multiplexed digital pathology images can represent the complexity involved in visually examining the intensity of multiple staining agents that co-localize within cells. Demixing multiplexed images becomes more difficult when multiple biomarkers, such as more than three or four, co-localize. For example, input real / synthetic triple images may include multiple different staining agents configured for uptake by progesterone receptor (PR), human epidermal growth factor receptor (HER), and estrogen receptor (ER). Additionally, real and / or synthetic multiplexed images may include signals from counterstaining biomarkers and / or hematoxylin configured to stain cell nuclei. For staining, PR can be stained with carboxytetramethylrhodamine (TAMRA), HER2 with green, and ER with benzyl sulfonyl (dansyl) and a blue counterstaining IHC marker, which utilizes hematoxylin for nuclear staining.
[0176] Estrogen is a hormone that can act as a contributing factor, particularly in breast and endometrial cancers. Estrogen binds to its estrogen receptor (ER), triggering a series of cellular responses involved in the proliferation and differentiation of specific cells. The estrogen receptor (ER) and progesterone receptor (PR) are biomarkers used in cancer pathology to assess the presence of receptors for estrogen and progesterone in tumor cells. ER and PR are nuclear receptors primarily located in the nucleus of cancer cells. Staining patterns of ER and PR can help identify the subcellular localization of these biomarkers. For ER, the commonly used antibody is ER-α. The staining agent is typically visualized using a chromogen (e.g., DAB). Progesterone staining may involve the use of PR antibodies, and the resulting staining agent can also be visualized using DAB.
[0177] Regarding unmixing, in one aspect of this disclosure, constraints can be introduced to simplify staining analysis, thereby reducing the complexity involved in staining unmixing. This technique can facilitate higher accuracy, precision, and / or reliability in generating synthetic single-weighted images from a given multiple image depicting a sample stained with, for example, three or more dyes / stains. In the disclosed technique, each pixel of the multiple image can be mapped to a location within a multidimensional color map. Pixels within specific portions of the color map (e.g., quadrants, portions defined by y-values greater than / less than x-values, wedge regions, etc.) can be assigned pixel-specific color vectors that predict the expression level of a first biomarker corresponding to this portion (e.g., based on grayscale optical density for the specific portion) and a "0" (or other predefined expression level) for each other biomarker corresponding to the multiple image. For pixels outside this specific portion, the unmixing technique can predict the expression levels of other biomarkers, maintaining predefined expression levels, such as a "0" for the first biomarker or other predefined numbers. In some cases, a particular part can be defined by inequalities with respect to the x and y coordinates (such as x > 25 and y < -15).
[0178] To extract specific portions from a color space, a GUI can be provided that interactively offers a toolset for defining portions of multiple images mapped to a multidimensional color space. These tools may include, but are not limited to, wedge regions, facets, exteriors, cylindrical regions, curves, elliptical regions, brush tools, or freeform selections. For example, a wedge region can be defined by selecting a center point and angle; an exterior can allow selection of points along the boundaries of the target region; a brush tool can define portions by allowing direct painting onto a chromaticity map, adjusting the brush size and shape to select target regions with different granularity levels. Tools can also be provided to incorporate thresholding techniques, where the user can specify thresholds for the x and y values used to define the portion. Additionally, freeform tools can provide the flexibility to define portions where predefined shapes may not adequately capture the target region. Once a portion is defined or selected within the color space, the GUI can be configured to perform actions such as assigning specific values to the remaining portions. The GUI can be configured to provide a corresponding matrix to apply unmixing techniques (such as those disclosed) to the remaining portions where the extracted portion is assigned "0".
[0179] In some embodiments, the multidimensional color space includes the International Commission on Illumination (CIE) color space, also known as the CIE XYZ color space. This color space is a standardized system for representing colors based on human perception. This standardized system defines three primary colors: X, Y, and Z, where Z represents luminance (brightness), and X and Y represent chromaticity (e.g., hue and saturation). For applications such as staining or color analysis, only XY can be used.
[0180] As an illustrative example, Figure 9A shows an example of an ER-PR-HER2 triple image, where each pixel of the triple image is mapped to a position within a cx-cy map (used as a multidimensional color map). To illustrate the disclosed technique, Figure 9A shows an example of an ER-PR-HER2 triple image 910 and a corresponding cx-cy map 915. In the distribution map 915, pixels 915a, 915b, 915c, and 915d represent color vectors associated with dansyl (ER), TAMRA (PR), green (HER2), and hematoxylin. It can be observed from Figure 915 that hematoxylin, TAMRA, dansyl, and HER2 are distributed in the first, second, third, and fourth quadrants, respectively. Therefore, by utilizing the disclosed constraint method, the optical density of different color distributions can be separated. These extractions can be performed using different constraints, such as linear regions, cylindrical regions, wedge regions, etc. For example, in Figure 920, only the green dye is extracted in the fourth quadrant. Depending on the constraint method, linear regions or other constraints can be used to extract dansyl, TAMRA, and hematoxylin signals from the triple image 910.
[0181] Figure 9B is an illustration of staining unmixing of an exemplary triple ER-PR-HER2 image 910 from Figure 9A according to some embodiments of the present disclosure. As shown in cx-cy figure 925 of Figure 9B, the residual distribution of ER-PR-HER2 includes TAMRA, tansyl sulfonate, and hematoxylin, thus producing a dual image. This figure can be achieved by using a bifaceted wedge region 960b from the constraint toolbox 960, which separates green from the remaining staining agents. In cx-cy figure 930, the hematoxylin signal can be extracted from the residual distribution of ER-PR-HER2 using facets (lines) 960d that connect the color vectors associated with TAMRA pixel 915b and green 915c. Similarly, in cx-cy figure 935, the residual distribution is the same as that of Figure 925, where different facets 960d connect tansyl sulfonate and hematoxylin. The resulting distribution can be seen in Figure 940, which can be achieved by applying the coloring and unmixing technique 335 described above. Using the constraint toolbox 960, the distribution can be divided into four quadrants by selecting the xy quadrant separation constraint 960a, as shown in Figure 915.
[0182] Figure 9C shows an example of staining unmixing results using the disclosed constraint technique for ER-PR-HER2 triplet images and one or more singlet images. The ER-PR-HER2 triplet image 962, stained with dansyl, TAMRA, green, and hematoxylin (a counterstain for the nucleus), is shown in Figure 9C. By utilizing the constraint technique, the triplet image 962 was unmixed into singlet images constituting dansyl (ER) 964, TAMRA (PR) 966, green (HER2) 968, and hematoxylin 970. Similarly, using the disclosed constraint technique, the dansyl singlet image 974, TAMRA singlet image 976, and green singlet image 978 were unmixed with adjacent registered singlet images in the bottom row of Figure 9C. These results demonstrate that this technique can be effectively used to obtain staining unmixing for different expression levels of low / medium / high HER2 (green).
[0183] Figure 9D shows an example of the color remixing results of the ER-PR-HER2 triple image and one or more single images using the disclosed constraint technique. Color remixing can be performed via process 345 as indicated above in Figure 3D. The top row 980 shows the remixing results of the ER-PR-HER2 triple image, and the bottom row 982 shows the baseline ground truth triple image and adjacent registered true single images used for comparison.
[0184] Figure 9E illustrates the staining remixing results of triple ER-PR-HER2 according to some embodiments of the present disclosure. In this example, the disclosed constraint technique is used to demix the triple ER-PR-HER2 image 992 to its constituent chromatographs. The disclosed constraint technique can extract individual signals by applying constraints (e.g., linear region, wedge region, cylindrical region constraints) from the provided interface. The extracted staining signals can then be remixed with the counterstain hematoxylin to obtain synthetic remixed dansyl 994, synthetic remixed TAMRA 996, and synthetic remixed green 998, as shown in Figure 9E.
[0185] Figure 10A illustrates an exemplary flowchart of the staining unmixing process 1000-A. Constraint techniques can support more accurate, precise, and / or reliable generation of synthetic single images from multiple images (e.g., slices depicting samples stained with three or more dyes or four or more dyes). Constraints can be added for staining unmixing, which can have the effect of reducing the complexity of potential color analysis.
[0186] At box 1005, a color vector is determined for each of at least four digital pathology stains. One or more color vectors may be determined using the technique described relative to box 405 of process 400 (or another technique disclosed herein). The color vectors may be adjusted according to the technique indicated above in Figure 3A. In some cases, each pixel in the digital pathology image may be mapped to a location within a multidimensional color space. Among the four stains, a specific stain may be selected at box 1010 such that at box 1015, that specific stain is attributable to a portion of the color space (e.g., quadrants, portions defined by being greater than / less than a certain y value and greater than / less than a certain x value, wedge-shaped regions, cylindrical regions, etc.). A specific stain may be a stain for which it is predicted that it will not be co-expressed with one, more, or all of the other four stains. For example, a specific stain may include a stain configured for uptake by the cell nucleus (e.g., having a given biological property), while other stains may be configured for uptake by the cell membrane (e.g., having corresponding other biological properties). As another example, a particular staining agent may include a staining agent configured to be absorbed by the cell membrane (e.g., having a given biological property), while other staining agents may be configured to be absorbed by the cell nucleus (e.g., having corresponding other biological properties).
[0187] At box 1020, a true multiplex image is accessed, depicting a sample (e.g., a tissue slide) stained with at least three digital pathological stains. At box 1025, each pixel of the true multiplex image can be mapped to a point in multidimensional space. At box 1030, for each pixel, a pixel-specific vector can be generated that predicts the expression level of each of the at least four stains in the portion of the biopsy slide depicted at that pixel. Finally, at box 1035, one or more synthetic single-multix images can be generated using the pixel-specific color vector.
[0188] Figure 10B further illustrates an exemplary flowchart of component 1030 from Figure 10A. At box 1030a, each pixel in a first subset of the pixel set is mapped to a point within a specific portion of the color space. At box 1030b, for each pixel associated with the specific portion, the expression level of a biomarker associated with a specific stain (associated with that portion) is predicted based on the pixel's optical density. For example, this portion of the color space could be a quadrant or wedge region associated with the green channel, and the predicted expression of the biomarker associated with the green stain assigned to each pixel in the quadrant or wedge region could be defined as the pixel's optical density. In some cases, the predicted expression level of a biomarker for each of at least four other stains could be set to zero or another constant.
[0189] At box 1030c, a second subset of the pixel set is defined, wherein each pixel in the second subset is mapped to a location outside this portion of the color space. At box 1030d, for the pixels in the second subset, a demixing technique (such as NMF) is performed to predict the expression level of each biomarker associated with other staining agents among at least four staining agents (excluding staining agents associated with this portion). In some cases, the expression of biomarkers associated with staining agents (associated with this portion of the color space) can be defined as zero.
[0190] In multiplex immunohistochemistry (mIHC), digital pathological images can be termed, for example, single, double, triple, etc., depending on the number of different markers or staining agents used for staining. For instance, single staining allows visualization of a specific target or protein on a tissue section using a single marker or staining agent and a counterstain. Similarly, in double and triple staining, two and three different markers and counterstains, respectively, can be applied to simultaneously detect corresponding numbers of different antigens (target proteins) within a single tissue sample. This technique can be used to study multiple biomarkers or antigens in the same tissue section, providing comprehensive information on the cell interactions, heterogeneity, location, function, and visualization of these antigens. This type of multiple staining involves: multiple primary antibodies, each recognizing a specific target; and then visualization using corresponding secondary antibodies labeled with different chromogens or fluorophores. Furthermore, multiple staining (e.g., triple staining) saves time (compared to three simple staining methods), preserves valuable samples using less material, and allows detection to be performed on the same tissue section.
[0191] Exemplary implementation:
[0192] Exemplary implementations of the disclosed techniques are provided for staining and unmixing of multiplex digital pathology images 110a to 110n or singlex images 108a to 108m. In the following example, stained slides were scanned at 20x magnification on a VENTANA DP200 scanner and annotated with ten fields of view (FOV) per slide using HALO image analysis software. All FOVs were quality controlled (QC) by independent team members to maintain consistency in FOV placement throughout the slide.
[0193] As previously noted, the color vector (initial W matrix) 315 obtained from Nonnegative Matrix Factorization (NMF) 310 may not perform well for staining unmixing. Figure 11A depicts a comparison of staining unmixing 335 of a dual image 1105a using the initial color matrix 315 and the adjusted color matrix according to an exemplary implementation. The first row 1105 of Figure 11A represents the unmixing performance using the conventional NMF 310 method. Blanks or noise can be observed in the synthetic TAMRA 1105b (e.g., weak / blurred cell nuclei (hematoxylin) issues visible in the synthetic TAMRA 1105b). The second row 1110 of Figure 11A represents the unmixing performance of the dual image 1105a using the adjusted color matrix 325.
[0194] Higher-resolution nuclear depiction can be observed in other synthetic TAMRA 1110a, achieved by shifting the dansyl vector to the left or away from the hematoxylin vector in cx-cy space using the disclosed technique. This color vector modification enhances the intensity of nuclear hematoxylin and provides better nuclear signal (e.g., visibility of nucleoli, chromatin, etc.). The improved nuclear signal in the synthetic images is comparable to the signal quality in the baseline ground truth image 1110b and H&E image 1110c. It is understood that the baseline ground truth singlet / multiplex images are derived from consecutive tissue slices representing corresponding adjacent singlet / multiplex images. For these baseline ground truth images, tissue morphology does not mismatch due to the fact that the images are derived from adjacent slides rather than the same slide. Therefore, tissue morphology differences still exist.
[0195] Figure 11B depicts a comparison of staining unmixing of another dual image 1115e and a single image 1115a using an initial color matrix and an adjusted color matrix. The color vectors of detected TAMRA stains (e.g., 1115d (image pointed to by the red arrow)) are investigated while the single tansyl 1115a and dual image 1115e are unmixed with very weak TAMRA, as illustrated in the illustrative examples in the first row of Figure 11B. The original color vectors of TAMRA are adjusted until a color vector is obtained that unmixes the single tansyl image 1115a with low or zero TAMRA signal (i.e., showing a very low TAMRA background). “Very low” means an intensity close to that of tissue where no cells are present. As shown in the second row 1120 of Figure 11B, the same color vector can be used to accurately detect TAMRA signals in images of samples where such signals are present. In these examples, the single dansyl image 1115a consists of a blue (countercolorant, such as hematoxylin 1115b) and a yellow channel 1115c, and this single dansyl image is expected to have a very small TARMA signal (1115d). Starting from the initial color vector, as a semi-automatic method, the color vector is calibrated or fine-tuned via an interactive graphical user interface (GUI) 112 to adjust the color vector so that it can demix the multiple images with good quality. The second row 1120 of Figure 11B shows the improved background noise in the TARMA channel using the adjusted color vector.
[0196] Figure 12A shows an example of a dual image 1202, which overlays candidate species at each nucleus (marked with a red dot) detected by automatic nucleus segmentation. In this example, automatic nucleus segmentation is performed based on an iterative modified radial symmetry method, Parvin et al., 2007 (see reference). The algorithm is applied to the hematoxylin image 1206 channels after demixing the dual image 1202. Dual image 1202a provides a magnified view of a fragment from 1202, while dual image 1202b shows 1202a with marked candidate species at each nucleus indicated by red dots. Candidate species can serve as initial markers or reference points for nucleus segmentation. As shown in Figure 12A, these candidate species and species markers from dual image 1202 are detected and segmented using the demixed hematoxylin 1206 (e.g., 1212 provides the segmented image). Then, the intensities of dansyl 1210 and TAMRA 1208 were attached to each candidate species. A simple filter was applied to remove some stromal cells or cells with very low dansyl 1210 and TAMRA 1208 intensities. The intensities of dansyl 1210 and TAMRA 1208 were measured for each FOV.
[0197] Figure 12B depicts a comparison of cell nucleus segmentation results for hematoxylin images obtained by demixing dual images using linear deconvolution (e.g., 1214) and NMF (e.g., 1216) techniques. It can be observed from the images that the hematoxylin image channels demixed using the linear deconvolution method (e.g., 1214) are blurred, and the cell regions are not well defined. In contrast, NMF more accurately demixes the cell nucleus regions, producing improved cell nucleus definitions that are separated from the background (as shown in 1216).
[0198] Linear deconvolution and NMF (e.g., 1218 and 1220, respectively) were also used on single-weighted slides, and it was determined that the nucleus segmentation results obtained from both unmixing methods showed comparable performance, as shown in the second row of Figure 12B. Table 1 lists the number of nuclei obtained from double and single images using linear deconvolution and NMF methods (with fine-tuning), respectively. This shows that the number of nuclei obtained from double images using linear deconvolution is significantly larger than that using NMF methods, while the table does not show a significant difference in single images using the two methods. A first double image was generated using the linear unmixing method to produce a synthetic single image. As shown in Table 1, 784 nuclei were detected in the synthetic single image. Meanwhile, the real adjacent single images depicted 563 nuclei, indicating a substantial inconsistency. Simultaneously, a second double image was generated using the NMF method (which includes unmixing) to produce a synthetic single image. As shown in Table 1, 624 nuclei were detected in the synthetic single image. Meanwhile, the true adjacent singlet image depicts 533 cell nuclei. Therefore, the estimated NMF results are more accurate than those obtained using the linear unmixing method.
[0199] Table 1. Number of cell nuclei obtained from double and single images using linear deconvolution and NMF methods, respectively.
[0200]
[0201] Figure 13 illustrates an exemplary graphical user interface (GUI) for generating composite pixels. As previously noted, the interface can be configured to interactively blend two or more dye colors at different ratios and display them in a cx-cy diagram. Additionally, the interface can be used to adjust a given color vector associated with a specific color primord. For example, in exemplary interface 1300, a set of four color primors (1305) and their associated RGB values are shown, such as dansyl [0.7108, 0.5888, 0.3849], TAMRA [0.9082, 0.3621, 0.21], cyan [0.244, 0.8821, 0.403], and hematoxylin [0.145, 0.2969, 0.9438]. The corresponding color vector is plotted as pixels in cx-cy Figure 1310, where the marker (“X”) represents the initial color vector for tansyl, which can be adjusted by various adjustment options from interface 1300. For example, the concentration ratio (amount) for each color source in color source 1305 can be selected via interface 1312. In addition to the color source amount, interface 1300 can also implement hue saturation adjustment 1314 for tansyl. By utilizing the adjustment options 1312 and 1314 for tansyl, the adjusted tansyl 1315 can be observed with the updated color vector [0.7882, 0.6784, 0.3686], where the concentration ratios from each color source are [1, 0.05, 0.05, 0.05, 0.05, 0.05]. By further adjusting the concentration ratio to [1, 0, 0, 0] by selecting an option from 1312, tansyl 1320 can be produced, with the updated RGB values being [0.8667, 0.7412, 0.3882].
[0202] When adjusting the amount of dye / color source, the adjusted amount is multiplied by the corresponding color vector in the OD space and then converted back to the RGB space for display. In interface 1300, the adjusted pixel for dansyl, indicated by ('*'), is shown in the cx-cy diagram. The positions of two pixels (e.g., the initial dansyl ('X') and the adjusted dansyl ('*')) in the cx-cy diagram show how close the two pixels are in hue and saturation. Increasing the dye dosage while maintaining the relative ratio of the color sources (e.g., keeping the composition constant) does not change the position of the pixel in the cx-cy diagram, but it does change the appearance of the composite pixel. This is consistent with the design of the cx-cy space, which counts only hue and saturation while keeping density constant.
[0203] This user interface enables: (1) visual inspection of a range of colors generated from a specific combination of chromogens determined from biomarkers for both pathologists and algorithm developers; (2) providing a baseline truth for staining unmixing since the components of each chromogen that produces the synthetic color stain are known, thus enabling color unmixing for a set of synthetic pixels and comparison of the results with known settings used to generate these synthetic pixels; (3) investigating possible unmixing errors (e.g., missing staining signals in some parts of the unmixed image) when various regularizations (such as wedge region constraints) for NMF-based unmixing are applied; and (4) aiding in the selection and comparison of chromogens by assessing which chromogens are more suitable for unmixing.
[0204] Figure 14A illustrates an exemplary GUI for evaluating a range of colors from the blending of multiple chromogens using composite pixels. In multi-images, different chromogen colors can be blended when multiple biomarkers stain the same or proximal structures in a tissue (e.g., a portion of a tissue expressing multiple proteins detected by assay). This color blending can generate a range of colors depending on the nature of the chromogens and the relative amount of each chromogen deposited in the tissue structure. Figure 14A includes a cx-cy diagram (e.g., 1402) representing pixels associated with the constituent chromogens of a triple image. In this Figure 1402, an exemplary color is generated by blending green and QM-dansyl (green row 1408 of the composite pixels), which is represented in the cx-cy diagram 1402 as the marker “*” (pixel 1). Another exemplary color exists, generated by mixing cyan and QM-dansyl (cyan row 1410 of the composite pixel), which is represented by the marker "x" (pixel 2) in cx-cy Figure 1402. Different ranges of green and cyan can be achieved by adjusting these color vectors (associated with composite pixel 1 and pixel 2) via various adjustment options in the interface, as shown in cx-cy Figures 1404 and 1406 together with the color vectors generated in rows 1408 and 1410, respectively.
[0205] Furthermore, as discussed in process 700 of Figure 7, the composite pixel generation interface can facilitate the selection and comparison of color sources by assessing which color sources are more suitable for demixing. Specifically, the color mixing of the prefix color source set with another color source in the test can be examined, and then quantitatively (by calculating the proximity of the mixed colors in the cx-cy space) and qualitatively (by visual inspection) determined which color source generates a color range that supports accurate color demixing. For example, if the mixed color is similar to another color source when a candidate color source is mixed with other color sources, then that candidate color source can be the next choice.
[0206] Figure 14A further illustrates the recommended colors for determining the third color source for triple determination. In the example where cyan and green are chosen as the third color source for triple determination (excluding yellow QM-dansyl and purple TAMRA), it can be observed that mixing cyan (cyan) or green with QM-dansyl (purple) can generate pixels with different appearances corresponding to a wide range of hues and saturations, but mixing with cyan produces a staining color similar to hematoxylin, as shown in row 1410 of Figure 14A. This color similarity to hematoxylin pixels may increase the difficulty of staining demixing because demixing errors may occur when the mixed pixel value is close to hematoxylin, as the algorithm incorrectly demixes such pixels as cyan and QM-dansyl instead of hematoxylin. The results shown in the exemplary implementation of Figure 14A indicate support for green instead of cyan as the third color source.
[0207] Figure 14B illustrates a comparison of one or more blended colors synthesized from dyes from different reagent sources according to an exemplary implementation. In this example, the same type of chromogen (“green”) from different reagent sources has been examined. The interface displays the associated colors of the synthesized pixels generated from the user interface as discussed above and in Figure 13. Figure 14B includes a color representation of pure hematoxylin 1420, which is a blend of TAMRA and green from source 1 (batch number: H27689) and source 2 (batch number: H35597) in boxes 1415a and 1415b, respectively. Chromogens from different reagent sources can produce dyes with slightly different colors, and the disclosed method can help select the most reliable and advantageous reagent source.
[0208] Figure 14C illustrates an example of generating a range of colors by blending two or more chromogens according to an exemplary implementation. Possible unmixing errors may include, for example, missing chromogen signals in the unmixed single image when various regularizations for NMF-based unmixing are applied. Using the constraint toolkit 960 for NMF, triple images can be unmixed based on biomarker localization. Figure 14C shows an exemplary triple image 1424 of MET-PDL1-EGFR and the associated cx-cy image 1422. Using wedge-shaped region constraints of the NMF method, the hematoxylin signal 1422a can be separated from other positive biomarker chromogens in the first quadrant of the cx-cy space 1422. The remaining signals can be unmixed into dansyl, TAMRA, and green using 3-color unmixing in all other quadrants. The wedge-shaped region constraint applied to hematoxylin and the pixels allocated within the wedge region can be seen in Figure 1425, represented by the red triangle in Figure 14C (cx-cy). Using the synthesized pixels to generate a user interface, the color range allocated to hematoxylin within the wedge region can be evaluated. These color ranges can be the result of incorporating TAMRA, green, and QM-dansyl.
[0209] Specifically, in cx-cy diagram 1430 of Figure 14C, the mixing of TAMRA and blue-green can generate a range of colors that can lie on the line connecting the TAMRA and green color vectors (the dark red line in Figure 1430). A small amount of QM-dansyl can pull the color inwards into the wedge-shaped region. Such mixed colors can be assigned to hematoxylin, thus producing unmixing errors. The degree of error can depend on the relative amount of each chromogen, which corresponds to the expression level of each biomarker detectable by biomarker assays.
[0210] Figure 14D illustrates the evaluation of a range of colors assigned to hematoxylin with wedge-shaped region constraints. In Figure 14D, an interface allows visualization of one or more exemplary colors that might be incorrectly assigned to hematoxylin. Quantitatively, the range of colors assigned to hematoxylin can be calculated using, for example, the L2 norm of their cx-cy values between the two colors at the intersection of the wedge-shaped region line and the line connecting the TAMRA and green color vectors (gray arrows in Figure 14C).
[0211] Some embodiments of this disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein. Some embodiments of this disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein.
[0212] This description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, this description of the preferred exemplary embodiments will provide those skilled in the art with a feasible description for implementing various embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope set forth in the appended claims.
[0213] Specific details are set forth in this description to provide a thorough understanding of the embodiments. However, it should be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as parts in block diagram form to avoid obscuring the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments.
Claims
1. A computer-implemented method, the method comprising: determining, for each of at least three digital pathology stains, a color vector representing the stain; providing an interface to a user device, wherein the interface includes: a representation of each of the determined color vectors; a real multi-plex digital pathology image, the real multi-plex digital pathology image depicting a biopsy slide stained with two or more of the at least three digital pathology stains; at least one synthetic single-plex image, wherein each of the at least one synthetic single-plex image is generated by filtering the real multi-plex digital pathology image using a single one of the determined color vectors; and one or more color-vector adjustment tools, wherein each of the one or more color- vector adjustment tools is configured to receive user input corresponding to an adjustment to a color vector representing a corresponding one of the at least three digital pathology stains; detecting input received via interaction with the interface, the input corresponding to a particular adjustment to a color vector representing a particular one of the at least three digital pathology stains; and in response to detecting the input, automatically updating the interface.
2. The computer-implemented method of claim 1, wherein the representation of each of the determined color vectors includes a representation of a location within an optical density space.
3. The computer-implemented method of claim 1, wherein the updated interface further includes the at least one synthetic single-plex image.
4. The computer-implemented method of claim 1, wherein the one or more color- vector adjustment tools include at least three color adjustment tools.
5. The computer-implemented method of claim 1, wherein determining the color vectors includes processing one or more single-stain images, the one or more single-stain images depicting the same biopsy slide or other biopsy slides that have been stained with only one of the at least three digital pathology stains.
6. The computer-implemented method of claim 1, wherein one or more single-stain images include a marker overlaid at a particular location within a multi-dimensional color space, and wherein the color vector is defined based on the particular location.
7. The computer-implemented method of claim 1, wherein the real multi-plex digital pathology image depicts the biopsy slide stained with at least four stains.
8. The computer-implemented method of claim 1, wherein the determined color vectors are within a two-dimensional color space, and wherein the method further comprises: determining a portion of the color space that is predicted to be attributable to a prominent signal corresponding to a particular one of the at least three digital pathology stains; wherein the automatic updating of the interface is performed using a unmixing technique that selectively focuses on the at least three digital pathology stains and excludes the particular stain.
9. The computer-implemented method of claim 1, wherein the determining of the color vectors is performed using non-negative matrix factorization.
10. The computer-implemented method of claim 1, further comprising: receiving a new multi-stain image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stain image based on the new multi-stain image and the adjusted color vectors; and outputting the new synthetic single-stain image.
11. The computer-implemented method of claim 1, wherein the real multi-stain digital pathology image is filtered using the color vectors and a machine learning model.
12. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: determining, for each of at least three digital pathology stains, a color vector representing the stain; providing an interface to a user device, wherein the interface includes: a representation of each of the determined color vectors, wherein the representation of each of the determined color vectors includes a representation of a location within an optical density space; a real multi-stain digital pathology image depicting a biopsy slide stained with two or more of the at least three digital pathology stains; at least one synthetic single-stain image, wherein each of the at least one synthetic single-stain image is generated by filtering the real multi-stain digital pathology image using a single color vector of the determined color vectors; and one or more color-vector adjustment tools, wherein each of the one or more color- vector adjustment tools is configured to receive user input corresponding to an adjustment to a color vector representing a corresponding stain of the at least three digital pathology stains; detecting input received via interaction with the interface, the input corresponding to a particular adjustment to a color vector representing a particular stain of the at least three digital pathology stains; and in response to detecting the input, automatically updating the interface, wherein the updated interface further includes the at least one synthetic single-stain image.
13. The system of claim 12, wherein determining the color vectors includes processing one or more single-stain images depicting the same biopsy slide or other biopsy slides that have been stained with only one of the at least three digital pathology stains.
14. The system of claim 12, wherein one or more single-stain images include a marker overlaid at a particular location within a multi-dimensional color space, and wherein the color vector is defined based on the particular location.
15. The system of claim 12, wherein the determined color vectors are within a two- dimensional color space, and wherein the operations further comprise: determining a portion of the color space that is predicted to be attributable to a prominent signal corresponding to a particular stain of the at least three digital pathology stains; wherein the automatic updating of the interface is performed using an unmixing technique that selectively focuses on the at least three digital pathology stains and excludes the particular stain, and wherein the automatic updating of the interface is performed using non-negative matrix factorization.
16. The system of claim 12, wherein the operations further comprise: receiving a new multi-stain image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stain image based on the new multi-stain image and the adjusted color vectors; and outputting the new synthetic single-stain image.
17. The system of claim 12, wherein the real multi-stain digital pathology image is filtered using the color vectors and a machine learning model.
18. A computer program product tangibly embodied in a non-transitory machine- readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform operations comprising: determining, for each stain of at least three digital pathology stains, a color vector representing the stain; providing an interface to a user device, wherein the interface includes: a representation of each of the determined color vectors, wherein the representation of each of the determined color vectors includes a representation of a location within an optical density space; a real multi-stain digital pathology image depicting a biopsy slide stained with two or more of the at least three digital pathology stains; at least one synthetic single-stain image, wherein each of the at least one synthetic single-stain image is generated by filtering the real multi-stain digital pathology image using a single color vector of the determined color vectors; and one or more color-vector adjustment tools, wherein each of the one or more color- vector adjustment tools is configured to receive user input corresponding to an adjustment to a color vector representing a corresponding stain of the at least three digital pathology stains; detecting input received via interaction with the interface, the input corresponding to a particular adjustment to a color vector representing a particular stain of the at least three digital pathology stains; and in response to detecting the input, automatically updating the interface, wherein the updated interface further includes the at least one synthetic single-stain image.
19. The computer program product of claim 18, wherein determining the color vectors includes processing one or more single-stain images depicting the same biopsy slide or other biopsy slides that have been stained with only one of the at least three digital pathology stains.
20. The computer program product of claim 18, wherein the operations further comprise: receiving a new multi-stain image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stain image based on the new multi-stain image and the adjusted color vectors; and outputting the new synthetic single-stain image. generating a new synthetic single-stain image based on the new multi-stain image and the adjusted color vector; and outputting the new synthetic single-stain image.
21. A computer-implemented method, the method comprising: determining, for each stain of at least three digital pathology stains, a color vector representing the stain; accessing a real multi-stain digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-stain digital pathology image using the color vector representing a second stain of the at least one second stain; generating a metric characterizing a signal property in the filtered output; identifying, using the metric and a spatial traversal technique, an adjustment to the color vector representing the second stain; receiving a new multi-stain image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stain image based on the new multi-stain image and the adjusted color vector representing the second stain; and outputting the new synthetic single-stain image.
22. The computer-implemented method of claim 21, wherein, for each stain of the at least three digital pathology stains, the color vector is a vector in an optical density space.
23. The computer-implemented method of claim 21, wherein the spatial traversal technique comprises a gradient descent technique.
24. The computer-implemented method of claim 21, wherein the spatial traversal technique comprises a Monte Carlo technique.
25. The computer-implemented method of claim 21, wherein the metric comprises a mean intensity, a median intensity, or a mode intensity.
26. The computer-implemented method of claim 21, wherein the metric characterizes a staining level across all or a portion of the filtered output.
27. The computer-implemented method of claim 21, wherein the filtered output is generated by using a machine learning model.
28. A system, the system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: determining, for each stain of at least three digital pathology stains, a color vector representing the stain; accessing a real multi-stain digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-stain digital pathology image using the color vector representing a second stain of the at least one second stain; generating a metric characterizing a signal property in the filtered output; identifying, using the metric and a spatial traversal technique, an adjustment to the color vector representing the second stain; using the metric and a space traversal technique to identify an adjustment to the color vector representing the second stain; receiving a new multi-stained image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stained image based on the new multi-stained image and the adjusted color vector representing the second stain; and outputting the new synthetic single-stained image.
29. The system of claim 28, wherein for each of the at least three digital pathology stains, the color vector is a vector in optical density space.
30. The system of claim 28, wherein the space traversal technique comprises a gradient descent technique.
31. The system of claim 8, wherein the space traversal technique comprises a Monte Carlo technique.
32. The system of claim 28, wherein the metric comprises a mean intensity, a median intensity, or a mode intensity.
33. The system of claim 28, wherein the metric characterizes a staining level across all or a portion of the filtered output.
34. The system of claim 28, wherein the filtered output is generated by using a machine learning model.
35. A computer program product tangibly embodied in a non-transitory machine- readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform operations comprising: for each of at least three digital pathology stains, determining a color vector representing the stain; accessing a real multi-stained digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-stained digital pathology image using the color vector representing a second stain of the at least one second stain; generating a metric characterizing a signal property in the filtered output; using the metric and a space traversal technique to identify an adjustment to the color vector representing the second stain; receiving a new multi-stained image stained with at least one of the at least three digital pathology stains; generating a new synthetic single-stained image based on the new multi-stained image and the adjusted color vector representing the second stain; and outputting the new synthetic single-stained image.
36. The computer program product of claim 35, wherein for each of the at least three digital pathology stains, the color vector is a vector in optical density space.
37. The computer program product of claim 35, wherein the space traversal technique comprises a gradient descent technique.
38. The computer program product of claim 35, wherein the space traversal technique comprises a Monte Carlo technique.
39. The computer program product of claim 35, wherein the metric comprises a mean intensity, a median intensity, or a mode intensity.
40. The computer program product of claim 35, wherein the metric characterizes a staining level across all or a portion of the filtered output, and wherein the filtered output is generated by using a machine learning model.
41. A computer-implemented method comprising: determining, for each of at least two digital pathology stains, a color vector representing the stain; accessing a real multi-stained digital pathology image, the real multi-stained digital pathology image depicting a biopsy slide stained with the at least two digital pathology stains; identifying a recommended color vector representing a potential additional stain by: identifying an initial color vector; generating a filtered output by filtering the real multi-stained digital pathology image using the initial color vector; generating a metric characterizing a signal property in the filtered output; and using the metric and a spatial traversal technique to identify the recommended color vector; and outputting the recommended color vector.
42. The computer-implemented method of claim 41, wherein the spatial traversal technique is performed to include minimizing a signal in the filtered output as one or more objectives in the traversal.
43. The computer-implemented method of claim 42, wherein minimizing the signal in the filtered output includes minimizing a mean, median, or mode intensity of a corresponding filtered output.
44. The computer-implemented method of claim 41, wherein the determining of color vectors is performed using non-negative matrix factorization.
45. The computer-implemented method of claim 41, wherein, for each of the at least two digital pathology stains, the color vector is a vector in an optical density space.
46. The computer-implemented method of claim 41, wherein the filtered output is generated by using a machine learning model.
47. The computer-implemented method of claim 41, wherein the spatial traversal technique includes a gradient descent technique.
48. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: determining, for each of at least two digital pathology stains, a color vector representing the stain; accessing a real multi-stained digital pathology image, the real multi-stained digital pathology image depicting a biopsy slide stained with the at least two digital pathology stains; identifying a recommended color vector representing a potential additional stain by: identifying an initial color vector; generating a filtered output by filtering the real multi-stained digital pathology image using the initial color vector; generating a metric characterizing a signal property in the filtered output; and using the metric and a spatial traversal technique to identify the recommended color vector; and outputting the recommended color vector.
49. The system of claim 48, wherein the spatial traversal technique is performed to include minimizing a signal in the filtered output as one or more objectives in the traversal.
50. The system of claim 49, wherein minimizing the signal in the filtered output includes minimizing a mean, median, or mode intensity of the corresponding filtered output.
51. The system of claim 48, wherein the determination of a color vector is performed using non-negative matrix factorization.
52. The system of claim 48, wherein for each of the at least two digital pathology stains, the color vector is a vector in an optical density space.
53. The system of claim 48, wherein the filtered output is generated by using a machine learning model.
54. A computer program product tangibly embodied in a non-transitory machine- readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform operations comprising: determining, for each of at least two digital pathology stains, a color vector representing the stain; accessing a real multi-stain digital pathology image, the real multi-stain digital pathology image depicting a biopsy slide stained with the at least two digital pathology stains; identifying a recommended color vector representing a potential additional stain by: identifying an initial color vector; generating a filtered output by filtering the real multi-stain digital pathology image using the initial color vector; generating a metric characterizing a signal property in the filtered output; and using the metric and a spatial traversal technique to identify the recommended color vector; and outputting the recommended color vector.
55. The computer program product of claim 54, wherein the spatial traversal technique is performed to include minimizing a signal in the filtered output as one or more objectives in the traversal.
56. The computer program product of claim 55, wherein minimizing the signal in the filtered output includes minimizing a mean, median, or mode intensity of the corresponding filtered output.
57. The computer program product of claim 54, wherein for each of the at least two digital pathology stains, the color vector is a vector in an optical density space.
58. The computer program product of claim 54, wherein the filtered output is generated by using a machine learning model.
59. The computer program product of claim 54, wherein the determination of a color vector is performed using non-negative matrix factorization.
60. The computer program product of claim 54, wherein the spatial traversal technique includes a gradient descent technique.
61. A computer-implemented method, the method comprising: determining, for each of at least three digital pathology stains, a color vector representing the stain; accessing a real multi-plex digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three digital pathology stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-plex digital pathology image using the color vector representing a second stain of the at least one second digital pathology stain; generating a performance prediction score representing a predicted degree to which the at least three digital pathology stains are sufficiently separable in practice to reliably support generation of synthetic single-plex images; and outputting the performance prediction score.
62. The computer-implemented method of claim 61, wherein the performance prediction score is generated using the filtered output.
63. The computer-implemented method of claim 61, wherein the performance prediction score comprises a mean, median, or mode intensity of corresponding filtered outputs.
64. The computer-implemented method of claim 61, wherein the performance prediction score comprises a correlation coefficient between each pair of synthetic single-plex images associated with the filtered output.
65. The computer-implemented method of claim 61, wherein the color vector is a vector in optical density space for each stain of the at least three digital pathology stains.
66. The computer-implemented method of claim 61, wherein the filtered output is generated by using a machine learning model.
67. The computer-implemented method of claim 61, wherein the color vector is adjusted based on the performance prediction score via a graphical user interface (GUI).
68. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: determining, for each stain of at least three digital pathology stains, a color vector representing the stain; accessing a real multi-plex digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three digital pathology stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-plex digital pathology using the color vector representing a second stain of the at least one second digital pathology stain; generating a performance prediction score representing a predicted degree to which the at least three digital pathology stains are sufficiently separable in practice to reliably support generation of synthetic single-plex images; and outputting the performance prediction score.
69. The system of claim 68, wherein the performance prediction score is generated using the filtered output. 70. The system of claim 68, wherein the performance prediction score comprises a mean, median, or mode intensity of a corresponding filtered output.
71. The system of claim 68, wherein the performance prediction score comprises a correlation coefficient between each pair of synthetic monochromatic images associated with the filtered output.
72. The system of claim 68, wherein the filtered output is generated by using a machine learning model.
73. The system of claim 68, wherein for each of at least two digital pathology stains, the color vector is a vector in an optical density space.
74. The system of claim 68, wherein the color vector is adjusted based on the performance prediction score via a graphical user interface (GUI).
75. A computer program product tangibly embodied in a non-transitory machine- readable storage medium, the computer program product including instructions configured to cause one or more data processors to perform operations comprising: determining, for each of at least three digital pathology stains, a color vector representing the stain; accessing a real multi-stain digital pathology image depicting a biopsy slide stained with at least a first stain of the at least three digital pathology stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least three stains; generating a filtered output by filtering the real multi-stain digital pathology using the color vector representing a second stain of the at least one second stain; generating a performance prediction score representing a predicted degree to which the at least three digital pathology stains are sufficiently separable in practice to reliably support generation of synthetic monochromatic images; and outputting the performance prediction score.
76. The computer program product of claim 75, wherein the performance prediction score is generated using the filtered output.
77. The computer program product of claim 75, wherein for each of at least two digital pathology stains, the color vector is a vector in an optical density space.
78. The computer program product of claim 75, wherein the performance prediction score comprises a correlation coefficient between each pair of synthetic monochromatic images.
79. The computer program product of claim 75, wherein the filtered output is generated by using a machine learning model.
80. The computer program product of claim 75, wherein the color vector is adjusted based on the performance prediction score via a graphical user interface (GUI).
81. A computer-implemented method, the method comprising: determining, for each of at least four digital pathology stains, a color vector representing the stain, wherein the determined color vectors are within a multi-dimensional color space; selecting a particular stain of the at least four digital pathology stains; determining, for each of at least two digital pathology stains, a color vector representing the stain, wherein the color vector is a vector in an optical density space; accessing a real multi-stain digital pathology image depicting a biopsy slide stained with at least a first stain of the at least two digital pathology stains, wherein the depicted biopsy slide is unstained with at least a second stain of the at least two stains; generating a filtered output by filtering the real multi-stain digital pathology using the color vector representing a second stain of the at least one second stain; generating a performance prediction score representing a predicted degree to which the at least two digital pathology stains are sufficiently separable in practice to reliably support generation of synthetic monochromatic images; and outputting the performance prediction score. determining a portion of the color space that is predicted to be attributable to a prominent signal corresponding to the particular stain; accessing a real multi-plex digital pathology image, the real multi-plex digital pathology image depicting a biopsy slide stained with at least three of the at least four digital pathology stains, wherein the real multi-plex digital pathology image comprises a set of pixels; mapping each pixel in the set of pixels in the real multi-plex digital pathology image to a point within the multi-dimensional color space; for each pixel in the set of pixels, generating a pixel-specific color vector that predicts, for each of the at least four digital pathology stains, a degree of expression of the stain in a portion of the biopsy slide depicted at the pixel, wherein generating the pixel-specific color vector comprises: determining that each pixel in a first subset of the set of pixels is mapped to a point within the portion of the color space; for each pixel in the first subset of pixels, determining an optical density, wherein the pixel-specific color vector for the pixel identifies a degree of expression of the particular stain corresponding to an optical density; determining that each pixel in a second subset of the set of pixels is mapped to a point outside of the portion of the color space; and performing an unmixing technique to predict, for each pixel in the second subset and for each of some of the at least four digital pathology stains, a degree of expression of the stain in the portion of the biopsy slide depicted at the pixel, wherein the some of the at least four digital pathology stains do not include the particular stain, and wherein the unmixing technique uses the color vectors determined to represent each of the some of the at least four digital pathology stains; and generating one or more synthetic single-plex images using the pixel-specific color vectors.
82. The computer-implemented method of claim 81, wherein the particular stain is selected based on information about which portions of a cell each of the at least four digital pathology stains is configured to stain.
83. The computer-implemented method of claim 81, wherein the color space comprises an International Commission on Illumination (CIE) color space.
84. The computer-implemented method of claim 81, wherein the portion of the color space comprises a wedge-shaped region.
85. The computer-implemented method of claim 81, wherein the portion of the color space comprises a portion of space defined based on an inequality with respect to an x-coordinate and an inequality with respect to a y-coordinate.
86. The computer-implemented method of claim 81, wherein the portion of the color space comprises a combination of primitives.
87. The computer-implemented method of claim 81, wherein performing the unmixing technique comprises using non-negative matrix factorization (NMF).
88. The computer- implemented method of claim 81, wherein the color vector is determined based on one or more user inputs received using one or more color- vector adjustment tools available within an interface.
89. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising: determining, for each of at least four digital pathology stains, a color vector representing the stain, wherein the determined color vectors are within a multidimensional color space; selecting a particular stain of the at least four digital pathology stains; determining a portion of the color space predicted to be attributable to a prominent signal corresponding to the particular stain; accessing a real multi-plex digital pathology image depicting a biopsy slide stained with at least three of the at least four digital pathology stains, wherein the real multi-plex digital pathology image comprises a set of pixels; mapping each pixel of the set of pixels in the real multi-plex digital pathology image to a point within the multidimensional color space; generating, for each pixel of the set of pixels, a pixel-specific color vector predicting, for each of the at least four digital pathology stains, an extent of expression of the stain in a portion of the biopsy slide depicted at the pixel, wherein generating the pixel-specific color vector comprises: determining that each pixel of a first subset of the set of pixels is mapped to a point within the portion of the color space; determining, for each pixel of the first subset of pixels, an optical density, wherein the pixel-specific color vector for the pixel identifies an extent of expression of the particular stain corresponding to the optical density; determining that each pixel of a second subset of the set of pixels is mapped to a point outside the portion of the color space; and performing an unmixing technique to predict, for each pixel of the second subset and for each of some of the at least four digital pathology stains, an extent of expression of the stain in the portion of the biopsy slide depicted at the pixel, wherein the some of the at least four digital pathology stains do not include the particular stain, and wherein the unmixing technique uses the color vectors determined to represent each of the some of the at least four digital pathology stains; and generating one or more synthetic monoplex images using the pixel-specific color vectors.
90. The system of claim 89, wherein the particular stain is selected based on information about which portions of a cell each of the at least four digital pathology stains is configured to stain.
91. The system of claim 89, wherein the color space comprises an International Commission on Illumination (CIE) color space. 92. The system of claim 89, wherein the portion of the color space comprises a wedge-shaped region or a combination of primitives.
93. The system of claim 89, wherein the portion of the color space comprises a portion of space defined based on an inequality with respect to an x-coordinate and an inequality with respect to a y-coordinate.
94. The system of claim 89, wherein performing the unmixing technique comprises using non-negative matrix factorization (NMF).
95. The system of claim 89, wherein the color vector is determined based on one or more user inputs received using one or more color-vector adjustment tools available within the interface.
96. A computer program product tangibly embodied in a non-transitory machine- readable storage medium, including instructions configured to cause one or more data processors to perform operations comprising: determining, for each of at least four digital pathology stains, a color vector representing the stain, wherein the determined color vectors are within a multidimensional color space; selecting a particular stain of the at least four digital pathology stains; determining a portion of the color space predicted to be attributable to a prominent signal corresponding to the particular stain; accessing a real multi-stain digital pathology image, the real multi-stain digital pathology image depicting a biopsy slide stained with at least three of the at least four digital pathology stains, wherein the real multi-stain digital pathology image comprises a set of pixels; mapping each pixel of the set of pixels in the real multi-stain digital pathology image to a point within the multidimensional color space; for each pixel of the set of pixels, generating a pixel-specific color vector predicting, for each of the at least four digital pathology stains, a degree of expression of the stain in a portion of the biopsy slide depicted at the pixel, wherein generating the pixel-specific color vector comprises: determining that each pixel of a first subset of the set of pixels is mapped to a point within the portion of the color space; for each pixel of the first subset of pixels, determining an optical density, wherein the pixel-specific color vector for the pixel identifies a degree of expression of the particular stain corresponding to the optical density; determining that each pixel of a second subset of the set of pixels is mapped to a point outside the portion of the color space; and performing an unmixing technique to predict, for each pixel of the second subset and for each of some of the at least four digital pathology stains, a degree of expression of the stain in the portion of the biopsy slide depicted at the pixel, wherein the some of the at least four digital pathology stains do not include the particular stain, and wherein the unmixing technique uses the color vectors determined to represent each of the some of the at least four digital pathology stains; and generating one or more synthetic single-stain images using the pixel-specific color vectors.
97. The computer program product of claim 96, wherein the particular stain is selected based on information about which portions of a cell each of the at least four digital pathology stains is configured to stain.
98. The computer program product of claim 96, wherein the portion of the color space comprises a wedge-shaped region, a combination of primitives, or a portion of space defined based on an inequality with respect to an x-coordinate and an inequality with respect to a y-coordinate.
99. The computer program product of claim 96, wherein performing the unmixing technique comprises using non-negative matrix factorization (NMF).
100. The computer program product of claim 96, wherein the color vector is determined based on one or more user inputs received using one or more color-vector adjustment tools available within the interface.