Expedited deep learning-based interactive ground-truth collection for cell detection / classification model development in multiplex brightfield

The method of generating synthetic singleplex images from multiplex images using color unmixing and machine-learning models addresses the challenge of multiplex brightfield immunohistochemistry annotation inefficiencies, enhancing diagnostic accuracy and scalability.

WO2025226430A1PCT designated stage Publication Date: 2025-10-30VENTANA MEDICAL SYSTEMS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/023298
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2025-04-04
Publication Date
2025-10-30

Smart Images

  • Figure US2025023298_30102025_PF_FP_ABST
    Figure US2025023298_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to techniques for generating annotated singleplex images from multiplex images, aimed at enhancing the efficiency of biomarker annotation in digital pathology. This process involves decomposing a multiplex image into synthetic singleplex images, each representing a specific biomarker or counterstain used during staining. A machine-learning model is then employed to generate initial annotated images for each synthetic singleplex image, which enable a user (e.g., pathologist) to review and quickly correct misidentified cells or modify biomarker expression statuses. Additionally, the initial annotations can be automatically validated using statistical techniques that involve overlaying a binary mask of one synthetic singleplex image onto a pure stain mask of another, producing a ground-truth annotated image. This overlay serves as a reference for cross-referencing the initial annotations, enabling automatic correction and refining of the image for accurate biomarker expression identification.
Need to check novelty before this filing date? Find Prior Art

Description

EXPEDITED DEEP LEARNING-BASED INTERACTIVE GROUND-TRUTH COLLECTION FOR CELL DETECTION / CLASSIFICATION MODEL DEVELOPMENT IN MULTIPLEX BRIGHTFIELDCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and the priority to U.S. Provisional Patent Application Number 63 / 637,262, filed on April 22, 2024, which is hereby incorporated by reference in its entirety for all purposes.BACKGROUND

[0002] Digital histopathology involves examination of diseased tissue and its characteristics by performing a microscopic analysis of stained tissue sections in a digital format. Histological staining, a method enhancing contrast and visibility of various tissue cells, may employ specific dyes, chromogens or markers applied to these sections. This technique may effectively prove invaluable for diagnosing and prognosing a range of diseases, notably cancer. In digital pathology, visualizing multiple disease biomarkers simultaneously improves diagnostic accuracy and tissue processing efficiency. However, distinguishing between the diverse dyes, stains, and markers that are used to identify these biomarkers may present a significant challenge. Multiplex brightfield immunohistochemistry (IHC) enables the detection of multiple biomarkers within a single tissue section, but the overlapping signals, variations in staining intensity, and potential artifacts make it difficult to disambiguate individual markers. Even expert pathologists may struggle to reliably annotate, interpret and score these images, as distinguishing between coexpressed biomarkers at a single-cell level requires careful analysis. Misinterpretation or inconsistencies in scoring can lead to diagnostic variability and inaccuracy.

[0003] Additionally, the process of scoring or annotating multiplex images is slow and labor-intensive. Pathologists typically review multiplex IHC images (e.g., singleplex images unmixed from multiplex images) to analyze individual cell locations, often manually clicking on single cells to assess biomarker expression. This process is tedious and exhaustive, significantly delaying diagnostics and research. Furthermore, pathologists are a limited and expensive resource, and their time is inefficiently consumed by such repetitive manual tasks. Theseinefficiencies not only increase costs but also limit the scalability of multiplex IHC in clinical and research applications.SUMMARY

[0004] Certain aspects and features of the present disclosure relate to techniques for expediting, facilitating and increasing efficiency with which a user can annotate presence of biomarkers in a digital pathology space. The disclosed techniques may include accessing a multiplex digital pathology image, hereinafter as multiplex image, that depicts a biological specimen labeled with multiple different markers. These multiple different markers (e.g., dyes or chromogens) are selected to identify (or stain) different biomarkers and / or cellular components relative to each other. The multiplex image received may be processed using various color unmixing techniques (e.g., spectral deconvolution or machine-learning based generative models) to separate individual biomarkers. These unmixing techniques are configured to generate multiple synthetic singleplex images from the multiplex image, where a synthetic singleplex image of the multiple synthetic singleplex images may correspond to a marker that was used to stain the multiplex image.

[0005] The multiple synthetic singleplex images are devoid of signals from other markers and may thus be scored more easily by either additional computer algorithms or by a pathologist. In some examples, a graphical user interface may be generated including the multiple synthetic singleplex images, which may enable a user (e.g., a pathologist) to annotate these images. Thus, creating a dataset of ground-truth images representing manually annotated singleplex images that may be used to train a machine-learning model for assisting the pathologists in expediting annotation process.

[0006] In some aspects, techniques may be employed to detect cells of a specific type, a particular biomarker, or a cellular component within the multiple synthetic singleplex images. For example, a machine-learning model (or a classification model) may be configured to predict a biomarker status indicating whether a given cell or the cellular component expresses the particular biomarker (e.g., positive or negative expression) within the synthetic singleplex image. Additionally, the classification model may assess an intensity of that expression (e.g., with an intensity score), indicating whether the positivity is strong, medium or weak. The classificationmodel may be trained using the dataset of manually annotated singleplex images, where expert pathologists provide ground-truth images that are labeled for expression of the particular biomarker. Once trained, the classification model may evaluate staining patterns, intensity distributions, and spatial relationships for multiple synthetic singleplex images to generate corresponding initially annotated images associated with the multiple different biomarker. An initial annotated singleplex image may include a set of predicted annotations that corresponds to a set of predicted spatial boundaries (e.g., cell membrane) or positions (e.g., cell nuclei) within the synthetic singleplex image. The set of predicted annotations may identify a depiction of a biological component that is not a tumor cell (e.g., immune cells, stromal cells). Additionally, the set of predicted annotations may include a presence (e.g., ER+, PR+), a relative level of presence (e.g., 1+, 2+ or 3+ for HER2), or an absence (e.g., ER-, PR-) of the particular biomarker associated with the given cell or the cellular component.

[0007] The initially annotated images associated with each biomarker may be further validated by a pathologist to manually correct misclassified cells via the GUI. Alternatively, the multiple initially annotated images may be validated dynamically to perform automatic corrections. To this end, a binary mask that represents a set of first predicted spatial boundaries or positions may be identified within a first synthetic singleplex image of the multiple synthetic singleplex images. Each of the set of first predicted spatial boundaries or positions are identified so as to predict a boundary or position of a first biomarker or the particular cellular component corresponding to the first synthetic singleplex image. Further, a pure stain mask (or image) corresponding to a second biomarker may be generated from a second synthetic singleplex image of the multiple synthetic singleplex images. The pure stain mask comprises a set of second predicted spatial boundaries or positions within the second synthetic singleplex image. Each of the set of second predicted spatial boundaries or positions are identified so as to predict a boundary or position of the second particular biomarker or another particular cellular component.

[0008] Based on the binary mask and the pure stain image, an overlay may be generated that serves as a reference for validating the initially annotated image based on a set of predefined designation rules. In some aspects, the overlay may represent a ground-truth annotated image identifying the depiction of the biological component that is not a tumor cell (e.g., immune cells, stromal cells), the presence (e.g., ER+, PR+), the relative level of presence (e.g., 1+, 2+ or 3+ for HER2), or the absence (e.g., ER-, PR-) of the particular biomarker associated with the particularcell or the cellular component. Accordingly, the graphical user interface (GUI) may be generated or updated to show the overlay and the multiple initial annotated images. The generated or updated GUI may be output or displayed to a user device.

[0009] In some examples, it may be identified that at least one predicted annotation of the set of predicted annotations within the initially annotated singleplex image does not correspond to the particular biomarker, the particular cell or the cellular component. The identification may involve comparing the overlay, including a set of third predicted spatial boundaries or positions, with the set of predicted annotations based on the set of predefined designation rules. Upon identifying a misclassified annotation as a result of the validation or comparison, corrections may be performed to the initially annotated image. For example, the corrections may be performed manually by a user (e.g., an expert pathologist) via the GUI that includes a set of tools enabling the user to adjust the corresponding biomarker status, as well as to rectify whether these cells or cell clusters are to be assigned a given label (e.g., tumor-positive or tumor-negative).Alternatively, or additionally, based on the overlay and the set of predefined designation rules, an autocorrected image may be generated automatically that includes a change or a corrected annotation corresponding to the at least one predicted annotation. For the autocorrected image, the GUI may be updated to depict the annotated image before and after the automatic correction for a comparative analysis and further verification by the user.

[0010] In some aspects, techniques may be used to convert or translate staining modalities between singleplex images to facilitate analyses that are dependent on specific modalities. For example, a singleplex HER2 image may be converted to a synthetic DAB (3,3'- Diaminobenzidine) image. The cross-modality translation may be performed by applying various techniques such as image processing (e.g., contrast enhancement, intensity normalization, and histogram matching) or Al-based techniques such as generative adversarial networks (GANs) and / or convolutional neural networks (CNNs). For instance, a generative model may be employed to translate a given synthetic singleplex image corresponding to one staining modality to a translated singleplex image corresponding to different staining modality than the multiple different staining modalities of the multiplex image. The generative model may be trained in a cyclic adversarial manner such that relevant features of the original image are preserved while adapting it to the targeted staining modality.

[0011] The GUI may enable the user to view any given image, such as the generated binary mask, translated singleplex image, initial annotated image, autocorrected image, overlay, or any synthetic singleplex image derived from the multiplex image, to assist in their designations or annotations. Additionally, the GUI may be configured to select individual cells, groups of cells, or specific regions within the given image using various selection techniques, including freehand (lasso tool), shape-based selections (e.g., circles, rectangles), or region-based selections (e.g., drawing around a specific area of interest). For example, a communication may be received from the GUI of the user device that defines a user-identified boundary within an initial annotated image by using a selection technique. The communication may indicate that the at least one predicted annotation within the initial annotated singleplex image or the user-defined boundary does not correspond to the particular biomarker, the particular cell or the cellular component.

[0012] Subsequently, the at least one predicted annotation may be detected that correspond to one or more predicted spatial boundaries or positions of the set of predicted spatial boundaries or positions that are within the user-identified boundary. Once the predicted annotations are detected, a number of operations may be applied to refine or manipulate the predicted spatial boundaries. For example, a deletion operation may be performed to remove unwanted predicted boundaries from the set, eliminating incorrect or irrelevant predictions. Alternatively, modification or editing may be performed by changing the at least one predicted annotation that corresponds to a boundary or a position of the particular biomarker or the particular cellular component within the initial annotated image.

[0013] These tools may provide flexibility and precision for focusing on areas of interest, whether adjusting individual cell annotations or modifying larger, contiguous regions of the image. In some examples, the GUI may also provide advanced image enhancement features to adjust brightness, contrast, and color settings of the given image for improved visibility and detail. Additionally, the user may modify the color codes or intensity of multiple predicted annotations, enabling finer control over the display of different biological components or markers.

[0014] Some aspects of the present disclosure are generalized to various forms of imaging that are commonly used in digital pathology. For example, the multiplex image may be collected by e.g., brightfield immunohistochemistry (IHC) imaging, fluorescence IHC (HHC), multispectral imaging (MSI), hyperspectral imaging (HSI), or any other technique where themultiple different markers (e.g., dyes or chromogens) are selected to identify (or stain) different biomarkers and / or cellular components relative to each other.

[0015] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.

[0016] In some embodiments, a computer-program product tangibly embodied in a non- transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods or processes disclosed herein.

[0017] In some embodiments, a system is provided that includes one or more means to perform part or all of one or more methods or processes disclosed herein.

[0018] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. The present disclosure is described in conjunction with the appended figures:

[0020] FIG. 1 shows an exemplary network of a digital pathology image generation system in accordance with some aspects of the present disclosure.

[0021] FIG. 2A shows an exemplary network illustrating color unmixing of a multiplex image in accordance with some aspects of the present disclosure.

[0022] FIG. 2B shows an illustrative example of the color unmixing of a triplex image in accordance with some aspects of the present disclosure.

[0023] FIG. 3A depicts an illustrative example of a classification model for detecting cells of a specific type or a particular biomarker and determining corresponding biomarker status based on a singleplex image, in accordance with some aspects of the present disclosure.

[0024] FIG. 3B shows an illustrative example of a systematic decision pathway for HER2 (human epidermal growth factor receptor 2) intensity quantification using established guidelines provided by College of American Pathologists (CAP) based on tumor cell membrane staining characteristics.

[0025] FIG. 4 shows an example illustration of a validation framework to perform validation and autocorrection of an initial annotated singleplex image, in accordance with some aspects of the present disclosure.

[0026] FIG. 5 illustrates an exemplary workflow of a cross-modality translation model (CMTM) for translation between various staining modalities, in accordance with some aspects of the present disclosure.

[0027] FIG. 6 shows an illustrative example of a synthetic singleplex image of green HER2- Low that is generated by processing a triplex image of ER-PR-HER2-Low via the unmixing model.

[0028] FIG. 7 shows an illustrative example of a graphical user interface (GUI) including a singleplex estrogen receptor (ER) image and a singleplex progesterone receptor (PR) image, along with their corresponding masks.

[0029] FIG. 8 shows an example of how masks may be used to validate accuracy of the initial annotated singleplex image, in accordance with some aspects of the present disclosure.

[0030] FIG. 9A shows an illustrative example of the graphical user interface (GUI) enabling a flexible selection of regions of interest, in accordance with some aspects of the present disclosure.

[0031] FIG. 9B shows an example of the GUI enabling a shape-based selection of the regions of interest, in accordance with some aspects of the present disclosure.

[0032] FIG. 9C illustrates the GUI enabling the shape-based selection for an annotated singleplex ER image, in accordance with some aspects of the present disclosure.

[0033] FIG. 9D shows an illustrative example of the GUI including an annotated singleplex PR image and corresponding zoomed regions of interest highlighting selection.

[0034] FIG. 10 shows an illustrative example of validating accuracy of one or more initial annotated singleplex images, in accordance with some aspects of the present disclosure.

[0035] FIG. 11 shows an illustrative example of three initial annotated images.

[0036] FIG. 12 shows an illustrative example of an automatic correction process applied to an initially annotated ER image.

[0037] FIG. 13 illustrates an exemplary workflow to generate annotated singleplex images based on a multiplex image in accordance with some aspects of the present disclosure.DETAILED DESCRIPTION

[0038] In the present disclosure, the term “biomarker” may be defined as a characteristic of tissue such as a presence of a particular cell type e.g., immune cells, particularly the one indicative of a medical condition. The identification of the biomarker may involve the presence of a particular molecule, for example a protein within the tissue feature.

[0039] The term ’’marker” may be understood in the sense of a stain, dye, or a tag that facilitates the differentiation of a biomarker from surrounding tissue or other biomarkers. The tag, which may be an antibody - specifically the one with an affinity for a protein associated with a particular biomarker - can be used for staining or dyeing. A marker may exhibit an affinity for a particular biomarker, e.g., to a particular molecule / protein / cell structure / cell (indicative of a particular biomarker). The biomarker to which a marker has an affinity may be specific / unique for the respective marker. A marker may mark tissue, i.e., a biomarker in the tissue, with a color. The color of tissue marked by a respective marker may be specific / unique to the respective marker.

[0040] The term “sample” may be understood as cellular material derived from a biological organism, comprising but not limited to hair, skin samples, tissue samples, cultured cells, cultured cell media, and biological fluids. The term “tissue” refers to a mass of interconnected cells (e.g., central nervous system (CNS) tissue, neural tissue, or eye tissue) derived from a human or other animal and includes the connecting material and the liquid material in association with the cells. In the context of histopathology, the term “slide” refers to a glass microscope slidecarrying a thin section of tissue that has been stained for microscopic examination. The term “sample” also includes media containing isolated cells. One skilled in the art may determine the quantity of samples required to obtain a reaction by standard laboratory techniques.

[0041] In multiplex immunohistochemistry (IHC), a digital pathology image may be classified as singleplex, duplex, or triplex, depending on the number of different markers or stains used. This approach enables a combined analysis by simultaneously detecting multiple targets, providing insights into the molecular composition of the tissue sample. Using multiple markers in one experiment saves time compared to performing separate individual stains and preserves valuable tissue samples by reducing material usage. Moreover, multiplex staining enables more data collection in a single sample, enhancing the understanding of complex biological processes. However, multiplex images are typically difficult for a human / pathologist to reliably score or annotate. Pathologists may need to use synthetic IHC images unmixed from multiplex images to review the cell locations. Moreover, clicking the single cells in each synthetic image is an expensive and exhaustive, slow procedure, which may delay the entire development process.

[0042] The present disclosure relates to techniques for generating annotated singleplex images based on a multiplex image, which may facilitate and increase efficiency with which a user (e.g., a pathologist) can annotate the presence of biomarkers in a digital pathology space. The generation of annotated singleplex images may involve decomposing the multiplex image into one or more synthetic singleplex images. Each synthetic singleplex image correspond to a single marker or antibody to visualize a specific target, often a protein or biomarker, resulting in an image with two colors: one for the general tissue stain (e.g., Hematoxylin, which appears blue or purple) and another for the target molecule (e.g., a protein, antigen, or biomarker) visualized with a second marker such as Tamra (a fluorophore). Similarly, duplex staining involves two distinct markers, producing an image with three colors, while triplex staining involves three distinct markers, producing an image with four colors: one for the general tissue stain and others for the separate markers.

[0043] The multiplex image, including one or more biomarkers, may be decomposed by leveraging various color unmixing techniques that separate individual biomarkers. For example, for color decomposition, linear unmixing techniques (e.g., spectral deconvolution or matrix inversion methods) or non-linear unmixing techniques (e.g., deep learning-based generativemodels such as convolutional autoencoders or generative adversarial networks) may be employed. These synthetic singleplex images may enable the deployment of disclosed classification and validation techniques to efficiently identify cells or cellular components expressing patterns that are associated with a particular biomarker within the synthetic singleplex image.

[0044] In some examples, the synthetic singleplex images may be presented in a graphical user interface (GUI) configured to enable a user (e.g., a pathologist) to annotate or label a biomarker status of individual cells or groups of cells, as well as to rectify whether these cells or cell clusters are to be assigned a given label (e.g., tumor-positive or tumor-negative). One application of these annotated images is to use them as ground-truth images for training a classification model, enabling automated generation of annotated images. This approach may significantly reduce the time and effort required by pathologists, eliminating a need to manually annotate each individual cell.

[0045] In some aspects of the present disclosure, a synthetic singleplex image of the one or more synthetic singleplex images may be processed by a variety of techniques (e.g., statistical or machine-learning techniques) for detecting cells of a specific type, a particular biomarker, or a cellular component within the synthetic singleplex image. For example, the classification model may be configured to predict a biomarker status indicating whether a given cell or the cellular component expresses the particular biomarker (e.g., positive or negative expression) within the synthetic singleplex image. Additionally, the classification model may assess an intensity of that expression (e.g., with an intensity score), indicating whether the positivity is strong, medium or weak. By leveraging this classification model, an initial annotated image may be generated based on the synthetic singleplex image including a set of predicted annotations that corresponds to a set of predicted spatial boundaries (e.g., cytoplasm) or positions (e.g., nuclei) labeled as e.g., tumor-negative, non-tumor, or tumor-positive (0, 1+, 2+, or 3+), where 3+ represents the highest expression and 1+ the lowest.

[0046] As an illustrative example, three synthetic singleplex images — corresponding to estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2) — may first be derived from an ER-PR-HER2 triplex image using a color unmixing model. The classification technique may then predict whether each of one or more cells depicted in each synthetic singleplex image are positive for each or any of one or more biomarkers (e.g.,ER, PR, or their HER2 designation). HER2 designation may follow the College of American Pathologists (CAP) designations of 0 and 1, 2+, or 3+ to indicate the relative level of HER2 protein abundance, such that a prediction may include a particular one of multiple designations representing protein abundance. The HER2 designation need not be based only on the intensity of the HER2 labeling, but it can also be based on the stain intensity within the cell membrane. The completeness of labeling within a cell membrane can inform a HER2 designation. The predicted biomarker status of cells within a sample in singleplex images (or the synthetic singleplex images) may be used by a pathologist to aid in diagnosis or to make recommendations regarding disease management.

[0047] The classification model may be trained using the dataset of the manually annotated singleplex images, where expert pathologists provide ground-truth labels for biomarker expressions. This training process typically involves supervised learning, enabling the model to associate staining patterns, intensity distributions, and spatial relationships with the appropriate biomarker classifications. To achieve this, deep learning architectures such as convolutional neural networks (CNNs) or transformer-based models may be employed, along with traditional machine-learning techniques such as support vector machines (SVMs) or decision trees. Additionally, data augmentation techniques such as rotation, scaling, and color jittering can enhance the robustness of the model by helping it adapt to variations in staining intensity, imaging conditions, and tissue morphology. These strategies also serve to expand the size of the training dataset, which is particularly beneficial when the available dataset of manually annotated singleplex images is limited in size. In some implementations, a unified classification model may be trained to handle multiple biomarkers (e.g., ER, PR, and HER2) within the multiplex image using multi-task learning, shared feature extraction, and separate classification heads for each biomarker. Alternatively, multiple independent models may be used, each specializing in the classification of a single biomarker.

[0048] The graphical user interface (GUI) may be generated or updated including the initial annotated singleplex images to a user device, where the user may perform manual corrections to the predicted annotations. For example, the GUI may enable the user to correct binary choices as to whether a cell expresses a specific biomarker (e.g., ER+ or ER-). GUI may further assist in making more nuanced or graded changes in regards to a specific marker. For instance, the breast cancer biomarker HER2 is often designated as 0 and +1, 2+, or 3+ to indicate the relative level ofHER2 protein abundance around the cell membrane, as mentioned above. The GUI allows the user to change such non-binary designations as well. The user may first select a cell or group of cells and subsequently select a grading or label to change their labeling. The GUI may provide tools for the user to rapidly make corrections in misidentified cells and to change the status of biomarker expression. The results of these tools may then be used to inform physicians with regards to disease diagnosis, prognosis, and management. The proofread and corrected images may also serve as input to next-generation machine-learning algorithms to develop future tools with even greater accuracy. As pathology slices often include hundreds and even up to thousands of cells, it can be appreciated how automated tools and powerful user interfaces, such as disclosed, can dramatically improve the accuracy of quantifying digital pathology and improve patient health outcomes.

[0049] To automate the validation and correction process, the synthetic singleplex images may optionally be preprocessed (e.g., by applying contrast enhancement, normalization, and noise reduction) to improve image quality and reduce artifacts from staining. These enhancements may contribute to making the target signals more distinct, which in turn may result in accurate biomarker identification. Following preprocessing, a binary mask corresponding to a target biomarker (e.g., HER2, ER, or PR) is generated by performing cell segmentation, which separates cells from background pixels. Various segmentation techniques, such as thresholding, region growing, or deep learning-based methods such as UNet or VNet, may be used to delineate cell boundaries. The result of the segmentation is a binary mask that may represent a set of first predicted spatial boundaries (e.g., cell membrane) within a synthetic singleplex image. Each of the set of first predicted spatial boundaries are identified so as to predict a boundary of a particular biomarker or a particular cellular component corresponding to the synthetic singleplex image.

[0050] After segmentation, the centroids of the identified regions are computed, serving as the geometric centers of each cell. These centroids provide precise coordinates for the cell nuclei, which are used as references for further analysis. Therefore, the binary mask may further include a set of first predicted positions (e.g., cell nuclei) within the synthetic singleplex image. Once the cell nuclei are identified from the synthetic singleplex image (e.g., HER2 image), their positions are cross-referenced with generated pure stain images from other markers (e.g., either separate or combined ER, PR masks). To generate pure stain images (or masks), overlapping signals (e.g.,hematoxylin (blue) and Tamra (pink) or Dabsyl (yellow)) may be separated to isolate the individual contributions of each marker. Techniques such as linear color unmixing, Gaussian mixture models (GMM), or clustering methods may be used to achieve this separation. A pure staining mask typically defines regions or areas of interest within an image where staining (or biomarker expression) is present, usually indicating the boundaries of cells or tissues based on the staining. This pure stain image may then be overlaid with the binary mask, generating an overlay, to localize cell regions and analyze biomarker expression.

[0051] The overlay provides a basis for validating the accuracy of the initial annotations based on a set of predefined designation rules that determine whether a given cellular component e.g., nucleus exhibits the expected biomarker expression. For example, if a cell is annotated as positive for a biomarker in the initial annotated image but lacks a corresponding signal in the pure biomarker mask, the annotation is corrected to negative. Similarly, in the overlay, presence of cell nuclei within pure stain masks may suggest tumor cells, aiding in accurate classification. Co-localization of multiple biomarkers is also assessed for correct classification. By iteratively comparing the initially annotated images with the corresponding overlay, the system can automatically refine or autocorrect the biomarker classifications. This autocorrection process improves the reliability, efficiency and accuracy of the annotations. The disclosed validation framework may be appended into an overall machine-learning pipeline, including the unmixing model and the classification model, thus enabling the generation of autocorrected singleplex images for more robust and scalable digital pathology applications.

[0052] It will be appreciated that there are several manners in which an overlap between a binary mask and the pure stain image may be calculated and that certain approaches may introduce small biases towards type 1 or type 2 errors. One example of how the mask may be used comprises superimposing the cell nuclei from the binary mask of the particular biomarker (e.g., HER2) onto pure stain image of any other synthetic singleplex image (e.g., ER or PR). In this example, a window comprising a 5><5 pixel region around each identified nucleus may be considered. This mask may be used to define a neighborhood around the seed (or nucleus) upon which certain designation rules may be applied. For example, if the seed’s label falls inside the PR (Tamra) or ER (Dabsyl) staining regions, then the label stays the same, retaining the original label of the initial annotated image. Similarly, when the corresponding location in the Tamra / Dabsyl masks indicates PR+ / ER+, any estimations within the 5x5 neighborhood that arewrongly labeled will be automatically corrected. It can be appreciated that the size of the window may be adjusted to compensate for imaging resolution, magnification, or cell size. Additionally, this process may be applied to multiple singleplex images highlighting different biomarkers.

[0053] In some embodiments, techniques may be used to convert modalities between singleplex images to facilitate analyses that are dependent on specific modalities. For example, a singleplex HER2 image may be converted to a synthetic DAB (3,3'-Diaminobenzidine) image prior to classification or validation process. The brown color produced by the DAB reaction is easily visible against a counterstained background (often hematoxylin), which stains cell nuclei blue. This contrast enables the clear visualization of the presence, location, and quantity of the antigen, facilitating the diagnosis and research into various diseases. Additionally, various algorithms (image processing, machine vision, or Al-based algorithms) have been developed to analyze DAB-stained images in pathology. The conversion from any synthetic singleplex modality image to a synthetic DAB image will allow the use of such tools. It will be appreciated that this feature is not limited to conversion of images synthetic DAB image but rather has the flexibility to convert between multiple singleplex image modalities. The results of algorithms and tools used to analyze such synthetic DAB images may then be easily overlayed back onto any synthetic singleplex image or the original multiplex image.

[0054] The GUI provides flexibility by offering a variety of selection tools that enable users to precisely define regions of interest within an annotated image. The user may choose from several selection methods, including a lasso tool, which enables freehand drawing around irregular areas, or shape-based tools such as circles or rectangles for more structured selections. Additionally, users may select individual cells or clusters of cells for performing various operations e.g., deleting, modifying, or editing predicted annotations, making it easier to correct classifications or adjust labels across different regions. This versatility enables the users to adjust or fine-tune the analysis based on their specific needs, whether it's selecting a small group of cells or a larger, more complex area.

[0055] The GUI also includes image enhancement functionalities, enabling users to adjust the brightness, contrast, and even change the color of specific annotations to improve visibility and clarity. Furthermore, the user may perform modifications and adjustments in correcting misclassified cells — such as immune cells wrongly categorized as tumor cells — and making corrections to classifications. For instance, pathologists may select a cell or a collection of cellsvia a selection method (e.g., lasso, shape-based or single selection) to correct a binary classification (e.g., HER2 -negative to HER2-positive) or adjust graded expressions (e.g., HER2 1+ to 3+). The GUI may further provide a comparative analysis mode that allows users to review the original and corrected annotations side by side so that the refinements improve diagnostic accuracy. The flexibility of this GUI makes it an invaluable tool for fine-tuning automated cell classifications or annotations, thereby enhancing the overall reliability of digital pathology analysis.

[0056] FIG. 1 shows an exemplary network of a digital pathology image generation system in accordance with some aspects of the present disclosure. The digital pathology images may be generated by an image generation system 102. A fixation / embedding system 104 fixes and / or embeds a tissue sample (e.g., a liquid fixing agent, such as formaldehyde solution) and / or an embedding substance (e.g., a historical wax, such as paraffin wax and / or one or more resins, such as styrene or polyethylene). Each slice may be fixed by exposing the slice to a fixating agent for a predefined period of time (e.g., at least 3 hours) and by then dehydrating the slice (e.g., via exposure to an ethanol solution and / or a clearing intermediate agent). The embedding substance can infiltrate the slice when it is in liquid state (e.g., when heated). A tissue slicer 106 then slices the fixed and / or embedded tissue sample (e.g., a sample of a tumor) to obtain a series of sections, with each section having a thickness of, for example, 4-5 microns. Such sectioning can be performed by first chilling the sample and then slicing the sample in a warm water bath. The tissue can be sliced using (for example) a vibratome or compresstome.

[0057] Because the tissue sections and the cells within them are virtually transparent, preparation of the slides typically includes staining (e.g., automatically staining) the tissue sections to render relevant structures more visible. In some instances, the staining is performed manually. In some instances, the staining is performed semi-automatically or automatically using a staining system 108. The staining can include exposing an individual section of the tissue to one or more different stains (e.g., consecutively, or concurrently) to express different characteristics of the tissue. For example, each section may be exposed to a predefined volume of a staining agent for a predefined period of time. The staining agent can include (for example) an RNA probe, protein probe (e.g., nuclear-protein probe or cytoplasm-protein probe), an immunohistochemistry stain, a probe for a secreted substance, etc. In some instances, the staining agent is one that stains for KAPPA mRNA or LAMBDA mRNA.

[0058] One exemplary type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes) to stain tissue structures. Histochemical staining may be used to indicate general aspects of tissue morphology and / or cell microanatomy (e.g., to distinguish cell nuclei from cytoplasm, to indicate lipid droplets, etc.). One example of a histochemical stain is hematoxylin and eosin (H&E). Other examples of histochemical stains include trichrome stains (e.g., Masson's Trichrome), Periodic Acid-Schiff (PAS), silver stains, and iron stains. The molecular weight of a histochemical staining reagent (e.g., dye) is typically about 500 kilodaltons (kD) or less, although some histochemical staining reagents (e.g., Alcian Blue, phosphomolybdic acid (PMA)) may have molecular weights of up to two or three thousand kD. One case of a high-molecular-weight histochemical staining reagent is alpha-amylase (about 55 kD), which may be used to indicate glycogen.

[0059] Another type of tissue staining is immunohistochemistry (IHC, also called "immunostaining"), which uses a primary antibody that binds specifically to the target antigen of interest (biomarker). IHC may be direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or fluorophore). In indirect IHC, the primary antibody is first bound to the target antigen, and then a secondary antibody that is conjugated with a label (e.g., a chromophore or fluorophore) is bound to the primary antibody. The molecular weights of IHC reagents are much higher than those of histochemical staining reagents, as the antibodies have molecular weights of about 150 kD or more. The sections may then be individually mounted on corresponding slides, which an imaging system 110 can then scan to generate raw multiplex digital-pathology images 112a-n. Each section may be mounted on a slide, which is then scanned to create a digital image that may be subsequently examined by digital pathology image analysis and / or interpreted by a human pathologist (e.g., using image viewer software). The imaging may include capturing bright-field image of the slide section.

[0060] In some instances, a pathologist may review and manually annotate the digital image of the slides (e.g., tumor area, necrosis, etc.). In some instances, annotations of regions of interest are performed automatically using the disclosed techniques. Some of the digital-pathology images 1 12a-n may be used by an unmixing model.

[0061] A digital histopathology image (e.g., 112a) typically includes an array, usually a rectangular matrix, of pixels. Each ‘pixel’ is one picture element and is a digital quantity that is a value that represents some property of the image at a location in the array corresponding to aparticular location in the image. Typically, in continuous tone black and white images the pixel values represent a gray scale value. Pixel values for a digital image typically conform to a specified range. For example, each array element may be one byte (i.e., eight bits) representing pixel values in the range of 0 to 255. In a gray scale image, a 255 may represent absolute white and Zero total black (or visa-versa). Color images consist of three-color planes, generally corresponding to red, green, and blue (RGB). For a particular pixel, there is one value for each of these color planes, (i.e., a value representing the red component, a value representing the green component, and a value representing the blue component). By varying the intensity of these three components, all colors in the color spectrum typically may be created.

[0062] FIG. 2A shows an exemplary network 200-A illustrating color unmixing of a multiplex image in accordance with some aspects of the present disclosure. Color unmixing for multiplexed stain images can be both linear and non-linear, depending on staining characteristics, stain specificity, and the unmixing technique used. Linear unmixing techniques assume that the spectral response of each stain is independent and additive. Techniques such as least squares regression, non-negative matrix factorization (NNMF), independent component analysis (ICA), and singular value decomposition (SVD) may be employed to separate individual stain contributions based on known spectral profiles. In contrast, non-linear unmixing techniques are employed when staining involves complex interactions between chromogens or fluorophores, as seen in multiplex IHC images. Such interactions in multiplex staining may arise due to factors such as spectral overlap of different chromogens, complex staining patterns, spatial heterogeneity or non-linear relationship between staining intensity and target concentrations. To address these challenges, non-linear unmixing techniques such as deep learning-based models (e.g., convolutional neural networks (CNNs), generative adversarial networks (GANs), autoencoders, and artificial neural networks (ANNs)) may be used. These methods may model intricate dependencies between stains and improve accuracy in separating biomarker signals.

[0063] An unmixing model 202, including one or more trained models (e.g., deep learning models), may receive the multiplex image from the image generation system 102 and generate multiple synthetic singleplex images e.g., 212a-c. For example, the unmixing model 202 may include multiple models that may be trained and / or utilized to transform a multiplex image into distinct types of singleplex images. In this approach, each model of the unmixing model 202 may be configured to generate a specific type of singleplex image corresponding to a particular targetof interest. For instance, for a triplex image 112, three synthetic singleplex images (e.g., 212a-c) may be generated by a first model of the unmixing model 202 trained to identify and extract a first biomarker from the multiplex image, a second model of the unmixing model 202 to identify and extract a second biomarker, and a third model of the unmixing model to identify and extract a third biomarker. Each model may be configured to process the multiplex image and isolate the specific biomarker of interest, thereby generating separate singleplex images corresponding to each biomarker.

[0064] Alternatively, the unmixing model 202 may include a single unified model that may be trained and / or used to transform a single multiplex image into multiple synthetic singleplex images. In this implementation, the unmixing model 202 may be configured to simultaneously generate multiple singleplex images e.g., 212a-c from a single multiplex image e.g., 112. For example, the model may extract and generate separate singleplex images for multiple biomarkers (e.g., first biomarker, second biomarker, and third biomarker) or different tissue components (e.g., nuclear staining, membrane staining, cytoplasm staining) in a single operation. Each synthetic singleplex image is devoid of signals from other markers and may thus be scored or annotated more easily by either additional computer algorithms or by a pathologist. This approach may be more efficient than the current standard of staining sequential slices of a sample with different markers, as multiplex images do not require registration between images and thus simplifies co-localization analyses.

[0065] The architecture of the unmixing model 202 may leverage deep learning models such as pretrained convolutional networks e.g., UNet, VNet, generative adversarial networks (GANs), or other advanced unmixing methods. For example, the unmixing model 202 may utilize an autoencoder architecture comprising two primary components: an encoder 204 and a decoder 206. The encoder 204 is responsible for reducing the dimensionality of the input multiplex image, transforming the complex input data into a more compact and meaningful representation (or an embedding 210). The encoder 204 may include multiple layers, such as convolutional layers, max-pooling layers, and other down-sampling operations. These layers are designed to progressively extract hierarchical features, enabling the model to learn local patterns from the input multiplex image. Additionally, or alternatively, self-attention mechanisms may be used in encoding layers to help the model focus on significant spatial relationships in the multiplex image. While convolution layers learn local patterns through sliding filters, self-attention layers enable the model to focus on long-range dependencies by considering the relationship between all pixels in an image, regardless of their spatial proximity. In a hybrid model, convolution layers and self-attention mechanisms may complement each other, with convolutions handling local feature extraction and self-attention capturing global contextual information.

[0066] The decoder 206 of the unmixing model 202 takes the embedding 210 generated by the encoder 204 and attempts to reconstruct the original input in the form of separate singleplex images e.g., 212a-c. The decoder 206 typically involves transposed convolutional layers (also called deconvolutional layers) that perform up-sampling, reversing the dimensionality reduction done by the encoder. These layers progressively increase the spatial resolution of the feature maps, enabling the model to generate a high-resolution output. The decoder 206 also often incorporates skip connections 208 from the encoder 204, enabling the model to retain finegrained details from the input image that contribute to accurate reconstruction. In more advanced implementations, self-attention mechanisms may also be integrated into the decoder 206 to help improve the generation of higher-quality reconstructions by focusing on relevant image regions. The combination of both the encoder 204 and decoder 206 may enable the unmixing model 202 to separate multiplexed data into individual, unmixed signals, such as biomarker or tissue component images, while maintaining the overall structure and fine details of the original input.

[0067] FIG. 2B shows an illustrative example of unmixing of a triplex image 214 in accordance with some aspects of the present disclosure. The term "triplex" refers to a combined image that simultaneously captures the expression of three distinct biomarkers or molecules of interest, which in this case are the estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor Receptor 2 (HER2). Estrogen is a hormone that can be a contributing factor, particularly in breast and endometrial cancer. Estrogen binds to an estrogen receptor (ER) triggering a series of cellular responses that involve proliferation and differentiation of the specific cells. Estrogen receptors (ER) and progesterone receptors (PR) are the biomarkers used in cancer pathology to assess the presence of receptors for estrogen and progesterone in tumor cells. ER and PR represent nuclear receptors primarily located within the nucleus of a cancer cell. The staining patterns of ER and PR may help identify the subcellular localization of these biomarkers. For ER, a commonly used antibody is ER-a. The stain isusually visualized with a chromogen e.g., DAB. Progesterone staining may involve the use of PR antibodies, and the resulting stain may also be visualized with DAB.

[0068] In the triplex image 214, the three biomarkers are captured together in a single multiplexed image. Using the disclosed unmixing technique, an unmixing model 202 processes this triplex image 214, effectively separating the overlapping signals of ER, PR, and HER2 into their individual components. This results in synthetic singleplex images for each biomarker: a synthetic singleplex HER2 image 216, a synthetic singleplex ER image 218, and a synthetic singleplex PR image 220. These separate images can then be analyzed individually, enabling a more detailed and accurate assessment of the expression of each biomarker in the tissue sample.

[0069] FIG. 3A depicts an illustrative example 300-A of techniques for detecting cells of a specific type or a particular biomarker and determining a biomarker status based on a singleplex image, in accordance with some aspects of the present disclosure. These techniques may help in the detection of specific cell types, such as immune cells, tumor-positive, or tumor-negative cells, based on characteristic intensity patterns or pixel-level features associated with specific biomarkers. The disclosed techniques may include receiving a singleplex image that may be collected from the image generation system 102 or the unmixing model 202. For the illustrative example 300-A, three singleplex images (e.g., 212a, 212b, and 212c) are shown, each representing a single biomarker (e.g., ER, PR, or HER2) and utilizing two distinct stains: hematoxylin (Hema), which typically stains cell nuclei blue, and a biomarker-specific stain — Dabsyl (yellow) for ER, Tamra (pink) for PR, or green for HER2. For example, the singleplex image 212a includes information about biomarker presence, expressed through the yellow signal (Dabsyl) and the blue signal (hematoxylin) that highlights the nuclei.

[0070] To detect cells of the specific type, the particular biomarker, or a cellular component within the singleplex images e.g., 212a-c, a variety of techniques — such as statistical analysis, and machine-learning — may be utilized. These techniques aim to identify whether a given cell expresses a specific biomarker (e.g., positive or negative expression) and to assess the intensity of that expression (e.g., with an intensity score), indicating whether the positivity is strong, weak, or absent. To enable accurate classification, the classification model 302 may be trained using a dataset of manually annotated singleplex images, where expert pathologists provide ground-truth images 310 labeled for biomarker expression. Training may involve supervised learning techniques, where the model learns to associate staining patterns, intensity distributions, andspatial relationships with corresponding biomarker classifications. The training process may utilize deep learning architectures such as convolutional neural networks (CNNs) or transformerbased models, alongside traditional machine-learning classifiers (e.g., support vector machines (SVMs) or random forests). Additionally, data augmentation techniques may be applied to enhance model robustness against variations in staining intensity, imaging conditions, and tissue morphology.

[0071] Once trained, the classification model 302 may process incoming singleplex images (e.g., 212a-212c) to analyze biomarker expression at the cellular and subcellular levels. The trained model evaluates staining patterns, intensity distributions, and spatial relationships within the singleplex images to generate corresponding initially annotated singleplex images (e.g., 304, 306, and 308), which may then be used for diagnostic or prognostic assessment.

[0072] In some implementations, the classification model 302 is a single, unified model trained to classify multiple biomarkers (e.g., ER, PR, and HER2) within the singleplex images by utilizing multi-task learning, shared feature extraction, and distinct classification heads corresponding to each biomarker. Alternatively, in other implementations, multiple independent classifiers may be utilized, with each classifier specifically trained to evaluate a single biomarker. For example, a dedicated ER classifier may be configured to distinguish ER-positive versus ER-negative cells, while a separate PR classifier determines PR expression levels, and a HER2 classifier assesses membrane staining characteristics in accordance with CAP standards. The selection between a unified model or multiple specialized classifiers may depend on factors such as dataset availability, model interpretability, computational efficiency, and clinical validation requirements.

[0073] The disclosed classification techniques may be employed to predict whether one or more cells depicted in the singleplex image are positive or negative for each or any of a number of biomarkers, such as the estrogen receptor (ER), progesterone receptor (PR), or HER2. For example, cells positive for ER, PR, or HER2 may exhibit certain intensity patterns, or unique spatial distributions, that distinguish them from negative cells. The image 304 depicts an ER singleplex image, processed using classification model 302 to determine ER status at the cellular level. Cells are classified and labeled (color-coded) according to the ER expression status. For example, ER-positive (ER+) cells expressing a presence of ER expression are labeled in red, and ER-negative (ER-) cells expressing an absence of ER expression are labeled in yellow.Additional detected structures, such as stromal cells or artifacts are categorized as others that are labeled in black. The image 306 demonstrates a PR singleplex image that is similarly processed using the disclosed techniques (i.e., classification model 302). PR-positive (PR+) cells, which exhibit PR expression are marked in red, while PR-negative (PR-) cells, which lack PR expression are labeled yellow. Structures that do not fall into these categories e.g., including non- epithelial cells or background noise are classified as others and marked in black.

[0074] The image 308 represents an HER2 singleplex image, which is analyzed using the disclosed techniques to assess HER2 expression in tumor regions. HER2 is a transmembrane receptor protein that contributes to cell growth and proliferation. Overexpression of HER2 is associated with HER2-positive breast cancer, which is characterized by uncontrolled tumor growth and aggressive behavior. For designation or intensity quantification of HER2, the classification model 302 may follow the established guidelines provided by the College of American Pathologists (CAP), as illustrated in an exemplary flowgraph 300-B of FIG. 3B. A systematic decision pathway for HER2 intensity quantification using IHC is illustrated in FIG. 3B, which is based on tumor cell membrane staining characteristics. The disclosed techniques classify HER2 expression by analyzing tumor cell membrane staining intensity, following a structured quality control process. This classification may indicate the degree of HER2 protein expression within a particular tissue sample, which may further contribute to determining appropriate treatment options, such as whether a patient might benefit from HER2 -targeted therapies.

[0075] The exemplary flowgraph 300-B for HER2 expression evaluation (or intensity quantification) begins at 310, where IHC staining is performed on a tissue specimen using a validated assay, thereby enabling standardized biomarker detection. Subsequently, at 312, quality control (QC) verification is performed in which batch controls check whether the staining reagents and procedure are performing correctly across multiple tests. Additional on-slide controls may validate that the specific slide under evaluation exhibits acceptable staining quality by utilizing internal tissue controls, such as known HER2 -positive and HER2 -negative regions. If QC verification 312 fails, the staining process is reevaluated or repeated before proceeding with HER2 interpretation. Alternatively, tumor cell membrane staining analysis 314 is conducted, where HER2 designation prediction is based not only on the intensity of HER2 labeling but also on the location and completeness of staining within the cell membrane. Tumorcells exhibiting characteristics at 316, such as circumferential, complete, and intense membrane staining (as shown in 324) in more than 10% of cells, may be classified as HER2 overexpression or IHC 3+ positive.

[0076] Tumor cells displaying characteristics at 318, including weak to moderate but complete membrane staining (as shown in 326) in more than 10% of cells, may be classified as IHC 2+ equivocal. These cells may require confirmatory reflex testing via in situ hybridization (ISH) or repeat IHC. Similarly, tumor cells with characteristics at 320, such as incomplete, faint, or barely perceptible membrane staining (as shown in 328) in more than 10% of cells, may be classified as IHC 1+ negative. The tumor cells that exhibit characteristics at 322 such as depicting no staining or faint / incomplete membrane staining in more than 10% of the tumor cells, may be classified as IHC 0 negative 330. Following the HER2 intensity quantification characteristics, the singleplex HER2 image 212c may be analyzed using the disclosed techniques to assess HER2 expression in tumor regions, a higher HER2 designation (e.g., 3+) labeled in brown may be associated with a stronger and complete staining across the membrane, suggesting a higher abundance of HER2 protein, whereas moderate designation (e.g., 2+) labeled in red may indicate less robust or incomplete membrane staining and a low HER2 (e.g., 1+) expression labeled in green may indicate faint membrane staining. Non-tumor regions, such as benign tissue or stromal areas are labeled yellow, while tumor-negative regions, where no detectable HER2 expression is observed, are labeled orange.

[0077] FIG. 4 shows an example illustration of a validation framework 400 to perform validation and autocorrection 402 of the initially annotated images (e.g., 304-308), in accordance with some aspects of the present disclosure. To begin validation, a statical analysis for autocorrection 402 may be performed in which the synthetic singleplex images e.g., 212a-c may optionally undergo preprocessing 404, where various techniques such as contrast enhancement, normalization, and noise reduction may be applied to improve the quality of the image and enabling the signals to be clear and distinguishable. These preprocessing techniques may contribute to accurate identification of biomarker signals and reducing artifacts from staining. Subsequent to preprocessing 404, a binary mask 410 corresponding to a target singleplex image (e.g., HER2, ER, PR, or any other biomarker) may be generated. The binary mask 410 separates the cells from the background pixels in the image by applying cell segmentation 406. The process of cell segmentation 406 may involve identifying regions of interest where the cells arelocated, typically based on intensity differences or texture features that distinguish the cells from the surrounding background. Cell segmentation 406 may include e.g., thresholding, region growing, or deep learning-based methods (e.g., UNet or VNet) that delineate individual cells from the background and tissue structures.

[0078] In thresholding, a specific intensity value (threshold) is set for the image. For example, pixels in the image with intensity values above a certain threshold (indicating potential cell presence) may be considered part of a cell, while those below the threshold may be discarded as background. The thresholding may be effective in images where the cells have significantly different intensities compared to the background, such as when the biomarker or nuclear stain is strong and distinct. The output of thresholding may be the binary mask 410 where pixels inside a cell are marked as part of the cell, and the rest are labeled as background. The cell segmentation 406 process may identify and mark each cell, using both the hematoxylin (blue) and targeted biomarker e.g., HER2 (green) signal to accurately detect the cell boundaries.

[0079] The region growing technique starts with a seed point (often inside a cell) and then iteratively adds neighboring pixels based on a predefined similarity criterion, such as intensity or texture, until the boundary of the cell is reached. This approach is particularly useful in cases where cells are irregularly shaped. The output of region growing is a segmented region for each cell, where the pixels representing the cell are marked, and the background is excluded. Deep learning-based approaches e.g., convolutional neural networks (CNNs) may be used for cell segmentation 406. These models may be trained on large datasets to recognize and separate cells from background tissue by learning spatial patterns in the image. The output of a deep learningbased method may represent a highly accurate segmentation mask, which assigns a label to each pixel (cell or background) based on learned features from the training data. Such deep-learning models may handle complex cell shapes, overlapping cells, and noisy images effectively. The output of the cell segmentation 406 may include the binary mask 410 or segmented regions (e.g., labeled as cells and background) of cells, identifying the spatial boundaries of each cell in the singleplex image. Here, the binary mask 410 would represent a set of first predicted spatial boundaries for that marker.

[0080] Upon creation of binary mask 410, the centroids of the identified regions may be calculated as the geometric center of each individual segmented cell region. This is typically done by computing the average position of all the pixels that belong to a particular segmentedcell. Mathematically, the centroid (Cx, Cy) of a region is determined by finding the weighted average of the x and y coordinates of all pixels within the segmented area, where coordinates of each pixel are weighted by its intensity value (if applicable). The centroids of the positive responses are then used to determine the exact locations of the cell nuclei, serving as a precise reference for subsequent analysis or annotation. Here, the cell nuclei may represent a set of first predicted spatial positions for that marker.

[0081] Once the cell nuclei are identified from the synthetic singleplex image (e.g., HER2 image), their locations may be cross-referenced with generated pure stain images (or masks) 412 of other markers (e g., ER / PR masks). For generating pure stained images 412, a process— pure stain unmixing 408 may be performed on the synthetic singleplex images from other biomarkers to separate the two different staining signals. For a particular example of singleplex PR image 212b, there are two dyes: hematoxylin (blue) and Tamra (pink), each representing different components (nuclei and PR biomarker). The goal of the pure stain unmixing 408 is to separate these overlapping signals so that the contribution of each stain may be distinguished independently. The unmixing may be performed by a linear color unmixing technique— a statistical technique to separate mixed signals based on predefined spectral characteristics. It assumes that the observed pixel values are a linear combination of the underlying components (in this case, hematoxylin and Tamra (or Dabsyl)). In this technique, the contribution of each dye (blue for hematoxylin and pink for Tamra) may be estimated in each pixel. The unmixing process may use the known spectral signatures (intensity profiles) of the two dyes to calculate their individual contributions and separate the signals. Here, the pure stain image such as 412a may correspond to a set of second predicted spatial boundaries or positions.

[0082] The other unmixing techniques may include Gaussian mixture model (GMM), clustering, or non-negative matrix factorization (NNMF) that factorize the image data into two non-negative matrices representing the contributions from the two stains. GMM is a probabilistic model used to separate the color channels by assuming that the pixel intensities follow a Gaussian distribution for each stain. GMM estimates the contribution of each dye in a given pixel, effectively isolating the two signals in mixed pixels. Clustering algorithms (e.g., K-means) may separate the signals based on pixel intensity patterns, grouping similar pixels into clusters and assigning them to either of the stains present in the singleplex image (e.g., hematoxylin or Dabsyl stain). The output of the pure stain unmixing 408 may represent two unmixed imagesproviding two distinct intensity maps of the input singleplex image. For example, for the singleplex image 212a, one shows the isolated signal from hematoxylin (nuclei, blue) 412b and the other shows the isolated signal from Tamra (cytoplasm, pink) 412a.

[0083] Based on the intensity maps (or pure stains) e.g., 412a from the pure stain unmixing 408 and binary mask 410 from the cell segmentation 406 (or the extracted cell nuclei), biomarker detection may be performed by image overlaying 414. The binary mask 410 or the nuclei locations tells where the individual cells are located within the singleplex image, while the intensity maps or the pure stain images (e.g., 412a) may provide the intensity for the biomarker expression in each segmented cell. For biomarker detection, the segmented regions (output from cell segmentation 406) may be overlaid onto the pure stain image e.g., 412a to localize each cell region and analyze its biomarker signal.

[0084] For validating and autocorrecting initially annotated singleplex images e.g., 304, 306, 308, predefined designation rules may be applied to determine whether a given cell nucleus exhibits the expected biomarker expression. For example, if a cell is annotated as positive for a biomarker in the target singleplex image but lacks a corresponding signal in the pure biomarker mask, its designation may be corrected to negative. Similarly, for validating HER2 classification, presence of cell nuclei within ER / PR masks may suggest tumor cells. Other cases involving colocalization of multiple biomarkers may be assessed to confirm accurate classification. By iteratively cross-referencing the initial annotated singleplex image with the corresponding overlay, the system may automatically refine (or autocorrect) the biomarker classification to improve accuracy and reliability. Additionally, the validation framework 400 as illustrated in reference to FIG. 4 may be integrated into the disclosed machine-learning pipeline (e.g., including unmixing model 202, and classification model 302) subsequent to the generation of initially annotated singleplex images e.g., 304-308 to generate autocorrected singleplex images e.g., 416. Thus, enabling a robust and scalable solution for digital pathology applications.

[0085] The predicted biomarker status of cells within a tissue sample, as determined through the singleplex image, may be invaluable for pathologists in clinical settings. The predictions, which may include classifications of cells as positive or negative for specific biomarkers (e.g., ER, PR, or HER2), may assist in making diagnostic decisions regarding the presence or absence of certain diseases, such as breast cancer. Furthermore, the HER2 designation and other biomarker statuses may inform decisions regarding disease management,such as recommending targeted therapies or determining patient prognosis. In this manner, the use of classification techniques, informed by advanced image analysis methods, can aid pathologists in diagnosing complex conditions.

[0086] In some implementations, an interactive graphical user interface (GUI) may be configured to enable users (e.g., pathologists) to adjust the biomarker status of individual cells or groups of cells, as well as to rectify whether these cells or cell clusters are to be assigned a given label (e.g., tumor-positive or tumor-negative). The GUI may allow the user to view any synthetic singleplex image generated from the multiplex image to assist in their designations. For instance, the algorithm can generate a synthetic singleplex HER2 image with various cells labeled as tumor negative. By concurrently viewing corresponding synthetic singleplex images for ER and PR, the user may choose to change the tumor status of one or more cells in the singleplex HER2 image. Because the singleplex images are derived from the same multiplex image, they are inherently aligned and require no complicated registration steps to assess co-localization.

[0087] FIG. 5 illustrates an exemplary workflow 500 of a cross-modality translation model (CMTM) for translation between various staining modalities according to one aspect of the present disclosure. In some examples, techniques may be used to convert staining modalities between singleplex images to facilitate analyses that are dependent on specific modalities. For cross-modality conversion, the CMTM may be employed that transforms one modality of a singleplex image into another such that relevant features of the original image are preserved while adapting it to the targeted staining modality. In the present disclosure, a staining modality refers to a specific method or technique used to apply stains or dyes to biological tissues or cells to enhance contrast and visualize structures under a microscope. Different staining modalities (e.g., hematoxylin and eosin (H&E) staining, fluorescent staining, or DAB staining) highlight distinctive features of the tissue, such as specific proteins, cell types, or cellular structures that help to detect diseases such as cancer, infections, and organ abnormalities. DAB staining is one of the commonly used staining techniques in immunohistochemistry (IHC) that facilitates the spatial and quantitative visualization of target antigens at high resolution, particularly advantageous in clinical and research settings.

[0088] DAB images refer to histopathology images stained with 3, 3 '-Diaminobenzidine (DAB), a brown chromogen used in immunohistochemistry. The brown color produced by the DAB reaction is easily visible against a counterstained background (often hematoxylin), whichstains cell nuclei blue. This contrast enables the clear visualization of the presence, location, and quantity of the antigen, facilitating the diagnosis and research into various diseases. One aspect of the present disclosure may involve translating a singleplex image of one staining modality (e.g., an HERZ singleplex image) into a corresponding singleplex image of another staining modality (e.g., a synthetic DAB image) to facilitate detailed analysis of histopathology images from different perspectives.

[0089] For translating one staining modality into another, various techniques exist such as image processing, machine vision, or Al-based techniques. The image processing techniques such as contrast enhancement, intensity normalization, and histogram matching may be employed to prepare the input image. Additionally, or alternatively, deep learning models such as generative adversarial networks (GANs) and / or convolutional neural networks (CNNs) may be utilized for domain translation. For instance, CycleGAN, a variant of GAN, may be employed to train a generative model (e.g., 504) for performing domain translation, such as translating a singleplex image of first staining modality into a synthetic singleplex image of a second staining modality or vice versa. To train the generative model Gxy504, cycleGAN may be leveraged in which two sets of training data may be used, the first set comprises real singleplex images of dataset X 513 (e.g., first staining modality), and the second set comprises real singleplex images of dataset Y 511 (e.g., second staining modality). CycleGAN may use two generators and two discriminators to perform image translation between these domains.

[0090] The first generator Gxy504, may be trained on the dataset X 513 to translate a singleplex image from domain X to domain Y (e.g., converting an HER2 singleplex image into a synthetic DAB image), while the second generator Gyx508 may perform the inverse transformation, mapping a synthetic singleplex image from domain Y back to domain X (e.g., generating a synthetic HER2 singleplex image from a synthetic DAB image). Each generator, paired with a corresponding discriminator, which evaluates the realism of generated images. The discriminator Dx512 and Dy510 may function as the first convolutional neural network, responsible for categorizing and / or assessing the generated images. Meanwhile, the generators Gxy504 and Gyx508, functioning as the second convolutional neural network, particularly used to perform image-to-image translation between domains.

[0091] Images in Y' domain 506 are the generated images (e.g., synthetic DAB images) from the generator Gxy504, provided as an input to the discriminator Dy510. The discriminatorDy510 may access real DAB images from the dataset Y 511 during the training process to distinguish between real and the generated synthetic DAB images. Both the discriminators Dx512 and Dy510 may process the synthetic images and output a scalar value ranging from 0 to 1, representing a probability that the input image is either real or fake. For instance, a value close to one suggests the input image is real, while a value close to zero suggests that the input image is fake.

[0092] The discriminator may comprise multiple layers, including an input layer, convolutional layer, down-sampling layer, dense layer, and output layer, to assess image realism. It receives an input image — either real or generated — formatted as a 3D tensor (height, width, and channels, e.g., RGB). The convolutional layer extracts spatial features such as textures, edges, and patterns using filters to analyze local structures. To abstract image content, a downsampling layer (e.g., max-pooling or strided convolution) reduces spatial dimensions while preserving essential information. The dense layer then processes these compressed features to classify the image as real or fake. Finally, the output layer, typically a single neuron with a sigmoid activation function, maps the result to a probability between 0 (fake) and 1 (real).

[0093] To generate synthetic singleplex images that closely resemble real ones, the deep learning model relies on a feedback system to iteratively refine its output. One key mechanism is adversarial loss 514, which trains the generator to produce images that can successfully deceive the discriminator into classifying them as real. Additionally, cycle consistency loss 516 helps preserve the structural integrity of the generated images for retaining the essential features and properties of real singleplex images. This loss enforces that an image translated from one domain to another can be reverted to its original form with minimal distortion, maintaining consistency and fidelity.

[0094] At the beginning of the training process, the generator Gxy504 may not generate the output (e.g., synthetic DAB image) that closely resembles the real DAB image. However, through feedback from the discriminator Dy510, the generator Gxy504 gradually learns to adjust its parameters to generate the synthetic DAB images that appear more realistic. The generator processes the input image through various processes, such as encoding, transforming, and decoding, to generate the synthetic DAB image. For example, the input image may be first processed through multiple convolutional layers extracting features such as textures, shapes, and edges, while compressing the image and increasing the number of channels. For example, a256x256x3 image may be reduced to a 64x64x256 representation after encoding, further processing through one or more residual blocks (depending on the image size) may help retain major details of an image. Finally, the transformed image is up-sampled using deconvolution layers, restoring the input image back to its original size. It will be appreciated that this feature is not limited to conversion of images synthetic DAB image, but rather has the flexibility to convert between multiple singleplex image modalities.

[0095] FIG. 6 shows an illustrative example of a synthetic singleplex image, representing synthetic green HER2-Low 604, which is generated by processing a triplex image, representing ER-PR-HER2-Low 602, by leveraging the unmixing model 202. For cross-modality translation of the synthetic green HER2-Low image 604 into a synthetic DAB-stained image 606, the contrast of the original singleplex image may be enhanced to match the specific color intensity ranges found in typical DAB staining. Intensity normalization may help the pixel values to align with the expected distribution for the new modality, while histogram matching adjusts the overall intensity distribution of the image to resemble that of a DAB-stained image, making the transformation visually consistent. Alternatively, or additionally, CMTM 504 may take the synthetic green HER2-Low image 604 and generate the corresponding synthetic DAB-stained image 606. The initial annotated images 608 and 610 represent the annotated outputs of the classification model 302, corresponding to the input HER2-Low image 604 and DAB-stained image 606, respectively. The annotated outputs include predictions with respect to five classes, including 1+, 2+, 3+, non-tumor and tumor negative labeled in green, red, brown, yellow and orange, respectively.

[0096] In some embodiments, a graphical user interface may enable a user to generate and view masks, associated with various biomarkers to assist in validating the accuracy of the annotated images. For example, a mask may be created for the chromophore Dabsyl (yellow) or the fluorescent dye Tamra (pink). In this instance, Dabsyl may be associated with ER while PR may be associated with Tamra. However, it will be appreciated that these designations are only used as examples, and that a given biomarker may be labeled with a broad range of markers. The masking may generate a binary image with pure Dabsyl or Tamra pixels associated with a certain intensity to be labeled in the mask. The pixels in the mask representing positive staining may represent the set of first predicted spatial boundaries or positions for that marker.

[0097] FIG. 7 shows an illustrative example of a GUI 700 including a singleplex ER image 702 and a singleplex PR image 708, along with their corresponding masks. The GUI 700 may include a toolbar 716 enabling a user to perform various image processing operations on to a selected image (e.g., crop, erase or delete, change color, image enhancements such as adjusting brightness or contrast and adding label to the image). For performing other general operations, the GUI 700 may offer a menu bar 714 for saving, print, sharing, or inserting new image or shape. These singleplex ER and PR images (e.g., 702 and 708) may be generated either from a triplex image by applying the unmixing model 202 or image generation system 102. Each of these singleplex images may comprise two distinct colors (or stains). For instance, Hematoxylin, a general stain that typically appears blue or purple, is used to highlight tissue structure, while specific markers such as Tamra (pink) visualize PR expression, and Dabsyl (yellow) are used to highlight ER expression.

[0098] By applying pure stain unmixing 408 or non-linear unmixing methods (such as the disclosed unmixing model 202), the pure Dabsyl ER image 704 and the pure TAMRA-based PR image 710 may be extracted from the singleplex ER image 702 and singleplex PR image 708, respectively. This capability to generate pure stain images from a given singleplex image is accessible through the GUI 700, where the user can select the “Unmix Image” option 720, which triggers the application of the unmixing process (pure stain unmixing 408). Once the pure stain images are generated, these may be optionally processed to create corresponding binary masks, such as ER mask 706 and PR mask 712. This functionality may be achieved by selecting the “Generate Mask” option 718 in the GUI 700, which in turn invoke the cell segmentation 406. Additionally, the GUI 700 may enable the creation of annotated images, by selecting “Generate Annotated Image” 722, which triggers the application of the classification model 302 to the selected image, providing detailed insights into the biomarker status for further analysis.

[0099] As discussed before, the cell segmentation 406 involve generation of the binary masks by segmenting the images through thresholding techniques that differentiate areas of interest — such as regions with significant biomarker expression — from the surrounding background (e.g., similar to the process of cell segmentation 406). This process may enable precise identification of cells expressing the targeted biomarkers for downstream analysis and interpretation. For example, in case of the pure Dabsyl ER image 704, a predefined intensity threshold may be applied to isolate regions where the ER marker is significantly present, creatinga binary mask (e.g., 706) that highlights only the cells with high ER expression. Similar techniques are applied to the pure Tamra PR image 710, where the Tamra staining intensity for PR expression is thresholded to create a corresponding binary mask for PR-expressing cells.

[0100] In some aspects, the GUI may present an overlay image e.g., by overlaying a binary mask of the synthetic singleplex HER2 image onto the generated mask (e.g., ER / PR masks 706 or 712). It may be validated automatically or via a user (e.g., a pathologist user) based on the overlaid image and the set of predefined designation rules that the assigned labels or annotations from the initially annotated singleplex images are correct or not. Upon identifying any wrong labels, label corrections may be initiated manually or automatically. For example, for validating HER2 classification, presence of cell nuclei within ER / PR masks may suggest tumor cells.

[0101] FIG. 8 shows an example of how masks may be used to validate accuracy of an annotated singleplex image and to perform correction, in accordance with some aspects of the present disclosure. The singleplex image (e.g., synthetic singleplex HER2 804) may be derived from a multiplex image (e.g., a triplex image 802 representing ER-PR-HER2). By selecting “Unmix Image” 720 from a GUI 800, which triggers the application of the unmixing model 202, the multiplex image may be unmixed to its constituent synthetic singleplex images. Similarly, by selecting “Generate Annotated Image” 722 from the GUI 800, classification techniques (i.e., 302) may be triggered to generate an initially annotated HER2 image 804. Additionally, the GUI 800 may support and display a comparative analysis mode, allowing the user to easily compare the initial annotated singleplex image 804 (before correction) with an autocorrected or manual corrected singleplex image 806, along with an overlay image 808.

[0102] Overlaying binary mask from HER2 and pure stain mask from ER / PR, an overlay image 808 may be generated, where the ER mask is labeled in black, and the PR mask is labeled in grey). The nuclei located within the biomarker-positive regions defined by the binary masks provide the basis for classifying the corresponding cells as both tumor cells and biomarkerpositive. The tumor classification can be inferred based on the co-localization of the nuclei within these biomarker-positive areas. In the specific example of FIG. 8, the presence of cell nuclei within the ER / PR masks suggests that these cells express the relevant biomarkers and should be classified as ER / PR-positive tumor cells. Subsequently, the superimposed results (e.g., 808) of both the ER / PR mask and the binary mask from HER2 identifying cell nuclei locations are used to generate an autocorrected image 806. For example, the red rectangle in the ER / PRmask highlights two tumor-negative cells (labeled in orange) that do not align with the initially annotated HER2 image 804, where these two cells are mistakenly labeled as non-tumor (or immune) cells in yellow. However, since these cells are located within the ER (Dabsyl) and PR (Tamra) positive regions, these cells should be reclassified as tumor cells. As a result, the system automatically updates the labels, resulting in the autocorrected image 806.

[0103] Additionally, the GUI 800 may enable users to modify various annotations manually by selecting a data point or a group of data points from the annotated image (e.g., 806) and changing its designated label from a dropdown list or other interface options (e.g., “Edit Label” 810g or “Add Label” 81 Oh). This includes both binary choices (e.g., positive or negative for a biomarker) as well as more nuanced or graded annotations, such as those for biomarkers with varying levels of expression. For example, the HER2 biomarker in breast cancer diagnostics is commonly categorized into levels such as 0, 1+, 2+, or 3+, indicating the intensity of expression around the cell membrane. The GUI 700 and / or 800 may enable users to adjust both binary and graded labels for individual cells or groups of cells. These cells may be selected individually by clicking on them, or in bulk by encircling them with a selection tool (e.g., 810a) that encompasses a region of interest, thus enabling editing of cell annotations.

[0104] Furthermore, the toolbar 716 may include tools to modify the intensity of expression, allowing for finer granularity. For example, pathologists may adjust the intensity scale of a biomarker such as HER2, providing greater insight into the varying levels of expression across a tissue sample. The inclusion of a "Brush Tool" 810d may let users paint over regions of interest, gradually increasing or decreasing the intensity, while a "Flood-Fill" 810c tool may help users quickly label large, contiguous areas of cells with the same grade or expression level. The GUI 800 may also include image enhancement features 81 Of for adjusting the brightness and contrast of the annotated image to enhance visibility and detail. Users can modify the visual properties of the image using sliders or other controls to fine-tune the display for better clarity. Additionally, the GUI may provide an option to change the color of specific data points (e.g., from red to yellow) to improve contrast or visibility. This can be done by selecting from a color palette (e.g., 81 Oe), allowing users to choose a color that best highlights the relevant markers or regions for easier analysis.

[0105] FIG. 9A shows an illustrative example of a GUI 900-A enabling a flexible selection of region of interest, in accordance with some aspects of the present disclosure. Alternately, oraddition to initial automatic correction, the user (or pathologist) may perform a screening to review the annotated images e.g., 902 generated by applying the disclosed techniques. Subsequently, percentages for different classifications, such as HER2 1+, 2+, 3+ and tumornegative cells may be calculated. For the annotated image 902, representing a synthetic singleplex HER2 image, the GUI 900-A displays a table 904 showing the percentages of cells categorized into these four classes, allowing the pathologist to assess the accuracy of the automated annotations. During this review, the user may identify wrongly labeled cells, such as immune cells that have been incorrectly classified as HER2 -negative cells. To facilitate this process, the GUI 900-A offers tools for more precise and flexible identification and correction. For instance, the lasso selection tool 906 allows the pathologist to draw custom, freehand shapes around areas of interest, such as groups of cells or specific regions of the image e.g., 902a and 902b that need revision. Once selected, the user may modify the labels or classifications of those cells, whether correcting binary classifications or adjusting graded designations such as HER2 1+, 2+, or 3+

[0106] FIG. 9B shows an example of the GUI 900-B enabling shape-based selection of regions of interest, in accordance with some aspects of the present disclosure. For the synthetic singleplex PR image 908, the GUI 900-B displays a table 910 with the percentages of cells categorized into binary classes, helping the user assess the accuracy of automated annotations. The user may identify misclassified cells, such as cells misclassified tumors as non-tumor for regions of interest (e.g., 908a, 908b, 908c and 908d). Using the shape select tool 912, predefined shapes (e g., circles, triangles) can be drawn around the regions of interest (e g., 908a, 908b) to revise labels or classifications, correcting binary classifications such as PR+ or PR-.

[0107] FIG. 9C illustrates the GUI 900-C enabling shape-based selection for the synthetic singleplex ER image 914, in accordance with some aspects of the present disclosure. The GUI 900-C displays a table 916 showing the percentages of cells categorized as ER+ or ER-, aiding the user in reviewing the accuracy of automated annotations. Misclassified cells, such as those incorrectly identified as ER- or ER+, can be easily corrected. Using the shape select tool 912 or Lasso selection tool 906, a shape may be drawn around regions of interest to revise labels or classifications, correcting binary choices such as ER+ or ER-.

[0108] FIG. 9D illustrates an example of a GUI 900-D featuring an annotated synthetic singleplex PR image 918, along with a corresponding zoomed-in regions of interest (ROI) 920and 922 both before and after selection, respectively. Initially, the user selects the region of interest 918a using a free-hand lasso selection tool 906. Once the selection is made, the cells within the defined area are highlighted, as depicted by circles around them in 922. The zoomedin images 922 provides a closer view, where the selected cells are distinctly marked with red circles, visually emphasizing the highlighted regions of interest. This allows for a clearer and more focused view of the cells that are relevant to the user analysis.

[0109] To detect or identify cell nuclei, the biomarkers present in a multiplex image may be processed collectively or separately based on the localization of the biomarker, biological function, and clinical relevance. Biomarkers localized in similar cellular compartments, such as nuclear biomarkers (e.g., ER and PR), can be processed together in a unified detection model (e.g., a duplex model) due to their similar staining distribution patterns. In contrast, biomarkers that localize in different compartments — such as membrane-bound biomarkers such as HER2 — require specialized models tailored to their distinct staining characteristics and localization patterns. These differences necessitate the use of separate computational techniques to accurately capture their expression. For example, nuclear biomarkers such as ER and PR are typically analyzed using thresholding and region-based segmentation, while membrane-bound biomarkers such as HER2 require methods such as edge detection or membrane segmentation.

[0110] FIG. 10 shows an illustrative example of automatic correction of initially annotated images, in accordance with some aspects of the present disclosure. Following the disclosed techniques, a triplex ER-PR-HER2-Low image 602 may be unmixed and fed to the classification model 302 for generating annotated images corresponding to synthetic singleplex ER image 1002, synthetic singleplex PR image 1004 and synthetic singleplex HER2 image. The nuclei centers (or seeds) from binary mask of HER2 may be used for extracting the set of predicted spatial positions to validate or automatically correct the initial annotations of synthetic singleplex ER / PR images (e.g., 1002, and 1004). To this end, the ER mask 1008 and PR mask 1010 may be generated separately or combined, such as by overlaying the pure stain images from both biomarkers to generate a unified ER / PR mask for further analysis (e.g., as shown in 808 of FIG. 8).

[0111] In this particular example, ER mask 1008 and PR mask 1010 are generated and overlaid onto the binary mask separately. For classification, if a seed or cell nucleus from the HER2 model lies within ER / PR masks, it is reclassified as tumor cells and assigned ER / PRpositive based on the mask it belongs to. Conversely, if the seed is outside the ER / PR mask, it is classified as ER / PR negative. Furthermore, if an ELER2 non-tumor estimation is found within the ER / PR masks, it is reclassified as a corresponding positive cell. In FIG. 10, triangle regions in each image highlight seed locations and their assigned labels, which are corrected in the respective ER (Dabsyl) mask 1008 and PR (Tamra) mask 1010, where red cells indicate positive classifications, while yellow cells denote negative ones.

[0112] Based on predefined classification rules, the annotated synthetic singleplex ER 1002 include misclassified ER+ labels in triangle region 1002a. The corrected ER mask 1008 updates this to ER- in triangle region 1008a, as the corresponding ER mask 1008 lacks the ER expression within triangle 1008a. Similarly, the annotated synthetic singleplex PR 1004 includes a misclassified PR+ label in triangle region 1004a. The corrected PR mask 1010 adjusts this to PR- in triangle region 1010a, since the PR mask 1010 lacks the PR expression in the corresponding triangle region 1010a and no cell nuclei are marked in the corresponding HER2 triangle region 1006a.

[0113] FIG. 11 shows an illustrative example of three annotated images, including the duplex ER / PR 1102, synthetic singleplex ER 1104, and synthetic singleplex PR 1106 images. In this process, the initially annotated synthetic singleplex ER 1104 and synthetic singleplex PR image 1106 are generated to identify individual biomarkers (ER and PR) within the tissue sample. These initial annotations provide preliminary classifications for ER+ and PR+ cells based on thresholding techniques. However, due to potential inaccuracies in the initial classification, the duplex ER / PR 1102 model is then applied to refine the seed classifications. This model simultaneously processes both ER and PR biomarkers, updating the initial annotations based on more precise, pre-trained analysis. Specifically, it corrects the classifications of cells, assigning the correct ER / PR positive or negative status, as represented by the distinct color codes: red for ER+PR+, green for ER+PR-, blue for ER-PR+, and yellow for ER-PR- cells. This updated classification process, as seen in FIG. 11, enables that the system provides an accurate mapping of the biomarker status across the image, facilitating improved tumor analysis and correct labeling of the ER / PR positive and negative regions within the annotated image.

[0114] It will be appreciated that there are several manners in which an overlap between a mask and a cell may be calculated and that certain approaches may introduce small biasestowards type 1 or type 2 errors. One example of how the mask may be used is shown in flowchart 1202 of FIG. 12. First, cell nuclei are detected using automated techniques, as described in reference to FIG. 4. After the seed locations are identified, a 5x5 mask (which is a small square or window of size 5x5 pixels centered on a seed) may be computed. This mask may be used to define a neighborhood around the seed where the surrounding pixels will be evaluated. For each label or nucleus, it may be determined that the nuclei lie inside the 5x5 mask, at 1204 e.g., at the center of 5x5 pixels.

[0115] Upon validation, it may be determined whether a label (e g., ER+ or PR+) from the seed locations is correctly labeled or is consistent with the corresponding staining information, at 1208. To this end, the Tamra staining information, which highlights PR expression, may be obtained from singleplex PR image that generates pure PR mask, and the Dabsyl staining information, which highlights ER expression, may be obtained from the singleplex ER image generating pure ER mask. If the seed’s label falls inside the PR (Tamra) or ER (Dabsyl) staining regions, then the label stays the same, retaining the original label. At 1210, when the corresponding location in the Tamra / Dabsyl masks indicates PR+ / ER+, any estimations within the 5x5 neighborhood that are wrongly labeled will be automatically corrected. It may be appreciated that the size of this second mask may be adjusted to compensate for imaging resolution, magnification, or cell size. Additionally, this process may be applied to multiple singleplex images highlighting different biomarkers.

[0116] FIG. 12 further illustrates an automatic correction process applied to an initially annotated synthetic singleplex estrogen receptor (ER) image 1212. This image represents the original annotation, where specific regions (1212a, 1212b, and 1212c) have been identified as misclassified based on the applied correction methodology. These misclassified regions are indicated by triangular markers and correspond to the classification categories of ER-positive (ER+, red dots), ER-negative (ER-, yellow dots), and other (black dots). The correction process is performed using the window-based approach that incorporates duplex-generated seed locations to compute 5><5 masks. If the original labels fall within these 5x5 masks, they are retained. However, when the corresponding locations in the Tamra / Dabsyl fields of view (FOVs) indicate PR+ / ER+, any estimations within the 5x5 neighborhood that are misclassified are corrected accordingly.

[0117] For example, in FIG. 12, region 1212b, originally annotated as ER- (yellow dot), is corrected to ER+ (red dot) in the corresponding corrected annotations (1214b). This correction is justified by the 5><5 mask surrounding the cell nuclei, which includes staining information indicating ER positivity. Similarly, region 1212c, originally classified as ER-, is reclassified based on the algorithm, which assigns ER+ when Tamra / Dabsyl FOVs indicate ER staining within the neighborhood. The image 1214 depicts the revised annotations after the automatic correction process, where previously misclassified regions (1214a, 1214b, and 1214c) have been corrected to reflect their appropriate classifications.

[0118] FIG. 13 illustrates an exemplary workflow 1300 to generate annotated singleplex images based on a multiplex image in accordance with some aspects of the present disclosure. The blocks in the workflow 1300 are illustrated in a specific order, while the order can be modified, for example, some blocks may be performed before others, and some blocks may be performed simultaneously. The blocks can be performed by hardware, software, or a combination thereof. For analyzing a particular biomarker, a process at block 1302 may include accessing multiplex digital -pathology images 112a-n from image generation system 102 that depicts a biological specimen labeled with multiple different markers. These multiple different markers are selected to identify different biomarkers and / or cellular components relative to each other. Multiplex images provide detailed depictions of various cellular structures or biomarkers.However, analyzing these images to identify specific cell types (e g., immune cells) or staining patterns can be challenging for pathologists due to their complexity.

[0119] At block 1304, multiple synthetic singleplex images (e.g., 212a, 212b, 212c) may be generated from a single multiplex image to analyze the pattern, location, presence of individual biomarker (e.g., HER2, PR, ER) precisely. These synthetic singleplex images, representing the depiction of individual biomarkers, can be easily annotated by pathologists or computational algorithms, may be used for further applications. One application of these annotated images is to use them as ground-truth images 310 for training classification model 302, enabling automated generation of annotated images. This approach may significantly reduce the time and effort required by pathologists, eliminating the need to manually annotate each individual cell.

[0120] At block 1306, the process may involve generating initial annotated singleplex images from synthetic singleplex images using the trained classification model. The initial annotated singleplex image, including predictions e.g., classifications of cells as positive ornegative for particular biomarkers (e.g., ER, PR, or HER2), may assist in making diagnostic decisions regarding the presence or absence of certain diseases, such as breast cancer. These initial annotated images may be validated manually e.g., by configuring the GUI for the user to adjust the annotations or alternatively, by performing a statistical analysis that involves various operations. The statistical analysis may process the synthetic singleplex images and undergo operations including preprocessing 404, cell segmentation 406, pure stain unmixing 408 and / or image overlaying 414 for generating accurate ground-truth annotated singleplex images. Cell segmentation 406 involves separating the cells from background pixels by identifying the specific region where particular type of cells are located to generate the binary mask 410 corresponding to a target singleplex image, at block 1308. Subsequent to cell segmentation 406, the synthetic singleplex images may be fed to the pure stain unmixing model 408 to unmix or separate pure stained images, at block 1310.

[0121] The binary mask 410 may correspond to a first biomarker of the multiple different biomarkers, identifying a set of first predicted spatial boundaries or positions within a first synthetic singleplex image of the multiple synthetic singleplex images. The process at block 1310 may further generate an overlay based on the binary mask and the pure stain mask. At block 1312, an initial annotated singleplex image may be validated based on the overlay with respect to a set of predefined designation rules. The graphical user interface (GUI) may be generated or updated to show the overlay and the initial annotated singleplex images. At block 1314, the process may include outputting the generated or updated GUI to a user device, where pathologist / user can review the resulted images and may perform manual corrections of the wrongly predicted annotations. The GUI may offer various tools / options to let pathologists make corrections and adjustments of wrong predictions.

[0122] Although specific aspects have been described, various modifications, alterations, alternative constructions, and equivalents are possible. Embodiments are not restricted to operation within certain specific data processing environments but are free to operate within a plurality of data processing environments. Additionally, although certain aspects have been described using a particular series of transactions and steps, it should be apparent to those skilled in the art that this is not intended to be limiting. Although some flowcharts describe operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process may have additional steps notincluded in the figure. Various features and aspects of the above-described aspects may be used individually or jointly.

[0123] Further, while certain aspects have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are also possible. Certain aspects may be implemented only in hardware, or only in software, or using combinations thereof. The various processes described herein can be implemented on the same processor or different processors in any combination.

[0124] Where devices, systems, components or modules are described as being configured to perform certain operations or functions, such configuration can be accomplished, for example, by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, such as by executing computer instructions or code, or processors or cores programmed to execute code or instructions stored on a non-transitory memory medium, or any combination thereof. Processes can communicate using a variety of techniques including but not limited to conventional techniques for inter-process communications, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0125] Specific details are given in this disclosure to provide a thorough understanding of the aspects. However, aspects may be practiced without these specific details. For example, well- known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail in order to avoid obscuring the aspects. This description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of other aspects. Rather, the preceding description of the aspects can provide those skilled in the art with an enabling description for implementing various aspects. Various changes may be made in the function and arrangement of elements.

[0126] The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. It can, however, be evident that additions, subtractions, deletions, and 1 other modifications and changes may be made thereunto without departing from the broader spirit and scope as set forth in the claims. Thus, although specific aspects have been described, these are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method comprising: accessing a multiplex image that depicts a biological specimen labeled with multiple different staining modalities, wherein the multiple different staining modalities are selected to identify multiple different biomarkers and / or cellular components relative to each other; generating multiple synthetic singleplex images using the multiplex image; generating multiple initial annotated images that correspond to multiple different biomarkers based on the multiple synthetic singleplex images via a classification model, wherein: the classification model is configured to predict a biomarker status of a particular cell or a cellular component and an associated intensity of presence of a particular biomarker within a synthetic singleplex image of the multiple synthetic singleplex images, and an initial annotated image of the multiple initial annotated images includes a set of predicted annotations that corresponds to a set of predicted spatial boundaries or positions within the synthetic singleplex image of the multiple synthetic singleplex images; generating a binary mask that corresponds to a first biomarker of the multiple different biomarkers by identifying a set of first predicted spatial boundaries or positions within a first synthetic singleplex image of the multiple synthetic singleplex images, wherein each of the set of first predicted spatial boundaries or positions are identified so as to predict a boundary or position of the first biomarker, the particular cell or particular cellular component corresponding to the first synthetic singleplex image; generating a pure stain mask that corresponds to a second biomarker of the multiple different biomarkers from a second synthetic singleplex image of the multiple synthetic singleplex images; generating, based on the binary mask and the pure stain mask, an overlay;validating the initial annotated image based on the overlay with respect to a set of predefined designation rules; generating or updating a graphical user interface to show the overlay and the initial annotated image; and outputting the graphical user interface to a user device.

2. The computer-implemented method of claim 1, further comprising: identifying, as a result of the validation, that at least one predicted annotation of the set of predicted annotations within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component; and generating an autocorrected image that includes a change or a corrected annotation corresponding to the at least one predicted annotation.

3. The computer-implemented method of claim 1, further comprising: generating, based on the first synthetic singleplex image, a translated image by applying a generative model that is trained in a cyclic adversarial manner, wherein the translated image corresponds to a different staining modality than the multiple different staining modalities of the multiplex image.

4. The computer-implemented method of claim 1, further comprising: receiving, from the user device, a communication that identifies or changes at least one predicted annotation of the set of predicted annotations that corresponds to a boundary or a position of the particular biomarker or the particular cellular component within the initial annotated image, wherein the communication indicates that the at least one predicted annotation within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component.

5. The computer-implemented method of claim 4, wherein the communication identifies a user-identified boundary, and wherein the computer-implemented method further comprises:detecting one or more predicted annotations of the set of predicted annotations corresponding to one or more predicted spatial boundaries or positions of the set of predicted spatial boundaries or positions that are within the user-identified boundary; and removing the one or more predicted spatial boundaries or positions from the set of predicted spatial boundaries or positions.

6. The computer-implemented method of claim 1, wherein the pure stain mask comprises a set of second predicted spatial boundaries or positions within the second synthetic singleplex image, wherein each of the set of second predicted spatial boundaries or positions are identified so as to predict a boundary or a position of the second biomarker or another particular cellular component.

7. The computer-implemented method of claim 1, wherein the multiple synthetic singleplex images are generated using a linear unmixing technique or a non-linear unmixing technique.

8. The computer-implemented method of claim 1, wherein the set of predicted annotations identifies: a depiction of a biological component that is not a tumor cell, and a presence, a relative level of presence, or an absence of the particular biomarker associated with the particular cell or the cellular component.

9. A system comprising: one or more data processors; and a non-transitory computer readable storage medium containing instruction which, when executed on the one or more data processors, cause the one or more data processors to perform a set of operations including: accessing a multiplex image that depicts a biological specimen labeled with multiple different staining modalities, wherein the multiple different staining modalities are selected to identify multiple different biomarkers and / or cellular components relative to each other;generating multiple synthetic singleplex images using the multiplex image; generating multiple initial annotated images that correspond to multiple different biomarkers based on the multiple synthetic singleplex images via a classification model, wherein: the classification model is configured to predict a biomarker status of a particular cell or a cellular component and an associated intensity of presence of a particular biomarker within a synthetic singleplex image of the multiple synthetic singleplex images, and an initial annotated image of the multiple initial annotated images includes a set of predicted annotations that corresponds to a set of predicted spatial boundaries or positions within the synthetic singleplex image of the multiple synthetic singleplex images; generating a binary mask that corresponds to a first biomarker of the multiple different biomarkers by identifying a set of first predicted spatial boundaries or positions within a first synthetic singleplex image of the multiple synthetic singleplex images, wherein each of the set of first predicted spatial boundaries or positions are identified so as to predict a boundary or position of the first biomarker, the particular cell or particular cellular component corresponding to the first synthetic singleplex image; generating a pure stain mask that corresponds to a second biomarker of the multiple different biomarkers by identifying a set of second predicted spatial boundaries or positions within a second synthetic singleplex image, wherein each of the set of second predicted spatial boundaries or positions are identified so as to predict a boundary or a position of the second biomarker or another particular cellular component corresponding to the second synthetic singleplex image; generating, based on the binary mask and the pure stain mask, an overlay; validating the initial annotated image based on the overlay with respect to a set of predefined designation rules; generating or updating a graphical user interface to show the overlay and the initial annotated image; and outputting the graphical user interface to a user device.

10. The system of claim 9, wherein the set of operations further comprising: identifying, as a result of the validation, that at least one predicted annotation of the set of predicted annotations within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component; and generating an autocorrected image that includes a change or a corrected annotation corresponding to the at least one predicted annotation.

11. The system of claim 9, wherein the set of operations further comprising: generating, based on the first synthetic singleplex image, a translated image by applying a generative model that is trained in a cyclic adversarial manner, wherein the translated image corresponds to a different staining modality than the multiple different staining modalities of the multiplex image.

12. The system of claim 9, wherein the set of operations further comprising: receiving, from the user device, a communication that identifies or changes at least one predicted annotation of the set of predicted annotations that corresponds to a boundary or a position of the particular biomarker or the particular cellular component within the initial annotated image, wherein the communication indicates that the at least one predicted annotation within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component.

13. The system of claim 9, wherein the set of operations further comprising: receiving, from the user device, a communication that identifies a user-identified boundary; detecting one or more predicted annotations of the set of predicted annotations corresponding to one or more predicted spatial boundaries or positions of the set of predicted spatial boundaries or positions that are within the user-identified boundary; and removing the one or more predicted spatial boundaries or positions from the set of predicted spatial boundaries or positions.

14. The system of claim 9, wherein the multiple synthetic singleplex images are generated using a linear unmixing technique or a non-linear unmixing technique.

15. The system of claim 9, wherein the set of predicted annotations identifies: a depiction of a biological component that is not a tumor cell, and a presence, a relative level of presence, or an absence of the particular biomarker associated with the particular cell or the cellular component.

16. A computer-program product for image editing tangibly embodied in a non- transitory machine readable storage medium, including instructions configured to cause one or more data processors to perform a set of operations comprising: accessing a multiplex image that depicts a biological specimen labeled with multiple different staining modalities, wherein the multiple different staining modalities are selected to identify multiple different biomarkers and / or cellular components relative to each other; generating multiple synthetic singleplex images using the multiplex image; generating multiple initial annotated images that correspond to multiple different biomarkers based on the multiple synthetic singleplex images via a classification model, wherein: the classification model is configured to predict a biomarker status of a particular cell or a cellular component and an associated intensity of presence within a synthetic singleplex image of the multiple synthetic singleplex images, and an initial annotated image of the multiple initial annotated images includes a set of predicted annotations that corresponds to a set of predicted spatial boundaries or positions within the synthetic singleplex image of the multiple synthetic singleplex images; generating a binary mask that corresponds to a first biomarker of the multiple different biomarkers by identifying a set of first predicted spatial boundaries or positions within a first synthetic singleplex image of the multiple synthetic singleplex images, wherein each of the set of first predicted spatial boundaries or positions are identified so as to predict a boundary or position of the first biomarker, the particular cell or particular cellular component corresponding to the first synthetic singleplex image;generating a pure stain mask that corresponds to a second biomarker of the multiple different biomarkers by identifying a set of second predicted spatial boundaries or positions within a second synthetic singleplex image, wherein each of the set of second predicted spatial boundaries or positions are identified so as to predict a boundary or a position of the second biomarker or another particular cellular component; generating, based on the binary mask and the pure stain mask, an overlay; validating the initial annotated image based on the overlay with respect to a set of predefined designation rules; generating or updating a graphical user interface to show the overlay and the initial annotated image; and outputting the graphical user interface to a user device.

17. The computer-program product of claim 16, wherein the set of operations further including: identifying, as a result of the validation, that at least one predicted annotation of the set of predicted annotations within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component; and generating an autocorrected image that includes a change or a corrected annotation corresponding to the at least one predicted annotation.

18. The computer-program product of claim 16, wherein the set of operations further including: receiving, from the user device, a communication that identifies or changes at least one predicted annotation of the set of predicted annotations that corresponds to a boundary or a position of the particular biomarker or the particular cellular component within the initial annotated image, wherein the communication indicates that the at least one predicted annotation within the initial annotated image does not correspond to the particular biomarker, the particular cell or the cellular component.

19. The computer-program product of claim 16, wherein the set of operations further including: receiving, from the user device, a communication that identifies a user-identified boundary; detecting one or more predicted annotations of the set of predicted annotations corresponding to one or more predicted spatial boundaries or positions of the set of predicted spatial boundaries or positions that are within the user-identified boundary; and removing the one or more predicted spatial boundaries or positions from the set of predicted spatial boundaries or positions.

20. The computer-program product of claim 16, wherein the set of operations further including: generating, based on the first synthetic singleplex image, a translated image by applying a generative model that is trained in a cyclic adversarial manner, wherein the translated image corresponds to a different staining modality than the multiple different staining modalities of the multiplex image.

Citation Information

Patent Citations

  • Quantitation of Signal in Stain Agrregates

    US20220148176A1