Deep learning enabled slide image analysis for prediction of treatment response

A deep learning system for analyzing whole slide images improves treatment response prediction by precisely identifying biomarkers in biological samples, addressing the limitations of current methods and enabling personalized therapy selection.

WO2026096551A1PCT designated stage Publication Date: 2026-05-07MERCK SHARP & DOHME LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MERCK SHARP & DOHME LLC
Filing Date
2025-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current methods for predicting treatment response in patients are inadequate, particularly in determining the presence of biomarkers in biological samples to guide therapy selection, lacking precision and efficiency in identifying relevant features for treatment efficacy.

Method used

A deep learning-enabled system and method for analyzing whole slide images of biological samples using segmentation models, feature extraction, and response prediction models to identify and classify cells based on target protein staining patterns, enabling precise prediction of treatment response.

Benefits of technology

Enhances the accuracy of predicting treatment response by identifying relevant biomarkers, allowing for personalized therapy selection and improved patient eligibility assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052983_07052026_PF_FP_ABST
    Figure US2025052983_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A method includes applying a segmentation model to segment an image depicting a biological sample. The segmentation model generates a cell map in which each pixel is assigned a value indicative of whether a corresponding pixel in the image depicts a cell. A feature extraction model is applied to determine a feature set based on the cell map. The feature set includes primitive features indicative of a presence or absence of a target protein. The feature set further includes higher order features derived from the primitive features. A response prediction model is applied to determine, based on the feature set, a response prediction for a patient associated with the biological sample. The response prediction indicates a likelihood of the patient responding to the therapeutic. The patient is identified as eligible for receiving the therapeutic based on the response prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No. 14463-679-228

[0002] DEEP LEARNING ENABLED SLIDE IMAGE ANALYSIS FOR PREDICTION OF TREATMENT RESPONSE

[0003] CROSS REFERENCE TO RELATED APPLICATION

[0004] [1] This application claims priority to U.S. Provisional Application No. 63 / 713,005, entitled “IMAGE ANALYSIS OF IMMUNOHISTOCHEMICALLY STAINED TISSUE SECTIONS FOR PREDICTION OF RESPONSE TO AN ANTIBODY DRUG CONJUGATE” and filed on October 28, 2024, the disclosure of which is incorporated herein by reference in its entirety.

[0005] TECHNICAL FIELD

[0006] [2] The present disclosure relates generally to digital pathology and more specifically to deep learning based techniques for analyzing tissue slide images to predict treatment response. INTRODUCTION

[0007] [3] Assessing a patient's likelihood of responding to a particular therapy is an important step in determining an appropriate treatment regimen for the patient. This assessment may include analyzing biological samples to identify biomarkers (e.g., histological biomarkers, molecular biomarkers, and / or the like) for determining one or more medically actionable classifications for the patient including, for example, disease diagnosis, disease stage, disease subtype, clinical response, and / or the like.

[0008] SUMMARY

[0009] [4] In one aspect, there is provided a system, a method, and a computer program product for a deep learning enabled pathological analysis workflow in which one or more deep learning models are applied to analyze slide images for predicting treatment response. For example, in some embodiments, a patient’s clinical response to a therapeutic may be determined based on an analysis of an image (e.g., a whole slide image (WSI)) depicting a biological sample (e.g., tissue sample, bodily fluids, and / or the like) to detect the presence (or absence) of one or more relevant biomarkers, such as a tumor associated antigen (TAA). In instances where the therapeutic is an anticancer drug binding to a target protein (e.g., TROP 2 or PD-L1) expressed by tumor cells, the image of the biological sample may be analyzed for the presence of one or more features indicative of the presence of the target protein. In some cases, the deep learning based pathological analysis workflow may include segmenting an image of a biological sample (e.g., tissue sample, bodily fluids, and / or like) to identify one or more regions of interest (ROIs). In some cases, each region of interest (ROI) may be further segmented to identify and classify one or more constituent cells (e.g., tumor cells, immune cells, stromal cells, and / or the like).

[0010] 1

[0011] NAI-5005412604vl [5] In some examples, a method for classifying cell membranous staining patterns includes (a) assessing a cell’s target protein staining intensity, (b) determining whether the cell contains differentiable edge signals and, if so, determining whether the edge signals can be differentiated from a nuclear boundary, (c) classifying the cell as target protein negative if no target protein staining is detected, (d) classifying the cell as having cytoplasmic-only staining if target protein staining is detected but no differentiable edge signals are present, (e) classifying the cell as having membranous staining if target protein expression is detected and the cell contains differentiable edge signals that are differentiable from the nuclear boundary, and (f) classifying the cell as being in a gray zone if target protein staining is detected but the cell contains differentiable edge signals that cannot be differentiated from the nuclear boundary.

[0012] [6] In some examples, a system for classifying cell membranous staining patterns includes a detector for detecting target protein staining of a cell, a classifier configured to categorize the cell based on a presence or absence of differentiable edge signals and a relationship of the edge signals to a nuclear boundary, and a processor for executing classification of the cell as one of (i) target protein negative if no target protein staining is detected, (ii) cytoplasmic-only staining if the target protein staining is detected and no differentiable edge signals are present, (iii) membranous staining if target protein staining is detected, and the cell contains differentiable edge signals that are differentiable from the nuclear boundary, or (iv) a gray zone if target protein staining is detected but the cell contains differentiable edge signals that cannot be differentiated from the nuclear boundary.

[0013] [7] In some examples, a method for detecting and classifying cells in a biological sample includes identifying ROIs within an image of the biological sample automatically using a digital pathology foundation model, detecting cell nuclei and classifying cells within the ROIs using a machine learning model trained on image data, and outputting the classification of the cells within segmented ROIs.

[0014] [8] In some examples, a system for cell detection and classification in a biological sample includes a digital pathology module configured to identify ROIs within an image of the biological sample and segment these regions based on tumor and non-tumor areas, a machine learning classifier for detecting and classifying cells within the segmented regions based on their nuclear and membrane features, and a visualization tool for generating a predictive mask of tumor cell classification for validation by a pathologist.

[0015] [9] In some examples, a method for categorizing tumor tissue specimens as ‘biomarker positive’ or ‘biomarker negative’ includes obtaining a digital image of a biological sample, extracting a plurality of cell and whole-image level tissue features from the digital image and

[0016] 2

[0017] NAI-5005412604vl assigning biomarker status for the tissue specimen based on a dichotomous phenotype that is derived from the cell and whole-image level tissue features.

[0018]

[0010] The method for characterizing tumor tissue specimens as ‘biomarker positive’ or ‘biomarker negative’ includes obtaining a digital image of a biological sample, extracting a plurality of local features that are centered on individual cells and their vicinities or defined local regions of the whole slide image (WSI) from the digital image, and assigning biomarker status for the tissue specimen based on a dichotomous phenotype that is derived from the local features.

[0019]

[0011] In some examples, a method for characterizing tumor tissue specimens as ‘biomarker positive’ or ‘biomarker negative’ includes obtaining a digital image of a biological sample, extracting a plurality of features from the digital image and in some instances assigning biomarker status for the tissue specimen based on a dichotomous phenotype that is derived from the plurality of both local as well as global (or whole image) features.

[0020]

[0012] In some examples, a method for extracting imaging features from an image (e.g., a whole slide image (WSI)) in digital pathology includes (a) dividing the image into tiles of a predetermined size, (b) deploying one or more pretrained vision foundation models (VFMs) to extract feature vectors from each tile, (c) collecting the extracted feature vectors into clusters to reduce dimensionality, and (d) selecting the most representative tiles from each cluster for biological significance evaluation by a pathologist.

[0021]

[0013] In some examples, a method for predicting clinical response using a composite machine learning based multivariate model includes obtaining a digital image of a biological sample, extracting a plurality of features from the digital image that include at least one local feature, at least one global (or whole image) feature and at least one de novo feature guided by clinical response, analyzing these features within the multivariate model, classifying the specimen as ‘positive’ or ‘negative’ for a biomarker based on the model-derived score in relation to cutpoints as determined in a clinical study, and predicting clinical response in patients based on the biomarker status classification.

[0022]

[0014] In some examples, a method includes setting a cut-off on a digital feature or composite as a component of an assay applied to tumor tissue for selecting a subgroup of patients for enrollment into a clinical trial, or for designating a subset of an enrolled population, assessing an investigational therapy for a higher likelihood of demonstrating superior efficacy to standard of care in the subgroup thus selected.

[0023] 3

[0024] NAI-5005412604vl

[0015] In some examples, a method includes setting a cut-off on a digital feature or a composite as a component of a diagnostic device that identifies patients for whom a cancer therapy has been approved based on the benefit / risk established from clinical trials that used the diagnostic.

[0025]

[0016] In another aspect, there is provide a system for deep learning enabled slide image analysis for prediction of treatment response. In some example embodiments, the system may include at least one data processor and at least one memory storing instructions that result in operations when executed by the at least one data processor. The operations may include: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; and identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction.

[0026]

[0017] In another aspect, there is provide a computer implemented method for deep learning enabled slide image analysis for prediction of treatment response. In some example embodiments, the method may include: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the

[0027] 4

[0028] NAI-5005412604vl one or more primitive features present in the image; applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; and identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction.

[0029]

[0018] In another aspect, there is provide a computer program product for deep learning enabled slide image analysis for prediction of treatment response. In some example embodiments, the computer program product may include a non-transitory computer readable medium storing instructions that cause operations when executed by at least one data processor. The operations may include: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; and identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction.

[0030]

[0019] In another aspect, there is provided a method of treating cancer involving administering to a cancer patient a therapeutic that is useful for treating cancer. The method may include: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction

[0031] 5

[0032] NAI-5005412604vl model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction; and administering the therapeutic in response to the patient being identified as eligible for receiving the therapeutic.

[0033]

[0020] In another aspect, there is provided a method of treating cancer involving administering to a cancer patient a therapeutic that is useful for treating cancer. The method may include: administering the therapeutic to a patient in response to the patient being identified as eligible for receiving the therapeutic, where the patient is identified as eligible for receiving the therapeutic based at least on a response prediction of the patient, and where the response prediction of the patient is determined by at least applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample associated with the patient, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; and applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic.

[0034]

[0021] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0035]

[0022] In some variations, the image is of a whole slide image (WSI).

[0036] 6

[0037] NAI-5005412604vl

[0023] In some variations, the target protein is TROP2 or PD-L1.

[0038]

[0024] In some variations, the therapeutic is an antibody drug conjugate (ADC), an alkylating agent, an antimetabolite, an enzyme inhibitor, an anti-angiogenesis agent, a microtubule disruptor, and / or an immunotherapy agent.

[0039]

[0025] In some variations, the biological sample depicted in image is treated with a stain sensitive to the target protein.

[0040]

[0026] In some variations, the biological sample depicted in the image is also treated with an additional stain sensitive to at least one other subcellular compartment.

[0041]

[0027] In some variations, the one or more primitive features include a location or a type of a cell present in the image or a region of interest (ROI) within the image.

[0042]

[0028] In some variations, the one or more primitive features include a completeness of a target protein sensitive stain around a cell present in the image or a region of interest (ROI) within the image.

[0043]

[0029] In some variations, the completeness of the stain is quantified by a proportion of (i) a first quantity of pixels depicting a membrane of the cell with the target protein sensitive stain and (ii) a second quantity of pixels depicting the membrane of the cell without the target protein sensitive stain.

[0044]

[0030] In some variations, the one or more primitive features include a location of a target protein sensitive stain around a cell present in the image or a region of interest (ROI) within the image.

[0045]

[0031] In some variations, the location of the target protein sensitive stain comprises one or more cellular compartments of the cell in which the target protein sensitive stain is present.

[0046]

[0032] In some variations, the one or more primitive features include a morphology of a cell present in the image or a region of interest (ROI) within the image.

[0047]

[0033] In some variations, the morphology of the cell comprises a differentiability of one or more cellular compartments of the cell.

[0048]

[0034] In some variations, the feature extraction model determines the one or more primitive features by at least: determining, for each cell present in the cell map, whether an area around the cell exhibits a signal from a protein sensitive stain sensitive to a target protein of the therapeutic; classifying the cell as negative for the target protein where the signal from the protein sensitive stain is absent from the area around the cell; where the area around the cell exhibits the signal from the protein sensitive stain, classifying the cell as having cytoplasmic only staining where the cell lacks a differentiable cell membrane, and where the area around the cell exhibits the signal from the protein sensitive stain and the cell exhibits the differentiable

[0049] 7

[0050] NAI-5005412604vl cell membrane, classifying the cell as having membranous staining where the cell membrane is differentiable from a nuclear boundary of the cell, and classifying the cell in a gray zone where the cell membrane is indifferentiable from the nuclear boundary of the cell.

[0051]

[0035] In some variations, the signal from the protein sensitive stain indicates an intensity of the staining.

[0052]

[0036] In some variations, a differentiability of a cell membrane of the cell is determined based at least on a signal from a stain sensitive to a cytoplasm and extracellular matrix.

[0053]

[0037] In some variations, a differentiability between the cell membrane of the cell and the nuclear boundary of the cell is determined based on the signal from the stain sensitive to the cell membrane and a stain sensitive to a nucleus of the cell.

[0054]

[0038] In some variations, the at least one higher order feature includes a regional feature localized to a portion of the image.

[0055]

[0039] In some variations, the regional feature comprises a tessellated feature extracted from a portion of the image.

[0056]

[0040] In some variations, the portion of the image comprises a tessellation that is 150 microns to 600 microns in size.

[0057]

[0041] In some variations, the regional feature comprises a local feature extracted from a cell- anchored neighborhood defined by a radius around a cell.

[0058]

[0042] In some variations, the cell-anchored neighborhood is defined by a radius of 8 microns to 60 microns.

[0059]

[0043] In some variations, the at least one higher order feature includes an additional regional feature having a different scale or resolution.

[0060]

[0044] In some variations, the additional regional feature comprises an additional tessellated feature extracted from a different size portion of the image.

[0061]

[0045] In some variations, the additional regional feature comprises an additional local feature extracted from a cell-anchored neighborhood defined by a different radius around the cell.

[0062]

[0046] In some variations, the regional feature includes a distribution of cells positive for the target protein of the therapeutic within the portion of the image.

[0063]

[0047] In some variations, the regional feature includes an intensity of a stain sensitive to the target protein within the portion of the image.

[0064]

[0048] In some variations, the regional feature includes a heterogeneity of a stain sensitive to the target protein within the portion of the image.

[0065] 8

[0066] NAI-5005412604vl

[0049] In some variations, the regional feature includes at least one of a clustering pattern of cells positive for the target protein and / or cells negative for the target protein within the portion of the image.

[0067]

[0050] In some variations, the regional feature includes a type of cell anchoring a neighborhood comprising the portion of the image.

[0068]

[0051] In some variations, the regional feature includes a spatial heterogeneity of cells positive for the target protein and cells negative for the target protein within the portion of the image.

[0069]

[0052] In some variations, the regional feature identifies a type of region comprising the portion of the image as one or more of total tissue, invasive region, or in situ region.

[0070]

[0053] In some variations, the regional feature identifies a stain intensity quantification technique as one or more of optical density (OD) based, proportion based, or pixel based.

[0071]

[0054] In some variations, the regional feature identifies a size of the portion of the image.

[0072]

[0055] In some variations, the at least one higher order feature includes a summary statistic computed for at least a portion of the image.

[0073]

[0056] In some variations, the summary statistic is computed for one or more quantiles.

[0074]

[0057] In some variations, the summary statistic is computed for multiple quantiles in quantile increments.

[0075]

[0058] In some variations, the feature extraction model determines, for inclusion in the feature set, one or more predetermined human interpretable features (HIFs) extracted from the image.

[0076]

[0059] In some variations, the feature extraction model determines, for inclusion in the feature set, one or more de novo features extracted from the image.

[0077]

[0060] In some variations, the feature extraction model determines the feature set by at least: extracting, from the image, one or more human interpretable features (HIFs); identifying, based at least on the one or more human interpretable features (HIFs), one or more cells positive for a target protein of the therapeutic; extracting, for inclusion in the feature set, one or more de novo features from a neighborhood anchored by each cell identified as positive for the target protein.

[0078]

[0061] In some variations, the response prediction model comprises a univariate model trained to determine, based on each individual feature in the feature set, the response prediction for the patient.

[0079]

[0062] In some variations, the response prediction model comprises a multivariate model trained to determine, based on multiple features from the feature set, the response prediction for the patient.

[0080] 9

[0081] NAI-5005412604vl

[0063] In some variations, the response prediction model comprises one or more of a random forest model, a decision tree, a gradient-boosted decision tree, or an elastic net.

[0082]

[0064] In some variations, the one or more segmentation model are applied to generate a segmentation map localizing one or more regions of interest (ROIs) present in the image. The one or more segmentation models are applied to generate, based at least on the segmentation map, the cell map such that the cell map localizes cells present in each region of interest (ROI).

[0083]

[0065] In some variations, the one or more regions of interest (ROIs) include tumor regions and excludes non-tumor regions.

[0084]

[0066] In some variations, the one or more regions of interest (ROIs) include tumor regions and non-tumor regions.

[0085]

[0067] In some variations, each pixel in the segmentation map is assigned a value indicative of whether a corresponding pixel in the image depicts a region of interest (ROI).

[0086]

[0068] In some variations, the segmentation map includes one or more color channels.

[0087]

[0069] In some variations, the one or more regions of interest (ROIs) include one or more of a total tissue, an invasive region, or an in situ region.

[0088]

[0070] In some variations, the cell map includes one or more color channels.

[0089]

[0071] In some variations, each pixel in the cell map is further assigned a value indicative of a type of cell depicted by the corresponding pixel in the image.

[0090]

[0072] In some variations, the at least one higher order feature comprises a global feature.

[0091]

[0073] In some variations, the global feature includes an overall spatial heterogeneity of cells positive for the target protein and cells negative target protein across the image.

[0092]

[0074] In some variations, the overall spatial heterogeneity is quantified based on one or more of a quantity, uniformity, and degree of clustering of the cells positive for the target protein and the cells negative target protein.

[0093]

[0075] In some variations, the overall spatial heterogeneity is quantified based on one or more of a colocalization hotspot score, tumor cell hotspot score, and Morisita-Horn index bivariate correlation.

[0094]

[0076] In some variations, the global feature includes an overall proportion of cells exhibiting each of a plurality of different intensity levels of a stain sensitive to the target protein of the therapeutic.

[0095]

[0077] In some variations, the plurality of different intensity levels include a weak level of stain intensity, a moderate level of stain intensity, and a strong level of stain intensity.

[0096]

[0078] In some variations, the global feature includes a metric quantifying an intensity of a stain sensitive to the target protein and proportion of cells positive for the target protein.

[0097] 10

[0098] NAI-5005412604vl

[0079] In some variations, the metric comprises a digital H-score.

[0099]

[0080] In some variations, the feature set is further determined by at least: performing dimensionality reduction to reduce a quantity of features extracted from the image by the feature extraction model, where the dimensionality reduction identifies groups of features exhibiting sufficient covariance and correlation to be added to the feature set collectively as a single feature instead of as multiple individual features.

[0100]

[0081] In some variations, the one or more segmentation models include a region of interest (ROI) segmentation model trained to localize, within the image, one or more regions of interest (ROIs).

[0101]

[0082] In some variations, the one or more regions of interest (ROI) include a total tissue region, an invasive tumor region, or an in situ tumor region.

[0102]

[0083] In some variations, the one or more segmentation models further include a cell segmentation model trained to localize, within each region of interest (ROI), the one or more cells.

[0103]

[0084] In some variations, the cell segmentation model includes a foundation model that has been pretrained on a corpus of unlabeled or weakly labeled images of biological samples.

[0104]

[0085] In some variations, the pretrained foundation model incrementally finetuned using annotated samples over multiple successive timesteps. The finetuned foundation model is applied to localize, within each region of interest (ROI), the one or more cells.

[0105]

[0086] In some variations, the feature extraction model determines the feature set by at least: partitioning, into a plurality of tiles, the image; extracting, from each tile of the plurality of tiles, a feature vector including one or more features present in the tile.

[0106]

[0087] In some variations, a tile-level representation of the tile is determined based at least on the feature vector of each tile of the plurality of tiles. The response prediction model to is applied to determine, based at least on the tile-level representation of each tile of the plurality of tiles, the response prediction.

[0107]

[0088] In some variations, an image-level representation of the image is determined based at least on the feature vector extracted from each tile of the plurality of tiles. The response prediction model is applied to determine, based at least on the image-level representation of the image, the response prediction.

[0108]

[0089] In some variations, the image-level representation of the image is generated by applying, to the feature vector of each tile of the plurality of tiles, a geospatially adjusted attention weight computed based on the tile and one or more neighboring tiles.

[0109] 11

[0110] NAI-5005412604vl

[0090] In some variations, the feature extraction model determines the presence or absence of the target protein by at least classifying the cell as positive for the target protein or negative for the target protein.

[0111]

[0091] In some variations, the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by least identifying, within the image of the biological sample, an expansion region including a set of pixels depicting a cytoplasm and a membrane of the cell, identifying, within the expansion region, one or more pixels positive for a target protein sensitive stain used to treat the image, dividing the expansion region into a plurality of partitions, determining a stain-positive proportion of the plurality of partitions, wherein a partition is stain-positive if a threshold quantity of pixels within the partition are positive for the stain, and classifying, based at least on the stain-positive proportion of the plurality of partitions, the cell as positive or negative of the target protein.

[0112]

[0092] In some variations, the feature extraction model identifies the expansion region by at least performing a radial expansion from a nucleus of the cell.

[0113]

[0093] In some variations, the radial expansion identifies the expansion region to be proportional in size to a size of the nucleus

[0114]

[0094] In some variations, the radial expansion identifies the expansion region to be proportional to a relative crowdedness of a neighborhood of the cell, and the relative crowdedness of the neighborhood of the cell corresponds to a ratio between a geometric area of the neighborhood and a quantity of cells occupying the neighborhood.

[0115]

[0095] In some variations, the cell is an immune cell that is classified as positive for the target protein if the cell exhibits either cytoplasmic staining or membranous staining.

[0116]

[0096] In some variations, the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least identifying a plurality of pixels associated with the cell, determining whether a threshold quantity of the plurality of pixels associated with the cell is positive for a target protein sensitive stain used to treat the image, and classifying the cell as negative for the target protein without the threshold quantity of the plurality of pixels associated with the cell being positive for a target protein sensitive stain used to treat the image.

[0117]

[0097] In some variations, the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the threshold quantity of the plurality of pixels associated with the cell is determined to be positive for the target protein sensitive stain, determining whether a membrane of the cell is differentiable, and where the

[0118] 12

[0119] NAI-5005412604vl membrane of the cell is determined to be indifferentiable, classifying the cell as exhibiting only cytoplasmic staining.

[0120]

[0098] In some variations, a differentiability of the membrane of the cell is determined by at least determining a quantity of cytoplasmic pixels between an outer perimeter of a nucleus of the cell and an outer perimeter of the cell, and determining the cell as exhibiting an indifferentiable membrane if the quantity of cytoplasmic pixels fails to satisfy one or more thresholds.

[0121]

[0099] In some variations, the differentiability of the membrane of the cell is further determined by least in response to the quantity of cytoplasmic pixels satisfying the one or more thresholds, determining a gradient shift corresponding to a ratio between high intensity pixels and low intensity pixels between an outer perimeter of a nucleus of the cell and an outer perimeter of the cell, and determining the cell as exhibiting a differentiable membrane if the gradient shift satisfies one or more thresholds.

[0122]

[0100] In some variations, the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the membrane of the cell is determined to be differentiable, determining whether the membrane of the cell is differentiable from a nuclear boundary of the cell, and where the membrane and the nuclear boundary of the cell are determined to be indifferentiable, classifying the cell as exhibiting indifferentiable membranous staining and cytoplasmic staining.

[0123]

[0101] In some variations, the cell exhibiting the indifferentiable membranous staining and cytoplasmic staining is further classified as being in a gray zone of cells positive for the target protein

[0124]

[0102] In some variations, the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the membrane and the nuclear boundary of the cell are determined to be differentiable, classifying the cell as exhibiting membranous staining, in response to the cell being determined to exhibit membranous staining, determining an intensity and a completeness of the membranous staining, and classifying the cell as being positive for the target protein if the intensity and the completeness of the membranous staining satisfies one or more thresholds.

[0125]

[0103] In some variations, the completeness of the membranous staining is determined by at least dividing the cell into a plurality of slices, and determining the completeness of the membranous staining to correspond to a ratio of a quantity of stain positive slices relative to a total quantity of the plurality of slices.

[0126] 13

[0127] NAI-5005412604vl

[0104] In some variations, the membranous staining is attributed to the cell instead of one or more adjacent cells when the intensity and completeness of the membranous staining satisfies the one or more thresholds.

[0128]

[0105] In some variations, the plurality of pixels associated with the cell are identified by at least deconvolving the image of the biological sample to separate a nuclear channel corresponding to a nuclei sensitive stain and a membrane channel corresponding to a target protein sensitive stain, applying the one or more segmentation models to identify, within the nuclear channel of the image, one or more pixels depicting a nucleus of the cell , and applying the one or more segmentation models to identify, within the membrane channel of the image, one or more pixels corresponding to a boundary of the cell.

[0129]

[0106] In some variations, one or more pixels depicting a cytoplasm of the cell depicting a cytoplasm of the cell are identified within the image of the biological sample. The one or more pixels depicting the cytoplasm of the cell comprise one or more pixels between the nucleus and the boundary of the cell.

[0130]

[0107] In some variations, the one or more segmentation models are based at least on a training dataset of images of a stain sensitive to a different target protein.

[0131]

[0108] In some variations, the one or more segmentation models include a membrane segmentation model trained to infer a boundary of cells negative for the target protein.

[0132]

[0109] In some variations, the cell is a tumor cell that is classified as positive for the target protein if the cell exhibits membranous staining.

[0133]

[0110] In some variations, the image is preprocessed to detect one or more pixels depicting non-specific staining. The one or more pixels depicting non-specific staining are excluded from being used by the one or more segmentation models to generate the cell map.

[0134]

[0111] In some variations, the preprocessing of the image further includes deconvolving the image to separate a first channel corresponding to a stain sensitive to the target protein from a second channel corresponding to a stain sensitive to one or more subcellular compartments, applying, to the first channel, one or more color filters comprising one or more thresholds on at least one of hue, saturation, and value of a plurality of pixels comprising the image of the biological sample, and excluding, from being used by the one or more segmentation models to generate the cell map, one or more pixels whose hue, saturation, and / or value fails to satisfy the one or more thresholds.

[0135]

[0112] In some variations, the preprocessing of the image further includes excluding, from being used by the one or more segmentation models to generate the cell map, one or more clusters of adjacent pixels whose size fails to satisfy one or more thresholds.

[0136] 14

[0137] NAI-5005412604vl

[0113] In some variations, the first channel corresponds to a hematoxylin dye and the second channel corresponds to an immunohistochemistry (IHC) stain.

[0138]

[0114] In some variations, the feature set includes (i) 90thquantile of target protein positive neighborhood pixel based proportion, (ii) 50thquantile target protein positive neighborhood cell based membrane stain intensity, (iii) 50thquantile target protein positive neighborhood cell based cytoplasm stain intensity, (iv) 90thquantile target protein positive neighborhood cell based cytoplasm stain intensity, (v) 50thquantile 150-micron regional target protein cytoplasm stain intensity, (vi) 90thquantile 150-micron regional target protein cytoplasm stain intensity, and (vii) multivariate model composite score of (i)-(vi).

[0139]

[0115] In some variations, the feature set includes a subset of the following plurality of features: percentage of optical density (OD) 1+ tumor cells based on pathologist matched cut points, percentage of OD 2+ tumor cells based on pathologist matched cut points, percentage of OD 3+ tumor cells based on pathologist matched cut points, H-score based on pathologist matched cut points, density of OD 1+ tumor cells based on pathologist matched cut points, density of OD 2+ tumor cells based on pathologist matched cut points, density of OD 3+ tumor cells based on pathologist matched cut points, percentage of OD 1+ tumor cells based on data driven cut points, percentage of OD 2+ tumor cells based on data driven cut points, percentage of OD 3+ tumor cells based on data driven cut points, H-score based on data driven cut points, density of OD 1+ tumor cells based on data driven cut points, density of OD 2+ tumor cells based on data driven cut points, and density of OD 3+ tumor cells based on data driven cut points, 10th percent quantile of the neighborhood- wise average target protein sensitive stain optical density (OD) values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest, 90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest, 10th percent quantile of average target protein sensitive stain OD values (OD based calculation)

[0140] 15

[0141] NAI-5005412604vl among the tessellated regions defined by 150um * 150 um, 50th percent quantile of average target protein sensitive stain OD values (OD based calculation) among the tessellated regions defined by 150um * 150 um, 90th percent quantile of average target protein sensitive stain OD values (OD based calculation) among the tessellated regions defined by 150um * 150 um, 10th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest, 90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest, 10th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um, 50th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um, 90th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um, 10th percent quantile of the neighborhoodwise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 90th

[0142] 16

[0143] NAI-5005412604vl percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest, 10th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest, 50th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest, and 90th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest.

[0144]

[0116] In some variations, a therapeutically effective quantity of the therapeutic is administered once a week, once every 2 weeks, once every 3 weeks, once every 4 weeks, or once every 5 weeks.

[0145]

[0117] In some variations, the therapeutically effective quantity of the therapeutic comprises a dose of 1 to 30 mg / kg, 1 to 20 mg / kg, 2-12 mg / kg, 2-5 mg / kg, or 4-7 mg / kg.

[0146]

[0118] In some variations, the therapeutically effective quantity of the therapeutic comprises a dose of 2 mg / kg, 2.5 mg / kg, 3 mg / kg, 3.5 mg / kg, 4 mg / kg, 4.5 mg / kg, 5 mg / kg, 5.5 mg / kg, 6 mg / kg, 6.5 mg / kg, 7 mg / kg, 7.5 mg / kg, 8 mg / kg, 8.5 mg / kg, 9 mg / kg, 9.5 mg / kg, 10 mg / kg, 11 mg / kg, or 12 mg / kg.

[0147]

[0119] In some variations, the therapeutic targets the target protein or a different protein.

[0148]

[0120] In some variations, the therapeutic targets a protein in cancer.

[0149]

[0121] In some variations, the therapeutic targets a protein on a cell surface or a cell interior.

[0150]

[0122] In some variations, the target protein is TROP2.

[0151]

[0123] In some variations, the therapeutic is sacituzumab tirumotecan.

[0152]

[0124] In some variations, the sacituzumab tirumotecan is used to treat the patient for breast cancer, non-small cell lung cancer, ovarian cancer, cervical cancer, endometrial cancer, gastric cancer, prostate cancer, esophageal cancer, or urothelial cancer.

[0153] 17

[0154] NAI-5005412604vl

[0125] In some variations, the tissue sample is a tumor tissue sample.

[0155]

[0126] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0156]

[0127] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the pharmacophore-based generative design of drug molecules, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.

[0157] BRIEF DESCRIPTION OF THE DRAWINGS

[0158]

[0128] Various embodiments of the presently disclosed methods are disclosed herein with reference to the drawings, wherein:

[0159]

[0129] FIG. 1A is a system diagram illustrating an example of a digital pathology system implementing a deep learning enabled pathological analysis workflow for analyzing images of biological samples.

[0160]

[0130] FIG. IB is a block diagram illustrating an example of a pathological analysis engine for extracting features from an image of a biological sample and determining a response prediction based on the features.

[0161] 18

[0162] NAI-5005412604vl

[0131] FIG. 1C is a block diagram illustrating another example of a pathological analysis engine for extracting features from an image of a biological sample and determining a response prediction based on the features.

[0163]

[0132] FIG. ID is a block diagram illustrating an example of a deep learning enabled pathological analysis workflow for analyzing an image of a biological sample.

[0164]

[0133] FIG. IE is a flowchart illustrating an example of a process for deep learning enabled pathological analysis for prediction of treatment response.

[0165]

[0134] FIG. IF is a block diagram illustrating an example of a process for extracting features from an image of a biological sample.

[0166]

[0135] FIG. 1G is a schematic illustration of a computing device that performs the analysis, including typical hardware components like processors, input / output devices, storage, and network communication capabilities.

[0167]

[0136] FIG. 2 is an illustration of an IHC image that has undergone automated segmentation with tumor region identification and ROI segmentation, highlighting evaluable and tumor regions: Left image: original IHC image, Top right: ground truth annotation mask for the tumor ROI by a pathologist, Bottom right: tumor ROI prediction probability (ranging from 0 to 1) heatmap by the model.

[0168]

[0137] FIG. 3 shows the performance of the tumor cell detection model in two different tumor types: non-small cell lung cancer (NSCLC, left) and triple negative breast cancer (TNBC, right). The cell prediction masks (yellow in NSCLC and red in TNBC) by the tumor cell detection models are overlaid on top of the original image.

[0169]

[0138] FIGS. 4A-4B show side-by-side images showing the performance of the cell detection model, comparing the raw image with ALgenerated tumor cell masks.

[0170]

[0139] FIGS. 5-7 illustrate two models used for membrane detection and segmentation of tumor cells, using (a) a previously trained IHC membrane model (Top) and (b) a multichannelbased cell boundary model (Bottom).

[0171]

[0140] FIG. 8 illustrates a composite model combining nuclear and membrane detection, highlighting a cell with nucleus, membrane, and cytoplasm features.

[0172]

[0141] FIG. 9 is a flowchart for classifying target protein (e.g., TROP2 or PD-L1) staining patterns into one of four categories.

[0173]

[0142] FIGS. 10A-C and 11-13 are patches of tumor cells where the center tumor cell has various target protein (e.g., TROP2 or PD-L1) staining patterns that belong to the different predicted categories.

[0174] 19

[0175] NAI-5005412604vl

[0143] FIG. 14 is a diagram illustrating quantitative assessment of target protein (e.g., TROP2 or PD-L1) staining completeness using one or more membrane stain metrics, such as positive pixel proportion (PPP).

[0176]

[0144] FIG. 15 illustrates a performance of the nuclear detection model with a prediction mask to enable pathologist validation.

[0177]

[0145] FIG. 16 illustrates the use of a square tessellation method to divide an image (e.g., a WSI) into squares for analyzing the spatial distribution of target protein positive and negative cells.

[0178]

[0146] FIG. 17 illustrates a Voronoi tessellation method to divide an image (e.g., a WSI) into polygons for analyzing the spatial distribution of target protein positive and negative cells.

[0179]

[0147] FIGS. 18A-18B illustrate one example of the spatial distribution between different tumor clusters and different types of spatial features that can be extracted.

[0180]

[0148] FIG. 19 illustrates the use of a hotspot feature extraction process.

[0181]

[0149] FIGS. 20-21 illustrate the segregation and co-localization of cancer and immune cells with respective Pearson Correlation and Morisita-Hom Index calculations.

[0182]

[0150] FIG. 22 illustrates methods of partitioning target protein negative tumor cells into 3 classes based on their proximity to the target protein positive tumor cell surface.

[0183]

[0151] FIG. 23 A depicts the use of image features.

[0184]

[0152] FIG. 23B illustrates the use of tessellated and cell-centered neighborhood analysis techniques to generate cell level and tessellation level digital features.

[0185]

[0153] FIG. 23C illustrates a method of analyzing pixel-level digital features.

[0186]

[0154] FIG. 24 illustrates the types of whole image level digital features that mimic features generated by pathologists.

[0187]

[0155] FIG. 25 illustrates the correlation of global features derived from both pathologistbased and data-driven cut-points.

[0188]

[0156] FIG. 26 is an overview of local feature extraction using tumor cells as anchor cells to evaluate target protein (e.g., TROP2 or PD-L1) expression in neighboring cells.

[0189]

[0157] FIG. 27 is a heatmap of Spearman correlations for a full collection of initial local digital features.

[0190]

[0158] FIG. 28 illustrates the correlation structure of local features across cell compartments.

[0191]

[0159] FIG. 29 illustrates local feature correlations across different regions of interest (ROIs).

[0192]

[0160] FIG. 30 illustrates correlations across various neighborhood radii for local analysis.

[0193]

[0161] FIG. 31 highlights features with minimal biological variability that were removed from further analysis.

[0194] 20

[0195] NAI-5005412604vl

[0162] FIG. 32 illustrates local features that were removed after correlation analysis.

[0196]

[0163] FIG. 33 illustrates a heatmap for the 35 selected local and global features, using continuous Spearman correlation.

[0197]

[0164] FIGS. 34-39 are images highlighting exemplary images showing the most extreme values for the local digital features for visual inspection of these features.

[0198]

[0165] FIG. 40 is a comparison showing correlation between local features and global features.

[0199]

[0166] FIG. 41 is a table comparing the distribution of various characteristics in a treated study population vs. the subset of patients expected to have both digital image and IHC data available.

[0200]

[0167] FIG. 42 is a scatterplot displaying the relationship between the pathologist IHC score and a strongly response predictive local digital feature, color coded by different best overall response (BOR) status.

[0201]

[0168] FIG. 43 shows the significance of the top-ranked local digital features and their independent statistical significance after adjusting for a strongly response-predictive local digital feature.

[0202]

[0169] FIG. 44 shows the area under the ROC curve (AUROC) distribution of the entire set of local digital features in BOR prediction.

[0203]

[0170] FIGS. 45-46 illustrate the effects of subcellular compartments and ROIs on the prediction performance of the local digital features for BOR.

[0204]

[0171] FIGS. 47A-47B illustrate predictive power when the target protein (e.g., TROP2 or PD- Ll) expression is high.

[0205]

[0172] FIG. 48 illustrates a workflow using VFMs to extract de novo imaging features for clinical response prediction.

[0206]

[0173] FIG. 49 shows original images (e.g., WSIs), target protein positive cells, and sampled non-redundant neighborhood patches, prior to VFM based feature extraction.

[0207]

[0174] FIG. 50 shows a Uniform Manifold Approximation and Projection (UMAP) plot of embeddings and a Leiden clustering of high dimensional VFM extracted features.

[0208]

[0175] FIG. 51 is an example of the VFM extracted feature clusters frequencies.

[0209]

[0176] FIG. 52 demonstrates that a cluster that is associated with a lack of lymphocytes (immune cell desert) in the stroma correlates with non-response in target protein positive populations.

[0210]

[0177] FIG. 53 illustrates the untargeted approach, where images (e.g., WSIs) are split into non-overlapping tiles regardless of target protein (e.g., TROP2 or PD-L1) positivity to explore tumor microenvironments and their morphological clusters.

[0211] 21

[0212] NAI-5005412604vl

[0178] FIGS. 54-58 are representative image tiles from images (e.g., WSIs) of various clusters of lymphocyte clusters, target protein negative tumor cells (TCs), lymphocyte and stromal colocalized clusters, target protein positive TCs with stromal + lymphocytes, and target protein positive TCs, respectively.

[0213]

[0179] FIG. 59 provides a visualization of ten distinct clusters identified from the analysis.

[0214]

[0180] FIGS. 60A-60B show original images and cluster overlays for a non-responder and a responder, illustrating how stromal tumor-infiltrating lymphocyte (sTIL) abundance may correlate with improved clinical outcomes.

[0215]

[0181] FIG. 61 is a scatter plot showing little correlation between the digital sTIL feature and a top local target protein (e.g., TROP2 or PD-L1) optical density (OD) feature.

[0216]

[0182] FIG. 62 provides formal statistical modeling showing the significance of the digital sTIL feature as an independent predictor of clinical response.

[0217]

[0183] FIG. 63 is a scatter plot illustrating the relationship between the Morisita-Horn Index (x-axis) might provide additional prediction power of best overall response (BOR) in addition to a top local target protein (e.g., TROP2 or PD-L1) OD feature, in the target protein (e.g., TROP2 or PD-L1) high subpopulation.

[0218]

[0184] FIG. 64 summarizes the 27 input features used to construct a multivariate clinical response prediction model based on a combination of human interpretable features (HIFs), global spatial heterogeneity features, and de novo imaging features like sTIL.

[0219]

[0185] FIG. 65 illustrates the final subset of features that were selected from the 27 input feature set via cross validation.

[0220]

[0186] FIG. 66 illustrates the prediction performance of the multivariate model via nested cross validation.

[0221]

[0187] FIG. 67 depicts the inaccurate deconvolution of an example of an image depicting a biological sample treated with hematoxylin and 3, 3 -diaminobenzidine (DAB).

[0222]

[0188] FIG. 68 depicts an example of the DAB channel of an image after the application of color filters and the removal of small pixel clusters.

[0223]

[0189] FIG. 69A depicts examples of valid stains.

[0224]

[0190] FIG. 69B depicts examples of invalid, non-specific stains corresponding to artifactual patterns not related to the presence of a target protein of an IHC stain.

[0225]

[0191] FIGS. 70A-70D depict examples of the DAB channel from an image depicting a biological sample treated with an immunohistochemistry (IHC) stain prior and subsequent to artifact removal.

[0226] 22

[0227] NAI-5005412604vl

[0192] FIGS. 71 A-71B depict the results of detecting IHC staining within an expansion region determined using a modified nucleus-based radial expansion approach on various examples of IHC images.

[0228]

[0193] FIG. 72 depicts a schematic diagram illustrating a cell that has been divided into 16 equal-sized, non-overlapping partitions (or slices) for partition-based classification of biomarker (e.g., target protein) status.

[0229]

[0194] FIG. 73 depicts a flowchart illustrating an example of a process for partition-based classification of biomarker status.

[0230]

[0195] FIG. 74 depicts a flowchart illustrating an example of a process for biomarker status classification that differentiates between different subcellular compartments.

[0231]

[0196] FIG. 75A depicts an example of an image in which neighboring cells appear to share the same membrane that is positive for a target protein-sensitive stain.

[0232]

[0197] FIG. 75B depicts a schematic diagram of a cell that has been partitioned into slices for the determination of membranous staining completeness.

[0233]

[0198] FIG. 76 depicts an example of an output of a membrane segmentation model trained to infer, in an image of a biological sample treated with a target protein sensitive stain, the boundaries of both cells positive for the target protein and cells negative for the target protein.

[0234]

[0199] FIG. 77A depicts a table summarizing the incidence of true positives, true negatives, false positives, and false negatives in the output of an example of a biomarker status classifier described herein.

[0235]

[0200] FIG. 77B depicts a graph illustrating the relationship between the false positive rate and the true positive rate in the output of an example of a biomarker status classifier described herein.

[0236]

[0201] Various embodiments are described below with reference to the appended drawings. It is to be appreciated that these drawings depict only some embodiments of the disclosure and are therefore not to be considered limiting of its scope.

[0237] DETAILED DESCRIPTION OF THE INVENTION

[0238]

[0202] Pathological analysis, including the examination of images depicting biological samples, is an essential component of the medical diagnostic process. For example, in some cases, the analysis of biological samples may include the analysis of one or more microscopy images, such as whole slide images (WSI), depicting the biological samples. In some cases, the one or more images may depict biological samples, such as biological samples, that have been treated with a stain to enhance the visualization of certain features of interest present in the biological samples. For instance, the stain may be sensitive to specific proteins such that the

[0239] 23

[0240] NAI-5005412604vl images of the biological samples (e.g., biological samples) treated with the stain may be analyzed to detect the presence (or absence) of cells expressing these proteins within the biological samples.

[0241]

[0203] Protein sensitive stains may include antibodies capable of binding to specific proteins. The binding interaction between the antibodies and the proteins may trigger a visually perceptible chemical reaction (e.g., color change, fluorescence, and / or the like) that enables a differentiation between cells that express a specific protein (biomarker positive) and those that do not (biomarker negative). Examples of protein sensitive stains include immunohistochemistry (IHC) stains. Chromogenic immunohistochemistry (IHC) stains use enzymes and colored precipitates while fluorescent immunohistochemistry (IHC) stains use fluorescently labeled antibodies. The immunohistochemistry (IHC) staining process employs a primary antibody directed at a target protein of interest and an enzyme-linked secondary antibody or detection system that couples recognition of the primary antibody with deposition of a chromogenic substrate to ‘stain’ regions of target protein expression. The tissue can also be stained with a counterstain (e.g., hematoxylin) to enable visualization of individual cells and other histologic features. When the target protein is associated with response to a specific therapeutic, the immunohistochemistry (IHC) assay can be used on patient samples to identify those more likely to respond to that therapeutic. For example, if the target protein is also the target of an anticancer therapeutic, such as targeted chemotherapy, immunotherapy, or antibody drug conjugate (ADC), the immunohistochemistry (IHC) assay can identify tumor cells that express (or are “positive for”) the target protein. In some cases, where a correlation exists between the proportion of tumor cells expressing the target protein and the likelihood of responding to the anticancer therapeutic, the result of immunohistochemistry (IHC) assay can inform whether a given patient is a suitable candidate for receiving the anticancer therapeutic. In this context, an image depicting a biological sample (e.g., tissue sample, bodily fluids, and / or the like) that has been treated with an IHC stain may effectively be an image of the IHC stain. As such, as used herein, the terms “image depicting a biological sample treated with an IHC stain” and “IHC image” may be used interchangeably.

[0242]

[0204] As noted, in some cases, an image of a biological sample treated with one or more stains (e.g., protein sensitive stains) may be examined for features associated with the presence (or absence) of certain biomarkers (e.g., histological biomarkers, molecular biomarkers, and / or the like), such as those that can be used to determine one or more medically actionable patient classifications (e.g., disease diagnosis, disease stage, disease subtype, clinical response, and / or the like). In some cases, pathological analysis may be preferable to transcriptome based

[0243] 24

[0244] NAI-5005412604vl diagnostics (e.g., RNA sequence based molecular subtyping), which may not always be suitable due to its prohibitive cost and the uncertain availability of patient transcriptome data. Nevertheless, despite significant advancements in both equipment and technique, conventional pathological analysis suffers from a number of shortcomings. For instance, conventional pathological analysis, which relies on pathologist interpretation of the features present in images of biological samples, is prone to inter-reader variability. Inter-reader variability may be especially high for ambiguous features, leading to inconsistent and unreliable patient diagnoses. Conventional pathological analysis is also limited to human-interpretable features (HIFs), which are those features perceptible to pathologists. In some instances, non-human interpretable features that evade conventional pathological analysis may exhibit better correlation to the presence of certain biomarkers. As such, the inability to detect non-human interpretable features further limits the diagnostic prowess of conventional pathological analysis.

[0245]

[0205] Various example embodiments of the present disclosure overcome the limitations of conventional pathological analysis by leveraging deep learning based methodologies. The deep learning based methodologies described herein permit detailed, quantitative measurements and biomarker analyses on a continuous scale. For example, the deep learning models methodologies described herein utilize deep learning models trained for feature extraction and response prediction, which are capable of operating without the same inter-reader variability plaguing conventional pathological analysis. Instead, the deep learning models described herein may detect biomarkers within images of biological samples and predict patient response to therapeutics with greater consistency and reliability than conventional pathological analysis. Furthermore, unlike conventional pathological analysis, which is limited to the detection of human interpretable features (HIFs), various example embodiments of the deep learning based methodologies described herein are also capable discerning so-called de novo features. Unlike human interpretable features (HIFs), which are often predetermined, de novo features are those identified by deep learning models and may include non-human interpretable features that generally evade detection by conventional pathological analysis. As such, it should be appreciated that the deep learning based pathological analysis methodologies described herein improve upon existing pathological analysis solutions.

[0246]

[0206] Moreover, as described in more details below, various example embodiments of the deep learning based methodologies described herein include the extraction of a variety of novel higher order features not found in conventional pathological analysis methodologies. Instead, conventional pathological analysis methodologies, even computational based pathological

[0247] 25

[0248] NAI-5005412604vl analysis techniques, rely on primitive features with insufficient or even spurious correlation to treatment response. For example, in some cases, the presence (or absence) of a particular biomarker (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1) alone is not sufficiently informative of the likelihood of response (or non-response) to a therapeutic (e.g., anticancer therapeutic). For many indications, such as cancer, treatment response may be highly contingent upon the broader environment (e.g., tumor microenvironment) in which the treatment is administered. For instance, in some cases, treatment response may depend on distribution of a certain biomarker (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1) within the environment in which the treatment is administered. Neither conventional pathologist-based analysis nor conventional computational-based analysis support adequate characterization of this broader environment (e.g., tumor microenvironment). Accordingly, in some cases, these higher order features may include regional features spanning a range of scales or resolutions, such as tessellated features extracted from different size portions of a biological sample image, local features extracted from cell-anchored neighborhoods defined by different radii, and / or the like. In some cases, the higher order features may further include spatial features derived from one or more regional features. It should be appreciated that when used as a basis for predicting patient response to a therapeutic, these novel higher order features may further enhance the accuracy and reliability of the response prediction derived therefrom.

[0249]

[0207] According to some example embodiments, a digital pathology platform may implement a deep learning enabled pathological analysis workflow in which a deep learning based feature extraction model is applied to extract, from an image depicting a biological sample, one or more features. In some cases, the one or more features may be pixel-based features, cell-based features, tessellated features, neighborhood features, spatial features, and / or the like. In some cases, the one or more features may include human-interpretable (HIF) features, de novo features, or a combination of both. As described in more details below, in some cases, the pathological analysis workflow may further include applying a response prediction model to determine, based at least on the one or more features, a clinical response of a patient associated with the biological sample. For instance, in some cases, the response prediction model may be a univariate model trained to determine the clinical response of the patient based on individual features extracted from the image depicting the biological sample. Alternatively and / or additionally, the response prediction model may be a multivariate model trained to determine the clinical response of the patient based on multiple features extracted from the image of the biological sample. In some cases, instead of the feature extraction model being separate from the response prediction model, the digital pathology platform may include an end-to-end model

[0250] 26

[0251] NAI-5005412604vl trained to operate on an image depicting a biological sample and determine a response prediction for the patient associated with the biological sample.

[0252]

[0208] In some example embodiments, the digital pathology platform may perform the pathological analysis workflow to determine, for a patient associated with a biological sample depicted in an image, a patient classification. Examples of the patient classification include disease diagnosis, disease stage, disease subtype, clinical response, and / or the like. As noted, in some cases, the pathological analysis workflow may include the digital pathology platform applying the combination of the feature extraction model and the response prediction model or, alternatively, the end-to-end model. In instances where the patient classification is clinical response (e.g., responder or non-responder) to a particular therapy, the pathological analysis workflow may be performed to identify, within an image of a biological sample associated with the patient, one or more biomarkers exhibiting a correlation to clinical response to the therapy. For instance, in cases where the therapy is an anticancer drug targeting a particular protein (e.g., a tumor-associated antigen (TAA) such as TROP2 or PD-L1), the one or more biomarkers may include the presence (or absence) of that protein in the biological sample associated with the patient.

[0253]

[0209] In some example embodiments, the digital pathology platform may perform the pathological analysis workflow to determine, for a patient associated with a biological sample depicted in an image, a patient classification identifying the patient as a responder or a non- responder to a therapy. In some cases, the biological sample is a tissue sample (e.g., tumor tissue sample) or a fluid sample (e.g., blood) and the therapy is an anticancer drug. For instance, in some cases, the therapy may be a monotherapy targeting a protein expressed by tumor cells. Monotherapies are being evaluated for patients with locally advanced, unresectable, or metastatic solid tumors. For example, sacituzumab tirumotecan is an antibody drug conjugate (ADC) in which a monoclonal antibody targeting trophoblast cell surface protein 2 (TROP2) is linked to a belotecan-derived payload. First-in-human (FIH) studies are currently underway for sacituzumab tirumotecan as monotherapy in patients who have a locally advanced, unresectable, or metastatic solid tumor that is refractory to standard therapies. In some cases, pathological analysis of an image of tumor tissue of sufficient quantity and quality for evaluation of TROP2 expression may enable assessment of the association between clinical response to sacituzumab tirumotecan and tumor expression of TROP2 as indicated by the level of staining present in the image of the tumor tissue treated by a protein sensitive stain such as an immunohistochemistry (IHC) stain. Within the study, tumor expression of TROP2 was assessed by pathologists according to standard methods incorporating the fraction of total

[0254] 27

[0255] NAI-5005412604vl tumor cells staining at low, intermediate, and high intensity. While pathologist assessment remains the standard in clinical practice, it is inherently subjective and subject to the interreader variability described earlier. Moreover, as noted, pathologist assessment may have limitations in the types of features that can be discerned while the full diversity of biomarker information relevant to clinical performance under treatment with sacituzumab tirumotecan and other therapies will elude such techniques.

[0256]

[0210] It should be appreciated that various example embodiments of the systems, methods, and techniques described in the present disclosure are exemplary. For example, though TROP2 will be described in some detail, it will be understood that various example embodiments of the systems, methods and techniques described within the present disclosure may be applied to identify other target proteins (e.g. or PD-L1) as well as other types of biomarkers in general. Likewise, references to antibody-drug conjugates (ADCs) are also exemplary and various example embodiments of the systems, methods and techniques described within the present disclosure may be applied to other therapeutics as well as other types of treatments in general.

[0257]

[0211] FIG. 1 A is a system diagram illustrating an example of a digital pathology system 100 implementing a deep learning enabled pathological analysis workflow for analyzing images of biological sample, in accordance with some example embodiments. According to some example embodiments, the digital pathology system 100 may implement a deep learning enabled pathological analysis workflow to identify one or more biomarkers present in an image of a biological sample associated with a patient and classify the patient based on the one or more biomarkers. For example, in some cases, the patient classification may identify the patient as likely to respond to a therapeutic if one or more biomarkers associated with a target (e.g., target protein) of the therapeutic drug is present in the image of the biological sample.

[0258]

[0212] Referring again to FIG. 1A, the digital pathology system 100 may include a digital pathology platform 110, an imaging system 120, a treatment system 121, and a client device 130. As shown in FIG. 1A, in some cases, the digital pathology platform 110, the imaging system 120, and the client device 130 may be communicatively coupled via a network 140. In some cases, the imaging system 120 may include one or more imaging devices including, for example, a microscope 122, a digital camera 124, a whole slide scanner 126, and / or the like. In some cases, the treatment system may include one or more devices for administering a treatment such as a dispensing cabinet 123, an infusion pump 125, and / or the like. In some cases, the client device 130 may be a processor-based device including, for example, a smartphone, a tablet computer, a wearable apparatus, a virtual assistant, an Internet-of-Things (loT) appliance, and / or the like. In some cases, the network 140 may be a wired network and / or

[0259] 28

[0260] NAI-5005412604vl a wireless network including, for example, a wide area network (WAN), a local area network (LAN), a virtual local area network (VLAN), a public land mobile network (PLMN), the Internet, and / or the like. In some cases, the digital pathology platform 110, the imaging system 120, and / or the client device 130 may be contained within and / or operate on a same platform and / or device. For example, in some cases, the client device 130 may form a part of, include, and / or be coupled to the one or more imaging devices 120.

[0261]

[0213] In some example embodiments, the digital pathology platform 110 may implement a deep learning enabled pathological analysis workflow in which an image 101 is analyzed to determine a response prediction 105. In some cases, the image 101 may be a whole slide image (WSI) generated by the imaging system 120. In some cases, the image 101 may depict a biological sample associated with a patient. For example, in some cases, the image 101 may depict a tissue sample (e.g., tumor tissue sample) associated with the patient and the response prediction 105 may indicate a likelihood of the patient being a responder (or non-responder) for an anti-cancer therapeutic targeting a protein whose presence (or absence) is discernable as one or more features present in the image 101.

[0262]

[0214] Accordingly, in some example embodiments, the deep learning enabled pathological analysis workflow may be performed to extract the one or more features from the image 101 and determine the response prediction 105 based on the one or more features. For example, in some cases, the deep learning enabled pathological analysis workflow may include a segmentation engine 112 generating a segmentation map localizing one or more regions of interest (ROIs) present in the image 101 of the biological sample. In some cases, each pixel in the segmentation map may be assigned one or more values (e.g., a binary value, a value from a range such as [0,1], and / or the like) indicating the probability of the corresponding pixel in the image 101 depicting at least a portion of a region of interest (ROI) present in the image 101. In some cases, the segmentation map may include one or more channels, each of which corresponding to a color (e.g., grayscale, RGB, CMYK, and / or the like) forming the image 101. In some cases, the segmentation engine 112 may be further applied to generate, based at least on the one more regions of interest (ROIs) identified in the segmentation map, a cell map 103 localizing one or more cells (or subcellular compartments) present in each region of interest (ROI). For instance, in some cases, each pixel in the cell map 103 may be assigned one or more values (e.g., a binary value, a value from a range such as [0,1], and / or the like) indicating the probability of the corresponding pixel in the image 101 depicting at least a portion of a cell (or subcellular compartment) present in the image 101. In some cases, the cell map 103 may also include one or more color channels (e.g., grayscale, RGB, CMYK, and / or

[0263] 29

[0264] NAI-5005412604vl the like) corresponding to the individual colors present in the image 101. Furthermore, in some cases, the cell map 103 may identify different types of cells (or subcellular compartments) present in the image 101 (e.g., tumor cells, immune cells, stromal cells, and / or the like).

[0265]

[0215] In some example embodiments, the segmentation engine 112 may apply a region of interest (ROI) segmentation model to generate the segmentation map localizing one or more regions of interest (ROIs) present in the image 101. In some cases, the region of interest (ROI) segmentation model may be implemented in accordance with a segmentation framework to enable annotation-efficient, interpretable segmentation of regions of interest (ROIs) in images of biological samples (e.g., whole slide images (WSIs)) without model retraining. In some cases, the segmentation framework may include the identification of features that enable the localization one or more regions of interest (ROI) present in an image of a biological sample. In some cases, such features may include lower level features, such as patch embeddings, that capture fine-grained, localized details (e.g., edges, textures, color patterns, and / or the like). Alternatively and / or additionally, such features may include higher-level features, such as observable morphological patterns, with more semantic meaning than the lower level features. In some cases, the segmentation framework may identify a higher-level feature by aggregating multiple lower-level features.

[0266]

[0216] As described in more details below, the segmentation engine 112 may further apply a cell segmentation model to localize, within the one or more regions of interest (ROIs) in the segmentation map generated by the region of interest segmentation model, one or more cells used for further downstream biomarker detection and response prediction. Accordingly, the region of interest (ROI) segmentation model may be trained to localize clinically significant regions in the image, which contain morphological features that are frequently observed but also highly discriminative, for example, of target protein positive (or negative) cells within tumor tissue. For example, for oncological applications, precise delineation of tumor boundaries may be critical for targeted biomarker (e.g., target protein) assessment, prognostic stratification, and response prediction (e.g., therapy selection). To do so, the region of interest (ROI) segmentation model may be trained to learn fine-grained, cell-type-specific features to enable the precise identification of distinct cell types.

[0267]

[0217] In some example embodiments, upon applying the region of interest (ROI) segmentation model to generate the segmentation map localizing the one or more regions of interest (ROIs) present in the image 101, the segmentation engine 112 may further apply a cell segmentation model to localize and identify (e.g., the type) of one or more cells present in each region of interest. For example, in some cases, the cell segmentation model may operate on the

[0268] 30

[0269] NAI-5005412604vl segmentation map generated by the region of interest segmentation model to generate the cell map 103 in which each pixel is assigned one or more values (e.g., a binary value, a value from a range such as [0,1], and / or the like) indicating the probability of the corresponding pixel in the image 101 depicting at least a portion of a cell (or subcellular compartment) or, alternatively, different types of cells (or subcellular compartments) present in the image 101 (e.g., tumor cells, immune cells, stromal cells, and / or the like).

[0270]

[0218] In some cases, the cell segmentation model may be trained using an adaptative cell detection framework in which a pretrained model finetuned in a few shot manner with a limited set of expert annotated images (e.g., with expert identification of cells) serving as ground truth labels. For example, in some cases, the cell segmentation model may be a foundation model f that has been pretrained on a large corpus of unlabeled or weakly labeled images (e.g., whole slide images (WSIs)) of biological samples. The pretraining of the foundation model fe provides a strong initialization for downstream cell detection tasks. In some cases, the pretrained foundation model fe then undergoes incremental updates (or finetuning), for example, at successive timestepsf, using a small quantity of annotated samples t For instance, in some cases, at each timestep the pretrained foundation model fe implementing the cell segmentation model may be updated by at least reducing (or minimizing) the loss function Lt with the regularization term

[0271] Lt(0t)=Ldet(9t>Dt) + A-5?(0t,0t— i) wherein L-det denotes the detection loss (e.g., focal loss, generalized intersection over union (loU) loss) quantifying the failure to identify pixels depicting cells or correctly identify celltype, denotes the current model parameters (e.g., weights, biases, and / or the like), and is a regularization term (e.g., elastic weight consolidation, knowledge distillation, or L2 penalty) that penalizes excessive deviation from the previous model parameters It should be appreciated that the inclusion of the regularization term may preserve knowledge from earlier training steps and prevent catastrophic forgetting. In some cases, the small number of annotated samples used for few-shot finetuning of the pretrained foundation model fe may be strategically selected based on model uncertainty or morphological diversity to maximize information gain with each incremental update of the pretrained foundation model fe.

[0272]

[0219] Unlike traditional approaches that require full model retraining with each batch of new data, this aforementioned adaptive cell detection framework may achieve efficient adaptation

[0273] 31

[0274] NAI-5005412604vl with minimal updates. For example, in some cases, the finetuning of the pretrained foundation model fe, which requires a small number of strategically selected annotated samples may include updating a subset of layers in the foundation model fe or leveraging memory-aware fine-tuning strategies. Doing so may reduce both the computational overhead and annotation cost of updating the foundation model fe for better generalization across tissue types, staining conditions, and cellular phenotypes. As a result, the performance of the cell segmentation model, including cell detection accuracy and precision may be continuously improved with minimal supervision and resources.

[0275]

[0220] In some example embodiments, the deep learning enabled pathological analysis workflow may include applying a pathological analysis engine 114 to determine, based at least on the cell map 103, one or more features present in the image 101. In some cases, the one or more features may include human-interpretable (HIF) features, de novo features, or a combination of both. In some cases, the one or more features may include primitive features and higher order features derived from the primitive features. In some cases, the pathological analysis engine 114 may determine, based at least on the one or more features, the response prediction 105. As described in more details below, the one or more features may be extracted from tiles (e.g., overlapping or nonoverlapping partitions) within the image 101. In some cases, the pathological analysis engine 114 may generate, based at least on the one or more features, an image-level representation (e.g., embedding) and the response prediction 105 may be determined based on the image-level representation (e.g., embedding) instead of the individual tiles. In some cases, the image-level presentation (e.g., embedding), which aggregates features across tiles in the image 101, may be generated in geospatially aware manner that preserves the spatial relationships present amongst neighboring tiles within the image 101.

[0276]

[0221] FIG. IB depicts a block diagram illustrating an example of the pathological analysis engine 114 that includes a feature extraction model 152 and a response prediction model 154. In this example of the pathological analysis engine 114, the feature extraction model 152 may be a deep learning model trained to extract, from the cell map 103 localizing one or more cells present in the image 101 (or in one or more regions of interest (ROI) identified within the image 101), one or more features for inclusion in a feature set 153. For example, in some cases, the feature set 153 may include one or more primitive features extracted from the image 101. Alternatively and / or additionally, the feature set 153 may include one or more higher order features derived from the one or more primitive features. Examples of primitive features, some of which are described in Table 2 below, may include features indicative of the presence (or

[0277] 32

[0278] NAI-5005412604vl absence) of a certain biomarker (e.g., a tumor-associated antigen (TAA) such as TROP2 or PD- Ll), as indicated by the intensity of a biomarker sensitive stain (e.g., immunohistochemistry (IHC) stain) used to treat the image 101. In some cases, the presence (or absence) of the biomarker may be further dictated by the location of the biomarker sensitive stain. For example, a cell may identified as positive (or negative) for the biomarker if the biomarker sensitive stain is detected in one or more specific subcellular compartments (e.g., nucleus, cytoplasm, membrane, and / or the like). As described in more details below, examples of higher order features derived from primitive features may include regional features spanning a range of scales or resolutions, such as tessellated features extracted from different size portions of the image 101, local features extracted from cell-anchored neighborhoods defined by different radii, and / or the like. In some cases, the higher order features may also include spatial features derived from one or more regional features. For instance, in instances where the biological sample depicted in the image 101 is treated with a stain that is sensitive to the target protein (e.g., a tumor-associated antigen (TAA) such as TROP2 or PD-L1) of a therapeutic, the primitive features may quantify the intensity and completeness of the staining on the cells (e.g., tumor cells) identified as positive for expressing the target protein. In some cases, the higher order features may include the distribution (e.g., density, proportion, and / or the like) of the cells (e.g., tumor cells) expressing the target protein across the image 101 as a whole or certain portions thereof (e.g., subcellular compartments).

[0279]

[0222] Referring again to FIG. IB, in some example embodiments, the feature extraction model 152 may be applied to extract human interpretable features (HIFs). For example, with the human interpretable feature (HIF) approach, the feature extraction model 152 may first be applied to extract, from the image 101, one or more primitive features. As noted, in some cases, the one or more primitive features may be associated with the presence (or absence) of one or more biomarkers. Where the biological sample depicted in the image 101 was treated with a protein sensitive stain (e.g., immunohistochemistry (IHC) stain), for example, the one or more primitive features may indicate the presence (or absence) of a target protein (e.g., a tumor- associated antigen (TAA) such as TROP 2). Accordingly, in some cases, the extraction of human interpretable features (HIFs) may include extracting, based at least on the cell map 103 generated by the segmentation engine 112, the one or more primitive features. As noted, in some cases, the cell map 103 may localize and classify the different types of cells (e.g., tumor cells, immune cells, stromal cells, and / or the like) present in the image 101 or in various regions of interest (ROIs) therein. In some cases, the one or more primitive features may be identified by at least classifying (e.g., by a biomarker classification model) the cells identified by the cell

[0280] 33

[0281] NAI-5005412604vl map 103 as exhibiting (or failing to exhibit) the one or more biomarkers (e.g., a tumor- associated antigen (TAA) such as TROP2 or PD-L1). In some cases, upon identifying the one or more primitive features, the feature extraction model 152 may derive one or more higher order human interpretable features (HIFs) from the one or more primitive features. For instance, in some cases, a set of predetermined human interpretable features (HIFs) may be identified based on the one or more primitive features.

[0282]

[0223] In some cases, the human interpretable features (HIFs) may include pixel -based features, cell-based features, and tessellate features. Some human interpretable features (HIF) may be regional features associated with portions of the image 101 or cell-anchored neighborhoods within the image 101. Those regional features may vary in size such as different radii (e.g., 8- / m, 15- / m, 25-Zm, 50- / m, and / or the like), tessellations (e.g., 150- *m, 300- / m, 450- / *m, 600-m, and / or the like), and / or the like. In some cases, a human interpretable feature (HIF) may be further categorized by the type of anchor cell (e.g., tumor cell, immune cell, stromal cell, and / or the like), the measurement method (e.g., optical density (OD), proportion, and / or the like), subcellular compartment being evaluated (e.g., membrane, cytoplasm, total cell, and / or the like), image (or slide) level region being evaluated (e.g., total tissue, invasive tumor region, in situ tumor region, and / or the like), percentile value used to generate summary statistics (e.g., image level summary statistics), and / or the like. In some cases, the feature extraction model 152 may perform correlation analyses to reduce redundant features or those features whose marginal distribution lacks biological variations in the sample set. In some cases, this reduction may reduce the quantity of features present in the feature set 153 (e.g., from >2174 features to approximately 35 features in one instance) used by the response prediction model 154 to determine the response prediction 105.

[0283]

[0224] In some example embodiments, one example of human interpretable features (HIFs) generated for inclusion in the feature set 153 are pixel- or cell-based spatial features. In some cases, a spatial feature may be determined for a cell-anchored neighborhood, such as a neighborhood of a certain radius around a cell. In some cases, the cell anchoring the neighborhood may be a tumor cell that is either positive or negative for a certain biomarker (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1). In some cases, the spatial feature for the cell-anchored neighborhood may be determined based on the level of staining (e.g., signal from the immunohistochemistry (IHC) stain) present in the cell-anchored neighborhood. In some cases, the level of staining may correspond to stain intensity normalized either by only stained (or biomarker positive) pixels or by both stained (or biomarker positive)

[0284] 34

[0285] NAI-5005412604vl and unstained (or biomarker negative) pixels. In some cases, the level of staining present in the cell-anchored neighborhood may be determined based on cell-based measurements from cells within the radius of the neighborhood and specific subcellular compartments of interest (e.g., membrane, cytoplasm, and / or the like. Alternatively, the level of staining present in the cell- anchored neighborhood may correspond to a pixel-wise summation of all cells within the radius of the neighborhood and, in some cases, pixels within a certain size (e.g., 2-m) expansion from all cells (e.g., tumor cells) within the neighborhood. In some cases, the feature set 153 may include, for the image 101 or portions thereof, one or more summary statistics of the various spatial features computed for various quantiles (e.g., 10thquantile, 50thquantile, 90thquantile, and / or the like) and in quantile increments.

[0286]

[0225] Alternatively, instead of human interpretable features (HIFs), the feature extraction model 152 may be applied to extract de novo features. In this context, de novo features may be identified by the feature extraction model 152 without any predetermination, as is the case with human interpretable features (HIFs). In many instances, de novo features may include nonhuman interpretable features that evade detection by conventional pathological analysis. In some cases, de novo features may be identified by the feature extraction model 152 (e.g., a vision foundation model (VFM)) trained on sample images with known status (or ground truth labels) as being positive (or negative) for certain biomarkers (e.g., tumor associated antigen (TAA) such as TROP2 or PD-L1). For example, in some cases, the feature extraction model

[0287] 152 may be a pretrained vision foundation model (VFM) that is then fine-tuned for specific indications. In the case of tumor tissue samples, the feature extraction model 152 may undergo few shot learning in which a limited number of sample images exhibiting the phenotypes of specific tumor types are used to finetune the feature extraction model 152. In some cases, the de novo features extracted by the feature extraction model 152 may be subjected to dimensionality reduction in order to reduce the number of features included in the feature set

[0288] 153 used by the response prediction model 154 to determine the response prediction 105. For instance, in some cases, the de novo features extracted by the feature extraction model 152 may be clustered, with each resulting cluster corresponding to features with sufficient covariance and correlation to be added to the feature set 153 collectively as a single feature instead of multiple individual features. In some cases, feature clusters may undergo pathologist verification, for example, to ascertain biological significance and clinical utility. In some cases, the feature set 153 may include feature clusters with verified biological significance and clinical utility.

[0289] 35

[0290] NAI-5005412604vl

[0226] In some example embodiments, the feature extraction model 152 may implement a hybrid approach to feature extraction. With the hybrid approach, the feature extraction model 152 may first classify the cells present in the image 101. For example, the feature extraction model 152 may operate on the cell map 103 to extract one or more human interpretable features (HIF s) including, for example, primitive features, higher order features, and / or the like. In some cases, the identification of human interpretable features (HIFs) may include identifying cells (e.g., tumor cells) positive for one or more biomarkers (e.g., tumor associated antigens (TAAs) such as TROP2 or PD-L1). In some cases, the hybrid feature extraction approach may further include applying a vision foundation model to one or more local neighborhoods anchored around cells positive for the one or more biomarkers. For example, in some cases, the vision foundation model may be applied to extract de novo features from one or more local neighborhoods within certain radii around cells (e.g., tumor cells) positive for the one or more biomarkers (e.g., tumor associated antigens (TAAs) such as TROP2 or PD-L1).

[0291]

[0227] Referring again to FIG. IB, in the example of the pathological analysis engine 114 shown in FIG. IB, the response prediction model 154 may determine, based at least on the feature set 153, the response prediction 105. As noted, in some cases, the response prediction 105 may indicate the likelihood of the patient associated with the biological sample depicted in the image 101 being a responder (or non-responder) for a particular therapeutic. In some cases, the response prediction model 154 may be a univariate model trained to determine, based on each individual feature from the feature set 153, the response prediction 105. Alternatively and / or additionally, the response prediction model 154 may be a multivariate model trained to determine the response prediction 105 based on multiple features from the feature set 153. It should be appreciated that the response prediction model 154 may be implemented using a variety of machine learning architectures including, for example, a random forest model, a decision tree, a gradient-boosted decision tree, an elastic net, and / or the like.

[0292]

[0228] In some example embodiments, the feature extraction model 152 may be applied to extract, from each tile within the image 101, one or more features for inclusion in the feature set 153. For example, in some cases, the image 101 may be partitioned into a set of tiles (e.g., 100 / rm X 100 / zm tiles). In some cases, the set of tiles may include overlapping tiles or, alternatively, non-overlapping tiles. In some cases, the feature extraction model 152 may extract, from each tilext in the set of tiles (e.g.,xteX), one or more features for inclusion in the feature set 153. For instance, in some cases, the feature extraction model 152 may generate, for each tilext in the set of tiles X (e.g.,xteX), a feature vector including the one or more

[0293] 36

[0294] NAI-5005412604vl features extracted therefrom. In some cases, the feature extraction model 152 may reduce the dimensionality of the feature set 153 by at least clustering the feature vector extracted from each tile and consolidating different features in each cluster into individual features at least because the features in a cluster exhibit sufficient covariance and correlation. The resulting tilelevel representations (e.g., embeddings) of the features present in the image 101 may include a fewer quantity of features than is originally present in the feature vectors extracted from each tile in the image 101. In some cases, the response prediction model 154 may be applied to determine, based at least on the tile-level feature representations in the feature set 153, the response prediction 105.

[0295]

[0229] In some cases, instead of the feature vectors of individual tiles, the feature extraction model 152 may determine an image-level representation (e.g., embedding) of the features present in the image 101 and the response prediction model 154 may operate on this imagelevel representation when determining the response prediction 105. However, conventional approaches to generating the image-level representation (e.g., embedding) of the features present in the image 101 treats each tilext in the set of tiles (e.g.,xtGX) independently. For example, where f denotes a pretrained encoder generating the image-level representation (e.g., embedding), the conventional approach may compute attention weightsat solely based on the tile-level representation hi=f(xi of each tilext in the set of tiles (e.g.,xiGX) The resulting image-level representation (e.g., embedding) of the features present in the image 101 may be computed asz = “A. However, the image-level representation (e.g., embedding) computed in this manner ignores the spatial relationships present amongst neighboring tiles in the image 101. Those spatial relationships, which capture tumor morphology and microenvironment context, may be crucial for the accuracy of the response prediction 105 determined based on the image-level representation (e.g., embedding) of the image 101. As noted, determining the response prediction 105 based on higher order features (e.g., regional features of different scales and resolutions, spatial features derived from regional features, and / or the like) extracted from the image 101 may increase the accuracy of the response prediction 105 at least because these higher order features better characterize the tumor microenvironment in which a treatment (e.g., antibody-drug conjugate (ADC), alkylating agents, antimetabolites, enzyme inhibitors, anti-angiogenesis agents, microtubule disruptors, immunotherapy agents, and / or the like). An image-level representation (e.g., embedding) of the features present in the image 101, including primitive features and higher order features, that preserves spatial relationships may similarly increase the accuracy of the response prediction 105.

[0296] 37

[0297] NAI-5005412604vl

[0230] In some cases, multiple tile-level representations, which are embeddings of tile-level features, may be aggregated and summarized at the image level (or slide level). In some cases, these image-level representations may be further combined with the human-interpretable features (HIFs) used, for example, by the response prediction model 154 to determine the response prediction 105. In some cases, the tile-level representations may be summarized at the image level by at least selecting, for example, based on a tile attention heatmap, a subset of tiles correspond to the most important histological features (or the histological features with the highest correlation to clinical response). Those tiles may then be summarized at the image level to form the image-level representation. For example, the attention heatmap may indicate stromal tumor-infiltrating lymphocytes (sTILs) are response predictive. During subsequent analysis, the abundance of sTILs may be determined directly at the image level and combine with other image level features rather than extracted from individual tiles. Alternatively, tilelevel representations may be clustered and the composition of each cluster may be determined at the image level. In some cases, the percentage cluster frequency may be combined with HIFs at the image level.

[0298]

[0231] In some cases, an image-level representation (e.g., embedding) of the features present in the image 101 may be generated by the use of a spatially contextual attention mechanism in which the attention for each tilexi in the set of tiles X (e.g.,xteX) isinfluenced by one or more of its neighbors For example, in some cases, the spatially contextual attention mechanism may include applying, to each tilext in the set of tiles X (e.g., a geospatially adjusted attention weightat defined as follows: wherein denotes the dot product similarity between the feature embedding hi of the tilext and the feature embedding hj of one or more neighboring tilesXJ, andSii denotes a spatial proximity weight.

[0299]

[0232] In some cases, the spatial proximity weightSii may be defined as follows:

[0300]

[0233] In some cases, the image-level representation (e.g., embedding) of the image 101 may be defined as follows:

[0301] 38

[0302] NAI-5005412604vl

[0234] It should be appreciated that the foregoing formulation enables the feature extraction model 152 to emphasize, when generating the image-level representation (e.g., embedding) of the image 101, tiles that are not only locally discriminative but also contextually supported by their spatially adjacent regions. Moreover, this geographically aware tile aggregation approach may maintain computational efficiency without requiring additional supervision (e.g., expert annotations) beyond image-level labels, rendering it well-suited for weakly supervised settings while still improving robustness in the presence of histological heterogeneity.

[0303]

[0235] FIG. 1C is a block diagram illustrating another example of the pathological analysis engine 114 in which feature extraction and response prediction are performed by a single end- to-end model 156 instead of the feature extraction model 152 and the response prediction model 154 shown in FIG. IB. In some cases, the end-to-end model 156 may combine the functionalities of the feature extraction model 152 and the response prediction model 154. For example, in some cases, the end-to-end model 156 may implement an attention based multiple instance learning (MIL) framework. Accordingly, in some cases, the end-to-end model 156 may directly identify features in the image 101 with sufficient correlation to treatment response (e.g., response to a therapeutic) and apply these features to the generation of the response prediction 105.

[0304]

[0236] FIG. ID depicts a block diagram illustrating an example of the deep learning enabled pathological analysis workflow implemented by the digital pathology platform 110. In some cases, the deep learning enabled pathological analysis workflow may be performed to determine, based at least on features extracted from the image 101, the response prediction 105 for a patient associated with the biological sample depicted in the image 101. In some cases, the image 101 may be generated by the imaging system 120. In some cases, the biological sample (e.g., tissue sample, bodily fluids, and / or the like) depicted in the image 101 depicting may be treated with a stain sensitive to one or more target proteins of a therapeutic. For a target protein such as a tumor associated antigen (e.g., TROP2 or PD-L1), for example, the biological sample in the image 101 may be treated using 3, 3 -diaminobenzidine (DAB). As shown in FIG. ID, the digital pathology platform 110 may operate on the image 101, including by performing the deep learning enabled pathological analysis workflow. For example, in some cases, the segmentation engine 112 may segment the image 101 to localize one or more regions of interest (ROI) and one or more cells (e.g., cells positive or negative for certain biomarkers) present within the image 101. In some cases, the feature extraction model 152 may extract one or more features, including primitive features and higher order features derived from primitive features before such features are used by the response prediction model 154 to generate the response

[0305] 39

[0306] NAI-5005412604vl prediction 105. For the tumor associated antigen (TAA) example, the digital pathology platform 110 may operate on the image 101 to quantify TAA expression for predicting a clinical response to a therapeutic targeting the TAA (e.g., antibody drug conjugates (ADCs), alkylating agents, antimetabolites, enzyme inhibitors, anti-angiogenesis agents, microtubule disruptors, immunotherapy agents, and / or the like). As described in more details below, the example of the deep learning enabled pathological analysis workflow shown in FIG. ID utilizes the feature extraction model 152 for feature extraction and the response prediction model 154 for response prediction. In some cases, the feature extraction model 152, by leveraging deep learning techniques, may be capable of extracting relevant features directly from the image 101 to characterize various aspects of the tumor microenvironment. In some cases, these features may be de novo features that would otherwise evade detection with conventional pathological analysis methodologies. In some cases, these features may serve as input to the response prediction model 154 such that the response prediction 105 may be generated based thereupon. However, as noted, it is also possible for the deep learning enabled pathological analysis workflow to utilize the end-to-end model 156 (shown in FIG. 1 C) for both feature extraction and response prediction.

[0307]

[0237] FIG. IE is a flowchart illustrating an example of a process 170 for deep learning enabled pathological analysis for prediction of treatment response, in accordance with some example embodiments. In some example embodiments, the process 170 may be performed by the digital pathology platform 170 including, for example, the segmentation engine 112 and the pathological analysis engine 114. In some cases, the process 170 may be performed in order to analyze the biological sample (e.g., tissue sample, bodily fluids, and / or the like) depicted in the image 101, which may be treated (e.g., by a protein sensitive stain such as an immunohistochemistry (IHC) stain) to enable visualization of one or more biomarkers (e.g., tumor associated antigen (TAA)) present in the biological sample (e.g., tissue sample, bodily fluids, and / or the like) depicted in the image 101. In some cases, the process 170 may include the extraction of the feature set 153, which may include primitive features and higher order features derived from the primitive features. In some cases, the process 170, which leverages deep learning models, may be capable of identifying de novo features that evades conventional pathological analysis methodologies. Moreover, in some cases, the higher order features may include novel regional and spatial features characterizing the broader environment (e.g., tumor microenvironment) in which the therapeutic is administered. Such higher order features, which are not supported by conventional pathological analysis methodologies, enable more accurate and reliable response predictions at least because higher order features characterizing the

[0308] 40

[0309] NAI-5005412604vl broader environment encountered by the therapeutic exhibit better correlation to the likelihood of response (or non-response) to the therapeutic. In some cases, the process 170 (or at least a portion thereof) may be performed in order to identify or select a patient for a particular treatment, such as a cancer treatment. In some cases, a patient selected for the treatment may be administered the treatment, for example, using the treatment system 121.

[0310]

[0238] At 172, an image depicting a biological sample is received. In some example embodiments, the image may be received from an imaging system, such as the imaging system 120 shown in FIG. 1A including the microscope 122, the digital camera 124, and the whole slide scanner 126. In some cases, the biological sample depicted in the image may be treated with a stain to enable visualization of one or more biomarkers present in the biological sample. For example, in instances where the biomarker is a protein in a cancer. In some cases, the protein may be present on a cell surface or a cell interior. In some cases, the protein may be targeted by a therapeutic (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1).

[0311]

[0239] In some cases, where the biomarker is a protein, the biological sample may be treated with a protein sensitive stain (e.g., immunohistochemistry (IHC) stain). In some cases, the biological sample depicted in the image may also be treated with a stain to enable visualization of one or more cells (or subcellular compartments) present in the image. For instance, in some cases, the biological sample in the image may be treated with a hematoxylin stain to enable visualization of cell nuclei, which are stained blue by the hematoxylin dye.

[0312]

[0240] At 174, one or more segmentation models are applied to segment the image. In some example embodiments, the one or more segmentation models may be applied to localize, within the image, one or more regions of interest (ROI). For example, in some cases, the one or more segmentation models may include a region of interest (ROI) segmentation model trained to segment the image by at least generating a segmentation map in which each pixel is assigned a value (or label) indicative of the likelihood of the pixel depicting a region of interest (ROI). In some cases, the segmentation map may include one or more channels, each of which corresponding to a color (e.g., grayscale, RGB, CMYK, and / or the like) present in the image. In instances where the biological sample is a tissue sample (e.g., tumor tissue sample), the one or more regions of interest (ROI) may include tumor regions exclusively or, alternatively, a combination of tumor regions as well as relevant non-tumor regions, Alternatively and / or additionally, the one or more segmentation models may be applied to localize and, in some cases, classify, one or more cells present in the image (or within each regions of interest (ROI)). For instance, in some cases, the one or more segmentation models may include a cell segmentation model trained to segment the image (or each region of interest (ROI) in the

[0313] 41

[0314] NAI-5005412604vl image) by at least generating a cell map (e.g., the cell map 103 shown in FIGS. 1A-1C) in which each pixel is assigned a value (or label) indicative of the likelihood of the pixel depicting a cell (or a subcellular compartment). In some cases, each pixel in the cell map may be assigned a value indicative of the type of cell (or type of subcellular compartment) depicted by the pixel. Moreover, in some cases, the cell map may include one or more channels, each of which corresponding to a color (e.g., grayscale, RGB, CMYK, and / or the like).

[0315]

[0241] At 176, a feature extraction model is applied to determine a feature set containing one or more features extracted from the image. In some example embodiments, the feature extraction model, such as the feature extraction model 152 shown in FIG. IB, may be applied to determine, based at least on the cell map, one or more primitive features present in the image. Moreover, in some cases, the feature extraction model 152 may derive at least one higher order feature based on the one or more primitive features present in the image. In some cases, the feature extraction model may be trained to extract, from the image, one or more predetermined human interpretable features (HIFs). Alternatively, the feature extraction model may be trained to identify de novo features, which may include non-human interpretable features that would generally evade detection by pathologists. In some cases, the feature extraction model may also implement a hybrid approach that combines aspects of human interpretable feature (HIF) approach and de novo feature discovery approach. For example, in some cases, the feature extraction model implementing the hybrid approach may first extract human interpretable features (HIFs) that enable the identification of cells that are positive (or negative) for one or more biomarkers. Thereafter, the feature extraction model may be applied to extract de novo features from neighborhoods around each cell identified as positive for the one or more biomarkers.

[0316]

[0242] In some example embodiments, the one or more primitive features may include the location and the types of cells present in the image (or in each region of interest (ROI) identified within the interest). In some cases, the one or more primitive features may include the location and types of cells that are positive (or negative) for the one or more biomarkers, such as the target protein of the therapeutic. In some cases, the determination of the one or more primitive features may include the identification of biomarker positive (e.g., target protein positive) and biomarker negative (e.g., target protein negative) cells. Examples of processes for classifying a cell as positive for a biomarker (e.g., a target protein such as TROP2 or PD-L1) or negative for the biomarker is shown in FIG. 9. As described in more detail below, whether a cell is positive (or negative) for a biomarker (e.g., a target protein such as TROP2 or PD-L1) may be determined based on the intensity of a stain indicating the presence of the biomarker and one

[0317] 42

[0318] NAI-5005412604vl or more subcellular compartments (e.g., membrane, cytoplasm, nucleus, and / or the like) exhibiting the staining.

[0319]

[0243] As noted, in some cases, one or more higher order features may be derived from the one or more primitive features. Examples of higher order features include regional features spanning a range of scales or resolutions, such as tessellated features extracted from different size portions of the image, local features extracted from cell-anchored neighborhoods defined by different radii, and / or the like. In some cases, the higher order features included in the feature set may further include spatial features, which characterize the distribution of primitive or other higher order features across the image (or portions thereof). In some cases, the higher order features included in the feature set may include summary statistics computed for various scales, resolutions, quantiles, and / or the like.

[0320]

[0244] At 178, a response prediction model is applied to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample. In some example embodiments, the response prediction model (e.g., the response prediction model 154 shown in FIG. IB) may be a univariate model trained determine a response prediction based on individual features extracted from the image. Alternatively, the response prediction model may be a multivariate model capable of determining the response prediction based on a combination of multiple features extracted from the image. In some cases, the response prediction may indicate a likelihood of the patient associated with the biological sample being a responder (or non-responder) for a therapeutic. In the case of the therapeutic targeting a particular protein expressed by tumor cells, the response prediction may be generated based on at least some features characterizing the environment (e.g., tumor microenvironment) in which the therapeutic is administered. Accordingly, it should be appreciated that the response prediction generated by various example embodiments of the deep learning enabled pathological analysis workflow described herein may be more accurate and reliable than those generated by conventional pathological analysis methodologies that overlook or are incapable of accounting for such features.

[0321]

[0245] At 180, the patient is identified as eligible for receiving the therapeutic based at least on the response prediction. In some example embodiments, the response prediction of the response prediction model may indicate the likelihood of the patient responding to (or benefitting from) the therapeutic (e.g., an antibody drug conjugate (ADC), alkylating agents, antimetabolites, enzyme inhibitors, anti-angiogenesis agents, microtubule disruptors, immunotherapy agents, and / or the like). In some cases, the patient may be identified as eligible for receiving the therapeutic if the likelihood of the patient responding to (or benefitting from)

[0322] 43

[0323] NAI-5005412604vl the therapeutic (e.g., an ADC such as sacituzumab tirumotecan) satisfies one or more thresholds. Conversely, the patient may be identified as ineligible for receiving the therapeutic if the likelihood of the patient responding to (or benefitting from) the therapeutic (e.g., an ADC, alkylating agent, antimetabolite, enzyme inhibitor, anti-angiogenesis agent, microtubule disruptor, immunotherapy agent, and / or the like) fails to satisfy the one or more thresholds. In some cases, where the patient is identified or selected as eligible for receiving the therapeutic (e.g., an ADC such as sacituzumab tirumotecan), a therapeutically effective quantity of the therapeutic may be administered to the patient. Alternatively, where the patient is identified as ineligible for receiving the therapeutic, a different therapeutic may be selected for the patient and a therapeutically effective quantity of the different therapeutic may be administered instead. In some cases, the administering of the therapeutic may be performed using a treatment system, such as the treatment system 121 (e.g., including the dispensing cabinet 123, the infusion pump 125, and / or the like) shown in FIG. 1 A.

[0324]

[0246] In some cases, the therapeutic may be used to treat the patient for a cancer including, for example, breast cancer, non-small cell lung cancer, ovarian cancer, cervical cancer, endometrial cancer, gastric cancer, prostate cancer, esophageal cancer, urothelial cancer, and / or the like. In some cases, the therapeutic may be an antibody drug conjugate (ADC). For example, in some cases, the therapeutic may be sacituzumab tirumotecan.

[0325]

[0247] In some cases, where the therapeutic is an antibody drug conjugate (ADC), such as sacituzumab tirumotecan, the therapeutically effective quantity of the therapeutic may be administered intravenously, intradermally, subcutaneously, intramuscularly, or intraperitoneally. In some cases, the therapeutically effective quantity of the therapeutic may be administered in accordance to a dosing schedule such as once a week, once every 2 weeks, once every 3 weeks, once every 4 weeks, or once every 5 weeks. In some cases, the therapeutically effective quantity of the therapeutic may include a dose of 1 to 30 mg / kg, 1 to 20 mg / kg, 2-12 mg / kg, 2-5 mg / kg, or 4-7 mg / kg. In some cases, the therapeutically effective quantity of the therapeutic may include a dose of 2 mg / kg, 2.5 mg / kg, 3 mg / kg, 3.5 mg / kg, 4 mg / kg, 4.5 mg / kg, 5 mg / kg, 5.5 mg / kg, 6 mg / kg, 6.5 mg / kg, 7 mg / kg, 7.5 mg / kg, 8 mg / kg, 8.5 mg / kg, 9 mg / kg, 9.5 mg / kg, 10 mg / kg, 11 mg / kg, or 12 mg / kg. It should be appreciated that other types of therapeutics may be administered in addition to or instead of the ADC, such as alkylating agents, antimetabolites, enzyme inhibitors, anti-angiogenesis agents, microtubule disruptors, immunotherapy agents, and / or the like.

[0326]

[0248] FIG. IF is a block diagram illustrating an example of a process 180 for extracting features from the image 101. In some cases, the process 180 shown in FIG. IF may implement

[0327] 44

[0328] NAI-5005412604vl at least a portion of operations 172 (for segmentation) and 174 (for feature extraction) of the process 170 shown in FIG. IE. In this illustration, parallel branches indicate that multiple options are available for that component of the pipeline, and the symbols indicate the pipeline component that is parallelizable for increased computational efficiency. The upper section (above the dashed line) shows that the feature extraction process includes the image 101 undergoing preprocessing prior to feature extraction. In some cases, the preprocessing of the image 101 may include region of interest (ROI) segmentation 182 (to localize regions of interest (ROI) in the image 101) and cell segmentation 186 (to localize and classify cells within each region of interest (ROI)), both of which may be performed by the segmentation engine 112. In some cases, in addition to the segmentation of the image 101, the preprocessing of the image 101 may also include channel standardization 184, an operation in which the different color channels present in the image 101 (e.g., for hematoxylin staining, H4C staining, and / or the like) are standardized for consistency across samples. Once the preprocessing of the image 101 is complete, feature extraction 188 may be performed, for example, by the feature extraction model 152, to identify relevant features present in the image 101.

[0329]

[0249] The lower section of the flowchart (below the dashed line) shows additional details of each of the aforementioned operations, and some of these will be progressively described in greater detail below. For example, in some cases, the image 101 may be annotated by a pathologist to identify regions of interest (ROIs) for further analysis. This may include pathologist-directed selection and classification of prototype tiles for analysis by a downstream pre-trained foundation model and / or Otsu threshold (or similar thresholding techniques for image binarization) to enhance segmentation accuracy. Dual-Path Fusion Module (DP-FM) in FIG. IF is a component of the Dual-Path U-Net Segmentation Diffusion segmentation model that may be used for segmenting the image 101. In FIG. IF, IHC ROI may follow where regions of interest (ROIs) are identified for downstream analysis based on the presence of immunohistochemistry (IHC) signal for the protein of interest. Channel decomposition may include tools for further processing. In this context, hematoxylin stains (or in some cases (4',6- diamidino-2-phenylindole) (DAPI) may be used to isolate the stained nuclei from the rest of the image while 3,3-diaminobenzidine (DAB) may be the specific type of immunohistochemistry (IHC) stain used to visualize the protein of interest or target protein (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1). Specifically, both the standard hematoxylin / DAB decomposition (stain decomposition matrix estimated from a previously trained IHC model for PD-L1 expression assessment) and deep learning-inferred multiplex immunofluorescence (DeepLIIF) may be implemented to extract both cellular signal

[0330] 45

[0331] NAI-5005412604vl (or subcellular compartment signal) and target protein signal (e.g., 3,3-diaminobenzidine (DAB) signal) channels from the image 101. Following stain decomposition, the nuclear channel signal may be used to perform cell detection and cell type classification, and the 3,3- diaminobenzidine (DAB) channel signal may be used to evaluate target protein (e.g., TROP2 or PD-L1) expression.

[0332]

[0250] In some example embodiments, where the image 101 is derived from a biological sample treated with hematoxylin (e.g., to visualize subcellular compartments such as nucleus) or an IHC stain (e.g., to visualize biomarkers such as target proteins), the image 101 may include different channels, each of which corresponding to a color of the hematoxylin stain or the IHC stain. Accordingly, in some cases, the preprocessing of the image 101 may include deconvoluting the image 101 to separate the different channels associated with the hematoxylin stain or the IHC stain. For example, the image 101 derived from hematoxylin stain may give rise to one color channel (e.g., blue or purple) associated with the hematoxylin dye. In some cases, the image 101 derived from an IHC stain may add additional color channels. For instance, the image 101 may be derived from an IHC stain with a 3,3-diaminobenzidine (DAB) chromogen. This type of IHC stain may give rise to an additional color channel (e.g., brown).

[0333]

[0251] Unlike the deconvolution of standard RGB color channels, deconvoluting the image 101 may generate channels representative of the intensity of each stain (e.g., hematoxylin, DAPI, DAB, and / or the like), thereby allowing for the isolation and quantification of the individual stains. In some cases, the deconvolution of the image 101 may be performed based on the known color profiles (e.g., optical absorbance) of the tissue stains from which the image 101 is derived (e.g., hematoxylin, DAPI, DAB, and / or the like). For example, in some cases, the deconvolution of the image 101 may include determining, based on the corresponding color profiles (e.g., optical absorbances), the contribution of each stain (e.g., hematoxylin, DAPI, DAB, and / or the like). Moreover, in some cases, the deconvolution of the image 101 may include generating one or more new channels, each of which is representative of the intensity of a single stain (e.g., hematoxylin, DAPI, DAB, and / or the like). In some cases, downstream pathological analysis may be performed based on the intensities of the individual stains. For instance, in some cases, the presence (or absence) of a target protein may be determined based at least on the intensity of the IHC stain (e.g., DAB) present in the tissue from which the image 101 is derived.

[0334]

[0252] Accordingly, it should be appreciated that the accuracy of subsequent pathological analysis may depend on the accuracy of the deconvolution. Nevertheless, conventional deconvolution methodologies yield suboptimal results due to the complex optical properties of

[0335] 46

[0336] NAI-5005412604vl the stains as well as variabilities in the staining process. For example, when generating the channel of a single stain, conventional deconvolution methodologies may misclassify the color (e.g., color signal) of some pixels. Where the image 101 is derived from a biological sample treated with 3,3-diaminobenzidine (DAB), a chromogen that is deposited as a brown precipitate in the presence of a target protein (e.g., TROP2 or PD-L1), conventional deconvolution methodologies are prone to misclassifying non-brown pixels in the image 101 as depicting the stain. To further illustrate, FIG. 67 depicts the inaccurate deconvolution of an example of an image 670 derived from tissue treated with both hematoxylin and DAB. The hematoxylin colors cell nuclei a blue or purple color while the DAB colors regions containing a target protein (e.g., TROP2 or PD-L1) a brown color. As shown in FIG. 67, the image 670 may be deconvoluted to yield a hematoxylin channel 671 and a DAB channel 672. Since the pixels associated with DAB are brown and the pixels associated with hematoxylin are non-brown (e.g., black, grey, green, and purple), the hematoxylin channel 671 should include a minimal quantity of non-brown pixels if the image 670 is deconvoluted correctly. Nevertheless, with conventional deconvolution methodologies, non-brown pixels, such as black, grey, green, and purple pixels, are prone to be misclassified as containing a valid DAB signal. The box highlights an area of the image 670 in which signal from the hematoxylin stain is inaccurately captured in the DAB channel 672 due to the misclassification of non-brown pixels.

[0337]

[0253] Various example embodiments of the present disclosure overcome the limitations of conventional deconvolution methodologies, including the aforementioned misclassification of pixels. In some example embodiments, one or more color filters may be applied to the original image 670 to exclude, from the DAB channel 672, pixels that fail to pass the color filters. For example, in some cases, subsequent to the deconvolution of the image 670 to extract the DAB channel 672, the pixels that passed this initial color deconvolution may undergo additional color filters on the hue (e.g., base color), saturation (e.g., color purity or intensity), and value (e.g., color brightness or lightness) (HSV) space. In some cases, these additional color filters may include one or more thresholds, such as a set of lower bound HSV values, a set of upper bound HSV values, and / or the like. In some cases, one or more pixels from the image 670 that passed the initial color deconvolution may be excluded, from the DAB channel 672, if the HSV values of these pixels fail to satisfy the one or more thresholds (e.g., the set of lower bound HSV values, the set of upper bound HSV values, and / or the like). In some cases, the application of the additional color filters in the HSV space may remove pixels that are in the range of grey and black, rather than the true brown associated with the DAB channel 672, in the original RGB space with better generalizability.

[0338] 47

[0339] NAI-5005412604vl

[0254] In some cases, in addition to (or instead of) the color filters, clusters of pixels (or groups of adjacent pixels) whose size fails to satisfy one or more thresholds (e.g., includes fewer than a threshold quantity of pixels) may be excluded from the DAB channel 672. For example, in some cases, a small pixel cluster may be a set of contiguous (or fully connected) pixels containing fewer than five pixels. In some cases, valid DAB stained pixels indicative of the presence of the target protein (e.g., PD-L1) are more likely to appear in higher prevalence whereas invalid, non-specific DAB stains are more likely to appear as random “salt-and- pepper” specks. Accordingly, in some cases, the removal of small pixel clusters from the DAB channel 672 may reduce the incidence of false positive signals not attributable to the valid DAB stains related to the presence of the protein target (e.g., PD-L1) of the DAB stain.

[0340]

[0255] FIG. 68 depicts an example of the DAB channel 685 of an image 680 after the application of color filters and the removal of small pixel clusters. As shown in FIG. 68, the application of the color filters and the removal of small pixel clusters reduce the quantity of misclassified pixels in the DAB channel 685.

[0341]

[0256] In some example embodiments, the preprocessing of the image 101 may also include the detection of pixels depicting non-specific staining. In some cases, non-specific staining is a type of artifact and is not related to the actual presence of the protein target of the IHC stain. FIG. 69A depicts examples of valid stains without excessive non-specific staining. Contrastingly, in FIG. 69B, the examples of stains enclosed in the rectangular boxes are nonspecific and are therefore categorized as invalid. It should be appreciated that pixels depicting non-specific staining may confound downstream pathological analysis. As such, in some cases, pixels identified as depicting non-specific staining may be excluded from downstream pathological analysis. However, conventional image processing techniques fail to detect and eliminate non-specific staining, thus compromising accuracy of downstream pathological analysis.

[0342]

[0257] Accordingly, various example implementations of the present disclosure increase the accuracy and precision of non-specific stain detection by applying a variance filter to individual portions (or patches) of the image 101. For example, in some cases, the variance filter may be applied, subsequent to deconvolution of the image 101 , to one or more individual stain channels (e.g., DAB channel) of the image 101. In some cases, the variance filter may be applied to nonoverlapping portions (or patches) of the image 101, for example, in a sliding window manner. In some cases, the variance filter may be the Laplacian filter shown below.

[0343] 48

[0344] NAI-5005412604vl rO 1 Oi 1 -4 1 .0 1 0.

[0345]

[0258] In some cases, the variance filter may be applied to a portion (or patch) of the image 101 to determine the local variance present within that portion (or patch) of the image 101. In some cases, the portion (or patch) of the image 101 may be determined to depict non-specific staining based at least on whether the local variance within the portion (or patch) of the image 101 satisfies one or more thresholds. For example, in some cases, a high local variance may indicate the presence of distinct edges and background, meaning that the staining depicted in the portion (or patch) of the image 101 is specific (e.g., consistent with an expected or reasonable pattern of protein target expression for the IHC stain). As such, in some cases, pixels forming portions (or patches) of the image 101 having a high variance (or above threshold variance) may be preserved for further pathological analysis. Contrastingly, in some cases, a low local variance may indicate the presence of spread and lack of distinct edges, which are characteristics of non-specific staining or artefactual patterns inconsistent with the actual presence of the target protein of the IHC stain. As such, in some cases, pixels forming portions (or patches) of the image 101 having a low variance (or below threshold variance) may be excluded from further pathological analysis. FIGS. 70A-70D depict examples of the 3,3- diaminobenzidine (DAB) channel from an image derived from a biological sample treated with an immunohistochemistry (IHC) stain before and after the removal of pixels depicting nonspecific staining.

[0346]

[0259] Primitive Detection in Hematoxylin and Eosin (H&E) Images

[0347]

[0260] In some example embodiments, one or more primitive features may be extracted from an image (e.g., a whole slide image (WSI)) depicting a biological sample (e.g., tissue sample, bodily fluids, and / or the like) treated with a hematoxylin and eosin (H&E) stain. In some cases, the biological sample in the image may be treated with an H&E stain to enable visualization of subcellular compartments, such as nucleus, cytoplasm, membrane, and / or the like. For example, in some cases, cell nuclei may be stained blue by the hematoxylin dye while the cytoplasm and extracellular matrix are stained various shades of pink by the eosin dye. In some cases, analysis of H&E stained images to extract primitive features may include algorithms in which regions of an image are assessed to localize individual cells, analyze cell morphology, and detect other cellular features. As an example, two cell segmentation algorithms, Stardist, a UNET based encoder-decoder, and CellViT, a transformer deep learning model with a Hovernet decoder, may be used to evaluate cellular features from these H&E regions to create

[0348] 49

[0349] NAI-5005412604vl a cell map (e.g., the cell map 103 shown in FIGS. 1 A-1C). Each of these may be used alone, or in combination with each other and / or with other cell segmentation algorithms. From the cell map and region of interest (ROI) map (e.g., from the dual-path fusion module (DP-FM) segmentation model in FIG. IF), one or more features, including primitive features, may be extracted. In some cases, embedded imaging features, such as those from ROI segmentation (e.g., by the DP-FM segmentation model in FIG. IF) can also be extracted. Finally, these final extracted features may be used to predict clinical responses, such as how patients with specific biomarker profiles respond to ADC therapies or other treatments such as alkylating agents, antimetabolites, enzyme inhibitors, anti-angiogenesis agents, microtubule disruptors, immunotherapy agents, and / or the like.

[0350]

[0261] From this overview, certain techniques and results may be useful and these may include a) membrane segmentation with a multi-channel deep learning approach, where the annotations (e.g., the Na-K channel, a membrane marker) are provided by another imaging assay (multiplexed immunofluorescence, mIF) instead of by a pathologist, (b) the ability to transfer learning from a model trained to detect one target protein (e.g., PD-L1 model) to the detection of a different target protein (e.g., TROP2 or PD-L1) model for stain deconvolution and DAB negativity / positivity classification, c) membrane extraction by digital pathology foundation model, d) positive membrane assignment in a group of tightly clustered positive cells, and / or e) detailed categorization of membranous staining patterns to include “gray zone” (or cells that may be either cytoplasmic or membrane positive, but the subcellular compartments cannot be separated definitively). When considering this overview, it will be understood that this workflow is exemplary, that at least some of the steps are optional, and that certain steps may be repeated or substituted with others as known in the art.

[0351]

[0262] Primitive Detection in Immunohistochemistry (IHC) Images

[0352]

[0263] In some example embodiments, one or more primitive features may be extracted from an image (e.g., a whole slide image (WSI)) depicting a biological sample (e.g., tissue sample, bodily fluids, and / or the like) treated with an immunohistochemistry (IHC) stain. In some cases, the IHC stain may be sensitive to a biomarker, such as a target protein (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1) such that treating the biological sample with the IHC stain may enable visualization of the biomarker in the image of the biological sample. In some cases, the one or more primitive features extracted from the IHC image may include cells positive for the biomarker (e.g., target protein positive cells) and cells negative for the biomarker (e.g., target protein negative cells). In some cases, whether a cell is positive (or negative) for the biomarker (e.g., target protein) may be contingent on the presence of the

[0353] 50

[0354] NAI-5005412604vl biomarker (e.g., target protein) in one or more specific subcellular compartments. In some cases, for different types of cell (or cell types), positive (or negative) status for the biomarker (e.g., target protein) may be determined based on the presence (or absence) of the IHC stain in different combinations of subcellular compartments. For example, in some cases, a tumor cell may be identified positive for a biomarker (e.g., a target protein such as TROP2 or PD-L1) if an IHC stain sensitive to the biomarker is detected in the membrane of the tumor cell. Contrastingly, a tumor cell may be identified as negative for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the IHC stain is absent from the cell membrane, even if the IHC stain is detected in other subcellular compartments (e.g., cytoplasm, nucleus) of the tumor cell. In some cases, unlike tumor cells, an immune cell may be identified as positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the IHC stain is present in either the membrane or the cytoplasm of the immune cell. In some cases, an immune cell with IHC stain in the nucleus of the immune cell may be identified as negative for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the IHC stain is absent from both the membrane and cytoplasm of the immune cell.

[0355]

[0264] Table 1 below summarizes the classification of tumor cells and immune cells as positive (or negative) for a target protein (e.g., PD-L1) based on the presence (or absence) of IHC staining in the membrane, cytoplasm, and nucleus.

[0356]

[0265] Table 1

[0357] 51

[0358] NAI-5005412604vl

[0266] Although the identification of biomarker positive and biomarker negative cells require the localization of individual subcellular compartments (e.g., membrane, cytoplasm, nucleus, and / or the like), treating an image with an immunohistochemistry (IHC) stain does not enable the visualization of subcellular compartments. For example, in some cases, in an image depicting a biological sample treated with an IHC stain sensitive to a biomarker (e.g., target protein such as TROP2 or PD-L1), cell membranes without the biomarker may be invisible. Conventional approaches to localizing subcellular compartments in an IHC image include nuclear-based radial expansion in which fixed-width morphological dilation is performed to estimate, based at least the location of the cell nucleus, the area of the cytoplasm and membrane of the cell. However, the assumptions underlying conventional nuclear-based radial expansion, such as the nucleus and membrane of cells having identical shape and uniform distance therebetween, are invalid. As such, nuclear-based radial expansion fails to localize the cytoplasm and membrane of cells in an IHC image with sufficient accuracy.

[0359]

[0267] Various example embodiments of the present disclosure includes various improvements upon existing nuclear-based radial expansion. For example, in some example embodiments, instead of a fixed-width, nucleus-based radial expansion may be performed as a function of the size of individual nuclei and relative crowdedness within the neighborhood of each cell. That is, in some cases, the estimated region (or area) for the cytoplasm and membrane (also called the expanded region or area) may be proportional to the size of the nucleus upon which the nucleus-based radial expansion is performed. Furthermore, in some cases, false positive classifications of biomarker positive (e.g., target protein positive) cells may be reduced by limiting the detection of IHC staining within the expanded region (e.g., between the boundary of the nucleus and the boundary of the membrane). For instance, in some cases, the aforementioned modified nucleus-based radial expansion may be performed to identify, within an image depicting a biological sample treated with an IHC stain, one or more pixels depicting the cytoplasm and / or membrane of a cell. In some cases, these pixels may form an expansion region (or area). In some cases, the result of the modified nucleus-based radial expansion may include a mask in which each pixel is assigned a value (e.g., a binary value) indicating whether the pixel is part of an expansion region (or area). In some cases, a pixel that is identified to be a part of an expansion region (or area) may undergo further evaluation to determine whether the pixel is positive for the IHC stain. As described in more detail below, whether a cell is positive for the biomarker (e.g., target protein such as TROP2 or PD-L1) may be determined based at least on the pixels identified as positive for the IHC stain.

[0360] 52

[0361] NAI-5005412604vl

[0268] To further illustrate, FIG. 71 A depicts several examples of images of biological samples treated with an IHC stain. The original IHC images are shown in the leftmost column of FIG. 71 A. As shown in FIG. 71 A, the images treated with the IHC stain first undergo cell detection to localize the cell nuclei present in each image, with the results of being shown in the second column from left. Thereafter, nucleus-based radial expansion is performed as a function of the size of each nucleus and the relativeness crowdedness of the neighborhood, thus generating expansion regions (or areas) shown in the second column from right of FIG. 71A. Finally, as shown in the rightmost column of FIG. 71 A, the presence of a biomarker (e.g., a target protein such as TROP2 or PD-L1) is detected based on the presence of the IHC stain (e.g., 3,3- diaminobenzidine (DAB)) within the expansion regions (or areas). FIG. 7 IB shows additional examples of images of biological samples treated with an IHC stain that undergoes nucleusbased radial expansion to determine an expansion region (or area) corresponding to the cytoplasm and membrane of each cell. Similarly to the examples of IHC images shown in FIG. 71A, in FIG. 71B, the detection of an IHC stain (e.g., 3, 3 -diaminobenzidine (DAB)) was performed within the expansion region (or area).

[0362]

[0269] In some example embodiments, a cell may be classified as positive (or negative) for a biomarker, such as a target protein (e.g., a tumor associated antigen (TAA) such as TROP2 or PD-L1) based at least on the presence (or absence) of a biomarker-sensitive stain in one or more specific subcellular compartments of the cell. Nevertheless, irregularities in the staining in and around cells may increase the likelihood of inconsistent and inaccurate classifications (e.g., false positives, false negatives). Various example embodiments of the present disclosure increase the consistency and accuracy of classifying a cell as positive (or negative) for a biomarker, such as a target protein (e.g., a TAA such as TROP2 or PD-L1), by at least employing a partition-based classification approach. For example, in some cases, the expanded region (or area) of a cell, which includes the cytoplasm and the membrane of the cell, may be portioned into multiple partitions (or slices). In some cases, the expanded region (or area) of the cell may be portioned into equal-sized, non-overlapping partitions (or slices). In some cases, a proportion of partitions (or slices) of the expanded region (or area) containing a signal from the biomarker-sensitive stain (e.g., an IHC stain) may be determined. In some cases, in some cases, a bin (or slice) of the expanded region (or area) may be identified as containing a signal from the biomarker-sensitive stain (e.g., the IHC stain) if one or more pixels of the expanded region (or area) is identified as being positive for the biomarker-sensitive stain. In some cases, a cell may be classified as positive for a biomarker (e.g., a target protein such as TROP2 or PD-L1) if the proportion of stain positive partitions (or slices) in the expanded region

[0363] 53

[0364] NAI-5005412604vl (or area) satisfies one or more thresholds. In some cases, the one or more thresholds may be determined based at least on the type of cell (or cell type). For instance, in some cases, one class of cell (e.g., an immune cell) may be identified as positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the proportion of stain positive partitions (or slices) of the cell satisfies one threshold (e.g., 15%) whereas another class of cell (e.g., a tumor cell) may be identified as positive for the biomarker if the proportion of stain positive partitions (or slices) of the cell satisfies a different threshold (e.g., 50%).

[0365]

[0270] To further illustrate the aforementioned partition-based approach to classifying cells as being positive (or negative) for a biomarker (e.g., a target protein such as TROP2 or PD-L1), FIG. 72 depicts a schematic diagram illustrating a cell 720 that has been divided into 16 equalsized, non-overlapping partitions (or slices). As shown in FIG. 72, the presence of a stain (e.g., an IHC stain sensitive to the biomarker) is detected in some but not all of the partitions (or slices). In some cases, the cell 720 may be classified as positive (or negative) for the biomarker (e.g., a target protein such as TROP2 or PD-L1) based at least on the proportion of partitions (or slices) positive for the stain (e.g., the IHC stain). For example, in cases where the cell 720 is a first type of cell (e.g., an immune cell), the cell 720 may be identified as positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the proportion of partitions (or slices) positive for the stain (e.g., the IHC stain) satisfies one threshold (e.g., 15%). Alternatively and / or additionally, in instances where the cell 720 is a second type of cell (e.g., a tumor cell), the cell 720 may be identified as positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) if the proportion of partitions (or slices) positive for the stain (e.g., the IHC stain) satisfies a different threshold (e.g., 50%).

[0366]

[0271] FIG. 73 depicts a flowchart illustrating an example of a process 730 for partition-based classification of biomarker status, in accordance with some example embodiments. In some example embodiments, the process 730 may be performed by a pathology analysis engine (e.g., the pathological analysis engine 114 shown in FIG. 1A). In some cases, the process 730 may be more performed by a feature extraction model of the pathology analysis engine (e.g., the feature extraction model 152 in FIG. IB) in order to determine, based at least on a cell map (e.g., the cell map 103) of an image of a biological sample (e.g., the image 101 in FIG. 1 A), a feature set (e.g., the feature set 153 in FIG. IB). In some cases, the process 730 may be performed as part of the process to extract one or more features (e.g., primitive features) from the image of the biological sample (e.g., the image 101 in FIG. 1 A). As described in more detail below, in some cases, the process 730 may be performed to identify, for use as a primitive feature, one or more biomarker positive (e.g., target protein positive) cells and / or biomarker

[0367] 54

[0368] NAI-5005412604vl negative (e.g., target protein negative) cells in the image of the biological sample (e.g., the image 101 in FIG. 1A).

[0369]

[0272] Referring again to FIG. 73, at 732, an expansion region originating from a nucleus of a cell that includes a set of pixels depicting a cytoplasm and a membrane of a cell is identified within an image of a biological sample. In some cases, nucleus-based radial expansion may be performed to identify, starting from the nucleus of the cell (e.g., the pixels the outer perimeter of the cell nucleus) the expansion region (or area). In some cases, the expansion region (or area) may include the pixels depicting the cytoplasm and the membrane of the cell. In some cases, instead of fixed-width nucleus-based radial expansion, the dimensions of the expansion region (or area) may be proportional to the size of the nucleus and the relative crowdedness of the surrounding neighborhood. In this context, the “crowdedness” of the neighborhood of the cell may correspond to a ratio between the geometric area of the neighborhood and the quantity of cells occupying the neighborhood. In some cases, the neighborhood of the cell may be an 8- , 15-, 25-, or 50-micron area around the cell.

[0370]

[0273] At 734, one or more pixels positive for a biomarker-sensitive stain are identified within the expansion region. In some cases, a pixel may be identified as positive (or negative) for the biomarker-sensitive stain (e.g., an immunohistochemistry (IHC) stain sensitive to a target protein such as TROP2 or PD-L1) based at least on the color and / or intensity of the pixel.

[0371]

[0274] At 736, the expansion region is divided into multiple partitions. For example, in some cases, the expansion region (or area) may be divided into equal-sized, non-overlapping partitions.

[0372]

[0275] At 738, a stain-positive proportion of partitions containing a threshold quantity of pixels positive for the biomarker-sensitive stain is determined. In some example embodiments, the proportion of stain-positive partitions (or slices) may correspond to a ratio of the quantity of partitions (or slices) in which a threshold quantity of pixels (e.g., at least one pixel) are positive for the biomarker sensitive stain (e.g., an immunohistochemistry (IHC) stain sensitive to a target protein such as TROP2 or PD-L1) relative to the total quantity of partitions (or slices). However, it should be appreciated that a stain-negative proportion of partitions (or slices) may be determined instead of the stain-positive proportion of partitions (or slices) and used in a corresponding manner to classify the biomarker status of the cell. A partition (or slice) may be identified as stain-negative if the quantity of stain-positive pixels within the partition (or slice) fails to satisfy one or more thresholds (e.g., no stain-positive pixels).

[0373]

[0276] At 740, the cell is classified as positive or negative for the biomarker based at least on the stain-positive proportion of partitions. For example, in some cases, the cell may be

[0374] 55

[0375] NAI-5005412604vl classified as positive (or negative) for the biomarker (e.g., a target protein such as TR0P2 or PD-L1) if the stain-positive proportion of partitions (or slices) satisfies one or more thresholds. In some cases, the cell may be classified as positive (or negative) for the biomarker (e.g., a target protein such as TR0P2 or PD-L1) if the stain-negative proportion of partitions (or slices) fails to satisfy the one or more thresholds or, alternatively, satisfies one or more different thresholds. In some cases, the one or more thresholds may be dependent on the type of cell (or cell type). For instance, where the cell is a first type of cell (e.g., an immune cell), the cell may be identified as positive for the biomarker (e.g., a target protein such as TR0P2 or PD-L1) if the proportion of partitions (or slices) positive for the stain (e.g., the IHC stain) satisfies one threshold (e.g., 15%). Alternatively and / or additionally, in instances where the cell is a second type of cell (e.g., a tumor cell), the cell may be identified as positive for the biomarker (e.g., a target protein such as TR0P2 or PD-L1) if the proportion of partitions (or slices) positive for the stain (e.g., the IHC stain) satisfies a different threshold (e.g., 50%).

[0376]

[0277] In some example embodiments, the process 730 may be performed to classify a cell as positive or negative for a biomarker (e.g., a target protein such as TROP2 or PD-L1) without differentiating between the membrane and cytoplasm of the cell. Thus, the process 730 may be suitable for determining the biomarker status of some types of cells (or cell types) but not others. In particular, the process 730 may be inadequate for determining the biomarker status of cell types that require a differentiation between membranous staining and cytoplasmic staining. For example, referring back to Table 1, whereas an immune cell may be classified as positive for a target protein (e.g., PD-L1) if either the membrane or the cytoplasm of the immune cell is positive for the target protein sensitive stain (e.g., IHC stain), a tumor cell cannot be classified as positive for the target protein (e.g., PD-L1) unless the membrane of the tumor cell is positive for the target protein sensitive stain. That is, if the cytoplasm of a tumor cell is positive for a target protein sensitive stain (e.g., IHC stain) but the membrane of the tumor cell is not, the tumor cell is classified as negative for the target protein (e.g., PD-L1). As such, while the process 730 may be performed to determine the biomarker status of tumor cells, it may be inadequate for purposes of determining the biomarker status of immune cells.

[0377]

[0278] In some example embodiments, a different classification protocol may be implemented to determine the biomarker status in instances where the biomarker status is contingent on a more precise localization of the biomarker-sensitive stain. For example, in some cases, a different classification protocol that differentiates between the cytoplasmic staining and membranous staining may be implemented to determine the biomarker status of tumor cells at least because a tumor cell cannot be classified as positive for the biomarker (e.g., a target

[0378] 56

[0379] NAI-5005412604vl protein such as TR0P2 or PD-L1) unless the biomarker-sensitive stain (e.g., a target protein sensitive immunohistochemistry (IHC) stain) is present in the membrane of the tumor cell.

[0380]

[0279] FIG. 74 depicts a flowchart illustrating an example of a process 740 for biomarker status classification that differentiates between different subcellular compartments, in accordance with some example embodiments. In some example embodiments, the process 740 may be performed by a pathology analysis engine (e.g., the pathological analysis engine 114 shown in FIG. 1 A). In some cases, the process 730 may be performed by a feature extraction model of the pathology analysis engine, such as the feature extraction model 152 shown in FIG. IB. In some cases, the process 730 may be performed to determine, based at least on a cell map, a feature set. For example, in some cases, the process 730 may be performed, based on the cell map 103, the feature set 153 in FIG. IB. In some cases, the process 740 may be performed as part of the process to extract one or more features (e.g., primitive features) from the image of the biological sample (e.g., the image 101 in FIG. 1A). As described in more detail below, in some cases, the process 740 may be performed to identify, for use as a primitive feature, one or more biomarker positive (e.g., target protein positive) cells and / or biomarker negative (e.g., target protein negative) cells in the image of the biological sample (e.g., the image 101 in FIG. 1A).

[0381]

[0280] Referring again to FIG. 74, at 742, one or more pixels depicting a nucleus of a cell are identified within an image of a biological sample. In some example embodiments, a nucleus segmentation model may be applied to identify the one or more pixels depicting the nucleus of the cell. In some cases, the nucleus segmentation model may be a hematoxylin channel -based cell detection and classification model trained on images of biological samples (e.g., tissue samples such as tumor tissue samples) from multiple types of tissue. In some cases, the images may be annotated, for example, with pixel-wise labels that identify those depicting cell nuclei. For example, in some cases, the nucleus segmentation model may be trained on a pan-tumor dataset, which includes nuclei labels across multiple tissue types. Furthermore, the cell nuclei in the pan-tumor dataset are categorized into different classes, such as neoplastic (tumor), inflammatory, connective / soft tissue, dead, epithelial cells, and / or the like.

[0382]

[0281] At 744, one or more pixels depicting a membrane of the cell are identified within the image of the biological sample. In some example embodiments, a membrane segmentation model may be applied to identify the one or more pixels depicting the membrane of the cell. In some cases, the membrane segmentation model may be a cell segmentation model that has been pretrained on multiplexed immunofluorescence (mIF) images before being applied to the 3,3'- Diaminobenzidine (DAB) channel of immunohistochemistry (IHC) images. In some cases, the

[0383] 57

[0384] NAI-5005412604vl IHC images may undergo deconvolution to separate the membrane channel (e.g., DAB channel) in which stains sensitive to the target protein from the nuclear channel (e.g., 4', 6- diamidino-2-phenylindole (DAPI) or hematoxylin channel) in which stains sensitive to cell nuclei are visible. It should be appreciated that cell membranes positive for the target protein may be visible on the membrane channel (e.g., DAB channel) whereas cell membranes negative for the target protein are not. As such, in some cases, the membrane segmentation model may be trained with a deep learning model to infer the cell boundaries of those cells whose membranes are negative for the target protein and are therefore non-visible on the membrane channel (e.g., DAB channel), using the nuclear region image of those cells as input to the model. The ability to infer the boundaries of target protein negative cells as well as determining the membrane intensities of these cells despite those membranes being non-visible on the membrane channel (e.g., DAB channel) is advantageous for a number of critical reasons. Notably, the membrane intensity distribution of target protein negative cells may inform the membrane intensity threshold differentiating between target protein negative and target protein positive cells. For instance, in some cases, the membrane intensity threshold for classifying a cell as positive for the target protein may be set to a certain quantile (e.g., 95thquantile) of the membrane intensities of the target protein negative cells.

[0385]

[0282] In some cases, the membrane segmentation model may be trained on annotated images of cells positive for one biomarker (e.g., a first target protein such as TR0P2 or PD-L1) to localize the membranes of cells in biological samples treated with a stain that is sensitive to a different biomarker (e.g., a second target protein such as TR0P2). In some cases, the membrane segmentation model may be trained on annotated images of cells positive for one biomarker (e.g., a first target protein such as TR0P2 or PD-L1) to localize the membranes of both cells that are positive for the different biomarker (e.g., a second target protein such as TR0P2) and cells that are negative for the different biomarker. It should be appreciated that the membranes of cells negative for a biomarker may be non-visible in images of stains sensitive to the biomarker. For example, as noted, the membranes of target protein negative cells may be invisible in images of a target protein sensitive stain (e.g., IHC images). Accordingly, in some cases, the membrane segmentation model, through training and exposure to the membranes of target protein positive cells, may be capable of inferring the boundaries of target protein negative cells despite the absence of the target protein sensitive stain demarcating the boundaries of target protein negative cells. To further illustrate, FIG. 76 depicts an example of an image of an IHC image that has been deconvolved into a nuclear channel (or hematoxylin channel) and a membrane channel (or DAB channel). The output of the membrane

[0386] 58

[0387] NAI-5005412604vl segmentation model may include the predicted cell boundaries of both target protein positive and target protein negative cells.

[0388]

[0283] At 746, one or more pixels depicting a cytoplasm of the cell are identified within the image of the biological sample. In some example embodiments, the one or more pixels depicting the cytoplasm of the cell may include one or more pixels disposed between the pixels depicting the nucleus (e.g., the pixels depicting the outer perimeter of the nucleus) and the pixels depicting the membrane (e.g., the pixels depicting the inner perimeter of the membrane).

[0389]

[0284] At 747, whether a threshold quantity of pixels associated with the cell are positive for a biomarker stain is determined. In some example embodiments, the pixels associated with the cell may include the pixels depicting the membrane, the cytoplasm, and the nucleus of the cell. In some cases, whether the cell is positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) may be determined based at least on whether a threshold quantity (e.g., at least one) of the pixels depicting the membrane, the cytoplasm, and the nucleus of the cell is positive for the biomarker-sensitive stain (e.g., a target protein sensitive IHC stain).

[0390]

[0285] At 747-N, not a threshold quantity of pixels associated with the cell are determined to be positive for the biomarker stain. As such, at 748, the cell is classified as negative for the biomarker. In some example embodiments, where the quantity of pixels depicting the membrane, the cytoplasm, and the nucleus of the cell that is positive for the biomarker-sensitive stain (e.g., a target protein sensitive IHC stain) fails to satisfy one or more thresholds (e.g., at least one stain positive pixel), the cell may be identified as being negative for the biomarker (e.g., a target protein such as TR0P2 or PD-L1).

[0391]

[0286] Alternatively, at 747-Y, a threshold quantity of pixels associated with the cell are determined to be positive for the biomarker stain. As such, at 751, whether the cell exhibits a differentiable membrane is determined. At 751-N, the cell is determined to not exhibit a differentiable membrane. As such, at 752, the cell is classified as exhibiting only cytoplasmic staining.

[0392]

[0287] Alternatively, at 751-Y, the cell is determined to exhibit a differentiable membrane. As such, at 753, whether the membrane of the cell is differentiable from the nuclear boundary of the cell is determined. In some example embodiments, where the cell is determined to exhibit a differentiable membrane (at operation 751), further image analysis may be performed to determine whether the membrane of the cell is differentiable from the nucleus of the cell (e.g., the outer perimeter of the cell nucleus). In some cases, the membrane of the cell may be differentiable from the cell nucleus if a sufficiently large (or wide) cytoplasmic compartment is interposed therebetween. For example, in some cases, the membrane of the cell and the

[0393] 59

[0394] NAI-5005412604vl nucleus boundary of the cell may be differentiable if the quantity of pixels between the membrane of the cell (e.g., the inner perimeter of the membrane of the cell) and the cell nucleus (e.g., the outer perimeter of the cell nucleus) satisfies one or more thresholds. In some cases, where the quantity of pixels between the membrane of the cell (e.g., the inner perimeter of the membrane of the cell) and the cell nucleus (e.g., the outer perimeter of the cell nucleus) fails to satisfy the one or more thresholds, the process 740 continues at operation 754 described in more detail below. Alternatively, where the quantity of pixels between the membrane of the cell (e.g., the inner perimeter of the membrane of the cell) and the cell nucleus (e.g., the outer perimeter of the cell nucleus) satisfies one or more thresholds, the process 740 continues at operation 756 described in more detail below.

[0395]

[0288] At 753-N, the membrane of the cell and the nuclear boundary of the cell are determined to be indifferentiable. As such, at 754, the cell is classified as exhibiting indifferentiable membranous staining and cytoplasmic staining. In some example embodiments, the cell may be classified as exhibiting indifferentiable cytoplasmic and membranous staining if, at operation 753, the cell is determined to exhibit an insufficient quantity of pixels associated with the cytoplasmic compartment. In some cases, where the cell is determined to exhibit indifferentiable membrane and nuclear boundary, the cell may be classified as being in a “gray zone.” In some cases, there may be at least two variations of cells in the gray zone, both of which may be positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1) for certain cell types (e.g., tumor cells). For example, in some cases, cells in the gray zone may exhibit strong membranous and cytoplasmic staining. Alternatively, cells in the gray zone may exhibit strong membranous staining but weak (or non-visible) staining in the cytoplasmic region. In some cases, for tumor cells, both types of cells in the gray zone may be considered positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1). As such, in some cases, where the cell is a tumor cell having an indifferentiable membrane and nuclear boundary, the cell may be classified as positive for the biomarker (e.g., a target protein such as TROP2 or PD-L1).

[0396]

[0289] Alternatively, at 753-Y, the membrane of the cell and the nuclear boundary of the cell are determined to be differentiable. As such, at 756, the cell may be classified as exhibiting membranous staining. In some example embodiments, where the cell is determined to exhibit membranous staining, further image analysis may be performed to determine the biomarker status of the cell (e.g., positive or negative for a target protein such as TROP2 or PD-L1). For example, in some cases, it may be possible for the stained membrane to be shared between two or more adjacent cells and the cell with the stained membrane is the one that should be

[0397] 60

[0398] NAI-5005412604vl classified as positive for the biomarker (e.g., a target protein such as TR0P2 or PD-L1). An example of this phenomenon is shown in FIG. 75A, which depicts an example of an image of a target protein sensitive IHC stain treated biological sample in which adjacent cells appear to share the same IHC stained membrane. The lack of visual differentiation between the membranes of adjacent cells may give rise to false positives when identifying cells positive for the biomarker (e.g., a target protein such as TR0P2 or PD-L1). Accordingly, to reduce the incidence of false positives in which the cell is incorrectly attributed the stain positive membrane of a neighboring cell, the cell is classified as positive for the biomarker (e.g., a target protein such as TR0P2 or PD-L1) if the completeness of the membrane stain fail satisfies one or more thresholds.

[0399]

[0290] In this context, membranous staining completeness may be metric corresponding to the proportion of the partitions of the membrane positive for the biomarker-sensitive stain (e.g., a target protein sensitive IHC stain). To further illustrate, FIG. 75B depicts a schematic diagram of a cell 750 that has been partitioned into slices. In some cases, the completeness of the membranous staining may correspond to a ratio of the quantity of stain positive slices relative to the total quantity of slices. In some cases, whether the cell exhibits a differentiable membrane signal may be determined based at least on whether the proportion of stain positive slices satisfies one or more thresholds. For example, the differentiable membranous staining identified at operation 753 may be attributed to the cell, thus conferring the cell with a biomarker positive status, if the cell exhibits sufficiently complete membranous staining, as indicated by proportion of stain positive slices present in the cell satisfying the one or more thresholds.

[0400]

[0291] In some example embodiments, the result of the process 740 may include the classification of an individual cell as either positive or negative for a biomarker (e.g., a target protein such as TROP2 or PD-L1). In some cases, one or more additional biomarker metrics may be computed based on the results of the process 740. For example, in some cases, the biomarker status of a patient may be determined based on the biomarker status of multiple individual cells present in an image of a biological sample associated with the patient. In some cases, the biomarker status of the patient may be determined based on a score (e.g., a combined positive score (CPS)), which may correspond to a ratio of biomarker positive cells (e.g., target protein positive cells) relative to another quantity of cells (e.g., total number of tumor cells) present in the biological sample. In some cases, the patient may be identified as positive for the biomarker when the CPS of the patient satisfies the one or more thresholds and negative for the biomarker when the CPS of the patient fails to satisfy the one or more thresholds.

[0401] 61

[0402] NAI-5005412604vl

[0292] The biomarker status of individual cells determined by the process 740 was applied towards determining the CPS and corresponding biomarker status of patients whose biomarker status was also determined using pathologist classifications of biomarker positive and negative cells. The results shown in FIGS. 77A-77B compares the performance of the process 740 as applied towards determining patient-level biomarker status classification, including sensitivity (quantified by positive percent agreement (PPA)) and specificity (quantified by negative percent agreement (NPA)), was evaluated against ground-truth biomarker status derived from pathologist annotations. The table in FIG. 77A summarizes the incidence of true negatives (e.g., predicted negative for true negative), false negatives (e.g., predicted negative for true positive), false positives (e.g., predicted positive for true negative), and true positives (e.g., predicted positive for true positive) in the output of the process 740. The process 740 achieved a PPA (for sensitivity) of 0.853 and an NPA (for specificity) of 0.875. The true positive rate and false positive rates of the process 740 are also shown graphically in FIG. 77B. The process 740 achieved an AUROC of 0.923, which is significantly higher than the AUROC of 0.5 associated with random guessing and close to the AUROC of 1.0 of a perfect classifier model.

[0403]

[0293] Exemplary Features

[0404]

[0294] Table 2 below includes examples of features from different scales (or resolutions). The examples of features shown in Table 2 include primitive features, which may be human interpretable features (HIFs), de novo features, and / or hybrid features extracted from an image (e.g., a whole slide image (WSI)) depicting a biological sample (e.g., a tissue sample, bodily fluids, and / or the like). In some cases, the primitive features may be extracted from specific regions (e.g., regions of interest (ROIs)) in the image, such as total tissue, invasive tumor region, in situ tumor region, and / or the like. The examples of features shown in Table 2 also include higher order features. As noted, in some cases, higher order features may be derived from one or more primitive features. Moreover, higher order features may include local (or regional) features extracted from overlapping and / or non-overlapping regions of the image (e.g., whole slide image (WSI)), such as cell-anchored neighborhoods, tessellations, and / or the like. In some cases, higher order features may be extracted from different size regions of the image, such as different radii for cell-anchored neighborhoods. Furthermore, in some cases, higher order features may also include human interpretable features (HIFs), de novo features, and / or hybrid features extracted from the image (e.g., a whole slide image (WSI)).

[0405]

[0295] Table 2

[0406] 62

[0407] NAI-5005412604vl

[0408] 63

[0409] NAI-5005412604vl

[0410]

[0296] Computing Platform

[0411]

[0297] A computing platform may be used to carry out at least some of the steps of algorithm 100, and FIG. 1G illustrates an example of a computing device 140 capable of doing so. Device 140 can be a host computer connected to a network. Device 140 can be a client computer or a server. As shown in FIG. 1G, device 140 can be any suitable type of microprocessor-based

[0412] 64

[0413] NAI-5005412604vl device, such as a personal computer, workstation, server or handheld computing device (portable electronic device) such as a phone or tablet. The device can include, for example, one or more of processor 141, input device 142, output device 143, storage 144, and communication device 146. Input device 142 and output device 143 can generally correspond to those described above and can either be connect to or integrated with the computer. Input device 142 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 143 can be any suitable device that provides output, such as a screen, touch screen, haptics device, or speaker.

[0414]

[0298] Storage 144 can be any suitable device that provides storage, such as an electrical, magnetic or optical memory including a RAM, cache, hard drive, or removable storage disk. Communication device 146 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus or wirelessly.

[0415]

[0299] Software 145, which can be stored in storage 144 and executed by processor 141, can include, for example, the programming that embodies the functionality of the present disclosure (e.g., as embodied in the devices as described above). Software 145 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage 144, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.

[0416]

[0300] Software 145 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic or infrared wired or wireless propagation medium.

[0417]

[0301] Device 140 may be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable

[0418] 65

[0419] NAI-5005412604vl communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines. Device 140 can implement any operating system suitable for operating on the network. Software 145 can be written in any suitable programming language. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a Web browser as a Web-based application or Web service, for example.

[0420]

[0302] Extraction of Tumor ROI from IHC Image and Cell Detection

[0421]

[0303] As shown in FIG. ID, the segmentation engine 112 may identify one or more relevant regions of interest (ROIs) present in the image 101. In some cases, the identification of regions of interest (ROIs) may include generating a segmentation map in which pixels in the image 101 depicting a portion of a region of interest(ROI) are assigned a first label (or a first value) and pixels in the image 101 not depicting a portion of a region of interest (ROI) are assigned a second label (or a second value). FIG. 2 shows an input image on the left, a ground truth segmentation mask annotated for tumor regions of interest (ROI) by a pathologist on the upper right, and a segmentation mask localizing of tumor regions of interest (ROI) generated by a deep learning model on the lower right. In sone cases, each segmentation mask may include annotations across one or more channels, each of which corresponding to a different color (e.g., grayscale, RGB, CMYK, and / or the like). In this example, the regions of interest (ROIs) are exclusively tumor regions (excluding major artifacts). However, in other examples, it may be desirable to extract features from the non-tumor regions for their potential impact on therapeutic response. After localizing the one or more regions of interest (ROIs), one or more primitive features may be extracted from the regions of interest (ROIs) for analysis. In some cases, the one or more primitive features may include the location and types of cells present in each region of interest (ROI). In some cases, additional machine learning enabled segmentation may be performed to localize and classify, within each region of interest (ROI), one or more cells, subcellular compartments, and / or the like. In some examples, the segmentation may be performed by applying a hematoxylin channel-based pan-tumor cell detection / classification model. The model may, for example, be trained on a public dataset such as the PanNuke public dataset, which includes nuclei labels across 19 different tissue types with over 0.2 million cell nuclei categorized into 5 classes: Neoplastic (Tumor), Inflammatory, Connective / Soft tissue, Dead and Epithelial Cells. Representative image tiles of the performance of PanNuke

[0422] 66

[0423] NAI-5005412604vl pretrained tumor cell detection model in two different cancer types: NSCLC (left) and TNBC (right) are shown in FIG. 3.

[0424]

[0304] FIG. 4 A are two fields from whole slide images (WSIs) side-by-side which show visualization of the model performance. Specifically, the two regions of interest (ROIs) show tumor cell nuclei detection / classification performance on a digital histologic image taken of a non-small cell lung cancer (NSCLC) specimen. The left-hand figure shows the raw image, and the right-hand figure shows the same image with a predictive mask (nuclei shown in yellow) of the tumor cells, overlaid on top of the image. In this example, a pathologist may study the model performance on the right-hand side in terms of the cells being predicted as tumor cells and other locations and check against the original image on the left-hand side to determine whether the model is performing adequately. FIG. 4B shows another example of two side-by- side images showing the use of the tumor detection model.

[0425]

[0305] After detecting tumor cells using the cell detection model, the next step is to classify whether a tumor cell is positive or negative for a target protein (e.g., TROP2 or PD-L1) based on a membranous staining model. Specifically, membrane boundaries may be useful in tumor cell target protein (e.g., TROP2 or PD-L1) quantification. In this example, the tumor cell membrane detection model was trained using a combination of images (e.g., whole slide images (WSIs) of biological samples treated with a target protein sensitive stain (e.g., an immunohistochemistry (IHC) sensitive to PD-L1), and a pre-trained cell segmentation model was also trained on multiplex immunofluorescent (mIF) images where the multiplex immunofluorescent (mIF) images were comprised of two channels: 1) a nuclear DAPI channel and 2) Na-K ATPase membrane channel to enable the delineation of the nuclear and membrane boundaries of a new input cell image. As shown in FIG. 5, two models are presented. The top row illustrates the use of a PD-L1 immunohistochemistry (IHC) model with a series of images (i.e., an immunohistochemistry (IHC) image, PD-L1 membrane detection and the PD-L1 overlay on the image), and the bottom row illustrates the use of a segmentation algorithm (Mesmer) membrane detection model (i.e., a multiplex immunofluorescent (mIF) patch, Mesmer membrane detection algorithm and Mesmer segmentation). Specifically, any multichannel deep learning cell segmentation algorithm may be used where an IF image is taken as an input and returns both a cell nucleus and cell membrane map as outputs. In order to perform the multi-channel deep learning cell segmentation algorithm on IHC images, an IHC image may be first deconvoluted to separate the Hematoxylin channel from the DAB channel, and the Hematoxylin channel may be used as the nuclear DAPI channel and the DAB channel as the membrane channel, respectively, to enable analysis of the IHC image in an analogous fashion

[0426] 67

[0427] NAI-5005412604vl to the 2D immunofluorescence image. FIGS. 6-7 show several other examples of comparisons between a PD-L1 model and a Mesmer membrane detection algorithm.

[0428]

[0306] Composite Model

[0429]

[0307] A composite model for cell detection and membrane detection may include a combination of the nuclei-based model (e.g., PanNuke-based model) with the cell boundary model (e.g., the Multiplex-trained model to recognize the membrane). Combining the nuclear model with the membrane model, a composite model may be developed which detects cells but also delineates the boundary of the cell membrane in one single approach as shown in FIG. 8. As depicted in FIG. 8, features of ROI 800 depicting multiple individual cells and their nuclei, membranes, and cytoplasmic compartments, as detected by the models, can be evaluated after being analyzed with the composite model, and a resulting image shows the cell nucleus 810 represented with a yellow region, the cell membrane 812 delineated by a black cell boundary, and the cell cytoplasm 814 as the area between the cell membrane and nucleus.

[0430]

[0308] Classification of Biomarker Positive and Biomarker Negative Cells

[0431]

[0309] In some example embodiments, the image 101 may undergo a qualitative assessment. Through quantification of the intensity of the protein sensitive stain (e.g., TROP2 or PD-L1 sensitive stain) and assessment of the stain’s spatial distribution, tumor cells can be classified as either positive or negative for the target protein (e.g., TROP2 or PD-L1). The positive / negative classification may be based on a transfer learning approach using an internal IHC model previously developed for PD-L1. In some examples, target protein negative cells may include tumor cells which contain zero DAB stains and / or tumor cells which contain DAB stains, but no clear “differentiable” signals that outline the cell membrane. In some examples, target protein positive cells may include tumor cells which contain edge signals that can be differentiated from the detected nuclear boundary, and / or tumor cells which contain edge signals that cannot be differentiated from cytoplasmic signals due to the homogeneity of stain intensity between the cytoplasmic and membranous cell compartments or small cell cytoplasm area.

[0432]

[0310] More specifically, a sub-algorithm may be used to classify or assign the membranous staining into one of four different classes of patterns. A flow chart illustrating an example of this process is shown in FIG. 9. As shown, the image 101 may be assessed to determine whether a cell area in the image 101 contains any target protein (e.g., TROP2 or PD-L1) signals or signals from a stain (e.g., immunohistochemistry stain) sensitive to the target protein. If there is no signal at all, then the cell is classified as target protein negative (e.g., class #1). If the cell contains a signal from the target protein (e.g., TROP2 orPD-Ll), then the image 101 is assessed

[0433] 68

[0434] NAI-5005412604vl to determine whether the cell contains differentiable “edge” signals. A cell that does not contain differentiable “edge” signals can be classified as a cell having cytoplasmic only staining (e.g., class #2). A cell that does contain differentiable “edge” signals will then be further interrogated to see if the edge signals can be differentiated from the nuclear boundary. A cell that has edge signals that can be differentiated from the nuclear boundary is a cell with membranous staining (e.g., class #3), and if the edge signals cannot be differentiated from the nuclear boundary, then then cell will be classified in a gray zone (e.g., class #4). In some cases, the gray zone may represent a cell that cannot be differentiated or if it is difficult to decipher whether the staining is cytoplasmic only or membranous as well. FIGS. 10A-10C are representative image tiles showing visual examples of tumor cells (shown in the center) with target protein signals (e.g., TROP2 or PD-L1 signals). Specifically, FIG. 10A shows cytoplasmic only staining (class #2), FIG. 10B shows membranous only staining (class #3), and FIG. IOC shows the gray zone (class #4). Additional representative image tiles are presented which show target protein negative cells (class #1) in FIG. 11, membranous positive cells (class #3) in FIG. 12, and gray zone cells (class #4) in FIG. 13.

[0435]

[0311] Quantitative Assessment of Protein Sensitive Staining Intensity

[0436]

[0312] In some cases, the protein sensitive stain (e.g., TROP2 or PD-L1 sensitive stain) intensities present in cells (e.g., tumor cells) may be quantified per subcellular compartments (e.g., membrane, cytoplasm, and entire cell) using an optical density (OD) and / or a proportionbased approach. In the optical density-based approach, the intensity of the protein sensitive stain may be quantified by averaging the intensity values of pixels positive for the protein sensitive stain in each individual subcellular compartment. In some cases, with the optical density-based approach, the quantification of the intensity of the protein sensitive stain may expressly exclude contribution from those pixels negative for the protein positive stain. Alternatively, with a proportion-based approach, the intensity of the protein sensitive stain may be quantified by summing the intensity values of pixels in each individual cellular compartment before dividing the value by the compartment’s area. Unlike the optical density -based method, the proportion-based approach yields a metric that integrates total signal from the stain by combining both stain intensity (or pixel intensity value) and area of the regions where signals are (or are expected to be) observed.

[0437]

[0313] In some cases, in addition to classifying a cell into the four classes or types shown in FIG. 9, it may also be useful to determine whether there is complete membranous staining or a partial membranous staining. FIG. 14 shows an example of this type of classification in which region 1425 is the nuclear region and the region 1450 outside of (or surrounding) region 1425

[0438] 69

[0439] NAI-5005412604vl is the non-nuclear region of a target protein positive tumor cell 1400. In this figure, the rim 1430 represents the protein sensitive stain (e.g., target protein sensitive 3,3-diaminobenzidine (DAB) stain), and the radial lines indicate the partitioning of the non-nuclear area into different slices, and the % completeness is calculated based on the number of slices containing a signal from the protein sensitive stain (e.g., the overlap between yellow rim and the non-nuclear area) divided by the total number of slices. It may also be possible to determine on a continuous scale the proportion of completeness of the membrane stain. In one example, to compute this proportion or the completeness measure, the number of pixels that has staining in the membrane compartment is evaluated and then normalized by the total area in the non-nuclear region. This metric may be referred to as the positive pixel proportion (PPP). This may be a useful metric to be used to compare the outcome to pathologist results or to make additional observations (e.g., to determine when cells are touching one another or to associate a particular stain with a particular cell). For example, a group of target protein positive cells may be tightly clustered and share membranes, and it may be challenging to correctly assign positivity to these cells. By estimating the PPP for a pair of connected positive cells in the cluster, it may be possible to assign the positivity to the cell (or cells) that exceed a prespecified threshold for the positive proportion / circumferential staining metric. In some examples, the PPP threshold for assigning membrane to co-localized positive cells may be in the range of 30%-50%. In this example, a 100% region staining implies complete membranous staining.

[0440]

[0314] FIG. 15 illustrates performance evaluation of the tumor cell classification model. On the left is a representative image tile of the approximate tile size to be validated, which is approximately 125 microns by 125 microns, and on the right is a computer-generated image of the predicted tumor cell nuclei. In this example, a pathologist can review the two images to determine if the performance is adequate. In this case, each cell is classified into one of the four classes discussed above (i.e., negative, membrane positive, cytoplasmic positive and gray zone), and the pathologist may re-classify the predicted cells into the correct or appropriate categories. For example, by moving the incorrectly predicted cells into the correct categories, the current model can be fine-tuned with a pathologist’s feedback. Additionally, it may be possible to evaluate neighboring cells that share a membrane. In order to do that for each of the positive cells in a cluster, the percent of the membrane completeness may be calculated, and the positivity may be assigned to the cell that has the most circumferential or the most complete membranous staining.

[0441]

[0315] Thus, from the cell map 103 generated by the segmentation engine 112 shown in FIG.

[0442] 1 A, it may be possible to produce target protein classification (i.e., positive v. negative) as well

[0443] 70

[0444] NAI-5005412604vl as differentiation between cytoplasmic positivity and membrane positivity. In addition, membrane information, including stain intensity and membrane completeness (e.g., partial v. complete), may also be measured at a continuous scale. Finally, local features including a) clustering patterns of positive / negative cells, b) number, size, types, and / or geographic distribution of cells, c) clustering patterns in relation to other cell types, d) diffusion, payload, and / or linker correlation from a later response study, and / or e) distribution of the target protein (e.g., TR0P2 or PD-L1) at pixel level, may also be measured.

[0445]

[0316] Spatial Heterogeneity Features

[0446]

[0317] In some example embodiments, the feature extraction model 152 (or the end-to-end model 156) may extract, from the cell map 103 of the image 101, one or more spatial heterogeneity features. In some cases, spatial heterogeneity features may include spatial heterogeneity features from overlapping and / or non-overlapping portions of the image 101, in which case these spatial heterogeneity features are also local (or regional) features of the image 101. In some cases, the spatial heterogeneity features may be extracted from one or more tessellations of the image 101. Examples of these tessellation features, which may be extracted from individual squares (e.g., about 250 microns by 250 microns, or other dimensions) in the image 101 are shown in FIG. 16. For example, in some cases, each tessellation may corresponding to a non-overlapping square of 250 microns by 250 microns created by at least partitioning the image 101. In some cases, the tessellations may include those that contain at least one target protein positive cell while excluding those that do not include any target protein positive cells. In the example shown in FIG. 16, an analysis may be performed to determine the relationship of target protein positive cells and target protein negative cells within that square. Those relationships may yield one or more spatial heterogeneity features. In some cases, alternate and / or additional spatial heterogeneity features may be generated by further aggregating information across two or more tessellations (e.g., squares). In some cases, the aggregation of information from multiple tessellations may yield global (or whole image) spatial heterogeneity features.

[0447]

[0318] In another example, shown in FIG. 17, a Voronoi tessellation may be used to define neighborhoods into polygonal regions and evaluate the local spatial relationship between one target protein positive cell and other target protein positive as well as negative cells. Sample features that may be calculated from this approach include the average number of target protein negative cells in each target protein positive neighborhood, the uniformity of target protein negative cells across neighborhoods, the degree of cluster and spread of target protein positive cell neighborhoods and / or the uniformity of neighborhood areas.

[0448] 71

[0449] NAI-5005412604vl

[0319] A third approach may include the use of hotspot features to examine the interaction between target protein positive and target protein negative cell clusters, and specifically to examine the clustering patterns among target protein positive cells and among target protein negative cells, and then identify the overlapping region of the target protein positive clusters and the target protein negative clusters as shown in FIG. 18A-18B. In some examples, the hotspot feature may represent the proportion of the tumor area that is occupied by the colocalization of the target protein positive and target protein negative clusters. FIG. 19 illustrates an example of hotspot feature extraction, and it will be understood that certain metrics may be derived such as (a) target protein positive and / or target protein negative colocalization hotspot score (i.e., the fractional area within the ROI, with an overlap of target protein positive and target protein negative cell clusters), (b) target protein positive tumor cell hotspot score, (c) target protein negative tumor cell hotspot score and / or (d) Morisita-Horn index bivariate correlation between target protein positive and target protein negative cells across tumor regions. Metrics (a)-(c) can be determined by the ratio of the region areas that contain overlapping target protein positive / negative cell clusters, the target protein positive cell clusters, and target protein negative cell clusters normalized by the total tumor area. And for metric (d), the Morisita-Horn Index Mcanbecalculated using the following equation:

[0450] I _ xlt cX1 I wherein—xi and Pi ~ ~c,xt denotes the number of times species i of lymphocytes I is represented in the total number of lymphocytes in the biological sample depicted in the C ‘ image, andxi denotes the number of times species1of cancer cellcis represented in the total number of cancer cellscin the biological sample depicted in the image.

[0451]

[0320] FIG. 20 illustrates this difference in highly segregated (left) and highly co-localized (right) target protein positive and target protein negative tumor cells. Another example is shown FIG. 21 of the co-localization of lymphocytes and cancer cells. Similar to this example, an analysis can be conducted to evaluate how dispersed or co-localized the target protein negative cells and target protein positive cells are on the image (e.g., whole slide image (WSI)) as a whole. In FIG. 21, the co-localization hotspots, as well as the spatial clusters of immune and cancer cells, are shown. To perform this analysis, for each cell type the method identifies the percentage of regions on the image (e.g., whole slide image (WSI)) that it is non-randomly clustered, which is z-score based. Co-localization of cell clusters such as lymphocyte clusters

[0452] 72

[0453] NAI-5005412604vl and cancer cell clusters (vs. single cells) can then be analyzed. Thez-score for a region1may indicate whether statistically significant clusters of specific cell types are found in the region, where hotspot regions are defined by false discovery rate (FDR) adjusted P < 0-^5 according to the following equation: wherein n denotes the quantity of grids, Cj denotes the quantity of target protein positive or target protein negative cells in grid j, w^j is 1 where grid i and grid j are neighbors, j is 0 where grid i and grid j are not neighbors, and c denotes the average quantity of target protein positive or target protein negative cells across all grids.

[0454]

[0321] Another statistical metric that may be useful to determine is the distribution of target protein negative tumor cells based on their proximity to the target protein positive tumor cells. In this example, for each image (e.g., whole slide image (WSI)), a method may include generating a kernel density map for target protein positive tumor cells according to the following equation:

[0455] In the above expression, for each target protein negative cell with the spatial coordinates x, (x) measures the spatially weighted density of the target protein positive cells surrounding that target protein negative cell. K denotes the kernel function (e.g., Gaussian kernel), h denotes the bandwidth parameter, and n denotes the total quantity of cells.

[0456]

[0322] For each target protein negative tumor cell with spatial coordinate 7, f(y) then measures the spatially weighted target protein positive tumor density surrounding that target protein negative tumor cell. In this example, f ) can be interpreted as the proximity of the target protein negative tumor cell to target protein positive tumor cell clusters. Subsequently, using Gaussian mixture clustering on / (y) , target protein negative tumor cells may be classified into three classes: intra target protein positive cluster, adjacent target protein positive

[0457] 73

[0458] NAI-5005412604vl cluster, and distal target protein positive cluster. Similarly, FIG. 22 illustrates certain techniques used in this analysis including statistical summary of the three classes of target protein negative cells, such that their prevalence distribution can be used to construction additional features for clinical response prediction.

[0459]

[0323] Feature Extraction Approaches

[0460]

[0324] Referring again to the feature extraction model 152 shown in FIG. ID, in some cases, the feature extraction model 152 may be configured to implement three different approaches to feature extraction: (1) the human interpretable (HIF) approach 162, (2) a data driven de novo imaging feature discovery approach 164, and (3) a hybrid approach 166 that combines aspects of the human interpretable feature (HIF) and the de novo imaging feature discovery approaches. It will be understood that these approaches are merely examples and that others are possible. Additionally, as illustrated by the hybrid approach, other combinations of approaches for feature extraction are also possible.

[0461]

[0325] Human Interpretable Feature (HIF) Approach

[0462]

[0326] In the HIF approach 162, one or more segmentation models trained for region of interest (ROI) segmentation and cell (or subcellular compartment) segmentation / classification may be first applied to the image 101 (e.g., a post quality control (QC) whole slide image (WSI)) to detect and classify different cell types (e.g., tumor, immune, stromal) present in each region of interest (ROI) in the image 101. An independent target protein positive / negative model may be subsequently used to determine the biomarker status (e.g., target protein positive or target protein negative) of each cell (e.g., tumor cell) or constituent subcellular compartments. Doing so may generate one or more primitive features. Once these primitive features are detected, the feature extraction model 152 may further extract one or more predetermined human interpretable features (HIFs) from the image 101.

[0463]

[0327] In one example, an image (e.g., a whole slide image (WSI)) from the MK-2870-001 TNBC cohort was assessed for adequacy, requiring a minimum of 1 square millimeter (mm2) of evaluable tumor area and at least 100 tumor cells in a contiguous span without obscuring artifacts. Expert annotations were then performed to identify the total evaluable tissue area, invasive ROIs, and ductal carcinoma in situ (DCIS), if present. Digital parameters were extracted to develop hypothesis-driven features, as well as higher-order features that captured local characteristics and expression heterogeneity. A targeted approach was adopted to curate non-redundant digital pathology features to test for relationships between a parsimonious set of features and clinical response in this modestly sized cohort.

[0464] 74

[0465] NAI-5005412604vl

[0328] Certain digital parameters may be evaluated, and these include the feature type (cellbased, pixel-based, or tessellated), definitions of the anchor cells for local neighborhoods, signal measurement method (OD vs proportion), size of the neighborhood to analyze (8, 15, 25, and 50 microns for cell-anchored neighborhoods or 150, 300, 450, and 600 microns for tessellation), cell compartment in which to evaluate for target protein staining (membrane, cytoplasm, total cell), image-level region in which to evaluate target protein staining (total tissue, invasive tumor region, invasive and DCIS region), and percentiles in 10% increments to give image level summary scores.

[0466]

[0329] These basic parameters also enable construction of higher order features combining multiple parameters. This approach enables development of local features of interest, which may relate to heterogeneity of target protein expression, clustering patterns, and / or stain intensity, among other possibilities, and may ultimately relate to response to an anticancer therapeutic (e.g., an antibody drug conjugate (ADC), alkylating agent, antimetabolite, enzyme inhibitor, anti-angiogenesis agent, microtubule disruptor, immunotherapy agent, and / or the like) targeting the target protein (e.g., TROP2 or PD-L1). These combinatorial features can also be defined at a global level and enable digital-based analyses that are analogous to approaches taken by pathologist whole-slide scoring, but with more granularity and reproducibility in measurement than can be achieved by pathologist scoring. These approaches may also make it possible to benchmark the performance of both pathologist scoring and digital (or computational) image analysis that approximates pathologist scoring and better understand the potential added value of complex extracted digital features in relation to clinical response.

[0467]

[0330] Global or Whole Image Features

[0468]

[0331] FIG. 23 A illustrates the use of global or whole image features to evaluate a sample. In standard clinical pathology practice, the assessment of tumor cells often involves various metrics and classifications. Positive tumor cells are those that react to a specific test or stain. The intensity of this reaction is then categorized into partitions (e.g., 1+, 2+, and 3+, and / or the like) corresponding to various levels of staining (e.g., weak, moderate, and strong staining). These intensity levels are determined by the pathologist's visual assessment, though data-driven methods like optical density (OD) quantiles may also be used to quantify staining.

[0469]

[0332] Method for determining Percent Positive Binning 2+,3+ Cutoff Points:

[0470]

[0333] In some examples, for each image (e.g., WSI), a method may compute OD values of all target protein positive tumor cells. The method may then summarize the above information to a image level statistic by estimating the 5%, 10%, 15%,. . .,90%, 95% percentile optical density (OD) values (19 data points). This operation may be performed to ensure equal contribution of

[0471] 75

[0472] NAI-5005412604vl each image to the combined histogram as described in step 3 below. In step 3, the method may concatenate all percentile values from 66 TNBC images (19*66 data points). From a histogram generated using these data points, the 1st Quartile and 3rd Quartile values are obtained. Cells with OD values above QI but less than Q3 will be considered as OD 2+, and cells with OD values above Q3 will be considered as 3+ OD. Additionally, visual approximations provided by pathologists may be used to derive l+,2+,3+ OD cutoff points to compare with quartile cutoff points derived from the TNBC dataset.

[0473]

[0334] Digital H-Score

[0474]

[0335] An H-score may be calculated, which combines both the intensity and the proportion of positive tumor cells to provide an overall score. In some examples, the empirically determined cutoff points QI and Q3 calculated above may be used and the following percentages may be calculated: A. Percentage of Tumor cells exhibiting mean membrane OD above +1, B. Percentage of Tumor cells exhibiting mean membrane OD above +2, and C. Percentage of Tumor cells exhibiting mean membrane OD above +3. A digital H-Score may be calculated as (1*A + 2*B + 3*C). Density / intensity composite features based on individual cell phenotypes may also be calculated by computing the number of target protein positive cells based on their respective percent positive score at each threshold (+l,+2,+3 defined using empirical cutoff points from QI and Q3) / Area of Tumor ROI (invasive tumor region) measured in millimeters squared.

[0475]

[0336] Local features may also be studied to assess the invasive tumor regions of interest (ROIs). In FIG. 23B, the total tissue is outlined in yellow, the carcinoma in situ is identified in blue, and the invasive tumor is identified in red. From this image, two methods can be used to study the ROIs. In a first technique, tessellated neighborhoods can be assessed where a location is broken into quadrants or sectors of a particular size of, for example, 150 to 600 microns. In a second technique, a cell-centered neighborhoods technique can be used where the radial distance from a location of interest is studied (e.g., 8 microns to 50 microns). Moreover, each cellular compartment may be assessed with either an intensity v. proportion scoring system and / or cell v. area signal summation system.

[0476]

[0337] One type of human interpretable feature (HIF) is the cell / pixel-based local spatial feature, which is centered on individual tumor cells, that may be positive or negative for the target protein (e.g., TROP2 or PD-L1) based on the target protein positive / negative model classification results. These local features include local neighborhoods defined by a radius extending from each anchor point tumor cell, for a defined set of radii. The stain signal (e.g., immunohistochemistry (IHC) signal) within the circle defined by such a radius is determined

[0477] 76

[0478] NAI-5005412604vl in one of two ways: (1) by cell-based measurements that only take into consideration signal associated with a cell that falls within the circle defined by the radius, and only includes signal assigned to the subcellular compartment of interest (e.g. “membrane,” “cytoplasm”); or (2) pixel-based summation which considers all cell-associated signal (including the 2 micron expansion from the outer membrane of all tumor cells) in the circle defined by the radius. FIG. 23C is an illustration of a target cell with a non-nuclear area as well as a 2-micron radial expansion therefrom. In some examples, the pixel information may be summarized by dividing the sum of all non-zero pixel values in the 2-micron expansion, the non-nuclear area, and the nucleus by the number of pixels comprising the nuclear region. Quantiles (e.g., 10thpercentile, 50thpercentile, 90thpercentile) in 10-%tile increments of the neighborhood-wise signal from the same image (e.g., WSI) are used as the image level summary score for the feature. Two quantitation methods were used to calculate these local spatial features: (a) DAB stain intensity normalized by DAB positive pixels only and (b) DAB stain intensity normalized by both DAB positive pixels and DAB negative pixels.

[0479]

[0338] FIG. 24 illustrates the use of features extracted via the examples of techniques described herein. In this example, whole slide image (WSI) features are derived using OD cut-points for scoring bin such as l+,2+,3+ for tumor characterization. To accurately evaluate tumor characterization, several features may be used. These include the percentage of tumor cells exhibiting different OD scores: OD+1, OD+2, and OD+3, which represent varying levels of staining intensity based on prespecified OD thresholds that correspond to a pathologist’s interpretation of visually estimated 1+, 2+, or 3+ scores, respectively. The H-Score integrates both the intensity and proportion of positively stained tumor cells, providing a comprehensive measure of tumor characterization expression. Additionally, the density of tumor cells with each OD score per square millimeter of tissue is assessed: the number of +1, +2, and +3 tumor cells per square millimeter (mm2). Together, these metrics offer a detailed overview of tumor cell expression and distribution, facilitating precise characterization and evaluation of biological (e.g., tissue samples, bodily fluids, and / or the like), wo primary methods can be utilized for determining these cut-points. The first method involves aligning with cut-points established by pathologists, ensuring consistency with clinical assessments. The second method relies on data-driven cut-points, where OD intensity is divided into quartiles within the dataset to establish thresholds. Both approaches aim to standardize scoring and enhance the accuracy of tumor characterization. Using seven variables and the two methods for cut point determination, a set of 14 WSI digital pathology features was derived, which are as follows:

[0480] Percentage of OD 1+ tumor cells based on pathologist matched cut points

[0481] 77

[0482] NAI-5005412604vl • Percentage of OD 2+ tumor cells based on pathologist matched cut points

[0483] • Percentage of OD 3+ tumor cells based on pathologist matched cut points

[0484] • H-score based on pathologist matched cut points

[0485] • Density of OD 1+ tumor cells based on pathologist matched cut points

[0486] • Density of OD 2+ tumor cells based on pathologist matched cut points

[0487] • Density of OD 3+ tumor cells based on pathologist matched cut points

[0488] • Percentage of OD 1+ tumor cells based on data driven cut points

[0489] • Percentage of OD 2+ tumor cells based on data driven cut points

[0490] • Percentage of OD 3+ tumor cells based on data driven cut points

[0491] • H-score based on data driven cut points

[0492] • Density of OD 1+ tumor cells based on data driven cut points

[0493] • Density of OD 2+ tumor cells based on data driven cut points

[0494] • Density of OD 3+ tumor cells based on data driven cut points

[0495]

[0339] FIG. 25 shows that the correlation among whole slide image (WSI) features is mostly weak to moderate when comparing pathologist-based cut points (left chart) and data-driven cut points (right chart). This correlation may be useful in identifying certain features that are strongly correlated and that might be trimmed for being redundant and / or duplicative. Due to the mostly weak-to-moderate correlation between this set of features, a method may include evaluating the full set of 14 image-scale (e.g., whole slide image (WSI) scale) features noted above.

[0496]

[0340] Local Features

[0497]

[0341] In addition to the whole slide image (WSI) features, other features may be evaluated based on a tumor cell’s local neighborhood. FIG. 26 shows an overview of certain features that may be extracted via these local techniques. Specifically, using tumor cells as anchor cells, each tumor cell’s local neighborhood may be assessed to quantify the level of target protein expression (e.g., TROP2 or PD-L1 expression) in a given local neighboring cell based on a given expansion radius surrounding each tumor cell (e.g., 8 microns to 50 microns).

[0498]

[0342] These features include the region of interest (ROI) in which to evaluate stain intensity (e.g., total tissue, invasive tumor region, in situ tumor region), the anchor cell type (e.g., target protein positive tumor cells or target protein negative tumor cells), the stain intensity quantification method (e.g., optical density-based, proportion-based, pixel-based, and / or the like), the cellular compartments being studied (e.g., membrane, cytoplasm or whole cell), the size of the neighborhood to analyze (e.g., 8 microns, 15 microns, 25 microns, or 50 microns

[0499] 78

[0500] NAI-5005412604vl for cell-anchored neighborhoods or 150, 300, 450, or 600 microns for tessellation), and the neighborhood aggregation method across the whole slide in percentiles of 10% increments (e.g., 10%, 20%,. ..80%, or 90%). A total of 2, 174 digital features were initially extracted from the image (e.g., WSI), and the Spearman correlation heatmap is presented in FIG. 27. Absolute values of pair-wise Spearman correlation were first calculated, and then for visualization purpose, continuous correlation was further categorized into 5 distinct partitions: 1) 0-0.2, 2) 0.2-0.4, 3) 0.4-0.6, 4) 0.6-0.8, and 5) 0.8-1. As shown from the heatmap, a large fraction of the features is highly correlated, forming distinct groups of redundant information. To avoid testing across highly redundant features, a curation of the set of digital parameters mentioned above was performed, looking at two main criteria: (a) eliminating digital parameters that do not actually lead to distinct patterns of scores and hence convey no independent information; and (b) eliminating features whose marginal distributions do not show biological variation across study subjects. It was hypothesized that removing highly correlated features would be beneficial because a smaller set of features would reduce the multiplicity penalty during hypotheses testing, and careful evaluation of the correlation structure among features might provide additional biological insights.

[0501]

[0343] FIG. 28 shows an example of evaluating the correlation structure of local features across cell compartments. Specifically, correlations between the cellular compartments (e.g., membrane, cytoplasm or whole cell) were studied, and heatmaps were generated to illustrate the various relationships between them. It was noted that across all subgroups, features measured from the membrane and cell compartments exhibit very high correlations (0.98- 0.99). Based on this study, the cytoplasm features and the whole cells features were eliminated.

[0502]

[0344] Similarly, FIG. 29 shows an example of evaluating the correlation structure of local features across ROIs (e.g., total tissue, invasive tumor region, invasive and DCIS region). In this example, correlation measured between regions were all above 0.96 across all subgroups, showing that it is sufficient to look at only the invasive tumor regions, for example, as the ROI. Likewise, FIG. 30 shows an example of evaluating the correlation structure of local features across different radial expansion or size of the neighborhood to analyze (e.g., 8, 15, 25, and 50 microns for cell-anchored neighborhoods or 150, 300, 450, and 600 microns for tessellation). From this analysis, a radial expansion of 25 microns was selected.

[0503]

[0345] Moreover, as previously noted, certain features were eliminated whose marginal distributions did not show biological variation across study subjects. FIG. 31 shows the histograms that summarize the distribution of how certain features, which show minimal biological variability, were removed. In particular, two graphs which show tessellated 150-

[0504] 79

[0505] NAI-5005412604vl micron invasive OD membrane 0.1 and tessellated 150-micron invasive proportion scoring membrane 0.1 were removed due to lack of variability.

[0506]

[0346] After performing this process across the feature set, certain features were eliminated. FIG. 32 shows the same feature collection that were extracted via local techniques, from FIG. 26, and shows the features, marked with an x, that were removed after performing the curation exercise. In particular, invasive tumor regions were selected as the ROI, and the regions comprising invasive, in situ, and total tissue regions were removed. Likewise, the membrane was chosen as the cellular compartment and the cytoplasm and whole cell features were removed. For radial expansion, 25 microns was selected, and the other radii were removed. Additionally, the 10%, 50% and 90% neighborhood aggregation methods were selected, and all others were removed.

[0507]

[0347] The exercise reduced the preliminary collection of features to 35 selected features and FIG. 33 shows a correlation heatmap for the 35 selected features, with continuous Spearman correlation being presented. Note, that for the global digital features, the “pathologist” prefix denotes those features created to putatively align with cut points that define 1+, 2+ or 3+ as derived externally to the study, but intended to align with how a pathologist might visually estimate these boundaries whereas the “data” prefix denotes those digital features using cut points that map to the quartiles of the empirical distribution of the target protein staining in this data set to define 1+, 2+ or 3+

[0508]

[0348] As previously noted, because the image-scale (e.g., whole slide image (WSI) scale) features showed weak to moderate correlation (See, FIG. 25), all 14 global-scale (e.g., whole image scale) features were kept: In addition to those 14 image-scale (e.g., whole slide image (WSI) scale) features, 21 local features make up the remainder of the 35 selected features. These include 7 local optical density -based digital pathology features, 8 local proportion-based digital pathology features, and 6 local pixel-based digital pathology features. Table 3 below provides a description of the 35 selected features used to determine a clinical response prediction.

[0509]

[0349] Table 3

[0510] 80

[0511] NAI-5005412604vl

[0512] 81

[0513] NAI-5005412604vl

[0514] 82

[0515] NAI-5005412604vl

[0516]

[0350] In some example embodiments, through the evaluation of the correlation structure of local features extracted across (a) different cell compartments, (b) different tissue regions of interest, and / or (c) different quantile values used to summarize neighborhood optical density (OD) based metrics on the same image (e.g., whole slide image (WSI); and through the evaluation of the biological variation of the features via histogram and % coefficient of variation analysis; an initial set of 2,174 features was reduced to the set of 35 features shown in Table 3, which includes 14 global or whole image-scale features, and 21 local features. Such feature reduction was blinded to both the clinical outcome and the pathologist scoring coming from the clinical data set.

[0517]

[0351] In some example embodiments, the reduced set of features, such as the suite of 35 features shown in Table 3, may form a baseline feature set that can undergo statistical clinical outcome association evaluation and independent validation for use in digital pathology analysis workflow for specific biomarkers (e.g., target proteins), indications, and / or therapies (e.g., antibody drug conjugate). For example, in some cases, the suite of baseline features may be analyzed while blinded from clinical outcome such that only the values of the individual

[0518] 83

[0519] NAI-5005412604vl features are passed to the next stage. In some cases, clinical outcome may be unblinded in the next stage to guide the selection of a subset of features from the baseline feature set as well as the discovery of additional de novo features, for example, using a feature extraction model (e.g., a vision foundation model (VFM)). The combination of the subset of features from the baseline feature set and the additional de novo features may form a prioritized feature set. In some cases, the prioritized feature set may undergo independent validation using holdout samples. For example, the performance of a response prediction model (e.g., a multivariate model) trained to use the prioritized feature set to determine a response prediction may be validated using patient samples not included in the training dataset used to train the response prediction model.

[0520]

[0352] The 35 selected features are provided herein by way of example and are not intended to be limiting of the scope of this disclosure. For example, the selected feature set may include more than 35 or less than 35 features. Certain features that are noted may be optional, and other features may be added to this selected set of features. Additionally, where a high correlation is present, it may be possible to substitute or select one variable instead of others (e.g., to select invasive + DCIS or all regions instead of invasive only regions). Notably, pathologist-based IHC scores may be used for comparison.

[0521]

[0353] FIGS. 34-39 are representative images from certain images (e.g., whole slide images (WSIs)) used to confirm the proper extraction of features. For each image (e.g., WSI), a score was calculated, and a pathologist was asked to confirm whether the image was consistent with that score. FIG. 34 shows an example of images filtered by density of OD 3+ tumor cells based on data driven cut points: high (left panel) vs low (right panel) values. FIG. 35 shows an example of images (e.g., WSIs) filtered by 90thpercent quantile of the neighborhood- wise average target protein DAB intensity proportion values of tumor cells inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive ROI: high (left panel) vs low (right panel) values. FIG. 36 shows an example of images filtered by 50thpercent quantile of the neighborhood-wise average target protein DAB OD values of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive ROI: high (left panel) vs low (right panel) values. FIG. 37 shows an example of images filtered by 10thpercent quantile of the neighborhood-wise average target protein DAB intensity at the pixel level (including DAB positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive ROI: high (left panel) vs low (right

[0522] 84

[0523] NAI-5005412604vl panel) values. FIG. 38 shows an example of images filtered by 50thpercent quantile of the neighborhood-wise average target protein DAB intensity at the pixel level (including DAB positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive ROI: high (left panel) vs low (right panel) values. FIG. 39 shows an example of images filtered by 90thpercent quantile of average target protein DAB OD among the tessellated regions defined by 150um * 150 um: high (left panel) vs low (right panel) values. In these examples, the color outlines represent pathologist annotations for different ROIs and are not used for evaluating local feature extraction performance.

[0524]

[0354] FIG. 40 shows a comparison of the image-scale (e.g., whole slide image (WSI) scale) feature capturing the density of OD 3+ tumor cells based on pathologist-matched cut-point, versus each of the 21 local features. As seen from these comparisons, local features are moderately positively correlated with the global feature. However, in the lower range of the global feature score, there are considerable variabilities in the local features. Variability of local features vs. global features may suggest opportunity for additional information leveraged by the digital approach.

[0525]

[0355] Based on the 35 selected features, an analysis was conducted with certain objectives as shown below. Letting “objective response” represent best overall response (BOR) of complete or partial response (complete response (CR) or partial response (PR), respectively), the present analyses had seven primary objectives and four secondary objectives. The primary objectives and their associated hypothesis were as follows: 1) To test whether 14 image-scale (e.g., whole slide image (WSI) scale) features are separately associated with objective response to sacituzumab tirumotecan monotherapy, with the hypothesis that each feature will be positively associated with improved objective response to sacituzumab tirumotecan monotherapy, 2) To test whether 7 local optical density -based features are separately associated with objective response to sacituzumab tirumotecan monotherapy, with the hypothesis that each feature will be positively associated with improved objective response to sacituzumab tirumotecan monotherapy, 3) To test whether 8 local proportion-based features are separately associated with objective response to sacituzumab tirumotecan monotherapy, with the hypothesis that each feature will be positively associated with improved objective response to sacituzumab tirumotecan monotherapy, 4) To test whether 6 local pixel-based features are separately associated with objective response to sacituzumab tirumotecan monotherapy, with the hypothesis that each feature will be positively associated with improved objective response to

[0526] 85

[0527] NAI-5005412604vl sacituzumab tirumotecan monotherapy, 5) To test whether pathologist-based immunohistochemistry (IHC) scores are separately associated with objective response to sacituzumab tirumotecan monotherapy, with the hypothesis that higher pathologist-based immunohistochemistry (IHC) score will be associated with improved objective response to sacituzumab tirumotecan monotherapy, 6) Within families of digital features identified in Objectives 1-4, to descriptively evaluate the correlation structure between digital features and pathologist-based immunohistochemistry (IHC) scores, and 7) To descriptively explore the digital features for their association with baseline prognostic variables, such as prior therapy, dose administered, and baseline tumor size.

[0528]

[0356] In addition to these seven primary obj ectives, four secondary obj ectives were identified as follows: 1) To test whether any of the statistically significant local and global (or image scale) digital feature associations identified above are independent of the predictive effects of pathologist scored immunohistochemistry (IHC) staining in a multivariate model, 2) To construct a multivariate model, digital pathology-based signature for measuring target protein expression for future validation in an independent testing set along with objective assessments of the accuracy of any proposed model using cross-validation methods, 3) To examine and compare the clinical utility (including positive predictive value / negative predictive value of predicting overall response (OR)) of statistically significant digital feature associations identified above and pathologist scored IHC, 4) To test whether there is an improvement in area under the receiver operating characteristic (AUROC) curve of each statistically significant digital pathology feature identified above for objective response to sacituzumab tirumotecan monotherapy relative to the most associated pathologist IHC measure.

[0529]

[0357] Experimental Examples

[0530]

[0358] Images of biological samples, in this instance whole slide images (WSIs) of tumor tissue samples, were evaluated for adequacy (minimum 1 mm2of evaluable tumor area and minimum 100 viable tumor cells, spanning a contiguous analyzable region of the whole slide image (WSI) without obscuring artifacts); one sample was excluded as it did not meet minimum size criteria for adequacy. The remaining samples were annotated by a breast pathologist to identify analyzable tissue regions, invasive carcinoma regions, and DCIS (if present), using criteria agreed upon by a team of pathologists and biometrics scientists. The total analyzable tissue encompassed the tumor areas and adjacent fibrotic stroma and normal tissue entrapped or involved by the carcinoma, as per the pathologist’s judgment. Distant normal breast tissue, other uninvolved normal structures, or adipose tissue uninvolved by carcinoma were not included. Areas of artifact (e.g., tissue folds, air bubbles, optical

[0531] 86

[0532] NAI-5005412604vl distortions, necrotic tissue) were excluded. The invasive carcinoma region was annotated as a subset of the total analyzable tissue region based on morphologic evidence of invasion. Glands involved by DCIS were annotated if present; DCIS was defined according to standard morphologic criteria as an intraductal neoplastic process spanning at least two duct spaces and at least 2 mm in greatest linear extent, or an intraductal lesion of any size with intermediate- or high-grade nuclear atypia. The determination of an intraductal process was made on morphologic evidence for a myoepithelial cell layer; if there was doubt as to the nature of the process (intraductal v. invasive), it was not annotated as DCIS. Other ancillary findings (e.g., atypical ductal hyperplasia, lobular neoplasia), if present, were not specifically annotated.

[0533]

[0359] An analysis was conducted using the subset of patients (N = 58 with both digital images and IHC data passing quality control). All patients included in this analysis were enrolled in the study long enough to allow for at least two post-baseline efficacy scans to determine confirmed best overall response (BOR). Baseline characteristics of the analysis population are discussed with reference to FIG. 41. As shown, a subset of N = 59 subjects was expected to have digital image, immunohistochemistry (IHC), and clinical data available for analysis (approximately 92% of the study population).

[0534]

[0360] FIG. 41 compares the distribution of various characteristics in the sacituzumab tirumotecan TNBC treated study population (N = 64) vs. the subset of TNBC patients expected to have both digital image and IHC data available (N = 59). In general, the subset of patients with digital image and IHC data available to date appears to be representative of the overall treated study population with respect to the characteristics shown in the table.

[0535]

[0361] Statistical Methods

[0536]

[0362] For the primary objectives, the association of biomarkers (image-scale features, local optical density (OD) based features, local proportion-based features, local pixel-based features, pathologist-based IHC score) with best overall response (BOR) of complete response (CR) or partial response (PR) was evaluated separately using logistic regression and the AUROC with corresponding 95% confidence interval (CI). Logistic regression models included a term for Eastern Cooperative Oncology Group (ECOG) status and the biomarker: one of 14 image-scale (e.g., whole slide image (WSI) scale) features (continuous), one of 7 local optical density -based features (continuous), one of 8 local proportion-based features (continuous), one of 6 local pixel-based features (continuous), or pathologist-based IHC scores (continuous). Both nominal and adjusted p-values were reported for testing related to the primary objectives, with multiplicity adjustment as determined by the families defined above. For the secondary objectives, joint logistic modeling of any local or image-scale (e.g., whole slide image (WSI)

[0537] 87

[0538] NAI-5005412604vl scale) feature found to significantly associate with clinical outcome and the pathologist-based IHC score was conducted to assess whether digital feature scores were independent of the predictive effects of the pathologist-based IHC score, and the nominal p-values may be reported for the digital and pathologist features of interest in the multivariable models.

[0539]

[0363] Other digital features whose association with clinical outcome is also high, similar to the statistically significant after applying multiplicity correction (Primary Objectives 1-4), may be good candidates for inclusion in a multivariate, digital pathology-based signature for measuring target protein expression. Candidate prediction models may be developed (including use of cross validation to get an objective assessment of their predictive ability) and may be validated using independent datasets.

[0540]

[0364] Multiplicity

[0541]

[0365] The primary objective hypotheses were tested within the 5 families of scores identified above (and within any set of models evaluating covariate sensitivity) at a one-sided alpha of 0.05 adjusted by the Hochberg step-up procedure. Nominal p-values were reported for the secondary objective around assessment of independent associations for digital features with clinical outcome when adjusting for pathologist features; these analyses were considered exploratory and expository and no multiplicity adjustment was used. Both nominal and multiplicity-adjusted p-values were reported for the secondary objective around testing superiority of area under the receiver operating characteristic (AUROC) curve for each digital feature vs. pathologist score (one-sided alpha of 0.05), with the degree of multiplicity correction dependent on the number of statistically significant digital feature scores being evaluated as result of the primary testing of families.

[0542]

[0366] Subsequent to the formal testing on the limited feature set, this cohort was used to train a multivariate predictive digital model yielding a composite score to be prospectively validated for its predictive ability using independent datasets. It is also possible that the key features driving the predictive model may be relevant to other anticancer agents and other IHC assays targeting other proteins expressed by tumor cells.

[0543]

[0367] Analysis of HIF Feature and Response

[0544]

[0368] FIG. 42 is a scatterplot showing the relationship between human interpretable features (HIFs) and the most predictive digital feature identified among the 35 selected features, labeled by patient best overall response (BOR) status. In this scatterplot, each circle indicates a nonresponder, while each triangle indicates a responder. From this scatterplot, we can see that a percentage of patients are non-responders. Moreover, the x-axis represents the pathologist score percent and the y-axis is the most predictive prespecified digital feature identified by the

[0545] 88

[0546] NAI-5005412604vl 35 selected features, which has an AUROC of approximately 0.76, and there is a modest correlation between the two. It can be seen that when the pathologist score is relatively small, there is a high likelihood of nonresponse. However, when the pathologist score is relatively high and discrete, the digital feature shows a range of target protein quantification within the discrete values of the pathologist score and further segregates the responders towards a higher range of that quantification, leading to an improvement in AUROC over the target protein expression quantified by the pathologist score. When the two features of FIG. 42 were jointly modeled, the most predictive digital feature retained statistical significance (p = 0.012) while the pathologist score did not (p = 0.455).

[0547]

[0369] To potentially improve on any findings coming out of the 35 originally selected features, the original collection of 1,831 digital features were again examined as a full training set of features for their association with response. FIG. 43 shows a portion of the evaluation of the independent statistical significance (relative to a model incorporating the information from the most predictive of the 35 selected features) on the complete set of 1,831 digital features. In this image, the “mainEffectP” is the univariate significance of the feature in predicting best overall response (BOR), and the 2DF P is the statistical significance of these features after adjusting for the most predictive prespecified digital feature among the 35 selected features, (i.e., ANOVA p-value that compares the full model with the stated feature and the prioritized feature vs. reduced model with only the most predictive of the 35 prioritized feature). Features that passed FDR at the main effect level, do not show independent statistical significance after adjusting for the prespecified feature in the joint model. Hence, it may point to the importance of the original down-selection of features via triaging the correlation structure and using subject matter input to prioritize the 35 features.

[0548]

[0370] In addition to looking at the statistical inference or the p-value based analysis described in FIG. 43, FIG. 44 shows the distribution of the AUROC of the complete set of 1,831 features. The histogram shows roughly 10% out of 1,831 features have an AUROC larger than the AUROC of the most predictive feature in the 35 selected feature set. In addition, it was observed that most of the higher AUROC features are cytoplasmic target protein positive neighborhood features. These features were previously trimmed due to their correlation with the membrane features.

[0549]

[0371] FIG. 45 illustrates an evaluation of the relationship of ROI / cell compartment / radius / quantile and the univariate AUROC in response prediction for positive cell neighborhood features. In this illustration, the area under the receiver operating characteristics (AUROC) curve is shown with the more intense gray color being more

[0550] 89

[0551] NAI-5005412604vl predictive of clinical response. A consistent trend that was observed in the data set is the cytoplasmic features tend to be more predictive of clinical response than the membrane-only or whole cell features. FIG. 46 shows a similar exercise being carried out for negative cell neighborhood features.

[0552]

[0372] Data Driven De Novo Imaging Feature Discovery Approach

[0553]

[0373] In addition to the HIF approach, other techniques may be collected from an image and used in feature extraction module 110. For example, in the data driven de novo imaging feature discovery approach, pretrained VFMs may be leveraged to extract imaging features directly from the image (e.g., WSI) through model embedding. A VFM is a machine learning framework specifically designed for digital pathology that is trained on a vast amount of publicly available pathology images, allowing it to extract the most relevant features from histological data. Due to the extensive volume of training data, the VFM becomes a powerful tool for feature extraction in pathology and is effective in identifying key patterns and details within histology images. One or more VFMs may be used to extract features. Generally, the image (e.g., WSI) may be divided into a group of tiles (e.g., 100 microns in width and 100 microns in length), and within each tile, a digital pathology VFM may be deployed to extract numeric feature vectors. These numeric feature vectors were subsequently clustered for dimension reduction, and the most representative tiles from each cluster were visually inspected by pathologist for interpretation of biological significance. Finally, only those feature clusters that are potentially biologically meaningful were included in the multivariate model feature set.

[0554]

[0374] Two basic approaches may be used to identify the tiles of interest on each image. In a first approach, the targeted approach, tiles are selected where each tile has a target protein positive tumor cell in the middle, and these tiles are passed to the VFM based feature extractor. As previously noted, pathologist prespecified digital features do not provide good predictive power when the target protein (e.g., TROP2 or PD-L1) expression is high (See, FIGS. 47A- 47B). However, digital pathology VFMs can extract additional morphological features from patches that might help to further classify the responders from non-responders in the high target protein (e.g., TROP2 or PD-L1) patient category. In some embodiments, digital pathology VFM can also be modified as a cell segmentation model to further identify the microenvironment of certain regions if they are considered more related with response / non- response.

[0555]

[0375] The overall workflow for using VFM to aid in clinical response prediction with de novo image features is shown in FIG. 48. As shown, the image (e.g., WSI) is split into overlapping

[0556] 90

[0557] NAI-5005412604vl tiles in targeted or non-targeted approach, and the VFM is deployed to individual tiles to extract the numeric features that can be operated downstream for statistical analysis in an embedding procedure. For the targeted approach, there may be overlapping tiles due to the positions of the cells relative to one another. Thus, an optional step may be employed to clean up the redundant tiles and only ensure that the non-redundant tiles are extracted from the overlapping target protein positive tumor cell clusters. FIG. 49 shows representative images with annotation to show the original image (e.g., WSI), all the target protein positive cells and the sampled target protein positive cells (non-redundant neighborhood patches). In some examples, redundant tiles are those whose areas overlap more than a predetermined area threshold (e.g., more than 80%, more than 85%, more than 90% or more than 95% overlapping). Once those numeric features are extracted from the images then clustering may be performed to reduce dimension of these features into a smaller set of clusters. In some examples, pathologists may be consulted to visualize the most representative patch from each cluster to understand if there is sufficient biological meaning or significance of the identified clusters from these embeddings. Finally, it may be possible to correlate these clusters of features with clinical response to further understand which imaging features based on the clustering are associated with clinical response to sacituzumab tirumotecan.

[0558]

[0376] In some examples, it is possible for VFMs to transfer patches into embeddings, reduce the dimension of the embeddings with Principal Component Analysis (PCA) and UMAP and cluster the reduced embeddings using the Leiden cluster algorithm. FIG. 50 shows the UMAP plot of all embeddings on the left and the Leiden clustering of 2D points on the right. From these clusters, it may be possible to calculate the proportional table for each patch in the image (e.g., whole slide image) against the clusters as shown in FIG. 51. Based on the proportional table, it is possible identify clusters that are most associated with response / non-response and further examine the patches patterns for that cluster. One promising cluster that was identified from the targeted approach is the lack of lymphocytes in a region is negatively correlated with non-response. Briefly, immune cell desert in the stromal region were identified as an imaging feature that is associated with non-responders in the target protein positive population. In the plot shown in FIG. 52, target protein positive TC neighborhood cell-based intensity 50thpercentile summary score is provided on the y-axis (beginning at 170) and the x- axis represents the cluster frequency for the lack of TIL stromal region frequencies.

[0559]

[0377] The second approach, the untargeted approach, is illustrated in FIG. 53. In this approach, the entire image (e.g., WSI) is cropped into non-overlapping tiles based on sliding windows (e.g., 100 microns by 100 microns), regardless of whether the tile contains a target protein

[0560] 91

[0561] NAI-5005412604vl positive cell. In some examples, it may be possible to digest the TNBC patients' tumor microenvironment and their compositions into the unique morphological clusters / prototypes that define that microenvironment. For example, an image (e.g., WSI) can then be summarized as being composed of 20% TILs, 30% Immune Clusters, and 50% Stroma. By examining these different findings, it may be possible to find morphological prototypes that could be associated with target protein response, which could provide supplementary information to the pathologist features.

[0562]

[0378] Training the De Novo Model

[0563]

[0379] The De Novo model was trained on a dataset from clinical trial protocol images (e.g., WSIs) that passed first line of Image QC (N = 231). This set was chosen due to having more samples, which allowed the model to capture wider range of morphologies. Almost 1 million patches were collected and used to generate prototype clusters. In this example, 40 initial clusters were identified, and these were expanded / collapsed according to certain criteria. First, patches assigned to each discovered cluster are visually examined and annotated, providing interpretability. If patches within the same cluster are considered heterogeneous (high intervariability), the cluster is broken into subclusters. If patches coming from different clusters are very similar (low intra-variability), clusters are merged into a single cluster. This merging / splitting of clusters is repeated until clustering performance generalizes across images (e.g., WSIs). In some examples, a machine learning technique, such as few shot learning and / or prototype-based approach (requiring 3-5 prototype tiles per phenotype class), may be taken to fine tune the VFM for each tumor type. These imaging features may then be clustered for dimension reduction. The feature clusters may be visually inspected by pathologists for potential biological significance, and those feature clusters that passed inspection were subsequently associated with clinical response for clinical utility assessment.

[0564]

[0380] FIGS. 54-58 show representative image tiles of various clusters. Specifically, FIG. 54 shows a collection of lymphocyte clusters, FIG. 55 shows a collection of target protein negative TCs, FIG. 56 shows a collection of lymphocyte and stromal co-localized clusters, FIG. 57 shows target protein positive TCs with stromal + lymphocytes and FIG. 58 shows target protein positive TCs. Ten different clusters were identified as shown in FIG. 59:

[0565] • target protein positive Tumor Cells (90% TC’s)

[0566] • target protein positive TC’s with stromal cells

[0567] • target protein positive TC’ s with Immune Cells (TIL’ s)

[0568] • target protein positive TC’s with stromal and immune cells

[0569] 92

[0570] NAI-5005412604vl • target protein negative TC’s

[0571] • Stromal Cells

[0572] • Immune Cells

[0573] • sTILs

[0574] • QC and irrelevant patches (Necrosis, region with very few cells)

[0575] • Unclear Class (Mostly QC regions mixed in with some other patches)

[0576] Note that some of these clusters may be further compartmentalized as needed. Additionally, for each patient, it may be possible to generate a pie-chart or other graphical aid to understand the distributional frequency differences of these ten clusters from patient to patient, and that may be used to examine a potential clinical response.

[0577]

[0381] Each cluster can be associated with a particular color, and the clustering can be visualized based on the coloring scheme of FIG. 59. FIGS. 60A-60B are representative images showing the original image on the left and the overlaid cluster representation on the left for two images (e.g., WSIs) for a non-responder (FIG. 60A) and a responder (FIG. 60B). Notably, both approaches, the targeted and untargeted, may be used to identify sTIL cluster abundance to be an imaging feature that is associated with improved clinical response.

[0578]

[0382] Association Between Stromal Tumor-Infiltrating Lymphocytes (sTIL) and best overall response (BOR)

[0579]

[0383] Stromal tumor-infiltrating lymphocytes (sTILs) are a specific type of tumor-infiltrating lymphocytes (TILs) found within the stroma of a tumor, rather than intermixed among the tumor cells themselves, and clusters of sTILs may correlate with clinical response. FIGS. 61- 62 illustrate formally testing the association of the sTIL cluster digital feature and clinical outcome using statistical modeling, and testing the unique or independent prediction power feature sTIL cluster digital feature offers independent of the most predictive feature that was previously validated in the 35 selected HIF set. FIG. 61 is a scatter plot between the sTIL feature and the most significant pathologist-prespecified HIF, and the Spearman correlation indicates that there is low correlation between the sTIL feature and the local target protein (e.g., TROP2 or PD-L1) abundance feature. FIG. 62 illustrates a more formal evaluation of joint model significance of the most significant pathologist-prespecified HIF and sTIL into the same model. Thus, the sTIL feature was considered as a candidate feature for a subsequent multivariate model based on these findings.

[0580]

[0384] Summary of Feature Extraction Module

[0581] 93

[0582] NAI-5005412604vl

[0385] As described above, a feature extraction module 110 (152 in FIG. IB) may include various categories of digital features, and may include for example, the HIFs (e.g., the set of the pathologist-prespecified digital image-scale (e.g., whole slide image (WSI) scale) features or local features), global spatial heterogeneity features and clinical response-guided de novo features discovered from the IHC images directly. The set of HIFs were discussed as a large collection (e.g., 1,831 features), which were then curated to a smaller set of 35 features based on independent prediction and biological variability. Both univariate statistical significance (p- value) and prediction performance (AUROC) suggest that certain features do not provide significant orthogonal signals to the 35 selected features. However, through in-depth exploratory analysis, it was discovered that cytoplasmic features, although highly correlated with membranous features, appear to consistently have a high prediction performance. Thus, the cytoplasmic features were selected to be included in a subsequent multivariate model’s candidate input feature set. The global spatial heterogeneity features of the tumor microenvironment between target protein positive and target protein negative tumor cells may be summarized using two spatial metrics: (1) the Morisita-Horn index, which quantifies the colocalization of target protein positive and target protein negative tumor cells, and (2) the percent target protein positive / target protein negative hotspot overlap index, which quantified the colocalization of target protein positive and target protein negative tumor cell clusters. Initial analysis indicates that there appears to be an interaction between the local membrane target protein positive OD feature (most significant feature from the selected features) and the Morisita-Horn index. More specifically, among patients with high local membrane target protein positive OD, subjects with a lower Morisita-Horn index are more likely to be responders to the treatment (AUROC = 0.77). FIG. 63 is a scatter plot between Morisita-Horn Index (x-axis) and the top local target protein OD feature identified from the prespecified feature analysis. Subjects are color coded by the best overall response (BOR) responder status (0 and circles indicating non-responder, 1 and triangles indicating responder). The horizontal reference line corresponds to the Youden index of the top feature from the clinical study data analysis. Subjects above the line may correspond to the high target protein population, and subjects below the line correspond to the low target protein population. Though the sample size is small, it appears that the spatial heterogeneity feature might offer additional information to the local target protein abundance features.

[0583]

[0386] Feature Selection and Multivariate Model Construction

[0584]

[0387] A combination of the three categories of features (e.g., HIFs or the set of the pathologist-prespecified digital features or local features, global spatial heterogeneity features,

[0585] 94

[0586] NAI-5005412604vl and clinical response-guided de novo features such as sTIL features discovered from the IHC images directly) may be used to construct the pathological analysis engine 114 (e.g., the response prediction model 154 or the end-to-end model 156) for the prediction of clinical response to a treatment, such as an antibody-drug conjugate (ADC), alkylating agent, antimetabolite, enzyme inhibitor, anti-angiogenesis agent, microtubule disruptor, immunotherapy agent, and / or the like. In some examples, at least one feature from one category of features is used to predict a clinical response. In some examples, at least one feature from any two categories of features may be used to predict a clinical response. In some examples, at least one feature from each category of features may be used to predict a clinical response.

[0587]

[0388] In some examples, all three categories of features: local target protein positive abundance digital features, spatial heterogeneity features, and sTIL were combined into one collective feature set and were fed into machine learning models such as elastic net / LASSO to construct a multivariate best overall response (BOR) prediction model. With the caveat of a small sample size, it was possible to identify a multivariate model including features from each of the three feature categories through cross validation. The cross validated AUROC of the final model is 0.8. In some examples, the machine learning models used for the prediction model include random forest, boosting, and elastic net.

[0588]

[0389] In some examples, the composite multivariate model includes 27 input features, which represent a combination of pathologist-prespecified digital features, global spatial heterogeneity features, and clinical response-guided de novo features. FIG. 64 shows these 27 features, which are as follows:

[0589] 95

[0590] NAI-5005412604vl

[0591] 96

[0592] NAI-5005412604vl

[0593] 97

[0594] NAI-5005412604vl

[0595]

[0390] In this example, the model includes the sTIL feature from the de novo feature set. It also includes pathologist-prespecified features (features 2 to 12) and then a collection of spatial heterogeneity features (features 13 to 27) that correspond to the Hom index as well as the hotspot features. FIGS. 65-66 illustrate LASSO regression coefficients on mean centered features and a test set AUROC v. Lambda to under the predictive power of certain features. By evaluating the magnitude of the LASSO regression coefficient, features from all three categories of features were selected in the final model. In addition, the two local features have the strongest power to predict clinical response (i.e., the largest regression coefficient), with the sTIL feature and several spatial heterogeneity features having somewhat lower predictive power.

[0596]

[0391] Triple Negative Breast Cancer (TNBC)

[0597]

[0392] A sample set of n=58 TNBC from MK -2870-001 was utilized for digital pathology proof of concept development given the largest availability (at the time) of digital images,

[0598] 98

[0599] NAI-5005412604vl MEDx IHC scores, mature response data, and an association between pathologist immunohistochemistry (IHC) scoring and clinical response. Data from other MK-2870-001 cohorts (other tumor types) had fewer n (or datapoints) and less mature response data at the time of initiation of this effort. 35 HIF features (Table 3), organized into five feature families, were evaluated against clinical outcome.

[0600]

[0393] Table 4 below summarizes the statistical significance of the association between biomarkers and clinical outcomes measured using best overall response (BOR) based on best objective response per RECIST (Response Evaluation Criteria in Solid Tumors) for the local optical density based digital feature family. These biomarkers include the 10thand 50thquantiles of the neighborhood cell based intensity of target protein positive tumor cells (TCs), the 10thand 50thquantiles of the neighborhood cell based intensity of target protein negative tumor cells (TCs), and the 10th, 50th, and 90thquantiles of the 150mregional intensity of the target protein. The significance of association is quantified using multiplicity adjusted P-value, which is the probability that the observed difference between a responder group and a nonresponder group occurred by random chance if the null hypothesis (no real difference) were true, after adjusting for the number of tests within each feature family. A low p-value, such as one below 0.05, indicates a statistically significant result that suggests the difference between the responder group and the non-responder group is attributable to the digital feature and not just random chance. As shown in Table 4, the 50thquantile of the neighborhood cell based intensity of target protein positive (e.g., Target Protein+) tumor cells (TC) demonstrated statistically significant association (adjusted p-value <0.05) with the clinical outcome (e.g., best overall response (BOR)) of the MK-2870-001 TNBC cohort.

[0601]

[0394] Table 4

[0602] 99

[0603] NAI-5005412604vl

[0395] After the completion of testing all 35 HIFs, having demonstrated proof-of-concept that the digital pipeline generates features that may lead to improved associations relative to pathologist scoring, the AI / ML team was un-blinded to clinical outcome as a training set to further examine where some of the strongest associations may be generated, and selected a subset of seven prioritized HIFs that demonstrated the strongest prediction of clinical outcome (by AUROC value and the associated 95% confidence intervals). During this stage of the analysis, VFM based de novo feature discovery was also performed as described previously. Two additional features were identified through the de novo feature discovery: stromal TIL, and a multivariate model including sTIL.

[0604]

[0396] Table 5 below depicts a summary of the AUROC associated with these seven prioritized HIFs extracted from the MK-2870-001 TNBC cohort dataset using the deep learning enabled pathological analysis workflow described herein.

[0605]

[0397] Table 5

[0606]

[0398] In the next stage of the analysis, the seven prioritized HIFs and de novo features were evaluated against clinical outcomes (both BOR and PFS) in an independent TNBC cohort from MK-2870-001 (N=34) for signal confirmation.

[0607]

[0399] Tables 6-7 below depict summaries of the AUROC of the seven prioritized digital pathology features and the de novo features for BOR prediction in the independent signal confirmation cohort. These results confirm the robustness and the clinical utility of the prioritized digital pathology features and the de novo features.

[0608]

[0400] Table 6

[0609] 100

[0610] NAI-5005412604vl

[0611]

[0401] Table ?

[0612]

[0402] Tables 8-9 below depict summaries of the Harrel’s C indices of the seven prioritized digital pathology features and the de novo features for PFS prediction in the independent signal confirmation cohort.. These results confirm the robustness and the clinical utility of the prioritized digital pathology features and the de novo features.

[0613]

[0403] Table 8

[0614]

[0404] Table 9

[0615]

[0405] As illustrated above, the analysis results on these additional 34 independent TNBC holdout samples confirmed the association between clinical response and both the seven prioritized HIFs as well as the VFM based de novo features.

[0616]

[0406] It is to be understood that the embodiments described herein are merely illustrative of the principles and applications of the present disclosure. It is therefore to be understood that numerous modifications may be made to the illustrative embodiments and that other

[0617] 101

[0618] NAI-5005412604vl arrangements may be devised without departing from the spirit and scope of the present disclosure as defined by the appended claims.

[0619]

[0407] It will be appreciated that the various dependent claims and the features set forth therein can be combined in different ways than presented in the initial claims. It will also be appreciated that the features described in connection with individual embodiments may be shared with other of the described embodiments.

[0620] 102

[0621] NAI-5005412604vl

Claims

1. CLAIMS1. A computer-implemented method, comprising: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; and identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction.

2. A method of treating cancer involving administering to a cancer patient a therapeutic that is useful for treating cancer, the method comprising: applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image;103NAI-5005412604vlapplying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic; identifying, using the one or more processors, the patient as eligible for receiving the therapeutic based at least on the response prediction; and administering the therapeutic in response to the patient being identified as eligible for receiving the therapeutic.

3. A method of treating cancer involving administering to a cancer patient a therapeutic that is useful for treating cancer, the method comprising: administering the therapeutic to a patient in response to the patient being identified as eligible for receiving the therapeutic, where the patient is identified as eligible for receiving the therapeutic based at least on a response prediction of the patient, and where the response prediction of the patient is determined by at least applying, using one or more processors, one or more segmentation models to segment an image depicting a biological sample associated with the patient, where the one or more segmentation models are trained to generate a cell map localizing one or more cells present in the image, where each pixel in the cell map is assigned a value indicative of whether a corresponding pixel in the image depicts a cell; applying, using the one or more processors, a feature extraction model to determine a feature set containing one or more features present in the image, where the feature extraction model is trained to determine, based at least on the cell map, one or more primitive features, where at least one primitive feature is indicative of a presence or absence of a target protein, and where the feature extraction model is further trained to generate, for inclusion in the feature set, at least one higher order feature derived from the one or more primitive features present in the image; and applying, using the one or more processors, a response prediction model to determine, based at least on the feature set, a response prediction for a patient associated with the biological sample, where the response prediction indicates a likelihood of the patient responding to the therapeutic.

4. The method of any of claims 1 to 3, wherein the image is of a whole slide image (WSI).

5. The method of claim 4, wherein the target protein is TROP2 or PD-L1.104NAI-5005412604vl6. The method of any of claims 4 to 5, wherein the therapeutic is an antibody drug conjugate (ADC), an alkylating agent, an antimetabolite, an enzyme inhibitor, an antiangiogenesis agent, a microtubule disruptor, and / or an immunotherapy agent.

7. The method of any of claims 1 to 6, wherein the biological sample depicted in image is treated with a stain sensitive to the target protein.

8. The method of claim 7, wherein the biological sample depicted in the image is also treated with an additional stain sensitive to at least one other subcellular compartment.

9. The method of any of claims 1 to 8, wherein the one or more primitive features include a location or a type of a cell present in the image or a region of interest (ROI) within the image.

10. The method of any of claims 1 to 9, wherein the one or more primitive features include a completeness of a target protein sensitive stain around a cell present in the image or a region of interest (ROI) within the image.

11. The method of claim 10, wherein the completeness of the stain is quantified by a proportion of (i) a first quantity of pixels depicting a membrane of the cell with the target protein sensitive stain and (ii) a second quantity of pixels depicting the membrane of the cell without the target protein sensitive stain.

12. The method of any of claims 1 to 11, wherein the one or more primitive features include a location of a target protein sensitive stain around a cell present in the image or a region of interest (ROI) within the image.

13. The method of claim 12, wherein the location of the target protein sensitive stain comprises one or more cellular compartments of the cell in which the target protein sensitive stain is present.

14. The method of any of claims 1 to 13, wherein the one or more primitive features include a morphology of a cell present in the image or a region of interest (ROI) within the image.

15. The method of claim 14, wherein the morphology of the cell comprises a differentiability of one or more cellular compartments of the cell.

16. The method of any of claims 1 to 15, wherein the feature extraction model determines the one or more primitive features by at least: determining, for each cell present in the cell map, whether an area around the cell exhibits a signal from a protein sensitive stain sensitive to a target protein of the therapeutic; classifying the cell as negative for the target protein where the signal from the protein sensitive stain is absent from the area around the cell;105NAI-5005412604vlwhere the area around the cell exhibits the signal from the protein sensitive stain, classifying the cell as having cytoplasmic only staining where the cell lacks a differentiable cell membrane, and where the area around the cell exhibits the signal from the protein sensitive stain and the cell exhibits the differentiable cell membrane, classifying the cell as having membranous staining where the cell membrane is differentiable from a nuclear boundary of the cell, and classifying the cell in a gray zone where the cell membrane is indifferentiable from the nuclear boundary of the cell.

17. The method of claim 16, wherein the signal from the protein sensitive stain indicates an intensity of the staining.

18. The method of any of claims 16 to 17, wherein a differentiability of a cell membrane of the cell is determined based at least on a signal from a stain sensitive to a cytoplasm and extracellular matrix.

19. The method of claim 18, wherein a differentiability between the cell membrane of the cell and the nuclear boundary of the cell is determined based on the signal from the stain sensitive to the cell membrane and a stain sensitive to a nucleus of the cell.

20. The method of any of claims 1 to 19, wherein the at least one higher order feature includes a regional feature localized to a portion of the image.

21. The method of claim 20, wherein the regional feature comprises a tessellated feature extracted from a portion of the image.

22. The method of any of claims 20 to 21, wherein the portion of the image comprises a tessellation that is 150 microns to 600 microns in size.

23. The method of any of claims 20 to 22, wherein the regional feature comprises a local feature extracted from a cell-anchored neighborhood defined by a radius around a cell.

24. The method of claim 23, wherein the cell-anchored neighborhood is defined by a radius of 8 microns to 60 microns.

25. The method of any of claims 20 to 24, wherein the at least one higher order feature includes an additional regional feature having a different scale or resolution.

26. The method of claim 25, wherein the additional regional feature comprises an additional tessellated feature extracted from a different size portion of the image.

27. The method of any of claims 25 to 26, wherein the additional regional feature comprises an additional local feature extracted from a cell-anchored neighborhood defined by a different radius around the cell.106NAI-5005412604vl28. The method of any of claims 20 to 27, wherein the regional feature includes a distribution of cells positive for the target protein of the therapeutic within the portion of the image.

29. The method of any of claims 20 to 28, wherein the regional feature includes an intensity of a stain sensitive to the target protein within the portion of the image.

30. The method of any of claims 20 to 29, wherein the regional feature includes a heterogeneity of a stain sensitive to the target protein within the portion of the image.

31. The method of any of claims 20 to 30, wherein the regional feature includes at least one of a clustering pattern of cells positive for the target protein and / or cells negative for the target protein within the portion of the image.

32. The method of any of claims 20 to 31, wherein the regional feature includes a type of cell anchoring a neighborhood comprising the portion of the image.

33. The method of any of claims 20 to 32, wherein the regional feature includes a spatial heterogeneity of cells positive for the target protein and cells negative for the target protein within the portion of the image.

34. The method of any of claims 20 to 33, wherein the regional feature identifies a type of region comprising the portion of the image as one or more of total tissue, invasive region, or in situ region.

35. The method of any of claims 20 to 34, wherein the regional feature identifies a stain intensity quantification technique as one or more of optical density (OD) based, proportion based, or pixel based.

36. The method of any of claims 20 to 35, wherein the regional feature identifies a size of the portion of the image.

37. The method of any of claims 1 to 36, wherein the at least one higher order feature includes a summary statistic computed for at least a portion of the image.

38. The method of claim 37, wherein the summary statistic is computed for one or more quantiles.

39. The method of any of claims 37 to 38, wherein the summary statistic is computed for multiple quantiles in quantile increments.

40. The method of any of claims 1 to 39, wherein the feature extraction model determines, for inclusion in the feature set, one or more predetermined human interpretable features (HIFs) extracted from the image.107NAI-5005412604vl41. The method of any of claims 1 to 40, wherein the feature extraction model determines, for inclusion in the feature set, one or more de novo features extracted from the image.

42. The method of any of claims 1 to 41, wherein the feature extraction model determines the feature set by at least: extracting, from the image, one or more human interpretable features (HIFs); identifying, based at least on the one or more human interpretable features (HIFs), one or more cells positive for a target protein of the therapeutic; extracting, for inclusion in the feature set, one or more de novo features from a neighborhood anchored by each cell identified as positive for the target protein.

43. The method of any of claims 1 to 42, wherein the response prediction model comprises a univariate model trained to determine, based on each individual feature in the feature set, the response prediction for the patient.

44. The method of any of claims 1 to 43, wherein the response prediction model comprises a multivariate model trained to determine, based on multiple features from the feature set, the response prediction for the patient.

45. The method of any of claims 1 to 44, wherein the response prediction model comprises one or more of a random forest model, a decision tree, a gradient-boosted decision tree, or an elastic net.

46. The method of any of claims 1 to 45, further comprising: applying the one or more segmentation model to generate a segmentation map localizing one or more regions of interest (ROIs) present in the image; and applying the one or more segmentation models to generate, based at least on the segmentation map, the cell map such that the cell map localizes cells present in each region of interest (ROI).

47. The method of claim 46, wherein the one or more regions of interest (ROIs) include tumor regions and excludes non-tumor regions.

48. The method of any of claims 46 to 47, wherein the one or more regions of interest (ROIs) include tumor regions and non-tumor regions.

49. The method of any of claims 46 to 48, wherein each pixel in the segmentation map is assigned a value indicative of whether a corresponding pixel in the image depicts a region of interest (ROI).

50. The method of any of claims 46 to 49, wherein the segmentation map includes one or more color channels.108NAI-5005412604vl51. The method of any of claims 46 to 50, wherein the one or more regions of interest (ROIs) include one or more of a total tissue, an invasive region, or an in situ region.

52. The method of any of claims 1 to 51, wherein the cell map includes one or more color channels.

53. The method of any of claims 1 to 52, wherein each pixel in the cell map is further assigned a value indicative of a type of cell depicted by the corresponding pixel in the image.

54. The method of any of claims 1 to 53, wherein the at least one higher order feature comprises a global feature.

55. The method of claim 54, wherein the global feature includes an overall spatial heterogeneity of cells positive for the target protein and cells negative target protein across the image.

56. The method of claim 55, wherein the overall spatial heterogeneity is quantified based on one or more of a quantity, uniformity, and degree of clustering of the cells positive for the target protein and the cells negative target protein.

57. The method of any of claims 55 to 56, wherein the overall spatial heterogeneity is quantified based on one or more of a colocalization hotspot score, tumor cell hotspot score, and Morisita-Horn index bivariate correlation.

58. The method of any of claims 55 to 57, wherein the global feature includes an overall proportion of cells exhibiting each of a plurality of different intensity levels of a stain sensitive to the target protein of the therapeutic.

59. The method of claim 58, wherein the plurality of different intensity levels include a weak level of stain intensity, a moderate level of stain intensity, and a strong level of stain intensity.

60. The method of any of claims 55 to 59, wherein the global feature includes a metric quantifying an intensity of a stain sensitive to the target protein and proportion of cells positive for the target protein.

61. The method of claim 60, wherein the metric comprises a digital H-score.

62. The method of any of claims 1 to 61, wherein the feature set is further determined by at least: performing dimensionality reduction to reduce a quantity of features extracted from the image by the feature extraction model, where the dimensionality reduction identifies groups of features exhibiting sufficient covariance and correlation to be added to the feature set collectively as a single feature instead of as multiple individual features.109NAI-5005412604vl63. The method of any of claims 1 to 62, wherein the one or more segmentation models include a region of interest (ROI) segmentation model trained to localize, within the image, one or more regions of interest (ROIs).

64. The method of claim 63, wherein the one or more regions of interest (ROI) include a total tissue region, an invasive tumor region, or an in situ tumor region.

65. The method of any of claims 63 to 64, wherein the one or more segmentation models further include a cell segmentation model trained to localize, within each region of interest (ROI), the one or more cells.

66. The method of any of claims 63 to 65, wherein the cell segmentation model includes a foundation model that has been pretrained on a corpus of unlabeled or weakly labeled images of biological samples.

67. The method of claim 66, further comprising: incrementally finetuning, over multiple successive timesteps, the pretrained foundation model using annotated samples; and applying the finetuned foundation model to localize, within each region of interest (ROI), the one or more cells.

68. The method of any of claims 1 to 67, wherein the feature extraction model determines the feature set by at least: partitioning, into a plurality of tiles, the image; extracting, from each tile of the plurality of tiles, a feature vector including one or more features present in the tile.

69. The method of claim 68, further comprising: determining, based at least on the feature vector of each tile of the plurality of tiles, a tile-level representation of the tile; and applying the response prediction model to determine, based at least on the tile-level representation of each tile of the plurality of tiles, the response prediction.

70. The method of any of claims 68 to 69, further comprising: determining, based at least on the feature vector extracted from each tile of the plurality of tiles, an image-level representation of the image; and applying the response prediction model to determine, based at least on the image-level representation of the image, the response prediction.

71. The method of claim 70, wherein the image-level representation of the image is generated by applying, to the feature vector of each tile of the plurality of tiles, a geospatially adjusted attention weight computed based on the tile and one or more neighboring tiles.110NAI-5005412604vl72. The method of any of claims 1 to 71, wherein the feature extraction model determines the presence or absence of the target protein by at least classifying the cell as positive for the target protein or negative for the target protein.

73. The method of claim 72, wherein the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by least identifying, within the image of the biological sample, an expansion region including a set of pixels depicting a cytoplasm and a membrane of the cell, identifying, within the expansion region, one or more pixels positive for a target protein sensitive stain used to treat the image, dividing the expansion region into a plurality of partitions, determining a stain-positive proportion of the plurality of partitions, wherein a partition is stain-positive if a threshold quantity of pixels within the partition are positive for the stain, and classifying, based at least on the stain-positive proportion of the plurality of partitions, the cell as positive or negative of the target protein.

74. The method of claim 73, wherein the feature extraction model identifies the expansion region by at least performing a radial expansion from a nucleus of the cell.

75. The method of claim 74, wherein the radial expansion identifies the expansion region to be proportional in size to a size of the nucleus76. The method of any of claims 74 to 75, wherein the radial expansion identifies the expansion region to be proportional to a relative crowdedness of a neighborhood of the cell, and wherein the relative crowdedness of the neighborhood of the cell corresponds to a ratio between a geometric area of the neighborhood and a quantity of cells occupying the neighborhood.

77. The method of any of claims 73 to 76, wherein the cell is an immune cell that is classified as positive for the target protein if the cell exhibits either cytoplasmic staining or membranous staining.

78. The method of any of claims 72 to 77, wherein the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least identifying a plurality of pixels associated with the cell, and determining whether a threshold quantity of the plurality of pixels associated with the cell is positive for a target protein sensitive stain used to treat the image, and111NAI-5005412604vlclassifying the cell as negative for the target protein without the threshold quantity of the plurality of pixels associated with the cell being positive for a target protein sensitive stain used to treat the image79. The method of claim 78, wherein the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the threshold quantity of the plurality of pixels associated with the cell is determined to be positive for the target protein sensitive stain, determining whether a membrane of the cell is differentiable, and where the membrane of the cell is determined to be indifferentiable, classifying the cell as exhibiting only cytoplasmic staining.

80. The method of claim 79, wherein a differentiability of the membrane of the cell is determined by at least determining a quantity of cytoplasmic pixels between an outer perimeter of a nucleus of the cell and an outer perimeter of the cell, and determining the cell as exhibiting an indifferentiable membrane if the quantity of cytoplasmic pixels fails to satisfy one or more thresholds.

81. The method of claim 80, wherein the differentiability of the membrane of the cell is further determined by least in response to the quantity of cytoplasmic pixels satisfying the one or more thresholds, determining a gradient shift corresponding to a ratio between high intensity pixels and low intensity pixels between an outer perimeter of a nucleus of the cell and an outer perimeter of the cell, and determining the cell as exhibiting a differentiable membrane if the gradient shift satisfies one or more thresholds.

82. The method of any of claims 78 to 81, wherein the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the membrane of the cell is determined to be differentiable, determining whether the membrane of the cell is differentiable from a nuclear boundary of the cell, and where the membrane and the nuclear boundary of the cell are determined to be indifferentiable, classifying the cell as exhibiting indifferentiable membranous staining and cytoplasmic staining.

83. The method of claim 82, wherein the cell exhibiting the indifferentiable membranous staining and cytoplasmic staining is further classified as being in a gray zone of cells positive for the target protein112NAI-5005412604vl84. The method of any of claims 82 to 83, wherein the feature extraction model classifies the cell as positive for the target protein or negative for the target protein by at least where the membrane and the nuclear boundary of the cell are determined to be differentiable, classifying the cell as exhibiting membranous staining, in response to the cell being determined to exhibit membranous staining, determining an intensity and a completeness of the membranous staining, and classifying the cell as being positive for the target protein if the intensity and the completeness of the membranous staining satisfies one or more thresholds.

85. The method of claim 84, wherein the completeness of the membranous staining is determined by at least dividing the cell into a plurality of slices, and determining the completeness of the membranous staining to correspond to a ratio of a quantity of stain positive slices relative to a total quantity of the plurality of slices.

86. The method of any of claims 84 to 85, wherein the membranous staining is attributed to the cell instead of one or more adjacent cells when the intensity and completeness of the membranous staining satisfies the one or more thresholds.

87. The method of any of claims 78 to 86, wherein the plurality of pixels associated with the cell are identified by at least deconvolving the image of the biological sample to separate a nuclear channel corresponding to a nuclei sensitive stain and a membrane channel corresponding to a target protein sensitive stain, applying the one or more segmentation models to identify, within the nuclear channel of the image, one or more pixels depicting a nucleus of the cell , and applying the one or more segmentation models to identify, within the membrane channel of the image, one or more pixels corresponding to a boundary of the cell.

88. The method of claim 87, further comprising: identifying, within the image of the biological sample, one or more pixels depicting a cytoplasm of the cell, wherein the one or more pixels depicting the cytoplasm of the cell comprise one or more pixels between the nucleus and the boundary of the cell.

89. The method of any of claims 87 to 88, further comprising: training, based at least on a training dataset of images of a stain sensitive to a different target protein, the one or more segmentation models.113NAI-5005412604vl90. The method of claim 89, wherein the one or more segmentation models include a membrane segmentation model trained to infer a boundary of cells negative for the target protein.

91. The method of any of claims 78 to 90, wherein the cell is a tumor cell that is classified as positive for the target protein if the cell exhibits membranous staining.

92. The method of any of claims 1 to 91, further comprising preprocessing the image to detect one or more pixels depicting non-specific staining; and excluding the one or more pixels depicting non-specific staining from being used by the one or more segmentation models to generate the cell map.

93. The method of claim 92, wherein the preprocessing of the image further includes deconvolving the image to separate a first channel corresponding to a stain sensitive to the target protein from a second channel corresponding to a stain sensitive to one or more subcellular compartments, applying, to the first channel, one or more color filters comprising one or more thresholds on at least one of hue, saturation, and value of a plurality of pixels comprising the image of the biological sample, and excluding, from being used by the one or more segmentation models to generate the cell map, one or more pixels whose hue, saturation, and / or value fails to satisfy the one or more thresholds.

94. The method of claim 93, wherein the preprocessing of the image further includes excluding, from being used by the one or more segmentation models to generate the cell map, one or more clusters of adjacent pixels whose size fails to satisfy one or more thresholds.

95. The method of any of claims 93 to 94, wherein the first channel corresponds to a hematoxylin dye and the second channel corresponds to an immunohistochemistry (IHC) stain.

96. The method of any of claims 1 to 95, wherein the feature set includes(i) 90thquantile of target protein positive neighborhood pixel based proportion,(ii) 50thquantile target protein positive neighborhood cell based membrane stain intensity,(iii) 50thquantile target protein positive neighborhood cell based cytoplasm stain intensity,114NAI-5005412604vl(iv) 90thquantile target protein positive neighborhood cell based cytoplasm stain intensity,(v) 50thquantile 150-micron regional target protein cytoplasm stain intensity,(vi) 90thquantile 150-micron regional target protein cytoplasm stain intensity, and(vii) multivariate model composite score of (i)-(vi).

97. The method of any of claims 1 to 96, wherein the feature set includes a subset of the following plurality of features: percentage of optical density (OD) 1+ tumor cells based on pathologist matched cut points, percentage of OD 2+ tumor cells based on pathologist matched cut points, percentage of OD 3+ tumor cells based on pathologist matched cut points, H-score based on pathologist matched cut points, density of OD 1+ tumor cells based on pathologist matched cut points, density of OD 2+ tumor cells based on pathologist matched cut points, density of OD 3+ tumor cells based on pathologist matched cut points, percentage of OD 1+ tumor cells based on data driven cut points, percentage of OD 2+ tumor cells based on data driven cut points, percentage of OD 3+ tumor cells based on data driven cut points, H-score based on data driven cut points, density of OD 1+ tumor cells based on data driven cut points, density of OD 2+ tumor cells based on data driven cut points, and density of OD 3+ tumor cells based on data driven cut points.10th percent quantile of the neighborhood-wise average target protein sensitive stain optical density (OD) values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest,115NAI-5005412604vl90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (OD based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest,10th percent quantile of average target protein sensitive stain OD values (OD based calculation) among the tessellated regions defined by 150um * 150 um,50th percent quantile of average target protein sensitive stain OD values (OD based calculation) among the tessellated regions defined by 150um * 150 um,90th percent quantile of average target protein sensitive stain OD values (OD based calculation) among the tessellated regions defined by 150um * 150 um,10th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest,90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein positive tumor cell, in the invasive tumor region of interest,90th percent quantile of the neighborhood-wise average target protein sensitive stain OD values (proportion based calculation) of tumor cells inside the neighborhood defined by the radial expansion by 25um radius to the anchor target protein negative tumor cell, in the invasive tumor region of interest,10th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um,50th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um,116NAI-5005412604vl90th percent quantile of average target protein sensitive stain OD values (proportion based calculation) among the tessellated regions defined by 150um * 150 um,10th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest,90th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein positive tumor cell, in the invasive tumor region of interest,10th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest,50th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest, and90th percent quantile of the neighborhood-wise average target protein sensitive stain intensity at the pixel level (including stain positive pixels both within the cell and in the 2 microns expanded region from the outer boundary of the cell) inside the neighborhood defined by the radial expansion of 25um to the anchor target protein negative tumor cell, in the invasive tumor region of interest.

98. The method of any of claims 1 to 97, further comprising:117NAI-5005412604vladministering a therapeutically effective quantity of the therapeutic once a week, once every 2 weeks, once every 3 weeks, once every 4 weeks, or once every 5 weeks.

99. The method of claim 98, wherein the therapeutically effective quantity of the therapeutic comprises a dose of 1 to 30 mg / kg, 1 to 20 mg / kg, 2-12 mg / kg, 2-5 mg / kg, or 4-7 mg / kg.

100. The method of any of claims 98 to 99, wherein the therapeutically effective quantity of the therapeutic comprises a dose of 2 mg / kg, 2.5 mg / kg, 3 mg / kg, 3.5 mg / kg, 4 mg / kg, 4.5 mg / kg, 5 mg / kg, 5.5 mg / kg, 6 mg / kg, 6.5 mg / kg, 7 mg / kg, 7.5 mg / kg, 8 mg / kg, 8.5 mg / kg, 9 mg / kg, 9.5 mg / kg, 10 mg / kg, 11 mg / kg, or 12 mg / kg.

101. The method of any of claims 1 to 100, wherein the therapeutic targets the target protein or a different protein.

102. The method of any of claims 1 to 101, wherein the therapeutic targets a protein in cancer.

103. The method of any of claims 1 to 102, wherein the therapeutic targets a protein on a cell surface or a cell interior.

104. The method of any of claims 1 to 103, wherein the target protein is TROP2.

105. The method of any of claims 1 to 104, wherein the therapeutic is sacituzumab tirumotecan.

106. The method of claim 105, wherein the sacituzumab tirumotecan is used to treat the patient for breast cancer, non-small cell lung cancer, ovarian cancer, cervical cancer, endometrial cancer, gastric cancer, prostate cancer, esophageal cancer, or urothelial cancer.

107. The method of any of claims 1 to 106, wherein the tissue sample is a tumor tissue sample.

108. A method for treating cancer involving administering a cancer patient a therapeutic that is useful for treating cancer, the method comprising: selecting a patient having cancer, identified as eligible for receiving the therapeutic using the method of any of claims 1 and 4 to 107; and treating the patient by administering a therapeutically effective amount of the therapeutic to the patient.

109. A system, comprising: at least one data processor; and at least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1, 3 to 98, and 101 to 107.118NAI-5005412604vl110. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1, 3 to 98, and 101 to 107.119NAI-5005412604vl

Citation Information

Patent Citations

  • Systems and methods for automatically determining 3-dimensional object information and for controlling a process based on automatically-determined 3-dimensional object information

    US20080068379A1

  • Systems and methods to process electronic images to provide localized semantic analysis of whole slide images

    US20220108446A1

  • Machine learning prediction of therapy response

    US20230049979A1