Spatial feature analysis of digital pathology images
By analyzing the spatial distribution of biological objects using a digital pathological imaging system, indicators are generated to predict biological status and treatment effects, solving the problem of missing spatial features of biological objects in existing technologies and improving the accuracy of diagnosis and treatment selection.
Patent Information
- Application Number
- JP2025170142
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-09-11
- Filing Date
- 2025-10-08
- Publication Date
- 2026-02-25
AI Technical Summary
Existing technologies cannot effectively capture the spatial features of biological objects in digital pathology image analysis, resulting in insufficient detection of microenvironment-dependent cell type activities, which affects the accuracy of diagnosis, prognosis, and treatment selection.
By using a digital pathological imaging system to detect a representative set of biological objects and generate spatial distribution indicators, the relative positions of different biological objects are analyzed, and machine learning models are used to predict the biological state or treatment effect.
It improves the predictive ability of cell type activity, enhances the accuracy of diagnosis, prognosis and treatment selection, and in particular, supports the screening of clinical trials by analyzing the spatial distribution of lymphocytes and tumor cells.
Smart Images

Figure 2026031924000001_ABST
Abstract
Description
[Technical Field]
[0001] Priority This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 63 / 077,232, filed September 11, 2020, and U.S. Provisional Patent Application No. 63 / 026,545, filed May 18, 2020.
[0002] This application relates generally to image processing of digital pathology images to generate outputs that characterize the spatial information of particular types of objects within the image. More specifically, digital pathology images may be processed to generate metrics that characterize the spatial distribution and interrelationships of representations of one or more types of biological objects across all or a portion of the image. [Background technology]
[0003] Image analysis involves processing individual images to generate image-level results. For example, the results may be binary results corresponding to an assessment of whether the image contains a particular type of object. As another example, the results may include an image-level count of the number of a particular type of object detected in the image. In the context of digital pathology, the results may include the number of a particular type of cell detected in the image of the sample, the ratio of the number of one type of cell to the number of another type of cell across the entire image, and / or the density of a particular type of cell.
[0004] This image-level approach can be advantageous because it can facilitate simple metadata storage and provide a ready understanding of how results were generated. However, this image-level approach can remove detail from the image, which can hinder detection of details of the depicted situation and / or environment. This simplification can be particularly impactful in the context of digital pathology, as the current or potential future activity of certain cell types can be highly dependent on the microenvironment.
[0005] It would therefore be beneficial to develop techniques for processing digital pathology images to produce outputs that reflect the spatial characteristics of the depicted biological objects. Summary of the Invention
[0006] In some embodiments, a computer-implemented method is provided that includes a digital pathology imaging system accessing a digital pathology image depicting a cross-section of a biological sample from a subject. The digital pathology imaging system detects a first set of biological object representations and a second set of biological object representations within the digital pathology image. Each of the first set of biological object representations depicts a first biological object of a first type of biological object. Each of the second set of biological object representations depicts a second biological object of a second type of biological object. The digital pathology imaging system uses the first set of biological object representations and the second set of biological object representations to generate a spatial distribution metric that characterizes the positions of the first set of biological object representations relative to the second set of biological object representations. The digital pathology imaging system uses the spatial distribution metric to generate subject-level results for a predicted biological state of the subject or a potential treatment for the subject. The digital pathology imaging system generates a display screen including the subject-level results. In certain embodiments, the first type of biological object includes a first type of cell, and the second type of biological object includes a second type of cell. In certain embodiments, the first type of biological object comprises lymphocytes, and the second type of biological object comprises tumor cells. In certain embodiments, the digital pathology image depicts a biological sample from the subject after treatment with one or more stains, each of which enhances the appearance of one or more of the first type of biological object or the second type of biological object. In certain embodiments, the digital pathology imaging processing system generates the spatial distribution metric by, for each first biological object representation of the one or more first biological object representations, identifying a first point location in the digital pathology image corresponding to the first biological object representation; for each second biological object representation of the one or more second biological object representations, identifying a second point location in the digital pathology image corresponding to the second biological object representation; and determining the spatial distribution metric based on the first point locations and the second point locations.In certain embodiments, the first point location within the digital pathology image indicates the location of the first biological object representation. In certain embodiments, the first point location within the digital pathology image is selected by calculating a mean point location, a centroid point location, a median point location, or a weighted point location for the first biological object representation. In certain embodiments, the digital pathology image processing system generates a spatial distribution metric by calculating, for each of at least some first biological object representations of the one or more first biological object representations and for each of at least some second biological object representations of the one or more second biological object representations, a distance between the first point location corresponding to the first biological object representation and the second point location corresponding to the second biological object representation. In certain embodiments, the digital pathology image processing system generates a spatial distribution metric by identifying, for each of at least some first biological object representations of the one or more first biological object representations, one or more of the second biological object representations associated with a distance between the first biological object representation and the second biological object representation. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by defining a spatial grid configured to divide the region of the digital pathology image into a set of image regions, assigning each first biological object representation of the one or more first biological object representations to an image region of the set of image regions, assigning each second biological object representation of the one or more second biological object representations to an image region of the set of image regions, and generating the spatial distribution metric based on the image region assignments. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by determining a first set of one or more image regions of the set of image regions that have a higher probability of containing the first biological object representation than adjacent image regions, determining a second set of one or more image regions of the set of image regions that have a higher probability of containing the second biological object representation than adjacent image regions, and further determining the spatial distribution metric based on the first set of image regions and the second set of image regions.In certain embodiments, the digital pathology image processing system generates the spatial distribution metric by determining a third set of one or more image regions of the set of image regions that have a higher probability of containing both the first biological object representation and the set of biological object representations than adjacent image regions, and further determining the spatial distribution metric based on the third set of image regions. In certain embodiments, the digital pathology image processing system generates a subject-level result for a predicted biological state of a subject or a potential treatment for a subject by comparing the spatial distribution metric generated for the digital pathology image with a previous spatial distribution metric generated for a previous digital pathology image, and outputting a subject-level result generated for the previous digital pathology image based on the comparison. In certain embodiments, the digital pathology image processing system generates a subject-level result by using a trained machine learning model to determine a diagnosis, prognosis, treatment recommendation, or treatment eligibility assessment for the subject based on processing the spatial distribution metric and the first set of biological object representations and the second set of biological object representations. In certain embodiments, the spatial distribution metric includes a metric defined based on a K-nearest neighbor analysis, a metric defined based on Ripley's K-function, a Morisita-Horn index, a Moran index, a metric defined based on a correlation function, a metric defined based on hot spot / cold spot analysis, or a metric defined based on kringing-based analysis. In certain embodiments, the spatial distribution metric is a first type of metric. The digital pathology imaging system uses the first set of biological object description and the second set of biological object description to generate a second spatial distribution metric that characterizes the positions of the first set of biological object description relative to the second set of biological object description. The second spatial distribution metric is a second type of metric different from the first type of metric. The subject-level result is further generated using the second spatial distribution metric.In certain embodiments, the digital pathology imaging system receives user input data from a user device including an identifier for the subject or the digital pathology image. The digital pathology image is accessed based on the received user input data. The digital pathology imaging system provides subject-level results for display by providing the subject-level results to the user device. In certain embodiments, the digital pathology imaging system outputs a clinical assessment to the user device for the subject. The clinical assessment may include a diagnosis, prognosis, treatment recommendation, or treatment eligibility assessment for the subject.
[0007] In some embodiments, a method is provided that includes accessing, by a digital pathology imaging system, a digital pathology image representing a cross-section of a biological sample taken from a subject with a given medical condition. The digital pathology imaging system detects a set of biological object representations within the digital pathology image. The set of biological object representations includes a first set of biological object representations of a first class of biological objects and a second set of biological object representations of a second class of biological objects. The digital pathology imaging system generates relative position representations of the one or more biological object representations. Each of the one or more relative position representations indicates a position of the first biological object representation relative to the second biological object representation. The digital pathology imaging system uses the one or more relative position representations to determine a spatial distribution metric that characterizes the degree to which at least some of the biological object representations of the first set are depicted as interspersed with at least some of the biological object representations of the second set. Based on the spatial distribution metric, the digital pathology imaging system generates a result corresponding to a prediction of the degree to which a given treatment that modulates the immune response will effectively treat the given medical condition in the subject. Based on the result, the digital pathology imaging system determines that the subject is eligible for the clinical trial. The digital pathology imaging system generates a display screen including an indication that the subject is eligible for the clinical trial. In certain embodiments, the spatial distribution metric includes a metric defined based on a K-nearest neighbor analysis, a metric defined based on Ripley's K-function, a Morisita-Horn index, a Moran index, a metric defined based on a correlation function, a metric defined based on hot spot / cold spot analysis, or a metric defined based on kringing-based analysis. In certain embodiments, the spatial distribution metric is a first type of metric, and the digital pathology imaging system uses the one or more relative position representations to determine a second spatial distribution metric that characterizes the degree to which at least some of the biological object representations in the first set are depicted as interspersed with at least some of the biological object representations in the second set.The second spatial distribution metric is a second type of metric different from the first type of metric. The result is further generated based on the second spatial distribution metric. In certain embodiments, generating the result includes the digital pathology image processing system processing the first spatial distribution metric and the cross-sectional spatial distribution metric using a trained machine learning model. The trained machine learning model is trained using a set of training elements. Each set of training elements corresponds to a different subject who has received a specific treatment associated with the clinical trial. Each set of training elements includes a different set of spatial distribution metrics and a responsiveness value indicating the degree to which the given treatment activated an immune response in the other subject. In certain embodiments, generating the result includes comparing the value of the spatial distribution metric to a threshold. In certain embodiments, the given medical condition is a type of cancer, and the given treatment is an immune checkpoint blockade treatment. In certain embodiments, the one or more associated location representations include, for a set of biological object representations, a set of coordinates identifying the location of the biological object representations within the digital pathology image. In certain embodiments, generating one or more related position representations of the biological object representations includes: for each biological object representation of the first set of biological object representations, identifying a first point location in the digital pathology image corresponding to the biological object representation; for each biological object representation of the second set of biological object representations, identifying a second point location in the digital pathology image corresponding to the biological object representation; and comparing the first point location and the second point location. In certain embodiments, the first point location in the digital pathology image is selected by calculating a mean point location, a centroid point location, a median point location, or a weighted point location for a biological object representation of one of the first set of biological object representations.In certain embodiments, the digital pathology imaging system determines the spatial distribution metric for each of at least some of the first set of biological object representations and for each of at least some of the second set of biological object representations by calculating a distance between a first point location corresponding to the biological object representation in the first set and a second point location corresponding to the biological object representation in the second set. In certain embodiments, the digital pathology imaging system determines the spatial distribution metric for each of at least some of the first set of biological object representations by identifying one or more of the second set of biological object representations associated with a distance between a first point location corresponding to the biological object representation in the first set and a second point location corresponding to the biological object representation in the second set. In certain embodiments, the one or more associated position representations include, for each of a set of image regions in the digital pathology image, a representation of an absolute or relative amount of biological object representations of a first class of biological object identified to be located in the region and a representation of an absolute or relative amount of biological object representations of a second class of biological object identified to be located in the region. In certain embodiments, the one or more associated location representations include a distance-based probability of a biological object representation of a first set of biological object representations being depicted as being located within a given distance from a biological object representation of a second set of biological object representations. In certain embodiments, the digital pathology imaging system accesses genetic sequencing or radiology image data of the subject, and the result is generated further based on characteristics of the genetic sequencing or radiology image data. In certain embodiments, the first class of biological objects are tumor cells and the second class of biological objects are immune cells. In certain embodiments, the digital pathology imaging system receives user input data from a user device including an identifier for the subject and accesses the digital pathology image in response to receiving the identifier. The digital pathology imaging system provides an indication to the user device that the subject is eligible for the clinical trial, thereby generating a display screen including an indication that the subject is eligible for the clinical trial.In certain embodiments, the digital pathology imaging system receives an indication that a subject is enrolled in a clinical trial. In certain embodiments, the digital pathology imaging system generates a display screen including an indication that the subject is eligible for the clinical trial by informing the subject of the trial eligibility determination.
[0008] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0009] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0010] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0011] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]
[0012] The present disclosure is described in conjunction with the accompanying drawings, in which:
[0013] [Figure 1] FIG. 1 illustrates an interactive system for generating and processing digital pathology images to characterize relative spatial information of biological objects, according to some embodiments.
[0014] [Figure 2] FIG. 2 illustrates an exemplary system for processing object-description data to generate a spatial distribution metric, according to some embodiments.
[0015] [Figure 3A-3B] 3A and 3B illustrate a process for providing a health-related assessment based on spatially specific image processing of digital pathology images according to some embodiments.
[0016] [Figure 4] FIG. 4 illustrates a process for processing an image using a landscape-based spatial point process analysis framework, according to some embodiments.
[0017] [Figures 5A-5C]5A-5C show example processed images using a discriminant-based spatial point process analysis framework according to some embodiments.
[0018] [Figures 6A-6D] 6A-6D illustrate exemplary distance- and intensity-based metrics for characterizing the spatial location of object depictions in exemplary images, according to some embodiments.
[0019] [Figure 7] FIG. 7 illustrates a process for processing an image using a lattice-based spatial-domain analysis framework, according to some embodiments.
[0020] [Figure 8] FIG. 8 illustrates a process for processing an image using Moran's Index according to some embodiments.
[0021] [Figure 9] FIG. 9 illustrates a process for processing an image using a hotspot-based spatial area analysis framework, according to some embodiments.
[0022] [Figure 10] FIG. 10 illustrates a process for processing images using a geostatistical analysis framework, according to some embodiments.
[0023] [Figure 11] FIG. 11 shows receiver operating curves characterizing the performance of a trained logistic regression model for predicting the occurrence of microsatellite instability based on processing of digital pathology images, according to some embodiments.
[0024] [Figure 12] FIG. 12 illustrates the process of assigning a predicted outcome label to each subject in the study cohort using a nested Monte Carlo cross-validation modeling strategy.
[0025] [Figure 13] FIG. 13 shows the Kaplan-Meir plot for subjects in the analysis of the two subject cohorts.
[0026] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. When only a first reference label is used in this specification, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION
[0027] Digital images are increasingly being used in medical settings to facilitate clinical evaluations such as diagnosis, prognosis, treatment selection, and treatment evaluation, among various other uses. In the field of digital pathology, digital pathology images can be processed to estimate whether a given image contains representations of a particular type or class of biological object. For example, a tissue sample section can be stained so that representations of a particular type of biological object (e.g., a particular type of cell, a particular type of organelle, or a blood vessel) preferentially absorb the stain and are therefore represented with a higher intensity of a particular color. The tissue sample can be imaged according to the techniques disclosed herein. The digital pathology image can then be processed to detect representations of biological objects. Detection of biological object representations can be based on biological objects meeting certain criteria, such as a specified range of size, a specified type of shape, or at least a specified amount of contiguous high-intensity pixels in an analysis corresponding to a staining profile. In certain embodiments, clinical evaluations or recommendations can be made based on whether representations of a particular type or class of object are observed and / or the amount of representations of one or more particular types or classes of objects.
[0028] With advances in image processing technology, digital imaging of tumor tissue slides is becoming a routine clinical procedure for managing many types of conditions. Digital pathology images can capture multiple objects of a given type or class at high resolution. It may be advantageous to characterize the degree of spatial heterogeneity of biological objects captured in digital pathology images, as well as the degree to which objects of a given type are spatially aggregated and / or dispersed relative to each other and / or to objects of different types. The current or potential activity or function of biological objects can change dramatically depending on the biological object's microenvironment. Objectively characterizing the location of depictions of a particular type of biological object can substantially affect the quality of current diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility assessment. Similarly, objectively characterizing the relationships of multiple types of biological objects within a digital pathology image or a region of a digital pathology image can substantially affect analysis results. The locations and relationships of depictions of biological objects in a digital pathology image can be correlated with the locations and relationships of corresponding biological objects in a target tissue sample. As disclosed herein, such objective spatial characterization can be performed by detecting a set of biological object depictions from a digital pathology image. The objects may be represented according to one or more spatial analysis frameworks, including, but not limited to, a spatial point process analysis framework, a spatial surface analysis framework, a geostatistical analysis framework, a graph-based framework, etc. In some embodiments, each detected biological object representation is associated with a specific point location within the image and may be further associated with an identifier for a particular type of object. In some embodiments, each of a set of regions within the image and one or more particular types of objects may be associated with an identifier for a particular type of object. For each, metadata can be stored that indicates the amount or density of representation of each particular type of biological object predicted or determined to be located within the region.
[0029] Spatial aggregation may involve measuring how objects within a digital pathology image are spatially aggregated or dispersed across the entire digital pathology image or across regions of the digital pathology image. For example, it may be advantageous to determine the degree to which one type or class of biological object (e.g., lymphocytes) spatially intermingles with another type or class of biological object (e.g., tumor cells). To illustrate, intratumoral tumor-infiltrating lymphocytes (TILs) are located within the tumor and interact directly with tumor cells, whereas interstitial TILs reside in the tumor stroma and do not interact directly with tumor cells. Not only do intratumoral TILs have different activity patterns than interstitial TILs, but each cell type may be associated with a different type of microenvironment, further influencing the behavioral differences between TIL types. When lymphocytes are detected in a particular location (e.g., within the tumor), the fact that they have infiltrated the tumor may convey information about the activity of the lymphocytes and / or tumor cells. Furthermore, the microenvironment may influence the current and future activity of lymphocytes. Identifying the relative locations of particular types of biological objects can be particularly useful for predictive applications such as identifying prognosis and treatment options, assessing patient eligibility for clinical trials, and typing the immunological characteristics of subjects and their conditions.
[0030] As another form of objective characterization of the locations and relationships of detected biological object representations, the detected biological object representations can be used to generate one or more spatial distribution metrics, which can characterize the extent to which biological objects of a given type or class are predicted to be interspersed with biological objects of another type or class, clustered with other objects of the same type, and / or clustered with biological objects of another given type, at the region, image, and / or object level. For example, a digital pathology image processing system can detect a first set of biological object representations and a second set of biological object representations in a digital pathology image. The system can predict that each of the first set of biological object representations depicts a first type of biological object (e.g., lymphocytes) and that each of the second set of biological object representations depicts a second type of biological object (e.g., tumor cells). The digital pathology imaging system may perform a distance-based assessment to generate a spatial distribution metric indicative of the degree to which individual biological object representations within a first set of biological object representations are spatially integrated or separated from individual biological object representations within a second set of biological object representations, and / or the degree to which the first set of biological object representations (e.g., collectively) are spatially integrated or separated from the second set of biological object representations (e.g., collectively). As disclosed herein, various spatial distribution metrics have been developed and applied for this purpose.
[0031] Principles and quantitative methods from advanced analytics (e.g., spatial statistics) can be applied to generate novel solutions that meet these needs. The technology provided herein can be used to process digital pathology images to generate results characterizing the spatial distribution and / or spatial pattern of one or more specific types or classes of depicted objects (e.g., biological objects). The digital pathology images can include digital images of stained sections of a sample. Processing can include detecting depictions of biological objects of each of a plurality of specific types (e.g., corresponding to each of a plurality of types of biological cells). Biological object detection can include detecting one or more of a first set of biological object depictions corresponding to a first biological object type and a second set of biological object depictions corresponding to a second biological object type. Additionally or alternatively, object detection can include, for each region of the set of regions in the digital pathology image and each of a plurality of specific biological object types, identifying a higher-order metric defined to depend on and correlate with the amount of biological object or a lower-order metric (e.g., a number, density, or image intensity estimated to represent the amount of a specific type of biological object represented in the corresponding image region). Additionally, spatial distribution metrics may be used in combination with other metrics (e.g., RNA sequencing, radiological imaging (CT, MRI, etc.)) to improve predictive capabilities or discover novel biomarkers for unmet medical needs.
[0032] Image locations of one or more biological object representations may be determined. The image locations may be determined and represented according to one or more spatial analysis frameworks, such as a spatial point process analysis framework, a spatial surface analysis framework, a geostatistical analysis framework, or a graph-based analysis framework. For example, a biological object may be associated with a single point location within the digital pathology image. Even if the biological object representation spans multiple pixels or voxels, the single point location may indicate or be selected as representative of the location of the biological object representation within the digital pathology image. As another example, the biological object representation may be collectively represented with or indicated by one or more other biological object representations that contribute to the number of objects detected within a particular region of the image, the density of the biological objects detected within a particular region of the image, the pattern of the biological objects detected within a particular region of the image, etc.
[0033] Digital pathology imaging systems may use spatial distribution metrics to facilitate, for example, diagnosis, prognosis, treatment evaluation, treatment selection, and / or treatment eligibility (e.g., the eligibility of a subject to be admitted to or recommended for a clinical trial or a particular group of clinical trials). For example, a particular prognosis may be identified in response to detecting some degree of infiltration of a set of biological objects of a first type or class within a set of biological objects of a second type or class, and a more relevant and accurate prognosis may be identified in response to detecting higher lymphocytic infiltration within individual tumors and / or metastatic tumor nests. As another example, a diagnosis of tumor or cancer stage may be informed based on the degree to which immune cells are spatially integrated with cancer cells (e.g., higher integration generally corresponds to a lower stage). As yet another example, therapeutic efficacy may be determined to be higher if the spatial proximity of lymphocytes to tumor cells is less after treatment initiation compared to before treatment or compared to predicted proximity based on one or more prior assessments performed on a given subject.
[0034] Biological object detection can be used to generate results that may include or be based on spatial distribution metrics that can indicate the proximity between depictions of the same or different types of biological objects and / or the degree of colocalization of depictions of one or more types of biological objects. Colocalization of depictions of biological objects can represent similar locations of multiple cell types in each of one or more regions of a digital pathology image. The results can indicate and / or predict interactions between different biological objects and types of biological objects that may occur within the microenvironment of a subject's or patient's structure, as indicated by a sample taken from the subject or patient. Such interactions may support and / or be essential for biological processes, such as tissue formation, homeostasis, regenerative processes, or immune responses. Thus, the spatial information conveyed by the results can be informative regarding the function and activity of specific biological structures and, therefore, can be used, for example, as a quantitative basis for characterizing disease states and prognosis. Results indicating where specific biological objects are located within a biological microenvironment can be used to select a treatment predicted to be effective for a particular subject (e.g., compared to other treatment options) or to predict outcomes for other subjects.
[0035] In certain embodiments, multiple spatial distribution metrics may be generated. In particular, one or more metrics may be generated, each corresponding to one or more metric types. For example, one or more first metrics may be generated using a spatial point process analysis framework. The first metric may be based on the distance between representations of different types of biological objects. For example, the first metric may use the Euclidean distance between biological object representations corresponding to tumor cells and lymphocytes. Other distance metrics may also be used. One or more second metrics may be generated using a spatial domain analysis framework. The second metric may characterize the count or density of representations of a first type of biological object in various image regions relative to the number or density of other representations of a second type of biological object.
[0036] Machine learning models or rules may be used to generate results corresponding to, for example, diagnosis, prognosis, treatment evaluation, treatment selection, eligibility for treatment (e.g., eligibility to be accepted or recommended for a clinical trial or a particular group in a clinical trial), and / or prediction of genetic mutations, genetic alterations, biomarker expression levels (including, but not limited to, genes or proteins), etc. One or more metrics, each corresponding to one or more metric types, may be used to generate results. Machine learning models may include, for example, but are not limited to, classification, regression, decision tree, or neural network techniques trained to learn one or more weights to use when processing metrics to generate results.
[0037] The digital pathology image processing system may further learn to identify and recognize patterns of location and relationship of the detected biological object depictions based in part on one or more spatial distribution metrics. For example, the digital pathology image processing system may detect patterns of location and relationship of the detected biological object depictions in the digital pathology image of the first sample. The digital pathology image processing system may generate a mask or other pattern storage data structure from the recognized patterns.
[0038] The digital pathology imaging system may use the spatial distribution metrics described herein to predict a diagnosis, prognosis, treatment evaluation, treatment selection, and / or treatment eligibility. The digital pathology imaging system may store the predicted prognosis, etc. in association with the detected pattern and / or generated mask. The digital pathology imaging system may receive a subject's outcome to verify the predicted prognosis, etc. Then, when processing a second digital pathology image from a second sample, the digital pathology image processing system may detect a pattern of positions and relationships of the detected biological object depictions in the second digital pathology image. The digital pathology image processing system may recognize a similarity between the pattern of positions and relationships detected in the second digital pathology image and the mask or stored detected pattern from the first digital pathology image. The digital pathology image processing system may inform a predicted prognosis, treatment recommendation, or treatment eligibility determination based on the recognized similarity and / or the subject's outcome. As an example, the digital pathology image processing system may compare the stored mask with the pattern of positions and relationships of the detected biological object depictions in the second digital pathology image. The digital pathology image processing system may determine one or more spatial distribution metrics of the second digital pathology image and base a comparison of the recognized pattern from the second digital pathology image with the stored mask based on a comparison of the spatial distribution metrics of the detected biological object depictions in the first digital pathology image and the second digital pathology image.
[0039] The pattern detected from the first digital pathology image processing system may be associated in many ways with the locations and relationships of one or more first biological object representations of one or more types. For example, the pattern may be associated with the locations and relationships of a first biological object representation of a first type within the digital pathology image without the context of other biological object representations within the digital pathology image. The pattern may be associated with an abstract representation of the locations and / or relationships of the biological object representations within the boundaries of the digital pathology image (e.g., evaluating the coordinates of the detected biological object representations, potentially lacking their context as biological object representations). As another example, the pattern may be associated with the locations and relationships of the first type of biological object representation relative to all of the other biological object representations within the digital pathology image. As yet another example, the pattern may be associated with the locations and relationships of one or more biological object representations of a first type relative to the locations and relationships of one or more biological object representations of a second type.
[0040] Patterns detected from digital pathology images may be associated with context including, for example, the type of sample the digital pathology image depicts (e.g., biopsy method including, but not limited to, lung biopsy, liver tissue sample, blood sample, formalin-fixed paraffin-embedded specimen, frozen specimen, cell preparation obtained from surgical evacuation, core needle biopsy fine needle aspirate from various organs, tumors, and / or metastases, etc.), the sample preparation method (e.g., type of stain used, age of the sample, etc.), the number and specific types of biological objects depicted throughout the sample or incorporated into the pattern (e.g., sample cell type, structures—e.g., glands, tumor lobules, sheets of cells, blood vessels, etc.—individual cells—e.g., tumor cells, immune cells, mitotic cells, stromal cells, endothelial cells, etc.—and cellular components—e.g., nucleus, cytoplasm, membrane, cilia, mucus excretion, etc.), the number and type of spatial distribution metrics used to detect or prepare the pattern, the type of subject-level result associated with the pattern, representation within the type of subject-level result, the degree of validation of the subject-level result, and many other factors that go towards characterizing patterns detected from digital pathology images. This context can be used to improve pattern recognition and application to future digital pathology images.
[0041] In some embodiments, a pattern may only apply to the same type of sample, the same type of biological object description, the same type of spatial distribution metric, subject-level results for a sample type, etc., but the digital pathology image processing system may be trained to apply pattern recognition methodologies across types. For example, based on analysis of digital pathology images corresponding to different types of tissue samples, the digital pathology image processing system may be trained to recognize the broad applicability of a pattern regarding lymphocyte infiltration and placement in tissue sample cells and provide similar subject-level results. The ability to reference and apply a pattern may be based on the applicability of a spatial distribution metric associated with different types of detected biological object description and can be applied across digital pathology images of different tissue sample types. The spatial distribution metric provides an objective, quantifiable measure for multi-dimensional comparisons.
[0042] Additionally or alternatively, the digital pathology imaging system may further use spatial distribution metrics to facilitate treatment selection. For example, immunotherapy or immune checkpoint therapy may be selectively recommended upon detecting an output indicating that lymphocytes are spatially integrated with tumor cells. As another example, upon detecting an output indicating that lymphocytes are spatially integrated with tumor cells, atezolizumab + bevacizumab + carboplatin + paclitaxel (ABCP) or atezolizumab + carboplatin + paclitaxel (ACP) may be selectively recommended over another chemotherapy treatment. The other chemotherapy treatment may include or be bevacizumab + carboplatin + paclitaxel (BCP). Other approaches may use other biological objects, or cellular components or compartments, to predict diagnosis, biomarker expression, or treatment response (e.g., vascular distribution, distribution of specific nuclear features in lymphoma, etc.).
[0043] Facilitating the identification of a diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility may include automatically generating a potential diagnosis, prognosis, treatment assessment, and / or treatment selection. The automatic identification may be based on one or more learned and / or static rules. The rules may have an if-then format, which may include, for example, inequalities and / or one or more thresholds in the condition, which may indicate that a metric exceeding the threshold is associated with suitability for a particular treatment. Alternatively or additionally, the rules may include functions, such as functions relating a numerical metric to a disease severity score or a quantified score of eligibility for a treatment. The digital pathology imaging system may output the potential diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility determination as a recommendation and / or prediction. For example, the digital pathology imaging system may provide output to a locally coupled display, transmit output to a remote device or access terminal, store results in local or remote data storage, etc. In this manner, a human user (e.g., a physician and / or healthcare provider) may use the automatically generated output or form a separate assessment informed by the quantitative metrics described herein.
[0044] Facilitating the identification of a diagnosis, prognosis, treatment evaluation, treatment selection, and / or treatment eligibility determination may include outputting a spatial distribution metric consistent with the disclosed subject matter. For example, the output may include a subject identifier (e.g., the subject's name), stored clinical data associated with the subject (e.g., past diagnoses, potential diagnoses, current treatments, symptoms, test results, and / or vital signs), and the determined spatial distribution metric. The output may include the digital pathology image from which the spatial distribution metric was derived and / or a modified version thereof. For example, the modified version of the digital pathology image may include an overlay and / or markings identifying each biological object representation detected in the digital pathology image. The modified version of the digital pathology image may further provide information about the detected biological object representation. For example, for each biological object representation, an interactive overlay may provide a specific object category corresponding to the object. A human user (e.g., a physician and / or healthcare provider) may then use the output including the spatial distribution metric to identify a diagnosis, prognosis, treatment evaluation, treatment selection, or treatment eligibility determination.
[0045] In certain embodiments, multiple types of spatial distribution metrics are generated using biological object depictions detected from a single digital pathology image. Multiple types of spatial distribution metrics may be used in combination according to the subject matter disclosed herein. The multiple types of spatial distribution metrics may, for example, correspond to different or the same frameworks for how the location of each biological object depiction is characterized. The multiple types of spatial distribution metrics may include different variable types (e.g., calculated using different algorithms) and may be presented on different value scales. The multiple types of spatial distribution metrics may be processed together using rules or machine learning models to generate labels. The labels may correspond to predicted diagnoses, prognoses, treatment assessments, treatment selections, and / or treatment eligibility determinations.
[0046] In certain embodiments, a computer-implemented method is provided. A digital pathology imaging system may access one or more digital pathology images. Each of the one or more digital pathology images may depict a cross-section of a biological sample from a subject. The depicted cross-section may include one stained with one or more stains. The digital pathology imaging system detects a first set of biological object representations and a second set of biological object representations in each of the one or more digital pathology images. Each of the first set of biological object representations may depict a first type of biological object. Each of the second set of object representations may depict a second type of biological object. The digital pathology imaging system uses the first set of biological object representations and the second set of biological object representations to generate one or more spatial distribution metrics of a first type of spatial distribution metrics. Each of the one or more first spatial distribution metrics characterizes a position of the first set of biological object representations relative to the second set of biological object representations. The digital pathology imaging system uses the first set of biological object representations and the second set of biological object representations to generate one or more spatial distribution metrics of a second type. The second type of spatial distribution metric characterizes the positions of the first set of biological object representations relative to the second set of biological object representations. The digital pathology imaging system can use the one or more first spatial distribution metrics and the one or more second spatial distribution metrics to generate subject-level results corresponding to a predicted biological state of the subject or a potential treatment for the subject. The digital pathology imaging system provides the subject-level results for display. In addition to providing the subject-level results, the digital pathology imaging system can provide a clinical evaluation for the subject based on the subject-level results. The clinical evaluation can include a diagnosis, a prognosis, a treatment evaluation, a treatment selection, and / or treatment eligibility.
[0047] The spatial distribution metric characterizing the positions of the first set of biological object representations may be determined based on, for example and without limitation, a point process, a surface / grid process, a geostatistical process, etc. In certain embodiments, the first type of biological object may include a first type of cell, and the second type of biological object may include a second type of cell. As an example, the first type of biological object may include lymphocytes, and the second type of biological object may include tumor cells. As another example, the first type of biological object may include macrophages, and the second type of biological object may include fibroblasts. In certain embodiments, the first type of biological object may include, for example, a first class of biological object defined by a first type of characteristic characteristic (e.g., size, shape, color, expected behavior, texture of the biological object or a component or section of the biological object), and the second type of biological object may include a second class of biological object defined by, for example, a second type of characteristic characteristic or a characteristic characteristic of a first type of variation. It will be appreciated that the subject matter disclosed herein may be equally applicable to any biological object that may be represented as a point corresponding to a location in a digital pathology image.
[0048] In certain embodiments, generating one or more spatial distribution metrics of a first type may include identifying a first point location in one or more digital pathology images for each first biological object representation of the one or more first biological object representations. The first point location may correspond to the location of the depicted first biological object. Generating one or more spatial distribution metrics of a first type may further include identifying a second point location in one or more digital pathology images for each second biological object of the one or more second biological objects. The second point location may correspond to the location of the depicted second biological object. Generating one or more spatial distribution metrics of a first type may further include determining one or more spatial distribution metrics of the first type based on the first point locations and the second point locations. In certain embodiments, generating the one or more spatial distribution metrics may include performing a distance-based technique to evaluate, for each first biological object of at least some of the one or more first biological objects and each second biological object of at least some of the one or more second biological objects, a distance between a first point location corresponding to the first biological object and a second point location corresponding to the second biological object.
[0049] In certain embodiments, generating the one or more spatial distribution metrics of the second type may include defining a spatial grid configured to divide a region of the digital pathology image into a set of image regions, and generating the one or more spatial distribution metrics of the second type may include assigning each second biological object of the one or more second biological objects to an image region of the set of image regions.
[0050] Generating one or more spatial distribution metrics of the second type may include generating one or more spatial distribution metrics of the second type based on image region assignments of each second biological object of the one or more second biological objects. Generating the target-level results may include processing one or more spatial distribution metrics of a first type and one or more spatial distribution metrics of a second type using a trained machine learning model. The trained machine learning model may include, by way of example and not limitation, a regression model, a decision tree model, or a neural network model. The first type of metric may be one of a set of metric types. The second type of metric may be another of a set of metric types. The set of metric types may include metrics defined based on K-nearest neighbor analysis, metrics defined based on Ripley's K-function, Moriscia-Horn index, Moran's index, Geary's C-index, G-function, metrics defined based on correlation functions, metrics defined based on hot spot analysis or cold spot analysis, or metrics defined based on kriging-based analysis.
[0051] In certain embodiments, a method is provided that includes sending a request communication from a client computing system to a remote computing system to process one or more digital pathology images depicting a particular portion of a biological sample from a patient, wherein in response to receiving the request communication from the client computing system, the remote computing system accesses the one or more digital pathology images and performs an analysis according to the subject matter disclosed herein.
[0052] According to the subject matter disclosed herein, in certain embodiments, there is provided the use of subject-level results in treating a subject. Subject-level results can be provided according to the subject matter disclosed herein.
[0053] In certain embodiments, a method is provided. A digital pathology image is accessed in a digital pathology imaging system. The digital pathology image depicts a tissue slide stained with one or more stains, the tissue of the tissue slide having been collected from a subject with a particular medical condition. The digital pathology image includes a representation of one or more biological objects. The one or more biological objects may include a set of cells. The set of cells may include a set of tumor cells and a set of other cells. The set of other cells may be a set of immune cells or a set of stromal cells. The digital pathology imaging system may identify a set of locations within the digital pathology image corresponding to one or more biological objects, such as tumor cell locations. Each tumor cell location within the set of tumor cell locations may correspond to a tumor cell within the set of tumor cells. The digital pathology imaging system may identify a set of other locations within the digital pathology image corresponding to one or more other biological objects, such as other cell locations. Each other cell location within the set of other cell locations may correspond to a cell within the other set of cells. The digital pathology imaging system may generate one or more relational location representations. Each of the one or more relational location representations may indicate the positions of at least some of a first set of cells relative to the positions of at least some of a second set of cells. The digital pathology imaging system may use the one or more relational location representations to determine a set of spatial distribution metrics. Each spatial distribution metric of the set of spatial distribution metrics may characterize the extent to which at least some of the other sets of cells are indicated as interspersed with at least some of the set of tumor cells. The digital pathology imaging system may generate a result based on the set of spatial distribution metrics. The result corresponds to predicting whether and / or to what extent a particular treatment that modulates the immune response will effectively treat a particular condition in the subject. Based on the result, the subject is determined to be eligible for the clinical trial. An indication that the subject is eligible for the clinical trial is output.
[0054] Generating the results may include processing the set of spatial heterogeneity metrics using a trained machine learning model. The trained machine learning model may have been trained using a set of training elements. Each of the sets of training elements may correspond to a different subject who has received a particular treatment associated with the clinical trial. Each of the sets of training elements may include a different set of spatial heterogeneity metrics and a response value indicating whether and / or to what extent the particular treatment has activated an immunological response in the subject.
[0055] In certain embodiments, the medical condition may be a type of cancer, and / or the specific treatment may be an immune checkpoint blockade treatment. The one or more relational location representations may include, for each cell of a set of cells, a set of coordinates identifying the location of a representation of the cell within the digital pathology image. The one or more relational location representations may include, for each of a set of regions within the digital pathology image, a representation of the absolute or relative quantity of tumor cells, stromal cells, and / or immune cells identified as being located within that region. The one or more relational location representations may indicate a distance-based probability that a first type of cell is indicated as being located within a certain distance from a second type of cell. Each of the first and second types may correspond to immune cells, stromal cells, or tumor cells. Genetic sequencing and / or radiological imaging data may be collected for the subject. The results may further depend on characteristics of the genetic sequencing and / or radiological imaging data.
[0056] The term "biological object representation" as used herein refers to a specific portion of an image (e.g., one or more pixels, a defined region of an image, etc.) that has been identified or is being identified as corresponding to a particular type of biological object. A biological object representation may depict a biological object (e.g., a cell). A biological object representation may include one or more pixels and / or one or more voxels. A pixel or voxel of a biological object representation may correspond, for example, to the centroid, edge, center of mass, or entirety of what is predicted to be the biological object representation. A biological object representation may be identified using machine learning algorithms, one or more static rules, and / or computer vision techniques. As applied to digital pathology images, the image may depict a stained section, and the stain may be selected to be preferentially absorbed by a particular type of biological object of interest, such that identifying the biological object representation may include an intensity-based evaluation.
[0057] The term "biological object" as referred to herein may refer to a biological unit. Biological objects may include, by way of example and not limitation, cells, organelles (e.g., nuclei), cell membranes, stroma, tumors, or blood vessels. It will be understood that biological objects may include three-dimensional objects, and that a digital pathology image may capture only a single two-dimensional slice of the object and need not even extend completely through the entire object along the plane of the two-dimensional slice. Nevertheless, such a captured portion may be referred to herein as depicting the biological object.
[0058] The term "type of biological object" or "type of biological object" as referred to herein may refer to a category of biological units. By way of example and not limitation, a type of biological object may refer to cells (generally), a specific type of cell (e.g., lymphocytes or tumor cells), cell membranes (generally), etc. Some disclosures may refer to detecting a biological object representation corresponding to a first type of biological object and another biological object representation corresponding to a second type of biological object. The first and second types of biological objects may have similar, the same, or different levels of specificity and / or generality. For example, the first and second types of biological objects may be identified as lymphocyte and tumor cell types, respectively. As another example, the first type of biological object may be identified as a lymphocyte, and the second type of biological object may be identified as a tumor.
[0059] The term "spatial distribution metric" as referred to herein may refer to a metric that characterizes the spatial arrangement of particular biological object representations within an image relative to one another and / or other particular biological object representations. A spatial distribution metric may characterize the extent to which one type of biological object (e.g., lymphocytes) infiltrates another type of biological object (e.g., tumors), is interspersed with another type of object (e.g., tumor cells), is physically proximate to another type of object (e.g., tumor cells), and / or is co-localized with another type of object (e.g., tumor cells).
[0060] FIG. 1 illustrates an interactive system or network 100 (e.g., a specially configured computer system) that may be used in accordance with the disclosed subject matter to generate and process digital pathology images to characterize relative spatial information of biological objects, according to some embodiments.
[0061] The digital pathology imaging system 105 may generate one or more digital images corresponding to a particular sample. For example, an image generated by the digital pathology imaging system 105 may include a stained section of a biopsy sample. As another example, an image generated by the digital pathology imaging system 105 may include a slide image of a liquid sample (e.g., a blood film). As another example, an image generated by the digital pathology imaging system 105 may include a fluorescent microscopy slide image depicting fluorescent in situ hybridization (FISH) after a fluorescent probe binds to a target DNA or RNA sequence.
[0062] Some types of samples (e.g., biopsies, solid samples, and / or tissue-containing samples) may be processed by the sample preparation system 110 to fix and / or embed the sample. The sample preparation system 110 may facilitate infiltrating the sample with a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding substance (e.g., histological wax). For example, the fixation subsystem may fix the sample by exposing the sample to a fixative for at least a threshold time (e.g., at least 3 hours, at least 6 hours, or at least 12 hours). The dehydration subsystem may dehydrate the sample (e.g., by exposing the fixed sample and / or a portion of the fixed sample to one or more ethanol solutions) and potentially clear the dehydrated sample using a clearing intermediate (e.g., including ethanol and histological wax). The embedding subsystem may infiltrate the sample with heated (e.g., liquid) histological wax (e.g., one or more times for a corresponding predetermined period of time). The histological wax may include paraffin wax and potentially one or more resins (e.g., styrene or polyethylene). The sample and wax may then be cooled to block the wax-infiltrated sample.
[0063] The sample slicer 115 may receive a fixed and embedded sample and create a set of sections. The sample slicer 115 may expose the fixed and embedded sample to a cold or low temperature. The sample slicer 115 may then cut the cooled sample (or a trimmed version thereof) to create a set of sections. Each section may have a thickness of (e.g.) less than 100 μm, less than 50 μm, less than 10 μm, or less than 5 μm. Each section may have a thickness of (e.g.) more than 0.1 μm, more than 1 μm, more than 2 μm, or more than 4 μm. Cutting of the cooled sample may be performed in a warm water bath (e.g., at a temperature of at least 30° C., at least 35° C., or at least 40° C.).
[0064] The automated staining system 120 may facilitate one or more staining of sections of a sample by exposing each section to one or more stains (e.g., hematoxylin and eosin, immunohistochemistry, or special stains). Each section may be exposed to a predetermined amount of stain for a predetermined period of time. In certain embodiments, a single section is exposed to multiple stains simultaneously or sequentially.
[0065] Each of the one or more stained sections may be presented to an image scanner 125, which may capture a digital image of the section. The image scanner 125 may include a microscope camera. The image scanner 125 may capture digital images at multiple magnifications (e.g., using a 10x objective, a 20x objective, a 40x objective, etc.). The images may be manipulated to capture selected portions of the sample at a desired magnification range. The image scanner 125 may further capture annotations and / or morphemes identified by a human operator. In certain embodiments, after one or more images have been captured, the sections are returned to the automated staining system 120 so that they may be washed, exposed to one or more other stains, and imaged again. When multiple stains are used, the stains may be selected to have different color profiles so that a first region of the image corresponding to a first section that has absorbed a large amount of a first stain can be distinguished from a second region of the image (or a different image) corresponding to a second section that has absorbed a large amount of a second stain.
[0066] It will be appreciated that one or more components of digital pathology imaging system 105 may, in certain embodiments, be operated in conjunction with a human operator. For example, the human operator may move samples through various subsystems (e.g., sample preparation system 110 or digital pathology imaging system 105) and / or initiate or terminate the operation of one or more subsystems, systems, or components of digital pathology imaging system 105. As another example, some or all of one or more components of the digital pathology imaging system (e.g., one or more subsystems of sample preparation system 110) may be partially or entirely replaced by the actions of a human operator.
[0067] Furthermore, while the various described and illustrated features and components of the digital pathology imaging system 105 relate to processing solid and / or biopsy samples, it will be understood that other embodiments may relate to liquid samples (e.g., blood samples). For example, the digital pathology imaging system 105 may be configured to receive a liquid sample (e.g., blood or urine) slide including a base slide, a contaminated liquid sample, and a cover. The image scanner 125 may then capture an image of the sample slide. Further embodiments of the digital pathology imaging system 105 may relate to capturing images of the sample using advanced imaging techniques, such as FISH, as described herein. For example, once fluorescent probes are introduced into the sample and allowed to bind to target sequences, appropriate image processing may be used to capture an image of the sample for further analysis.
[0068] A given sample may be associated with one or more users (e.g., one or more physicians, laboratory technicians, and / or healthcare providers). Associated users may include the person who ordered the test or biopsy that produced the sample being imaged and / or the person authorized to receive the results of the test or biopsy. For example, a user may correspond to a physician, pathologist, clinician, or subject (from whom the sample was taken). A user may use one or more devices 130 to initially submit one or more requests (e.g., identifying the subject) for a sample to be processed by digital pathology image generation system 105 and the resulting images to be processed by digital pathology image processing system 135 (for example).
[0069] In certain embodiments, digital pathology image generation system 105 transmits the digital pathology image generated by image scanner 125 back to user device 130, which communicates with digital pathology image processing system 135 to initiate automated processing of the digital pathology image. In certain embodiments, digital pathology image generation system 105 provides the digital pathology image generated by image scanner 125 directly to digital pathology image processing system 135, for example, at the direction of a user of user device 130. Although not shown, other intermediate devices (e.g., a data store on a server connected to digital pathology image generation system 105 or digital pathology image processing system 135) may be used. Furthermore, for simplicity, only one digital pathology image processing system 135, digital pathology image generation system 105, and user device 130 are shown in network 100. The present disclosure contemplates the use of one or more of each type of system and its components without necessarily departing from the teachings of the present disclosure.
[0070] The digital pathology image processing system 135 may be configured to identify spatial characteristics of the image and / or characterize the spatial distribution of depictions of biological objects. The section aligner subsystem 140 may be configured to align multiple digital pathology images and / or regions of the digital pathology images corresponding to the same sample. For example, the multiple digital pathology images may correspond to the same section of the same sample. Each image may depict a section stained with a different stain. As another example, each of the multiple digital pathology images may correspond to a different portion of the same sample (e.g., each corresponding to the same stain, or different subsets of the images corresponding to different stains). For example, alternating sections of the sample may be stained with different stains.
[0071] The section aligner subsystem 140 may determine whether and / or how each digital pathology image is translated, rotated, magnified, and / or stretched so that the digital pathology images corresponding to a single sample and / or a single section are aligned. The alignment may be determined using (for example) a correlation evaluation (e.g., to identify an alignment that maximizes correlation).
[0072] The biological object detector subsystem 145 may be configured to automatically detect depictions of one or more specific types of objects (e.g., biological objects) in each of the aligned digital pathology images. The object types may include, for example, cells of a type of biological structure. For example, a first set of biological objects may correspond to a first type of cell (e.g., immune cells, white blood cells, lymphocytes, tumor-infiltrating lymphocytes, etc.), and a second set of biological objects may correspond to a second type of cell (e.g., tumor cells, malignant tumor cells, etc.) or type of biological structure (e.g., tumor, malignant tumor, etc.). The biological object detector subsystem 145 may detect depictions of one or more types of respective biological objects from the aligned digital pathology images. The digital pathology images may depict various stains in a single digital pathology image. Such a digital pathology image may include a single image that may correspond to sections of a sample stained with each of multiple stains. For example, the biological object detector subsystem 145 may detect depictions of lymphocytes and tumor cells from a single digital pathology image. The biological object detector 145 may detect depictions of biological objects from different digital pathology images corresponding to different stains, for example.
[0073] For example, representations of lymphocytes may be detected in a first digital pathology image, and representations of tumor cells may be detected in a second digital pathology image. The first digital pathology image may depict an image of a section of a sample stained with a first stain, and the second digital pathology image may depict the same section stained with a second stain and re-imaged. The biological object detector subsystem 145 may detect representations of a first specific type of biological object in the first digital pathology image, which may correspond to a section of the sample stained with the first stain. The biological object detector subsystem 145 may detect representations of a second specific type of biological object shown in the second digital pathology image, which may correspond to the same section stained with a second stain or a different section of the sample stained with the second stain. Furthermore, the biological object detector subsystem 145 may detect one or more biological objects of one or more types of biological objects in one or more digital pathology images not associated with the same sample for the purpose of generating spatial distribution metrics and subject-level results.
[0074] The biological object detector subsystem 145 may detect and characterize biological objects using static rules and / or trained models. Rule-based biological object detection may include detecting one or more edges, identifying a subset of edges whose shapes are sufficiently connected and closed, and / or detecting one or more high-intensity regions or pixels. For example, if the area of the region within the closed edge is within a predetermined range and / or if the high-intensity region has a size within a predetermined range, a portion of the digital pathology image may be determined to depict a biological object. Detecting depictions of biological objects using a trained model may include using a neural network, such as a convolutional neural network, a deep convolutional neural network, and / or a graph-based convolutional neural network. The model may be trained using annotated images that include annotations indicating the location and / or boundaries of the object. The annotated images may be received from a data repository (e.g., a public data store) and / or from one or more devices associated with one or more human annotators. The model may be trained using generic or natural images (e.g., not just images commonly captured for digital pathology or medical use). This may extend the ability of the model to distinguish between different types of biological objects, which may have been trained using a specialized training set of images, such as digital pathology images, selected to train the model to detect specific types of objects.
[0075] Rule-based biological object detection and trained model biological object detection may be used in any combination. For example, rule-based biological object detection may detect one type of biological object representation, and the trained model is used to detect another type of biological object representation. Another example may include using the biological objects output by the trained model to validate results from rule-based biological object detection, or validating the results of a trained model using a rule-based approach. Yet another example may include using rule-based biological object detection as initial object detection, followed by using the trained model for more sophisticated biological object analysis, or applying a rule-based object detection approach to images after an initial set of biological object representations has been detected via the trained network.
[0076] Biological object detection may also include (for example) preprocessing the digital pathology image. Preprocessing may convert the resolution of the digital pathology image to a target resolution, apply one or more color filters, and / or normalize the digital pathology image for use by the rule-based biological object detection method or trained model. For example, a color filter may be applied that passes colors corresponding to the color profile of the stain used by the automated staining system 120. Rule-based biological object detection or trained model biological object detection may be applied to the preprocessed image.
[0077] For each detected biological object, the biological object detector subsystem 145 may identify and store a representative location of the depicted biological object (e.g., a centroid or midpoint), a set of pixels or voxels corresponding to an edge of the depicted object, and / or a set of pixels or voxels corresponding to a region of the depicted biological object. This biological object data may be stored along with biological object metadata, which may include, by way of example and not limitation, a biological object identifier (e.g., a numerical identifier), an identifier of the corresponding digital pathology image, an identifier of the corresponding region within the corresponding digital pathology image, a corresponding subject identifier, and / or an object type identifier.
[0078] The biological object detector subsystem 145 may generate an annotated digital pathology image that includes the digital pathology image and further includes one or more overlays that identify where in the image the detected biological objects are depicted. In certain embodiments in which multiple types of biological objects are detected, for example, different colors may be used to represent the different types of annotations.
[0079] The biological object distribution detector subsystem 150 may be configured to generate and / or characterize a spatial distribution of one or more objects. The distribution may be generated (for example) by using one or more static rules (e.g., identifying how to apply a distance-based metric of a point-location representation of the biological objects, using absolute or smoothed counts or densities of biological objects within a grid region of the digital pathology image, etc.) and / or by using a trained machine learning model (e.g., capable of predicting that initial object depiction data should be adjusted to account for the predicted quality of one or more digital pathology images). For example, the characterization may indicate the extent to which certain types of biological objects are depicted densely together, the extent to which depictions of certain types of biological objects are spread across all or a portion of an image, the extent to which the proximity of depictions of certain types of biological objects compares to the proximity of depictions of other types of biological objects (relative to each other), the proximity of depictions of one or more certain types of biological objects to depictions of one or more other types of biological objects, and / or the extent to which depictions of one or more certain types of biological objects are within and / or proximate to a region defined by one or more of one or more other types of biological objects. As described in further detail below in connection with FIG. 2 , the biological object distribution detector subsystem 150 may initially generate a representation of the biological object using a particular framework (e.g., a spatial point process analysis framework, a spatial domain analysis framework, or a geostatistical analysis framework).
[0080] The subject-level label generation subsystem 155 may generate one or more subject-level labels using spatial distribution metrics. Subject-level labels may include labels determined for individual subjects (e.g., patients), defined subject groups (e.g., patients with similar characteristics), clinical trial groups, etc. Labels may correspond, for example, to possible diagnoses, prognoses, treatment assessments, treatment recommendations, or treatment eligibility determinations. In certain embodiments, labels may be generated using predefined or learned rules. For example, a rule may dictate that spatial distribution metrics above a predetermined threshold should be associated with a particular disease state (e.g., as a potential diagnosis), while metrics below the threshold are not associated with the particular disease state. As another example, a rule may indicate that a particular treatment should be recommended when the spatial distribution metric is within a predetermined range (e.g., but not otherwise). Illustratively, checkpoint immunotherapy may be recommended if a distance-based metric (e.g., characterizing how far the center of gravity of a lymphocyte representation is from the center of gravity of a tumor cell representation) is below a predetermined threshold. As yet another example, the rules may identify different bands of treatment effectiveness based on the ratio of a spatial distribution metric corresponding to a recently acquired digital pathology image to a stored baseline spatial distribution metric corresponding to a less recently acquired digital pathology image.
[0081] The object-level label generation subsystem 155 may further use one or more patterns or masks, for example, in conjunction with a spatial distribution metric, to generate one or more object-level labels. In certain embodiments, the object-level label generator subsystem 155 may retrieve or provide one or more patterns or masks associated with previous labels and / or object results (which may help validate the labels). In certain embodiments, the object-level label generator subsystem 155 may derive masks according to one or more rules or using a trained model. For example, a rule may indicate that a particular mask or subset of masks should be retrieved and compared with the digital pathology image in response to a determination of one or more types of one or more biological objects depicted in the digital pathology image. As another example, a rule may indicate that a particular mask or subset of masks should be retrieved and compared with the digital pathology image in response to a determination of a spatial distribution metric that meets or does not meet a threshold, or that occupies or does not occupy a threshold range. Values associated with the rules may be learned by the object-level label generation subsystem 155. In certain embodiments, one or more machine learning processes described herein may be used to train a model to identify patterns for searching and applying digital pathology images based on the overall characteristics of the digital pathology image, the data derived therefrom, and the metadata associated therewith.
[0082] The digital pathology imaging system 135 may output the generated spatial distribution metrics, object-level labels, and / or annotated images. Output may include local presentation or transmission (e.g., to a user device 130).
[0083] Each component and / or system in Figure 1 may include (for example) one or more computers, one or more servers, one or more processors, and / or one or more computer-readable media. In certain embodiments, a single computing system (having one or more computers, one or more servers, one or more processors, and / or one or more computer-readable media) may include multiple components shown in Figure 1. For example, digital pathology imaging system 135 may include a single server and / or a collection of servers that collectively implement all of the functionality of section aligner subsystem 140, biological object detector subsystem 145, biological object distribution detector subsystem 150, and object-level label generator subsystem 155.
[0084] It will be understood that various alternative embodiments are contemplated. For example, digital pathology image processing system 135 may not have an object-level label generator subsystem 155 and / or may not generate object-level labels. Rather, annotated images (with annotations generated by biological object detector subsystem 145) and / or one or more spatial distribution metrics (generated by biological object distribution detector subsystem 150) may be output by digital pathology image processing system 135. A user may then consider the output data to identify a label (e.g., corresponding to a diagnosis, prognosis, treatment assessment, or treatment recommendation).
[0085] 2 illustrates an exemplary biological object pattern computation system 200 for processing object data to generate a spatial distribution metric, according to some embodiments of the present invention. The biological object distribution detector subsystem 150 may include part or all of system 200.
[0086] The biological object pattern computation system 200 includes multiple subsystems: a point processing subsystem 205, a region processing subsystem 210, and a geostatistics subsystem 215. Each subsystem corresponds to a different framework, a point processing analysis framework 225, a region analysis framework 230, or a geostatistics framework 235, and uses them to generate spatial distribution metrics or their constituent data. The point processing analysis framework 225 may have an object-specific focus, for example, identifying a point location for each detected biological object depiction. The region analysis framework 230 may be a framework in which data (e.g., the locations of depicted biological objects) are indexed using coordinates and / or spatial grids rather than by individual biological object depictions. The geostatistics analysis framework 235 may provide predictions of the prevalence and / or observation probability of a particular type of biological object depiction at each of a set of locations. Each framework may support the generation of one or more metrics characterizing the spatial patterns and / or distributions formed across one or more biological object depictions of each of one or more types.
[0087] For example, the point processing subsystem 205 may use a point processing analysis framework 225 that may represent each biological object representation as a point location within the image. In certain embodiments, the point location may be the centroid, midpoint, center of mass, or the like of the biological object representation. In some embodiments, the point location is detected upon detection of the biological object representation (e.g., by the biological object detector subsystem 145). In some embodiments, the point processing subsystem 205 determines the location of the biological object representation (e.g., based on a location relative to an edge and / or region of the depicted biological object). The point processing subsystem 205 may include a distance detector 245 for detecting and processing one or more distances between the biological object representations, a point-based cluster generator 250 and a correlation detector 255 for characterizing cross-correlation and / or autocorrelation between one or more biological object representations of each of one or more types, and a three-dimensional landscape generator 260 that scales the computational complexity of the biological object representations across a two-dimensional space corresponding to the dimensionality of the image (e.g., the third dimension of the landscape indicates the computational complexity). Cross-correlation and auto-correlation can identify the probability, as a function of distance, that a point representing a first type of biological object representation (and thus a biological object in a sample) is located away from the observed biological object representation. In the case of cross-correlation, the probability is calculated for the second type of biological object. In the case of auto-correlation, the probability is calculated for the first type of biological object. Cross-correlation or auto-correlation can include one-dimensional representations (e.g., with the x-axis set to distance) or two-dimensional representations (e.g., with the x-axis set to horizontal distance and the y-axis set to vertical distance).
[0088] The distance detector 245 may detect points and the locations of each point within the image. For each of one or more pairs of points (e.g., "point pairs"), the distance (e.g., Euclidean distance) between the locations of the points associated with that pair is calculated. Each of the one or more pairs of points may correspond to depictions of the same type of biological object or depictions of different types of biological objects. For example, for a given depicted lymphocyte, the distance detector 245 may identify the distance between the location of the depicted lymphocyte and other depicted lymphocytes, and the distance detector 245 may identify the distance between the location of the depicted lymphocyte and each depicted tumor cell. The distance detector 245 may generate one or more spatial distribution metrics based on statistics. For example, the spatial distribution metric may be defined as and / or based on the mean, median, and / or standard deviation of the distances between depictions of a given type of biological object and / or between depictions of one or more different types of biological object. For example, the distance between the locations of all depicted lymphocytes may be detected, and then the average distance may be calculated. Similar calculations may be performed based on the distance between each lymphocyte-tumor-cell pair. The spatial distribution metric may be based on a first statistic generated based on the distance between representations of a first type of biological object and a second statistic generated based on the distance between representations of a second type of biological object.
[0089] The point-based cluster generator 250 may use distances to perform cluster analysis (e.g., multi-distance spatial cluster analysis such as Ripley's K function). For example, a K value generated using Ripley's K function may represent an estimated degree to which the spatial distribution of biological object depictions corresponds to a spatially random distribution (e.g., as opposed to a distribution with one or more spatial clusters).
[0090] The correlation detector 255 may use the distances and / or point locations to generate one or more correlation-based metrics. A correlation-based metric may indicate the degree to which the presence of a given type of biological object representation at one location predicts whether another biological object representation of the given type or another type is present at another location. The other locations may be designated, for example, based on predetermined spatial increments or target regions surrounding the biological object representation. For example, a cross-correlogram may identify the probability of observing a tumor cell representation within each of various distance ranges from a lymphocyte representation. The metric may identify the sum of probabilities over distances from zero to a specific distance. The correlation-based metric may include a randomized dependence coefficient or correlation coefficient. In certain embodiments, the correlation-based metric indicates a distance value associated with the maximum value of the cross-correlogram.
[0091] The landscape generator 260 may use the point locations of the biological object representations of a given type to generate a three-dimensional "landscape" data structure (e.g., a landscape map) that indicates the probability that a representation of the given type of object will be observed for each horizontal and vertical position in the image. The landscape data structure may be identified by applying one or more algorithms. For example, a data structure (or other peak structure) configured to represent zero, one, or more Gaussian distributions may be applied. The landscape generator 260 may be configured to compare a landscape data structure generated for a given type of biological object with another landscape data structure generated for another type of biological object. For example, the landscape generator 260 may compare the position, amplitude, and / or width of one or more peaks in a landscape corresponding to a given type of biological object with the position, amplitude, and / or width of one or more peaks in another landscape data structure corresponding to another type of biological object. When the landscape is represented in three dimensions and visualized, it may include peaks that indicate a high probability that certain types of objects are present in the corresponding regions. While the landscape data representation depicts the density and / or number of objects through three dimensions, the same data may alternatively be conveyed using other visualization techniques (e.g., via a heat map). Exemplary landscape data structures generated by landscape generator 260 are shown in FIG. 4 as landscape representations 420a and 420b.
[0092] While the point process analysis framework 225 may index the data by individual depictions of biological objects, the surface analysis framework 230 may index the data in a more abstract sense using coordinates and / or spatial grids. The region processing subsystem 210 may apply the region analysis framework 230 to identify density (or count) for each set of coordinates and / or regions associated with image regions. The density may be identified using one or more of a grid-based divider 265, a grid-based cluster monitor, and / or a hotspot monitor 275.
[0093] The grid-based divider 265 can impose a spatial grid on the image that includes a representation of the locations of biological objects depicted on the image. The spatial grid, which includes a set of rows and a set of columns, can define a set of regions, each region corresponding to a row-column combination. Each row can have a specified height and each column can have a specified width, such that each region of the spatial grid can have a specified area.
[0094] The grid-based divider 265 may determine an intensity metric using the spatial grid and point locations of the biological object representations. For example, for each grid region, the intensity metric may indicate and / or be based on the quantity of one or more types of respective biological object representations that have point locations within the region (e.g., for at least a threshold portion of the biological object representations). In certain embodiments, the intensity metric may be normalized and / or weighted based on the total number of biological objects (e.g., of a given type) detected within the digital pathology image and / or for the sample, based on the count of biological objects of a given type detected in other samples, and / or based on the scale of the digital pathology image. In certain embodiments, the intensity metric is smoothed and / or otherwise transformed. For example, the initial count may be thresholded so that the final intensity metric is binary. For example, a binary metric may include determining whether the grid region is associated with several biological object representations that meet a threshold (e.g., whether there are at least five tumor cells assigned to the region). In certain embodiments, the grid-based divider 265 may use the surface data to generate one or more spatial distribution metrics by (for example) comparing intensity metrics across different types of biological objects.
[0095] The grid-based cluster generator 270 may generate one or more spatial distribution metrics based on the cluster-association data related to one or more types of biological objects. For example, for each of one or more types of biological objects, clustering and / or fitting techniques may be applied to determine the extent to which depictions of that type of biological object are spatially clustered, e.g., with each other and / or with depictions of other types of biological objects. Clustering and / or fitting techniques may be further applied to determine the extent to which the depictions of biological objects are spatially dispersed and / or randomly distributed. For example, the grid-based cluster generator 270 may determine a Morita-Horn index and / or a Moran index. For example, a single metric may indicate the extent to which depictions of one type of biological object are spatially clustered and / or in close proximity to depictions of another type of object.
[0096] The hot spot / cold spot monitor 275 may perform analysis to detect "hot spot" locations where representations of one or more particular types of biological objects are likely to be present, or "cold spot" locations where representations of one or more particular types of biological objects are likely to be absent. In certain embodiments, a gridded intensity metric may be used to (for example) identify local intensity extrema (e.g., maxima or minima) and / or fit one or more peaks that may be characterized as hot spots or one or more valleys that may be characterized as cold spots. In certain embodiments, a Getis-Ord hot spot algorithm may be used to identify any hot spots (e.g., intensities over a set of adjacent pixels that are sufficiently high to be significantly different compared to other intensities in the digital pathology image) or any cold spots (e.g., intensities over a set of adjacent pixels that are sufficiently low to be significantly different compared to other intensities in the digital pathology image). In certain embodiments, "significantly different" may correspond to a determination of statistical significance. Once object type-specific hot spots and cold spots are identified, the hot spot / cold spot monitor 275 may compare the location, amplitude, and / or width of any hot spots or cold spots detected for one type of biological object with the location, amplitude, and / or width of any hot spots / cold spots detected for another type of biological object.
[0097] The geostatistical subsystem 215 may use the geostatistical analysis framework 235 to estimate the underlying smoothed distribution based on the discrete samples. The geostatistical analysis framework 235 may be configured to convert data corresponding to one dimension and / or resolution to two dimensions and / or resolution. For example, the locations of biological object depictions may first be defined using a 1 mm resolution across the digital pathology image. The location data may then be fitted to a continuous function that is not constrained to mm resolution. As another example, the locations of biological object depictions initially defined as two-dimensional coordinates may be transformed to generate a data structure including the number of biological object depictions in each set of row and column combinations. The geostatistical analysis framework 235 may be configured to fit a function (for example) using multiple data points identifying the locations of specific biological object depictions (of a given type). For example, a variogram may be generated for each type of biological object, indicating, for each of a series of distances, whether two biological objects of the same type were detected at a distance apart. A single type of object may be more likely to be detected at short separation distances compared to longer distances. A semivariogram may then be generated by fitting the variogram data. The observed biological objects and semivariograms may then be used by the geostatistical subsystem 215 to generate an image map that predicts the prevalence and / or probability of observing depictions of particular types of biological objects at each of a set of locations. The resolution and / or size of the image map may be higher and / or larger, respectively, as compared to the one or more digital pathology images that were processed to initially detect the biological object depictions.The geostatistical subsystem 215 may generate one or more spatial distribution metrics using the geostatistical data by (for example) comparing predicted biological object values (e.g., predicted prevalence and / or observation probability) across different types of biological objects, characterizing spatial correlation of predicted biological object values between different types of biological objects, characterizing spatial autocorrelation using predicted object values of individual types of biological objects, and / or comparing the locations of spatial clusters (or hot spots / cold spots) of predicted object values across different types of objects.
[0098] It will be understood that various subsystems may include components not shown and may perform processing not explicitly described. For example, the region processing subsystem 210 may generate a spatial distribution metric corresponding to an entropy-based mutual information measure to indicate the extent to which information about the location of a representation of a first type of biological object within a given region reduces uncertainty about whether a representation of another biological object (of the same or other type) is present at that location within another region. For example, the mutual information metric may indicate that the location of one type of biological object provides information about the location of another type of biological object (thus reducing entropy). Such mutual information may potentially be associated with cases where one type of cell is interspersed with another type of cell (e.g., tumor-infiltrating lymphocytes interspersed within tumor cells).
[0099] As another example, the point processing subsystem 205 may generate nearest neighbor distance metrics based on the distances (or distance statistics) between individual biological object detection points of a given type of biological object and one or more other nearest points corresponding to biological object depictions of the same type of biological object and / or another type of biological object. To illustrate, for each depiction of a biological object, the intra-object type distance value may refer to the average distance between the location of the biological object depiction and the location of the nearest number of depictions of the same type of biological object. The intra-object type distance statistic value of a biological object may refer to (for example) the average or median of the intra-object type distance values of all biological object depictions of the object type. The inter-object type distance value may refer to the average distance between the location of the biological object depiction and the location of the nearest number of depictions of a different type of object. The inter-object distance statistic may be (for example) the average or median of the inter-object distance values. A small / low inter-object type distance statistic may indicate that depictions of different types of biological objects are close to each other. The intra-object type distance statistic may be used (for example) for normalization purposes or to assess general clustering of biological objects of a given type.
[0100] As yet another example, the point processing subsystem 205 may generate a correlation-based metric based on a cross- and / or autocorrelation function, such as a pairwise correlation (cross-type) function or a mark correlation function. The correlation function may include (for example) a correlation value as a function of distance. The baseline correlation value may correspond to a random distribution. The metric may include the spatial distance at which the correlation function (or a smoothed version of the correlation function) crosses the baseline correlation value (or some adjusted version of the baseline correlation value, such as a threshold calculated by adding a fixed amount to the baseline correlation value and / or multiplying the baseline correlation value by a predefined coefficient).
[0101] The biological object pattern computation system 200 may generate results (which may themselves be spatial distribution metrics) using a combination of multiple (e.g., two or more, three or more, four or more, or five or more) spatial distribution metrics (e.g., such as those disclosed herein) of various types. The multiple spatial distribution metrics may include metrics generated using different frameworks (e.g., two or more, three or more, or all of the point process analysis framework 225, the area analysis framework 230, and the geostatistics framework 235) and / or metrics generated by different subsystems (e.g., two or more, three or more, or all of the point processing subsystem 205, the area processing subsystem 210, and the geostatistics subsystem). For example, a spatial distribution metric may be generated using a distance-based metric (generated using the spatial point process analysis framework) and a Morisita-Horn index metric (generated using the spatial area analysis framework).
[0102] In particular embodiments, multiple metrics may be combined using one or more user-defined and / or predefined rules and / or using a trained model. For example, the machine learning (ML) model controller 295 may train a machine learning model to learn one or more parameters (e.g., weights) that specify how various lower-level metrics should be processed together to generate an integrated spatial distribution metric. The integrated spatial distribution metric may be more accurate overall than the individual parameters alone. The architecture of the machine learning model may be stored in the ML model architecture data store 296. For example, the machine learning model may include logistic regression, linear regression, decision tree, random forest, support vector machine, or neural network (e.g., feedforward neural network), and the ML model architecture data store 296 may store one or more equations that define the model. Optionally, the ML model hyperparameter data store 297 stores one or more hyperparameters used to define the model and / or its training but that are not learned. For example, the hyperparameters may identify the number of hidden layers, dropout, learning rate, etc. The learned parameters (e.g., corresponding to one or more weights, thresholds, coefficients, etc.) may be stored in the ML model parameters data store 298.
[0103] In particular embodiments, some or all of one or more subsystems are trained using some or all of the same set of training data used to train the ML model (thereby learning the ML model parameters store in ML model parameter data store 298). In particular embodiments, a different training dataset is used to train one or more subsystems as compared to the ML model controlled by ML model controller 295. Similarly, when multiple frameworks, subsystems, and / or subsystem components are used to generate metrics that are combined to generate a spatially distributional metric, individual frameworks, subsystems, and / or subsystem components may be trained using non-overlapping, partially overlapping, fully overlapping, or the same training dataset with respect to other training datasets.
[0104] 2, the biological matter pattern computation system 200 may further include one or more components for aggregating spatial distribution metrics across slices of a sample of interest to generate one or more aggregated spatial distribution metrics. Such aggregated metrics may be generated (for example) by a component within a subsystem (e.g., hot spot monitor 275), by a subsystem (e.g., by point processing subsystem 205), by the ML model controller 295, and / or by the biological matter pattern computation system 200. The aggregated spatial distribution metric may include (for example) a sum, median, mean, maximum, or minimum value of a set of slice-specific metrics.
[0105] 3A and 3B show processes 300a and 300b for providing a health-related assessment based on image processing of a digital pathology image using spatial distribution metrics, according to some embodiments. More specifically, the digital pathology image may be processed, for example, by a digital pathology imaging system to generate one or more metrics that characterize the spatial pattern and / or distribution of one or more cell types, which may then inform a diagnosis, prognosis, treatment evaluation, or treatment eligibility determination. The process begins at step 310, where a subject-associated identifier may be received by the digital pathology imaging system (e.g., digital pathology imaging system 135). The subject-associated identifier may include an identifier for the subject, sample, section, and / or digital pathology image. The subject-associated identifier may be provided by a user (e.g., the subject's healthcare provider and / or physician). For example, the user may provide the identifier as an input to a user device, which may transmit the identifier to digital pathology imaging system 135.
[0106] In step 315, the digital pathology image processing system 135 may access one or more digital pathology images of the stained tissue sample associated with the identifier. For example, a local or remote data store may be queried using the identifier. As another example, a request including the identifier may be sent to another system (e.g., a digital pathology image generation system), and the response may include an image. The image may depict a stained section of the sample from the subject. In certain embodiments, a first digital pathology image shows a section stained with a first stain, and a second digital pathology image shows a section stained with a second stain. In certain embodiments, a single digital pathology image shows sections stained with multiple stains. In certain embodiments, the digital pathology image may be separated into regions or tiles before or during the analysis step 300a. The separation may be based on user-directed focus on a particular region, detected regions of interest (e.g., detected according to rules based on machine learning methods, etc.), or the like.
[0107] In step 320, a first set of biological object representations of a first type and a second set of biological object representations of a second type may be detected from the digital pathology image. In certain embodiments, the first type of object may correspond to biological object associated with a first stain, and the second type of object may correspond to biological object associated with a second stain. The first type of object may correspond to a first type of biological object (e.g., a first cell type), and the second type of object may correspond to a second type of biological object (e.g., a second cell type).
[0108] Each biological object may be associated with location metadata that indicates where the object is depicted in the digital pathology image. The location metadata may include (for example) a set of coordinates corresponding to a point in the image, coordinates corresponding to an edge or boundary of the biological object depiction, and / or coordinates corresponding to an area of the depicted object. For example, the detected biological object depiction may correspond to a 5x5 square of pixels in the image under analysis. The location metadata may identify all 25 pixels of the biological object depiction, 16 pixels along the boundary, or a single representative point. The single representative point may be (for example) the midpoint or may be generated by pre-weighting each of the 25 pixels using intensity values and then calculating a weighted center point. Other weightings may also be applied, such as weighting that takes content or context into account.
[0109] In step 325, a data structure is generated based on the biological object representations detected in step 320. The data structure may include object information characterizing the biological object representations. For each detected biological object representation, the data structure may identify, for example, the centroid of the biological object representation, pixels corresponding to the perimeter of the biological object representation, or pixels corresponding to the region of the biological object representation. For each biological object representation, the data structure may further identify the type of biological object (e.g., lymphocyte, tumor cell, etc.) corresponding to the depicted biological object.
[0110] In step 330, one or more spatial distribution metrics are generated. The spatial distribution metrics characterize the relative positions of the biological object representations. In some cases, step 330 may include generating a spatial distribution metric based on the detected biological object representations and object types of example step 320. For example, the spatial distribution metric may characterize how close and / or clustered representations of a particular type of object are to each other and / or to representations of another particular type of object.
[0111] In step 335, the spatial distribution metrics generated in step 330 are output to a storage entity / database, a user interface, or a service platform. The service platform may use the output spatial distribution metrics to provide further analysis. The spatial distribution metrics may be transmitted to a user device (which may present the metrics to a user) and / or may be presented locally via a user interface. In certain embodiments, images and / or annotations corresponding to the detected biological object depictions are additionally output (e.g., transmitted and / or output).
[0112] In certain embodiments, a user may use spatial distribution metrics to inform a subject's diagnosis, prognosis, treatment recommendation, or treatment eligibility determination. For example, immunotherapy and / or checkpoint immunotherapy may be identified as treatment recommended if the spatial distribution metrics indicate that lymphocytes are close to and / or co-localized with tumor cells. If (for example) a metric representing the distance between lymphocytes and tumor cells is similar (e.g., less than 300%, less than 200%, less than 150%, or less than 110%) to a metric representing the distance between the same cell type (e.g., lymphocytes or tumor cells), lymphocytes may be determined to be close to tumor cells or tumor cells may be interspersed. If intensity values representing the amount of each cell type assigned to individual regions within the image are similar, lymphocytes may be determined to be close to tumor cells and / or tumor cells may be interspersed. For example, the analysis may determine whether intensity values indicate that cell types are densely located within the same or similar subsets of image regions.
[0113] A user may provide a diagnosis, prognosis, etc. to a subject. For example, the diagnosis, prognosis, etc. may be communicated to the subject verbally and / or transmitted from the user's device to the subject's device (e.g., via a secure portal). The user may further use the user device to update the subject's electronic health record to include the diagnosis, prognosis, etc.
[0114] As a result of the recommendation, a subject's treatment may be initiated, modified, or stopped. For example, in response to a diagnosis of a subject with a particular disease, a recommended treatment may be initiated and / or an approved treatment for a particular disease may be initiated.
[0115] 3B shows another process 300b for providing a health-related assessment based on image processing of a digital pathology image using a spatial distribution metric, according to some embodiments. Steps 305-330 of process 300b are generally similar to steps 305-330 of process 300a. However, in certain embodiments, the digital pathology image processing system 135 may use the spatial distribution metric to predict a diagnosis, prognosis, treatment recommendation, or treatment eligibility determination for a subject (e.g., in step 347). The prediction may be generated using one or more rules that identify one or more thresholds and / or ranges for the metric. The prediction may include a result that represents a diagnosis, prognosis, or treatment recommendation. The outcome can be (for example) a binary value (e.g., predicting whether a subject has a particular disease state): a categorical value (e.g., predicting tumor stage or identifying a particular treatment among a set of potential treatments), or a numeric value (e.g., identifying the probability that a subject has a given condition, predicting the probability that a given treatment will slow disease progression, and / or predicting the time until the condition progresses to the next stage). Treatment recommendations can include the use of checkpoint blockade therapy or immunotherapy (e.g., if metrics indicate that tumor cells are lymphocyte-infused).
[0116] The results may be generated by a trained machine learning model, such as, by way of example and not limitation, a trained regression, decision tree, or neural network model. In certain embodiments, the spatial distribution metrics include multiple different types of metrics, and the model is configured to process multi-type data. For example, the set of metric types may include metrics defined based on K-nearest neighbor analysis, metrics defined based on Ripley's K-function, Morisita-Horn index, Moran index, metrics defined based on correlation functions, metrics defined based on hotspot analysis, and metrics defined based on Kriging interpolation (e.g., ordinary Kriging or indicator Kriging), and the results may be generated based on at least two, at least three, or at least four metrics of the set of metric types.
[0117] In step 348, the digital pathology imaging system 135 may output the prediction (which may include outputting the results) to a storage entity / database, a user interface, or a service platform. For example, the prediction may be presented locally and / or transmitted to a user device (which may, for example, display or present the prediction). The digital pathology imaging system 135 may further output (and the user may further receive) annotation data identifying spatial distribution metrics, digital images, and / or detected biological object depictions.
[0118] The user may then identify a definitive diagnosis, prognosis, treatment recommendation, or determination of treatment eligibility. The confirmed diagnosis, prognosis, etc. may be consistent and / or correspond to the predicted diagnosis, prognosis, etc. The prediction (and / or other data) generated by the digital pathology imaging system may inform the user's decision regarding which diagnosis, prognosis, or treatment recommendation is identified. In certain embodiments, feedback may be provided from the user to the digital pathology imaging system, the feedback indicating whether the user-identified diagnosis, prognosis, or treatment recommendation is consistent with that of the prediction. Such feedback may be used to train models and / or update rules relating spatial distribution metrics to prediction outputs.
[0119] Figure 4 illustrates various stages in identifying spatial patterns and distribution metrics. For example, Figure 4 illustrates an initial digital pathology image, detection results of biological object features from the received image, point process analysis of the image based on the detected biological object features, and spatial distributions (shown as landmark assessments) showing the locations / intensities of the biological object features detected in the received image. The spatial distributions are shown as landmark assessments, and the detected objects are lymphocytes and tumor cells.
[0120] FIG. 4 shows a digital pathology image 405 of an exemplary stained section of a subject's tissue biopsy. The tissue biopsy was collected, fixed, embedded, and sectioned. Each section may be stained with H&E stain and imaged. Hematoxylin in the stain may stain certain cellular structures (e.g., cell nuclei) a first color, while eosin in the stain stained the extracellular matrix and cytoplasm pink. The digital pathology image 405 was processed (using a deep neural network) to detect depictions of two types of objects: lymphocytes and tumor cells. The object data was processed according to various image processing frameworks and techniques (described below) to generate spatial distribution metrics (described below).
[0121] Some embodiments include new and modified frameworks and metrics, as well as new uses of the frameworks and metrics, for processing digital pathology images.
[0122] Table 410 shown in FIG. 4 includes exemplary biological object data that identifies, for each of a plurality of biological object representations, an object identifier associated with the biological object, the type of stain used to stain the sample prior to imaging, the type of biological object (e.g., lymphocyte or tumor cell), and the coordinates of the center of the biological object representation in the digital pathology image. An object detector (e.g., biological object detector subsystem 145) was used to create table 410 and identify a single point location for each biological object representation. The single point location was defined as the centroid point of the biological object representation. Based on table 410, a point process analysis framework was implemented.
[0123] Lymphocyte point image 415a shows lymphocyte point representations 417a in tumor cell coordinates for all detected lymphocyte representations, and tumor cell point image 415b shows point representations 417b in point coordinates for all detected tumor cell representations.
[0124] The exemplary landscape representations 420a and 420b graphically depict three-dimensional landscape data of feature types of biological objects, in this case lymphocyte and tumor cell feature types, respectively.
[0125] Three-dimensional landscape data for landscape representations 420a and 420b may be generated using point data for each of two types of biological objects (e.g., as shown in table 410). The x-axis and y-axis of landscape representation 420a may correspond to the x-axis and y-axis of image 405 and lymphocyte point image 415a (for example). In certain embodiments, the x-axis and y-axis of landscape representation 420b may correspond to the x-axis and y-axis of digital image 405 and tumor cell point image 415b. The landscape data may further include z-values that characterize the computational complexity of biological object representations of a given type detected within an area corresponding to the (x,y) coordinates. Each (x,y) coordinate pair in the landscape data corresponds to a range of x-values and a range of y-values. Thus, the z-value may be determined based on the number of biological object representations of a given type located across an area defined by the range of x-values (corresponding to a portion of the total width of the landscape) and the range of y-values (corresponding to a portion of the total length of the landscape).
[0126] The three-dimensional representation facilitates determining how the density of representations of one type of biological object compares to the density of representations of another type of biological object in a given portion of an image, in that the heights of the peaks can be visually compared. For example, landscape data may be generated for each of one or more types of biological object, such as lymphocytes and tumor cells. Thus, a peak in the lymphocyte landscape data may indicate a high number of lymphocytes in the region of the digital pathology image corresponding to the location of the peak, and a peak in the tumor cell landscape data may indicate a high number of tumor cells in the region of the digital pathology image corresponding to the location of the peak. Observing a peak for a first type of biological object compared to a peak for a second type of biological object may indicate a relationship between and / or the representation of the biological object types. For example, observing a peak in the tumor cell landscape in an area corresponding to an area with a lymphocyte peak may indicate lymphocytes interspersed with tumor cells. For example, peak 425a in landscape representation 420a may correspond to peak 425b in landscape representation 420b, and peak 430a may correspond to peak 430b. The peaks in landscape representation 420a and landscape representation 420b are generally in the same location, thus indicating scattering between types of biological objects. A comparison of the peaks shows less scattering at the locations of peaks 430a and 430b compared to the scattering at the locations of peaks 425a and 425b. In some cases, the digital pathology locations corresponding to the locations of peaks 430a and 430b may be of interest and may generate a prompt to collect more digital pathology image data or additional biological samples corresponding to that image location.
[0127] Ripley's K-function can be used as an estimator to detect deviations from spatial uniformity in a set of points (e.g., points corresponding to point-representative image locations of a biological object depiction) and can be used to assess the degree of spatial clustering or dispersion at many distance scales. The K-function (or more specifically, a sample-based estimate thereof) can be defined as follows: TIFF2026031924000002.tif16170(in the formula, d ij denotes the pairwise Euclidean distance between the i-th and j-th biological object representations out of a total of n biological object representations, r is the search radius, λ is the average density of the biological object representations (e.g., n / A, where A is the area of the tissue that encompasses all biological object representations), and I(·) is the distance between d ij is an indicator function that has 1 if ≦r, and w ij is an edge correction function to avoid biased estimation due to edge effects.
[0128] To design an efficient machine learning scheme, the entire K function can be summarized by formulating the following metric: 1. Area under the curve: distance between biological objects r, r max A clinically meaningful maximum value of r is identified, and 0 ≤ r ≤ r max The area between the observed K function and the theoretical value (e.g., under the null hypothesis that biological objects of the same or different types are spatially independent) can be calculated. 2.r=r max Point estimate of the difference between the observed and theoretical values of Ripley's K function at The above features may be derived separately for a first type of biological object and a second type of biological object (e.g., tumor cells and lymphocytes). Additionally, a cross-type Ripley's K function may be derived similarly. The Ripley's K function may be used to estimate and output the degree of spatial clustering or dispersion of the biological objects, thereby understanding this clustering among depictions of the biological objects (e.g., indicating the interpenetration or separation of the first type of biological object and the second type of biological object).
[0129] To identify nearest neighbor metrics, distances between the locations of various pairs of detected biological object representations may be determined. Each distance may be calculated for each pair of biological object representations of different types (e.g., between each tumor cell / lymphocyte pair). For a given biological object representation (e.g., a representation of an individual lymphocyte), a subset of nearest neighbor object representations may be defined as those identified as being of a given type and depicted as closest to the given biological object representation. For example, for a given lymphocyte, the nearest neighbor subset may identify n tumor cells depicted closest to the given lymphocyte compared to other tumor cells depicted in the image, where n may be a programmable, user-defined, or machine-learned value. For each subset, a centroid of the biological object representation locations of the subset may be calculated. A nearest neighbor distance metric between the centroid and the location of the given biological object representation may be determined therefrom.
[0130] 5A and 5B show two exemplary nearest neighbor subsets. The locations of exemplary biological object representations are represented by open circle data points in each of FIGS. 5A and 5B. For each biological object representation (e.g., lymphocyte), one or more nearest neighbor biological object representations of a second type (e.g., a predetermined number of nearest tumor biological object representations) may be identified. In the illustrated example, five other nearest neighbor biological object representations were identified. These nearest neighbor locations are represented by solid data points in FIGS. 5A and 5B. For the nearest neighbor locations, nearest neighbor centroids may be calculated. A midpoint may be calculated, for example, as the mean, median, weighted average, center of mass, etc., for the nearest neighbor locations. In the illustrated example, the centroid locations are represented by the end positions of the lines extending from the open circles. The nearest neighbor distance metric between the exemplary biological object locations and centroids is represented by the lines extending from the open circles in FIGS. 5A-5B.
[0131] Thus, for a given biological object, a nearest neighbor distance metric may be calculated for the nearest subset of biological objects of a second type. The distance metric may be used to classify the biological objects. As an example, if a first biological object is a lymphocyte and the nearest neighboring biological object is a tumor cell, the classification may be either an adjacent tumor lymphocyte or an intratumoral lymphocyte. The classification may be based on a learned or rule-based evaluation of nearest neighbor distance. For example, a lymphocyte may be classified as an adjacent tumor lymphocyte if the distance metric exceeds a threshold, and as an intratumoral lymphocyte if the distance metric does not exceed the threshold. The threshold may be fixed or defined based on a distance metric associated with one or more digital pathology images. In certain embodiments, the threshold may be calculated by fitting a two-component Gaussian mixture model to the distance metrics associated with all biological objects depicted in the digital pathology images. Figure 5C shows an exemplary characterization of biological objects using this discriminant analysis, depending on the process context (e.g., the identity of the biological object depiction, the number of biological object depictions, the identity of the type of biological object depiction, the number of types of biological object depictions, the absolute and relative values of the nearest neighbor distance, etc.). In the example shown in Figure 5C, black dots represent tumor cell depictions, blue dots represent lymphocytes classified as intratumoral lymphocytes, and green dots represent lymphocytes classified as tumor-adjacent lymphocytes.
[0132] The cross-type pair correlation function (PCF-cross) is another statistical measure of spatial dependence between points (e.g., points corresponding to point-representative image locations of biological object representations) in a spatial point process. In certain embodiments, the PCF-cross function may quantify how a biological object representation of a first type (e.g., lymphocytes) is surrounded by a biological object representation of a second type (e.g., tumor cells). The PCF-cross may be expressed as: TIFF2026031924000003.tif16170(λ, ω ij and d ij is similarly defined as a K function of replay, k h (·) is a smoothing kernel with smoothing bandwidth h>0)
[0133] The overall PCF cross can be summarized by formulating the following metric: 1. Area under the curve: distance r from biological object to biological object, r max A clinically meaningful maximum value of r can be selected, and the area between the observed PCF cross and the theoretical (e.g., under the null hypothesis that biological objects of the same or different types are spatially independent) PCF cross for 0≦r≦rmax was calculated. 2.r=r max Point estimate of the difference between the observed and theoretical PCF crossovers in .
[0134] The mark correlation function (MCF) facilitates determining whether the locations of biological object representations are more or less similar than expected with respect to the locations of nearby biological object representations (e.g., of different types), or whether their locations are independent (e.g., random) from the second type of biological object representation. In other words, whether the locations and presence of the second type of biological object representation affect the locations and presence of the first type of biological object representation. The mark correlation function may be defined as follows: TIFF2026031924000004.tif18170(where E(s i , s j ) is the distance r, M(s i ), M(s j ) Digital pathology image position s i and s j (In the denominator, M and M' are biological object types drawn randomly and independently from their marginal distributions, and I(m1;m2) is defined as 1 if m1 == m2.
[0135] We summarized the overall MCF by formulating the following metrics: 1. Area under the curve: distance between biological objects r, r max Select the clinically meaningful maximum value of , and set 0 ≤ r ≤ r maxThe area between the observed MCF and the theoretical (e.g., under the null hypothesis that biological objects of the same or different types are spatially independent) MCF was calculated. 2.r=r max Point estimate of the difference between observed and theoretical MCF in .
[0136] Further evaluation of biological object depictions can be based on a comparison of the prevalence of one or more types of biological object depictions. For example, features can be derived from a comparison of the amount of a first type of biological object depiction and a second type of biological object depiction. Furthermore, features can be strengthened by comparing biological object depictions with a specific classification (e.g., a first type or a second type).
[0137] For example, classification of lymphocyte depictions based on statistical analysis of tumor spatial heterogeneity may be characterized by the intratumoral lymphocyte ratio (ITLR), which may characterize lymphocyte depiction location relative to tumor cell density. In some embodiments, assessment may be guided by the use of digital pathology image annotations, such as annotation of regions of interest (e.g., tumor regions). Within each of these regions, each lymphocyte depiction may be characterized as an adjacent tumor lymphocyte or an intratumoral lymphocyte based on a Euclidean distance measure (described herein). The n nearest tumor cells may be identified for each lymphocyte depiction (e.g., using a nearest neighbor technique, such as the technique described in Section VI.A.3). Here, n is a definable parameter related to the number of neighbors used. Second, the centroid coordinates of the convex hull region formed by the n nearest tumor cell depictions may be derived. The distance from each lymphocyte depiction to the nearest tumor cell depiction and the centroid of the convex hull may then be calculated, and a two-component Gaussian mixture model may be fitted to further distinguish lymphocytes into adjacent tumor lymphocytes or intratumoral lymphocytes. If lymphocytes are infiltrating the tumor core region, the distance to the center of mass should be small. In contrast, if lymphocytes are still migrating to the tumor core region, the distance is likely to be larger. The characteristics of ITLR were defined as follows: JPEG2026031924000005.jpg14170 (in the formula, N腫瘍内リンパ球 represents the total number of intratumoral lymphocytes, and N 腫瘍細胞 (where σ represents the total number of tumor cells.) Although described in the context of a specific classification of a particular type of biological object, BOR can be extended using similar principles to other biological object descriptions that have their own context-dependent properties.
[0138] The G-cross function calculates a probability distribution of the distance from a first type of biological object representation to the nearest second type of biological object representation within any given distance. Specifically, the G-cross function can be viewed as a spatial distance distribution metric that represents the probability of finding at least one biological object representation (e.g., of a specified type) within a circle of radius r centered at a given point (e.g., a point location representation of a biological object representation in a digital pathology image). These probability distributions can be applied to quantify the relative proximity of any two types of biological object representation. Thus, for example, the G-cross function can be a quantitative proxy for penetration determination. Mathematically, the G-cross function is expressed as follows: TIFF2026031924000006.tif16170
[0139] (In the ceremony TIFF2026031924000007.tif9170, j represents the index of the first kind of biological object depiction, and I(·) represents the d i is an indicator function that has 1 if ≦r, and n lym is the total number of biological objects.)
[0140] Similarly, the entire G-cross function can be summarized by formulating the following metric: 1. Area under the curve: distance between biological objects r, r max Select the clinically meaningful maximum value of , and set 0 ≤ r ≤ r max The area between the observed G-cross function and the theoretical (e.g., under the null hypothesis that biological objects of the same or different types are spatially independent) G-cross function was calculated. 2.r=r max Point estimate of the difference between the observed and theoretical G-cross functions at .
[0141] 6A-6D show exemplary distance- and intensity-based metrics characterizing the spatial arrangement of biological object depictions in exemplary digital pathology images, according to some embodiments. For each of four types of spatial feature metrics derived based on the digital pathology images, statistical values are plotted across a range of r values. FIG. 6A shows the observed G-cross function calculated from the sample (thin dashed line) and the theoretical G-cross function (thick dashed line) under the null hypothesis that the first and second types of biological objects are spatially independent. The G-cross function may be calculated as described herein. FIG. 6B shows the difference (solid line) between the K-function calculated for the biological object depictions of the first species pair and the K-function calculated for the biological object depictions of the second species. The K-function was calculated as described herein. Figure 6C shows a crossed pair correlation function (dotted line) calculated under the null hypothesis that the first type of biological object and the second type of biological object are spatially independent, or a crossed pair correlation function (solid line) calculated by comparing the positions of the first type of illustrated biological object with the second type of illustrated biological object. The pair correlation was calculated as described herein. Figure 6D shows a mark correlation function (dotted line) calculated under the null hypothesis that the first type of biological object and the second type of biological object are spatially independent, or a mark correlation function (solid line) calculated by comparing the positions of the first type of illustrated biological object with the second type of illustrated biological object. The mark correlation was calculated as described herein.
[0142] The plots in Figures 6A-6D show that, in this example, the depictions of the first and second types of biological objects are spatially correlated based on objective measures. Further quantitative features may be derived based on the algorithms disclosed herein.
[0143] 7 illustrates the application of the region analysis framework 230. In particular, the region analysis framework 230 was used to process a digital pathology image 405 of a stained sample section. As described above in connection with the spatial point process analysis framework, depictions of specific types of biological objects (e.g., lymphocytes and tumor cells) were detected. The region analysis framework 230 further generates biological object data, an example of which is shown in Table 410.
[0144] A spatial grid with a defined number of columns and rows may be used to divide the digital pathology image 405 into regions. As an example, as shown in FIG. 7, the spatial grid was used to divide the digital pathology image 405 into 22 columns and 19 rows. The spatial grid includes 418 regions. Each biological object representation may be assigned to a region. In certain embodiments, a region may be a region that includes the midpoint or other representation of the biological object representation. For each type of biological object and each grid region, several biological object representations of the type of biological object assigned to the region may be identified. For each type of biological object, a collection of region-specific biological object counts may be defined as grid data for that particular type of biological object. FIG. 7 shows certain embodiments of grid data 715a for a first type of biological object representation and grid data 715b for a second type of biological object representation, each overlaid on a representation of the digital pathology image 405 of a stained section. The grid data may be defined to include, for each region in the grid, a prevalence value defined as the region's equal count divided by the total count across all regions. Thus, regions where no biological matter of a given type is present will have a prevalence value of 0, and regions where at least one biological matter of a given type is present will have a positive non-zero prevalence value.
[0145] The same amount of biological objects (e.g., lymphocytes) in two different contexts (e.g., tumors) does not imply a characteristic or degree of characteristic (e.g., the same immune infiltration). Instead, how a representation of a first type of biological object is distributed relative to a representation of a second type of biological object may indicate a functional state in some cases. Therefore, characterizing the proximity of biological object representations of the same type and different types may reflect more information. The Morisita-Horn index is an ecological measure of similarity (e.g., overlap) in biological systems or ecosystems. In certain embodiments, the Morisita-Horn index (MH), which characterizes the bivariate relationship between two populations (e.g., of two types) of biological object representations, may be defined as follows: TIFF2026031924000008.tif16170 (in the formula, (TIFF2026031924000009.tif8170 show the prevalence of a first type of biological object depiction and a second type of biological object depiction, respectively, in square grid i.) In FIG. 7, grid data 715a shows an exemplary distribution of first type of biological object depictions across grid points. TIFF2026031924000010.tif9170, and grid data 715b is an exemplary representation of a second type of biological object across grid points. Showing TIFF2026031924000011.tif9170.
[0146] The Morisita-Horn index is defined to be 0 when an individual grid region does not contain a representation of both types of biological matter (indicating that the distributions of the different types of biological matter are spatially separated). For example, the index is 0 when considering the exemplary spatially separate distributions shown in exemplary first grid data 720a. The Morisita-Horn index is defined to be 1 when the distribution of a first type of biological matter across a grid region matches (or is a scaled version of) the distribution of a second type of biological matter across the grid region. For example, the index is close to 1 when considering the exemplary highly co-localized distributions shown in exemplary second grid data 720b.
[0147] 7, the Morisita-Horn index calculated using grid data 715a and grid data 715b was 0.47. A high index value indicates a high degree of co-localization of the depictions of the first and second types of biological objects.
[0148] The Jaccard index (J) and the Sorensen index (L) are similar and closely related to each other. In certain embodiments, the indices may be defined as follows: TIFF2026031924000012.tif31170(in formula TIFF2026031924000013.tif9170 represent the prevalence of the first type of biological object depiction and the second type of biological object depiction in square grid i, respectively, and min(a, b) returns the minimum value between a and b. In certain embodiments, another metric that can characterize the spatial distribution of biological object representations is the Moran's Index, which is a measure of spatial autocorrelation. Generally, the Moran's Index statistic is the correlation coefficient for the relationship between a first variable and a second variable in adjacent spatial units.
[0149] In certain embodiments, the first variable may be defined as the prevalence of a first type of biological object depiction, and the second variable may be defined as the prevalence of a second type of biological object depiction, thereby quantifying the degree to which depictions of the two types of biological objects are interspersed in the digital pathology image. In some embodiments, the Moran Index I may be defined as: TIFF2026031924000014.tif18170(in the formula, x i , y j represents the standardized prevalence of biological object representations of a first type (e.g., tumor cells) in area unit i and the standardized prevalence of biological object representations of a second type (e.g., lymphocytes) in area unit j.)ω ijwhere is the binary weight of areal units i and j, where the weight is 1 if the two units are adjacent and 0 otherwise, and a first-order scheme may be used to define the neighborhood structure. Moran's I may be derived separately for biological object delineations of different types of biological objects.
[0150] As shown in Figure 8, the Moran index is defined to be equal to -1 when biological object representations are perfectly dispersed across the lattice (thus having negative spatial autocorrelation; "colocalization scenario" 820a), and to be equal to 1 when biological object representations are densely clustered (thus having positive autocorrelation; "segregation scenario" 820b).
[0151] The Moran's Index is defined as 0 when the distribution of objects matches a random distribution. Therefore, an area representation of a particular type of biological object representation makes it easy to generate a grid that supports the calculation of the Moran's Index for each type of biological object. The Moran's Index calculated using the lattice data 715a was 0.50. The Moran's Index calculated using the lymphocyte lattice data 715b was 0.22. The difference between the Moran's Index calculated for each of the two biological object representations can provide an indication of collocation (e.g., a difference close to 0 indicates collocation).
[0152] Geary's C, also known as Geary's continuity ratio, is a measure of spatial autocorrelation, or an attempt to determine whether adjacent observations of the same phenomenon are correlated. Geary's C is inversely related to Moran's I, but is not identical. Moran's I is a measure of global spatial autocorrelation, while Geary's C is more sensitive to local spatial autocorrelation. TIFF2026031924000015.tif17170(in the formula, z i is a square lattice i, ω i、j (These terms represent the prevalence of either the first or second type of biological object depiction in the United States, as defined above.)
[0153] In certain embodiments, the grid data 715a and grid data 715b may be further processed to generate hotspot data 915a corresponding to detected depictions of a first type of biological object and hotspot data 915b corresponding to detected depictions of a second type of biological object, respectively. In FIG. 9 , the hotspot data 915a and hotspot data 915b indicate regions determined to be hotspots for each type of detected biological object depiction. Regions detected as hotspots are indicated by red symbols, while regions determined not to be hotspots are indicated by black symbols. The hotspot data 915a, 915b are defined for each region associated with a non-zero object count. The hotspot data 915a, 915b may also include a binary value indicating whether a given region has been identified as a hotspot. In addition to the hotspot data and analysis, coldspot data and analysis may be performed.
[0154] For biological object representations, hotspot data 915a, 915b may be generated for each type of biological object by determining the Getis-Ord local statistic for each region associated with a non-zero object count. Getis-Ord hotspot / coldspot analysis may be used to identify statistically significant hotspots / coldspots of tumor cells or lymphocytes. Here, a hotspot is an area unit with a statistically significantly higher value of the prevalence of a biological object representation compared to neighboring area units, and a cold spot is an area unit with a statistically significantly lower value of the prevalence of a biological object representation compared to neighboring area units. The values and determinations that make a region a hotspot / coldspot compared to neighboring areas may be selected according to user preference, and in certain embodiments, may be selected according to a rule-based approach or a trained model. For example, the number and / or type of biological objects detected, the absolute number of representations, and other factors may be considered. The Getis-Ord local statistic is a z-score, which may be defined for a square grid i as follows: TIFF2026031924000016.tif22170, where i represents an individual region (a particular row-column combination) in the lattice, and n is the number of row and column combinations (i.e., the number of regions) in the lattice; TIFF2026031924000017.tif7170 is the spatial weight between i and j, and z j is the prevalence of a given type of biological object depiction in the domain, TIFF2026031924000018.tif6170 is the average object prevalence of a given type across the region.) TIFF2026031924000019.tif17170
[0155] In certain embodiments, the Getis-Ord local statistics may be converted to binary values by determining whether each statistic exceeds a threshold. For example, the threshold may be set to 0.16. The threshold may be selected according to user preference, and in certain embodiments, may be set according to rules based on a machine learning approach.
[0156] In certain embodiments, a logical AND function may be used to identify regions identified as hotspots for two or more types of biological object representations. For example, colocalization hotspot data 920 shows regions identified as hotspots for two types of biological object representations (shown with red symbols). A high ratio of the number of regions identified as colocalization hotspots to the number of hotspot regions identified for a given type of object (e.g., for tumor cell objects) may indicate that a given type of biological object representation shares spatial characteristics with other types of objects. On the other hand, a low ratio at or near zero may be consistent with spatial separation of different types of biological objects.
[0157] Geostatistics is a collection of mathematical and statistical methods originally developed to predict the probability distribution of spatially stochastic processes in the mining industry. Geostatistics has been widely applied in diverse fields, including petroleum geology, geosciences, agriculture, soil science, and environmental exposure assessment. In the field of geostatistics, variograms can be used to represent the spatial continuity of data. To generate features from variogram fitting, an empirical variogram can first be calculated as a discrete function using a measure of variability between pairs of points (e.g., representative locations in a biological object depiction) separated by various distances. Second, the empirical variogram can be estimated and fitted to a theoretical variogram. In certain embodiments, the Matern function can be used as a theoretical variogram model. Consider JPEG2026031924000020.jpg8170, where Z(s) is the prevalence of tumor cells or lymphocytes at location s, and D represents the set of sample points s, s, ...s. The empirical variogram may be calculated as follows: JPEG2026031924000021.jpg17170
[0158] In the example of Figure 10, an empirical variogram was generated based on a depiction of biological objects detected in an H&E stained image 405 (shown in Figure 10 as points on a theoretical variogram plot). A theoretical variogram 1015 was then generated by fitting a Matern function to the empirical variogram.
[0159] In the above calculation, the sum is calculated only for N(h) pairs of observations (e.g., pairs of biological object depictions) separated by a Euclidean distance h. Parameters from the Matern function can be used as features from this method. Features can be obtained separately from variogram fitting of detected depictions of a first type of biological object (e.g., tumor cells) and detected depictions of a second type of biological object (e.g., lymphocytes). Alternatively, index variogram fitting can be performed when combining detected biological object depictions by type.
[0160] The variograms and point locations of the detected biological object estimates may then be used to generate, for each region (e.g., pixel) of the digital pathology image 405, a probability that a particular type of biological object is depicted in that region. The Kriging map 1020 shown in Figure 10 indicates, for each of a plurality of regions in the digital pathology image 405, the probability that a particular type of biological object (e.g., tumor cell) is depicted in that region.
[0161] In certain embodiments, a regression machine learning model can be trained to process digital pathology images, for example, of biopsy sections from a subject, to predict an assessment of the subject's condition from the digital pathology images. As an example, a regression machine learning model can be trained to predict whether a cancer exhibits microsatellite stability in tumor DNA (versus microsatellite instability in tumor DNA) based on digital pathology images of biopsy sections from a subject diagnosed with colorectal cancer. Microsatellite instability can be associated with a relatively large number of mutations within microsatellites.
[0162] Biopsies may be collected from each of a plurality of subjects with a disease, in this example, colorectal cancer. The samples may be fixed, embedded, sliced, stained, and imaged in accordance with the subject matter disclosed herein. Representations of designated types of biological objects, such as tumor cells and lymphocytes, may be detected, for example, using biological object detector subsystem 145. In certain embodiments, biological object detector subsystem 145 may recognize and identify biological object representations using a trained deep convolutional neural network.
[0163] For each of a plurality of subjects, a label may be generated to indicate whether the condition (e.g., cancer) exhibited a specified characteristic (e.g., microsatellite stability vs. microsatellite instability). Ground truth labels may be generated based on pathologist evaluation and assay-based test results. For each subject, an input vector may be defined to include a set of spatial distribution metrics. The set of spatial distribution metrics may include a selection of metrics described herein. As an example, metrics included in the input vector may include: - the area between the observed and theoretical K functions for distances between biological objects ranging from 0 to the maximum observed distance; -Point estimate of the difference between the observed Ripley's K function and the theoretical Ripley's K function at the maximum biological object distance; Area under the curve of the G-cross function for distances between biological objects ranging from -0 to the maximum observation distance; -point estimate of the difference between the observed G-cross function and the theoretical G-cross function at the maximum biological object-to-object distance; Area under the curve of the pair correlation function (cross type) for distances between biological objects ranging from -0 to the maximum observation distance; - point estimate of the difference between the observed and theoretical pair correlation functions at the maximum biological object distance (cross type); Area under the curve of the mark correlation function (cross type) for distances between biological objects ranging from -0 to the maximum observation distance; - point estimate of the difference between the observed and theoretical mark correlation functions at the maximum biological object distance (cross type); -Percentage of intratumoral lymphocytes; -Morisita-Horn index; -Jacquard index; -Sorensen index; -Moran's index; -Geary's C; - the ratio of non-local spots (e.g., hot spots, cold spots, insignificant spots) for a type of biological object description over the number of spots (e.g., hot spots, cold spots, insignificant spots) for a first type of biological object description, using spots (e.g., hot spots, cold spots, insignificant spots) defined using Getis-Ord local statistics; and - Features obtained by variogram fitting of two types of biological objects (e.g., tumor cells and lymphocytes) representations.
[0164] The selected metric corresponds to multiple frameworks (point process analysis framework, area process analysis framework, and geostatistical framework). In certain embodiments, for each subject, a label may be defined indicating whether the displayed feature (e.g., microsatellite stability) is observed. The L1 regularized logistic regression model may be trained and tested using paired input data and labels, with repeated 5-fold cross-validation by lasso. Specifically, for each of the five data folders, the model may be trained on the remaining four folders and tested on the remaining folder to calculate the area under the ROC.
[0165] Figure 11 shows an exemplary median receiver operating curve (ROC) generated using 5-fold cross-validation. In the example described, the median area under the ROC generated using the validation set was 0.931. The 95% confidence interval was (0.88, 0.96). The variables from the input dataset most frequently selected by the L1 regularized logistic regression model can be identified to indicate which metrics were deemed most predictive of a particular feature of the subject's condition. For example, the most frequently selected metrics could be the area under the curve of the pair correlation function and hotspot ratio calculated using Getis-Ord local statistics, indicating that these metrics are most predictive of microsatellite instability. Processing digital pathology images can serve as a reliable alternative to certain laborious and expensive tests. For example, in the example discussed herein, the digital pathology image processing system can demonstrate that processing can mirror or exceed DNA analysis in determining whether a given subject's tumor exhibits microsatellite instability. Thus, use of image-based techniques according to the presently disclosed subject matter may eliminate the need to collect additional biopsy samples from the subject to collect DNA, further saving time and money in performing DNA analysis.
[0166] In certain embodiments, digital pathology images of stained biopsy sections are accessed for each of a first subject and a second subject. A first type of biological object representation and a second type of biological object (e.g., lymphocytes and tumor cells) representation can be detected in each image according to the techniques described herein. An input vector as described herein can be generated for each subject. The input vectors can be separately processed by a trained logistic regression model as described herein.
[0167] The model outputs a first label in response to processing an input vector associated with the first subject, which may correspond, for example, to a prediction that the first subject's cancer exhibits microsatellite instability.
[0168] The model outputs a second label in response to processing the input vector associated with the second subject, which may correspond, for example, to a prediction that the second subject's cancer does not exhibit microsatellite stability.
[0169] Each of the first label and the second label can be processed (separately) according to a treatment recommendation rule. The rule can be configured to recommend a specific treatment, such as an immunotherapy (or immune checkpoint therapy) treatment, upon detecting a specific feature of the subject's condition, such as microsatellite instability, or to not recommend the use of another treatment, such as an immunotherapy (or immune checkpoint therapy) treatment, upon detecting a specific feature of the subject's condition. The results from rule processing can indicate, for example, that an immunotherapy treatment is recommended for a first subject but not for a second subject.
[0170] In certain embodiments, digital pathology images can depict the tumor microenvironment, including the spatial structure of tissue components and their microenvironmental interactions, which can have profound effects on tissue formation, homeostasis, regenerative processes, immune responses, and the like.
[0171] Non-small cell lung cancer (NSCLC) is a major global health problem and the leading cause of cancer-related deaths worldwide. Despite the wide range of available treatment options, chemotherapy remains the mainstay of treatment for patients with metastatic (EGFR- and ALK-negative / unknown) NSCLC. However, immune checkpoint inhibitors are revolutionizing the treatment algorithm for this subpopulation.
[0172] Digital pathology images can be used to calculate spatial statistics (e.g., spatial distribution metrics) to determine how well they predict overall survival for various treatments. Clinical study groups can be established to test the effectiveness of various treatments. An exemplary clinical trial was conducted to evaluate the safety and efficacy of atezolizumab (an engineered anti-programmed death-ligand 1 [PD-L1] antibody) in combination with carboplatin and paclitaxel (e.g., "group ACP") with or without bevacizumab (e.g., "group ABCP") compared with carboplatin, paclitaxel, and bevacizumab (e.g., "group CPB") in chemotherapy-naive participants with stage IV non-squamous NSCLC. Participants were randomized in a 1:1:1 ratio to group ACP, group ACPB, or group CPB (control group).
[0173] Tissue samples were collected at baseline. Digital pathology (e.g., H&E pathology) images of the baseline tissue samples were captured for each subject in each treatment group. H&E-stained slides of the tissue samples were scanned and digitized to generate the types of digital pathology images described herein. Regions associated with one or more depictions of biological objects on the digital pathology images (also referred to as whole slide images or "WSIs") were annotated. Depictions of specific types of biological objects, including tumor cells, immune cells, and other stromal cells, were detected. For example, according to the subject matter disclosed herein, location coordinates for each depiction of each type of biological object were generated. In one example, while investigating the efficacy of different test groups, the focus may be on, for example, lymphocytes and tumor cells to investigate immune infiltrates, tumor resource distribution, and cell-cell interactions.
[0174] For each image, a wide variety of spatial features may be derived based on the state of the detected biological objects and / or their respective relative locations based on the spatial statistics (e.g., spatial distributional metrics) algorithms described herein, including, for example, spatial point processing methods (e.g., Ripley's K function features, G function features, pair correlation function features, mark correlation function features, and intratumoral lymphocyte ratio), spatial grid processing methods (e.g., Morisita-Horn index, Jaccard index, Sorensen index, Moran's I, Geary's C, and Getis-OrdHotspot), and geostatistical processing methods (ordinary Kriging mechanism, indicator Kriging mechanism).
[0175] Additionally, the objective of the trial, the outcome variable, for example, overall survival of the subject, may be identified.
[0176] Generally, the analyses performed in this example were conducted to determine whether the difference in overall survival between the ACP and BCP cohorts would be more pronounced if only a portion of each cohort were considered, with that portion selected as individuals predicted to have longer survival compared to other subjects in the cohort. Predictions may be based, for example, on one or more of the spatial distribution metrics described herein, generated on digital pathology images of samples taken from subjects. In certain embodiments, the first analysis involved comparing the intention-to-treat populations of ACP versus BCP with overall survival. The second analysis involved using a model-based predictive enrichment strategy to investigate the association between derived spatial features and overall survival (OS). Predictive enrichment of clinical trials, including NSCLC clinical studies, involves identifying responder subpopulations within the overall patient population Ω0 that have a greater-than-average response to treatment, as measured, for example, by odds ratio (OR), relative risk (RR), or hazard ratio (HR). Focusing on this subpopulation has the advantage of increasing the efficiency or feasibility of the study and enhancing the benefit-risk relationship for subjects in the subpopulation compared to the overall population. One enrichment strategy is an open-label, single-arm study followed by randomization. In this design, the clinical trial treatment is administered to all subjects, and responders identified by pre-specified criteria (e.g., study endpoints or biomarkers) are randomized to a placebo-controlled trial.
[0177] Model-based methods can be used to address, for example, the problem of predictive enrichment. In particular, when a clinical trial has already been conducted, an enrichment model can be developed retrospectively. To retrospectively develop an enrichment model, data can be split into training, validation, and test sets by 60:20:20 in each group (e.g., according to the subject matter disclosed herein). Training set by treatment group can be used, for example, to simulate an open-label pre-randomization stage in an empirical design. A Cox model or objective response model that inputs spatial statistical features can be fitted with L1 or L2 regularization on the training set of the treatment group, e.g., ACP. The predicted risk score or predicted response probability from the fitted Cox model can be calculated as a response score. JPEG2026031924000022.jpg5170, and the responder criteria may be specified in the form of a subset condition. TIFF2026031924000023.tif11170 (in the formula, S q denotes the q quantile of the response score, and x represents the subject-level covariate characterized by the feature vector.) A validation set combining the treatment and control groups can be used to simulate the subject group recruited before randomization. To implement subsetting, quantiles with the same q can be calculated for the treatment and control groups in the validation set, respectively, and subsets are taken for the treatment and subjects in the validation set, respectively, using the formula above. In this example, JPEG2026031924000024.jpg9170It can be estimated by assessing q towards the most significant difference between treatment and control using either a log-rank test for survival data or a permutation test for objective response data, both subsets relative to the responder subgroup in the validation set. JPEG2026031924000025.jpg8170It can also be estimated using a pre-specified response threshold q. The enrichment conditions using JPEG2026031924000026.jpg7170 are JPEG2026031924000027.jpg9170, which can then be evaluated in the same way as the test set for hazard ratios or odds ratios.
[0178] In embodiments with limited sample size, nested Monte Carlo cross-validation (nMCCV) may be used to evaluate model performance. The same enrichment procedure is repeated B times by randomly splitting the training, validation, and test sets in equal proportions to determine the score function and threshold. JPEG2026031924000028.jpg10170. For the i subject, the ensembled responder status may be estimated by averaging the responder group membership for i among the iterations in which i is randomized to the test set and thresholding at 0.5. Hazard ratios or odds ratios, along with 95% confidence intervals and p-values, may be calculated for the aggregated test subjects.
[0179] The overall workflow of the predictive analysis is summarized in the flowchart in Figure 12. More specifically, to assign a label to each subject in the test cohort, a nested Monte Carlo cross-validation (nMCCV) modeling strategy was used to overcome overfitting.
[0180] Specifically, for each subject, in block 1205, the dataset may be split into training, validation, and test data portions in a 60:20:20 ratio. In block 1210, a 10-fold cross-validated Ridge-Cox (L2 regularized Cox model) may be performed using the training set to generate 10 models (having the same model architecture). A specific model from among the 10 generated models may be selected and stored based on the 10-fold training data. In block 1215, the specific model may be applied to the validation set to adjust specified variables. For example, the variable may identify a risk score threshold. Then, in block 1220, the threshold and the specific model may be applied to an independent test set to generate a vote for the subject that predicts whether the subject will be stratified into a longer or shorter survival group. The data splitting, training, cutoff identification, and vote generation (blocks 1205-1220) may be repeated N times (e.g., N = 1000). Thereafter, in block 1225, the subject is assigned to either the longer survival group or the shorter survival group based on the table. For example, the step in block 1225 may include assigning the subject to the longer survival group or the shorter survival group by determining which group is associated with the majority of the table. Thereafter, in block 1230, a survival analysis of the subjects in the longer / shorter survival groups may be performed. It will be appreciated that similar procedures for applying a variety of labels to data based on the desired outcome may be applied to any suitable clinical evaluation or qualification test.
[0181] In contrast to the primary finding when comparing the intent-to-treat populations for ACP versus BCP, with an overall survival hazard ratio (HR) of 0.85 (95% CI 0.71-1.03), the proposed approach yielded a clear separation between the ACP and BCP cohorts, with an overall survival hazard ratio of 0.64 (95% CI 0.45-0.91; Figure 13). Note that an overall survival hazard ratio of 1.0 indicates statistically similar survival between the cohorts. Thus, in this described example, the lower hazard ratio obtained using the second analytical approach (during which statistics were calculated only for the portion of the cohort predicted to have longer survival based on spatial statistics and / or spatial distribution metrics) suggests that the second analysis was able to better identify subjects who would benefit from treatment (ACP treatment). Therefore, the use of spatial distribution metrics represents an improvement over previous approaches.
[0182] The comprehensive model based on spatial statistics and spatial distribution metrics used in this example analysis enhanced an analytical pipeline that generates systems-level knowledge, in this case of the spatial heterogeneity of the tumor microenvironment, by modeling histopathology images as spatial data. The results demonstrate that spatial statistics-based methods can stratify subjects who benefit from atezolizumab treatment compared with standard of care. This effect is not limited to the specific treatment evaluation discussed in this example. Using spatial statistics to characterize histopathology and other digital pathology images may be useful in clinical settings to predict treatment outcomes and thus inform treatment selection.
[0183] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0184] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.
[0185] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.
[0186] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
Claims
1. 1. A computer-implemented method by a digital pathology image processing system, comprising: accessing a digital pathology image representing a section of a biological sample from a subject; Within the digital pathology image, a first set of biological object representations, each of the first set of biological object representations depicting a first biological object of a first type of biological object; detecting a second set of biological object representations, each of the second set of biological object representations depicting a second biological object of a second type of biological object; generating a spatial distribution metric using the first set of biological object representations and the second set of biological object representations, the spatial distribution metric characterizing a position of the first set of biological object representations relative to the second set of biological object representations; using the spatial distribution metric to generate a subject-level outcome corresponding to a predicted biological state of the subject or a potential treatment for the subject; generating a display including the subject-level results; 11. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the first type of biological matter comprises a first type of cell and the second type of biological matter comprises a second type of cell.
3. The computer-implemented method of claim 2 , wherein the first type of biological matter comprises lymphocytes and the second type of biological matter comprises tumor cells.
4. 2. The computer-implemented method of claim 1, wherein the digital pathology image depicts a biological sample from the subject after treatment with one or more stains, each of which enhances the appearance of one or more of the first type of biological object or the second type of biological object.
5. generating the spatial distribution metric comprises: for each first biological object representation of the one or more first biological object representations, identifying a first point location within the digital pathology image corresponding to the first biological object representation; for each second biological object representation of the one or more second biological object representations, identifying a second point location within the digital pathology image corresponding to the second biological object representation; determining the spatial distribution metric based on the first point locations and the second point locations; The computer-implemented method of claim 1 , comprising:
6. The computer-implemented method of claim 5 , wherein the first point location in the digital pathology image indicates a location of the first biological object representation.
7. 7. The method of claim 6, wherein the first point location in the digital pathology image is selected by calculating a mean point location, a centroid point location, a median point location, or a weighted point location for the first biological object representation.
8. 6. The computer-implemented method of claim 5, wherein generating the spatial distribution metric further comprises calculating, for each of at least some first biological object representations of the one or more first biological object representations and for each of at least some second biological object representations of the one or more second biological object representations, a distance between the first point location corresponding to the first biological object representation and the second point location corresponding to the second biological object representation.
9. 9. The computer-implemented method of claim 8, wherein generating the spatial distribution metric further comprises identifying, for each of at least some first biological object representations of the one or more first biological object representations, one or more of the second biological object representations that are associated with a distance between the first biological object representation and the second biological object representation.
10. generating the spatial distribution metric comprises: defining a spatial grid configured to divide an area of the digital pathology image into a set of image regions; assigning each of the one or more first biological object representations to an image region of the set of image regions; assigning each of the one or more second biological object representations to an image region of the set of image regions; generating the spatial distribution metric based on the image region assignments; The computer-implemented method of claim 1 , comprising:
11. generating the spatial distribution metric comprises: determining a first set of one or more image regions of the set of image regions that have a higher probability of containing a first biological object representation than adjacent image regions; determining a second set of one or more image regions of the set of image regions that have a higher probability of containing a second biological object representation than adjacent image regions; determining the spatial distribution metric based on the first set of image regions and the second set of image regions; The computer-implemented method of claim 10 further comprising:
12. generating the spatial distribution metric comprises: determining a third set of one or more image regions of the set of image regions that have a higher probability of containing both the first biological object representation and the set biological object representation than adjacent image regions; determining the spatial distribution metric based on the third set of image regions; The computer-implemented method of claim 11 further comprising:
13. using the first spatial distribution metric to generate a subject-level result corresponding to a predicted biological state of the subject or a potential treatment for the subject, comparing the spatial distribution metric generated for the digital pathology image with a previous spatial distribution metric generated for a previous digital pathology image; outputting object-level results generated for the previous digital pathology image based on the comparison; The computer-implemented method of claim 1 , comprising:
14. generating said subject-level results, 2. The computer-implemented method of claim 1, comprising using a trained machine learning model to determine a diagnosis, prognosis, therapy recommendation, or treatment eligibility assessment for the subject based on processing the spatial distribution metrics and the first set of biological object depictions and the second set of biological object depictions.
15. The spatial distribution metric is: A metric defined based on K-nearest neighbor analysis, A metric defined based on Ripley's K function, Morisita-Horn index, Moran's index, Metrics defined based on correlation functions, Metrics defined based on hot spot / cold spot analysis, or Metrics defined based on kling-based analysis The computer-implemented method of claim 1 , comprising:
16. the spatial distribution metric is a first type of metric; the computer-implemented method further comprising using the first set of biological object representations and the second set of biological object representations to generate a second spatial distribution metric characterizing a position of the first set of biological object representations relative to the second set of biological object representations, the second spatial distribution metric being a second type of metric different from the first type of metric; The computer-implemented method of claim 1 , wherein the subject-level results are generated further using the second spatial distribution metric.
17. receiving user input data from a user device including an identifier of the subject or the digital pathology image, wherein the digital pathology image is accessed based on the received user input data; The computer-implemented method of claim 1 , wherein providing the subject-level results for display comprises providing the subject-level results to the user device.
18. 10. The computer-implemented method of claim 1, further comprising outputting a clinical assessment to a user device for the subject, the clinical assessment comprising a diagnosis, prognosis, therapy recommendation, or treatment eligibility assessment for the subject.
19. one or more data processors; a computer-readable non-transitory storage medium communicatively coupled to the one or more data processors and comprising instructions that, when executed by the one or more data processors, cause the one or more data processors to perform one or more of the following operations: accessing a digital pathology image representing a section of a biological sample from a subject; Within the digital pathology image, a first set of biological object representations, each of the first set of biological object representations depicting a first biological object of a first type of biological object; detecting a second set of biological object representations, each of the second set of biological object representations depicting a second biological object of a second type of biological object; generating a spatial distribution metric using the first set of biological object representations and the second set of biological object representations, the spatial distribution metric characterizing a position of the first set of biological object representations relative to the second set of biological object representations; using the first spatial distribution metric to generate a subject-level result corresponding to a predicted biological state of the subject or a potential treatment for the subject; generating a display including the subject-level results.
20. One or more computer-readable non-transitory storage media containing instructions that, when executed by one or more data processors, cause the one or more data processors to perform the following operations: accessing a digital pathology image representing a section of a biological sample from a subject; Within the digital pathology image, a first set of biological object representations, each of the first set of biological object representations depicting a first biological object of a first type of biological object; detecting a second set of biological object representations, each of the second set of biological object representations depicting a second biological object of a second type of biological object; generating a spatial distribution metric using the first set of biological object representations and the second set of biological object representations, the spatial distribution metric characterizing a position of the first set of biological object representations relative to the second set of biological object representations; using the first spatial distribution metric to generate a subject-level result corresponding to a predicted biological state of the subject or a potential treatment for the subject; generating a display including the subject-level results; and one or more computer-readable non-transitory storage media.