Pathology prediction based on spatial feature analysis
By generating spatial distribution metrics from digital pathology images, the method addresses the limitations of image-level analysis, providing accurate clinical assessments and treatment recommendations based on the spatial relationships of biological objects.
Patent Information
- Application Number
- JP2022571162
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-11
- Filing Date
- 2021-05-17
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2041-05-17
AI Technical Summary
Current image-level analysis in digital pathology simplifies spatial information, hindering the detection of detailed interactions between biological objects and their microenvironment, which are crucial for accurate diagnosis, prognosis, and treatment decisions.
A computer-implemented method processes digital pathology images to generate spatial distribution metrics by analyzing the relative positions of different biological object representations, using techniques like spatial point process analysis and machine learning, to predict biological states and treatment outcomes.
Enhances the accuracy of diagnosis, prognosis, and treatment decisions by objectively characterizing the spatial distribution of biological objects, enabling more precise clinical assessments and treatment recommendations.
Smart Images

Figure 0007803880000026 
Figure 0007803880000027 
Figure 0007803880000028
Abstract
Description
[Technical Field]
[0001] Priority This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 077,232, filed September 11, 2020, and U.S. Provisional Patent Application No. 63 / 026,545, filed May 18, 2020.
[0002] This application relates generally to image processing of digital pathology images to generate output that characterizes spatial information of particular types of objects within the image. More particularly, digital pathology images can be processed to generate metrics that characterize the spatial distribution and correlation in the depiction of one or more types of biological objects across all or a portion of the image. [Background technology]
[0003] Image analysis involves processing individual images to generate image-level results. For example, the results may be binary results corresponding to an assessment of whether the image contains a particular type of object. As another example, the results may include an image-level count of the number of a particular type of object detected in the image. In the context of digital pathology, the results may include a count of a particular type of cell detected in the image of the sample, a ratio of the count of one type of cell to another type of cell across the entire image, and / or the density of a particular type of cell.
[0004] This image-level approach can be useful because it can facilitate the storage of simple metadata, making it easier to understand how the results were generated. However, this image-level approach can remove detail from the image, which can hinder the detection of details of the depicted situation and / or environment. This simplification can be particularly impactful in the context of digital pathology, as the current or possibly future activity of certain types of cells can be highly dependent on their microenvironment.
[0005] It would therefore be advantageous to develop techniques for processing digital pathology images to produce an output that reflects the spatial characterization of the depicted biological objects. Summary of the Invention
[0006] In some embodiments, a computer-implemented method is provided, including a digital pathology imaging system accessing a digital pathology image depicting a section of a biological sample from a subject. The digital pathology imaging system detects a first set of biological object representations and a second set of biological object representations within the digital pathology image. Each of the first set of biological object representations depicts a first biological object of a first type of biological object. Each of the second set of biological object representations depicts a second biological object of a second type of biological object. The digital pathology imaging system uses the first set of biological object representations and the second set of biological object representations to generate a spatial distribution metric that characterizes the position of the first set of biological object representations relative to the second set of biological object representations. The digital pathology imaging system uses the spatial distribution metric to generate subject-level results corresponding to a predicted biological state of the subject or a potential treatment for the subject. The digital pathology imaging system generates a display including the subject-level results. In certain embodiments, the first type of biological object includes a first type of cell, and the second type of biological object includes a second type of cell. In certain embodiments, the first type of biological object comprises lymphocytes and the second type of biological object comprises tumor cells. In certain embodiments, the digital pathology image depicts a biological sample from the subject after treatment with one or more stains, each of which enhances the appearance of one or more of the first type of biological object or the second type of biological object. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by, for each first biological object representation among the one or more first biological object representations, identifying a first point location in the digital pathology image corresponding to the first biological object representation; for each second biological object representation among the one or more second biological object representations, identifying a second point location in the digital pathology image corresponding to the second biological object representation; and determining the spatial distribution metric based on the first and second point locations.In certain embodiments, the first point location in the digital pathology image indicates the location of the first biological object representation. In certain embodiments, the first point location in the digital pathology image is selected by calculating an average point location, a centroid point location, a median point location, or a weighted point location for the first biological object representation. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by calculating, for each of at least some of the one or more first biological object representations and each of at least some of the one or more second biological object representations, a distance between a first point location corresponding to the first biological object representation and a second point location corresponding to the second biological object representation. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by identifying, for each of at least some of the one or more first biological object representations, one or more second biological object representations associated with a distance between the first biological object representation and the second biological object representation. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by defining a spatial grid configured to divide an area of the digital pathology image into a set of image regions, assigning each first biological object representation among the one or more first biological object representations to an image region in the set of image regions, assigning each second biological object representation among the one or more second biological object representations to an image region in the set of image regions, and generating the spatial distribution metric based on the image region assignments. In certain embodiments, the digital pathology imaging system generates the spatial distribution metric by determining one or more image regions in a first set of image regions that have a higher probability of containing the first biological object representation than adjacent image regions in the set of image regions, determining one or more image regions in a second set of image regions that have a higher probability of containing the second biological object representation than adjacent image regions in the set of image regions, and determining the spatial distribution metric further based on the first set of image regions and the second set of image regions.The digital pathology imaging system generates the spatial distribution metric by determining a third set of one or more image regions from the set of image regions that have a higher probability of containing both the first biological object depiction and the set of biological object depictions than adjacent image regions, and determining the spatial distribution metric further based on the third set of image regions. In certain embodiments, the digital pathology imaging system generates a subject-level result corresponding to a predicted biological state of the subject or a potential treatment for the subject using the first spatial distribution metric by comparing the spatial distribution metric generated for the digital pathology image with a previous spatial distribution metric generated for a previous digital pathology image, and outputting a subject-level result generated for the previous digital pathology image based on the comparison. In certain embodiments, the digital pathology imaging system generates the subject-level result by determining a diagnosis, prognosis, treatment recommendation, or treatment eligibility assessment for the subject based on processing the spatial distribution metric and the first set of biological object depictions and the second set of biological object depictions using a trained machine learning model. In certain embodiments, the spatial distribution metric includes a metric defined based on a K-nearest neighbor analysis, a metric defined based on Ripley's K-function, a Morisita-Horn index, a Moran index, a metric defined based on a correlation function, a metric defined based on hot spot / cold spot analysis, or a metric defined based on a Kriging-based analysis. In certain embodiments, the spatial distribution metric is of a first type of metric. The digital pathology imaging system uses the first set of biological object depictions and the second set of biological object depictions to generate a second spatial distribution metric that characterizes the positions of the first set of biological object depictions relative to the second set of biological object depictions. The second spatial distribution metric is of a second type of metric that is different from the first type of metric. Object-level results are further generated using the second spatial distribution metric. In certain embodiments, the digital pathology imaging system receives user input data from a user device including an identifier of an object or a digital pathology image.The digital pathology images are accessed based on the received user input data. The digital pathology imaging system displays and provides subject-level results by providing the subject-level results to a user device. In certain embodiments, the digital pathology imaging system outputs a clinical assessment to the user device for the subject. The clinical assessment includes a diagnosis, prognosis, treatment recommendation, or treatment eligibility assessment for the subject.
[0007] In some embodiments, a method is provided that includes accessing, by a digital pathology imaging system, a digital pathology image depicting a section of a biological sample collected from a subject having a given medical condition. The digital pathology imaging system detects a set of biological object representations within the digital pathology image. The set of biological object representations includes a first set of biological object representations of a first class of biological objects and a second set of biological object representations of a second class of biological objects. The digital pathology imaging system generates one or more relative location representations of the biological object representations. Each of the one or more relative location representations indicates a location of the first biological object representation relative to the second biological object representation. The digital pathology imaging system uses the one or more relative location representations to determine a spatial distribution metric that characterizes the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations. Based on the spatial distribution metric, the digital pathology imaging system generates a result corresponding to a prediction regarding how effectively a given treatment that modulates an immune response will treat the given medical condition in the subject. The digital pathology imaging system determines that the subject is eligible for an outcome-based clinical trial. The digital pathology imaging system generates a display including an indication that the subject is eligible for the clinical trial. In certain embodiments, the spatial distribution metric includes a metric defined based on a K-nearest neighbor analysis, a metric defined based on Ripley's K-function, a Morisita-Horn index, a Moran index, a metric defined based on a correlation function, a metric defined based on hot spot / cold spot analysis, or a metric defined based on a Kriging-based analysis. In certain embodiments, the spatial distribution metric is of a first type of metric, and the digital pathology imaging system uses the one or more relational location representations to determine a second spatial distribution metric that characterizes the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations.The second spatial distribution metric is of a second type of metric different from the first type of metric. The result is further generated based on the second spatial distribution metric. In certain embodiments, generating the result includes a digital pathology image processing system processing the first spatial distribution metric and the second spatial distribution metric using a trained machine learning model. The trained machine learning model is trained using a set of training elements. Each set of training elements corresponds to a different subject who has received a particular treatment associated with the clinical trial. Each set of training elements includes a different set of spatial distribution metrics and a responsiveness value indicating the extent to which the given treatment activated an immune response in the different subject. In certain embodiments, generating the result includes comparing the value of the spatial distribution metric to a threshold. In certain embodiments, the given condition is a type of cancer and the given treatment is an immune checkpoint inhibitor treatment. In certain embodiments, the one or more relational location representations include a set of coordinates identifying a location of the biological object representation within the digital pathology image for each biological object representation in the set of biological object representations. In certain embodiments, generating one or more relational position representations of the biological object representations includes, for each biological object representation from the first set of biological object representations, identifying a first point location in the digital pathology image corresponding to the biological object representation, for each biological object representation from the second set of biological object representations, identifying a second point location in the digital pathology image corresponding to the biological object representation, and comparing the first and second point locations. In certain embodiments, the first point locations in the digital pathology image are selected by calculating a mean point location, a centroid point location, a median point location, or a weighted point location for the biological object representations from the first set of biological object representations.In certain embodiments, the digital pathology imaging system determines the spatial distribution metric by calculating, for each of at least some of the first set of biological object representations and each of at least some of the second set of biological object representations, a distance between a first point location corresponding to the biological object representation in the first set and a second point location corresponding to the biological object representation in the second set of biological object representations. In certain embodiments, the digital pathology imaging system determines the spatial distribution metric by identifying, for each of at least some of the first set of biological object representations, one or more of the second set of biological object representations associated with a distance between a first point location corresponding to the biological object representation in the first set and a second point location corresponding to the biological object representation in the second set of biological object representations. In certain embodiments, the one or more relative location representations include, for each set of image regions in the digital pathology image, a representation of the absolute or relative amount of biological object representations of a first class of biological objects identified as being located in the region and the absolute or relative amount of biological object representations of a second class of biological objects identified as being located in the region. In certain embodiments, the one or more relational location representations include a distance-based probability that a biological object representation from the first set of biological object representations is depicted as being located within a given distance from a biological object representation from the second set of biological object representations. In certain embodiments, the digital pathology imaging system accesses genetic sequencing or radiological imaging data related to the object, and the result is generated further based on characteristics of the genetic sequencing or radiological imaging data. In certain embodiments, the first class of biological objects are tumor cells and the second class of biological objects are immune cells. In certain embodiments, the digital pathology imaging system receives user input data from a user device including an identifier for the object and accesses a digital pathology image in response to receiving the identifier.The digital pathology imaging system generates a display including an indication that the subject is eligible for the clinical trial by providing an indication that the subject is eligible for the clinical trial to a user device. In certain embodiments, the digital pathology imaging system receives an indication that the subject is enrolled in the clinical trial. In certain embodiments, the digital pathology imaging system generates a display including an indication that the subject is eligible for the clinical trial by notifying the subject of a determination of eligibility for the clinical trial.
[0008] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.
[0009] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0010] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0011] The terms and expressions which have been employed are used as conditions of description rather than of limitation, and no intention is intended in the use of such terms and expressions to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.
[0012] The present disclosure is described in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 illustrates an interactive system for generating and processing digital pathology images to characterize relative spatial information of biological objects, according to some embodiments. [Figure 2] FIG. 1 illustrates an exemplary system for processing object depiction data to generate a spatial distribution metric, according to some embodiments. [Figure 3A] FIG. 1 illustrates a process for providing a health-related assessment based on spatially specific image processing of digital pathology images, according to some embodiments. [Figure 3B] FIG. 1 illustrates a process for providing a health-related assessment based on spatially specific image processing of digital pathology images, according to some embodiments. [Figure 4] FIG. 1 illustrates a process for processing an image using a landscape-based spatial point process analysis framework, according to some embodiments. [Figure 4A] FIG. 1 illustrates a process for processing an image using a landscape-based spatial point process analysis framework, according to some embodiments. [Figure 4B]FIG. 1 illustrates a process for processing an image using a landscape-based spatial point process analysis framework, according to some embodiments. [Figure 4C] FIG. 1 illustrates a process for processing an image using a landscape-based spatial point process analysis framework, according to some embodiments. [Figure 5A-5B] FIG. 1 illustrates an exemplary processing of an image using a discriminant-based spatial point process analysis framework, according to some embodiments. [Figure 5C] FIG. 1 illustrates an exemplary processing of an image using a discriminant-based spatial point process analysis framework, according to some embodiments. [Figures 6A-6D] 1A-1C illustrate exemplary distance and intensity-based metrics for characterizing the spatial arrangement of object depictions in an exemplary image, according to some embodiments. [Figure 7] FIG. 1 illustrates a process for processing an image using a grid-based spatial area analysis framework, according to some embodiments. [Figure 7A] FIG. 1 illustrates a process for processing an image using a grid-based spatial area analysis framework, according to some embodiments. [Figure 7B] FIG. 1 illustrates a process for processing an image using a grid-based spatial area analysis framework, according to some embodiments. [Figure 7C] FIG. 1 illustrates a process for processing an image using a grid-based spatial area analysis framework, according to some embodiments. [Figure 8] FIG. 1 illustrates a process for processing an image using Moran's index, according to some embodiments. [Figure 8A] FIG. 1 illustrates a process for processing an image using Moran's index, according to some embodiments. [Figure 8B] FIG. 1 illustrates a process for processing an image using Moran's index, according to some embodiments. [Figure 8C]FIG. 1 illustrates a process for processing an image using Moran's index, according to some embodiments. [Figure 9] FIG. 1 illustrates a process for processing an image using a hotspot-based spatial area analysis framework, according to some embodiments. [Figure 9A] FIG. 1 illustrates a process for processing an image using a hotspot-based spatial area analysis framework, according to some embodiments. [Figure 9B] FIG. 1 illustrates a process for processing an image using a hotspot-based spatial area analysis framework, according to some embodiments. [Figure 9C] FIG. 1 illustrates a process for processing an image using a hotspot-based spatial area analysis framework, according to some embodiments. [Figure 10] FIG. 1 illustrates a process for processing images using a geostatistical analysis framework, according to some embodiments. [Figure 10A] FIG. 1 illustrates a process for processing images using a geostatistical analysis framework, according to some embodiments. [Figure 10B] FIG. 1 illustrates a process for processing images using a geostatistical analysis framework, according to some embodiments. [Figure 11] FIG. 10 shows receiver operating curves characterizing the performance of a trained logistic regression model to predict the occurrence of microsatellite instability based on processing of digital pathology images, according to some embodiments. [Figure 12] FIG. 1 illustrates the process of assigning predicted outcome labels to each subject in a study cohort using a nested Monte Carlo cross-validation modeling strategy. [Figure 13] FIG. 1 shows Kaplan-Meir plots for subjects in an analysis of two subject cohorts.
[0014] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by a dash following the reference label and a second label that distinguishes among the similar components. When only a first reference label is used in the specification, the description is applicable to any similar component having the same first reference label, regardless of the second reference label. DETAILED DESCRIPTION OF THE INVENTION
[0015] Digital images are increasingly being used in medical contexts to facilitate clinical evaluations such as diagnosis, prognosis, treatment selection, and treatment assessment, among other uses. In the field of digital pathology, digital pathology images can be processed to estimate whether a given image contains a representation of a particular type or class of biological object. For example, a section of a tissue sample can be stained so that representations of a particular type of biological object (e.g., a particular type of cell, a particular type of organelle, or a blood vessel) preferentially absorb the stain and are therefore depicted in a particular, more intense color. The tissue sample can be imaged according to the techniques disclosed herein. The digital pathology image is then processed to detect biological object representations. Detection of biological object representations can be based on biological objects meeting certain criteria under analysis corresponding to a staining profile, such as having a size within a specified range, a specified type of shape, or at least a specified amount of consecutive high-intensity pixels. In certain embodiments, clinical evaluation or recommendations can be made based on whether a representation of a particular type or class of object is observed and / or the amount of representation of one or more particular types or classes of objects.
[0016] With the evolution of imaging technology, digital imaging of tumor tissue slides has become a routine clinical procedure for managing many types of conditions. Digital pathology images can capture multiple objects of a given type or class at high resolution. It may be advantageous to characterize the degree of spatial heterogeneity of biological objects captured in digital pathology images, as well as the degree to which objects of a given type are spatially aggregated and / or distributed relative to each other and / or to objects of different types. The current or potential activity or function of biological objects can vary significantly depending on the biological object's microenvironment. Objectively characterizing the location of a particular type of biological object representation can substantially affect the quality of current diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility determination. Similarly, objectively characterizing the relationship of multiple types of biological objects within a digital pathology image or a region of a digital pathology image can substantially affect the results of an analysis. The location and relationship of biological object representations in a digital pathology image can be correlated with the location and relationship of corresponding biological objects in a subject's tissue sample. As disclosed herein, such objective spatial characterization can be performed by detecting a set of biological object representations from a digital pathology image. The objects may be represented according to one or more spatial analysis frameworks, including, but not limited to, a spatial point process analysis framework, a spatial area analysis framework, a geostatistical analysis framework, a graph-based framework, etc. In some embodiments, each detected biological object representation is associated with a particular point location within the image and may be further associated with an identifier for a particular type of object. In some embodiments, for each set of regions within the image, and for each of one or more particular types of objects, metadata may be stored that indicates the amount or density of biological object representations of each particular type that are predicted or determined to be located within the region.
[0017] Spatial aggregation can include a measure of how objects in a digital pathology image are spatially aggregated or distributed across the entire digital pathology image or across a region of the digital pathology image. For example, it may be advantageous to determine the extent to which one type or class of biological object (e.g., lymphocytes) is spatially intermixed with another type or class of biological object (e.g., tumor cells). To illustrate, intratumoral tumor-infiltrating lymphocytes (TILs) are located within the tumor and have direct interactions with tumor cells, while stromal TILs are located within the tumor stroma and do not have direct interactions with tumor cells. Not only do intratumoral TILs have different activity patterns than stromal TILs, but each cell type can be associated with a different type of microenvironment, further influencing the behavioral differences between TIL types. When lymphocytes are detected in a particular location (e.g., within a tumor), the fact that the lymphocytes were able to infiltrate the tumor can convey information about the activity of the lymphocytes and / or tumor cells. Furthermore, the microenvironment can affect the current and future activity of lymphocytes. Identifying the relative location of particular types of biological objects can be particularly useful for predictive applications, such as identifying prognosis and treatment options, assessing patient eligibility for clinical trials, and characterizing the immunological characteristics of subjects and their conditions.
[0018] As another form of objective characterization of the locations and relationships of the detected biological object representations, the detected biological object representations can be used to generate one or more spatial distribution metrics that can characterize, at the region, image, and / or object level, how interspersed a biological object of a given type or class is with another type or class of biological object, how clustered it is with other objects of the same type, and / or how clustered it is with another given type of biological object. For example, a digital pathology image processing system can detect a first set of biological object representations and a second set of biological object representations in a digital pathology image. The system can predict that each of the first set of biological object representations depicts a first type of biological object (e.g., lymphocytes) and each of the second set of biological object representations depicts a second type of biological object (e.g., tumor cells). The digital pathology imaging system can perform a distance-based assessment to generate a spatial distribution metric that indicates the extent to which individual biological object representations in a first set of biological object representations are spatially integrated or separated from individual biological object representations in a second set of biological object representations, and / or the extent to which the first set of biological object representations (e.g., collectively) are spatially integrated or separated from the second set of biological object representations (e.g., collectively). As disclosed herein, various spatial distribution metrics have been developed and applied for this purpose.
[0019] Principles and quantitative methods from advanced analysis (e.g., spatial statistics) can be applied to generate novel solutions that meet these needs. The techniques provided herein can be used to process digital pathology images to generate results that characterize the spatial distribution and / or spatial pattern of one or more specific types or classes of depicted objects (e.g., biological objects). The digital pathology images can include digital images of stained sections of the sample. Processing can include detecting depictions of biological objects of each of a plurality of specific types (e.g., corresponding to biological cells of each of a plurality of types). Biological object detection can include detecting one or more of a first set of biological object depictions corresponding to a first biological object type and each of a second set of biological object depictions corresponding to a second biological object type. Additionally or alternatively, object detection can include identifying, for each of a set of regions in the digital pathology image and for each of a plurality of specific biological object types, a higher-order metric defined as quantity-dependent or quantity-correlated, or a lower-order metric of the biological object (e.g., count, density, or image intensity inferred to represent the amount of a particular type of biological object represented in the corresponding image region). Additionally, spatial distribution metrics can be used in combination with other metrics (e.g., RNA sequencing, radiological imaging (CT, MRI, etc.)) to improve the predictive ability of the metrics or to find novel biomarkers for unmet medical needs.
[0020] Image locations of one or more biological object representations can be determined. The image locations can be determined and represented according to one or more spatial analysis frameworks, such as a spatial point process analysis framework, a spatial area analysis framework, a geostatistical analysis framework, or a graph-based analysis framework. For example, a biological object can be associated with a single point location within the digital pathology image. Although the biological object representation may extend across multiple pixels or voxels, the single point location can be selected to indicate or represent the location of the biological object representation within the digital pathology image. As another example, the biological object representation can be collectively represented or represented by one or more other biological object representations as contributing to the count of objects detected within a particular region of the image, the density of biological objects detected within a particular region of the image, the pattern of biological objects detected within a particular region of the image, etc.
[0021] Digital pathology imaging systems can use spatial distribution metrics to facilitate, for example, diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility (e.g., whether a subject is eligible for or recommended for a clinical trial or a particular group of clinical trials). For example, a particular prognosis can be identified in response to detecting a particular degree of infiltration of a set of biological objects of a first type or class within a biological object of a second type or class, and a more relevant and accurate prognosis can be identified in response to detecting greater lymphocytic infiltration within individual tumors and / or metastatic tumor nests. As another example, a diagnosis of tumor or cancer stage can be informed based on the extent to which immune cells are spatially integrated with cancerous cells (e.g., with a higher degree of integration generally corresponding to an earlier stage). As yet another example, greater therapeutic efficacy can be determined when lymphocytes are less spatially proximate to tumor cells after treatment is initiated compared to before treatment or compared to projected proximity based on one or more prior assessments performed on a given subject.
[0022] Biological object detection can be used to generate results that may include or be based on spatial distribution metrics, and that can indicate the proximity between representations of the same or different types of biological objects and / or the degree to which representations of one or more types of biological objects colocalize. The colocalization of representations of biological objects can represent the similar location of multiple cell types in each of one or more regions of a digital pathology image. The results can be indicative of and / or predictive of interactions between different biological objects and biological object types that may occur within the microenvironment of a subject's or patient's internal structure as represented by a sample collected from the subject or patient. Such interactions may support and / or be essential for biological processes such as tissue formation, homeostasis, regenerative processes, or immune responses. Therefore, the spatial information conveyed by the results can be informative regarding the function and activity of specific biological structures and thus can be used as quantitative support, for example, to characterize disease states and prognoses. Results indicating where specific biological objects are located within a biological microenvironment can be used to select a treatment predicted to be effective for a particular subject (e.g., compared to other treatment options) or to predict outcomes for other subjects.
[0023] In certain embodiments, multiple spatial distribution metrics can be generated. In particular, one or more metrics can be generated, each corresponding to a metric type among one or more metric types. For example, one or more first metrics can be generated using a spatial point process analysis framework. The first metric can be based on the distance between representations of different types of biological objects. For example, the first metric can use the Euclidean distance between a biological object representation corresponding to a tumor cell and a biological object representation corresponding to a lymphocyte. Other distance metrics can also be used. One or more second metrics can be generated using a spatial area analysis framework. The second metric can characterize the count or density of representations of a first type of biological object within various image regions relative to the count or density of other representations of a second type of biological object.
[0024] The machine learning model or rules can be used to generate results corresponding to diagnosis, prognosis, treatment assessment, treatment selection, treatment eligibility (e.g., eligibility to be accepted or recommended for a clinical trial or a particular group of a clinical trial), and / or prediction of genetic mutations, genetic alterations, biomarker expression levels (including, but not limited to, genes or proteins), etc., using, for example, one or more metrics each corresponding to a metric type of one or more metric types. Machine learning models can include, by way of example and not limitation, classification, regression, decision tree, or neural network techniques that are trained to learn to use one or more weights when processing metrics to generate results.
[0025] The digital pathology imaging system can further identify and learn to recognize patterns in the locations and relationships of the detected biological object depictions based in part on one or more spatial distribution metrics. For example, the digital pathology imaging system can detect patterns in the locations and relationships of the detected biological object depictions in the digital pathology image of the first sample. The digital pathology imaging system can generate a mask or other pattern storage data structure from the recognized pattern. The digital pathology imaging system can use the spatial distribution metrics as described herein to predict a diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility decision. The digital pathology imaging system can store a predicted prognosis, etc., in association with the detected pattern and / or the generated mask. The digital pathology imaging system can receive the subject's results and verify the predicted prognosis, etc.
[0026] The digital pathology image processing system can then, when processing a second digital pathology image from a second sample, detect patterns of location and relationship of the detected biological object depictions in the second digital pathology image. The digital pathology image processing system can recognize similarities between the detected patterns of location and relationship in the second digital pathology image and the mask or stored detected pattern from the first digital pathology image. The digital pathology image processing system can inform a predicted prognosis, treatment recommendation, or treatment eligibility determination based on the recognized similarities and / or the subject outcome. As an example, the digital pathology image processing system can compare the stored mask with the patterns of location and relationship of the detected biological object depictions in the second digital pathology image. The digital pathology image processing system can determine one or more spatial distribution metrics for the second digital pathology image and compare the spatial distribution metrics of the detected biological object depictions in the first digital pathology image and the second digital pathology image based on a comparison of the stored mask and the recognized pattern from the second digital pathology image.
[0027] The pattern detected from the first digital pathology image processing system can relate in many ways to the location and relationship of one or more first biological object representations of one or more types. For example, the pattern can relate to the location and relationship of the first biological object of the first type in the digital pathology image without the context of other biological object representations in the digital pathology image. The pattern can relate to an abstract representation of the location and / or relationship of the biological object representation within the boundaries of the digital pathology image (e.g., assessing the coordinates of the detected biological object representation, potentially without its context as a biological object representation). As another example, the pattern can relate to the location and relationship of the biological object representation of the first type relative to all of the other biological object representations in the digital pathology image. As yet another example, the pattern can relate to the location and relationship of one or more biological object representations of the first type relative to the location and relationship of one or more biological object representations of a second type.
[0028] The patterns detected from the digital pathology images may be used to identify, for example, the type of specimen the digital pathology images depict (e.g., lung biopsies, liver tissue samples, blood samples, formalin-fixed paraffin-embedded specimens, frozen specimens, surgical excisions, etc., from various organs, tumors, and / or metastatic sites, etc.). Patterns detected from digital pathology images can be associated with a context including the cell preparation obtained from the sample (e.g., biopsy procedures including, but not limited to, cell preparations obtained from exhale (e.g., biopsy procedures including, but not limited to, core needle biopsy and fine needle aspiration), the method of sample preparation (e.g., type of stain used, age of the sample, etc.), the number and specific types of biological objects depicted throughout the sample or incorporated into the pattern (e.g., sample cell types, structures (glands, tumor islands, cell layers, blood vessels, etc.), individual cells (tumor cells, immune cells, mitotic cells, stromal cells, endothelial cells, etc.), and cellular components (nucleus, cytoplasm, membrane, cilia, mucus excreta, etc.)), the number and type of spatial distribution metrics used to detect or prepare the pattern, the type of object-level result associated with the pattern, instructions within the type of object-level result, the degree of validation of the object-level result, and many other factors used to characterize the pattern. The context can be used to improve pattern recognition and application to future digital pathology images.
[0029] In some embodiments, a pattern may only be applicable to the same type of sample, the same type of biological object delineation, the same type of spatial distribution metric, sample-type object-level results, etc., although a digital pathology image processing system can be trained to apply pattern recognition methodologies across multiple types. For example, a digital pathology image processing system can be trained to recognize the broad applicability of a pattern associated with lymphocyte infiltration and location within tissue sample cells and provide similar object-level results based on analysis of digital pathology images corresponding to different types of tissue samples. The ability to reference and apply a pattern can be based on the applicability of a spatial distribution metric associated with different types of detected biological object delineations and across digital pathology images of different tissue sample types. The spatial distribution metric provides an objective and quantifiable measure for multi-dimensional comparisons.
[0030] Additionally or alternatively, the digital pathology imaging system can further use spatial distribution metrics to facilitate the identification of treatment options. For example, detecting an output indicating that lymphocytes are spatially integrated with tumor cells can selectively recommend immunotherapy or immune checkpoint therapy. As another example, detecting an output indicating that lymphocytes are spatially integrated with tumor cells can selectively recommend atezolizumab + bevacizumab + carboplatin + paclitaxel (ABCP) or atezolizumab + carboplatin + paclitaxel (ACP) over another chemotherapy treatment. The other chemotherapy treatment can include or be bevacizumab + carboplatin + paclitaxel (BCP). Other strategies can use other biological objects or cellular components or compartments to predict diagnosis, biomarker expression, or treatment response (e.g., vascular distribution, distribution of specific nuclear morphologies in lymphoma, etc.).
[0031] Facilitating the identification of a diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility can include automatically generating a potential diagnosis, prognosis, treatment assessment, and / or treatment selection. The automatic identification can be based on one or more learned and / or static rules. The rules can have an if-then format, where conditions can include inequality and / or one or more thresholds that, for example, may indicate that a metric above the threshold is associated with suitability for a particular treatment. The rules can alternatively or additionally include functions, such as functions relating a numerical metric to a severity score for a disease or a quantified score of eligibility for a treatment. The digital pathology imaging system can output the potential diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility determination as a recommendation and / or prediction. For example, the digital pathology imaging system can provide output to a locally combined display, transmit the output to a remote device or an access terminal of a remote device, store the results in local or remote data storage, etc. In this manner, a human user (e.g., a physician and / or healthcare provider) can use the automatically generated output or form a different assessment informed by the quantitative metrics discussed herein.
[0032] Facilitating the identification of a diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility determination may include outputting a spatial distribution metric consistent with the disclosed subject matter. For example, the output may include a subject identifier (e.g., the subject's name), stored clinical data associated with the subject (e.g., past diagnoses, possible diagnoses, current treatments, symptoms, test results, and / or vital signs), and the determined spatial distribution metric. The output may include the digital pathology image from which the spatial distribution metric was derived, and / or a modified version thereof. For example, the modified version of the digital pathology image may include an overlay and / or markings identifying each biological object representation detected in the digital pathology image. The modified version of the digital pathology image may further provide information about the detected biological object representation. For example, for each biological object representation, an interactive overlay may provide a particular object category corresponding to the subject. A human user (e.g., a physician and / or healthcare provider) may then use the output including the spatial distribution metric to identify a diagnosis, prognosis, treatment assessment, treatment selection, or treatment eligibility determination.
[0033] In certain embodiments, multiple types of spatial distribution metrics are generated using detected biological object depictions from a single digital pathology image. Multiple types of spatial distribution metrics can be used in combination according to the subject matter disclosed herein. The multiple types of spatial distribution metrics can correspond to different or the same frameworks, for example, related to how the location of each biological object depiction is characterized. The multiple types of spatial distribution metrics can include different variable types (e.g., calculated using different algorithms) and can be presented on different value scales. The multiple types of spatial distribution metrics can be collectively processed using rules or machine learning models to generate labels. The labels can correspond to predicted diagnoses, prognoses, treatment assessments, treatment selections, and / or treatment eligibility decisions.
[0034] In certain embodiments, a computer-implemented method is provided. A digital pathology imaging system can access one or more digital pathology images. Each of the one or more digital pathology images can depict a section of a biological sample from a subject. The depicted section can include sections stained with one or more stains. The digital pathology imaging system detects a first set of biological object representations and a second set of biological object representations in each of the one or more digital pathology images. Each of the first set of biological object representations can depict a first type of biological object. Each of the second set of object representations can depict a second type of biological object. Using the first set of biological object representations and the second set of biological object representations, the digital pathology imaging system generates one or more spatial distribution metrics of a first type of spatial distribution metrics. Each of the one or more first spatial distribution metrics characterizes a position of the first set of biological object representations relative to the second set of biological object representations. Using the first set of biological object representations and the second set of biological object representations, the digital pathology imaging system generates one or more spatial distribution metrics of a second type. The second type of spatial distribution metric characterizes the positions of the first set of biological object representations relative to the second set of biological object representations. Using the one or more first spatial distribution metrics and the one or more second spatial distribution metrics, the digital pathology imaging system can generate subject-level results corresponding to a predicted biological state of the subject or a potential treatment for the subject. The digital pathology imaging system displays and provides the subject-level results. In addition to providing the subject-level results, the digital pathology imaging system can provide a clinical evaluation for the subject based on the subject-level results. The clinical evaluation can include diagnosis, prognosis, treatment assessment, treatment selection, and / or treatment eligibility.
[0035] The spatial distribution metric characterizing the locations of the first set of biological object depictions can be determined based on, for example and without limitation, a point process, an area / grid process, a geostatistical process, etc. In certain embodiments, the first type of biological object can include a first type of cell, and the second type of biological object can include a second type of cell. As an example, the first type of biological object can include lymphocytes, and the second type of biological object can include tumor cells. As another example, the first type of biological object can include macrophages, and the second type of biological object can include fibroblasts. In certain embodiments, the first type of biological object can include a first class of biological object defined, for example, by characteristic properties of the first type (e.g., size, shape, color, expected behavior, texture of the biological object or a component or compartment of the biological object), and the second type of biological object can include a second class of biological object defined, for example, by characteristic properties of the second type or a variant of the first type. It will be appreciated that the subject matter disclosed herein may be equally applicable to any biological object that can be represented as a point corresponding to a location in a digital pathology image.
[0036] In certain embodiments, generating one or more spatial distribution metrics of a first type can include identifying a first point location in the one or more digital pathology images for each first biological object representation among the one or more first biological object representations. The first point location can correspond to a location of the depicted first biological object. Generating one or more spatial distribution metrics of a first type can further include identifying a second point location in the one or more digital pathology images for each second biological object among the one or more second biological objects. The second point location can correspond to a location of the depicted second biological object. Generating one or more spatial distribution metrics of a first type can further include determining one or more spatial distribution metrics of the first type based on the first and second point locations. In certain embodiments, generating the one or more spatial distribution metrics may include implementing a distance-based technique to evaluate, for each of at least some first biological objects of the one or more first biological objects and for each of at least some second biological objects of the one or more second biological objects, a distance between a first point location corresponding to the first biological object and a second point location corresponding to the second biological object.
[0037] In certain embodiments, generating one or more spatial distribution metrics of the second type can include defining a spatial grid configured to divide an area of one of the digital pathology images into a set of image regions. Generating the one or more spatial distribution metrics of the second type can include assigning each second biological object of the one or more second biological objects to an image region of the set of image regions. Generating the one or more spatial distribution metrics of the second type can include generating the one or more spatial distribution metrics of the second type based on the image region assignment of each second biological object of the one or more second biological objects.
[0038] Generating object-level results may include processing one or more spatial distribution metrics of a first type and one or more spatial distribution metrics of a second type using a trained machine learning model. The trained machine learning model may include, by way of example and not limitation, a regression model, a decision tree model, or a neural network model. The first type of metric may be one of a set of metric types. The second type of metric may be another of the set of metric types. The set of metric types may include metrics defined based on K-nearest neighbor analysis, metrics defined based on Ripley's K-function, Morisita-Horn index, Moran index, Geary's C-index, G-function, metrics defined based on correlation function, metrics defined based on hot spot analysis or cold spot analysis, or metrics defined based on Kriging-based analysis.
[0039] In certain embodiments, a method is provided that includes sending a request communication from a client computing system to a remote computing system to process one or more digital pathology images depicting specific sections of a biological sample from a subject, wherein in response to receiving the request communication from the client computing system, the remote computing system accesses the one or more digital pathology images and performs analysis in accordance with the subject matter disclosed herein.
[0040] According to the presently disclosed subject matter, in certain embodiments, there is provided the use of subject-level results in treating a subject. Subject-level results can be provided according to the presently disclosed subject matter.
[0041] In certain embodiments, a method is provided. A digital pathology image is accessed by a digital pathology imaging system. The digital pathology image depicts a tissue slide stained with one or more stains, the tissue of the tissue slide being collected from a subject with a particular pathology condition. The digital pathology image includes a depiction of one or more biological objects. The one or more biological objects may include a set of cells. The set of cells may include a set of tumor cells and a set of other cells. The set of other cells may be a set of immune cells or a set of stromal cells. The digital pathology imaging system may identify a set of locations within the digital pathology image corresponding to the one or more biological objects, such as tumor cell locations. Each tumor cell location within the set of tumor cell locations may correspond to a tumor cell within the set of tumor cells. The digital pathology imaging system may identify a set of other locations within the digital pathology image corresponding to one or more other biological objects, such as other cell locations. Each other cell location within the set of other cell locations may correspond to a cell within the other set of cells. The digital pathology imaging system may generate one or more relational location representations. Each of the one or more relational location representations can indicate the positions of at least some of a first set of cells relative to the positions of at least some of a second set of cells. Using the one or more relational location representations, the digital pathology imaging system can determine a set of spatial distribution metrics. Each spatial distribution metric in the set of spatial distribution metrics can characterize the extent to which at least some of the other sets of cells are described as interspersed with at least some of the set of tumor cells. The digital pathology imaging system can generate a result based on the set of spatial distribution metrics. The result corresponds to a prediction of whether and / or how effectively a particular treatment that modulates the immune response will effectively treat a particular condition in the subject. Based on the result, the subject is determined to be eligible for the clinical trial. An indication that the subject is eligible for the clinical trial is output.
[0042] Generating the results may include processing the set of spatial heterogeneity metrics using a trained machine learning model. The trained machine learning model may have been trained using a set of training elements. Each set of training elements may correspond to a different subject who received a particular treatment associated with the clinical trial. Each set of training elements may include a different set of spatial heterogeneity metrics and a responsiveness value indicating whether and / or to what extent the particular treatment activated an immune response in the subject.
[0043] In certain embodiments, the disease state can be a type of cancer, and / or the specific treatment can be an immune checkpoint inhibition treatment. The one or more relational location representations can include, for each cell in the set of cells, a set of coordinates identifying the location of the cell's depiction in the digital pathology image. The one or more relational location representations can include, for each of a set of regions in the digital pathology image, a representation of the absolute or relative amount of tumor cells, stromal cells, and / or immune cells identified as being located in the region. The one or more relational location representations can indicate a distance-based probability that a first type of cell is depicted as being located within a distance from a second type of cell. The first type and the second type can correspond to immune cells, stromal cells, or tumor cells, respectively. Gene sequencing and / or radiological imaging data can be collected for the subject. The results can further depend on characteristics of the gene sequencing and / or radiological imaging data.
[0044] The term "biological object representation," as referred to herein, can refer to a specific portion of an image (e.g., one or more pixels, a defined region of an image, etc.) that is identified or has been identified as corresponding to a particular type of biological object. The biological object representation can depict a biological object (e.g., a cell). The biological object representation can include one or more pixels and / or one or more voxels. A pixel or voxel of the biological object representation can correspond, for example, to a centroid, edge, center of mass, or the entirety of what is predicted to be a representation of the biological object. The biological object representation can be identified using machine learning algorithms, one or more static rules, and / or computer vision techniques applied to the digital pathology image. The image can depict stained sections, and the stain can be selected to be preferentially absorbed by biological objects of a particular type of interest, such that identification of the biological object representation can include an intensity-based evaluation.
[0045] The term "biological object," as referred to herein, may refer to a biological unit. Biological objects may include, by way of example and not limitation, a cell, an organelle (e.g., a nucleus), a cell membrane, a stroma, a tumor, or a blood vessel. It will be understood that biological objects may include three-dimensional objects, and that a digital pathology image may capture only a single two-dimensional slice of the object, which need not extend entirely throughout the object along the plane of the two-dimensional slice. Nevertheless, references herein may refer to such a captured portion as depicting the biological object.
[0046] The term "type of biological object" or biological object type, as referred to herein, can refer to a category of biological units. By way of example and not limitation, a type of biological object can refer to a cell (generally), a specific type of cell (e.g., a lymphocyte or tumor cell), a cell membrane (generally), etc. Some disclosures can refer to detecting a biological object representation corresponding to a first type of biological object and another biological object representation corresponding to a second type of biological object. The first and second types of biological objects can have similar, the same, or different levels of specificity and / or generality. For example, the first and second types of biological objects can be identified as lymphocyte type and tumor cell type, respectively. As another example, the first type of biological object can be identified as a lymphocyte, and the second type of biological object can be identified as a tumor.
[0047] The term "spatial distribution metric," as referred to herein, can refer to a metric that characterizes the spatial arrangement of particular biological object representations within an image relative to one another and / or relative to other particular biological object representations. A spatial distribution metric can characterize how one type of biological object (e.g., lymphocytes) infiltrates another type of biological object (e.g., tumors), how interspersed it is with another type of object (e.g., tumor cells), how physically proximate it is to another type of object (e.g., tumor cells), and / or how co-localized it is with another type of object (e.g., tumor cells).
[0048] FIG. 1 illustrates an interactive system or network 100 of interacting systems (e.g., specially configured computer systems) that can be used in accordance with the disclosed subject matter for generating and processing digital pathology images that characterize relative spatial information of biological objects, according to some embodiments.
[0049] The digital pathology imaging system 105 can generate one or more digital images corresponding to a particular sample. For example, an image generated by the digital pathology imaging system 105 can include a stained section of a biopsy sample. As another example, an image generated by the digital pathology imaging system 105 can include a slide image of a liquid sample (e.g., a blood smear). As another example, an image generated by the digital pathology imaging system 105 can include fluorescence microscopy, such as a slide image depicting fluorescence in situ hybridization (FISH) after a fluorescent probe binds to a target DNA or RNA sequence.
[0050] Some types of samples (e.g., biopsies, solid samples, and / or tissue-containing samples) can be processed by the sample preparation system 110 to fix and / or embed the sample. The sample preparation system 110 can facilitate infiltrating the sample with a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding material (e.g., tissue wax). For example, the fixation subsystem can fix the sample by exposing the sample to a fixative for at least a threshold amount of time (e.g., at least 3 hours, at least 6 hours, or at least 12 hours). The dehydration subsystem can dehydrate the sample (e.g., by exposing the fixed sample and / or a portion of the fixed sample to one or more ethanol solutions) and potentially wash the dehydrated sample using a wash intermediate (e.g., including ethanol and tissue wax). The embedding subsystem can infiltrate the sample (e.g., one or more times corresponding to a predefined period of time) with heated (e.g., therefore liquid) tissue wax. The tissue wax can include paraffin wax and potentially one or more resins (e.g., styrene or polyethylene). The sample and wax can then be cooled, and the wax-infiltrated sample can then be blocked.
[0051] The sample slicer 115 can accept a fixed, embedded sample and create a set of sections. The sample slicer 115 can expose the fixed, embedded sample to cold or low temperatures. The sample slicer 115 can then cut the cooled sample (or a trimmed version of it) to create a set of sections. Each section can have a thickness that is (for example) less than 100 μm, less than 50 μm, less than 10 μm, or less than 5 μm. Each section can have a thickness that is (for example) greater than 0.1 μm, greater than 1 μm, greater than 2 μm, or greater than 4 μm. Cutting of the cooled sample can be performed in a warm water bath (e.g., at a temperature of at least 30° C., at least 35° C., or at least 40° C.).
[0052] The automated staining system 120 can facilitate staining one or more of the sample compartments by exposing each compartment to one or more stains (e.g., hematoxylin and eosin, immunohistochemistry, or specialized stains). Each compartment can be exposed to a predefined amount of stain for a predefined time. In certain embodiments, a single compartment is exposed to multiple stains simultaneously or sequentially.
[0053] Each of the one or more stained sections can be presented to an image scanner 125, which can capture a digital image of the section. The image scanner 125 can include a microscope camera. The image scanner 125 can capture digital images at multiple magnification levels (e.g., using a 10x objective, a 20x objective, a 40x objective, etc.). Image manipulation can be used to capture selected portions of the sample at a desired magnification range. The image scanner 125 can also capture annotations and / or morphometry identified by a human operator. In certain embodiments, the sections are returned to the automated staining system 120 after one or more images have been captured so that the sections can be washed, exposed to one or more other stains, and imaged again. When multiple stains are used, the stains can be selected to have different color profiles so that a first region of an image corresponding to a portion of a first section that has absorbed a large amount of a first stain can be distinguished from a second region of an image (or a different image) corresponding to a portion of a second section that has absorbed a large amount of a second stain.
[0054] It will be understood that one or more components of digital pathology imaging system 105 can, in certain embodiments, operate in conjunction with a human operator. For example, the human operator can move a sample through various subsystems (e.g., of sample preparation system 110 or of digital pathology imaging system 105) and / or initiate or terminate the operation of one or more subsystems, systems, or components of digital pathology imaging system 105. As another example, one or more components of the digital pathology imaging system (e.g., one or more subsystems of sample preparation system 110) can partially or wholly replace the actions of a human operator.
[0055] Additionally, while various described and illustrated features and components of the digital pathology imaging system 105 relate to processing solid and / or biopsy samples, it will be understood that other embodiments may relate to liquid samples (e.g., blood samples). For example, the digital pathology imaging system 105 can be configured to accept a liquid sample (e.g., blood or urine) slide, including a base slide, a smeared liquid sample, and a cover. The image scanner 125 can then capture an image of the sample slide. Further embodiments of the digital pathology imaging system 105 can relate to capturing images of the sample using advanced imaging techniques, such as FISH, as described herein. For example, once fluorescent probes are introduced into the sample and allowed to bind to target sequences, appropriate imaging can be used to capture and further analyze the image of the sample.
[0056] A given sample can be associated with one or more users (e.g., one or more physicians, laboratory technicians, and / or healthcare providers). Associated users can include those who created the sample to be imaged, ordered the test or biopsy, and / or have authorization to receive the test or biopsy results. For example, a user can correspond to a physician, pathologist, clinician, or subject (from whom the sample was taken). A user can (for example) initially submit one or more requests (e.g., identify a subject) using one or more devices 130 to process a sample through the digital pathology image generation system 105 and to process the resulting images through the digital pathology image processing system 135.
[0057] In certain embodiments, the digital pathology image generation system 105 transmits the digital pathology image created by the image scanner 125 back to the user device 130, which communicates with the digital pathology image processing system 135 to initiate automated processing of the digital pathology image. In certain embodiments, the digital pathology image generation system 105 utilizes the digital pathology image created by the image scanner 125 directly to the digital pathology image processing system 135, e.g., at the direction of a user of the user device 130. Although not shown, other intermediate devices (e.g., a data store on the digital pathology image generation system 105 or a server connected to the digital pathology image processing system 135) may also be used. Additionally, for simplicity, only one digital pathology image processing system 135, digital pathology image generation system 105, and user device 130 are shown in the network 100. The present disclosure contemplates the use of one or more of each type of system and its components without necessarily departing from the teachings of the present disclosure.
[0058] The digital pathology image processing system 135 can be configured to identify spatial characteristics of the images and / or characterize the spatial distribution of biological object depictions. The section aligner subsystem 140 can be configured to align multiple digital pathology images and / or regions of digital pathology images corresponding to the same sample. For example, multiple digital pathology images can correspond to the same section of the same sample. Each image can depict a section stained with a different stain. As another example, multiple digital pathology images can each correspond to a different section of the same sample (e.g., each corresponding to the same stain, or different subsets of the images corresponding to different stains). For example, alternating sections of the sample can be stained with different stains.
[0059] The section aligner subsystem 140 can determine whether and / or how to translate, rotate, scale, and / or distort each digital pathology image so that the digital pathology images corresponding to a single sample and / or a single section are aligned. The alignment can be determined using (for example) a correlation evaluation (e.g., to identify an alignment that maximizes correlation).
[0060] The biological object detector subsystem 145 can be configured to automatically detect representations of one or more specific types of objects (e.g., biological objects) in each of the registered digital pathology images. The object types can include, for example, types of biological structures, such as cells. For example, a first set of biological objects can correspond to a first cell type (e.g., immune cells, white blood cells, lymphocytes, tumor-infiltrating lymphocytes, etc.), and a second set of biological objects can correspond to a second cell type (e.g., tumor cells, malignant tumor cells, etc.) or a type of biological structure (e.g., tumor, malignant tumor, etc.). The biological object detector subsystem 145 can detect representations of each of one or more specific types of biological objects from the registered digital pathology images. The digital pathology images can depict various stains in a single digital pathology image. Such a digital pathology image can include a single image that can correspond to a section of a sample stained with each of multiple stains. For example, the biological object detector subsystem 145 can detect representations of lymphocytes and tumor cells from a single digital pathology image. The biological object detector 145 can detect depictions of biological objects from different digital pathology images corresponding to different stains, for example.
[0061] For example, representations of lymphocytes can be detected in a first digital pathology image, and representations of tumor cells can be detected in a second digital pathology image. The first digital pathology image can depict an image of a section of a sample stained with a first stain, and the second digital pathology image can depict the same section stained with a second stain and re-imaged. The biological object detector subsystem 145 can detect representations of a first particular type of biological object in the first digital pathology image, which can correspond to the section of the sample stained with the first stain. The biological object detector subsystem 145 can detect representations of a second particular type of biological object shown in the second digital pathology image, which can correspond to the same section stained with a second stain or a different section of the sample stained with the second stain. Furthermore, the biological object detector subsystem 145 can detect one or more biological objects of one or more biological object types in one or more digital pathology images that are not related to the same sample, for the purpose of generating spatial distribution metrics and object-level results.
[0062] The biological object detector subsystem 145 can detect and characterize biological objects using static rules and / or trained models. Rule-based biological object detection can include detecting one or more edges, identifying a subset of edges that are well connected and closed in shape, and / or detecting one or more high-intensity regions or pixels. For example, a portion of a digital pathology image can be determined to depict a biological object if an area of a region within a closed edge is within a predefined range and / or if a high-intensity region has a size within a predefined range. Detecting biological object depictions using a trained model can include using a neural network, such as a convolutional neural network, a deep convolutional neural network, and / or a graph-based convolutional neural network. The model can be trained using annotated images that include annotations indicating the location and / or boundaries of the object. The annotated images can be received from a data repository (e.g., a public data store) and / or from one or more devices associated with one or more human annotators. The model can be trained using general-purpose or natural images (e.g., not just images captured for digital pathology applications or medical applications overall), which can expand the model's ability to discriminate between different types of biological objects, or it can be trained using a dedicated training set of images, such as digital pathology images, that have been selected to train the model to detect specific types of objects.
[0063] Rule-based biological object detection and trained model biological object detection can be used in any combination. For example, rule-based biological object detection can detect representations of one type of biological object, and the trained model is used to detect representations of another type of biological object. Another example can include using the biological objects output by the trained model to validate the results from the rule-based biological object detection, or using a rule-based strategy to validate the results of the trained model. Yet another example can include using rule-based biological object detection as initial object detection and then using the trained model for more advanced biological object analysis, or applying a rule-based object detection strategy to images after an initial set of representations of biological objects has been detected via the trained network.
[0064] Biological object detection can also include (for example) preprocessing the digital pathology image. Preprocessing can be converting the resolution of the digital pathology image to a target resolution, applying one or more color filters, and / or normalizing and using the digital pathology image with a rule-based biological object detection method or trained model. For example, a color filter can be applied that passes colors corresponding to the color profile of the stains used by the automated staining system 120. Rule-based biological object detection or trained model biological object detection can be applied to the preprocessed image.
[0065] For each detected biological object, the biological object detector subsystem 145 may identify and store a representative location (e.g., centroid or midpoint) of the depicted biological object, a set of pixels or voxels corresponding to an edge of the depicted object, and / or a set of pixels or voxels corresponding to an area of the depicted biological object. This biological object data may be stored along with metadata about the biological object, which may include, by way of example and not limitation, an identifier of the biological object (e.g., an identification number), an identifier of the corresponding digital pathology image, an identifier of the corresponding region within the corresponding digital pathology image, an identifier of the corresponding object, and / or an identifier of the type of object.
[0066] The biological object detector subsystem 145 can generate an annotated digital pathology image that includes the digital pathology image and further includes one or more overlays that identify where the detected biological objects are depicted within the image. In certain embodiments in which multiple types of biological objects are detected, different colors, for example, can be used to represent the different types of annotations.
[0067] The biological object distribution detector subsystem 150 can be configured to generate and / or characterize a spatial distribution of one or more objects. The distribution can be generated (for example) by using one or more static rules (e.g., identifying how to apply a distance-based metric of a point location representation of the biological objects, identifying how to use absolute or smoothed counts or densities of the biological objects within a grid region of the digital pathology image, etc.) and / or by using a trained machine learning model (e.g., capable of predicting that initial object delineation data should be adjusted in light of the predicted quality of one or more digital pathology images). For example, the characterization may indicate the extent to which biological objects of a particular type are depicted as densely clustered relative to one another, the extent to which representations of biological objects of a particular type are spread across all or a portion of the image, how the proximity of representations of biological objects of a particular type compares to the proximity of representations of biological objects of another type, the proximity of representations of one or more biological objects of a particular type to representations of one or more other types of biological objects, and / or the extent to which representations of one or more biological objects of a particular type are within and / or proximate to a region defined by one or more representations of one or more other types of biological objects. As described in more detail below in connection with FIG. 2, the biological object distribution detector subsystem 150 may initially generate representations of the biological objects using a particular framework (e.g., a spatial point process analysis framework, a spatial area analysis framework, or a geostatistical analysis framework, etc.).
[0068] The subject-level label generator subsystem 155 can generate one or more subject-level labels using the spatial distribution metrics. Subject-level labels can include labels determined for individual subjects (e.g., patients), defined groups of subjects (e.g., patients with similar characteristics), clinical study populations, etc. The labels can correspond, for example, to potential diagnoses, prognoses, treatment assessments, treatment recommendations, or treatment eligibility determinations. In certain embodiments, the labels can be generated using predefined or learned rules. For example, the rules can indicate that a spatial distribution metric above a predefined threshold is associated with a particular disease state (e.g., as a potential diagnosis), and a metric below the threshold is not associated with the particular disease state. As another example, the rules can indicate that a particular treatment is recommended when the spatial distribution metric is within a predefined range (and, for example, not recommended otherwise). By way of illustration, checkpoint immunotherapy can be recommended when a distance-based metric (e.g., characterizing how far the centroid of a lymphocyte representation is from the centroid of a tumor cell representation) is below a predefined threshold. As yet another example, the rules may identify different bands of treatment efficacy based on the ratio of a spatial distribution metric corresponding to a recently acquired digital pathology image to a stored baseline spatial distribution metric corresponding to an earlier acquired digital pathology image.
[0069] The object-level label generator subsystem 155 can further generate one or more object-level labels using one or more patterns or masks, for example, in conjunction with a spatial distribution metric. In certain embodiments, the object-level label generator subsystem 155 can obtain or provide one or more patterns or masks associated with previous labels and / or object results (which can help validate the labels). In certain embodiments, the object-level label generator subsystem 155 can obtain the masks according to one or more rules or using a trained model. For example, a rule can indicate that a particular mask or subset of masks is to be obtained and compared to the digital pathology image in response to determining one or more types of one or more biological objects depicted in the digital pathology image. As another example, a rule can indicate that a particular mask or subset of masks is to be obtained and compared to the digital pathology image in response to determining a spatial distribution metric that meets or does not meet a threshold, or that occupies or does not occupy a threshold range. Values associated with the rules can be learned by the object-level label generator subsystem 155. In certain embodiments, a model can be trained using one or more machine learning processes described herein to identify patterns to be acquired and applied to digital pathology images based on a holistic characterization of the digital pathology images, data derived therefrom, and metadata associated therewith.
[0070] The digital pathology imaging system 135 can output the generated spatial distribution metrics, object-level labels, and / or annotated images. Output can include local presentation or transmission (e.g., to the user device 130).
[0071] Each component and / or system in Figure 1 may include (for example) one or more computers, one or more servers, one or more processors, and / or one or more computer-readable media. In certain embodiments, a single computing system (having one or more computers, one or more servers, one or more processors, and / or one or more computer-readable media) may include multiple components shown in Figure 1. For example, digital pathology imaging system 135 may include a single server and / or a collection of servers that collectively implement all of the functionality of compartment aligner subsystem 140, biological object detector subsystem 145, biological object distribution detector subsystem 150, and object-level label generator subsystem 155.
[0072] It will be appreciated that various alternative embodiments are contemplated. For example, digital pathology imaging system 135 may not necessarily include object-level label generator subsystem 155 and / or may not necessarily generate object-level labels. Instead, annotated images (with annotations generated by biological object detector subsystem 145) and / or one or more spatial distribution metrics (generated by biological object distribution detector subsystem 150) may be output by digital pathology imaging system 135. A user may then identify a label (e.g., corresponding to a diagnosis, prognosis, treatment assessment, or treatment recommendation) in light of the output data.
[0073] 2 illustrates an exemplary biological object pattern computation system 200 for processing object data to generate spatial distribution metrics, according to some embodiments of the present invention. The biological object distribution detector subsystem 150 may include some or all of the system 200.
[0074] The biological object pattern computation system 200 includes multiple subsystems, such as a point processing subsystem 205, an area processing subsystem 210, and a geostatistical subsystem 215. Each subsystem corresponds to and uses a different framework to generate spatial distribution metrics or component data of spatial distribution metrics, such as a point process analysis framework 225, an area analysis framework 230, or a geostatistical framework 235. The point process analysis framework 225 can have an object-specific focus, for example, identifying point locations for each detected biological object depiction. The area analysis framework 230 can be a framework in which data (e.g., locations of depicted biological objects) are indexed using coordinates and / or spatial grids rather than by individual biological object depictions. The geostatistical analysis framework 235 can provide predictions of the prevalence and / or probability of observing a particular type of biological object depiction at each set of locations. Each framework can support the generation of one or more metrics that characterize the spatial patterns and / or distributions formed across one or more biological object depictions of one or more respective types.
[0075] For example, the point processing subsystem 205 may employ a point process analysis framework 225 that may represent each biological object representation as a point location within the image. In certain embodiments, the point location may be a centroid, midpoint, center of mass, or the like, of the biological object representation. In some embodiments, the point location is detected (e.g., by the biological object detector subsystem 145) when detecting the biological object representation. In some embodiments, the point processing subsystem 205 determines the location of the biological object representation (e.g., based on a location associated with the edge and / or area of the depicted biological object). The point processing subsystem 205 may include a distance detector 245 that detects and processes one or more distances between the biological object representations, a point-based cluster generator 250 and a correlation detector 255 that characterize cross-correlation and / or auto-correlation between one or more biological object representations of each of one or more types, and a landscape generator 260 that generates a three-dimensional landscape corresponding to the calculated quantities of the biological object representations across a two-dimensional space corresponding to the dimensions of the image (e.g., with the third dimension of the landscape indicating the calculated quantities). Cross-correlation and auto-correlation can identify the probability, as a function of distance, that a point representing a biological object representation of a first type (and thus a biological object in a sample) is located at a distance away from the observed biological object representation. In the case of cross-correlation, the probability is calculated with respect to the biological object of a second type. In the case of auto-correlation, the probability is calculated with respect to the biological object of the first type. Cross-correlation or auto-correlation can include one-dimensional representations (e.g., with the x-axis set to distance) or two-dimensional representations (e.g., with the x-axis set to horizontal distance and the y-axis set to vertical distance).
[0076] The distance detector 245 can detect points within the image and their respective locations. For each pair or pairs of points (e.g., "point pairs"), a distance (e.g., Euclidean distance) between the point locations associated with the pair is calculated. Each of the one or more point pairs can correspond to representations of the same type of biological object or representations of different types of biological object. For example, for a given depicted lymphocyte, the distance detector 245 can identify the distance between the location of the depicted lymphocyte and each other depicted lymphocyte, and the distance detector 245 can identify the distance between the location of the depicted lymphocyte and each depicted tumor cell. The distance detector 245 can generate one or more spatial distribution metrics based on statistics. For example, the spatial distribution metric can be defined as and / or based on the mean, median, and / or standard deviation, etc., of the distances between depictions of a given type of biological object and / or the distances between depictions of one or more different types of biological object. By way of example, the distance between the locations of all depicted lymphocytes can be found and then the average distance can be calculated, and similar calculations can be performed based on the distance between each lymphocyte-tumor cell pair. The spatial distribution metric can be based on a first statistic generated based on the distance between depictions of a first type of biological object and a second statistic generated based on the distance between depictions of a second type of biological object.
[0077] The point-based cluster generator 250 can use the distances to perform a cluster analysis (e.g., a multi-distance spatial cluster analysis such as Ripley's K function). For example, a K value generated using Ripley's K function can represent an estimate of the extent to which the spatial distribution of biological object depictions corresponds to a spatially random distribution (e.g., as opposed to a distribution having one or more spatial clusters).
[0078] The correlation detector 255 can generate one or more correlation-based metrics using the distances and / or point locations. The correlation-based metrics can indicate how well the presence of a biological object representation of a given type at one location predicts whether another biological object representation of the given type or another type is present at another location. The other location can be specified, for example, based on a predefined spatial increment or target region surrounding the biological object representation. For example, the cross-correlation diagram can identify the probability of observing a tumor cell representation within each of various distances from a lymphocyte representation. The metric can identify the sum of probabilities over distances from zero to a specific distance. The correlation-based metric can include a probabilistic dependence coefficient or correlation coefficient. In certain embodiments, the correlation-based metric indicates a distance value associated with a maximum value in the cross-correlation diagram.
[0079] The landscape generator 260 can use the point locations of one or more given type biological object representations to generate a three-dimensional “landscape” data structure (e.g., a landscape map) that indicates the probability of observing a given type of object representation for each horizontal and vertical position in the image. The landscape data structure can be identified by fitting one or more algorithms. For example, a data structure configured to represent zeros can fit one or more Gaussian distributions (or other peak structures). The landscape generator 260 can be configured to compare a landscape data structure generated for a given biological object type with another landscape data structure generated for another biological object type. For example, the landscape generator 260 can compare the position, amplitude, and / or width of one or more peaks in a landscape corresponding to a given biological object type with the position, amplitude, and / or width of one or more peaks in another landscape data structure corresponding to another biological object type. When visualized, the landscape may include a three-dimensional representation in which peaks represent high probabilities that objects of a given type are located within the corresponding region. The landscape data representation represents object density and / or count through the third dimension, although the same data may alternatively be conveyed using other visualization strategies (e.g., via a heat map). Exemplary landscape data structures generated by landscape generator 260 are shown in FIG. 4 as landscape representations 420a and 420b.
[0080] While the point process analysis framework 225 can index the data by individual depictions of biological objects, the area analysis framework 230 can index the data in a more abstract sense using coordinates and / or spatial grids. The area processing subsystem 210 can apply the area analysis framework 230 to identify densities (or counts) for each set of coordinates and / or regions associated with the image area. Densities can be identified using one or more of a grid-based partitioner 265, a grid-based cluster monitor, and / or a hotspot monitor 275.
[0081] The grid-based partitioner 265 can impart a spatial grid onto the image, the spatial grid including a representation of the location of depicted biological objects on the image. The spatial grid, including a set of rows and a set of columns, can define a set of regions, each region corresponding to a combination of a row and a column. Each row can have a defined height, and each column can have a defined width, such that each region of the spatial grid can have a defined area.
[0082] The grid-based partitioner 265 can determine an intensity metric using the spatial grid and the point locations of the biological object representations. For example, for each grid region, the intensity metric can indicate the amount of each of one or more types of biological object representations having point locations within the region (e.g., for at least a threshold portion of the biological object representations) and / or can be based on the amount of biological object representations. In certain embodiments, the intensity metric can be normalized and / or weighted based on the total number of biological objects (e.g., of a given type) detected in the digital pathology image and / or for the sample, the count of biological objects of a given type detected in other samples, and / or the scale of the digital pathology image. In certain embodiments, the intensity metric is smoothed and / or otherwise transformed. For example, the initial count can be thresholded so that the final intensity metric is binary. For example, the binary metric can include determining whether a grid region is associated with a number of biological object representations that meets a threshold (e.g., whether there were at least five tumor cells assigned to the region). In particular embodiments, the grid-based partitioner 265 can use the area data to generate one or more spatial distribution metrics by (for example) comparing intensity metrics across different types of biological objects.
[0083] The grid-based cluster generator 270 can generate one or more spatial distribution metrics based on the cluster-related data for one or more biological object types. For example, for each of the one or more biological object types, clustering and / or fitting techniques can be applied to determine how spatially clustered representations of biological objects of that type are, for example, with each other and / or with representations of biological objects of another type. Clustering and / or fitting techniques can be further applied to determine how spatially dispersed and / or randomly distributed the representations of the biological objects are. For example, the grid-based cluster generator 270 can determine a Morisita-Horn index and / or a Moran index. For example, a single metric can indicate how spatially clustered representations of one type of biological object are and / or how closely they resemble representations of another type of object.
[0084] The hot spot / cold spot monitor 275 can perform analysis to detect either "hot spot" locations where representations of one or more particular types of biological objects are likely to be present, or either "cold spot" locations where representations of one or more particular types of biological objects are likely to be absent. In certain embodiments, a gridded intensity metric can be used to identify (for example) local intensity extremes (e.g., maxima or minima) and / or fit one or more peaks that can be characterized as hot spots or one or more valleys that can be characterized as cold spots. In certain embodiments, a Getis-Ord hot spot algorithm can be used to identify either hot spots (e.g., intensities over a set of adjacent pixels that are sufficiently high to be significantly different compared to other intensities in the digital pathology image) or cold spots (e.g., intensities over a set of adjacent pixels that are sufficiently low to be significantly different compared to other intensities in the digital pathology image). In certain embodiments, "significantly different" can correspond to a determination of statistical significance. Once object type-specific hot spots and cold spots are identified, the hot spot / cold spot monitor 275 can compare the location, amplitude, and / or width of any hot spots or cold spots detected for one biological object type with the location, amplitude, and / or width of any hot spots / cold spots detected for another biological object type.
[0085] The geostatistical subsystem 215 can use the geostatistical analysis framework 235 to estimate an underlying smoothed distribution based on discrete samples. The geostatistical analysis framework 235 can be configured to convert data corresponding to a first dimensionality and / or resolution to a second dimensionality and / or resolution. For example, the locations of biological object representations can be first defined using a 1 mm resolution across the entire digital pathology image. The location data can then be fitted to a continuous function that is not constrained to mm resolution. As another example, the locations of biological object representations, initially defined as two-dimensional coordinates, can be transformed to create a data structure that includes counts of biological object representations in each set of row and column combinations. The geostatistical analysis framework 235 can be configured to fit a function using multiple data points that identify (for example) the locations of specific biological object representations (of a given type). For example, for each specific type of biological object, a variogram can be generated that indicates, for each set of distances, whether two biological objects of the same type, separated by a distance, were detected. A single type of object may be more likely to be detected at short separation distances compared to longer distances. A semivariogram can then be generated by fitting the variogram data. The observed biological objects and semivariogram can then be used by the geostatistical subsystem 215 to generate an image map that predicts the prevalence and / or probability of observing a particular type of biological object depiction at each of a set of locations. The resolution and / or size of the image map can be higher resolution and / or larger size, respectively, compared to the one or more digital pathology images that were originally processed to detect the biological object depictions.The geostatistical subsystem 215 can generate one or more spatial distribution metrics using the geostatistical data by (for example) comparing predicted biological object values (e.g., predicting prevalence and / or probability of observation) across different types of biological objects, characterizing spatial correlation of predicted biological object values between different types of biological objects, characterizing spatial autocorrelation using predicted object values for individual types of biological objects, and / or comparing the locations of spatial clusters (or hot spots / cold spots) of predicted object values across different types of objects.
[0086] It will be appreciated that various subsystems may include components not shown and may perform processing not explicitly described. For example, the area processing subsystem 210 may generate a spatial distribution metric corresponding to an entropy-based mutual information index to indicate the extent to which information about the location of a representation of a first type of biological object within a given region reduces uncertainty about whether a representation of another biological object (of the same or other type) is present at a location within another region. For example, the mutual information metric may indicate that the location of one biological object type provides information about the location of another biological object type (and thus reduces entropy). Such mutual information may potentially be associated with instances in which cells of one cell type are interspersed with cells of another cell type (e.g., tumor-infiltrating lymphocytes are interspersed among tumor cells).
[0087] As another example, the point processing subsystem 205 can generate a neighborhood distance metric based on the distance (or distance statistics) between individual biological object detection points of a given biological object type and one or more other closest points corresponding to biological object depictions of the same biological object type and / or another biological object type. To illustrate, for each depiction of a biological object, the intra-object-type distance value can refer to the average distance between the location of the depiction of the biological object and the locations of the nearest number of depictions of biological objects of the same type. The intra-object-type distance statistics for a biological object type can refer (for example) to the average or median of the intra-object-type distance values for all biological object depictions of the object type. The intra-object-type distance value can refer to the average distance between the location of the depiction of the biological object and the locations of the nearest number of depictions of objects of different types. The intra-object-type distance statistics can be (for example) the average or median of the intra-object-type distance values. A small / low intra-object-type distance statistic can indicate that different types of biological object depictions are close to each other. The intra-object-type distance statistics can be used (for example) for normalization purposes or to assess the overall clustering of biological objects of a given type.
[0088] As yet another example, the point processing subsystem 205 can generate correlation-based metrics based on cross- and / or auto-correlation functions, such as pairwise (cross-type) or mark correlation functions. The correlation function can include (for example) a correlation value as a function of distance. The baseline correlation value can correspond to a random distribution. The metric can include a spatial distance at which the correlation function (or a smoothed version of the correlation function) intersects the baseline correlation value (or some adjustment to the baseline correlation value, such as a threshold calculated by adding a fixed amount to the baseline correlation value and / or multiplying the baseline correlation value by a predefined factor).
[0089] The biological object pattern computation system 200 can generate results (which may themselves be spatial distribution metrics) using a combination of multiple (e.g., two or more, three or more, four or more, or five or more) spatial distribution metrics (e.g., such as those disclosed herein) of various types. The multiple spatial distribution metrics can include metrics generated using different frameworks (e.g., two or more, three or more, or all of the point process analysis framework 225, the area analysis framework 230, and the geostatistics framework 235) and / or metrics generated by different subsystems (e.g., two or more, three or more, or all of the point processing subsystem 205, the area processing subsystem 210, and the geostatistics subsystem). For example, a spatial distribution metric can be generated using a distance-based metric (generated using a spatial point process analysis framework) and a Morisita-Horn index metric (generated using a spatial area analysis framework).
[0090] In particular embodiments, multiple metrics can be combined using one or more user-defined and / or predefined rules and / or using a trained model. For example, the machine learning (ML) model controller 295 can train a machine learning model to learn one or more parameters (e.g., weights) that specify how various low-level metrics are collectively processed to generate an integrated spatial distribution metric. The integrated spatial distribution metric may be more accurate in aggregate than the individual parameters alone. The architecture of the machine learning model can be stored in the ML model architecture data store 296. For example, the machine learning model can include logistic regression, linear regression, decision tree, random forest, support vector machine, or neural network (e.g., forward propagation neural network), and the ML model architecture data store 296 can store one or more equations that define the model. Optionally, the ML model hyperparameter data store 297 stores one or more hyperparameters used to define the model and / or its training but that have not yet been learned. For example, the hyperparameters can identify the number of hidden layers, dropout, learning rate, etc. The learned parameters (e.g., corresponding to one or more weights, thresholds, coefficients, etc.) can be stored in the ML model parameters data store 298.
[0091] In particular embodiments, some or all of one or more subsystems are trained using some or all of the same set of training data used to train the ML model (thereby learning the ML model parameters stored in ML model parameter data store 298). In particular embodiments, one or more subsystems are trained using a different training dataset compared to the ML model controlled by ML model controller 295. Similarly, when multiple frameworks, subsystems, and / or subsystem components are used to generate metrics that are combined to create spatially distributed metrics, individual frameworks, subsystems, and / or subsystem components can be trained using training datasets that are non-overlapping, partially overlapping, fully overlapping, or the same as other training datasets.
[0092] 2, the biological object pattern computation system 200 may further include one or more components that aggregate spatial distribution metrics across plots of the sample of interest to generate one or more aggregated spatial distribution metrics. Such aggregated metrics may be generated (for example) by a component within a subsystem (e.g., by the hotspot monitor 275), by a subsystem (e.g., by the point processing subsystem 205), by the ML model controller 295, and / or by the biological object pattern computation system 200. The aggregated spatial distribution metrics may include (for example) the sum, median, mean, maximum, or minimum of a set of plot-specific metrics.
[0093] 3A and 3B illustrate processes 300a and 300b for providing a health-related assessment based on image processing of a digital pathology image using spatial distribution metrics, according to some embodiments. More specifically, the digital pathology image can be processed, for example, by a digital pathology imaging system, to generate one or more metrics that characterize the spatial pattern and / or distribution of one or more cell types, which can then inform a diagnosis, prognosis, treatment assessment, or treatment eligibility determination. The process begins at step 310, where an identifier associated with a subject can be obtained by the digital pathology imaging system (e.g., digital pathology imaging system 135). The identifier associated with the subject can include identifiers for the subject, sample, section, and / or digital pathology image. The identifier associated with the subject can be provided by a user (e.g., a healthcare provider and / or the subject's physician). For example, the user can provide the identifier as input to a user device, and the user device can transmit the identifier to digital pathology imaging system 135.
[0094] In step 315, the digital pathology image processing system 135 can access one or more digital pathology images of the stained tissue sample associated with the identifier. For example, the identifier can be used to query a local or remote data store. As another example, a request including the identifier can be sent to another system (e.g., a digital pathology image generation system), and the response can include an image. The image can depict stained sections of the sample from the subject. In certain embodiments, a first digital pathology image depicts a section stained with a first stain, and a second digital pathology image depicts a section stained with a second stain. In certain embodiments, a single digital pathology image depicts sections stained with multiple stains. In certain embodiments, the digital pathology image can be separated into regions or tiles before or during the analysis process 300a. The separation can be based on user-directed focus on specific regions, detected regions of interest (e.g., detected according to rules based on machine learning schemes, etc.).
[0095] At step 320, a first set of representations of a first type of biological object and a second set of representations of a second type of biological object may be detected from the digital pathology image. In certain embodiments, the first type of object may correspond to a biological object associated with a first stain, and the second type of object may correspond to a biological object associated with a second stain. The first type of object may correspond to a first type of biological object (e.g., a first cell type), and the second type of object may correspond to a second type of biological object (e.g., a second cell type).
[0096] Each biological object can be associated with location metadata that indicates where the object is depicted within the digital pathology image. The location metadata can include (for example) a set of coordinates corresponding to a point within the image, coordinates corresponding to an edge or boundary of the biological object depiction, and / or coordinates corresponding to the area of the depicted object. For example, a detected biological object depiction can correspond to a 5x5 square of pixels within the image under analysis. The location metadata can identify all 25 pixels of the biological object depiction, 16 pixels along the boundary, or a single representative point. The single representative point can be (for example) the midpoint or can be generated computationally by pre-weighting each of the 25 pixels using intensity values and then calculating a weighted center point. Other weighting indices can also be applied, including weighting indices that are content- or context-dependent.
[0097] In step 325, a data structure is generated based on the biological object representations detected in step 320. The data structure may include object information characterizing the biological object representations. For each detected biological object representation, the data structure may identify, for example, the center of gravity of the biological object representation, pixels corresponding to the perimeter of the biological object representation, or pixels corresponding to the area of the biological object representation. The data structure may further identify, for each biological object representation, a type of biological object (e.g., lymphocyte, tumor cell, etc.) corresponding to the depicted biological object.
[0098] At step 330, one or more spatial distribution metrics are generated. The spatial distribution metrics characterize the relative positions of the biological object representations. In some cases, step 330 may include generating a spatial distribution metric based on the detected biological object representations and object types of example step 320. For example, the spatial distribution metric may characterize how close and / or clustered representations of objects of a particular type are relative to each other and / or to representations of other objects of a particular type.
[0099] At step 335, the spatial distribution metrics generated at step 330 are output to a storage entity / database, a user interface, or a service platform. The service platform can use the output spatial distribution metrics to provide further analysis. The spatial distribution metrics can be transmitted to a user device (where the metrics can be presented to the user) and / or presented locally via a user interface. In certain embodiments, images and / or annotations corresponding to the detected biological object depictions are further output (e.g., transmitted and / or output).
[0100] In certain embodiments, a user can use spatial distribution metrics to inform a diagnosis, prognosis, treatment recommendation determination, or treatment eligibility determination for a subject. For example, if spatial distribution metrics indicate that lymphocytes are close to and / or co-localized with tumor cells, immunotherapy and / or checkpoint immunotherapy can be identified as a recommended treatment. If (for example) a metric representing the distance between lymphocytes and tumor cells is similar (e.g., less than 300%, less than 200%, less than 150%, or less than 110%) to a metric representing the distance between the same cell type (e.g., lymphocytes or tumor cells), lymphocytes can be determined to be close to or interspersed with tumor cells. If intensity values representing the amount of each cell type assigned to individual regions within the image are similar, lymphocytes can be determined to be close to and / or interspersed with tumor cells. For example, the analysis can determine whether intensity values indicate that cell types are densely located within the same or similar subsets of image regions.
[0101] A user can provide a diagnosis, prognosis, etc. to a subject. For example, the diagnosis, prognosis, etc. can be communicated to the subject verbally and / or transmitted from the user's device to the subject's device (e.g., via a secure portal). The user can further use the user device to update the subject's electronic medical record to include the diagnosis, prognosis, etc.
[0102] As a result of the recommendation, a subject's treatment can be initiated, modified, or discontinued. For example, in response to a diagnosis of a subject with a particular disease, a recommended treatment can be initiated and / or an approved treatment for a particular disease can be initiated.
[0103] 3B illustrates another process 300b for providing a health-related assessment based on image processing of a digital pathology image using a spatial distribution metric, according to some embodiments. Steps 305-330 of process 300b are generally similar to steps 305-330 of process 300a. However, in certain embodiments, the digital pathology image processing system 135 can use the spatial distribution metric to predict (e.g., at step 347) a diagnosis, prognosis, treatment recommendation, or treatment eligibility determination for the subject. The prediction can be generated using one or more rules that identify one or more thresholds and / or ranges for the metric. The prediction can include a result that represents a diagnosis, prognosis, or treatment recommendation. The result can be (for example) a binary value (e.g., predicting whether a subject has a particular condition), a categorical value (e.g., predicting a tumor stage or identifying a particular treatment among a set of potential treatments), or a numeric value (e.g., identifying the probability that a subject has a given symptom, predicting the probability that a given treatment will slow disease progression, and / or predicting the time until symptoms progress to the next stage). Treatment recommendations can include using checkpoint inhibitor therapy or immunotherapy (e.g., if metrics indicate that tumor cells are scattered in lymphocytes).
[0104] The results can be generated by a trained machine learning model, such as, by way of example and not limitation, a trained regression, decision tree, or neural network model. In particular embodiments, the spatial distribution metrics include multiple different types of metrics, and the model is configured to process multiple types of data. For example, the set of metric types can include metrics defined based on K-nearest neighbor analysis, metrics defined based on Ripley's K-function, Morisita-Horn index, Moran index, metrics defined based on correlation functions, metrics defined based on hotspot analysis, and metrics defined based on Kriging interpolation (e.g., regular Kriging or directed Kriging), and the results can be generated based on at least two, at least three, or at least four metrics of the set of metric types.
[0105] At step 348, the digital pathology imaging system 135 may output the prediction (which may include outputting the results) to a storage entity / database, a user interface, or a service platform. For example, the prediction may be presented locally and / or sent to a user device (e.g., where the prediction may be displayed or otherwise presented). The digital pathology imaging system 135 may further output (and the user may further receive) annotation data identifying spatial distribution metrics, digital images, and / or detected biological object depictions.
[0106] The user can then identify a confirmed diagnosis, prognosis, treatment recommendation, or treatment eligibility determination. The confirmed diagnosis, prognosis, etc. can be consistent with and / or correspond to the predicted diagnosis, prognosis, etc. The prediction (and / or other data) generated by the digital pathology imaging system can inform the user's decision regarding which diagnosis, prognosis, or treatment recommendation was identified. In certain embodiments, the user can provide feedback to the digital pathology imaging system indicating whether the user-identified diagnosis, prognosis, or treatment recommendation is consistent with the prediction. Such feedback can be used to train models and / or update rules relating spatial distribution metrics to predicted outputs.
[0107] Figure 4 illustrates various stages in identifying spatial patterns and distribution metrics. Figures 4A, 4B, and 4C show enlarged versions of the images in Figure 4. For example, Figure 4 illustrates an initial digital pathology image, the results of detecting biological object features from the received image, point process analysis of the image based on the detected biological object features, and a spatial distribution (depicted as landmark assessment) showing the location / intensity of the biological object features detected in the received image. The spatial distribution is depicted as landmark assessment, and the detected objects are lymphocytes and tumor cells.
[0108] FIG. 4 shows a digital pathology image 405 of exemplary stained sections of a subject's tissue biopsy. The tissue biopsy was collected, fixed, embedded, and sectioned. Each section can be stained with H&E stain and imaged. Hematoxylin in the stain can stain specific cellular structures (e.g., cell nuclei) a first color, while eosin in the stain stained the extracellular matrix and cytoplasm pink. The digital pathology image 405 was processed (using a deep neural network) to detect depictions of two types of objects: lymphocytes and tumor cells. The object data was processed according to various image processing frameworks and techniques (as described below) to generate spatial distribution metrics (as described below).
[0109] Some embodiments include new and modified frameworks and metrics, as well as new applications of the frameworks and metrics for processing digital pathology images.
[0110] Table 410 shown in FIG. 4 includes example biological object data identifying, for each of a plurality of biological object representations, an object identifier associated with the biological object, the type of stain used to stain the sample prior to imaging, the type of biological object (e.g., lymphocyte or tumor cell), and the coordinates of the center of the biological object representation in the digital pathology image. Table 410 was created using an object detector (e.g., biological object detector subsystem 145) to identify a single point location for each biological object representation. The single point location was defined to be the centroid point for the biological object representation. A point process analysis framework was implemented based on table 410.
[0111] Lymphocyte point image 415a depicts lymphocyte point representations 417a in tumor cell coordinates for all detected lymphocyte representations. Tumor cell point image 415b depicts point representations 417b in point coordinates for all detected tumor cell representations.
[0112] The exemplary landscape representations 420a and 420b graphically illustrate three-dimensional landscape data for feature types of biological objects, in this case lymphocyte and tumor cell feature types, respectively.
[0113] Three-dimensional landscape data for landscape representations 420a and 420b can be generated using point data for each of two types of biological objects (e.g., as shown in table 410). The x-axis and y-axis of landscape representation 420a can correspond to the x-axis and y-axis of image 405 and lymphocyte point image 415a (for example). In certain embodiments, the x-axis and y-axis of landscape representation 420b can correspond to the x-axis and y-axis of digital image 405 and tumor cell point image 415b. The landscape data can further include z-values, which characterize the calculated amount of a given type of biological object representation detected within the area corresponding to the (x,y) coordinates. Each (x,y) coordinate pair in the landscape data corresponds to a range of x-values and a range of y-values. Thus, a z-value can be determined based on the number of biological object representations of a given type located over an area defined by an x-value range (corresponding to a portion of the overall width of the landscape) and a y-value range (corresponding to a portion of the overall length of the landscape).
[0114] The three-dimensional representation facilitates determining how the density of representations of one type of biological object compares to the density of representations of another type of biological object in a given portion of the image, in that the heights of the peaks can be visually compared. Landscape data can be generated for each of one or more types of biological object, such as lymphocytes and tumor cells. Thus, a peak in the lymphocyte landscape data can indicate a high number of lymphocytes in the region of the digital pathology image corresponding to the location of the peak, and a peak in the tumor cell landscape data can indicate a high number of tumor cells in the region of the digital pathology image corresponding to the location of the peak. Observing a peak of a first biological object type compared to a peak of a second biological object type can indicate a relationship between the biological object types and / or their representations. For example, observing a tumor cell landscape peak in a region corresponding to a lymphocyte peak can indicate that tumor cells are interspersed with lymphocytes. For example, peak 425a in landscape representation 420a can correspond to peak 425b in landscape representation 420b, and peak 430a can correspond to peak 430b. The peaks in landscape representation 420a and landscape representation 420b are at approximately the same locations, thus indicating interspersion between biological object types. A comparison of the peaks indicates less interspersion at the locations of peaks 425a and 425b compared to the interspersion at the locations of peaks 430a and 430b. In some cases, the digital pathology locations corresponding to the locations of peaks 430a and 430b may be of interest, and instructions may be generated to collect more digital pathology image data or additional biological samples corresponding to that image location.
[0115] Ripley's K-function can be used as an estimator to detect deviations from spatial homogeneity in a set of points (e.g., points corresponding to image locations representing points of a biological object depiction) and can be used to assess the amount of spatial clustering or dispersion at many distance scales. The K-function (or more specifically, a sample-based estimate thereof) can be defined as follows: JPEG0007803880000001.jpg15170, d ij refers to the pairwise Euclidean distance between the i-th and j-th biological object representations among a total of n biological object representations, r is the search radius, λ is the average density of the biological object representations (e.g., n / A, where A is the area of the tissue that encompasses all biological object representations), and I(·) is the distance between the i-th and j-th biological object representations, d ij An indicator function, w, that is 1 if ≦r ij is an edge correction function that avoids biased estimation due to edge effects.
[0116] To design efficient machine learning schemes, we can summarize the entire K-function by formulating the following metric: 1. Area under the curve: r, the maximum clinically meaningful value of the biological object distance r max is identified, and the observed K function and 0≦r≦r max One can calculate the area between a theoretical function (eg, based on the null hypothesis that biological objects of the same or different types are spatially independent) for 2. Observed Ripley's K function and r=r max Point estimate of the difference between the theoretical Ripley's K function in The above features can be derived separately for a first type of biological object and a second type of biological object (e.g., tumor cells and lymphocytes). In addition, a cross-type Ripley's K function can be derived in a similar manner. The Ripley's K function can be used to estimate and output the degree of spatial clustering or dispersion of the biological objects to provide an understanding of this clustering within a representation of the biological objects (e.g., indicating the infiltration or separation of the first type of biological object from the second type of biological object).
[0117] To identify a neighborhood metric, distances between the locations of various pairs of detected biological object representations can be determined. Each distance can be calculated for each pair of biological object representations of different types (e.g., between each tumor cell / lymphocyte pair). For a given biological object representation (e.g., a representation of an individual lymphocyte), a subset of neighboring object representations can be defined as those identified as being of a given type and depicted as closest to the given biological object representation. For example, for a given lymphocyte, the neighborhood subset can identify n tumor cells depicted closest to the given lymphocyte relative to other tumor cells depicted in the image, where n can be a programmable, user-specified, or machine-learned value. For each subset, a center of gravity of the locations at the biological object representation locations of the subset can be calculated. A neighborhood distance metric between the center of gravity of the given biological object representation and the location can be determined therefrom.
[0118] 5A and 5B show two exemplary neighborhood subsets. The locations of the exemplary biological object representations are represented in FIGS. 5A and 5B by open circle data points. For each biological object representation (e.g., lymphocytes), one or more neighboring biological object representations of a second type (e.g., a predefined number of neighboring tumor biological object representations) can be identified. In the illustrated example, five other neighboring biological object representations were identified. The locations of these neighborhoods are represented in FIGS. 5A and 5B by closed circle data points. A neighborhood centroid can be calculated for the neighboring locations. The midpoint can be calculated, for example, as the mean, median, weighted mean, center of mass, etc., for the neighboring locations. In the illustrated example, the centroid location is represented by the location of the end of the line extending from the open circle. The neighborhood distance metric between the exemplary biological object locations and the centroid is represented in FIGS. 5A-5B by the line extending from the open circle.
[0119] Thus, for a given biological object, a neighborhood distance metric can be calculated for a neighborhood subset of biological objects of a second type. The distance metric can be used to classify the biological objects. As an example, if the first biological object is a lymphocyte and the neighboring biological objects are tumor cells, the classification can be tumor margin lymphocytes or intratumoral lymphocytes. The classification can be based on a learned or rule-based evaluation of neighborhood distances. For example, a lymphocyte can be classified as a tumor margin lymphocyte if the distance metric exceeds a threshold, and as an intratumoral lymphocyte if the distance metric does not exceed the threshold. The threshold can be fixed or specified based on a distance metric associated with one or more digital pathology images. In certain embodiments, the threshold can be calculated by fitting a two-component Gaussian mixture model to the distance metric associated with all biological objects depicted in the digital pathology images. Figure 5C shows an example characterization of biological objects according to this discriminant analysis and depending on the context of the process (e.g., biological object representation identity, number of biological object representations, biological object representation type identity, number of biological object representation types, absolute and relative values of neighborhood distances, etc.). In the example shown in Figure 5C, black dots represent tumor cell representations. Blue dots represent lymphocyte representations classified as intratumoral lymphocytes. Green dots represent lymphocyte representations classified as tumor margin lymphocytes.
[0120] The cross-type pair correlation function (PCF cross) is another statistical measure of spatial dependence between points in a spatial point process (e.g., points corresponding to image locations representing points of a biological object representation). In certain embodiments, the PCF cross function can quantify how a first type of biological object representation (e.g., lymphocytes) is surrounded by a second type of biological object representation (e.g., tumor cells). The PCF cross can be expressed as: JPEG0007803880000002.jpg15170In formula, λ, ω ij , and d ijis defined similarly to Ripley's K function, and k h (·) is the smoothing kernel with smoothing bandwidth h>0.
[0121] The overall PCF cross can be summarized by formulating the following metrics: 1. Area under the curve: r, the maximum clinically meaningful value of the biological object distance r max can be selected, and the observed PCF crossover and 0 ≤ r ≤ r max The area between the theoretical PCF cross (based on the null hypothesis that biological objects of the same or different types are spatially independent) for 2. Observed PCF cross and r = r max Point estimate of the difference between the theoretical PCF crosses in.
[0122] The mark correlation function (MCF) facilitates determining whether the location of a biological object representation is substantially similar to that expected with respect to the locations of nearby (e.g., different types of) biological object representations, or whether the location is independent (e.g., random) of the second type of biological object representation. In other words, whether the location and presence of the second type of biological object representation influences the location and presence of the first type of biological object representation. The mark correlation function can be defined as follows: JPEG0007803880000003.jpg21170In formula, JPEG0007803880000004.jpg9170 is a set of digital pathology image positions S separated by a distance r. i and S j We present an empirical conditional expectation, assuming that there is a biological object description in M(s i ),M(s j ) denotes the biological object type associated with these two biological object depictions. In the denominator, M, M' are biological object types depicted randomly and independently of their marginal distributions, and I(m1;m2) is defined as 1 when m1 == m2.
[0123] The entire MCF was summarized by formulating the following metrics: 1. Area under the curve: r, the maximum clinically meaningful value of the biological object distance r max Select the observed MCF and 0 ≤ r ≤ r max The area between the theoretical MCF (based on the null hypothesis that biological objects of the same or different types are spatially independent) and the 2. Observed MCF and r = r max Point estimate of the difference between the theoretical MCF in
[0124] Further evaluation of the biological object depictions can be based on a comparison of the prevalence of one or more types of biological object depictions. For example, the features can be derived from a comparison of the amount of a first type of biological object depiction and a second type of biological object depiction. Furthermore, the features can be improved by comparing biological object depictions having a particular classification (e.g., a first type or a second type).
[0125] For example, categorization of lymphocyte depictions based on statistical analysis of tumor spatial heterogeneity can be characterized by the intratumoral lymphocyte ratio (ITLR), which can characterize lymphocyte depiction location in relation to tumor cell density. In some embodiments, evaluation can be guided by the use of digital pathology image annotations, such as annotation of areas of interest (e.g., tumor area). Within each of these areas, each lymphocyte depiction can be characterized as being a tumor margin lymphocyte or an intratumoral lymphocyte based on a Euclidean distance measure (as described herein). For each lymphocyte depiction, the n nearest tumor cells can be identified (e.g., using a neighborhood technique such as that described in Section VI.A.3), where n is a definable parameter for the number of neighbors used. Second, the centroid coordinates of the convex hull region formed by the n nearest tumor cell depictions can be derived. The distances from each lymphocyte depiction to the nearest tumor cell depiction and to the centroid of the convex hull can then be calculated, and a binary Gaussian mixture model can be fitted to further differentiate lymphocytes into tumor margin lymphocytes or intratumoral lymphocytes. If lymphocytes have infiltrated the tumor core region, the distance to the center of mass should be short. In contrast, if lymphocytes are still migrating to the tumor core region, the distance tends to be long. The ITLR feature was defined as follows: JPEG0007803880000005.jpg13170In formula, N intra-tumor lymphocyte indicates the total number of intratumoral lymphocytes, and N tumor cell denotes the total number of tumor cells. Although described in the context of a specific classification of a particular biological object type, BOR can be extended using similar principles to other biological object delineations with their own context-dependent characterizations.
[0126] The G-cross function calculates a probability distribution of the distance from a biological object representation of a first type to the nearest biological object representation of a second type within any given distance. Specifically, the G-cross function can be considered a spatial distance distribution metric that represents the probability of finding at least one biological object representation (e.g., of a specified type) within an r-radius circle centered on a given point (e.g., a point location representation of a biological object representation in a digital pathology image). These probability distributions can be applied to quantify the relative proximity of any two types of biological object representation. Thus, for example, the G-cross function can be a quantitative surrogate for invasion determination. Mathematically, the G-cross function is expressed as follows: JPEG0007803880000006.jpg13170
[0127] During the ceremony, JPEG0007803880000007.jpg7170 denotes the index of the first type of biological object depiction, and I(·) is the i An indicator function, n, that is 1 if ≦r lym is the total number of biological subjects.
[0128] Similarly, the entire G-cross function can be summarized by formulating the following metric: 1. Area under the curve: r, the maximum clinically meaningful value of the biological object distance r max Select the observed G cross function and 0 ≤ r ≤ r max We calculated the area between the theoretical G cross function (based on the null hypothesis that biological objects of the same or different types are spatially independent) for 2. Observed G cross function and r = r max Point estimate of the difference between the theoretical G cross function in
[0129] 6A-6D show exemplary distance- and intensity-based metrics characterizing the spatial arrangement of biological object representations in exemplary digital pathology images, according to some embodiments. Statistics are shown plotted across a range of r values for each of four types of spatial feature metrics derived based on the digital pathology images. FIG. 6A shows the G-cross function (thin dashed line) for an observed G-cross function calculated from a sample and a theoretical G-cross function (thick dashed line) based on the null hypothesis that the first type of biological object and the second type of biological object are spatially independent. The G-cross function can be calculated as described herein. FIG. 6B shows the difference (solid line) between the K-function calculated for the first type of biological object representation and the K-function calculated for the second type of biological object representation. The K-function was calculated as described herein. Figure 6C shows cross-type pair correlation functions calculated based on the null hypothesis that the first type of biological object and the second type of biological object are spatially independent (dotted line) or by comparing the positions of the first type of depicted biological object and the second type of depicted biological object (solid line). The pair correlations were calculated as described herein. Figure 6D shows mark correlation functions calculated based on the null hypothesis that the first type of biological object and the second type of biological object are spatially independent (dotted line) or by comparing the positions of the first type of depicted biological object and the second type of depicted biological object (solid line). The mark correlations were calculated as described herein.
[0130] The plots in Figures 6A-6D show that, in this example, the first and second types of biological object delineations are spatially correlated based on objective metrics. Additional quantitative features can be derived based on the algorithms disclosed herein.
[0131] Figure 7 illustrates the application of the area analysis framework 230. Figures 7A, 7B, and 7C show enlarged versions of the image in Figure 7. In particular, the area analysis framework 230 was used to process a digital pathology image 405 of a stained specimen section. As described above in connection with the spatial point process analysis framework, depictions of specific types of biological objects (e.g., lymphocytes and tumor cells) were detected. The area analysis framework 230 further generates biological object data, examples of which are shown in table 410.
[0132] A spatial grid having a specified number of columns and rows can be used to divide the digital pathology image 405 into regions. As an example, shown in FIG. 7, the spatial grid was used to divide the digital pathology image 405 into 22 columns and 19 rows. The spatial grid includes 418 regions. Each biological object representation can be assigned to a region. In certain embodiments, the region can be the region that includes the midpoint or other representation point of the biological object representation. For each biological object type and each grid region, the number of biological object representations of the biological object type assigned to the region can be identified. For each biological object type, a collection of region-specific biological object counts can be defined as grid data for the biological object type. FIG. 7 shows certain embodiments of grid data 715a for a first type of biological object representation and grid data 715b for a second type of biological object representation, each overlaid on a representation of the digital pathology image 405 of a stained section. The grid data can be defined to include, for each region of the grid, a prevalence value equal to the coefficient for the region divided by the total counts across all regions. Thus, regions with no biological objects of a given type will have a prevalence value of 0, and regions with at least one biological object of a given type will have a positive non-zero prevalence value.
[0133] The presence of the same amount of biological objects (e.g., lymphocytes) in two different contexts (e.g., tumors) does not imply characterization or the degree of characterization (e.g., the same immune infiltration). Instead, how a first type of biological object representation is distributed relative to a second type of biological object representation may indicate functional status. Thus, characterizing the proximity of biological object representations of the same and different types can reflect more information. The Morisita-Horn index is an ecological indicator of similarity (e.g., overlap) in biological or ecological systems. In certain embodiments, the Morisita-Horn index (MH), which characterizes the bivariate relationship between two populations of biological object representations (e.g., of two types), can be defined as follows: JPEG0007803880000008.jpg13170In formula, JPEG0007803880000009.jpg8170 respectively show the prevalence of a first type of biological object depiction and a second type of biological object depiction in a square grid i. In FIG. 7, grid data 715a shows example prevalence values of a first type of biological object depiction across grid points. JPEG0007803880000010.jpg8170, where grid data 715b shows example prevalence values for a representation of a second type of biological object across grid points. This shows JPEG0007803880000011.jpg8170.
[0134] If an individual grid region does not contain both types of biological object depictions (indicating that the distributions of different biological object types are spatially separated), the Morisita-Horn index is defined as 0. For example, considering the exemplary spatially separate distributions shown in exemplary first grid data 720a, the index would be 0. If the distribution of a first biological object type across a grid region matches (or is a scaled version of) the distribution of a second biological object type across the grid region, the Morisita-Horn index is defined as 1. For example, considering the exemplary largely co-localized distributions shown in exemplary second grid data 720b, the index would be close to 1.
[0135] 7, the Morisita-Horn index calculated using grid data 715a and grid data 715b was 0.47. A high index value indicates that the depictions of the first and second types of biological objects were significantly co-localized.
[0136] The Jaccard index (J) and the Sorensen index (L) are similar and closely related to each other. In certain embodiments, they can be defined as follows: JPEG0007803880000012.jpg26170In formula, JPEG0007803880000013.jpg8170 respectively indicate the prevalence of the first type of biological object depiction and the second type of biological object depiction in the square grid i, and min(a,b) returns the minimum value between a and b.
[0137] In certain embodiments, another metric that can characterize the spatial distribution of biological object representations is the Moran index, which is an index of spatial autocorrelation. Generally, the Moran index statistic is the correlation coefficient for the relationship between a first variable and a second variable in adjacent spatial units. In certain embodiments, to quantify the extent to which two types of biological object representations are scattered in a digital pathology image, the first variable can be defined as the prevalence of a first type of biological object representation, and the second variable can be defined as the prevalence of a second type of biological object representation. In some embodiments, the Moran index I can be defined as follows: JPEG0007803880000014.jpg15170, x i、 y j denotes the standardized prevalence of biological object representations of a first type (e.g., tumor cells) in areal unit i and the standardized prevalence of biological object representations of a second type (e.g., lymphocytes) in areal unit j. ij is a binary weight for areal units i and j, where the weight is 1 if the two units are adjacent and 0 otherwise, and a first-order scheme can be used to define the adjacent structure. Moran's I can be derived separately for different types of biological object delineations.
[0138] As shown in FIG. 8 (and corresponding FIGS. 8A-8C, which show enlarged images of FIG. 8), the Moran index is defined to be equal to −1 when the biological object representations are perfectly dispersed across the grid (and thus have negative spatial autocorrelation, a “colocalization scenario” 820a) and 1 when the biological object representations are tightly clustered (and thus have positive autocorrelation, a “segregation scenario” 820b). When the symmetric distribution corresponds to a random distribution, the Moran index is defined to be 0. Thus, an area representation of a particular biological object representation type facilitates generating a grid that supports the calculation of the Moran index for each biological object type.
[0139] The Moran's index calculated using the lattice data 715a was 0.50. The Moran's index calculated using the lymphocyte lattice data 715b was 0.22. The difference between the Moran's index calculated for each of the two types of biological object representations can provide an indication of co-location (e.g., having a difference near zero indicating co-location).
[0140] Geary's C, also known as Geary's contact ratio, is a measure of spatial autocorrelation, or an attempt to determine whether adjacent observations of the same phenomenon are correlated. Geary's C is inversely proportional to Moran's I, but is not identical. Moran's I is a measure of global spatial autocorrelation, while Geary's C is more affected by local spatial autocorrelation. JPEG0007803880000015.jpg15170, z i denotes the prevalence of either the first or second type of biological object depiction in square grid i, and ω ij is the same as defined above.
[0141] In certain embodiments, the grid data 715a and grid data 715b can be further processed to generate hotspot data 915a corresponding to detected depictions of a first type of biological object and hotspot data 915b corresponding to detected depictions of a second type of biological object, respectively. In FIG. 9 (and corresponding FIGS. 9A-9C, which show enlarged images of FIG. 9), the hotspot data 915a and hotspot data 915b indicate regions determined to be hotspots for each type of detected depiction of a biological object. Regions detected as hotspots are indicated with red symbols, and regions determined not to be hotspots are indicated with black symbols. The hotspot data 915a, 915b was defined for each region associated with a non-zero object count. The hotspot data 915a, 915b can also include a binary value indicating whether a given region was identified as a hotspot. In addition to the hotspot data and analysis, coldspot data and analysis can be performed.
[0142] For biological object depictions, hotspot data 915a, 915b can be generated for each biological object type by determining a Getis-Ord local statistic for each region associated with a non-zero object coefficient for that biological object type. Getis-Ord hotspot / coldspot analysis can be used to identify statistically significant hotspots / coldspots of tumor cells or lymphocytes, where a hotspot is an areal unit with a statistically significant higher value of the prevalence of the biological object depiction compared to adjacent areal units, and a cold spot is an areal unit with a statistically significant lower value of the prevalence of the biological object depiction compared to adjacent areal units. The values and determination of hotspot / coldspot regions relative to adjacent regions can be selected according to user preference, and in certain embodiments, according to a rule-based strategy or trained model. For example, the number and / or type of biological object depictions detected, the absolute number of depictions, and other factors can be considered. The Getis-Ord local statistic is a z-score, which can be defined for a square grid i as follows: JPEG0007803880000016.jpg23170where i represents an individual region (a particular row and column combination) of the lattice, n is the number of row and column combinations in the lattice (i.e., the number of regions), and ω ij is the spatial weight between i and j, z j is the prevalence of a given type of biological object depiction in the domain, JPEG0007803880000017.jpg8170 is the average subject prevalence of a given type across regions. JPEG0007803880000018.jpg19170
[0143] In certain embodiments, the Getis-Ord local statistics can be converted to binary values by determining whether each statistic exceeds a threshold. For example, the threshold can be set to 0.16. The threshold can be selected according to user preference and, in certain embodiments, can be set according to a rule-based or machine-learned strategy.
[0144] In certain embodiments, a logical AND function can be used to identify regions identified as hotspots for more than one type of biological object representation. For example, colocalization hotspot data 920 shows regions identified as hotspots for two types of biological object representation (denoted by red symbols). A high ratio between the number of regions identified as colocalization hotspots and the number of hotspot regions identified for a given object type (e.g., for tumor cell objects) can indicate that a given type of biological object representation shares spatial characteristics with other object types. On the other hand, a low ratio of zero or near zero can be consistent with spatial separation of different types of biological objects.
[0145] Geostatistics is a corpus of spatial mathematical / statistical methods originally developed to predict probability distributions of spatial stochastic processes for mining operations. Geostatistics is widely applied in diverse disciplines, including petroleum geology, earth and environmental sciences, agriculture, soil science, and environmental exposure assessment. In the field of geostatistics, variograms can be used to describe the spatial continuity of data. To generate features from variogram fits, first, an empirical variogram can be calculated as a discrete function using measures of variability between pairs of points (e.g., representative locations in a biological object delineation) separated by various distances. Second, a theoretical variogram can be fitted to estimate the empirical variogram. In certain embodiments, the Matern function can be used as a theoretical variogram model. Consider a spatial model {Z(s):s∈D}, where Z(s) is the prevalence of tumor cells or lymphocytes at location s, and D denotes the set of sample points s, s, ..., sn. The empirical variogram can be calculated as follows: JPEG0007803880000019.jpg16170
[0146] In the example of Figure 10 (and corresponding Figures 10A and 10B, which show enlarged versions of the image of Figure 10), an empirical variogram was generated based on depictions of biological objects detected in H&E stained image 405 (shown in Figure 10 as points on the theoretical variogram plot). A theoretical variogram 1015 was then generated by fitting a Matern function to the empirical variogram.
[0147] In the above calculation, summation is only for N(h) pairs of observations (e.g., pairs of biological object depictions) separated by a Euclidean distance h. Parameters from the Matern function can be used as features from this method. Features can be obtained separately from variogram fits of detected depictions of a first type of biological object (e.g., tumor cells) and a second type of biological object (e.g., lymphocytes). Alternatively, when combining detected depictions of biological objects across types, indicator variogram fits can also be performed.
[0148] The variogram and point locations of the detected depictions of the biological object estimator can then be used to generate, for each region (e.g., pixel) in the digital pathology image 405, a probability that a particular type of biological object is depicted in the region. The Kriging map 1020 shown in Figure 10 illustrates, for each of a plurality of regions in the digital pathology image 405, the probability that a particular type of biological object (e.g., tumor cell) is depicted in the region.
[0149] In certain embodiments, a regression machine learning model can be trained to process, for example, digital pathology images of biopsy sections from a subject to predict a subject's symptom assessment from the digital pathology images. As an example, based on digital pathology images of biopsy sections from a subject diagnosed with colorectal cancer, a regression machine learning model can be trained to predict whether the cancer exhibits microsatellite stability in tumor DNA (versus microsatellite instability in tumor DNA). Microsatellite instability can be associated with a relatively high number of mutations in microsatellites.
[0150] Biopsies can be collected from each of a plurality of subjects having a condition, in this example, colon cancer. The samples can be fixed, embedded, sliced, stained, and imaged according to the subject matter disclosed herein. Depictions of specified types of biological objects, e.g., tumor cells and lymphocytes, can be detected using, for example, the biological object detector subsystem 145. In certain embodiments, the biological object detector subsystem 145 can recognize and identify biological object depictions using a trained deep convolutional neural network. For each subject of the plurality of subjects, a label can be generated to indicate whether the condition (e.g., cancer) exhibited a specified characteristic (e.g., microsatellite stability vs. microsatellite instability). Ground truth labels can be generated based on a bacteriologist's evaluation and assay-based testing results.
[0151] For each subject, an input vector can be defined to include a set of spatially distributed metrics. The set of spatially distributed metrics can include a selection of metrics described herein. By way of example, metrics included in the input vector can include: The area between the observed and theoretical K functions for distances between biological objects from 0 to the maximum observed distance Point estimate of the difference between observed and theoretical Ripley's K function at maximum biological inter-subject distance Area under the curve of the G-cross function for the distance between biological objects from 0 to the maximum observed distance Point estimate of the difference between the observed and theoretical G-cross functions at the maximum biological intersubject distance Area under the curve of the pair correlation function (cross type) for the distance between biological objects from 0 to the maximum observed distance Point estimate of the difference between the observed and theoretical pair correlation functions (cross-type) at the maximum biological inter-object distance Area under the curve of the Mark correlation function (cross type) for distances between biological objects from 0 to the maximum observed distance Point estimate of the difference between the observed and theoretical Mark correlation functions (cross-type) at the maximum inter-biological object distance Intratumoral lymphocyte ratio Morisita-Horn index Jacard index Sorensen index Moran index Geary's C The ratio of co-localized spots (e.g., hot spots, cold spots, non-significant spots) for a type of biological object depiction to the number of spots (e.g., hot spots, cold spots, non-significant spots) for a first type of biological object depiction and spots (e.g., hot spots, cold spots, non-significant spots) defined using Getis-Ord local statistics. Features obtained by variogram fitting of two types of biological object representations (e.g., tumor cells and lymphocytes)
[0152] The selected metric corresponds to multiple frameworks (point process analysis framework, area process analysis framework, and geostatistical framework). In certain embodiments, for each subject, a label can be defined to indicate whether the indicated feature (e.g., microsatellite stability) is observed or not. An L1 regularized logistic regression model can be trained and tested with paired input data and labels using repeated nested 5-fold cross-validation with Lasso. Specifically, for each of the five data splits, a model can be trained on the remaining four splits and tested on the remaining splits to calculate the area under the ROC.
[0153] FIG. 11 shows an exemplary median receiver operating curve (ROC) generated using 5-fold cross-validation. In the described example, the median area under the ROC generated using the validation set was 0.931. The 95% confidence interval was (0.88, 0.96). The variables from the input dataset most frequently selected by the L1 regularized logistic regression model can be identified to show which metrics were considered most predictive of the specified characteristics of the target condition. For example, the most frequently selected metrics were the area under the curve of the pairwise correlation function and the hotspot ratio calculated using the Getis-Ord locality statistic, and it can be shown that these metrics are most predictive of microsatellite instability. Digital pathology image processing can serve as a reliable replacement for certain paid and expensive tests. For example, in the example discussed herein, a digital pathology image processing system can show that processing can mirror or surpass DNA analysis in determining whether a given subject's tumor exhibits microsatellite instability. Thus, using an image-based approach according to the presently disclosed subject matter can eliminate the need to collect an additional biopsy sample from the subject to collect DNA, further saving the time and expense of performing DNA analysis.
[0154] In certain embodiments, digital pathology images of stained biopsy sections are accessed for each of a first subject and a second subject. Representations of a first type of biological object and a second type of biological object (e.g., lymphocytes and tumor cells) can be detected in each image according to techniques described herein. Input vectors can be generated for each subject as described herein. The input vectors can be processed separately by trained logistic regression models as described herein.
[0155] The model outputs a first label in response to processing an input vector associated with a first subject, which may correspond, for example, to a prediction that the first subject's cancer exhibits microsatellite instability.
[0156] The model outputs a second label in response to processing the input vector associated with the second subject, which can correspond, for example, to a prediction that the cancer of the second subject does not exhibit microsatellite stability.
[0157] The first label and the second label can each be processed (separately) according to treatment recommendation rules. The rules can be configured to recommend a particular treatment, e.g., an immunotherapy (or immune checkpoint therapy) treatment, upon detecting a particular characteristic of the target condition, e.g., microsatellite instability, or to recommend not using another treatment, e.g., an immunotherapy (or immune checkpoint therapy) treatment, upon detecting a particular characteristic of the target condition. The results from rule processing can, for example, indicate that an immunotherapy treatment is recommended for a first subject but not for a second subject.
[0158] In certain embodiments, digital pathology images can depict the tumor microenvironment, including the spatial structure of tissue components and their microenvironment interactions, which can be highly influential with respect to tissue formation, homeostasis, regenerative processes, immune responses, and the like.
[0159] Non-small alveolar carcinoma (NSCLC) is a major global health problem and the leading cause of cancer-related mortality worldwide. Despite the availability of a wide range of treatment options, chemotherapy remains the mainstay of treatment for patients with metastatic (EGFR- and ALK-negative / unknown) NSCLC. However, immune checkpoint inhibitors are transforming the treatment algorithm for this subpopulation.
[0160] Spatial statistics (e.g., spatial distribution metrics) can be calculated using digital pathology images to determine how well the statistics predict overall survival for various treatments. Clinical study arms can be established to test the efficacy of various treatments. An exemplary clinical trial was conducted to assess the safety and efficacy of atezolizumab (an engineered anti-programmed death-ligand 1 [PD-L1] antibody) in combination with carboplatin and paclitaxel (e.g., the "ACP arm"), with or without bevacizumab (e.g., the "ABCP arm"), compared with treatment with carboplatin, paclitaxel, and bevacizumab (e.g., the "CPB arm") in chemotherapy-naive participants with stage IV non-squamous NSCLC. Participants were randomized in a 1:1:1 ratio to the ACP arm, the ACPB arm, or the CPB control arm.
[0161] Tissue samples were collected at baseline. Digital pathology (e.g., H&E pathology) images of the baseline tissue samples can be captured for each subject in each treatment group. H&E-stained slides of the tissue samples were scanned and digitized to generate digital pathology images of the type described herein. Regions associated with one or more depictions of biological objects in the digital pathology images (also called whole slide images or "WSIs") were annotated. Representations of specific types of biological objects were detected, including tumor cells, immune cells, and other stromal cells. Position coordinates for each depiction of each type of biological object were generated, for example, according to the subject matter disclosed herein. In one example, the efficacy of different study groups can be investigated while focusing, for example, on lymphocytes and tumor cells, to investigate immune infiltration, tumor resource distribution, and cell-cell interactions.
[0162] For each image, a wide variety of spatial features can be derived based on the detected biological target symptoms and / or their respective associated locations based on spatial statistical (e.g., spatial distribution metrics) algorithms discussed herein, including, for example, spatial point process methods (e.g., Ripley's K-function feature, G-function feature, pair correlation function feature, Mark correlation function feature, and intratumoral lymphocyte ratio), spatial lattice process methods (e.g., Morisita-Horn index, Jaccard index, Sorensen index, Moran's I, Geary's C, and Getis-Ord hotspot), and geostatistical process methods (ordinary Kriging feature, directed Kriging feature).
[0163] Additionally, for clinical research purposes, outcome variables can be identified, such as overall subject survival.
[0164] Generally, the analysis performed in this example was performed to determine whether the difference in overall survival between the ACP and BCP cohorts would be more pronounced if only a portion of each cohort were considered—a selected portion of individuals predicted to survive longer relative to other subjects in the cohort. Predictions can be based on one or more of the spatial distribution metrics discussed herein, for example, generated on digital pathology images of samples from the subjects. In certain embodiments, the first analysis involved comparing the overall survival of the ACP versus BCP intent-to-treat population. The second analysis involved using a model-based predictive enrichment strategy to investigate the correlation between derived spatial features and overall survival (OS). Predictive enrichment for clinical studies, including NSCLC clinical studies, involves identifying a responder subpopulation Ω in the total patient population Ω that has a greater than average response to treatment, as measured, for example, by odds ratio (OR), relative risk (RR), or hazard ratio (HR). Focusing on this subpopulation has the advantage of increasing the efficiency or feasibility of the study and improving the benefit-risk relationship for this subset of subjects compared to the overall population. One possible enhancement strategy is an open-label, single-arm trial prior to randomization. In this design, the investigated treatment is given to all subjects, and responders identified by pre-specified criteria (e.g., study endpoints or biomarkers) are randomized to a placebo-controlled trial.
[0165] Model-based methodologies can be used to address, for example, the problem of predictive enrichment. In particular, when a clinical study has already been conducted, an enrichment model can be developed retrospectively. To retrospectively develop an enrichment model, data can be split into training, validation, and test sets in a 60:20:20 ratio in each group (e.g., according to the subject matter disclosed herein). The training set of a treatment group can be used to simulate, for example, an open-label pre-randomization phase in an empirical design. A Cox model or objective response model with spatial statistical features as input can be fitted with L1 or L2 regularization for the training set of a treatment group, e.g., ACP. The predicted risk score or predicted response probability from the fitted Cox model can be used as the response score S, and the responder criteria can be specified in the form of a subset condition of the following formula: JPEG0007803880000020.jpg11170In formula, S q denotes the q quantile of the response score, and x represents a subject-level covariate characterized by the feature vector. A validation set of combined treated and control patients can be used to simulate the subject group to be recruited before randomization. To implement the subset condition, quantiles can be calculated for each treatment and control group with the same q in the validation set, and the above equations are used to obtain the subsets for the treatment and control in the validation set, respectively. The subsets in this example JPEG0007803880000021.jpg9170 can be estimated by assessing q using a log-rank test for survival data or a permutation test for objective response data, towards the most significant difference between treatment and control, both of which are subsets of the responder subgroup in the validation set. JPEG0007803880000022.jpg8170 can also be estimated using a pre-specified response threshold q. The enhancement condition with JPEG0007803880000023.jpg8170 is JPEG0007803880000024.jpg10170, which can then be assessed using the same method in the test set for hazard ratios or odds ratios.
[0166] In embodiments with limited sample size, nested Monte Carlo cross-validation (nMCCV) can be used to assess model performance, with equal random splits between training, validation, and test sets, and a score function and threshold. The same enrichment procedure can be repeated B times by creating an ensemble of JPEG0007803880000025.jpg9170. For the i subject, the ensembled responder status can be assessed by averaging the responder group membership for i across iterations in which i is randomized to the test set and thresholding at 0.5. Both hazard ratios or odds ratios, with 95% confidence intervals and p-values, can be calculated for the aggregated test subjects.
[0167] The overall workflow of the predictive analysis is summarized in the flowchart in Figure 12. More specifically, to assign a label to each subject in the study cohort, a nested Monte Carlo cross-validation (nMCCV) modeling strategy was used to overcome overfitting.
[0168] Specifically, for each subject, in block 1205, the dataset can be split into training, validation, and test data portions in a 60:20:20 ratio. In block 1210, a 10-fold cross-validated Ridge-Cox (L2 regularized Cox model) can be performed using the training set to generate 10 models (having the same model architecture). Based on the 10-fold training data, a specific model can be selected and stored among the 10 generated models. In block 1215, the specific model can then be applied to the validation set to adjust specified variables. For example, the variable can identify a threshold for the risk score. In block 1220, the threshold and the specific model can then be applied to an independent test set to generate a vote for the subject that predicts whether the subject will be stratified into a long-term or short-term survival group. The data splitting, training, cutoff identification, and vote generation (blocks 1205-1220) can be repeated N (e.g., =1000) times. At block 1225, the subjects are then assigned to one of the long-term or short-term survival groups based on the votes. For example, the step of block 1225 may include assigning subjects to a long-term or short-term survival group by determining which group is associated with a majority of the votes. At block 1230, a survival analysis of the subjects in the long-term / short-term survival groups may then be performed. It will be appreciated that similar procedures for assigning a variety of labels to data based on outcomes of interest may be applied to any suitable clinical assessment or eligibility study.
[0169] In contrast to the primary finding when comparing the intention-to-treat populations for ACP versus BCP, which yielded an overall survival hazard ratio (HR) of 0.85 (95% CI 0.71–1.03), the proposed strategy resulted in a clear separation between the identified groups of ACP and BCP cohorts, in this example HR = 0.64 (95% CI 0.45–0.91; Figure 13). Note that an overall survival hazard ratio of 1.0 indicates that survival is statistically similar across cohorts. Therefore, in this described example, the low hazard ratio secured using the second analysis strategy (during which statistics were calculated for only the portion of the cohort predicted to have long survival based on spatial statistics and / or spatial distribution metrics) suggests that the second analysis was able to identify subjects for whom treatment (ACP treatment) would be effective. Therefore, the use of spatial distribution metrics represents an improvement over previous strategies.
[0170] The comprehensive model based on spatial statistics and spatial distribution metrics used in this example analysis powered an analytical pipeline that generated systems-level knowledge of the spatial heterogeneity of the tumor microenvironment, in this case by modeling histopathology images as spatial data. Results demonstrate that spatial statistics-based methods can stratify subjects who benefit from atezolizumab treatment compared to standard of care. This effect is not limited to the specific treatment assessment considered in this example. Using spatial statistics to characterize histopathology images, as well as other digital pathology images, may be useful in clinical settings to predict treatment outcomes and therefore inform treatment selection.
[0171] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0172] The terms and expressions which have been employed are used as conditions of description rather than of limitation, and no intention is intended in the use of such terms and expressions to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the invention as claimed has been specifically disclosed by embodiments and optional features, it is to be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.
[0173] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.
[0174] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
Claims
1. accessing, by a computing system, a digital pathology image depicting a section of a biological sample collected from a subject having a given pathology; detecting a set of biological object representations within the digital pathology image, the set comprising a first set of biological object representations of a first class of biological object and a second set of biological object representations of a second class of biological object; generating one or more relative position representations of a first biological object representation, each representing a position of the first biological object representation relative to a second biological object representation; using the one or more relative location representations to determine a spatial distribution metric that characterizes the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations; generating a result corresponding to a prediction as to how effectively a given treatment that modulates an immune response will treat the given condition in the subject based on the spatial distribution metric; determining that the subject is eligible for the clinical trial based on the results; generating a display including an indication that the subject is eligible for the clinical trial; A method comprising: the spatial distribution metric is a first spatial distribution metric, the first spatial distribution metric being of a first type of metric; the method further includes using the one or more relational location representations to determine a second spatial distribution metric characterizing the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations, the second spatial distribution metric being of a second type of metric different from the first type of metric; the results are generated further based on the second spatial distribution metric; The method further includes accessing genetic sequencing or radiological imaging data relating to the subject, wherein the results are generated further based on characteristics of the genetic sequencing or radiological imaging data.
2. The spatial distribution metric is a metric defined based on K-nearest neighbor analysis; A metric defined based on Ripley's K function, Morisita-Horn index, Moran index, a metric defined based on a correlation function; Defined metrics based on hot spot / cold spot analysis, or including metrics defined based on Kriging-based analysis; The method of claim 1.
3. 2. The method of claim 1, wherein generating the result comprises processing the first spatial distribution metric and the second spatial distribution metric using a trained machine learning model, the trained machine learning model being trained using a set of training elements, each of the sets of training elements corresponding to a different subject receiving a particular treatment associated with the clinical trial, each of the sets of training elements including a different set of spatial distribution metrics and a responsiveness value indicative of how well a given treatment activated an immune response in the other subjects.
4. The method of claim 1 , wherein generating the result comprises comparing the value of the spatial distribution metric to a threshold.
5. 10. The method of claim 1, wherein the given condition is a type of cancer and the given treatment is an immune checkpoint inhibitor treatment.
6. The method of claim 1 , wherein the one or more relational position representations include, for each biological object representation in the set of biological object representations, a set of coordinates that identify the position of the biological object representation within the digital pathology image.
7. generating the one or more relational positional representations of the biological object depictions, For each biological object depiction of the first set of biological object depictions, identifying a first point location within the digital pathology image that corresponds to the biological object depiction; for each biological object representation of the second set of biological object representations, identifying a second point location within the digital pathology image that corresponds to the biological object representation; comparing the first point location with the second point location; The method of claim 1 , comprising:
8. 8. The method of claim 7, wherein the first point location in the digital pathology image is selected by calculating a mean point location, a centroid point location, a median point location, or a weighted point location for the biological object depiction in the first set of biological object depictions.
9. 8. The method of claim 7, wherein determining the spatial distribution metric includes, for each of at least some of the first set of biological object depictions and each of at least some of the second set of biological object depictions, calculating a distance between the first point location corresponding to the biological object depiction in the first set of biological object depictions and the second point location corresponding to the biological object depiction in the second set of biological object depictions.
10. 8. The method of claim 7, wherein determining the spatial distribution metric further comprises, for each of at least some of the first set of biological object depictions, identifying one or more of the second set of biological object depictions associated with a distance between the first point location corresponding to the biological object depiction in the first set of biological object depictions and the second point location corresponding to the biological object depiction in the second set of biological object depictions.
11. The method of claim 1, wherein the one or more relational location representations include, for each of a set of image regions in the digital pathology image, a representation of the absolute or relative amount of biological object depictions of the first class of biological objects identified as being located within the region, and a representation of the absolute or relative amount of biological object depictions of the second class of biological objects identified as being located within the region.
12. 2. The method of claim 1, wherein the one or more relational location representations include a distance-based probability that a biological object representation from the first set of biological object representations is depicted as being located within a given distance from a biological object representation from the second set of biological object representations.
13. 10. The method of claim 1, wherein the first class of biological objects is tumor cells and the second class of biological objects is immune cells.
14. receiving user input data from a user device including an identifier of the subject, wherein the computing system accesses the digital pathology image in response to receiving the identifier; generating the display including the indication that the subject is eligible for the clinical trial includes providing the indication that the subject is eligible for the clinical trial to the user device. The method of claim 1.
15. 15. The method of claim 14, further comprising receiving an indication that the subject is enrolled in the clinical trial.
16. 10. The method of claim 1, wherein generating the display including the indication that the subject is eligible for the clinical trial comprises notifying the subject of the determination of eligibility for the clinical trial.
17. one or more data processors; a non-transitory computer-readable storage medium communicatively coupled to the one or more data processors and including instructions that, when executed by the one or more data processors, cause the one or more data processors to perform one or more operations, the one or more operations comprising: accessing a digital pathology image depicting a section of a biological sample collected from a subject having a given pathology; detecting a set of biological object representations within the digital pathology image, the set comprising a first set of biological object representations of a first class of biological object and a second set of biological object representations of a second class of biological object; generating one or more relative position representations of a first biological object representation, each representing a position of the first biological object representation relative to a second biological object representation; using the one or more relative location representations to determine a spatial distribution metric that characterizes the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations; generating a result corresponding to a prediction as to how effectively a given treatment that modulates an immune response will treat the given condition in the subject based on the spatial distribution metric; determining that the subject is eligible for the clinical trial based on the results; generating a display including an indication that the subject is eligible for the clinical trial; Including, the spatial distribution metric is a first spatial distribution metric, the first spatial distribution metric being of a first type of metric; the operations further include determining a second spatial distribution metric using the one or more relational location representations to characterize the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations, the second spatial distribution metric being of a second type of metric different from the first type of metric; the results are generated further based on the second spatial distribution metric; The system, wherein the operations further include accessing genetic sequencing or radiological imaging data relating to the subject, and the results are generated further based on characteristics of the genetic sequencing or radiological imaging data.
18. One or more computer-readable non-transitory storage media containing instructions that, when executed by one or more data processors, cause the one or more data processors to perform operations, the operations including: accessing a digital pathology image depicting a section of a biological sample collected from a subject having a given pathology; detecting a set of biological object representations within the digital pathology image, the set comprising a first set of biological object representations of a first class of biological object and a second set of biological object representations of a second class of biological object; generating one or more relative position representations of a first biological object representation, each representing a position of the first biological object representation relative to a second biological object representation; using the one or more relative location representations to determine a spatial distribution metric that characterizes the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations; generating a result corresponding to a prediction as to how effectively a given treatment that modulates an immune response will treat the given condition in the subject based on the spatial distribution metric; determining that the subject is eligible for the clinical trial based on the results; generating a display including an indication that the subject is eligible for the clinical trial; Including, the spatial distribution metric is a first spatial distribution metric, the first spatial distribution metric being of a first type of metric; the operations further include determining a second spatial distribution metric using the one or more relational location representations to characterize the extent to which at least some of the first set of biological object representations are depicted as interspersed with at least some of the second set of biological object representations, the second spatial distribution metric being of a second type of metric different from the first type of metric; the results are generated further based on the second spatial distribution metric; A computer-readable non-transitory storage medium, wherein the operations further include accessing genetic sequencing or radiological imaging data relating to the subject, and the results are generated further based on characteristics of the genetic sequencing or radiological imaging data.
Citation Information
Patent Citations
Scoring of tumor infiltration by lymphocytes
US20170365053A1
Method for scoring pathology images using spatial analysis of tissues
US20180089495A1
Method for scoring pathology images using spatial statistics of cells in tissues
US9865053B1
Method for scoring pathology images using spatial analysis of tissues
WO2019108230A1
Distance-based tissue state determination
WO2020083970A1