Inference of three-dimensional cell positioning and full-cell correction of highly multiplexed imaging data

The method and system use machine learning models to correct biomarker expression and predict biomarker locations in three dimensions from two-dimensional images, addressing spatial heterogeneity issues in biological imaging for accurate tissue and cell analysis.

WO2026090759A1PCT designated stage Publication Date: 2026-05-07LUNENFELD-TANENBAUM RESEARCH INSTITUTE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LUNENFELD-TANENBAUM RESEARCH INSTITUTE
Filing Date
2025-11-03
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing imaging technologies for three-dimensional biological samples are costly, limited in sample processing, and provide inaccurate results due to spatial heterogeneity, leading to erroneous analyses and diagnoses.

Method used

A method and system using computational machine learning models to infer three-dimensional cell data from two-dimensional images by correcting biomarker expression through z-level regression and predicting the presence and location of biomarkers outside the image plane, utilizing semi-supervised and unsupervised learning to account for spatial heterogeneity.

Benefits of technology

Accurately determines the state of individual cells and tissues by predicting biomarker expression and location across the z-axis, providing a more complete analysis of biological samples without the need for extensive sample preparation, thus improving diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025051465_07052026_PF_FP_ABST
    Figure CA2025051465_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A method of determining presence of biomarkers in a sample comprises receiving an image of the sample, the image comprising a plurality of cell-cross segments; dividing the image into a plurality of two-dimensional grid segments, each of the plurality of two-dimensional grid segments comprising a plurality of cell segments; determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of a cell contained within the segment; correcting a segment-quantified biomarker expression for each of the cells contained within each of the plurality of grid segments by regressing out z-level information; and returning a corrected expression approximating a complete biomarker expression for each of the cells contained within each of the plurality of grid segments as if a biomarker expression of an entirety of each of the cells contained within each of the plurality of the grid segments had been measured.
Need to check novelty before this filing date? Find Prior Art

Description

DocketNo.: 064987-501001 WOINFERENCE OF THREE-DIMENSIONAL CELL POSITIONING AND FULL-CELL CORRECTION OF HIGHLY MULTIPLEXED IMAGING DATACROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 715,369, filed November 1, 2024, which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to inference of three- dimensional cell data based on data extracted from highly multiplexed images.BACKGROUND

[0003] The medical field uses various different types of imaging technology across many disciplines. Commonly, biological tissue are prepared for imaging and analysis by sectioning the biological material into extremely thin slices. However, such procedures only measure a small subsample of the tissue or can be very time-intensive and costly, if many sections are produced in the analysis of the sample, and procedures for preparing the sections for imaging can be a long and variable process involving hazardous chemicals.

[0004] Alternatively, some imaging modalities use serial sections or technologies such as confocal microscopy, attempt to “stitch” together multiple adjacent planar digital images of a sample in order to reconstruct a three-dimensional image of the sample. However, many such imaging modalities can be costly, as the imaging equipment is large and expensive, and is limited in what samples can be processed, as it is important to use a tissue through which light can pass in order to obtain the images.

[0005] New methods of evaluating three-dimensional biological samples are needed.SUMMARY

[0006] Systems, methods, and articles of manufacture are provided for inferring three-Docket No.: 064987-501001 WO dimensional cell data based on data extracted from single plane images or spatially resolved molecular measurements.

[0007] According to some aspects, a method for determining whether one or more biomarkers of interest are present in a sample comprises: receiving an image of a portion of the sample, the image comprising a plurality of cell-cross segments; dividing the image into a plurality of two-dimensional grid segments, each of the plurality of two-dimensional grid segments comprising a plurality of cell segments; determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of a cell contained within the segment; correcting a segment-quantified biomarker expression for each of the cells contained within each of the plurality of grid segments by regressing out z-level information; and returning a corrected expression approximating a predicted biomarker expression for each of the plurality of cell segments as if a biomarker expression of an entirety of each of the cells contained within each of the plurality of the grid segments had been measured. In some aspects, the method further comprises determining a location of the one or more biomarkers in the sample. In some aspects, the method further comprises determining that one or more biomarkers are present in the sample outside of a plane of the sample defined by the image. In some aspects, the method further comprises determining a state of the sample based on the presence of the one or more biomarkers in the sample. In some aspects, the method is performed using a trained computational machine learning model. In some aspects, the computational machine learning model is a semi-supervised machine learning model. In some aspects, the method further comprises determining a risk for a subject to develop a condition based on the determination of a whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject. In some aspects, the method further comprises determining whether a subject has a condition based on the determination of whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject. In some aspects, the condition is a cancer, a tumor, a pro- inflammatory condition, or a neurodegenerative disease. In some aspects, a system for determining the presence of one or more biomarkers in a sample comprises a processor and a non-transitory memory, the non-transitory memory comprising instructions that, when executed by the processor, perform any method disclosed herein.

[0008] In some aspects, a system for determining the presence or absence of one or more biomarkers of interest in a sample comprises: a processor; a non-transitory memory comprisingDocket No.: 064987-501001 WO instructions that, when executed by the processor to perform steps comprising: dividing an image of a portion of a sample into a plurality of two-dimensional grid segments, each of the plurality of two-dimensional grid segments comprising a plurality of cell segments; determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of a cell contained within the segment; correcting a segment-quantified biomarker expression for each of the cells contained within each of the plurality of grid segments by regressing out z-level information; and returning a corrected expression approximating a predicted biomarker expression for each of the plurality of cell segments as if a biomarker expression of an entirety of each of the cells contained within each of the plurality of the grid segments had been measured. In some aspects, the steps further comprise determining whether a subject has a condition based on the determination of whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject. In some aspects, the condition is a cancer, a tumor, a pro- inflammatory condition, or a neurodegenerative disease. In some aspects the system further comprises a display for displaying the corrected expression or information reflecting the determination whether a subject has the condition. In some aspects, the non-transitory memory comprises a trained computational machine learning model.

[0009] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes, it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,

[0011] FIG. 1 depicts a flow diagram for the assessment of biomarker expression in aDocket No.: 064987-501001 WO sample comprising a cell in accordance with described embodiments;

[0012] FIG. 2 depicts a visualization of a cell segment in accordance with described embodiments;

[0013] FIGs. 3 A-3C depict an assessment of a cell’s z-level in accordance with described embodiments;

[0014] FIG. 4 depicts an exemplary distribution of z-levels for cells within a given tissue sample in accordance with described embodiments;

[0015] FIG. 5A depicts a visualization of a step for segmenting an image of a cell in accordance with described embodiments;

[0016] FIG. 5B depicts a visualization of protein of interest feature extraction by z-level in accordance with described embodiments;

[0017] FIG. 6 shows method validation data evaluating accuracy of predicted cell feature of interest location compared to actual cell position in accordance with described embodiments;

[0018] FIG. 7 depicts a flow diagram for the assessment of biomarker expression in a sample comprising a biological tissue in accordance with described embodiments;

[0019] FIG. 8 shows method validation data evaluating accuracy of predicted tissue feature of interest location compared to actual location of the feature of interest in a tissue in accordance with described embodiments.

[0020] FIG. 9 is a system diagram illustrating one embodiment of a computing device configured to assess the biomarker expression in a sample comprising a cell.

[0021] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0022] Disclosed herein are systems and methods for determining a spatial relationship of a first biological feature to a second biological feature in a three-dimensional space (e.g., a three-dimensional portion of biological tissue sample) based partially or entirely on a two- dimensional, image-based input, wherein the image-based input represents a portion of a biologicalDocketNo.: 064987-501001 WO sample. As described herein, such spatial relationships of features, which can be or can represent structures or molecules of a biological cell or tissue, can be used to determine a condition or state of the cell or tissue. In some embodiments, the condition or state of a cell or tissue determined according to the methods and systems described herein can be used to determine the presence of - or a risk of developing - a condition in a subject from which the biological sample used to create the image was derived.

[0023] In embodiments, conclusions as to the spatial relationships of features (e.g., structures and / or molecules of interest) in a biological tissue sample or portion thereof can advantageously be determined according to the methods and systems described herein without prior knowledge of the orientation of the features represented in the imaged portion of the biological sample relative to the three-dimensional biological tissue sample from which the imaged area was derived. In some embodiments, spatial distance(s) of a first set of features of interest present in the two-dimensional image can be determined relative to a second set of features of interest (e.g., a second set of structures and / or molecules of interest) that are present in the three- dimensional biological tissue but which are not present in the two-dimensional image. In this way, spatial distances between a first set of features of interest in a biological tissue and a second set of features of interest in a biological tissue, which can be useful in determining a condition or state of a biological tissue or a subject, can be determined even in when the two-dimensional image fails to capture both the first and second sets of features of interest.

[0024] As described herein, features of a two-dimensional image representing structures or molecules of interest of at least a portion of one or more cells of a biological tissue can be extracted and then analyzed using a trained computational learning model to determine (e.g., based on the presence or absence and relative positions of the features in the image) a position of a second set of structures or molecules (e.g., located beyond the boundaries of the image in a third z- dimension that is perpendicular to the two-dimensional x-y plane of the image or in a direction within the x-y plane that extends beyond the in-plane edges of the image) with respect to the actual, three-dimensional cell(s) through which the two-dimensional image is a slice. Based on the determined presence, absence, and / or relative positions of features that are present in the image, properties of the complete three-dimensional cell(s) and / or tissue(s) can be predicted and / or reconstructed, including the presence, absence, and / or relative spatial location or distance of the second set of structural or molecular features. Quantification and analysis of the complete three-Docket No.: 064987-501001 WO dimensional cells in the tissue sample, including complete analysis of biomarker expression, can then be performed based on the predicted data, for instance to determine a state or condition of the cell(s), tissue(s), and / or subject from which the image was derived.

[0025] In many cases, existing technologies for determining the presence of biomarkers or cells of interest in a sample, an image of a sample, or a subject (e.g., a patient having or suspected of having a condition, such as cancer) provide inaccurate results, even when analyzed by hand by trained technicians. This can arise from artifacts of the imaging and / or sample preparation processes, for example, because many common forms of sample analysis are performed on portions (e.g., sections) of a sample that are assumed to be representative of the sample as a whole when, in fact, most biological tissues from which samples are taken - and even individual cells of the samples that are harvested - are spatially heterogeneous. As disclosed herein, it is possible to accurately determine not only the state of individual cells of an analyzed biological tissue sample but, in many cases, the state or condition of the tissue itself and / or the subject from which the tissue is derived.

[0026] One of the technical challenges addressed by the technology disclosed herein resides in the logistical limitations of analyzing biological tissue samples. It is frequently impossible or inadvisable to harvest, prepare, and analyze the entirety of the biological tissue of interest. As a result, it is often necessary to select only a subset (e.g., one or more portions or samples) of the tissue for analysis, for example, by obtaining biopsy samples of the subject’s tissue and slicing the samples into extremely thin sections (e.g., as thin as 10 micrometers or less). The thin slices (e.g., portions) of the biological tissue samples, which are often treated with stains, enzymes, and / or immunohistochemical agents to highlight structures and molecules of interest (including proteins, nucleic acids, carbohydrates, and the like) before the prepared samples are imaged or direct observed and analyzed by a technician, may or may not actually be representative of the cell or tissue as a whole or of a condition and may or may not accurately reflect a phenotype of all or a portion of the tissue due to the spatial heterogeneity of the biological tissue sample (or cells thereof). Analysis of biological tissue samples and, where applicable, diagnosis of a subject based on said analysis using existing technologies assumes that the analyzed sample(s) and / or image(s) are representative of the sample as a whole and that the sample as a whole is representative of the condition of the biological tissue of the patient. Therefore, if the specific portion(s) or section(s) of the biological tissue that are analyzed happen not to include one or moreDocket No.: 064987-501001 WO structures or molecules of interest (e.g., one or more proteins, nucleic acids, carbohydrates, etc.) or happen to include a portion of the cell(s) or tissue(s) that happens to have a particularly high spatial concentration of features (e.g., structures or molecules) of interest, analysis of the sample can overestimate, underestimate, or completely miss one or more parameters critical to the formation of conclusions regarding the condition of the cell, tissue, and / or subject, which can result in erroneous analyses and diagnoses of subjects and / or biological tissue samples.METHODS

[0027] Methods described herein (which include methods for determining the presence, absence, or relative spatial position of one or more biological features of interest and methods for determining a state or condition of a biological cell, a biological tissue, or a subject) can comprise receiving an image of a sample or portion thereof (e.g., a portion of a biological tissue sample), dividing the image into a plurality of two-dimensional grid segments, and determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of one or more features of interest present in the segment. In embodiments, determination of the presence, absence or location of one or more biological features of interest (e.g., biomarkers) can be accomplished using a trained computational machine learning model (e.g., a trained, semi-supervised computational machine learning model). For example, identification, extraction, and / or analysis of aspects of a first set of features of interest (e.g., biomarkers) from an image (e.g., a two-dimensional image) and prediction of the presence, absence, and / or location of second set of biological features of interest (e.g., a second set of biomarkers, either within the plane and boundaries of the image or entirely or partially outside of the plane and / or planar boundaries of the image) can be performed by a trained (e.g., semi-supervised) computational machine learning model. Aspects of the first set of biological features of interest (e.g., biomarkers) that can be identified, extracted, or analyzed as part of one or more steps of a method disclosed herein can comprise the presence, absence, relative intensity, and / or positioning of one or more biological features of interest of the first set of biological features of interest.

[0028] A plane in which a portion of a biological sample can define an x-y plane. In some embodiments, the plane can represent a thin slice of a biological sample, such as a single, essentially planar section of a tissue sample prepared with a microtome, or an image thereof). InDocketNo.: 064987-501001 WO some embodiments, a portion of a cell or tissue can be captured in the cross-sectional, planar region in the x-y plane defined by the section of the biological sample (or by the image thereof). Accordingly, an axis perpendicular to the tissue sample and running through a cell, tissue, or portion thereof in the captured portion of the cell or tissue can be considered to be a z-axis of the cell or biological tissue of the sample. Further, the position along a cell’s or tissue’s z-axis at which the tissue slice intersects (or the position along the z-axis of the cell or tissue represented in the image) the cell may be quantified in terms of a “z-level”. As explained further with respect to the Figures, a cell’s z-level can represent the distance of the sampled / intersected portion of the cell or tissue from either the top or bottom of the cell or tissue (e.g., whichever is closer) along the z-axis. The z-level of a cell is a relative measure on a per-cell, per-tissue, or per-image basis. A portion of a cell or tissue having a z-level of 0 can be defined to sit exactly centrally (in a z-axis direction) relative to the portion of the cell or tissue represented in the sectioned slice or image, while a portion of a cell or tissue having a z-level of 1 (or other non-zero value) would be located (e.g., in the unsectioned cell or tissue) partially or entirely out of the plane of tissue sample section or image (e.g., in a z-axis direction). It is important to note that a z-level of 0 (which can be defined by the plane of the section of the biological sample or image thereof) does not necessarily have to be the center of the cell or tissue (e.g., as measured by a maximum length, width, height, or centroid of the cell or tissue) and does not have to be centered on a biological feature of interest, such as a molecule of interest or structure of interest like a nucleus. Embodiments of the methods disclosed herein can provide analysis and predictive results regardless of the orientation of the cell or tissue in the sample section (or image thereof) and with relatively little requirements for what specific biomarkers must be captured in the imaged section of biological sample. As would be understood by persons of skill in the art, this is a major advantage over existing technologies, as it can be exceedingly technically difficult, expensive, and / or time-consuming to ensure that an imaged section of a biological sample captures a desired portion (e.g., wherein the desired portion is known to include specific features or structures of interest) or orientation of the sample during the sample preparation and imaging process.

[0029] Methods described herein and systems configured to execute such methods can be used to predict the presence of one or more features of interest that are not present in the sectioned portion of the biological sample or the image thereof. In some embodiments, a computational machine learning model (e.g., which has been trained to recognize the three-Docket No.: 064987-501001 WO dimensional distribution of one or more sets of biomarkers in a cell or tissue having a particular condition or state) can be used to predict the presence, absence, and / or location of one or more additional biomarkers that are not observed in the sectioned portion of the biological sample (or the image thereof). For example, a machine learning algorithm (e.g., a supervised computational machine learning model) can be trained on a set of training data comprising cells and / or tissues having a condition or state of interest (e.g., a set of images depicting biomarker distribution in a cell or tissue having a pro-inflammatory phenotype, a cancerous phenotype, an apoptotic phenotype, etc.) and used to determine a condition or state of the biological sample represented in an image of a biological sample of unknown condition or state. Using the information determined by the first computational machine learning module (e.g., a trained supervised computational machine learning model), which can include identification of structures and / or molecules of interest in the image, spatial distributions of the features of interest in the image, and / or a determined condition or state of the biological sample represented in the image of the portion of the biological sample, a second computation machine learning module (e.g., an unsupervised computational machine learning model) can be used to predict the presence or absence of one or more additional biological features of interest (e.g., biomarkers) that are not represented in the image of the biological sample. In some cases, the second (e.g., unsupervised) computational machine learning module can predict the presence, absence, or spatial location of one or more biological features of interest (e.g., structures or molecules of interest) that likely exist in a z-level outside of the plane of the biological sample represented in the analyzed image. In some cases, the second (e.g., unsupervised) computational machine learning module can predict the presence, absence, or spatial location of one or more biological features of interest (e.g., structures or molecules of interest) that likely exist at a location in the x-y plane of the biological sample outside of the boundaries of the analyzed image. Indeed, methods described herein can also be used to predict the existence and / or location of features of interest (e.g., structures or molecules of interest) that are likely present in the biological sample within the area and plane of the analyzed image but which have not been directly identified in the image (e.g., as a result of a failed or ineffective immunohistochemical or immunofluorescent agent or because no attempt was made to image the predicted feature(s) of interest), using the trained computational model’s “understanding” of what features of interest should be present in a cell or tissue of a condition or state that matches the condition or state of the cell or tissue represented in the image. As will be understood by personsDocket No.: 064987-501001 WO skilled in the art, this can provide a more complete picture of the molecules and structures of interest in a biological sample without the need to use time and money staining for those features of interest.

[0030] A method disclosed herein (or a system configured to perform said method) can be used to predict the presence, absence, or spatial location of one or more features of interest in a cell, wherein only a portion of the cell is represented in an image analyzed using the method (or system) and wherein the predicted information is not observable in the analyzed image.

[0031] Methods and systems described herein, which can be used to predict the z-level of a portion of a cell in a sample, as measured in a two-dimensional multiplexed imaging assay, can comprise segmenting the area of an image of a portion of a biological sample (e.g., after the image of the sample comprising the portion of the cell is received by the system). For example, a method step can comprise dividing (e.g., segmenting) a two-dimensional cell sample into segments in the x-y defined by a sectioned portion of the sample (or an image thereof), using the x-y coordinate of each cell segment as a surrogate data for z-level information. Based on the predicted z-levels for the cells in the sample, the expression profile of the entirety of each cell in the sample can be predicted or reconstructed from the partial cell segments in the imaged portion of the sample by training a class of models to link the cell segment expression to x-y-z data based on the surrogate data.

[0032] Such prediction systems and methods can be combined with models that account for variation in expression that remains once variation due to z-level has been accounted for to yield a prediction for the expression that any particular segment of a cell type captured in a tissue sample can be. For each of the original cell segments present in a tissue sample, it can then be possible to estimate the expression of the original full cells by modeling the unobserved other segments of the cell.

[0033] A method (or system configured to execute such a method) can comprise using a computational machine learning algorithm to determine how markers may be differentially measured between cell segments. In some cases, the machine learning algorithm can be used to predict the full biomarker expression of an entire cell from the two-dimensional image(s) of a biological sample comprising the cell.Docket No.: 064987-501001 WO

[0034] FIG. 1 demonstrates an exemplary workflow or method 100 for a whole-cell model of biomarker expression prediction that is performed on a two-dimensional image of a tissue sample. Method 100 may be performed, in some embodiments, by a processor of a computing system.

[0035] In a step 102 of the method 100, an image of a cell culture sample or tissue sample can received by the processor. The image may be generated using any HMI technology, including Xenium, CosMx, MERFISH, IMC, MIBI, CycIF, mIF, and 4i. Indeed, the image may be generated using any imaging technologies described above or any image production method that produces a two-dimensional image. The image of the cell culture sample or tissue sample may define an x-y plane in which portions of a cell of the sample are present. As described above, the cell of the sample at least partially captured in the area of the image may comprise portions of the cell that exist in x-y planes at various z-levels, and each of the z-levels of the cell may contain various and potentially differing quantities and types of features of interest, including molecules and structures of interest that may comprise portions of a nucleus, a cytoplasm, and / or a cell membrane.

[0036] In step 104 of the method 100, the processor divides the received image into a plurality of two-dimensional grid segments. The two-dimensional grid segments may divide the x-y plane defined by the image into a subgrid comprising a plurality of grid segments. The resolution of the subgrid may vary as appropriate depending on the density of cells and / or molecules of interest in the imaged portion of the tissue sample. Each grid segment of the subgrid may contain at least a portion of a first cell segment. In some embodiments, the number of grid segments and / or the distribution of grid segments into which the image is subdivided can be a hyperparameter. In some cases, the number of grid segments can be determined automatically by the computational machine learning model. In some cases, the number of grid segments and / or the distribution of grid segments across the image (or a portion thereof) can be definable by a user. In some embodiments, the number of grid segments per cell or region of a tissue can be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or greater than 50.

[0037] At step 106 of method 100, the processor may predict, for each cell segment in each grid segment of the subgrid, a z-level. The z-level represents the distance of the cell segmentDocket No.: 064987-501001 WO from the top or bottom of the cell, whichever is nearer. In other words, by predicting the z-level for a particular cell segment, the position of the segment within the original cell is predicted. The z-level may be predicted using a model trained on data from the x-y plane because, assuming cells are relatively spherically symmetric or isotropic, data about the cell in the x-y plane can be used to model data in the z-direction.

[0038] At step 108, the correct segment-quantified expression for the z-level of each cell is regressed out. Probabilistic models including standard linear regression, using robust regression variants, or more complex models including deep learning can be used to determine how biomarker expression varies with the combination of representation and predicted z-level. Thus, systems and methods described herein can predict complete biomarker expression in three dimensions for every cell having a portion of a segment in a given tissue sample, given only the two-dimensional image data.

[0039] At step 110 of the predictive model returns the corrected expression approximating the segment biomarker expression as if the entire cell had been measured.

[0040] Step 110 or a subsequent step in the method can comprise modeling (e.g., predicting) the presence, absence, and / or location of one or more biological features of interest (e.g., biomarker expression) over one or more additional z-levels for one or more cell segment or grid segment (e.g., using an unsupervised computational machine learning module).

[0041] Application of such a method can advantageously remove a ubiquitous source of bias and inaccuracy in single-cell data coming from studies based on thin tissue samples, and will allow for a more accurate analysis of images and tissue samples. Such advantages are to be particularly pronounced when analyzing rare cell populations that, by their natures, may be disproportionately affected by the distortions of sample preparation, as well as when performing spatial analyses where the proper identification and quantification of individual cells becomes more important. Systems and methods described herein could thus significantly increase the sensitivity of cell type quantification and, thus subsequently increase the accuracy of biomarker detection from multiplexed imaging. While the systems and methods described herein generally relate to antibody-based imaging, they could be applied to similar methods for measuring RNA (including, for example, 10X Xenium system) without departing from the spirit of the disclosure.

[0042] FIG. 2 illustrates a tissue sample 200 having segments 210 and 220 of twoDocket No.: 064987-501001 WO respective cells 211 and 221 captured therein. In some embodiments, one or more dimensions of a segment are manually defined by a user during analysis. In some embodiments, one or more dimensions of a segment are automatically defined by a computational machine learning model. A first layer 202 of FIG. 2 shows three-dimensional representations of cells 211 and 221 having respective nuclei 212 and 222, cytoplasm 214 and 224, and membranes 216 and 226. The tissue sample 200 is taken along the bottom surface 203 of layer 202. As shown in layer 204, the tissue sample 200 intersects cells 211 and 221 at varying distances along their vertical axes (e.g., at varying z-levels, as described further below with respect to FIG. 3). In some embodiments, cell 211 can have a z-level characterized by distance zi from the top of cell 211. In some embodiments, cell 221 can have a z-level characterized by distance Z2 from the top of cell 221.

[0043] Cell 211 may be positioned relative to plane 200 such that corresponding cell segment 210 in the tissue sample 200 comprises only cytoplasm 214. Accordingly, none of nucleus 212 is captured in the cell segment 211 contained in the tissue sample 200. Cell 221, however, is positioned such that corresponding cell segment 220 in the tissue sample 200 comprises both cytoplasm 224 and nucleus 222.

[0044] As described further with respect to FIGs. 3A-3C, the cell components of cells 210 and 220 that are present in respective cell segments 210 and 220 differ due to the z-levels of the cells. In general, the extent of a cell’s nucleus and cytoplasm captured in a given tissue sample will vary with the z-level, as described further below with respect to FIG. 3. Based on an assumption of roughly spherical and symmetrical cells, cells would show a larger area and a greater nuclear intensity in central segments compared to edge segments. Cells violating this assumption may be improperly registered in quantification of biomarker expression.

[0045] FIG. 3A shows a cell 300 having a nucleus 212, cytoplasm 214, and membrane 216. Cell 300 of FIG. 3 A may be, for example, either of cells 211 or 221 of FIG. 2. Cell 300 of FIG. 3 A may be assumed to have spherical symmetry and / or isotropy without loss of generality.

[0046] FIG. 3B shows the division of cell 300 into a plurality of partial cell segments 300-0, 300-1, 300-2, and 300-3. The cell segments of FIG. 3B are taken along the vertical or z- axis of cell 300, but aside described further below, cell segments can be taken along any of the axes of cell 300. As shown in FIG. 2B, segments 300-0 and 300-3 are spherical caps comprising only cytoplasm 214 and membrane 216, while segments 300-1 and 300-2 are spherical sectionsDocket No.: 064987-501001 WO comprising cytoplasm 214 and a portion of the cell nucleus 216.

[0047] As shown in FIG. 3C, the z-level of a cell segment represents the distance of the segment from the nearest spherical end cap. For example, the top of segment 300-3 and the bottom of segment 0 have z-levels of 0. The segment of the cell taken along the center of the nucleus 212 has a maximum z-level. A z-level may be a dimensionless number on a scale of 0 to 1. As such, the segment of the cell taken along the center of the nucleus 212 may have a z-level of 0.5.

[0048] While the cell segments 300-0, 300-1, 300-2, and 300-3 of FIG. 3B are shown with respect to the z-axis of cell 300, a cell can be divided into cell segments along any of its axes. For example, upon receiving an image of a tissue sample comprising a plurality of cell segments in the x-y plane, the cell segments can be divided into cell segments along their x axes and along their y axes. Because cells present in tissue samples may be generally spherically symmetric and / or isotropic (within reasonable assumption), the z-level of a cell segment present in a tissue sample can be predicted based on the cell segments in the x and y directions.

[0049] FIG. 4 demonstrates a distribution of z-levels of cells within cell clusters captured in an exemplary tissue sample to which the HMI technologies and present embodiments may be applied. As shown by FIG. 6, the z-levels of cells within given clusters can vary from 0 to 0.5, with means ranging from 0 to approximately 0.325. Accordingly, biomarker expression analysis gleaned from such cells may vary as certain biomarkers may present differently based on the z- level of the corresponding cell in which they are present in the tissue sample.

[0050] FIG. 5A and FIG. 5B illustrate method steps for determining the presence of features of interest in an image of a portion of a cell or tissue. As shown in FIG. 5A, a two- dimensional image of a portion of a cell can be divided into a plurality of grid segments for analysis. As shown in FIG. 5B, the grid segments can be assigned positional labels to determine spatial relationships of the grid segments to one another (e.g., center: 0, cytoplasm: 1, cell edge: 2), and an intensity of signal can be recorded for a plurality of individual biomarkers (e.g., wherein each differently colored column of FIG. 5B represents a signal intensity of a different protein of interest) in each grid segment. In some embodiments, these can be performed iteratively as a part of a supervised computational machine learning model training process. In some embodiments, observed signal intensities and locations for the assayed biomarkers can be used (e.g., by a trained computational machine learning model) to predict the presence of one or more biological featuresDocket No.: 064987-501001 WO of interest (e.g., biomarkers) that were not present in the image or that were not assayed in the imaging process.

[0051] A machine-learning model (MLM) that is thus able to learn what a cell looks like given data about the cell’s appearance in the x-y plane should thus be able to predict what the cell looks like in the z-axis based on the x-y plane data. In some implementations, cell segment data taken along the x and / or y directions can be used to train an ensemble learning model to predict z- level data of cell segments. For example, the x and y segment data can be used to train at least one of standard regression model, a convolutional neural network (CNN), and a decision tree model including a random forest model.

[0052] The x and y segment data can be subject to any number of modes of preprocessing prior to being fed to the computational machine learning model. The x and y segment data may be subject to any preprocessing method reasonably applied to IMC data or HMI data in general, including arcsinh normalization and standardization that shifts all marker channels into the same numerical range.

[0053] Feature extraction can be performed on cells divided into segments along the x and y axes. These features are put into a supervised learning model which can then be trained to predict the subsegments’ x and y levels. For feature extraction, a binary mask is constructed for each cell with 1 representing pixels contained in the cell and 0 otherwise. For each segment within the cell, a segment-specific binary mask is then created, where all pixels that are not both within the cell segment and within the cell are set to 0. The features for machine learning model training are computed by either extracting summary statistics (e.g., mean, median, max, min) of the quantified expression corresponding to pixels in the cell and segment or by returning the highdimensional expression image subset to pixels (or a bounding box around the pixels) that are nonzero in the computed mask.

[0054] Corresponding features extracted from segments of three-dimensional whole cells would then be able to be input into the model so trained to predict the z-level of the cell segment. For example, a random forest model or other machine learning model may be trained to predict the mean expression of a cell segment represented in a tissue sample based on the x- and y-level data. The model can taken an image or an expression vector as an input. In some implementations, cell-level prediction on the original tissue image is performed via the sameDocket No.: 064987-501001 WO procedure described above (i.e., with the creation of binary masks). Because prediction needs to be performed at the cell level, features may optionally be pooled over all segments in the cell before prediction via a pooling function such as mean, max, or mix.

[0055] FIG. 6 shows data evaluating the accuracy of predicted positions of biomarker (Ki67 expression in HER2 -positive breast cancer cells) expression at the edge, cytosol, or nuclear locations of a cell using method steps disclosed herein based on a trained computational machine learning model. The correspondence between the predicted position of the biomarker in the cell (y-axis) to the known actual position of the biomarker in the cell (x-axis), which can be used to assess the accuracy of the method and system, was strong for all conditions. Data showed accurate discernment of Ki67 expression at the edge of the cell versus in the cytosol of the cell (e.g., predicted versus true location of the biomarker at the edge of the cell vs. in the cytosol), with a p- value of 6.64 x 10'162. Data showed accurate discernment of Ki67 expression in the cytosol of the cell versus in the nucleus of the cell (e.g., predicted versus true location of the biomarker in the cytosol of the cell vs. in the nucleus), with a p-value of 7.79 x IO’182. Data showed accurate discernment of Ki67 expression at the edge of the cell versus in the nucleus of the cell (e.g., predicted versus true location of the biomarker at the edge of the cell vs. in the nucleus), with a p- value of 9.41 x 10'33.

[0056] FIG. 7 shows exemplary steps in determine the presence, absence, and / or location of one or more biological features of interest in a tissue sample (e.g., based on analysis of a two- dimensional image). A method disclosed herein, which can include one or more steps of method 700, can comprise a step of obtaining or receiving one or more images of a portion of a biological sample (e.g., step 702). A method can include a step of identifying one or more structures of interest in the one or more images (e.g., via manual analysis or via analysis by a trained computational model). A method can include a step of dividing the image into grid segments (e.g., step 704), for example, based a defined number, shape, and / or size of grid segments per image defined by a user or automatically by a computational model. A method can comprise determining for each of the grid segments whether the grid segment comprises a structure of interest (e.g., step 706). A method can comprise determining a distance to the structure of interest for each grid segment determined not to comprise a structure of interest and assigning the determined distance to the grid segment (step 708). A method can comprise training a supervised computational machine learning model to recognize expression characteristics of one or more biomarkers in aDocket No.: 064987-501001 WO tissue having a given condition or state. In some embodiments, a method can comprise training a computational machine learning model to predict a distance or proximity between a structure of interest in a tissue having a given condition or state and a second biological feature of interest in the tissue. A method can comprise predicting the presence, absence, and / or location of one or more features of interest relative to a structure of interest in a tissue based on analysis of a given image of a portion of a biological tissue sample, for example, wherein the predicted aspect(s) of the one or more features of interest are outside of the x-y plane of the tissue represented by the image and / or wherein the predicted aspect(s) of the one or more features of interest are outside of the boundaries of the image but still in the x-y plane of tissue represented by the image.

[0057] FIG. 8 shows data evaluating the accuracy of predicted positions of a tertiary lymphoid structure (TLS) to a feature of interest in a tissue using method steps disclosed herein based on a trained computational machine learning model. The correspondence between the predicted distance of the TLS in the tissue relative to the feature of interest (y-axis) to the known actual distance of the TLS in the cell relative to the feature of interest (x-axis) was strong. Data showed accurate correlation (p<0.001) between the predicted versus true distance of the TLS to the feature of interest.SYSTEMS

[0058] Systems described herein can comprise a processor and a non-transitory memory, the non-transitory memory comprising instructions that, when executed by the processor, perform one or more steps of a method disclosed herein. For instance, a system can comprise a computer running a computational machine learning algorithm trained to identify one or more first features of interest (e.g., first set of biomarkers) from a two-dimensional image and determine the presence or absence of one or more second features of interest (e.g., second set of biomarkers) not represented in the image (e.g., wherein the one or more second features of interest lie in a z-level not represented in the image) and / or a distance from the one or more first features of interest to the one or more second features of interest. In some embodiments, the system can perform one or more steps of a method disclosed herein for training the computational machine learning algorithm to identify and / or extract one or more features of interest from a two-dimensional image. In some embodiments, the system can perform one or more steps of a method disclosed herein to train theDocket No.: 064987-501001 WO computational machine learning algorithm to determine a spatial distance between a first set of biological features of interest (e.g., biomarkers) and a second set of biological features of interest that may not lie within the boundaries of the image.

[0059] A computational machine learning model of a system described herein, which can be stored or implemented using a system described herein, can comprise a supervised machine learning algorithm. In some cases, the supervised machine learning algorithm can be used to perform method steps of segmenting an image into portions (e.g., which may include z-level segmentation described herein) and / or identifying biological features of interest in a given segment of the analyzed image. In some embodiments, the supervised machine learning algorithm can be a linear regression model. In some cases, a machine learning algorithm can be a polynomial regression model.

[0060] A computational machine learning model of a system described herein can comprise an unsupervised machine learning algorithm or modality. For example, a system described herein can operate in an unsupervised machine learning regime after the model has been trained in a supervised machine learning context, e.g., to predict the presence, absence, or spatial location of one or more features of interest no present in a given image of a portion of a sample. Thus, in embodiments, a system described herein can comprise a semi-supervised machine learning configuration.

[0061] In some cases, a machine learning model of a method or system described herein can be or can comprise a polynomial regression model, a lasso regression model, a ridge regression model, a decision tree machine learning model, a random forest model, a support vector machine model, a k-nearest neighbors model, a neural network model, or a Bayesian network model.

[0062] In some cases, a system can comprise an input / output module for configured to receive and, optionally, store one or more images (e.g., one or more two-dimensional images of a portion of a biological sample). In some cases, the one or more images can comprise a plurality of images for training the computational machine learning algorithm with respect to identifying one or more biological features of interest (e.g., biomarkers). In some cases, the one or more images can comprise one or more images of portions of samples derived from a subject, which may be received from an imaging device. In some cases, a system can comprise an imaging device, which can comprise one or more of a camera, a microscope, and / or a light source (e.g., visible light, orDocket No.: 064987-501001 WO fluorescent light).

[0063] FIG. 9 is a system diagram illustrating one embodiment of a computing device configured to assess the biomarker expression in a sample comprising a cell. As illustrated, system 980 includes processor 910, memory 920, storage component 930, input interface 950, output interface 960, communication interface 970, and bus 940.

[0064] Bus 940 includes a component that permits communication among the components of system 980. In some implementations, processor 910 can be implemented in hardware, software, or a combination of hardware and software. In some examples, processor 910 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), and / or the like), a microphone, a digital signal processor (DSP), and / or any processing component (e.g., a field-programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or the like) that can be programmed to perform at least one function. Memory 920 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, optical memory, and / or the like) that stores data and / or instructions for use by processor 910.

[0065] Storage component 930 stores data and / or software related to the operation and use of system 980. In some examples, storage component 930 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, a solid state disk, and / or the like), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, a CD-ROM, RAM, PROM, EPROM, FLASH-EPROM, NV-RAM, and / or another type of computer readable medium, along with a corresponding drive.

[0066] Input interface 950 includes a component that permits system 980 to receive information, such as via user input (e.g., a touchscreen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, a camera, and / or the like). Additionally or alternatively, in some implementations the input interface 950 includes a sensor that senses information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, an actuator, and / or the like). Output interface 960 includes a component that provides output information from system 980 (e.g., a display, a speaker, one or more light-emitting diodes (LEDs), and / or the like).

[0067] In some implementations, communication interface 970 includes a transceiverlike component (e.g., a transceiver, a separate receiver and transmitter, and / or the like) that permitsDocket No.: 064987-501001 WO system 980 to communicate with other devices via a wired connection, a wireless connection, or a combination of wired and wireless connections. In some examples, communication interface 970 permits system 980 to receive information from another system and / or provide information to another system. In some examples, communication interface 970 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a cellular network interface, and / or the like.

[0068] In some implementations, system 980 performs one or more processes described herein. System 980 performs these processes based on processor 910 executing software instructions stored by a computer-readable medium, such as memory 920 and / or storage component 930. A computer-readable medium (e.g., a non-transitory computer readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes memory space located inside a single physical storage device or memory space spread across multiple physical storage devices.

[0069] In some implementations, software instructions are read into memory 920 and / or storage component 930 from another computer-readable medium or from another device via communication interface 970. When executed, software instructions stored in memory 920 and / or storage component 930 cause processor 910 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry can be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, embodiments described herein are not limited to any specific combination of hardware circuitry and software unless explicitly stated otherwise.

[0070] Memory 920 and / or storage component 930 includes data storage or at least one data structure (e.g., a database and / or the like). System 980 can be capable of receiving information from, storing information in, communicating information to, or searching information stored in the data storage or the at least one data structure in memory 920 or storage component 930. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0071] In some implementations, system 980 can be capable of executing software instructions that are either stored in memory 920 and / or in the memory of another device (e.g.,Docket No.: 064987-501001 WO another device that is the same as or similar to system 980). As used herein, the term “module” refers to at least one instruction stored in memory 920 and / or in the memory of another device that, when executed by processor 910 and / or by a processor of another device (e.g., another device that is the same as or similar to system 980) cause system 980 (e.g., at least one component of system 980) to perform one or more processes described herein. In some implementations, a module can be implemented in software, firmware, hardware, and / or the like.

[0072] The number and arrangement of components illustrated in FIG. 9 are provided as an example. In some implementations, system 980 can include additional components, fewer components, different components, or differently arranged components than those illustrated in FIG. 9. Additionally or alternatively, a set of components (e.g., one or more components) of system 980 can perform one or more functions described as being performed by another component or another set of components of system 980.SAMPLES

[0073] Methods and systems described herein can use biological samples or portions thereof as an input, e.g., to determine a condition or state of a cell, tissue, or subject. In some cases, a biological sample can be or can comprise a portion of a tissue obtained from a subject (e.g., via biopsy, collection of biological secretions, exudates, excretions, or other cell-containing liquids and solids, A method or system described herein can comprise obtaining and / or preparing a biological sample for imaging. In some cases, preparing a sample for imaging can comprise fixation of all or a portion of the sample (e.g., using formalin, paraformaldehyde, or the like). In some cases, a sample can be unfixed during imaging, so a method or system described herein may, in some cases, not comprise fixation of a portion of a sample being imaged. In some cases, preparing a sample for imaging can comprise freezing all or a portion of the sample.

[0074] In embodiments, one or more portions of a biological sample (e.g., a tissue or plurality of cells) is used for imaging. As discussed herein, the portion(s) of the biological sample can be thinly sliced sections of the sample (e.g., sections having a z-axis thickness of less than 5 micrometers, less than 10 micrometers, less than 15 micrometers, or less than 20 micrometers and x- and / or y- axis dimensions that are greater than the z-axis thickness, for instance, greater than 20 micrometers, greater than 100 micrometers, greater than 500 micrometers, greater than 1Docket No.: 064987-501001 WO millimeter, greater than 5 millimeters, greater than 10 millimeters, or greater than 20 millimeters) A biological sample can be sectioned using a microtome or a cryotome, for example, after freezing or fixation of the sample or embedding the sample in paraffin. In some cases, preparation of a sample can comprise deparaffinization. Preparation of a sample for imaging can comprise one or more pathology, immunohistochemical (IHC), or immunofluorescent imaging steps. For example, a sample can be exposed to a fixative or dehydrating agent such as formalin, ethanol, xylene, or paraformaldehyde. Preparation of a sample can comprise rehydration of a portion of the sample (e.g., after fixation). Preparation of a sample can comprise an antigen retrieval step, for example, to improve binding of an immunohistochemical imaging agent to a feature of interest in the sample. In some cases, preparation of a sample can comprise permeabilization of a cell or tissue, for example, comprising exposing the cell or tissue to a detergent such as Triton-X or Tween20 or treating with a permeabilization agent such as saponin. Preparation of a sample can comprise incubation of a portion of the sample with a molecular binding agent (e.g., an antibody, nanobody, aptamer, capture oligonucleotide, or probe) capable of binding specifically to a molecule of interest. In some cases, the antibody can be conjugated directly to probe or tag (e.g., comprising a polynucleotide or polypeptide used for detection and / or amplification), a chromophore, a fluorophore, an enzyme, or a colorimetric substrate. In some cases, the antibody can be a primary antibody. In some cases, preparation of a sample can comprise incubation of a portion of the sample with a secondary antibody capable of specifically recognizing and biding to the primary antibody. In some cases, preparation of a sample can comprise washing a portion of a sample, for instance, to remove one or more agents used in preparing the sample prior to the addition of new agents and cyclic reptition of molecular detection. Preparing a sample can comprise exposing a portion of the sample to a blocking agent (e.g., bovine serum albumin).

[0075] A sample can be obtained from a subject. The subject from which the sample is obtained can have or can be suspected of having a condition, such as a disease. For instance, the subject can have or be suspected of having a tumor or cancer. In some cases, the cancer can be a breast cancer, a brain cancer (e.g., a glioblastoma), a gastrointestinal cancer (e.g., an esophageal cancer, a gastric cancer, a colorectal cancer), a pancreatic cancer, a liver cancer, a bone cancer, a sarcoma, a cancer of the skin (e.g., melanoma), an ovarian cancer, a prostate cancer, or a lung cancer. In some cases, a sample or portion thereof (e.g., a sectioned portion of a biological tissue sample) can comprise one or more cells. A sample or portion thereof (e.g., a sectioned portion ofDocket No.: 064987-501001 WO a biological tissue sample) can comprise one or more tumor cells and / or one or more cancer cells. In some cases, the one or more tumor cells and / or one or more cancer cells can express one or more biomarkers indicative of or specific to a tumor or cancer cell phenotype or class, such as a cell surface marker (e.g., HER2, EGFR, ABCG2, CD 133, or carcinoembryonic antigen (CD66e)) or a proliferation marker (e.g., Ki67, PCNA, MCM2, Cyclin Bl, Cyclin DI), or a pro- or antiapoptosis marker such as Bcl-2, Mcl-1, Bcl-xL, p53, or Caspase-3. In some cases, the specific or relative location of one or more of these molecules in a cell can be indicative of a cancer or tumor phenotype. In some cases, a sample or portion thereof (e.g. a sectioned portion of a biological tissue sample) can comprise one or more immune cells, such as one or more T-cells, one or more B-cells, one or more natural killer (NK) cells, one or more macrophages, one or more monocytes, one or more neutrophils, and or one or more glial cells. In some cases, the one or more immune cells can express one or more biomarkers indicative of or specific to an immune cell phenotype or class, such as CD3, CD4, CD8, CD19, CDl lb, CD14, CD29, or CD45. In some cases, the biomarker can consist of a combination of probe measurements, or a gene signature. As described herein, the one or more biomarkers identifying a cell or a phenotype of a cell can be used in a method or system described herein to determine a condition or state of a cell, a tissue, or a subject, for instance, by using one or more aspects of the biomarker (e.g., spatial location relative to one or more other biomarkers and / or intensity of signal) as parameter inputs to a trained computer learning model.IMAGE-BASED INPUTS

[0076] Methods and systems described herein can comprise analysis of one or more images. The one or more images can be two-dimensional images (e.g., extending in an x-direction and a y-direction of an x-y plane). The image can represent all or a portion of a sample of cells or a biological tissue, e.g., wherein the cells or tissue are obtained from a subject or an in vitro cell culture. The image can be captured or generated from a sample (e.g., wherein the sample is a sectioned portion of a biological tissue) using an analog or digital camera. The image can comprise a plurality of regions comprising a plurality of pixels of differing color or intensity, which can define one or more features (e.g., features of interest) of one or more cells, structures, or molecules of a biological tissue. In some cases, regions of differing color or intensity can be the result of aDocket No.: 064987-501001 WO colorimetric signal (e.g., as the result of a dye such as hematoxylin and eosin, an enzymatic reaction such as generated using a horseradish peroxidase-catalyzed reaction, and / or a fluorescent signal such as generated by a fluorophore). In some cases, one or more aspects of the feature can be informed or defined by an intensity of the pixel(s). For instance, a location of a structure (e.g., a nucleus or other intracellular organelle, a center of a cell, an edge of a cell (e.g., as represented by a cell membrane or cell wall), a cell-cell interface) or molecule can be identified by a local maximum of pixel intensity, a local minimum of pixel intensity, a gradient or sudden change in pixel intensity in the image. Identification of features of interest (e.g., structures or molecules of interest) by identifying differences in color or intensity of pixels of an image can comprise or contribute to determination of a parameter of the feature, such as the presence or absence of the feature, the quantity of the feature that is present at a location of the image, or a distance from the feature to another structure or molecule.

[0077] Given the large degree of spatial heterogeneity in a given biomarker that may be present in the cells of a sample, measured biomarker expressions of cells contained in such samples may vary depending on which cell segment (i.e., which cross-section in the x-y plane taken along the vertical axis) of each cell is present in the sample. Such variation can inhibit accurate cellular interpretations in downstream analysis including cell type identification, especially in clustering for single cell phenotyping, because any quantification of properties of the cell segments in the tissue only reflects the segments of the cell present in the tissue sample, rather than properties of the whole cell.

[0078] In particular, this variation can inhibit subsequent quantification of cellular neighborhoods, communities, or niches, and can worsen associations with patient-level data including patient responses and outcomes. Further, the spatial heterogeneity of biomarkers in a sample can worsen quantification of intracellular distributions (e.g., of cell polarities or subcellular components including membranes, cytoplasm, nuclei, and organelles) and whole cell quantification. Absolute counts of spots also need to be corrected for the amount of cell sampled in a given tissue sample. Because the large degree of spatial heterogeneity in a given biomarker that may be present in the cells of a tissue sample complicates these analyses, it is common to throw out any imaging data that does not capture a full cell. Doing so further inhibits downstream analysis by excluding partial cell data and limiting the data on which such analyses are based. These difficulties can be circumvented using three-dimensional imaging techniques, but suchDocket No.: 064987-501001 WO techniques are often cost- and computation-intensive.

[0079] Histological analysis of tissue samples can be used to analyze tissue samples from patients, for instance, to determine a condition of a cell or biological tissue (e.g., the presence of cancerous cells). Hematoxylin and eosin stain (H&E stain) imaging can be used to examine biological tissues in histological preparations of a biological tissues. Low multiplexed imaging technologies including single biomarker immunohistochemistry (such as Opal Vectra 7-color detection) and immunofluorescence imaging are techniques that can be used for imaging in clinical applications.

[0080] Highly multiplexed imaging (HMI) technologies such as Imaging Mass Cytometry (IMC), Multiplexed Ion-Beam Imaging (MIBI), Cyclic Immunofluorescence (CycIF), Multiplexed Immunofluorescence (mIF), and Iterative Indirect Immunofluorescence Imaging (4i) can be used to study of biomarkers in tissue samples. These HMI technologies can image a single sample at subcellular resolution (200 - 1,000 nm) with many more biomarkers than is possible with conventional immunofluorescence methods. Spatial transcriptomics technologies such as probe and beads or spot capture, including Xenium, CosMx, Visium, Vizgen MERFISH, RNAscope, GeoMx, and LCM-mass spectrometry, MALDI-mass spectrometry, DESI-mass spectrometry or other mass spectrometry can also be used to determine the presence of biomarkers in tissue samples.

[0081] Such imaging capabilities allow for high parameter, single-cell measurements of DNA, RNA, proteins, drugs, and metabolites that can be placed and preserved in a spatial context. As one example, IMC can use time-of-flight mass cytometry to measure heavy metal ions tagged onto antibodies with a resolution of one squared micron and an upper limit of approximately forty biomarkers. IMC offers the advantage of measuring all such biomarkers in a single acquisition, without requiring cycles of sample processing and tissue removal, like most immunofluorescencebased methods. HMI and IMC technologies generally operate on thin slices of tissue. In some instances, the slices of tissue on which these technologies operate can have thicknesses of between 2 and 5 micrometers (pm). These technologies can be useful in cancer and immunology research, for example, as the advantages just described can be useful in studying tumor microenvironments, identifying immune subtypes, and developing biomarkers.

[0082] HMI, IMC, and other imaging technologies can be applied to tissue samples thatDocket No.: 064987-501001 WO usually have thicknesses thinner than the diameters of most of the cells present in the sample. These thin tissue slices define an x-y plane that intersects the cells contained therein at different positions along the axes of the cells’ that are perpendicular to the plane defined by the tissue sample (referred to herein as the cell’s vertical or z axes). The cells in the samples are thus divided into partial cell segments at different levels along their vertical axes (i.e., as used herein, the terms “segment”, “cell segment”, and “partial cell segment” describe a cross-section of a cell at any point along the z-axis of the cell). Thus, in embodiments of methods and systems described herein, segments of cells, rather than complete three-dimensional cells, can be captured of the sample and used for imaging.

[0083] As will be understood by persons of skill in the art, an advantageous distinction of the methods and systems described herein, for example, compared to 3 -dimensional imaging modalities and techniques is that the methods and systems described herein do not require multiple images to be combined (e.g., “stacked” or “stitched together”) to create a three-dimensional representation of the tissue. Indeed, one major advantage of the methods and systems described herein, as compared to existing technologies, is that accurate predictions and conclusions regarding the presence or absence one or more features of interest not present in the analyzed image(s) and / or spatial distance(s) between one or more features of interest present in an analyzed image and one or more features of interest that are not present in the analyzed image (or group of analyzed images) can still be made. For instance, one or more features of interest present in an image analyzed using a method or system described herein can be used to determine the presence or absence of a feature of interest that may be in a different z-level as compared to the x-y plane of the analyzed image (or at a location in the x-y plane that exists beyond the edge of the image’s x-axis or y-axis boundaries) and / or to determine a distance between the one or more features of interest present in the analyzed image and the one or more features of interest that are not in the analyzed image.BIOLOGICAL FEATURES OF INTEREST

[0084] Methods and systems described herein can utilize data pertaining to biological feature(s) of interest present in or extracted from an (e.g., two-dimensional) image in determining one or more biomarker parameters. The one or more biomarker parameters that can be determined from data pertaining to the biological feature(s) of interest in an image can include one or more of:Docket No.: 064987-501001 WO the presence or absence of other biological features of interest in a sample, a distance between a first set of biological features present in an image and a second set of biological features present in an image, and / or a distance between one or more biological features observed in an image and one or more biological features not present in the image. Once determined (e.g., using a trained computational machine learning model, as described herein) one or more of these biological features of interest (e.g., parameters) can be used to determine a condition or state of a cell, a biological tissue, or a patient from which the imaged and analyzed sample was derived. For example, the presence, absence, and / or spatial location of one or more biological features of interest (e.g., biomarkers) in a (e.g., two-dimensional) image can be extracted and analyzed as described herein to determine one or more biomarker parameters (e.g., the presence, absence, and / or relative location of one or more other biological features of interest that cannot be observed in the image), and, optionally, a determination of a condition or state of a cell, biological tissue, or subject can be determined from the one or more biomarker parameters. Some examples of conditions or states of cells that can be (e.g., accurately and unambiguously) determined in this way using methods and systems disclosed herein can include: cellular states of quiescence, proliferation, hyperproliferation, apoptosis, epithelial-mesenchymal transition (EMT), cell cycle state, inflammatory activation (e.g., of immune cells), inflammatory recruitment, or phagocytosis. Some examples of conditions or states of tissues that can be (e.g., accurately and unambiguously) determined in this way using methods and systems disclosed herein can include: pro-inflammatory environments, tissue necrosis, the presence or progression of multi-cellular aggregates or lesions (e.g., the presence or formation of tumors, cancers, or tertiary lymphoid structures), extravasation of cells from a blood vessel, tissue necrosis, metastatic cancer phenotypes, invasion or penetration of a first cell type (such as an immune cell) into a structure (such as a tumor or cancer lesion), graft versus host (GVH) phenomena, and / or pro-angiogenic environments. Some examples of conditions or states of subjects that can be (e.g., accurately and unambiguously) determined in this way using methods and systems disclosed herein can include: the presence of a disease or condition in the subject (e.g., cancer, arthritis, cardiovascular disease (including peripheral arterial disease), autoimmune disorders, and / or neurodegenerative disease such as amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Parkinson’s disease, or Alzheimer’s disease), a risk of developing a disease or condition, and / or the extent of progression of a disease or condition in the subject. In some embodiments of the methods and systems disclosed herein, the determination ofDocket No.: 064987-501001 WO a condition or state of a cell, a tissue, or a subject can be used to determine a treatment for a patient in need thereof.

[0085] Prediction of the presence, the absence, or a location of one or more features of interest (e.g., one or more features of interest in three-dimensional space based on analysis of a two-dimensional image) can comprise evaluation of a plurality of features of interest present in an analyzed two-dimensional image. In some embodiments, one or more aspects, such as relative spatial location in the image, intensity, or color of at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, at least 1000, or more than 1000 individual features of interest (e.g., individual instances of the features, wherein the features may include multiple instances of the same type of molecule or structure) present in the two-dimensional image can be determined (e.g., identified, extracted as a parameter value, calculated, or predicted) by a trained computational machine learning model of a method or system disclosed herein. In some embodiments, one or more aspects, such as relative spatial location in the image, intensity, or color of at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, at least 1000, or more than 1000 different types of features of interest (e.g., different types of molecules or structures) present in the two-dimensional image can be determined (e.g., identified, extracted as a parameter value, or calculated) can be processed by a trained computational machine learning model of a method or system disclosed herein. In embodiments, the spatial orientation of a cell or tissue, or of a first set of biomarkers to a second set of biomarkers in a training data set does not have to be the same as the orientation of the cell, tissue, or arrangement of biomarkers in the analyzed image or the predicted (e.g., reconstructed) cell, tissue, or biomarker data. For example, an image to be analyzed using a method or system disclosed herein can be independent of the orientation of data (e.g., images) of the data set used to train the computational machine learning model(s) described herein.

[0086] Biological features of interest (e.g., biomarkers) can include a molecule that is incorporated into, contained inside of, or associated with a biological cell or tissue. In some cases,Docket No.: 064987-501001 WO a biological feature of interest (e.g., a biomarker) can be selected from a protein (e.g., enzymes, hormones, cytokines, or structural proteins such as collagen, elastin, or tubulin), a nucleic acid (e.g., DNA which can include genomic nuclear DNA, mitochondrial DNA cell-free DNA and / or complementary DNA, RNA which can include mRNA, miRNA, rRNA, tRNA, shRNA, and / or siRNA, or oligonucleotides), a carbohydrate, a glycoprotein, or a lipid (e.g., phospholipids, steroids, etc.). In embodiments, a biological feature of interest can be a cell surface molecule (e.g., a protein, lipid, or carbohydrate associated with or incorporated into a cell membrane). In embodiments, a biological feature of interest (e.g., biomarker) can be a molecule associated with a cell’s nucleus. In embodiments, a biological feature of interest (e.g., biomarker) can be a trafficked protein (e.g., a protein transported within a cell). In embodiments, a biological feature of interest (e.g., biomarker) can be a signaling molecule. In embodiments, a biological feature of interest (e.g., biomarker) can be a cleavage product (e.g., a product of enzymatic cleavage). In embodiments, a biological feature of interest (e.g., biomarker) can be a region of increased or decreased pH (e.g., a low-pH region within a lysosome). In some embodiments, a biological feature of interest (e.g., a biomarker) can be a modified molecule (e.g., a post-transcriptionally modified molecule, such as an acetylated protein, a phosphorylated protein, a glycosylated protein, a glucosylated protein, a ubiquitinated protein, an alkylated protein, a methylated protein, an aminated protein, or a hydroxylated protein). In some embodiments, a biological feature of interest (e.g., a biomarker) can be a splice variant of an mRNA or a protein. In some embodiments, a biological feature of interest (e.g., a biomarker) can be a combination of multiple molecular measurements. In some embodiments, a biological feature of interest (e.g., biomarker) can be a gene signature or genetic expression pattern. A gene signature or genetic expression pattern can be directly observed in the input data (e.g., the image(s) being analyzed) or, in some embodiments, a gene signature or genetic expression pattern can be determined or derived from analysis of one or more observed or measured biomarkers in the input data (e.g., one or more biomarkers identified and measured in the analyzed image(s)). Thus, in some embodiments, it is possible to use analysis of one or more biomarkers (e.g., structures and / or molecules of interest present in the image(s)) to determine a gene signature or genetic expression pattern, which may be itself used as a biomarker in determining the state or condition of a cell, tissue, or subject. Biomarkers comprising or consisting of such gene signature or gene expression pattern information can also be associated with a spatial location, e.g., within the image as an input metric for analysis by the methods andDocket No.: 064987-501001 WO systems disclosed herein or at a position predicted as an output by the methods and systems described herein.

[0087] Biological features of interest (e.g., biomarkers) can include a structure of a biological cell or tissue. A structure of a biological sample can be a blood vessel (e.g., an edge of a cell of a blood vessel, such as a vascular smooth muscle cell membrane or vascular endothelial cell membrane). A structure of a biological sample can be a plurality of cells associated with one another, such as an aggregate of cells. In some cases, the aggregate of cells can be homogeneous with respect to the cell type making up the aggregate. In some cases, the aggregate of cells can be heterogeneous with respect to the cell type making up the aggregate. In some embodiments, the composition of an aggregate of cells (e.g., the quantity of cells and / or the degree of heterogeneity or homogeneity of cell phenotypes, for example, as assessed by expression of molecular features of the cells of the aggregate) can be a feature of interest (e.g., biomarker) useful in a method or system disclosed herein. In some embodiments, a structural feature of interest of an image of a tissue can be non-biological. For instance, in some embodiments (e.g., testing of biocompatibility of implanted materials), a non-biological structural feature can be a metal (or metal alloy), a polymer (e.g., a synthetic biopolymer), or a ceramic material.

[0088] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. For example, the logic flows may include different and / or additional operations than shown without departing from the scope of the present disclosure. One or more operations of the logic flows may beDocket No.: 064987-501001 WO repeated and / or omitted without departing from the scope of the present disclosure. Other implementations may be within the scope of the following claims.

[0089] Further, while “upwards”, “downwards,” or other directions are described herein, the directions are referred to as such for ease of description. The directions are not limited to upwards, downwards, and / or the like.

Claims

1. Docket No.: 064987-501001 WOCLAIMSWhat is claimed is:

1. A method of determining whether one or more biomarkers of interest are present in a sample, the method comprising: receiving an image of a portion of the sample, the image comprising a plurality of cellcross segments; dividing the image into a plurality of two-dimensional grid segments, each of the plurality of two-dimensional grid segments comprising a plurality of cell segments; determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of a cell contained within the segment; correcting a segment-quantified biomarker expression for each of the cells contained within each of the plurality of grid segments by regressing out z-level information; and returning a corrected expression approximating a predicted biomarker expression for each of the plurality of cell segments as if a biomarker expression of an entirety of each of the cells contained within each of the plurality of the grid segments had been measured.

2. The method of claim 1, further comprising determining a location of the one or more biomarkers in the sample.

3. The method of claim 2, further comprising determining that one or more biomarkers are present in the sample outside of a plane of the sample defined by the image.

4. The method of any one of claims 1-3, further comprising determining a state of the sample based on the presence of the one or more biomarkers in the sample.

5. The method of any one of claims 1-4, wherein the method is performed using a trained computational machine learning model.

6. The method of claim 5, wherein the computational machine learning model is a semisupervised machine learning model.

7. The method of any one of claims 1-6, further comprising determining a risk for a subjectDocket No.: 064987-501001 WO to develop a condition based on the determination of a whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject.

8. The method of any one of claims 1-6, further comprising determining whether a subject has a condition based on the determination of whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject.

9. The method of claim 7 or claim 8, wherein the condition is a cancer, a tumor, a pro- inflammatory condition, or a neurodegenerative disease.

10. A system for determining the presence of one or more biomarkers in a sample, the system comprising a processor and a non-transitory memory, the non-transitory memory comprising instructions that, when executed by the processor, perform the method of any one of the preceding claims.

11. A system for determining the presence or absence of one or more biomarkers of interest in a sample, the system comprising: a processor; a non-transitory memory comprising instructions that, when executed by the processor to perform steps comprising: dividing an image of a portion of a sample into a plurality of two-dimensional grid segments, each of the plurality of two-dimensional grid segments comprising a plurality of cell segments; determining, for each of the plurality of two-dimensional grid segments, a corresponding z-level of a cell contained within the segment; correcting a segment-quantified biomarker expression for each of the cells contained within each of the plurality of grid segments by regressing out z-level information; and returning a corrected expression approximating a predicted biomarker expression for each of the plurality of cell segments as if a biomarker expression of an entirety of each of the cells contained within each of the plurality of the grid segments had been measured.Docket No.: 064987-501001 WO12. The system of claim 11, wherein the steps further comprise determining whether a subject has a condition based on the determination of whether one or more biomarkers of interest are present in the sample, wherein the sample is derived from the subject.

13. The system of claim 12, wherein the condition is a cancer, a tumor, a pro-inflammatory condition, or a neurodegenerative disease.

14. The system of claim 12 or claim 13, further comprising a display for displaying the corrected expression or information reflecting the determination whether a subject has the condition.

15. The system of any one of claims 11-14, wherein the non-transitory memory comprises a trained computational machine learning model.

Citation Information

Patent Citations

  • Method for scoring pathology images using spatial analysis of tissues

    US20180089495A1