Graph construction and visualization of multiplex immunofluorescence images

JP7920158B2Active Publication Date: 2026-09-14ASTRAZENECA AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023541733
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-12
Filing Date
2022-01-11
Publication Date
2026-09-14
Estimated Expiration
2042-01-11

AI Technical Summary

Benefits of technology

【0005】 いくつかの非限定的な実施形態では、システムは、画像分析及びデータ可視化のための汎用コンピューティングデバイス又はより専用的なデバイスの中で実施されるパイプラインであってもよい。システムは、命令がその中に記憶されたメモリ及び/又は非一時的コンピュータ可読記憶デバイスを含み得る。少なくとも1つのコンピュータプロセッサによって実行されるときに、MIF画像を分析し、MIF画像内の細胞を表す対話型可視化を生成するために、様々な動作がローカル又はリモートで実行され得る。本明細書に開示される技術の実施態様では、対話型可視化は、MIF画像内の細胞に関連付けられたデータを操作するためのインターフェースを提供する。このようにして、MIF画像内のデータは、従来の分析を通して以前は可能でなかった画像内の細胞的洞察を明らかにするために処理され得る。対話型可視化は、医療提供者が仮説及びデータ主導の研究を行い得る方法でそれらの洞察を示し、仮説及びデータ主導の研究は、疾患のより正確な診断、医療成績及び治療に対する反応のより正確な予測、並びに現在の治療に細胞がどのように反応するかについてのより良い理解につながり得る。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007920158000001
    Figure 0007920158000001
  • Figure 0007920158000002
    Figure 0007920158000002
  • Figure 0007920158000003
    Figure 0007920158000003
Patent Text Reader

Abstract

Provided herein are embodiments of systems, devices, articles of manufacture, methods, and / or computer program products, and / or combinations and subcombinations thereof, for providing interactive exploration and analysis of a cellular environment depicted in a MIF image. The embodiments include a pipeline configured to generate an interactive visualization having selectable icons representing cells in the MIF image by identifying cells in the MIF image and generating a graph of the MIF image based on coordinates and properties of the identified cells, where each node in the graph corresponds to a cell and a neighboring cell. The graph may be converted into an embedding, and an interactive visualization of the graph may be generated based on the embedding. The selectable icons in the interactive visualization correspond to the nodes in the graph.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure generally aims to construct a graph representation of multiplex immunofluorescence images for interactively exploring cell relationships within multiplex immunofluorescence images and generating predictions regarding treatment outcome and therapeutic effect. [Background Art]

[0002] Multiplex immunofluorescence (MIF) is a molecular histopathology tool for antigen detection in biological samples using labeled antibodies. MIF has emerged as a useful tool for enabling simultaneous detection of biomarker expression in tissue sections and providing insights into cell composition, function, and interactions. One of the advantages of MIF is that it captures complex and extensive information about cells in the cellular microenvironment. Although the depth and breadth of data provided by MIF images are useful, the enormous complexity and data volume present challenges for interpretation and visualization. In other words, analyzing data in MIF images can be a difficult task. [Summary of the Invention] [Means for Solving the Problems]

[0003] Cellular microenvironments are known to be complex. They can comprise millions of cells and many different cell types. Each of these cells can have hundreds of potential interactions. While analysis of biomarker expression in MIF images provides a useful starting point for analyzing cells and their interactions, manual analysis of MIF images is practically impossible given the number of cells involved. The technology described in the present disclosure converts the data provided by MIF images into an intuitive graphical format, in which cells in MIF images can be manipulated, filtered, queried, and utilized for relevant medical prediction.

[0004] This specification provides embodiments of systems, apparatus, products, methods, and / or computer program products, as well as combinations and partial combinations thereof, for providing interactive exploration and analysis of the cellular environment represented in MIF images.

[0005] In some non-limiting embodiments, the system may be a pipeline implemented within a general-purpose computing device or a more specialized device for image analysis and data visualization. The system may include memory and / or non-temporary computer-readable storage devices in which instructions are stored. Various operations may be performed locally or remotely to analyze MIF images and generate interactive visualizations representing cells within the MIF images when executed by at least one computer processor. In embodiments of the technology disclosed herein, the interactive visualization provides an interface for manipulating data associated with cells within the MIF images. In this way, data within the MIF images can be processed to reveal cellular insights within the images that were previously not possible through conventional analysis. The interactive visualization presents these insights in a way that enables healthcare providers to conduct hypothesis- and data-driven research, which can lead to more accurate diagnosis of disease, more accurate prediction of medical outcomes and responses to treatment, and a better understanding of how cells respond to current treatments.

[0006] Embodiments are intended to describe embodiments of systems, apparatus, products, methods, and / or computer program products, and / or combinations and partial combinations thereof, for generating interactive visualizations having selectable icons representing cells in MIF images. These embodiments may include identifying cells in an MIF image. Each cell may be associated with coordinates and properties. Embodiments may further include generating a graph of the MIF image based on coordinates and properties, wherein the graph includes nodes corresponding to cells and neighboring cells. The graph may further include edges connecting the nodes and encoding properties about each cell, such as information about neighboring cells around that cell. Embodiments may further include converting the graph into an embedding, which is a mathematical vector representation of the graph including nodes, edges, and properties. Embodiments may further include providing an interactive visualization of the embedding-based graph.

[0007] It should be understood that the section describing embodiments for carrying out the invention, rather than the section describing the summary and abstract of the invention, is intended to be used in interpreting the claims. The section describing the summary and abstract of the invention may describe some, but not all, possible embodiments of the enhanced densification technique for providing interactive visualization of MIF data described herein, and is therefore not intended to limit the scope of the appended claims in any way.

[0008] The attached drawings are incorporated herein and form part of the specification. [Brief explanation of the drawing]

[0009] [Figure 1] An overview of some exemplary embodiments is provided below. [Figure 2A] Several embodiments of a multiplex immunofluorescence imaging system pipeline are shown as examples for processing multiplex immunofluorescence images. [Figure 2B]An alternative embodiment of a multiplex immunofluorescence imaging system pipeline is shown as an example. [Figure 3] This document presents an example method for providing interactive visualization of multiplex immunofluorescence images using several embodiments. [Figure 4A] The following process flows illustrate examples of interactive visualization of multiplex immunofluorescence images using several embodiments. [Figure 4B] This document presents an example process flow for generating cell vicinity plots from interactive visualization of multiplex immunofluorescence images using several embodiments. [Figure 5] An example computer system useful for implementing various embodiments is shown. [Modes for carrying out the invention]

[0010] In drawings, similar reference numbers generally indicate the same or similar elements. Additionally, the leftmost digit of a reference number generally identifies the drawing in which that reference number first appears.

[0011] The modes for carrying out the following inventions of this disclosure refer to the accompanying drawings illustrating exemplary embodiments consistent with this disclosure. The exemplary embodiments fully illustrate the general nature of this disclosure so that others, by applying the knowledge of those skilled in the art, may readily modify and / or adapt such exemplary embodiments for various uses without departing from the spirit and scope of this disclosure, without unnecessary experimentation. Accordingly, such adapted and modified forms are intended to be within the meaning and scope of the exemplary embodiments, based on the teachings and guidance presented herein. It should be understood that the terms or technical terms used herein are for illustrative purposes only, not limitation, so that they should be interpreted by those skilled in the art in light of the teachings provided herein. Therefore, the modes for carrying out the inventions do not mean to limit this disclosure.

[0012] The embodiments described and references in the specification such as “one embodiment,” “an embodiment,” and “an exemplary embodiment” indicate that the embodiments described may include certain features, structures, or characteristics, but not all embodiments necessarily include such features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, when certain features, structures, or characteristics are described in relation to an embodiment, it should be understood that it is within the knowledge of those skilled in the art that such features, structures, or characteristics may be obtained in relation to other embodiments, whether or not they are explicitly described.

[0013] Figure 1 is an overview diagram of an example process flow 100. As described herein, the example process flow 100 simply provides a schematic description of the features for providing interactive visualization of MIF image data. The following description describes interactive visualization in the context of MIF and MIF image data, but interactive visualization may also be based on other types of image data that provide cellular information, such as multiplex immunohistochemistry (MIHC) image data and mass spectrometry imaging. Thus, the following description of processing MIF images to create graphs and embeddings and subsequently generating interactive visualization may also apply to processing image data from MIHC and mass spectrometry images. Further details of the process and interactive visualization are described below with reference to Figures 2 to 4.

[0014] As shown in Figure 1, an example process flow 100 may begin with an MIF image. The cellular environment in an MIF image is typically very complex, and the information provided by the biomarker representation of the MIF results in a large dataset to be interpreted. Process flow 100 describes transforming the MIF image into a visualization that allows for the manipulation and utilization of these datasets in a more automated manner. Process flow 100 may involve analyzing the MIF image to identify data points associated with the cellular environment in the image. In embodiments, the analysis may include identifying tissue cells in the MIF image and segmenting the identified tissue cells into biologically significant regions. Examples of such biologically significant regions include tumor centers and tumor stroma.

[0015] Examples of functions for performing the splitting step may include a training-region-based random forest classifier. For example, cell splitting may involve identifying round objects of a given size with sufficient contrast eccentricity for nuclear biomarker channels. As another example, Voronoi bifurcations may be used to estimate the cell membrane and cytoplasm of a cell in an MIF image, and membrane biomarkers may be used to refine these estimates. Biomarker positivity is generally assessed by integrating the intensity of each channel across a cellular area or intracellular compartment.

[0016] The results of MIF analysis provide information on image data points that indicate the cellular environment, such as cell location, cellular phenotype (which may consider combinations of biomarkers associated with the cell), intercellular relationships, and immunofluorescence intensity. Biomarkers used in MIF may include molecules that can be measured on or inside a cell and used to characterize a cell or tissue. For example, the expression of a certain biomarker may indicate the identity, function, or activity of a cell. Different diseases (cancer, heart disease) may be associated with biomarkers that can be used to diagnose the disease, and responses to medical treatment may cause changes in biomarker expression. Examples of biomarkers include DNA or proteins.

[0017] Positioning information can be stored in a tabular format describing the coordinates and phenotype of each cell identified in the MIF image. Cell positioning may be based on the centroid of the nucleus or the entire cell. In embodiments, positioning may be performed as cell X / Y coordinates on a whole-slide coordinate system. Image information may be recorded in a log file and used in the next step in process flow 100, which is graph construction.

[0018] The graph may be constructed based on image information. For example, this step may involve extracting cell locations (X / Y coordinates) from a log file and extracting biomarker information such as the immunofluorescence intensity of each cell. Other examples of biomarker information include auxiliary features such as cell diameter, cell area, and cell eccentricity. The graph may be constructed based on this extracted information by performing cell pair identification. In embodiments, a selectable distance threshold may be used to identify cell pairs. Identified pairs within the range of the selectable distance threshold may be connected in the graph by edges. The graph may be constructed when all pairs of cells have been identified and added to the graph as nodes connected by edges. Thus, the constructed graph consists of nodes corresponding to cells shown in the MIF image and edges connecting some nodes. The edges correspond to selectable distance thresholds between any nodes in the graph.

[0019] Once the graph is constructed, the next step in process flow 100 involves converting the graph into a graph embedding, which is a numerical or feature vector representing the encoded information in the constructed graph. The embedding can numerically represent, for example, properties or features of the constructed graph in a vector space where nearby embeddings represent similar nodes, edges, and subgraphs, for example, in a matrix. Thus, the embedding can capture information about the graph at various levels of granularity, including the graph's nodes, graph edges, and subgraphs of the entire graph.

[0020] The embedding can represent information about each cell represented as a node in a graph, such as cell position, immunofluorescence intensity, cell phenotype, and cell neighborhood. A cell neighborhood may be centered on the node corresponding to that cell and any surrounding nodes within a certain number of hops from the central node, which correspond to neighboring cells shown in the MIF image. The embedding may also be trained to more accurately represent cell information or neighborhood information.

[0021] The embedding can be generated by a graph training algorithm to learn nodes and edges in a graph. Application of the graph training algorithm may include the initial steps of selecting a number of hops to determine the neighborhood size, and selecting an embedding size or type of embedding output. In some embodiments, the graph training algorithm is unsupervised. Examples include Deep Graph Infomax (DGI) and GraphSAGE. The size of the embedding represents the amount of graph information that can be captured by the embedding. A larger embedding size allows more information to be represented in the embedding space, and thus potentially more information from the graph can be captured by the embedding. Types of embedding output include node embedding, edge embedding, hybrids of both node and edge embedding, subgraph embedding, and full graph embedding, to name a few.

[0022] Converting graphs to vector representations provides several advantages. Data in graphs are not easily manipulated because they are composed of edges and nodes, and this structure also reduces the effectiveness of machine learning algorithms. In comparison, vector representations are easier to manipulate and provide more options for machine learning algorithms that can be applied to them.

[0023] After generating the graph embedding, the process flow 100 next generates an interactive visualization based on the information in the graph embedding. An interactive visualization is a tool for exploring information encoded in an embedding. In an embodiment, the interactive visualization provides a two-dimensional visual projection of the embedding. A scatter plot is an example of such a two-dimensional projection. Nodes in the interactive visualization can be manipulated in various ways such as cell selection, biomarker filtering, or neighborhood plots, to name a few. Thereby, several real-world diagnostic benefits are expected, including clustering nodes for a particular patient, identifying patients who responded to treatment based on their nodes, and generating predictions of how a patient may respond to treatment based on extrapolated estimation of data based on the nodes.

[0024] Interactive visualization enables insight into cellular data encoded in MIF images, which typically require laborious manual analysis and effort, since MIF images can represent millions of cells. The graph embedding provides information not only about each cell, but also the surrounding neighborhood of the cell and cell relationships including inter-cell similarity, inter-cell difference, cell clusters, and temporal information, which may similarly be provided in an interactive and visual manner. Because it shows interactions at the cell and cell neighborhood levels, interactive visualization enables advanced hypothesis and data-driven research based on MIF images for several possible uses, including understanding cellular mechanisms regarding cell responses to different medical treatments and different responses and outcomes across different patients.

[0025] Process flow 100 will now be described in further detail with reference to FIGS. 2A, 2B, and 3 to 5.

[0026] Figure 2A shows a multiplex immunofluorescence (MIF) system pipeline 200A as an example for processing MIF images according to several embodiments. The MIF system pipeline 200A may be implemented as part of a computer system, such as the computer system 500 in Figure 5. The MIF system pipeline 200A may include various components for processing MIF images and providing interactive visualization of MIF images. As shown in Figure 2A, the MIF system pipeline 200A may include an MIF image analyzer 202A, a graph generator 204A, an embedding generator 206A, and a visualizer 208A.

[0027] The MIF image analyzer 202A analyzes MIF images to identify cells in the images, extracts information from the MIF images, and converts the information into tabular data. In some embodiments, the MIF image analyzer 202A may use software to perform the conversion to tabular data. An example of such software includes a digital pathology platform for quantitative tissue analysis.

[0028] Tabular data may include the coordinates and properties of each cell. The cell coordinates can be used to identify the cell's position in the MIF image and to plot the cell relative to other cells in the MIF image. In embodiments, the coordinates may be implemented as X / Y coordinates. The coordinates may be based on different aspects of the cell. For example, the MIF image analyzer 202A may identify the cell nucleus relative to the cell, and the center of the cell nucleus may be used as the cell's X / Y coordinate. As another example, the entire cell may be used as the X / Y coordinate.

[0029] Tabular data may further include other information extracted from the MIF image, such as cell properties. The set of cell properties may also be characterized as a cell phenotype, and these properties may include any number of auxiliary features of the cell determined from image analysis performed by the MIF image analyzer 202A. These auxiliary features may include cell diameter, cell size, cell eccentricity, cell membrane, cell cytoplasm, and cellular biomarker positivity. Each cell in the MIF image may be associated with an immunofluorescence intensity. In embodiments, the MIF image analyzer 202A may also normalize the immunofluorescence intensities identified in the MIF image. Normalization may be performed by batch normalization of batches of immunofluorescence intensities. A batch may represent a subset of multiple immunofluorescence intensities identified in the MIF image. Immunofluorescence intensities may be used as input for training a graph, and batch normalization of intensities helps improve training efficiency by standardizing the intensities. In this embodiment, the MIF image analyzer 202A can also normalize auxiliary features to generate normalized auxiliary features.

[0030] In the embodiment, image analysis involves dividing the image into sections and labeling the divided sections based on biological regions. For example, in an MIF image containing cancerous tissue, the MIF image analyzer 202A may identify cells in the image that correspond to tumor centers or tumor stroma based on the characteristics of the cells in the image.

[0031] Based on the coordinates and properties identified by the MIF image analyzer 202A, the graph generator 204A may generate a graph representing the MIF image. The generated graph may include nodes corresponding to cells identified in the MIF image. The number of nodes may directly correspond to the number of cells identified in the MIF image. The generated graph may also include edges connecting some nodes. Nodes connected by edges may be based on a selectable distance threshold (e.g., 10 microns), where nodes within the distance threshold range are connected by edges in the graph, and nodes outside the distance threshold are not connected. In embodiments, the graph generator 204A may perform pairwise modeling between cells by identifying node pairs based on distance thresholds and by placing edges between node pairs based on the distance between nodes. The distance between nodes in the graph may be calculated by determining the distance between coordinates associated with the corresponding cells and based on whether the distance is within a predetermined threshold range.

[0032] In embodiments, the distance threshold may be determined based on the cell type of cells in the MIF image. Certain cell types may be known to have specific interactions within a certain distance range. If those cell types are identified in the MIF image, the distance threshold may be selected based on that specific distance. In embodiments, different distance thresholds may exist established when constructing the graph. For example, if cell types A and B are known to interact within a 10-micron range, and cell types A and C are known to interact within a 20-micron range, the graph generator 204A may utilize these different distances when determining whether to connect nodes in the graph. As a result, in this example, a node corresponding to cell type A may be connected to a node corresponding to cell type B only when the node is within a 10-micron range. A node corresponding to cell type A may be connected to a node corresponding to cell type C only when the node is within a 20-micron range.

[0033] In the embodiment, the distance threshold may be determined based on the biomarkers of cells in the MIF image. For example, different thresholds may exist for all pairs of intracellular biomarkers.

[0034] The generated graph shows the relationships between pairs of nodes and provides a graphical representation of the cells identified in the MIF image. The connectivity of nodes in the graph can be used to characterize cell neighborhoods.

[0035] In the embodiment, the graph generator 204A may perform a partial sampling step before the graph is converted into an embedding by the embedding generator 206A. Partial sampling may involve dividing the graph into multiple subgraphs and then feeding one or more of the subgraphs (instead of the entire graph) to the embedding generator 206A. Partial sampling of the graph may be done to improve the efficiency of calculating the embedding by allowing the embedding generator 206A to operate on a smaller portion of the graph. For example, the entire graph may be irrelevant, and therefore the embedding may be generated based only on the subgraph in question.

[0036] The embedding generator 206A trains graph embeddings (or subgraph embeddings if the graph is partially sampled) to generate embeddings that are mathematical vector representations of information in a graph, including information about each node (e.g., X / Y coordinates, biomarker representation) and the neighborhoods of each node (e.g., a subset of nodes within a certain distance range from each node). In embodiments, a graph training algorithm is applied to the graph to train the graph embeddings. The algorithm may involve selecting nodes from the graph and selecting the hop count associated with each selected node. The hop count is defined as the hop count representing the maximum number of edges that can be continuously traversed, and the center of each neighborhood, around the node in the graph that has each selected node. In this scheme, each neighborhood is a subset of nodes that includes the selected node and any nodes within the hop count range relative to the selected node. The algorithm may also involve selecting an embedding size that represents the amount of graph-related information to be held within each embedding. In embodiments, the embedding size is a fixed-length mathematical vector, such as a value of 32 or 64.

[0037] In this embodiment, the selected node and the selected hop count are inputs to a variable aggregate function. When generating an embedding for the selected node, the aggregate function extracts embedding information from the surrounding nodes and uses the extracted embedding information to generate an embedding for the selected node. As a result, the aggregate function allows each embedding to be a numerical representation of the node and its surrounding nodes (depending on the selected hop count).

[0038] In this embodiment, the graph training algorithm is an unsupervised algorithm such as DGI or Graph-Sage. The hyperparameters associated with the graph training algorithm may be selected so that the generated embeddings give an accurate representation of the nodes in the graph. Examples of hyperparameters relevant to the application of the graph training algorithm include the learning rate, dropout, optimizer, weight decay, and edge weighting. Tuning the values ​​for the hyperparameters ensures that the graph training algorithm produces an accurate embedding that represents the graph. The result of the graph training algorithm is an embedding for each node in the graph.

[0039] Embeddings may also include normalized auxiliary features if those features have been extracted from the MIF image by the MIF image analyzer 202A. Each cell identified in the MIF image has a corresponding embedding, each embedding encoding all information about the cell in a numerical format. An example of an embedding is a vector. Embeddings for nodes can be generated based on embedding information from a number of surrounding nodes. In this scheme, the embedding for a node incorporates information about the node and its neighbors.

[0040] The embeddings may also encode temporal information associated with each cell. For example, the temporal information may relate to the characteristics of cells at different points in time during a medical procedure (e.g., the second dose at week 4, the fourth dose at week 6) or during the progression of a particular disease (e.g., the patient's liver at week 1, the patient's liver at week 2). Cellular characteristics such as size or eccentricity may indicate how the cells are responding to the medical procedure or the progression of the disease. As an example, the size of cells at the second dose of a medical procedure may be compared to the size of cells at the fourth dose of a medical procedure. Changes in size (e.g., increase, decrease, no change) may be encoded in the embeddings.

[0041] The visualizer 208A provides an interactive visualization of the graph based on the embeddings generated by the embedding generator 206A. The visualizer 208A may reduce the number of dimensions of the embeddings to two or three dimensions so that the embedding information can be visually displayed on the graph or elsewhere in the interactive visualization.

[0042] Interactive visualizations may include several selectable icons, which are visual representations of the numerical information provided in the embedding. The interactive visualization is a user interface that displays the selectable icons and allows interaction with them. The arrangement and visual configuration of the icons are based on the embedding. Therefore, each selectable icon in the interactive visualization can be considered to correspond to a specific node in the graph generated by the graph generator 204A and its vicinity. In some embodiments, the interactive visualization is a two-dimensional representation of the embedding.

[0043] The visual properties of selectable icons can be adjusted to reflect different encoded information in the embedding. Selectable icons may have different colors, sizes, shapes, or borders, each of which may be configured to reflect a particular property of the embedding and its corresponding node in the graph. For example, the icon color may be used to reflect different cell phenotypes (e.g., a blue icon may be associated with cell phenotype A, and a red icon with cell phenotype B), and the icon border may be used to reflect biomarkers associated with the cell (e.g., a thick border may be associated with biomarker A, and a thin border with biomarker B). Interactive visualization capabilities may include enabling, querying, and filtering selectable icons, plotting / graphing embeddings, generating neighborhood graphs, and generating statistical summaries. Users may provide a selection of a subset of selectable icons by using the cursor to draw a lasso tool over the icons or by drawing a box. The selected subset of selectable icons represents the cell neighborhoods corresponding to the embeddings associated with the selectable icons. Interactive visualization may provide a graphical plot containing a subset of selectable icons, which can be configured so that the plot represents cell neighborhoods. The graphical plot may display information about the selected subset, including highlighting selected cells within the subset; an infographic with details about each of the selected cells; and other information given to the embedding of the selected cells. The infographic may include biomarker information, temporal information, and labeling information.

[0044] In embodiments, when temporal information is encoded in the embedding, interactive visualization can provide information about the characteristics of each cell at different points in time, such as different medical treatments and different weeks. Interactive visualization can provide filters that allow selection of these different times and comparison of cells at those selected times.

[0045] Figure 2B shows an embodiment of the MIF system pipeline 200B for processing MIF images, which may be distributed across multiple computers in a network or cloud-based environment. For example, the MIF system pipeline 200B may be implemented as part of a cloud-based environment, and as components of the MIF system pipeline 200B distributed across multiple computers connected to each other within the environment. The MIF system pipeline 200B may include an MIF image analyzer 202B, a graph generator 204B, an embedding generator 206B, and a visualizer 208B. These components may perform the same functions as the respective components of the MIF system pipeline 200A described above, but may be implemented on the same or different computers and connected within the MIF system pipeline 200B via network connections 210A-210C in a network or cloud-based environment.

[0046] In some embodiments, the MIF image analyzer 202B, graph generator 204B, embedding generator 206B, and visualizer 208B may be run on different computers. The MIF image analyzer 202B may communicate the results of its image analysis to the graph generator 204B via network connection 210A. Similarly, the graph generator 204B may communicate the trained graph generated based on the image analysis results to the embedding generator 206B via network connection 210B. Furthermore, the embedding generator 206B may communicate the embeddings generated from the trained graph to the visualizer 208B via network connection 210C.

[0047] In some embodiments, one or more components may be implemented on the same device within a network or cloud-based network. For example, the MIF image analyzer 202B may be implemented on the same device as the graph generator 204B, and the embedding generator 206B may be implemented on the same device as the visualizer 208B.

[0048] Further details regarding interactive visualization are explained in Figures 4A and 4B.

[0049] Figure 3 shows Method 300 as an example for providing interactive visualization of multiplex immunofluorescence images according to several embodiments. As a non-limiting example with reference to Figure 2, one or more processes described with reference to Figure 3 may be performed by an MIF system (e.g., the MIF system pipeline 200 in Figure 2) to transform the MIF images into interactive visualization of the cellular environment shown in the MIF images. In such embodiments, the MIF system pipeline 200 may execute code in memory to perform certain steps of Method 300. Although Method 300 is described below as being performed by the MIF system pipeline 200, other devices may store the code and thus perform Method 300 by directly executing the code. Accordingly, the following description of Method 300 simply refers to Figure 2 as an exemplary non-limiting embodiment of Method 300. For example, Method 300 may be performed on any computing device, such as a computer system and / or hardware (e.g., circuits, proprietary logic, programmable logic, microcode, etc.), software (e.g., instructions to be executed on a processing device), or a combination thereof, as described, for example, with reference to Figure 5. Furthermore, it should be understood that not all steps are required to carry out the disclosures given herein. Additionally, some steps may be performed simultaneously or in a different order than those shown in Figure 3, as will be understood by those skilled in the art.

[0050] In 310, the MIF image analyzer 202 performs image analysis on the MIF image, which includes identifying cells within the MIF image. The result of this analysis is tabular data representing information about each identified cell.

[0051] In 320, the MIF image analyzer 202 can extract cell coordinates and biomarker information from tabular data. In some embodiments, cell coordinates may include the X / Y coordinates of each cell, and biomarker information may include the immunofluorescence intensity associated with all cells. In some embodiments, the extracted information may further include the auxiliary features described above.

[0052] In step 330, the graph generator 204 constructs a graph based on the extracted information. In some embodiments, the graph generator 204 may construct a graph using information about each node, where the nodes in the graph correspond to identified cells in the image. The graph facilitates the detection of relationships between cells in the image by representing the image data as more easily processed information about each identified cell in the image. In some embodiments, the graph connects nodes based on a distance threshold, where the size of the neighborhood in the graph is defined by the number of hops from a particular node. For example, if the number of hops is set to 3, the neighborhood of a particular node includes all nodes in the graph that are 3 hops (or connected) from that node. At this stage, each node in the graph contains information corresponding to a single cell in the MIF image.

[0053] In 340, the embedding generator 206 trains a graph and generates embeddings from the trained graph. Training the graph results in each node containing information about the corresponding cell and its neighbors. Different machine learning algorithms may be used to train the graph. In some embodiments, the embedding generator 206 may use an unsupervised or self-supervised machine learning algorithm to train the graph to identify embeddings in the current graph using unlabeled data. In other embodiments, the embedding generator 206 may use a supervised machine learning algorithm to train the current graph using a previously labeled graph and identify embeddings based on the previously labeled graph.

[0054] In embodiments utilizing unsupervised or self-supervised machine learning algorithms, the embedding generator 206 may use correlation rules to discover patterns or relationships between features of neighboring cells (also called "nodes") within a predefined distance threshold range. Trainable parameters may exist within the trainable weight matrix applied to the features, but there are no target predictions defined, for example, by a previously labeled graph. The advantage of unsupervised or self-supervised machine learning in these embodiments is that larger amounts of unlabeled data can be leveraged for pre-training, thereby increasing the likelihood that the generated embeddings for the target graph contain more useful domain information.

[0055] In embodiments utilizing supervised machine learning algorithms, the embedding generator 206 may use classification algorithms to recognize and group cells based on previously labeled graphs or data. The advantages of supervised machine learning in these embodiments are that node classification may be more accurate, and graphs can be trained to satisfy specific classifications (based on previously labeled data).

[0056] Embeddings are generated to capture information about all cells and neighboring cells identified within the MIF image. Nodes in the graph represent cells and related information, including biomarker representations of cells. Every node is associated with an embedding that encodes information about the cell and its neighborhood (defined by the number of selectable hops from the cell). For example, information about the cell and its neighborhood, such as cell features, neighbor cell features, neighbor cell identification, and connectivity structures, can be aggregated together into a single feature vector using a trainable weight matrix. In some embodiments, this feature vector constitutes the embedding.

[0057] In 350, the visualizer 208 generates an interactive visualization into which selectable icons are populated based on the embeddings. The interactive visualization provides an interface for receiving input to manipulate the selectable icons. The visualizer 208 may apply learning techniques to reduce the dimensionality of the embeddings. In embodiments, the embeddings may be reduced to a two-dimensional visual representation (such as a plot). Current examples of such learning techniques include homogeneous manifold approximation and projection (UMAP) and t-distributed stochastic neighbor embeddings (t-SNE). The visualizer 208 may, in some cases, partially sample the embeddings used to generate the visualization when it is not practical to perform a dimensionality reduction algorithm on all embeddings due to computational complexity. The partial sampling algorithm may be random, stratified random, or any alternative method.

[0058] In 360, the MIF system pipeline 200 may additionally supply embeddings to a neural network machine learning algorithm to generate predictive models for therapeutic response and survival predictions. The output of the machine learning algorithm may be a prediction of how a patient may respond to a particular treatment (at the cellular level) or about their chances of survival. Information within the embedding, including any temporal information, may be used as input to the machine learning algorithm. In embodiments, the machine learning algorithm is supervised. The embedding may be provided as input to a neural network, which may, based on the embedding, generate predictions associated with cells identified in the MIF image. In embodiments, the predictions relate to predicted medical or therapeutic outcomes associated with cells. For example, the MIF image may show cancer cells, and the machine learning algorithm may use the embedding to predict the outcome of the cells. The predictions may take the form of interactive visualizations updated with selectable cell visual properties updated to reflect the predicted outcomes. In other words, the MIF image may show cells associated with a medical condition, and the predicted outcome is about the medical condition. In another embodiment, the predictions may relate to predicted responses associated with multiple cells. For example, an MIF image may show cells being treated for cancer, and a machine learning algorithm may use the embedding to predict how the cells might respond to the treatment. The prediction may take the form of an interactive visualization updated with selectable cell visual properties updated to reflect the predicted treatment response. In other words, an MIF image may show cells associated with a medical condition, and the predicted response is about a patient's response to treatment for that medical condition.

[0059] In embodiments, applying a machine learning algorithm may involve an embedding aggregation procedure, or selecting how information within an embedding, including temporal information, should be processed into a single vector representation. Examples of procedures include averaging the embedding information, providing cell-by-cell predictions, providing predictions based on a region (or group of cells) of interest, or taking the average or maximum value of a grid of cells. In another embodiment, the machine learning algorithm may include a multilayer perceptron (MLP) classifier or a convolutional neural network.

[0060] In some embodiments, the performance of a predictive model can be evaluated by comparing the prediction with the actual results or patient responses. For example, the prediction of a patient's response to a particular medical procedure may be compared with the actual response, and this comparison may be used to further refine the predictive model to improve accuracy.

[0061] Figure 4A shows a process flow 400A as an example of utilizing interactive visualization 410 of multiplex immunofluorescence images in several embodiments.

[0062] The interactive visualization 410 may provide an interface for graph queries 411. In embodiments, graph queries 411 may involve receiving graph queries associated with nodes in the graph as input. Queries may focus on searching for specific cells or cell neighborhoods in the graph that match search criteria. Queries may include parameters or thresholds for identifying nodes in the graph. For example, a query may be used to search for cells with cell neighborhoods that have a certain biomarker expressed by a threshold of 50% or more of the cells. Graph queries 411 may result in displaying the results of the query, such as cells or cell neighborhoods that match the query. In embodiments, matching cells or cell neighborhoods may be highlighted in the interactive visualization 410.

[0063] The interactive visualization 410 may also provide an interface for performing filtering 412, which may include data parameters for filtering selectable icons within the interactive visualization 410. Examples of parameters include whole-slide imaging (WSI), patient, label, and temporal information such as pre- and post-operative and response to treatment. Parameters may be provided by drop-down boxes. The interactive visualization 410 filters the selectable icons by displaying only those selectable icons that match the selected option. For example, a specific digital file (WSI) may be selected, and some examples include a specific patient (or multiple specific patients), one or more labels assigned to nodes in the graph, a label identifying cells from a pre-operative patient, a label identifying cells from a post-operative patient, a label identifying cells from a pre-treatment patient, a label identifying cells from a post-treatment patient, and a response label identifying patients who responded to or did not respond to a particular treatment. Each of these parameters may have additional sub-options. For example, patients may be divided into different categories, such as patients who responded to a particular treatment and patients who did not.

[0064] The interactive visualization 410 may also provide an interface for plots or graph embeddings 413, which involves projecting the embeddings onto a two-dimensional plot and supplying the two-dimensional plots to an interface that can receive user input for manipulation and interaction. Icons in the two-dimensional plot (corresponding to nodes in the graph represented by the embeddings) may be configured to visually represent the properties of cells and their neighborhoods. Examples include different colors, different shapes, and different sizes. In embodiments, icons may be configured such that the interface provides a heatmap (e.g., a change in color) to represent temporal information about the corresponding cells. For example, a heatmap may be used to show changes in cells from before to after surgery. The interface may allow box / lasso selection input to select a subset of selectable icons. The interface may also allow gradients on the plot to be highlighted to indicate temporal information or cell neighborhood heterogeneity. The interface may also provide options for the user to input commands for manipulating groups of icons in the interactive visualization 410, such as panning, rotating around an axis, or zooming to a particular cell or cell neighborhood.

[0065] The interactive visualization 410 may also provide an interface for graphing cell neighborhoods 414. After selecting a subset of selectable icons (e.g., by box / lasso selection), the identifier of the corresponding cell may be determined based on the embedding associated with the selected subset. These identifiers may be used to generate a plot of cell neighborhoods for the selected subset. In embodiments, graphing cell neighborhoods 414 may involve configuring the neighborhood plot to represent the selectable icons based on the corresponding biomarkers (or other cellular information) for the cells corresponding to the icons shown in the interactive visualization 410, using different visual properties (e.g., borders, colors, shapes).

[0066] An example of a cell proximity plot is explained with respect to Figure 4B.

[0067] The interactive visualization 410 may also provide an interface for generating a statistical summary 415. After receiving a selection of selectable icons (e.g., box / lasso selection) from the interactive visualization 410, the interface may generate a statistical summary associated with the selected icons. For example, the interface may report the number of cells in the selection that have an immunofluorescence intensity above a certain threshold. The interface may also report the ratio of icons in the selection that exceed the threshold for each biomarker. In another embodiment, the interface may report the average immunofluorescence intensity per biomarker for the selected icons.

[0068] Other examples of statistical summaries include reporting biomarker proportions and counts, providing subgraph (e.g., cell neighborhood) densities, ranking biomarker vectors / cell phenotypes based on prevalence (e.g., showing counts and / or proportions) within selected icons, providing temporal reports such as summaries showing before and after images of selected icons (e.g., comparing previous cell information with current cell information), and providing comparisons between neighbors by highlighting changes in neighbor biomarker expression, neighborhood density, or cell phenotype.

[0069] Figure 4B shows an example process flow 400B for generating cell vicinity plots 420 from interactive visualization 410 of multiplex immunofluorescence images, according to several embodiments.

[0070] The cell neighborhood plot 420 may include a plot 421 showing the neighborhood around the target cell 427 and a legend providing information about visual indicators of icons within the plot. In this embodiment, the hop count is any selectable value from 1 to 6, which represents all cells within 1 to 6 hops from the target cell 427. The visual indicators may show properties of the corresponding cells. For example, the legend may associate cell color 422 with cell phenotype, boundary width 423 with biomarker information, and shape / boundary filter 424 with multiple cell phenotypes.

[0071] The icons in the cell neighborhood plot 421 may also be interactive and selectable. Similar to the icons in the interactive visualization 410, the icons in the cell neighborhood plot 421 also correspond to specific trained embeddings (representing the cell and its surrounding neighborhood). Interacting with the icons may provide information about the corresponding node and neighborhood.

[0072] The cell proximity plot 420 may also provide a view based on filtering visual indicators such as the cell phenotype view 425 or the biomarker view 426.

[0073] Examples of computing systems The various embodiments and / or components described herein may be implemented using one or more computer systems, such as the computer system 500 shown in Figure 5. The computer system 500 may be any computer or computing device capable of performing the functions described herein. For example, one or more computer systems 500 may be used to implement any embodiment and / or any combination or partial combination thereof of Figures 1 to 4.

[0074] The following computer systems or multiple instances thereof, as examples, may be used to implement, according to several embodiments, the method 300 in Figure 3, the system shown in Figure 2, or any component thereof.

[0075] Various embodiments may be implemented using one or more well-known computer systems, such as the computer system 500 shown in Figure 5. One or more computer systems 500 may be used, for example, to implement any of the embodiments described herein, as well as combinations and partial combinations thereof.

[0076] The computer system 500 may include one or more processors (also called a central processing unit, i.e., a CPU), such as a processor 504. The processor 504 may be connected to a bus or a communication infrastructure 506.

[0077] The computer system 500 may also include user input / output devices 505 such as a monitor, keyboard, and pointing device, which can communicate with a communication infrastructure 506 through a user input / output interface 502.

[0078] One or more of the processors 504 may be graphics processing units (GPUs). In embodiments, the GPU may be a processor which is a dedicated electronic circuit designed to process mathematically intensive applications. The GPU may have a parallel structure that is efficient for parallel processing of large data blocks such as mathematically intensive data common to computer graphics applications, images, videos, vector processing, array processing, etc., and for cryptography, including, for example, brute-force cracking, generating cryptographic hashes or hash sequences, solving partial hash inversion problems, and / or producing the results of other proof-of-work calculations for some blockchain-based applications. Due to the general-purpose computing capabilities on graphics processing units (GPGPUs), GPUs may be particularly useful in at least the feature extraction and machine learning embodiments described herein.

[0079] Additionally, one or more of the processors 504 may include other embodiments of a coprocessor or logic for accelerating cryptographic computations or other specialized mathematical functions, including a hardware accelerated cryptographic coprocessor. Such an accelerating processor may further include an instruction set for acceleration, which uses the coprocessor and / or other logic to facilitate such acceleration.

[0080] The computer system 500 may also include main memory or primary memory 508, such as random access memory (RAM). The main memory 508 may include one or more levels of cache. The main memory 508 may store control logic (i.e., computer software) and / or data.

[0081] The computer system 500 may also include one or more secondary storage devices or secondary memory 510. The secondary memory 510 may include, for example, a main storage drive 512 and / or a removable storage device or drive 514. The main storage drive 512 may be, for example, a hard disk drive or a solid-state drive. The removable storage drive 514 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.

[0082] The removable storage drive 514 may interact with the removable storage unit 518. The removable storage unit 518 may include a computer-enabled or readable storage device on which computer software (control logic) and / or data are stored. The removable storage unit 518 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 514 may read from and / or write to the removable storage unit 518.

[0083] The secondary memory 510 may include other means, devices, components, tools, or other methods for making computer programs and / or other instructions and / or data accessible by the computer system 500. Such means, devices, components, tools, or other methods may include, for example, a removable storage unit 522 and an interface 520. Examples of the removable storage unit 522 and interface 520 may include a program cartridge and cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.

[0084] The computer system 500 may further include a communication or network interface 524. The communication interface 524 may enable the computer system 500 to communicate with and interact with any combination of external devices, external networks, external entities, etc. (referenced individually and collectively by reference number 528). For example, the communication interface 524 may enable the computer system 500 to communicate with an external or remote device 528 via a communication path 526. The communication path 526 may be wired and / or wireless (or a combination thereof) and may include any combination of LAN, WAN, Internet, etc. Control logic and / or data may be transmitted to and from the computer system 500 via the communication path 526.

[0085] The computer system 500 may also be, to give some non-limiting examples, a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smartwatch or other wearable, an electrical appliance, part of the Internet of Things (IoT), and / or an embedded system, or any combination thereof.

[0086] The computer system 500 may be a client or server that accesses or hosts any application and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions, local or on-premises software (e.g., “on-premises” cloud-based solutions), “as a service” models (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS), Database as a Service (DBaaS), etc.), and / or hybrid models that include any combination of the aforementioned examples or other service or delivery paradigms.

[0087] Any applicable data structure, file format, and scheme may be derived from standards including, but not limited to, JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), MessagePack, XML User Interface Language (XUL), or any other functionally similar expression, either alone or in combination. Alternatively, a proprietary data structure, format, or scheme may be used exclusively or in combination with known or open standards.

[0088] Any suitable data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in human-readable formats such as numeric, text, graphic, or multimedia formats, further including various types of markup languages ​​among other possible formats. Alternatively, or in combination with the above formats, data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted formats, or in any other machine-readable format.

[0089] Interfaces or interconnections between various systems and layers may employ any number of mechanisms, including but not limited to, open or proprietary mechanisms capable of achieving similar functions and results, such as the Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Pattern (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG), HTML5 Web Messaging, Representational State Transfer (REST) ​​or RESTful Web Services, Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other open or proprietary mechanisms.

[0090] Such interfaces or interconnections may also utilize uniform resource identifiers (URIs), which may further include uniform resource locators (URLs) or uniform resource names (URNs). Other forms of uniforms and / or unique identifiers, locators, or names may be used exclusively or in combination with the forms described above.

[0091] Any of the protocols or APIs described above may interface with, or be implemented using, any procedural, functional, or object-oriented programming language, and may be compiled or sequentially interpreted. Non-limiting examples include, but are not limited to, Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Express.js, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, and many other non-limiting examples, C, C++, C#, Objective-C, Java, Swift, Go, Ruby, Perl, Python, JavaScript, WebAssembly, or substantially any other language in any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, together with any other libraries or schemes.

[0092] In some embodiments, a tangible non-temporary device or product, including a tangible non-temporary computer-usable or readable medium on which control logic (software) is stored, may also be referred to herein as a computer program product or program storage device. This includes, but is not limited to, tangible products embodying the computer system 500, the main memory 508, the secondary memory 510, and the removable storage units 518 and 522, as well as any combination thereof. Such control logic may be executed by one or more data processing devices (such as the computer system 500) to cause such data processing devices to operate as described herein.

[0093] Based on the teachings contained herein, it will be apparent to those skilled in the art how to create and use embodiments of the disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 5. In particular, embodiments may operate with software, hardware, and / or operating system embodiments other than those described herein.

[0094] conclusion It should be understood that the provision describing the modes for carrying out the invention, rather than any other provision, is intended to be used in interpreting the claims. Other provisions may describe one or more exemplary embodiments, but not all, that the inventors may consider, and are therefore not intended to limit the scope of this disclosure or the appended claims in any way.

[0095] While this disclosure describes exemplary embodiments for exemplary fields and uses, it should be understood that this disclosure is not limited thereto. Other embodiments and modifications thereof are possible and fall within the scope and spirit of this disclosure. For example, without limiting the universality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or entities shown in the drawings and / or described herein. Furthermore, embodiments (whether expressly described herein or not) may have significant utility in fields and uses other than those described herein.

[0096] Embodiments are described herein with the help of function-building blocks that illustrate embodiments of specific functions and their relationships. The boundaries of these function-building blocks are arbitrarily defined herein for the sake of clarity. Alternative boundaries may be defined insofar as the specific functions and relationships (or their equivalents) are adequately performed. Furthermore, alternative embodiments may perform function-building blocks, steps, operations, methods, etc., in a different order than those described herein.

[0097] References to “one embodiment,” “embodiment,” “example embodiment,” “several embodiments,” or similar phrases herein indicate that the embodiments described may include certain features, structures, or characteristics, but not all embodiments necessarily include certain features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same embodiment.

[0098] Furthermore, when certain features, structures, or characteristics are described in relation to embodiments, it is assumed that incorporating such features, structures, or characteristics into other embodiments, whether or not they are explicitly mentioned or described herein, is within the knowledge of those skilled in the art. In addition, some embodiments may be described using the expressions “connected” and “linked” together with their derivatives. These terms are not necessarily intended to be synonymous with each other. For example, some embodiments may be described using the terms “linked” and / or “linked” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “linked” may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.

[0099] The scope and width of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely by the following claims and their equivalents.

Claims

1. A method for characterizing intercellular interactions from multiplex immunofluorescence imaging, Identifying and determining multiple cells in the multiplex immunofluorescence image, and determining the coordinates and properties of each identified cell among the multiple cells, wherein the multiple cells include at least a first cell and a second cell. The method for generating a graph of the multiplex immunofluorescence image based on the coordinates and properties of each of the multiple cells, wherein the graph includes a plurality of nodes representing cells identified from the plurality of cells, and a plurality of edges connecting each pair of nodes from the plurality of nodes, wherein the plurality of nodes include at least a first node representing the first cell and a second node representing the second cell, and the plurality of edges include at least edges between the first node and the second node. The transformation involves converting the graph into multiple embeddings, wherein the multiple embeddings include node embeddings of the multiple nodes representing cells identified from the multiple cells. Using dimensionality reduction techniques, the number of dimensions of the node embedding is reduced to either two or three dimensions. To provide an interactive visualization that includes, for each node representing a cell identified from the aforementioned plurality of cells, a plurality of selectable icons representing each dimensionality-reduced node embedding, Methods that include...

2. The method according to claim 1, further comprising generating normalized immunofluorescence intensities for a plurality of cells using batch normalization of a batch of immunofluorescence intensities, wherein the generated properties for each of the plurality of cells include the normalized immunofluorescence intensities.

3. The method according to claim 1, further comprising generating the graph by setting the plurality of edges based on an evaluation of the distance between the respective coordinates of the plurality of nodes.

4. The method according to claim 1, wherein determining each of the properties of the first cell comprises extracting an auxiliary feature of the first cell, the auxiliary feature comprising at least one of the diameter of the first cell, the area of ​​the first cell, or the eccentricity of the first cell.

5. The method according to claim 1, wherein determining the respective properties of each identified cell of the plurality of cells comprises generating normalized auxiliary features using normalization of a plurality of auxiliary features of the plurality of cells, and the plurality of embeddings are generated based in part on the normalized auxiliary features.

6. Converting the aforementioned graph into the aforementioned multiple embeddings is Selecting a number of hops to define the neighborhood of a selected node in the graph, wherein the neighborhood of the selected node includes a subset of the plurality of nodes, which includes the selected node and any node of the plurality of nodes that are within the range of the number of hops relative to the selected node. Selecting an embedding size that defines the amount of information associated with the graph to be held within the aforementioned multiple embeddings, The method according to claim 1, further comprising:

7. Converting the aforementioned graph into the aforementioned multiple embeddings is Applying the training algorithm to the graph based on the number of hops and the embedding size, Based on applying the aforementioned training algorithm, the plurality of embeddings are generated, The method according to claim 6, further comprising:

8. The method according to claim 7, wherein the training algorithm is an unsupervised graph training algorithm, and the application of the training algorithm to the graph is further based on at least one selected hyperparameter.

9. The interactive visualization is configured to receive a selection of a subset of selectable icons representing dimensionally reduced node embeddings for a subset of nodes representing a subset of the plurality of cells, wherein the subset of the plurality of cells is associated with one or more cell neighborhoods, and the method The method according to claim 1, further comprising providing a graphical plot of the vicinity of one or more cells.

10. The aforementioned interactive visualization, Receiving queries relating to at least a subset of the properties of each of the identified cells represented by the plurality of nodes, wherein the queries include a threshold for identifying a subset of the plurality of nodes in the graph. Filtering the plurality of selectable icons in the interactive visualization based on data parameters, wherein the data parameters include at least one of patient, response label, image label, or cell type. Receiving the selection of the subset of the multiple selectable icons within the interactive visualization, Receiving commands for operating the aforementioned interactive visualization, or To provide a statistical summary of biomarkers or cellular phenotypes associated with the aforementioned plurality of selectable icons, The method according to claim 1, configured to perform at least one of the following.

11. The aforementioned multiple embeddings are provided as input to the neural network, The neural network generates predictions associated with the plurality of cells based on the plurality of embeddings, wherein the prediction is one of the prediction results associated with the plurality of cells or the prediction response associated with the plurality of cells. The method according to claim 1, further comprising:

12. The method according to claim 11, wherein the multiplex immunofluorescence image is associated with a medical condition, and the prediction result is for the medical condition.

13. The method according to claim 11, wherein the multiplex immunofluorescence image is associated with a medical condition, and the predicted response is a patient response to treatment for the medical condition.

14. Before converting the aforementioned graph into the aforementioned multiple embeddings, The aforementioned graph is partially sampled into multiple subgraphs, Converting the aforementioned multiple subgraphs into the aforementioned multiple embeddings, The method according to claim 1, further comprising:

15. A computing system, At least one non-temporary computer-readable medium, At least one processor, The computing system comprises, and the program instructions stored in the at least one non-temporary computer-readable medium, wherein when the program instructions are executed by the at least one processor, the computing system In a multiplex immunofluorescence image, multiple cells are identified, and for each identified cell among the multiple cells, the coordinates and properties of each cell are determined, and the multiple cells include at least a first cell and a second cell. A graph of the multiplex immunofluorescence image is generated for each of the multiple cells based on the respective coordinates and properties of each of the multiple cells, wherein the graph includes a plurality of nodes representing cells identified from the plurality of cells, and a plurality of edges connecting each pair of nodes from the plurality of nodes, wherein the plurality of nodes includes at least a first node representing the first cell and a second node representing the second cell, and the plurality of edges includes at least edges between the first node and the second node. The graph is converted into multiple embeddings, the multiple embeddings including node embeddings of the multiple nodes representing cells identified from the multiple cells, Using dimensionality reduction technology, the number of dimensions of the node embedding is reduced to either two or three dimensions. For each node representing a cell identified from the aforementioned group of cells, an interactive visualization is provided that includes multiple selectable icons, each representing a dimensionally reduced node embedding. Computing system.

16. A non-temporary computer-readable medium on which instructions are stored, wherein, when the instructions are executed by at least one processor, the computing system... Identifying and determining multiple cells in a multiplex immunofluorescence image, determining the coordinates and properties of each identified cell among the multiple cells, wherein the multiple cells include at least a first cell and a second cell. The method for generating a graph of the multiplex immunofluorescence image based on the coordinates and properties of each of the multiple cells, wherein the graph includes a plurality of nodes representing cells identified from the plurality of cells, and a plurality of edges connecting each pair of nodes from the plurality of nodes, wherein the plurality of nodes include at least a first node representing the first cell and a second node representing the second cell, and the plurality of edges include at least edges between the first node and the second node. The transformation involves converting the graph into multiple embeddings, wherein the multiple embeddings include node embeddings of the multiple nodes representing cells identified from the multiple cells. Using dimensionality reduction techniques, the number of dimensions of the node embedding is reduced to either two or three dimensions. To provide an interactive visualization that includes, for each node representing a cell identified from the aforementioned plurality of cells, a plurality of selectable icons representing each dimensionality-reduced node embedding, A non-temporary computer-readable medium that enables the execution of actions including [specific actions].

17. The method according to claim 1, wherein identifying the plurality of cells in the multiplex immunofluorescence image includes applying cell division to the multiplex immunofluorescence image.

18. The method according to claim 1, wherein each of the node embeddings encodes information about an identified cell, the information comprising (i) the coordinates of each of the identified cells, (ii) at least a subset of the properties of each of the identified cells, and (iii) neighborhood information about the identified cell.

19. The method according to claim 1, wherein each of the plurality of selectable icons has a visual property determined based on the respective properties of the identified cell.

20. The method according to claim 19, wherein the visual property determined based on each of the properties of the identified cell includes at least one visual property determined based on each of the cellular phenotypes of the identified cell.

Citation Information

Patent Citations

  • Improved Differentiation Methods

    JP2019527068A

  • Systems and methods for analyzing mixed cell populations

    JP2020527946A

  • Classification and mutation prediction from histopathology images using deep learning

    US20200184643A1