Context-dependent spatial multicellular network motifs for single-cell spatial biology
By employing context-dependent spatial motifs and machine learning, the method addresses the challenge of inferring high-dimensional intercellular interactions in tissue organization, providing accurate disease prediction and spatial context analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BG NEGEV TECHNOLOGIES & APPLICATIONS LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
AI Technical Summary
Traditional tissue-scale bulk readouts and single-cell omics techniques fail to explain tissue-scale cell organization due to masking of single-cell type and state heterogeneity, and high-dimensional intercellular interactions are difficult to infer systematically.
Adaptation and extension of network motifs to characterize fine-scale multicellular spatial networks, using context-dependent spatial motifs to identify discriminative motifs associated with disease states, and applying machine learning for prediction and interpretation.
Discriminative motifs provide a higher layer of abstraction in multicellular organization, enabling accurate disease state prediction and revealing spatial contexts associated with disease, with machine learning validation outperforming state-of-the-art methods.
Smart Images

Figure IL2025051041_28052026_PF_FP_ABST
Abstract
Description
CONTEXT-DEPENDENT SPATIAL MULTICELLULAR NETWORK MOTIFS FOR SINGLE-CELL SPATIAL BIOLOGYCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is a PCT International Application claiming the benefit of priority of U. S. Patent Application No. 63 / 723,204, titled: “CONTEXT-DEPENDENT SPATIAL MULTICELLULAR NETWORK MOTIFS FOR SINGLE-CELL SPATIAL BIOLOGY” filed November 21, 2024, which is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates generally to determining and interpreting a condition of a biological tissue sample. More specifically, the present invention relates to using context-dependent spatial multicellular network motifs for analyzing single-cell spatial biology.BACKGROUND
[0003] The clinical state of diseased tissue may be caused by complex intercellular processes that go beyond pairwise cell-cell interactions and are difficult to infer due to the combinatorial explosion of such high-dimensionality. Traditional tissue-scale bulk readouts cannot explain tissue scale cell organization because population averages mask the single cell types and cell states heterogeneity. Single cell omics techniques overcome these limitations but cannot resolve the information regarding the tissue's spatial organization.
[0004] Recent spatial single cell omics techniques enable determination of cell type and state for every cell and mapping back to physiological tissue context in situ. These advances bring technical challenges of systematically representing and bridging the different spatial scales of tissue organization from individual cells to physiologically meaningful tissue state.SUMMARY
[0005] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0006] Most analysis efforts are currently focused on measuring cell type proportions, pairwise interactions between cell types, cellular microenvironments and graph modeling approaches. However, an intermediate spatial scale of tissue organization may exist involving a few locally interacting cells that define modular components of different clinical states of diseased tissues. Such modular components may define a new layer of abstraction that can help reveal biologically meaningful higher-order intercellular interactions that are associated with tissue state, thus bridging the gap between pairwise interactions and tissue organization.
[0007] The number of potential high-dimensional intercellular interactions grows exponentially as a function of the number of cells involved in these interactions, and thus are difficult to explicitly infer systematically. Embodiments of the invention may address these challenges by adapting and extending the concept of "network motifs", reoccurring sub-networks that define "building blocks" of complex networks, to characterize the fine-scale organization of multicellular spatial networks of human disease tissues.
[0008] Network motifs have been established in transcription networks of multiple experimental systems and organisms, suggestive of their role as modular and universal building blocks of transcription networks that are repeatedly selected by evolution for their biological function. Motifs also appeared in other biological and non-biological networks. Embodiments of the invention may introduce context-dependent identification of spatial motifs that identifies "discriminative" motifs in the spatial network that are associated with biological conditions such as disease states.
[0009] For example, embodiments of the invention may demonstrate that modular structures composed of as few as 3-5 cells and their relative spatial organization can encode differences in clinical disease states in cohorts of triple-negative breast cancer and melanoma patients. Machine Learning (ML) validation may indicate that discriminative motifs outperform state-of-the-art methods for disease state prediction while enabling interpretation of which interactions in what spatial context are associated with these predictions.
[0010] Context-dependent discriminative motifs may define a higher layer of abstraction and modularity in multicellular organization and function with broad applicability in the domain of spatial single cell omics and beyond. The discriminative motifs may be used as features for accurate machine learning prediction, and their localization in the tissue may reveal the spatial contexts for which motifs are associated with the disease state.
[0011] Embodiments of the invention may reveal that the specific relative spatial organization of cells within the motif contributes to disease discrimination beyond the motifs' cell type composition. The same patient-derived motifs may be differentially analyzed under varying disease contexts.
[0012] For example, in metastasis-free sentinel lymph nodes of stage II melanoma patients, the inventors have demonstrated that embodiments may identify discriminative motifs consisting of B cells and CD8 T cells that appear in the outer layer of B cells surrounding the germinal centers that progressed in five years to remote metastases. In triple-negative breast cancer patients, embodiments may identify discriminative motifs consisting of interactions between tumor cells and CD8 T cells or CD4 T cells that are associated with long-term survival.
[0013] Context-dependent discriminative motifs may therefore define a new spatial layer of abstraction and modularity in multicellular function with broad applicability in the domain of spatial single cell analysis and beyond.
[0014] The foregoing general description of the illustrative embodiments and the following detailed description thereof are merely exemplary aspects of the teachings of this disclosure and are not restrictive.
[0015] According to an aspect of the present invention, a method of determining a condition of a biological tissue sample may be provided. The method may include receiving a target image depicting spatial single-cell omics data of cells in the tissue sample. The method may include constructing a multicellular network of interconnected nodes, wherein each node may represent a cell in the target image. The method may include enumerating subgraphs within the multicellular network to identify a set of spatial motifs, wherein each spatial motif may represent a multicellular pattern of interconnected cells of specific types. The method may include analyzing the set of spatial motifs to generate one or more motif-based features, characterizing the biological tissue sample. The method may include applying one or more trained machine learning (ML) models to the motif-based features to predict the condition of the biological tissue sample.
[0016] According to other aspects of the present invention, the method may include one or more of the following features. Constructing the multicellular network may include segmenting the target image to identify individual cells and determine spatial coordinates for each cell. The method may include classifying the individual cells according to their respective cell types based on the spatial single-cell omics data. The method may include representing each cell as a node inthe multicellular network, wherein each node may be attributed with its respective cell type and spatial location. The method may include connecting spatially adjacent cells with edges based on spatial proximity, wherein spatial proximity may be determined using a predetermined neighborhood criterion.
[0017] According to some embodiments, enumerating subgraphs within the multicellular network may include examining subgraphs that may include predetermined numbers of interconnected nodes within the multicellular network. For each examined subgraph, the method may include determining a canonical pattern based on the cell types of the nodes and their connectivity structure. The method may include calculating a frequency of occurrence for each canonical pattern across the multicellular network. The method may include selecting canonical patterns whose frequency may exceed a predetermined threshold, to obtain the set of spatial motifs.
[0018] According to some embodiments, the method may include obtaining a context-dependent subset of the set of spatial motifs, wherein the subset may include discriminative motifs which may distinguish between different biological conditions pertaining to that context. The method may include providing calculated frequencies of the context-dependent discriminative motifs as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
[0019] According to some embodiments, the method may include obtaining a context-dependent subset of the set of spatial motifs, wherein the subset may include discriminative motifs which may distinguish between different biological conditions pertaining to that context. The method may include mapping occurrence of discriminative motifs within the multicellular network. Based on the mapping, the method may include calculating spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample. The method may include providing the calculated spatial distribution patterns as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
[0020] According to some embodiments, the method may include obtaining a context-dependent subset of the set of spatial motifs, wherein the subset may include discriminative motifs which may distinguish between different biological conditions pertaining to that context. For one or more discriminative motifs, the method may include (a) examining pairwise connections between cells of specific types within that discriminative motif, and (b) for each pairwiseconnection, calculating a pairwise frequency value, representing frequency of occurrence of that pairwise connection in the multicellular network. The method may include providing the calculated pairwise frequency values as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
[0021] According to some embodiments, obtaining the context-dependent discriminative motifs may include receiving a training dataset that may include a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image may be annotated according to a biological condition pertaining to that context. The method may include identifying a preliminary set of spatial motifs based on the plurality of training images. The method may include performing an iterative training process, wherein each iteration may include (i) training an ML model of the one or more ML models to predict the biological condition based on a subset of the preliminary set of spatial motifs, (ii) evaluating precision of said prediction based on said training, and (iii) altering the subset for training the ML model in a subsequent training iteration. The method may include selecting an optimal subset of the preliminary set of spatial motifs based on the evaluated precision of said iterations as the context-dependent discriminative motifs.
[0022] According to some embodiments, obtaining the context-dependent discriminative motifs may include receiving a training dataset that may include a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image may be annotated according to a biological condition pertaining to that context. The method may include identifying a preliminary set of spatial motifs based on the plurality of training images. The method may include identifying a subset of the preliminary set of spatial motifs that may appear in tissue samples associated with a first biological condition and may not appear in tissue samples associated with a second biological condition. The method may include selecting the identified subset as the context-dependent discriminative motifs for distinguishing between the first biological condition and the second biological condition.
[0023] According to some embodiments, (a) the context may pertain to cancer metastatic potential, and the biological conditions may include metastasis-free patients versus patients who may develop distant metastases within a predetermined time period, (b) the context may pertain to cancer survival prognosis, and the biological conditions may include short-term survivors versus long-term survivors, (c) the context may pertain to treatment response assessment, and thebiological conditions may include treatment-responsive patients versus treatment-resistant patients, or (d) the context may pertain to disease progression monitoring, and the biological conditions may include patients with stable disease versus patients with progressive disease.
[0024] According to some embodiments, calculating spatial distribution patterns may include obtaining one or more tissue structures within the biological tissue sample. The method may include determining localization of the discriminative motifs relative to the one or more tissue structures. The method may include characterizing the spatial distribution patterns based on the determined localization.
[0025] According to some embodiments, the one or more tissue structures may be selected from a list consisting of: (a) germinal centers, wherein the discriminative motifs may be localized at edges of B cell clusters surrounding the germinal centers, (b) tumor boundaries, wherein the discriminative motifs may be localized at interfaces between tumor regions and immune microenvironment; and (c) immune cell clusters, wherein the discriminative motifs may be localized within or adjacent to specific immune cell aggregations.
[0026] According to some embodiments, the method may include obtaining a context-dependent subset of the set of spatial motifs, wherein the subset may include discriminative motifs which may distinguish between different biological conditions pertaining to that context. The method may include generating motif-based features by (a) calculating frequencies of the context-dependent discriminative motifs, (b) calculating spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample, and (c) calculating pairwise frequency values representing frequency of occurrence of pairwise connections between cells of specific types within the discriminative motifs in the multicellular network. The method may include applying an ML model of the one or more ML models to the motif-based features to predict the condition of the biological tissue sample.
[0027] According to some embodiments, the one or more trained ML models may include a Random Forest classifier that may facilitate interpretability by calculating feature importance scores for individual motif-based features, representing contribution of these motif-based features to prediction of the biological condition. The Random Forest classifier may rank the motif-based features based on the calculated feature importance scores to identify top-contributing features for the prediction. The Random Forest classifier may generate a layered presentation that may includethe target image, the multicellular network, and top-scoring motif-based features via a user interface, to highlight contributors to the prediction.
[0028] According to some embodiments, the method may include determining potential therapeutic targets for drug development based on the spatial distribution patterns of the discriminative motifs and the predicted biological condition.
[0029] According to another aspect of the present invention, a system for determining a condition of a biological tissue sample may be provided. The system may include a processor, and memory storing instructions that, when executed by the processor, may cause the processor to receive a target image depicting spatial single-cell omics data of cells in the tissue sample. The processor may be configured to construct a multicellular network of interconnected nodes, wherein each node may represent a cell in the target image. The processor may be configured to enumerate subgraphs within the multicellular network to identify a set of spatial motifs, wherein each spatial motif may represent a multicellular pattern of interconnected cells of specific types. The processor may be configured to analyze the set of spatial motifs to generate one or more motif-based features, characterizing the biological tissue sample. The processor may be configured to apply one or more trained machine learning (ML) models to the motif-based features to predict the condition of the biological tissue sample.BRIEF DESCRIPTION OF FIGURES
[0030] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best be understood by reference to the following detailed description when read with the accompanying drawings in which:
[0031] Fig. 1 illustrates a computing device for determining a condition of a biological tissue sample, according to aspects of the present invention;
[0032] Fig. 2 depicts a block diagram of a system for determining a condition of a biological tissue sample, according to aspects of the present invention;
[0033] Figs. 3A-3D illustrate a workflow for context-dependent identification of spatial motifs for analyzing single-cell spatial omics data, according to aspects of the present invention;
[0034] Fig. 4 depicting a method for determining a condition of a biological tissue sample using spatial motif analysis at least one processor;
[0035] Figs. 5A-5D depict context-dependent identification of spatial motifs as predictors of melanoma disease state, according to aspects of the present invention;
[0036] Figs. 6A-6D illustrate context-dependent motifs-induced pairwise interactions as predictors of melanoma disease state, according to aspects of the present invention;
[0037] Figs. 7A-7D demonstrate spatial structures of discriminative motifs as contributors to disease state classification, according to aspects of the present invention;
[0038] Figs. 8A-8D present mapping to spatial coordinates for characterizing microenvironment of abundant discriminative motifs, according to aspects of the present invention;
[0039] Figs. 9A-9F display context-dependent identification of spatial motifs analysis comparing different melanoma patient groups, according to aspects of the present invention; and
[0040] Figs. 10A-10F exhibit context-dependent identification of spatial motifs analysis of a triple negative breast cancer cohort, according to aspects of the present invention.DETAILED DESCRIPTION
[0041] The following description sets forth exemplary aspects of the present invention. It should be recognized, however, that such description is not intended as a limitation on the scope of the present invention. Rather, the description also encompasses combinations and modifications to those exemplary aspects described herein.
[0042] Reference is now made to Fig. 1, which is a block diagram depicting a computing device, which may be included within an embodiment of a system for determining and interpreting a condition of a biological tissue sample, according to some embodiments.
[0043] Computing device 1 may include a processor or controller 2 that may be, for example, a central processing unit (CPU) processor, a chip or any suitable computing or computational device, an operating system 3, a memory 4, executable code 5, a storage system 6, input devices 7 and output devices 8. Processor 2 (or one or more controllers or processors, possibly across multiple units or devices) may be configured to carry out methods described herein, and / or to execute or act as the various modules, units, etc. More than one computing device 1 may be included in, and one or more computing devices 1 may act as the components of, a system according to embodiments of the invention.
[0044] Operating system 3 may be or may include any code segment (e.g., one similar to executable code 5 described herein) designed and / or configured to perform tasks involving coordination, scheduling, arbitration, supervising, controlling or otherwise managing operation of computing device 1, for example, scheduling execution of software programs or tasks or enabling software programs or other modules or units to communicate. Operating system 3 may be a commercial operating system. It will be noted that an operating system 3 may be an optional component, e.g., in some embodiments, a system may include a computing device that does not require or include an operating system 3.
[0045] Memory 4 may be or may include, for example, a Random- Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Memory 4 may be or may include a plurality of possibly different memory units. Memory 4 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM. In one embodiment, a non-transitory storage medium such as memory 4, a hard disk drive, another storage device, etc. may store instructions or code which when executed by a processor may cause the processor to carry out methods as described herein.
[0046] Executable code 5 may be any executable code, e.g., an application, a program, a process, task, or script. Executable code 5 may be executed by processor or controller 2 possibly under control of operating system 3. For example, executable code 5 may be an application that may determine and / or interpret a condition of a biological tissue sample as further described herein. Although, for the sake of clarity, a single item of executable code 5 is shown in Fig. 1, a system according to some embodiments of the invention may include a plurality of executable code segments similar to executable code 5 that may be loaded into memory 4 and cause processor 2 to carry out methods described herein.
[0047] Storage system 6 may be or may include, for example, a flash memory as known in the art, a memory that is internal to, or embedded in, a micro controller or chip as known in the art, a hard disk drive, a CD-Recordable (CD-R) drive, a Blu-ray disk (BD), a universal serial bus (USB) device or other suitable removable and / or fixed storage unit. Data pertaining to tissue sample images may be stored in storage system 6 and may be loaded from storage system 6 into memory 4 where it may be processed by processor or controller 2. In some embodiments, some of the components shown in Fig.1 may be omitted. For example, memory 4 may be a non-volatile memory having the storage capacity of storage system 6. Accordingly, although shown as a separate component, storage system 6 may be embedded or included in memory 4.
[0048] Input devices 7 may be or may include any suitable input devices, components, or systems, e.g., a detachable keyboard or keypad, a mouse and the like. Output devices 8 may include one or more (possibly detachable) displays or monitors, speakers and / or any other suitable output devices. Any applicable input / output (I / O) devices may be connected to Computing device 1 as shown by blocks 7 and 8. For example, a wired or wireless network interface card (NIC), a universal serial bus (USB) device or external hard drive may be included in input devices 7 and / or output devices 8. It will be recognized that any suitable number of input devices 7 and output device 8 may be operatively connected to Computing device 1 as shown by blocks 7 and 8.
[0049] A system according to some embodiments of the invention may include components such as, but not limited to, a plurality of central processing units (CPU) or any other suitable multi-purpose or specific processors or controllers (e.g., similar to element 2), a plurality of input units, a plurality of output units, a plurality of memory units, and a plurality of storage units.
[0050] The term neural network (NN) or artificial neural network (ANN), e.g., a neural network implementing a machine learning (ML) or artificial intelligence (AI) function, may be used herein to refer to an information processing paradigm that may include nodes, referred to as neurons, organized into layers, with links between the neurons. The links may transfer signals between neurons and may be associated with weights. A NN may be configured or trained for a specific task, e.g., pattern recognition or classification. Training a NN for the specific task may involve adjusting these weights based on examples. Each neuron of an intermediate or last layer may receive an input signal, e.g., a weighted sum of output signals from other neurons, and may process the input signal using a linear or nonlinear function (e.g., an activation function). The results of the input and intermediate layers may be transferred to other neurons and the results of the output layer may be provided as the output of the NN. Typically, the neurons and links within a NN are represented by mathematical constructs, such as activation functions and matrices of data elements and weights. At least one processor (e.g., processor 2 of Fig. 1) such as one or more CPUs or graphics processing units (GPUs), or a dedicated hardware device may perform the relevant calculations.
[0051] Referring to Fig. 2, a system 100 for determining a condition of a biological tissue sample may be configured to automatically analyze images 20 depicting spatial, single-cell omics data and predict biological conditions through spatial motif analysis.
[0052] The spatial single-cell omics data may include various molecular profiling approaches that characterize single cell phenotypes. For example, the omics data may include proteomics, which involves analysis of protein expression levels and modifications within individual cells to determine cellular function and state. In another example, the omics data may include genomics, which encompasses analysis of DNA sequences, mutations, and genetic variations at the singlecell level to identify cellular identity and potential alterations. In another example, the omics data may include transcriptomics, which involves measurement of RNA expression profiles and gene activity patterns within individual cells to characterize cellular processes and responses. In yet another example, the omics data may include organellomics, which involves comprehensive analysis of organelles within individual cells, including their composition, distribution, interactions, and functional states to provide insights into cellular metabolism and organellespecific functions that contribute to overall cell phenotype.
[0053] Additionally, or alternatively, image 20 may include any bioimage-related spatial data that characterizes biological structures and their spatial organization. For example, image 20 may depict cells in live microscopy images, where dynamic cellular behaviors and interactions are captured over time to reveal spatial patterns of cellular movement, division, or response to stimuli. In another example, image 20 may represent graphs of organelles as nodes 120N and their interactions as edges 120E within a cell, where the multicellular network 120NW may be constructed at the subcellular level to analyze organellar spatial organization and functional relationships.
[0054] System 100 may receive an image 20 depicting spatial single-cell omics data of cells in the tissue sample. In some cases, the spatial single-cell omics data may be acquired using MIBI-TOF (Multiplexed Ion Beam Imaging by Time-Of-Flight) technology with 36-channel or 38-channel protein detection capabilities.
[0055] The system 100 may include a preprocessing module 110 that processes the image 20 to prepare the image 20 for analysis. Preprocessing module 110 may comprise a segmentation module 111 that identifies individual cells in the image 20 and generates segments 111S representing the identified cells along with location 111L data indicating spatial coordinates foreach cell. In some cases, the spatial coordinates may have a resolution of 0.5 × 0.5 μm per pixel for subcellular resolution imaging. Segmentation module 111 may use an appropriate algorithm for single cell segmentation from multichannel spatial omics data.
[0056] Preprocessing module 110 may further include a cell type classification module 113 that may assign each identified cell to a cell type class 113C based on the spatial single-cell omics data. Cell type classification module 113 may perform cell type classification using an appropriate algorithm for automated cell type assignment.
[0057] System 100 may further include a graph generation module 120, configured to construct a multicellular network 120NW of nodes 120N, interconnected by edges 120E based on the cell type classification 113C and location 111L. Each node 120N of multicellular network 120NW may represent a cell in image 20.
[0058] The graph generation module 120 may connect nodes which correspond to spatially adjacent cells with edges 120E based on spatial proximity criteria. In some cases, the edges 120E between nodes 120N may be determined using Delaunay triangulation, based on location data 111L, to connect spatially adjacent cells. Additionally, or alternatively, graph generation module 120 may apply a spatial proximity threshold to exclude edges 120E between nodes 120N, that correspond to cells located more than a predetermined distance (e.g., 50 pm) apart.
[0059] As shown in Fig. 2, a motif extraction module 130 may enumerate subgraphs 131 within the multicellular network 120NW to identify a set of spatial motifs 133.
[0060] As used herein, spatial motifs 133 may represent a multicellular pattern of interconnected cells of specific types. A spatial motif 133 may be determined as a subgraph 131 within the multicellular network 120NW that occurs more frequently than expected by chance, where the subgraph includes a predetermined number (e.g., 3, 4 or 5) of nodes which represent cells, and connected by edges which representing spatial adjacencies in a specific connectivity structure. A spatial motif 133 may be characterized by both the cell types of its constituent nodes and the pattern of connections between them. These motifs 133 may serve as building blocks of multicellular organization, representing functional modules where specific cell types interact in stereotypic spatial arrangements that are associated with particular biological states or processes.
[0061] As explained herein, spatial motifs 133 may provide an intermediate layer of abstraction for analyzing biological samples, bridging the gap between simple pairwise cell-cellinteractions and complex tissue-scale organization by capturing local multicellular patterns that are more informative than individual cell pairs but more tractable than whole-tissue analysis.
[0062] Motif extraction module 130 may examine subgraphs 131 comprising predetermined numbers of interconnected nodes 120N within the multicellular network 120NW and determine canonical patterns 137 based on cell types and connectivity structures. The inventors have experimentally used spatial motifs 133 that include 3-cell, 4-cell, or 5-cell subgraphs, representing different sizes of multicellular patterns. In these experiments, subgraphs 131 have been enumerated using up to 128 colored networks (128 cell types) for motif detection.
[0063] Motif extraction module 130 may determine canonical patterns 137 based on cell types and connectivity structures. According to some embodiments, motif extraction module 130 may generate a unique canonical representation for each examined subgraph 131 that remains invariant regardless of different organizations of nodes 120N and edges 120E within the subgraph 131. This canonical representation may ensure that subgraphs 131 having identical cell type compositions and connectivity structures are recognized as the same canonical pattern 137, even when they appear in different orientations or spatial arrangements within the multicellular network 120NW.
[0064] The canonical pattern determination process may utilize algorithmic approaches to encode each subgraph 131 as a standardized bit sequence representation. Each canonical pattern 137 may be characterized by both the specific cell types of its constituent nodes 120N and the particular arrangement of edges 120E connecting these nodes 120N. According to some embodiments, the canonical pattern representation may be global in nature, enabling consistent identification and classification of multicellular patterns across different tissue samples and patient cohorts. This standardized representation may allow system 100 to compare and match spatial motifs 133 between tissues, facilitating cross-tissue analysis and pattern recognition. Through this canonical pattern determination, motif extraction module 130 may systematically identify and categorize recurring multicellular organizational structures within the tissue sample, enabling the enumeration of spatial motifs 133 that represent biologically meaningful patterns of cell interactions and spatial arrangements.
[0065] According to some embodiments, motif extraction module 130 may calculate a frequency 135 of occurrence for each canonical pattern 137 (e.g., how many subgraphs 131 are associated to each canonical pattern 137), across the multicellular network 120NW with respect tothe total number of corresponding subgraphs of the same size in the patient's sample. Motif extraction module 130 may subsequently select canonical patterns 137 whose frequency exceeds a predetermined threshold and are statistically significant compared to obtaining the pattern by chance in the specific tissue, to obtain a set of spatial motifs. This statistical significance requirement may distinguish spatial motifs 133 from regular frequent graph patterns by ensuring that the identified patterns occur more frequently than would be expected in random networks with similar characteristics. In other words, motif extraction module 130 may look for canonical patterns 137 who are most frequently manifested in image 20 and demonstrate statistical significance, and select those as the set of spatial motifs 133. In this context, the terms spatial motifs 133 and canonical pattern or representation 137 may be used interchangeably.
[0066] As shown in Fig. 2, system 100 may further include a feature extraction module 140 configured to analyze the set of spatial motifs 133 to generate motif-based features 140F characterizing the biological tissue sample.
[0067] For example, feature extraction module 140 may include a machine-learning (ML) based discrimination model 141 that may be configured to identify discriminative motifs 142 which distinguish between different biological conditions, as elaborated herein. Additionally, or alternatively, the discriminative motifs 142 may be selected based on a discrimination stringency parameter that defines the fraction of patients of a given disease state that share a discriminating motif 142. As explained herein, discriminative motifs 142 may be used as motif-based features 140F for predicting, or classifying a biological condition of biological tissue samples depicted in image 20.
[0068] In another example, feature extraction module 140 may further include a pairwise analysis module 145 that examines pairwise connections between cells of specific types that are included in the discriminative motifs 142. Pairwise analysis module 145 may focus specifically on the pooled set of discriminative motif instances to calculate pairwise frequency 146 values representing frequency of occurrence of these pairwise connections within the discriminative motifs 142 rather than across the entire pool of spatial motifs 133. Feature extraction module 140 may subsequently use pairwise frequency 146 values as motif-based features 140F for predicting, or classifying biological conditions of the biological tissue of target image 20.
[0069] In yet another example, feature extraction module 140 may further include a spatial distribution analysis module 143 that may be configured to calculate spatial distribution patterns144, characterizing locations of specific discriminative motifs 142 within the biological tissue sample. As explained herein, embodiments of the invention may use spatial distribution patterns 144 as motif-based features 140F for predicting, or classifying biological conditions manifested in image 20.
[0070] According to some embodiments, spatial distribution patterns 144 may describe distribution of discriminative motifs 142 in relation to landmarks or biological structures depicted in image 20. For example, system 100 may include a cellular structure identification module 150, configured to identify tissue structures (also referred to as "biological structures") 150S within the biological tissue sample based on cell type 113C clustering patterns, spatial organization of specific cell populations, and morphological features derived from the spatial single-cell omics data 24. The cellular structure identification module 150 may analyze the spatial coordinates (e.g., location 111L) and cell type classifications 113C to detect characteristic cellular arrangements or structures 150S such as germinal centers surrounded by B cell clusters, tumor boundaries defined by transitions between tumor and immune cell populations, or immune cell aggregations with distinct spatial signatures. Cellular structure identification module 150 may characterize, or augment spatial distribution patterns 144 by determining localization of the discriminative motifs 142 relative to these identified cellular tissue structures 150S through spatial proximity analysis, calculating distances between motif instances and structure boundaries, and mapping the stereotypic positioning of discriminative motifs 142 within or adjacent to specific tissue architectural features or structures 150S such as germinal centers, tumor boundaries, immune cell clusters, and the like. Feature extraction module 140 may characterize, or generate the spatial distribution patterns 144 based on the determined localization, as elaborated herein. System 100 may subsequently use these spatial distribution patterns 144, augmented by biological structure 150S information, as motif-based features 140F for predicting, or classifying biological conditions in image 20.
[0071] According to some embodiments, system 100 may include context-specific, or application- specific logic (denoted herein "application 160") that may be configured to apply one or more application-specific, trained machine learning models 161 to the motif-based features 140F, to predict the condition of the biological tissue sample. In some embodiments, application 160 may include an ML based classification model 161 that generates a biological condition classification 161C based on the motif-based features 140F.
[0072] The inventors have experimentally configured application 160 for various clinical contexts, such as predicting melanoma metastatic potential by distinguishing between negativenegative (NN) patients who remain metastasis-free and negative -positive (NP) patients who develop distant metastases within a predetermined time period, or analyzing triple-negative breast cancer survival prognosis by classifying patients as short-term survivors versus long-term survivors based on survival time thresholds. Additionally, application 160 was adapted for treatment response assessment, distinguishing between treatment-responsive and treatmentresistant patients, or for disease progression monitoring by differentiating patients with stable disease from those with progressive disease. The results of these embodiments are discussed further herein.
[0073] The application 160 may further include an interpretation module 163, configured to generate an interpretation 163P of the prediction, facilitating understanding of contributors to the biological condition classification 161C. The interpretation module 163 may calculate feature importance scores for individual motif-based features, representing their contribution to the biological condition classification 161C, and rank the motif-based features based on these scores to identify top-contributing features. In some embodiments, the interpretation module 163 may generate a layered presentation comprising the target image 20, the multicellular network 120NW, and top-scoring motif-based features via a user interface, to visually highlight spatial locations and types of discriminative motifs 142 that most significantly contribute to the prediction of the biological condition.
[0074] Reference is now made to Fig. 4, depicting a method for determining a condition of a biological tissue sample using spatial motif analysis at least one processor (e.g., processor 2 of Fig.1).
[0075] Embodiments of the invention may include a sequential process that transforms spatial single-cell omics data into predictive biological condition classifications through network construction, motif identification, feature extraction, and machine learning-based prediction.
[0076] Embodiments of the method may begin at a step SI 005, which involves receiving a target image depicting spatial single-cell omics data of cells in the tissue sample. Step S1005 may correspond to the functionality provided by system 100 when receiving image 20 depicting spatial single-cell omics data. In some embodiments, target image 20 may contain multichannelspatial omics data 24 acquired through MIBI-TOF technology, where each cell may be characterized by multiple protein markers that define cell types and functional states.
[0077] In step S 1010, the at least one processor may construct a multicellular network 120NW of interconnected nodes 120N, where each node 120N represents a cell in the target image. Step S1010 may be implemented through the preprocessing module 110 and graph generation module 120 of system 100. During step S1010, the segmentation module 111 may identify, or segment individual cells and determine spatial coordinates through location 11 IL data, while the cell type classification module 113 may assign each cell to a cell type class 113C. Graph generation module 120 may then represent each cell as a node 120N in the multicellular network 120NW, where each node is attributed with its respective cell type 113C and spatial location 111L. Graph generation module 120 may further connect spatially adjacent cells with edges 120E based on spatial proximity, or neighborhood criteria.
[0078] In step S1015, the at least one processor may enumerate subgraphs within the multicellular network 120NW to identify a set of spatial motifs, where each spatial motif represents a multicellular pattern of interconnected cells of specific types. Step S 1015 may be performed by the motif extraction module 130 of system 100. During step S 1015, the motif extraction module 130 may examine subgraphs 131 comprising predetermined numbers, or a range of numbers (e.g., 3, 4, 5) of interconnected nodes 120N and determine canonical patterns 137 based on cell types 113C and connectivity structures. Motif extraction module 130 may calculate frequency of occurrence for each canonical pattern 137 and select patterns whose frequency exceeds a predetermined threshold, to obtain the set of spatial motifs 133.
[0079] In other words, motif extraction module 130 may measure the frequency of an n-node graph as the number of occurrences of a specific canonical pattern divided by the sum across all n-node unique canonical subgraphs within the tissue sample. The criteria for determining whether a pattern qualifies as a motif may be defined by statistical testing (e.g., a p-value) for chance occurrence, where motif extraction module 130 may randomly shuffle the nodes of graphs multiple times to generate thousands of randomized networks for comparison. System 100 may subsequently filter out non-significant motifs based on these statistical tests (e.g., p-value), ensuring that only patterns occurring more frequently than expected by chance are retained. This filtering process may occur before feature extraction module 140 begins the discriminative motif selection process. According to some embodiments, the frequency calculation may represent aratio between the specific subgraph occurrence and the sum across all n-node unique canonical subgraphs within the tissue, providing a normalized measure of pattern prevalence that accounts for the total diversity of possible multicellular arrangements of the same size.
[0080] In step S1020, the at least one processor may analyze the set of spatial motifs to generate one or more motif-based features 140F characterizing the biological tissue sample. Step S1020 may be implemented through the feature extraction module 140 of system 100.
[0081] For example, during step S1020, the feature extraction module 140 may identify discriminative motifs 142 that distinguish between different biological conditions, and select these discriminative motifs 142 as motif-based features 140F, to be subsequently analyzed for biological condition classification 161C. In another example, spatial distribution analysis module 143 may calculate spatial distribution patterns 144 characterizing locations of specific discriminative motifs 142 within the biological tissue sample (e.g., in relation to recognized cellular structures 150S), and select these spatial distribution patterns 144 as motif-based features 140F for classifying biological conditions 161C. In yet another example, pairwise analysis module 145 may be configured to examine pairwise connections between cells of specific types within the discriminative motifs 142 and calculate pairwise frequency 146 values. Feature extraction module 140 may subsequently select these pairwise frequency 146 values as motif-based features 140F for classifying biological conditions 161C.
[0082] In step S1025, embodiments of the invention may apply one or more trained machine learning models 161 to the motif-based features 140F to predict, or classify the condition 161C of the biological tissue sample. Step S1025 may be performed by the context specific, or application 160 specific logic of system 100. During step SI 025, ML based classification model 161 may process the motif -based features 140F generated in step SI 020 to generate a biological condition classification 161C. Additionally, or alternatively, interpretation module 163 may generate an interpretation 163P of the prediction, providing insights into which discriminative motifs 142 spatial distribution patterns 144 and pairwise interrelations contribute most significantly to the biological condition classification 161C.
[0083] The biological condition classification 161C and interpretation 163P may be presented through an interactive User Interface (UI), such as input elements 7 and output elements 8 of Fig.1, that may enable researchers and physicians to visualize and understand the predictive analysis results, thereby providing practical application in clinical and research settings.
[0084] The UI may display the target image 20 with overlaid multicellular network 120NW representations, where discriminative motifs 142 are highlighted using, for example, color-coding, boundary markers, or interactive selection tools. The UI may provide layered visualization capabilities, allowing users to toggle between different views such as raw tissue imagery, cell type classifications, network connectivity, and motif-based feature importance scores. Interactive elements may enable users to click on specific discriminative motifs 142 to access detailed information about their composition, spatial distribution patterns 144, and contribution scores to the biological condition classification 161C.
[0085] The interpretation 163P may be presented through complementary visualization panels that display ranked lists of top-contributing motif-based features, statistical significance indicators, and confidence metrics associated with the biological condition classification 161C. The UI may include comparative analysis tools that allow researchers to examine how different discriminative motifs 142 vary across patient cohorts, disease states, or treatment responses. Heat maps, bar charts, and network diagrams may provide intuitive representations of pairwise frequency 146 values and spatial distribution patterns 144, enabling users to identify patterns and correlations that may not be apparent through traditional analysis methods, thereby improving technology for spatial tissue analysis.
[0086] This technological implementation advances the field of computational pathology by providing practical application of context-dependent motif-based feature analysis that transforms complex spatial single-cell data into interpretable predictive insights. The determination and display of discriminative motifs 142 and their spatial contexts improve technology for tissue analysis by enabling automated identification of biologically meaningful multicellular patterns that distinguish between disease states, thereby improving technology for diagnostic applications and clinical research through enhanced spatial understanding of cellular organization in biological tissues.
[0087] According to some embodiments, system 100 may obtain a context-dependent subset of spatial motifs 133, which may include discriminative motifs 142 that distinguish between different biological conditions pertaining to that context.
[0088] For example, the context may pertain to cancer metastatic potential, and the biological conditions 161C may include metastasis-free patients versus patients who develop distant metastases within a predetermined time period. In another example, the context may pertain tocancer survival prognosis, and the biological conditions 161C may include short-term survivors versus long-term survivors based on survival time thresholds. Additionally, the context may pertain to treatment response assessment, and the biological conditions 161C may include treatment-responsive patients versus treatment-resistant patients based on therapeutic outcomes. In yet another example, the context may pertain to disease progression monitoring, and the biological conditions 161C may include patients with stable disease versus patients with progressive disease based on longitudinal clinical assessments. Other types of contexts and corresponding, distinguishing biological conditions may also be possible.
[0089] The context-dependent subset of discriminative motifs 142 may include discriminative motifs 142 that appear with different frequencies across patient populations having distinct biological conditions. Additionally, or alternatively, the context-dependent subset of discriminative motifs 142 may be determined by the distance between the motif frequency distributions under the different biological conditions, where greater distance between the frequency distributions may indicate stronger separation between the conditions and enhanced discriminative capability for distinguishing between the biological states.
[0090] According to some embodiments, discriminative motifs 142 may be selected based on a discrimination stringency parameter that may define the fraction of patients of a given disease state that share a discriminating motif 142. For example, feature extraction model 140 may identify discriminative motifs 142 that appear in at least a predetermined fraction of patients with one disease state while not appearing in patients with another disease state.
[0091] For example, system100 may receive a training dataset 30 that includes a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, as depicted in Fig. 2. Each training image in training dataset 30 may be annotated with annotations 31 according to a biological condition pertaining to a specific context.
[0092] System 100 may employ motif extraction module 130 to identify a preliminary set of spatial motifs 133 based on the plurality of training images. Feature extraction module 140 may identify a subset of the preliminary set of spatial motifs 133 that appear in tissue samples associated with a first biological condition and do not appear in tissue samples associated with a second biological condition. Feature extraction module 140 may subsequently select the identified subset as the context-dependent discriminative motifs 142 for distinguishing between the first biological condition and the second biological condition pertaining to that context.
[0093] Additionally, or alternatively, feature extraction module 140 may include an ML-based discrimination model 141 that may be configured to identify discriminative motifs 142 which are prevalent in distinguishing between different biological conditions. ML based discrimination model 141 may analyze spatial motifs 133 enumerated by motif extraction module 130 to determine which motifs are associated with, or significantly contribute to classifying specific biological states or disease conditions, pertaining to that context.
[0094] For example, feature extraction module 140 may obtain the context-dependent subset of discriminative motifs 142 through an iterative training and selection process. According to some embodiments, system 100 may receive training data set 30 that includes a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples. Each training image in training dataset 30 may be annotated with annotations 31, representing a biological condition pertaining to a specific context.
[0095] System 100 may identify a preliminary set of spatial motifs 133 based on the plurality of training images through motif extraction module 130 as elaborated herein. Feature extraction module 140 may perform an iterative training process to select optimal discriminative motifs 142 from the preliminary set of spatial motifs 133. Each iteration of the training process may include training ML based discrimination model 141 (which may be the same as ML based classification model 161) to predict the biological condition based on a subset of the preliminary set of spatial motifs 133. System 100 may evaluate precision of the prediction based on the training, and feature extraction module 140 may alter the subset of spatial motifs 133 for training ML based discrimination model 141 (that may be the same as ML based classification model 161) in a subsequent training iteration.
[0096] According to some embodiments, ML based discrimination model 141 may systematically vary the subset of spatial motifs 133 used as features across multiple training iterations, testing different combinations and configurations of motifs to identify those that provide optimal predictive performance. The iterative process may involve adjusting the subset of selected motifs and / or parameters such as motif size, discrimination stringency, or motif frequency thresholds to arrive at an optimized selection of discriminative motifs 142. Additionally, ML based discrimination model 141 may consider selecting the top-n frequencies of discriminative motifs 142 to evaluate in each iteration, where n represents a predetermined number of highest-ranking motifs based on their frequency values. It may occur that the set of selected motifs may changebetween iterations. However, the final list of discriminative motifs 142 may represent the aggregated pool of the selected motifs across all iterations, providing a comprehensive set of spatial patterns that consistently demonstrate discriminative capability throughout the cross-validation process.
[0097] In other words, feature extraction module 140 may select an optimal subset of the preliminary set of spatial motifs 133 based on the evaluated precision across the iterations as the context-dependent discriminative motifs 142. The selected discriminative motifs 142 may represent those spatial patterns that consistently demonstrate superior predictive capability for distinguishing between the biological conditions specified in annotations 31.
[0098] Feature extraction module 140 may calculate frequencies 142F of the context-dependent discriminative motifs 142 across multicellular network 120NW. These calculated frequencies 142F may represent how often each discriminative motif 142 occurs within the biological tissue sample relative to the total number of corresponding subgraphs 131 of the same size. According to some embodiments, each patient may be represented by a sparse feature vector of discriminative motif frequencies 142F, where each element of the vector may correspond to the frequency 142F of occurrence of a specific discriminative motif 142.
[0099] The calculated frequencies 142F of discriminative motifs 142 may be provided as motif-based features 140F to ML based classification model 161, as shown in Fig. 2. ML based classification model 161 may be configured to process these motif-based features 140F to generate biological condition classification 161C. According to some embodiments, ML based classification model 161 may apply trained machine learning algorithms to the frequency -based features to predict the condition of the biological tissue sample. The prediction process may involve analyzing patterns in discriminative motif frequencies that are associated with specific biological conditions, enabling system 100 to classify 161C the tissue sample according to its predicted biological state.
[0100] The calculated frequencies 142F of discriminative motifs 142 may be provided as motif-based features 140F to ML based classification model 161, as shown in Fig. 2. ML based classification model 161 may be configured to process these motif-based features 140F to generate biological condition classification 161C. According to some embodiments, ML based classification model 161 may apply trained machine learning algorithms to the frequency -based features to predict the condition of the biological tissue sample. The prediction process mayinvolve analyzing patterns in discriminative motif frequencies that are associated with specific biological conditions, enabling system 100 to classify 161C the tissue sample according to its predicted biological state.
[0101] According to some embodiments, ML based classification model 161 may be trained using training dataset 30 that includes a plurality of training images with corresponding annotations 31 indicating biological conditions, as depicted in Fig. 2. During the training process, system 100 may extract discriminative motifs 142 from the training images and calculate frequencies 142F for each discriminative motif 142 across the training samples. ML based classification model 161 may employ a supervised learning scheme, as known in the art, to learn associations between the calculated frequencies 142F of discriminative motifs 142 and the corresponding biological condition annotations 31, enabling the model to identify patterns in motif-based features 140F that are predictive of specific biological conditions.
[0102] For example, when ML based classification model 161 includes a deep neural network, the training may be performed using supervised learning with backward propagation to adjust network weights based on prediction errors. Alternatively, when ML based classification model 161 includes a Random Forest classifier, the training may be performed using ensemble learning algorithms that construct multiple decision trees and combine their predictions for improved classification accuracy.
[0103] The trained ML based classification model 161 may subsequently apply these learned associations to analyze new tissue samples and generate biological condition classification 161C based on the frequency patterns of discriminative motifs 142 observed in the target image 20.
[0104] According to some embodiments, spatial distribution analysis module 143 may map occurrence of discriminative motifs 142 within multicellular network 120NW based on location information 11 IL. In other words, spatial distribution analysis module 143 may determine spatial coordinates and positioning of each instance of discriminative motifs 142 across the tissue sample.
[0105] Additionally, spatial distribution analysis module 143 may analyze this mapping to calculate spatial distribution patterns 144 characterizing locations of specific discriminative motifs 142 within the biological tissue sample, as depicted in Fig. 2. The spatial distribution patterns 144 may describe how discriminative motifs 142 are distributed throughout different regions of the biological tissue, including their location, their density, clustering patterns, and relative positioning with respect to tissue landmarks, biological structures and / or boundaries.
[0106] According to some embodiments, spatial distribution patterns 144 may be, or may include data structures (e.g., tables, linked lists and the like) that quantify the spatial heterogeneity of each type of discriminative motif 142, to identify regions where specific motifs are preferentially located. For example, the data structures may encode that discriminative motifs 142 comprising B cells and CD8 T cells are preferentially located at the edges of B cell clusters surrounding germinal centers, as demonstrated in the melanoma cohort analysis. In another example, the data structures may quantify that specific discriminative motifs 142 involving interactions between tumor cells and immune cells are preferentially positioned at tumor boundaries or within mixed tumor-immune regions, as observed in the triple-negative breast cancer analysis where certain motifs appeared in patients with "mixed" rather than compartmentalized tumor-immune phenotypes.
[0107] The calculated spatial distribution patterns 144 may be provided as motif-based features 140F to ML based classification model 161 for predicting the condition of the biological tissue sample. ML based classification model 161 may process these spatial distribution-based features to identify associations between the spatial organization of discriminative motifs 142 and specific biological conditions. According to some embodiments, the spatial context information captured in spatial distribution patterns 144 may enhance the predictive capability of ML based classification model 161, to produce prediction of biological condition 161C based on incorporating not only the presence and frequency of discriminative motifs 142, but also their spatial arrangement pattern 144 within the tissue architecture.
[0108] According to some embodiments, ML based classification model 161 may be trained using training dataset 30 to learn associations between spatial distribution patterns 144 and biological condition annotations 31. During the training process, system 100 may employ a supervised learning scheme, as known in the art, to calculate spatial distribution patterns 144 characterizing the locations of discriminative motifs 142 within each training sample.
[0109] Spatial distribution analysis module 143 may analyze the spatial positioning, density patterns, and localization preferences of discriminative motifs 142 relative to tissue structures across the training samples. ML based classification model 161 may learn correlations between the calculated spatial distribution patterns 144 and the corresponding biological condition annotations 31, enabling the model to identify spatial organizational features that are predictive of specific biological conditions. The trained ML based classification model 161 may subsequentlyapply these learned spatial associations to analyze new tissue samples and generate biological condition classification 161C, further based on the spatial arrangement patterns of discriminative motifs 142 observed in the target image 20.
[0110] According to some embodiments, pairwise analysis module 145 may obtain a context-dependent subset of discriminative motifs 142 and examine pairwise connections 145P between cells of specific types within these discriminative motifs 142. Pairwise analysis module 145 may calculate pairwise frequency values 146 representing frequency of occurrence of these pairwise connections 145P in multicellular network 120NW, as depicted in Fig. 2.
[0111] The calculated pairwise frequency values 146 may be provided as motif-based features 140F to ML based classification model 161 for predicting the condition of the biological tissue sample. ML based classification model 161 may process these pairwise interaction -based features to identify associations between specific cell-cell interaction patterns and biological conditions 161C.
[0112] According to some embodiments, the pairwise frequency 146 information may enhance the predictive capability of ML based classification model 161, to predict biological condition 161C by capturing fine-grained cell-cell interaction signatures that are embedded within the discriminative motifs 142 and are predictive of specific biological states.
[0113] The inventors have shown that pairwise frequency 146 values may provide refined analysis that differs from conventional pairwise cell interaction studies. Rather than examining arbitrary pairs of cell types across the entire tissue sample, pairwise analysis module 145 may focus specifically on cell type pairs that have been identified as components of discriminative motifs 142. This targeted approach may provide enhanced discriminative power because the pairwise connections 145P analyzed are derived from spatial motifs 133 that have been determined to distinguish between different biological conditions. According to some embodiments, this context-dependent selection of pairwise interactions 145P may reveal biologically meaningful cell-cell relationships that are embedded within the discriminative multicellular patterns, thereby providing more specific and relevant features for biological condition prediction compared to general pairwise interaction analysis across all possible cell type combinations.
[0114] According to some embodiments, ML based classification model 161 may be trained using training dataset 30 to learn associations between pairwise frequency values 146 and biological condition annotations 31. During the training process, system 100 may extractdiscriminative motifs 142 from the training images and calculate pairwise frequency values 146 for cell-cell interactions within these motifs across the training samples.
[0115] ML based classification model 161 may employ a supervised learning scheme, as known in the art, to learn correlations between the calculated pairwise frequency values 146 and the corresponding biological condition annotations 31, enabling the model to identify interaction patterns (e.g., frequency values 146) that are predictive of specific biological conditions. The trained ML based classification model 161 may subsequently apply these learned pairwise interaction associations to analyze new tissue samples and generate biological condition classification 161C based on the pairwise frequency patterns observed in the target image 20.
[0116] According to some embodiments, system 100 may include a random forest classifier 170 which may be the same as, or integrated within ML based classification model 161, as depicted in Fig. 2. Random Forest classifier 170 may be configured to facilitate interpretability of classification 161C by allowing detailed analysis of feature contributions to the biological condition prediction process.
[0117] For example, random forest classifier 170 may calculate feature importance scores 140S for individual motif-based features 140F, representing the contribution of these motif-based features 140F to the prediction of the biological condition. The feature importance scores 140S may quantify how significantly each motif-based feature 140F influences the biological condition classification 161C generated by Random Forest classifier 170.
[0118] According to some embodiments, classification model 161 may calculate feature importance score 140S based on reduction in prediction error or information gain achieved by each motif-based feature 140F during the decision tree construction process within Random Forest classifier 170.
[0119] Random Forest classifier 170 may rank the motif-based features 140F based on the calculated feature importance scores 140S to identify top-contributing features for the prediction. The ranking process may order the motif-based features 140F according to their relative importance in distinguishing between different biological conditions, enabling identification of the most discriminative spatial patterns and cellular interactions.
[0120] According to some embodiments, interpretation module 163 may collaborate with interpretation module 163, to generate a layered presentation via a user interface 163UI (e.g., including input element 7 and output element 8 of Fig. 1). This layered presentation may bereferred to herein as interpretation 163P. It may include, for example, target image 20, multicellular network 120NW, top-scoring motif-based features 140F and any combination thereof, to highlight contributors to the prediction.
[0121] The presentation of layered interpretation 163P via UI 163UI may visually integrate the spatial context of the tissue sample with the analytical results, enabling researchers to understand which discriminative motifs 142, spatial distribution patterns 144, and / or pairwise interactions 145 contribute most significantly to the biological condition classification 161C. Additionally, or alternatively, the presentation of layered interpretation 163P via UI 163UI may provide interactive visualization capabilities that may allow users to explore the relationship between spatial motif locations and their predictive importance scores.
[0122] According to some embodiments, system 100 may include a target identification module 167 that may be configured to determine potential therapeutic targets for drug development based on the spatial distribution patterns 144 of discriminative motifs 142 and the predicted biological condition classification 161C. Target identification module 167 may receive spatial distribution patterns 144 from spatial distribution analysis module 143, which characterize the locations of specific discriminative motifs 142 within the biological tissue sample, and biological condition classification 161C from ML based classification model 161.
[0123] Target identification module 167 may analyze the spatial distribution patterns 144 to identify regions where discriminative motifs 142 are consistently localized in association with specific biological conditions. According to some embodiments, target identification module 167 may correlate the spatial positioning of discriminative motifs 142 with the predicted biological condition to identify cellular interaction patterns that may be causally related to disease progression, treatment resistance, or other clinically relevant outcomes. The module may examine how the spatial context of discriminative motifs 142 relates to tissue architecture and disease pathology to prioritize potential therapeutic interventions.
[0124] According to some embodiments, the potential therapeutic targets determined by target identification module 167 may comprise specific cell-cell interaction patterns represented by discriminative motifs 142 that are associated with the biological condition. Target identification module 167 may identify these interaction patterns by analyzing the connectivity structures and cell type compositions within discriminative motifs 142 that demonstrate strong spatial associations with adverse biological conditions. For example, target identification module 167 mayidentify discriminative motifs 142 containing interactions between tumor cells and immune cells that are spatially localized at tumor boundaries and associated with poor prognosis, designating these specific cellular interactions as potential targets for therapeutic modulation.
[0125] Target identification module 167 may generate, and display (e.g., via UI 163UI) notifications or reports indicating the identified therapeutic targets, including detailed information about the spatial localization patterns 144, cell type compositions, and predicted biological significance of the target interactions. According to some embodiments, target identification module 167 may provide recommendations for drug development strategies based on the spatial distribution patterns 144 and biological condition predictions, enabling researchers to focus therapeutic development efforts on cellular interactions that are both spatially relevant and clinically significant.
[0126] Figs. 3A-3D illustrate a workflow for Context-dependent Identification of Spatial Motifs (CISM) that may be implemented by system 100 of Fig. 2, for analyzing single-cell spatial omics data and predicting biological conditions through spatial motif analysis, according to some embodiments of the invention.
[0127] Fig. 3A depicts the initial data processing steps that may be implemented by preprocessing module 110 of Fig. 2. The figure shows the progression from multichannel spatial omics imaging 20 of cell populations through cell segmentation 111 to cell type classification 113. Segmentation module 111 may identify individual cell segments 111S and cell type classification module 113 may categorize each cell segment according to its respective cell type 113S. The spatial single-cell omics data 24 may be processed to generate segments 111S and cell type classes 113C with corresponding spatial coordinates 11 IL.
[0128] Fig. 3B demonstrates an unsupervised motif extraction process that may be performed by motif extraction module 130 of Fig. 2. Patient tissue sample images 20 may be converted into spatial multicellular networks 120NW represented as graphs, where cells serve as nodes 120N connected by edges 120E based on spatial proximity. The motif extraction module 130 applies algorithms to extract and enumerate recurring spatial patterns (spatial motifs 133) from each multicellular network 120NW, and calculating the relative frequency 135 of occurrence for each motif structure 133 within the network 120NW.
[0129] Fig. 3C outlines the supervised machine learning validation pipeline that may be implemented by one or more ML-based models (e.g., 161, 141) of Fig. 2 using a leave-one-patient-out cross-validation approach. The process includes holding out one patient as a test case, extracting discriminative motifs 142 from the remaining patients based on disease state context using feature extraction module 140, training ML based model (161 / 141) using these discriminative motifs 142 as motif-based features 140F, and evaluating the model's performance (e.g., calculating an ROC curve) on the held-out patient through biological condition classification 161C analysis. This process may be repeated iteratively for all patients to generate comprehensive performance metrics.
[0130] Fig. 3D presents two applications of the identified discriminative motifs 142 that may be implemented by system 100, according to some embodiments. The first application (left) involves spatial interpretation, where spatial distribution analysis module 143 may map discriminative motifs 142 back onto the tissue image and / or network 120NW to reveal their stereotypic localization patterns and spatial distribution patterns 144. The second application (right) involves motif-induced pairwise interaction analysis that may be performed by pairwise analysis module 145, which may generate disease signatures by examining the probability of edges 120E between different cell types 113C within the pooled discriminative motifs 142, displayed as interaction matrices showing connectivity patterns that differ between biological conditions.
[0131] The inventors have experimentally demonstrated system 100's capability of analyzing melanoma lymph node samples, to predict metastatic potential through spatial motif analysis. In some cases, the biological conditions may include melanoma metastatic potential classification between negative-negative (NN), and negative-positive (NP) patients based on five-year followup outcomes. The NN classification may correspond to patients with tumor-free lymph nodes who remain metastasis-free five years following treatment, while the NP classification may correspond to patients with tumor-free lymph nodes who develop widespread metastases in distant organs within the five-year follow-up period.
[0132] During experimentation, the experimental cohort included 76 tissue sections from lymph nodes of 38 stage I-II melanoma patients, where the spatial single-cell omics data 24 was acquired using 36-channel MIBI-TOF technology with subcellular resolution.
[0133] The cell types 113C were organized in hierarchical lineage relationships and reduced from 26 full-resolution cell types to 16 cell types for improved generalization. This hierarchical organization allowed system 100 to balance discrimination capability with computationaltractability, where the reduced cell type resolution provided better generalization across patient cohorts while maintaining biologically meaningful distinctions between cell populations.
[0134] As shown in Fig. 2, system 100 may further include a cell type resolution optimization module 180. Cell type resolution optimization module 180 may employ an iterative approach to determine the optimal cell type resolution that maximizes separation between distinct biological conditions 161C. According to some embodiments, module 180 may systematically evaluate different levels of cell type hierarchical resolution, e.g., starting from fine-grained cell type classifications and progressively consolidating related cell types into broader categories. At each iteration, cell type resolution optimization module 180 may assess the discriminative performance of the resulting cell type resolution by measuring the ability of discriminative motifs 142 to distinguish between biological conditions 161C. The module may calculate performance metrics such as classification accuracy, area under the curve scores, or statistical significance measures to evaluate how well each resolution level enables feature extraction module 140 to identify meaningful discriminative motifs 142. The iterative process may continue until convergence criteria are met, such as achieving maximum discriminative performance or reaching a predetermined balance between discrimination capability and computational tractability.
[0135] Cell type resolution optimization module 180 may be trained, for example, using supervised learning techniques with training dataset 30 containing labeled tissue samples with known biological conditions specified in annotations 31. According to some embodiments, the training process may involve exposing module 180 to various cell type resolution configurations across multiple patient cohorts, allowing it to learn optimal resolution parameters for different biological contexts. During training, the module may learn to balance the trade-off between using more detailed cell types 131C for better biological specificity versus using consolidated cell type categories for improved statistical power and reduced computational complexity. The training may incorporate feedback from ML based classification model 161 performance to guide the selection of resolution levels that optimize downstream biological condition prediction accuracy.
[0136] Applications of cell type resolution optimization may include enhancing disease diagnosis accuracy by identifying the most informative level of cellular detail for specific pathological conditions. The optimization may improve drug discovery research by determining optimal cell type granularity for identifying therapeutic targets and understanding drug mechanisms of action. In personalized medicine approaches, module 180 may enablecustomization of cell type resolution based on individual patient characteristics or specific disease subtypes. The optimization may also support biological research studies requiring precise cell type differentiation, such as developmental biology investigations, immune system analysis, and tissue engineering applications.
[0137] The graph generation module 120 constructed the multicellular network 120NW using Delaunay triangulation to connect spatially adjacent cells, with edges 120E excluded between cells located more than 50 pm apart. The motif extraction module 130 enumerated hundreds to thousands of 3-cell, 4-cell, and 5-cell motifs correspondingly, examining subgraphs 131 of predetermined sizes within the multicellular network 120NW.
[0138] Feature extraction module 140 performed context-dependent motif selection according to a discrimination stringency parameter, which defined the fraction of patients of a given disease state that share a discriminating motif 142. Feature extraction module 140 employed ML based discrimination model 141 to identify discriminative motifs 142 that appear in at least a predetermined fraction of patients with one disease state while not appearing in patients with the other disease state.
[0139] The machine learning validation was performed using Random Forest-based leave-one -patient-out cross-validation, where each patient was represented by a sparse feature vector of discriminative motif frequencies. The ML based classification model 161 achieved optimal performance with 4-cell discriminative motifs 142 that appear in at least 46% of patients from one disease state, reaching a maximal area under the curve (AUC) score of 0.88 for melanoma metastatic potential prediction.
[0140] Under these conditions, the discriminative motifs approach of the present invention has achieved superior performance in comparison to alternative approaches, including whole graph neural networks and pairwise cell type machine learning methods. The graph neural network approach has reached an AUC score of 0.53, while the pairwise cell type machine learning approach has achieved an AUC score of 0.56, both substantially lower than the discriminative motifs approach of the present invention. This superior performance may demonstrate that discriminative motifs 142 of four cells may serve as multicellular modular signatures of melanoma disease state.
[0141] The inventors further validated the prediction results through permutation testing, where patient disease state labels were randomly permuted and the complete classification pipelinewas repeated. These permutation tests further confirmed that the discriminative motifs approach provides statistically significant predictive capability rather than overfitting to the training data.
[0142] In other words, the inventors have managed to demonstrate that motif-based features 140F may capture biologically meaningful multicellular patterns that distinguish between NN and NP disease states, where the specific spatial organization of cells within the discriminative motifs 142 may contribute to disease state classification beyond simple cell type composition analysis.
[0143] Figs. 5A-5D include diagrams depicting context-dependent identification of spatial motifs as predictors of melanoma disease state, according to some embodiments of system 100 as depicted in Fig. 2.
[0144] Fig. 5A illustrates the melanoma patient cohort structure analyzed by system 100, showing 38 stage I- II melanoma patients divided into two groups based on five-year follow-up outcomes. The cohort includes 20 Negative-Negative (NN) patients with tumor-free lymph nodes who remain metastasis-free, and 18 Negative-Positive (NP) patients with tumor-free lymph nodes who develop widespread metastases in distant organs within the five-year follow-up period. According to some embodiments, system 100 may analyze spatial single-cell omics data 24 from these patient groups to identify discriminative motifs 142 that distinguish between the two biological conditions.
[0145] Fig. 5B displays a comparative analysis of cell type distributions between NP and NN patient groups, as determined by cell type classification module 113 of Fig. 2. The horizontal bar chart shows fractions of various immune cell populations including Memory CD4 T cells, CD4 T cells, Germinal Center B cells, Macrophages, and other cell types 113C present in the lymph node microenvironment. According to some embodiments, the cell type distribution analysis may reveal similar overall compositions between NN and NP groups, with subtle differences in specific immune cell populations that may contribute to disease state discrimination.
[0146] Fig. 5C demonstrates the relationship between discrimination stringency parameter and the number of discriminative motifs 142 identified by feature extraction module 140 of Fig.2. The logarithmic plot shows exponential decay patterns for both 3-cell and 4-cell spatial motifs 133 as the discrimination stringency parameter increases from 0.0 to 0.7. According to some embodiments, motif extraction module 130 may enumerate hundreds to thousands of spatial motifs 133, with the number of discriminative motifs 142 decreasing as the stringency requirements become more restrictive.
[0147] Fig. 5D presents machine learning validation results of melanoma disease state discrimination with CISM’s discriminative motifs representation, obtained by ML based classification model 161 of Fig. 2. Leave one patient out cross-validation (LOOCV) model performance (AUC) (y-axis) over the discrimination stringency parameter (x-axis) (green - 3-cell motifs, purple-4-cell motifs). The pink, cyan, and black dashed horizontal lines represent the LOOCV AUC score of pairwise, GNN, and a random null model, correspondingly. The discriminative motifs derived from 4-cell motifs with the discrimination stringency parameter of 0.46 for discriminative motifs selection were used for further analysis.To explore the landscape of the nineteen 4-node discriminative motifs selected over the cross-validation iterations of the machine learning analysis the inventors visualized their cell type compositions (Fig. 6A). The inventors calculated their induced cell type distribution across the pooled set of instances of discriminative motifs associated with NN or NP across the patient samples (Fig. 6B). Intriguingly, the discriminative motifs-induced cell type distribution was very different from the general cell type distribution (compare Fig. 6B to Fig. 5B). First, cell types that were common in the general distribution were less common in the discriminative motifs-induced distribution. For example, Memory CD4 T cells, that composed over 25% of the general cell type distribution in both NN and NP patients, appeared in less than 10% of the NN-associated discriminative motifs-induced cell type distribution and did not appear at all in NP-associated motifs. Second, cell types that were less prevalent in the general distribution, were more abundant in the discriminative motifs-induced distribution. For example, macrophages and CD8 T cells, that appeared in ~10% of the cohort’s general cell population, accounted for 16-29% of the NN- and NP-associated discriminative motifs-induced cell type distributions. Third, while the general cell type distributions showed marginal differences between NN and NP patients, the motifs-induced distribution showed dramatic differences. For example, CD8 T cells and B cells were almost twice as probable and five-times as probable, correspondingly, to participate in NP-associated discriminative motifs, while CD4 T cells appeared almost 4-fold more in NN patients. Also, each of the Memory CD4 T cells, Neutrophils and NK cells, exclusively appeared in ~10% of the NN-associated motifs cell type distribution.
[0148] Figs. 6A-6D include diagrams depicting context-dependent motifs-induced pairwise interactions as predictors of melanoma disease state, according to some embodiments of system 100 depicted in Fig. 2.
[0149] Fig. 6A illustrates cell type composition of the discriminative motifs associated with NN (n = 8 motifs) and NP (n = 11 motifs) patients, generated by pairwise analysis module 145 of Fig. 2.
[0150] Fig. 6B displays cell type distribution induced from the pooled set of discriminative motifs’ instances across the cohort.
[0151] Fig. 6C presents the distribution of motifs-induced pairwise interactions. The cell type pairwise edge probabilities in the cohorts’ patients pooled set of instances of discriminative motifs associated with NN (left) and NP (right) patients. Spearman correlation between NN and NP pairwise interactions was -0.0157 with p-value ~ 0.87.
[0152] Fig. 6D presents the top three highly ranked motifs-induced pairwise interaction.
[0153] Next, the inventors examined the distribution of discriminative-motifs induced pairwise cell-cell interactions according to the cell type pairwise edge probabilities in the pooled set of instances of discriminative motifs associated with NN or NP across the patient samples (Fig.6C). This analysis revealed potential pairwise interactions associated with one disease state and not the other. The three most abundant (motif-induced) pairwise interactions in NN and in NP are depicted in Fig. 6D: Neutrophils with CD4 T-cells, Macrophages with CD8 T-cells, and CD8 T-cells with NK cells, versus B cells with CD8 T-cells, CD8 T-cells with CD8 T-cells, and Macrophages with CD8 T-cells in NN- versus NP-associated discriminative motifs. Intriguingly, the interaction between CD8 T-cells and Macrophages appeared in both NN- and NP-associated discriminative motifs, in different contexts: the (n = 3) corresponding NN-motifs included interactions between CD8 T Cells with Neutrophil, NK cells, or Stroma cells while the (n =2) corresponding NP-motifs included interactions between CD8 T Cells with CD8 or CD4 T cells. The motifs-induced disease state-associated pairwise interactions were robust to the stringency of the motifs’ discrimination criterion and to the motif’s size. These results suggest that the strict context-dependent discriminative motif selection refines the huge and noisy landscape of cell types and all possible cell-cell interactions in the full spatial multicellular network.
[0154] As discussed herein, a motif may be defined according to its cell types and the edges between them. The inventors have speculated whether the discriminative motifs' cell type composition, i.e., the cell types without considering their relative spatial organization to one another, is sufficient to discriminate between the NN and NP disease states, or in other words -whether the spatial structure of discriminative motifs' matter for disease state classification. Toanswer this question, the inventors repeated the steps of context-dependent motifs' selection and machine learning analysis, but instead of using the discriminative motifs the inventors used the discriminative cell type composition representations. The cell type composition of a motif was represented as a sparse vector that encodes the number of occurrences of each cell type in the motif.
[0155] For example, a 4-cell motif that includes two edges from a HEV to two memory CD4 T-cells and two edges from a B-cell to the same two memory CD4 T-cells is transformed to a sparse vector encoding a single B-cell, one HEV cell, and two memory CD4 T-cells (Fig. 7A). This cell type composition representation was less specific and thus induced fewer discriminative cell type compositions with respect to the (more specific) discriminative motifs. Machine learning evaluation confirmed that the cell type composition representation failed to generalize to discriminate between NN and NP disease states.
[0156] To directly assess which of the discriminative motifs' spatial organizations are most associated with the disease state, the inventors have devised an approach that measures the gain in classification performance attributed to the inclusion of each discriminative motifs spatial organization instead of its corresponding cell type composition.
[0157] First, the inventors transformed the representation from the discriminative motifs space to the cell type composition space by replacing each discriminative motif with its corresponding cell type composition (Fig. 7B). Note that multiple discriminative motifs can be mapped to the same cell type composition, and that a cell type composition can include instances of motifs that did not meet the strict discrimination stringency parameter used to determine the discriminative motifs. This discriminative motifs-derived cell type composition representation reached a reduced AUC score of 0.655, which was used as a baseline (Fig. 7B, middle).
[0158] Second, beginning from this cell type composition representation, the inventors iteratively introduced back the spatial information for each cell type composition by replacing it with the discriminative motifs that induced it. This replacement created a hybrid representation that contained both cell type composition and discriminative motif features (Fig. 7B, right). This hybrid representation was used for machine learning assessment, where the contribution of the corresponding motifs' spatial organization was attributed to the classification gain with respect to the cell type composition baseline. The classification gain for each cell type composition wasranked, showing that the vast majority (17 / 19) of the motifs' spatial structure contributed to disease state prediction (Fig. 7C).
[0159] The three highest ranked cell type compositions included "unidentified" cell types, i.e., cells without a specific classification, and thus were not further analyzed. The 4thhighest ranked cell type composition contributed 0.11 to the cell type composition baseline AUC, and was composed of one B cell, one high-endothelial venule (HEVs) cell, and two memory CD4 T cells (Fig. 7C, motif #1 in green, this motif was used as an example in Fig. 7A-B). The 5thhighest ranked cell type composition contributed 0.09 to the baseline AUC and was composed of a CD4 T cell, a pair of Macrophage cells, and a vessel cell (Fig. 7B, motif #11 in orange; this motif was used as an example in Fig. 7B.
[0160] To further verify that the spatial structures of these motifs were more discriminate than their corresponding cell type compositions, the inventors counted the number of instances of these motifs and their corresponding cell type compositions across the cohort's patients. While the number of instances of the cell type compositions induced by motifs I and II did not show clear discrimination between the disease states (Fig. 7D, top), the prevalence of motifs I and II clearly discriminated between NN and NP (Fig. 7D, bottom). Thus, distilling the contribution of the discriminative motifs' spatial structures indicates that the specific intra-motif direct cell-cell interactions are more sensitive as markers for disease state with respect to the corresponding motifs cell type composition.
[0161] Figs. 7A-7D are diagrams, explaining the spatial structures of discriminative motifs as contributors to disease state classification, according to some embodiments of system 100 depicted in Fig. 2.
[0162] Fig. 7A depicts an example of mapping a discriminative motif (left) to its corresponding cell type composition (right). The cell type composition representation is 16 -dimensional, counting the number of instances of each cell type in the input discriminative motif.
[0163] Fig. 7B reflects the assessment of discriminative contribution of the spatial structure of discriminative motifs. First, all the discriminative motifs (left) are mapped to their corresponding cell type compositions (middle). This discriminative motifs-derived cell type composition representation is used to train a baseline machine learning model to predict the disease state. Second, the contribution of the spatial structure is measured as the classification performancegain by replacing one cell type composition with its corresponding discriminative motifs, one at a time (right).
[0164] Fig. 7C reflects the ranked classification gain for each cell type composition of the n = 19 discriminative motifs (x-axis). Specifically, each column represents a cell type composition, each row represents a cell type, and the color encodes the number of cell type instances in the corresponding cell type composition. For example, the cell type composition #1 corresponds to one B cell, one HEV, and two memory CD4 T cells (see Fig. 7A). The deviation of the AUC of a cell type composition baseline (yellow dashed line, AUC = 0.655 ) with respect to its corresponding hybrid representation (blue bar) is the classification gain. The spatial structure of 17 / 19 of the motifs contributed to the disease state prediction. Two motifs and their corresponding cell type compositions (I - green, II - orange), whose spatial structure contributed the most while not containing (less interpretable) "unidentified" cell types, were highlighted for further analysis.
[0165] In Fig. 7D, the number of instances of the cell type compositions I (green) and II (orange) are shown on the top panel, and their corresponding motifs is shown at the bottom panel. The outlier motifs II (patient #45) and I (patient #8) were selected when the corresponding patient was used for the test (i.e., was selected as discriminative when not considering that patient). The cell type compositions did not discriminate between NN and NP (Mann-Whitney U test: cell type compositions I (green): p -value = 0.6398, cell type compositions II (orange): p -value = 0.5586 ), while the motifs significantly discriminate between NN and NP (motifs I (green): p -value = 0.0029, motifs II (orange): p -value = 0.0032).
[0166] One of the most appealing attributes of CISM is the ability to locate instances of a discriminative motif in the multicellular network and characterize the stereotypic spatial contexts associated with its localization and the disease state. Observing that a discriminative motif is consistently localized in a specific microenvironment can drive the interpretation of potential mechanisms regarding the motifs physiological roles in the context of the disease. To decide which discriminative motifs to spatially interpret, the inventors have balanced between two desired properties: (1) generalization - discriminative motifs that appeared in many patients, and (2) interpretability - discriminative motifs with many instances per patient. Placing the 19 NN-NP discriminative motifs on this generalization-interpretability axes highlighted three motifs that appeared in many patients with many instances per patient, on average (Fig. 8A). Motif I (green)was NN-associated and included a Macrophage, Neutrophil, and a pair of CD4 T cells, motif II (magenta) was also NN-associated and comprised of an NK cell, a CD8 Tcell, a Memory CD4 T cell, and an Unidentified cell, and motif III (yellow) was NP-associated and comprised of a Macrophage, a pair of CD8 T cells, and a B cell (Fig. 8B). Motif III was found to be the most abundant in several NP patients and was thus used for spatial interpretation. Localizing motif III in the tissue revealed a stereotypical spatial pattern appearing at the edges of B cell clusters that surround the germinal center (Fig. 8D). The motifs B cell was part of the germinal center associated B cell cluster that was adjacent to two CD 8 T cells that were adjacent to a macrophage, perhaps serving as a specific "adaptor" of the germinal center to the rest of the immune microenvironment that associates with long-term metastasis progression. These results highlight the possibility of providing physiologically meaningful interpretation by observing the spatial localization of CISM-derived discriminative motifs.
[0167] Figs. 8A-8D are diagrams, explaining the process of mapping to spatial coordinates for characterizing microenvironment of abundant discriminative motifs, according to some embodiments of system 100 depicted in Fig. 2.
[0168] Fig. 8A shows the mean per-patient number of discriminative motif instances ( y -axis) in relation to the number of patients that contained that motif ( x -axis). Each dot represents a discriminative motif ( n = 19 ). Note motifs I (green), II (magenta), and III (yellow) that appear in many patients and in many instances per patient.
[0169] Fig. 8B shows the three discriminative motifs that were marked in Fig. 8A. The edge colors are consistent across the other figure panels.
[0170] Fig. 8C shows the number of instances of motifs I-III (top-bottom, log scale) across patients. Only patients with at least one instance were listed.
[0171] Fig. 8D shows localizing motif III in the tissue of NP patients #10 (left) and #15 (right) and in additional patients. Each dot represents a cell colored according to its cell type: B cell, CD 8 T cell, Germinal Center B cells, and macrophages (other cell types are not shown). The germinal center (GC) is surrounded by a dashed white circle. Instances of motif III are marked with dashed yellow rectangles (larger rectangles indicate multiple instances of the motif).
[0172] CISM's context dependency implies that the same motifs extracted from a specific tissue can be differentially analyzed according to the question at hand - via context dependent selection of the discriminative motifs. To test how context can alter the motifs-induced pairwiseinteractions disease signature when considered relative to different biological contexts, and the spatial interpretation of the motifs' localization, the inventors turned to the extended melanoma cohort that included 13 patients that had metastatic sentinel lymph nodes that did not develop distant metastases within a follow-up of at least 5 years (PN). The inventors then evaluated the motifs that discriminated between NP patients and PN patients (Fig. 9A).
[0173] Focusing on changes in the organization of the lymph node immune microenvironment, the inventors excluded all cells inside the tumor cell clusters in PN patients using a refined version of the convex hull algorithm. Cell type classification showed a similar cell type distribution for NP and PN patients. CISM analysis of NP versus PN found the optimal tradeoff between classification performance and sufficient discriminative motifs and their instances for interpretability in 5-cell motifs and discrimination stringency of 0.4. CISM reached an AUC of 0.81, surpassing pairwise cell type machine learning with an AUC close to random and GNN with an AUC of 0.76. The motif-induced pairwise interactions showed clear differences distinguishing between the disease states (Fig. 9B). Most of the NP-associated discriminative motifs’ induced pairwise interactions appearing in either NP-vs-NN (Fig. 6C) or in NP-vs-PN (Fig.9B), but not in both. The only motif-induced pairwise interaction that appeared in NP in both contexts was between B cells and CD8 T-cells, suggesting these interactions as putative specific markers for future metastases for stage II melanoma.
[0174] To interpret the NP-versus-PN discriminative motifs, the inventors selected three prominent discriminative motifs that appeared in many patients with many instances per patient (Fig. 9C-D). Two of these motifs were NP-associated (motif I in green and III in yellow) and one was PN-associated (motif II in magenta) (Fig. 9E). Following the discriminative motifs’ localization at the boundaries of B cell zones that surrounded the germinal center in the context of NN-vs-NP (Fig. 8D), the inventors examined the localization of NP-vs-PN discriminative motifs and identified two out of the three prominent motifs (II and III) that appeared near germinal centers (Fig. 9F). Motif II was PN-associated and appeared in 10 / 13 of PN patients adjacent to the germinal center, connecting it with the B cell follicles surrounding it (Fig. 9F). Motif III was NP-associated, included an edge between a B-cell and a CD8 T-cell and appeared in 6 / 18 of NP patients stereotypically at the edges of B cell follicles surrounding the germinal center (Fig. 9F, right). This latter specific B- and CD8 T-cell interaction and its spatial pattern aligned with our earlier observation in the context of NN-vs-NP (Fig. 8D). Altogether, interactions between B- and CD8T-cells at the boundaries of B cell follicles surrounding the germinal center could be an early specific marker of future widespread metastasis in tumor-free lymph node stage II melanoma patients.
[0175] Figs. 9A-9F are diagrams, explaining CISM analysis of NP versus PN, as may be performed by embodiments of system 100 depicted in Fig. 2.
[0176] In Fig. 9A, the melanoma cohort includes 18 patients diagnosed with a primary tumor with no apparent tumor cells spreading to the lymph node (i.e., melanoma stage I-II) with widespread melanoma metastasis five years following treatment (Negative-Positive, NP) as described in Fig. 5 A. The cohort was extended with 13 stage III melanoma patients who healed, i.e., patients with metastatic sentinel lymph nodes, and five years follow up prognosis of tumor-free lymph nodes and no metastases (Positive-Negative, PN).
[0177] In Fig. 9B, the top three highly ranked motifs-induced pairwise interaction.
[0178] In Fig. 9C The mean per-patient number of discriminative motif instances (y-axis) in relation to the number of patients that contained that motif (x-axis). Each dot represents a discriminative motif (n = 210). Note motifs I (NP-associated, green), II (PN-associated, magenta), and III (NP-associated, yellow) that appeared in many patients and in many instances per patient.
[0179] Fig. 9D shows the three discriminative motifs that were marked in Fig. 9C. The edge colors are consistent across the other figure panels.
[0180] In Fig. 9E the number of instances of motifs I-III (top-bottom, log scale) across patients. Only patients with at least one instance were listed.
[0181] In Fig. 9F shows localizing motif II in the tissue of PN patient #86 (left) and motif III in the tissue of NP patient #20 (right). Each dot represents a cell colored according to its cell type: B cell, CD8 T-cell, CD4 T-cell, Germinal Center cells, and Stroma (other cell types are not shown). The germinal center (GC) is surrounded by a dashed white circle. Instances of motifs II and III are marked with dashed magenta and yellow rectangles (larger rectangles indicate multiple instances of the motif) respectively.
[0182] To demonstrate generalization, the inventors applied CISM to analyze a MIBI-TOF-acquired human TNBC cohort consisting of 38 tissue sections of 37 patients. The inventors partitioned the patients into short-term (less than 1000 days, N = 7 patients) and long-term (more than 1000 days, N = 30) survivors according to the time a person lived after diagnosis. Using the cell type classification the inventors applied CISM to extract 3- and 4-cell motifs and then selectedthose that discriminated between short- and long-term survivors with a smaller number of shortterm associated discriminative motifs attributed to the imbalance between the disease states. Unlike the melanoma cohort, in which the lymph nodes did not necessarily include tumor cells (NN and NP), in the TNBC cohort the inventors did consider the tumor cells and their interactions with cells in the tumor microenvironment. Machine learning validation reached an AUC of 0.85, surpassing pairwise cell type machine learning and GNN (Fig. 10A). Discriminative motifs-induced pairwise interactions revealed that long-term survivors were associated with the motif-induced pairwise interactions of DC / MONO cells with CD4 T cells, CD8 T cells with CD4 T cells, and between CD4 T cells, while the motif-induced pairwise interactions associated with short-term survival were of mesenchymal cells with CD8 T cells and Neutrophils with CD8 T cells (Fig.10B). These results propose interactions of CD4 T cells with other immune cells as a marker for long-term survival in contrast to interactions of CD8 T cells with neutrophils and mesenchymal cells. For spatial interpretation, the inventors selected three discriminative motifs that were associated with long-term survivors (Fig. 10C-E), but did not find a clear stereotypic spatial localization of these motifs (Fig. 10F). Cumulatively, these results demonstrated that CISM is a general method to identify local cell structures associated with a disease state in single cell spatial data.
[0183] Figs. 10A-10F are diagrams, explaining CISM analysis of triple negative breast cancer cohort, as may be performed by embodiments of system 100 depicted in Fig. 2.
[0184] In Fig. 10A, validation of CISM’s discriminative motifs representation of TNBC shortterm versus long-term survival discrimination. Leave one patient out cross-validation (LOOCV) model performance (AUC) (y-axis) over the discrimination stringency parameter (x-axis) (3 -and 4-cell motifs in green and purple correspondingly). The pink, cyan, and black dashed horizontal lines represent the LOOCV AUC score of pairwise, GNN, and a random null model, correspondingly. The discriminative motifs derived from 3-cell motifs with the discrimination stringency parameter of 0.45 were used for further analysis.
[0185] Fig. 10B shows the top three highly ranked interactions of motifs induced pairwise interaction. Note that short-term survival had a single discriminative motif and thus had only two motif-induced pairwise interactions.
[0186] Fig. 10C shows the mean per-patient number of discriminative motif instances (y-axis) in relation to the number of patients that contained that motif (x-axis). Each dot represents adiscriminative motif (n = 33). Note that motifs I (green), II (magenta), and III (yellow) are all associated with long-term survival and appeared in many patients and in many instances per patient.
[0187] Fig. 10D shows the three discriminative motifs that were marked in Fig. 10C. The edge colors are consistent across the other figure panels.
[0188] Fig. 10E shows the number of instances for motifs I-III (top-to-bottom, log scale) across patients. Only patients with at least one instance were listed.
[0189] Fig. 10F shows localizing motifs I-III in the tissue of the long-term survivor patient #1. Each dot represents a cell colored according to its cell type (right legend; other cell types are not colored). Instances of motifs I, II, and III are marked with dashed green, pink, and yellow rectangles correspondingly. Larger rectangles indicate multiple instances of the motif.
[0190] System 100 may provide a two-step method for identification and characterization of fine-grained spatial inter-cellular modules associated with physiological tissue states. First, using motif extraction module 130 with an enhanced implementation, system 100 may extract, in an unsupervised manner, all spatial motifs 133 in the patients' multicellular spatial networks 120NW. Second, feature extraction module 140 may select, in a supervised manner, discriminative motifs 142 according to the patients' context of interest. This approach may distill discriminative motifs 142 from the landscape of all possible cell-cell interactions. These discriminative motifs 142 may then be used for several downstream applications that the inventors have demonstrated in melanoma and triple-negative breast cancer cohorts: machine learning prediction of tissue's physiological context using ML based classification model 161, identifying signature discriminative motifs 142 and motifs-derived pairwise interactions that associate with the physiological context, and spatial interpretation through spatial distribution analysis module 143 to characterize discriminative motifs' stereotypic localization patterns.
[0191] System 100 may outperform graph neural networks and pairwise interactions representations in predicting disease tissue state, indicating that discriminative motifs 142 may serve as powerful representations of tissue disease state. According to some embodiments, the spatial organization of cells within the motifs, in terms of their specific cell-cell adjacencies as indicators of potential interactions, may contribute to the discrimination between different biological conditions 161C. These results may suggest that the local organization of a few cells asdiscriminative motifs 142 are emergent properties that may define an intermediate spatial scale driving tissue function.
[0192] Although graph neural networks have shown promise in classifying tissue states across multiple diseases, their success may come at the cost of poor interpretability. This inherent difficulty in providing human meaningful explanations for neural network decisions may hamper the possibility of generating specific mechanistic hypotheses. The approach of feature extraction module 140 for discriminative motif 142 selection may resemble more interpretable machine learning approaches, with engineered features that are easier to decipher. The exhaustive enumeration of all motifs by motif extraction module 130 may provide joint properties of discrimination power along with interpretability.
[0193] According to some embodiments, system 100 may offer several layers of interpretability. First, pairwise analysis module 145 may generate discriminative motifs-induced pairwise interactions disease signatures that highlight specific cell-cell interactions that associate with the tissue state. Second, spatial distribution analysis module 143 may enable mapping discriminative motifs 142 back to their location in the tissue space to link different spatial scales by characterizing the fine scale motifs localization in relation to specific spatial contexts. The inventors have demonstrated these interpretability capacities by associating interactions between B and CD8 T-cells at the edges of B cell clusters surrounding germinal centers in patients who developed metastases in both contexts analyzed.
[0194] CISM approach of feature (i.e., discriminative motif) selection resembles the more interpretable classic machine learning approach, with “engineered” features that are easier to decipher. The exhaustive enumeration of all potential motifs followed by the selection of those motifs that discriminate between the disease states, provides these desired joint properties of discrimination power along with interpretability. Specifically, CISM offers several layers of interpretability (Fig. 3D). First, discriminative motifs-induced cell distribution and pairwise interactions highlight specific cell-cell interactions that, when considered in context, associate with the tissue state. Second, the ability to map discriminative motifs back to their location in the tissue space enables to link the two different spatial scales by characterizing the (fine scale) motif’s localization in relation to a (coarse scale) specific spatial context, the inventors demonstrated these interpretability capacities by associating interactions between B and CD8 T-cells at the edges of B cell follicles surrounding the germinal center in melanoma NP patients when comparing to eitherPN or NN contexts. Some of the melanoma motifs included ’’Unidentified” cells. While these cells did not have distinguishable markers in the protein data, concomitant transcriptomics could suggest that some of these may be plasma blasts. The role of these interactions should be further evaluated in functional studies.
[0195] When applying CISM to TNBC, the inventors associated long-term survival with interactions of CD4 T cells with dendritic cells aligning with a recent study suggesting that interactions between dendritic cells and T cells induce a more effective immune response. Moreover, the discriminative motif that appeared in the vast majority (23 / 30) of long-term survival patients was composed of edges between a CD4 T cell, a dendritic cell and a tumor cell, aligning with another study that argued that interactions between CD4 T cells and dendritic cells suppress tumor cells. For short-term survivors, the inventors identified associated interactions between Neutrophil cells and CD8 T cells that were proposed as protumorigenic by inhibiting CD8 T cells proliferation and more generally neutrophils have been previously suggested to modulate adaptive immune responses by inhibiting T cells.
[0196] Beyond enabling biological insight at the spatial scale of a few cells, CISM benefits from several additional advantages. The strategy of unsupervised motif enumeration followed by supervised context-dependent selection enables us to use the same motifs to answer different questions as shown for melanoma NP versus NN / PN. Moreover, this approach is suitable for characterizing differences in small cohorts, for example, in matched healthy-versus-disease tissue of the same single patient. However, CISM also has some limitations. The fine resolution of a few cells and their spatial arrangement inherently implies that CISM is less forgiving of errors in cell segmentation and in cell type classification. Also, CISM has multiple parameters that must be carefully configured. First, the cell type lineage resolution. Using more cell types provides better discrimination and more specific biological insight at the cost of combinatorially increasing the size of the motifs space. Thus, beyond increasing the required computational resources, this implies a dramatic decrease in the number and prevalence of discriminative motifs. Similar considerations take place regarding the choice of the motif’s size and the discrimination stringency parameter. Increasing these parameters should lead in principle to better discrimination, but is limited by the lower probabilities to generalize across patients. Practically, CISM users should balance these parameters according to quantitative readouts that include the number of discriminative motifs (Fig. 5C) and the discrimination performance (Fig. 5D) as a function of thediscrimination stringency parameter. Automated parameter calibration is a possible direction for future research. The strict definition of a discriminative motif, of appearing in a sufficient fraction of patients of one disease state but excluded from all patients of the other disease state, could be too restrictive, especially with the expected increase in cohorts’ sizes in upcoming years, which inherently decreases the probability for the presence of discriminative motifs. This requirement for “hard” discrimination can lead to situations where motifs appear in many patients and even in many instances of patients of one disease state but also in very small fractions of patients and in small numbers of instances of the other disease state. These “almost discriminative” motifs can have great discriminative and interpretative value but are currently entirely excluded from our analysis. To remedy this limitation, CISM could incorporate a “softer” discriminative criterion such as the distance between distributions of motif frequencies of each disease state (e.g., the Wasserstein distance. The implementation of “soft“ discrimination could also be useful to extend CISM to multiple disease states and to incorporate motifs with finer resolution in their cell types and cell states.
[0197] The inventors analyzed an existing-acquired cohort of 51 human melanoma patients from Leeat Keren’s lab at the Weizmann Institute of Science. The patients’ disease state was determined according to the clinical diagnostic of the sentinel lymph nodes (SLN) and then at least five years following treatment, except for one living patient who had a two-year follow-up. Patients with or without the existence of tumor cells in their sLNs were defined as “positive” or “negative,” respectively. Patients who were diagnosed with the prognosis of widespread metastases in distant organs five years following treatment were defined as “negative -positive” (NP), while those who did not were defined as “negative-negative” (NN) or “positive-negative” (PN) according to their initial diagnosis. Cumulatively, the inventors analyzed three patient groups: NN (N patients = 20, n tissue sections FOVs = 40), NP (N = 18, n = 36), and PN (N = 13, n = 42). The dataset includes 39-channel MIBLTOF of sLNs with a resolution of 0.45X0.45 pm per pixel. In the original manuscript, each cell was classified into one of 26 cell types, the inventors reduced the cell type resolution to ’’tumor” and 15 non-tumor lymph node microenvironment cell types according to their lineage to increase the number of discriminative motifs and their occurrences, the inventors included ’’unidentified” cells in our machine learning analyses and excluded them from post-processing biological interpretability for selecting the three mostabundant pairwise co-localized cell types (Fig. 6D, Fig. 9B) and for selecting an example for spatial arrangement analysis compared to the cell type (Fig. 7C).
[0198] The inventors analyzed a published MIBI-TOF-acquired human cohort TNBC dataset that included 37 patients and 38 tissue sections with 38 protein channels and a spatial resolution of 0.5x0.5 pm per pixel, the inventors used cell segmentation masks and assignment of each cell to one of 16 cell types as in the original study. The cell types used in this dataset include Tumor, Endothelial, Mesenchyme, Tregs, CD4 T cells, CD8 T cells, CD3 T cells, NK cells, B cells, Neutrophils, Macrophages, DC, DC / Mono, Mono / Neu, Immune other and Unidentified, the inventors used the number of survival days since diagnosis as the clinical readout where patients who survived less than 1,000 days were defined as “short-term” survivors, and those who survived at least 1,000 days were defined as “long-term” survivors.
[0199] The inventors defined the spatial multicellular network, where single cells define the network nodes and are “colored” according to their cell type. Close-adjacent pairs of cells were connected with an edge using the Delaunay triangulation, excluding those edges between cells that distant more than 50 pm away from one another.
[0200] In the melanoma cohort, the inventors focused on disease-associated alterations of the lymph node immune microenvironment, specifically because NN and NP patients did not have tumor cells in their lymph nodes. Thus, when handling PN patients, the inventors had to exclude tumor cells, the inventors opted to exclude from the multicellular network all cells (tumor and nontumor) in regions enriched with tumor cells to avoid potential tumor-immune interactions, as an obvious bias indicative of PN patients, which could serve as a “shortcut” for classification. Thus, the inventors used the “alpha shape”, a generalization of the convex hull for a set of points in space (in our case - cell localizations) based on parameter a that controls the shape’s sensitivity. For a equal to zero, the alpha shape of a point set is the same as its convex hull. As a increases, the alpha shape shrinks with concave regions (derived from the point set) beginning to appear, gradually “cutting out” voids and indentations in the shape, the inventors empirically set a to 0.01 after manual calibration. Overall, the inventors adjusted the multicellular network using the following pipeline: (1) identify all tumor cells, (2) divide them into connected components in the multicellular network, (3) calculate the alpha shape, and (4) exclude all cells inside the alpha shape from further analysis.
[0201] The inventors adapt and extend the concept of “network motifs”, reoccurring subnetworks that define “building blocks” of complex networks, to characterize the fine-scale organization of multicellular spatial networks of human disease tissues. One of the important characteristics of the motif is a statistical measurement of its significance. A subgraph is considered a motif if its frequency in the observed graph precedes its frequency in random graphs generated by edge-switching that preserves the graph’s degree distribution. There are several existing tools to extract network motifs, but all have limitations regarding the motif size (i.e., number of cells) and the number of node “colors”, and graph type support (i.e., multigraph, directed / undirected graph). The only method that the inventors found that supports node coloring is called FANMOD. However, FANMOD has limitations on the number of colors as a function of the motifs’ size that it can analyze. Specifically, the maximum number of node colors that FANMOD can analyze is 15, and this number decreases as the motif size increases: at most, 15 colors for 4-node motifs and no coloring for 8-node motifs. This is a severe limitation for spatial single cell multiplexed data characterized by multiple cell types, the inventors realized that this was an implementation (rather than algorithmic) limitation and thus adjusted the FANMOD source code, using the C++ boost library, to enable larger motif detection.
[0202] In more detail. To enumerate the set of all unique subgraphs within the network, FANMOD generates a unique representation of each subgraph, known as the “canonical representation”, that is invariant for different organizations of nodes and edges (i.e., isomorphic graphs are mapped to the same representation). Specifically, FANMOD is using the Nauty algorithm to generate a canonical representation that is encoded as a sequence of bits. FANMOD+ increases the space (i.e., number of bits) for the canonical representation from 64-bit fixed allocation to 128-bit dynamic allocation (i.e., physical memory allocation extended when used). Thus, FANMOD+ subgraph canonical representation enables the encoding of more colors and larger motifs, supporting the analysis of modern spatial multiplex single cell omics data. For instance, with 0-32 colors, FANMOD+ can support a maximal motif size of 5 nodes; with 128 colors, FANMOD+ can support a maximal motif size of 4 nodes. Note that increasing the motif size or the number of colors still has a computational cost.
[0203] To assess how the execution time increases as a function of more colors and larger motifs, the inventors performed benchmarking using randomly generated graphs. Since, on average for the Melanoma dataset, there were ~ 10,000 cells in each patient’s sample, the inventorsgenerated random graphs with n = 10,000 cells, each represented as a pixel, randomly spread on a grid of size 10,000 x 10,000 pixels and connected nearest neighbors cells using Delaunay triangulation with —30,000 edges (~3n following the Euler rule for planar graph when n > 3). The inventors generated five random graphs for each combination of motif size (3-5) and colors (in the range of 8-128). Due to RAM usage limitations of 64GB, the inventors limited our benchmarking to five nodes and 32 colors. The benchmarking was executed on an AMD Ryzen 9 5900HX CPU with 64GB RAM. The average runtimes for 3 and 4 -node motifs with 128 colors were 30 seconds and 1846 seconds (i.e., —30 minutes) correspondingly. For 5-node motifs with 32 colors, the average runtime was 1539 seconds. Overall, the runtime increases exponentially with the number of colors.
[0204] The inventors applied FANMOD+ to each patient’s tissue sample to extract and enumerate thousands, tens-of-thousands and hundreds-of-thousands of motifs of size 3-to-5 cells from the space of hundreds-of-thousands, a million and a few millions candidate subgraphs, correspondingly (Table SI). The inventors calculated each n-cells motif’s frequency with respect to the total number of corresponding subgraphs of size n in the patient’s sample. The inventors used 1000 random graphs for the statistical test and included only the significant motifs for the analysis (p-value < 0.05). To illustrate the definition of what defines a subgraph as a motif, the inventors provide a specific example. Consider a network in a single field of view that contains 10,000 nodes (cells), each with 1 of 10 colors (cell types) and 30,000 edges (cell-cell interactions) (Euler rule for planner graph). For a 4-node subgraph with 3 edges each, the probability for a given cell type composition is, and the probability for the cells arrangement is a subgraph is. Thus, the estimated expected number of subgraph instances is E(4 node candidates) * P(structure) * P(cell types) ~ (410000) * 2.16 * 10-10* 10-4- 4.17 * 1014* 2.16 * 10-10* 10-4- 9. Thus, in this specific example, the number of instances in the original network would need to significantly exceed 9 instances to be considered a motif.
[0205] Context-dependent discriminative motifs are those motifs that sufficiently appear in patients of one physiological context and do not appear at all in patients of the other physiological context. The degree of presence of a motif in each context is measured as the fraction of patients in that context that had at least one instance of that motif. Discriminative motifs were selected based on a tunable discrimination stringency parameter that defines a threshold on the degree ofpresence in a given context with the constraint that the motif does not appear in any of the patients of the other context.
[0206] To reduce the computational complexity, improve interpretability and remove redundant features, the inventors selected the top ndisc_motifs discriminative motifs in each leave-one-out iteration. The ranking of the discriminative motifs was determined according to their frequencies - the ratio between the number of motif instances to the total number of corresponding subgraphs (of the motif’s size) in the patient’s sample. Our assumption was that motifs appearing in more patients and in more instances per patient are a preferable representation of the physiological context.
[0207] To assess the discriminative capacity of motifs’ representation of the disease tissue state, the inventors performed Random Forest-based leave-one-patient-out cross-validation for binary classification between two disease contexts. The reason for using a Random Forest model was its simplicity and stable performance even with default settings. The reason for the leave-one-patient-out cross-validation was the small cohorts sizes that are usually limited to tens of patients. Thus, the inventors performed iterations of training and testing, each iteration with one patient at the test and the rest used for training. Each iteration followed the processes depicted in Fig.3C (left-to-right): (1) keeping one patient out, (2) extracting motifs from all remaining patients, (3) selecting discriminative motifs and ranking the top n discriminative motifs (i.e., ndisc_motifs), (4) representing a patient by the (sparse) feature vector encoding the per-patient frequencies of these highest ranked discriminative motifs. The inventors used the Random Forest binary classification model with default parameters because of its simplicity, efficient and fast execution, and being suitable for interpretability. From the aggregated patient classification probabilities, the inventors computed the receiver operating characteristic curve (ROC) and calculated the corresponding area under the curve (AUC) score. The inventors set the number of features (ndisc_motifs) to n = 30 in all classification tasks throughout this manuscript because it provided a nice balance between stable performance and rich representation. By stable performance, the inventors refer to a representation that was not susceptible to fluctuations in the AUC scores as a function of the discrimination stringency parameter that typically occurs when using a small number of features. By rich representation, the inventors refer to a sufficient number of discriminative motifs for high discriminative performance along with a manageable number of discriminative motifs for downstream evaluation and interpretation. Since the heuristic approachof ranking the motif’s importance and selecting a subset to train the model, the AUC as a function of the discrimination stringency parameter is not necessarily monotonic.
[0208] To assess if the prediction results and the selected features are significant and not overfitted to the data, the inventors conduct a permutation test by permutating the patient’s disease states, i.e., their classification labels. The inventors executed the complete evaluation pipeline and measured the ROC-AUC score for each permutation. The inventors did not consider permutations that led to no discriminative motifs in at least one of the iterations. This permutation statistical test’s p-value was determined as the ratio between the number of permutations that resulted with AUC scores higher than the observed labels, and the total number of permutation trials.The inventors used machine learning to evaluate the discriminative performance of cell type distribution. The inventors followed the same process, i.e., the Random Forest leave-one-out- cross-validation. The inventors computed the receiver operating characteristic curve (ROC) from the aggregated patient classification probabilities and calculated the corresponding area under the curve (AUC) score.
[0209] The inventors measure the probability of an edge connecting two cell types across patients for each disease state. First, the inventors created a pairwise cell-type interactions matrix with the probability of an edge connecting to two cell types for each field of view (FOV) from a patient at the given disease state:Where I is a field of view and i, j are cell types. Then, we summed the overall frequencies across the field of views and normalize:‘ Type PairWhere I is a field of view and a, b, i,j are cell types.
[0210] The inventors used machine learning to evaluate the discriminative capacity of pairwise interactions. The inventors represented a patient with the probabilities of the pairwise cell-type interactions:goundou)H H £ ( &0Where I is a field of view, L(p) is the patient’s FOVs and i, j are cell types. The inventors followed the Random Forest leave-one-out-cross-validation. From the aggregated patient classification probabilities, the inventors computed the receiver operating characteristic curve (ROC) and calculated the corresponding area under the curve (AUC) score.
[0211] The discriminative motif-based cell type distribution was derived from the pool of discriminative motifs. For a given disease state the inventors pooled all the discriminative motif instances across patients’ field of views (FOVs) from that disease state and calculated the motifs- induced cell type distribution:CellTypeDistribution(i)Where i is cell type, m is a discriminative motif, DM is the set of all discriminative motifs derived from all patients of a given disease state, and wmis the total count of discriminative motif instances across patients.
[0212] The discriminative motif-based pairwise interactions were derived from the pool of discriminative motifs. For a given disease state the inventors pooled all the discriminative motif instances across patients’ field of views (FOVs) from that disease state and calculated the motifs- induced pairwise interactions:InteractionMatrix(i,j) =Where i and j are cell types, m is a discriminative motif, DM is the set of all discriminative motifs derived from all patients of a given disease state, and wmis the mean motif frequency across patients.
[0213] To measure the contribution of the motif structure compared to only considering the cell type composition, the inventors replaced each discriminative motif with its corresponding cell type composition vector and used machine learning to evaluate the derived cell type composition representation in comparison to the corresponding motifs’ representation. Next, the inventors incorporated motif information back to the cell type composition representation, one motif at a time, to evaluate the residual contribution of each motif to the cell type composition. More specifically, the cell type composition vector derived from a motif was represented as a vector of a fixed length according to the number of cell types, where each entry holds the number ofinstances of the corresponding cell type in the motif. For example, cell type composition I in Fig.7C (green column) is composed of a B-cell, a HEVs cell, and two Memory CD4 T-cells. Note that different motifs can be mapped to the same cell type composition. For instance, two motifs with different inner structures that comprise 2 * x, 1 * y, and 2 * z cell types are mapped to the same cell type composition vector of (2, 1, 0, 2), given cell types (x, y, t, z). Thus, the evaluation was performed as follows. First, the inventors replaced all discriminative motifs with their corresponding cell type composition vectors and measured the AUC of the cell type composition representation using Random Forest leave-one -patient-out as the baseline. Note that each cell type composition can correspond to multiple motifs, where some of these motifs do not meet the strict discrimination criterion used to define the discriminative motifs.
[0214] Second, starting from this cell type composition representation, the inventors introduced back the spatial information of one cell type composition by replacing the cell type composition feature with its corresponding discriminative motifs. This replacement induced a hybrid representation that contains both cell type composition and discriminative motif features. This representation was used for machine learning assessment, where the classification gain with respect to the cell type composition baseline was attributed to the contribution of the corresponding motifs’ spatial arrangement. This process was iteratively repeated for each cell type composition, that were then ranked according to the contribution of the corresponding motifs’ spatial arrangement. The difference between the residual gain and the baseline score captures the contribution of the specific set of discriminative motifs for prediction compared to the generalization of the cell type composition.
[0215] To evaluate the motifs cell type composition information without the specific inner structure, the inventors repeated the steps of context-dependent motifs’ selection and machine learning analysis, but instead of using the discriminative motifs, the inventors used the discriminative cell type composition representations, i.e., cell type compositions that sufficiently appear in patients of one physiological context and do not appear at all in patients of the other physiological context. Each patient was represented by the (sparse) vector of its cell type composition frequencies. For instance, if multiple motifs are mapped to a single cell type composition, then the frequency is calculated as overall number of instances across the corresponding motifs count with respect to the total corresponding subgraphs. The cell typecomposition representation is less specific than the motif representation and, therefore, expects to generate a smaller number of discriminate cell type compositions.
[0216] The inventors performed a whole graph classification to predict the disease state using a two-layer graph convolutional network (GCN), and a fully-connected layer at the end (classification head). The spatial multicellular network defines the input for the GCN where the attribute of each node was its corresponding cell type. The inventors aggregated the information across nodes after the final classification layer, using global pooling (average between the global max pooling and the global mean pooling), to get a single disease state prediction per sample. The model was optimized with the Adam optimizer and using the cross-entropy loss function:L(y, ŷ) = −∑
[0217] The inventors performed leave-one -patient-out cross-validation. The training was as follows: each patient was held out for test, and the remaining samples were split into trainvalidation sets (75% train, 25% validation). After 500 epochs, the inventors selected the optimal weights according to the epochs’ validation losses, and the inventors used the corresponding model for the test. The reason for the train-validation-test splits was to avoid overfitting due to the relatively small dataset.
Claims
CLAIMS1. A method of determining a condition of a biological tissue sample, the method comprising: receiving a target image depicting spatial single-cell omics data of cells in the tissue sample;constructing a multicellular network of interconnected nodes, wherein each node represents a cell in the target image;enumerating subgraphs within the multicellular network to identify a set of spatial motifs, wherein each spatial motif represents a multicellular pattern of interconnected cells of specific types;analyzing the set of spatial motifs to generate one or more motif-based features, characterizing the biological tissue sample; andapplying one or more trained machine learning (ML) models to the motif-based features to predict the condition of the biological tissue sample.
2. The method of claim 1, wherein constructing the multicellular network comprises:segmenting the target image to identify individual cells and determine spatial coordinates for each cell;classifying the individual cells according to their respective cell types based on the spatial single-cell omics data;representing each cell as a node in the multicellular network, wherein each node is attributed with its respective cell type and spatial location; andconnecting spatially adjacent cells with edges based on spatial proximity, wherein spatial proximity is determined using a predetermined neighborhood criterion.
3. The method according to any one of claims 1 and 2, wherein enumerating subgraphs within the multicellular network comprises:examining subgraphs comprising predetermined numbers of interconnected nodes within the multicellular network;for each examined subgraph, determining a canonical pattern based on the cell types of the nodes and their connectivity structure;calculating a frequency of occurrence for each canonical pattern across the multicellular network; andselecting canonical patterns whose frequency exceeds a predetermined threshold, to obtain the set of spatial motifs.
4. The method of claim 3, further comprising:obtaining a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context; andproviding calculated frequencies of the context-dependent discriminative motifs as motifbased features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
5. The method according to any one of claims 3 and 4, further comprising:obtaining a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;mapping occurrence of discriminative motifs within the multicellular network; based on the mapping, calculating spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample; andproviding the calculated spatial distribution patterns as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
6. The method according to any one of claims 3-5, further comprising:obtaining a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;for one or more discriminative motifs (a) examining pairwise connections between cells of specific types within that discriminative motif, and (b) for each pairwise connection, calculating a pairwise frequency value, representing frequency of occurrence of that pairwise connection in the multicellular network; andproviding the calculated pairwise frequency values as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
7. The method according to any one of claims 4-6, wherein obtaining the context-dependent discriminative motifs comprises:receiving a training dataset comprising a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image is annotated according to a biological condition pertaining to that context;identifying a preliminary set of spatial motifs based on the plurality of training images; performing an iterative training process, wherein each iteration comprises (i) training an ML model of the one or more ML models to predict the biological condition based on a subset of the preliminary set of spatial motifs, (ii) evaluating precision of said prediction based on said training, and (iii) altering the subset for training the ML model in a subsequent training iteration; andselecting an optimal subset of the preliminary set of spatial motifs based on the evaluated precision of said iterations as the context-dependent discriminative motifs.
8. The method according to any one of claims 4-7, wherein obtaining the context-dependent discriminative motifs comprises:receiving a training dataset comprising a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image is annotated according to a biological condition pertaining to that context;identifying a preliminary set of spatial motifs based on the plurality of training images; identifying a subset of the preliminary set of spatial motifs that appear in tissue samples associated with a first biological condition and do not appear in tissue samples associated with a second biological condition; andselecting the identified subset as the context-dependent discriminative motifs for distinguishing between the first biological condition and the second biological condition.
9. The method according to any one of claims 1-8, wherein (a) the context pertains to cancer metastatic potential, and the biological conditions comprise metastasis-free patients versuspatients who develop distant metastases within a predetermined time period, (b) the context pertains to cancer survival prognosis, and the biological conditions comprise short-term survivors versus long-term survivors, (c) the context pertains to treatment response assessment, and the biological conditions comprise treatment-responsive patients versus treatment-resistant patients, or (d) the context pertains to disease progression monitoring, and the biological conditions comprise patients with stable disease versus patients with progressive disease.
10. The method according to any one of claims 5-9, wherein calculating spatial distribution patterns comprises:obtaining one or more tissue structures within the biological tissue sample; determining localization of the discriminative motifs relative to the one or more tissue structures; andcharacterizing the spatial distribution patterns based on the determined localization.
11. The method of claim 10, wherein the one or more tissue structures are selected from a list consisting of: (a) germinal centers, wherein the discriminative motifs are localized at edges of B cell clusters surrounding the germinal centers, (b) tumor boundaries, wherein the discriminative motifs are localized at interfaces between tumor regions and immune microenvironment; and (c) immune cell clusters, wherein the discriminative motifs are localized within or adjacent to specific immune cell aggregations.
12. The method according to any one of claims 3-11, further comprising:obtaining a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;generating motif-based features by (a) calculating frequencies of the context-dependent discriminative motifs, (b) calculating spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample, and (c) calculating pairwise frequency values representing frequency of occurrence of pairwise connections between cells of specific types within the discriminative motifs in the multicellular network; andapplying an ML model of the one or more ML models to the motif-based features to predict the condition of the biological tissue sample.
13. The method according to any one of claims 4-12, wherein the one or more trained ML models comprise a Random Forest classifier that facilitates interpretability by:calculating feature importance scores for individual motif-based features, representing contribution of these motif-based features to prediction of the biological condition;ranking the motif-based features based on the calculated feature importance scores to identify top-contributing features for the prediction; andgenerating a layered presentation comprising the target image, the multicellular network, and top-scoring motif-based features via a user interface, to highlight contributors to the prediction.
14. The method according to any one of claims 5-13, further comprising:determining potential therapeutic targets for drug development based on the spatial distribution patterns of the discriminative motifs and the predicted biological condition.
15. A system for determining a condition of a biological tissue sample, the system comprising: a processor, and memory storing instructions that, when executed by the processor, cause the processor to:receive a target image depicting spatial single-cell omics data of cells in the tissue sample; construct a multicellular network of interconnected nodes, wherein each node represents a cell in the target image;enumerate subgraphs within the multicellular network to identify a set of spatial motifs, wherein each spatial motif represents a multicellular pattern of interconnected cells of specific types;analyze the set of spatial motifs to generate one or more motif-based features, characterizing the biological tissue sample; andapply one or more trained machine learning (ML) models to the motif-based features to predict the condition of the biological tissue sample.
16. The system of claim 15, wherein the processor is further configured to:segment the target image to identify individual cells and determine spatial coordinates for each cell;classify the individual cells according to their respective cell types based on the spatial single-cell omics data;represent each cell as a node in the multicellular network, wherein each node is attributed with its respective cell type and spatial location; andconnect spatially adjacent cells with edges based on spatial proximity, wherein spatial proximity is determined using a predetermined neighborhood criterion.
17. The system according to any one of claims 15-16, wherein the processor is further configured to:examine subgraphs comprising predetermined numbers of interconnected nodes within the multicellular network;for each examined subgraph, determine a canonical pattern based on the cell types of the nodes and their connectivity structure;calculate a frequency of occurrence for each canonical pattern across the multicellular network; andselect canonical patterns whose frequency exceeds a predetermined threshold, to obtain the set of spatial motifs.
18. The system of claim 17, wherein the processor is further configured to:obtain a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context; andprovide calculated frequencies of the context-dependent discriminative motifs as motifbased features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
19. The system according to any one of claims 17-18, wherein the processor is further configured to:obtain a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;map occurrence of discriminative motifs within the multicellular network;based on the mapping, calculate spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample; andprovide the calculated spatial distribution patterns as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
20. The system according to any one of claims 17-19, wherein the processor is further configured to:obtain a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;for one or more discriminative motifs (a) examine pairwise connections between cells of specific types within that discriminative motif, and (b) for each pairwise connection, calculate a pairwise frequency value, representing frequency of occurrence of that pairwise connection in the multicellular network; andprovide the calculated pairwise frequency values as motif-based features to an ML model of the one or more ML models, to predict the condition of the biological tissue sample.
21. The system according to any one of claims 18-20, wherein obtaining the context-dependent discriminative motifs comprises the processor being configured to:receive a training dataset comprising a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image is annotated according to a biological condition pertaining to that context;identify a preliminary set of spatial motifs based on the plurality of training images; perform an iterative training process, wherein each iteration comprises (i) training an ML model of the one or more ML models to predict the biological condition based on a subset of the preliminary set of spatial motifs, (ii) evaluating precision of said prediction based on said training, and (iii) altering the subset for training the ML model in a subsequent training iteration; andselect an optimal subset of the preliminary set of spatial motifs based on the evaluated precision of said iterations as the context-dependent discriminative motifs.
22. The system according to any one of claims 18-21, wherein obtaining the context-dependent discriminative motifs comprises the processor being configured to:receive a training dataset comprising a plurality of training images depicting spatial single-cell omics data in a respective plurality of tissue samples, wherein each training image is annotated according to a biological condition pertaining to that context;identify a preliminary set of spatial motifs based on the plurality of training images; identify a subset of the preliminary set of spatial motifs that appear in tissue samples associated with a first biological condition and do not appear in tissue samples associated with a second biological condition; andselect the identified subset as the context-dependent discriminative motifs for distinguishing between the first biological condition and the second biological condition.
23. The system according to any one of claims 15-22, wherein (a) the context pertains to cancer metastatic potential, and the biological conditions comprise metastasis-free patients versus patients who develop distant metastases within a predetermined time period, (b) the context pertains to cancer survival prognosis, and the biological conditions comprise short-term survivors versus long-term survivors, (c) the context pertains to treatment response assessment, and the biological conditions comprise treatment-responsive patients versus treatment-resistant patients, or (d) the context pertains to disease progression monitoring, and the biological conditions comprise patients with stable disease versus patients with progressive disease.
24. The system according to any one of claims 19-23, wherein calculating spatial distribution patterns comprises the processor being configured to:obtain one or more tissue structures within the biological tissue sample;determine localization of the discriminative motifs relative to the one or more tissue structures; andcharacterize the spatial distribution patterns based on the determined localization.
25. The system of claim 24, wherein the one or more tissue structures are selected from a list consisting of: (a) germinal centers, wherein the discriminative motifs are localized at edges of B cell clusters surrounding the germinal centers, (b) tumor boundaries, wherein the discriminative motifs are localized at interfaces between tumor regions and immune microenvironment; and (c) immune cell clusters, wherein the discriminative motifs are localized within or adjacent to specific immune cell aggregations.
26. The system according to any one of claims 17-25, wherein the processor is further configured to:obtain a context-dependent subset of the set of spatial motifs, wherein the subset comprises discriminative motifs which distinguish between different biological conditions pertaining to that context;generate motif-based features by (a) calculating frequencies of the context-dependent discriminative motifs, (b) calculating spatial distribution patterns characterizing locations of specific discriminative motifs within the biological tissue sample, and (c) calculating pairwise frequency values representing frequency of occurrence of pairwise connections between cells of specific types within the discriminative motifs in the multicellular network; andapply an ML model of the one or more ML models to the motif-based features to predict the condition of the biological tissue sample.
27. The system according to any one of claims 18-26, wherein the one or more trained ML models comprise a Random Forest classifier that facilitates interpretability by the processor being configured to:calculate feature importance scores for individual motif-based features, representing contribution of these motif-based features to prediction of the biological condition;rank the motif-based features based on the calculated feature importance scores to identify topcontributing features for the prediction; andgenerate a layered presentation comprising the target image, the multicellular network, and top-scoring motif-based features via a user interface, to highlight contributors to the prediction.
28. The system according to any one of claim 19, wherein the processor is further configured to determine potential therapeutic targets for drug development based on the spatial distribution patterns of the discriminative motifs and the predicted biological condition.
29. The system of claim 28, wherein the potential therapeutic targets comprise specific cell-cell interaction patterns represented by the discriminative motifs that are associated with the biological condition.
Citation Information
Patent Citations
Systems and methods for tissue classification
US20130051650A1
Pathology case review, analysis and prediction
US20190087954A1
Predicting overall survival in early stage lung cancer with feature driven local cell graphs (FEDEG)
US20220405931A1
A computer-implemented, graph-based method of analysing an image of a tissue specimen
WO2024052703A1