Disease characterization

By constructing individual networks for each patient and integrating them into similarity matrices, the method addresses the limitations of single-modality disease subtyping, enhancing predictive power for cancer outcomes through flexible multi-modal data integration.

CN120283286APending Publication Date: 2025-07-08F HOFFMANN LA ROCHE & CO AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380081733.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing methods for disease subtype analysis usually rely on a single data mode, fail to fully utilize the complementarity of multimodal data, and existing methods may reduce the signal-to-noise ratio or have high computational complexity when data is integrated, making it difficult to effectively characterize complex diseases.

Method used

Using individual-specific network construction methods, the interactions between multiple biological factors are integrated by generating individual networks and calculating similarity matrix between individuals, and using machine learning models to predict disease diagnosis or prognosis.

Benefits of technology

It improves the accuracy and prediction ability of disease subtype classification, can flexibly integrate multiple modal data, enhances the ability to capture individual heterogeneity, and supports precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283286A_ABST
    Figure CN120283286A_ABST
Patent Text Reader

Abstract

A computer-implemented method of providing disease diagnosis or prognosis for a patient is described. These include generating individual networks, each individual network comprising a plurality of nodes and edges between pairs of nodes, where each node indicates a biological factor in biological data of an individual and each edge indicates a relationship between a pair of biological factors corresponding to the node to which the edge is connected in a respective individual; determining values of one or more similarity indicators between one or more individual networks generated for the patient and one or more individual networks generated for other individuals of the plurality of individuals; and predicting the diagnosis or prognosis of the patient using a machine learning model using the values of the one or more similarity indicators as inputs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods for analyzing biological samples from a subject using machine learning and subject-specific graphs representing multiple biological factors and relationships between biological factors. In particular, the present invention relates to methods for using such methods to provide prognosis, diagnosis, treatment recommendations, or patient selection, and to related systems and devices. Background Art

[0002] Disease subtyping refers to the identification of homogeneous groups of patients. It can be used to detect the severity of a disease or to anchor the treatment regimen with the highest probability of success. Since the molecular types and severities of cancers are highly diverse, disease subtyping is crucial in cancer research. Many methods for disease subtyping analysis rely on only a single data modality. However, it is unlikely that a single modality will have sufficient informational richness to capture the overall complexity of a complex disease. In addition, large datasets are available, making multimodal integration feasible. For example, several studies have examined the benefits of combining image and genomic data (Ash et al. 2021; Schneider et al. 2022).

[0003] Data from multiple modalities can be integrated at different stages of a predictive method. Using late integration, data sources are used independently to obtain classifications, and then the classification results are combined. The main drawback of this method is that it does not exploit the possible complementarity of the modalities. Alternatively, using early fusion, data from multiple data sources can be concatenated before applying a machine learning model. While this solution is easy to implement, concatenation reduces the signal-to-noise ratio of each data modality. In the past decade, alternatives have been proposed that combine data between the start and end steps to address these issues.

[0004] For example, iCluster (Shen et al., 2009) applies data fusion and dimensionality reduction simultaneously to multiple genomic data types. This method uses a Gaussian latent variable model (i.e., jointly estimates latent tumor subtypes from different genomic modalities, assuming that each subtype is linearly related to the latent variable through its own model) and a lasso-type penalty term to induce sparsity in the coefficient matrix for feature selection. A drawback of this method is its high computational complexity. An alternative is affinity aggregation for spectral clustering (Huang et al., 2012). The main idea is to compute a similarity matrix between samples for each data source. Then, these multiple affinity matrices are clustered via spectral clustering using a weighted linear combination, where the weights are optimized using multiple kernel learning. Similarly, similarity network fusion (SNF) (Wang et al., 2014) is implemented to combine multiple similarity matrices into a single matrix by iteratively updating the matrices to make them increasingly similar until the algorithm converges. This final matrix becomes the new input to the classification algorithm. Subsequently, regularized unsupervised multiple kernel learning (Speicher and Pfeifer, 2015) was introduced. This extends multiple kernel learning for dimensionality reduction (projecting samples into a lower-dimensional subspace for further analysis) by adding constraints that result in regularization of the vector controlling the kernel combination to avoid overfitting during the optimization process.

[0005] Despite these advances, there remains a need for improved methods for analyzing biological data to characterize a subject's disease. SUMMARY OF THE INVENTION

[0006] Broadly speaking, the inventors of the present invention postulate that existing methods for personalized disease characterization do not utilize all the information contained in biological datasets, at least because they consider only one variable at a time without considering the interactions between variables. The inventors postulate that this can be addressed by using networks (graphs) to consider the interactions between feature pairs. They further postulate that in this context, using individual (i.e., subject-specific) networks rather than group-based networks would be beneficial because such networks represent the individual relationships between variables for each person and are thus particularly useful for precision medicine such as disease subtype classification. Accordingly, the inventors have developed a multi-step pipeline to predict outcomes via graphs. First, one or more networks are constructed for each individual based on the raw biological data regarding the subject. Then the distance between each pair of subject-specific graphs is calculated to estimate the degree of similarity between individuals. The results are aggregated in a similarity matrix (inter-individual similarity), which serves as the input to a machine learning model. In other words, the inventors consider the similarity to a reference panel as a new variable. Then, the outcome is predicted based on these similarities to the reference set. The inventors have shown that a graph-based method for predicting disease subtype and severity achieves competitive or improved performance compared to methods that do not consider the raw data in a graph-based context. They further show that the process is advantageously capable of flexibly integrating multiple modalities with different characteristics, and intermediate integration is often beneficial in this regard.

[0007] Accordingly, a first aspect provides a method of characterizing a disease of a patient, the method comprising: obtaining biological data for each of a plurality of individuals including the patient, the biological data comprising values of a plurality of biological factors; generating, for each of the plurality of individuals, one or more individual networks, each individual network comprising a plurality of nodes and edges between pairs of the nodes, wherein each node indicates a biological factor in the biological data of the individual, and each edge indicates a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the respective individual; determining values of one or more similarity metrics between the one or more individual networks generated for the patient and the one or more individual networks generated for other individuals of the plurality of individuals; and predicting a diagnosis or prognosis of the patient using a machine learning model configured to predict a diagnosis or prognosis of the disease of the patient, wherein the machine learning model has been trained to take as input values of one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output.

[0008] The method according to this aspect may have one or more of the following optional features.

[0009] Determining the value of one or more similarity metrics can include determining the value of one or more similarity matrices, each similarity matrix including the value of the similarity metric between the individual networks of pairs of the plurality of individuals.

[0010] For each pair of the plurality of pairs of individual networks, the one or more similarity metrics can include the similarity between edges in the individual network, the similarity between edges in the individual network, and the similarity between nodes in the individual network, or a similarity that combines the similarity between edges in the individual network and the similarity between nodes in the individual network.

[0011] Each node in the individual network can have a value that is the value of a biological factor in the biological data of the corresponding individual. Each edge in the individual network can have a value that is the product of the values of the nodes connected by the edge in the corresponding individual. Each edge in the individual network can have a value that is the difference between the edge values of a network obtained using the plurality of individuals in the absence or presence of the corresponding individual. Each edge in the individual network can have a value e x ij = N * (e α ij - e α-x ij ) + e α-x ij , where e α ij is the weight of the edge between nodes i and j in a network modeled on all N individuals in the plurality of individuals, and e α-x ij is the weight of the edge in a network modeled on all samples except the corresponding individual x.

[0012] For each pair in the multiple pairs of individual networks, the one or more similarity metrics may include the similarity between nodes in the individual network, which is obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the respective individual. For each pair in the multiple pairs of individual networks, the one or more similarity metrics may include a similarity that combines the similarity between edges in the individual network and the similarity between nodes in the individual network, where the similarity between nodes in the individual network is obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the respective individual. For each pair in the multiple pairs of individual networks, the one or more similarity metrics may include the similarity between nodes in the individual network that is obtained as a Spearman correlation coefficient, or a similarity that combines the similarity between edges in the individual network and the similarity between nodes in the individual network that is obtained as a Spearman correlation coefficient. For each pair in the multiple pairs of individual networks, the one or more similarity metrics include the similarity between edges in the individual network, which is obtained as an Euclidean distance, a Jaccard distance, an edge difference distance, a DeltaCon distance, a spectral distance, a graphlet-based metric, a Hamming distance, a shortest path kernel, a k-step random walk kernel, a graph diffusion distance, and a feature portrait divergence.

[0013] For each pair in the multiple pairs of individual networks, the one or more similarity metrics include a similarity that combines the similarity between nodes in the individual network and the similarity between edges in the individual network, where the similarity between edges in the individual network is obtained as an Euclidean distance, a Jaccard distance, an edge difference distance, a DeltaCon distance, a spectral distance, a graphlet-based metric, a Hamming distance, a shortest path kernel, a k-step random walk kernel, a graph diffusion distance, and a feature portrait divergence.

[0014] For each pair in the multiple pairs of individual networks, the one or more similarity metrics may include the similarity between edges in the individual network that is obtained as an edge difference distance, or a similarity that combines the similarity between nodes in the individual network and the similarity between edges in the individual network that is obtained as an edge difference distance. The edge difference distance may be obtained as the Frobenius norm of the difference between a pair of matrices, the pair of matrices including the values of the edges in the respective individual network for which the similarity is obtained.

[0015] The method may further include generating a report of a diagnosis or prognosis of the patient's disease. The method may further include generating a machine learning model configured to predict a diagnosis or prognosis of the patient's disease.

[0016] The biological data for each of the plurality of individuals may include values of a plurality of biological factors, the plurality of biological factors including multiple sets of factors obtained using respective data modalities, wherein the biological data includes biological data obtained using multiple data modalities. The biological data for each of the plurality of individuals may include values of a plurality of biological factors derived from at least one of transcriptomics, proteomics, metabolomics, microbiome, clinical, medical imaging, demographic, or histopathology data. Optionally, the biological data for each of the plurality of individuals includes values of a plurality of biological factors derived from transcriptomics or proteomics data and values of a plurality of biological factors obtained from histopathology data.

[0017] Obtaining one or more individual networks for each of the plurality of individuals may include: using the values of a plurality of biological factors, obtaining at least one individual network for each of the plurality of individuals, the plurality of biological factors including biological factors obtained using at least two different data modalities.

[0018] The one or more similarity metrics between one or more individual networks include one or more similarity metrics derived from the individual networks, the individual networks being obtained from data including values of biological factors obtained using at least two different data modalities. Obtaining one or more individual networks for each of the plurality of individuals may include obtaining a plurality of individual networks for each of the plurality of individuals, each individual network being obtained using the values of a corresponding plurality of biological factors. Optionally, each individual network is obtained using the values of a corresponding plurality of biological factors, the corresponding plurality of biological factors being obtained using the same data modality, and the plurality of individual networks include individual networks obtained using at least two different data modalities.

[0019] The one or more similarity metrics between one or more individual networks may include: a first set of one or more similarity metrics derived from individual networks obtained from data including values of biological factors obtained using a first set of data modalities; and a second set of one or more similarity metrics derived from individual networks obtained from data including values of biological factors obtained using a second set of data modalities, wherein the first set is different from the second set. The one or more similarity metrics between one or more individual networks may include one or more similarity metrics obtained by combining multiple similarity metrics for a pair of individuals, each similarity metric being derived from a pair of individual networks of the corresponding individuals obtained from data including values of biological factors obtained using a different set of one or more data modalities.

[0020] A machine learning model can include multiple machine learning models, each machine learning model being configured to predict a diagnosis or prognosis of a disease of a patient, wherein each machine learning model has been trained to take as input values of a corresponding subset of one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output, wherein the corresponding subset of similarity metrics is sourced from individual networks generated from biological factor values obtained using corresponding data modalities, and wherein providing a diagnosis or prognosis for a patient includes combining the outputs of multiple machine learning models.

[0021] The machine learning model includes a classification or regression model. The machine learning model can include a support vector machine model.

[0022] Providing a diagnosis or prognosis for a patient can include combining predictions of disease subtype or severity. Such a disease can be cancer. Providing a diagnosis or prognosis for a patient can include predicting the Gleason score of a patient diagnosed with prostate cancer, classifying a patient diagnosed with brain cancer between a first class corresponding to low-grade glioma (lgg) of the brain and a second class corresponding to glioblastoma multiforme (gbm), or classifying a patient diagnosed with lung cancer between a first class corresponding to lung adenocarcinoma (luad) and a second class corresponding to lung squamous cell carcinoma (lusc).

[0023] Biological factors can include gene or protein expression levels and optionally histopathological data. In an embodiment, the disease is prostate cancer and the biological factor includes the expression level of MAP7. In an embodiment, the disease is brain cancer and the biological factor includes the expression level of GTP2 and / or HIPK2. In an embodiment, the disease is lung cancer and the biological factor includes the expression level of TGM2 and / or DUSP4.

[0024] Biological factors can include latent variables of a trained machine learning model applied to image data, optionally wherein the image data is histopathological data. The trained machine learning model can be a machine learning model that has been trained in a supervised manner to take histopathological data as input and provide a disease type label as output, optionally a neural network.

[0025] At least one of the one or more individual networks, and optionally all of the one or more individual networks, may include nodes selected using a feature selection process and / or edges selected using a feature selection process. Generating one or more individual networks for each of the plurality of individuals may include applying a feature selection process to a plurality of nodes, each of the plurality of nodes indicating a biological factor in the biological data of an individual, and / or applying a feature selection process to a plurality of edges, the edges indicating a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the respective individual.

[0026] Generating one or more individual networks for each of the plurality of individuals may include selecting a plurality of nodes to include in each respective individual network, each of the plurality of nodes indicating a biological factor in the biological data of an individual, wherein the selection is performed separately for each individual or uniformly for the plurality of individuals. Selecting the plurality of nodes may include selecting a plurality of biological factors that are different between an individual and a reference set of individuals, or selecting a plurality of nodes that have variability satisfying one or more predetermined criteria across the plurality of individuals.

[0027] Generating one or more individual networks for each of the plurality of individuals may include selecting a plurality of edges, the edges indicating a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the respective individual, wherein the selection is performed separately for each individual or uniformly for the plurality of individuals. Selecting the plurality of edges may include selecting a plurality of edges that are different between an individual and a reference set of individuals, or selecting a plurality of edges that are different between a plurality of subsets of the plurality of individuals.

[0028] Generating one or more individual networks for each of the plurality of individuals may include selecting a plurality of nodes to include in each respective individual network, each of the plurality of nodes indicating a biological factor in the biological data of an individual, wherein selecting the plurality of nodes includes selecting a plurality of nodes that have variability satisfying one or more predetermined criteria across the plurality of individuals.

[0029] Generating one or more individual networks for each of a plurality of individuals may include selecting a plurality of edges that indicate a relationship between a pair of biological factors in a corresponding individual that correspond to the nodes connected by the edge, wherein selecting the plurality of edges includes selecting a plurality of edges associated with a difference between: (a) a first edge value obtained for a pair of nodes in a first subset of the plurality of individuals, and (b) a second edge value obtained for the same pair of nodes in a second subset of the plurality of individuals, the difference satisfying a predetermined criterion. The first edge value may be a correlation between pairs of nodes across the first subset of the plurality of individuals, while the second edge value may be a correlation between pairs of nodes across the second subset of the plurality of individuals. The predetermined criterion may be that the difference is within a predetermined threshold or alternatively among the top x differences among all possible edges between nodes in the individual network after node selection, where x is a predetermined value, such as 3, 5, 10.

[0030] Unless the context otherwise indicates, the methods described herein are computer-implemented. In fact, the size of the biological data (e.g., omics data) typically used for the purposes of this method, as well as the size of the networks to be compared (which typically include hundreds of nodes and hundreds to thousands of edges), means that the complexity of the process of determining the similarity between networks and training machine learning models to classify subjects based on these similarities far exceeds the capabilities of mental analysis.

[0031] Thus, according to a first aspect, a method of diagnosing a disease in a patient is also described, the method comprising:

[0032] Obtaining biological factors of a plurality of individuals, at least some of the biological factors being related to the disease;

[0033] For each of the plurality of individuals, generating an individual graph of nodes and edges between the nodes, each node being related to one of the biological factors, and wherein the edges between the nodes are related to the relationships between the biological factors of the corresponding individual;

[0034] Calculating one or more similarity matrices representing the similarity between the individual graphs;

[0035] Generating a machine learning model for predicting a diagnosis or prognosis of a disease in a patient, the machine learning model being trained using the one or more similarity matrices and biological factors;

[0036] Predicting a diagnosis or prognosis of the disease in the patient based on the machine learning model and biological factors obtained from the patient; and

[0037] Generating a report of the diagnosis or prognosis of the disease in the patient.

[0038] The one or more similarity matrices can be generated based on the similarity between the individual graphs and the similarity between biological factors independent of the graphs. The one or more similarity matrices can be based on at least one of Spearman calculation and node product or lioness calculation. The biological factors on which the similarity matrix is based can include at least one of gene or protein expression and histopathological readings, and predicting the diagnosis or prognosis of the disease can include predicting cancer diagnosis or prognosis. Predicting the cancer diagnosis or prognosis can include determining at least one of the type or severity of cancer. Determining at least one of the type or severity of cancer can include calculating the Gleason score. Determining at least one of the type or severity of cancer can include differentiating low-grade glioma (lgg) and glioblastoma multiforme (gbm). Determining at least one of the type or severity of cancer can include differentiating lung adenocarcinoma (luad) and lung squamous cell carcinoma (lusc). Cancer diagnosis or prognosis can include prostate cancer, and at least one gene or protein expression can include MAP7. Cancer diagnosis or prognosis can include brain cancer, and at least one gene or protein expression can include GTP2 or HIPK2. Cancer diagnosis or prognosis can include lung cancer, and at least one gene or protein expression can include TGM2 or DUSP4.

[0039] According to a second aspect, there is provided a computer-implemented method for obtaining a tool for characterizing a patient's disease, the method comprising: obtaining, for each of a plurality of training individuals, biological data and a diagnosis or prognosis label associated with the individual, the biological data comprising values of a plurality of biological factors of the individual; generating, for each of the plurality of individuals, one or more individual networks, each individual network comprising a plurality of nodes and edges between the nodes, wherein each node indicates a biological factor in the biological data of the individual, and each edge indicates a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the respective individual; determining values of one or more similarity metrics between the one or more individual networks generated for the patient and the one or more individual networks generated for other individuals in the plurality of individuals; and generating a machine learning model configured to predict the diagnosis or prognosis of the patient's disease, wherein the machine learning model takes as input the values of the one or more similarity metrics between the individual networks and produces a diagnosis or prognosis as output. The method according to this aspect can have any of the features described in relation to the first aspect. In particular, the method can include any of the steps described herein in relation to the method of characterizing a patient's disease, such as a feature selection step, a step of obtaining biological factors from imaging data, a step of obtaining similarity metrics, a step of obtaining individual networks, and the like.

[0040] According to a third aspect, there is provided a computer-implemented method for providing treatment recommendations for a patient suffering from a disease, the method comprising: characterizing the patient's disease using the method of any embodiment of the first aspect, and selecting the patient to be treated with a treatment associated with the predicted diagnosis or prognosis. The method may further comprise treating the patient with the selected treatment.

[0041] According to any aspect, obtaining biological data including values ​​of a plurality of biological factors may include receiving data from a database, a computer readable memory, or a user interface. According to any aspect, obtaining biological data including values ​​of a plurality of biological factors may include measuring values ​​of one or more biological factors in a sample previously obtained from an individual.

[0042] According to a fourth aspect, a computer-implemented method for performing quality control on biological data about a patient suffering from a disease is provided, the method comprising: using biological data about the patient to characterize the patient's disease using the method described in any embodiment of the first aspect, the biological data comprising values ​​of multiple biological factor subsets obtained using corresponding different data modalities; using biological data about the patient comprising only values ​​of a first biological factor subset to characterize the patient's disease using the method described in any embodiment of the first aspect; and comparing a predicted diagnosis or prognosis obtained using the multiple biological factor subsets and the first biological factor subset, wherein a difference in the predicted diagnosis or prognosis between the first biological factor subset and the multiple biological factor subsets indicates a poor quality of the biological data including the first biological factor subset.

[0043] According to a further aspect, there is provided a system comprising: a processor; and a computer-readable medium comprising instructions, which when executed by the processor causes the processor to perform the (computer-implemented) steps of the method of any of the preceding aspects. According to a further aspect, there is provided one or more non-transitory computer-readable media or media comprising instructions, which when executed by at least one processor causes the at least one processor to perform the method of any embodiment of any aspect described herein. According to a further aspect, there is provided a computer program comprising code, which when executed on a computer causes the computer to perform the method of any embodiment of any aspect described herein.

[0044] Embodiments of the invention will now be described by way of example and not limitation with reference to the accompanying drawings. However, various further aspects and embodiments of the invention will be apparent to those skilled in the art in view of this disclosure.

[0045] The present invention includes combinations of the aspects and preferred features described, unless such combinations are clearly impermissible or are stated to be specifically avoided. These and further aspects and embodiments of the present invention will be described in further detail below with reference to the accompanying examples and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart showing a method of characterizing a disease of a subject as described herein.

[0047] Figure 2 shows an embodiment of a system for characterizing a disease of a subject as described herein.

[0048] Figure 3 Schematically shows the data integration and prediction workflow used in embodiments of the present disclosure.

[0049] Figure 4 Shows the evaluation results (macro F1 score (%)) of the classification performance obtained by classifying using PPNs constructed using different multimodal methods from IN. A. Classification of lung cancer severity. B. Classification of brain cancer types. C. Classification of lung cancer types. The greener, the better the prediction, while the redder, the worse the prediction. Untested methods are left blank. The first three rows refer to methods based on the nodes of an individual network. Rows 4 and 5 use the edge weights of the individual network. Rows 6 and 7 combine the individual nodes and edges. The first two columns focus on a single data modality. Columns 3 to 6 involve data integration.

[0050] Figure 5 Shows a comparison of the classification results of a graph-based model described herein with various classification algorithms applied to the original features. The models are ranked according to their prediction performance. The smaller the area enclosed by the colored lines, the better. (a) Shows the average ranking of each model in the datasets (prostate, brain, lung) and data types (RNAseq or histopathology). (b) Shows the average ranking of each model on the datasets for the combined (i.e., concatenated) data types. For each analysis, the best graph method is presented. Since the method based on the nodes of IN does not use any graph structure information, only the method based on the edges of IN and the method based on the nodes and edges of IN are considered graph methods. (a) Shows that for a single data type, in six analyses conducted, the graph-based method is superior to other models in four cases (i.e., 2 / 3 of the cases). (b) Shows that for the combined data types, the graph-based method performs best in the lung use case and the brain use case.

[0051] Figure 6Shows the results of LIMMA analysis on the 50 edges with the largest co-expression differences between groups (see Example 4). Genes with an absolute t-statistic < 1.5 are shown in white. In the prostate use case (a), if an edge / node has a higher coefficient in the Gleason pattern 3 group, it is red (pattern 4 is blue). This shows that the most relevant gene includes MAP7, which can predict the survival of stage II colon cancer patients. In the brain use case (b), if an edge / node has a higher coefficient in low-grade glioma of the brain, it is red (blue in glioblastoma multiforme). In the lung use case (c), if an edge / node has a higher coefficient in lung adenocarcinoma, it is red (blue in lung squamous cell carcinoma). The thicker the edge, the higher the log fold change.

[0052] Figure 7 Shows the principles of the person-to-person network (PPN) and the individual network (IN). (a) Multimodal fusion of the person-to-person network. The nodes are individuals, and the edges show the proximity of two individuals. (b) The individual network. The nodes and / or edges are individual-specific.

[0053] Figure 8 Shows the evaluation results of the classification performance (macro F1 score (%)) obtained by using the PPN from the IN for classification, using multiple data transformations to calculate the similarity between graphs (see Example 2). Top: IN from RNAseq data. Bottom: IN from histopathology data. No graph information is used when inferring the similarity matrix using Spearman correlation, Euclidean distance, or Gaussian kernel. Graph information is only studied when calculating the similarity from individual graphs constructed using the node product or LIONESS algorithm. When combining the similarity matrices obtained using Spearman correlation and node product and using Spearman correlation and Lioness, both the original data and graph information are analyzed.

[0054] Figure 9 shows the evaluation results of the classification performance (macro F1 score (%)) obtained by using the PPN from the IN for classification, where graphs of different data modalities from different levels are fused (see Example 3). A: Graph constructed using the node product method. B: Graph constructed using the Lioness method. For RNAseq and histopathology values, only one data type was used. The combinations of data types are as follows: at the early stage via the linkage of two databases (early), at the intermediate stage via the average of the similarity matrix (average) or the SNF program (SNF), and at the late stage via majority voting (late). Note that late integration is only performed when using a combination of the original data and graph data in order to apply majority voting to more than 2 outcomes.

[0055] Figure 10 Shows the results of gene set enrichment analysis of IN (see Example 4). The figure shows the top 10 enriched gene sets from the largest component obtained from the LIMMA analysis, where the feature selection for the prostate cancer subject set (upper figure) and the lung cancer subject set (lower figure) is as described in Example 1. No enriched gene sets were detected in the brain cancer use case (see Example 1). The size of the pathway represents the number of genes in the pathway after removing genes not present in the largest component. For prostate cancer, the most significantly enriched set is the Chandran metastasis pathway. Metastasis is the most adverse outcome of cancer. Detailed implementation

[0056] In describing the present invention, the following terms will be used and are intended to be defined as indicated below.

[0057] As used herein, "and / or" is considered a specific disclosure of each of two designated features or components with or without the other. For example, "A and / or B" will be considered to specifically disclose (i) A, (ii) B, and (iii) each of A and B, as if each were set forth separately herein.

[0058] Personalized screening before treatment paves the way for improving the diagnostic accuracy and treatment outcomes of various diseases including cancer. However, most methods are limited to single - type data and do not consider the interactions between features, ignoring the complementary insights that multimodality and systems biology can provide. The inventors demonstrated the use of graph theory, and in particular subject - specific networks, for this purpose. Networks are powerful tools that consider the interactions between feature pairs and can thus better utilize all the information available in the dataset compared to only considering the values of the features individually.

[0059] The present disclosure relates at least in part to methods of utilizing a subject-specific network (also referred to as an "individual network", IN). An IN is a network in which the values and / or presence of nodes and / or edges are specific to an individual. A "network" or "graph" is a data structure G that includes a set of nodes V and a set of edges E between the nodes (G = (V, E)). In the context of the present disclosure, the edges in a graph can be described by an adjacency matrix A, where the coefficient A(i, j) indicates the presence (e.g., A(i, j) is 0 or 1 when there is no edge or when there is an edge, respectively) or weight of the relationship (edge) between node i and node j. When an edge is associated with a weight, the network is referred to as a weighted network. In the context of the present disclosure, a network is generally a weighted network where the edge weight indicates the relationship between the two nodes it connects. The relationship between nodes (i.e., the edge) is quantified using any metric known in the art for quantifying the relationship between variables, such as, for example, the node product between the values of two variables corresponding to the nodes connected by the edge (the product of the (optionally normalized) values of the two nodes connected by the edge, where the value of a node is generally equal to the value of the variable corresponding to the node connected by the edge), correlation, mutual information, or other or covariation metrics (such as, for example, the context-dependent likelihood described in Akhand et al. 2015), or path / message-related metrics (such as, for example, the weights obtained using the PANDA (Passing Attributes between Networks for Data Assimilation) algorithm described in Glass et al. 2013), and metrics derived from any of the above (such as the increase or decrease in the value of these metrics when an individual is added to or removed from a group of subjects from which an aggregated network is obtained). In an individual network, all node and edge values are associated with a specific individual. However, depending on the way the edges are obtained, they can already be determined using data about a group of individuals (such as, for example, when an edge represents an increase or decrease in an edge weight obtained when comparing networks obtained for a group that includes an individual with the same group without that individual), or using individual data alone (e.g., when an edge is determined using the node product method). Figure 7(b) schematically shows an individual network. For each individual, the nodes are variables (also called "biological factors") (such as genes), and the edges show the links (also called "relationships") between these variables for that individual. Most existing graph analysis methods for complex diseases aggregate information across the entire group but cannot detect individual characteristics. The inventors hypothesized that using individual-specific interactions rather than group-level systems would help capture the heterogeneity between individuals and enhance the identification of new biomarkers for precision medicine. This hypothesis emphasizes the choice of using individual networks (INs). Since INs represent individual relationships between variables, the inventors hypothesized that they could be easily used in precision medicine. Individual networks can be inferred via multiple methods. For example, the variable values of an individual (e.g., gene expression) can be overlaid onto a reference network obtained from external knowledge (e.g., protein interactions), as described in Menche et al. 2017. By such a method, only the node values differ between individuals, while the graph topology does not. Another option is linear interpolation for obtaining network estimates for a single sample (LIONESS, described in Kuijier et al. 2019). LIONESS calculates edge weights by using the difference in edge weights between a network constructed using all samples and a network reconstructed using all samples except the sample of interest. Another option is the single-sample network based on the Pearson correlation (ssPCC) algorithm described in Liu et al. 2016. These individual networks result from the perturbation of Pearson correlations caused by adding a specific individual to a given set of samples. Both LIONESS and ssPCC use a reference set / group of samples. Alternatively, edge weights can be calculated without a reference set / group by adding the Z scores of the log-transformed values of two associated nodes, as described in Koh et al. 2019, or by making repeated measurements of each variable for each individual (and, for example, calculating the correlation between variable pairs over the repeated measurements).

[0060] This disclosure is in part directed to methods of combining supervised data integration and individual networks. In particular, the methods provided by this disclosure include: obtaining individual networks of a plurality of individuals including disease subjects and a reference subject set, calculating the similarity between the obtained plurality of individual networks (i.e., determining the value of a similarity metric), and using a machine learning model to predict a diagnosis or prognosis of a disease subject, where the machine learning model has been trained to use the similarity between the plurality of individual networks of the reference subject set as an input to predict the diagnosis or prognosis. To the inventors' knowledge, such methods have never been attempted or evaluated. The inventors of the present invention have shown that such methods can predict disease subtypes and severity from patient data by using data from two or more modalities and combining these modalities at various stages of the method. In fact, individual networks can be obtained together for multiple modalities (i.e., the IN includes nodes associated with different modalities), or individual networks can be obtained separately for each modality (single-modal IN). The former results in a multi-modal IN and can be referred to as "early fusion". When using single-modal INs, these single-modal INs can be combined by calculating the similarity between multiple INs of the same modality (single-modal similarity), and then the single-modal similarities can be combined into a multi-modal similarity. This can be referred to as "intermediate fusion". This is illustrated in Figure 7 (a), where interpersonal networks are used to represent the similarity between individuals for each modality, and these similarities can be mapped to each other at the individual level and combined across multiple modalities. Alternatively, single-modal INs can be used to calculate single-modal similarities, which in turn can be used to obtain single-modal predictions using corresponding machine learning models trained on similarity inputs from a single modality. These single-modal predictions can then be combined into a multi-modal prediction (late fusion). Note that all combinations of the above methods are also possible and conceivable for various subsets of multiple data modalities. For example, the method can include obtaining one or more single-modal INs for one or more corresponding modalities, obtaining one or more multi-modal INs, each multi-modal IN combining data from multiple modalities (early fusion), obtaining corresponding single-modal and multi-modal similarities, combining at least some of the obtained similarities (e.g., one or more subsets of single-modal similarities) into a multi-modal similarity (intermediate fusion), obtaining predictions for each of the resulting similarities, and combining these predictions (if the similarities include multiple similarities, i.e., not all similarities were combined in the previous step; late fusion).

[0061] The term "data modality" refers to data that has been obtained (i.e., recorded or measured) from a subject, and the data is from a specific source (also referred to as "type"). For example, a data modality can be gene expression data (transcriptomic data, data on the presence and / or level of one or more transcripts) obtained from one or more samples of a subject, protein expression data (proteomic data, data on the presence and / or level of one or more proteins) obtained from one or more samples of a subject, metabolomic data (data on the concentration of one or more metabolites) obtained from one or more samples of a subject, genomic data (data on the presence and / or characteristics of one or more genomic features, such as somatic mutations (any type, including single-base substitutions, insertions / deletions (indels), and rearrangements), polymorphisms, epigenetic marks, copy number variations, etc.) obtained from one or more samples of a subject, demographic data about the subject (e.g., data about the subject's age, gender, race, etc.), histopathological data obtained from one or more samples of a subject, medical imaging data obtained from the subject (e.g., MRI, endoscopy, X-ray, ultrasound, CT scan, etc.), clinical data about the subject (e.g., comorbidities, treatment history, exposure factors, etc.), microbiome data (e.g., the presence and / or abundance of one or more microbial taxa in one or more samples from the subject) obtained from one or more samples of a subject. Each such data can be obtained from a sample, can be recorded in one or more databases, or may have been previously obtained from a sample and recorded in one or more databases from which the data can be retrieved for the purpose of performing the methods described herein. For each of one or more data modalities, a data set includes biological factors that are values of multiple biological characteristics, or data from which these values can be obtained. An individual data set can be obtained separately for an individual sample, or data from multiple samples from the same subject can be combined. Individual data sets are typically obtained using the same measurement technique. For example, a single gene expression data set can include gene expression data (e.g., RNA sequencing data) obtained from one or more samples of a subject. A single histopathological data set can include whole-slide images or portions thereof obtained from one or more samples of a subject. These images can all be from the same sample (e.g., the same tissue section) and can correspond to different regions of the sample or different measurement channels (e.g., corresponding to different markers). Biological factors can be values of one or more characteristics measured for the subject or from one or more samples from the subject. These values may have been normalized, standardized, filtered, etc. one or more times before use. For example, in the context of transcriptomics, a characteristic can be the expression of individual genes, and a biological factor can be the expression level of individual genes.The same principle can be applied to any data modality for measuring the presence or level of any biomolecule or entity (transcript, metabolite, protein, microbial taxon, mutation, etc.). For clinical data, the features can be exposure to one or more risk factors, the presence or number of comorbidities, the presence of a specific treatment history, etc., and the biological factors can be values indicating the presence or degree of presence of one or more risk factors, the number or presence of comorbidities, the presence or absence of a specific treatment history, etc. The biological factors can be values of one or more features derived from the feature values measured on a subject or from one or more samples from the subject. For example, the biological factors can be values of one or more features derived from the measured values by applying one or more data reduction methods (e.g., values of one or more principal components obtained from a set of measured values by applying principal component analysis) and / or one or more machine learning models, where the biological factors are latent variables of such machine learning models. Thus, the biological factors can be values of a latent variable of a machine learning model that takes as input a set of measured values on a subject or from one or more samples from the subject. The set of measured values can be, for example, pixel values of one or more images (e.g., medical images or histopathological images). In such a case, the machine learning model can be any machine learning model that has been trained in image processing. The machine learning model can have been trained at least in part (e.g., trained from scratch or fine-tuned) using training data that includes images of the same type as the images to be analyzed (e.g., histopathological images when histopathological images are used to generate biological features). The machine learning model can be a model that has been trained for a general image object recognition task. The machine learning model can be a model that has been trained (either fully or by fine-tuning a pre-trained model) for a specific classification or regression task associated with the image type of the images to be analyzed. For example, in the context of histopathological images of tumor samples, the machine learning model can be a machine learning model that has been trained to classify histopathological images among multiple cancer types or subtypes. The task can be different from the prediction task that is the ultimate goal of the methods described herein (i.e., the task can be different from the diagnosis or prognosis prediction performed as part of the methods described herein). In the context of the present disclosure text, edges are generally undirected.

[0062] Gene expression data refers to data on the expression levels of gene expression products. Within the context of the present disclosure, the expression levels of genes of interest are preferably determined at the nucleic acid level and, in particular, at the mRNA level. Thus, a reference to "gene expression level" may refer to a transcriptomic expression metric. Multiple gene expression levels may be referred to as an expression profile or collectively as a gene expression data set. "Gene expression data" means a data set related to the expression levels of multiple genes in an individual. Determination of gene expression levels may involve determining the amount of mRNA of a specific gene or a group of genes in a sample. Methods for achieving this are well known to those skilled in the art. Any conventional method may be used to determine gene expression levels in a sample, such as using nucleic acid microarrays, using nucleic acid synthesis methods (such as quantitative PCR (qPCR), also known as qRT-PCR), using molecular counting assays, or using RNA sequencing (including bulk RNA sequencing and single-cell RNA sequencing). For example, the NanoString nCounter analysis system may be used to determine gene expression levels (see, e.g., US 7,473,767). When using single-cell RNAseq, pseudo-bulk RNA expression data for the whole or one or more parts of the sample may be obtained (such as, for example, one or more combined RNA expression levels may be obtained based on the expression levels of multiple respective cells in respective specific cell populations). For example, a gene expression data set obtained from single-cell RNA sequencing may include a first set of gene expression levels of cells in a first population (e.g., normal cells) and a second set of gene expression levels of cells in a second population (e.g., cancer cells). According to the methods of the present disclosure, the expression levels of each gene represented in the first set and each gene represented in the second set (optionally after feature selection) may be used as the values of the nodes. The same concept may be extended to any number of cell populations.

[0063] The present disclosure relates to a method comprising the step of determining the similarity between individual networks. The terms "similarity" and "distance" are used interchangeably to refer to metrics that quantify the degree to which two networks are similar or dissimilar to each other. These include similarity metrics (which quantify the degree to which two networks are similar to each other) and distance metrics (which quantify the degree to which two networks are dissimilar to each other). A similarity metric can be converted to a distance metric using a function such as s(x,y) = u / (1 + d(x,y)), and vice versa, where s is the similarity metric, d is the distance metric, x and y are the two networks to be compared, and u is an upper bound, or for kernel-based similarity measures, d(x,y) = s(x,x) + s(y,y) - 2*s(x,y). The distance / similarity metrics for network edges and network nodes can be calculated separately. A distance / similarity metric calculated based only on the nodes of two networks does not reflect the similarity between the networks because it does not take into account the edges. However, a distance / similarity metric calculated based only on the edges of a pair of networks does reflect the similarity between the networks. In addition, a distance / similarity metric calculated based on the edges of a pair of networks and the nodes of the pair of networks does reflect the similarity between the networks. Such a distance / similarity metric can be obtained by combining (e.g., summing or averaging) the distance / similarity metric obtained based on the edges of the pair of networks and the distance / similarity metric obtained based on the nodes of the pair of networks.

[0064] The edge distance / similarity metric can be selected from: Euclidean distance, Jaccard distance, edge difference distance, DeltaCon, spectral distance, motif-based metric, Hamming distance, shortest path kernel, k-step random walk kernel, graph diffusion distance, and feature portrait divergence. Euclidean distance, Jaccard distance, edge difference distance, and DeltaCon distance are applicable when there is a node correspondence between the two networks being compared. This is the case when all nodes are labeled and there are the same nodes in the two networks being compared. In an embodiment, the edge difference distance metric is used to calculate the similarity between INs. This distance takes two adjacency matrices (matrix A includes coefficients A(I,i), equal to the weight of the edge between each pair of nodes i,j) and calculates the Frobenius norm of their difference. This is found to achieve a good balance between information richness (leading to good prediction performance) and computational efficiency. Spectral distance, motif-based metric, feature portrait divergence, Hamming distance, shortest path kernel, k-step random walk kernel, and graph diffusion distance are applicable even when the correspondence between nodes is not known, but compare the global structure of the networks and are thus not suitable for comparing fully connected networks with node correspondence.

[0065] The node distance / similarity metric can be selected from: Euclidean distance, affinity matrix, Gaussian kernel, cosine similarity, and Spearman correlation. The Spearman correlation coefficient has been found to achieve a good balance between information richness (resulting in good prediction performance) and computational efficiency. When classifying subjects based only on nodes, using the Spearman correlation coefficient leads to the highest classification accuracy on the test data. The affinity matrix distance can be determined by a distance matrix such as, for example, the Euclidean distance between each pair of nodes (the corresponding nodes in the two individuals to be compared). The Gaussian kernel distance can be calculated as where v x and v y are the node value vectors of individual x and individual y, ||v x -v y || is the Euclidean distance between these vectors, and σ 2 is a parameter corresponding to the bandwidth of the kernel. This value can be set empirically. For example, σ = 1000 has been found to be appropriate. The cosine similarity between vectors v x and v y is calculated as where n is the magnitude of vectors v x and v y .

[0066] The network distance / similarity metric can be selected from: edge distance / similarity metrics, and combinations of edge distance / similarity metrics and node distance / similarity metrics. Multiple similarity metrics can be combined, for example, by summing or averaging.

[0067] The method of the present disclosure may include one or more feature selection steps applied to the IN. Feature selection refers to the process of selecting nodes and / or edges of the IN for further analysis. Feature selection can be data-driven and / or based on prior knowledge. For example, nodes (i.e., variables) that are known to be more likely to be informative for a particular purpose can be included in the IN to be analyzed (or used for training a machine learning model), while nodes (i.e., variables) that are known to be less likely to be informative can be excluded. The same principle applies to edges. For omics data, edges can be selected, for example, based on pathway information from one or more databases, such as by including edges between genes or proteins known to interact with each other and excluding edges between genes or proteins known not to interact with each other. Data-driven feature selection refers to the process of selecting edges and / or nodes purely based on the data available for an individual or a group of individuals. Feature selection based on a group of individuals results in the same set of nodes and edges for all individuals. Feature selection based on a single individual can result in a different set of nodes and / or edges being selected for each individual in the group. Examples of feature selection based on a single individual include selecting nodes that are significantly different in an individual compared to a reference or control group (e.g., a group of individuals considered to be normal). Examples of feature selection based on a group of individuals include selecting nodes that have variability across the group that meets one or more predefined criteria. For example, the top x nodes with the highest standard deviation across individuals in the group can be selected. The value of x can be selected, for example, based on computational requirements (since a lower value will reduce the computational load required to implement the method) and / or based on the predictive performance evaluated on a test dataset. For example, x can be a value such that when used for selecting nodes to include in the IN for the purposes of the present method, that value results in predictions with the highest accuracy (e.g., highest macro F1 score) in a group of test samples. Alternatively, nodes with a standard deviation higher than a threshold y can be selected, where the threshold y can be selected as described above for the value x. Alternatively or in addition, feature selection based on a group of individuals can include obtaining a first network of a first subset of the group and a second network of a second subset of the group, and selecting the top x edges with the greatest difference between the two networks, or edges with a difference between the two networks higher than a threshold y and / or edges with a statistically significant difference between the two networks. The values of the thresholds x and y can be selected as described above. The first and second networks can be obtained by calculating the correlation coefficient (e.g., Pearson correlation coefficient) between each pair of nodes across individuals in the first and second subsets, respectively. Any method known in the art can be used to obtain a network aggregating data from a group of subjects. The individuals in the first and second subsets can have different known diagnoses or prognoses. These can match the diagnosis or prognosis labels that the method is intended to predict. In such cases, feature selection can be performed on training data including subjects with known diagnoses / prognoses using a group of individuals.

[0068] The methods of the present disclosure include using a machine learning model to obtain a prediction for a subject. The machine learning model is trained using supervised learning. This uses training data (also referred to as "reference data") that includes the biological characteristics of multiple individuals and ground truth labels indicative of the diagnosis or prognosis to be predicted. The machine learning model can be a classification model or a regression model. The machine learning model can be any machine learning model that can take as input a similarity vector between an individual and a set of reference individuals and produce an output indicative of the diagnosis or prognosis to be predicted. This can be in the form of a class label, the probability of belonging to one or more classes, or a predicted estimate of the diagnosis or prognosis value to be predicted. The machine learning model can be a support vector machine (SVM), a naive Bayes classifier, a k-nearest neighbor classifier, a classification or regression tree, or a neural network. A classification model can be particularly suitable for predicting disease subtypes, or discrete severity scores or categories corresponding to a severity score range. A regression model can be applicable for predicting continuous values, such as a continuous severity score, a viability metric (e.g., OS, DFS, PFS, as explained below), a survival probability, etc. The machine learning model can be an SVM. The SVM is applicable to both classification tasks and regression tasks. The SVM algorithm identifies a hyperplane in an N-dimensional space (where N is the number of features associated with each instance to be classified) that can optimally classify the instances between predefined classes. This is performed by identifying the most similar instances from different classes to each other. Kernel-based SVM operates directly on a transformation of the set of original feature vectors, which represents the similarity between the instances represented by these original features. Many machine learning algorithms can be represented by the dot product between the vectors to be compared, and any such machine learning algorithm can be used in the present method. This includes, for example, SVM, logistic regression, perceptron, etc.

[0069] The present disclosure relates in part to methods for providing a disease diagnosis or prognosis for a subject. The diagnosis can be the identification of a disease subtype. Disease subtyping refers to the identification of homogeneous groups of patients, i.e., patients who share molecular characteristics, histological characteristics, and / or clinical characteristics. For example, many cancers include multiple subtypes that are associated with different etiologies (e.g., tissue of origin), different histological features (e.g., lung adenocarcinoma vs. lung squamous cell carcinoma), different molecular features (e.g., driver mutations, gene expression patterns, etc.), and / or different phenotypic features (e.g., hormone-dependent cancers vs. non-hormone-dependent cancers).

[0070] Prognosis can be the identification of the severity of a disease (e.g., grade, stage, or severity score) or the likely outcome (e.g., predicting whether the prognosis of a subject is good or bad / poor, belonging to the group of subjects with good prognosis or the group of subjects with poor prognosis). The severity of a disease can be evaluated using a disease severity score, grade, or stage. These are typically disease-specific and are evaluated using multiple criteria. For example, the Gleason score is used to evaluate the severity of prostate cancer and is evaluated based on histopathological data. It is obtained by examining the appearance of cancerous cells in a biopsy: biopsies that include cells similar in appearance to normal prostate tissue are designated as grade 1, biopsies that include mostly cells similar in appearance to normal cells are designated as grade 2, and biopsies that include tumor cells are designated as one of grades 3 to 5 (depending on the degree of abnormality in the cell appearance). The score is associated with the prognosis because the likely growth of the cancer is related to the score (grade 1 cancer may grow very slowly, and grade 5 cancer may grow very rapidly).

[0071] Whether the prognosis is considered good or bad can vary depending on the disease context (e.g., cancer type, stage of the disease, etc.). Generally speaking, a good prognosis means that the overall survival (OS), disease-free survival (DFS), and / or progression-free survival (PFS) are longer compared to a control group or value, such as for example the average of that stage and cancer type, or the average of a control group of subjects (e.g., a group of separately aggregated subjects). If the OS, DFS, and / or PFS are lower than the control group or value (such as for example the average of that stage and cancer type, or the average of a control group of cancers), the prognosis can be considered poor. Thus, generally speaking, a "good prognosis" means that the survival (OS, DFS, and / or PFS) and / or disease stage of an individual patient are advantageous compared to the expected survival and / or disease stage of a group of patients in a comparable disease setting. Similarly, a "poor prognosis" means that the survival (OS, DFS, and / or PFS) of an individual patient is lower (or the disease stage is worse) compared to the expected survival of a group of patients in a comparable disease setting.

[0072] As used herein, a "sample" can be a cell or tissue sample (e.g., a biopsy), or an extract from which biological material can be obtained for analysis (such as transcriptomic analysis (whole transcriptome sequencing or targeted (also known as "panel") sequencing), genomic analysis (e.g., genomic sequencing), proteomic analysis, histopathological analysis). For example, the sample can be a tumor sample or a blood sample. In the context of histopathology, the sample can be a tissue sample, such as a tumor sample. In the context of cancer prognosis or diagnosis, the sample can be a tumor sample or a biofluid sample, e.g., including circulating tumor DNA or tumor cells. The sample can be a sample freshly obtained from a subject, or a sample that has been processed and / or stored prior to performing an assay (e.g., frozen, fixed, or subjected to one or more purification, enrichment, or extraction steps). Specifically, the sample can be a cell or tissue culture sample derived from a tumor. Thus, a sample as described herein can refer to any type of sample that includes biological material from which a biological characteristic can be determined. In addition, the sample can be transported and / or stored, and collection can occur at a location remote from the location of biological data collection (e.g., sequencing), and / or computer-implemented method steps can occur at a location remote from the sample collection location and / or at a location remote from the location of biological data collection (e.g., sequencing) (e.g., computer-implemented method steps can be performed by a networked computer, such as through a "cloud" provider). A "tumor sample" refers to a sample that contains tumor cells or genetic material derived from tumor cells. A tumor sample can be a cell or tissue sample (e.g., a biopsy) obtained directly from a tumor.

[0073] As used herein, "treatment" and "therapy" refer to reducing, alleviating, or eliminating one or more symptoms of a disease being treated, relative to the symptoms prior to treatment.

[0074] A subject or individual according to the present disclosure is preferably a mammal (including a human or a model animal, such as a mouse, a rat, etc.), preferably a human. The terms "patient", "subject", and "individual" are used interchangeably. A patient can be a patient who has been diagnosed with or is likely to have a disease. Thus, providing a diagnosis can include confirming a diagnosis of a disease, or providing a diagnosis of a subtype of a disease that the patient has been diagnosed with (including a molecular subtype, a histopathological subtype, a phenotypic subtype, a treatment response group, a severity group, or any other distinction of the patient or disease group, etc.).

[0075] In addition to the structural components and user interactions described herein, the systems and methods described herein can be implemented in a computer system. As used herein, the term "computer system" includes the hardware, software, and data storage devices used to embody the system or execute the method according to the above-described embodiments. For example, a computer system can include a processing unit (such as a central processing unit (CPU) and / or a graphics processing unit (GPU)), an input device, an output device, and a data storage device, which can be embodied as one or more connected computing devices. Preferably, the computer system has a display or includes a computing device having a display for providing a visual output display. The data storage device can include RAM, a disk drive, or other computer-readable media. The computer system can include multiple computing devices that are connected by a network and capable of communicating with each other via the network. It is expressly contemplated that the computer system can be composed of or include cloud computers.

[0076] The methods described herein can be provided as a computer program or as a computer program product or a computer-readable medium carrying the computer program, which is arranged to execute the methods described herein when run on a computer. As used herein, the term "computer-readable medium" includes, but is not limited to, any one or more non-transitory media that can be directly read and accessed by a computer or computer system. The media can include, but are not limited to, magnetic storage media, such as floppy disks, hard disk storage media, and magnetic tapes; optical storage media, such as optical discs or CD-ROMs; electrical storage media, such as memories, including RAM, ROM, and flash memories; and mixtures and combinations of the above storage media, such as magnetic / optical storage media.

[0077] Method for characterizing a subject

[0078] Figure 1is a flow chart that schematically shows a method for characterizing a disease subject according to the present disclosure. At optional step 10, one or more samples are obtained from the subject. At optional step 12, the samples are analyzed to obtain multiple biological factors for each of one or more data modalities. This can include, for example, obtaining gene expression data (i.e., transcriptomic data) from a sample previously obtained from the subject using RNA sequencing. This can include, for example, obtaining histopathological images from a sample previously obtained from the subject. Other data modalities and their combinations are possible and are expressly contemplated to include, for example, demographic data about the subject (such as age, gender, race), clinical data about the subject (such as comorbidities, exposures such as, for example, smoking history), medical imaging data (including histopathology, MRI, X-rays, etc.), microbiome data about the subject (such as the presence and / or amount of one or more microbial populations (such as microbial taxa) in a sample previously obtained from the subject), metabolomic data (such as the amount of one or more metabolites and / or the value of one or more metabolic fluxes in a sample previously obtained from the subject), genomic data (such as the presence of one or more genomic features such as mutations (including single-base substitutions, multi-base substitutions, insertions, deletions, and rearrangements), copy number variations, and / or chromosomal instability), proteomic data (such as the presence or amount of one or more proteins in a sample previously obtained from the subject—including targeted and untargeted assays such as, for example, the measurement of a specific cytokine in a subject sample), physiological data about the subject (such as from a wearable device), etc. The data preferably includes at least one omics modality (such as transcriptomics, proteomics, metabolomics, genomics) and / or one or more imaging modalities (such as histopathological images). The values of the biological factors can be values that have been previously subjected to one or more transformations (such as normalization, standardization, log transformation, etc.). For example, the node values (i.e., the values of the biological factors assigned to the nodes in an individual graph) can be normalized using a min-max normalization algorithm. The normalization can be performed separately for each individual graph. The multiple biological factors can include at least some biological factors related to the disease. The biological factors on which the similarity metric (such as a similarity matrix) is based can include at least one gene or protein expression and histopathological readings.

[0079] Biological factors can include latent variables of a trained machine learning model applied to image data, optionally where the image data is histopathological data. The trained machine learning model can be a machine learning model that has been trained in a supervised manner to take histopathological data as input and provide disease type labels as output, optionally a neural network. The trained machine learning model can be a computer vision model. The trained machine learning model can be a deep neural network, such as ResNet. The machine learning model can have been trained using multiple histopathological images from samples of multiple different cancer types. The machine learning model can have been trained to predict the cancer type of a histopathological image. The multiple different cancer types can include the cancer type of the patient whose prognosis or diagnosis is being predicted. The multiple different cancer types can include at least 10 or at least 20 different cancer types.

[0080] At step 14, using the biological data obtained at step 12, one or more individual networks are generated for each of the multiple individuals, each individual network including a plurality of nodes and edges between the nodes, where each node indicates a biological factor in the biological data of the individual, and each edge indicates a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the corresponding individual. The one or more individual networks can be referred to as individual graphs. They can include or consist of a set of nodes and a set of edges between the nodes. Each node can be associated with one of the biological factors. The edges between the nodes can be associated with the relationships between the biological factors of the corresponding individual.

[0081] Generating one or more individual networks for each of the plurality of individuals may include selecting a plurality of nodes to include in each respective individual network, where each node of the plurality of nodes indicates a biological factor in the biological data of the individual, and where the selection is performed either individually for each individual or uniformly for the plurality of individuals. Optionally, where selecting the plurality of nodes includes selecting a plurality of biological factors that are different between an individual and a reference set of individuals, or selecting a plurality of nodes that have variability meeting one or more predetermined criteria across the plurality of individuals. Generating one or more individual networks for each of the plurality of individuals may include selecting a plurality of edges, where the edges indicate relationships between pairs of biological factors corresponding to the nodes to which the edges are connected in the respective individual, and where the selection is performed either individually for each individual or uniformly for the plurality of individuals. Selecting the plurality of edges may include selecting a plurality of edges that are different between an individual and a reference set of individuals, or selecting a plurality of edges that are different between multiple subsets of the plurality of individuals. Selecting the plurality of edges may include selecting a plurality of edges that are different between a first subset of the plurality of individuals and a second subset of the plurality of individuals. The first and second subsets may be subsets associated with a first and second prognosis or diagnosis to be predicted. Similarly, the plurality of subsets may be subsets associated with a plurality of different prognoses or diagnoses to be predicted (e.g., different cancer types, different groups of cancer severity, etc.). Edges that are different between different subsets of individuals may be edges for which the difference between the networks obtained for the respective subsets is above a predetermined threshold. Different nodes / edges may refer to a difference above a predetermined threshold or to the top x most different nodes / edges, where x is a predetermined value. Selecting a plurality of nodes that have variability meeting one or more predetermined criteria across the plurality of individuals may include selecting nodes having variability above a predetermined threshold (e.g., standard deviation), or selecting the top x most variable nodes, where x is a predetermined value.

[0082] At step 16, determine the values of one or more similarity metrics between one or more individual networks generated for a patient and one or more individual networks generated for other individuals among multiple individuals. Determining the values of one or more similarity metrics may include calculating one or more similarity matrices representing the similarity between individual graphs. For each pair of multiple pairs of individual networks, one or more similarity metrics may include the similarity between edges in the individual networks, the similarity between edges in the individual networks and the similarity between nodes in the individual networks, or a similarity that combines the similarity between edges in the individual networks and the similarity between nodes in the individual networks. The similarity between edges in the individual networks may be referred to as the similarity between individual networks (graphs). The similarity between nodes in the individual networks may be referred to as the similarity between biological factors independent of the graph (individual network). For each pair of multiple pairs of individual networks, one or more similarity metrics may include: the similarity between nodes in the individual networks, where the similarity between nodes in the individual networks is obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the respective individuals; or, a similarity that combines the similarity between edges in the individual networks and the similarity between nodes in the individual networks, where the similarity between nodes in the individual networks is obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the respective individuals. A similarity metric obtained as a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the respective individuals may be calculated as where v x and v y are vectors of node values for individuals x and y, ||v x - v y || is the Euclidean distance between these vectors, and σ 2 is a parameter corresponding to the bandwidth of the kernel.

[0083] One or more similarity metrics between one or more individual networks can include one or more similarity metrics obtained by combining multiple similarity metrics for a pair of individuals, each similarity metric being sourced from a pair of individual networks of corresponding individuals obtained from data including biological factor values obtained using a different set of one or more data modalities. For example, one or more similarity metrics between individual i and individual j can include a similarity metric obtained by combining: (i) a similarity metric between the individual network obtained for individual i and the individual network obtained for individual j using a first data modality (e.g., gene or protein expression); and (ii) a similarity metric between the individual network obtained for individual i and the individual network obtained for individual j using a second data modality (e.g., histopathology). Any number of similarity metrics obtained from INs sourced from any number of data modalities can be used. Summation or averaging can be used to perform the combination of similarity metrics. Such a process can be referred to as intermediate fusion. In an embodiment, for a pair of individuals, all similarity metrics can be obtained by combining multiple similarity metrics, each similarity metric being sourced from a pair of individual networks of corresponding individuals obtained from data including biological factor values obtained using a different set of one or more data modalities.

[0084] At step 18, a machine learning model configured to predict a diagnosis or prognosis of a patient's disease is used to predict the diagnosis or prognosis of the patient, wherein the machine learning model has been trained to take as input the value of one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output. Using the machine learning model to predict the diagnosis or prognosis of the patient can include predicting the diagnosis or prognosis of the patient's disease based on the machine learning model and based on biological factors obtained from the patient. In other words, the trained machine learning model can be used in conjunction with similarity metrics (e.g., a similarity matrix) obtained from the patient's biological factors to predict the diagnosis or prognosis.

[0085] The machine learning model can include multiple machine learning models, each machine learning model being configured to predict a diagnosis or prognosis of the patient's disease, wherein each machine learning model has been trained to take as input the value of a corresponding subset of one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output, wherein the corresponding subset of similarity metrics is sourced from individual networks generated from biological factor values obtained using a corresponding data modality, and wherein providing a diagnosis or prognosis for the patient includes combining the outputs of the multiple machine learning models. Such a method can be referred to as late fusion. Combining the outputs of the multiple machine learning models can be performed by averaging (e.g., when the output is continuous) or by majority voting (e.g., when the output is categorical).

[0086] The result of step 18 can be used to select patients for a specific course of treatment based on any of the prognostic or diagnostic features described above, to select patients for a clinical trial based on the features of a sample from said patient to identify that the patient may respond to a therapy, or to provide a prognosis or diagnosis associated with a predictive feature (e.g., a prognosis associated with predicting a disease subtype). Thus, at step 18, a subject can be classified as having a good prognosis or a poor prognosis. Alternatively or in addition, a subject can be selected to participate in a clinical trial. Alternatively or in addition, a subject can be classified at step 18 as likely or unlikely to respond to a particular course of treatment. At optional step 20, a particular course of treatment (which can include one or more different individual therapies) can be identified based on the result of step 18. For example, a subject who has been identified at step 18 as unlikely to respond to a particular course of treatment can be identified as likely to benefit from a therapy different from said particular course of treatment. Alternatively, a subject who has been identified at step 18 as likely to respond to a particular course of treatment can be identified as likely to benefit from a therapy that includes said particular course of treatment. As another example, a subject who has been identified at step 18 as having a poor prognosis can be identified as likely to benefit from a more aggressive course of treatment compared to a subject who has been identified at step 18 as having a good prognosis. As yet another example, a subject who has been identified at step 18 as having a first type of disease can be identified as likely to benefit from a treatment directed at that first subtype of disease. At optional step 22, the subject can be treated with the therapy identified at step 20.

[0087] At optional step 24, the result of any one or more of steps 12 to 20 can be provided to the user.

[0088] The subject is preferably a human patient. The subject can be a subject diagnosed with cancer. Thus, the characterized disease can be cancer. The cancer can be ovarian cancer, breast cancer, endometrial cancer (uterus / womb cancer), renal cancer (renal cell), lung cancer (small cell, non-small cell, and mesothelioma), brain cancer (glioma, astrocytoma, glioblastoma), melanoma, Merkel cell carcinoma, clear cell renal cell carcinoma (ccRCC), lymphoma, gastrointestinal cancer (e.g., colorectal cancer), small intestine cancer (duodenum and jejunum), leukemia, pancreatic cancer, hepatobiliary tumors, liver cancer (e.g., hepatocellular carcinoma), germ cell carcinoma, prostate cancer, head and neck cancer, bladder cancer, thyroid cancer, esophageal cancer, melanoma (e.g., uveal melanoma), cutaneous squamous cell carcinoma, and sarcoma. For example, the cancer can be head and neck squamous cell carcinoma (HNSCC), hepatocellular carcinoma (HCC), colorectal cancer (CRC), different types of lung cancer (LC), clear cell renal cell carcinoma (ccRCC), prostate cancer (PC), breast cancer (BC), bladder urothelial carcinoma (BUC), esophageal squamous cell carcinoma (ESCC), uveal melanoma (UV), and cutaneous squamous cell carcinoma (cSCC). The cancer can be brain cancer, lung cancer, or prostate cancer.

[0089] The prognostic or diagnostic feature predicted at step 18 can be a cancer diagnosis or prognosis. The prognostic or diagnostic feature predicted at step 18 can be a disease subtype (e.g., cancer type) or disease severity (e.g., cancer severity or grade). The cancer severity or grade can be a score calculated using any severity metric known in the art. The cancer severity metric predicted at step 18 can be the Gleason score. The disease severity can be a risk score, such as the risk of metastasis / recurrence. The cancer subtype can be any cancer subtype known in the art. For example, in the context of lung cancer, the cancer subtype can be selected from lung adenocarcinoma (luad) and lung squamous cell carcinoma (lusc). Thus, step 18 can include classifying a subject (who has or is suspected of having lung cancer) as having lusc or luad (i.e., classifying the subject between a first class that includes subjects with lusc and a second class that includes subjects with luad). As another example, in the context of brain cancer, the cancer subtype can be selected from low-grade glioma (lgg) and glioblastoma multiforme (gbm). Thus, step 18 can include classifying a subject (who has or is suspected of having brain cancer) as having lgg or gbm (i.e., classifying the subject between a first class that includes subjects with lgg and a second class that includes subjects with gbm).

[0090] The method may further include an optional step 17 of generating a machine learning model configured to predict a diagnosis or prognosis of a disease of a patient. The machine learning model may have been trained using one or more similarity matrices and biological factors. Accordingly, the method may include training the machine learning model using biological data, the biological data including values of a plurality of biological factors of a plurality of individuals (optionally excluding the patient for whom the prediction is being made). The machine learning model may have been trained using training data or may be trained as part of the method, the training data including values of a plurality of biological factors of a plurality of individuals and known prognoses or diagnoses of all individuals except the patient for whom the prediction is being made. Generating the model may include: obtaining, for each of a plurality of training individuals, biological data and a diagnosis or prognosis label associated with the individual, the biological data including values of a plurality of biological factors of the individual; generating, for each of the plurality of individuals, one or more individual networks, each individual network including a plurality of nodes and edges between the nodes, wherein each node indicates a biological factor in the biological data of the individual and each edge indicates a relationship between a pair of biological factors corresponding to the nodes connected by the edge in the respective individual; determining values of one or more similarity metrics between one or more individual networks generated for the patient and one or more individual networks generated for other individuals of the plurality of individuals; and generating a machine learning model configured to predict a diagnosis or prognosis of a disease of a patient, wherein the machine learning model takes as input the values of one or more similarity metrics between the individual networks and produces a diagnosis or prognosis as output. The method may be performed in the context of performing quality control on biological data of a patient having a disease. Accordingly, methods are also described herein that include characterizing a disease of a patient using biological data of the patient as described with respect to steps 10 to 18, the biological data including values of a plurality of subsets of biological factors obtained using respective different data modalities.

[0091] Characterizing a disease of a patient using biological data of the patient (the biological data including only values of a first subset of biological factors) as described with respect to steps 10 to 18; and performing quality control on the data at an optional step 23 by comparing a predicted diagnosis or prognosis obtained using a plurality of subsets of biological factors and the first subset of biological factors, wherein the predicted diagnosis or prognosis of the first subset of biological factors is different compared to the plurality of subsets of biological factors, indicating poor quality of the biological data including the first subset of biological factors.

[0092] System

[0093] Figure 2An embodiment of a system for characterizing a subject and / or for providing prognostic, diagnostic, or treatment recommendations in accordance with the present disclosure is shown. The system includes a computing device 1, which includes a processor 101 and a computer-readable memory 102. In the illustrated embodiment, the computing device 1 further includes a user interface 103, which is shown as a screen, but may include any other device for communicating information to the user, such as, for example, via auditory or visual signals. The computing device 1 is communicatively connected to a biological data acquisition device 3 (such as, for example, via a network), such as, for example, a sequencer, a microscope, a mass spectrometer, etc., and / or is connected to one or more databases 2 storing biological data (values of a plurality of biological factors). The one or more databases 2 may further store one or more of the following: one or more machine learning algorithms, training data, parameters (such as, for example, parameters of a machine learning model, a feature selection algorithm, an IN calculation method, etc.), clinical and / or sample-related information, etc. The computing device may be a smart phone, a tablet computer, a personal computer, or other computing device. The computing device is configured to implement a method for characterizing a diseased subject as described herein. In an alternative embodiment, the computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to implement a method for characterizing a diseased subject as described herein. In such a case, the remote computing device may also be configured to send the results of the method to the computing device. Additionally, the various steps of the method described herein may be divided between the computing device 1 and the remote computing device. Communication between the computing device 1 and the remote computing device may be via a wired or wireless connection, and may occur over a local or public network 6 (such as, for example, over the public Internet). The biological data acquisition device may be wired to the computing device 1, or may be capable of communicating via a wireless connection, such as, for example, via WiFi and / or the public Internet, as shown. The connection between the computing device 1 and the biological data acquisition device 3 may be direct or indirect (such as, for example, via a remote computer). The biological data acquisition device 3 is configured to acquire biological data including values of a plurality of biological factors from a sample previously obtained from a subject. The biological data acquisition device 3 may include a gene expression data acquisition device (such as a next-generation sequencer) and / or a histopathology data acquisition device (such as a microscope).

[0094] It is presented below by way of example and should not be construed as limiting the scope of the claims.

[0095] Example

[0096] Personalized cancer screening prior to treatment paves the way for improved diagnostic accuracy and treatment outcomes. Most methods are limited to a single data type and do not consider interactions between features, overlooking the complementary insights that multimodality and systems biology can provide. In these embodiments, the inventors demonstrate data integration via individual networks using graph theory, where nodes and edges are individual-specific. They showcase the results of early, intermediate, and late fusion of graph-based RNAseq data and histopathological whole-slide images in predicting cancer subtypes and severity. The methods demonstrated are as follows: 1) create individual networks; 2) calculate similarities between individuals from these graphs; 3) train a model on the similarity matrix; 4) evaluate performance using the macro F1 score. The pros and cons of the pipeline elements are evaluated based on publicly available real-world datasets. The inventors show that graph-based methods can improve performance compared to methods that do not study interactions. Additionally, combining multiple data sources tends to improve classification compared to single-data-based models, especially through intermediate fusion. The proposed workflow is demonstrated in the context of cancer but can be adapted to other disease contexts to accelerate and enhance personalized healthcare.

[0097] Example 1 - Predicting Outcomes via Individual Graphs

[0098] In this embodiment, the inventors describe a newly developed multi-step workflow (see Figure 3 ) to predict outcomes via individual graphs. First, a network is constructed for each individual: nodes and / or edges are individual-specific. From these individual networks, we calculate a similarity matrix, which we call the pairwise proximity network (PPN): nodes are individuals and edges represent the degree of similarity between individuals. Information from various levels of the individual graphs is used to construct the pairwise proximity network: nodes, edges, or both nodes and edges. The pairwise proximity network becomes the input to a machine learning model. In other words, we treat the similarities to a reference set as variables. Then, outcomes are predicted based on these similarities to the reference set.

[0099] Method

[0100] Data. The results regarding the impact of using or not using individual graphs are based on publicly available real-world data generated by the TCGA Research Network

[49] (https: / / www.cancer.gov / tcga). For each patient, two types of data are used: images (specifically, histopathological whole-slide images) and genomic data (specifically, gene expression data obtained via RNA sequencing). For the images, features are extracted as described below. For the RNAseq data, the feature is the read count for each gene.

[0101] These embodiments focus on three use cases: prostate cancer severity using Gleason scores, differentiation of low-grade glioma (lgg) and glioblastoma multiforme (gbm) in the brain, and differentiation of lung adenocarcinoma (luad) and lung squamous cell carcinoma (lusc). For each use case, we analyzed two data modalities: RNAseq data and whole-slide histopathology images (WSIs). Prostate cancer is the most common cancer in men, and prostate cancer stage is typically described according to Gleason scores, which helps to evaluate prognosis (Egevad et al. 2002). The score is derived from the appearance of cancerous cells and can correspond to 5 patterns (from normal to tumor cells). Grade 1 cells are indistinguishable from normal prostate tissue; grade 5 corresponds to tumor cells. Thus, the higher the Gleason score, the more severe the cancer. Doctors determine the Gleason score by examining biopsy samples and assign a grade to the dominant pattern (primary Gleason score). Typically, a second Gleason grade is given to the second dominant pattern, and the two grades are added to set the secondary Gleason score. These embodiments focus on the primary Gleason score, and particularly patterns 3 and 4. In this work, the inventors examined whether the newly developed workflow could highlight the differences between these two patterns. The database contains 297 individuals in the training set (130 pattern 3, 167 pattern 4) and 71 individuals in the test set (34 pattern 3, 37 pattern 4).

[0102] Low-grade glioma (lgg) in the brain is a cancerous brain tumor. It originates from the supportive cells in the brain. Glioblastoma multiforme (gbm) is an invasive cancer in the brain or spinal cord. Studies have identified variations between these two tumors, such as sex-specific molecular differences. Here, the inventors investigated whether IN and the combination of RNAseq and histopathology data could help to identify these two brain tumors. The training set contains 344 individuals (282 lgg, 62 gbm), and the test set contains 156 individuals (122 lgg, 34 gbm).

[0103] Lung adenocarcinoma (luad) and lung squamous cell carcinoma (lusc) are the most common lung cancer subtypes and are both considered non-small cell lung cancer (NSCLC). They have different biological signatures, but despite recent research progress, these variations in their biological mechanisms remain to be clarified [6]. The training data contains 603 patients (232 luad, 371 lusc), and the test data has 140 patients (50 luad and 90 lusc).

[0104] Gene expression data processing. Publicly available gene expression data from TCGA (portal.gdc.cancer.gov) has been used. Only samples from primary tumor sites were selected. We downloaded gene expression data in fragments per kilobase of exon per million reads mapped (FPKM), with 56,602 Ensemble gene identifiers. We normalized the FPKM data to TPM (transcripts per million) using the following equation: Only protein-coding genes were included, using Ensembl annotations. Genes with an expression value of zero were excluded.

[0105] Feature extraction - Histopathology. Feature extraction was performed on histopathology whole-slide images (WSIs). We considered the complete TCGA dataset with 30 cancer types, including the types to be classified in the above 3 use cases (i.e., brain cancer, lung cancer, and prostate cancer), but excluded individuals in the test set. A pre-trained neural network model was applied to distinguish between cancer types. Specifically, we used Resnet18 (He et al., 2015) and Attention MIL (Ilse et al., 2018), trained for 100 epochs on all TCGA slides, with 128 random tiles sampled per slide per epoch. The ReNet18 classifier pre-trained on imageNet was pre-trained and then used to create embeddings, which were used by the trained Attention MIL model to provide classification based on the embeddings. 512 features contained in the N-1 layer were selected as new variables. Each feature is a vector of length equal to the number of individuals and contains discriminative information about the cancer type. We hypothesized that differences in cancer types would provide relevant information for distinguishing groups in our 3 use cases. Therefore, a table composed of individuals in rows and neural network features in columns was used as the input data for histopathology information.

[0106] Similarity between individuals - Single data source. Three categories of methods for calculating similarity between individuals were tested and are explained in more detail below. In the first category of methods, similarity is obtained only based on node values. This is the baseline in the sense that network information is never used. In the second category of methods, similarity is obtained only based on the interactions between variables (edges). In the third category of methods, similarity is obtained based on both node (raw data) values and edge values.

[0107] Single data source - At the level of individual nodes. We created a baseline that uses only the node weights of the individual network (i.e., only the raw feature values). We call them node-level methods. This can also be called the "raw data" method. This method can be considered not to depend on the individual network structure because the interactions between features are not used in the model. Three ways of obtaining the interpersonal network PPN using these individual graphs were tested. n PPNn (x, y) indicates the similarity degree between individuals x and y. The first option is to use the Euclidean distance between the features of each pair of individuals and calculate an affinity matrix representing the neighborhood graph of the individuals. This is performed using the function affinityMatrix in the SNFtools package (Wang et al. 2021). This function accepts three parameters: a distance matrix (obtained through the Euclidean distance in this case), the parameter K, and σ. K is the number of neighbors, where the affinity outside the neighborhood is set to zero and the affinity within the neighborhood is normalized. σ is the hyperparameter of the scaled exponential similarity kernel used for the actual affinity calculation. These parameters are chosen empirically (K = 20, σ = 0.05). The variation is to apply a Gaussian kernel. This is performed using the function gausskernel from the KLRS package (Ferwerda et al. 2017). Given two vectors v x and v y , the Gaussian kernel is defined as where ||v x - v y || is the Euclidean distance, and σ 2 (here, σ = 1000) is the bandwidth of the kernel. The third option is to calculate the Spearman correlation between each pair of patients. This is performed using the function rcorr from the Hmisc package (Harrell 2021).

[0108] Single data source - at the individual edge level. The second type of interpersonal network PPN e is calculated to measure the influence considering the interactions between variables. Specifically, we constructed a network for each individual, where the nodes are variables and the edges represent the links between these variables. Since the individual networks become very large when the number of variables in each data source increases, we perform feature selection at the node and edge levels. It enables us to focus on relevant signals, remove noise, and reduce the computation time. Alternatively, this can be done only for gene expression data and imaging data, and the number of variables can be controlled at the step of obtaining features (e.g., based on the architecture of a model that specifies the number of latent variables, such as using regularization). First, we select the data features (nodes) with the highest standard deviation across individuals in the training set (i.e., the k genes with the largest expression variations). Then, we create condition-specific networks for each subgroup for prediction. We calculated the differences between the adjacency matrices of these condition-specific networks and the selected edges, where the selected edges have a large absolute difference at their co-expression level (calculated as the product of nodes, see below). We select the edges with a Pearson R correlation coefficient difference of at least l. To define the thresholds k and l, we performed stratified 5-fold cross-validation within the training set and selected the parameters that on average bring a higher macro F1 score.

[0109] We obtained individual edge weights based on two methods. In the first method, which we call "node product", we applied the min-max normalization algorithm across the variable values to scale them between 0 and 1: where i and i' are the variable and its normalized version respectively. Then, for individual x, the weight of the edge between nodes i and j is e x ij =i ′x *j ′x , where i' and j' are the normalized versions of variables i and j. Since this method is computationally efficient, it can be used to construct an IN for feature selection.

[0110] The second method for creating edge weights is the LIONESS algorithm (Kuijer et al. 2019). The overall idea is to study the difference between the network composed of all individuals and the network composed of all individuals except one. If there is a difference, it must be due to the omission of the said individual. The LIONESS equation is as follows: e x ij =N(e α ij -e α-x ij )+e α-x ij , where e α ij is the weight of the edge between nodes i and j in the network modeled on all N samples, and e α-x ij is the weight of this edge in the network modeled on all samples except the sample x of interest. For each individual, we used the lionessR function (Kuijer, 2022) to obtain the edge weights. It should be noted that these edge weights are specific to the reference set used to calculate e α ij .

[0111] Regardless of the method used, we reduced the obtained IN to the previously identified edge selection. Finally, we applied the min-max scaling algorithm to the IN so that the edge weights are between 0 and 1. Specifically, we considered the minimum and maximum weights across all INs in the scaling so that the order of the weights between individuals remains unchanged. We used the similarity to the individuals in the training set as a new variable for prediction.

[0112] Using the degree of similarity between an individual and a reference set (here the training set) as a predictor variable requires defining a measure of distance between individuals. Many methods have been developed to compare graphs. The particularity of our context limits the choice of distance. In fact, even on large graphs, the measure should be computed in a reasonable amount of time and should handle undirected and weighted networks. Since the same variables (e.g., the same genes) are used for all individuals, there is a node correspondence between different INs (node 1 in IN x1 corresponds to node 1 in IN x2 where x and x2 are different individuals). This type of graph is called a multiplex. In fact, we are comparing multiple layers of the same graph, which correspond to different individuals. From the way we construct the INs, in addition to having the same nodes, individuals also have the same edges. It's just that the edge weights vary from person to person. This means that without additional filters, all graph distances based on structural differences will not allow us to identify whether some graphs are more similar than others. For example, Euclidean distance, Jaccard distance, edge difference distance, or DeltaCon distance are suitable for our context. We used the edge difference distance in this project because it has good computational properties. This distance takes two adjacency matrices and computes the Frobenius norm of their difference. We applied this distance to each pair of individual networks to obtain a similarity matrix between individuals.

[0113] Individual network inference for new patients is performed by the LIONESS algorithm. By the node product method, it is straightforward to calculate the degree of similarity between a new individual and each individual in the training set. However, with the LIONESS algorithm, when we consider a new individual, the INs of the reference set change. Moreover, the derivation of the IN for the new individual depends on this set. Therefore, to create an IN for a new individual, we use all individuals from the training set and one new individual at a time. If we directly add all new individuals to the individuals from the training set (reference set), we can significantly modify the relationships between the INs of the training set. Adding only one individual to the training set at a time minimizes the impact on the similarity between individuals from the training set. In summary, with the LIONESS algorithm, for each individual x from the test set, we create a new database containing the reference set and x, and we construct the INs for all patients in this temporary database. From these new INs, we calculate the similarity between x and each individual in the training set. Therefore, the complexity of creating individual networks for individuals in the test set by the LIONESS algorithm motivates the use of the node product method to start the analysis.

[0114] Single data source—combination of individual nodes and individual edges. There is no reason to assume that the information of individual nodes and edges cannot be complementary. Therefore, we also investigate the combination of interpersonal networks constructed by individual nodes and individual edges. We construct interpersonal networks independently for both methods, and we average their corresponding adjacency matrices to merge them.

[0115] Data Integration—Early Fusion. One option to integrate information from multiple databases is to concatenate the original data (early fusion) and apply the pipeline as in the single-data process. When using individual edges or edges and nodes, we include an additional step: for each data source, the top variables are selected as described above. Multiple edge correlation thresholds (0.25, 0.5, and 0.75) were tested to further reduce IN. Here, nodes can be variables from any of the original databases, and edges can therefore represent associations between variables of any type.

[0116] Data integration - intermediate fusion. Fusion can also be performed at the PPN level (intermediate fusion). PPNs (similarity matrices) are obtained separately for each data source and merged to benefit from their potential complementarity. Simple methods can be applied to this task, such as calculating the average of different PPNs (average similarity matrices). More advanced methods include similarity network fusion (SNF) (Wang et al. 2014). SNF has been shown to be effective in combining multiple data, such as mRNA expression, DNA methylation, and microRNA expression data for cancer data. In this project, we tested both the average algorithm and the similarity network fusion algorithm. Then, the SVM model was applied as described below.

[0117] Data Integration — Late Fusion. The last alternative considered is late fusion, where the data is merged after each data source (a and b) has been analyzed independently. For continuous outcomes, this can be calculated by summing or averaging. Since we are doing classification, we used a majority voting approach. In most of our application settings, we only consider two data modalities, so majority voting does not provide additional knowledge. Therefore, we only apply late fusion on results obtained from interpersonal networks (derived from the combination of individual nodes and edges), where we consider predictions from four outcomes: PPN n,a 、PPN n,b 、PPN e,a and PPN e,b When two labels are equally predicted for an individual, the final label is randomly assigned to one of them.

[0118] Prediction and performance evaluation. Similarity between individuals (personal network, PPN(x,y)–IN x with IN yThe similarity between (where x and y are two individuals) is normalized by the following transformation: Using the kernlab package (Karatzoglou et al., 2004), a support vector machine (SVM) model is trained using the normalized PPN. The SVM classification method operates directly on the similarity matrix rather than on the original sample representation in the original dimensional space. We tested multiple options for the parameter C (10 k , where k = 1, 2, 3, 4, 5), and we selected the C that brought the best performance. Note that feature selection and hyperparameter tuning were performed on the same training set. Then, the evaluation was performed on an independent test set (as described above). We obtained the performance by comparing with the true labels of the test set data. Due to the imbalance between groups, we used the macro F1 score to evaluate the performance, where where l is the label index and L is the number of labels.

[0119] Results

[0120] A new data integration workflow was proposed, as Figure 3 shown. Three inputs were considered: the RNAseq modality, the histopathology images (input for mid- and late-stage fusion), and the concatenation of these two modalities (input for early-stage fusion). An individual network was constructed separately for each input and for each individual in the training set. Through these individual networks, an interpersonal network was constructed, where the nodes are individuals and the edges represent the proximity between two individuals. Nodes, edges, or both nodes and edges in the individual graphs can be used to construct the interpersonal network. The interpersonal network was used to train a support vector machine (SVM) model for each of the three prediction tasks (classifying prostate cancer individuals as Gleason score 3 or 4, classifying brain cancer individuals as low-grade glioma or glioblastoma multiforme, and classifying lung cancer individuals as lung adenocarcinoma or lung squamous cell carcinoma). Then, the individual networks of the test set were calculated. The similarity between the individuals from the test set and the individuals from the training set (reference set) was calculated to create the interpersonal network of the test set. The SVM model was applied to these similarities with the reference set, and the macro F1 score was used to determine the performance of the classification. All analyses (i.e., different types of modality integration and different levels of information used for calculating the interpersonal network) were compared, as explained in Example 2 and Example 3 below.

[0121] Example 2 - Single Data Source: Utilizing the Influence of Nodes, Edges, or Both Nodes and Edges in Individual Graphs

[0122] Methods

[0123] See Example 1.

[0124] Model comparison. We compared our graph-based method with multiple classification methods applied to the raw features. Namely, we used penalized logistic regression, classification tree, random forest, AdaBoost, and Naive Bayes methods. These algorithms were applied individually to each data type (RNAseq and histopathological features) and to the combined dataset (concatenated RNAseq and histopathological features). For each algorithm, we computed the associated macro F1 score to show how our model and its variants compared to the standard and state-of-the-art classification methods. Note that these five models were only compared with the graph-based methods on the edges of IN, and the nodes and edges of IN, since the method based on the nodes of IN does not use any graph structure in the process.

[0125] Data preprocessing was as follows: Constant variables and correlated variables (|r| > 0.75) were removed. For penalized logistic regression, we used the function cv.glmnet from the package glmnet (Friedman et al. 2010), with the option alpha = 1, lambda = NULL. For random forest, we applied the function randomForest from the randomForest package (Liaw and Wiener, 2002), with the option ntree = 500. For AdaBoost, we used the function boosting from the adabag package (Alfaro et al., 2013), with the options boos = TRUE, and mfinal = 50. For classification tree, we applied the function rpart from the package rpart (Therneau and Atkinson, 2019), with the default options. For the Naive Bayes method, we used the function naiveBayes from the package e1071 (Meyer et al. 2022), with the default options.

[0126] Results

[0127] We compared the prediction performance obtained using different levels of information in the individual graphs (i.e., using and not using the individual graphs). To create IN, feature selection was performed as described in Example 1. It brought 700 genes (edge weight percentile threshold tedge = 0.25) and 300 histopathological features (tedge = 0.25) for the prostate use case. In the brain cancer dataset, 600 genes (tedge = 0.5) and 300 image features (tedge = 0.75) were retained. For the classification of two types of lung cancer, 600 genes (tedge = 0.5) and 500 histopathological features (tedge = 0.75) were considered.

[0128] Figure 4The results of this comparison are shown. The first two columns of each heatmap show the impact on each data modality using the nodes (rows 1 to 3), edges (rows 4 and 5), or both nodes and edges (rows 6 and 7) of the individual networks. Figure 8 Additional visualizations are presented in Figure 8 . Among the three methods of constructing similarity at the node level (only raw data), Spearman correlation performs best in two-thirds of the scenarios. This motivated the choice of Spearman correlation to combine node-level and edge-level information.

[0129] We observed that using more than just node information (i.e., using the graph) increased the macro F1 scores for the prostate use case (maximum F1 = 0.71) and brain use case (maximum F1 = 0.99) leveraging RNAseq data, as well as for the lung use case (maximum F1 = 0.94) leveraging histopathology data. Using individual edges or individual nodes led to equivalent performance in the context of lung cancer from RNAseq data (maximum F1 = 0.94). Pipelines based on node-level information achieved higher predictions through prostate histopathology data (maximum F1 = 0.82) and brain classification (maximum F1 = 0.94).

[0130] Among the two methods of constructing the similarity matrix via individual edges (i.e., using the graph), the node product outperformed the LIONESS algorithm in all cases except for predicting prostate cancer severity using RNAseq data. However, when combining individual nodes and edges, the LIONESS method produced higher results in half of the cases. Overall, for single data, in two-thirds of the cases, classification based on individual edge weights (with or without combination with individual node weights) was better or equivalent than prediction based only on individual nodes (i.e., without individual graph structure). In other words, the data show that in most cases, classification based on individual graphs (with or without combination with raw data) was better or equivalent than raw data prediction. Thus, these results highlight the high potential of individual graphs for disease subtype typing.

[0131] In addition, we separately compared the graph-based models with multiple classification algorithms applied to the raw features for each data type: penalized logistic regression, classification trees, random forests, AdaBoost, and the naive Bayes method. These models were ranked based on their macro F1 scores, where the best model was ranked 1 and the worst model was ranked 6. In Figure 5(a) shows the results, where the smaller the area enclosed by the colored lines indicates better performance. Among the six analyses conducted, the graph-based method outperformed other models in four of them. Overall, the graph-based method showed the best performance, with an average rank of 1.75. Following closely, the adaboost algorithm achieved the second-best performance, with an average rank of 2.5. These results demonstrate the great potential of individual graphs for disease subtype classification.

[0132] Example 3 - Multi-Data Integration: Influence of Early, Mid, and Late Integration

[0133] In this example, the inventors studied the impact of using graphs on the predictive performance of multimodal integration. Specifically, they focused on three different fusions that occur at different stages of the pipeline: early, middle, or late.

[0134] Method

[0135] See Example 1 and Example 2.

[0136] Results

[0137] The first objective of this analysis was to study whether it was possible to use graphs to predict disease subtypes and severity based on patient data. The results showed that in brain cancer, the workflow achieved perfect prediction. We also obtained high performance (macro F1 score = 0.97) in the lung cancer use case. Detecting the severity of prostate cancer was more difficult, with a maximum macro F1 score of 0.82. The second objective of this study was to examine which modality produced the best prediction. The answer varied depending on the use case. Histopathological images were the most informative data in the prostate scenario, but RNAseq data achieved better results in the lung and brain scenarios.

[0138] The third objective was to use the results of using IN and PPN to combine databases with information at each step. Figure 4 The impact of using the edge weights of individual graphs for multimodal integration is shown in rows 4 and 5 of the heatmap. An alternative visualization is presented in Figure 9. Since the interpersonal network only originates from individual edge weights (i.e., no node weights), the fusion of the two data sources provided better results for prostate cancer (maximum F1 = 0.75) and lung cancer (maximum F1 = 0.96). For brain cancer, there was no difference between one modality or the fusion of two modalities (maximum F1 = 0.97). Therefore, for an interpersonal network that only originates from individual edge weights, combining multiple modalities is beneficial. In other words, just through the graph, the data shows the benefits of combining multiple modalities. There was no clear winner between the LIONESS method and the node product method.

[0139] When considering the interpersonal network calculated based on the combination of individual edge weights and individual node weights ( Figure 4, for the 6th and 7th rows of each heatmap), the fusion of RNAseq and histopathology data yielded improved predictions for prostate cancer (maximum F1 = 0.79) and brain cancer (maximum F1 = 1). No differences were observed for lung cancer as multiple pipelines produced the best macro F1 score of 0.97. Thus, by combining individual node and edge weights, we also observed improved performance in two-thirds of cases when fusing the two data sources. In our example, intermediate fusion via the mean similarity matrix (mid-fusion) outperformed early or late fusion.

[0140] Finally, we compared the results across all analyses: the use of nodes and / or edges in individual graphs, and the use on one or both data modalities (the entire heatmap— Figure 4 ). For prostate cancer, the two best results were obtained from the Spearman correlation of histopathology data and the mean mid-fusion of the combination of the two data types with node-level and edge-level information. Since doctors determine the primary Gleason score by observing biopsy samples, the good performance observed via histopathology images was expected. For brain cancer, the best results were achieved via the mean mid-fusion of the two data modalities with Spearman correlation and via the mid-fusion of the combination of node-level and edge-level information. Note that brain classification was already perfect (macro F1 score = 1) using only node-level information, and thus no improvement could be obtained using individual edges.

[0141] The ideal use case requires complementary data, with each data providing partial information. For lung cancer, the maximum macro F1 score was obtained from six different settings, involving Spearman correlation and the combination of Spearman correlation and an individual-edge-based method. For both brain cancer and lung cancer, the good performance observed using the combination of individual nodes and edges may be mainly attributed to the nodes. Thus, no single method outperforms others in all contexts, and no general rule can be drawn.

[0142] In addition, we conducted further analyses to evaluate the added value of our graph-based models. Specifically, we compared these models with several classification algorithms applied to the raw features on the combined data types, where features from different data types were concatenated. The outcomes of these analyses are shown in Figure 5(b). Among the three analyses performed (brain, prostate, and lung data), the graph-based method performed best in two of them. When considering the overall performance, both the graph-based method and personalized logistic regression showed the best results, with an average rank of 1.67. Then, the adaboost and naive Bayes algorithms achieved sub-optimal performance, with an average rank of 3.67. These findings further emphasize that the graph-based model provides valuable insights and demonstrates its effectiveness in handling combined data types.

[0143] In summary, the data show that graphs (with or without fusion) achieve highly competitive performance and are often beneficial even on a single data source. Thus, the data show that it is beneficial to consider individual networks for disease subtyping because the performance will be as good as or better than using only the raw data, and even if the performance is not better, the graph-based method still provides additional value (opportunity), such as interpretability, explainability, and flexibility. In fact, the methods described herein can be easily extended to other types of data. Thus, the data show the benefits of considering graph-based methods for supervised learning and especially for multimodal classification.

[0144] Example 4 - Interpretability

[0145] Graphs have important properties in terms of interpretability. For example, when the nodes are genes, the network can be easily overlaid with external knowledge or compared with the results of independent analyses. In this embodiment, the inventors propose to associate the predictions with complementary methods, such as LIMMA (Ritchie et al. 2015) and pathway analysis (Subramanian et al. 2005) to exploit the full potential of the graph.

[0146] Method

[0147] See Example 1.

[0148] LIMMA and Gene Set Enrichment Analysis on Graphs. Initially, LIMMA is the analysis of gene expression data which uses linear models to simultaneously evaluate differential expression among many targets. We applied LIMMA analysis to the RNAseq individual networks as genes are interpretable units of analysis. We selected the co-expressed gene pairs with the largest differences and colored the edges according to their values in the classes to be predicted. It identified the edges with significant differences in weights between groups. At the same time, we also used LIMMA to test for significant differences in gene expression levels between groups. Specifically, we colored the nodes based on the t-statistics from the LIMMA analysis. Thus, in the resulting network, the edges were colored based on whether the weights of the edges were high or low in patients from different groups. Genes with absolute t-statistics < 1.5 were shown in white, and red / blue genes had higher expression in patients from group a / b, respectively. The thicker the edge, the higher the log fold change. The resulting network obtained by applying LIMMA analysis to the selected features described in Example 1 can be further analyzed by Gene Set Enrichment Analysis (BenGuebila et al., 2022). We focused on the largest connected component. Since all genes in the module are connected, they can indicate broader biological mechanisms leading to group differences. For pathway analysis, we performed LIMMA analysis on the selected features described in Example 1, and we used the fgsea package (Sergushichev 2016) to perform Gene Set Enrichment Analysis, with a minimum gene size of 10 and 5000 permutations. Two inputs are required: a ranked gene list and a gene set list for testing enrichment. For the former, we used the gene t-values of the genes in the largest component from the LIMMA analysis. It represents the gene statistical differences between the two groups being compared. For the latter, we downloaded all ontology and curated Molecular Signatures Database (MSigDB version 7) gene sets (Liberzon et al., 2015). We applied an FDR cutoff of 0.05 for significance assessment.

[0149] Results

[0150] We applied LIMMA analysis and Gene Set Enrichment Analysis to the RNAseq data to illustrate the potential of individual networks for understanding biological mechanisms. To detect which edge weights were significantly different between classes (e.g., Gleason score 3 vs. 4), the top 50 edges with the largest co-expression differences were selected and colored (see Figure 6a - b - c). Based on the t - statistics from the LIMMA analysis, nodes with significantly different gene expressions between groups were also identified. This visualization outlines the organization of the most relevant gene pairs, differentiates the individual groups, and highlights specific nodes and interesting modules. For example, we can examine the genes with the most connections (more than 5 connected neighbors). The Gleason score classification indicates that MAP7 can predict the survival of stage II colon cancer patients (Blum et al. 2008). In brain cancer prediction, GTP2 and HIPK2 were identified. GTP2 is associated with neurological diseases, encephalopathies, and microcephaly (e.g., Hengel et al. 2018), and HIPK2 is associated with tumor progression and malignancy (e.g., Garufi et al. 2019). In lung cancer differentiation, we detected TGM2 and DUSP4. The loss of DUSP4 was observed in EGFR - mutant tumors (Chitale et al. 2009). Thus, the graphical method helps to distinguish target genes and gene pairs between the two analyzed groups.

[0151] We performed gene set enrichment analysis to examine the biological mechanisms associated with the differences between subtypes. It was based on the LIMMA analysis, which included all the identified features explained in the method. That is, 700 genes were analyzed for prostate cancer (each percentile edge - weight threshold tedge = 0.25), 600 genes for brain cancer (tedge = 0.5), and 600 genes for lung cancer (tedge = 0.5). 36 gene sets were enriched in the prostate cancer use case (see Figure 10 ), and 10 gene sets were enriched in the lung cancer use case. No enriched pathways were detected for the two brain cancers. The most significant gene set in the prostate analysis was Chandran metastasis. In prostate cancer, metastasis represents the worst outcome, and it is hypothesized that the genes associated with this pathway play a role in the biology of metastatic disease (Chandran et al. 2007). We also identified the Liu prostate cancer set (Liu et al. 2006), which is related to the study showing that sex - determining region Y - box 4 is a transforming oncogene in human prostate cancer cells. In the lung cancer scenario, the most significantly enriched gene set was Shedden lung cancer good survival a4 (Shedden et al. 2008), originating from the study of lung adenocarcinoma survival prediction based on gene expression. Thus, these results highlight the relevance of the graph in identifying the biological processes involved in differentiating cancer subtypes.

[0152] Example 3 - Discussion and Conclusions

[0153] Methods

[0154] See Example 1 and Example 2.

[0155] Results

[0156] Despite the increasing amount of human data, the research on data modality integration methods remains insufficient. Usually, late integration is performed manually and relies on prior knowledge of the disease under study. Moreover, biological mechanisms are often organized into complex systems. Assigning networks to each individual can model such interactions while taking into account individual specificities. Starting from these observations, we analyzed the added value of individual graphs for cancer subtype classification. We integrated data in the individual space using the similarity network between patients rather than measurement results (e.g., gene expression). First, we evaluated the benefit of converting input data into individual graphs on individual data. We showed that treating features as a connection system can improve prediction performance. Second, we demonstrated that combining individual networks and multimodal integration can yield better performance. Finally, as we illustrated with cancer data, one advantage of graph-based methods is the ability to visualize and gain insights into the causal factors that lead to differences between disease subtypes. Although we focused here on RNAseq and histopathology data, our framework is applicable to any multiplex data. In clinical studies, it provides an opportunity to integrate various measurements such as demographic, microbiome, and metabolomics data. Using the interpersonal network (i.e., the similarity matrix between individuals) as the input for machine learning is advantageous because such an interpersonal network, although derived from the original data, is completely independent of the original data. This means that any type of data can be combined to obtain such a PPN, and then the PPN can be used for prediction as shown.

[0157] We have made several choices for constructing individual networks representing each patient. For example, in the lioness algorithm, our edge weights are based on Pearson correlation. Some studies have shown that double weights can also yield good results on RNAseq data. One drawback of lioness is that the calculation of edge weights is based on a reference set. Therefore, the computational efficiency of this option is not as good as the node product because we need to recalculate all individual networks when adding a new individual. Moreover, since we are interested in the difference between the network obtained through the population and the network obtained through the population excluding one individual, this method takes into account the loss of information for creating the IN. Another possibility could be to alternatively consider the gain of information to construct such a network.

[0158] Another limitation of our individual networks is that they all have the same structure: the same nodes, the same edges, and only the edge weights vary between different patients. This limits the distance choices for us to evaluate the degree of difference between two networks. To address this issue, the IN can be filtered. One of the simplest and most commonly used methods for sparse networks is to set a threshold, such as a quantile, and only consider the edges with weights higher than the said threshold. This quantile can be calculated for each individual (selecting the top edges of each individual) or across individuals. If such an additional filter is used on the individual networks to obtain different structures, then metrics such as spectral distance, graphlet-based metrics, feature portrait divergence, or graph kernel-based measurements can be tested. Focusing on the relevant modules in each individual network not only benefits from being able to apply more advanced distances between graphs, but also from being able to focus on predictive signals and remove noise. In our application, we consider the node weights (from the original data) and edge weights (from lioness or node products) separately, and we combine the original data and graph data at the similarity matrix level by calculating the average. Another possibility is to use the distance between graphs that simultaneously considers node weights and edge weights. Thus, the combination will be at the level of the individual graphs rather than at the level of the similarity matrix.

[0159] In this study, one data type in the cancer use case is histopathology whole-slide images. This data was pre-converted to transform the images into image features with continuous values. To achieve this transformation, we considered a dataset of individuals including but not limited to our use case, and we applied a neural network model to predict the cancer type. From this model, we obtained the features that distinguish the cancer types. Although it may seem more straightforward to apply a neural network to each use case and train it based on the label of interest (such as Gleason score), we observed that distinguishing cancer types produced better results (data not shown). One possible explanation is that an image model trained purely for group identification results in less generalizable embeddings. In addition, a more general cancer type model contains more individuals, which can lead to better embeddings for identifying basic features.

[0160] Future enhancements include data integration strategies that leverage graph specificities. In this work, we studied the impact of combining individual graphs and data integration, but we did not use network properties in the integration itself. The data is integrated either before calculating the individual graphs (early integration) or after deriving the similarity matrix (mid and late integration). An alternative would be to combine the data within the process of creating the individual networks. In Figure 3Among these, this would correspond to the intermediate integration occurring at the level of the second box (“individual network”). For example, a method can be developed to select predictive features in the individual network obtained using the first database (e.g., RNAseq) from the second dataset (e.g., histopathological data). Such a method can allow focusing on interpretable variables while incorporating knowledge from additional databases. Secondly, it would sparsify the individual network, enabling the use of more advanced graph distances to compare individuals. Additionally, the method can accommodate features from data modalities whose nodes cannot be easily mapped into the individual graph. In such a case, the relationships between individuals can be directly modeled in the similarity matrix.

[0161] The proposed method is flexible and not specific to one machine learning model. This embodiment uses a support vector machine model because the method operates directly on the similarity matrix. Another option is to create an embedding of the individual network and apply another machine learning model, such as a random forest of neural networks. It is worth noting that due to our small sample size, neural networks tend to be less interpretable and may provide lower performance. The graph-based method proposed here advantageously enables multiple options related to the fusion stage, namely early, late, or mid-term. In contrast, most prior art methods are only able to accommodate early or late fusion. Additionally, our model takes as input data the degree of similarity of individuals to reference data. This already differs from the most common models applied to RNAseq data, which tend to consider the upregulation and downregulation of individual genes compared to the average distribution. With our model, not only must the weights of the machine learning model be known to predict the group of new individuals, but the data of the reference set must remain accessible.

[0162] This example further shows that the advantage of a graph-dependent protocol is its interpretability feature. In this project, the inventors used LIMMA analysis to visualize the genes and gene pairs that play the most significant roles in the differentiation of the test group. In these examples, we always distinguish only two groups, but we may encounter situations where we need to analyze more than two classes. In such cases, LIMMA networks can be represented for each pair of groups to highlight the genes and gene interactions that lead to the differences between each pair of groups. For example, if we compare three Gleason score groups (e.g., patterns 3, 4, and 5), then we can represent the LIMMA networks for patterns 3-4, 3-5, and 4-5. It will show the genes that lead to the transitions and thus the genes that contribute to the development and severity of prostate cancer. Additionally, we decided to apply gene set enrichment analysis to the largest components of the LIMMA networks. In fact, since all the genes in the module are connected, they can indicate broader biological mechanisms that lead to group differences. Another possibility is to apply path analysis specifically developed for networks, such as the network neighborhood search protocol (see Duroux et al. 2022), which uses the shortest paths between the genes under study and a reference biological network to account for the network's topology.

[0163] Conclusion. Despite the recent focus on disease subtype classification research, individual treatment decisions remain a challenging issue. Leveraging the complementarity of multiple data sources can help provide more precise subtypes. Ongoing research on multimodal integration mainly considers one variable at a time while ignoring their interactions. The fusion of individual graphs based on considering these interactions can bring additional information. In this study, we demonstrated the potential of graph theory. In particular, we highlighted the advantages of this approach in the context of prostate cancer, brain cancer, and lung cancer subtype classification and severity assessment. We observed that even on a single data source, individual graphs can be beneficial, and we highlighted that intermediate integration is often one of the best-performing. Graph-based methods achieved competitive performance while bringing additional explainability features. We identified biologically relevant genes, gene interactions, and pathways for different use cases. The presented workflow is flexible and can be easily applied to other data modalities. The results inspire more research on the development of individual network methods for precision medicine.

[0164] References

[0165] Esteban Alfaro, Gamez,and Noelia adabag: An R package for classification with boosting and bagging. Journal of Statistical Software, 54(2): 1–35, 2013.

[0166] Jerome Friedman, Trevor Hastie, and Rob Tibshirani. Regularization paths for generalized linear models via coordinate descent. Journal of statistical software, 33(1): 1, 2010.

[0167] Andy Liaw and Matthew Wiener. Classification and regression by random forest. R News, 2(3): 18–22, 2002

[0168] David Meyer, Evgenia Dimitriadou, Kurt Hornik, Andreas Weingessel, and Friedrich Leisch. e1071: Misc Functions of the Department of Statistics, Probability Theory Group (Formerly: E1071), TU Wien, 2022. R package version 1.7 - 11.

[0169] Terry Therneau and Beth Atkinson. rpart: Recursive Partitioning and Regression Trees, 2019. R package version 4.1 - 15.

[0170] Diane Duroux, Héctor Climente - González, Chloé - Agathe Azencott, Kristel Van Steen, Interpretable network - guided epistasis detection, GigaScience, Volume 11, 2022, giab093.

[0171] M.A.H.Akhand,R.N.Nandi,S.M.Amran and K.Murase,"Context likelihood ofrelatedness with maximal information coefficient for Gene Regulatory Networkinference,"2015 18th International Conference on Computer and InformationTechnology(ICCIT),Dhaka,Bangladesh,2015,pp.312-316

[0172] Glass K,Huttenhower C,Quackenbush J,Yuan GC.Passing messages betweenbiological networks to refine predicted interactions.PLoS One.2013 May 31;8(5):e64832.

[0173] Ash et al.2021.Jordan T Ash,Gregory Darnell,Daniel Munro,and BarbaraE Engelhardt.Joint analysis of expression levels and histologi cal imagesidentifies genes associated with tissue morphology.Nature communications,12(1):1–12,2021.

[0174] Schneider et al. 2022. Lucas Schneider, Sara Laiouar-Pedari, Sara Kuntz, Eva Krieghoff-Henning, Achim Hekler, Jakob N Kather, Timo Gaiser, Stefan Fröhling, and Titus J Brinker. Integration of deep learning-based image analysis and genomic data in cancer pathology: A systematic review. European Journal of Cancer, 160:80–91, 2022.

[0175] Wang et al. 2014. Bo Wang, Aziz M Mezlini, Feyyaz Demir, Marc Fiume, Zhuowen Tu, Michael Brudno, Benjamin Haibe-Kains, and Anna Goldenberg. Similarity network fusion for aggregating data types on a genomic scale. Nature methods, 11(3):333–337, 2014.

[0176] Speicher and Pfeifer, 2015. Nora K Speicher and Nico Pfeifer. Integrating different data types by regularized unsupervised multiple kernel learning with application to cancer subtype discovery. Bioinformatics, 31(12):i268–i275, 2015.

[0177] Glass et al. 2013. Glass K, Huttenhower C, Quackenbush J, Yuan GC. Passing messages between biological networks to refine predicted interactions. PLoS One. 2013 May 31; 8(5): e64832.

[0178] Kuijjer et al. 2019. Marieke L Kuijjer, Ping-Han Hsieh, John Quackenbush, and Kimberly Glass. lionessr: single sample network inference in r. BMC cancer, 19(1): 1–6, 2019.

[0179] Menche et al. 2017. Jörg Menche, Emre Guney, Amitabh Sharma, Patrick J Branigan, Matthew J Loza, Frédéric Baribaud, Radu Dobrin, and Albert László Barabási. Integrating personalized gene expression profiles into predictive disease-associated gene pools. NPJ systems biology and applications, 3(1): 1–10, 2017.

[0180] Egevad et al. 2002. Lars Egevad, T Granfors, L Karlberg, A Bergh, and Per Stattin. Prognostic value of the gleason score in prostate cancer. BJU international, 89(6): 538–542, 2002.

[0181] Ilse et al. 2018. Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. In Jennifer Dy and Andreas Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 2127–2136. PMLR, 10–15 Jul 2018.

[0182] He et al. 2015. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. CoRR, abs / 1512.03385, 2015.

[0183] Wang et al. 2021. Bo Wang, Aziz Mezlini, Feyyaz Demir, Marc Fiume, Zhuowen Tu, Michael Brudno, Benjamin Haibe-Kains, and Anna Goldenberg. SNFtool: Similarity Network Fusion, 2021. R package version 2.3.1.

[0184] Ferwerda et al. 2017. Jeremy Ferwerda, Jens Hainmueller, and Chad J. Hazlett. Kernel-based regularized least squares in R(KRLS) and Stata(krls). Journal of Statistical Software, 79(3): 1–26, 2017.

[0185] Kuijer, 2022. Marieke Lydia Kuijjer. lionessR: Modeling networks for individual samples using LI ONESS, 2022. R package version 1.0. Karatzoglou et al. 2004. Alexandros Karatzoglou, Alex Smola, Kurt Hornik, and Achim Zeileis. kernlab - an s4 package for kernel methods in r. Journal of statistical software, 11(9): 1–20, 2004.

[0186] Ritchie et al. 2015. Matthew E Ritchie, Belinda Phipson, DI Wu, Yifang Hu, Charity W Law, Wei Shi, and Gordon K Smyth. limma powers differential expression analyses for rna - sequencing and microarray studies. Nucleic acids research, 43(7): e47–e47, 2015.

[0187] Ben Guebila et al. 2022. Marouen Ben Guebila, Tian Wang, Camila M Lopes - Ramos, Viola Fanfani, Deborah Weighill, Rebekka Burkholz, Daniel Schlauch, Joseph N Paulson, Michael Altenbuchinger, Abhijeet Sonawane, et al. The network zoo: a multilingual package for the inference and analysis of biological networks. bioRxiv, 2022.

[0188] Sergushichev 2016. Alexey Sergushichev. An algorithm for fast preranked gene set enrichment analysis using cumulative statistic calculation. bioRxiv, 2016.

[0189] Liberzon et al. 2015. Arthur Liberzon, Chet Birger, Helga Thorvaldsdóttir, Mahmoud Ghandi, Jill P. Mesirov, and Pablo Tamayo. The molecular signatures database hallmark gene set collection. Cell Systems, 1(6):417–425, December 2015.

[0190] Blum et al. 2008. Craig Blum, Amanda Graham, Matt Yousefzadeh, Jessica Shrout, Katie Benjamin, Murli Krishna, Raza Hoda, Rana Hoda, David J Cole, Elizabeth Garrett-Mayer, et al. The expression ratio of map7 / b2m is prognostic for survival in patients with stage ii colon cancer. International journal of oncology, 33(3):579–584, 2008.

[0191] Hengel et al. 2018. Holger Hengel, Reinhard Keimer, Werner Deigendesch, Angelika Rieß, Hiyam Marzouqa, Jimmy Zaidan, Peter Bauer, and Ludger Schöls. Gpt2 mutations cause developmental encephalopathy with microcephaly and features of complicated hereditary spastic paraplegia. Clinical Genetics, 94(3 - 4):356–361, 2018.

[0192] Chitale et al. 2009. Dhananjay Chitale, Yixuan Gong, Barry S Taylor, Stephen Broderick, Cameron Brennan, Romel Somwar, Benjamin Golas, Lu Wang, Noriko Motoi, Janos Szoke, et al. An integrated genomic analysis of lung cancer reveals loss of dusp4 in egfr - mutant tumors. Oncogene, 28(31):2773–2783, 2009.

[0193] Chandran et al. 2007. Uma R Chandran, Changqing Ma, Rajiv Dhir, Michelle Bisceglia, Maureen Lyons Weiler, Wenjing Liang, George Michalopoulos, Michael Becich, and Federico A Monzon. Gene expression profiles of prostate cancer reveal involvement of multiple molecular pathways in the metastatic process. BMC cancer, 7(1):1–21, 2007.

[0194] Shedden et al. 2008. Kerby Shedden, Jeremy MG Taylor, Steve A Enkemann, Ming S Tsao, Timothy J Yeatman, William L Gerald, Steve Eschrich, Igor Jurisica, Seshan E Venkatraman, Matthew Meyerson, et al. Gene expression-based survival prediction in lung adenocarcinoma: a multi-site, blinded validation study: Director’s challenge consortium for the molecular classification of lung adenocarcinoma. Nature medicine, 14(8):822, 2008.

[0195] All of the references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application were specifically and individually indicated to be incorporated by reference in its entirety.

[0196] The specific embodiments described herein are provided by way of example and not limitation. Any subheadings in this document are included for convenience only and should not be construed as limiting the disclosure in any way.

Claims

1. A computer-implemented method of characterizing a patient's disease, the method comprising: obtaining biological data for each of a plurality of individuals including the patient, the biological data including values of a plurality of biological factors; generating, for each of the plurality of individuals, one or more individual networks, each individual network including a plurality of nodes and edges between pairs of the nodes, wherein each node indicates a biological factor in the biological data of the individual, and each edge indicates a relationship between a pair of biological factors in the corresponding individual corresponding to the nodes connected by the edge; determining values of one or more similarity metrics between the one or more individual networks generated for the patient and the one or more individual networks generated for other individuals of the plurality of individuals; and predicting the diagnosis or prognosis of the patient using a machine learning model configured to predict the diagnosis or prognosis of the patient's disease, wherein the machine learning model has been trained to take as input values of one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output.

2. The method according to claim 1, wherein determining values of one or more similarity metrics includes determining values of one or more similarity matrices, each similarity matrix including values of similarity metrics between individual networks of pairs of the plurality of individuals.

3. The method according to any one of the preceding claims, wherein for each pair of a plurality of pairs of individual networks, the one or more similarity metrics include similarity between edges in the individual networks, similarity between edges in the individual networks and similarity between nodes in the individual networks, or a similarity combining similarity between edges in the individual networks and similarity between nodes in the individual networks.

4. The method according to any one of the preceding claims, wherein each node in an individual network has a value that is the value of a biological factor in the biological data of the corresponding individual.

5. The method according to any one of the preceding claims, wherein each edge in the individual network has a value which is the product of the values of the nodes connected by the edge in the corresponding individual, or the difference between the edge values of the network obtained using the plurality of individuals in the absence or non - existence of the corresponding individual. Optionally, wherein each edge in the individual network has a value e x ij = N*(e α ij - e α-x ij ) + e α-x ij , where e α ij is the weight of the edge between nodes i and j in the network modeled on all N individuals in the plurality of individuals, and e α-x ij is the weight of the edge in the network modeled on all samples except the corresponding individual x.

6. The method according to any one of the preceding claims, wherein for each pair of a plurality of pairs of individual networks, the one or more similarity metrics include: similarity between nodes in the individual network, the similarity between nodes in the individual network being obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the corresponding individual, or a similarity combining similarity between edges in the individual network and similarity between nodes in the individual network, the similarity between nodes in the individual network being obtained as a Spearman correlation coefficient, an affinity matrix, or a Gaussian kernel using a distance metric between vectors corresponding to the nodes in the corresponding individual.

7. The method according to any one of the preceding claims, wherein for each pair of a plurality of pairs of individual networks, the one or more similarity metrics include similarity between nodes in the individual network obtained as a Spearman correlation coefficient, or a similarity combining similarity between edges in the individual network and similarity between nodes in the individual network obtained as a Spearman correlation coefficient.

8. The method according to any one of the preceding claims, wherein for each pair in the plurality of pairs of individual networks, the one or more similarity metrics comprise: similarity between edges in the individual network, the similarity between edges in the individual network being obtained as Euclidean distance, Jaccard distance, edge difference distance, DeltaCon distance, spectral distance, motif-based metric, Hamming distance, shortest path kernel, k-step random walk kernel, graph diffusion distance, and feature portrait divergence, or similarity combining similarity between nodes in the individual network and similarity between edges in the individual network, the similarity between edges in the individual network being obtained as Euclidean distance, Jaccard distance, edge difference distance, DeltaCon distance, spectral distance, motif-based metric, Hamming distance, shortest path kernel, k-step random walk kernel, graph diffusion distance, and feature portrait divergence.

9. The method according to any one of the preceding claims, wherein for each pair of the plurality of pairs of individual networks, the one or more similarity metrics comprise: similarity between edges in the individual network obtained as edge difference distance, or similarity combining similarity between nodes in the individual network and similarity between edges in the individual network obtained as edge difference distance, optionally, wherein the edge difference distance is obtained as the Frobenius norm of the difference between a pair of matrices, the pair of matrices comprising values of edges in the corresponding individual network for which similarity is obtained.

10. The method according to any one of the preceding claims, the method further comprising generating a report on the diagnosis or prognosis of the disease of the patient.

11. The method according to any one of the preceding claims, the method further comprising generating the machine learning model configured to predict the diagnosis or prognosis of the disease of the patient.

12. The method according to any one of the preceding claims, wherein the biological data of each of the plurality of individuals comprises values of a plurality of biological factors, the plurality of biological factors comprising multiple sets of factors obtained using respective data modalities, wherein the biological data comprises biological data obtained using multiple data modalities.

13. The method according to claim 12, wherein the biological data of each of the plurality of individuals comprises values of a plurality of biological factors derived from at least one of transcriptomics, proteomics, metabolomics, microbiome, clinical, medical imaging, demographics, or histopathology data, optionally, wherein the biological data of each of the plurality of individuals comprises values of a plurality of biological factors derived from transcriptomics or proteomics data and values of a plurality of biological factors obtained from histopathology data.

14. The method according to any one of the preceding claims, wherein obtaining one or more individual networks for each of the plurality of individuals comprises: obtaining, for each of the plurality of individuals, at least one individual network using values of a plurality of biological factors, the plurality of biological factors comprising biological factors obtained using at least two different data modalities; and / or wherein the one or more similarity metrics between one or more individual networks comprise one or more similarity metrics derived from individual networks, the individual networks being obtained from data comprising values of biological factors obtained using at least two different data modalities.

15. The method according to any one of the preceding claims, wherein obtaining one or more individual networks for each of the plurality of individuals comprises obtaining a plurality of individual networks for each of the plurality of individuals, each individual network being obtained using values for a corresponding plurality of biological factors, optionally wherein each individual network is obtained using values for a corresponding plurality of biological factors, the corresponding plurality of biological factors being obtained using the same data modality, and the plurality of individual networks comprising individual networks obtained using at least two different data modalities.

16. The method according to any one of the preceding claims, wherein the one or more similarity metrics between one or more individual networks include: A first set of one or more similarity metrics, which are derived from individual networks obtained from data comprising values of biological factors obtained using a first set of data modalities; and a second set of one or more similarity metrics, which are derived from individual networks obtained from data comprising values of biological factors obtained using a second set of data modalities, wherein the first set is different from the second set.

17. The method according to any one of the preceding claims, wherein the one or more similarity metrics between one or more individual networks comprise one or more similarity metrics obtained by combining a plurality of similarity metrics for a pair of individuals, each similarity metric being derived from a pair of individual networks of a corresponding individual obtained from data comprising values of biological factors obtained using a different set of one or more data modalities.

18. The method according to any one of the preceding claims, wherein the machine learning model comprises a plurality of machine learning models, each machine learning model being configured to predict a diagnosis or prognosis of the disease of the patient, wherein each machine learning model has been trained to take as input the values of a corresponding subset of the one or more similarity metrics between individual networks and produce a diagnosis or prognosis as output, wherein the corresponding subset of similarity metrics is derived from individual networks generated from values of biological factors obtained using the corresponding data modality, and wherein providing a diagnosis or prognosis for the patient comprises combining the outputs of the plurality of machine learning models.

19. The method according to any one of the preceding claims, wherein the machine learning model comprises a classification or regression model, and / or wherein the machine learning model comprises a support vector machine model.

20. The method according to any one of the preceding claims, wherein providing a diagnosis or prognosis for the patient comprises combining predictions of disease subtype or severity, and / or wherein the disease is cancer.

21. The method according to any one of the preceding claims, wherein providing a diagnosis or prognosis for the patient comprises: predicting the Gleason score of a patient diagnosed with prostate cancer, classifying patients diagnosed with brain cancer into two categories: a first category corresponding to low-grade glioma (lgg) of the brain, and a second category corresponding to glioblastoma multiforme (gbm), or classifying patients diagnosed with lung cancer into two categories: a first category corresponding to lung adenocarcinoma (luad), and a second category corresponding to lung squamous cell carcinoma (lusc).

22. The method according to any one of the preceding claims, wherein the biological factors comprise gene or protein expression levels and optionally histopathological data, and wherein: The disease is prostate cancer and the biological factor includes the expression level of MAP7; The disease is brain cancer, and the biological factor includes the expression level of GTP2 and / or HIPK2; or The disease is lung cancer and the biological factor includes the expression level of TGM2 and / or DUSP4.

23. The method according to any one of the preceding claims, wherein the biological factor includes latent variables of a trained machine learning model applied to the image data, optionally, wherein the image data is histopathological data and / or wherein the trained machine learning model is a machine learning model that has been trained in a supervised manner to take histopathological data as input and provide disease type labels as output, optionally a neural network.

24. The method according to any one of the preceding claims, wherein at least one of the one or more individual networks, optionally all of the one or more individual networks, includes nodes selected using a feature selection process and / or edges selected using a feature selection process, and / or wherein generating one or more individual networks for each of the plurality of individuals includes applying a feature selection process to a plurality of nodes, each of the plurality of nodes indicating a biological factor in the biological data of an individual, and / or applying a feature selection process to a plurality of edges, the edges indicating the relationship between a pair of biological factors corresponding to the nodes connected by the edge in the corresponding individual.

25. The method according to any one of the preceding claims, wherein generating one or more individual networks for each of the plurality of individuals includes selecting a plurality of nodes to be included in each corresponding individual network, each of the plurality of nodes indicating a biological factor in the biological data of an individual, wherein the selection is performed separately for each individual or uniformly for the plurality of individuals, optionally wherein selecting a plurality of nodes includes selecting a plurality of biological factors that are different between an individual and a reference set of individuals, or selecting a plurality of nodes having variability that meets one or more predetermined criteria across the plurality of individuals.

26. The method according to any one of the preceding claims, wherein generating one or more individual networks for each of the plurality of individuals includes selecting a plurality of edges, the edges indicating the relationship between a pair of biological factors corresponding to the nodes connected by the edge in the corresponding individual, wherein the selection is performed separately for each individual or uniformly for the plurality of individuals, optionally, wherein selecting a plurality of edges includes selecting a plurality of edges that are different between an individual and a reference set of individuals, or selecting a plurality of edges that are different between a plurality of subsets of the plurality of individuals.

27. The method according to any one of the preceding claims, wherein generating one or more individual networks for each of the plurality of individuals includes selecting a plurality of nodes to be included in each corresponding individual network, each of the plurality of nodes indicating a biological factor in the biological data of an individual, wherein selecting a plurality of nodes includes selecting a plurality of nodes having variability that meets one or more predetermined criteria across the plurality of individuals, and / or Generating one or more individual networks for each of the plurality of individuals includes selecting a plurality of edges, the edges indicating a relationship between a pair of biological factors in the corresponding individual that corresponds to the nodes connected by the edge, wherein selecting the plurality of edges includes selecting a plurality of edges associated with a difference between: (a) a first edge value obtained for a pair of nodes of a first subset of the plurality of individuals, and (b) a second edge value obtained for the same pair of nodes of a second subset of the plurality of individuals, the difference satisfying a predetermined criterion, optionally, wherein the first edge value is a correlation between pairs of nodes across the first subset of the plurality of individuals, and the second edge value is a correlation between pairs of nodes across the second subset of the plurality of individuals, and / or wherein the predetermined criterion is that the difference is within a predetermined threshold or optionally among the top x differences among all possible edges between nodes in the individual network after node selection.

28. A computer-implemented method for obtaining a tool for characterizing a disease of a patient, the method comprising: Obtaining biological data and a diagnosis or prognosis label associated with the individual for each of a plurality of training individuals, the biological data including values of a plurality of biological factors of the individual; Generating one or more individual networks for each of the plurality of individuals, each individual network including a plurality of nodes and edges between pairs of the nodes, wherein each node indicates a biological factor in the biological data of the individual, and each edge indicates a relationship between a pair of biological factors in the corresponding individual that corresponds to the nodes connected by the edge; Determining values of one or more similarity metrics between one or more individual networks generated for the patient and one or more individual networks generated for other individuals of the plurality of individuals; and Generating a machine learning model configured to predict a diagnosis or prognosis of the disease of the patient, wherein the machine learning model takes as input the values of the one or more similarity metrics between the individual networks and produces a diagnosis or prognosis as output.

29. A computer-implemented method for providing treatment recommendations for a patient suffering from a disease, the method comprising: Characterizing the disease of the patient using the method according to any one of claims 1 to 27, and Selecting the patient for treatment with a treatment associated with the predicted diagnosis or prognosis.

30. A computer-implemented method for performing quality control on biological data regarding a patient suffering from a disease, the method comprising: Using the biological data regarding the patient, characterizing the disease of the patient using the method according to any one of claims 1 to 27, the biological data including values of a plurality of subsets of biological factors obtained using respective different data modalities; Using biological data regarding the patient that includes only values of a first subset of biological factors, characterizing the disease of the patient using the method according to any one of claims 1 to 27; And Compare the predicted diagnosis or prognosis obtained using the multiple subsets of biological factors and the first subset of biological factors, wherein a difference in the predicted diagnosis or prognosis between the first subset of biological factors and the multiple subsets of biological factors indicates poor quality of the biological data including the first subset of biological factors.

31. A system, the system comprising: at least one processor; and at least one non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 30.

32. A non-transitory computer-readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of claims 1 to 30.

Citation Information

Patent Citations

  • Methods for detection and quantification of analytes in complex mixtures

    US7473767B2