Assessing the robustness and transferability of predictive signatures across molecular biomarker datasets
A supervised learning system addresses the challenge of transferring genetic signatures across diverse datasets by using quantile normalization and pairwise comparisons, ensuring robustness and transferability for precision medicine applications.
Patent Information
- Application Number
- JP2022570234
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-21
- Filing Date
- 2021-01-21
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2041-01-21
AI Technical Summary
Existing methods struggle to transfer genetic signatures across different datasets due to inconsistencies in data generation technologies, patient demographics, and batch effects, limiting their applicability and regulatory approval in precision medicine.
A supervised learning system that autonomously builds gene signatures by training classification or regression models on one dataset and applies them to others, using quantile normalization and pairwise comparisons to ensure robustness and transferability across diverse datasets.
Enables the transfer of gene signatures across different data generation technologies and patient samples, enhancing their reliability and suitability for regulatory approval and clinical application.
Smart Images

Figure 0007741104000001 
Figure 0007741104000002 
Figure 0007741104000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 62 / 963,735, filed January 21, 2020, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Embodiments of the present disclosure relate to the analysis of genetic and other molecular biomarker signatures, and more particularly to assessing the robustness and transferability of predictive signatures across genomic, proteomic, or metabolomic datasets. Summary of the Invention
[0003] In embodiments of the present disclosure, methods and computer program products are provided for identifying transferable molecular biomarker signatures. In various embodiments, at least one signature is read. Each signature associates a first plurality of molecular biomarkers with one of a plurality of output classes. For each of a plurality of datasets, the expression values of each of the first plurality of molecular biomarkers are normalized to each of the plurality of output classes to obtain a plurality of normalized expressions, each associated with one of the first plurality of molecular biomarkers, one of the plurality of output classes, and one of the plurality of datasets. For each of the first plurality of molecular biomarkers, pairwise comparisons are performed between the normalized expressions associated with that molecular biomarker. Each pairwise comparison is performed between normalized expressions associated with the same output class and a different dataset, thereby determining a transferability score for each of the plurality of molecular biomarkers. The first plurality of molecular biomarkers are ranked based on their transferability scores. A second plurality of molecular biomarkers is generated from the first plurality of molecular biomarkers by applying a transferability score threshold to the first plurality of molecular biomarkers.
[0004] In some embodiments, each of the first plurality of molecular biomarkers is a gene. In some embodiments, each of the first plurality of molecular biomarkers is a protein. In some embodiments, each signature comprises a mapping function. In some embodiments, each signature comprises a plurality of synaptic weights. In some embodiments, each output classification comprises a phenotype. In some embodiments, the phenotype is a disease phenotype. In some embodiments, the normalization comprises quantile normalization. In some embodiments, the normalization is relative to a predetermined reference distribution. In some embodiments, performing the pairwise comparison comprises calculating a Kolmogorov-Smirnov statistic.
[0005] In some embodiments, determining the transferability score comprises calculating an average of the pairwise comparisons. In some embodiments, the plurality of datasets comprises at least one dataset from each of a plurality of platform technologies. In some embodiments, the platform technologies comprise microarrays and RNA sequencing. In some embodiments, the platform technologies comprise mass spectrometry, ELISA, antibody arrays, peptide fingerprinting, and / or protein barcoding. In some embodiments, each of the plurality of datasets is derived from the same biological sample.
[0006] In an embodiment of the present disclosure, a computing node is provided that includes a computer-readable storage medium having program instructions embedded therein. The program instructions are executable by a processor of the computing node to cause the processor to perform the following method: A first signature is read. The first signature associates a first plurality of molecular biomarkers with a first plurality of output classifications. For each of the plurality of datasets, the expression values of each of the first plurality of molecular biomarkers are normalized to each of the plurality of output classifications to obtain a plurality of normalized expressions, each associated with one of the first plurality of molecular biomarkers, one of the plurality of output classifications, and one of the plurality of datasets. For each of the first plurality of molecular biomarkers, pairwise comparisons are performed between the normalized expressions associated with that molecular biomarker. Each pairwise comparison is performed between normalized expressions associated with the same output classification and different datasets, thereby determining a transferability score for each of the plurality of molecular biomarkers. The first plurality of molecular biomarkers are ranked based on their transferability scores. The second plurality of molecular biomarkers is generated from the first plurality of molecular biomarkers by applying a transferability score threshold to the first plurality of molecular biomarkers.
[0007] In some embodiments, each of the first plurality of molecular biomarkers is a gene. In some embodiments, each of the first plurality of molecular biomarkers is a protein. In some embodiments, each signature comprises a plurality of synaptic weights. In some embodiments, each signature comprises a mapping function. In some embodiments, each output classification comprises a phenotype. In some embodiments, the phenotype is a disease phenotype. In some embodiments, the normalization comprises quantile normalization. In some embodiments, the normalization is relative to a predetermined reference distribution. In some embodiments, performing the pairwise comparison comprises calculating a Kolmogorov-Smirnov statistic.
[0008] In some embodiments, determining the transferability score comprises calculating an average of the pairwise comparisons. In some embodiments, the plurality of datasets comprises at least one dataset from each of a plurality of platform technologies. In some embodiments, the platform technologies comprise microarrays and RNA sequencing. In some embodiments, the platform technologies comprise mass spectrometry, ELISA, antibody arrays, peptide fingerprinting, and / or protein barcoding. In some embodiments, each of the plurality of datasets is derived from the same biological sample.
[0009] In various embodiments, a computer-readable storage medium having program instructions embodied therein is provided, the program instructions being executable by a processor to cause the processor to perform the following method: At least one signature is read; each signature associates a first plurality of molecular biomarkers with one of a plurality of output classes; for each of the plurality of datasets, the expression value of each of the first plurality of molecular biomarkers is normalized to each of the plurality of output classes to obtain a plurality of normalized expressions, each associated with one of the first plurality of molecular biomarkers, one of the plurality of output classes, and one of the plurality of datasets; for each of the first plurality of molecular biomarkers, pairwise comparisons are performed between the normalized expressions associated with that molecular biomarker; each pairwise comparison is performed between normalized expressions associated with the same output class and different datasets, thereby determining a transferability score for each of the plurality of molecular biomarkers; and the first plurality of molecular biomarkers are ranked based on their transferability scores. The second plurality of molecular biomarkers is generated from the first plurality of molecular biomarkers by applying a transferability score threshold to the first plurality of molecular biomarkers.
[0010] In some embodiments, each of the first plurality of molecular biomarkers is a gene. In some embodiments, each of the first plurality of molecular biomarkers is a protein. In some embodiments, each signature comprises a plurality of synaptic weights. In some embodiments, each signature comprises a mapping function. In some embodiments, each output classification comprises a phenotype. In some embodiments, the phenotype is a disease phenotype. In some embodiments, the normalization comprises quantile normalization. In some embodiments, the normalization is relative to a predetermined reference distribution. In some embodiments, performing the pairwise comparison comprises calculating a Kolmogorov-Smirnov statistic.
[0011] In some embodiments, determining the transferability score comprises calculating an average of the pairwise comparisons. In some embodiments, the plurality of datasets comprises at least one dataset from each of a plurality of platform technologies. In some embodiments, the platform technologies comprise microarrays and RNA sequencing. In some embodiments, the platform technologies comprise mass spectrometry, ELISA, antibody arrays, peptide fingerprinting, and / or protein barcoding. In some embodiments, each of the plurality of datasets is derived from the same biological sample.
[0012] In embodiments of the present disclosure, methods and computer program products are provided for evaluating the robustness and transferability of predictive signatures across datasets. In various embodiments, the method reads at least one signature. Each signature associates a first plurality of molecular biomarkers with one of a plurality of output classifications. For each of a plurality of datasets, each pair of datasets is derived from a different platform technology and biological sample, and a correlation coefficient is determined for each of the first plurality of molecular biomarkers between the pair of datasets. For each of the plurality of output classifications, a class-specific correlation coefficient is determined for each of the first plurality of molecular biomarkers between the pair of datasets. The first plurality of molecular biomarkers are ranked based on their respective correlation coefficients and class-specific correlation coefficients. A second plurality of molecular biomarkers is generated from the first plurality of molecular biomarkers. A transferable signature is provided that associates the second plurality of molecular biomarkers with the first plurality of output classifications. [Brief explanation of the drawings]
[0013] [Figure 1A] 1 illustrates exemplary groups of molecular biomarkers and associated groupings according to embodiments of the present disclosure. [Figure 1B] 1 illustrates exemplary groups of molecular biomarkers and associated groupings according to embodiments of the present disclosure. [Figure 2A] 1 shows RNA extraction and quantification of gene expression according to an embodiment of the present disclosure. [Figure 2B] 1 shows RNA extraction and quantification of gene expression according to an embodiment of the present disclosure. [Figure 3] 1 illustrates a method for ensuring gene transferability according to an embodiment of the present disclosure. [Figure 4] 1 illustrates the effect of quantile transformation on the distribution of expression values across samples in a given dataset according to embodiments of the present disclosure. [Figure 5A] 1 shows a distribution of exemplary gene expression values grouped by phenotypic label according to embodiments of the present disclosure. [Figure 5B] 1 shows a distribution of exemplary gene expression values grouped by phenotypic label according to embodiments of the present disclosure. [Figure 5C] 1 shows a distribution of exemplary gene expression values grouped by phenotypic label according to embodiments of the present disclosure. [Figure 6] 1 shows a comparison between phenotypic labels and datasets according to an embodiment of the present disclosure. [Figure 7] 1 illustrates pairwise Kolmogorov-Smirnov statistics according to an embodiment of the present disclosure. [Figure 8] 1 is a flowchart illustrating the calculation of a metric for feature transferability according to an embodiment of the present disclosure. [Figure 9] 1 is a graph of cumulative probability reflecting the reordering of genes by rank according to an embodiment of the present disclosure. [Figure 10] 1 is a flowchart illustrating a method for determining feature transferability according to an embodiment of the present disclosure. [Figure 11] 10 is a rank plot of Spearman correlation coefficients between microarray and RNA-Seq TPM expression by sample number according to embodiments of the present disclosure. [Figure 12] FIG. 10 is a rank plot using Spearman correlation coefficient as a transferability metric between microarray and RNA-seq TPM expression according to embodiments of the present disclosure. [Figure 13A] 1 is a rank plot of genes using Spearman correlation coefficient as a transferability metric between microarray and RNA-Seq TPM expression according to embodiments of the present disclosure. [Figure 13B] 1 is a rank plot of genes using Spearman correlation coefficient as a transferability metric between microarray and RNA-Seq TPM expression according to embodiments of the present disclosure. [Figure 14] 1 is a plot of Spearman correlation coefficient between microarray and RNA-Seq TPM expression according to an embodiment of the present disclosure. [Figure 15A] 1 is a plot of exemplary transferability statistics by gene rank according to embodiments of the present disclosure. [Figure 15B] 1 is a plot of exemplary transferability statistics by gene rank according to embodiments of the present disclosure. [Figure 16A] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 16B] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 17] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 18] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 19] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 20] 1 is a plot of an exemplary transferability statistic versus gene rank according to embodiments of the present disclosure. [Figure 21] 1 illustrates a compute node according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] A gene signature (or gene expression signature) is a single gene or group of genes in a cell that has a unique characteristic gene expression pattern resulting from an altered or unaltered biological process or pathogenic condition. A gene signature further requires the relationship between genes defined by some set of parameters, weights, values or measures.
[0015] Figure 1 illustrates these associations. In Figure 1A, exemplary gene groups are shown. In Figure 1B, a tree diagram is provided that relates several exemplary genes to groups of interest via exemplary values.
[0016] Gene signatures are important in precision medicine, where gene signatures for particular diseases are used as biomarkers and are particularly useful in diagnosing the presence of disease, classifying disease types, and predicting which patients are most likely to respond to particular treatments.
[0017] Gene signatures can be determined from datasets that assess gene expression—usually the abundance of messenger RNA (mRNA)—from biological samples. Figure 2A shows RNA extraction from cells. These can include experimental samples or patient-derived samples, such as cells collected from blood draws or tumor biopsies. Various mathematical techniques within the fields of bioinformatics and biostatistics can be used to determine gene signatures for specific datasets. Gene signatures can be generated using software tools such as GSEA (Gene Set Enrichment Analysis) or by differential gene expression analysis or pathway analysis. Such tools depend on a specific gene expression dataset as a starting point. Alternatively, genes can be manually enumerated based on a hypothesized mechanism of action.
[0018] Gene expression datasets can be generated from platform technologies such as microarray or RNA-seq analysis, or their derivatives. Figure 2B illustrates several approaches to quantifying gene expression after genetic material has been extracted. However, gene signatures established for one dataset may not exhibit the same expression distribution or pattern when considered for other datasets. Several factors, alone or in combination, can limit the ability to transfer gene signatures between datasets, for example: 1. Processing raw biological samples into sequencing libraries can introduce inconsistencies and biases due to material handling, library chemistry, composition, etc.; 2. The sequencing or array platform technology used to generate the data may result in incompatibilities in direct data comparisons; 3. Demographic (e.g., age, sex) or experimental characteristics of patients / biological specimens may introduce confounding factors; 4. General batch effects resulting from unintended variations in any of the above or other factors.
[0019] Thus, a genetic signature cannot be applied to a different dataset and expected to retain its usefulness without taking steps to ensure its applicability to the new dataset. In other words, a genetic signature cannot be transferred from one dataset to another without assessing and modifying its transferability.
[0020] This creates problems for the validation and commercialization of diagnostic, prognostic, and predictive gene signatures: without the ability to generalize a gene signature to newly generated data sets (e.g., new patient samples), the gene signature will be essentially useless and certainly not worthy of regulatory approval or clinical application.
[0021] Approaches to this problem can be divided into manual and semi-manual approaches. The former rely on curation by domain experts to perform sanity checks and smell tests (i.e., empirical heuristics) on results when gene signatures are transferred to new datasets. This is highly subjective and prone to errors and bias. Furthermore, such manual approaches are not applicable on a commercial scale, nor are they suitable for regulatory approval of diagnostic products. Instead, various mathematical techniques can be employed to reduce this reliance on biased manual input. For example, principal component analysis (PCA)-based methods can be used to convert gene signatures into summary scores that can be compared across datasets. However, such methods have a fundamental limitation: PCA does not work well for composite signatures that describe multiple events. For complex diseases such as cancer, gene signatures often arise from the interactions of many cellular, genetic, and chemical factors, and therefore PCA-based methods may not be appropriate. Another approach uses zero-sum regression signatures learned on high-content data. In this case, the weights are retained from one data set to the next.
[0022] Therefore, precision medicine requires methods for transferring gene signatures from one dataset to another that are robust across data generation technologies and patient sample sources. Such methods must make minimal assumptions about data origin and distribution characteristics and be applicable to gene signatures that represent complex biology.
[0023] To address these and other shortcomings of alternative approaches, the present disclosure provides supervised learning systems and methods that autonomously build gene signatures by training classification or regression models on one or more gene expression datasets (so that the models are blind to dataset technology, processing of raw biological samples, and other batch effects), and can be applied to other, different datasets for predictive tasks.
[0024] In various embodiments, gene expression is assumed to be measured using any transcriptomics platform technology, including but not limited to, RNA sequencing by Illumina or IonTorrent, HTG Edge-seq, Nanostring, qPCR, or microarray. It is further assumed that the expression value for each gene in a specific gene set (or all genes in a genome) is calculated using standard bioinformatics programs (for example, RNA-Seq methods and pipelines known in the art, including those provided by Genialis, Inc.).
[0025] Similarly, although various examples provided below relate to gene expression data, the techniques described herein are generally applicable to molecular biomarkers, including genes, proteins, and metabolites. For example, in some embodiments relating to proteomics data, it is assumed that protein expression has been measured using any proteomics platform technology, including but not limited to mass spectrometry, ELISA, antibody arrays, peptide fingerprinting, protein barcoding, or other similar methods for inferring the protein sequence of multiple proteins from biological samples. It is further assumed that the value for each protein in a particular signature (or all proteins in a proteome) has been calculated using standard bioinformatics programs (e.g., proteomics methods and pipelines known in the art, including those provided by Genialis, Inc.).
[0026] In various embodiments of the supervised learning systems and methods, the inputs include an expression matrix from a dataset and a list of genes (e.g., up to several hundred genes) or other molecular biomarkers such as proteins, and the output is a gene signature function or other signature function related to the molecular biomarkers.
[0027] The signature function is inferred from labeled training data consisting of a set of training samples. Each sample is a pair consisting of an input object (e.g., a gene expression vector) and a desired output value (which may be discrete or continuous). It will be understood that one or more continuous value outputs can be converted into classifications by binning, thresholding, winner-take-all, and various other methods. The training data is analyzed to generate an inferred function, which can be used to map new samples from other, different datasets. The inferred gene signature function can take various forms according to the particular machine learning method employed. For example, the signature function can be a matrix operator that can be applied to the input expression matrix from the sample. In another example, the signature function can be a set of synaptic weights for an artificial neural network.
[0028] In various embodiments, supervised learning techniques such as artificial neural networks, random forests, support vector machines, and logistic regression analysis are employed. It will be understood that various additional supervised learning techniques are suitable for use with the present disclosure. Ensemble techniques such as stacking are used in various embodiments to improve accuracy. Particular care must be taken to avoid overfitting, especially in parameter tuning. The training and test datasets must contain separate, non-overlapping sample sets. Samples may be split using cross-validation, bagging (bootstrap aggregation), or other techniques.
[0029] In some embodiments, the feature vector is provided to a learning system. Based on the input features, the learning system generates one or more outputs. In some embodiments, the output of the learning system is a feature vector.
[0030] In some embodiments, the learning system comprises an SVM. In other embodiments, the learning system comprises an artificial neural network. In some embodiments, the learning system is pre-trained using training data. In some embodiments, the training data is retrospective data. In some embodiments, the retrospective data is stored in a data repository. In some embodiments, the learning system may be additionally trained by manual curation of previously generated outputs.
[0031] In some embodiments, the output of the learning system is a trained classifier. In some embodiments, the trained classifier is a random decision forest. However, it will be appreciated that a variety of other classifiers are suitable for use with the present disclosure, including linear classifiers, support vector machines (SVMs), or neural networks such as recurrent neural networks (RNNs).
[0032] Suitable artificial neural networks include, but are not limited to, feedforward networks, radial basis function networks, self-organizing maps, learning vector quantization, recurrent neural networks, Hopfield networks, Boltzmann machines, echo state networks, long short-term memory, bidirectional recurrent neural networks, hierarchical recurrent neural networks, probabilistic neural networks, modular neural networks, associative neural networks, deep neural networks, deep belief networks, convolutional neural networks, convolutional deep belief networks, large memory storage and retrieval neural networks, deep Boltzmann machines, deep stacking networks, tensor deep stacking networks, spike and slab restricted Boltzmann machines, compound hierarchical-deep models, deep coding networks, multi-layer kernel machines, or deep Q-networks.
[0033] Referring to Figure 3, a method for ensuring gene transferability according to an embodiment of the present disclosure is shown. At 301, quantile normalization of expression values is performed. At 302, calculation of feature transferability statistics is performed. At 303, features (e.g., genes) are filtered by a transferability threshold.
[0034] For illustrative purposes, the following example utilizes exemplary data. It will be understood that the present disclosure is applicable to a variety of datasets and indicators, and this example is illustrative rather than limiting. In this example, gene expression data is obtained from the following datasets: Asian Cancer Research Group (ACRG); The Cancer Genome Atlas (TCGA); and Singapore Cohort (SING).
[0035] Individual samples in these datasets are further labeled as the following phenotype classes: phenotype 1, phenotype 2, phenotype 3, phenotype 4.
[0036] Quantile normalization is a technique for making the statistical properties of two distributions the same. Figure 4 shows the effect of quantile transformation on the distribution of expression values across samples in a given dataset. The dataset is normalized to a reference distribution, which is one of the standard statistical distributions, such as a uniform distribution, a Gaussian distribution, or a Poisson distribution. The reference distribution can be generated by randomly or regularly sampling from the cumulative distribution function of the distribution. Any reference distribution can be used.
[0037] All gene expression datasets are then normalized to the same reference distribution. The transformation is applied independently to each feature (expression value of one gene). First, an estimate of the feature's cumulative distribution function is used to map the original values to a uniform distribution. The resulting values are then mapped to the desired output distribution using the associated quantile function.
[0038] The robustness of the procedure increases logarithmically with the number of samples. Several dozen samples (approximately 30 or more) per dataset are required to ensure a base level of performance for the gene signature. The overall performance of the gene signature gradually increases and levels off as the number of quantile-normalized samples reaches several hundred.
[0039] In various embodiments, quantile normalization is used as a preprocessing step in supervised learning, and therefore special care must be taken to avoid overfitting. Quantile normalization parameters must be fitted to a training set of samples and then used to transform test and validation samples. Test and validation samples must be excluded from fitting the quantile normalization parameters.
[0040] Transferable features (genes) should have similar distributions of gene expression values across datasets given the target variable. However, some should differ significantly and be excluded from the gene signature. Differences can be due to technology (e.g., RNA-Seq vs. microarray), experimental bias, population bias, and other influences.
[0041] In Figures 5A-C, the distribution of exemplary gene expression values is grouped by four phenotypic labels (in the legend). Top panel: gene CCL3; bottom panel: gene IFNa2. Figures 5A-C are for the ACRG, TCGA, and SING datasets, respectively. Expression values are quantile normalized to a uniform distribution (within each dataset). Distribution estimates of gene expression for CCL3 are consistent across datasets, but not for IFNA2.
[0042] The present disclosure provides a metric for feature transferability, defined as a reduced set of test statistics obtained from pairwise comparisons of distributions of gene expression datasets.
[0043] The test statistic should be selected based on whether the target variable is categorical, continuous, or other. In the example case below, the metadata is categorical (phenotypes 1-4). Feature transferability is calculated from aggregation, e.g., the arithmetic mean of pairwise Kolmogorov-Smirnov tests of phenotype-specific distributions of gene expression between datasets. This process is illustrated in Figure 6, where four phenotypic labels are compared in a pairwise fashion between the first and second datasets and between the first and third datasets. Aggregation can also be achieved by considering median or min-max range characteristics; the most appropriate type of aggregation can be calculated empirically.
[0044] The Kolmogorov-Smirnov (KS) test is a nonparametric test for the equality of continuous one-dimensional probability distributions that quantifies the distance between the empirical distribution functions of two samples. The KS statistic is defined as the maximum difference between the two joint cumulative distribution functions. The arithmetic mean of the KS statistic represents the average distance between the distributions of expression values grouped by four phenotypes.
[0045] Figure 7 shows the pairwise Kolmogorov-Smirnov statistics. The light and dark lines correspond to the empirical distribution functions, respectively, and the black arrows indicate the difference in distributions obtained by the KS statistics.
[0046] This metric can be used to reduce the bias of a dataset by removing features with inconsistent gene expression distributions. For each gene, multiple KS statistics are calculated, one for each phenotype-dataset pair combination. To obtain a single transferability score for each gene, the KS statistics need to be aggregated across phenotype-dataset pairs. Among common aggregation methods, the arithmetic mean performed well for these exemplary datasets. However, it will be understood that alternative methods such as median, minimum, and maximum may also be used in some embodiments.
[0047] Referring to Figure 8, the computation of a metric for feature transferability according to an embodiment of the present disclosure is shown. At 801, a series of KS tests are computed for a) gene expression values across all phenotype / outcome classes; and b) two dataset pairs (ACRG-TCGA, TCGA-SING). In this example, the phenotype / outcome classes include Phenotype 1, Phenotype 2, Phenotype 3, and Phenotype 4. At 802, the average of the eight KS tests is computed.
[0048] At 803, the K-S statistic is plotted and ranked for all genes in a particular signature. At 804, the ranked gene loadings are thresholded. In some embodiments, thresholding is performed by selecting a point (on the X-axis) just before the start of the rapidly increasing tail of the K-S statistic. Genes with low K-S statistic (ranked closest to 1) are considered the most transferable. In some embodiments, thresholding is performed by converting the K-S statistic to a p-value using a standard conversion table to select a p-value cutoff (setting the threshold on the Y-axis instead of the X-axis). After correcting for multiple hypothesis testing, a useful p-value threshold can be reliably selected.
[0049] At 805, genes that do not meet the KS or p-value threshold are removed from the signature.
[0050] 9, a graph of cumulative probability is provided that reflects the sorting of genes by rank. In this case, a threshold fix value is set at gene rank 98 (out of 125 genes in the example signature) based on the rapid increase in the slope of the curve at values greater than 98. Thus, genes ranked between 99 and 125 may be classified as "non-transferable" and removed from the model.
[0051] The threshold may be estimated automatically by determining the second derivative of the diversion probability curve to identify the inflection point. It will be appreciated that various techniques are known for locating such a threshold. For example, in some embodiments, an average is obtained using a sliding window. In some embodiments, the threshold is set by a predetermined change in the slope of the curve. In some embodiments, the threshold is determined empirically based on the distribution of slope changes.
[0052] The methods described herein are applicable in any pharmaceutical or diagnostic research and development context where gene expression data is being evaluated for predictive potential. For example, the transferable gene signature output from this method can form the companion diagnostic (CDX) or laboratory-developed test (LDT) standard for a drug. Thus, the transferable gene signature can form the standard for approved diagnostic tests deployed at the point of care by clinical practitioners. Alternatively, the transferable gene signature can constitute a list of potential drug targets for early drug discovery research and development. Because the transferable gene signature is robust to patient demographics, it can be used to evaluate drug repositioning. Finally, this method can be used to guide indication expansion, i.e., the identification of new disease areas for testing the efficacy of specific drugs or therapies.
[0053] As described above, methods are provided for determining whether the genes of a gene expression signature that serve as features of a model behave consistently across datasets having different origins (e.g., different data-generating technology platforms, diseases, patient cohorts, etc.).
[0054] In some cases, gene expression data generated by two different technology platforms will be available for the same biological specimen. For example, certain cell line libraries (e.g., the Cancer Cell Line Encyclopedia (CCLE) by Broad / Novartis) have been profiled by both gene expression microarrays and RNA-seq analysis. Similarly, archival tumor biopsies previously analyzed by microarrays can be newly analyzed by RNA-seq (particularly The Cancer Genome Atlas (TCGA)). The challenge in applying gene signatures or predictive models derived from microarray data to newly generated RNA-seq data is determining whether the gene features are transferable across these technologies. Overcoming this challenge is essential to utilize potentially useful previously collected data or any data and analyses performed with previous, older-generation expression technologies. Given the rapid pace of change in "omics profiling," important datasets are at risk of becoming outdated every few years. They can be resurrected and advanced using the methods described herein to determine feature transferability.
[0055] With reference to FIG. 10, an exemplary method is provided for assessing the impact of technology platform and biological variation (e.g., disease type) on feature transferability given gene signatures and paired gene expression datasets generated by microarrays and RNA-Seq.
[0056] In 1001, the concordance between samples analyzed by different technology platforms is determined. For each sample pair, the Spearman correlation coefficient is calculated between the microarray and RNA-Seq expression of signature genes. Samples are sorted in descending order by Spearman correlation coefficient. For each sample pair, the Spearman correlation coefficient is plotted as a function of sample rank. Samples that have a concordance below a certain threshold can be excluded or examined individually to determine the source of variation. At this stage, all samples are processed together regardless of disease type.
[0057] The exemplary dataset includes a 170-gene signature and microarray and RNA-Seq data from 140 paired cell line samples from CCLE. These 140 sample pairs represent three different cancer types: 110 gastric cancers, 22 sarcomas, and 8 mesotheliomas.
[0058] Referring to Figure 11, a rank plot by sample number is provided based on the Spearman correlation coefficient between microarray and RNA-Seq TPM (Transcript Per Million normalization) expression. This includes all considered disease types: gastric cancer, sarcoma, and mesothelioma. This analysis reveals a relatively high concordance between microarray and RNA-Seq TPM expression for nearly all samples and all included disease types. The Spearman correlation coefficient for the samples was 0.05, consistent with industry standards, R S = roughly close to 0.8.
[0059] Upon visual inspection, one may consider removing samples below 0.75 because they deviate significantly from the rest, although it will be appreciated that various statistical methods may be used to determine the cutoff value, as discussed above.
[0060] At 1002, the genes that show the greatest concordance across all sample pairs are determined. For each gene, the Spearman correlation coefficient is calculated between the microarray and RNA-Seq expression of the paired samples. The genes are sorted in descending order by Spearman correlation coefficient. For each gene, the Spearman correlation coefficient is plotted as a function of gene rank.
[0061] Referring to Figure 12, a rank plot of 170 genes is provided using the Spearman correlation coefficient as a transferability metric between microarray and RNA-seq TPM expression. Each point represents a gene. Each correlation coefficient is calculated across all sample pairs (in this example, gastric cancer, sarcoma, and mesothelioma subjects (140 pairs in total)). The left Y-axis (large circle) corresponds to the Spearman correlation coefficient between microarray and RNA-seq TPM expression calculated across the above samples. The right Y-axis (small circle) corresponds to the median of raw RNA-seq counts + 1 calculated across the above samples.
[0062] The correlation for genes between microarray-derived expression and RNA-Seq decreases linearly for approximately the top 125 genes, then drops off rapidly. The genes with the lowest ranks have the greatest correlation (in this dataset, CXCL8 (R S =0.98). A threshold can be set on the left vertical axis where the linear slope changes (to superlinear or exponential decay). In the example above, this inflection point is approximately R S = 0.60, and therefore all genes with a rank above about 125 can be removed from the analysis.
[0063] The correlation between microarray and RNA-Seq TPM expression can be partially explained by the expression level of the genes. Most poorly expressed genes with median raw RNA-Seq counts below 10 had a correlation R S On the other hand, the expression of genes with median raw counts above 100 often correlates well between microarray and RNA-Seq (R S>0.6). This overlay therefore allows the determination of a minimum gene expression threshold below which specific genes can be removed.
[0064] At 1003, the contribution of biological factors (rather than technology platform) to gene / sample rank is determined. For each gene, the Spearman correlation coefficient is calculated between microarray and RNA-Seq expression of paired samples, separately for each disease. In this example, the diseases included are gastric cancer, sarcoma, and mesothelioma. Genes are sorted in descending order by Spearman correlation coefficient. For each gene, the Spearman correlation coefficient is plotted as a function of the gene rank for the disease type that the majority of samples have (in this case, gastric cancer is the most prevalent type).
[0065] Referring to Figure 13A, a gene rank plot is provided using the Spearman correlation coefficient as a transferability metric between microarray and RNA-seq TPM expression. Each point represents a gene. Each correlation coefficient is calculated pairwise separately across samples from each biological condition or disease (in this example, gastric cancer, sarcoma, and mesothelioma).
[0066] The Spearman correlation coefficient calculation is repeated using gene ranks based on all disease types rather than the most common.
[0067] Referring to Figure 13B, another plot is provided in which genes are ranked on the X-axis based on correlation across subjects in all three indications rather than simply being the most common.
[0068] The scatter in Figure 13B compared to Figure 13A indicates that the degree of variation in agreement is driven by biological state, an important observation if the goal of gene signature development is to generate a versatile feature set that can function across states, for example, as a pan-cancer diagnostic gene panel.
[0069] At 1004, agreement between correlation coefficients is examined across disease indications. For each gene, the Spearman correlation coefficient is calculated between the microarray and RNA-Seq expression of paired samples, as in step 1003. The correlation coefficients for samples representing states (B, C, . . . Z) are plotted as a function of the correlation coefficient for state A. In this example, B = sarcoma, C = mesothelioma, and A = gastric cancer. If one of these states is clearly the most common, it can serve as the independent variable. If the states are more evenly distributed, the analysis should be repeated, alternating which state serves as the independent variable.
[0070] Referring to Figure 14, the Spearman correlation coefficient between the expression of microarray and RNA-Seq TPMs for states B and C (sarcoma and mesothelioma) is shown as a function of the same correlation coefficient for state A (gastric cancer). Each point corresponds to a gene.
[0071] Genes that are consistently most highly correlated between sample pairs cluster in the upper right. The box drawn at (X, Y = 0.6, 0.6) gates informative features across biological states (e.g., disease). This analysis validates the thresholding approach of step 1002.
[0072] In some embodiments, the consistently most highly correlated genes (or other molecular biomarkers) among the input signatures are retained to obtain a transferable signature at 1005. However, the above-described matching methods may be combined with the above-described transferability statistic (KS) methods. For example, a transferability statistic may be calculated at 1006 for each of the highly correlated biomarkers determined at 1005. Alternatively, signatures using each method may be calculated in parallel at 1005, 1006, and then combined into an aggregate signature at 1007. The aggregate signature may be determined by taking the sum or product of the two input signatures.
[0073] The expression of each gene across all samples is quantile transformed to a uniform distribution. For each gene, the Kolmogorov-Smirnov test statistic is calculated across all sample pairs for all biological conditions (e.g., gastric cancer, sarcoma, and mesothelioma) using the distribution of quantile-normalized expression. Genes are sorted in ascending order by the Kolmogorov-Smirnov statistic. For each gene and disease indication combination, the Kolmogorov-Smirnov statistic is plotted as a function of gene rank.
[0074] Referring to Figure 15A, a plot of the Kolmogorov-Smirnov statistic by gene rank is provided, showing the transferability of expression distribution by gene between the AB, AC, and BC (gastric cancer, sarcoma, and mesothelioma) subsets of samples.
[0075] The best gene transferability is consistently achieved between AB (gastric cancer and sarcoma). Transferability between AC (gastric cancer and mesothelioma) is similar to transferability between BC (sarcoma and mesothelioma). The KS statistic as a function of gene rank is approximately linear. The value of the KS statistic increases fairly rapidly and enters a region where transferability is questionable at best (in this example, KS>0.5). As mentioned above, rather than setting a cutoff at an inflection point, it can be set based on a predetermined or empirical transferability statistic. In addition, it will be understood that the KS statistic can be converted to a p-value or other probability value to set a threshold.
[0076] Referring to Figure 15B, a plot of the Kolmogorov-Smirnov statistic by gene rank is provided, showing the transferability of gene expression distributions between the AB, AC, and BC (gastric cancer, sarcoma, and mesothelioma) subsets of samples for the expanded input gene set. Cross-disease transferability between gastric cancer and sarcoma is observed / confirmed with this expanded feature set.
[0077] Evidence of the utility of quantile normalization is provided with reference to Figures 16A-B. In these examples, the same KS-rank method is applied as described above for the disease comparison of AB (gastric cancer vs. sarcoma). Three expression preprocessing methods are compared: TPM normalization, z-score (TPM+1), and quantile transformation of TPM-normalized expression.
[0078] FIG. 16A shows the transferability of gene-by-gene expression distribution between gastric cancer and sarcoma for the three expression pretreatment methods.
[0079] Figure 16B shows the transferability of gene-by-gene expression distribution between gastric cancer and sarcoma for the three expression preprocessing methods using the expanded feature set.
[0080] The quantile transformation (1603) shows superior performance, followed by the z-score (1602) and no preprocessing (1601). The above results are reproducible across all pairwise state comparisons.
[0081] An additional utility of this method is to estimate transferability between samples of different diseases based on their therapeutic phenotype. For example, one can ask whether genes predicting drug sensitivity are more transferable than genes predicting drug resistance. Thus, input samples are stratified by phenotypic marker, and transferability statistics are calculated between two conditions (hereafter, gastric cancer and sarcoma) as described above.
[0082] Referring to Figure 17, a graph is provided showing the transferability of expression distribution by gene between gastric cancer and sarcoma, separately for each response group of samples.
[0083] The observation that genes (traits) are more transferable to cell lines with a "resistant" phenotype suggests that the biological pathways involved in drug resistance are conserved between disease states (gastric cancer vs. sarcoma), whereas the biological pathways contributing to drug sensitivity are more heterogeneous.
[0084] Thus, feature transferability methods allow estimation of which drug response phenotypes are most reliably predicted from a given set of features.
[0085] As mentioned above, the feature transferability methods provided herein are broadly applicable. Some additional examples are provided below.
[0086] Transferability across data generation platforms In the first example, transferability between microarray and RNA-Seq platforms from separate patient subpopulations at different time points with different treatment histories is assessed.
[0087] The dataset used in this example was: 1) ACRG (Asian Cancer Research Group) Gastric cancer subjects (N=300) were second-line or later-line patients who had received prior chemotherapy and / or irradiation. Affymetrix microarrays; GEO GSE62254, GSE62717; Cristescu et al. 2015 2)TCGA(The Cancer Genome Atlas) Gastric cancer patients (N=388) were a mix of multiple treatment options RNA-Seq; portal.gdc.cancer.gov data; Cancer Genome Atlas Research Network 2014 3) Singapore cohort Gastric cancer patients (N=192) were a mix of multiple treatment options Affymetrix microarray platform; GEI (GSE15459); Lei et al. 2013
[0088] Referring to Figure 18, a plot of the KS statistic versus gene rank is provided. We calculated the KS statistic for the 125 signature genes. When sorted by rank, an initial increase in the slope of the KS statistic at rank 98 can be observed. Therefore, the remaining 27 genes are considered non-transferable and can be removed from the model.
[0089] Data platform, transferability across disease histologies In this example, the transferability between ovarian / gynecological cancer and anti-VEGF datasets is evaluated using the following platforms: exome RNA-Seq and total RNA-Seq; histology types: ovarian / gynecological cancer and gastric cancer.
[0090] The dataset used in this example was: 1) Original clinical trials (anti-VEGF / DLL4 therapy, ovarian cancer and gynecological cancer) A single-arm phase 1b study of 4+-selected platinum-resistant ovarian cancer patients treated with anti-VEGF / anti-DLL4 dual specific plus paclitaxel RNA-Seq (subset N=30); unpublished data 2) ACRG (Asian Cancer Research Group) Gastric cancer subjects (N=300) were second-line or later and had received prior chemotherapy and / or irradiation. Affymetrix microarrays; GEO GSE62254, GSE62717; Cristescu et al. 2015 3) Unique gastric VEGF Gastric and GEJ cancer subjects, mixed prior treatment history, 100% Asian demographic Treatment with anti-VEGF ramucirumab RNA-Seq (N=48); unpublished data 4) ICON7 Ovarian cancer target Treated with chemotherapy + bevacizumab (anti-VEGF) Microarray (N=380); GEO accession number GSE140082)
[0091] Referring to Figure 19, a plot of the KS statistic against gene rank is provided. We calculated the KS statistic for 160 signature genes (98 genes from above and 62 genes from another signature). When sorted by rank, an initial increase in the slope of the KS statistic at rank 136 can be observed. Therefore, the remaining 26 genes are considered "non-transferable" and can be removed from the model. Figure 20 similarly shows thresholds (e.g., located at inflection points) for the transferability statistic.
[0092] 21, a schematic diagram of an example computational node is shown. Computational node 10 is merely one example of a suitable computational node and is not intended to suggest any limitation as to the scope of use or functionality of the embodiments described herein. Regardless, computational node 10 may implement and / or perform any of the functionality previously described herein.
[0093] Compute node 10 includes computer system / server 12 that is operable in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0094] The computer system / server 12 may be described in the general context of computer system-executable instructions, such as program modules, being executed by the computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system / server 12 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including memory storage devices.
[0095] 21, the computer system / server 12 in the compute node 10 is shown in the form of a general-purpose computing device. Components of the computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 connecting various system components including the system memory 28 to the processors 16.
[0096] Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and without limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe), and an Advanced Microcontroller Bus Architecture (AMBA).
[0097] Computer system / server 12 typically includes a variety of computer system-readable media, which can be any available media that can be accessed by computer system / server 12 and includes both volatile and nonvolatile media, removable and non-removable media.
[0098] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may provide for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, typically referred to as a "hard drive"). Also, not shown, may be a magnetic disk drive that reads from or writes to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive that reads from or writes to optical media, such as a CD-ROM, DVD-ROM, or the like. In such cases, each may be connected to the bus 18 by one or more data media interfaces. As further depicted and described below, the memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present disclosure.
[0099] A program / utility 40 having a set (at least one) of program modules 42, as well as, by way of example and not limitation, an operating system, one or more application programs, other program modules, and program data, may be stored in memory 28. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may comprise an implementation of a network environment. The program modules 42 typically perform the functions and / or methods of the embodiments as described herein.
[0100] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, display 24; one or more devices that allow a user to interact with the computer system / server 12; and / or any device (e.g., a network card, modem, etc.) that allows the computer system / server 12 to communicate with one or more other computer devices. Such communication may occur via an input / output (I / O) interface 22. Furthermore, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. Although not shown, it should be understood that other hardware and / or software components may be used in combination with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data archive storage systems, etc.
[0101] The present disclosure may be embodied as a system, method, and / or computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present disclosure.
[0102] A computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination thereof. As used herein, computer-readable storage media should not be construed as ephemeral signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals passing through electrical wires.
[0103] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium into each computer / processing device, or can be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.
[0104] The computer-readable program instructions for carrying out the operations of the present disclosure may be either source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, may execute as part of the user's computer as a standalone software package, may execute as part of the user's computer and on a remote computer, or may execute entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), and the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may utilize state information of computer-readable program instructions to personalize the electronic circuitry and execute the computer-readable program instructions to carry out aspects of the present disclosure.
[0105] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0106] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, whereby the instructions, executed by the processor of the computer or other programmable data processing apparatus, create an apparatus that creates means for performing the functions / acts specified in the flowchart and / or block diagram blocks or blocks. These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other apparatus to function in a particular manner, whereby a computer-readable storage medium having instructions stored therein includes an article of manufacture containing instructions that implement aspects of the functions / acts specified in the flowchart and / or block diagram blocks.
[0107] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to create a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device may perform the functions / acts specified in the flowchart and / or block diagram blocks.
[0108] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for performing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently or may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified function or function or executes a combination of special-purpose hardware and computer instructions.
[0109] The descriptions of various embodiments of the present disclosure are provided for illustrative purposes and are not intended to be exhaustive or limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best describe the principles of the embodiments, technical improvements to practical applications or technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for determining a signature transferable between a plurality of datasets based on a plurality of signatures associating biomarkers with a plurality of output classifications, each of the plurality of datasets including an expression value associated with each molecular biomarker of a plurality of said molecular biomarkers, the method comprising: reading a first signature of the plurality of signatures, the first signature associating a first plurality of molecular biomarkers with a first output classification of the plurality of output classifications; for each of the plurality of datasets and each of the plurality of output classifications, normalizing the expression value of each of the first plurality of molecular biomarkers associated with that output classification to obtain a plurality of normalized expression values, wherein each of the plurality of normalized expression values is associated with one of the first plurality of molecular biomarkers, one of the plurality of output classifications, and one of the plurality of datasets; for each of the first plurality of molecular biomarkers, performing pairwise comparisons between the normalized expression values associated with that molecular biomarker, each pairwise comparison being between the first and second normalized expression values, the first and second normalized expression values being associated with a same output classification of the plurality of output classifications and different datasets of the plurality of datasets, thereby determining a transferability score for each of the plurality of molecular biomarkers; ranking the first plurality of molecular biomarkers based on a transferability score for each molecular biomarker; generating a second plurality of molecular biomarkers from the first plurality of molecular biomarkers by applying a transferability score threshold to the first plurality of molecular biomarkers; and providing a transferable signature, the transferable signature relating the second plurality of molecular biomarkers to the first output classification of the plurality of output classifications; A method comprising:
2. 10. The method of claim 1, wherein each of the first plurality of molecular biomarkers is a gene.
3. 10. The method of claim 1, wherein each of the first plurality of molecular biomarkers is a protein.
4. The method of claim 1 , wherein each signature includes a mapping function.
5. The method of claim 1 , wherein each signature comprises a plurality of synaptic weights.
6. The method of claim 1 , wherein each output classification comprises a phenotype.
7. The method of claim 6 , wherein the phenotype is a disease phenotype.
8. The method of claim 1 , wherein the normalization comprises quantile normalization.
9. The method of claim 1 , wherein the normalization is relative to a predetermined reference distribution.
10. The method of claim 1 , wherein performing the pairwise comparisons comprises calculating a Kolmogorov-Smirnov statistic.
11. The method of claim 1 , wherein determining the transferability score comprises calculating an average of the pairwise comparisons.
12. The method of claim 1 , wherein the plurality of datasets includes at least one dataset from each of a plurality of platform technologies.
13. The method of claim 12 , wherein the platform technologies include microarrays and RNA sequencing.
14. 13. The method of claim 12, wherein the platform technology comprises mass spectrometry, ELISA, antibody array, peptide fingerprinting, and / or protein barcoding.
15. 13. The method of claim 12, wherein each of the plurality of datasets is derived from the same biological sample.
16. 16. A system including a computing node including a computer readable storage medium having program instructions embedded therein, the program instructions being executable by a processor of the computing node to cause the processor to perform a method including the method of any one of claims 1 to 15.
17. 16. A computer program product for determining transferable molecular biomarker signatures, said computer program product comprising a computer readable storage medium having program instructions embodied therein, said program instructions executable by a processor to cause the processor to perform a method comprising the method of any one of claims 1 to 15.
18. A method for determining a signature transferable between a plurality of datasets based on a plurality of signatures associating biomarkers with a plurality of output classifications, each of the plurality of datasets including an expression value associated with each molecular biomarker of a plurality of said molecular biomarkers, the method comprising: reading a first signature of the plurality of signatures, the first signature associating a first plurality of molecular biomarkers with a first output classification of the plurality of output classifications; For each dataset pair of the plurality of datasets, each dataset pair is from a different platform technology and each dataset pair is from the same biological sample, determining a correlation coefficient for each of the first plurality of molecular biomarkers between said dataset pair; determining, for each output classification of the plurality of output classifications, a class-specific correlation coefficient for each of the first plurality of molecular biomarkers between the pair of datasets; ranking the first plurality of molecular biomarkers based on each of the correlation coefficients and class-specific correlation coefficients; generating a second plurality of molecular biomarkers from the first plurality of molecular biomarkers by applying a rank threshold to the first plurality of molecular biomarkers; and providing a transferable signature, the transferable signature relating the second plurality of molecular biomarkers to the first output classification of the plurality of output classifications; A method comprising:
19. 20. The method of claim 18, wherein each of the first plurality of molecular biomarkers is a gene.
20. 20. The method of claim 18, wherein each of the first plurality of molecular biomarkers is a protein.
21. The method of claim 18 , wherein each signature includes a mapping function.
22. The method of claim 18 , wherein each signature comprises a plurality of synaptic weights.
23. 20. The method of claim 18, wherein each output classification comprises a phenotype.
24. 24. The method of claim 23, wherein the phenotype is a disease phenotype.
25. 20. The method of claim 18, wherein the platform technologies include microarrays and RNA sequencing.
26. 20. The method of claim 18, wherein the platform technology comprises mass spectrometry, ELISA, antibody array, peptide fingerprinting, and / or protein barcoding.
27. 27. A system including a computing node including a computer readable storage medium having program instructions embedded therein, the program instructions being executable by a processor of the computing node to cause the processor to perform a method including the method of any one of claims 18 to 26.
28. 27. A computer program product for determining transferable molecular biomarker signatures, said computer program product comprising a computer readable storage medium having program instructions embodied therein, said program instructions executable by a processor to cause the processor to perform a method comprising the method of any one of claims 18 to 26.
29. Determining a first transferable signature according to the method of claim 1; Determining a second transferable signature according to the method of claim 18; determining a third reusable signature by multiplying the first and second reusable signatures; A method comprising:
30. Determining a first transferable signature according to the method of claim 1; Determining a second transferable signature according to the method of claim 18; determining a third reusable signature by summing the first and second reusable signatures; A method comprising:
31. Determining a first transferable signature according to the method of claim 1; applying the method of claim 18 to the first reusable signature to determine a second reusable signature; A method comprising:
32. Determining a first transferable signature according to the method of claim 18; applying the method of claim 1 to the first reusable signature to determine a second reusable signature; A method comprising:
Citation Information
Patent Citations
GENE EXPRESSION DATA MANAGEMENT SYSTEM AND METHOD
JP2004535612A
Systems and methods related to network-based biomarker signatures
JP2015525412A
Predictive gene signature for metastatic disease
JP2018527886A
Methods and machine learning systems for predicting the likelihood or risk of having cancer
US20180068083A1
Disease development determination device, disease development determination method, and disease development determination program
WO2018079840A1