Systems and methods for development of proteomic models

EP4802514A1Pending Publication Date: 2026-09-09SOMALOGIC OPERATING CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024808775
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-10-28
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

The development of predictive proteomic models is challenging due to the assumption of independent contributions of proteins to diagnostic or prognostic results, and the complexity of co-expression of multiple proteins.

Method used

The use of non-negative factorization to generate proteomics-based diagnostic or prognostic models, which involves obtaining a training protein expression dataset, generating sample and topic definition datasets, and applying these datasets to develop a predictive model.

Benefits of technology

This approach allows for the identification of characteristic proteins and topics that are biologically significant, leading to improved performance of predictive models in diagnostics and prognostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000037_0001
    Figure IMGF000037_0001
  • Figure IMGF000039_0001
    Figure IMGF000039_0001
  • Figure 00000071_0000
    Figure 00000071_0000
Patent Text Reader

Abstract

Methods, systems, and computer-readable media can use topic modeling in developing predictive proteomics models. Topic modeling can be used to generate topics that can be included as input features in a predictive model. The biological significance of topics can be identified using a sample definition dataset that permits the association of topics with health information, or using a topic definition dataset the permits the association of topics with proteins (or sets of proteins) having known biological significance. The contribution of proteins to topics can be evaluated using tests that can robustly detect proteins that contribute to multiple topics, while being computationally efficient and avoiding false correlations. When developing a model, topics can be selected for inclusion (or potential inclusion) based on assessed biological significance. Similarly, proteins can be selected for inclusion (or potential inclusion) based on the assessed biological significance of the topics to which they contribute.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 15988.0009-00304 SYSTEMS AND METHODS FOR DEVELOPMENT OF PROTEOMIC MODELS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present disclosure claims priority to and the benefits of priority to U.S. Provisional Patent Application No.63 / 594,334, filed on October 30, 2023. The provisional application is incorporated herein by reference in its entirety. BACKGROUND

[0002] Aptamer-based assays can simultaneously measure protein levels for thousands of proteins in samples drawn from patients. Predictive proteomic models can use measured protein levels for diagnostic or prognostic purposes. Such models may assume that the contribution of each protein to a diagnostic or prognostic result is independent from the contribution of other proteins. However, the co-expression of multiple proteins may have diagnostic or prognostic value. Furthermore, the selection of proteins for inclusion in predictive proteomic models can be challenging. SUMMARY

[0003] Certain embodiments of the present disclosure relate to development of proteomics models using non-negative factorization.

[0004] The disclosed embodiments include a proteomics-based diagnostic or prognostic method. The method can include obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples. The training protein expression dataset can include, for the set of training samples, training protein expression values for a set of proteins. The method can further include generating a sample definition dataset and a topic definition dataset using the training protein expression dataset. The sample definition datasetAttorney Docket No.: 15988.0009-00304 can include topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples. The topic definition dataset can include protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics. The method can further include obtaining health information corresponding to the training samples. The method can further include generating a diagnostic or prognostic model using the sample definition dataset and the health information. The method can further include obtaining patient protein expression values for the set of proteins. The patient protein expression values can be generated using the aptamer-based assay. The method can further include generating patient topic weights using the topic definition dataset and the patient protein expression values. The method can further include generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.

[0005] The disclosed embodiments further include another method. The method can include obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples. The training protein expression dataset can include, for the set of training samples, training protein expression values for a set of proteins. The method can further include generating a sample definition dataset and a topic definition dataset using the training protein expression dataset. The sample definition dataset can include topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples. The topic definition dataset can include protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics. The method can further include identifying a set of characteristic proteins for a first topic of the set of topics. The identification can include determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics. The identification can further include determining, using the topic definition dataset, protein value differences. The protein value differences can be based on the median protein value for the first protein andAttorney Docket No.: 15988.0009-00304 the protein weights for the first protein across the set of topics. The identification can further include determining a median of the protein value differences and including the first protein in the set of characteristic proteins for the first topic. The first protein in the set of characteristic proteins for the first topic can be included based on the median of the protein value differences and a protein value of the first protein for the first topic. The method can further include generating a diagnostic or prognostic model using the identified set of characteristic proteins.

[0006] The disclosed embodiments can further systems including processor(s) and non- transitory computer readable media containing instructions that, when executed by the processor(s) of the systems, cause the systems to perform the above-disclosed methods.

[0007] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims. Other systems, methods, and computer-readable media are also discussed within. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the disclosed principles. In the drawings:

[0009] FIG.1 depicts an example pipeline for developing, validating, and deploying proteomics models for predicting health information using aptamer-based blood tests, consistent with disclosed embodiments.

[0010] FIGs.2A to 2H provide a high-level depiction of stages in an example aptamer-based serum or plasma assay, consistent with disclosed embodiments.

[0011] FIG.3 depicts an example process for generating and applying a proteomics model, consistent with disclosed embodiments.Attorney Docket No.: 15988.0009-00304

[0012] FIGs 4A to 4D depict an assessment of a protein expression dataset for a cancer cell line study, consistent with disclosed embodiments.

[0013] FIGs.5A to 5B concern evaluation of the biological significance of topics identified in the protein expression dataset obtained in the cancer cell line study, consistent with disclosed embodiments.

[0014] FIGs.5C to 5E concern evaluation of the biological significance of topics identified in a training protein expression dataset obtained from nonalcoholic fatty liver disease samples, consistent with disclosed embodiments. DETAILED DESCRIPTION

[0015] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed example embodiments. However, it will be understood by those skilled in the art that the principles of the example embodiments may be practiced without every specific detail. Well-known methods, procedures, and components have not been described in detail so as not to obscure the principles of the example embodiments. Unless explicitly stated, the example methods and processes described herein are neither constrained to a particular order or sequence nor constrained to a particular system configuration. Additionally, some of the described embodiments or elements thereof can occur or be performed simultaneously, at the same point in time, or concurrently. Reference will now be made in detail to the disclosed embodiments, examples of which are illustrated in the accompanying drawings.

[0016] Proteomics can concern or involve the quantitative assessment of proteins present in a sample obtained from a human or non-human animal. Proteomics can concern classifications or predictions drawn from large-scale (e.g., concerning hundreds to thousands) measurementAttorney Docket No.: 15988.0009-00304 of protein levels (e.g., presence, amount, concentration, or the like). In this manner, proteomics can be similar to genomics, but applied to the realm of proteins.

[0017] A predictive model can be model that maps a set of inputs to an output value or category. In some embodiments, the predictive models can be machine-learning models, which can be trained using training data to generate appropriate outputs. The predictive models can be trained using supervised methods, semi-supervised methods, unsupervised methods, or reinforcement learning methods. In some embodiments, the predictive models can be statistical models that estimate a relationship, using training data, between independent variables (e.g., protein levels in a biological sample) and dependent variables (e.g., health information values). Exemplary predictive models include regression models (e.g., penalized regression models), support vector machines, decision trees or forests, linear discriminant analysis, clustering models, nearest neighbor models, neural network models, ensemble models including one or more of any of the foregoing models, or the like.

[0018] A proteomics model can be a predictive model configured to receive protein level measurements and output classifications or predictions based at least in part on those measurements.

[0019] Health information can include a probability of a health outcome (e.g., cardiovascular disorder, dementia, kidney disease, or the like) occurring within a particular time frame (e.g., within six months, a year, four years, a decade, two decades, or longer), whether a health outcome occurred within a particular time frame, an indication of health status (e.g., a diagnosis of a disease, disorder, or injury; percentage body fat; basal metabolic rate; lean muscle mass; aerobic fitness; visceral fat content; glomerular filtration rate; presence or absence of excess liver fat; glucose tolerance; or the like), a behavioral health prediction (e.g., predicted weekly alcohol consumption, or the like), or the like.Attorney Docket No.: 15988.0009-00304

[0020] Consistent with disclosed embodiments, providing models, data, or instructions can include direct (e.g., between source and target) or indirect (e.g., through intermediaries) provision of such models, data, or instructions. Providing a model, data, or instructions can include providing by value the model, data, or instructions or providing by reference to a location storing the model, data, or instructions.

[0021] Consistent with disclosed embodiments, performance measures can be used to quantify the performance of a predictive model. In some embodiments, performance measures can quantify the agreement between two predictions (e.g., to evaluate reproducibility or the effect of differences in input data on the predicted output). Such performance measures can include Lin’s concordance correlation coefficient, Pearson’s correlation coefficient, or other suitable correlation measures. In some embodiments, performance measures can quantify the agreement between predictions and a ground truth. In some embodiments, such performance measures can quantify this agreement in terms of a relationship between true positive rate and false positive rate, as parameters of the predictive model are varied (e.g., a discrimination threshold, or the like). Such measures can include a receiver operating characteristic (ROC) curve, or derivative performance measures, such as area under the ROC curve (AUC). In some embodiments, such performance measures can include or depend upon numbers of true positives, true negative, false positives, or false negatives for a classification task. In some embodiments, such measures can include sensitivity, specificity, precision, false negative rate, false positive rage, prevalence, accuracy, F1 score, or other such measures. In some embodiments, such performance measures can include or depend upon a per-class comparison between predictions and a ground truth, such as a confusion matrix or the like.

[0022] A proteomics dataset can include many (e.g., tens to thousands) of observations. Each observation can include hundreds to thousands of protein-level measurements. CausalAttorney Docket No.: 15988.0009-00304 relationships may be difficult to detect in such a high-dimensional dataset. Furthermore, the size of the proteomics dataset may hinder analysis. Accordingly, dimensionality reduction techniques can be used to express the proteomics dataset in a lower dimensional space.

[0023] Suitable dimensionality reduction techniques can include topic modeling. In the context of a proteomics dataset, a topic model can express the biological relationships, pathways, or structures that cause the observed protein concentration measurements in terms of topics. As may be appreciated, biological pathways may be partially overlapping or otherwise non-orthogonal. Topic modeling can support non-orthogonal topics, unlike other dimensionality reduction techniques such as principal component analysis. The topics defining a lower dimensional space generated by topic modeling may therefore better represent the biological relationships, pathways, or structures that cause the observed protein concentration measurements in the original, high-dimensional proteomics dataset.

[0024] A proteomics dataset can be divided into a topic definition matrix that expresses topics in terms of protein weights and a sample definition matrix that expresses samples in terms of topic weights. In some embodiments, the proteomics dataset can be divided using non-negative factorization (e.g., using the sklearn.decomposition.nmf class, or another suitable approach).

[0025] In some embodiments the sample definition matrix can be used to generate a predictive model. The information can be associated with the samples. The predictive model can be trained to predict the health information for a sample given the topic weights for the sample. In such embodiments, the biological significance of each topic may not be identified. The topics can represent complicated patterns of co-occurrence of proteins. Such co- occurrences may be more biologically predictive than expression levels of individual proteins. Accordingly, proteomics models that use topics as input features may exhibit improved performance (e.g., according to some relevant performance measure), as comparedAttorney Docket No.: 15988.0009-00304 to proteomics models that use individual protein expression levels. In this manner, the disclosed embodiments provide an improved way to generate predictive models.

[0026] In some embodiments, the topic definition matrix can be used to identify the biological, diagnostic, or prognostic significance of a topic. Proteins contributing to a topic can be identified using the topic definition matrix. These proteins may have known biological functions. For example, the proteins may be known growth factors (e.g., VEGF-D, epidermal growth factor-like protein 6, fibroblast growth factor 22, fibroblast growth factor 6, or the like), or have a known association with particular tissue type (e.g., neural tissue). Similarly, the proteins may have known associations with a diagnosis or prognosis. For example, the corresponding gene expression levels of proteins captured by a topic may be associated with neuroblastoma severity (e.g., neuromodulin, DOPA decarboxylase, dihydropyrimidinase- related protein 1, or the like). The biological, diagnostic, or prognostic significance of the topic can be inferred or assessed from the biological, diagnostic, or prognostic significance of the proteins contributing to the topic.

[0027] In some embodiments, topics can be used to select proteins for inclusion in a proteomics model. The biological, diagnostic, or prognostic significance of a topic can be determined using some proteins that contribute to the topic. Another protein that contributes to the model can then be selected for inclusion in a proteomics model based on its contribution to the topic. For example, a topic can be identified as generally diagnostic of cancer in a sample based on the contribution of several cancer-associated proteins to the topic. A cancer diagnostic model can be constructed using, at least as initial input features, expression levels of all the proteins that contribute to the topic. Thus proteins that might not have been initially considered diagnostic of cancer can be included in a cancer diagnosis model.Attorney Docket No.: 15988.0009-00304

[0028] A protein can be identified as contributing to a topic based on a score. The score can be generated using protein weights in the topic definition matrix. In some embodiments, the score can be or depend on a statistical distance, such as a Kullback–Leibler divergence. For example, a protein can be ranked within a topic by its minimum KL divergence (e.g., maximizing the minimum KL divergence of a cluster, such as a topic, against all other clusters, such as other topics, while assuming an underlying Poisson model for KL divergence measurement). Generating scores using the minimum KL divergence can be computationally efficient (e.g., requiring less than 4 seconds to process 207 samples, each sample including about seven thousand protein measurements). Furthermore, the generated scores can be robust against topic misattribution, which can occur when two topics are correlated in membership across samples. However, this method may underestimate the contribution of a protein to a topic when that protein contributes to multiple topics.

[0029] In some embodiments, a score can be or depend on weighted differential protein expressions (e.g., fold-changes of proteins, or the like) using an underlying Poisson distribution. Such a method may avoid underestimating the contribution of a protein to a topic when that protein contributes to multiple topics. However, such a method may be computationally inefficient (e.g., such a method may require approximately 14 minutes to process 207 samples, each sample including about seven thousand protein measurements). Furthermore, when two topics are correlated in membership across samples, a protein contributing to one of the two topics may be incorrectly scored as contributing to both topics.

[0030] Consistent with disclosed embodiments, a score can be calculated using median protein expression values. Generation of such a score can be computationally efficient (e.g., requiring less than 4 seconds to process 207 samples, each sample including about seven thousand protein measurements). Furthermore, the generated scores can be robust against topic misattribution and may avoid underestimating the contribution of a protein to a topicAttorney Docket No.: 15988.0009-00304 when that protein contributes to multiple topics. The envisioned embodiments can therefore provide an improved way of determining the contribution of proteins to topics.

[0031] FIG.1 depicts an exemplary pipeline 100 for developing, validating, and deploying proteomics models for predicting health information using aptamer-based blood tests, consistent with disclosed embodiments. Pipeline 100 can include components from which data is originally obtained, such as measurement system 101 or records 103. Pipeline 100 can include components, such as data input engine 110 or dataset generator 120, for the collection and preparation of the obtained data. Pipeline 100 can include components, such as data storage 140 and model storage 130, for the storage of prepared data and machine learning models. Pipeline 100 can include training engine 160 for generating trained machine learning models using obtained data and models. Pipeline 100 can include prediction engine 170 for using trained machine learning models to predict health information. Components of pipeline 100 can be managed and configured through a user device 180. User device 180 can also be used to display outputs of other components, original or prepared data, trained or untrained models, or predicted health information. A user can interact with components of pipeline 100 to perform model validation and development, consistent with disclosed embodiments. In some embodiments, a user can interact with training engine 160 (or another suitable component of pipeline 100) to perform model development as described with regards to FIG. 4. Likewise, in some embodiments, a user can interact with prediction engine 170 (or another suitable component of pipeline 100) to apply a proteomics model (e.g., a proteomics model developed using training engine 160, or received from another system) as described with regards to FIG.4. Overall, pipeline 100 can provide a convenient, scalable, platform for developing, validating, and deploying proteomics models, as disclosed herein.

[0032] Measurement system 101 can be a device suitable for obtaining data indicating protein presence or concentration in a biological sample. For convenience of discussion,Attorney Docket No.: 15988.0009-00304 measurement system 101 is described herein as a microarray scanner. Such a scanner can be configured to measure fluorescence at predetermined locations corresponding to different proteins on a microarray. The intensity of the fluorescence can indicate a concentration of the protein in a biological sample used to prepare the microarray. In some instances, multiple locations can correspond to the same protein. Fluorescence intensity data for these multiple locations can be combined to better estimate the concentration of the protein in the sample. As may be appreciated, pipeline 100 is not limited to embodiments in which measurement system 101 is a microarray scanner. Furthermore, in some embodiments, rather than providing data directly to data input engine 110, measurement system 101 can provide data to record(s) 103. Data input engine 110 can then obtain this data from record(s) 103.

[0033] Consistent with disclosed embodiments, record(s) 103 can include one or more storage locations for data usable by pipeline 100 to predict healthcare outcomes. In some embodiments, such data can include fluorescence intensity data generated by measurement system 101 from biological samples. In various embodiments such data can include medical record information from the patients from which the biological samples were obtained. Such medical record information can include medical records, case notes, request or requisition information (e.g., pertaining to the sample or to the prediction to be performed by pipeline 100) provided by a physician or other clinician. In some embodiments, the medical record information can include class or data label information corresponding to the samples (e.g., for use in generating training datasets).

[0034] Consistent with disclosed embodiments, the medical record information can be associated with health information. For example, in a cancer screening or diagnosis setting, the medical record information can include data associated with a cancer type, a cancer prevalence, a cancer prognosis, or a cancer stage.Attorney Docket No.: 15988.0009-00304

[0035] Consistent with disclosed embodiments, the medical record information can include information suitable for use in generating personalized predictive models. For example, the medical record information can indicate patient demographic or health characteristics, such as age, gender, race / ethnicity, height / weight of the patient. As an additional example, the medical record information can indicate patient medical history, such as a history or medical treatment, clinical visits, or surgical history. As an additional example, the medical record information can indicate behavioral factors that could affect patient health, such as smoking history, alcohol consumption, drug use, diet type, etc. As an additional example, the medical record information can include data obtained from a biological sampling from the patient, such as data relating to a sample of blood, plasma, serum, or urine. As an additional example, the medical record information can indicate family history data, genetic data, or immunological data.

[0036] Consistent with disclosed embodiments, data input engine 110 can be configured to retrieve data from a variety of data sources (e.g., measurement system 101, record(s) 103, or other suitable sources) and process the data for use by other components of pipeline 100. In some embodiments, data input engine 110 can include data extractor 111, data transformer 113, and data loader 115.

[0037] Consistent with disclosed embodiments, data extractor 111 can receive or retrieve data from measurement system 101, record(s) 103, or other suitable sources. As described herein, measurement system 101 can be a diagnostic system configured to obtain fluorescence values corresponding to protein presence or concentrations in a biological sample. Similarly, record(s) 103 can be one or more databases or data storage locations containing data concerning the biological samples used to generate the fluorescence values. The disclosed embodiments are not limited to any particular format of the obtained data, or method for obtaining this data. For example, the obtained data can be or include structuredAttorney Docket No.: 15988.0009-00304 data or unstructured data. Data extractor 111 can interact with the various data sources, receive or retrieve the relevant data, and provide that data to data transformer 113.

[0038] Consistent with disclosed embodiments, data transformer 113 can receive data from data extractor 111 and process the data into standard formats. In some embodiments, data transformer 113 can normalize data such as dates or numerical values based on specific units of measure. For example, measurement system 101 can store dates in day-month-year format, while records 103 can include records that store dates in year-month-day format, or various records that store measurement in differing units (e.g., body weights in kilograms or pounds, fluorescence in absolute magnitude or log10 magnitude, or the like). In this example, data transformer 113 can modify the data provided through data extractor 111 into a consistent date format or standardized unit format, respectively. Accordingly, data transformer 113 can effectively clean the data provided through data extractor 111 so that all of the data, although originating from a variety of different sources, has a consistent format.

[0039] Moreover, data transformer 113 can extract additional data points from the data. For example, data transformer 113 can process a date in year-month-day format by extracting separate data fields for the year, the month, and the day. Data transformer 113 can also perform other linear and non-linear transformations (e.g. logarithmic transformations of continuous data, or the like) and extractions on categorical and numerical data, such as normalization and centering the data about the mean of the data. Data transformer 113 can provide the transformed or extracted data to data loader 115.

[0040] Consistent with disclosed embodiments, data loader 115 can receive the normalized data from data transformer 113. Data loader 115 can merge the data into varying formats depending on the specific requirements of dataset generator 120. Data loader 115 can then provide the processed data to dataset generator 120 (or to a suitable data storage, from which dataset generator 120 can retrieve the data).Attorney Docket No.: 15988.0009-00304

[0041] Consistent with disclosed embodiments, dataset generator 120 can be configured to generate datasets from data processed by data input engine 110. In some embodiments, training engine 160 or prediction engine 170 can be configured to expect datasets having a particular structure. Dataset generator 120 can be configured to format data into that particular structure. For example, dataset generator 120 can be configured to collect observations, associate the observations with corresponding training labels or metadata, and store the observations and metadata in data storage 140.

[0042] In some embodiments, dataset generator 120 can be configured to extract features from data received from data input engine 110. A feature can be a property or characteristics of a phenomenon. For example, the presence or absences of a fluorescent value in excess of a threshold at a location corresponding to a particular protein can be a feature. As an additional example, the average fluorescence value over samples on a biochip can be a feature. Features can be determined based on the domain, data type of a category, or many other factors associated with data stored in a data structure. Additionally, a feature can represent information about multiple data records in a data set or information about a single category in a data record. Moreover, multiple features can be produced to represent the same data.

[0043] In some embodiments, dataset generator 120 can include a classifier 121 and an annotator 123. Consistent with disclosed embodiments, classifier 121 can be configured to create classes from data provided by dataset generator 120. In some embodiments, such classes can be data labels used for training a predictive model. For example, a set of fluorescence data indicating protein levels can be associated with a medical record of a patient. In this example, the classifier can extract health information from the medical record. For example, the classifier can determine that a patient experienced a cardiovascular event within a particular amount of time following acquisition of a blood sample used to generate the fluorescence data. As an additional example, the classifier can identify a glomerularAttorney Docket No.: 15988.0009-00304 filtration rate of the patient from the medical record. In some embodiments, the classifier can be or include a natural language processing engine.

[0044] Consistent with disclosed embodiments, classifier 121 can be configured to accept classifications provided by a user through user device 180. For example, classifier 121 can be configured to provide data (or metadata concerning the data) received from data input engine 110 to user device 180 for display. In response, classifier 121 can receive class information. For example, classifier 121 can provide medical records associated with a set of fluorescence data for display to user device 180. In response, a user can interact with user device 180 to provide indications of health information based on the provided medical records. As may be appreciated, such functionality is not limited to medical records, but could be another information used to generate training labels for the data.

[0045] Consistent with disclosed embodiments, annotator 123 can be configured to receive data and associated value or class information from dataset generator. Annotator 123 can be configured to create a suitably formatted entry that associates the value or class information with the data. For example, given an array or matrix of fluorescence values and a glomerular filtration rate, annotator 123 can create an object including a “response_value” key and an “protein_data_input” key. The glomerular filtration rate can be stored with the “response_value” key and the array or matrix of fluorescence values “protein_data_input” key. In this example, training engine 160 or prediction engine 170 can expect, or be configured to expect, observations having such a format. As an additional example, given sets of fluorescence values obtained from blood samples drawn from patients and dates the patients suffered heart attacks, annotator 123 can create a relational database with rows corresponding to patients, one column storing whether the patient suffered a heart attack, another column storing when the patient suffered the heart attack, and the remaining columns storing the fluorescence values.Attorney Docket No.: 15988.0009-00304

[0046] Consistent with disclosed embodiments, model storage 130 can be a storage location for predictive models usable by training engine 160 or prediction engine 170. The disclosed embodiments are not limited to any particular implementation of data storage 140. Consistent with disclosed embodiments, data storage 140 can be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0047] Consistent with disclosed embodiments, data storage 140 can be a storage location for prepared datasets usable by training engine 160 or prediction engine 170. The disclosed embodiments are not limited to any particular implementation of data storage 140. Consistent with disclosed embodiments, data storage 140 can be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0048] Consistent with disclosed embodiments, data / model selector 150 can be configured to access model storage 130 or data storage 140 to retrieve predictive models or datasets, respectively. In some embodiments, data / model selector 150 can provide an abstraction layer for training engine 160 or prediction engine 170. In some embodiments, data / model selector 150 can be configured to control access to model storage 130 or data storage 140.

[0049] Consistent with disclosed embodiments, training engine 160 can be configured to train, or create and train, predictive models. Training engine 160 can be configured to obtain existing models from model storage 130 or training datasets from data storage 140. In some embodiments, training engine 160 can be configured to interact with data / model selector 150 to obtain the existing models or training datasets. Training engine 160 can be configured to store trained predictive models in model storage 130. In some embodiments, training engine 160 can be configured to interact with data / model selector 150 to store the trained predictive models in model storage 130.Attorney Docket No.: 15988.0009-00304

[0050] Consistent with disclosed embodiments, training engine 160 can include model trainer 161 and model evaluation 163. Training engine 160 can be configured to train a predictive model using model trainer 161 and then determine performance measure values for the predictive model using model evaluation 163. In some embodiments, training engine 160 can automatically update the predictive model being trained based on the performance measure values. In various embodiments, training engine 160 can update the predictive model being trained in response to user input provided through user device 180. Updating the predictive model can include one or more of performing additional training (e.g., using the existing training dataset or another training dataset), modifying the model (e.g., changing the input features used by the model, changing the architecture of the model, or the like), or changing the training environment (e.g., changing training hyperparameters, changing a division of the training dataset into training, cross-validation, and holdout portions, or the like).

[0051] Consistent with disclosed embodiments, as described herein, training engine 160 can be configured to determine performance measure values for a model using data obtained in different contexts. Such performance measure values can be displayed to a user through user device 180. The user can then interact through user device 180 with training engine 160 to update the model, as described herein.

[0052] Consistent with disclosed embodiments, model trainer 161 can create or train predictive models. Model trainer 161 can create or train predictive models as instructed by training engine 160. For example, training engine 160 can instruct model trainer 161 to create and train a support vector machine using a training portion of a training dataset. Model trainer 161 can then create and train this support vector machine, returning the trained support vector to the training engine 160. As an additional example, training engine 160 can instruct model trainer 161 to create and train a penalized regression model. The training engine can configure model trainer 161 with a type of the penalized regression model (e.g., ridgeAttorney Docket No.: 15988.0009-00304 regression, lasso regression, elastic net, or another suitable type), parameter values for the penalized regression (e.g., a lambda value that weights the sum of squared coefficient values), and a training portion of a training dataset. Model trainer 161 can then create and train the penalized regression model, returning the trained penalized regression model to the training engine 160. As an additional example, training engine 160 can instruct model trainer 161 to train a random forest model. Training engine 160 can provide hyperparameters, such as the size of each bootstrap sample, the number of features to consider at each split, the depth of each decision tree, and the number of decision trees in the random forest. Training engine 160 can provide a training portion of the training dataset. Model trainer 161 can then create and train the random forest model, returning the trained random forest model to the training engine 160.

[0053] Consistent with disclosed embodiments, model evaluation 163 can evaluate models trained by model trainer 161. Training engine 160 can provide model evaluation 163 a model and a cross-validation or holdout portion of the training dataset. In some embodiments, training engine 160 can specify one or more performance measures for evaluation by model evaluation 163. In various embodiments, model evaluation 163 can be configured with a predetermined or default set of performance measures. In some embodiments, the performance measures can include confusion matrices, mean-squared-error, mean-absolute- error, sensitivity or selectivity, receiver operating characteristic curves or area under such curves, precision and recall, F-measure, or any other suitable performance measure.

[0054] Consistent with disclosed embodiments, prediction engine 170 can be configured to predict health information using a patient dataset and a trained predictive model. In some embodiments, prediction engine 170 can obtain the trained predictive model from model storage 130. In some embodiments, prediction engine 170 can obtain the patient dataset from data storage 140. In some embodiments, prediction engine 170 can obtain the patient datasetAttorney Docket No.: 15988.0009-00304 (or a portion thereof) from another data storage location. This alternative data storage location can be associated with another entity or user. For example, prediction engine 170 can receive or retrieve the patient dataset from a healthcare system separate from the entity that controls prediction engine 170. In some embodiments, prediction engine 170 can obtain the model or data using data / model selector 150.

[0055] Consistent with disclosed embodiments, prediction engine 170 can apply the patient dataset to the trained predictive model to predict health information for the patient. The health information (or an indication thereof) can be provided by prediction engine 170 to user device 180. The health information can be stored on a computing device associated with pipeline 100 or provided to another system.

[0056] Consistent with disclosed embodiments, user device 180 can provide a user interface for interacting with other components of pipeline 100. The user interface can be a graphical user interface. The user interface can enable a user to configure data input engine 110 to extract, transform, and load data according to user specification. The user interface can enable the user to specify how the transformed data received by dataset generator 120 is converted into labeled training data (or patient data suitable for predictions). In some embodiments, the user interface can enable the user to interact with dataset generator 120 to manually or semi-manually label or annotate the training data. In some embodiments, the user interface can enable a user to interact with data / model selector 150 to manage data or models stored in model storage 130 or data storage 140. Such management can include deleting or creating models, deleting datasets, or restricting access by training engine 160 or prediction engine 170 to models or datasets. In some embodiments, the user interface can enable a user to interact with data / model selector 150 to push data or models to training engine 160 for training, or to prediction engine 170 for prediction. In some embodiments, the user interface can enable a user to interact with training engine 160 to create or select aAttorney Docket No.: 15988.0009-00304 predictive model for training, create or select a dataset for use in training the model, or select training parameters or hyperparameters. In some embodiments, the user interface can enable a user to interact with training engine 160 to display information related to training of the model (e.g., performance measure values, a change in loss function values during training, or other training information). In some embodiments, the user interface can enable a user to interact with prediction engine 170 to select a training model and patient data for use in predicting health information. In some embodiments, the user interface can enable a user to interact with prediction engine 170 to display the health information, store the health information on a computing device, or transmit the health information to another system.

[0057] Components of pipeline 100 can be implemented using one or more computing devices. Such computing devices can include tablets, laptops, desktops, workstations, computing clusters, or cloud computing platforms. In some embodiments, components of pipeline 100 can be implemented using cloud computing platforms. For example, one or more of data input engine 110, dataset generator 120, data / model selector 150, training engine 160, and prediction engine 170 can be implemented on a cloud computing platform. In some embodiments, components of pipeline 100 can be implemented using on-premises systems. For example, measurement system 101, records 103, or user device 180 can be, or be hosted on, on-premises systems. As an additional example, model storage 130 or data storage 140 can be, or be hosted on, on-premises systems.

[0058] Components of pipeline 100 can communicate using any suitable method. In some embodiments, two or more components of pipeline 100 can be implemented as microservices or web services. Such components can communicate using messages transmitted on a computer network. The messages can be implemented using SOAP, XML, HTTP, JSON, RCP, or any other suitable format. In some embodiments, two or more components of pipeline 100 can be implemented as software, hardware, or combined software / hardwareAttorney Docket No.: 15988.0009-00304 modules. Such components can communicate using data or instructions written to or read from a memory (e.g., a shared memory), function calls, or any other suitable communication method.

[0059] As may be appreciated, the particular structure of pipeline 100 is not intended to be limiting. Consistent with disclosed embodiments, any two or more of record(s) 103, model storage 130, or data storage 140 can be combined, or hosted on the same computing device. Consistent with disclosed embodiments, data input engine 110 and dataset generator 120 can be omitted from pipeline 100. In such embodiments, datasets formatted and configured for use by training engine 160 or prediction engine 170 can be deposited in data storage 140 by another system or using another method. Consistent with disclosed embodiments, data input engine 110 and dataset generator 120 can be combined. In such embodiments, data extraction, transformation, and loading can be combined with feature extraction, annotation, and classification. Consistent with disclosed embodiments, data / model selector 150 can be combined with one or more of training engine 160 and prediction engine 170. For example, training engine 160 or prediction engine 170 can include functionality for retrieving selected data or models from model storage 130 or data storage 140.

[0060] Though shown with one user device 180, pipeline 100 could have multiple user devices. Different user devices could be associated with different entities or different users having different roles. For example, user device 180 could be associated with a software engineer or data scientist who is developing the test, while another user device could be associated with a clinician who is using the test.

[0061] User device 180 can be combined with one or more other components of pipeline 100. In some embodiments, user device 180 and at least one of data / model selector 150, training engine 160, or prediction engine 170 can be implemented by the same computing device. InAttorney Docket No.: 15988.0009-00304 various embodiments, user device 180 and at least one of model storage 130 or data storage 140 can be implemented by the same computing device.

[0062] As may be appreciated, pipeline 100 can be integrated into a method for treating patients with a particular health condition. Prediction engine 170 can use a trained predictive model and input data obtained from a patient sample to determine whether the patient is experiencing a negative health outcome. For example, the trained predictive model and input data obtained from a patient sample can be used to support a diagnosis (e.g., a diagnosis of cancer, non-alcoholic fatty liver disease (NAFLD), or the like), indicate a disease state or state (e.g., a stage of NAFLD, or the like), a prognosis (e.g., a neuroblastoma prognosis), or the like. If the patient has a particular diagnosis, or a risk greater than (or potentially equal to) a health-outcome dependent threshold, then the patient can be treated or monitored according to a first, more-aggressive or intensive protocol. If the patient lacks the diagnosis, or has a risk less than (or potentially equal to) the health-outcome dependent threshold, then the patient can be treated or monitored according to a second, less-aggressive or intensive protocol.

[0063] FIGs.2A-2H provide a high-level depiction of stages in an exemplary aptamer-based serum or plasma assay 200, consistent with disclosed embodiments. Assay 200 can generate suitable input data for training a predictive model or predicting health information using a trained predictive model. Assay 200 can be performed using, at least in part, a testing system, such as measurement system 101 of pipeline 100.

[0064] Consistent with disclosed embodiments, assay 200 can quantitatively transform protein epitope availability in a biological sample into a specific DNA signal. In general, assay 200 can use SOMAmer® (Slow Off-rate Modified Aptamer) reagents that comprise short, single-stranded DNA sequences that incorporate hydrophobic modifications. Assay 200 can measure native proteins in complex matrices by transforming available binding sitesAttorney Docket No.: 15988.0009-00304 on individual proteins into a corresponding SOMAmer reagent concentration, which can then be quantified by hybridization to microarrays. In this manner, the test takes advantage of SOMAmer reagents’ dual nature as both protein affinity-binding reagents with defined three- dimensional structures and unique nucleotide sequences recognizable by specific DNA hybridization probes. Thus relative epitope concentrations can be converted into measurable nucleic acid signals that can be quantified using DNA-hybridization microarrays.

[0065] Consistent with disclosed embodiments, suitable versions of assay 200 can quantify relative levels of proteins in plasma spanning 10 logs in abundance. Such test versions can measure up to one thousand, three thousand, five thousand, seven thousand, ten thousand, or more unique protein analytes. Tests can be performed on small volume samples (e.g., samples greater than 10 microliters, 20 microliters, 40 microliters, 100 microliters, 200 microliters, 400 microliter, 1 milliliter, or greater).

[0066] As may be appreciated, SOMAmer reagents can be selected against proteins in their native folded conformations. Thus such reagents may require an intact, tertiary protein structure for binding. Accordingly, unfolded and denatured—and therefore presumably inactive—proteins may not be detected by SOMAmer reagents (or may be detected with reduced or varying sensitivity).

[0067] As depicted in FIG.2A, SOMAmer reagents can be synthesized with a fluorophore, photocleaveable linker, and biotin. Next, as depicted in FIG.2B, SOMAmer reagents bound to streptavidin beads can be used to capture proteins from a complex mixture of proteins in a biological sample (e.g., a serum or plasma sample). Next, as depicted in FIG.2C, unbound proteins can be washed away, and bound proteins can be tagged with biotin. Next, as depicted in FIG.2D, electromagnetic radiation (e.g., ultraviolet light, or the like) can be applied to the solution to break the photocleaveable linker, releasing the proteins complexes and bound SOMAmers back into solution. As shown in FIG.2E, non-specific complexes canAttorney Docket No.: 15988.0009-00304 disassociate from corresponding SOMAmers, while specific complexes remain bound. Next, as depicted in FIG.2F a polyanionic competitor can be added to the solution. The polyanionic competitor can prevent rebinding of non-specific complexes. As shown in FIG. 2G, the biotinylated proteins (and bound SOMAmer reagents) can then be captured on streptavidin beads. The beads and bound proteins can be separated from the solution or concentrated. Next, as depicted in FIG.2H, the SOMAmer reagents can be released from the protein complexes by denaturing the proteins. Fluorophores can be measured after hybridization to complementary sequences on a microarray chip. The fluorescence intensity detected on the microarray can be related to the amount of available epitope in the original sample.

[0068] As may be appreciated, assay 200 is intended to be exemplary. The disclosed system and methods are not limited to tests having these particular steps. In some embodiments, other aptamers (or even other classes of components) can be used to bind protein complexes. Alternative methods of inhibiting non-specific binding may be used in place of a polyanionic competitor. Alternative methods of separating protein-compound complexes may be used in place of capturing protein-compound complexes on streptavidin beads. Alternative indicia of protein levels can be measured in place of fluorophore measurements on a microarray chip. However, such alternative methods can still exhibit the technical challenges described herein. Therefore, such alternative systems and methods can benefit from the disclosed technical solutions.

[0069] FIG.3 depicts a process 300 for generating and applying a proteomics model, consistent with disclosed embodiments. Process 300 can generate the proteomics model using a training protein expression dataset and health information. The proteomics model can be a predictive model, as described herein. The proteomics model can be configured to accept as input topic weights and provide as output a diagnosis or prognosis. The proteomics modelAttorney Docket No.: 15988.0009-00304 can be used with patient protein expression data to generate a diagnosis or prognosis for the patient.

[0070] For convenience of description, process 300 is described herein as being performed by a pipeline for developing, validating, and deploying proteomics models for predicting health information, such as pipeline 100. For example, protein expression data and health information can be obtained using a data input engine (e.g., data input engine 110, or the like). The data input engine can process the data into a format expected by a dataset generator (e.g., dataset generator 120, or the like). The dataset generator can be configured to generate a training dataset and store the training dataset in a suitable storage location (e.g., data storage 140, or the like). A training engine (e.g., training engine 160, or the like) can generate a proteomics model using the training dataset. A prediction engine (e.g., prediction engine 170, or the like) can use the generated proteomics model to generate the diagnosis or prognosis for the patient. However, the disclosed embodiments are not so limited. Other systems or architectures can be used without departing from the envisioned embodiments.

[0071] Similarly, process 300 is described herein as being performed using protein expression data generated using an aptamer-based assay, such as the assay described above with regards to FIGs.2A to 2H. However, the disclosed embodiments are not so limited. Other assays capable of generating protein expression data can be used without departing from the envisioned embodiments.

[0072] In step 301 of process 300, the data input engine can obtain a training protein expression dataset, consistent with disclosed embodiments. The training protein expression dataset can include protein expression values for a set of training samples. The training samples can be obtained from patients. In some embodiments, the training samples can be plasma samples. In some embodiments, the training samples can be serum samples. In some embodiments, the training samples can be cell lysate samples. In some embodiments, the setAttorney Docket No.: 15988.0009-00304 of training samples (or a subset of the training sample) can share a common characteristic. For example, the set or subset of samples can be obtained from patients or patient tissues satisfying a particular criterion. As may be appreciated, this common characteristic can affect the generated topics.

[0073] In some embodiments, the set of training samples can be configured to favor generation of topics suitable for diagnosis of a particular health status. In some embodiments, the set of training samples can include target samples taken from patients having the health status and a control samples taken from patients lacking the health status. Structuring the training dataset in this manner can favor generation of topics that distinguish the target samples from the control samples. For example, the set of training samples can include samples taken from patients with and without NAFLD. The training samples taken from patients with NAFLD can be further stratified into subgroups having different degrees of NAFLD (e.g., early stage, late stage, or the like), or different patient histories (e.g., severity or rapidity of disease progression, or the like). Topics may then be identified that diagnose NAFLD, identify the degree of NAFLD, or predict the severity or rapidity of NAFLD progression.

[0074] In some embodiments, the set of training samples can include only samples taken from patients having the health status. The biological significance of topics can be evaluated using proteins contributing to the topic. Proteins contributing to biologically significant topic(s) can then be included in a predictive model. The predictive model can be configured to provide a diagnosis or prognosis regarding the health condition. For example, the set of training samples can include only samples from patients with neuroblastomas. Topics indicative of neuroblastoma sub-type can then be identified based on the proteins contributing to those topics. Other proteins also contributing to these topics can be selected for inclusion in a proteomics model configured to predict neuroblastoma prognosis based onAttorney Docket No.: 15988.0009-00304 neuroblastoma type. In this manner, known proteins can be used to identify topics, which can then be used to identify additional relevant proteins.

[0075] Consistent with disclosed embodiments, a measurement system (e.g., measurement system 101, or the like) of the pipeline can obtain protein expression values from the set of training samples using an aptamer-based assay. The data input engine can obtain the training protein expression dataset from the measurement system or from record(s) (e.g., records 103, or the like).

[0076] In step 303 of process 300, a component of the pipeline (e.g., the data input engine, the dataset generator, the training engine, or the like) can generate a sample definition dataset and a topic definition dataset, consistent with disclosed embodiments. The sample definition dataset can include topic weights. A topic weight for a combination of a sample and a topic can indicate an association of the topic with the sample. For example, the greater the topic weight, the stronger the association of the topic with the sample. The topic definition dataset can include protein weights. A protein weight for a combination of protein and topic can indicate an association of the protein with the topic. For example, the greater the protein weight, the stronger the association of the protein with the topic.

[0077] In some embodiments, the topic definition dataset can be a topic definition matrix. The topic definition matrix can include a first dimension corresponding to proteins in the training protein expression dataset and a second dimension corresponding to topics.

[0078] In some embodiments, the sample definition dataset can be a sample definition matrix. The sample definition matrix can include a first dimension corresponding to topics and second dimension corresponding to samples in the training protein expression dataset.

[0079] In some embodiments, the component of the pipeline can be configured to generate the topic definition matrix and the sample definition matrix using non-negative matrix factorization. The disclosed embodiments are not limited to any particular implementation ofAttorney Docket No.: 15988.0009-00304 non-negative factorization. For example, the Python sklearn.decomposition.nmf class can be used to generate a non-negative factorization of a the training protein expression dataset. In some embodiments, the number of topics (e.g., the size of the second dimension of the topic definition matrix and of the first dimension of the sample definition matrix) can be an input to the non-negative matrix factorization. This number of topics can be a default number or a parameter of process 300. When the number is a parameter of process 300, the number can be determined a-priori (e.g., based on previous results or using a heuristic), through an iterative process of generating non-negative factorizations of the training protein expression dataset and investigating the resulting topics, as described herein, or through another suitable method. In some embodiments, the number of topics can be greater than 1 and less than the number of features (e.g., the number of proteins). In some embodiments, a suitably informative number of topics can be within the range of 5 to 100 topics. In some instances, the number of topics can be selected based on the number of observations and nature of the data, an a-priori hypothesis, or the like. In some instances, the number of topics can be reached through an iterative exploration of the data.

[0080] In some embodiments, process 300 can include generation of a predictive model using topics identified in step 303. In such embodiments, the predictive model can be generated without identifying proteins contributing to the topics. Accordingly, in such embodiments, process 300 can proceed to step 309. In some embodiments, process 300 can include identification of proteins contributing to topics. Accordingly, in such embodiments, process 300 can proceed to step 305.

[0081] In step 305 of process 300, a component of the pipeline (e.g., the data input engine, the dataset generator, the training engine, or the like) can identify contributing proteins for topics, consistent with disclosed embodiments. The component can determine protein contribution values for each combination of protein and topic. A protein can then beAttorney Docket No.: 15988.0009-00304 identified as contributing to a topic based on the protein contribution value for that combination of protein and topic.

[0082] In some embodiments, protein contribution values can be determined using the topic definition dataset. The protein weights for a protein can be used to determine the protein contribution values for the protein to the topics in the topic definition dataset. In some embodiments, a first statistic of original protein weight values for the protein can be determined. In various embodiments, the first statistic can be an average, a median or percentile (e.g., the 25thpercentile, 50thpercentile, 75thpercentile, or the like), or the like.

[0083] In some embodiments, updated protein weight values can be determined for the protein using the original protein weight values and the first statistic. The updated protein weight value can be a protein value difference. The protein value difference can be a difference between the original protein weight value and the first statistic. For example, when the first statistic is the median of the original protein weight values for the protein, the updated protein weight value can be the difference between the original protein weight values can the median of the original protein weight values.

[0084] In some embodiments, a second statistic of updated protein weight values for the protein can be determined. The second statistic can be an average, a median or percentile (e.g., the 25thpercentile, 50thpercentile, 75thpercentile, or the like), or the like. The second statistic may be the same as the first statistic. For example, the first statistic can be the median of the original protein weights and the second statistic can be the median of the updated protein weights. However, in some embodiments, the second statistic may differ from the first statistic. For example, the first statistic can be the average of the original protein weights and the second statistic can be the 75% percentile of the updated protein weights.Attorney Docket No.: 15988.0009-00304

[0085] In some embodiments, the protein contribution values for the protein can be determined using the second statistic and the original protein weight values for the protein. The protein contribution values can be a function of the second statistic and the original protein weight values for the protein. For example, a protein contribution value can be the ratio (or percentage, or the like) of the original protein weight and the second statistic. For example, the protein contribution value can be the original protein weight divided by the median of the updated protein values.

[0086] Consistent with disclosed embodiments, a protein can be identified as a contributing protein for a topic based on the protein contribution value for that combination of protein and topic. In some embodiments, the protein contribution value can be compared to a threshold value. When the protein contribution value is greater than the contribution value, the protein can be identified as contributing to the topic.

[0087] In some embodiments, the threshold value can be predetermined. For example, the threshold value can be a default value. As an additional example, the threshold value can be determined by a user based on previous experience or a heuristic. In some embodiments, the threshold value can be between 5 and 50. In some instances, the threshold value can be selected (e.g., automatically, manually, or the like) based on the distribution of protein contribution values (e.g., at a location on the tail of the distribution).

[0088] In some embodiments, the threshold value can be determined based on the protein contribution values for the protein. The threshold value can be based on statistic(s) of the protein contribution values. For example, the threshold value can be the mean of the protein contribution values for the protein plus a constant times the standard deviation of the protein contribution values for the protein. The constant can be chosen to achieve a desired performance value (e.g., true positive rate, or the like). The threshold value can similarly beAttorney Docket No.: 15988.0009-00304 determined based on the protein contribution values for all of the proteins in the topic description database.

[0089] In some embodiments, the threshold value can be determined based on the topic weights in the topic description database. The topic weights can be randomly reassigned to different combinations of protein and topic. Protein contribution values determined for these randomized topic weights. The threshold can be based on statistic(s) of these randomized protein contribution values. For example, the threshold value can be the mean of the randomized protein contribution values plus a constant times the standard deviation of the randomized protein contribution values.

[0090] In some embodiments, process 300 can include generation of a predictive model using proteins identified as contributing to topics in step 305. In some such embodiments, the predictive model can be generated without identifying the biological significance of the topics to which the proteins contribute. For example, all proteins identifying as contributing to at least one topic can be used as features for constructing a predictive model. Accordingly, in such embodiments, process 300 can proceed to step 309. In some embodiments, process 300 can include identification of the biological significance of topics based on the proteins contributing to the topics. Accordingly, in such embodiments, process 300 can proceed to step 307.

[0091] In step 307 of process 300, the biological significance of one or more topics can be identified based on the proteins contributing to those topics, consistent with disclosed embodiments. The disclosed embodiments are not limited to any particular method of assessing the biological significance of a topic based on the proteins contributing to the topic.

[0092] In some embodiments, the assessment can be performed automatically. For example, a training engine (or another suitable component of the pipeline) can be configured to automatically obtain information about the proteins that contribute to the topic. SuchAttorney Docket No.: 15988.0009-00304 information can be obtained from another system, such as a repository of protein information (e.g., gene product ontology information from the Gene Ontology initiative, or the Protein Ontology database, or the like). The training engine of the pipeline can be configured to determine potential biological relationships, pathways, or structures using this retrieved information.

[0093] In some embodiments, the assessment can be performed manually or semi-manually. For example, the training engine can be configured to display indications of the contributing proteins to a computing device of a user (e.g., user device 180). The user can determine the biological significance of the topic based on the displayed indications.

[0094] As may be appreciated, based on the biological significance of a topic, a topic may be included in a predictive model. The topic may be automatically, manually, or semi-manually included in the predictive model. In some embodiments, a user can interact with the pipeline to include a topic in a predictive model, based on a biological significance of the topic. In various embodiments, a user can indicate relevant biological relationships, pathways, or structures, and the pipeline can automatically include topics concerning such relationships, pathways, or structures in the predictive model. For example, the user can include (or direct the pipeline to include) topics concerning biological pathways common to cancer in a predictive model for diagnosing cancer.

[0095] Similarly, based on the biological significance of a topic, proteins contributing to that topic can be included in a predictive model. In some embodiments, the predictive model can be configured to accept as input protein expression values (e.g., as opposed to topic weights) as inputs.

[0096] In step 309 of process 300, a data input engine (or another suitable component of the pipeline) can obtain health information, consistent with disclosed embodiments. As describedAttorney Docket No.: 15988.0009-00304 herein, the data engine can receive or retrieve the health information from record(s) associated with the pipeline, or from another system.

[0097] The health information can concern the samples corresponding to the training protein expression dataset. The health information can provide (or be usable to generate) ground truth labels for training the predictive model. As described herein, such health information can include whether a health outcome occurred, or is predicted to occur, within a particular time frame, an indication of a health status, a behavioral health prediction, or the like. For example, when process 300 is being used to develop a diagnostic model for a particular type of cancer, the training protein expression dataset can be generated using samples acquired from patients having the particular type of cancer. In this example, the health information can specify which samples were obtained from patients having the type of cancer and which samples were obtained from controls. Similarly, when process 300 is being used to develop a prognostic model of patient survival time for a particular type of cancer, the health information can include patient survival time.

[0098] As may be appreciated, step 309 can be performed prior to any of steps 301, 303, 305, or 307, without departing from the envisioned embodiments. Furthermore, step 309 can be combined with step 301. For example, a data input engine can be configured to retrieve medical records including patient health information and assay results. The assay results can be used to generate the training protein expression dataset, while the patient information can be used to generate ground truth labels for training the predictive model. In such embodiments, process 300 can proceed from step 303, 305, or 307 to step 311, rather than step 309.

[0099] In step 311 of process 300, a training engine (or another suitable component of the pipeline) can generate a predictive model, consistent with disclosed embodiments. The predictive model can output health information based on input features. The healthAttorney Docket No.: 15988.0009-00304 information obtained in step 309 can provide a ground truth for training. In some embodiments, after training is complete, the training engine can store the trained model in model storage 130, or another suitable location.

[0100] In some embodiments, the input features can be topics, or can depend upon topics. The biological significance of the topics may (e.g., in step 307) or may not have been determined as part of process 300. In such instances, the training engine can be configured to use the sample definition dataset to train the predictive model. The predictive model can be trained to generate an appropriate output for a sample based on the topic weights for that sample in the sample description dataset.

[0101] In some embodiments, the input features can be proteins, or can depend upon proteins. In some embodiments, the proteins may have been identified in step 305 as contributing to a topic, or identified in step 307 based on contributions to a particular topic having relevant biological significance. In such instances, the training engine can be configured to use the training protein expression dataset to train the predictive model. The predictive model can be trained to generate an appropriate output for a sample based on the protein expression values for the selected proteins in the training protein expression dataset.

[0102] In some embodiments, the predictive model can be a regression model (e.g., a linear or logistic regression model). The regression model can be a penalized regression model (e.g., an elastic net, lasso, ridge, or other suitable penalized regression model). The disclosed embodiments are not limited to any particular training implementation. Instead, the training implementation can depend on the type of predictive model, the intended output and use of the predictive model, the size and quality of the training protein expression dataset, and other known features or characteristics.

[0103] In step 313 of process 300, the prediction engine (or another suitable component of the pipeline) can apply the predictive model generated in step 311, consistent with disclosedAttorney Docket No.: 15988.0009-00304 embodiments. In some embodiments, the prediction engine can obtain the predictive model from a model storage of the pipeline, or from another suitable location. The prediction engine can obtain patient protein expression data. The patient protein expression data can be obtained a data storage associated with the pipeline, or another suitable location.

[0104] Consistent with disclosed embodiments, the patient protein expression data can be acquired from a patient sample using an aptamer-based assay. In some embodiments, the patient sample can be of the same type as the training samples used to generate the training protein expression dataset. For example, when the training samples are cell lysate samples, the patient sample can be a cell lysate sample. In some embodiments, the patient sample may differ in type from the training samples.

[0105] When the predictive model accepts topics as inputs, the prediction engine can generate an outcome measure as some function of patient topic weights using the sample definition dataset, or patient protein expression data for proteins defined as relevant to the outcome by the topics. The prediction engine can then provide the patient topic weights to the predictive model to generate a diagnosis or prognosis. When the predictive model accepts protein expression values as inputs, the prediction engine can provide the patient protein expression values to the predictive model to generate a diagnosis or prognosis.

[0106] Consistent with disclosed embodiments, the prediction engine can be configured to output the diagnosis or prognosis. In some embodiments, the prediction engine can provide the diagnosis or prognosis to the user device. The user device can display an indication of the diagnosis or prognosis to the user. In various embodiments, the prediction engine can provide the diagnosis or prognosis to a record(s) system associated with the pipeline. The record(s) system can be configured to store the diagnosis or prognosis (e.g., in a medical record of the patient, or another suitable location). In some embodiments, the prediction engine can be configured to provide the output to another system.Attorney Docket No.: 15988.0009-00304

[0107] FIGs 4A to 4D depict an assessment of a protein expression dataset for a cancer cell line study, consistent with disclosed embodiments. A protein expression dataset was obtained using cancer cell lysate samples for twenty-seven cancer cell lines (each including multiple samples). A topic definition dataset and a sample definition dataset were generated using non-negative factorization of the protein expression dataset. The non-negative factorization used fifteen topics. FIG.4A depicts a portion of the topic definition dataset, consistent with disclosed embodiments. This portion includes the protein weights for one of the proteins for each of 15 topics in the cancer cell line study. As may be seen, the protein weights for four topics are substantially elevated, as compared to the protein weights for the remaining topics. However, because the protein contributes to four topics, and not one, there is a risk that a divergence-based method (e.g., a minimum KL divergence-based method, or the like) will fail to identify the protein as contributing to any of these topics.

[0108] In contrast, in some embodiments, protein contributions to the topics can be calculated as follows:where ^ is the protein description matrix, ^^is a vector of original protein weights corresponding to the ithprotein, ^^^is the original protein weight for the ithprotein and the jthtopic, k is a constant scale factor, and ^^^is the contribution of the ithprotein to the jthtopic. The contribution ^^^can be described as the median absolute deviation from median of the ithprotein for the jthtopic. The constant k may be set to 1, or to approximately 1.4836 to ensure ^^^is a consistent estimator of standard deviation, or some other value. FIG.4B depicts the protein contributions (e.g., median absolution deviations from median) corresponding to the protein weights depicted in FIG.4A. In this example, a threshold value of 5 (identified in theAttorney Docket No.: 15988.0009-00304 figures as a dashed line) is used to identify proteins significantly contributing to a topic. This protein is therefore identified as significantly contributing to one topic, the topic for which the protein contribution value exceeds the threshold value.

[0109] FIGs.4C and 4D depict the sample definition dataset (e.g., sample dataset 401), consistent with disclosed embodiments. Sample dataset 401 shows the proportionate contribution to each topic of each sample. The columns of the sample definition dataset have been hierarchically clustered (e.g., as shown in hierarchical clustering 402) based on these contributions. This clustering has resulted in the grouping of cell samples (e.g., cell samples 403, each cell sample positioned above the corresponding column of the sample definition dataset 401) from the same cell line (e.g., cell line 404). In turn, cell lines are associated with tissue types (e.g., tissue type 405), which also exhibit a degree of clustering. The sample distribution dataset shows the relative contributions of the 15 topics (e.g., topics 406) to each sample in the cancer cell line study. As may be observed, certain topics contribute heavily to samples from certain tissues, but not to samples from other tissues. For example, topic K8 appears largely specific to samples from the autonomic ganglia, topic K1 appears largely specific to samples from the hematopoietic lymphoid, topic K15 appears largely specific to samples from the stomach, and topic K11 appears largely specific to samples from the endometrium. Furthermore, as shown by the hierarchical clustering of the columns of the sample dataset, samples from the same cell line cluster together, and samples from similar tissues may cluster together. For example, two cell lines from autonomic ganglia (e.g., SKNDZ and SKNFI) cluster together, as do two kidney cell lines (e.g., A704 and 769P). However, other cells lines drawn from the same tissue cluster apart, suggesting differences in those cell lines that may support different treatments or prognoses (e.g., HUNS1, RJ, and RL).Attorney Docket No.: 15988.0009-00304

[0110] FIGs.5A and 5B concern evaluation of the biological significance of topics identified in the protein expression dataset obtained in the cancer cell line study discussed above, consistent with disclosed embodiments. FIG.5A concerns a first one of the fifteen topics. The sample definition dataset was evaluated to identify the first topic, based on the first topic having elevated sample weights for many of the cancer cell lines. The proteins contributing to this topic were then evaluated using the formula described above with regards to FIG.4 to determine the contributionof each protein to this topic. A predetermined threshold value of 5 was used to identify proteins significantly contributing to the topic, for biological annotation of the topic. Out of approximately 7000 proteins in the original training protein expression dataset, 156 proteins were identified as contributing to this topic. Some of these proteins had known biological significance. For example, the contributing proteins including growth factors (e.g., VEGF-D, epidermal growth factor-like protein 6, fibroblast growth factor 22, fibroblast growth factor 6) and other cancer-related proteins (e.g., tumor protein 63).

[0111] Furthermore, clustering was performed on the top contributing proteins using a database of known and predicted protein-protein interactions (including direct physical interactions and functional interactions). FIG.5A depicts the results of this clustering: two distinct clusters of related proteins. Cluster 501 was associated with the WNT signaling pathway. Abnormal WNT signaling is associated with multiple cancer types. Cluster 503 was associated with inflammation.

[0112] Given the biological significance of the proteins associated with the topic, the topic can be deemed to involve a general cancer pathway. Thus the topic may be included in predictive models intended to diagnose cancer.

[0113] FIG.5B concerns a second one of the fifteen topics. The sample definition dataset was evaluated to identify the second topic, based on the topic having high sample weights forAttorney Docket No.: 15988.0009-00304 two neuroblastoma cancer cell lines (SK-N-FI and SK-N-DZ) and low sample weights for the remaining cancer cell lines. The proteins contributing to this topic were then evaluated using the formula described above with regards to FIG.4. A predetermined threshold (five) was used to identify proteins significantly contributing to the topic. Out of approximately 7000 proteins in the original training protein expression dataset, 156 proteins were identified as contributing to this topic. Some of these proteins had known biological significance. For example, the contributing proteins included proteins associated with neural tissue, such as synaptic signaling proteins, neurexin signaling proteins and nervous system development proteins. The top identified proteins included complexin 3 (a synaptic regulator, at rank 1), synaptotagmin (a synaptic vesicle membrane protein, at rank 5), and contactin (an oncogenic neuronal surface interaction mediator, at rank 6). Furthermore, the top-ranked proteins included proteins associated with genes whose upregulation is, in turn, associated with worse neuroblastoma outcomes (e.g., neuromodulin at rank 19, DOPA decarboxylase at rank 24, dihydropyrimidinase-related protein 1 at rank 151).

[0114] Clustering was performed on the top contributing proteins using a database of known and predicted protein-protein interactions (including direct physical interactions and functional interactions). FIG.5B depicts the results of this clustering: a large cluster of related proteins, showing that the top proteins contributing to the topic are functionally related.

[0115] Given the biological significance of the proteins associated with the topic, the topic can be deemed to involve neuronal tissues. The proteins having increased expression values in the training protein expression dataset were associated with synapses, synaptic signaling, neurexin family protein bind, nervous system development, dendrites, and neoplasms. Thus the topic may be included in predictive models intended to diagnose neuroblastomas.Attorney Docket No.: 15988.0009-00304

[0116] FIGs.5C and 5D concern evaluation of the biological significance of topics identified in a second training protein expression dataset, consistent with disclosed embodiments. The second training protein dataset may be obtained from blood serum samples of a disease- enriched cohort, such as nonalcoholic fatty liver disease (NAFLD). A predetermined number of topics (e.g. fifteen) could be used to generate a topic definition dataset and a sample definition dataset. Using the sample definition dataset, health information associated with samples could be associated with the generated topics.

[0117] FIG.5C depicts the association between a first topic in the sample description matrix and health information associated with a training sample. The training samples could be associated with a health information associated score ranging from 0 (least severe) to 2 (most severe), such as with cellular ballooning in NAFLD. For this first topic, higher topic weights could be predictive of a higher health information associated score, such as a cellular ballooning score. Furthermore, higher topic weights may also be predictive of other clinically relevant variables. For example, with NAFLD this includes lobular inflammation and steatosis scores.

[0118] FIG.5D depicts the association between a second topic in the sample description matrix and second health information (fibrosis score) associated with a training sample. The training samples are associated with fibrosis score, ranging from 0 (none) to 4 (most severe stage). For the second topic, higher topic weights are predictive of disease stage (e.g., fibrosis or cirrhosis). Different proteins contribute differently to the first topic and the second topic. In such a case where topics are associated with different clinically relevant variables of a disease, one or more topics may be used to generate a predictive model capable of assessing disease progression. As fibrosis is associated with later stages of NAFLD, the association of the second topic, but not the first topic, with fibrosis could be used to generate a predictive model capable of assessing NAFLD progression.Attorney Docket No.: 15988.0009-00304

[0119] Furthermore, the health information associated topics could have proteins having known biological relevance to the disease or endpoint of interest. For example, known NAFLD biomarkers are among the top-ranked proteins by protein contribution value for topic 11 (e.g. PTGR1 at rank 9, AKR1B10 at rank 11, ACY1 at rank 18). Furthermore, the top n proteins (where n=100, or some other number) may include many proteins related to the function of the disease-affected tissue or disease-related biological processes (here, being liver function, hepatic steatosis, and abnormalities of the liver).

[0120] Clustering was performed on the top contributing proteins using a database of known and predicted protein-protein interactions (PPI, including direct physical interactions and functional interactions). FIG.5E depicts the results of such clustering: the PPI network of the top defining proteins of the third topic of the NAFLD samples. An associated PPI shows that the top proteins contributing to the topic are functionally related, and that these top proteins are enriched for these functions.

[0121] Consistent with disclosed embodiments, the first and second topics (or the proteins contributing to the first and second topics) can be included in a predictive model, as described herein.

[0122] As used herein, unless specifically stated otherwise, the term “or” encompasses all possible combinations, except where infeasible. For example, if it is stated that a component may include A or B, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0123] Example embodiments are described above with reference to flowchart illustrations or block diagrams of methods, apparatus (systems) and computer program products. It will be understood that each block of the flowchart illustrations or block diagrams, and combinationsAttorney Docket No.: 15988.0009-00304 of blocks in the flowchart illustrations or block diagrams, can be implemented by computer program product or instructions on a computer program product. These computer program instructions may be provided to a processor of a computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart or block diagram block or blocks.

[0124] These computer program instructions may also be stored in a computer readable medium that can direct one or more hardware processors of a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the computer readable medium form an article of manufacture including instructions that implement the function / act specified in the flowchart or block diagram block or blocks.

[0125] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart or block diagram block or blocks.

[0126] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a non-transitory computer readable storage medium. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.Attorney Docket No.: 15988.0009-00304

[0127] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, IR, etc., or any suitable combination of the foregoing.

[0128] Computer program code for carrying out operations, for example, embodiments may be written in any combination of one or more programming languages, including an object- oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The program code may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0129] The flowchart and block diagrams in the figures illustrate examples of the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams or flowchart illustration, and combinations of blocks in the block diagrams or flowchart illustration, can be implementedAttorney Docket No.: 15988.0009-00304 by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0130] It is understood that the described embodiments are not mutually exclusive, and elements, components, materials, or steps described in connection with one example embodiment may be combined with, or eliminated from, other embodiments in suitable ways to accomplish desired design objectives.

[0131] The disclosed embodiments may further be described using the following clauses: 1. A proteomics-based diagnostic or prognostic system comprising: at least one processor; and at least one computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay; generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.Attorney Docket No.: 15988.0009-00304 The proteomics-based diagnostic or prognostic system of clause 1, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 3. The proteomics-based diagnostic or prognostic system of any one of clauses 1 to 2, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information. 4. The proteomics-based diagnostic or prognostic system of any one of clauses 1 to 2, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins. 5. The proteomics-based diagnostic or prognostic system of clause 4, wherein: determining protein contribution values for the first topic comprises: determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic.Attorney Docket No.: 15988.0009-00304 6. The proteomics-based diagnostic or prognostic system of any one of clauses 1 to 5, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples. 7. The proteomics-based diagnostic or prognostic system of any one of clauses 1 to 6, wherein: the health information includes: a cancer diagnosis or prognosis; or a non- alcoholic fatty liver disease diagnosis or prognosis. 8. The proteomics-based diagnostic or prognostic system of clause 7, wherein: the cancer diagnosis or prognosis concerns neuroblastoma. 9. A non-transitory, computer-readable medium containing instructions that, when executed by at least one processor of a system, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay; generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.Attorney Docket No.: 15988.0009-00304 10. The non-transitory, computer-readable medium of clause 9, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 11. The non-transitory, computer-readable medium of any one of clauses 9 to 10, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information. 12. The non-transitory, computer-readable medium of any one of clauses 9 to 10, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins. 13. The non-transitory, computer-readable medium of clause 12, wherein: determining protein contribution values for the first topic comprises: determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic.Attorney Docket No.: 15988.0009-00304 14. The non-transitory, computer-readable medium of any one of clauses 9 to 13, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples. 15. The non-transitory, computer-readable medium of any one of clauses 9 to 14, wherein: the health information includes: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis. 16. The non-transitory, computer-readable medium of clause 15, wherein: the cancer diagnosis or prognosis concerns neuroblastoma. 17. A method comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay; generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.Attorney Docket No.: 15988.0009-00304 18. The method of clause 17, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 19. The method of any one of clauses 17 to 18, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information. 20. The method of any one of clauses 17 to 18, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins. 21. The method of clause 20, wherein: determining protein contribution values for the first topic comprises: determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic. 22. The method of any one of clauses 17 to 21, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.Attorney Docket No.: 15988.0009-00304 23. The method of any one of clauses 17 to 22, herein: the health information includes: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis. 24. The method of clause 23, wherein: the cancer diagnosis or prognosis concerns neuroblastoma. 25. A system for training a proteomics-based diagnostic or prognostic model, comprising: at least one processor; and at least one computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in the set of characteristic proteins for the first topic based on the median of the protein value differencesAttorney Docket No.: 15988.0009-00304 and a protein value of the first protein for the first topic; and generating a diagnostic or prognostic model using the identified set of characteristic proteins. 26. The system of clause 25, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 27. The system of any one of clauses 25 to 26, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins. 28. The system of any one of clauses 25 to 26, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise including the first topic in the subset of topics. 29. The system of any one of clauses 25 to 28, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples. 30. The system of any one of clauses 25 to 29, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; or a non- alcoholic fatty liver disease diagnosis or prognosis. 31. A non-transitory, computer-readable medium containing instructions, that when executed by at least one processor of a system for training a proteomics-based diagnostic or prognostic model, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set ofAttorney Docket No.: 15988.0009-00304 training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in the set of characteristic proteins for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic; and generating a diagnostic or prognostic model using the identified set of characteristic proteins. 32. The non-transitory, computer-readable medium of clause 31, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 33. The non-transitory, computer-readable medium of any one of clauses 31 to 32, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins.Attorney Docket No.: 15988.0009-00304 34. The non-transitory, computer-readable medium of any one of clauses 31 to 32, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise including the first topic in the subset of topics. 35. The non-transitory, computer-readable medium of any one of clauses 31 to 34, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples. 36. The non-transitory, computer-readable medium of any one of clauses 31 to 35, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis. 37. A method comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in theAttorney Docket No.: 15988.0009-00304 set of characteristic proteins for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic; and generating a diagnostic or prognostic model using the identified set of characteristic proteins. 38. The method of clause 37, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non-negative factorization of the training protein expression dataset. 39. The method of any one of clauses 37 to 38, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins. 40. The method of any one of clauses 37 to 38, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the method further comprises including the first topic in the subset of topics. 41. The method of any one of clauses 37 to 40, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples. 42. The method of any one of clauses 37 to 41, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; or a non- alcoholic fatty liver disease diagnosis or prognosis.

[0132] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary from implementation to implementation. Certain adaptations and modifications of the described embodiments can be made. OtherAttorney Docket No.: 15988.0009-00304 embodiments can be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only. It is also intended that the sequence of steps shown in figures are only for illustrative purposes and are not intended to be limited to any particular sequence of steps. As such, those skilled in the art can appreciate that these steps can be performed in a different order while implementing the same method.

Claims

Attorney Docket No.: 15988.0009-00304 WHAT IS CLAIMED IS:

1. A proteomics-based diagnostic or prognostic system comprising: at least one processor; and at least one computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer- based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay; generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.Attorney Docket No.: 15988.0009-00304 2. The proteomics-based diagnostic or prognostic system of claim 1, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

3. The proteomics-based diagnostic or prognostic system of claim 1, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information.

4. The proteomics-based diagnostic or prognostic system of claim 1, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins.

5. The proteomics-based diagnostic or prognostic system of claim 4, wherein: determining protein contribution values for the first topic comprises: determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on:Attorney Docket No.: 15988.0009-00304 the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic.

6. The proteomics-based diagnostic or prognostic system of claim 1, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.

7. The proteomics-based diagnostic or prognostic system of claim 1, wherein: the health information includes: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis.

8. The proteomics-based diagnostic or prognostic system of claim 7, wherein: the cancer diagnosis or prognosis concerns neuroblastoma.

9. A non-transitory, computer-readable medium containing instructions that, when executed by at least one processor of a system, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer- based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins;Attorney Docket No.: 15988.0009-00304 generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay; generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.

10. The non-transitory, computer-readable medium of claim 9, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

11. The non-transitory, computer-readable medium of claim 9, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information.Attorney Docket No.: 15988.0009-00304 12. The non-transitory, computer-readable medium of claim 9, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins.

13. The non-transitory, computer-readable medium of claim 12, wherein: determining protein contribution values for the first topic comprises: determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic.

14. The non-transitory, computer-readable medium of claim 9, wherein: the training samples comprise: plasma samples;Attorney Docket No.: 15988.0009-00304 serum samples; or cell lysate samples.

15. The non-transitory, computer-readable medium of claim 9, wherein: the health information includes: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis.

16. The non-transitory, computer-readable medium of claim 15, wherein: the cancer diagnosis or prognosis concerns neuroblastoma.

17. A method comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; obtaining health information corresponding to the training samples; generating a diagnostic or prognostic model using the sample definition dataset and the health information; obtaining patient protein expression values for the set of proteins, the patient protein expression values generated using the aptamer-based assay;Attorney Docket No.: 15988.0009-00304 generating patient topic weights using the topic definition dataset and the patient protein expression values; and generating a diagnosis or prognosis by applying the patient topic weights to the diagnostic or prognostic model.

18. The method of claim 17, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

19. The method of claim 17, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model to diagnose or predict health information given topic weights using the sample definition dataset and the health information.

20. The method of claim 17, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the method further comprise: determining protein contribution values for a first topic of the set of topics; identifying a set of characteristic proteins for the first topic based on the determined protein contribution values; and including the first topic in the subset of topics based on the identified set of characteristic proteins.

21. The method of claim 20, wherein: determining protein contribution values for the first topic comprises:Attorney Docket No.: 15988.0009-00304 determining, using the topic definition dataset, a median of protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein weight value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and determining a first protein contribution value of the first protein for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic.

22. The method of any one of claim 17, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.

23. The method of claim 17, wherein: the health information includes: a cancer diagnosis or prognosis; or a non-alcoholic fatty liver disease diagnosis or prognosis.

24. The method of claim 23, wherein: the cancer diagnosis or prognosis concerns neuroblastoma.

25. A system for training a proteomics-based diagnostic or prognostic model, comprising: at least one processor; andAttorney Docket No.: 15988.0009-00304 at least one computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer- based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in the set of characteristic proteins for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic; andAttorney Docket No.: 15988.0009-00304 generating a diagnostic or prognostic model using the identified set of characteristic proteins.

26. The system of claim 25, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

27. The system of claim 25, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins.

28. The system of claim 25, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise including the first topic in the subset of topics.

29. The system of claim 25, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.

30. The system of claim 25, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; orAttorney Docket No.: 15988.0009-00304 a non-alcoholic fatty liver disease diagnosis or prognosis.

31. A non-transitory, computer-readable medium containing instructions, that when executed by at least one processor of a system for training a proteomics-based diagnostic or prognostic model, cause the system to perform operations comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in the set of characteristic proteins for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic; andAttorney Docket No.: 15988.0009-00304 generating a diagnostic or prognostic model using the identified set of characteristic proteins.

32. The non-transitory, computer-readable medium of claim 31, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

33. The non-transitory, computer-readable medium of claim 31, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins.

34. The non-transitory, computer-readable medium of claim 31, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the operations further comprise including the first topic in the subset of topics.

35. The non-transitory, computer-readable medium of claim 31, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.

36. The non-transitory, computer-readable medium of claim 31, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; orAttorney Docket No.: 15988.0009-00304 a non-alcoholic fatty liver disease diagnosis or prognosis.

37. A method comprising: obtaining a training protein expression dataset generated using an aptamer-based assay of a set of training samples, the training protein expression dataset including, for the set of training samples, training protein expression values for a set of proteins; generating a sample definition dataset and a topic definition dataset using the training protein expression dataset: the sample definition dataset including topic weights that indicate an association of each topic in a set of topics with each training sample in the set of training samples, and the topic definition dataset including protein weights that indicate an association of each protein in the set of proteins with each topic in the set of topics; identifying a set of characteristic proteins for a first topic of the set of topics, the identification comprising: determining, using the topic definition dataset, a median of the protein weights for a first protein across the set of topics; determining, using the topic definition dataset, protein value differences based on: the median protein value for the first protein; and the protein weights for the first protein across the set of topics; determining a median of the protein value differences; and including the first protein in the set of characteristic proteins for the first topic based on the median of the protein value differences and a protein value of the first protein for the first topic; andAttorney Docket No.: 15988.0009-00304 generating a diagnostic or prognostic model using the identified set of characteristic proteins.

38. The method of claim 37, wherein: the training protein expression dataset comprises a matrix; and the sample definition dataset and the topic definition dataset are generated using non- negative factorization of the training protein expression dataset.

39. The method of claim 37, wherein: the diagnostic or prognostic model comprises a regression model; and generating the diagnostic or prognostic model comprises training the regression model using protein expression values for proteins selected from the set of characteristic proteins.

40. The method of claim 37, wherein: the diagnostic or prognostic model is generated using a subset of the sample definition dataset corresponding to a subset of the set of topics; and the method further comprises including the first topic in the subset of topics.

41. The method of claim 37, wherein: the training samples comprise: plasma samples; serum samples; or cell lysate samples.

42. The method of claim 37, wherein: the diagnostic or prognostic model is trained to diagnose or predict: a cancer diagnosis or prognosis; orAttorney Docket No.: 15988.0009-00304 a non-alcoholic fatty liver disease diagnosis or prognosis.