Systems and methods for proteomic model validation

The technical solution addresses the issue of developing proteomic models by incorporating systems with processors and computer-readable media to adjust protein level measurements and generate validation datasets.

JP2025536258APending Publication Date: 2025-11-05SOMALOGIC OPERATING CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025520843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-10-11
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Aptamer-based assays for measuring protein levels in samples are prone to fluctuations due to changes in sample characteristics, processing, or assay protocols, leading to reduced performance of proteomic models and necessitating revalidation, which is hindered by the unavailability or difficulty of obtaining representative samples.

Method used

Developing proteomic models tolerant to changes in input data by estimating the effects of such fluctuations and using updated data for validation or training, incorporating systems with processors and computer-readable media to adjust protein level measurements and generate validation datasets.

Benefits of technology

Enhances the robustness of proteomic models by accounting for data variations, ensuring consistent performance across different sample contexts and protocols, improving the reliability of predictive models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536258000001_ABST
    Figure 2025536258000001_ABST
Patent Text Reader

Abstract

A method, system, and computer-readable medium for developing or validating a resilient proteomic model can estimate the impact of different contexts on input data on the model and use the estimated impact to develop or validate the proteomic model. An exemplary method for validating a proteomic model can predict protein level noise using a control dataset and a treatment dataset. The control dataset and the treatment dataset can be different datasets from the training dataset used to generate the model. A validation dataset can be generated using the training dataset and the estimated protein level noise. The validation dataset can include adjusted protein level measurements. A set of predictions can be generated by applying the validation dataset to a proteomic model trained using the training dataset. The set of predictions can be used to determine the validity of the proteomic model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Reference to Related Application) This application is the benefit of U.S. Provisional Application No. 63 / 415,978 (filing date: October 13, 2022). The benefit of which is claimed, the entire disclosure of which is incorporated herein by reference. [Background technology]

[0002] Aptamer-based assays can simultaneously measure the protein levels of multiple proteins in samples collected from patients. However, due in part to the nature of such assays, the measured protein levels may fluctuate with changes in sample characteristics, sample processing, or assay protocols. These changes may affect the results of proteomic models that use the measured protein levels as input features. Therefore, changes in sample characteristics, sample processing, and assay protocols may lead to reduced performance of such proteomic models.

[0003] Therefore, changes in sample characteristics, sample processing, or assay protocols may necessitate revalidation of the corresponding proteomic model. However, revalidation of a proteomic model may require samples representing the full range of expected inputs. Such samples may be unavailable or difficult to obtain. The difficulty of obtaining such samples may hinder the development of proteomic models that are tolerant to changes in sample characteristics, sample processing, or assay protocols. Summary of the Invention [Means for solving the problem]

[0004] Certain embodiments of the present disclosure relate to developing proteomic models that are tolerant to changes in input data, such as sample characteristics, sample processing, or assay protocols. According to embodiments of the present disclosure, the effects of such changes on the input data can be estimated, and the estimated effects can be used to generate updated data for validating or training the proteomic model.

[0005] An embodiment of the present disclosure includes a system for validating a proteomic model. The system includes at least one processor and at least one non-transitory computer-readable medium having instructions executed by the at least one processor. The computer-readable medium includes instructions that, when executed by the at least one processor, cause the system to perform operations. The operations include acquiring a control dataset including protein level measurements. The operations further include acquiring a treatment dataset including protein level measurements. The operations further include acquiring a training dataset including protein level measurements. The operations further include estimating protein level noise using the control dataset and the treatment dataset. The operations further include generating a validation dataset using the training dataset and the estimated protein level noise. The validation dataset includes adjusted protein level measurements. The operations further include generating a set of predictions by applying the validation dataset to a proteomic model trained using the training dataset. The operations further include determining a validity indicator using the set of predictions and providing the validity indicator.

[0006] An embodiment of the present disclosure includes a system for developing a proteomic model. The system includes at least one processor and at least one non-transitory computer-readable medium. The computer-readable medium includes instructions that, when executed by at least one processor, cause the system to perform operations. The operations include obtaining a trained proteomic model to generate predictions based on learning of protein level measurements generated from training aliquots acquired in a first context. The operations further include generating therapeutic protein level measurements from therapeutic aliquots or samples acquired in a plurality of second contexts. The operations further include determining a dependency of the proteomic model on differences between the plurality of second contexts. The operations further include updating the proteomic model based on the determined dependency.

[0007] Embodiments of the present disclosure further include corresponding methods for validating and / or developing proteomic models, and non-transitory computer-readable media containing executable instructions for performing such methods.

[0008] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the scope of the claims. Other systems, methods, and computer-readable media are also described herein. [Brief explanation of the drawings]

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles of the disclosure.

[0010] [Figure 1] FIG. 1 illustrates an exemplary pipeline for developing, validating, and deploying proteomic models for predicting health information using aptamer-based blood tests, according to an embodiment of the present disclosure. [Figure 2A] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2B]1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2C] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2D] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2E] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2F] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2G] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 2H] 1 provides a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay according to an embodiment of the present disclosure. [Figure 3A] 1 illustrates exemplary challenges that arise in validating proteomic models for predicting health information using aptamer-based serum or plasma tests, according to embodiments of the present disclosure. [Figure 3B] 1 illustrates exemplary challenges that arise in validating proteomic models for predicting health information using aptamer-based serum or plasma tests, according to embodiments of the present disclosure. [Figure 4A] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 4B] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 4C] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 4D] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 4E] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 4F] 1 illustrates an exemplary model resilience challenge that arises in training a proteomic model for predicting health information using an aptamer-based blood test, according to an embodiment of the present disclosure. [Figure 5] FIG. 5 illustrates an exemplary process for validating a proteomic model according to an embodiment of the present disclosure. [Figure 6] FIG. 6 illustrates an exemplary process for developing a proteomic model according to an embodiment of the present disclosure. [Figure 7] 7A-7D show the effect of differences in sample handling on the output of a binary endpoint proteomics model, according to an embodiment of the present disclosure. [Figure 8] 8A-8D show the effect of sample processing differences on the output of a continuous endpoint proteomics model according to an embodiment of the present disclosure. [Figure 9] 9A-9D show the effect of differences in sample processing on the predictive performance of predictions made by proteomic models, according to embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of exemplary embodiments of the present disclosure. However, it will be understood by those skilled in the art that the principles of the exemplary embodiments may be practiced without all specific details. Well-known methods, procedures, and components will not be described in detail herein so as not to obscure the principles of the exemplary embodiments. Unless explicitly stated, the exemplary methods and processes described herein are not constrained to a particular order or sequence, or to a particular system configuration. Furthermore, some of the embodiments or elements thereof described herein may co-occur or be performed simultaneously, at the same time. Reference will now be made to the details of the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings.

[0012] As used herein, proteomics may relate to or involve the quantitative evaluation of proteins present in samples obtained from humans or non-human animals. Proteomics may relate to classification or prediction derived from large-scale (e.g., hundreds to tens of thousands) measurements of protein levels (e.g., presence, amount, concentration, etc.). In this way, proteomics may be similar to genomics, but applied to the field of proteins.

[0013] As used herein, a predictive model may be a model that maps a set of inputs to an output value or category. In some embodiments, a predictive model may be a machine learning model that may be trained using training data to generate an appropriate output. A predictive model may be trained using supervised, semi-supervised, unsupervised, or reinforcement learning methods. In some embodiments, a predictive model may be a statistical model that uses training data to estimate the relationship between an independent variable (e.g., protein levels in a biological sample) and a dependent variable (e.g., health information value). Exemplary predictive models include regression models (e.g., penalized regression models), support vector machines, decision trees or forests, These include linear discriminant analysis, clustering models, nearest neighbor models, neural network models, and ensemble models that include any one or more of these models.

[0014] As used herein, a proteomic model may be a predictive model configured to receive protein level measurements and output a classification or prediction based at least in part on those measurements.

[0015] Health information may include the probability of a health outcome (e.g., cardiovascular disease, dementia, renal disease, etc.) occurring within a specific time frame (e.g., within 6 months, 1 year, 4 years, 10 years, 20 years, or more), indicators of current health status (e.g., body fat percentage, basal metabolic rate, lean muscle mass, aerobic fitness, visceral fat content, glomerular filtration rate, presence or absence of excess liver fat, glucose tolerance, etc.), behavioral health predictions (e.g., predicted weekly alcohol consumption, etc.), etc.

[0016] According to embodiments of the present disclosure, providing a model, data, or instructions may include providing such model, data, or instructions directly (e.g., between a source and a target) or indirectly (e.g., via an intermediate means). Providing a model, data, or instructions may include providing the model, data, or instructions by value or by referencing a memory location of the model, data, or instructions.

[0017] According to embodiments of the present disclosure, performance measures can be used to quantify the performance of a predictive model. In some embodiments, performance measures can quantify the agreement between two predictions (e.g., assessing reproducibility or the impact of differences in input data on the predicted output). Such performance measures can include Lin's concordance correlation coefficient, Pearson's correlation coefficient, or other suitable correlation measures. In some embodiments, performance measures can quantify the agreement between a prediction and a ground truth. In some embodiments, such performance measures can quantify this agreement in terms of the relationship between the true positive rate and the false positive rate, depending on variations in parameters of the predictive model (e.g., the discrimination threshold, etc.). Such measures can include a receiver operating characteristic (ROC) curve or a differential performance measure such as the area under the ROC curve (AUC). In some embodiments, such performance measures can include or depend on the number of true positives, true negatives, false positives, or false negatives for a classification task. In some embodiments, such measures can include sensitivity, specificity, precision, false negative rate, false positive rate, prevalence, accuracy, F1 score, or other such measures. In some embodiments, such performance measures may include or rely on per-class comparisons between predictions and a ground truth, such as a confusion matrix.

[0018] FIG. 1 illustrates an exemplary pipeline 100 for developing, validating, and deploying proteomic models for predicting health information using aptamer-based blood tests, according to an embodiment of the present disclosure. The pipeline 100 may include components from which data is acquired, such as a measurement system 101 or a record 103. The pipeline 100 may include components such as a data input engine 110 or a dataset generator 120 for collecting and preparing the acquired data. The pipeline 100 may include components such as a data storage 140 and a model storage 130 for storing the prepared data and machine learning models. The pipeline 100 may include a learning engine 160 for generating a trained machine learning model using the acquired data and model. The pipeline 100 may include a prediction engine 170 for using the trained machine learning model to predict health information. The components of the pipeline 100 may be managed and configured via a user device 180. The user device 180 may receive and use the output of other components, original data or prepared data, and other data. The pipeline 100 may be used to display trained or untrained models, as well as predicted health information. Embodiments of the present disclosure allow a user to interact with components of the pipeline 100 to perform model validation and development. In some embodiments, a user may interact with the learning engine 160 (or another suitable component of the pipeline 100) to perform model validation, as described with reference to FIG. 5, or model development, as described with reference to FIG. 6. Overall, the pipeline 100, as disclosed herein, may provide a convenient and scalable platform for developing, validating, and deploying proteomic models.

[0019] The measurement system 101 is a device suitable for obtaining data indicative of the presence or concentration of proteins in a biological sample. For ease of explanation, the measurement system 101 is described herein as a microarray scanner. The scanner can be configured to measure fluorescence at predetermined locations corresponding to different proteins on a microarray. The intensity of the fluorescence can indicate the concentration of the protein in the biological sample used to prepare the microarray. In some cases, multiple locations correspond to the same protein. Fluorescence intensity data from multiple locations can be combined to more accurately estimate the concentration of the protein in the sample. The pipeline 100 is not limited to embodiments in which the measurement system 101 is a microarray scanner. Furthermore, in some embodiments, rather than providing data directly to the data input engine 110, the measurement system 101 can provide data to records 103. The data input engine 110 can then retrieve this data from records 103.

[0020] In disclosed embodiments, record 103 includes one or more storage locations for data usable by pipeline 100 to predict healthcare outcomes. In some embodiments, this data can include fluorescence intensity data generated by measurement system 101 from a biological sample. In various embodiments, this data can include medical record information for the patient from whom the biological sample was obtained. This medical record information can include medical records, case notes, and request information (e.g., related to the sample or predictions performed by pipeline 100) provided by a physician, clinician, or the like. In some embodiments, this medical record information can include class or data label information corresponding to the sample (e.g., used to generate a training dataset).

[0021] According to embodiments of the present disclosure, medical record information can be associated with health information. For example, in a cancer screening or diagnostic setting, the medical record information can include data related to cancer type, cancer prevalence, cancer prognosis, or cancer stage.

[0022] According to embodiments of the present disclosure, the medical record information may include information suitable for use in generating a personalized predictive model. For example, the medical record information may indicate patient demographic or health characteristics, such as the patient's age, gender, race / ethnicity, height / weight, etc. As a further example, the medical record information may indicate the patient's medical history, such as medical history, medical treatments, clinical visits, and surgical history. As an additional example, the medical record information may indicate behavioral factors that may affect the patient's health, such as smoking history, alcohol consumption, medication use, and diet type. As a further example, the medical record information may include data obtained from biological sampling from the patient, such as data regarding blood, plasma, serum, or urine samples. As a further example, the medical record information may indicate family history data, genetic data, or immunological data.

[0023] According to an embodiment of the present disclosure, the data input engine 110 retrieves data from various data sources (e.g., measurement systems 101, records 103, or other suitable data sources) and processes the data for use by other components of the pipeline 100. In some embodiments, the data input engine 110 can include a data extractor 111, a data converter 113, and a data loader 115.

[0024] According to embodiments of the present disclosure, the data extractor 111 can receive or retrieve data from the measurement system 101, the record 103, or other suitable data sources. As described herein, the measurement system 101 can be a diagnostic system configured to acquire fluorescence values ​​corresponding to the presence or concentration of a protein in a biological sample. Similarly, the record 103 can be one or more databases or data storage locations containing data about the biological sample used to generate the fluorescence values. Embodiments of the present disclosure are not limited to any particular format of the acquired data or method for acquiring this data. For example, the acquired data can be or include structured or unstructured data. The data extractor 111 can interact with various data sources, receive or retrieve relevant data, and provide that data to the data converter 113.

[0025] According to embodiments of the present disclosure, the data converter 113 can receive data from the data extractor 111 and process the data into a standard format. In some embodiments, the data converter 113 can normalize data, such as dates or numeric values, based on specific units of measurement. For example, the measurement system 101 can store dates in a day-month-year format, while the records 103 can include records that store dates in a year-month-day format or various records that store measurements in different units (e.g., weight (kilograms or pounds), fluorescence (absolute or log10), etc.). In this example, the data converter 113 can modify the data provided via the data extractor 111 into a consistent date format or standardized unit format, respectively. Thus, the data converter 113 effectively straightens the data provided via the data extractor 111 so that all of the data, although originating from various sources, is obtained in a consistent format.

[0026] Additionally, the data converter 113 can extract additional data points from the data. For example, the data converter 113 can process dates in year-month-day format by extracting separate data fields for year-month-day. The data converter 113 can perform other linear and non-linear transformations (e.g., logarithmic transformation of continuous data), as well as extractions on categorical and numeric data, such as normalizing and centering the data about its mean value. The data converter 113 can provide the transformed or extracted data to the data loader 115.

[0027] According to embodiments of the present disclosure, data loader 115 can receive normalized data from data transformer 113. Data loader 115 can merge the data in various formats depending on the particular requirements of dataset generator 120. Data loader 115 can then provide the processed data to dataset generator 120 (or to a suitable data store from which dataset generator 120 can retrieve the data).

[0028] According to embodiments of the present disclosure, dataset generator 120 may be configured to generate a dataset from data processed by data input engine 110. In some embodiments, learning engine 160 or prediction engine 170 may be configured to expect a dataset having a particular structure. Dataset generator 120 may be configured to format the data into that particular structure. For example, dataset generator 120 may collect observations and associate the observations with corresponding learning labels or metadata. , observations and metadata can be configured to be stored in data storage 140.

[0029] In some embodiments, the dataset generator 120 can be configured to extract features from the data received from the data input engine 110. Here, a feature can be a property or characteristic of a phenomenon. For example, the presence or absence of a fluorescence value above a threshold at a location corresponding to a particular protein can be a feature. As a further example, the average fluorescence value across multiple samples on a biochip can be a feature. Features can also be determined based on domain, categorical data type, or many other factors associated with the data stored in the data structure. Furthermore, a feature can represent information about multiple data records in a dataset or information about a single category within a data record. Furthermore, multiple features can be generated to represent the same data.

[0030] In some embodiments, the dataset generator 120 may include a classifier 121 and an annotator 123. According to embodiments of the present disclosure, the classifier 121 may be configured to create classes from the data provided by the dataset generator 120. In some embodiments, such classes may be data labels used to train a predictive model. For example, a set of fluorescence data indicating protein levels may be associated with a patient's medical records. In this example, the classifier may extract health information from the medical records. For example, the classifier may determine that a patient experienced a cardiovascular event within a certain time period after acquisition of the blood sample used to generate the fluorescence data. As a further example, the classifier may identify a patient's glomerular filtration rate from the medical records. In some embodiments, the classifier is or may include a natural language processing engine.

[0031] According to embodiments of the present disclosure, the classifier 121 may be configured to accept classifications provided by a user via the user device 180. For example, the classifier 121 may be configured to provide data (or metadata about the data) received from the data input engine 110 to the user device 180 for display. In response, the classifier 121 may receive class information. For example, the classifier 121 may provide medical records associated with a set of fluorescence data for display on the user device 180. Such functionality is not limited to medical records, but may also be other information used to generate learned labels for the data.

[0032] According to embodiments of the present disclosure, the annotator 123 may be configured to receive data and associated value or class information from the dataset generator. The annotator 123 may be configured to create appropriately formatted entries that associate value or class information with the data. For example, given an array or matrix of fluorescence values ​​and glomerular filtration rate, the annotator 123 may create an object that includes a "response_value" key and a "protein_data_input" key. The glomerular filtration rate may be stored using the "response_value" key and the array or matrix of fluorescence values, "protein_data_input" key. In this example, the learning engine 160 or the prediction engine 170 may predict or be configured to predict observations having such a format. As a further example, given a set of fluorescence values ​​obtained from a blood sample taken from a patient and the date the patient suffered a heart attack, the annotator 123 may create a relational database with a column corresponding to the patient, a column storing whether the patient suffered a heart attack, another column storing when the patient suffered a heart attack, and remaining columns storing fluorescence values.

[0033] According to embodiments of the present disclosure, model storage 130 may be a storage location for predictive models that can be used by learning engine 160 or prediction engine 170. Embodiments of the present disclosure are not limited to any particular implementation of data storage 140. According to embodiments of the present disclosure, data storage 140 may be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0034] According to embodiments of the present disclosure, data storage 140 may be a storage location for prepared datasets that can be used by learning engine 160 or prediction engine 170. Embodiments of the present disclosure are not limited to any particular implementation of data storage 140. According to embodiments of the present disclosure, data storage 140 may be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0035] According to embodiments of the present disclosure, data / model selector 150 may be configured to access model storage 130 or data storage 140 to retrieve predictive models or datasets, respectively. In some embodiments, data / model selector 150 may provide an abstraction layer for learning engine 160 or prediction engine 170. In some embodiments, data / model selector 150 may be configured to control access to model storage 130 or data storage 140.

[0036] According to embodiments of the present disclosure, learning engine 160 may be configured to learn or create and learn a predictive model. Learning engine 160 may be configured to retrieve an existing model from model storage 130 or a training dataset from data storage 140. In some embodiments, learning engine 160 may be configured to interact with data / model selector 150 to retrieve an existing model or a training dataset. Learning engine 160 may be configured to store the learned predictive model in model storage 130. In some embodiments, learning engine 160 may be configured to interact with data / model selector 150 to store the learned predictive model in model storage 130.

[0037] According to embodiments of the present disclosure, learning engine 160 may include model trainer 161 and model evaluation 163. Learning engine 160 may be configured to train a predictive model using model trainer 161 and then determine values ​​of performance measures for the predictive model using model evaluation 163. In some embodiments, learning engine 160 may automatically update the trained predictive model based on the values ​​of the performance measures. In various embodiments, learning engine 160 may update the trained predictive model in response to user input provided via user device 180. Updating the predictive model may include one or more of performing additional training (e.g., using an existing training dataset or a different training dataset), modifying the model (e.g., changing input features used by the model, changing the model architecture, etc.), or changing the training environment (e.g., changing training hyperparameters, changing the training dataset to the training, cross-validation, holdout portion, etc.).

[0038] According to embodiments of the present disclosure, as described herein, the learning engine 160 determines the values ​​of the performance measures of the model using data acquired in different contexts. The values ​​of such performance measures may be displayed to a user through a user device 180, through which the user can interact with the learning engine 160 and update the model, as described herein.

[0039] According to embodiments of the present disclosure, model trainer 161 can create or train a predictive model. Model trainer 161 can create or train a predictive model as instructed by learning engine 160. For example, learning engine 160 can instruct model trainer 161 to create and train a support vector machine using the training portion of the training dataset. Model trainer 161 can then create and train a support vector machine and return the trained support vectors to learning engine 160. As a further example, learning engine 160 can instruct model trainer 161 to create and train a penalized regression model. The learning engine can configure model trainer 161 with a type of penalized regression model (e.g., ridge regression, lasso regression, elastic net, or another suitable type), parameter values ​​for the penalized regression (e.g., a lambda value that weights the sum of squared coefficient values), and the training portion of the training dataset. Model trainer 161 can then create and train the penalized regression model and return the trained penalized regression model to learning engine 160. As a further example, learning engine 160 can instruct model trainer 161 to train a random forest model. Learning engine 160 can provide hyperparameters such as the size of each bootstrap sample, the number of features to consider at each split, the depth of each decision tree, and the number of decision trees in the random forest. Learning engine 160 can provide a training portion of a training dataset. Model trainer 161 can then create and train a random forest model and return the trained random forest model to learning engine 160.

[0040] According to embodiments of the present disclosure, model evaluation 163 can evaluate the model trained by model trainer 161. Learning engine 160 can provide model and cross-validation or holdout portions of the training dataset to model evaluation 163. In some embodiments, learning engine 160 can specify one or more performance measures for evaluation by model evaluation 163. In various embodiments, model evaluation 163 can be configured with a predetermined or default set of performance measures. In some embodiments, the performance measures can include a confusion matrix, mean squared error, mean absolute error, sensitivity or selectivity, a receiver operating characteristic curve or area under such a curve, precision and recall, an F-measure, or any other suitable performance measure.

[0041] According to embodiments of the present disclosure, prediction engine 170 can be configured to predict health information using a patient dataset and a learned predictive model. In some embodiments, prediction engine 170 can retrieve the learned predictive model from model storage 130. In some embodiments, prediction engine 170 can retrieve the patient dataset from data storage 140. In some embodiments, prediction engine 170 can retrieve the patient dataset (or a portion thereof) from another data storage location. This alternative data storage location can be associated with another entity or user. For example, prediction engine 170 can receive or retrieve the patient dataset from a healthcare system separate from the entity controlling prediction engine 170. In some embodiments, prediction engine 170 can retrieve the model or data using data / model selector 150.

[0042] According to an embodiment of the present disclosure, the prediction engine 170 can apply the learned predictive model to a patient dataset to predict the patient's health information. The health information may be provided by the prediction engine 170 to the user device 180. The health information may be stored on a computing device associated with the pipeline 100 or provided to another system.

[0043] According to embodiments of the present disclosure, user device 180 can provide a user interface for interacting with other components of pipeline 100. The user interface can be a graphical user interface. The user interface allows a user to configure data input engine 110 to extract, transform, and load data according to user specifications. The user interface allows a user to specify how transformed data received by dataset generator 120 is converted into labeled training data (or patient data suitable for prediction). In some embodiments, the user interface allows a user to interact with dataset generator 120 to manually or semi-manually label or annotate training data. In some embodiments, the user interface allows a user to interact with data / model selector 150 to manage data or models stored in model storage 130 or data storage 140. In some embodiments, the user interface allows a user to interact with data / model selector 150 to push data or models to learning engine 160 for learning or to prediction engine 170 for prediction. In some embodiments, a user interface allows a user to interact with learning engine 160 to create or select a predictive model for training, to create or select a dataset for use in training a model, or to select training parameters or hyperparameters. In some embodiments, a user interface allows a user to interact with learning engine 160 to display information related to model training (e.g., values ​​of performance measures, changes in loss function values ​​during training, or other learning information). In some embodiments, a user interface allows a user to interact with prediction engine 170 to select a learning model and patient data for use in predicting health information.In some embodiments, a user interface allows a user to interact with the prediction engine 170 to display the health information, store the health information on a computing device, or transmit the health information to another system.

[0044] Components of pipeline 100 may be implemented using one or more computing devices. Such computing devices may include tablets, laptops, desktops, workstations, computing clusters, or cloud computing platforms. In some embodiments, components of pipeline 100 may be implemented using a cloud computing platform. For example, one or more of data input engine 110, dataset generator 120, data / model selector 150, learning engine 160, and prediction engine 170 may be implemented on a cloud computing platform. In some embodiments, components of pipeline 100 may be supplemented using on-premises systems. For example, measurement system 101, records 103, or user device 180 may be on-premises systems or may be hosted on on-premises systems. As a further example, model storage 130 or data storage 140 may be on-premises systems or may be hosted on on-premises systems.

[0045] The components of pipeline 100 may communicate using any suitable method. In some embodiments, two or more components of pipeline 100 may be implemented as microservices or web services. Such components may communicate using messages sent over a computer network. A message may be implemented using SOAP, XML, HTTP, JSON, RCP, or any other suitable format. In some embodiments, two or more components of pipeline 100 may be implemented as software, hardware, or combined software / hardware modules. Such components may communicate using data or instructions written to or read from memory (e.g., shared memory), function calls, or any other suitable communication method.

[0046] The particular structure of the pipeline 100 is not intended to be limiting. According to embodiments of the present disclosure, any two or more of the records 103, the model storage 130, or the data storage 140 may be combined or hosted on the same computing device. According to embodiments of the present disclosure, the data input engine 110 and the dataset generator 120 may be omitted from the pipeline 100. In such embodiments, datasets formatted and configured for use by the learning engine 160 or the prediction engine 170 may be stored in the data storage 140 by another system or using another method. According to embodiments of the present disclosure, the data input engine 110 and the dataset generator 120 may be combined. In such embodiments, data extraction, transformation, and loading may be combined with feature extraction, annotation, and classification. According to embodiments of the present disclosure, the data / model selector 150 may be combined with one or more of the learning engine 160 and the prediction engine 170. For example, the learning engine 160 or the prediction engine 170 may include functionality for retrieving selected data or models from the model storage 130 or the data storage 140.

[0047] Although one user device 180 is shown in the figure, multiple user devices may be provided. Different user devices may be associated with different entities or different users with different roles. For example, a user device 180 may be associated with a software engineer or data scientist developing a test, while another user device may be associated with a clinician using the test.

[0048] User device 180 may be combined with one or more other components of pipeline 100. In some embodiments, user device 180 and at least one of data / model selector 150, learning engine 160, or prediction engine 170 may be implemented by the same computing device. In some embodiments, user device 180 and at least one of model storage 130 or data storage 140 may be implemented by the same computing device.

[0049] Here, the pipeline 100 can be integrated into a method for treating patients with a particular health condition. The prediction engine 170 can use the learned predictive model and input data obtained from patient samples to determine the patient's risk of experiencing a negative health outcome (e.g., a cardiovascular event or recurrent cardiovascular events within the next four years, dementia within the next 20 years, or death within one year with stable heart failure and reduced ejection fraction or preserved ejection fraction). If the patient has a risk greater than (or potentially equal to) the dependent threshold for the health outcome, the patient can be treated or monitored according to a first, more aggressive, or intensive protocol. If the patient has a risk less than (or potentially equal to) the dependent threshold for the health outcome, the patient can be treated or monitored according to a second, less invasive, or intensive protocol.

[0050] 2A-2H show a high-level depiction of the stages in an exemplary aptamer-based serum or plasma assay 200 according to an embodiment of the present disclosure. The assay 200 is Suitable input data can be generated for training a predictive model or for predicting health information using the trained predictive model. Assay 200 can be performed, at least in part, using a test system, such as measurement system 101 of pipeline 100.

[0051] According to embodiments of the present disclosure, assay 200 can quantitatively convert the availability of protein epitopes in a biological sample into specific DNA signals. Generally, assay 200 can use SOMAmer® (Slow Off-rate Modified Aptamer) reagents, which contain short, single-stranded DNA sequences incorporating hydrophobic modifications. Assay 200 can measure native proteins in complex matrices by converting available binding sites on individual proteins into corresponding SOMAmer reagent concentrations, which are then quantified by hybridization to microarrays. In this way, this test exploits the dual nature of SOMAmer reagents as both protein affinity binding reagents with defined three-dimensional structures and unique nucleotide sequences recognizable by specific DNA hybridization probes. Thus, relative epitope concentrations can be converted into measurable nucleic acid signals that can be quantified using DNA hybridization microarrays.

[0052] According to embodiments of the present disclosure, a preferred version of assay 200 can quantify the relative levels of abundant proteins in plasma over a 10-log range. This test version can measure up to 1,000, 3,000, 5,000, 7,000, 10,000, or more unique protein analytes. This test can be performed on small sample volumes (e.g., 10 microliters, 20 microliters, 40 microliters, 100 microliters, 200 microliters, 400 microliters, 1 milliliter, or more).

[0053] Here, SOMAmer reagents may be selected against proteins in their native folded structure. Thus, such reagents may require an intact third protein structure for binding. Thus, unfolded and denatured, and therefore inactive, proteins may not be detected by SOMAmer reagents (or may be detected with reduced or variable sensitivity).

[0054] As shown in Figure 2A, SOMAmer reagents can be synthesized with a fluorophore, a photocleavable linker, and biotin. Then, as shown in Figure 2B, SOMAmer reagents bound to streptavidin beads can be used to capture proteins from a complex mixture of proteins in a biological sample (e.g., a serum or plasma sample). Next, as shown in Figure 2C, unbound proteins can be washed away, and bound proteins can be tagged with biotin. Next, as shown in Figure 2D, electromagnetic radiation (e.g., ultraviolet light) can be applied to the solution to cleave the photocleavable linker, releasing the protein complex and bound SOMAmer back into solution. As shown in Figure 2E, nonspecific complexes can dissociate from their corresponding SOMAmers, while specific complexes remain bound. Next, as shown in Figure 2F, a polyanionic competitor can be added to the solution. The polyanionic competitor can prevent rebinding of nonspecific complexes. As shown in Figure 2G, biotinylated proteins (and the bound SOMAmer reagent) can be captured on streptavidin beads. The beads and bound proteins can be separated from the solution or concentrated. The SOMAmer reagent can then be released from the protein complex by denaturing the protein, as shown in Figure 2H. The fluorophore can be measured after hybridization to a complementary sequence on a microarray chip. The fluorescence intensity detected on the microarray can be related to the amount of available epitope in the original sample.

[0055] Assay 200 is intended to be illustrative. The disclosed systems and methods are not limited to tests having these specific steps. In some embodiments, other aptamers (or even other classes of components) can be used to bind protein complexes. Alternative methods for inhibiting nonspecific binding can be used in place of polyanionic competitors. Alternative methods for separating protein-compound complexes can be used instead of capturing protein-compound complexes on streptavidin beads. Alternative indicators of protein levels can be measured in place of fluorescent measurements on microarray chips. However, such alternative methods may still present the technical challenges described herein. Therefore, such alternative systems and methods can benefit from the disclosed technical solutions.

[0056] 3A and 3B illustrate exemplary challenges that arise in validating a proteomic model for predicting health information using an aptamer-based serum or plasma test, according to embodiments of the present disclosure. In some embodiments, the technical challenge may be ensuring consistent predictions across contexts. This challenge may arise when a proteomic model developed using input data obtained in a first context needs to be validated for use with input data obtained in another context.

[0057] According to embodiments of the present disclosure, contextual differences herein include differences between samples, differences in sample processing, or differences in assay protocols. According to embodiments of the present disclosure, differences between samples may include the presence or absence of interfering agents in the sample, the use of citrate plasma versus EDTA plasma, the use of serum versus plasma, the fasting / fed state of the patient providing the sample, or other variations in sample characteristics that may affect the validity of the proteomic predictive model. Embodiments of the present disclosure are not limited to any particular interfering agent. In various embodiments, interfering agents may include nonsteroidal anti-inflammatory drugs (NSAIDs), birth control medications, blood pressure medications, mental health medications (e.g., antidepressants, antipsychotics, antianxiety or hypnotics, mood stabilizers, stimulants, etc.), cholesterol medications, asthma medications, diabetes medications, thyroid medications, antiviral or antibiotic medications, or other commonly used medications. According to embodiments of the present disclosure, differences in sample processing can include differences in the time between sample collection and sample freezing, the freezing temperature of the sample, the duration at the freezing temperature, the number of freeze / thaw cycles of the sample, the spin time of plasma samples, the clotting or decanting time of serum samples, or other variations in sample processing conditions that may affect the validity of a proteomic predictive model. According to embodiments of the present disclosure, differences in assay protocols can include differences in the reagents used, different dilutions of the same reagents, the addition or deletion of steps in the assay protocol, or differences in the device used to perform the assay protocol. For example, a first version of the assay protocol can be adapted to detect the concentration of 5,000 proteins, and a second version of the assay protocol can be adapted to detect the concentration of 7,000 proteins. These two versions of the assay protocol can use different sets of aptamers, different microarray chips, and potentially different microarray scanners.

[0058] According to embodiments of the present disclosure, appropriate input data acquired in the second context may not be available. The original samples used to generate the original input data may be missing, depleted, or deteriorated over time. Furthermore, currently available samples may differ from the original samples. For example, the original samples may include samples from patients with the health condition that the predictive proteomic model was developed to detect. For example, the original samples may be acquired over a long period of time through collaboration with a medical center that specializes in treating patients with the condition that the predictive proteomic model was developed to detect. However, such patients may be extremely rare in the general population. Also, currently available samples can be obtained from the general population. For example, currently available samples can be obtained from routine blood draws at community health centers. Therefore, currently available samples may not include patients with conditions that predictive proteomic models have previously been developed to detect.

[0059] FIG. 3A shows an exemplary correlation between the output of a predictive proteomics model for input data obtained according to two different assay protocols. In this example, the output indicates the risk of death within one year for individuals with a rare disease. The two different assay protocols are a 5000-protein protocol and a 7000-protein protocol. The 7000-protein protocol includes 5000 proteins and an additional 2000 proteins. A predictive proteomics model is developed using the input data obtained according to the 5000-protein protocol. The input data obtained according to the 7000-protein protocol can be input into the predictive proteomics model by truncating the input to include only the shared 5000 proteins.

[0060] In this example, a set of samples collected from patients with a rare disease is available. Two sets of input data can be generated for each sample: one set following a 5000-protein protocol and one set following a 7000-protein protocol. These sets of input data can be applied to a predictive proteomic model to generate a probability of death within one year. As can be seen in Figure 3A, these predicted probabilities are highly correlated, with Lin's concordance correlation coefficient (CCC) of 0.89.

[0061] Figure 3B shows an exemplary correlation between the output of the same predictive proteomics model for the same 5000-protein and 7000-protein protocols. In this example, a set of samples taken from patients with a rare disease is not available. Instead, samples from healthy normal patients are used. As described with respect to Figure 3A, two sets of input data can be generated for each sample: one set according to the 5000-protein protocol and one set according to the 7000-protein protocol. These sets of input data can be applied to the predictive proteomics model to generate a probability of death within one year. As can be seen from Figure 3B, these predicted probabilities are poorly correlated, with a CCC of 0.31. is.

[0062] Here, this lack of correlation arises from the severe limitation of the range of predicted probabilities. The predictive proteomics model correctly finds that a healthy normal patient has an extremely low probability of dying from the disease within one year. However, as a result, the range of predicted probabilities shown in Figure 3B is approximately 20 times smaller than the range of predicted probabilities shown in Figure 3A. Therefore, naive testing using a set of available samples may underestimate the reliability of the predictive proteomics model when used with the 7000-protein protocol.

[0063] 4A-4F illustrate exemplary model resilience challenges that arise in training a proteomic model for predicting health status using an aptamer-based blood test, according to embodiments of the present disclosure. As can be understood from the description of the exemplary SOMAmer-based blood test with respect to FIGS. 2A-2H, aptamer-based blood tests can be sensitive to changes in sample processing that can affect protein form or concentration (e.g., due to degradation over time, reaction with other sample components, etc.). Changes in protein levels have a clear impact on measured protein levels. Similarly, any denaturation of a protein affects the ability of an aptamer complex to bind to such a protein, affecting measured protein levels.

[0064] Therefore, sample processing conditions can affect the output of proteomic models that take protein levels as input. The dependence of the predicted output on sample processing conditions may be independent of the predictive power of the model. For example, two proteomic models may be able to predict similarly, but the first model may show substantial variance with respect to a particular variation in sample processing, while the second model may be resilient to this particular variation. Such differences may arise from the specific proteins relied upon by the different models. For example, two proteins may provide similar or correlated information about a health condition. A parsimonious model may rely on one protein but not both. However, one of the two proteins may be much more stable than the other. Therefore, even though two proteins may be equivalent from a predictive perspective, a resilient model may be designed to rely on the more stable and less stable protein.

[0065] 4A and 4B show the dependence of two proteomic models trained to predict the likelihood of having chronic kidney disease on sample processing conditions, according to an embodiment of the present disclosure. In this study, the variation in sample processing is between sample collection and shipping the sample on dry ice to the testing laboratory. Multiple healthy normal samples are collected, and six different times between sample collection and shipping are investigated. For each sample, multiple aliquots are prepared, each corresponding to one of six variations in shipping time. Two predictive models are trained here.

[0066] 4A shows the dependence of the predicted probability of having chronic kidney disease on the time to shipment of the first proteomic model. An increase in the predicted probability of having chronic kidney disease is observed with increasing time to shipment. There is an increase in the median predicted likelihood and the appearance of significant outliers in the predicted values. Here, the clinician does not indicate a delay in shipping the sample to the laboratory. Therefore, the laboratory may provide an erroneous prediction due to the dependency of the predicted risk on the shipping time.

[0067] Figure 4B shows the dependence of the predictability of chronic kidney disease on time to delivery for the second proteomic model. This second proteomic model is developed by updating the first proteomic model based on the determined sensitivity of the first proteomic model to differences in time to delivery. The contribution to the model from proteins showing sensitivity to differences in time to delivery is reduced compared to the first model. Here, the second predictive proteomic model shows a smaller increase in predicted risk than the first predictive proteomic model. Furthermore, the variability of the predicted risk is reduced, a change that appears on the plot as a slight increase in the interquartile range combined with a decrease in the number and magnitude of outliers.

[0068] Additionally, evaluating external samples in addition to variations resulting from sample processing may identify issues with overfitting of the learning model. Figure 4C shows the variation in predicted risk over time for the first proteomic model as chronic kidney disease progresses. In this example, the predicted risk is determined using repeated samples over time for the same patient. The samples are obtained at multiple intervals over a 12-month period. Here, patients who are inherently at elevated risk for developing chronic kidney disease have a stable risk across the test interval. As observed in Figure 4C, some patients experience dramatic changes in predicted risk over time, potentially due to overfitting in the original model, leading to unstable predictions for the external theta set.

[0069] FIG. 4D shows the variation in predicted relative risk of developing chronic kidney disease over time for the second proteomic model. The second model shows reduced variance in predicted risk. This model accounts for at least the sample-to-sample variation that contributes to the first prediction model. The first model can be generated by identifying one protein at a time. The contribution of the protein to the first model can be reduced. In some embodiments, the first model can be retrained to generate the second model.

[0070] In addition to the variation resulting from sample processing, the measured protein levels may exhibit process variation. Some measured protein levels may exhibit greater measurement variation, resulting in overfitting in the model. A proteomic model that relies on the protein level values ​​of such proteins may predict greater variation than a proteomic model that does not rely on the protein level values, even between aliquots taken from the same sample.

[0071] Figure 4E shows risk values ​​determined using the first proteomics model for two different aliquots of the same sample. In this example, disease diagnosis depends on the predicted risk value. Values ​​indicated by open circles give different diagnoses for different aliquots of the sample (e.g., the upper left quadrant and the lower right quadrant indicate a negative diagnosis for one aliquot and a positive diagnosis for another aliquot). Values ​​indicated by closed circles indicate that the same diagnosis is given for both aliquots of the sample.

[0072] Figure 4F shows risk values ​​determined using a second proteomic model for two different aliquots of the same sample. Compared to the predictions in Figure 4E, the number of samples showing differences in diagnosis between aliquots of the same sample is significantly lower. A second proteomic model can be generated from the first proteomic model by identifying proteins that contribute to the first proteomic model and show large intra-aliquot variability and reducing the contribution of these proteins in generating the second proteomic model.

[0073] FIG. 5 illustrates an exemplary process 500 for validating a proteomics model according to an embodiment of the present disclosure. The proteomics model can validate input data collection using different data collection protocols. Process 500 can be performed using a machine learning pipeline, such as pipeline 100 described in FIG. 1. For convenience of explanation, process 500 is described herein as being performed using learning engine 160. However, process 500 is not limited to such implementation. According to embodiments of the present disclosure, process 500 can be performed using other components of pipeline 100 or other machine learning systems. For example, process 500 can be performed using a standalone system or computing device for training a machine learning model.

[0074] Process 500 can provide a technical solution to the technical problem described in Figures 3A and 3B. In embodiments of the present disclosure, process 500 enables validation of a proteomic model for use in a second context once the proteomic model has been developed for use in a first context. In some embodiments, input data used to develop a proteomic model may be available, but the original samples from which this input data was generated may not be available. Also, the available samples may not cover the full range of potential inputs to the model. For example, such available samples may be derived largely or entirely from "healthy normal" patients. Thus, such available samples may not support validation of the full range of proteomic model outputs.

[0075] According to an embodiment of the present disclosure, process 500 can compensate for unavailable original samples by using available samples to determine protein-level noise. The protein-level noise is determined by the first context and the second context. The differences in protein-level measurements between text samples can be characterized. Protein-level noise can be used with the input data originally used to develop the proteomics model to generate a validation dataset. The validation dataset can then be used to validate the proteomics model in a second context. Thus, even if the available samples do not span the full range of potential inputs to the model, these samples can still support validation of the full range of the proteomics model's outputs.

[0076] At step 510 of process 500, the learning engine 160 can obtain control, treatment, and training datasets consistent with embodiments of the present disclosure. In some embodiments, these datasets can include protein level measurements (e.g., measurements of relative or absolute protein levels in aliquots of a sample). Such protein level measurements can be numerical. For example, the datasets can include numerical values ​​of relative fluorescence units (RFU) or fluorescence units (FLU). These values ​​can indicate protein levels. In some embodiments, the control and treatment datasets are or can include paired datasets. Such paired datasets can be generated using multiple aliquots from the same sample or paired samples collected from the same patient. The control data can be generated using a first context, and the treatment dataset can be generated using a second context. The training dataset can be generated using original samples in the first context and represents the same input as the data used to train the model.

[0077] In some embodiments, the first and second contexts can test for different sets of proteins. For example, the first context can be or include a previously developed assay, and the second context can be or include a next-generation assay. The next-generation assay can test for a superset of the proteins tested in the previously developed assay. For example, the previously developed assay can test for 5,000 proteins, while the next-generation assay can test for an additional 2,000 proteins in addition to these 5,000 proteins.

[0078] In some embodiments, different sample processing techniques can be employed in the first and second contexts. For example, the control data set can be generated from a citrated plasma sample, and the treatment data set can be generated from a potassium ethylenediaminetetraacetic acid (EDTA) plasma sample. The citrated plasma sample and the EDTA plasma sample can be obtained from the same patient. As a further example, the first and second contexts can use different times between blood collection and sample spinning, different times between sample spinning and sample decantation, different times between sample decantation and freezing, different freezing times or storage temperatures, different numbers of freeze / thaw cycles, etc.

[0079] In some embodiments, the first and second contexts can use different assay protocols. For example, such differences can include differences in the reagents used, different dilutions of the same reagents, the addition or subtraction of steps in the assay protocol, or differences in the device used to perform the assay protocol. As a further example, the first context can include a first set of aptamers, and the second context can include a second set of aptamers.

[0080] In some embodiments, the second context may include any interaction not present in the first context. For example, preparing an aliquot according to the second context can include adding an interfering agent to the aliquot. Thus, process 500 can be used to determine whether the presence of a commonly used drug in a patient's blood adversely affects the performance of a predictive proteomics model.

[0081] Embodiments of the present disclosure are not limited to any particular method of obtaining the control, treatment, and training datasets. In some embodiments, learning engine 160 can obtain these datasets from data storage 140 or another location. In various embodiments, learning engine 160 can obtain these datasets from another system (e.g., a healthcare system, an insurance system, etc.). In some embodiments, the control, treatment, and training datasets can be generated from aliquots or samples using a proteomic assay, such as assay 200. However, embodiments of the present disclosure are not limited to embodiments using assay 200. Additionally or alternatively, other proteomic assays or other types of assays can be used to generate control, treatment, and training datasets that can be used to validate predictive models in embodiments of the present disclosure.

[0082] At step 520 of process 500, learning engine 160 can estimate protein level noise according to embodiments of the present disclosure. In some embodiments, protein level noise can be estimated on a protein-by-protein basis. Learning engine 160 (or another component of pipeline 100, such as dataset generator 120) can determine pairwise differences between protein level measurements in the treatment dataset and corresponding protein level measurements in the control dataset. For example, a first entry in the control dataset and a second entry in the treatment dataset can correspond to aliquots taken from the same sample. The first entry and the second entry can include protein level measurements for the same protein. Learning engine 160 can subtract the protein level measurement from the first entry from the protein level measurement for the second entry to determine the pairwise difference for each protein. Learning engine 160 can determine such pairwise differences for each entry in the control dataset and each corresponding entry in the treatment dataset.

[0083] According to embodiments of the present disclosure, learning engine 160 can characterize the distribution of pairwise differences for each protein. Methods for determining the distribution characteristics can include estimating the distribution of pairwise differences (e.g., using a histogram of pairwise differences, estimating parameters of a parametric model of pairwise differences, etc.), estimating statistics (e.g., mean, median, mode, standard deviation, quartiles, percentiles, etc.) or moments (e.g., first moment, second moment, third moment, etc.) of the distribution of pairwise differences, maintaining a set of pairwise differences and resampling from the set, or other suitable methods for determining the distribution characteristics of pairwise differences.

[0084] Here, depending on the difference between the first context and the second context, the control dataset and the treatment dataset may contain protein level measurements for different sets of proteins. For example, the treatment dataset may contain protein level measurements for a superset of the proteins in the control dataset. In some embodiments, pairwise differences may be determined, and distributional characteristics may be determined for proteins present in both databases.

[0085] In step 530 of process 500, learning engine 160 can generate a validation data set in an embodiment of the present disclosure. The validation data set can include adjusted protein level values. The adjusted protein level values ​​can be used to estimate the protein level values. The validation dataset can be generated using noise and a training dataset for protein levels. In some embodiments, the validation dataset can include multiple entries. Each entry can correspond to an entry in the training dataset. In some embodiments, the learning engine 160 can generate adjusted protein level measurements for each protein in each entry.

[0086] According to embodiments of the present disclosure, the adjusted protein level values ​​may be a function (e.g., a sum, etc.) of the protein level values ​​in the training dataset and noise samples. In some embodiments, the learning engine 160 may generate the noise samples. The generation of the samples may depend on determining the characteristics of the distribution in step 520. In some embodiments, when the shape of the distribution is estimated (e.g., using a histogram), the distribution having the estimated shape may be sampled. For example, if an estimated mean and standard deviation are determined for a protein in the training dataset, a normal distribution with the estimated mean and standard deviation may be sampled to generate samples for that protein. In some embodiments, a set of pairwise differences for the protein may be resampled to generate samples for the protein.

[0087] In step 540 of process 500, the learning engine 160, in embodiments of the present disclosure, can generate a set of predictions by applying the validation dataset to a proteomics model. The proteomics model can be a model developed using the training dataset. Each entry in the validation dataset can be used to generate a corresponding prediction in the set of predictions. How the validation dataset is applied to the proteomics model to generate the set of predictions can depend on the implementation of the proteomics model. In some embodiments, the proteomics model can be implemented as an instance of an object. The object can specify how to train the model and how to use the model to make predictions. Additionally, generating predictions using the model involves invoking a method with an entry as input. For example, assume the function svm.svc(parameters) returns a support vector machine object with certain parameters. This object can have a fit(x,y) method to train data (using the x and y training data) and a predict(x) method to take a vector of samples (each containing a vector of features) and generate a vector of class labels. In this simple example, the generation, training, and prediction may be as follows: proteomis_model=svm.svc(parameters) poteomics_model.fit(x_training,y_training) classifications=proteomics_model.predict(x_test)

[0088] As a further example, the function linear_model.LogisticRegression(parameters) can return a penalized logistic regression object with specific parameters. Similar to the support vector machine object above, the object can specify a fit() method for training the object and a predict() method for generating predictions using input data.

[0089] Note that the above examples of support vector machine and penalized logistic regression objects are by way of example only and are not limiting. The particular manner in which the validation data set is applied to the proteomics model will depend on the particular implementation of the proteomics model.

[0090] In step 550 of process 500, learning engine 160 generates The set of generated predictions can be used to determine one or more performance measures of the proteomics model. In some embodiments, the learning engine 160 can determine a performance measure that depends on the match between two sets of predictions (the original set of predictions generated by the proteomics model using the training data and the set of predictions generated in step 540). In some embodiments, the performance measure can be Lin's concordance correlation coefficient, Pearson's correlation coefficient, or another suitable performance measure. In some embodiments, the learning engine 160 can determine the performance measure as a function of the match between a ground truth associated with the training data and the set of predictions generated in step 540. In some embodiments, the ground truth can be specified by a label associated with an entry in the training dataset. The label can specify the presence or absence of a condition, a binned survival time, or another appropriate ground truth regarding the patient's health information.

[0091] According to embodiments of the present disclosure, the learning engine 160 can determine an indication of validity based on one or more performance measures. In some embodiments, the indication of validity can be the value of one or more performance measures. In various embodiments, the indication of validity can depend on the value of one or more performance measures. For example, the values ​​of the performance measures can be binned or thresholded, with values ​​below one bin or a threshold being assigned a "warning" or "fail" indicator and values ​​above another bin or another threshold being assigned a "pass" or "valid" indicator.

[0092] According to embodiments of the present disclosure, process 500 may include optional operations not shown in FIG. 5 . This operation may be performed by learning engine 160 after obtaining the control dataset and the treatment dataset (e.g., in step 510). In some embodiments, learning engine 160 may determine pairwise differences between a control set of predictions generated by applying the control dataset to the proteomic model and a treatment set of predictions generated by applying the treatment dataset to the proteomic model. Here, if the control dataset and the treatment dataset include overlapping sets of proteins, the pairwise differences may be calculated only for proteins at the intersection of the sets. In some embodiments, learning engine 160 may determine a statistical value of the pairwise differences (e.g., mean, median, 75th percentile, 90th percentile, or other suitable threshold). In some embodiments, learning engine 160 may determine whether the statistical value meets an invalidity condition. The invalidity condition may depend on the characteristics of the assay 200. For example, the predictive model may have inherent variability resulting from noise inherent in the assay 200 (e.g., protein level noise determined based on the difference in test-retest values ​​of protein levels, etc.).

[0093] In some embodiments, the invalidity condition may be at least partially met if the pairwise difference statistic exceeds a function of the intrinsic variation statistic. For example, the intrinsic variation statistic may be the standard deviation of the intrinsic variation. Alternatively, the function may be a multiple of the statistic (e.g., a multiple selected in the range of 0.5 to 5, such as 3 or another suitable value). In some embodiments, the null condition may be met when the pairwise difference statistic exceeds a multiple of the standard deviation of the inherent variation. In some embodiments, the null condition may be met by further requiring that a statistical test (e.g., a paired t-test) reject the null hypothesis that the pairwise difference statistic is equal to zero.

[0094] According to an embodiment of the present disclosure, if the override condition is met, the process 500 may proceed to step 520; otherwise, the process 500 may end.

[0095] FIG. 6 illustrates an exemplary process for developing a proteomic model according to an embodiment of the present disclosure. 1 illustrates a process 600 for developing proteomic models. Development of proteomic models can be structured to enhance the resilience of such models to variations in input data. As described herein, input datasets can be generated in different contexts, and variations in input data can arise from differences between such contexts. Process 600 can be performed using a machine learning pipeline, such as the pipeline described in FIG. 1 . For ease of explanation, process 600 is described herein as being performed using learning engine 160. However, process 600 is not limited to such implementation. According to embodiments of the present disclosure, process 600 can be performed using other components of pipeline 100 or other machine learning systems. For example, process 600 can be performed using a standalone system or computing device for training a machine learning model.

[0096] Process 600 can provide a technical solution to the technical problem described in FIGS. 4A-4F. According to embodiments of the present disclosure, process 600 enables the development of proteomic models that are tolerant to differences in the context in which the input data sets are obtained. Here, the training data used in developing the proteomic models is obtained from samples that are handled strictly in accordance with sample processing guidelines. However, samples obtained from clinicians, healthcare systems, laboratories, or other users may not be handled strictly in accordance with sample processing guidelines. Therefore, according to embodiments of the present disclosure, proteomic models can be developed to be tolerant to variations in sample processing. Furthermore, proteomic models can be developed to be tolerant to variations in the assay process or in the device used to measure protein levels.

[0097] In some embodiments, process 600 can be combined with process 500 described above. For example, a proteomic model may have been developed using samples that are not currently available. Furthermore, testing the resilience of a proteomic model may require a large number of samples obtained under many varying conditions. According to process 500, training data used to develop a proteomic model can be used to test the resilience of the model. Applying process 500, multiple treatment data sets can be obtained with various variations in the context in which the data is obtained, while a control data set is obtained using the same context as the training data. Furthermore, the impact of these variations can be investigated by estimating noise in protein levels and generating validation data sets corresponding to each of the treatment data sets.

[0098] In some embodiments, process 600 can be performed as part of an iterative learning and development process. For example, a user may interact with learning engine 160 (e.g., via user device 180) to select and learn a proteomic model. The user may also interact with learning engine 160 to repeatedly perform process 600 until a proteomic model that meets their requirements is developed. This model may exhibit a desired degree of resilience across a determined set of contextual variations.

[0099] At step 610 of process 600, learning engine 160 can retrieve a proteomic model according to an embodiment of the present disclosure. The proteomic model may have been trained (e.g., by learning engine 160 or another system) to generate predictions based on measurements of protein levels. The proteomic model may have been trained using a training dataset obtained from training samples in a first context. Learning engine 160 may retrieve the proteomic model from a model storage device, such as model storage 130, or another suitable storage location.

[0100] At step 620 of process 600, learning engine 160 can acquire a therapeutic dataset including protein level measurements, according to embodiments of the present disclosure. In some embodiments, the therapeutic dataset can be generated by assaying therapeutic samples obtained in therapeutic settings. In some embodiments, each therapeutic setting can vary along one or more dimensions. For example, a dimension can be spin time, where the spin times for a therapeutic setting can be 0.5, 1.5, 3, 9, and 24 hours. As a further example, a dimension can be number of freeze / thaw cycles, where the freeze / thaw cycles for a therapeutic setting can be 2, 3, 4, 5, and 10 cycles. As described herein, the therapeutic dataset can be generated by pipeline 100 from the therapeutic sample using an assay, such as assay 200, and a measurement system, such as measurement system 101. Additionally or alternatively, the therapeutic dataset can be acquired by pipeline 100 (or learning engine 160) from another system.

[0101] In some embodiments, each treatment dataset can be applied to a proteomic model to generate a set of predictions, and these predictions can be analyzed in step 630 of process 600. In various embodiments, a control dataset can be obtained in step 620 of process 600. The control dataset can be generated from samples obtained in the same context as the samples used to generate the training dataset (e.g., the dataset used to train the proteomic model). As described in process 500, the control dataset, training dataset, and each treatment dataset can be used to generate validation datasets, each corresponding to a treatment dataset and generated using the treatment dataset.

[0102] In step 630 of process 600, the dependency of the proteomic model on the differences between the second contexts may be determined according to embodiments of the present disclosure. In some embodiments, the dependency may be determined based on one or more performance measures of the proteomic model. In various embodiments, the dependency may be determined based on displayed metrics for the set of predictions generated in step 620.

[0103] According to embodiments of the present disclosure, the learning engine 160 can determine one or more performance measures of the proteomics model. The learning engine 160 can determine one or more performance measures for one or more of the treatment (or validation) datasets. In some embodiments, the learning engine 160 can determine one or more performance measures for the control (or training) dataset. In some embodiments, the learning engine 160 can determine the performance measure as a function of the agreement between predictions generated by applying the control dataset (or training dataset) to the proteomics model and predictions generated by applying the treatment dataset (or validation dataset) to the proteomics model. In some embodiments, the learning engine 160 can determine the performance measure as a function of the agreement between a ground truth and predictions generated by applying the control dataset or the treatment dataset (or the training dataset or validation dataset) to the proteomics model.

[0104] According to embodiments of the present disclosure, learning engine 160 can provide a user interface accessible through user device 180. The user interface can display values ​​of one or more performance-related metrics generated by learning engine 160. In some embodiments, the user interface can be displayed as a box plot, qq plot, histogram, table, or other suitable display of the set of predictions generated in step 620. Examples of such displays include FIGS. 4A-4F, 7A-7D, 8A-8D, and 9A-9B.

[0105] According to embodiments of the present disclosure, the dependency of a proteomic model on differences between second contexts can be automatically determined. In some embodiments, the learning engine 160 can automatically determine whether one or more performance measures differ between predictions generated using the control and treatment datasets (or between the training and validation datasets). In some embodiments, the learning engine 160 can automatically determine whether such differences are statistically significant. In various embodiments, a user can determine the dependency based on a display of the values ​​of one or more performance measures. In some embodiments, the dependency of a proteomic model can also be determined semi-automatically. The learning engine 160 can automatically determine the values ​​of one or more performance measures, and a user can determine the dependency based on a display of the values ​​of one or more performance measures. In some embodiments, the dependency of a proteomic model can also be determined manually. The learning engine 160 can provide the set of predictions generated in step 620 for display, and a user can determine the dependency based on this display.

[0106] In step 640 of process 600, the proteomic model can be updated based on the dependencies determined in step 630, according to embodiments of the present disclosure. In some embodiments, proteins in the input dataset that are applied to the proteomic model can be identified, and the proteomic model can be updated in a manner that reduces the importance of proteins in the proteomic model.

[0107] According to embodiments of the present disclosure, proteins can be identified based on the protein's effect on the variability of a predictive model between a control dataset and a treatment dataset, or between a training dataset and a validation dataset (e.g., "variation effect"). The manner and method of identification of the effect can depend on the implementation of the proteomics model. In some embodiments, the variation effect can depend on the importance of the protein to the model. For example, the importance of a protein to a regression proteomics model can depend on the magnitude of the protein's coefficient in the regression proteomics model. As a further example, the importance of a protein to an SVM proteomics model can depend on the protein's feature weight in the SVM proteomics model. As a further example, the importance of a protein to a random forest model can be calculated using Gini importance, mean-reduced accuracy, permutation-based importance, Spree-value-based importance, or any other suitable method. In some embodiments, the variation effect can depend on the variation in protein levels between the control dataset and the treatment dataset (or between the training dataset and the validation dataset). In some examples, the greater this variance in measured protein levels, the greater the effect of the protein on the variability of the predictive model.

[0108] According to embodiments of the present disclosure, the proteomic model may be updated to reduce the effect of variation of the identified proteins. The effect of variation may be reduced by decreasing the magnitude of a weight or coefficient associated with the protein, excluding the protein from the input dataset, retraining the model, or by another suitable technique. In some embodiments, the learning engine 160 may automatically update the proteomic model. For example, the learning engine 160 may automatically identify proteins with the greatest effect of variation. The learning engine 160 may also update the model to reduce the effect of variation of one or more of these proteins. In some embodiments, the learning engine 160 may also semi-automatically update the proteomic model. For example, the learning engine 160 may automatically identify proteins with the greatest effect of variation and display indices of these proteins to the user. The user may then select one, more, or none of the identified proteins. The learning engine 160 may also update the proteomic model semi-automatically. For example, the learning engine 160 may automatically identify proteins with the greatest effect of variation and display indices of these proteins to the user. The user may then select one, more, or none of the identified proteins. Learning engine 160 can update the proteomic model to reduce the effect of variability of any selected proteins. In some embodiments, a user can interact with a user interface provided by learning engine 160 to identify proteins with the greatest effect of variability. For example, a user can examine the model weights or coefficients, the model architecture, or the calculated contributions of different proteins to the model's output. A user can also provide commands to learning engine 160 to update the proteomic model to reduce the effect of variability of any identified proteins.

[0109] Here, identifying proteins may involve more than simply removing proteins with the greatest variability from the model. In some cases, such proteins may provide information necessary for the functioning of the proteomics model. In these situations, rather than removing proteins, a user can determine contextual differences between the control and treatment datasets that are potentially detrimental to the performance of the proteomics model. For example, if delaying serum clotting affects the model output and removing effect-influencing proteins would unduly impair model performance, the instructions for collecting samples can be updated to emphasize that the instructions are followed to ensure serum clotting is not delayed.

[0110] Figures 7A-7D illustrate the effect of sample processing variations on the output of a proteomic model according to an embodiment of the present disclosure. In this example, the samples are serum samples. Sample processing is varied along four dimensions: number of freeze / thaw cycles, serum time to clot, serum time to decant, and serum freezing time at -80°C. A proteomic model is then developed using the training data. In this example, the proteomic model is trained using known significant aptamers and five randomly selected aptamers. Using steps 610-630 of processes 500 and 600, predictions are generated using the training data set as well as the control and treatment data sets. Boxplots of these predictions, broken down by class label, are shown in Figures 7A-7D. Here, the proteomic model may exhibit stability or monotonic change in variation across each of the four dimensions. If updating the model to exhibit such stability or monotonic change in the sample processing dimension unduly degrades model performance, the sample instructions for collection can emphasize the importance of this dimension of sample processing.

[0111] Figure 7A shows the dependence of predicted probability values ​​generated by applying entries in a training or validation data set to a proteomic model on variation in the number of freeze / thaw cycles of a sample. As shown, the number of outlier observations increases sharply between the baseline sample treatment value and two freeze / thaw cycles, and then remains high as the number of freeze / thaw cycles increases.

[0112] Figure 7B shows the dependence of predicted probability values ​​on the variation in sample clotting time. As shown, as clotting time increases, the median predicted probability value for negative class observations increases. The distribution of predicted probability values ​​for negative class observations shows an increasing curve toward higher inappropriate predicted probability values. For positive class observations, the number of outliers with inappropriately low predicted values ​​increases as clotting time increases.

[0113] Figure 7C shows the dependence of predicted probability values ​​on sample decanting time. As shown, as decanting time increases, the median predicted probability value for negative-class observations increases. The distribution of predicted probability values ​​for negative-class observations shows an increasing curve toward higher, inappropriate predicted probability values. For positive-class observations, the number of outliers with inappropriately low predicted values ​​increases as decanting time increases.

[0114] Figure 7D shows the dependence of predicted probability values ​​on variations in sample freezing time at -80° C. As can be seen, increasing freezing time beyond the baseline value immediately affects the distribution of predicted probability values, increasing outliers for both negative and positive class samples.

[0115] 8A-8D illustrate the effect of sample processing variations on the output of a proteomic model, according to an embodiment of the present disclosure. In this example, the samples are plasma samples. Sample processing is varied in four dimensions: number of freeze / thaw cycles, 24-day freezer storage temperature of the plasma, freezing time of the plasma at -80°C, and spin time of the plasma. A proteomic model is then developed using the training data. In this example, the proteomic model is trained using known significant aptamers and five randomly selected aptamers. According to steps 610-630 of process 500 and process 600, predictions are generated using the training data set as well as the control and treatment data sets. Boxplots of these predictions are shown in FIGS. 8A-8D.

[0116] Figure 8A shows the dependence of predicted output values ​​generated by applying entries in a training or validation data set to a proteomic model on the number of freeze / thaw cycles of the sample. As shown, the median predicted output value increases as the number of freeze / thaw cycles increases.

[0117] Figure 8B shows the dependence of predicted power on variations in plasma storage temperature (e.g., -80°C vs. -20°C) over 24 days of sample storage. As shown, predicted power increases when plasma is stored at higher temperatures.

[0118] Figure 8C shows the predicted output values ​​for samples with varying freezing times at -80°C. As can be seen, there is no obvious dependence of the predicted output values ​​at different freezing times.

[0119] Figure 8D shows the dependence of predicted output power on variation in sample spin time. As can be seen, increasing spin time from baseline to over 9 hours affects output power.

[0120] 9A-9D illustrate the effect of sample processing variations on the predictive performance of predictions made by a proteomic model, according to an embodiment of the present disclosure. In this example, the predictive performance is root mean square error (RMSE), and the samples are serum samples. Sample processing is varied in four dimensions: number of freeze / thaw cycles, serum clotting time, serum decanting time, and serum freezing time at -80°C. A proteomic model is created using training data. In this example, the proteomic model is trained using known significant aptamers and five randomly selected aptamers. According to steps 610-630 of process 500 and process 600, predictions are generated using the training data set as well as the control and treatment data sets. Bar graphs of the measured RMSE are shown in FIGS. 9A-9D.

[0121] Figure 9A shows the dependence of the RMSE value on the number of freeze / thaw cycles of the sample. As shown in the figure, the RMSE value does not increase significantly between 1 and 10 freeze / thaw cycles.

[0122] Figure 9B shows the dependence of the RMSE value on the variation in clotting time of the samples. As can be seen, the RMSE value increases significantly from baseline clotting time to times greater than 3 hours.

[0123] Figure 9C shows the dependence of the RMSE value on variation in the decant time of the sample. As shown, the RMSE value increases significantly between baseline decant and decant time, up to times exceeding 3 hours.

[0124] Figure 9D shows the dependence of RMSE values ​​on variation in freezing time for samples at -80° C. As can be seen, the RMSE values ​​increase between baseline freezing times and times greater than 3 hours, and increase substantially between baseline freezing times and 24 hours freezing times.

[0125] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if a component is described as including A or B, the component may include A, or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if a component is described as including A, B, or C, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless specifically stated otherwise or impracticable.

[0126] The exemplary embodiments have been described above with reference to flowcharts or block diagrams of methods, apparatuses (systems), and computer program products. Each block of the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, may be implemented by a computer program product or instructions on a computer program product. These computer program instructions may also be provided to a processor of a computer or other programmable data processing apparatus to generate a machine, such that the instructions, when executed by a processor of the computer or other programmable data processing apparatus, generate means for implementing the functions / acts identified in one or more blocks of the flowcharts or block diagrams.

[0127] These computer program instructions may also be stored on a computer-readable medium that can direct one or more hardware processors of a computer, other programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored on the computer-readable medium form a product including instructions that implement the functions / acts specified in one or more blocks of the flowchart or block diagram.

[0128] Also, computer program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device, creating a computer-implemented process such that the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / operations specified in one or more blocks of the flowchart or block diagram.

[0129] One or more computer-readable mediums may be used in any combination. A computer-readable medium may be a non-transitory computer-readable storage medium. As used herein, a computer-readable storage medium may be any tangible medium that can have or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0130] The program code embodied on the computer readable medium may be transmitted via any suitable means, including but not limited to wireless, wired, fiber optic cable, RF, IR, or any suitable combination thereof. The transmission may be carried out using any suitable medium that does not require the use of a

[0131] Computer program code for implementing operations, e.g., embodiments, may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional procedural programming languages, such as the C programming language or similar programming languages. The program code may execute completely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet by an Internet Service Provider).

[0132] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products in various embodiments. Each block in a flowchart or block diagram represents a module, segment, or code portion having one or more executable instructions that implement a specified logical function. In some alternative implementations, the functions shown in the blocks may be executed in a different order than shown in the figures. For example, two consecutive blocks may actually be executed substantially simultaneously or in the reverse order, depending on the functionality being performed. Each block in a block diagram or flowchart, and combinations of multiple blocks in a block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.

[0133] The above embodiments are not mutually exclusive, and elements, components, materials, or steps described in connection with one exemplary embodiment may be combined with, or excluded from, other embodiments in any suitable manner to achieve desired design objectives.

[0134] The above embodiment can be further explained by the following.

[0135] 1. A system for validating a proteomic model, comprising at least one processor.

[0136] The system also includes at least one non-transitory computer-readable medium having instructions that, when executed by the at least one processor, cause the system to perform operations including obtaining a control dataset including protein level measurements, obtaining a treatment dataset including protein level measurements, obtaining a training dataset including protein level measurements, estimating protein level noise using the control dataset and the treatment dataset, generating a validation dataset including adjusted protein level measurements using the training dataset and the estimated protein level noise, applying the validation dataset to a proteomic model trained using the training dataset to generate a set of predictions, determining a validity indicator using the set of predictions, and providing the validity indicator.

[0137] 2. The system according to claim 1, wherein acquiring the control data set comprises: The method includes assaying a first aliquot to obtain a first protein level measurement, and obtaining the therapeutic dataset includes assaying a second aliquot to obtain a second protein level measurement, wherein the first aliquot and the second aliquot are obtained from a plurality of samples having different sample characteristics, or the first aliquot and the second aliquot are obtained from a plurality of samples that have been subjected to different treatments, or the first aliquot and the second aliquot follow different assay protocols.

[0138] 3. The system according to paragraph 2 above, wherein the first aliquot and the second aliquot are Recoats are obtained from the plurality of samples having different sample characteristics, the sample characteristics including at least one of the following: citrated plasma or EDTA plasma composition, fed or fasted state of the patient providing the sample, serum versus plasma composition, presence or absence of interfering agents.

[0139] 4. The system according to claim 3, wherein the sample characteristics include the presence or absence of the interfering agent. Differently, the interfering agent comprises at least one of the following: a nonsteroidal anti-inflammatory drug (NSAID), a birth control drug, a blood pressure drug, a mental health drug, a cholesterol drug, an asthma drug, a diabetes drug, a thyroid drug, an antiviral drug, or an antibacterial drug.

[0140] 5. The system according to any one of the above items 2 to 4, wherein the first aliquot The first and second aliquots are obtained from a plurality of samples that have undergone different treatments, and the sample treatments differ in at least one of the following: time between sample collection and sample freezing, sample freezing temperature, duration of freezing temperature, number of freeze / thaw cycles for the sample, spin time if the plurality of differently treated samples are plasma samples, and clotting time if the plurality of differently treated samples are serum samples.

[0141] 6. The system according to any one of the above items 2 to 4, wherein the first aliquot The first aliquot and the second aliquot are subjected to different assay protocols, and the different assay protocols differ in at least one of the following: reagents used, dilutions of the reagents, steps in the assay protocol, equipment used to perform the assay protocol, or the number of proteins assayed by the assay protocol.

[0142] 7. The system according to any one of paragraphs 1 to 6, wherein the training data set The data set is acquired in a first context and the treatment data set is acquired in a second context different from the first context.

[0143] 8. The system according to any one of items 1 to 7, wherein the proteomics The model is trained to predict a health outcome, wherein the proportion of entries in the training dataset corresponding to patients experiencing the health outcome is greater than the proportion of entries in the treatment dataset corresponding to patients experiencing the health outcome.

[0144] 9. The system according to any one of the above items 1 to 8, wherein the control data set Using the control dataset and the treatment dataset to estimate protein level noise includes determining pairwise differences between protein level measurements in the control dataset and corresponding protein level measurements in the treatment dataset, and determining, for each protein in the control dataset, a characteristic of the distribution of the pairwise differences.

[0145] 10. The system according to any one of paragraphs 1 to 9, wherein the training data set generating the validation data set using the estimated protein level noise for a first protein in the training data set; and adding the sample to the protein level measurements to generate corresponding adjusted protein level measurements.

[0146] 11. The system described in paragraph 10 above, wherein the estimated noise in the protein level includes an estimated mean and standard deviation of the first protein, and generating the samples includes sampling a normal distribution having the estimated mean and standard deviation of the first protein.

[0147] 12. The system according to any one of paragraphs 1 to 11, wherein the prediction set Determining the measure of validity using the proteomics model includes quantifying the agreement between the set of predictions and a corresponding set of predictions generated using the proteomics model and the training dataset, and quantifying the agreement between the set of predictions and a ground truth associated with the training dataset.

[0148] 13. The system according to any one of the above items 1 to 12, wherein the operation is determining a pairwise difference statistic between a control set of predictions generated by applying the control dataset to the proteomics model and a treatment set of predictions generated by applying the treatment dataset to the proteomics model; and determining that the pairwise difference statistic satisfies a null condition, wherein the noise in the protein levels is estimated in response to the null condition being satisfied.

[0149] 14. A system for developing a proteomics model, comprising at least one processor and at least one non-transitory computer-readable medium having instructions for execution by the at least one processor, wherein when the instructions are executed, the system performs operations including obtaining a proteomic model trained to generate predictions based on learning of protein level measurements generated from training aliquots obtained in a first context; generating therapeutic protein level measurements from therapeutic aliquots or samples obtained in a plurality of second contexts; determining a dependency of the proteomic model on differences between the plurality of second contexts; and updating the proteomic model based on the determined dependency.

[0150] 15. The system according to claim 14, further comprising: Updating the proteomic model includes identifying proteins based on the variation effect of the proteins and reducing the variation effect of the proteins in the proteomic model.

[0151] 16. The system according to any one of items 14 to 15, wherein the proteo The mixture model includes a linear regression model, a logistic regression model, a survival model, a random forest model, a support vector machine model, or a linear discriminant analysis model.

[0152] 17. The system according to any one of items 14 to 16, wherein the learning ant The coat is obtained from a serum or plasma sample.

[0153] 18. The system according to any one of items 14 to 17, wherein the plurality of The two contexts have different sample characteristics, sample processing, or assay protocols.

[0154] 19. The system according to any one of items 14 to 18, wherein the plurality of The differences between the two contexts are freezing time, number of freeze / thaw cycles, spin time, fasting time, etc. The difference includes a difference in at least one of time, shipping time, freezer storage time, setting time, or decanting time.

[0155] 20. The system according to any one of items 14 to 20, wherein the plurality of Determining the dependence of the proteomic model on the difference between two contexts includes determining protein level noise using the therapeutic protein level measurements, generating validation protein level measurements using the training protein level measurements and the protein noise, and generating a prediction using the validation protein level measurements.

[0156] In the foregoing specification, the embodiments have been described with numerous specific details that vary from implementation to implementation. Several adaptations and modifications can be made to the described embodiments. From the description herein, one skilled in the art will understand that other embodiments are possible by practicing the present disclosure. The specification and examples provide merely examples. Furthermore, the order of execution of steps shown in the figures is for illustrative purposes and is not intended to limit the execution of the steps to any particular order. Thus, one skilled in the art will understand that these steps may be executed in different orders when implementing the same method.

Claims

1. 1. A system for validating a proteomics model, comprising: at least one processor; at least one non-transitory computer-readable medium having instructions for execution by said at least one processor; and When the instructions are executed, the system: obtaining a control dataset comprising measurements of protein levels; obtaining a therapeutic dataset including protein level measurements; obtaining a training dataset comprising measurements of protein levels; using the control dataset and the treatment dataset to estimate noise in protein levels; generating a validation dataset comprising adjusted protein level measurements using the training dataset and the estimated protein level noise; applying the validation dataset to a proteomic model trained using the training dataset to generate a set of predictions; determining a measure of validity using the set of predictions; provide an indication of the appropriateness of said Perform an action that includes A system for validating a proteomics model, comprising:

2. obtaining the control dataset includes assaying a first aliquot to obtain a first protein level measurement; obtaining the therapeutic dataset includes assaying a second aliquot to obtain a second protein level measurement; the first aliquot and the second aliquot are obtained from a plurality of samples having different sample characteristics; or the first aliquot and the second aliquot are obtained from differently treated samples, or The first aliquot and the second aliquot are subjected to different assay protocols.

2. The system of claim 1.

3. the first aliquot and the second aliquot are obtained from the plurality of samples having different sample characteristics; The sample characteristics include at least one of the following: Citrated plasma or EDTA plasma compositions the fed or fasting state of the patient providing said sample Serum vs. plasma composition Presence or absence of interference agents 3. The system of claim 2.

4. the sample characteristics differ in the presence or absence of the interfering agent; The interfering agent comprises at least one of the following: Nonsteroidal anti-inflammatory drugs (NSAIDs) birth control drugs blood pressure medication Mental health medications Cholesterol medication asthma medication diabetes medication thyroid medication Antiviral or antibiotic drugs 4. The system of claim 3.

5. the first aliquot and the second aliquot are obtained from a plurality of samples that have been subjected to different treatments; The sample treatments differ in at least one of the following: Time between sample collection and sample freezing Sample freezing temperature Duration of freezing temperatures The number of freeze / thaw cycles of the sample Spin times when the differently treated samples are plasma samples Clotting times when the differently treated samples are serum samples 3. The system according to claim 2, wherein:

6. the first aliquot and the second aliquot follow different assay protocols; The different assay protocols differ in at least one of the following: Reagents used Dilution of the reagent Steps of the assay protocol Devices used to carry out the assay protocol Number of proteins assayed by the assay protocol 3. The system of claim 2.

7. The training data set is obtained in a first context; The treatment data set is acquired in a second context different from the first context.

2. The system of claim 1.

8. the proteomic model is trained to predict a health outcome; a proportion of entries in the training dataset corresponding to patients experiencing the health-related outcome is greater than a proportion of entries in the treatment dataset corresponding to patients experiencing the health-related outcome; 2. The system of claim 1.

9. estimating noise in protein levels using the control dataset and the treatment dataset, determining pairwise differences between protein level measurements in the control dataset and corresponding protein level measurements in the treatment dataset; determining, for each protein in the control dataset, a characteristic of the distribution of pairwise differences; Contains 2. The system of claim 1.

10. generating the validation dataset using the training dataset and the estimated protein-level noise, For the measured protein level of a first protein in the training data set, generating a noise sample of the estimated protein level of the first protein; adding said sample to said protein level measurements to generate corresponding adjusted protein level measurements. Contains 2. The system of claim 1.

11. the estimated noise in the protein level comprises an estimated mean and standard deviation of the first protein; generating the sample includes sampling a normal distribution having the estimated mean and the standard deviation of the first protein; The system of claim 10.

12. Determining the measure of validity using the set of predictions comprises: quantifying the agreement between the set of predictions and a corresponding set of predictions generated using the proteomic model and the training dataset; Quantifying the agreement between the set of predictions and a ground truth associated with the training dataset. Contains 2. The system of claim 1.

13. The operation is determining a pairwise difference statistic between a control set of predictions generated by applying the control data set to the proteomic model and a treatment set of predictions generated by applying the treatment data set to the proteomic model; determining that the statistical value of the pairwise difference values ​​satisfies an invalidity condition; Including, The noise in the protein level is estimated according to satisfying the null condition.

2. The system of claim 1.

14. 1. A system for developing a proteomic model, comprising: at least one processor; at least one non-transitory computer-readable medium having instructions for execution by said at least one processor; and When the instructions are executed, the system: obtaining a proteomic model trained to generate predictions based on training protein level measurements generated from training aliquots obtained in a first context; generating therapeutic protein level measurements from the therapeutic aliquots or samples obtained in a plurality of second contexts; determining a dependency of the proteomic model on differences between the plurality of second contexts; and updating the proteomic model based on the determined dependencies. Perform an action that includes A system for developing a proteomics model characterized by:

15. updating the proteomic model based on the determined dependencies, identifying a protein based on the perturbation effect of said protein; Reducing the effect of variability of said proteins in said proteomic model. Contains 15. The system of claim 14.

16. 15. The system of claim 14, wherein the proteomic model comprises a linear regression model, a logistic regression model, a survival model, a random forest model, a support vector machine model, or a linear discriminant analysis model.

17. 15. The system of claim 14, wherein the training aliquots are obtained from serum or plasma samples.

18. 15. The system of claim 14, wherein the plurality of second contexts have different sample characteristics, sample processing, or assay protocols.

19. 15. The system of claim 14, wherein the differences in the plurality of second contexts include differences in at least one of freezing time, number of freeze / thaw cycles, spin time, fasting time, shipping time, freezer storage time, coagulation time, or decanting time.

20. Determining the dependence of the proteomic model on differences between the plurality of second contexts includes: determining protein level noise using said therapeutic protein level measurements; generating a validation protein level measurement using the training protein level measurements and the protein level noise; generating a prediction using said validation protein level measurements; Contains 15. The system of claim 14.