Systems and methods for validating proteomic models

By estimating protein level noise and generating verification data sets, the problem of performance degradation of proteomics models when sample characteristics, processing methods or assay schemes is solved, and the elasticity and stability of the model are improved.

CN120035864APending Publication Date: 2025-05-23SOMALOGIC OPERATING CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071982.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-10-11
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing proteomics models have deteriorated performance when sample characteristics, sample processing methods, or assay protocols are changed, and it is difficult to obtain samples representing all expected inputs, hindering the development of resilient models.

Method used

A system was developed to estimate protein level noise by obtaining control data sets, processing data sets, and training data sets, generating validation data sets, and applying them to trained proteomics models to generate prediction results to determine the indication of validity.

Benefits of technology

The verification and update of the proteomic model was achieved, which enhanced the model's adaptability to the changes in input data, and improved the stability and prediction performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035864A_ABST
    Figure CN120035864A_ABST
Patent Text Reader

Abstract

Methods, systems, and computer-readable media for developing or validating an elastic proteomic model may estimate effects of different scenarios on model input data, and then develop or validate a proteomic model using the estimated effects. An exemplary method for validating a proteomic model may estimate protein level noise using a control dataset and a processing dataset. The contrast dataset and the processing dataset may be different from the training dataset used to generate the model. A validation dataset may be generated using the training dataset and the estimated protein level noise. The validation dataset may include corrected protein level measurements. A set of predictors may be generated by applying the verification dataset to a proteomic model trained using the training dataset. The set of predictions may be used to determine the effectiveness of the proteomic model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 415,978, filed on October 13, 2022, the entire contents of which are incorporated herein by reference. Background Art

[0003] Aptamer-based assays can simultaneously measure protein levels of thousands of proteins in samples taken from patients. However, due in part to the nature of such assays, the measured protein levels can change as a result of changes in sample characteristics, changes in sample processing, or changes in assay protocols. These changes can affect the results of proteomic models that use the measured protein levels as input features. As a result, changes in sample characteristics, sample processing, and assay protocols can lead to reduced performance of such proteomic models.

[0004] Therefore, if sample characteristics, sample processing, or assay protocols change, it may be necessary to revalidate the corresponding proteomic model. However, revalidation of the proteomic model may require samples that are representative of all expected inputs. Such samples may no longer be available or difficult to obtain. Difficulty in obtaining such samples may also hinder the development of proteomic models that are resilient to changes in sample characteristics, sample processing, or assay protocols. Summary of the invention

[0005] Certain embodiments of the present disclosure relate to developing proteomics models that are resilient to changes in input data (such as sample characteristics, sample processing methods, or assay protocols). Consistent with the disclosed embodiments, the impact of these changes on the input data can be estimated. The estimated impact can then be used to generate updated data for validating or training the proteomics model.

[0006] The disclosed embodiments include a system for validating a proteomics model. The system may include at least one processor and at least one non-transitory computer-readable medium. The computer-readable medium may contain instructions that, when executed by the at least one processor, cause the system to perform operations. The operations may include obtaining a control data set including protein level measurements. The operations may also include obtaining a processed data set including protein level measurements. The operations may also include obtaining a training data set including protein level measurements. The operations may also include estimating protein level noise using the control data set and the processed data set. The operations may also include generating a validation data set using the training data set and the estimated protein level noise. The validation data set may include corrected protein level measurements. The operations may also include generating a set of prediction results by applying the validation data set to a proteomics model trained using the training data set. The operations may also include determining a validity indication using the set of prediction results, and providing the validity indication.

[0007] The disclosed embodiments include a system for developing a proteomics model. The system may include at least one processor and at least one non-transitory computer-readable medium. The computer-readable medium may contain instructions that, when executed by the at least one processor, cause the system to perform operations. The operations may include obtaining a proteomics model that is trained to generate predictions based on training protein level measurements generated by training aliquots obtained in a first scenario. The operations may also include generating treatment protein level measurements from treatment aliquots or samples obtained in a second scenario. The operations may also include determining a dependency of the proteomics model on differences between the second scenarios. The operations may also include updating the proteomics model based on the determined dependencies.

[0008] The disclosed embodiments also include corresponding methods for validating and / or developing proteomic models; and non-transitory computer-readable media containing executable instructions for performing such methods.

[0009] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims.Other systems, methods, and computer-readable media are also discussed. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles of the disclosure. In the drawings:

[0011] Figure 1An exemplary pipeline for developing, validating, and deploying a proteomic model for predicting health information using aptamer-based blood tests consistent with the disclosed embodiments is described.

[0012] FIG. 2A to FIG. 2H A high-level description of the stages in an exemplary aptamer-based serum or plasma assay consistent with the disclosed embodiments is provided.

[0013] Figure 3A and Figure 3B Exemplary challenges that arise in validating proteomic models that predict health information using aptamer-based serum or plasma assays consistent with the disclosed embodiments are described.

[0014] 4A to 4F Exemplary model-resilience challenges that arise in training proteomic models for predicting health information using aptamer-based blood tests consistent with the disclosed embodiments are described.

[0015] Figure 5 An exemplary process for validating a proteomic model consistent with the disclosed embodiments is described.

[0016] Figure 6 An exemplary process for developing a proteomic model consistent with the disclosed embodiments is described.

[0017] 7A to 7D The effects of differences in sample processing consistent with the disclosed embodiments on the output of a binary endpoint proteomics model are described.

[0018] FIG. 8A to FIG. 8D The effects of differences in sample processing consistent with the disclosed embodiments on the output of a continuous endpoint proteomics model are described.

[0019] 9A to 9D The effects of differences in sample processing consistent with the disclosed embodiments on the predictive performance of proteomics models are described. DETAILED DESCRIPTION

[0020] In the following detailed description, many specific details are set forth to provide a comprehensive understanding of the disclosed exemplary embodiments. However, those skilled in the art will appreciate that the principles of the exemplary embodiments can be implemented without each specific detail. Well-known methods, steps, and components are not described in detail to avoid obscuring the principles of the exemplary embodiments. Unless explicitly stated, the exemplary methods and processes described herein are neither limited to a specific order or sequence nor to a specific system configuration. In addition, some of the described embodiments or their elements may occur or execute simultaneously, at the same time point, or synchronously. The disclosed embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings.

[0021] As used herein, proteomics can involve or include the quantitative assessment of proteins present in a sample obtained from a human or non-human animal. Proteomics can involve the classification or prediction of protein levels (e.g., presence, amount, concentration, etc.) measured on a large scale (e.g., involving hundreds to thousands). In this way, proteomics can be similar to genomics, but applied to the field of proteins.

[0022] As used herein, a prediction model can be a model that maps an input set to an output value or category. In some embodiments, the prediction model can be a machine learning model that can be trained using training data to generate appropriate outputs. The prediction model can be trained using supervised methods, semi-supervised methods, unsupervised methods, or reinforcement learning methods. In some embodiments, the prediction model can be a statistical model that estimates the relationship between an independent variable (e.g., protein levels in a biological sample) and a dependent variable (e.g., health information value) using training data. Exemplary prediction models include regression models (e.g., penalized regression models), support vector machines, decision trees or decision forests, linear discriminant analysis, cluster models, nearest neighbor models, neural network models, integrated models including one or more of any of the foregoing models, etc.

[0023] As used herein, a proteomic model can be a predictive model that is configured to receive protein level measurements and output a classification or prediction result based at least in part on these measurements.

[0024] Health information may include the probability of developing a health outcome (e.g., cardiovascular disease, dementia, kidney disease, etc.) within a specific time frame (e.g., within six months, one year, four years, ten years, twenty years, or longer), indications of current health status (e.g., body fat percentage, basal metabolic rate, lean muscle mass, aerobic fitness, visceral fat content, glomerular filtration rate, presence or absence of excess liver fat, glucose tolerance, etc.), behavioral health predictions (e.g., predicted weekly alcohol consumption, etc.), etc.

[0025] Consistent with the disclosed embodiments, providing a model, data, or instruction may include providing such model, data, or instruction directly (e.g., between a source and a target) or indirectly (e.g., through an intermediary). Providing a model, data, or instruction may include providing by numerating the model, data, or instruction, or providing by referencing a location where the model, data, or instruction is stored.

[0026] In accordance with the disclosed embodiments, performance metrics can be used to quantify the performance of the prediction model. In some embodiments, the performance metric can quantify the consistency between two prediction results (for example, to evaluate the impact of repeatability or input data differences on the prediction output). Such performance metrics may include Lin's consistency correlation coefficient, Pearson correlation coefficient or other suitable correlation metrics. In some embodiments, the performance metric can quantify the consistency between the prediction result and the true value (ground truth). In some embodiments, such performance metrics can quantify this consistency according to the relationship between the true positive rate and the false positive rate when the parameters of the prediction model change (for example, discrimination threshold, etc.). Such metrics may include receiver operating characteristic (ROC) curves or derived performance metrics, such as the area under the ROC curve (AUC). In some embodiments, such performance metrics may include or depend on the number of true positive results, true negative results, false positive results or false negative results of the classification task. In some embodiments, such metrics may include sensitivity, specificity, precision, false negative rate, false positive rate, prevalence (prevalence), accuracy, F1 score or other such metrics. In some embodiments, such performance metrics may include or depend on a class-by-class comparison between the predicted results and the true values, such as a confusion matrix or the like.

[0027] Figure 1An exemplary pipeline 100 for developing, validating, and deploying a proteomic model for predicting health information using aptamer-based blood tests consistent with the disclosed embodiments is described. The pipeline 100 may include components from which data is initially obtained, such as a measurement system 101 or a recorder 103. The pipeline 100 may include components for collecting and preparing the obtained data, such as a data input engine 110 or a data set generator 120. The pipeline 100 may include components for storing prepared data and machine learning models, such as a data storage 140 and a model storage 130. The pipeline 100 may include a training engine 160 for generating a trained machine learning model using the obtained data and model. The pipeline 100 may include a prediction engine 170 for predicting health information using the trained machine learning model. The components of the pipeline 100 may be managed and configured through a user device 180. The user device 180 may also be used to display the output of other components, raw or prepared data, trained or untrained models, or predicted health information. Consistent with the disclosed embodiments, a user may interact with components of pipeline 100 to perform model validation and development. In some embodiments, a user may interact with training engine 160 (or another suitable component of pipeline 100) to perform Figure 5 The model validation or Figure 6 In summary, pipeline 100 can provide a convenient, scalable platform for the development, validation, and deployment of the proteomics models disclosed herein.

[0028] The measurement system 101 can be a device suitable for obtaining data indicating the presence or concentration of proteins in a biological sample. For ease of discussion, the measurement system 101 is described herein as a microarray scanner. Such scanners can be configured to measure fluorescence at predetermined positions corresponding to different proteins on a microarray. Fluorescence intensity can indicate the concentration of proteins in a biological sample for preparing a microarray. In some cases, multiple positions can correspond to the same protein. The fluorescence intensity data of these multiple positions can be combined to better estimate the concentration of proteins in the sample. It is understood that the pipeline 100 is not limited to the embodiment in which the measurement system 101 is a microarray scanner. In addition, in some embodiments, the measurement system 101 can provide data to one or more recorders 103, rather than directly providing data to the data input engine 110. Then, the data input engine 110 can obtain the data from one or more recorders 103.

[0029] Consistent with the disclosed embodiments, the one or more recorders 103 may include one or more storage locations for storing data that the pipeline 100 may use to predict health outcomes. In some embodiments, such data may include fluorescence intensity data generated by the measurement system 101 from a biological sample. In multiple embodiments, such data may include medical record information from a patient from whom the biological sample was obtained. Such medical record information may include medical records, case records, requests, or application information provided by a physician or other clinician (e.g., related to the sample or the prediction to be performed by the pipeline 100). In some embodiments, the medical record information may include category or data label information corresponding to the sample (e.g., for generating a training data set).

[0030] Consistent with the disclosed embodiments, medical record information can be associated with health information. For example, in a cancer screening or diagnosis scenario, medical record information can include data related to cancer type, cancer incidence, cancer prognosis, or cancer stage.

[0031] Consistent with the disclosed embodiments, medical record information may include information suitable for generating personalized prediction models. For example, medical record information may indicate demographic or health characteristics of a patient, such as the patient's age, gender, race / ethnicity, height / weight. As another example, medical record information may indicate a patient's medical history, such as a medical history or medical treatment, clinical visit, or surgical history. As another example, medical record information may indicate behavioral factors that may affect a patient's health, such as smoking history, alcohol consumption, medication use, diet type, and the like. As another example, medical record information may include data obtained from a patient's biological sample, such as data related to blood, plasma, serum, or urine. As another example, medical record information may indicate family history data, genetic data, or immunological data.

[0032] Consistent with the disclosed embodiments, the data input engine 110 can be configured to acquire data from a plurality of data sources (e.g., the measurement system 101, one or more recorders 103, or other suitable sources) and process the data for use by other components of the pipeline 100. In some embodiments, the data input engine 110 can include a data extractor 111, a data converter 113, and a data loader 115.

[0033] Consistent with the disclosed embodiments, the data extractor 111 can receive or acquire data from the measurement system 101, one or more recorders 103, or other suitable sources. As described herein, the measurement system 101 can be a diagnostic system configured to obtain a fluorescence value corresponding to the presence or concentration of a protein in a biological sample. Similarly, the one or more recorders 103 can be one or more databases or data storage locations containing data related to the biological sample used to generate the fluorescence value. The disclosed embodiments are not limited to any particular format of the obtained data or the method of obtaining the data. For example, the obtained data can be structured data or unstructured data, or include structured data or unstructured data. The data extractor 111 can interact with multiple data sources, receive or acquire relevant data, and provide the data to the data converter 113.

[0034] Consistent with the disclosed embodiments, data converter 113 can receive data from data extractor 111 and process the data into a standard format. In some embodiments, data converter 113 can normalize data such as dates or values ​​according to specific units of measurement. For example, measurement system 101 can store dates in day-month-year format, and recorder 103 can include a recorder that stores dates in year-month-day format, or multiple recorders that store measurements in different units (e.g., weight in kilograms or pounds, fluorescence in absolute magnitude or log10 magnitude, etc.). In this example, data converter 113 can modify the data provided by data extractor 111 into a consistent date format or standardized unit format, respectively. Therefore, data converter 113 can effectively clean up the data provided by data extractor 111, so that all data (although from multiple sources) have a consistent format.

[0035] In addition, the data converter 113 can also extract additional data points from the data. For example, the data converter 113 can process dates in the year-month-day format by extracting separate data fields for the year, month, and day. The data converter 113 can also perform other linear and nonlinear transformations (e.g., logarithmic transformations of continuous data, etc.) and extractions on categorical and numerical data, such as normalization and centering the data at the data mean. The data converter 113 can provide the transformed or extracted data to the data loader 115.

[0036] Consistent with the disclosed embodiments, the data loader 115 may receive the normalized data from the data converter 113. The data loader 115 may merge the data into different formats according to the specific requirements of the dataset generator 120. The data loader 115 may then provide the processed data to the dataset generator 120 (or a suitable data storage from which the dataset generator 120 may obtain the data).

[0037] Consistent with the disclosed embodiments, the data set generator 120 can be configured to generate a data set from data processed by the data input engine 110. In some embodiments, the training engine 160 or the prediction engine 170 can be configured to require a data set with a specific structure. The data set generator 120 can be configured to format the data into the specific structure. For example, the data set generator 120 can be configured to collect observations, associate the observations with corresponding training labels or metadata, and store the observations and metadata in the data storage 140.

[0038] In some embodiments, the data set generator 120 can be configured to extract features from the data received by the data input engine 110. A feature can be a property or characteristic of a phenomenon. For example, a feature can be the presence or absence of a fluorescence value exceeding a threshold at a position corresponding to a specific protein. As another example, a feature can be the average fluorescence value of a sample on a biochip. Features can be determined based on a domain, a category data type, or many other factors associated with the data stored in the data structure. In addition, a feature can represent information for multiple data loggers in a data set or information for a single category in a data logger. In addition, multiple features can be generated to represent the same data.

[0039] In some embodiments, the data set generator 120 may include a classifier 121 and an annotator 123. Consistent with the disclosed embodiments, the classifier 121 may be configured to create classes from the data provided by the data set generator 120. In some embodiments, these classes may be data labels for training a predictive model. For example, a set of fluorescence data indicating protein levels may be associated with a patient's medical records. In this instance, the classifier may extract health information from the medical records. For example, the classifier may determine that a patient experienced a cardiovascular event within a specific time after obtaining a blood sample for generating fluorescence data. As another example, the classifier may identify a patient's glomerular filtration rate from a medical record. In some embodiments, the classifier may be a natural language processing engine, or include a natural language processing engine.

[0040] Consistent with the disclosed embodiments, the classifier 121 can be configured to accept a classification provided by a user via the user device 180. For example, the classifier 121 can be configured to provide data (or metadata relating to the data) received from the data input engine 110 to the user device 180 for display. In response, the classifier 121 can receive category information. For example, the classifier 121 can provide a medical record associated with the fluorescence data set for display to the user device 180. In response, the user can interact with the user device 180 to provide an indication of health information based on the provided medical record. It will be appreciated that such functionality is not limited to medical records, but can also be other information used to generate data training labels.

[0041] Consistent with the disclosed embodiments, the annotator 123 can be configured to receive data and associated value or category information from a dataset generator. The annotator 123 can be configured to create appropriately formatted entries that associate the value or category information with the data. For example, given an array or matrix of fluorescence values ​​and glomerular filtration rate, the annotator 123 can create an object including a "response_value" key and a "protein_data_input" key. The glomerular filtration rate can be stored with the "response_value" key and the array or matrix of fluorescence values ​​"protein_data_input" key. In this instance, the training engine 160 or the prediction engine 170 can require or be configured to require observations in this format. As another example, given a set of fluorescence values ​​obtained from a blood sample drawn from a patient and the date the patient had a heart attack, the annotator 123 can create a relational database where rows correspond to patients, one column stores whether the patient experienced a heart attack, another column stores the time of the patient's heart attack, and the remaining columns store fluorescence values.

[0042] Consistent with the disclosed embodiments, model storage 130 can be a storage location for prediction models available to training engine 160 or prediction engine 170. The disclosed embodiments are not limited to any particular implementation of data storage 140. Consistent with the disclosed embodiments, data storage 140 can be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0043] Consistent with the disclosed embodiments, data store 140 can be a storage location for prepared data sets that can be used by training engine 160 or prediction engine 170. The disclosed embodiments are not limited to any particular implementation of data store 140. Consistent with the disclosed embodiments, data store 140 can be implemented using one or more relational databases, object-oriented or document-oriented databases, tabular data stores, graph databases, distributed file systems, or other suitable data storage options.

[0044] Consistent with the disclosed embodiments, the data / model selector 150 may be configured to access the model storage 130 or the data storage 140 to obtain a prediction model or a data set, respectively. In some embodiments, the data / model selector 150 may provide an abstraction layer for the training engine 160 or the prediction engine 170. In some embodiments, the data / model selector 150 may be configured to control access to the model storage 130 or the data storage 140.

[0045] Consistent with the disclosed embodiments, the training engine 160 can be configured to train a prediction model, or to create and train a prediction model. The training engine 160 can be configured to obtain an existing model from the model storage 130, or to obtain a training data set from the data storage 140. In some embodiments, the training engine 160 can be configured to interact with the data / model selector 150 to obtain an existing model or a training data set. The training engine 160 can be configured to store the trained prediction model in the model storage 130. In some embodiments, the training engine 160 can be configured to interact with the data / model selector 150 to store the trained prediction model in the model storage 130.

[0046] Consistent with the disclosed embodiments, the training engine 160 may include a model trainer 161 and a model evaluation 163. The training engine 160 may be configured to train a prediction model using the model trainer 161 and then determine a performance metric of the prediction model using the model evaluation 163. In some embodiments, the training engine 160 may automatically update the prediction model being trained based on the performance metric. In multiple embodiments, the training engine 160 may update the prediction model being trained in response to user input provided through the user device 180. Updating the prediction model may include performing additional training (e.g., using an existing training data set or another training data set), modifying the model (e.g., changing the input features used by the model, changing the architecture of the model, etc.), or changing the training environment (e.g., changing training hyperparameters, changing the training data set partition to a training part, a cross-validation part, and a retention part, etc.). One or more.

[0047] Consistent with the disclosed embodiments, as described herein, the training engine 160 can be configured to use data obtained under different circumstances to determine performance metrics of the model. These performance metrics can be displayed to the user via the user device 180. The user can then interact with the training engine 160 via the user device 180 to update the model, as described herein.

[0048] Consistent with the disclosed embodiments, the model trainer 161 can create or train a prediction model. The model trainer 161 can create or train a prediction model as instructed by the training engine 160. For example, the training engine 160 can instruct the model trainer 161 to create and train a support vector machine using a training portion of a training data set. The model trainer 161 can then create and train the support vector machine and return the trained support vector to the training engine 160. As another example, the training engine 160 can instruct the model trainer 161 to create and train a penalized regression model. The training engine can configure the model trainer 161 with the type of penalized regression model (e.g., ridge regression, lasso regression, elastic net, or other suitable type), the parameter values ​​of the penalized regression (e.g., the lambda value that weights the sum of the squared values ​​of the coefficients), and the training portion of the training data set. The model trainer 161 can then create and train the penalized regression model and return the trained penalized regression model to the training engine 160. As another example, the training engine 160 can instruct the model trainer 161 to train a random forest model. The training engine 160 can provide hyperparameters such as the size of each bootstrap sample, the number of features to consider at each split, the depth of each decision tree, and the number of decision trees in the random forest. The training engine 160 can provide the training portion of the training data set. The model trainer 161 can then create and train a random forest model and return the trained random forest model to the training engine 160.

[0049] Consistent with the disclosed embodiments, model evaluation 163 can evaluate a model trained by model trainer 161. Training engine 160 can provide a cross-validation or hold-out portion of the model and training data set to model evaluation 163. In some embodiments, training engine 160 can specify one or more performance metrics for evaluation by model evaluation 163. In various embodiments, model evaluation 163 can be configured with a predetermined or default set of performance metrics. In some embodiments, performance metrics can include a confusion matrix, mean square error, mean absolute error, sensitivity or selectivity, a receiver operating characteristic curve or an area under such a curve, precision and recall, F-score, or any other suitable performance metric.

[0050] Consistent with the disclosed embodiments, the prediction engine 170 can be configured to predict health information using a patient data set and a trained prediction model. In some embodiments, the prediction engine 170 can obtain a trained prediction model from the model storage 130. In some embodiments, the prediction engine 170 can obtain a patient data set from the data storage 140. In some embodiments, the prediction engine 170 can obtain a patient data set (or a portion thereof) from another data storage location. This alternative data storage location can be associated with another entity or user. For example, the prediction engine 170 can receive or obtain a patient data set from a healthcare system independent of the entity controlling the prediction engine 170. In some embodiments, the prediction engine 170 can obtain a model or data using the data / model selector 150.

[0051] Consistent with the disclosed embodiments, the prediction engine 170 can apply the patient data set to the trained prediction model to predict the patient's health information. The prediction engine 170 can provide the health information (or an indication thereof) to the user device 180. The health information can be stored on a computing device associated with the pipeline 100, or provided to another system.

[0052] Consistent with the disclosed embodiments, the user device 180 may provide a user interface for interacting with other components of the pipeline 100. The user interface may be a graphical user interface. The user interface may enable a user to configure the data input engine 110 to extract, transform, and load data according to user instructions. The user interface may enable a user to specify how to transform the transformed data received by the dataset generator 120 into labeled training data (or patient data suitable for predicting an outcome). In some embodiments, the user interface may enable a user to interact with the dataset generator 120 to manually or semi-manually label or annotate the training data. In some embodiments, the user interface may enable a user to interact with the data / model selector 150 to manage data or models stored in the model storage 130 or the data storage 140. Such management may include deleting or creating models, deleting datasets, or limiting access to models or datasets by the training engine 160 or the prediction engine 170. In some embodiments, the user interface may enable a user to interact with the data / model selector 150 to push data or models to the training engine 160 for training, or to the prediction engine 170 for prediction. In some embodiments, the user interface may enable a user to interact with the training engine 160 to create or select a prediction model for training, create or select a data set for training a model, or select training parameters or hyperparameters. In some embodiments, the user interface may enable a user to interact with the training engine 160 to display information related to model training (e.g., performance metrics, changes in loss function values ​​during training, or other training information). In some embodiments, the user interface may enable a user to interact with the prediction engine 170 to select a training model and patient data for predicting health information. In some embodiments, the user interface may enable a user to interact with the prediction engine 170 to display health information, store health information on a computing device, or transmit health information to another system.

[0053] The components of pipeline 100 can be implemented using one or more computing devices. Such computing devices may include tablet computers, laptop computers, desktop computers, workstations, computing clusters, or cloud computing platforms. In some embodiments, the components of pipeline 100 can be implemented using cloud computing platforms. For example, one or more of data input engine 110, data set generator 120, data / model selector 150, training engine 160, and prediction engine 170 can be implemented on a cloud computing platform. In some embodiments, the components of pipeline 100 can be implemented using a local deployment system. For example, measurement system 101, recorder 103, or user device 180 can be a local deployment system, or hosted in a local deployment system. As another example, model storage 130 or data storage 140 can be a local deployment system, or hosted in a local deployment system.

[0054] The components of pipeline 100 can communicate using any suitable method. In some embodiments, two or more components of pipeline 100 can be implemented as microservices or network services. Such components can communicate using information transmitted on a computer network. Information can be implemented using SOAP, XML, HTTP, JSON, RCP or any other suitable format. In some embodiments, two or more components of pipeline 100 can be implemented as software, hardware or combined software / hardware modules. Such components can communicate using data or instructions written to memory (e.g., shared memory) or read from memory, function calls or any other suitable communication methods.

[0055] It will be appreciated that the specific structure of the pipeline 100 is not intended to be limiting. Consistent with the disclosed embodiments, any two or more of one or more recorders 103, model storage 130, or data storage 140 may be combined or hosted on the same computing device. Consistent with the disclosed embodiments, the data input engine 110 and the data set generator 120 may be omitted from the pipeline 100. In such an embodiment, the data set formatted and configured for the training engine 160 or the prediction engine 170 may be stored in the data storage 140 by another system or using another method. Consistent with the disclosed embodiments, the data input engine 110 and the data set generator 120 may be combined. In such an embodiment, data extraction, conversion, and loading may be combined with feature extraction, annotation, and classification. Consistent with the disclosed embodiments, the data / model selector 150 may be combined with one or more of the training engine 160 and the prediction engine 170. For example, the training engine 160 or the prediction engine 170 may include the functionality of obtaining selected data or models from the model storage 130 or the data storage 140.

[0056] Although one user device 180 is shown, there may be multiple user devices. Different user devices may be associated with different entities or different users with different responsibilities. For example, user device 180 may be associated with a software engineer or data scientist who is developing a test, while another user device may be associated with a clinician who is using the test.

[0057] The user device 180 may be combined with one or more other components of the pipeline 100. In some embodiments, the user device 180 and at least one of the data / model selector 150, the training engine 160, or the prediction engine 170 may be implemented by the same computing device. In various embodiments, the user device 180 and at least one of the model storage 130 or the data storage 140 may be implemented by the same computing device.

[0058] It will be appreciated that the pipeline 100 can be integrated into a method for treating patients with a particular health condition. The prediction engine 170 can use a trained prediction model and input data obtained from a patient sample to determine a patient's risk of experiencing a negative health outcome (e.g., a cardiovascular event or recurrent cardiovascular event within the next 4 years; dementia within the next 20 years; death from stable heart failure with reduced ejection fraction or heart failure with preserved ejection fraction within one year, etc.). If the patient's risk is greater than (or likely to be equal to) a health outcome dependency threshold, the patient can be treated or monitored according to a first, more aggressive or intensive regimen. If the patient's risk is less than (or likely to be equal to) a health outcome dependency threshold, the patient can be treated or monitored according to a second, less aggressive or intensive regimen.

[0059] Figure 2A-2H A high-level description of the stages in an exemplary aptamer-based serum or plasma assay 200 consistent with the disclosed embodiments is provided. The assay 200 can generate suitable input data for training a predictive model or predicting health information using a trained predictive model. The assay 200 can be performed at least in part using a test system such as the measurement system 101 of the pipeline 100.

[0060] Consistent with the disclosed embodiments, assay 200 can quantitatively convert protein epitope availability in a biological sample into a specific DNA signal. Typically, assay 200 can use (Slow off-rate modified aptamer) reagents comprising short single-stranded DNA sequences incorporating hydrophobic modifications. Assay 200 can measure native proteins in complex matrices by converting available binding sites on individual proteins into corresponding SOMAmer reagent concentrations, which can then be quantified by hybridization to a microarray. In this way, the test takes advantage of the dual properties of SOMAmer reagents, which are both protein affinity binding reagents with defined three-dimensional structures and unique nucleotide sequences that can be recognized by specific DNA hybridization probes. Therefore, relative epitope concentrations can be converted into measurable nucleic acid signals that can be quantified using DNA hybridization microarrays.

[0061] Consistent with the disclosed embodiments, a suitable version of the assay 200 can quantify the relative levels of proteins in plasma whose abundance spans 10 log units. Such a version of the assay can measure up to one thousand, three thousand, five thousand, seven thousand, ten thousand, or more unique protein analytes. The test can be performed on small volume samples (e.g., samples greater than 10 microliters, 20 microliters, 40 microliters, 100 microliters, 200 microliters, 400 microliters, 1 milliliter, or more).

[0062] It is understood that SOMAmer reagents may be selected for the native folded conformation of proteins. Therefore, such reagents may require intact tertiary protein structure for binding. Therefore, SOMAmer reagents may not detect (or may detect with reduced or altered sensitivity) proteins that are unfolded and denatured and therefore may be inactive.

[0063] like Figure 2A As shown, SOMAmer reagents can be synthesized using a fluorophore, a photocleavable linker, and biotin. Figure 2B As shown, SOMAmer reagents bound to streptavidin beads can be used to capture proteins from a complex protein mixture in a biological sample (e.g., a serum or plasma sample). Figure 2C As shown, unbound proteins can be washed away and bound proteins can be labeled with biotin. Figure 2D As shown, electromagnetic radiation (e.g., ultraviolet light, etc.) can be applied to the solution to destroy the photocleavable linker fragments, releasing the protein complex and the bound SOMAmer back into the solution. Figure 2E As shown, nonspecific complexes can dissociate from the corresponding SOMAmer, while specific complexes remain bound. Figure 2F As shown, a polyanionic competitor can be added to the solution. The polyanionic competitor can prevent the reassociation of nonspecific complexes. Figure 2GAs shown, the biotinylated protein (and bound SOMAmer reagent) can then be captured onto streptavidin beads. The beads and bound protein can be separated or concentrated from the solution. Figure 2H As shown, SOMAmer reagents can be released from protein complexes by denaturing the protein. The fluorophore can be measured after hybridization to the complementary sequence on the microarray chip. The fluorescence intensity detected on the microarray can be correlated with the amount of available epitopes in the original sample.

[0064] It is understood that detection 200 is intended to be exemplary. The disclosed systems and methods are not limited to detection with these specific steps. In some embodiments, other aptamers (or even other types of components) can be used to bind protein complexes. Alternative methods for suppressing nonspecific binding can be used to replace polyanionic competitors. Alternative methods can be used to separate protein-compound complexes, instead of capturing protein-compound complexes on streptavidin beads. Alternative markers of protein levels can be measured, instead of fluorophore measurements on microarray chips. However, such alternative methods may still present technical challenges described herein. Therefore, such alternative systems and methods can benefit from the disclosed technical solutions.

[0065] Figure 3A and Figure 3B Described are exemplary challenges that arise in validating a proteomic model that predicts health information using aptamer-based serum or plasma tests consistent with the disclosed embodiments. In some embodiments, such a technical challenge can be to ensure consistency of prediction results across scenarios. Such a challenge may arise when a proteomic model developed using input data obtained in a first scenario must be validated before it can be used with input data obtained in another scenario.

[0066] In accordance with the disclosed embodiments, the situational differences may include differences between samples, differences in sample handling methods, or differences in assay protocols. In accordance with the disclosed embodiments, the differences between samples may include the presence or absence of interfering agents in the sample, the use of citrate plasma and EDTA plasma, the use of serum and plasma, the feeding / fasting state of the patient providing the sample, or other changes in sample characteristics that may affect the effectiveness of the proteomics prediction model. The disclosed embodiments are not limited to any specific interfering agents. In multiple embodiments, interfering agents may include nonsteroidal anti-inflammatory drugs (NSAIDs), contraceptives, antihypertensive drugs, mental health drugs (e.g., antidepressants, antipsychotics, antianxiety drugs or hypnotic drugs, mood stabilizers, stimulants, etc.), cholesterol-lowering drugs, asthma drugs, diabetes drugs, thyroid drugs, antiviral drugs or antibacterial drugs or other commonly used drugs. In accordance with the disclosed embodiments, the differences in sample handling methods may include the time difference between sample collection and sample freezing, sample freezing temperature, duration at freezing temperature, sample freeze-thaw cycle number, plasma sample centrifugation time, serum sample coagulation or decantation time, or other changes in sample handling method conditions that may affect the effectiveness of the proteomics prediction model. Consistent with the disclosed embodiments, the differences in the assay protocols may include differences in the reagents used, different dilutions of the same reagents, adding or removing steps in the assay protocol, or differences in the equipment used to perform the assay protocol. For example, a first version of an assay protocol may be used to detect the concentrations of 5,000 proteins, while a second version of an assay protocol may be used to detect the concentrations of 7,000 proteins. The two versions of the assay protocol may use different aptamer sets, different microarray chips, and possibly different microarray scanners.

[0067] Consistent with the disclosed embodiments, suitable input data obtained in the second scenario may not be available. The original sample used to generate the original input data may be lost, exhausted or decomposed over time. In addition, the currently available sample may be different from the original sample. For example, the original sample may contain patient samples suffering from health conditions for which the predictive proteomics model is developed for detection. For example, the original sample may have been obtained over the years through cooperation with a medical center that specializes in treating patients with conditions for which the predictive proteomics model is developed for detection. However, such patients may be extremely rare in the general population. The currently available sample may be obtained from the general population. For example, the currently available sample may be obtained from a routine blood draw at a community health center. Therefore, the currently available sample is unlikely to include patients suffering from conditions for which the predictive proteomics model is developed for detection.

[0068] Figure 3ADescribes an exemplary correlation between the output of a predictive proteomics model for input data obtained using two different assay protocols. In this small example, the output represents the risk of death within one year for a person with a rare disease. The two different assay protocols are a 5000 protein protocol and a 7000 protein protocol. The 7000 protein protocol includes the 5000 proteins and an additional 2000 proteins. The predictive proteomics model is developed using the input data obtained according to the 5000 protein protocol. The input data obtained using the 7000 protein protocol can be input into the predictive proteomics model by truncating the input to include only the common 5000 proteins.

[0069] In this example, a collection of samples collected from patients with a rare disease is available. Two input data sets can be generated for each sample, one based on a 5000 protein solution and the other based on a 7000 protein solution. These input data sets can be applied to a predictive proteomics model to generate the probability of death within one year. Figure 3A It is evident that these predicted probabilities are highly correlated, with a Lin's concordance coefficient (CCC) of 0.89.

[0070] Figure 3B An exemplary correlation between the output of the same predictive proteomics model for the same 5000 protein scenario and 7000 protein scenario is depicted. In this example, a sample set collected from patients with a rare disease was not available. Instead, samples from healthy normal patients were used. Figure 3A As described above, two input data sets can be generated for each sample, one set generated based on a 5000 protein solution and the other set generated based on a 7000 protein solution. These input data sets can be applied to a predictive proteomics model to generate a probability of death within one year. Figure 3B Obviously, the correlation between these predicted probabilities is not very high, with a CCC of 0.31.

[0071] Understandably, this lack of correlation is due to the severely restricted range of predicted probabilities. The predictive proteomics model correctly found that a healthy normal patient had a very low probability of dying from the disease within a year. Figure 3B The predicted probability range described in Figure 3A The range of predicted probabilities described in is about 20 times smaller. Therefore, a simple test using the available sample set would underestimate the reliability of the predictive proteomic model when using the 7000 protein scenario.

[0072] 4A to 4FExemplary model resilience challenges that arise in training a proteomic model for predicting health information using aptamer-based blood tests consistent with the disclosed embodiments are described. FIG. 2A to FIG. 2H As will be appreciated from the depiction of an exemplary SOMAmer-based blood test shown, aptamer-based blood tests are sensitive to changes in sample handling that may affect protein shape or concentration (e.g., by degradation over time, reaction with other sample components, etc.). Changes in protein levels will obviously affect the measured protein levels. Similarly, any denaturation of proteins will affect the ability of aptamer complexes to bind to those proteins, thereby affecting the measured protein levels.

[0073] Thus, sample processing conditions can affect the output of proteomic models that use protein levels as input. The dependence of the predicted output on sample processing conditions can be independent of the predictive power of the model. For example, two proteomic models may have similar predictive power, but the first model may show significant differences in response to a specific change in how the sample was processed, while the second model may be resilient to this specific change. This difference may arise from the specific proteins that the different models rely on. For example, two proteins may provide similar or related information about health conditions. A parsimonious model may rely on one of the proteins but not both. However, one of the two proteins may be much more stable than the other. Therefore, even though the two proteins may be equivalent from a prediction perspective, a resilient model can be designed to rely on the more stable protein and abandon the less stable protein.

[0074] Figure 4A and Figure 4B Dependence of two proteomic models trained to predict likelihood of chronic kidney disease on sample processing conditions consistent with the disclosed embodiments is described. In this study, sample processing varied in the time between sample collection and transportation of the samples on dry ice to a testing laboratory. Multiple healthy normal samples were collected and six different times between sample collection and transportation were studied. Multiple aliquots were prepared from each sample, each corresponding to one of the six different transportation times. Two prediction models were trained.

[0075] Figure 4A The dependence of the probability of chronic kidney disease predicted by the first proteomic model on the shipping time is described. As the shipping time increases, the predicted probability of chronic kidney disease increases. The median predicted probability increases, and significant outliers appear in the predicted values. It is understandable that clinicians may not point out delays in the shipping of samples to the laboratory. Therefore, the laboratory may provide erroneous predictions due to the dependence of the predicted risk on the shipping time.

[0076] Figure 4B Dependence of the likelihood of developing chronic kidney disease predicted by a second proteomic model on transit time is depicted. The second proteomic model was developed by updating the first proteomic model based on its determined sensitivity to transit time differences. The contribution of proteins that are sensitive to transit time differences to the model was reduced compared to the first model. As can be appreciated, the second predictive proteomic model predicts less increased risk compared to the first predictive proteomic model. In addition, the variability of the predicted risk was reduced, which is shown in the figure as a slight increase in the interquartile range and a decrease in the number and range of outliers.

[0077] In addition to changes caused by sample processing, evaluation on external samples can also identify overfitting issues in trained models. Figure 4C The variability of the risk of developing chronic kidney disease predicted by the first proteomic model over time is described. In this study, the predicted risk was determined using repeated samples from the same patient over time. These samples were obtained at intervals over a 12-month period. It is understood that patients who initially showed an increased risk of developing chronic kidney disease would have their risk remain constant throughout the testing interval. Figure 4C As shown, the predicted risk of some patients changed dramatically over time, which may be due to overfitting of the original model, which then provided unstable prediction results for the external dataset.

[0078] Figure 4D The variability of the relative risk of chronic kidney disease predicted by the second proteomic model over time is described. The second model shows a reduced variability in predicting risk. Such a model can be generated by identifying at least one protein that contributes to the first prediction model and that exhibits variability between samples. The contribution of the protein to the first model can be reduced. In some embodiments, the first model can be retrained to generate the second model.

[0079] In addition to changes caused by sample handling, measured protein levels may also exhibit process variability. Some measured protein levels may exhibit greater measurement variability and lead to model overfitting. Proteomic models that rely on protein level values ​​for such proteins may exhibit greater variability in predictions than proteomic models that do not rely on protein level values ​​for such proteins, even between aliquots taken from the same sample.

[0080] Figure 4EDescribed is a risk value determined for two different aliquots of the same sample using a first proteomic model. In this example, the disease diagnosis depends on the predicted risk value. The values ​​shown as open circles give different diagnoses for different aliquots of the sample (e.g., the upper left quadrant and the lower right quadrant indicate a negative diagnosis for one aliquot and a positive diagnosis for another aliquot). The values ​​shown as closed circles are the same for the two aliquots of the sample.

[0081] Figure 4F Depicts the risk values ​​determined for two different aliquots of the same sample using a second proteomic model. Figure 4E Compared with the prediction results in , the samples with diagnostic differences between different aliquots of the same sample are much less. A second proteomic model can be generated from the first proteomic model by identifying proteins that contribute to the first proteomic model and show high within-aliquot variability, and reducing the contribution of these proteins when generating the second proteomic model.

[0082] Figure 5 An exemplary process 500 for validating a proteomics model consistent with the disclosed embodiments is described. The proteomics model can be validated against input data collected using different data collection schemes. The process 500 can be performed using a machine learning pipeline, as described above with respect to Figure 1 The pipeline 100 described herein. For ease of description, process 500 is described herein as being performed using training engine 160. However, process 500 is not limited to this implementation. Consistent with the disclosed embodiments, process 500 may be performed using other components of pipeline 100 or other machine learning systems. For example, process 500 may be performed using an independently running system or computing device for training a machine learning model.

[0083] Process 500 can be as described above. Figure 3A and Figure 3B The technical problem described provides a technical solution. Consistent with the disclosed embodiments, when a proteomics model is developed for use in a first situation, process 500 can verify the proteomics model for use in a second situation. In some embodiments, the input data for developing the proteomics model may be available, but the original samples that generated the input data may no longer be available. Moreover, the available samples may not cover all potential inputs of the model. For example, these available samples may be mainly or entirely from "healthy normal" patients. Therefore, the available samples may not support the verification of all proteomics model outputs.

[0084] Consistent with the disclosed embodiments, process 500 can compensate for unavailable original samples by determining protein level noise using available samples. Protein level noise can characterize the difference in protein level measurements between a first scenario and a second scenario. Protein level noise can be used together with the input data originally used to develop the proteomics model to generate a validation data set. The validation data set can then be used to validate the proteomics model in the second scenario. Therefore, even if the available samples do not cover all potential inputs to the model, these samples can still support validation of all proteomics model outputs.

[0085] Consistent with the disclosed embodiments, in step 510 of process 500, the training engine 160 may obtain a control data set, a processed data set, and a training data set. In some embodiments, these data sets may include protein level measurements (e.g., measurements of relative or absolute protein levels in aliquots of a sample). Such protein level measurements may be numerical values. For example, a data set may include numerical values ​​in relative fluorescence units (RFU) or fluorescence units (FLU). These values ​​may indicate protein levels. In some embodiments, the control data set and the processed data set may be paired data sets, or include paired data sets. Such paired data sets may be generated using multiple aliquots of the same sample or paired samples collected from the same patient. The control data may be generated using a first scenario, while the processed data set may be generated using a second scenario. The training data set may be generated using the original sample in the first scenario and should represent the same input as the data used to train the model.

[0086] In some embodiments, the first scenario and the second scenario can detect different sets of proteins. For example, the first scenario can be a previously developed assay, or include a previously developed assay, and the second scenario can be a next generation assay, or include a next generation assay. The next generation assay can detect a superset of the proteins detected in the previously developed assay. For example, a previously developed assay can detect 5000 proteins, and the next generation assay can detect those 5000 proteins, plus an additional 2000 proteins.

[0087] In some embodiments, the first scenario and the second scenario can use different sample processing techniques. For example, the control data set can be generated from a citrate plasma sample, and the treatment data set can be generated from an ethylenediaminetetraacetic acid potassium (EDTA) plasma sample. The citrate plasma sample and the EDTA plasma sample can be obtained from the same patient. As another example, the first scenario and the second scenario can use different times from blood drawing to centrifugal sample, different times from centrifugal sample to decanted sample, different times from decanted sample to frozen sample, different freezing durations or storage temperatures, different freeze-thaw cycle times, etc.

[0088] In some embodiments, the first scenario and the second scenario can use different assay protocols. For example, such differences can include differences in reagents used, different dilutions of the same reagents, increasing or decreasing steps in the assay protocol, or differences in the equipment used to perform the assay protocol. As another example, the first scenario can include a first set of aptamers, and the second scenario can include a second set of aptamers. The second set can be at least in part a revised version of the first set of aptamers.

[0089] In some embodiments, the second scenario may include an interfering agent that is not present in the first scenario. For example, preparing an aliquot according to the second scenario may include adding an interfering agent to the aliquot. Thus, process 500 may be used to determine whether the presence of a commonly used drug in a patient's blood will negatively impact the performance of a predictive proteomics model.

[0090] The disclosed embodiments are not limited to any particular method of obtaining control data sets, processed data sets, and training data sets. In some embodiments, the training engine 160 can obtain these data sets from the data storage 140 or another location. In multiple embodiments, the training engine 160 can obtain these data sets from another system (e.g., a health care system, an insurance system, etc.). In some embodiments, the control data sets, processed data sets, and training data sets can be generated from aliquots or samples using proteomic assays (such as assay 200). However, the disclosed embodiments are not limited to embodiments using assay 200. In addition or alternatively, other proteomic assays or other types of assays can be used to generate control data sets, processed data sets, and training data sets that can be used to validate the prediction model, which is consistent with the disclosed embodiments.

[0091] In step 520 of process 500, training engine 160 can estimate protein level noise, which is consistent with the disclosed embodiments. In some embodiments, protein level noise can be estimated on a per-protein basis. Training engine 160 (or another component of pipeline 100, such as data set generator 120, etc.) can determine the paired differences between the protein level measurements of the processing data set and the corresponding protein level measurements of the control data set. For example, a first entry in the control data set and a second entry in the processing data can correspond to aliquots extracted from the same sample. The first entry and the second entry can include protein level measurements of the same protein. Training engine 160 can subtract the protein level measurements of the first entry from the protein level measurements of the second entry to determine the paired differences for each protein. Training engine 160 can determine this paired difference for each entry in the control data set and each corresponding entry in the processing data set.

[0092] Consistent with the disclosed embodiments, the training engine 160 can characterize the distribution of the pairwise differences for each protein. Characterizing the distribution can include estimating the distribution of the pairwise differences (e.g., using a pairwise difference histogram, estimating parameters of a parametric model of the pairwise differences, etc.), estimating statistics (e.g., mean, median, mode, standard deviation, quartiles, percentiles, etc.) or moments (e.g., first moment, second moment, third moment, etc.) of the distribution of the pairwise differences, maintaining a set of pairwise differences and resampling from the set, or using other suitable methods for characterizing the distribution of pairwise differences.

[0093] It is understood that the control data set and the treated data set may include protein level measurements for different sets of proteins, depending on the difference between the first scenario and the second scenario. For example, the treated data set may include protein level measurements for a superset of proteins in the control data set. In some embodiments, pairwise differences may be determined and the distribution of proteins present in the two databases may be characterized.

[0094] In step 530 of process 500, training engine 160 can generate a validation data set, which is consistent with the disclosed embodiments. The validation data set can include corrected protein level values. Corrected protein level values ​​can be generated using estimated protein level noise and the training data set. In some embodiments, the validation data set can include multiple entries. Each entry can correspond to an entry in the training data set. In some embodiments, training engine 160 can generate corrected protein level measurements for each protein in each entry.

[0095] Consistent with the disclosed embodiments, the corrected protein level values ​​can be a function (e.g., a sum, etc.) of the protein level values ​​in the training data set and the noise samples. In some embodiments, the training engine 160 can generate the noise samples. The generation of the samples can be based on the distribution characterization in step 520. In some embodiments, when the shape of the distribution is estimated (e.g., using a histogram), the distribution with the estimated shape can be sampled. In some embodiments, when a statistic or moment is estimated, the statistic or moment can be used to generate samples. For example, when the mean and standard deviation of the protein in the training data set are estimated, a normal distribution with the estimated mean and standard deviation can be sampled to generate a sample of the protein. In some embodiments, the paired difference set of the protein can be resampled to generate a sample of the protein.

[0096] In step 540 of process 500, the training engine 160 can generate a set of prediction results by applying the validation data set to the proteomics model, which is consistent with the disclosed embodiment. The proteomics model can be a model developed using a training data set. Each entry in the validation data set can be used to generate a corresponding prediction result in the prediction result set. How to apply the validation data set to the proteomics model to generate a set of prediction results depends on the implementation of the proteomics model. In some embodiments, the proteomics model can be implemented as an instance of an object. The object can specify a method for training the model and using the model to predict. Then, using the model to generate prediction results can include calling the method using the entry as input. For example, assume that the function svm.svc (parameters) returns a support vector machine object with certain parameters. This object can have a fit (x, y) method for training data (using x and y training data) and a predict (x) method for generating a class label vector using a sample vector (each sample vector includes a feature vector). In this simple example, generation, training and prediction can be:

[0097] proteomics_model=svm.svc(parameters);

[0098] proteomics_model.fit(x_training, y_training); and

[0099] classifications=proteomics_model.predict(x_test).

[0100] As another example, the function linear_model.LogisticRegression(parameters) can return a penalized logistic regression object with certain parameters. As with the support vector machine object above, this object can specify a fit() method for training the object and a predict() method for generating predictions using input data.

[0101] It will be appreciated that the above examples of support vector machines and penalized logistic regression objects are illustrative and are not intended to be limiting. The specific manner in which the validation dataset is applied to the proteomics model will depend on the specific implementation of the proteomics model.

[0102] In step 550 of process 500, the training engine 160 can determine one or more performance metrics of the proteomics model using the prediction result set generated in step 540. In some embodiments, the training engine 160 can determine a performance metric that depends on the consistency of two prediction result sets, including the original prediction result set generated by the proteomics model using the training data and the prediction result set generated in step 540. In some embodiments, the performance metric can be a Lin's consistency correlation coefficient, a Pearson correlation coefficient, or other suitable performance metrics. In some embodiments, the training engine 160 can determine a performance metric that depends on the consistency between the true value associated with the training data and the prediction result set generated in step 540. In some embodiments, the true value can be specified by a label associated with an entry in the training data set. Such labels can specify the presence or absence of symptoms, binned survival time, or other suitable true values ​​involving patient health information.

[0103] Consistent with the disclosed embodiments, the training engine 160 can determine the validity indication based on one or more performance metrics. In some embodiments, the validity indication can be the value of one or more performance metrics. In multiple embodiments, the validity indication can depend on the value of one or more performance metrics. For example, the values ​​of the performance metrics can be binned or thresholded, where values ​​in a certain bin or below a certain threshold are designated as "warning" or "failed" indications, and values ​​in another bin or above another threshold are designated as "passed" or "valid" indications.

[0104] Consistent with the disclosed embodiments, process 500 may include Figure 5 Optional operations not shown in . This operation can be performed by the training engine 160 after obtaining the control data set and the processing data set (for example, in step 510). In some embodiments, the training engine 160 can determine the paired differences between the control prediction result set generated by applying the control data set to the proteomics model and the processing prediction result set generated by applying the processing data set to the proteomics model. It can be understood that when the control data set and the processing data set include overlapping protein sets, only the paired differences of the proteins in the set intersection can be calculated. In some embodiments, the training engine 160 can determine the statistics of the paired differences (for example, the mean, median, 75th percentile, 90th percentile or other suitable thresholds). In some embodiments, the training engine 160 can determine whether the statistics meet the invalid condition. The invalid condition can depend on the characteristics of the detection 200. For example, the prediction model can have an inherent variability due to the inherent noise of the detection 200 (for example, the protein level noise determined by the test-retest value difference of the protein level, etc.).

[0105] In some embodiments, the invalid condition may be at least partially satisfied when the paired difference statistic exceeds a function of the inherent variability statistic. For example, the inherent variability statistic may be the standard deviation σ of the inherent variability. The function of the statistic may be a multiple of the statistic (e.g., a multiple selected in the range of 0.5 to 5, such as 3 or another suitable value). In some embodiments, the invalid condition may be satisfied when the paired difference statistic exceeds a multiple of the standard deviation of the inherent variability. In some embodiments, satisfying the invalid condition may also require that a statistical test (e.g., a paired t-test, etc.) reject the null hypothesis that the paired difference statistic is equal to zero.

[0106] Consistent with the disclosed embodiments, if the invalidation condition is satisfied, process 500 may proceed to step 520, otherwise process 500 may terminate.

[0107] Figure 6 An exemplary process 600 for developing proteomics models consistent with the disclosed embodiments is described. The development of proteomics models can be structured to enhance the resilience of these models to changes in input data. As described herein, input data sets can be generated under different circumstances, and changes in the input data may arise from differences between these circumstances. Process 600 can be performed using a machine learning pipeline, as described above. Figure 1 For ease of description, process 600 is described herein as being performed using training engine 160. However, process 600 is not limited to this implementation. Consistent with the disclosed embodiments, process 600 may be performed using other components of pipeline 100 or other machine learning systems. For example, process 600 may be performed using an independently running system or computing device for training a machine learning model.

[0108] Process 600 can be as described above. 4A to 4F The described technical problem provides a technical solution. Consistent with the disclosed embodiment, process 600 can enable the development of a proteomics model that is flexible to differences in the situations in which an input data set is obtained. It is understood that the training data for developing the proteomics model can be obtained from samples that are processed in strict accordance with the sample processing guidelines. However, samples received from clinicians, health care systems, laboratories, or other users may not be processed in strict accordance with the sample processing guidelines. Therefore, consistent with the disclosed embodiment, the proteomics model can be developed to be flexible to changes in the sample processing method. In addition, the proteomics model can be developed to be flexible to changes in the detection process or the equipment used to measure protein levels.

[0109] In some embodiments, process 600 can be combined with process 500 described above. For example, a proteomics model may be developed using samples that are no longer available. Testing the elasticity of a proteomics model may require obtaining many samples in many different situations. Process 500 can enable the training data used to develop a proteomics model to be used for elasticity testing. Applying process 500, a control data set can be obtained using the same situation as the training data, while a plurality of processed data sets can be obtained using changes in the situation in which the data is obtained. The effects of these changes can then be studied by estimating protein level noise and generating a validation data set corresponding to each processed data set.

[0110] In some embodiments, process 600 can be performed as part of an iterative process of training and development. For example, a user can interact with training engine 160 (e.g., via user device 180) to select and train a proteomics model. The user can then interact with training engine 160 to repeatedly perform process 600 until a satisfactory proteomics model is developed. The model can exhibit a desired degree of resilience across a determined set of situation variations.

[0111] In step 610 of process 600, the training engine 160 can obtain a proteomics model, which is consistent with the disclosed embodiments. The proteomics model can have been trained (e.g., by the training engine 160 or another system) to generate prediction results based on protein level measurements. The proteomics model can have been trained using a training data set obtained from the training samples in the first scenario. The training engine 160 can obtain the proteomics model from a model memory (e.g., model memory 130) or another suitable storage location.

[0112] In step 620 of process 600, training engine 160 can obtain a processing data set including protein level measurements, which is consistent with the disclosed embodiments. In some embodiments, the processing data set can be generated by detecting a processing sample obtained in a processing situation. In some embodiments, each processing situation can be different in one or more aspects. For example, the aspect can be the centrifugation time, and the centrifugation time of the processing situation can be 0.5, 1.5, 3, 9 and 24 hours. As another example, the aspect can be the number of freeze-thaw cycles, and the processing situation can have 2, 3, 4, 5 and 10 freeze-thaw cycles. As described herein, the processing data set can be generated from the processing sample by pipeline 100 using detection (such as detection 200) and measurement system (such as measurement system 101). Alternatively or in addition, the processing data set can be obtained from another system by pipeline 100 (or training engine 160).

[0113] In some embodiments, each processed data set can be applied to a proteomics model to generate a set of predicted results. Then, these predicted results can be analyzed in step 630 of process 600. In a plurality of embodiments, a control data set can also be obtained in step 620 of process 600. The control data set can be generated from a sample obtained under the same circumstances as the sample used to generate the training data set (e.g., a data set used to train the proteomics model). As described in process 500, the control data set, the training data set, and each processed data set can be used to generate a validation data set, each validation data set corresponding to the processed data set and generated using the processed data set.

[0114] In step 630 of process 600, a dependency of the proteomics model on the difference between the second scenario can be determined, consistent with the disclosed embodiments. In some embodiments, the dependency can be determined based on one or more performance metrics of the proteomics model. In multiple embodiments, the dependency can be determined based on a display metric related to the set of prediction results generated in step 620.

[0115] Consistent with the disclosed embodiments, the training engine 160 may determine one or more performance metrics for the proteomics model. The training engine 160 may determine one or more performance metrics for one or more processed (or validation) data sets. In some embodiments, the training engine 160 may determine one or more performance metrics for a control (or training) data set. In some embodiments, the training engine 160 may determine a performance metric that depends on the consistency between the prediction results generated by applying the control data set (or training data set) to the proteomics model and the prediction results generated by applying the processed data set (or validation data set) to the proteomics model. In some embodiments, the training engine 160 may determine a performance metric that depends on the consistency between the true value and the prediction results generated by applying the control data set or the processed data set (or training data set or validation data set) to the proteomics model.

[0116] Consistent with the disclosed embodiments, the training engine 160 may provide a user interface accessible via the user device 180. The user interface may enable the training engine 160 to display the values ​​of one or more performance metrics. In some embodiments, the user interface may display a box plot, qq plot, histogram, table, or other suitable display of the set of prediction results generated in step 620. Examples of such displays include 4A to 4F , 7A to 7D , FIG. 8A to FIG. 8D as well as FIG. 9A to FIG. 9B .

[0117] Consistent with the disclosed embodiments, the dependency of the proteomics model on the difference between the second scenario can be automatically determined. In some embodiments, the training engine 160 can automatically determine whether there is a difference between one or more performance metrics generated using the control data set and the processed data set (or between the training data set and the validation data set). In some embodiments, the training engine 160 can automatically determine whether these differences are statistically significant. In multiple embodiments, the user can determine the dependency based on the display of one or more performance metric values. In some embodiments, the dependency of the proteomics model can be semi-automatically determined. The training engine 160 can automatically determine the value of one or more performance metrics, and the user can determine the dependency based on the display of one or more performance metric values. In some embodiments, the dependency of the proteomics model can be manually determined. The training engine 160 can provide a display of the set of prediction results generated in step 620, and the user can determine the dependency based on the display.

[0118] In step 640 of process 600, the proteomics model can be updated according to the dependencies determined in step 630, consistent with the disclosed embodiments. In some embodiments, proteins in the input data set that are applied to the proteomics model can be identified. The proteomics model can then be updated in a manner that reduces the importance of the proteins in the proteomics model.

[0119] In accordance with the disclosed embodiments, the identification of proteins can be performed according to the impact of proteins on the variability of the prediction model between the control data set and the processed data set or between the training data set and the validation data set (e.g., "variability effect"). The mode of the impact and the mode of identification can depend on the implementation of the proteomics model. In some embodiments, the variability effect can depend on the importance of the protein to the model. For example, the importance of the protein to the regression proteomics model can depend on the coefficient size of the protein in the regression proteomics model. As a further example, the importance of the protein to the SVM proteomics model can depend on the feature weight of the protein in the SVM proteomics model. As a further example, the importance of the protein to the random forest model can be calculated using Gini importance, average accuracy decline, importance based on arrangement, importance based on Shapley value, or any other suitable method. In some embodiments, the variability effect can depend on the variance of the protein level between the control data set and the processed data set (or the training data set and the validation data set). In some cases, the greater the variance of the measured protein level, the greater the impact of the protein on the variability of the prediction model.

[0120] In accordance with the disclosed embodiments, the proteomics model can be updated to reduce the variability effect of the identified proteins. The variability effect can be reduced by reducing the size of the weight or coefficient associated with the protein, excluding the protein from the input data set, retraining the model, or another suitable way. In some embodiments, the training engine 160 can automatically update the proteomics model. For example, the training engine 160 can automatically identify the protein with the largest variability effect. Then, the training engine 160 can update the model to reduce the variability effect of one or more of these proteins. In some embodiments, the training engine 160 can semi-automatically update the proteomics model. For example, the training engine 160 can automatically identify the protein with the largest variability effect and display the tags of these proteins to the user. Then, the user can select zero, one or more identified proteins. Then, the training engine 160 can update the proteomics model to reduce the variability effect of any selected one or more proteins. In some embodiments, the user can interact with the user interface provided by the training engine 160 to identify the protein with the largest variability effect. For example, the user can check the model weights or coefficients, the model architecture, or the calculated contribution of different proteins to the model output. The user may then provide instructions to the training engine 160 to update the proteomics model to reduce the effects of variability in any identified protein or proteins.

[0121] It is understood that identification of proteins may involve more than simply removing the proteins with the largest variability effects from the model. In some cases, such proteins may provide information essential to the function of the proteomics model. In such cases, the user may determine that differences in conditions between the control and treatment datasets are identified as detrimental to the performance of the proteomics model, rather than removing proteins. For example, when serum clotting delays affect the output of the model, and removing the protein causing this effect would unduly impair the performance of the model, the instructions for collecting samples may be updated to emphasize that serum clotting should be performed as instructed and not delayed.

[0122] 7A to 7D The effects of differences in sample processing consistent with the disclosed embodiments on the output of a proteomics model are described. In this example, the samples are serum samples. The sample processing differs in four aspects: the number of freeze-thaw cycles, the serum coagulation time, the serum decantation time, and the serum freezing (at -80°C) time. A proteomics model is developed using training data. In this example, the proteomics model is trained using known important aptamers plus five randomly selected aptamers. Consistent with process 500 and steps 610 to 630 of process 600, prediction results are generated using the training data set, the control data set, and the processed data set. 7A to 7DBox plots of these predictions by class label are shown. It is understood that a proteomics model may exhibit stability or monotonic changes in each of the four aspects. If updating the model to exhibit such stability or monotonic changes in one aspect of sample processing unduly degrades the model's performance, then the sample collection instructions can emphasize the importance of this aspect of sample processing.

[0123] Fig. 7A Depicts the dependence of the predicted probability values ​​generated by applying the entries in the training or validation datasets to the proteomics model on the number of freeze-thaw cycles of the samples. It can be observed that the number of outlier observations increases rapidly between the baseline sample treatment value and two freeze-thaw cycles, and then remains high as the number of freeze-thaw cycles increases.

[0124] Figure 7B Dependence of the predicted probability values ​​on the variation of the sample clotting time is depicted. As shown, the median predicted probability value for negative class observations increases with increasing clotting time. The distribution of predicted probability values ​​for negative class observations also shows an increasing skew toward inappropriate, higher predicted probability values. For positive class observations, the number of outliers with too low predicted values ​​increases with increasing clotting time.

[0125] Figure 7C Dependence of the predicted probability values ​​on the change in the sample decantation time is depicted. As shown, the median predicted probability value for negative class observations increases with increasing decantation time. The distribution of predicted probability values ​​for negative class observations also shows an increasing skew toward inappropriate, higher predicted probability values. For positive class observations, the number of outliers with too low predicted values ​​increases with increasing decantation time.

[0126] Fig.7D Describes the dependence of the predicted probability values ​​on the change in the freezing time of the samples at -80°C. As shown in the figure, increasing the freezing time beyond the baseline immediately affects the distribution of the predicted probability values, increasing the number of outliers in the negative and positive class samples.

[0127] FIG. 8A to FIG. 8DThe effects of differences in sample processing consistent with the disclosed embodiments on the output of a proteomics model are described. In this example, the samples are plasma samples. Sample processing differs in four aspects: the number of freeze-thaw cycles, the plasma freezer temperature during 24-day storage, the time the plasma is frozen at -80°C, and the time the plasma is centrifuged. A proteomics model is developed using training data. In this example, the proteomics model is trained using known significant aptamers plus five randomly selected aptamers. Consistent with process 500 and steps 610 to 630 of process 600, prediction results are generated using the training data set, the control data set, and the processed data set. FIG. 8A to FIG. 8D Box plots of these predictions are shown.

[0128] Fig. 8A Describes the dependence of the predicted output values ​​generated by applying the entries in the training or validation dataset to the proteomics model on the number of freeze-thaw cycles of the samples. It can be observed that the median predicted output value increases with the number of freeze-thaw cycles.

[0129] Figure 8B The dependence of the predicted output value on the change in plasma storage temperature (e.g., -80°C vs. -20°C) during the 24-day storage of the sample is depicted. As shown, the predicted output value increases when the plasma is stored at a higher temperature.

[0130] Figure 8C The dependence of the predicted output on the time the sample is frozen at -80°C is depicted. As shown, there is no significant dependence of the predicted output on these different freezing times.

[0131] Fig.8D The dependence of the predicted output value on the change in the sample centrifugation time is depicted. As shown, increasing the centrifugation time from the baseline to 9 hours or more affects the output value.

[0132] 9A to 9D The effects of differences in sample processing consistent with the disclosed embodiments on the predictive performance of a proteomics model are described. In this example, the predictive performance is the root mean square error (RMSE) and the sample is a serum sample. The sample processing differs in four aspects: the number of freeze-thaw cycles, the serum coagulation time, the serum decantation time, and the serum freezing time (-80°C). A proteomics model was developed using training data. In this example, the proteomics model was trained using known important aptamers plus five randomly selected aptamers. Consistent with process 500 and steps 610 to 630 of process 600, prediction results are generated using the training data set, the control data set, and the processed data set. 9A to 9D A histogram of the measured RMSE is shown.

[0133] Fig. 9A Describes the dependence of the RMSE value on the number of freeze-thaw cycles of the sample. It can be observed that the RMSE value does not increase significantly between 1 and 10 freeze-thaw cycles.

[0134] Fig. 9B The dependence of the RMSE value on the variation of the sample coagulation time is depicted. It can be observed that the RMSE value increases significantly between the baseline coagulation time and the time exceeding 3 hours.

[0135] Fig. 9C The dependence of the RMSE value on the change in the sample decantation time is depicted. It can be observed that the RMSE value increases significantly between the baseline decantation time and the time exceeding 3 hours.

[0136] Fig.9D The dependence of the RMSE value on the variation of the time that the samples were frozen at -80°C is depicted. It can be observed that the RMSE value increases between the baseline freezing time and the time longer than 3 hours, and increases substantially between the baseline freezing time and the freezing time of 24 hours.

[0137] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, except where not feasible. For example, if it is stated that a component can include A or B, then, unless otherwise specifically stated or not feasible, the component can include A, or B, or A and B. As a second example, if it is stated that a component can include A, B, or C, then, unless otherwise specifically stated or not feasible, the component can include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0138] Exemplary embodiments are described above with reference to flowchart illustrations or block diagrams of methods, apparatus (systems) and computer program products. It will be appreciated that each block in the flowchart illustrations or block diagrams and combinations of blocks in the flowchart illustrations or block diagrams can be implemented by a computer program product or instructions on a computer program product. These computer program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create means for implementing the functions / behaviors specified by one or more blocks in the flowchart or block diagram.

[0139] These computer program instructions may also be stored in a computer-readable medium, which may instruct one or more hardware processors of a computer, other programmable data processing apparatus, or other device to operate in a specific manner, so that the instructions stored in the computer-readable medium form a manufactured product including instructions for implementing the functions / behaviors specified by one or more blocks in the flowchart or block diagram.

[0140] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operating steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / behaviors specified by one or more blocks in the flowchart or block diagram.

[0141] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a non-transitory computer-readable storage medium. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store a program for use by or in association with an instruction execution system, apparatus, or device.

[0142] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, IR, etc., or any suitable combination of the foregoing.

[0143] For example, the computer program code for performing the operations may be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or to an external computer (e.g., via the Internet using an Internet service provider).

[0144] The flowchart and block diagram in the figure show examples of the architecture, functions and operations of possible implementations of the system, method and computer program product according to multiple embodiments. In this regard, each block in the flowchart or block diagram can represent a module, a code segment or a partial code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative embodiments, the functions marked in the block may not appear in the order marked in the figure. For example, the two blocks shown in succession in the figure may actually be executed synchronously at the same time, or the blocks may sometimes be executed in the opposite order, depending on the functions involved. It will also be noted that each block in the block diagram or flowchart illustration, as well as the block combination in the block diagram or flowchart illustration, can be implemented by a hardware-based special-purpose system that performs a specified function or behavior, or a combination of special-purpose hardware and computer instructions.

[0145] It is to be understood that the described embodiments are not mutually exclusive, and elements, components, materials or steps described in conjunction with one exemplary embodiment may be combined with or removed from other embodiments in an appropriate manner to achieve the desired design goals.

[0146] The implementation scheme can be further described using the following terms:

[0147] 1. A system for validating a proteomics model, comprising: at least one processor; and

[0148] At least one non-transitory computer-readable medium, the non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations, the operations comprising: obtaining a control data set including protein level measurements; obtaining a processed data set including protein level measurements; obtaining a training data set including protein level measurements; estimating protein level noise using the control data set and the processed data set; generating a validation data set using the training data set and the estimated protein level noise, the validation data set including corrected protein level measurements; generating a set of prediction results by applying the validation data set to a proteomics model trained using the training data set; determining an effectiveness indication using the set of prediction results; and providing the effectiveness indication.

[0149] 2. A system as described in claim 1, wherein: obtaining the control data set includes measuring a first aliquot to obtain a first protein level measurement value; obtaining the processed data set includes measuring a second aliquot to obtain a second protein level measurement value; and wherein: the first aliquot and the second aliquot are taken from samples having different sample characteristics; the first aliquot and the second aliquot are taken from samples having different processing methods; or different measurement protocols are used for the first aliquot and the second aliquot.

[0150] 3. A system as described in claim 2, wherein: the first aliquot and the second aliquot are taken from samples having different sample characteristics; and the difference in the sample characteristics lies in at least one of the following: citrated plasma or EDTA plasma composition; the fed or fasting state of the patient providing the sample; serum versus plasma composition; or the presence or absence of an interfering agent.

[0151] 4. A system as described in claim 3, wherein: the difference in the sample characteristics lies in the presence or absence of the interfering agent; and the interfering agent includes at least one of the following: non-steroidal anti-inflammatory drugs (NSAIDs); contraceptive drugs; antihypertensive drugs; mental health drugs; cholesterol-lowering drugs; asthma drugs; diabetes drugs; thyroid drugs; or antiviral drugs or antibacterial drugs.

[0152] 5. A system as described in any of clauses 2 to 4, wherein: the first aliquot and the second aliquot are taken from samples processed in different ways; and the difference in the processing methods of the samples lies in at least one of the following: the time between sample collection and sample freezing; the sample freezing temperature; the duration at the freezing temperature; the number of freeze-thaw cycles of the sample; the centrifugation time, and the samples processed in different ways are plasma samples; or the coagulation time, and the samples processed in different ways are serum samples.

[0153] 6. A system as described in any of clauses 2 to 4, wherein: different assay protocols are applied to the first aliquot and the second aliquot; and the different assay protocols differ in at least one of the following: the reagents used; the dilution of the reagents used; the steps in the assay protocol; the equipment used to perform the assay protocol; or the amount of protein measured by the assay protocol.

[0154] 7. The system of any one of clauses 1 to 6, wherein: the training data set is obtained under a first situation; and the processing data set is obtained under a second situation different from the first situation.

[0155] 8. A system as described in any of clauses 1 to 7, wherein: the proteomic model is trained to predict a health outcome; and the proportion of entries in the training dataset corresponding to patients experiencing the health outcome is greater than the proportion of entries in the processing dataset corresponding to patients experiencing the health outcome.

[0156] 9. A system as described in any of clauses 1 to 8, wherein: estimating protein level noise using the control dataset and the processed dataset comprises: determining paired differences between protein level measurements of the control dataset and corresponding protein level measurements of the processed dataset; and characterizing the distribution of the paired differences for each protein in the control dataset.

[0157] 10. A system as described in any of clauses 1 to 9, wherein: generating the validation dataset using the training dataset and the estimated protein level noise comprises: generating an estimated protein level noise sample of a first protein in the training dataset for a protein level measurement of the first protein; and adding the sample to the protein level measurement to generate corresponding corrected protein level measurement.

[0158] 11. The system of clause 10, wherein: the estimated protein level noise comprises an estimated mean and standard deviation of the first protein; and generating the sample comprises sampling a normal distribution having the estimated mean and standard deviation of the first protein.

[0159] 12. A system as described in any of clauses 1 to 11, wherein: determining the effectiveness indication using the prediction result set includes: quantifying the consistency between the prediction result set and a corresponding prediction result set generated using the proteomics model and the training data set; or quantifying the consistency between the prediction result set and a true value associated with the training data set.

[0160] 13. A system as described in any of clauses 1 to 12, wherein: the operation further includes: determining a statistic of pairwise differences between: a control prediction result set generated by applying the control data set to the proteomics model, and a treatment prediction result set generated by applying the treatment data set to the proteomics model; and determining that the statistic of pairwise differences satisfies an invalid condition; and estimating the protein level noise in response to satisfying the invalid condition.

[0161] 14. A system for developing a proteomics model, comprising: at least one processor; and at least one non-transitory computer-readable medium, wherein the non-transitory computer-readable medium contains instructions that, when executed by the at least one processor, cause the system to perform operations, the operations comprising: obtaining a proteomics model, the proteomics model being trained to generate predictions based on training protein level measurements generated by training aliquots obtained in a first situation; generating processed protein level measurements from processed aliquots or samples obtained in a second situation; and determining a dependency of the proteomics model on differences between the second situations; and updating the proteomics model based on the determined dependency.

[0162] 15. The system of clause 14, wherein: updating the proteomic model according to the determined dependency relationship comprises: identifying the protein according to its variability effect; and reducing the variability effect of the protein in the proteomic model.

[0163] 16. The system of any one of clauses 14 to 15, wherein: the proteomic model comprises a linear or logistic regression model, a survival model, a random forest model, a support vector machine model, or a linear discriminant analysis model.

[0164] 17. The system of any one of clauses 14 to 16, wherein the training aliquots are taken from serum samples or plasma samples.

[0165] 18. The system of any one of clauses 14 to 17, wherein the second scenario has different sample characteristics, sample processing methods or assay protocols.

[0166] 19. The system of any one of clauses 14 to 18, wherein: the difference in the second situation includes a difference in at least one of freezing time, number of freeze-thaw cycles, centrifugation time, fasting time, transportation time, frozen storage time, clotting time or decantation time.

[0167] 20. A system as described in any of clauses 14 to 20, wherein: determining the dependency of the proteomics model on the difference between the second situations comprises: determining protein level noise using the processed protein level measurements; generating validation protein level measurements using the training protein level measurements and the protein level noise; and generating predictions using the validation protein level measurements.

[0168] In the above description, embodiments have been described with reference to many specific details, which may vary depending on the implementation. Certain adjustments and modifications may be made to the described embodiments. Other embodiments may be apparent to those skilled in the art by considering the description and practice of the invention disclosed herein. This description and the examples are intended to be considered as exemplary only. The sequence of steps shown in the figure is also intended to be used for illustrative purposes only and is not intended to be limited to any particular sequence of steps. Therefore, it will be appreciated by those skilled in the art that these steps may be performed in different orders when implementing the same method.

Claims

1. A system for validating proteomics models, include: at least one processor; as well as at least one non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a control data set including protein level measurements; obtaining a processed data set including protein level measurements; obtaining a training data set including protein level measurements; estimating protein level noise using the control dataset and the treated dataset; Generate a validation dataset using the training dataset and the estimated protein level noise, The validation data set includes corrected protein level measurements; generating a set of prediction results by applying the validation dataset to a proteomics model trained using the training dataset; Determining an effectiveness indication using the set of prediction results; and An indication of said effectiveness is provided.

2. The system of claim 1, in: Obtaining the control data set includes assaying a first aliquot to obtain a first protein level measurement; Obtaining the processed data set includes assaying a second aliquot to obtain a second protein level measurement; and in: The first aliquot and the second aliquot are taken from samples having different sample characteristics; The first aliquot and the second aliquot are taken from samples that have been processed in different ways; or Different assay protocols are used for the first aliquot and the second aliquot.

3. The system of claim 2, in: The first aliquot and the second aliquot are taken from samples having different sample characteristics; and The sample characteristics differ in at least one of the following: Citrate plasma or EDTA plasma composition; the fed or fasting status of the patient providing the sample; serum and plasma composition; or Presence or absence of interfering agents.

4. The system of claim 3, in: The difference in the characteristics of the samples is the presence or absence of the interfering agent; and The interfering agent includes at least one of the following: Nonsteroidal anti-inflammatory drugs (NSAIDs); Contraceptive drugs; Blood pressure medications; Mental health medications; Cholesterol-lowering drugs; Asthma medications; Diabetes medications; thyroid medication; or Antiviral or antibacterial medicines.

5. The system of claim 2, in: The first aliquot and the second aliquot are taken from samples that have been processed in different ways; and The samples are processed in a manner that differs in at least one of the following: the time between sample collection and sample freezing; Sample freezing temperature; Duration of exposure to freezing temperatures; the number of freeze-thaw cycles of the sample; Centrifugation time, the samples processed differently are plasma samples; or Coagulation time, the samples with different processing methods are serum samples.

6. The system of claim 2, in: applying different assay protocols to the first aliquot and the second aliquot; and The different assay protocols differ in at least one of the following: Reagents used; The dilution of the reagents used; the steps in the assay protocol; Equipment for carrying out said assay protocol; or The amount of protein measured by the assay protocol described.

7. The system of claim 1, in: The training data set is obtained in a first situation; and The processed data set is obtained under a second situation different from the first situation.

8. The system of claim 1, in: training the proteomic model to predict health outcomes; and A proportion of entries in the training dataset that correspond to patients experiencing the health outcome is greater than a proportion of entries in the processing dataset that correspond to patients experiencing the health outcome.

9. The system of claim 1, in: Estimating protein level noise using the control dataset and the processed dataset comprises: determining pairwise differences between the protein level measurements of the control data set and corresponding protein level measurements of the treated data set; and For each protein in the control dataset, the distribution of the pairwise differences is characterized.

10. The system of claim 1, in: Generating the validation dataset using the training dataset and the estimated protein level noise comprises: generating an estimated protein level noise sample of a first protein for the protein level measurement value of the first protein in the training data set; and The sample is added to the protein level measurements to generate corresponding corrected protein level measurements.

11. The system of claim 10, in: The estimated protein level noise comprises an estimated mean and standard deviation of the first protein; and Generating the sample includes sampling a normal distribution having an estimated mean and standard deviation of the first protein.

12. The system of claim 1, in: Using the set of prediction results to determine the effectiveness indication includes: quantifying the consistency between the set of predictions and a corresponding set of predictions generated using the proteomic model and the training dataset; or The consistency between the set of predictions and true values ​​associated with the training data set is quantified.

13. The system of claim 1, in: The operations also include: Determine a statistic for the pairwise differences between: a set of control prediction results generated by applying the control data set to the proteomics model; and a processed set of predictions generated by applying the processed data set to the proteomics model; and determining that the pairwise difference statistic satisfies a null condition; and The protein level noise is estimated in response to satisfying the null condition.

14. A system for developing a proteomics model, include: at least one processor; as well as at least one non-transitory computer-readable medium containing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: obtaining a proteomic model trained to generate predictions based on training protein level measurements generated from training aliquots obtained in a first instance; generating a treatment protein level measurement from the treatment aliquot or sample obtained at the second instance; and determining a dependency of the proteomic model on the difference between the second scenarios; and The proteomic model is updated according to the determined dependencies.

15. The system of claim 14, in: Updating the proteomics model according to the determined dependency relationship comprises: identifying the protein based on its variability effect; and The effect of variability of the protein in the proteomic model is reduced.

16. The system of claim 14, in: The proteomics model includes a linear or logistic regression model, a survival model, a random forest model, a support vector machine model or a linear discriminant analysis model.

17. The system of claim 14, in: The training aliquots were taken from serum samples or plasma samples.

18. The system of claim 14, in: The second scenario has different sample characteristics, sample processing methods or measurement protocols.

19. The system of claim 14, in: The difference in the second situation includes at least one difference in freezing time, number of freeze-thaw cycles, centrifugation time, fasting time, transportation time, frozen storage time, coagulation time or decantation time.

20. The system of claim 14, in: Determining the dependency of the proteomic model on the difference between the second situations comprises: determining protein level noise using said processed protein level measurements; generating validation protein level measurements using the training protein level measurements and the protein level noise; and Predictions are generated using the validation protein level measurements.