Systems and methods for generating surgical risk scores and uses thereof - Patents.com
Patent Information
- Application Number
- JP2023556770
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-03-18
- Filing Date
- 2022-03-18
- Publication Date
- 2025-11-06
AI Technical Summary
Existing risk prediction tools for surgical site complications (SSCs) are insufficient for accurately estimating individual patient risks, relying solely on clinical parameters and failing to incorporate biological markers that govern the developmental principles of SSCs.
A method using machine learning models that integrate multi-omics biological data, including genomic, transcriptomic, proteomic, and cytomic features, with clinical data to predict surgical outcomes by analyzing immune cell responses and plasma proteins before and after surgery, employing techniques like mass cytometry and plasma proteomics to generate a surgical risk score.
The method provides a robust and accurate prediction of SSCs, enabling personalized treatment plans to reduce complications, hospital stays, and readmission rates by identifying individual patient risks through integrated multi-omics and clinical data analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with Government support under Contracts GM137936 and GM138353 awarded by the National Institutes of Health. The Government has certain rights in this invention.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 162,912, entitled "Systems and Methods to Generate a Surgical Risk Score and Uses Thereof," by Gaudilliere et al., filed March 18, 2021, the disclosure of which is hereby incorporated by reference in its entirety.
[0003] FIELD OF THEINVENTION The present invention relates to predicting surgical outcomes, and more particularly to predicting surgical outcomes such as post-operative infections and surgical site complications from clinical and multi-omics data using machine learning models. [Background technology]
[0004] background Over 300 million surgeries are performed worldwide each year, and this number is expected to grow. Surgical complications, including infection, long-term pain, functional impairment, and end-organ damage, occur in 10-60% of surgeries, causing personal suffering, prolonged hospital stays, readmissions, and significant socioeconomic burden. After major abdominal surgery, surgical site complications (SSCs), including superficial or deep wound infections, organ space infections, anastomotic leaks, fascial dehiscence, and incisional hernias, are some of the most severe, costly, and common surgical complications, occurring in up to 25% of patients. (See, e.g., Healy MA et al. JAMA Surg 2016; 151(9):823-30, the disclosure of which is incorporated herein by reference in its entirety.) Accurate prediction of SSC risk for individual patients is crucial to guide high-quality surgical decision-making, including optimizing preoperative intervention and timing of surgery. Existing risk prediction tools are based on clinical parameters and are insufficient to estimate the risk of SSC for individual patients. (See, for example, Eamer G, et al. Am J Surg 2018; 216(3):585-594, and Cohen ME et al. J Am Coll Surg 2017; 224(5):787-795 e1, the disclosures of which are incorporated herein by reference in their entirety.) Therefore, there is a need in the art for a robust tool that predicts SSC with higher accuracy. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Healy MA et al. JAMA Surg 2016; 151(9):823-30 [Non-Patent Document 2] Eamer G, et al. Am J Surg 2018; 216(3):585-594 [Non-Patent Document 3] Cohen ME et al. J Am Coll Surg 2017; 224(5):787-795 e1 Summary of the Invention [Means for solving the problem]
[0006] Summary of the invention This Summary provides some examples and is not intended to limit the scope of the invention in any way. For example, any feature included in the examples of this Summary is not required in a claim unless the claim explicitly recites that feature. Various features and steps described elsewhere in this disclosure may be included in the examples summarized here, and features and steps described here and elsewhere may be combined in various ways.
[0007] In some aspects, the technology described herein relates to a method for determining a risk of a surgical complication for an individual following surgery, the method comprising obtaining or having obtained values of a plurality of features, the plurality of features including omics biological features and clinical features; calculating a surgical risk score for the individual based on the plurality of features using a model obtained via machine learning techniques; and providing an assessment of the patient's risk of developing a surgical complication based on the calculated surgical risk score.
[0008] In some aspects, the technology described herein relates to a method, wherein obtaining or having obtained values of a plurality of features includes obtaining or having obtained a sample for analysis from an individual undergoing surgery, and measuring or having measured values of a plurality of omics biological and clinical features.
[0009] In some aspects, the technology described herein relates to a method, wherein the plurality of characteristics further comprises demographic characteristics.
[0010] In some aspects, the technology described herein relates to a method, wherein the omics biological feature comprises at least one feature from the group consisting of a genomics feature, a transcriptomics feature, a proteomics feature, a cytomics feature, and a metabolomics feature.
[0011] In some aspects, the technologies described herein relate to a method, in which a machine learning model is trained using a bootstrap procedure on multiple individual data layers, each data layer representing one type of data from a plurality of features and at least one artificial feature.
[0012] In some aspects, the technology described herein relates to methods, wherein each type is selected from among genomics, transcriptomics, proteomics, cytomics, metabolomics, clinical, and demographics.
[0013] In some aspects, the techniques described herein relate to a method, wherein each data layer includes data for a population of individuals, and each feature includes feature values for all individuals in the population of individuals, and for each data layer, each artificial feature is obtained from a non-artificial feature of the plurality of features via a mathematical operation performed on the feature values of the non-artificial feature.
[0014] In some aspects, the technology described herein relates to a method, where the mathematical operation is selected from among replacement, sampling with replacement, sampling without replacement, combination, knock-off, and inference.
[0015] In some aspects, the technology described herein relates to a method, in which the model assigns weights (β) to a set of selected biological and clinical or demographic features. i ), and for each data layer during machine learning, at each bootstrap iteration, initial weights (w j ) is calculated, and the calculated initial weights (w jAt least one selected feature is determined for each data layer based on a statistical criterion that depends on
[0016] In some aspects, the techniques described herein relate to a method, wherein the initial statistical learning technique is selected from regression techniques and classification techniques.
[0017] In some aspects, the techniques described herein relate to a method, where the initial statistical learning technique is selected from a sparse technique and a non-sparse technique.
[0018] In some aspects, the techniques described herein relate to a method, wherein the sparse technique is selected from the Lasso technique and the Elastic Net technique.
[0019] In some aspects, the technology described herein relates to a method, in which the statistical criterion is a function of the calculated initial weights (w j ) depending on the effective weight of
[0020] In some aspects, the techniques described herein relate to a method, where the initial statistical learning technique is a sparse regression technique, the effective weights are non-zero weights.
[0021] In some aspects, the techniques described herein relate to a method, where the initial statistical learning technique is a non-sparse regression technique, the effective weights are those weights that are above a predefined weight threshold.
[0022] In some aspects, the technology described herein relates to a method for determining initial weights (w j ) is further calculated for multiple values of hyperparameters, where the hyperparameters are parameters whose values are used to control the learning process.
[0023] In some aspects, the techniques described herein relate to methods, where the hyper-parameters are regularization coefficients used according to their respective mathematical norms in the context of sparse initial techniques.
[0024] In some aspects, the technology described herein relates to a method, wherein the mathematical norm is a p-norm, where p is an integer.
[0025] In some aspects, the techniques described herein relate to a method, in which when the initial statistical learning technique is a Lasso technique, the hyperparameters are the initial weights (w j ), where the L1 norm refers to the sum of all the absolute values of the initial weights.
[0026] In some aspects, the techniques described herein relate to a method, in which when the initial statistical learning technique is an Elastic Net technique, the hyperparameters are the initial weights (w j ) and the initial weights (w j ), where the L1 norm refers to the sum of all the absolute values of the initial weights, and the L2 norm refers to the square root of the sum of all the squared values of the initial weights.
[0027] In some aspects, the technology described herein relates to a method, wherein the statistical criterion is based on a frequency of occurrence of valid weights.
[0028] In some aspects, the techniques described herein relate to a method where, for each feature, a unit occurrence frequency is calculated for each hyper-parameter value, where the unit occurrence frequency is equal to the number of valid weights for that feature for successive bootstrap iterations divided by the number of bootstrap iterations.
[0029] In some aspects, the technology described herein relates to a method, wherein the frequency of occurrence is equal to the highest unit occurrence frequency among unit occurrence frequencies calculated for a plurality of hyper-parameter values.
[0030] In some aspects, the techniques described herein relate to a method, wherein the statistical criterion is that each feature is selected if its frequency of occurrence is greater than a frequency threshold, the frequency threshold being calculated according to the frequency of occurrence obtained for the artificial features.
[0031] In some aspects, the technology described herein relates to a method, wherein the number of bootstrap iterations is between 50 and 100,000.
[0032] In some aspects, the techniques described herein relate to a method, wherein a plurality of hyper-parameter values are between 0.5 and 100 for the Lasso technique or the Elastic Net technique.
[0033] In some aspects, the technology described herein relates to a method, during machine learning, to adjust the weights (β i ) is further calculated using a final statistical learning technique on the data associated with the selected set of features.
[0034] In some aspects, the techniques described herein relate to a method, wherein the final statistical learning technique is selected from a regression technique and a classification technique.
[0035] In some aspects, the techniques described herein relate to a method, where the final statistical learning technique is selected from a sparse technique and a non-sparse technique.
[0036] In some aspects, the techniques described herein relate to a method, wherein the sparse technique is selected from the Lasso technique and the Elastic Net technique.
[0037] In some aspects, the technology described herein relates to a method, wherein during a use phase following machine learning, a surgical risk score is calculated according to an individual's measurements for a set of selected features.
[0038] In some aspects, the techniques described herein relate to a method, where the final statistical learning technique is a classification technique, the surgical risk score is calculated by multiplying the respective weights (β i ) is the probability calculated according to a weighted sum of measurements multiplied by
[0039] In some aspects, the technology described herein relates to a method, the surgical risk score being calculated according to the formula:
number
[0040] In some aspects, the technology described herein relates to a method, where Odd is the exponent of the weighted sum.
[0041] In some aspects, the techniques described herein relate to a method, where the final statistical learning technique is a regression technique, the surgical risk score is calculated by multiplying the respective weights (β i ) is a term that depends on a weighted sum of measurements multiplied by
[0042] In some aspects, the technology described herein relates to a method, wherein the surgical risk score is equal to the index of the weighted sum.
[0043] In some aspects, the techniques described herein relate to a method, wherein during the machine learning, the method further comprises generating additional values of a plurality of non-artificial features using data augmentation techniques based on the obtained values before obtaining the artificial features, and the artificial features are then obtained according to both the obtained values and the generated additional values.
[0044] In some aspects, the techniques described herein relate to a method, wherein the data augmentation technique is selected from among a non-compositing technique and a composite technique.
[0045] In some aspects, the techniques described herein relate to a method, wherein the data augmentation technique is selected from among the SMOTE technique, the ADASYN technique, and the SVMSMOTE technique.
[0046] In some aspects, the techniques described herein relate to a method where, for a given non-artificial feature, the fewer values obtained, the more additional values are generated.
[0047] In some aspects, the technology described herein relates to a method, wherein the omics biological feature is selected from one or more of a cytomic feature, a proteomic feature, a transcriptomic feature, and a metabolomic feature.
[0048] In some aspects, the technology described herein relates to methods, wherein the cytomic signature comprises surface and intracellular proteins at the single cell level in immune cell subsets, and the proteomic signature comprises circulating extracellular proteins.
[0049] In some aspects, the technology described herein relates to a method, wherein the sample comprises at least one sample obtained pre-operatively.
[0050] In some aspects, the technology described herein relates to methods whereby samples are obtained during the period from any time prior to surgery to before the surgical incision is made on the day of surgery.
[0051] In some aspects, the technology described herein relates to a method, wherein the post-operative sample comprises at least one sample obtained post-operatively.
[0052] In some aspects, the technology described herein relates to methods, wherein the post-operative sample is obtained approximately 24 hours after surgery.
[0053] In some aspects, the technology described herein relates to a method, wherein the sample is a blood sample, a peripheral blood mononuclear cell (PBMC) fraction of a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, or dissociated cells from a tissue sample.
[0054] In some aspects, the technology described herein relates to a method, wherein a sample is contacted ex vivo with an effective amount of an activating agent for a period of time sufficient to activate immune cells in the sample.
[0055] In some aspects, the technology described herein relates to methods, where the step of measuring or having measured a value comprises measuring a surface or intracellular protein at the single cell level in an immune cell subset by contacting the sample with an isotopically or fluorescently labeled affinity reagent specific for the surface or intracellular protein.
[0056] In some aspects, the technology described herein relates to methods, whereby surface or intracellular proteins at the single cell level in immune cell subsets are determined by flow or mass cytometry.
[0057] In some aspects, the technology described herein relates to a method, wherein the step of measuring or having measured a value comprises analyzing circulating proteins by contacting the sample with a plurality of isotopically or fluorescently labeled affinity reagents specific for the extracellular proteins.
[0058] In some aspects, the technology described herein relates to a method, wherein the affinity reagent is an antibody or an aptamer.
[0059] In some aspects, the technology described herein relates to a method, wherein the demographic or clinical characteristics include data selected from the group consisting of age, sex, body mass index (BMI), functional status, emergency case, American Society of Anesthesiologists (ASA) class, steroid use for chronic conditions, ascites, multiple cancers, diabetes, hypertension, congestive heart failure, dyspnea, smoking history, history of severe COPD, dialysis, acute renal failure.
[0060] In some aspects, the technology described herein relates to a method, where clinical features are obtained from patient medical records using machine learning algorithms.
[0061] In some aspects, the technology described herein relates to a method, wherein the surgical complication is a surgical site complication (SSC).
[0062] In some aspects, the technology described herein relates to a method, wherein the step of measuring or having measured the value comprises contacting the sample ex vivo with an effective amount of an activating agent for a period of time sufficient to activate immune cells in the sample, the activating agent being one or a combination of a TLR4 agonist (such as LPS), interleukin (IL)-2, IL-4, IL-6, IL-1β, TNFα, IFNα, PMA / ionomycin.
[0063] In some aspects, the technology described herein relates to methods, wherein the period of time is from about 5 minutes to about 240 minutes.
[0064] In some aspects, the technology described herein relates to methods, where the step of measuring or having measured a value comprises measuring a surface or intracellular protein at the single cell level in an immune cell subset by contacting the sample with an isotopically or fluorescently labeled affinity reagent specific for the surface or intracellular protein.
[0065] In some aspects, the technology described herein relates to a method, wherein immune cells are identified using a single cell surface or intracellular protein marker selected from the group consisting of CD235ab, CD61, CD45, CD66, CD7, CD19, CD45RA, CD11b, CD4, CD8, CD11c, CD123, TCRγδ, CD24, CD161, CD33, CD16, CD25, CD3, CD27, CD15, CCR2, OLMF4, HLA-DR, CD14, CD56, CRTH2, CCR2, and CXCR4.
[0066] In some aspects, the technology described herein relates to a method, wherein the intracellular protein of a single cell is selected from the group consisting of phospho(p)pMAPKAPK2 (pMK2), pP38, pERK1 / 2, p-rpS6, pNFκB, IκB, p-CREB, pSTAT1, pSTAT5, pSTAT3, pSTAT6, cPARP, FoxP3, and Tbet.
[0067] In some aspects, the technology described herein relates to a method, wherein intracellular protein levels are measured for neutrophils, granulocytes, basophils, CXCR4+ neutrophils, OLMF4+ neutrophils, CD14+CD16- classical monocytes (cMCs), CD14-CD16+ non-classical monocytes (ncMCs), CD14+CD16+ intermediate monocytes (iMCs), HLADR+CD11c+ myeloid dendritic cells (mDCs), HLADR+CD123+ plasmacytoid dendritic cells (pDCs), CD14+HLADR-CD11b+ monocytic myeloid-derived suppressor cells (M-MDSCs), CD3+CD56+ NK-T cells, CD7+CD19-CD3- NK cells, CD7+CD56loCD16hi NK cells, CD7+CD56hiCD16lo NK cells, CD19+ It is measured in immune cell subsets selected from B cells, CD19+CD38+ plasma cells, CD19+CD38- non-plasma B cells, CD4+CD45RA+ naive T cells, CD4+CD45RA- memory T cells, CD4+CD161+ Th17 cells, CD4+Tbet+ Th1 cells, CD4+CRTH2+ Th2 cells, CD3+TCRγδ+ γδ T cells, Th17 CD4+ T cells, CD3+FoxP3+CD25+ regulatory T cells (Tregs), CD8+ CD45RA+ naive T cells, and CD8+ CD45RA- memory T cells.
[0068] In some aspects, the technology described herein relates to a method, wherein a patient's risk of developing a surgical site complication correlates with increased pMAPKAPK2 (pMK2) in neutrophils, increased prpS6 in mDCs, or decreased IκB in neutrophils or decreased pNFκB in CD7+CD56hiCD16lo NK cells in response to ex vivo activation with LPS in samples collected prior to surgery.
[0069] In some aspects, the technology described herein relates to methods, wherein a patient's risk of developing a surgical site complication correlates with increased pSTAT3 in neutrophils, mDCs, or Tregs, increased prpS6 in CD56hiCD16lo NK cells or mDCs, increased pSTAT5 in mDCs or pDCs, or decreased IκB in CD4+Tbet+ Th1 cells, decreased pSTAT1 in pDCs in response to ex vivo activation with IL-2, IL-4, and / or IL-6 in samples collected prior to surgery.
[0070] In some aspects, the technology described herein relates to a method, wherein a patient's risk of developing a surgical site complication correlates with increased prpS6 in neutrophils or mDCs, increased pERK in M-MDSCs or ncMCs, increased pCREB in γδ T cells, or decreased IκB, pP38 or pERK in neutrophils, or decreased pCREB or pMAPKAPK2 in CD4+Tbet+ Th1 cells, or decreased pERK in CD4+CRTH2+ Th2 cells in response to ex vivo activation with TNFα in samples collected prior to surgery.
[0071] In some aspects, the technology described herein relates to a method, wherein a patient's risk of developing a surgical site complication correlates with increased pSTAT3 in neutrophils, M-MDSC, cMC, or ncMC, increased pSTAT5 in Treg or CD45RA- memory CD4+ T cells, increased pMAPKAPK2 in mDC, pCREB or IκB in CD4+Tbet+ Th1 cells, increased pSTAT6 in NKT cells, or decreased pERK in CD4+Tbet+ Th1 cells in unstimulated samples collected pre- and / or post-surgery.
[0072] In some aspects, the technology described herein relates to a method, wherein a patient's risk of developing a surgical site complication correlates with increased M-MDSC, G-MDSC, ncMC, Th17 cells, or decreased CD4+CRTH2+ Th2 cell frequencies collected pre- and / or post-surgery.
[0073] In some aspects, the technology described herein relates to a method, wherein a patient's risk of developing a surgical site complication is correlated with increased IL-1β, ALK, WWOX, HSPH1, IRF6, CTNNA3, CCL3, sTREM1, ITM2A, TGFα, LIF, ADA, or decreased ITGB3, EIF5A, KRT19, NTproBNP collected pre- and / or post-surgery.
[0074] In some aspects, the technology described herein relates to a system comprising a processor and a memory containing instructions that, when executed by the processor, direct the processor to perform any of the methods described above.
[0075] In some aspects, the technology described herein relates to a non-transitory machine-readable medium containing instructions that, when executed by a computer processor, direct the processor to perform any of the methods described above.
[0076] In some aspects, the technology described herein relates to a method, further comprising treating the individual according to an assessment of the individual's risk of developing a surgical site complication before surgery is performed.
[0077] In some aspects, the technology described herein relates to a method, further comprising treating the individual after the surgery has been performed according to an assessment of the individual's risk of developing a surgical site complication.
[0078] Other features and advantages of the present invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the invention.
[0079] The description and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the invention and should not be construed as a complete recitation of the scope of the invention. [Brief description of the drawings]
[0080] [Figure 1] 1 illustrates an exemplary method for predicting a patient's clinical outcome after surgery using a machine learning algorithm that integrates multi-omics biological (e.g., single cell immune response and plasma proteomic data) and clinical data, according to various embodiments. Various embodiments provide a method for guiding a surgeon's or medical personnel's clinical decision making by using a multi-omics bootstrap (MOB) machine learning algorithm to generate a predictive model of the probability that a patient will develop a surgical site complication (SSC).
[0081] [Diagram 2] FIG. 2 illustrates an exemplary methodology of a MOB machine learning model that integrates biological and clinical data to predict surgical outcomes, according to various embodiments.
[0082] [Figure 3A] 3A-3B show exemplary pseudocode for the MOB algorithm, according to various embodiments. [Figure 3B] 3A-3B show exemplary pseudocode for the MOB algorithm, according to various embodiments.
[0083] [Figure 4] FIG. 4 illustrates an exemplary workflow for identifying a predictive model for surgical site complications in patients undergoing abdominal surgery, according to various embodiments.
[0084] [Diagram 5] 5A-5C show an exemplary MOB predictive model for SSC derived from analysis of patient samples collected prior to abdominal surgery, according to various embodiments.
[0085] [Figure 6] FIG. 6 shows an exemplary MOB predictive model for SSC derived from integrated analysis of multi-omics biological data collected from patients, according to various embodiments.
[0086] [Figure 7] FIG. 7 illustrates an exemplary MOB predictive model of SSC derived from an analyzed patient sample collected 24 hours after abdominal surgery, according to various embodiments.
[0087] [Figure 8A] 8A-8D show exemplary single-cell immune response and proteomic signatures contributing to the DOS MOB predictive model of SSCs, according to various embodiments. [Figure 8B] 8A-8D show exemplary single-cell immune response and proteomic signatures contributing to the DOS MOB predictive model of SSCs, according to various embodiments. [Figure 8C] 8A-8D show exemplary single-cell immune response and proteomic signatures contributing to the DOS MOB predictive model of SSCs, according to various embodiments. [Figure 8D] 8A-8D show exemplary single-cell immune response and proteomic signatures contributing to the DOS MOB predictive model of SSCs, according to various embodiments.
[0088] [Figure 9-1] 9A-9N show exemplary features contributing to the POD1 MOB predictive model of SSCs according to various embodiments of the present invention. FIGs 9A-9G show single cell immune response features and FIGs 9H-9N show plasma proteomic features. [Figure 9-2] 9A-9N show exemplary features contributing to the POD1 MOB predictive model of SSCs according to various embodiments of the present invention. FIGs 9A-9G show single cell immune response features and FIGs 9H-9N show plasma proteomic features. [Figure 9-3] 9A-9N show exemplary features contributing to the POD1 MOB predictive model of SSCs according to various embodiments of the present invention. FIGs 9A-9G show single cell immune response features and FIGs 9H-9N show plasma proteomic features. [Figure 9-4]9A-9N show exemplary features contributing to the POD1 MOB predictive model of SSCs according to various embodiments of the present invention. FIGs 9A-9G show single cell immune response features and FIGs 9H-9N show plasma proteomic features.
[0089] [Figure 10] FIG. 10 illustrates an exemplary gating strategy for identifying immune cell subsets, according to various embodiments of the present invention.
[0090] [Figure 11] 11A-11B show an exemplary set of single cell immune responses and plasma proteins that are differentially expressed pre- and post-surgery, according to various embodiments of the present invention.
[0091] [Figure 12] FIG. 12 illustrates an exemplary patient enrollment according to the CONSORT criteria, according to various embodiments of the present invention.
[0092] [Figure 13] FIG. 13 shows a block diagram of components of a processing system within a computing device that may be used to generate a surgical risk score, according to one embodiment of the present invention.
[0093] [Figure 14] FIG. 14 illustrates a network diagram of a distributed system for generating a surgical risk score, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0094] Detailed Description As mentioned above, existing risk prediction tools are based on clinical parameters and are insufficient to estimate the risk of SSC in individual patients (see, for example, Eamer G, et al. cited above). Therefore, incorporating mechanisms reflecting biological parameters that govern the development of SSC is a promising approach to improve the accuracy of risk prediction.
[0095] Surgery involves significant tissue trauma, which triggers a programmed inflammatory response that engages the innate and adaptive branches of the immune system. Within hours of surgical incision, a highly diverse network of innate immune cells (including monocytes, neutrophils, and their subsets) is activated in response to circulating DAMPs (damage-associated molecular patterns) and inflammatory cytokines (e.g., HMGB1, TNFα, and IL-1β). Following the early innate immune response to surgery, a compensatory anti-inflammatory adaptive immune response has traditionally been described. However, recent transcriptomics and mass cytometry analyses suggest that the adaptive immune response is recruited in concert with the innate immune response and occurs concomitantly with the activation of a subset of specialized immunosuppressive immune cells, such as myeloid-derived suppressor cells (MDSCs). In the context of uncomplicated post-operative recovery, innate and adaptive responses act synergistically to direct the pro- and anti-inflammatory (pro-resolution) processes required for pathogen-defense tissue remodeling and resolution of pain and inflammation following injury. (See, e.g., Stoecklein VM et al. J Leukoc Biol 2012; Gaudilliere B et al. Sci Transl Med 2014; 6(255):255ral31, the disclosures of which are hereby incorporated by reference in their entireties.)
[0096] Complications, including infection, wound dehiscence, and ultimately end-organ damage, arise when the balance between proinflammatory and immunosuppressive responses is skewed. Thus, detailed characterization of immune mechanisms that differ between patients with and without surgical complications is a very promising approach to identify preoperative and postoperative biological events that contribute to and occur before surgical complications. Previous attempts to detect biological markers predictive of SSC risk have focused on secreted humoral factors, surface marker expression on select immune cells, or transcriptional analysis of pooled circulating leukocytes. However, the associations detected were insufficient to accurately predict the risk of SSC in individual patients.
[0097] One of the major obstacles has been the lack of high-volume, functional assays capable of characterizing the complex multicellular inflammatory response to surgery at single-cell resolution. In addition, there is a lack of analytical tools capable of integrating single-cell immune data with other omics and clinical data to predict the development of SSC. Thus, there is a need for improved measures for the diagnosis, prognosis, treatment, management, and therapeutic development of SSC after surgery.
[0098] High-throughput omics assays, including metabolomic, proteomic, and cytometric immunoassay data, can potentially capture complex mechanisms of disease and biological processes by providing thousands of systematically obtained measurements for each biological sample.
[0099] Analysis of mass cytometry immunoassays as well as other omics assays typically has two related goals that are analyzed in a dichotomous manner. The first goal is to identify biomarkers that predict the outcome of interest and are the best set of predictors of the outcome being studied, and the second goal is to identify potential pathways related to the disease that provide a deeper understanding of the underlying biology. The first goal is addressed by deploying machine learning methods to fit predictive models that typically select only a small number of the most informative biomarkers out of thousands of measurements. The second goal is typically addressed by performing univariate analysis of each measurement and determining the significance of that measurement with respect to the outcome by assessing its p-value, which is then adjusted for multiple hypothesis testing.
[0100] In the context of machine learning, omics data characterized by a large number of features p and a much smaller number of samples n fall into the scenario where p>>n. Exemplary machine learning methodologies for this scenario consist of the use of regularized regression or classification methods, specifically sparse linear models such as Lasso (see, e.g., Tibshirani, Robert. "Regression shrinkage and selection via the lasso." Journal of the Royal Statistical Society: Series B (Methodological) 58.1 (1996): 267-288, the disclosure of which is hereby incorporated by reference in its entirety) and Elastic Net. (see, e.g., Zou, Hui, and Trevor Hastie. "Regularization and variable selection via the elastic net." Journal of the royal statistical society: series B (statistical methodology) 67.2 (2005): 301-320, the disclosure of which is hereby incorporated by reference in its entirety). For example, consider the following linear model given by: Y=Xβ+∈ Where:
number
number
number
number
[0101] Instability is an inherent problem in feature selection for machine learning models. Since the model learning phase is performed on a finite data sample, any perturbation in the data may result in a slightly different set of selected variables. In a situation where performance is evaluated by cross-validation, this implies that Lasso will result in a slightly different set of selected biomarkers, making any biological interpretation of the results impossible. Consistent feature selection in Lasso is difficult because it is achieved only under constrained conditions. Most sparse techniques, such as Lasso, cannot provide a quantification of how far the selected model is from the correct model, nor can they quantify the variability of the selected features.
[0102] Another major limitation of existing methods is the difficulty of integrating various sources of biological information. Most machine learning algorithms use input data agnostic in the model learning process. The main challenge is to integrate multiple data sources with their differences in modality, size, and signal-to-noise ratio in the learning process. In the learning process, current methods are usually limited by bias in the assessment of the contribution of each individual data source when juxtaposed as a unique data set. Ultimately, it is important to use the identified informative features from each different layer together to optimize the predictive power of such algorithms. Most methods also lack the ability to assess the individual interactions between features when putting together the various results from the individual data sources, which is important for modeling the biological mechanisms involved.
[0103] Turning now to the drawings, systems and methods for generating a surgical risk score and their uses are provided. In many embodiments, configurations and methods are provided for prediction, classification, diagnosis, and / or theranosis of clinical outcomes after surgery in subjects based on integration of multi-omics biological data and clinical data using machine learning models (e.g., FIG. 1). Many embodiments provide methods for generating a predictive model of the probability that a patient will develop a surgical site complication (SSC). In many embodiments, the predictive model is obtained by quantifying certain biological and clinical features before or after surgery. Various embodiments use at least one omics (including, but not limited to, genomics, cytomics, proteomics, transcriptomics, metabolomics) feature in combination with clinical data to generate a predictive model. Various embodiments utilize machine learning models to integrate various clinical features and / or cytomics, proteomics, transcriptomics, or metabolomics features to generate a predictive model. In some embodiments, the clinical outcome is the development of an SSC (including surgical site infection, wound dehiscence, abscess, or fistula formation). A predictive model according to many embodiments can indicate a patient's risk of developing SSC.
[0104] Once a classification or prognosis has been made, it can be provided to the patient or caregiver. Classification can provide prognostic information to guide clinical decision-making of healthcare professionals and surgeons, such as delaying or adjusting timing of surgery, adjusting surgical technique, adjusting type and timing of antibiotic and immunomodulatory therapy, personalizing or tailoring a pre-rehabilitation health optimization program, planning for more time in the hospital before or after surgery, or planning for time spent in a managed care facility. Appropriate nursing care can reduce rates of SSC, length of hospital stay, and / or readmission rates for post-operative patients.
[0105] As shown in FIG. 1, various embodiments are directed to a method of predicting clinical outcomes of an individual (e.g., a patient) undergoing surgery. Many embodiments collect a patient sample at 102. Such a sample may be collected any time before or after surgery. In some embodiments, samples are collected up to one week (7 days) before or after surgery. In certain embodiments, samples are collected for 1, 2, 3, 4, 5, 6, or 7 days before surgery, and some embodiments collect samples for 1, 2, 3, 4, 5, 6, or 7 days after surgery. Additional embodiments collect samples on the day of surgery, including before and / or after surgery, including immediately before and / or after surgery. Certain embodiments collect multiple samples before, after, or before and after surgery, anesthesia, and / or any other treatment steps included in a particular surgical protocol or procedure.
[0106] At 104, many embodiments obtain omics data (e.g., proteomics, cytomics, and / or any other omics data) from the sample. Certain embodiments combine multiple omics data, such as plasma proteomics (e.g., analysis of plasma protein expression levels) and single-cell cytomics (e.g., single-cell analysis of circulating immune cell frequency and signaling activity) as multi-omics data. Certain embodiments obtain clinical data of the individual. Clinical data according to various embodiments includes one or more of medical history, age, weight, body mass index (BMI), sex / gender, current medications / supplements, functional status, emergency cases, steroid use for chronic conditions, ascites, multiple cancers, diabetes, hypertension, congestive heart failure, dyspnea, smoking history, history of severe chronic obstructive pulmonary disease (COPD), dialysis, acute renal failure, and / or any other applicable clinical data. Clinical data may also be derived from clinical risk scores, such as the American Society of Anesthesiologists (ASA) or American College of Surgeons (ACS) risk scores.
[0107] Additional embodiments generate a predictive model of surgical complications such as SSC at 106. Many embodiments utilize machine learning models as described herein. Various embodiments operate in a pipeline manner, such that acquired or collected data is immediately sent to the machine learning model to generate an integrated surgical risk score. Some embodiments house the machine learning model locally, such that the integrated risk score is generated without network communication, while some embodiments run the machine learning model on a server or other remote device, such that the clinical data and multi-omics data are transmitted over a network and the integrated surgical risk score is returned to a medical professional / practitioner at the local organization, clinic, hospital, and / or other medical facility.
[0108] At 108, further embodiments adjust the individual's treatment based on the combined surgical risk score. In various embodiments, this adjustment may include delaying surgery (e.g., until an improved combined surgical risk score is obtained), prescribing additional antibiotics to prevent infection, and / or adjusting the surgical procedure to counteract the increased risk identified by the combined surgical risk score. Using this approach, the treatment regimen is personalized and adapted according to the patient's predicted probability of developing SSC, thereby providing therapy that is individually appropriate.
[0109] It should be noted that the embodiment shown in FIG. 1 is illustrative of various steps, features, and details that may be implemented in various embodiments, but is not exhaustive or limiting to all embodiments. In addition, various embodiments may include additional steps not described herein and / or fewer steps than shown and / or described (e.g., certain steps may be omitted). Various embodiments may also repeat certain steps, such as repeating the generation of the predictive model 106 to identify whether the individual is more or less likely to develop a risk score or SSC, where additional data, predictions, or procedures may be updated for the individual. Further embodiments may obtain samples or clinical data from third parties, collaborating individuals, subordinates, or other individuals, and / or obtain stored or previously collected or obtained samples. Certain embodiments may even perform certain operations or functions in a different order than shown or described, and / or perform some operations or functions simultaneously or relatively simultaneously (e.g., one operation may begin or commence before another operation is finished or completed).
[0110] definition Many of the words used herein have the meaning that those skilled in the art would ascribe to them. Words specifically defined herein have the meaning given in the context of the entire teaching and as commonly understood by those skilled in the art. In the event of a discrepancy between the art-understood definition of a word or phrase and the definition of that word or phrase specifically taught in this specification, the present specification shall control.
[0111] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0112] It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0113] The terms "subject", "individual" and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals providing samples for analysis include canines, felines, equines, bovines, ovines, etc., and primates, particularly humans. Animal models, particularly small mammals, such as mice, rabbits, etc., may be used for experimental investigations. The methods of the present invention may be applied for veterinary purposes. The terms "biomarker", "marker", or "feature" for the purposes of the present invention refer to, without limitation, proteins and their associated metabolites, mutants, variants, polymorphisms, phosphorylations, modifications, fragments, subunits, degradation products, elements, and other analyte- or sample-derived measures. Markers may include intracellular or extracellular protein expression levels. Markers may also include any one or more combinations of the aforementioned measurements, including trends and differences over time. When used broadly, markers may also refer to immune cell subsets.
[0114] As used herein, the term "omics" or "-omics" data refers to data generated to quantify pools of biological molecules or processes leading to the structure, function, and dynamics of one or more living organisms. Examples of omics data include (but are not limited to) genomics, transcriptomics, proteomics, metabolomics, and cytomics data, among others.
[0115] As used herein, the term "cytomics" data refers to omics data generated using a technology or analysis platform that allows for the quantification of biological molecules or processes at the single-cell level. Examples of cytomics data include (but are not limited to) data generated using flow cytometry, mass cytometry, single-cell RNA sequencing, and cell imaging techniques, among others.
[0116] The term "inflammatory" response is the development of a humoral (antibody-mediated) and / or cellular response that may be mediated by innate immune cells (such as neutrophils or monocytes) or by antigen-specific T cells or their secretory products. An "immunogen" is capable of eliciting an immune response against itself when administered to a mammal or due to an autoimmune disease.
[0117] "Analyzing" includes determining a set of values associated with a sample by measuring a marker in the sample (e.g., the presence or absence of a marker, or the expression level of a component, etc.) and comparing the measurements to measurements in a sample or set of samples from the same subject or other control subjects. The markers of the present teachings may be analyzed by any of a variety of conventional methods known in the art. "Analyzing" may include performing statistical analysis, such as normalizing data, determining statistical significance, determining statistical correlation, clustering algorithms, etc.
[0118] A "sample" in the context of the present teachings refers to any biological sample isolated from a subject, generally a blood or plasma sample that may contain circulating immune cells. Samples may include, without limitation, bodily fluids, plasma, serum, aliquots of whole blood, PBMCs (white blood cells or leucocytes), tissue biopsies, dissociated cells from tissue samples, urine samples, saliva samples, synovial fluid, lymphatic fluid, peritoneal fluid, and interstitial or extracellular fluid. A "blood sample" may refer to whole blood or a fraction thereof, including blood cells, plasma, serum, white blood cells or leucocytes. Samples may be obtained from a subject by means including, but not limited to, venipuncture, biopsy, needle aspiration, lavage, scraping, surgical incision, or intervention, or means known in the art.
[0119] A "dataset" is a set of numerical values obtained from evaluating a sample (or a population of samples) under desired conditions. The values of a dataset may be obtained, for example, by experimentally obtaining measurements from the samples and constructing the dataset from those measurements, or alternatively, by obtaining the dataset from a service provider such as a laboratory, or from a database or server on which the dataset is stored. Similarly, the term "obtaining a dataset associated with a sample" encompasses obtaining a set of data determined from at least one sample. Obtaining a dataset encompasses obtaining a sample and processing the sample, for example, to measure antibody binding or otherwise quantify a signaling response, to experimentally determine the data. The phrase also encompasses receiving a set of data from a third party, for example, who has processed the sample to experimentally determine the dataset.
[0120] "Measuring" or "measurement" in the context of the present teachings refers to determining the presence, absence, quantity, amount, or effective amount of a substance in a clinical sample or a sample derived from a subject (including the presence, absence, or concentration level of such a substance) and / or assessing the value or categorization of a clinical parameter in a subject based on a control group, e.g., a baseline level of a marker.
[0121] Classification can be performed according to a predictive modeling method that sets a threshold value to determine the probability that a sample belongs to a given class.Probability is preferably at least 50%, or at least 60%, or at least 70%, or at least 80% or higher.Classification can also be performed by determining whether the comparison of the obtained data set with a reference data set produces a statistically significant difference.If it produces a significant difference, the source sample of the data set is classified as not belonging to the reference data set class.Conversely, if such comparison does not differ statistically significantly from the reference data set, the source sample of the data set is classified as belonging to the reference data set class.
[0122] The predictive ability of a model can be evaluated according to its ability to provide a quality indicator, such as the Area Under the Curve (AUC), or the accuracy of a particular value or range of values. In some embodiments, the desired quality threshold is a predictive model that classifies samples with an accuracy of at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, at least about 0.95, or higher. As an alternative measure, the desired quality threshold can refer to a predictive model that classifies samples with an AUC of at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, or higher.
[0123] As known in the art, the relative sensitivity and specificity of a predictive model can be "tuned" to favor either selectivity index or sensitivity index, and the two indices have an inverse relationship.The limitations of the model as described above can be adjusted to provide a selected sensitivity or specificity level according to the specific requirements of the test being performed.One or both of the sensitivity and specificity can be at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, or higher.
[0124] As used herein, the term "theranosis" refers to the use of results obtained from a prognostic or diagnostic method to guide the selection, maintenance, or modification of a therapeutic therapy, including, but not limited to, the selection of one or more therapeutic agents, modification of dosage, modification of administration schedule, modification of dosing mode, and modification of formulation. Diagnostic methods used to inform diagnostic therapy can include any method that provides information about the state of a disease, condition, or symptom.
[0125] The terms "therapeutic agent", "therapeutic agent", or "therapeutic agent" are used interchangeably and refer to a molecule, compound, or any non-pharmacological therapy that provides some beneficial effect when administered to a subject. Beneficial effects include enabling the determination of a diagnosis, alleviating a disease, sign, disorder, or pathological condition, reducing or preventing the onset of a disease, sign, disorder, or condition, and generally arresting a disease, sign, disorder, or pathological condition.
[0126] As used herein, "treatment" or "treating" or "alleviating" or "alleviating" are used interchangeably. These terms refer to an approach to obtain beneficial or desired results, including but not limited to therapeutic benefit and / or preventive benefit. Therapeutic benefit refers to any treatment-related improvement or effect on one or more diseases, symptoms, or signs being treated. For preventive benefit, the formulation may be administered to a subject at risk of developing a particular disease, symptom, or sign, or to a subject reporting one or more physiological signs of a disease, even though the disease, symptom, or sign may not yet be manifest.
[0127] The term "effective amount" or "therapeutically effective amount" refers to an amount of an agent sufficient to achieve a beneficial or desired result. The therapeutically effective amount varies depending on the subject and the disease condition being treated, the subject's weight and age, the severity of the disease condition, the mode of administration, etc., which can be easily determined by one of ordinary skill in the art. This term also applies to a dose that provides an image for detection by any one of the imaging methods described herein. The specific dose will vary depending on the particular agent selected, the dosing regimen employed, whether it is administered in combination with other compounds, the timing of administration, the tissue being imaged, and the bodily delivery system in which the agent is delivered.
[0128] "Suitable conditions" shall have the meaning according to the context in which the term is used. That is, when used in relation to an antibody, the term shall mean conditions that allow the antibody to bind to its corresponding antigen. When used in relation to contacting an agent with a cell, the term shall mean conditions that allow the agent capable of doing so to enter the cell and perform its intended function. In one embodiment, the term "suitable conditions" as used herein means physiological conditions.
[0129] The term "antibody" includes full length antibodies and antibody fragments and may refer to natural antibodies from any organism, modified antibodies, or antibodies recombinantly produced for experimental, therapeutic, or other purposes, as further defined below. Examples of antibody fragments known in the art include Fab, Fab', F(ab')2, Fv, scFv, or other antigen-binding portion sequences of antibodies, produced by modification of whole antibodies or synthesized de novo using recombinant DNA technology. The term "antibody" includes monoclonal and polyclonal antibodies. Antibodies may be antagonistic, agonistic, neutralizing, inhibitory, or stimulatory. They may be humanized, glycosylated, bound to a solid support, and have other variants. Machine learning methods for predicting surgical outcomes
[0130] To obtain a predictive model of clinical outcome after surgery, many embodiments use machine learning methods that integrate single-cell analysis of immune cell responses using mass cytometry with multiple assessments of inflammatory plasma proteins in blood samples collected from patients before or after surgery. Many embodiments use multi-omics bootstrap (MOB) machine learning methods to predict the onset of SSC after surgery. MOB, according to various embodiments, integrates one or more omics data categories (e.g., categories described herein) by extracting the most robust features from each data layer and then combining those features, ensuring the stability of features selected during statistical modeling of omics data sets.
[0131] The development of stability selection methods (see, for example, Nicolai Meinshausen. Peter Buhlmann. Ann. Statist. 34 (3) 1436 - 1462, June 2006, the disclosure of which is incorporated herein by reference in its entirety) is a key element in the development of the MOB algorithm. Although the problem of variability is inherent and cannot be completely overcome, stability selection is able to characterize this variability by considering the frequency with which each feature is selected when multiple Lasso models are obtained based on subsampled data. The selection frequency gives a quantitative measure of the importance of each feature that can be easily interpreted from a biological point of view. It has been shown that stability selection requires much weaker assumptions for asymptotically constant variable selection compared to Lasso. In other words, instead of selecting one model, stability selection repeatedly subsamples the data and selects stable variables, i.e., variables that appear in a large proportion of the resulting models. The selected stable variables are defined by having a selection frequency above a selected threshold, as follows:
number
number
[0132] One of the difficulties of the previous methods is the difficulty of assessing noise. Since the goal is to distinguish noisy variables from the variables to be predicted, the use of negative control features is a suitable approach to develop an internal noise filter in the learning process. Negative control features refer to synthetically created noisy features. One of the main contributions of this work is that, if properly constructed, it makes it possible to adapt the previously mentioned thresholds from a distribution of artificial features in the stability selection process. Two schemes for generating these artificial features are considered. Both techniques involve augmenting the initial input to eventually produce an input matrix
number
number
number
number
number
number
number
number
number
[0133] Machine learning models are typically trained using a bootstrap procedure, among other steps, on multiple individual data layers, each data layer representing one type of data from multiple possible features and at least one artificial feature, each feature being selected from the group consisting of, for example, genomics, transcriptomics, proteomics, cytomics, metabolomics, clinical, and demographic data.
[0134] Each data layer contains data about a population of individuals, and each feature contains feature values for all individuals in that population of individuals. During machine learning, the feature values obtained for the population of individuals for each data layer are typically arranged in a matrix X with n rows and p columns, where each row corresponds to a respective individual and each column corresponds to a respective feature. In other words, the matrix X is a concatenation of p vectors, each vector relating to a respective feature and containing n feature values, typically one feature value for each individual.
[0135] For each data layer, each artificial feature is obtained from a non-artificial feature of the plurality of features via a mathematical operation performed on the feature values of the non-artificial feature. The mathematical operation is, for example, selected from the group consisting of replacement, sampling, combination, knock-off method, and inference. The replacement is, for example, a full replacement without replacement of the feature values. The sampling is typically a sampling with partial replacement of the feature values or without replacement of the feature values. The combination is, for example, a linear combination of the feature values. The knock-off method is, for example, a model X knock-off applied to the feature values. The inference is typically a fitting of a statistical distribution of the feature values, such as a normal distribution, an exponential distribution, a uniform distribution, or a Poisson distribution, and then a random inference sampling therefrom. The acquisition of the artificial feature is also called spiking of the artificial feature, and corresponds to instruction 2 in the pseudocode of Figures 3A and 3B.
[0136] The model uses weights β for a selected set of biological and clinical or demographic features. i Such weights β i are the initial weights w that are typically changed repeatedly during machine learning of the model. j is derived from
[0137] For each data layer during machine learning, at each bootstrap iteration, initial weights w are calculated for the features and at least one artificial feature associated with that data layer using an initial statistical learning technique. j The generation of bootstrap samples and the initial weights w j The estimates of correspond to instructions 4 and 5, respectively, in the pseudocode of Figures 3A and 3B.
[0138] The initial statistical learning technique is typically a sparse technique or a non-sparse technique. The initial statistical learning technique is, for example, a regression technique or a classification technique. Thus, the initial statistical learning technique is preferably selected from the group consisting of sparse regression techniques, sparse classification techniques, non-sparse regression techniques, and non-sparse classification techniques.
[0139] By way of example, initial statistical learning techniques may thus include linear or logistic linear regression techniques with L1 or L2 regularization, such as the Lasso technique or the Elastic Net technique (see, e.g., Tibshirani and Zou and Hastie, cited above), model-fitted linear or logistic linear regression techniques with L1 or L2 regularization, such as the Bolasso technique (see, e.g., Bach, Francis R. "Bolasso: model consistent lasso estimation through the bootstrap." Proceedings of the 25th international conference on Machine learning. 2008, the disclosure of which is hereby incorporated by reference in its entirety), relaxed Lasso (see, e.g., Meinshausen, Nicolai. "Relaxed lasso." Computational Statistics & Data Analysis 52.1 (2007): 374-393, the disclosure of which is hereby incorporated by reference in its entirety), random Lasso techniques (see, e.g., Wang, Sijian, et al., the disclosure of which is hereby incorporated by reference in its entirety), and the like. "Random lasso." The annals of applied statistics 5.1 (2011): 468), grouped Lasso techniques (see, e.g., Friedman, Jerome, Trevor Hastie, and Robert Tibshirani. Applications of the lasso and grouped lasso to the estimation of sparse graphical models. Technical report, Stanford University, 2010, the disclosure of which is hereby incorporated by reference in its entirety), LARS techniques (see, e.g., Eyraud, Remi, Colin De La Higuera, and Jean-Christophe Janodet.The rewriting algorithm is selected from the group consisting of linear or logistic-linear regression techniques without L1 or L2 regularization, nonlinear regression or classification techniques with L1 or L2 regularization, decision tree techniques, random forest techniques, support vector machine techniques also known as SVM techniques, neural network techniques, and kernel smoothing techniques.
[0140] Next, the calculated initial weights w j At least one selected feature is determined for each data layer based on a statistical criterion that depends on the calculated initial weights w j The determination of the effective weights depends on the effective weights among the weights (e.g., non-zero weights when the initial statistical learning technique is a sparse regression technique, or weights above a predefined weight threshold when the initial statistical learning technique is a non-sparse regression technique). The determination of the effective weights corresponds to instruction 6 in the pseudocode of Figures 3A and 3B.
[0141] As an example, the effective weights are non-zero weights when the initial statistical learning technique is selected from the group consisting of: linear or logistic linear regression techniques with L1 or L2 regularization, such as the Lasso technique and the Elastic Net technique; model-fitting linear or logistic linear regression techniques with L1 or L2 regularization, such as the Bolasso technique, relaxed Lasso, random Lasso technique, grouped Lasso technique, and LARS technique; non-linear regression or classification techniques with L1 or L2 regularization; and kernel smoothing techniques.
[0142] A "non-zero weight" is one whose absolute value is below a very low predefined threshold, e.g. 10 -5 (also written as 1e-5). Therefore, a "nonzero weight" is usually a weight with an absolute value of 10 -5 Say the weight is greater than.
[0143] Alternatively, when the initial statistical learning technique is selected from the group consisting of a linear or logistic linear regression technique without L1 or L2 regularization, a decision tree technique, a random forest technique, a support vector machine technique, and a neural network technique, the effective weights are those weights that are above a predefined weight threshold. In the example of a neural network technique, the effective weights are those weights that are above a predefined weight threshold in the first layer of the corresponding neural network.
[0144] Those skilled in the art will recognize that the support vector machine technique is considered a sparse technique that uses support vectors, which leads to keeping only the support vectors. Those skilled in the art will also recognize that in the case of decision tree techniques, the aforementioned weights correspond to the importance of the features, and thus the effective weights are the features whose splits in the decision tree cause a certain degree of reduction in impurity.
[0145] Optionally, the initial weights w j is further calculated for multiple values of a hyperparameter λ, the hyperparameter λ being a parameter whose value is used to control the learning process. The hyperparameter λ is a regularization coefficient that is typically used in the context of sparse initial techniques according to a respective mathematical norm. The mathematical norm is, for example, the P-norm, where P is an integer.
[0146] As an example, the hyperparameter λ is the initial weight w when the initial statistical learning technique is the Lasso technique. j is an upper bound on the coefficients of the L1 norm of , where the L1 norm refers to the sum of all the absolute values of the initial weights.
[0147] As another example, the hyperparameter λ is the initial weight w when the initial statistical learning technique is the Elastic Net technique. j The sum of the L1 norms of and the initial weights w j is the upper bound of the sum of the L2 norms of both coefficients, where the L1 norm is defined above and the L2 norm refers to the square root of the sum of all the squared values of the initial weights.
[0148] For feature selection, the statistical criterion may depend, for example, on the frequency of occurrence of the effective weights. As an example, the statistical criterion is that each feature is selected if its frequency of occurrence is greater than a frequency threshold.
[0149] To determine the frequency of occurrence for each feature, a unit occurrence frequency is calculated for each value of the hyperparameter λ, where the unit occurrence frequency is equal to the number of effective weights associated with the feature for successive bootstrap iterations divided by the number of bootstrap iterations used for the feature. The occurrence frequency is then typically equal to the highest unit occurrence frequency among the unit occurrence frequencies calculated for all values of the hyperparameter λ. The determination of the occurrence frequency of each feature, also referred to as the selection frequency, corresponds to instructions 8 and 10 in the pseudocode of Figures 3A and 3B.
[0150] The frequency threshold is usually calculated according to the occurrence frequency obtained for the artificial features. This frequency threshold is, for example, two standard deviations from the mean or median of the occurrence frequencies obtained for the artificial features. Alternatively, the frequency threshold is three times the mean of the occurrence frequencies obtained for the artificial features. Still alternatively, the frequency threshold is equal to the maximum between one of the above examples of the calculated frequency threshold and a predefined frequency threshold. The calculation of the frequency threshold corresponds to instruction 11 in the pseudocode of Figures 3A and 3B.
[0151] Finally, feature selection is performed for each layer based on statistical criteria. For example, the features selected are those whose occurrence frequency is greater than a frequency threshold. Feature selection corresponds to instruction 12 in the pseudocode of Figures 3A and 3B.
[0152] As an example, each value of the hyperparameter λ is selected according to a predetermined scheme from values between the lower limit and the upper limit of the selected value range of the hyperparameter λ. As a variant, the values of the hyperparameter λ are evenly distributed between the lower limit and the upper limit of the selected value range of the hyperparameter λ. When the initial statistical learning technique is the Lasso technique or the Elastic Net technique, the hyperparameter λ is usually between 0.5 and 100.
[0153] For the bootstrapping process, the number of bootstrap iterations is typically between 50 and 100,000, preferably between 500 and 10,000, and more preferably equal to 10,000.
[0154] During machine learning, after feature selection, the model weights β i is further calculated using a final statistical learning technique on data associated with the selected set of features.
[0155] The final statistical learning technique is typically a sparse technique or a non-sparse technique. The final statistical learning technique is, for example, a regression technique or a classification technique. Therefore, the final statistical learning technique is preferably selected from the group consisting of sparse regression techniques, sparse classification techniques, non-sparse regression techniques, and non-sparse classification techniques.
[0156] By way of example, the final statistical learning technique is therefore selected from the group consisting of: linear or logistic linear regression techniques with L1 or L2 regularization, such as the Lasso technique and the Elastic Net technique; model-fitted linear or logistic linear regression techniques with L1 or L2 regularization, such as the bo-Lasso technique, the soft-Lasso technique, the random Lasso technique, the grouped Lasso technique, the LARS technique; linear or logistic linear regression techniques without L1 or L2 regularization; non-linear regression or classification techniques with L1 or L2 regularization; decision tree techniques; random forest techniques; support vector machine techniques, also known as SVM techniques; neural network techniques; and kernel smoothing techniques.
[0157] During the use phase following machine learning, a surgical risk score is calculated according to an individual's measurements against a set of selected features.
[0158] As an example, if the final statistical learning technique is a classification technique, the surgical risk score can be calculated by assigning each measurement a respective weight β for the set of selected features. i The probability is calculated according to a weighted sum of the values multiplied by
[0159] According to this example, the surgical risk score is typically calculated as follows:
number
[0160] As a further example, Odd is the exponent of the weighted sum. Odd is calculated, for example, according to the following formula:
number
[0161] Those skilled in the art will recognize that in the above formula, the weight β i and measurement X i It will be noted that can be negative or positive.
[0162] As another example, if the final statistical learning technique is a respective regression technique, the surgical risk score may be calculated by assigning each measurement a respective weight β for the set of selected features. i is a term that depends on a weighted sum of values multiplied by
[0163] Following this alternative example, the surgical risk score is equal to the index of the weighted sum, typically calculated by the formula above.
[0164] As an optional addition, during the machine learning, prior to obtaining the artificial features, additional values of the plurality of non-artificial features are generated using data augmentation techniques based on the obtained values. Following this optional addition, the artificial features are subsequently obtained according to both the obtained values and the generated additional values.
[0165] According to this optional addition, the data augmentation technique is typically a non-composite technique or a composite technique, for example selected from the group consisting of the SMOTE technique, the ADASYN technique, and the SVMSMOTE technique.
[0166] According to this addition on an as-needed basis, for a given non-artificial feature, the fewer values obtained, the more additional values are generated.
[0167] According to this optional addition, the generation of this additional value using a data augmentation technique is an optional additional step before the bootstrap process. According to the above, this generation makes it possible to "augment" the initial input matrix X and the corresponding output vector Y with a data augmentation algorithm, i.e. to increase the size of the matrix X and the vector Y, respectively. If the matrix X is of size (n,p), then the vector Y is of size (n). This generation step generates an X augmented and Y of size (n') augmented where n'>n.
[0168] This generation is preferably more sophisticated than a bootstrap process. The goal is to "augment" the input by creating synthetic samples constructed using the samples obtained, rather than random duplication of the samples. Indeed, if one simply duplicated the non-artificial feature values, the augmentation would not be substantially different from a bootstrap process in which the non-artificial feature values may already be oversampled and / or duplicated. With the addition of data augmentation as needed, the bootstrap process would therefore be supplied with new data points added to the original one.
[0169] For classification, the data augmentation technique is, for example, the SMOTE technique, also called the SMOTE algorithm or SMOTE. SMOTE first randomly selects a minority class instance A and finds its K nearest minority class neighbors (using the K nearest neighbor method). Then, a composite instance is created by randomly selecting one of the K nearest neighbors B and connecting A and B to form a line segment in the feature space. The composite instance is generated as a convex combination of the two selected instances. Those skilled in the art will recognize that this technique is also a means of artificially balancing between classes. As a variant, the data augmentation technique is the ADASYN technique or the SVMSMOTE technique.
[0170] In the case of surgical site complications, i.e., when the risk to be determined is SSC, the algorithm is applied to each layer independently. The layers used to determine the SSC are, for example: immune cell frequency (containing 24 cell frequency features), basal signaling activity of each cell subset (312 basal signaling features), signaling response capacity to each stimulation condition (6 data layers each containing 312 features), and plasma proteomics (276 proteomic features).
[0171] As an example, there are 41 samples for each layer. In other words, the number of feature values for each feature, n, is equal to 41 in this example. So, for the immune frequency layer, the dimension of the matrix X is 41 samples (n) x 24 features (p). For the base signaling case, the matrix X is of dimension 41 x 312. Y is a vector of outcome values, i.e. occurrences of SSC. This vector Y is in this case a vector of length 41. So, for each sample, one respective outcome value, i.e. one SSC value, is determined.
[0172] In this example, M is chosen to be equal to 10,000, which allows for sufficient sampling to derive estimates of selection frequencies for the artificial features.
[0173] The selected range value of the hyperparameter λ is between 0.5 and 100 when the statistical learning technique is the Lasso technique or the Elastic Net technique.
[0174] In this example, the frequency threshold is chosen to be equal to three times the average occurrence frequency obtained for the artificial features, thereby reducing variability and allowing tighter control over feature selection.
[0175] In the following examples of Figures 2, 3A and 3B, those skilled in the art will notice that the mathematical operations used to obtain the artificial features are replacement or sampling, and will understand that other mathematical operations are also applicable, including the other mathematical operations mentioned in the above description, namely combination, knock-off and inference. Similarly, in these examples, the statistical learning techniques used to calculate the initial weights are sparse regression techniques such as Lasso and Elastic Net, and those skilled in the art will understand that other statistical learning techniques may also be applicable, including the other statistical learning techniques mentioned in the above description, namely non-sparse techniques and classification techniques. In these examples, the effective weights are non-zero weights, and those skilled in the art will understand that other effective weights, such as weights above a predefined weight threshold, are also applicable, depending on the type of initial statistical learning technique, as explained above.
[0176] Turning to FIG. 2, a schematic diagram of the MOB algorithm used according to many embodiments is shown. In such an embodiment, at 202, subsets are obtained from the original cohort by a procedure using repeated sampling with or without replacement at the individual data layers. In many embodiments, artificial features are included by random sampling from the distribution of the original samples or with replacement and added to the original dataset. At 204, for each of the subsets, an individual model is calculated, for example using the Lasso algorithm, and features are selected based on their contribution in the model (in the case of Lasso, non-zero features are selected). At 206, using the features selected for each model, many embodiments obtain a stability path that displays the frequency of selection of each contributing feature (artificial or not) by hyperparameters. The distribution of the selection of the artificial features is then used to estimate the distribution of noise in the dataset. A cutoff for the relevant biological or clinical features is calculated based on the estimated distribution of noise in the dataset. The relevant features from each layer are then used and combined in a final model for the prediction of the relevant surgical outcome. At 208, the final synthesis of the model occurs, where each of the individual layers are combined through a process of selection similar to that described at 202-206. At 208, all the top features are combined and used as predictors in the final layer.
[0177] 3A-3B show exemplary pseudocode of the MOB algorithm of various embodiments. In many embodiments, MOB uses a multiple resampling procedure with or without replacement, called bootstrapping, for each data layer. At each data layer, simulated features are spiked in the original dataset at each bootstrap iteration to estimate the robustness of the biological feature selection compared to the artificial features. The optimal cutoff of the biological or clinical features is selected using the distribution of the artificial features used to estimate the behavior of noise on the robustness of the biological or clinical features from that data layer. The MOB algorithm then selects features at each layer that are above the optimal threshold calculated from the distribution of noise, and builds the final model using features from each data layer that pass the optimal robustness threshold. In many embodiments, the performance is benchmarked and the robustness of the feature selection is evaluated based on the simulated and biological data.
[0178] In the embodiment demonstrated in Figures 3A-3B, such an embodiment first obtains a subset from the original cohort in a procedure using repeated sampling with or without replacement at each data layer. For each bootstrap, an artificial feature is constructed by selecting one by one features (vectors of size p) of the original data matrix. To construct the artificial features, such an embodiment performs either random permutation (corresponding to randomly drawing all values of the vector without replacement) or random sampling (constructing a new vector of size p by randomly drawing p elements of the original feature with replacement). This process is repeated for each feature independently. Such an embodiment concatenates the artificial features with the real features and then draws samples from this concatenated dataset with or without replacement.
[0179] Then, for each of the subsets, an individual model is calculated, for example using the Lasso algorithm (Tibshirani, R. (1996). Journal of the Royal Statistical Society: Series B (Methodological), 58(1), 267-288.), and features are selected based on their contribution in the model (in the case of the Lasso, non-zero features are selected). At this stage of the process, contributing features have non-zero coefficients when fitting the Lasso. This is the same for any other technique that induces sparsity, such as Elastic Net. In the case of non-sparse regression techniques, an arbitrary contribution threshold needs to be defined. The algorithm can be adapted to suit the machine learning technique used. The Lasso is a well-known sparse regression technique, but other techniques that select a subset of the original features can be used. For example, Elastic Net (EN), a combination of Lasso and Ridge, would also work (Zou, H., & Hastie, T. (2005). Journal of the royal statistical society: series B (statistical methodology), 67(2), 301-320.).
[0180] Furthermore, using the features selected for each model, a stability path can be obtained by the hyperparameters, which displays the selection frequency of each contributing feature (artificial or real). The stability path is the output matrix of the process before any graphical transformation is performed. Its size is (p, #{Lambda}). Each value (feature_i, lambda_j) corresponds to the frequency of selection of feature_i with parameter lambda_j. From this matrix, such an embodiment can display the path of each feature (e.g., FIG. 2, 206), where each row corresponds to the frequency of selection of each feature across all lambdas examined. The distribution of the selection of the artificial features is then used to estimate the distribution of noise in the dataset. A cutoff for the relevant biological or clinical features is calculated based on the estimated distribution of noise in the dataset. Then, only the relevant features from each layer are used and combined in the final model for the prediction of the relevant surgical outcome. In the embodiment of FIG. 3B, the final model uses the selected features obtained at each data layer. Therefore, the input of the final model is of size (n, p_stable), where p_stable is the number of selected features (all layers included). p_stable is significantly smaller than the dimension of the original feature space. This reduced matrix is then trained for predicting outcomes.
[0181] The exemplary embodiment shown in FIG. 3B provides a wider range of hyperparameters. For example, in the exemplary embodiment shown in FIG. 3A, the optimal parameter selection is constrained to a leave-one-out cross-validation fit.
number
[0182] In addition, the exemplary embodiment of Fig. 3B allows for the use of a selection threshold based on the distribution of all artificial features; specifically, a cutoff is defined based on the entire distribution of artificial features. To define the cutoff, such an embodiment takes the maximum of the probability of selection of each artificial feature and then takes the average of these maximum values. From this average, such an embodiment can construct a threshold (e.g., 3 standard deviations from the mean). In contrast, in the embodiment shown in Fig. 3A, only the artificial features with the maximum selection frequency may be used.
[0183] Furthermore, the exemplary embodiment of FIG. 3B allows for the combination of artificial generation and bootstrapping procedures to simplify the algorithmic complexity.
[0184] More specifically, the embodiment as shown in FIG. 1. Iterating over a number of bootstrap iterations to get a fair assessment of the sampleability of artificial feature selection. We track the index to see how sampling from the original distribution or via replacement behaves over multiple trials. This corresponds to the first for loop in the algorithm and yields the results in lines 10-13. 2. A replacement or random sampling is taken from the original data set and the generated matrix is the juxtaposition of the original matrix and the new matrix of computed artificial features. The number of artificial features (p') can vary but is usually chosen to match the number of original features included in the algorithm. For computational purposes, if p is very large, a smaller number can be chosen for p'. 3. In order to properly explore the behavior of the selection with respect to the selected algorithm hyperparameters, a grid search type method is used to evaluate different combinations of hyperparameters and then plot the curves of the "stability path" (see Figure 2). This step is also a measure to avoid missing information if only a limited amount of hyperparameters are examined. The range of examined hyperparameters can be thoroughly explored to avoid artifacts (e.g., examining lambda=0 in the case of Lasso would result in selecting all features for all bootstrap steps, leading to cases where the maximum frequency of selections are all equal to 1). 4-6. For each selected value of the hyperparameters, at a given number of spikes, the resampling procedure allows the estimation of the model fit behavior and the selection of the most robust features to small changes in the dataset. By model fit behavior, the model refers to the assessment of the probability of selection by Lasso for a given value of the hyperparameter. Bootstrapping (the resampling procedure) allows to induce few perturbations in the original dataset, and only the more robust features are selected with a higher frequency than the others. The EN or Lasso algorithm is highly prone to small changes in the original cohort, especially in the sense that it tends to select less robust features, thus making biological interpretation and robustness to new cohorts difficult. In this situation, resampling produces small variations centered on the original cohort. This procedure allows a proper investigation of robustness in feature selection. 8. Coefficient extraction in conjunction with inducing sparsity through L1 regularization. This extraction uses a simple cutoff for non-zero coefficients (typically 1e-5 in absolute value) to select the top performing features at each step of the bootstrap procedure. This selection of top performing features at each iteration of the bootstrap procedure allows the model to derive the frequency of selection for each feature in the dataset. 10-12. Because the model includes spiked artificial features, it can use the stability path definition to estimate the distribution of typical "noise" in the dataset and use that distribution to calculate a cutoff for the relevant features. This cutoff is typically two standard deviations away from the mean or median stability path of the artificial features, or three times the average of the maximum probability of selection of the artificial features. An arbitrary fixed threshold can also be added to take the maximum between the constructed threshold and an arbitrary fixed threshold. Some embodiments construct a threshold by taking the maximum of the probability of selection of each artificial feature and then taking the average of these maximums (2 * , 3 * , or a combination of this threshold with any fixed threshold).
[0185] Turning now to FIG. 4, an exemplary method for generating multi-omics biological data and generating a predictive MOB model for SSC that integrates multi-omics biological data and clinical data is shown. At 402, certain embodiments obtain a biological sample from an individual. While FIG. 4 shows a blood draw (whole blood and plasma), various embodiments obtain biological samples from other tissues, fluids, and / or other biological sources. The biological sample may be obtained pre-operatively (including the day of surgery or "day of surgery") and / or post-operatively. Pre-operative samples may be obtained 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 1 day, and / or 0 days (i.e., before the first incision on the day of surgery), while post-operative samples may be obtained within 24 hours after surgery, including 0 hours, 1 hour, 3 hours, 6 hours, 8 hours, 10 hours, 12 hours, 16 hours, 18 hours, and / or 24 hours after surgery (i.e., Post Operative Day 1 (POD)). In many embodiments, multi-omics data is obtained from the biological sample at 404. Such multi-omics data may include cytomic data and plasma protein expression data obtained using mass cytometry. Further embodiments utilize additional forms of omics data to identify cytomic, proteomic, transcriptomic, and / or genomic data applicable to certain embodiments. In certain embodiments, a predictive MOB model is generated based on the omics (including multi-omics) data at 406, and / or clinical data is generated, and such models can be generated by the methods described herein.
[0186] 5A-5C, exemplary embodiments are shown that demonstrate the efficacy of embodiments to predict SSC after abdominal surgery. Specifically, FIG. 5A shows biological samples obtained pre-operatively (DOS) coupled with post-operative assessment 30 days after surgery. A summary of the data used is provided in Table 1. FIG. 5B shows exemplary data showing an AUC of 0.82 (95% confidence interval, CI [0.66-0.94], Mann-Whitney rank sum test) for a model trained on multi-omics data alone. However, many embodiments implement machine learning techniques that integrate multi-omics data and clinical data to derive a predictive model for SSC. Figure 5C shows an exemplary MOB model integrating preoperative clinical variables with cytomics and plasma proteomics variables collected at DOS, and this MOB model predicts SSC with better predictive performance (AUC = 0.92, 95% CI [0.84-0.99], Mann-Whitney rank sum test) than models built based on biological or clinical data alone.
[0187] In addition, Figures 6-7 show exemplary performance data of additional embodiments. Specifically, Figure 6 shows another exemplary DOS model predicting SSC with an AUC of 0.77, 95% CI [0.65-0.89], n=93, Mann-Whitney rank sum test, and a summary of the data used to generate Figure 6 is provided in Table 2. Additionally, Figure 7 shows an exemplary MOB prediction model of SSC derived from analyzed patient samples collected 24 hours after abdominal surgery (POD1) with an AUC of 0.86. Methods for generating multi-omics biological data
[0188] In many embodiments, methods for generating predictive models of surgical complications, such as SSC, rely on multi-omics analysis of biological samples (e.g., blood-based samples, tumor samples, and / or any other suitable biological samples) obtained from an individual before or after surgery to obtain, for example, determination of immune cell subset frequencies and signaling activities, and changes in plasma proteins.
[0189] The biological sample can be of any suitable type that allows for the analysis of one or more cells, proteins, preferably a blood sample. The sample can be obtained from an individual once or multiple times. Multiple samples can be obtained from different sites on an individual, from an individual at multiple different times, or any combination thereof.
[0190] According to certain embodiments, at least one biological sample is obtained pre-operatively (including the day of surgery or "DOS"). According to certain embodiments, at least one biological sample is obtained post-operatively. According to certain embodiments, at least one biological sample is obtained pre-operatively and at least one biological sample is obtained post-operatively. The pre-operative biological sample may be collected 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 1 day, and / or 0 days (i.e., before the first incision on the day of surgery). The post-operative biological sample may be obtained within 24 hours after surgery, including 0 hours, 1 hour, 3 hours, 6 hours, 8 hours, 10 hours, 12 hours, 16 hours, 18 hours, and / or 24 hours after surgery (i.e., POD1).
[0191] The biological sample can be from any source that contains immune cells. In some embodiments, the biological sample for analyzing immune cell responses is blood. However, the PBMC fraction of a blood sample can also be used. In some embodiments, the biological sample for proteomic analysis is the plasma fraction of a blood sample, but the serum fraction can also be used.
[0192] In some embodiments, the sample is activated ex vivo, which as used herein refers to contacting a sample, e.g., a blood sample or cells derived therefrom, with a stimulating agent (one example of which is shown in FIG. 4, 404) outside the body. In some embodiments, whole blood is preferred. The sample may be diluted or suspended in a suitable medium that maintains cell viability, e.g., minimal medium, PBS, etc. The sample may be fresh or frozen. Stimulating agents of interest include agents that activate innate or adaptive cells, e.g., TLR4 agonists such as LPS, and / or one or a combination of IL-1β, IL-2, IL-4, IL-6, TNFα, IFNα, or PMA / ionomycin. Generally, the activation of cells ex vivo is compared to a negative control, e.g., vehicle alone or an agent that does not induce activation. The cells are cultured in the biological sample for a period of time sufficient to activate the immune cells. For example, the activation time can be up to about 1 hour, up to about 45 minutes, up to about 30 minutes, up to about 15 minutes, and may be up to about 10 minutes or up to about 5 minutes. In some embodiments, the period is up to about 24 hours, or from about 5 to about 240 minutes. After activation, the cells are fixed for analysis.
[0193] In many embodiments, cytomics and proteomics features are detected using affinity reagents. "Affinity reagent" or "specific binding member" may be used to refer to affinity reagents, such as antibodies, ligands, etc., that selectively bind to the proteins or markers of the invention. The term "affinity reagent" includes any molecule, such as peptides, nucleic acids, small organic molecules. For some purposes, the affinity reagent selectively binds to cell surface or intracellular markers, such as CD3, CD4, CD7, CD8, CD11b, CD11c, CD14, CD15, CD16, CD19, CD24, CD25, CD27, CD33, CD45, CD45RA, CD56, CD61, CD66, CD123, CD235ab, HLA-DR, CCR2, CCR7, TCRγδ, OLMF4, CRTH2, and CXCR4. For other purposes, affinity reagents selectively bind to cellular signaling proteins, particularly those capable of detecting the activation state of one signaling protein relative to the activation state of another. Signaling proteins of interest include, without limitation, pSTAT3, pSTAT1, pCREB, pSTAT6, pPLCγ2, pSTAT5, pSTAT4, pERK1 / 2, pP38, prpS6, pNF-κB(p65), pMAPKAPK2(pMK2), pP90RSK, IκB, cPARP, FoxP3, and Tbet.
[0194] In some embodiments, proteomic features are measured, including measuring circulating extracellular proteins.Therefore, other affinity reagents of interest bind to plasma proteins.Particularly interesting plasma protein targets include IL-1β, ALK, WWOX, HSPH1, IRF6, CTNNA3, CCL3, sTREM1, ITM2A, TGFα, LIF, ADA, ITGB3, EIF5A, KRT19 and NTproBNP.
[0195] In some embodiments, cytomics features are measured, including measuring surface or intracellular proteins at the single cell level among immune cell subsets. Immune cell subsets include, for example, neutrophils, granulocytes, basophils, monocytes, dendritic cells (DCs), such as myeloid dendritic cells (mDCs) and plasmacytoid dendritic cells (pDCs), B cells, or T cells, such as regulatory T cells (Tregs), naive T cells, memory T cells, and NK-T cells. Immune cell subsets are more specifically neutrophils, granulocytes, basophils, CXCR4 + Neutrophil, OLMF4 + Neutrophil, CD14 + CD16 - Classical monocytes (cMCs), CD14 - CD16 + Non-classical monocytes (ncMCs), CD14 + CD16 + Intermediate monocytes (iMCs), HLADR + CD11c + Myeloid dendritic cells (mDC), HLADR + CD123 + Plasmacytoid dendritic cells (pDC), CD14 + HLADR - CD11b + Monocytic myeloid-derived suppressor cells (M-MDSC), CD3 + CD56 + NK-T cells, CD7 + CD19 - CD3 - NK cells, CD7 + CD56loCD16hi NK cells, CD7 + CD56 hi CD16 lo NK cells, CD19 + B cells, CD19 + CD38 + Plasma cells, CD19 + CD38 - non-plasma B cells, CD4 + CD45RA + Naive T cells, CD4 + CD45RA-memory T cells, CD4 + CD161 + Th17 cells, CD4 + Tbet +Th1 cells, CD4 + CRTH2 + Th2 cells, CD3 + TCRγδ + γδT cells, Th17 CD4 + T cells, CD3 + FoxP3 + CD25 + regulatory T cells (Treg), CD8 + CD45RA + Naive T cells, and CD8 + Includes CD45RA memory T cells.
[0196] In some embodiments, both proteomic and cytomic features are measured in a biological sample.
[0197] In some embodiments, the affinity reagent is a peptide, polypeptide, oligopeptide, or protein, particularly an antibody, or an oligonucleotide, particularly an aptamer and specific binding fragment, and variants thereof. A peptide, polypeptide, oligopeptide, or protein may be composed of naturally occurring amino acids and peptide bonds, or synthetic peptidomimetic structures. Thus, as used herein, "amino acid" or "peptide residue" includes both naturally occurring and synthetic amino acids. Proteins that contain non-naturally occurring amino acids may be synthesized or, in some cases, recombinantly produced; see van Hest et al., FEBS Lett 428:(l-2) 68-70 May 22, 1998, and Tang et al., Abstr. Pap Am. Chem. S218: U138 Part 2 Aug. 22, 1999. Both documents are expressly incorporated herein by reference.
[0198] Many antibodies have been produced that specifically bind to phosphorylated isoforms of proteins but not to non-phosphorylated isoforms of proteins, many of which are commercially available (see, for example, Cell Signaling Technology, www.cellsignal.com or Becton Dickinson, www.bd.com). Many such antibodies have been produced for the study of reversibly phosphorylated signaling proteins. In particular, many such antibodies have been produced that specifically bind to phosphorylated, activated isoforms of proteins and plasma proteins. Examples of proteins that can be analyzed using the methods described herein include, but are not limited to, phospho(p)rpS6, pNF-κB(p65), pMAPKAPK2(pMK2), pSTAT5, pSTAT1, pSTAT3, etc.
[0199] The methods of the invention may utilize affinity reagents that include a label, labeling element, or tag. By label or labeling element is meant a molecule that can be detected directly (i.e., a primary label) or indirectly (i.e., a secondary label), e.g., a label can be visualized and / or measured or otherwise identified such that its presence or absence can be known.
[0200] The compounds may be directly or indirectly conjugated to a label that provides a detectable signal, such as a non-radioactive isotope, a radioactive isotope, a fluorophore, an enzyme, an antibody, an oligonucleotide, a particle such as a magnetic particle, a chemiluminescent molecule, a molecule detectable by mass spectrometry, or a specific binding molecule. Specific binding molecules include pairs such as biotin and streptavidin, digoxin and antidigoxin, etc. Examples of labels include, but are not limited to, metal isotopes, optical fluorescent and chromogenic dye containing labels, labeling enzymes, and radioisotopes. In some embodiments of the invention, these labels can be conjugated to affinity reagents. In some embodiments, one or more affinity reagents are uniquely labeled.
[0201] Labels include optical labels such as fluorescent dyes or moieties. Fluorophores can be either "small molecule" fluorophores or protein fluorophores (e.g., green fluorescent protein and all its variants). In some embodiments, antibodies specific for activation states are labeled with quantum dots, as disclosed by Chattopadhyay et al (2006) Nat. Med. 12, 972-977. Quantum dot labeled antibodies can be used alone or in conjunction with organic fluorochrome conjugated antibodies to increase the total number of available labels. As the number of labeled antibodies increases, so does the ability to subclassify known cell populations.
[0202] Antibodies can be labeled using chelated or caged lanthanides as disclosed by Erkki et al. (1988) J. Histochemistry Cytochemistry, 36:1449-1451 and US Patent No. 7,018850. Other labels are tags suitable for Inductively Coupled Plasma Mass Spectrometer (ICP-MS) as disclosed by Tanner et al. (2007) Spectrochimica Acta Part B: Atomic Spectroscopy 62(3):188-195. Isotopic labels suitable for mass cytometry can be used as described, for example, in US Published Application No. 2012-0178183.
[0203] Alternatively, detection system based on FRET can be used.FRET has application in the present invention, for example, to detect activation state involving clustering or multimerization, where the proximity of two FRET labels is changed due to activation.In some embodiments, at least two fluorescent labels are used that are members of a fluorescence resonance energy transfer (FRET) pair.
[0204] When using fluorescently labeled components in the methods and compositions of the present invention, it will be appreciated that a variety of fluorescence monitoring systems, such as cytometric measurement device systems, can be used to practice the present invention. In some embodiments, flow cytometry systems are used, or systems dedicated to high throughput screening, such as 96 well or more microtiter plates, are used. Methods for assaying fluorescent substances are well known in the art and are described, for example, in Lakowicz, JR, Principles of Fluorescence Spectroscopy, New York: Plenum Press (1983); Herman, B., Resonance energy transfer microscopy, in: Fluorescence Microscopy of Living Cells in Culture, Part B, Methods in Cell Biology, vol. 30, ed. Taylor, DL & Wang, Y.-L., San Diego: Academic Press (1989), pp. 219-243; Turro, NJ, Modern Molecular Photochemistry, Menlo Park: Benjamin / Cummings Publishing Col, Inc. (1978), pp. 296-361.
[0205] The detecting, sorting, or isolating steps of the methods of the invention may involve fluorescence-activated cell sorting (FACS) techniques, where FACS is used to select cells from a population that contain a particular surface marker, or the selection step may involve the use of magnetically responsive particles as a retrievable support for target cell capture and / or background subtraction. A variety of FACS systems are known in the art and may be used in the methods of the invention (see, e.g., WO 99 / 54494, filed April 16, 1999, and U.S. Patent No. 20010006787, filed July 5, 2001, each of which is expressly incorporated herein by reference).
[0206] In some embodiments, a FACS cell sorter (e.g., a FACSVantage™ cell sorter, Becton Dickinson Immunocytometry Systems, San Jose, Calif.) is used to sort and collect cells (positive cells) based on their activation profile in the presence or absence of increased activation levels in signaling proteins in response to a modulator. Other commercially available flow cytometers include the LSR II and Canto II, both available from Becton Dickinson. For additional information regarding flow cytometers, see Shapiro, Howard M., Practical Flow Cytometry, 4 th Ed., John Wiley & Sons, Inc., 2003.
[0207] In some embodiments, cells are first contacted with a labeled activation state-specific affinity reagent (e.g., an antibody) directed to a particular activation state of a particular signaling protein. In such embodiments, the amount of affinity reagent bound on each cell can be measured by passing droplets containing the cells through a cell sorter. The cells can be separated from other cells by applying an electromagnetic charge to droplets containing positive cells. Cells selected as positive can then be collected in a sterile collection vessel. These cell sorting procedures are described in detail, for example, in the FACSVantage™ Training Manual, which is incorporated herein by reference in its entirety, see especially Sections 3-11 to 3-28 and 10-1 to 10-17. For detection systems, see the patents, applications, and articles referenced and incorporated above.
[0208] In some embodiments, the activation level of intracellular proteins is measured using inductively coupled plasma mass spectrometry (ICP-MS). Affinity reagents labeled with specific elements bind to the marker of interest. When cells are placed in the ICP, they are atomized and ionized. The elemental composition of cells containing the labeled affinity reagents bound to the signaling protein is measured. The presence and intensity of the signal corresponding to the label on the affinity reagent indicates the level of the signaling protein on the cell (Tanner et al. Spectrochimica Acta Part B: Atomic Spectroscopy, 2007 Mar;62(3):188-195.).
[0209] Mass cytometry has applications in analysis, for example as described in the examples provided herein. Mass cytometry, or CyTOF (DVS Sciences), is a variant of flow cytometry in which antibodies are labeled with heavy metal ion tags rather than fluorescent dyes. Readout is by time-of-flight mass spectrometry. This allows more antibody specificities to be combined in a single sample without significant spillover between channels. See, for example, Bodenmiller at a. (2012) Nature Biotechnology 30:858-867.
[0210] One or more cells or cell types or proteins can be isolated from a body sample. Cells can be separated from a body sample by red blood cell lysis, centrifugation, elutriation, density gradient separation, apheresis, affinity selection, panning, FACS, centrifugation using Hypaque, solid supports with antibodies attached (magnetic beads, beads in a column, or other surfaces), etc. A relatively homogeneous population of cells can be obtained by using antibodies specific to markers identified in a particular cell type. Alternatively, a heterogeneous population of cells can be used, such as circulating peripheral blood mononuclear cells.
[0211] In some embodiments, the phenotypic profile of a population of cells is determined by measuring the activation level of signaling proteins. The methods and compositions of the present invention can be used to examine and profile the state of any signaling protein or collection of such signaling proteins in a cellular pathway. Single or multiple individual pathways can be profiled (sequentially or simultaneously), or a subset of signaling proteins within a single pathway or across multiple pathways can be examined (sequentially or simultaneously).
[0212] In some embodiments, the basis for classifying cells is that the distribution of activation levels for one or more specific signaling proteins is different for different phenotypes. A certain activation level, or more generally a range of activation levels of one or more signaling proteins found in a cell or a population of cells, indicates that the cell or population of cells belongs to a specific phenotype. In addition to the activation level of a signaling protein, other measurements such as the cellular levels (e.g., expression levels) of biomolecules that may not contain signaling proteins can also be used to classify cells, and it will be appreciated that these levels also follow a distribution. Thus, one or more activation levels of one or more signaling proteins of a cell or population of cells can be used, optionally in conjunction with the levels of one or more biomolecules that may or may not contain signaling proteins, to classify a cell or population of cells into a class. It is understood that activation levels can exist as a distribution, and that the activation level of a particular element used to classify cells can be a particular point on the distribution, but more generally a portion of the distribution. In addition to the activation levels of intracellular signaling proteins, the levels of intracellular or extracellular biomolecules, e.g., proteins, can be used alone or in combination with the activation state of signaling proteins to classify cells. Furthermore, additional cellular elements, e.g., biomolecules or molecular complexes, such as RNA, DNA, carbohydrates, metabolites, etc., can be used in conjunction with activation states or expression levels in the classification of cells encompassed herein.
[0213] In some embodiments of the invention, a specific cell population (e.g., CD4 +To analyze the 100% IgG subpopulations (T cells only), various gating strategies can be used. These gating strategies can be based on the presence of one or more specific surface markers. Subsequent gating can distinguish between dead and live cells, and subsequent gating of live cells classifies them into, for example, myeloblasts, monocytes, and lymphocytes. Unequivocal comparisons can be performed by using 2-dimensional contour plot representations, 2-dimensional dot plot representations, and / or histograms. Exemplary gating strategies used for the analysis of patient samples are shown in FIG. 10.
[0214] The immune cells are analyzed for the presence of activated forms of signaling proteins of interest. Signaling proteins of interest include, without limitation, pMAPKAPK2 (pMK2), pP38, prpS6, pNF-κB (p65), IκB, pSTAT3, pSTAT1, pCREB, pSTAT6, pSTAT5, pERK. To determine whether the change is significant, the signal in the patient's reference sample can be compared to a reference scale from a cohort of patients with known outcome.
[0215] Samples may be taken at one or more time points. When a single time point sample is used, a comparison is made to a reference "baseline" level of the characteristic, which may be obtained from a normal control group, a pre-determined level obtained from an individual or a population of individuals, a negative control group for ex vivo activation, etc.
[0216] In some embodiments, the method includes the use of liquid handling components. The liquid handling system may include a robotic system that includes any number of components. In addition, any or all of the steps outlined herein may be automated, thus, for example, the system may be fully or partially automated. See US Patent Application No. 61 / 048,657. As will be understood by those skilled in the art, there are various components that can be used, including, but not limited to, one or more robotic arms; a plate handler for placing microplates; an automated lid or cap handler for removing and replacing the lids of wells on non-cross-contaminated plates; a tip assembly for sample distribution with disposable tips; a washable tip assembly for sample distribution; a 96-well loading block; a cooling reagent rack; a microtiter plate pipette position (optionally cooled); a plate and tip loading tower; and a computer system.
[0217] Fully robotic or microfluidic systems include automated liquid, particle, cell, and bioprocessing, including high throughput pipetting operations to perform all steps of screening applications. This includes liquid, particle, cell, and biomanipulations such as aspiration, dispensing, mixing, dilution, washing, precise volume transfer, removal, and discarding of pipette tips, as well as repetitive pipetting of the same volume for multiple deliveries from a single sample aspiration. These operations are cross-contamination-free liquid, particle, cell, and biomanipulation transfers. The instruments perform automated duplication of microplate samples to filters, membranes, and / or daughter plates, high density transfers, full plate sequential dilutions, and high volume operations.
[0218] In some embodiments, platforms for multi-well plates, multi-tubes, holders, cartridges, mini-tubes, deep well plates, microcentrifuge tubes, cryovials, square well plates, filters, chips, optical fibers, beads and other solid phase matrices, or platforms with various volumes, are housed in a modular platform that can be upgraded for additional volumes. The modular platform includes a variable speed orbital shaker, and a multi-position work deck for source samples, sample and reagent dilutions, assay plates, sample and reagent containers, pipette tips, and active wash stations. In some embodiments, the methods of the invention include the use of a plate reader.
[0219] In some embodiments, interchangeable pipette hands (single or multi-channel) with single or multiple magnetic probes, affinity probes, or pipettors robotically manipulate liquids, particles, cells, and biological organisms. Multi-well or multi-tube magnetic separators or platforms manipulate liquids, particles, cells, and biological organisms in single or multiple sample formats.
[0220] In some embodiments, the instrument includes a detector, which can be a variety of different detectors depending on the label and assay.In some embodiments, useful detectors include: microscopes with multiple fluorescent channels; plate readers that provide spectrophotometric detection of fluorescence, ultraviolet and visible light, with single and dual wavelength endpoints and kinetic capabilities, fluorescence resonance energy transfer (FRET), luminescence, quenching, two-photon excitation and intensity redistribution; CCD cameras that capture data and images and convert them into quantifiable formats; and computer workstations.
[0221] In some embodiments, the robotic device includes a central processing unit that communicates with a memory and a set of input / output devices (e.g., keyboard, mouse, monitor, printer, etc.) through a bus. Again, this may be in addition to or instead of the CPU of the multiplexing device of the present invention, as outlined below. The general interactions between the central processing unit, memory, input / output devices, and bus are known in the art. Thus, depending on the experiment to be performed, a variety of different procedures are stored in the CPU memory.
[0222] The differential presence of these markers is shown to allow for prognostic assessment to detect individuals who have time to onset of labor. In general, such prognostic methods involve determining the presence or level of activated signaling proteins in individual samples of immune cells. Detection can utilize one or a panel of specific binding members, for example, a panel or cocktail of binding members specific for one, two, three, four, five or more markers.
[0223] This invention incorporates information disclosed in other applications and texts. The following patents and other publications are incorporated herein by reference in their entirety: Alberts et al., The Molecular Biology of the Cell, 4th Ed., Garland Science, 2002; Vogelstein and Kinzler, The Genetic Basis of Human Cancer, 2d Ed., McGraw Hill, 2002; Michael, Biochemical Pathways, John Wiley and Sons, 1999; Weinberg, The Biology of Cancer, 2007; Immunobiology, Janeway et al. 7th Ed., Garland, and Leroith and Bondy, Growth Factors and Cytokines in Health and Disease, A Multi Volume Treatise, Volumes 1A and IB, Growth Factors, 1996.
[0224] Unless otherwise clear from the context, any element, step, or feature described herein can be used in any combination with the other elements, steps, or features.
[0225] General methods in molecular and cellular biochemistry can be found in standard textbooks such as: Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998). Reagents, cloning vectors, and kits for genetic manipulation referenced in this disclosure are available from commercial vendors such as BioRad, Stratagene, Invitrogen, Sigma-Aldrich, and ClonTech. Data analysis
[0226] In many embodiments, the method for generating a predictive model of a surgical complication such as SSC uses the MOB algorithm described herein, which integrates multi-omics biological data and / or clinical data. In other embodiments, a predictive model of a surgical complication such as SSC, or a signature pattern associated with a surgical complication such as SSC, can be generated from a biological sample using any convenient protocol, for example as described below. The readout can be the mean, average, median, or variance, or other statistically or mathematically derived value associated with the measurement. The marker readout can be further refined by direct comparison with a corresponding reference or control pattern. The binding pattern can be evaluated for several points to determine whether there is a statistically significant change at any point in the data matrix compared to the reference value, whether the change is an increase or decrease in binding, whether the change is specific to one or more physiological conditions, etc. The absolute values obtained for each marker under identical conditions display the variability that is inherent to a living biological system and also reflect the variability that is inherent between individuals.
[0227] Following obtaining a signature pattern from the sample to be assayed, the signature pattern can be compared to a reference or reference profile to make a prognosis regarding the phenotype of the patient from which the sample was obtained / derived. Additionally, the reference or control signature pattern can be a signature pattern obtained from a sample of a patient known to have a normal pregnancy.
[0228] In certain embodiments, the obtained signature pattern is compared with a single reference / control profile to obtain information about the phenotype of the patient being assayed.In still other embodiments, the obtained signature pattern is compared with two or more different reference / control profiles to obtain more detailed information about the phenotype of the patient.For example, the obtained signature pattern is compared with positive and negative reference profiles to obtain definitive information about whether the patient has the phenotype of interest.
[0229] Samples may be obtained from tissues or bodily fluids of an individual. For example, samples may be obtained from whole blood, tissue biopsies, serum, etc. Other sources of samples are bodily fluids such as lymphatic fluid, cerebrospinal fluid, etc. The term also includes derivatives and fractions of such cells and bodily fluids.
[0230] To identify profiles that indicate responsiveness, statistical tests can provide a confidence level for which changes in the levels of markers between the test and reference profiles are considered significant. The raw data can be first analyzed by measuring values for each marker, usually in duplicate, triplicate, quadruple, or 5-10 replicates per marker. The test data set is considered different from the reference data set if one or more of the profile's parameter values exceed a limit value corresponding to a predefined significance level.
[0231] To provide an ordering of significance, a false discovery rate (FDR) can be determined. First, a set of null distributions of dissimilarity values is generated. In one embodiment, the observed profile values are permuted to create a sequence of distributions of correlation coefficients obtained by chance, thereby creating a suitable set of null distributions of correlation coefficients (see Tusher et al. (2001) PNAS 98, 5116-21, incorporated herein by reference). This analysis algorithm is currently available as a software "plug-in" for Microsoft Excel called Significance Analysis of Microarrays (SAM). The set of null distributions is obtained by permuting the values of each profile for all available profiles, calculating pairwise correlation coefficients for all profiles, calculating the probability density function of the correlation coefficients for this permutation, and repeating this procedure N times, where N is a large number, typically 300. Use the distribution of N to calculate an appropriate measure (mean, median, etc.) of the number of correlation coefficient values that exceed a value obtained from the distribution of empirically observed similarity values at a given significance level.
[0232] The FDR is the ratio of the number of predicted spurious significant correlations (estimated from correlations in the randomized data set that are greater than this selected Pearson correlation) to the number of correlations in the experimental data that are greater than this selected Pearson correlation (significant correlations). This cutoff correlation value can be applied to correlations between experimental profiles.
[0233] In SAM, the Z-score represents another measure of dispersion in a data set, and is equal to the value of X minus the mean of X divided by the standard deviation. The Z-score indicates how a single data point corresponds to a normal data distribution. The Z-score not only demonstrates whether a data point is above or below the mean, but also how unusual the measurement is. The standard deviation is the average distance between each value in a data set and the mean of the values in the data set.
[0234] Using the distributions described above, a confidence level is selected for significance. This is used to determine the lowest value of the correlation coefficient that exceeds the possible outcome obtained by chance. This method is used to obtain a threshold for positive correlation, negative correlation, or both. Using this threshold, the user can filter the pairwise correlation coefficient observations and eliminate those that do not exceed the threshold. Furthermore, an estimate of the false positive rate can be obtained for a given threshold. For each of the individual "random correlation" distributions, it is possible to know how many observations are outside the threshold range. This procedure provides a sequence of counts. The mean and standard deviation of this sequence provide the average number of potential false positives and their standard deviation. Alternatively, any convenient method of statistical validation can be used.
[0235] The data can be subjected to unsupervised hierarchical clustering to reveal relationships between profiles. For example, hierarchical clustering can be performed, where Pearson correlation is used as the clustering index. One approach is to consider a patient-disease dataset as a "training sample" in a "supervised learning" problem. CART is standard in medical applications (Singer(1999) Recursive Partitioning in the Health Sciences, Springer), which converts arbitrary qualitative features into quantitative features and groups them into a set of features similar to Hotellinig's T 2 Sorting by the achieved significance level, assessed by the sample re-use method for statistics, and modification can be made by appropriate application of the Lasso method. A prediction problem can be transformed into a regression problem without losing predictive power by appropriate use of the Gini criterion for classification in assessing the quality of the regression.
[0236] Other analytical methods that may be used include logistic regression. One method of logistic regression is described in Ruczinski (2003) Journal of Computational and Graphical Statistics 12:475-512. Logistic regression is similar to CART in that its classifier can be viewed as a binary tree. It differs in that each node has a Boolean statement about the features that is more general than the simple "and" statements produced by CART.
[0237] Another approach is that of nearest shrunken centroids (Tibshirani (2002) PNAS 99:6567-72). This technique is similar to k-means, but has the advantage of automatically selecting features (similar to Lasso) to focus on a small number of informative features by shrinking cluster centers. This approach is available as a software "plug-in" for Microsoft Excel, the Prediction Analysis of Microarrays (PAM) software, and is widely used. Two further sets of algorithms are Random Forests (Breiman (2001) Machine Learning 45:5-32) and MART (Hastie (2001) The Elements of Statistical Learning, Springer). These two methods are already "committee methods"; thus, they involve predictors that "vote" on the outcome. Some of these methods are based on the "R" software developed at Stanford University, which provides a statistical framework that is constantly being improved and updated on an ongoing basis.
[0238] Other statistical analysis techniques include principal component analysis, recursive partitioning, predictive algorithms, Bayesian networks, and neural networks.
[0239] These tools and methods can be applied to several classification problems, for example, methods can be developed from comparing i) all cases with all controls, ii) all cases with non-responsive controls, iii) all cases with responsive controls.
[0240] In the second analytical approach, the variables selected in the cross-sectional analysis are used separately as predictors. Given the specific outcome, the random length of time each patient is observed, and the choice of proteomic and other features, the parametric approach to analyze reactivity can be better than the widely applied semi-parametric Cox model. The parametric fit of Weibull survival allows the hazard rate to be monotonically increasing, decreasing, or constant, and also has a proportional hazards representation (similar to the Cox model) and a speeded up failure time representation. All standard tools available for obtaining approximate maximum likelihood estimators of regression coefficients and their functions are available with this model.
[0241] In addition, Cox models can be used, especially since the reduction of the number of covariates to a manageable size using the Lasso greatly simplifies the analysis and allows the possibility of a fully non-parametric survival approach.
[0242] The analysis and database storage may be implemented in hardware or software, or a combination of both. In one embodiment of the present invention, a machine-readable storage medium is provided, which includes a data storage material having machine-readable data encoded thereon, which is capable of displaying any of the data sets and data comparisons of the present invention when using a machine programmed with instructions for using said data. Such data may be used for a variety of purposes, such as patient monitoring, early diagnosis, and the like. Preferably, the present invention is implemented as a computer program executed on a programmable computer comprising a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The program code is applied to the input data to perform the functions described above and generate output information. The output information is applied to one or more output devices in a known manner. The computer may be, for example, a personal computer, a microcomputer, or a workstation of conventional design.
[0243] Each program is preferably implemented in a high-level procedural or object-oriented programming language for communicating with a computer system. However, if necessary, the program may be implemented in assembly or machine language. In either case, the language may be a compiled or interpreted language. Each such computer program is preferably stored in a general or special purpose programmable computer-readable storage medium or device so as to configure and operate the computer to perform the procedures described herein when the storage medium or device is read by the computer. The system may also be considered to be implemented as a computer-readable storage medium configured by a computer program, in which case the storage medium so configured operates the computer in a specific, predefined manner to perform the functions described herein.
[0244] Various structural formats for the input and output means can be used to input and output information within the computer-based system of the present invention. One format for the output means is a test data set having different degrees of similarity to the trusted profile. Such a presentation provides a person skilled in the art with a ranking of similarities and identifies the degree of similarity contained in the test patterns.
[0245] The signature patterns and their databases may be provided in various media to facilitate their use. "Media" refers to an article of manufacture containing the signature pattern information of the present invention. The database of the present invention may be recorded on a computer readable medium, e.g., any medium that can be directly read and accessed by a computer. Such media include, but are not limited to, magnetic storage media such as floppy disks, hard disk storage media and magnetic tape, optical storage media such as CD-ROM, electrical storage media such as RAM and ROM, and hybrids of these categories such as magnetic / optical storage media. Those skilled in the art can readily recognize how any of the currently known computer readable media can be used to create an article of manufacture that includes a record of the database information of the present invention. "Recorded" refers to a process for storing information on a computer readable medium using any such method known in the art. Any convenient data storage structure can be selected based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, e.g., word processing text files, database formats, etc. Computer-Implemented Embodiments
[0246] The process of providing a method and system for generating a surgical risk score according to some embodiments is performed by a computing device or computing system, such as a desktop computer, a tablet, a mobile device, a laptop computer, a notebook computer, a server system, and / or any other device capable of performing one or more features, functions, methods, and / or steps described herein. Relevant components within a computing device, which may perform a process according to some embodiments, are shown in FIG. 13. Those skilled in the art will recognize that a computing device or system may include other components omitted for brevity without departing from the described embodiments. A computing device 1300 according to such an embodiment comprises a processor 1302 and at least one memory 1304. The memory 1304 may be a non-volatile memory and / or a volatile memory, and the processor 1302 is a processor, a microprocessor, a controller, or a combination of a processor, a microprocessor, and / or a controller that executes instructions stored in the memory 1304. Such instructions stored in the memory 1304, when executed by the processor, may instruct the processor to perform one or more features, functions, methods, and / or steps described herein. Any input information or data may be stored in memory 1304, either the same memory or a separate memory. According to various other embodiments, computing device 1300 may have hardware and / or firmware that includes instructions and / or is capable of performing those processes.
[0247] Certain embodiments may include a networking device 1306 that allows for communication (wired, wireless, etc.) with another device, such as through a network, near field communication, Bluetooth, infrared, radio frequency, and / or any other suitable communication system. Such a system may be useful for receiving data, information, or input (e.g., omics data and / or clinical data) from another computing device and / or transmitting data, information, or output (e.g., surgical risk scores) to another device.
[0248] Turning to FIG. 14, an embodiment using distributed computing devices is shown. Such an embodiment may be beneficial when computing power is not available at a local level and a central computing device (e.g., a server) performs one or more features, functions, methods, and / or steps described herein. In such an embodiment, a computing device 1402 (e.g., a server) is connected to a network 1404 (wired and / or wireless) that can receive input from one or more computing devices, including clinical data from a record database or repository 1406, omics data provided from a laboratory computing device 1408, and / or any other relevant information from one or more other remote devices 1410. Once the computing device 1402 performs one or more features, functions, methods, and / or steps described herein, any output may be sent to one or more computing devices 1406, 1408, 1410 for inputting into the record, for performing medical actions, including (but not limited to) prehabilitation, delaying surgery, giving antibiotics, and / or any other action related to the surgical risk score. Such actions may be sent directly to a medical professional for such action (e.g., via messaging such as email, SMS, voice / audio alert, etc.) and / or entered into the medical record.
[0249] According to yet other embodiments, the instructions for the processes may be stored on any of a variety of non-transitory computer-readable media suitable for a particular application. Exemplary embodiments
[0250] While the following embodiments provide details regarding certain specific embodiments of the invention, it should be understood that they are merely exemplary in nature and are not intended to limit the scope of the invention. EXAMPLES
[0251] Example 1 Combined plasma and single-cell proteomic analysis of the host immune response to major abdominal surgery Context: In this study, we used an integrative approach combining functional analysis of immune cell subsets using mass cytometry with highly multiplexed assessment of inflammatory plasma proteins to quantify dynamic changes in over 2,388 single cell and plasma proteomic events in patients before and after major abdominal surgery.
[0252] Methods: Forty-one patients undergoing abdominal surgery who met the inclusion criteria were enrolled preoperatively (Table 1, Figure 4). All patients underwent non-cancer major abdominal surgery with bowel resection. The primary outcome was the presence of postoperative surgical site complications (SSC) within 30 days after surgery, including surgical site infection (visceral space, deep, superficial), anastomotic leak, or wound dehiscence. The rationale for combining these three surgical site complications into one primary outcome is that anastomotic leak and wound dehiscence are closely related to the development principle of surgical site infection. Postoperative outcomes were investigated over a 30-day period after surgery. Eleven patients (27%) developed SSC, including superficial surgical site infection, peristomal mucocutaneous separation, and peristomal ulcer. Clinical and operative characteristics of patients who did not and did develop surgical site complications can be seen in Table 1. Patients who developed SSC had significantly higher BMI, operative duration, and estimated intraoperative blood loss.
[0253] For each study participant, blood samples were collected on the day of surgery (DOS, prior to induction of general anesthesia) and on postoperative day 1 (POD1). Blood samples were analyzed using a multimodal approach combining plasma proteomics (i.e., analysis of 274 plasma protein expression levels using the Olink platform (see, e.g., Assarsson E, et al. PLoS One 2014; 9(4):e95192, the disclosure of which is hereby incorporated by reference in its entirety)) and single-cell proteomics (i.e., single-cell analysis of circulating immune cells by mass cytometry, FIG. 4). For the mass cytometry analysis, a 39-parameter immunoassay was used to quantify the frequency and intracellular signaling activity of all major innate and adaptive immune cells. Single-cell analyses were performed using unstimulated blood samples to quantify immune cell subset frequencies and intrinsic signaling activities, as well as samples stimulated with a range of receptor-specific ligands that elicit key intracellular signaling responses involved in the host immune response to trauma / injury, including lipopolysaccharide (LPS), PMA / ionomycin (PI), interleukin (IL)-1β, interferon (IFN)-α, tumor necrosis factor (TNF)α, and combinations of IL-2,4,6.
[0254] To estimate the impact of major abdominal surgery on the human immune system, univariate analyses were performed to compare each plasma or single-cell proteomic feature before and after surgery. Differences between POD1 and DOS were calculated as log-fold changes of plasma proteomic features or as Arcsinh ratios of single-cell proteomic features and visualized as Volcano plots. Plasma and single-cell proteomic features were ranked according to the magnitude of their response to surgery.
[0255] Results: A total of 224 proteomic and 421 mass cytometry features were significantly different (FDR<0.05) after surgery (Figure 11A-B). Specifically, Figure 11A-B show the changes in innate and adaptive immune composition and function in plasma proteomics (Figure 11A) and single-cell mass cytometry (Figure 11B) in response to surgical trauma, depicted as volcano plots. Individual immune features with higher expression in DOS samples are shown on the left (i.e., negative log2 fold change), features with higher expression in POD1 samples are shown on the right (i.e., positive log2 fold change), and features with false discovery rates below 5% are shown above the horizontal dotted line (green points p<0.05, blue points log2FC, or red points both p<0.05 and log2FC). Consistent with previous transcriptomic and mass cytometric analyses of human immune responses to trauma, major abdominal surgery resulted in the simultaneous recruitment of innate and adaptive branches of the human immune system. Specifically, investigation of the top % differentially regulated signatures revealed a major activation of the innate immune response, including increased proinflammatory cytokines such as TNFα, IL-6, and members of the IL-1 superfamily, increased chemotactic proteins (including CCL23 and CX3Cl1), and increased canonical inflammatory signaling responses (such as JAK / STAT signaling) in innate myeloid cell subsets. Conversely, adaptive immune cell frequencies (including CD4+ and CD8+ T cell subsets), as well as adaptive immune responses to inflammatory stimuli (especially JAK / STAT signaling responses to IL2 / 4 / 6 stimulation) and concentrations of regulatory proteins (such as IL-10RA) were decreased on POD1 compared to DOS. We also observed a robust increase in the frequency and JAK / STAT signaling activity of monocytic myeloid-derived suppressor cells (M-MDSCs), a population of innate immune cells with immunosuppressive properties that accumulate in the context of malignancy, sepsis, and severe trauma, including surgery.
[0256] Conclusions: Overall, differential immune profiling of patients pre- and 24 hours post-surgery demonstrated that major abdominal surgery induces a complex inflammatory response involving both pro-inflammatory and immunosuppressive components of the innate and adaptive immune systems. Importantly, there was significant inter-patient variability in the magnitude of this immune response, which prompted further investigation into whether it reflects patient-specific differences that may predetermine the development of surgical complications.
[0257] Example 2 Predicting surgical site complications (SSCs) with integrated modeling of preoperative multi-omics biological and clinical data - Study 1 Background: Differential analysis of immune responses on POD1 versus DOS (Example 1) highlights biological aspects of the human immune response to trauma that may govern the development of SSC. However, the ability to identify preoperatively (i.e., to DOS) which patients will develop SSC is of greatest clinical interest, as it would allow preoperative risk stratification and personalization of preoperative interventions.
[0258] Methods: Figure 12 shows patient enrollment according to the CONSORT criteria used in this study. 41 patients were prospectively enrolled in Study 1, 11 patients developed SSC within 30 days of surgery, and 30 patients did not. Whole blood samples collected before incision on the day of surgery (DOS) and on postoperative day 1 (POD1) were stimulated with lipopolysaccharide (LPS), tumor necrosis factor (TNF)α, interleukin (IL)-2,4,6 cocktail, PMA / ionomycin (P / I), interferon (IFN)α, IL-1β, or left unstimulated ("unstimulated"). Whole blood samples were analyzed using a 47-parameter single-cell mass cytometry assay to quantify the abundance of all major innate and adaptive immune cell subsets and the intracellular activity of single cells of key signaling responses involved in the immune response to surgical trauma. Plasma samples were analyzed using the Olink multiplex proteomics platform (Study 1, 274 proteins were analyzed). Table 5 provides a list of the antibody panels used in this study.
[0259] To determine whether the immune status of patients with and without SSC differs before surgery, we applied an integrated multi-omics bootstrap (MOB) analysis pipeline (Figures 2-3) to the DOS immune dataset (derived from samples collected before induction of anesthesia and surgical incision). The method leverages the interconnectedness and multi-layered nature of the combined plasma and single-cell proteomics dataset, resulting in an integrated feature selection framework with robustness-based selection. The dataset contained nine unique data layers: immune cell frequency (containing 24 cell frequency features), basal signaling activity of each cell subset (312 basal signaling features), signaling response capacity to each stimulation condition (six data layers each containing 312 features), and plasma proteomics (276 proteomic features) (Figures 2, 4). The method integrates the nine data layers using several steps. First, at each layer, artificial features are introduced by replacing the original features, thus creating features that are not related to the outcome. Then, a bootstrap procedure is performed multiple times, in which the fitting of a machine learning model is repeated by resampling from this dataset with or without replacement. Typically, the machine learning models used are logistic or linear regression with L1 or L2 regularization, commonly denoted as Lasso, Ridge, or Elastic Net models. The repetition of this procedure allows the estimation of the distribution of the simulated noise and allows the description of its distribution. For each variable (artificial or not), a stability path is calculated, which is defined as the frequency of selection in the model from the non-zero features or the features with the highest importance in the model (e.g. the absolute value of the coefficient is the largest). The optimal cutoff of the biological or clinical features is selected using the distribution of the artificial features used to estimate the behavior of the noise on the robustness of the biological or clinical features from that data layer.
[0260] Multi-omics biological features utilized for MOB analysis were defined as follows: Single-cell proteomic features: 2,116 single-cell proteomic features, including cell frequency, intrinsic signaling, and signaling response to ex vivo stimulation, were derived from mass cytometry data as previously described. For each immune cell subset, immune cell frequency features were calculated from unstimulated samples. Mononuclear cell frequency was calculated as the percentage of viable singlet mononuclear cells (cPAPRs) and the percentage of viable singlet mononuclear cells (cPAPRs). - CD45 + CD66 - Granulocyte frequency was determined as a percentage of gated live singlet cells (cPARP - ) was determined as a percentage of the total intracellular signaling response. For single-cell signaling features, median expression of intracellular signaling proteomic markers was simultaneously quantified per cell for phospho(p)STAT-1, pSTAT-3, pSTAT4, pSTAT5, pSTAT6, pNfκB, total IκBα, pMAPKAPK2 (pMK2), pERK1 / 2, prpS6, pCREB, Ki67, and PD-1. Endogenous signaling activity was expressed as arcsinh-transformed values from unstimulated samples. Signaling responses to ex vivo stimulation were reported as the arcsinh-transformed median difference of stimulated values from endogenous values (asinh ratios). As previously described, a knowledge-based penalty matrix was applied to intracellular signaling response features in mass cytometry data based on mechanistic immunological knowledge. (See, e.g., N. Aghaeepour et al (2017). Sci Immunol 2, the disclosure of which is hereby incorporated by reference in its entirety.) Importantly, the mechanistic prior knowledge used in the penalty matrix is independent of immunological knowledge relevant to postoperative recovery. Plasma proteomic features were quantified using the Olink immune response panel, an inflammatory panel, and a metabolic panel to quantify concentrations of 272 unique plasma proteins. Relative levels of plasma proteins are reported in arbitrary units calculated from data normalized to internal control groups and are reported after log2 transformation.
[0261] Results: A robust MOB model was constructed that accurately distinguished patients with and without SSC (AUC=0.82, 95% CI [0.66-0.94], unpaired Mann-Whitney rank sum test on cross-validated values of the MOB model, Figure 5B). The predictive performance of the MOB model was superior to existing surgical outcome prediction models, such as the ACS NSQIP risk assessment score (see, e.g., Bilimoria KY et al. J Am Coll Surg 2013; 217(5):833-42 e1-3, the disclosure of which is incorporated herein by reference in its entirety) based on clinical variables (ACS AUC=0.73). Confounding factor analysis, including clinical and demographic variables that differed between the two patient groups, showed that the MOB model captured much more information when accounting for differences in age, BMI, preoperative diagnostic characteristics, and surgery type (Table 6). Comparison of generalized linear models with or without MOB prediction resulted in a much better fit for models using MOB values (p=8e-05, chi-squared test for deviation between fits). Finally, integration of preoperative clinical variables (i.e., age, sex, BMI, functional status, emergency case, American Society of Anesthesiologists (ASA) class, steroid use for chronic conditions, ascites, multiple cancers, diabetes, hypertension, congestive heart failure, dyspnea, smoking history, history of severe COPD, dialysis, acute renal failure) into single cell and plasma proteomic variables collected in DOS further improved the accuracy of the DOS model for predicting SSC (AUC=0.92, 95% CI[0.84-0.99], Figure 5C).
[0262] Conclusions: Taken together, these results suggest that integration of immune and clinical information collected preoperatively has strong potential to accurately identify patients at risk for postoperative SSC. The predictive performance of the MOB model suggests that sufficiently powerful predictive models can be developed to risk stratify individual patients and assign them to patient-specific treatment pathways aimed at reducing their risk of developing SSC.
[0263] Example 3 Predicting SSC with integrated modeling of preoperative multi-omics biological data - Study 2 Background: Results of prospective study 1 demonstrate that accurate risk estimates for the development of SSC can be derived from analysis of patients' immune status before surgery. However, clinical and demographic variables influence patients' immune status and act as confounders for the development of SSC. To determine the contribution of patients' preoperative immune status to the development of SSC, we performed a retrospective study (study 2) comparing two groups of patients undergoing major abdominal surgery, matched on key clinical and demographic variables. The primary outcome of this study was the development of SSC within 30 days of surgery.
[0264] Methods: From a larger cohort of 450 patients contained in the Stanford Surgical Biobank, 93 patients undergoing major abdominal surgery at Stanford Hospital were selected (Table 2). Sixteen patients developed SSC (cases) and 77 patients did not (controls). Cases and controls were matched using a frequency matching algorithm that ensured equal distribution between groups of clinical and demographic variables: age, sex, BMI, smoking history, surgical approach, and perioperative treatment therapy. Blood and plasma samples collected preoperatively at DOS were processed as described in Study 1 and analyzed using a multiomics combination of mass cytometry and multiplexed plasma proteomics. The plasma proteomics platform used in Study 2 is the aptamer-based platform Somalogic, which allows for the quantification of over 2400 circulating proteins (see, e.g., L. Gold et al., PloS one 5, e15004, 2010, the disclosure of which is incorporated herein by reference in its entirety). We applied the MOB predictive modeling pipeline to build a predictive model to distinguish patients with and without SSC.
[0265] Results: Application of the MOB method to the combined DOS dataset of mass cytometry and plasma proteomics collected preoperatively identified a multivariate model that classified patients who developed SSC from controls with high accuracy (AUC=0.77, 95% CI [0.66-0.89], unpaired Mann-Whitney rank sum test on cross-validated values of the MOB model, Figure 6).
[0266] Conclusions: The results of this independent retrospective study of an additional 93 patients confirm previous findings (Study 1) and suggest that pooled analysis of preoperative immune data using MOB can identify patients at risk of developing SSC after surgery. In addition, results obtained using data from a retrospective cohort of matched cases and controls suggest that a patient's preoperative immune status distinguishes patients at risk of developing SSC, independent of key clinical and demographic variables that may be associated with SSC.
[0267] Example 4 Integrated modeling of immune responses 24 hours after surgery accurately classifies patients with postoperative SSC Context: This study used an integrated predictive modeling approach to determine whether immune responses detectable on POD1, 24 hours after surgery, could distinguish patients who developed SSC at that time from those who had an uncomplicated postoperative recovery.
[0268] Methods: Peripheral blood and plasma samples were collected on POD1 after abdominal surgery from patients enrolled in Study 1 (Figure 4, Figure 7, Table 1). Samples were analyzed using a multi-omics combination of mass cytometry (to analyze immune cell frequencies and intracellular signaling responses) and plasma proteomics as described in Example 1. Predictive modeling of SSC was performed using the MOB pipeline.
[0269] Results: The predictive MOB model built on the POD1 immune dataset performed very well in classifying patients who developed SSC (AUC=0.86, p=2.48e-04, Mann-Whitney nonparametric unpaired test on cross-validated MD predictors, Figure 7). To account for confounding clinical and demographic variables, a post-hoc confounder analysis was performed on the cross-validated predictors of the model. Comparison of the generalized linear model with or without the MOB predictor resulted in a much better fit of the model using SG values (p=2e-07, chi-square test for deviation between fits, Table 7). In addition, when assessed with confounders one at a time in the linear model, the SG model still showed to be highly predictive of SSC when accounting for patient variability in either age, sex, surgery type, preoperative diagnosis, or length of surgery.
[0270] Conclusions: Analysis of the immune response to surgery on POD1 accurately classifies patients who developed SSC from those who did not, thereby identifying a predictive model that highlights biological differences in the response to trauma that may govern the development of SSC. Identifying a predictive MOB model of SSC on POD1 prior to its initiation is clinically relevant as it allows for preemptive intervention to prevent SSC.
[0271] Example 5 Single-cell immune response and plasma proteomic biological signatures contribute to an integrated predictive model of SSC Context:The multivariate MOB prediction pipeline yielded statistically robust models that accurately classified patients with and without SSC from analysis of biological and clinical data obtained before surgery (DOS model) or shortly after surgery (POD1 model).To understand the biological relevance of the high-dimensional MOB model, we investigated in more detail the individual MOB features that contributed most to the multivariate model.
[0272] Methods: Using an iterative "bootstrapping" procedure (i.e., 1000 iterations of resampling the data with replacement), individual MOB model features were ranked according to their relative contribution to the multivariate MOB model (Figures 2, 3A, and 3B). Features were ranked using an objective relative model contribution index (MCI) to objectively select the most informative single-cell immune response and plasma proteomic features (MCI [feature] > MCI [decoy feature]).
[0273] Results: Application of the iterative bootstrap MOB procedure to the multi-omics biological data obtained before surgery (Study 1 and Study 2) selected 55 features that contributed most to the multivariate MOB model (Figures 8A-8D, Table 3). Specifically, Figure 8A shows the informative DOS MOD model single cell immune features selected from the plasma proteomics data layer, Figure 8B shows the informative DOS MOD model single cell immune features selected from the LPS data layer, Figure 8C shows the informative DOS MOD model single cell immune features selected from the IL2 / 4 / 6 data layer, and Figure 8D shows the informative DOS MOD model single cell immune features selected from the TNFα data layer. The graphs on the left represent the probability of an individual feature being selected from the real or decoy dataset at each bootstrap iteration. The box and whisker plots on the right show examples of the most informative features for each single cell data layer. A list of informative features of the MOD model is provided in Table 3 (DOS model) and Table 4 (POD1 model).
[0274] Plasma proteomic signatures included 12 plasma proteins that were increased (IL-1β, ALK, WWOX, HSPH1, IRF6, CTNNA3, CCL3, sTREM1, ITM2A, TGFα, LIF, ADA) and 4 plasma proteins that were decreased (ITGB3, EIF5A, KRT19, NTproBNP) in patients who subsequently developed SSC. Single-cell immune response signatures distinguished patients who subsequently developed SSC from controls with 4 LPS response signatures (increased pMAPKAPK2 (pMK2) in neutrophils, increased prpS6 in mDCs, and decreased IκB and CD7 in neutrophils). + CD56 hi CD16 lo reduced pNFκB in NK cells), 9 IL-2 / IL-4 / IL-6 response signatures (increased pSTAT3, CD56 in neutrophils, mDCs, or Tregs) hi CD16 lo Increased prpS6 in NK cells or mDCs, increased pSTAT5 in mDCs or pDCs, and CD4 + Tbet + reduced IκB in Th1 cells, reduced pSTAT1 in pDCs), 11 TNFα response features (increased prpS6 in neutrophils or mDCs, increased pERK in M-MDSCs or ncMCs, increased pCREB in γδ T cells, or reduced IκB, pP38 or pERK, or CD4 in neutrophils). + Tbet + Decreased pCREB or pMAPKAPK2 in Th1 cells, or CD4 + CRTH2 + reduced pERK in Th2 cells), 10 non-stimulatory features (increased pSTAT3 in neutrophils, M-MDSCs, cMCs, or ncMCs, Tregs, or CD45RA - memory cd4 + Increased pSTAT5 in T cells, increased pMAPKAPK2 in mDCs, and CD4 + Tbet + pCREB or IκB in Th1 cells, increased pSTAT6 in NKT cells, or CD4 + Tbet +decreased pERK in Th1 cells), as well as five frequency characteristics (increased M-MDSC, G-MDSC, ncMC, Th17 cells, or decreased CD4 + CRTH2 + Th2 cells).
[0275] Application of the MOB procedure to multi-omics biological data acquired 24 hours after surgery (Study 1) selected 16 features that contributed most to the multivariate POD1 model (Figures 9A-9N, Table 4). Specifically, Figures 9A-9G show single cell immune response features that contribute to the POD1 MOB prediction model of SSC, while Figures 9H-9N show plasma proteomic features that contribute to the POD1 MOB prediction model of SSC.
[0276] Conclusions: Analysis of plasma-based and single-cell immune events before and shortly after surgery provided a systems-level view of immune mechanisms implicated in trauma associated with the development of SSC. Two major themes emerged characterizing the early immune response to surgery in patients who subsequently developed SCC: 1) intensified pro-inflammatory IL-6R and TLR-related signaling responses, and 2) an increased immunosuppressive cellular response, including M-MDSC and Treg responses.
[0277] Key elements of the POD1 SG model nicely integrate with previous knowledge of immune mechanisms predisposing to SSC. Previous reports have shown that elevated IL-6 plasma concentrations early after surgery correlate with increased risk of postoperative complications, including infection. Consistent with previous findings, increased STAT3 signaling activity in cMCs (canonically activated by IL-6) was one of the most informative single-cell features associated with SSC. Similarly, the intensified MyD88 signaling response to LPS in innate myeloid cells in patients who later developed SSC reiterates previous results, which indicate that unchecked systemic activation of proinflammatory innate immune cells in response to surgical site injury may contribute to the development of SSC. Thus, excessive local immune responses to inflammation may increase the release of DAMPs and PAMPs from the surgical site in a cycle of enhanced MyD88-associated TLR signaling, triggering barrier disruption, and additional tissue damage. In this context, it is also noteworthy that overstimulation of TLR signaling can result in a state of endotoxin tolerance, which can increase the patient's susceptibility to infection.
[0278] The single-cell resolution afforded by mass cytometry has provided new insights into cell type-specific responses that may contribute to the pathogenesis of SSCs. Increased STAT3 signaling in M-MDSCs and increased M-MDSC frequency 24 hours after surgery were among the most informative features of the POD1 model. This result is in line with previous studies of patients undergoing orthopedic surgery showing a strong correlation between STAT3 signaling in MDSCs and delayed postoperative recovery. MDSCs are a heterogeneous subset of immature myeloid cells with immunosuppressive functions that are recruited in the context of acute and chronic inflammatory diseases. Previous investigations of the immune response to trauma and sepsis have shown that MDSCs mediate an anti-inflammatory program that suppresses the adaptive immune system, particularly antigen-specific CD8 + and CD4 +In patients who subsequently developed SSC, elevated STAT3 signaling, which is required for MDSC proliferation and immunosuppressive function, may synergistically promote MDSC expansion, thus further exacerbating the state of immunosuppression.
[0279] We also observed an upregulation of endogenous STAT5 signaling in immunosuppressive Tregs in patients who developed SSC. In contrast, pSTAT5 responses to ex vivo stimulation with IL-2 / 4 / 6 were lower in patients who developed SSC, which may indicate that a higher endogenous pSTAT5 signaling tone may prevent further ex vivo activation. IL-2R-dependent activation of STAT5 in Tregs is essential for mature Tregs to maintain FoxP3 expression levels and exert their immunosuppressive functions. Reportedly, FoxP3 expression and Treg lineage-specific transcription are further promoted by the IL-6 family cytokine LIF. The regulatory function of LIF in inducing Treg development and maturation indicates an ambivalent role for IL-6 family cytokines in the context of inflammation and trauma. Overall, excessive endogenous Treg signaling may act synergistically with the observed enhanced MDSC responses to induce a persistent immunosuppressive state that blunts responses to invading pathogens in patients who develop SSC.
[0280] Whereas the POD1 model provided important information regarding surgery-induced mechanisms implicated in the pathogenesis of SSC, the DOS SG model showed single cell features and plasma proteomic factors that distinguished the two patient groups pre-surgery. The most informative features of the DOS SG model were the proteomic features IL-1β, sTREM1 and ITM2A. Our results showing that sTREM1 is elevated at DOS and POD1 in patients who later develop SSC recall previous studies showing increased sTREM1 plasma concentrations in patients with bacterial infection and sepsis. From a mechanistic perspective, sTREM1 is a metalloprotease cleavage product of membrane-bound TREM1, an amplifier of pattern recognition receptors on myeloid cells. sTREM1 can function as a decoy receptor that antagonizes TREM1. However, microbial products such as LPS can increase the membrane expression of TREM1 as well as stimulate the release of sTREM1, thereby increasing sTREM1 plasma concentrations. Whether elevated sTREM1 in patients with SSC is consistent with TREM1 expression on myeloid cells or results in functional suppression of TREM1 is an important question that warrants further investigation.
[0281] Another proteomic feature of the DOS model, ITM2A, is upregulated by PKA-CREB signaling, leading to the accumulation of autophagosomes and the suppression of autolysosome formation. Effective autophagy is essential for many physiological functions, including tissue differentiation, cell cycle regulation, and immune cell maturation, especially Th cell development. Other informative features of the DOS model included differences between multiple innate and adaptive cell subsets, such as neutrophils, pDCs, and Th2 cells. Notably, in patients who developed SSC, the signaling responses to multiple stimuli (including IL-1β, TNFα, and IFNα) were upregulated by CRTH2. + Th2-like CD4 +T cells, which play a key role in protective immunity against extracellular pathogens and tissue repair. Our results suggest that a patient's intrinsic immune status before surgery may increase the risk of developing SSC. Therefore, preoperative assessment of specific immune markers may help risk stratify patients along with applying interventions to mitigate the risk of developing SSC.
[0282] Doctrine of Equivalents Although several embodiments have been described, it will be recognized by those skilled in the art that various modifications, alternative configurations, and equivalents may be used without departing from the spirit of the present invention. Additionally, some well-known processes and elements have not been described to avoid unnecessarily obscuring the present invention. Therefore, the above description should not be construed as limiting the scope of the present invention.
[0283] Those skilled in the art will appreciate that the above examples and descriptions of various preferred embodiments of the present invention are merely illustrative of the entire invention, and that variations in the components or steps of the present invention may be made within the spirit and scope of the present invention. Accordingly, the present invention is not limited to the specific embodiments described herein, but rather is defined by the appended claims. [Table 1-1] [Table 1-2] [Table 2] [Table 3-1] [Table 3-2] [Table 4] [Table 5]
Table 6
Table 7
Claims
1. A method for analyzing values of a plurality of features as indicators of risk of surgical complications for an individual following surgery, comprising: obtaining or having obtained values of the plurality of features, the plurality of features including omics biological features and clinical features; calculating a surgical risk score for the individual based on the plurality of features using a model obtained via machine learning techniques; providing an assessment of the patient's risk of developing a surgical complication based on the calculated surgical risk score; a machine learning model is trained using a bootstrap procedure on a plurality of individual data layers, each data layer representing one type of data from the plurality of features and at least one artificial feature; each data layer contains data about a population of individuals; each feature comprises a feature value for every individual in said population of individuals; for each data layer, each artificial feature is obtained from a non-artificial feature of the plurality of features via a mathematical operation performed on feature values of the non-artificial feature; The method, wherein the mathematical operation is selected from the group consisting of replacement, sampling with replacement, sampling without replacement, combination, knock-off, and inference.
2. The step of obtaining or having obtained values of a plurality of features, obtaining or having obtained a sample for analysis from the individual undergoing surgery; measuring or having measured values of a plurality of omics biological and clinical features; Including, the samples preferably include at least one sample obtained pre-operatively; The sample is more preferably obtained any time before surgery until the day of surgery, before the surgical incision is made; the samples preferably include at least one sample obtained post-operatively; More preferably, the post-operative sample is obtained approximately 24 hours after surgery; the sample is preferably a blood sample, a peripheral blood mononuclear cell (PBMC) fraction of a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, or dissociated cells from a tissue sample; 10. The method of claim 1, wherein the sample is contacted, preferably ex vivo, with an effective amount of an activating agent for a period of time sufficient to activate immune cells in the sample.
3. The omics biological features include at least one feature from the group consisting of genomics features, transcriptomics features, proteomics features, cytomics features, and metabolomics features; the cytomics signature preferably comprises surface and intracellular proteins at the single cell level in immune cell subsets; said proteomic signature preferably comprises circulating extracellular proteins; the plurality of characteristics preferably further comprises demographic characteristics; 3. The method of claim 2, wherein the demographic or clinical characteristics preferably comprise data selected from the group consisting of age, sex, body mass index (BMI), functional status, emergency case, American Society of Anesthesiologists (ASA) class, steroid use for chronic conditions, ascites, multiple cancers, diabetes, hypertension, congestive heart failure, dyspnea, smoking history, history of severe COPD, dialysis, and acute renal failure.
4. A method according to any one of claims 1 to 3, wherein each type of data is selected from the group consisting of genomics, transcriptomics, proteomics, cytomics, metabolomics, clinical, and demographics.
5. The model includes weights (β i ) for a set of selected biological and clinical or demographic features; For each data layer during machine learning, at each iteration of the bootstrap, initial weights (w j ) are calculated for the plurality of features and the at least one artificial feature associated with that data layer using an initial statistical learning technique, and at least one selected feature is determined for each data layer based on a statistical criterion that depends on the calculated initial weights (w j ); said initial statistical learning technique is preferably selected from regression and classification techniques; The initial statistical learning technique is further preferably selected from the Lasso technique and the Elastic Net technique; the statistical criterion preferably depends on the effective weights among the calculated initial weights (w j ); said statistical criteria are preferably based on the frequency of occurrence of said valid weights; the statistical criterion is further preferably such that each feature is selected if its frequency of occurrence is greater than a frequency threshold, the frequency threshold being calculated according to the frequency of occurrence obtained for the artificial features; If the initial statistical learning technique is a sparse regression technique, the effective weights are further preferably non-zero weights; The method according to any of claims 1 to 4, wherein if the initial statistical learning technique is a non-sparse regression technique, the effective weights are more preferably weights that are above a predefined weight threshold.
6. The initial weights (w j ) are further calculated for multiple values of hyperparameters, the hyperparameters being parameters whose values are used to control the learning process; said hyperparameters are regularization coefficients preferably used according to their respective mathematical norms in the context of sparse initial techniques, the mathematical norm is more preferably a p-norm, where p is an integer; If the initial statistical learning technique is the Lasso technique, the hyperparameter is preferably an upper bound on the coefficient of the L1 norm of the initial weights (w j ), the L1 norm being the sum of all absolute values of the initial weights; When the initial statistical learning technique is the Elastic Net technique, the hyperparameter is preferably an upper limit of the coefficients of both the sum of the L1 norm of the initial weights (w j ) and the sum of the L2 norm of the initial weights (w j ), where the L1 norm refers to the sum of the absolute values of all the initial weights, and the L2 norm refers to the square root of the sum of the square values of all the initial weights. The method of claim 5.
7. For each feature, a unit occurrence frequency is calculated for each hyperparameter value, said unit occurrence frequency equal to the number of valid weights for said feature for successive bootstrap iterations divided by the number of bootstrap iterations; The occurrence frequency is preferably equal to the highest unit occurrence frequency among unit occurrence frequencies calculated for a plurality of the hyperparameter values. The method of claim 6.
8. During the machine learning, the weights (β i ) of the model are further calculated using a final statistical learning technique on data associated with the selected set of features; said final statistical learning technique is preferably selected from regression and classification techniques; The method according to any of claims 5 to 7, wherein the final statistical learning technique is further preferably selected from the Lasso technique and the Elastic Net technique.
9. The method of claim 8, wherein during a use phase following the machine learning, the surgical risk score is calculated according to the individual's measurements for the selected set of features; If the final statistical learning technique is a classification technique, the surgical risk score is preferably a probability calculated according to a weighted sum of the measurements multiplied by their respective weights (β i ) for the set of selected features; The surgical risk score is more preferably calculated according to the formula [Equation 20] is calculated according to where P represents the surgical risk score, Odd is a term that depends on the weighted sum, and Odd is preferably an exponent of the weighted sum; If the final statistical learning technique is the regression technique, the surgical risk score is preferably a term that depends on a weighted sum of the measurements multiplied by the respective weights (β i ) for the set of selected features; The surgical risk score is more preferably equal to the index of the weighted sum. The method according to any one of claims 5 to 8.
10. During the machine learning, before obtaining the artificial features, the method further comprises: generating additional values of the plurality of non-artificial features based on the obtained values using data augmentation techniques; The method according to any one of claims 4 to 9, wherein the artificial features are then obtained according to both the obtained values and the generated additional values.
11. The step of measuring or having measured said value comprises measuring said surface or intracellular protein at the single cell level in an immune cell subset by contacting said sample with an isotopically or fluorescently labeled affinity reagent specific for said surface or intracellular protein; The surface or intracellular proteins at the single cell level in immune cell subsets are preferably measured by flow cytometry or mass cytometry; The method of any one of claims 1 to 10 in combination with claim 2.
12. The method of claim 1, wherein the step of measuring or having measured said values comprises analyzing circulating proteins by contacting said sample with a plurality of isotopically or fluorescently labeled affinity reagents specific for extracellular proteins; The affinity reagent is preferably an antibody or an aptamer. The method of any one of claims 1 to 11 in combination with claim 2.
13. The surgical complication is a surgical site complication (SSC), the patient's risk of developing a surgical site complication is correlated with increased pMAPKAPK2 (pMK2) in neutrophils, increased prpS6 in mDCs, or decreased IκB in neutrophils, decreased pNFκB in CD7 + CD56 hi CD16 lo NK cells, preferably in response to ex vivo activation with LPS in samples collected before surgery; the patient's risk of developing a surgical site complication is correlated with increased pSTAT3 in neutrophils, mDCs, or Tregs, increased prpS6 in CD56 hi CD16 lo NK cells or mDCs, increased pSTAT5 in mDCs or pDCs, or decreased IκB in CD4 + Tbet + Th1 cells, decreased pSTAT1 in pDCs, in response to ex vivo activation with IL-2, IL-4, and / or IL-6, preferably in a sample collected before surgery; the patient's risk of developing a surgical site complication correlates with increased prpS6 in neutrophils or mDCs, increased pERK in M-MDSCs or ncMCs, increased pCREB in γδT cells, or decreased IκB, pP38, or pERK in neutrophils, or decreased pCREB or pMAPKAPK2 in CD4 + Tbet + Th1 cells, or decreased pERK in CD4 + CRTH2 + Th2 cells, preferably in response to ex vivo activation with TNFα in samples collected before surgery; the patient's risk of developing a surgical site complication is correlated with increased pSTAT3 in neutrophils, M-MDSCs, cMCs, or ncMCs, increased pSTAT5 in Treg or CD45RA − memory CD4 + T cells, increased pMAPKAPK2 in mDCs, pCREB or IκB in CD4 + Tbet + Th1 cells, increased pSTAT6 in NKT cells, or decreased pERK in CD4 + Tbet + Th1 cells, preferably in unstimulated samples collected before and / or after surgery; the patient's risk of developing a surgical site complication is correlated with increased M-MDSC, G-MDSC, ncMC, Th17 cell, or decreased CD4 + CRTH2 + Th2 cell frequency, preferably collected pre- and / or post-operatively; the patient's risk of developing a surgical site complication is correlated with increased IL-1β, ALK, WWOX, HSPH1, IRF6, CTNNA3, CCL3, sTREM1, ITM2A, TGFα, LIF, ADA, or decreased ITGB3, EIF5A, KRT19, NTproBNP, preferably collected pre- and / or post-operatively; The method according to any one of claims 1 to 12.
14. A method according to any one of claims 1 to 13 in combination with claim 2, wherein the step of measuring or having measured the value comprises contacting the sample ex vivo with an effective amount of an activator for a period of time sufficient to activate immune cells in the sample, wherein the activator is one or a combination of a TLR4 agonist (such as LPS), interleukin (IL)-2, IL-4, IL-6, IL-1β, TNFα, IFNα, PMA / ionomycin.
15. The method of claim 1, wherein the step of measuring or having measured said value comprises measuring a surface or intracellular protein at the single cell level in an immune cell subset by contacting said sample with an isotopically or fluorescently labeled affinity reagent specific for said surface or intracellular protein; the immune cells are identified using a single cell surface or intracellular protein marker, preferably selected from the group consisting of CD235ab, CD61, CD45, CD66, CD7, CD19, CD45RA, CD11b, CD4, CD8, CD11c, CD123, TCRγδ, CD24, CD161, CD33, CD16, CD25, CD3, CD27, CD15, CCR2, OLMF4, HLA-DR, CD14, CD56, CRTH2, CCR2, and CXCR4; the intracellular protein of the single cell is preferably selected from the group consisting of phospho(p)pMAPKAPK2 (pMK2), pP38, pERK1 / 2, p-rpS6, pNFκB, IκB, p-CREB, pSTAT1, pSTAT5, pSTAT3, pSTAT6, cPARP, FoxP3, and Tbet; The intracellular protein level is preferably selected from the group consisting of neutrophils, granulocytes, basophils, CXCR4 + neutrophils, OLMF4 + neutrophils, CD14 + CD16 − classical monocytes (cMC), CD14 − CD16 + non-classical monocytes (ncMC), CD14 + CD16 + intermediate monocytes (iMC), HLADR + CD11c + myeloid dendritic cells (mDC), HLADR + CD123 + plasmacytoid dendritic cells (pDC), CD14 + HLADR − CD11b + monocytic myeloid-derived suppressor cells (M-MDSC), CD3 + CD56 + NK-T cells, CD7 + CD19 − CD3 − NK cells, CD7 + CD56 lo CD16 hi NK cells, CD7 + CD56 hi CD16 lo NK cells, CD19 + B cells, CD19 + CD38 + plasma cells, CD19 + CD38- non-plasma B cells, CD4 + CD45RA + naive T cells, CD4 + CD45RA- memory T cells, CD4 + CD161 + Th17 cells, CD4 + Tbet + Th1 cells, CD4 + CRTH2 + Th2 cells, CD3 + TCRγδ + γδT cells, Th17 CD4 + T cells, CD3 + FoxP3 + CD25 + Regulatory T cells (Tregs), CD8 + CD45RA + naive T cells, and CD8 + CD45RA- measured in an immune cell subset selected from the group consisting of memory T cells; 15. The method of claim 13 or 14.
16. A system comprising a processor and a memory containing instructions which, when executed by the processor, instruct the processor to perform any of the methods of claims 1, 3-10, and 13.
17. A non-transitory machine-readable medium containing instructions which, when executed by a computer processor, instruct the processor to perform any of the methods of claims 1, 3-10, and 13.