A protein biomarker combination, kit and application for early screening of colorectal cancer
By detecting the combination of protein markers in plasma and combining machine learning models, the problem of insufficient sensitivity and specificity in early screening of colorectal cancer is solved, and efficient non-invasive screening and accurate diagnosis are achieved.
Patent Information
- Application Number
- CN202410576103.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-01
AI Technical Summary
The prior art has insufficient sensitivity and specificity in early screening of colorectal cancer, which is difficult to meet the screening needs of large-scale risk groups, and the accuracy of non-invasive detection methods is limited.
The protein marker combination, including LRG1, SERPINA1, ITIH3, CP, ORM1, C9, IGFBP2 and CNDP1, was used to detect the protein expression level in plasma by mass spectrometry, and combine machine learning models to predict, diagnose and prognosis of colorectal cancer.
It provides a non-invasive screening method with high sensitivity and high specificity, which can accurately predict the risk of colorectal cancer, assist in diagnosis and evaluation of treatment prognosis, reduce the rate of missed diagnosis, and improve detection accuracy.
Smart Images

Figure CN118518886B_ABST
Abstract
Description
[0001] Related patents
[0002] This application is a divisional application of a Chinese invention patent application with the application number 2023100498923, the application date of February 1, 2023, and the invention title of "A Protein Marker, Kit and Application for Early Screening of Colorectal Cancer". Technical field
[0003] The present invention belongs to the technical field of cancer proteomics detection. Specifically, it relates to a protein marker combination, kit and application for early screening of colorectal cancer. Background technique
[0004] Currently, the main clinical methods for colorectal cancer screening include colonoscopy, imaging examination, fecal occult blood test, DNA detection, protein marker detection such as CEA, etc. These conventional techniques are invasive or cause radiation damage. More importantly, their sensitivity is not high, making it difficult to be used for early screening of a large number of at-risk populations. Moreover, the tolerance and acceptance of colonoscopy among the general population are also relatively low. The only non-invasive detection method applied clinically is the chemical and immunological detection of fecal occult blood. However, the sensitivity of this type of detection is only 61-79% on the premise of 86-95% specificity for colorectal cancer. Although it is widely used clinically, the detection rate of early colorectal cancer is difficult to meet the clinical needs.
[0005] In recent years, liquid biopsy technology has developed rapidly, which has solved the problem of relatively low sensitivity of traditional detection techniques to a certain extent. For example, products for detecting the methylation of the Septin9 gene in plasma (Epi procolon), and early colorectal cancer screening products for detecting the methylation of BMP3 / NDRG4 combined with KRAS gene mutation plus FIT in feces (Cologuard). However, there is still a large room for improvement in the sensitivity and specificity of these detection techniques. For example, the Epi procolon detection has a specificity of 97.5%, but the sensitivity is only 79%, which will cause a large proportion of missed diagnoses. Cologuard can reach a sensitivity of 95.55%, but its specificity will be reduced to 87.1%. Improving both sensitivity and specificity can better improve the accuracy of detection and minimize the probability of missed diagnoses and misdiagnoses. In addition, for protein markers such as CEA detection, their sensitivity and specificity are even more limited.
[0006] In recent years, proteomics based on high-resolution mass spectrometers has not only greatly improved the detection accuracy but also increased the detection speed, and is gradually applicable to analyzing the proteome expression levels of large-scale clinical samples. After years of practice, the industry generally believes that early cancer screening methods with high sensitivity and high specificity need to shift from single protein markers to combined markers. Currently, there is no early screening and diagnostic kit for colorectal cancer based on protein markers. Summary of the Invention
[0007] To solve at least one of the above technical problems, the technical solutions adopted in the present invention are as follows:
[0008] In a first aspect of the present invention, there is provided a protein marker combination for predicting, diagnosing or prognosticating colorectal cancer, the protein marker combination comprising at least one selected from LRG1, SERPINA1, ITIH3, CP, ORM1, C9, IGFBP2, CNDP1.
[0009] ITIH3: The heavy chain H3 of inter-α-trypsin inhibitor, which can stabilize the extracellular matrix through its ability to bind hyaluronic acid. Polymorphisms of this gene may be associated with an increased risk of schizophrenia and major depressive disorder.
[0010] LRG1: Belonging to the leucine-rich repeat protein family, it plays an important role in protein-protein interaction, signal transduction, cell adhesion and development.
[0011] C9: This protein is the last component of the complement system and is involved in the formation of the membrane attack complex (MAC). The membrane attack complex plays a key role in innate and adaptive immune responses.
[0012] IGFBP2: This protein can bind insulin-like growth factors I and II (IGF-I and IGF-II). After being secreted into the blood, it can better bind IGF-I and IGF-II, and can also interact with different ligands inside cells. High expression of IGFBP2 can promote the growth of various tumors and can predict the prognosis of patients.
[0013] CNDP1: This protein is one of the members of the M20 metalloprotease family, a homodimeric dipeptidase specifically expressed in the brain, and the coding region of the gene contains a trinucleotide (CTG) repeat sequence.
[0014] SERPINA1: This protein is a serine protease inhibitor belonging to the serine superfamily, and its target sites include elastase, plasmin, thrombin, trypsin, chymotrypsin, and plasminogen activator. This protein is produced by lymphocytes and monocytes in the liver, bone marrow, and lymphoid tissues, as well as Paneth cells in the intestine. It has been reported that defects in this gene are associated with chronic obstructive pulmonary disease, emphysema, and chronic liver disease.
[0015] CP: This protein is a metalloprotein that can bind most of the copper in plasma and participate in the peroxidation reaction of iron (II) transferrin to iron (III) transferrin. Mutations in this gene can lead to acute plasminemia, resulting in iron accumulation and tissue damage, and are associated with diabetes and neurological abnormalities.
[0016] ORM1: This protein belongs to the acute-phase plasma proteins. Its expression level increases during acute inflammatory responses. The specific function of this protein is unknown and may be involved in immunosuppression.
[0017] In some embodiments of the present invention, the protein biomarker combination includes LRG1 and at least one of SERPINA1, ITIH3, CP, ORM1, C9, IGFBP2, and CNDP1.
[0018] In some other embodiments of the present invention, the protein biomarker combination includes C9 and at least one of LRG1, SERPINA1, ITIH3, CP, ORM1, IGFBP2, and CNDP1.
[0019] In some specific embodiments of the present invention, the protein biomarker combination includes ITIH3, LRG1, C9, IGFBP2, and CNDP1.
[0020] In some specific embodiments of the present invention, the protein biomarker combination includes CP, LRG1, C9, IGFBP2, and CNDP1.
[0021] In some specific embodiments of the present invention, the protein biomarker combination includes ITIH3, CP, LRG1, C9, and CNDP1.
[0022] In some specific embodiments of the present invention, the protein biomarker combination includes SERPINA1, LRG1, C9, IGFBP2, and CNDP1.
[0023] In some specific embodiments of the present invention, the protein biomarker combination includes SERPINA1, CP, LRG1, C9, and CNDP1.
[0024] In some specific embodiments of the present invention, the protein biomarker combination includes LRG1, ORM1, C9, IGFBP2, and CNDP1.
[0025] In some specific embodiments of the present invention, the protein biomarker combination includes LRG1, SERPINA1, CP, ORM1, C9, and CNDP1.
[0026] In some specific embodiments of the present invention, the protein biomarker combination includes LRG1, SERPINA1, ITIH3, CP, C9, and CNDP1.
[0027] In some specific embodiments of the present invention, the protein biomarker combination includes LRG1, SERPINA1, ITIH3, C9, IGFBP2, and CNDP1.
[0028] In some specific embodiments of the present invention, the protein biomarker combination includes SERPINA1, ITIH3, LRG1, C9, IGFBP2, and CNDP1.
[0029] In some specific embodiments of the present invention, the protein biomarker combination includes SERPINA1, ITIH3, LRG1, ORM1, C9, and CNDP1.
[0030] In the present invention, by detecting the expression levels of the proteins in the protein biomarker combination, it is possible to predict whether a subject has a risk of developing colorectal cancer, that is, it can be used for early screening of colorectal cancer; it can also diagnose whether a subject has colorectal cancer, and the diagnosis can be an auxiliary diagnosis, which is diagnosed by a clinician in combination with other clinical indicators; it can also evaluate the prognosis of a subject with colorectal cancer after treatment.
[0031] The second aspect of the present invention provides a polypeptide combination for predicting, diagnosing, or prognosticating colorectal cancer, and the polypeptide combination includes at least one polypeptide from each protein in any of the protein biomarker combinations of the first aspect of the present invention.
[0032] Optionally, the polypeptide from C9 includes the amino acid sequence shown in SEQ ID No. 1 or SEQ ID No. 2.
[0033] Optionally, the polypeptide from SERPINA1 includes the amino acid sequence shown in SEQ ID No. 3.
[0034] Optionally, the polypeptide from ITIH3 includes the amino acid sequence shown in SEQ ID No. 4.
[0035] Optionally, the polypeptide from CP includes the amino acid sequence shown in SEQ ID No. 5.
[0036] Optionally, the polypeptide from LRG1 comprises the amino acid sequence shown in SEQ ID No. 6 or SEQ ID No. 7.
[0037] Optionally, the polypeptide from IGFBP2 comprises the amino acid sequence shown in SEQ ID No. 8.
[0038] Optionally, the polypeptide from KNG1 comprises the amino acid sequence shown in SEQ ID No. 9.
[0039] Optionally, the polypeptide from ORM1 comprises the amino acid sequence shown in SEQ ID No. 10.
[0040] Optionally, the polypeptide from PRDX2 comprises the amino acid sequence shown in SEQ ID No. 11.
[0041] Optionally, the polypeptide from CNDP1 comprises the amino acid sequence shown in SEQ ID No. 12.
[0042] The third aspect of the present invention provides the use of a reagent for detecting the expression level of the protein biomarker combination according to any one of the first aspect of the present invention in the preparation of a kit for colorectal cancer prediction, diagnosis or prognosis.
[0043] In some embodiments of the present invention, the detection reagent detects the expression level of each protein in the protein biomarker combination based on a mass spectrometry method.
[0044] In some specific embodiments of the present invention, the expression level of each protein in the protein biomarker combination is detected by detecting the level of one or more polypeptides of each protein in the protein biomarker combination.
[0045] Optionally, the polypeptide from C9 comprises the amino acid sequence shown in SEQ ID No. 1 or SEQ ID No. 2.
[0046] Optionally, the polypeptide from SERPINA1 comprises the amino acid sequence shown in SEQ ID No. 3.
[0047] Optionally, the polypeptide from ITIH3 comprises the amino acid sequence shown in SEQ ID No. 4.
[0048] Optionally, the polypeptide from CP comprises the amino acid sequence shown in SEQ ID No. 5.
[0049] Optionally, the polypeptide from LRG1 comprises the amino acid sequence shown in SEQ ID No. 6 or SEQ ID No. 7.
[0050] Optionally, the polypeptide from IGFBP2 comprises the amino acid sequence shown in SEQ ID No. 8.
[0051] Optionally, the polypeptide from KNG1 comprises the amino acid sequence shown in SEQ ID No. 9.
[0052] Optionally, the polypeptide from ORM1 comprises the amino acid sequence shown in SEQ ID No. 10.
[0053] Optionally, the polypeptide from PRDX2 comprises the amino acid sequence shown in SEQ ID No. 11.
[0054] Optionally, the polypeptide from CNDP1 comprises the amino acid sequence shown in SEQ ID No. 12.
[0055] The fourth aspect of the present invention provides a kit for predicting, diagnosing or prognosticating colorectal cancer, comprising reagents for detecting the expression levels of any of the protein biomarker combinations of the first aspect of the present invention.
[0056] The fifth aspect of the present invention provides a method for predicting, diagnosing or prognosticating colorectal cancer, comprising the following steps:
[0057] S1, obtaining the expression level data of each protein in any of the protein biomarker combinations of the first aspect of the present invention for a subject;
[0058] S2, constructing a machine learning model using the expression level data of each protein in the protein biomarker combination in a population sample and information on whether each sample is from a colorectal cancer patient, and determining whether the subject has colorectal cancer or has a risk of developing colorectal cancer or has a good prognosis for colorectal cancer based on the machine learning model.
[0059] In some embodiments of the present invention, the machine learning model is trained using any one of the following algorithms:
[0060] Random forest algorithm, support vector machine algorithm, linear regression algorithm, logistic regression algorithm, Bayesian classifier, and neural network algorithm.
[0061] In some preferred embodiments of the present invention, the machine learning model is trained using the logistic regression algorithm.
[0062] Furthermore, a preset threshold is obtained based on the population sample using the machine learning model. For the model determination result of each subject sample, if it is higher than the preset threshold, it is determined that the subject has colorectal cancer or has a risk of developing colorectal cancer or has a poor prognosis for colorectal cancer. If it is not higher than the preset threshold, it is determined that the subject does not have colorectal cancer or does not have a risk of developing colorectal cancer or has a good prognosis for colorectal cancer.
[0063] In some embodiments of the present invention, in step S1, the blood sample of the subject is anticoagulated with EDTA to obtain plasma. After the plasma protein is denatured, reduced, and alkylated, trypsin is added for enzymatic digestion to obtain polypeptide fragments. After desalting and evaporation, liquid phase separation and mass spectrometry detection are performed, and the level of the protein biomarker combination is determined based on the level of the polypeptide.
[0064] In some embodiments of the present invention, the mass spectrometry detection is performed using a triple quadrupole mass spectrometry method.
[0065] The sixth aspect of the present invention provides a system for colorectal cancer prediction, diagnosis, or prognosis, including the following modules:
[0066] A data input module for inputting the expression level data of each protein in any of the protein biomarker combinations of the first aspect of the present invention for the subject.
[0067] A data storage module for storing the expression level data of each protein in the protein biomarker combination in a population sample and the information on whether each sample is from a colorectal cancer patient.
[0068] A colorectal cancer analysis module, which is respectively connected to the data input module and the data storage module, constructs a machine learning model using the expression level data of each protein in the protein biomarker combination in the population sample stored in the data storage module and the information on whether each sample is from a colorectal cancer patient, and determines whether the subject has colorectal cancer or has a risk of developing colorectal cancer or has a good prognosis of colorectal cancer based on the machine learning model.
[0069] In some embodiments of the present invention, the machine learning model is trained using any one of the following algorithms:
[0070] Random forest algorithm, support vector machine algorithm, linear regression algorithm, logistic regression algorithm, Bayesian classifier, and neural network algorithm.
[0071] In some embodiments of the present invention, the colorectal cancer analysis module further inputs the expression level data of each protein in the protein biomarker combination of the subject and the judgment result into the data storage module.
[0072] In some preferred embodiments of the present invention, the machine learning model is trained using a logistic regression algorithm.
[0073] Advantages of the present invention
[0074] Compared with the prior art, the present invention has the following advantages:
[0075] Simultaneously detect multiple protein markers in plasma based on targeted mass spectrometry and perform absolute quantification, with accurate results and saving the time cost of detection.
[0076] The protein marker combination of the present invention provides a non-invasive screening method for early colorectal cancer based on plasma.
[0077] Using the method and system of the present invention for colorectal cancer prediction, diagnosis or prognosis is non-invasive to patients, convenient for sampling, requires a small amount of plasma sample, has high sensitivity and specificity, and most importantly, fills the blank of no effective protein markers for early colorectal cancer.
[0078] The protein marker combination of the present invention has high accuracy in predicting early colorectal cancer. After a positive result is judged, it prompts the patient to undergo further confirmation, and in the long run, it can effectively reduce the mortality rate of colorectal cancer in the population.
[0079] Detecting the marker proteins in plasma using machine learning can achieve the purpose of dynamically monitoring the disease state of patients. Brief Description of the Drawings
[0080] Figure 1 Shows the receiver operating characteristic curve of a single protein marker LRG1. The areas under the curve (AUC) of the training set, test set and independent validation set are 0.904, 0.85 and 0.8 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and Falsepostive rate (1-specificty) represents the false positive rate (1-specificity).
[0081] Figure 2 Shows the receiver operating characteristic curve of a single protein marker SERPINA1. The areas under the curve (AUC) of the training set, test set and independent validation set are 0.837, 0.779 and 0.771 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1-specificty) represents the false positive rate (1-specificity).
[0082] Figure 3The receiver operating characteristic curve of a single protein marker ITIH3 is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.835, 0.921, and 0.79 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0083] Figure 4 The receiver operating characteristic curve of a single protein marker CP is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.823, 0.842, and 0.624 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0084] Figure 5 The receiver operating characteristic curve of a single protein marker ORM1 is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.818, 0.783, and 0.697 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0085] Figure 6 The receiver operating characteristic curve of a single protein marker C9 is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.875, 0.91, and 0.81 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and Falsepostive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0086] Figure 7The receiver operating characteristic curve of a single protein biomarker IGFBP2 is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.728, 0.738, and 0.737 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0087] Figure 8 The receiver operating characteristic curve of a combination of 5 protein biomarkers is shown. The area under the curve (AUC) of the training set, test set, and independent validation set are 0.956, 0.954, and 0.893 respectively, where train represents the training set, test represents the test set, and valid represents the independent validation set; True positive rate (sensitivity) represents the true positive rate (sensitivity), and False postive rate (1 - specificty) represents the false positive rate (1 - specificity).
[0088] Figure 9 The confusion matrix of a combination of 5 protein markers is shown. Among them, there are 121 colorectal cancer patients and 186 healthy people. 1 represents positive and 0 represents negative. Where train represents the training set, test represents the test set, and valid represents the independent validation set; Truth represents the truth and Prediction represents the prediction. Detailed implementation mode
[0089] Unless otherwise specified, implied from the context or in accordance with the convention of the prior art, all parts and percentages in this application are based on weight, and the test and characterization methods used are synchronized with the filing date of this application. Where applicable, any patents, patent applications, or published content referred to in this application are incorporated herein by reference in their entirety, and their equivalent family patents are also incorporated by reference, especially the definitions of relevant terms in this field disclosed in these documents. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition of the term provided in this application shall prevail.
[0090] The numerical ranges in this application are approximate values, so unless otherwise stated, they may include values outside the range. The numerical range includes all values from the lower limit value to the upper limit value increased by 1 unit, provided that there is an interval of at least 2 units between any lower value and any higher value. For ranges containing values less than 1 or fractions greater than 1 (such as 1.1, 1.5, etc.), 1 unit is appropriately regarded as 0.0001, 0.001, 0.01 or 0.1. For ranges containing single-digit numbers less than 10 (such as 1 to 5), 1 unit is usually regarded as 0.1. These are merely specific examples of what is intended to be expressed, and all possible combinations of the values between the lowest and highest values listed are considered to be clearly recorded in this application.
[0091] The terms "comprising", "including", "having" and their derivatives do not exclude the existence of any other components, steps or processes, and are independent of whether these other components, steps or processes are disclosed in this application. To eliminate any doubt, unless expressly stated, all compositions using the terms "comprising", "including" or "having" in this application may contain any additional additives, excipients or compounds. In contrast, except for those necessary for the operating performance, the term "consisting essentially of" excludes any other components, steps or processes from the scope described below any such term. The term "consisting of" does not include any components, steps or processes not specifically described or listed. Unless expressly stated, the term "or" refers to the individual members listed or any combination thereof.
[0092] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention clearer and more understandable, the following further details the present invention in conjunction with embodiments.
[0093] Embodiment
[0094] The following examples are used here to demonstrate the preferred embodiments of the present invention. Those skilled in the art will understand that the technologies disclosed in the following examples represent the technologies that the inventor has found can be used to implement the present invention, and thus can be regarded as the preferred solutions for implementing the present invention. However, those skilled in the art should understand from this specification that the specific embodiments disclosed here can be modified in many ways and still obtain the same or similar results without departing from the spirit or scope of the present invention.
[0095] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs, and all materials cited herein and their cited materials will be incorporated by reference.
[0096] Those skilled in the art will recognize or be able to ascertain through routine experimentation many equivalents to the specific embodiments of the inventions described herein. Such equivalents are intended to be encompassed by the claims.
[0097] Unless otherwise specified, the experimental methods in the following examples are all conventional methods. Unless otherwise specified, the instrumentation and equipment used in the following examples are all conventional laboratory instrumentation and equipment; unless otherwise specified, the test materials used in the following examples are all obtained from regular biochemical reagent stores.
[0098] Example 1 Discovery of Protein Biomarkers
[0099] The inventors collected fresh blood samples from 101 colorectal cancer patients and 89 healthy controls matched for gender and age for the discovery of protein biomarkers.
[0100] 1. Blood Sample Processing
[0101] After anticoagulation treatment, the fresh blood samples were centrifuged at 1000 g for 5 min to obtain plasma samples, which were stored long-term in a -70 °C refrigerator.
[0102] The plasma samples were diluted 50-fold and then assayed for concentration by the BCA method: BSA standards were serially diluted to concentration gradients of 2, 1, 0.5, 0.25, 0.125, 0.0625 mg / mL to calibrate the plasma concentration as a working curve. The diluted samples and standards were added to a 96-well plate, the pre-prepared BCA working solution was added, and the reaction was carried out at 37 °C for 30 min. The plasma protein concentration was measured at an absorbance of 562 nm.
[0103] 50 μg of protein was taken and ammonium bicarbonate solution was added to a final concentration of 50 mM. DTT was added to a final concentration of 10 mM and heated at 95 °C for 10 min. After returning to room temperature, IAA was added to a final concentration of 15 mM for a dark reaction for 30 min. 1 μg of trypsin was added to each sample, and the enzymatic digestion reaction was carried out overnight at 37 °C in a metal bath for 12 - 14 h. The next day, formic acid was added to a final concentration of 1% for acidification treatment to terminate the enzymatic digestion reaction.
[0104] 2. Differential Proteins and Polypeptides
[0105] The selection of targets is first based on finding differentially expressed proteins. The inventors performed mass spectrometry collection on 190 plasma samples with symmetric gender and age (89 healthy individuals and 101 colorectal cancer patients) using a data-independent acquisition mode (DIA). Further, the inventors analyzed the obtained expression data of proteins and polypeptides using DIA-NN software, and performed normalization analysis using the intensity of total proteins. A total of 714 proteins and 7,988 polypeptides were quantified. For proteins and polypeptides with a normal distribution of expression, the inventors used the t-test to find differentially expressed proteins and polypeptides. For proteins and polypeptides with a non-normal distribution of expression, the inventors used the Wilcoxon non-parametric test to find differentially expressed proteins and polypeptides. Finally, the inventors obtained a total of 96 differentially expressed proteins and 832 differentially expressed polypeptides. The differentially expressed polypeptides were integrated.
[0106] 3. Screening of biomarker proteins
[0107] Potential polypeptides capable of distinguishing colorectal cancer patients from healthy individuals were selected using the random forest method. The random forest calculated the average Gini coefficient of these targets and ranked them according to importance. Further, in combination with the biological functions of the proteins, 10 top-ranked proteins were finally obtained, namely LRG1, SERPINA1, ITIH3, CP, ORM1, C9, IGFBP2, CNDP1, KNG1, and PRDX2. The corresponding polypeptide sequences are shown in Table 1:
[0108] Table 1 Polypeptide sequences of candidate proteins
[0109]
[0110] Example 2 Establishment of machine learning model
[0111] C of the appropriate concentration for each polypeptide 13 and N 15 The heavy isotope-labeled polypeptides of C and N were added to the digested plasma samples after digestion, and after mixing, desalting and drying were performed using a 96-well SOLA solid-phase extraction device.
[0112] For each polypeptide, a calibration curve range with an appropriate concentration (9 calibration curve points) was configured, and an equal amount of internal standard was also added to each calibration curve point. Mass spectrometry detection was performed using an AB Sciex 5500 Qtrap mass spectrometer. The polypeptides were separated using a C18 chromatographic column (Phenomenex), the column temperature was set at 45 °C, and the standard sample was injected at 15 μL. 150 μL of 0.1% formic acid was added to the dried sample, mixed well, and 15 μL was injected for mass spectrometry detection. The liquid phase separation conditions are shown in Table 2:
[0113] Table 2 Liquid phase separation conditions
[0114]
[0115] Triple quadrupole targeted mass spectrometry detection was then performed, and the ion pair information for multiple reaction monitoring (MRM) is shown in Table 3.
[0116] Table 3 MRM monitoring information
[0117]
[0118] After mass spectrometry acquisition, the polypeptide concentrations corresponding to their respective protein markers were quantified and used for model establishment. 80% (152 cases) of 190 samples were randomly selected as the training set, and the remaining 20% (38 cases) were used as the test set. A logistic regression model was further established for 10 potential protein markers. The inventors found that 7 individual protein markers, namely LRG1, SERPINA1, ITIH3, CP, ORM1, C9, and IGFBP2, had very good predictive ability in both the training set and the test set, and their ROC curves are respectively as Figures 1 to 7 shown.
[0119] Example 3 Model verification
[0120] The inventors selected 121 colorectal cancer patients and 186 matched healthy individuals as the validation set for model verification. To more accurately quantify the polypeptides and reduce the errors caused by cumbersome experimental procedures, the inventors no longer performed the operation of removing high-abundance proteins, which could also greatly reduce the pre-treatment cost of the experiment. After protein extraction and concentration determination, liquid phase separation and mass spectrometry detection were performed.
[0121] Example 4 Establishment and verification of a model with a combination of multiple markers
[0122] The inventors further used the optimal combination of the aforementioned proteins - the concentrations of 5 protein markers (ITIH3, LRG1, SERPINA1, IGFBP2, and CDNP1) to establish a logistic regression model to well distinguish colorectal cancer patients from healthy individuals. Specifically, 77 colorectal cancer patients and 79 healthy individuals were used for logistic regression modeling to learn the discrimination effect of the 5 protein markers. The threshold in the logistic regression model was set at 0.34, and 44 colorectal cancer patients and 107 healthy individuals were used for independent verification of the model. Based on the model results of all 307 plasma samples, the threshold was set. For the model determination result of each sample, if it was higher than this threshold, it was judged as positive. If the model determination result of the sample was lower than this threshold, it was judged as negative.
[0123] The ROC curve is as Figure 8As shown, the areas under the curves (AUC) of the training set, test set, and independent validation set are 0.956, 0.954, and 0.893, respectively. The final sensitivity is 92%, the specificity is 81%, the negative predictive value is 94%, and the positive predictive value is 76%, as Figure 9 shown.
[0124] In addition, the inventors also presented 10 other combinations of protein markers that performed well during the machine learning process, and the results are shown in Table 4.
[0125] Table 4 Combinations of Protein Markers
[0126]
[0127] All documents mentioned in the present invention are incorporated herein by reference as if each individual document was specifically and individually incorporated by reference. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
Claims
1. A protein biomarker combination for colorectal cancer diagnosis, characterized in that, The protein biomarker combination: (1) Composed of CP, LRG1, C9, IGFBP2, and CNDP1; (2) Composed of SERPINA1, LRG1, C9, IGFBP2, and CNDP1; (3) Composed of SERPINA1, CP, LRG1, C9, and CNDP1; (4) Composed of LRG1, ORM1, C9, IGFBP2, and CNDP1; (5) Composed of LRG1, SERPINA1, CP, ORM1, C9, and CNDP1; (6) Composed of LRG1, SERPINA1, ITIH3, CP, C9, and CNDP1; (7) Composed of LRG1, SERPINA1, ITIH3, C9, IGFBP2, and CNDP1; (8) Composed of SERPINA1, ITIH3, LRG1, C9, IGFBP2, and CNDP1; or (9) Composed of SERPINA1, ITIH3, LRG1, ORM1, C9, and CNDP1.
2. Use of a reagent for detecting the expression level of the protein biomarker combination according to claim 1 in the preparation of a kit for diagnosing colorectal cancer.
3. The application according to claim 2, characterized in that, The detection reagent detects the expression levels of the proteins in the protein biomarker combination based on a mass spectrometry method.
4. A kit for colorectal cancer diagnosis, characterized in that, A detection reagent for the expression level of the protein biomarker combination according to claim 1.
5. A system for colorectal cancer diagnosis, characterized in that, Including the following modules: A data input module for inputting the expression level data of each protein in the protein biomarker combination of a subject, where the protein biomarker combination is as described in claim 1; A data storage module for storing the expression level data of each protein in the protein biomarker combination in a population sample and information on whether each sample is from a colorectal cancer patient; A colorectal cancer analysis module, connected to the data input module and the data storage module respectively, constructs a machine learning model using the expression level data of each protein in the protein biomarker combination in the population sample stored in the data storage module and information on whether each sample is from a colorectal cancer patient, and determines whether the subject has colorectal cancer based on the machine learning model.
6. The system according to claim 5, characterized in that, The machine learning model is trained using any one of the following algorithms: Random forest algorithm, support vector machine algorithm, linear regression algorithm, logistic regression algorithm, Bayesian classifier, and neural network algorithm.
7. The system according to claim 5 or 6, characterized in that, The colorectal cancer analysis module further inputs the expression level data of each protein in the protein biomarker combination of the subject and the judgment result into the data storage module.
Citation Information
Patent Citations
Methods and machine learning systems for predicting the likelihood or risk of having cancer
CN109036571A
Method for constructing mathematical model for detecting colorectal cancer in vitro, and application thereof
CN111584008A
Compositions and methods for diagnosis and prognosis of colorectal cancer
US20120149022A1
Methods for Detection and Treatment of Colorectal Cancer
US20170269089A1
Biomarker assay and uses thereof for diagnosis, therapy selection, and prognosis of cancer
WO2013152989A2