A marker group for intestinal cancer screening and diagnosis and use thereof
By combining specially modified peptide biomarkers and machine learning algorithms, the problems of insufficient sensitivity and specificity in colorectal cancer screening and diagnosis have been solved, achieving higher diagnostic accuracy and making it suitable for early screening and diagnosis of colorectal cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies lack sufficient sensitivity and specificity in colorectal cancer screening and diagnosis. Traditional methods are highly invasive or susceptible to external factors, and existing biomarkers such as CEA and CA19-9 have low specificity, making it difficult to meet clinical needs.
A diagnostic model is constructed using a biomarker set, including specifically modified peptides, through serum sample detection, combined with machine learning algorithms, to improve the sensitivity and specificity of auxiliary diagnosis and early screening for colorectal cancer.
It achieves higher sensitivity and specificity than existing technologies, is suitable for colorectal cancer screening and diagnosis, reduces false negative results, and improves the early detection rate of patients.
Smart Images

Figure CN121186358B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular biotechnology, specifically relating to a biomarker group for colorectal cancer screening and diagnosis and its uses. Background Technology
[0002] Because colorectal cancer often presents with no symptoms in its early stages, many patients are diagnosed at an advanced stage, missing the optimal treatment window. A high proportion of patients with advanced colorectal cancer (CRC) undergo colostomy, leading to a more uncomfortable later life. Therefore, early diagnosis and effective treatment of CRC patients are crucial for improving their clinical symptoms, quality of life, and prolonging their lifespan.
[0003] Colorectal cancer is a cancer suitable for early screening. Traditional diagnostic methods for CRC mainly include colonoscopy, tissue biopsy, imaging examinations, and tumor markers. However, these methods have certain limitations. Colonoscopy has high sensitivity and specificity for identifying lesions and can also remove or biopsie lesions, making it the recognized gold standard for colorectal cancer screening. However, it is highly invasive, leading to poor patient compliance. Furthermore, the sensitivity and specificity of colonoscopy can vary between different operators. While tissue biopsy can provide pathological evidence, it is also invasive and may be subject to sampling errors, resulting in false negatives. Imaging examinations such as CT and MRI can provide anatomical information about tumors, but they have limitations in detecting early, small lesions. Additionally, the fecal occult blood test (FIT) is low-cost, simple, and convenient, making it one of the most widely used and cost-effective screening methods. However, its results are easily affected by factors such as medication and diet, resulting in low specificity, and it has been gradually replaced by the immunochemical fecal occult blood test (FIT). If the FIT test result is positive in clinical practice, further colonoscopy is required.
[0004] Liquid biopsy, as an emerging diagnostic technique, can overcome the influence of tumor heterogeneity, providing more comprehensive molecular information about tumors and offering strong support for guiding clinical treatment and assessing prognosis. Carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9) are currently the most widely used serum biomarkers for CRC in clinical practice. CEA is a glycoprotein widely present in embryonic tissues, and its expression level is significantly elevated in the serum of CRC patients, with its concentration closely related to tumor progression and metastasis. However, CEA has low specificity and may also be elevated in patients with other types of cancer and non-tumor diseases. Although CA19-9 also has some diagnostic value in CRC patients, it is mainly used for the diagnosis of pancreatic cancer and biliary tract cancer, and its application in CRC is also limited by sensitivity and specificity. Therefore, although CEA and CA19-9 have certain auxiliary roles in the diagnosis of CRC, their use alone is insufficient to meet clinical needs. Summary of the Invention
[0005] In view of the above-mentioned shortcomings in the prior art, the present invention provides a set of biomarkers for colorectal cancer screening and diagnosis and their uses. When used for auxiliary diagnosis and early screening of colorectal cancer, it has excellent sensitivity and specificity and is expected to be applied to the diagnosis and treatment of colorectal cancer.
[0006] To achieve the above objectives, the technical solution adopted by the present invention to solve its technical problem is as follows:
[0007] A biomarker group for colorectal cancer screening and diagnosis includes at least four of the polypeptides shown in SEQ ID NO. 1-35, with specific sequences shown in Table 1.
[0008] Table 1. Peptide Sequences
[0009]
[0010] Among them, the first amino acid Q of peptide 6 is modified with Gln->pyro-Glu, and the sixth amino acid N is modified with Dehydrated; the second amino acid K of peptide 16 is modified with Acetyl; the third amino acid G of peptide 22 is modified with Phospho; the fourth amino acid G of peptide 31 is modified with Dehydrated; and the fourth amino acid G of peptide 32 is modified with Phospho.
[0011] Furthermore, the biomarker set includes the peptides shown in sequences 1 and 2, as well as the following combinations of peptides:
[0012] The polypeptide combination is one of sequence 3 and sequence 4; sequence 5 and sequence 6; sequence 11 and sequence 12; sequence 30 and sequence 35; sequence 3, sequence 4 and sequence 5; sequence 5, sequence 6 and sequence 7; sequence 11, sequence 12 and sequence 20; sequence 9, sequence 30 and sequence 35.
[0013] Furthermore, the biomarker set includes the peptides shown in sequences 7 and 8, as well as the following combinations of peptides:
[0014] The polypeptide combination is one of the following: sequence 3 and sequence 4; sequence 9 and sequence 10; sequence 3, sequence 4 and sequence 9; or sequence 9, sequence 10 and sequence 11.
[0015] Furthermore, the biomarker group includes peptides shown in sequences 1, 10, and 30, as well as peptides shown in sequences 35, 20, 35, and 9, or sequences 20 and 22.
[0016] Furthermore, the biomarker set includes the peptides shown in sequences 5 and 15, as well as the following combinations of peptides:
[0017] The polypeptide combination is one of the following: sequence 25 and sequence 35; sequence 2 and sequence 16; sequence 9, sequence 25 and sequence 35; sequence 2, sequence 16 and sequence 20.
[0018] Furthermore, the biomarker set includes the peptides shown in sequences 4 and 8, as well as the following combinations of peptides:
[0019] The polypeptide combination is the polypeptide shown in sequences 19 and 20; sequences 2 and 6; or sequences 19, 20 and 22.
[0020] Furthermore, the biomarker set includes the polypeptides shown in sequence 7, sequence 10, sequence 17, sequence 18 and / or sequence 22.
[0021] Furthermore, the biomarker group includes polypeptides as shown in sequences SEQ ID NO.1~35.
[0022] The use of the above biomarkers in the preparation of formulations for colorectal cancer screening and diagnosis.
[0023] Furthermore, colorectal cancer is specifically colorectal cancer.
[0024] The above biomarkers may be used in basic medical research for non-diagnostic / therapeutic purposes.
[0025] Further, basic medical research includes Western blotting, immunohistochemistry, or flow cytometry.
[0026] The beneficial effects of this invention are:
[0027] The biomarker combination constructed in this invention has superior sensitivity and specificity when used for the auxiliary diagnosis and early screening of colorectal cancer, and is expected to be applied to the diagnosis and treatment of colorectal cancer. Attached Figure Description
[0028] Figure 1 ROC curve for marker combination 1;
[0029] Figure 2 ROC curve for marker combination 2;
[0030] Figure 3 ROC curve for marker combination 3;
[0031] Figure 4 ROC curve for marker combination 4;
[0032] Figure 5 ROC curve for marker combination 5;
[0033] Figure 6 ROC curve for marker combination 6;
[0034] Figure 7 ROC curve for marker combination 7;
[0035] Figure 8 ROC curve for marker combination 8;
[0036] Figure 9 ROC curve for marker combination 9;
[0037] Figure 10 ROC curve for marker combination 10;
[0038] Figure 11 ROC curve for marker combination 11;
[0039] Figure 12 ROC curve for marker combination 12;
[0040] Figure 13 ROC curve for marker combination 13;
[0041] Figure 14 ROC curve for marker combination 14;
[0042] Figure 15 ROC curve for marker combination 15;
[0043] Figure 16 ROC curve for marker combination 16;
[0044] Figure 17 ROC curve for marker combination 17;
[0045] Figure 18 ROC curve for marker combination 18;
[0046] Figure 19 ROC curve for marker combination 19;
[0047] Figure 20 ROC curve for marker combination 20;
[0048] Figure 21 ROC curve for marker combination 21;
[0049] Figure 22 ROC curve for marker combination 22;
[0050] Figure 23 ROC curve for marker combination 23;
[0051] Figure 24 ROC curve for marker combination 24;
[0052] Figure 25 ROC curve for marker combination 25;
[0053] Figure 26 ROC curve for marker combination 26. Detailed Implementation
[0054] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0055] In this invention, all patient samples were collected from Zhongshan Hospital, affiliated with Fudan University, and passed ethical review.
[0056] The experimental methods used in this invention are as follows:
[0057] I. Serum Sample Collection
[0058] 1) Sample type: serum.
[0059] 2) Collection requirements: Fasting is required. Use a coagulation tube to draw 5 mL of venous blood, let it stand for 30 min, centrifuge at 3000 rpm for 15 min, and take out about 1 mL of serum and put it into a cryopreservation tube.
[0060] 3) Sample storage:
[0061] Use on the same day; store at 2-8℃.
[0062] If not used on the same day, store at -20℃ for up to 30 days;
[0063] If stored for an extended period (more than one month), it should be kept at -80°C.
[0064] The freeze-thaw cycle should not exceed 3 times.
[0065] II. Extraction of analytes from serum
[0066] 1) After calibrating the mass spectrometer, turn on the Solid Bio Fully Automated Sample Analysis System SPS1000 / SPS4000, and put in the consumables, matching reagent kits and the sample to be tested;
[0067] Select the procedure method "Concentrated loading";
[0068] Run the program:
[0069] a. Opening a hole;
[0070] b. Take at least 10µL of serum sample and activation reagent, mix them in a 1:1 ratio, and place them in the G-row pre-reserved well for later use;
[0071] c. Clean the custom pipette tip in cleaning reagent 1 and cleaning reagent 2 in sequence. Each time, aspirate at least 10µL of liquid and repeat the aspiration and dispensing process at least 3 times.
[0072] d. Process the serum mixture in the G-row wells using the cleaned custom pipette tips. Aspirate at least 10 µL of solution each time, repeating the process at least three times.
[0073] e. Clean the custom pipette tip after adsorbing the serum mixture using cleaning reagent 3. During cleaning, aspirate at least 10µL of liquid each time, repeating the aspiration and dispensing process at least 3 times.
[0074] f. Transfer no less than 10µL of buffer reagent into the H-row pre-reserved hole, and place the customized pipette tip after using cleaning reagent 3 into the liquid to draw no less than 10µL of liquid. Repeat the suction and aspiration at least 3 times.
[0075] g. Transfer at least 10µL of sample matrix solution into the H-row pre-reserved well to complete sample processing;
[0076] h. Spot 2.0 µL of the solution from well H onto the hydrophobic-coated biochip (Wuxi Pimo Technology Co., Ltd.).
[0077] i. Vacuum drying for 240 seconds.
[0078] The main components of each reagent are shown in Table 2.
[0079] Table 2 Reagent Composition
[0080]
[0081] III. Mass Spectrometry Data Acquisition and Upload
[0082] The hydrophobic coated biochip (Wuxi Pimo Technology Co., Ltd.) was vacuum dried and placed into a mass spectrometer;
[0083] Data acquisition is performed using the pre-defined SP1 voltage (target high voltage), SP2 voltage (pulse high voltage), focusing voltage (lens high voltage), detector voltage (MCP voltage), pulse delay time, acquisition card range, target diameter, laser frequency, calibration method, and laser intensity.
[0084] IV. Quality Control
[0085] 1) After data collection, the data will be uploaded to the "Mass Spectrometry Data Analysis Software";
[0086] 2) The software reads the sample information and signal spectrum, and judges whether the sample and sample pretreatment are qualified according to the quality control model; quality control failure may include a variety of possibilities, including the sample is not a colorectal cancer lesion sample, the signal spectrum intensity is not up to standard, etc.
[0087] 3) If the quality control fails, adjust the corresponding parameters according to the quality control results and repeat the serum analyte extraction process;
[0088] 4) If the quality control is qualified, proceed to the next process.
[0089] V. Establishment of Positive Criterion Value and Result Analysis
[0090] The study of positive cutoff values used colorectal cancer samples with clear diagnostic information and normal human samples, covering patients with benign intestinal diseases such as irritable bowel syndrome and benign polyps. The core algorithm is based on supervised learning of known colorectal cancer sample atlases. Through a series of processes such as smoothing, noise reduction, and baseline removal, characteristic peaks are screened, a classification model is constructed, and the similarity between hormone signal atlases and known hormone signal atlases stored in the software is calculated (Cannataro M, Guzzi PH, Mazza T, et al. Preprocessing, Management, and Analysis of Mass Spectrometry Proteomics Data[J]. 2005.). Finally, the similarity score positive cutoff value of the kit is determined by the Youden index maximization method. When the similarity score < positive cutoff value, the sample test result is negative; when the similarity score ≥ positive cutoff value, the sample test result is positive. When using the maximum similarity score as the positive cutoff value to assist in the diagnosis of colorectal cancer or to perform early screening for colorectal cancer, the sensitivity is calculated as: (number of true positives / (number of true positives + number of false negatives)) × 100%, and the specificity is calculated as: (number of true negatives / (number of true negatives + number of false positives)) × 100%.
[0091] Example 1: Screening and Identification of Biomarkers
[0092] This invention analyzed 330 healthy individuals (168 males (50.9%) and 162 females (49.1%), aged 25 to 70 years, with a mean age of 46.8 ± 11.7 years; specifically, the age distribution was: 25-34 years (53 cases), 35-44 years (107 cases), 45-54 years (118 cases), 55-64 years (78 cases), and 65-70 years (34 cases)) and 330 colorectal cancer samples (185 males (56.1%) and 145 females (43.9%), aged 38 to 78 years, with a mean age of 5 years). The mean age of the patients was 8.6 ± 10.3 years. The specific age distribution was as follows: 48 cases aged 38-47 years, 126 cases aged 48-57 years, 143 cases aged 58-67 years, and 63 cases aged 68-78 years. Time-of-flight mass spectrometry (TOF-MS) was performed. Through first-level mass spectrometry testing, the relative abundance differences of characteristic peak data in normal individuals and colorectal cancer patients were comprehensively considered, along with statistical differences (p < 0.05, t-test). A machine learning algorithm (random forest) was used to rank the influence factors for feature selection and to assess the matching degree of data in the database. This identified 35 blood peptides with diagnostic capabilities for colorectal cancer. The mass-to-charge ratio (m / z), relative abundance, and influence factors of these peptides are shown in Table 1. The relative abundance was normalized based on the normal human sample, i.e., ln(mean signal intensity of cancer patient samples / mean signal intensity of normal human samples). In machine learning, the selection of characteristic peaks was based on feature importance assessment using ensemble learning. Multiple decision trees were constructed to quantify the contribution of each mass-to-charge ratio (m / z) peak in classification / prediction. The feature importance score, or influence factor, is obtained by calculating the mean reduction in impurity caused by the feature when splitting across all tree nodes (Biau, G., Scornet, E. A random forest guided tour. TEST 25, 197–227 (2016). https: / / doi.org / 10.1007 / s11749-016-0481-7).
[0093] The specific parameters for first-order mass spectrometry are as follows:
[0094] Ionization method: Matrix-assisted laser desorption / ionization (MALDI), with α-cyano-4-hydroxycinnamic acid (CHCA) as the matrix.
[0095] Quality range: 100-4000 Da.
[0096] Resolution: 20000 (full quality range).
[0097] Laser energy: 30-40%.
[0098] Acquisition mode: Positive ion mode.
[0099] Calibration: External quality calibration was performed using the Bruker Peptide Calibration Standard.
[0100] Subsequently, the sequences of these 35 substances in the clinical serum were confirmed using secondary mass spectrometry (MS / MS or TOF / TOF is a peptide identification method recommended by the guidelines of the China Food and Drug Administration and the U.S. Food and Drug Administration). The secondary mass spectrometry data analysis process is as follows:
[0101] Data analysis methods: Mascot software (version 2.8) was used for database searching, with the UniProt Human Proteome Database (released in 2023) as the target database. Search parameters: Enzyme was set to "no digestion", parent ion mass error was allowed ±0.5 Da, fragment ion mass error was allowed ±0.3 Da, fixed modification was cysteine urea methylation, and variable modification was methionine oxidation.
[0102] Sequence confirmation criteria: The confirmation of peptide sequences is based on the matching of fragment ion spectra (b- and y- ions) with theoretical spectra. A Mascot score higher than 30 (p<0.05) is considered significant.
[0103] False positive exclusion: Specificity was verified through reverse database search, and the false positive rate was controlled to below 1%.
[0104] Secondary mass spectrometry parameters:
[0105] Collision-induced dissociation (CID).
[0106] Collision energy: 30 eV.
[0107] Fragment ion mass range: 100-3500 Da.
[0108] Data acquisition: Each sample is scanned at least 1000 times with lasers to improve the signal-to-noise ratio.
[0109] The sequences and specificity of these 35 biomarkers were confirmed by secondary mass spectrometry, ruling out false positives. The specific sequences are shown in Table 1.
[0110] Example 2: Validation of Marker Combinations
[0111] Based on the 35 polypeptide biomarkers identified and confirmed in Example 1 (sequences and mass-to-charge ratios are shown in Table 1), different biomarker combinations were formed as shown in Table 3, and the sensitivity and specificity of the biomarker combinations in the auxiliary diagnosis and early screening of colorectal cancer were verified. The analysis process is as follows: For all validation cohort samples, the same MALDI-TOFMS platform and parameters as in Example 1 were used for detection to obtain the mass spectrometry peak intensity data of all biomarkers in Table 1. The entire detection process was blinded, meaning that the experimental operators were unaware of the sample grouping information. The raw mass spectrometry data were processed by baseline correction, smoothing, and normalization (based on the internal standard peak intensity), and then the peak area or intensity value of each biomarker was extracted. Statistical analysis was performed using R software (version 4.0.2). The preprocessed biomarker intensity data were input into a logistic regression model. For auxiliary diagnostic validation, ten-fold cross-validation was performed using all samples from this cohort (Sun T, Liu J, Yuan H, Li X, Yan H. Construction of a risk prediction model for lung infection after chemotherapy in lung cancer patients based on the machine learning algorithm. Front Oncol. 2024 Aug 9;14:1403392. doi: 10.3389 / fonc.2024.1403392. PMID:39184040; PMCID: PMC11341396.). For early screening validation, given the imbalanced sample size, oversampling (SMOTE) was used to process the data before model training and testing (van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B. The harm of class imbalance corrections for riskprediction models: illustration and simulation using logistic regression. JAm Med Inform Assoc. 2022 Aug 16;29(9):1525-1534. doi: 10.1093 / jamia / ocac093.PMID: 35686364; PMCID: PMC9382395.).The data analysis process referenced general guidelines for constructing clinical prediction models (Zweig MH, Campbell G. Receiver-operating characteristic (ROC) plots: a fundamental evaluation tool in clinical medicine. Clin Chem. 1993 Apr;39(4):561-77. Erratum in: Clin Chem 1993 Aug;39(8):1589. PMID: 8472349.). Specific calculation results for performance indicators are shown in Table 3, and the corresponding ROC curves are shown below. Figures 1-26 .
[0112] 1) The use of various biomarker combinations for the auxiliary diagnosis of colorectal cancer was validated in 500 patients with colorectal cancer (281 males (56.2%), 219 females (43.8%), aged 36 to 76 years, with a mean age of 57.8 ± 9.6 years. Specifically, the distribution was: 63 patients aged 36-45, 127 aged 46-55, 186 aged 56-65, and 124 aged 66-76) and 500 healthy individuals (256 males (51.2%), 244 females (48.8%), aged 28 to 72 years, with a mean age of 47.3 ± 10.8 years. Specifically, the distribution was: 78 patients aged 28-37, 126 aged 38-47, 138 aged 48-57, 108 aged 58-67, and 50 aged 68-72).
[0113] 2) The early screening efficacy of various biomarker combinations for colorectal cancer was validated using a sample of 450 colorectal cancer patients (253 males (56.2%), 197 females (43.8%), aged 38 to 75 years, with a mean age of 56.3 ± 9.8 years. Specifically, the distribution was as follows: 58 patients aged 38-47, 134 patients aged 48-57, 163 patients aged 58-67, and 95 patients aged 68-75) and 2500 healthy individuals (1278 males (51.1%), 1222 females (48.9%), aged 30 to 74 years, with a mean age of 48.6 ± 11.2 years. Specifically, the distribution was as follows: 368 patients aged 30-39, 587 patients aged 40-49, 623 patients aged 50-59, 612 patients aged 60-69, and 310 patients aged 70-74).
[0114] Table 3 Sensitivity and specificity of biomarker combinations for assisted diagnosis and early screening
[0115]
[0116] (Continued from the table above)
[0117]
[0118] (Continued from the table above)
[0119]
[0120] According to Table 3 and Figures 1-26 The test results show that the discriminant peak set (biomarker combination) selected in this invention is superior to the FIT home screening kit (sensitivity and specificity of 73.8% and 94.9%, respectively) and lncRNA-based detection kit (sensitivity and specificity of 82.9% and 72.9%, respectively) currently recommended for use in the NHS (National Health Service) and CT colonography (sensitivity and specificity of 84% and 85%, respectively).
[0121] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A biomarker set for colorectal cancer screening and diagnosis, characterized in that, The biomarker group is selected from one of the following combinations of peptides: A) Sequence 1, Sequence 2, Sequence 3 and Sequence 4; B) Sequences 1, 2, 5, and 6; C) Sequence 1, Sequence 2, Sequence 11 and Sequence 12; D) Sequences 1, 2, 30, and 35; E) Sequence 1, Sequence 2, Sequence 3, Sequence 4 and Sequence 5; F) Sequence 1, Sequence 2, Sequence 5, Sequence 6 and Sequence 7; G) Sequence 1, Sequence 2, Sequence 11, Sequence 12 and Sequence 20; H) Sequence 1, Sequence 2, Sequence 9, Sequence 30 and Sequence 35; I) Sequences 2, 5, 15, and 16; J) Sequence 2, Sequence 5, Sequence 15, Sequence 16 and Sequence 20; K) Sequence 1~Sequence 35; L) Sequences 2, 6, 4, and 8; The amino acid sequences of sequences 1 to 35 are shown in SEQ ID NO. 1 to 35; In sequence 6, the first amino acid Q has a Gln->pyro-Glu modification, and the sixth amino acid N has a dehydrated modification; the second amino acid K in sequence 16 has an Acetyl modification; the third amino acid G in sequence 22 has a Phospho modification; the fourth amino acid G in sequence 31 has a dehydrated modification; and the fourth amino acid G in sequence 32 has a Phospho modification.
2. Use of the reagent for detecting the biomarker group of claim 1 in the preparation of formulations for colorectal cancer screening and diagnosis.
Citation Information
Patent Citations
Colorectal cancer marker polypeptide, and method for diagnosis of colorectal cancer
WO2009025080A1
System and method for diagnosing diseases
WO2009045552A1