Use of reagents for detecting biomarkers in the manufacture of a product for the diagnosis of colorectal cancer
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU UNIV
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-07
AI Technical Summary
[0007]尽管现有研究初步显示了抗原多肽在肿瘤自身抗体检测中的应用潜力,但目前尚缺乏基于全人类蛋白质组层面、系统性地针对结直肠癌相关“抗原多肽-自身抗体”生物标志物的研究
(1)本发明基于包含全人类蛋白质组肽库的PhIP-Seq技术,筛选出潜在的可用于检测或其他表征癌症的抗原多肽,通过多肽微阵列芯片技术筛选出结直肠癌的早期检测生物标志物,再经过ELISA实验进行验证,最终筛选出可用于结直肠癌筛查和诊断的一组结直肠癌联合检测生物标志物,包括CCT6B抗原多肽的自身抗体、第一HNRNPLL抗原多肽的自身抗体、DNMT1抗原多肽的自身抗体、第二HNRNPLL抗原多肽的自身抗体、LIN9抗原多肽的自身抗体、LMO7抗原多肽的自身抗体、NOB1抗原多肽的自身抗体、STAU2抗原多肽的自身抗体和TP53抗原多肽的自身抗体,可辅助用于结直肠癌的诊断检测,具有较好的参考价值;
Smart Images

Figure CN122525127A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to the application of reagents for detecting biomarkers in the preparation of products for the diagnosis of colorectal cancer. Background Technology
[0002] Colorectal cancer (CRC) is a common malignant tumor of the digestive system. According to global cancer statistics in 2022, it ranked third in incidence and second in mortality, accounting for 9.3% of all cancer deaths worldwide. In my country, there were approximately 517,000 new cases in 2022, ranking second among malignant tumors, and approximately 240,000 deaths, ranking fourth in mortality. However, studies have found that while the 5-year relative survival rate for stage I colorectal cancer patients can reach 90%, the early diagnosis rate is only about 37%; nearly half of patients already have distant metastases at initial diagnosis, and the 5-year survival rate for stage IV patients is only 13%–14%. Therefore, improving the early diagnosis rate of CRC is key to reducing its mortality rate and is also the core breakthrough for improving prognosis.
[0003] Colonoscopy is considered the gold standard for colorectal cancer screening, but its invasive nature and requirement for thorough bowel preparation result in low participation rates in organized screening in my country. Fecal immunochemical testing (FIT) has high sensitivity for colorectal cancer diagnosis, but its ability to identify precancerous lesions is limited, and patient acceptance is low. Multiple studies suggest that blood tests could be used to improve CRC screening participation rates in conjunction with fecal testing or colonoscopy. While CT colonography offers the advantage of being non-invasive and has high detection sensitivity for colorectal cancer and precancerous lesions, its clinical use is limited by strict bowel preparation requirements, insufficient equipment and skilled personnel, and radiation exposure risks. Furthermore, although several multi-target FIT-DNA assays have been approved by the National Medical Products Administration, their high cost remains a point of contention, preventing widespread use in large-scale colorectal cancer screening. Therefore, there is still a need to develop a safe, highly compliant, sensitive, specific, and cost-effective method for early colorectal cancer detection.
[0004] Blood-based testing is a popular and ideal method for colorectal cancer screening due to its convenient sample acquisition, high acceptance rate, and repeatability. Surveys have shown that blood-based screening is more readily accepted than fecal testing. Currently, clinically available biomarkers for colorectal cancer diagnosis include carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9). However, their sensitivity and specificity are not ideal in the early stages, limiting their use primarily for postoperative monitoring. Research indicates that the body's immune system can sense the biological changes accompanying tumor development, triggering an immune response and producing autoantibodies against tumor-associated antigens (TAAbs). These autoantibodies, as serum biomarkers, offer advantages such as high sensitivity, stability, ease of detection, and low cost, showing promising potential for application in tumor diagnosis. It has been reported that autoantibodies often appear months or even years before the clinical diagnosis of a tumor. Significantly elevated autoantibody levels have been observed before clinical diagnosis in various cancers, including liver cancer, lung cancer, and colorectal cancer. Currently, the EarlyCDT®-Lung lung cancer diagnostic kit, composed of seven autoantibodies, and the thirteen-antibody lung cancer-related antibody detection kit (flow cytometry immunoassay) developed by a research team at the Hangzhou Institute of Medical Sciences, Chinese Academy of Sciences, have both been approved by the National Medical Products Administration. Both have played an important clinical role in screening high-risk populations and in early, accurate diagnosis. In conclusion, TAAbs testing is expected to become an important supplementary method for tumor diagnosis, providing more options for the early detection of tumors.
[0005] In existing studies on colorectal cancer-related autoantibodies (TAAbs) as serological biomarkers, the sensitivity and specificity of various indicators fluctuate significantly, and only a limited number of indicators possess good diagnostic performance. This phenomenon may be attributed to the heterogeneity of tumor cells and the differences in genetic and epigenetic backgrounds among patients, leading to diverse antitumor humoral immune responses. The low detection frequency of a single indicator in serum thus limits its diagnostic value. In contrast, combined detection strategies can effectively compensate for this deficiency. Currently, most identified colorectal cancer-related autoantibodies and their combinations rely on recombinant proteins as target antigens, which presents challenges such as complex preparation, difficulty in ensuring purity, and inconsistencies between spatial structures and native conformations, potentially leading to the ineffective exposure of some actual antigenic epitopes.
[0006] Studies have shown that the antigen-binding site (complementary site) of an antibody typically binds to an antigenic epitope consisting of 15-22 amino acids, with only 2-5 residues contributing the majority of the binding energy. Compared to full-length proteins, short-sequence antigenic epitopes theoretically elicit a more direct and efficient immune response, and are easier to synthesize and less expensive. Currently, some studies have explored using antigenic epitope peptides to replace full-length proteins in TAAbs detection, demonstrating good diagnostic value.
[0007] While existing research has preliminarily demonstrated the potential of antigenic peptides in the detection of tumor autoantibodies, there is currently a lack of systematic studies targeting colorectal cancer-related "antigen-peptide-autoantibody" biomarkers at the whole-human proteome level. Therefore, systematically screening and identifying antigenic peptides and their corresponding autoantibodies that can be used for colorectal cancer detection could provide new insights for the serological diagnosis of colorectal cancer and further enrich the research foundation in the field of precision oncology. Summary of the Invention
[0008] The purpose of this invention is to provide the application of reagents for detecting biomarkers in the preparation of products for the diagnosis of colorectal cancer.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: The application of reagents for detecting biomarkers in the preparation of products for the diagnosis of colorectal cancer, wherein the biomarkers are one or a combination of two or more of the following: autoantibodies against CCT6B antigen peptide, first HNRNPLL antigen peptide, DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide.
[0010] The reagent is an antigen.
[0011] The antigen is one or a combination of two or more of the following: CCT6B antigenic peptide, first HNRNPLL antigenic peptide, DNMT1 antigenic peptide, second HNRNPLL antigenic peptide, LIN9 antigenic peptide, LMO7 antigenic peptide, NOB1 antigenic peptide, STAU2 antigenic peptide, and TP53 antigenic peptide.
[0012] The sequence of the CCT6B antigen polypeptide is LLDVARTSLQTKVHAELADV; The sequence of the first HNRNPLL antigen polypeptide is DYTKPYLGRRDRGKGRQRQA; The sequence of the DNMT1 antigen polypeptide is PLAPGSDWRDLPNIEVRLSD; The sequence of the second HNRNPLL antigen peptide is HPSSFRHDGYGSHGPLLPLP. The sequence of the LIN9 antigen peptide is VIGTKVTARLRGVHDGLFTG; The sequence of the LMO7 antigen peptide is LDEELMVLSSNSMSLTTREP; The sequence of the NOB1 antigen polypeptide is TIREVVTEIRDKATRRRLAV; The sequence of the STAU2 antigen polypeptide is TIARELLMNGTSSTAEAIGL; The sequence of the TP53 antigen peptide is PLSQETFSDLWKLLPENNVL.
[0013] The biomarkers are one or a combination of two or more of the following: autoantibodies against DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide.
[0014] The biomarkers are a combination of autoantibodies against DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide.
[0015] The product is a protein chip, reagent kit, or formulation.
[0016] The kit is an ELISA detection kit, which includes an antigen for detecting biomarkers, the antigen being coated on a solid-phase carrier.
[0017] The ELISA test kit further includes any one or more combinations of positive control serum, negative control serum, blocking solution, sample diluent, secondary antibody, secondary antibody diluent, washing solution, coating solution, colorimetric solution, or stop solution; the solid phase carrier is made of polyvinyl chloride, polystyrene, polyacrylamide, or cellulose.
[0018] The samples are serum, plasma, interstitial fluid, saliva, or urine.
[0019] The second antibody is a horseradish peroxidase-labeled mouse anti-human IgG monoclonal antibody.
[0020] The method for preparing the solid-phase carrier coated with antigen peptides in an ELISA detection kit includes the following steps: (1) Coating antigen peptides: DNMT1 antigen peptide, HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide and TP53 antigen peptide were diluted with coating buffer and added to solid-phase carrier at 50 μL / well and incubated at 4 °C for more than 14 h; the coating concentration of DNMT1 antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide and STAU2 antigen peptide was 2 μg / mL, the coating concentration of HNRNPLL antigen peptide was 1 μg / mL, and the coating concentration of NOB1 antigen peptide and TP53 antigen peptide was 0.5 μg / mL; (2) Blocking: After incubation, remove the coating solution and then add the blocking solution to the solid support at 100 μL / well, and block at 37℃ for 2 h; (3) After sealing, remove the liquid in the solid carrier, wash with washing solution, and pat dry to obtain the solid carrier coated with antigen peptide.
[0021] The washing conditions were 300 μL washing solution / well, 10 s / cycle, and repeated 3 times.
[0022] The detection method of the ELISA test kit includes the following steps: (1) Primary antibody incubation: Dilute the primary antibody (e.g., serum) with the sample diluent at a ratio of 1:100, mix well, and add 50 μL / well to the solid-phase carrier coated with the antigen peptide. At the same time, set up blank control wells (e.g., sample diluent) and incubate at 37°C for 1 h. (2) Secondary antibody incubation: After the first antibody incubation is completed, discard the liquid in the solid-phase carrier, wash and pat dry, dilute the second antibody (horseradish peroxidase-labeled mouse anti-human IgG monoclonal antibody) at 1:5000, add 50 μL / well to the solid-phase carrier, and incubate at 37℃ for 1 h. (3) Color development and termination: After the second antibody incubation is completed, discard the liquid in the solid support, wash and pat dry, add 50 μL of color development solution to the solid support, place at room temperature in the dark for 5-20 min, and then add 25 μL of termination solution to the solid support. (4) Scanning: Wipe the bottom of the solid-phase carrier clean, and read the absorbance at dual wavelengths of 450 nm and 620 nm using an ELISA reader. Calculate the difference (OD450nm - OD620nm). If the OD value of the blank control well exceeds 0.1, the result is invalid and needs to be repeated.
[0023] The washing conditions were 300 μL washing solution / well, 10 s / cycle, and repeated 5 times.
[0024] Compared with the prior art, the present invention has the following beneficial effects: (1) Based on PhIP-Seq technology containing a full human proteome peptide library, this invention screens out potential antigenic peptides that can be used to detect or otherwise characterize cancer. It screens out early detection biomarkers for colorectal cancer through peptide microarray chip technology, and then verifies them through ELISA experiments. Finally, it screens out a group of combined detection biomarkers for colorectal cancer that can be used for colorectal cancer screening and diagnosis, including autoantibodies against CCT6B antigen peptide, first HNRNPLL antigen peptide, DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide. These can be used to assist in the diagnosis and detection of colorectal cancer and have good reference value. (2) The expression levels of autoantibodies against CCT6B antigen peptide, first HNRNPLL antigen peptide, DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide in the serum of colorectal cancer patients were significantly higher than those in normal individuals, and the differences were statistically significant. The expression levels of anti-colorectal cancer-related antigen peptides DNMT1, HNRNPLL, LIN9, and TP53 in human serum were detected. The expression levels of autoantibodies for LMO7, NOB1, STAU2, and TP53 can effectively diagnose and differentiate colorectal cancer from normal individuals. Verification has shown that when this invention combines seven biomarkers—autoantibodies for DNMT1 antigen peptide, HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide—to diagnose and differentiate colorectal cancer from normal individuals, the AUC value of the ROC curve is 0.83, indicating good differentiation performance and its potential use in the auxiliary diagnosis of colorectal cancer. (3) The kit of the present invention detects the expression levels of autoantibodies against DNMT1 antigen peptide, HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide and TP53 antigen peptide in human serum by indirect ELISA. It can effectively distinguish colorectal cancer patients from healthy controls and is helpful for the clinical diagnosis of colorectal cancer. (4) The test sample of the kit of the present invention is a biological liquid such as serum, which can avoid invasive diagnosis. The risk of colorectal cancer can be obtained by taking serum through minimally invasive means. The detection method has the characteristics of high sensitivity, strong specificity and low cost. It is also simple and quick to operate, and can provide a basis for the diagnosis of colorectal cancer. Attached Figure Description
[0025] Figure 1 Heatmap of positive expression of 26 candidate antigenic peptides screened by PhIP-Seq technology in colorectal cancer and healthy controls; Figure 2 The expression of autoantibodies against 11 antigenic peptides screened by peptide microarray in colorectal cancer and healthy controls; where CRC represents the colorectal cancer group and HC represents the healthy control group; ** represents P<0.01, *** represents P<0.001, and **** represents P<0.0001. Figure 3 The expression of autoantibodies against 11 antigenic peptides screened by peptide microarrays in early colorectal cancer and healthy controls; where CRC represents the colorectal cancer group and HC represents the healthy control group. * represent P <0.05, ** represent P <0.01, *** represent P <0.001, **** represent P <0.0001; Figure 4 The graph shows the expression levels of autoantibodies against 11 antigenic peptides in the validation group; where CRC represents the colorectal cancer group, HC represents the healthy control group, Normalized OD represents the optical density value, * represents P<0.05, ** represents P<0.01, **** represents P<0.0001, and ns represents no statistically significant difference. Figure 5 ROC curves for the use of autoantibodies against nine antigenic peptides in the validation group to diagnose colorectal cancer; Figure 6 The graph shows the expression levels of autoantibodies against nine antigenic peptides in the test group; where CRC represents the colorectal cancer group, HC represents the healthy control group, Normalized OD represents the optical density value, and * indicates P<0.05. *** **** represents P < 0.001; **** represents P < 0.0001. Figure 7 ROC curves for the diagnosis of colorectal cancer using autoantibodies against nine antigenic peptides in the test group; Figure 8Example diagrams for using the CRCPepAb tool; the left side shows single-sample analysis and visualization, and the right side shows batch analysis of multiple samples; Figure 9 In the middle, from left to right, are the ROC curves of the colorectal cancer diagnostic model on the training set, validation set, and test set. Detailed Implementation
[0026] The present invention will be further described below with reference to specific embodiments and accompanying drawings.
[0027] Example 1: Screening of candidate antigen peptides related to colorectal cancer based on PhIP-Seq technology 1. Experimental Samples This study collected serum samples from 50 patients with colorectal cancer and 50 healthy controls, all from the First Affiliated Hospital of Zhengzhou University. The colorectal cancer group consisted of patients with pathologically confirmed primary colorectal cancer who had not received any anti-tumor treatments such as surgery, radiotherapy, chemotherapy, targeted therapy, immunotherapy, or traditional Chinese medicine. The diagnosis period was from June 2021 to December 2023. In this group, there were 25 males (50.0%) and 25 females (50.0%), with a median age of 60.5 years and an interquartile range of 58–66 years. The healthy control group consisted of 25 males (50.0%) and 25 females (50.0%), with a median age of 61 years and an interquartile range of 58–65 years. This study was approved by the hospital's ethics committee and institutional review committee, and informed consent was obtained from all patients.
[0028] Serum collection: 5-10 mL of whole blood was collected from all subjects using red-tipped blood collection tubes. After standing at room temperature for 1 h, the samples were centrifuged at 4℃ and 3500 rpm for 5 min. The supernatant was collected, and each sample was aliquoted, labeled, and stored in a -80℃ freezer to avoid repeated freeze-thaw cycles.
[0029] 2. PhIP-Seq (phage immunoprecipitation sequencing) technology (1) Experimental methods 1) Serum pretreatment: Place each serum sample at 4℃ and centrifuge at 12000 rpm for 20 min. Take the supernatant, dilute it with PBST buffer at a ratio of 1:1000, add it to a pre-sealed 96-well plate, and set up negative control, blank control and positive control wells in each 96-well plate.
[0030] 2) Phage display peptide library incubation: Diluted and centrifuged phage display peptide library (Human Proteome, Shanghai Kangmaxinrui Biotechnology Co., Ltd.) was added to each well, ensuring a library diversity of 16000× phages per well. The 96-well plate was then sealed with the matching sealing cap and incubated overnight at 4°C by rotation.
[0031] 3) Magnetic bead precipitation and enrichment: After incubation, 20 µL of washed Protein G magnetic beads were added to each well, and incubated at 4°C for 4 h. Afterwards, the mixture was diluted with 500 mL of water. g Centrifuge for 3 min, place the reaction plate on a magnetic rack to adsorb Protein G beads, and discard the original solution; add 200 µL of 0.1% IGEPAL to each well. ® Wash the CA-630 beads with TBS buffer, resuspend them by pipetting, and transfer them to a new 96-well PCR plate. Repeat the washing process twice, then resuspend the magnetic beads in 40 µL of enzyme-free water and transfer them to a new 96-well PCR plate.
[0032] 4) PCR library preparation: After sealing the plate, centrifuge briefly for 10 seconds and heat at 95℃ for 10 min in a PCR instrument. Two rounds of PCR were used for sequencing library preparation. The first round used universal primers matching the phage sequence, and the second round used primers with barcode sequences and sequencing adaptation sequences for amplification. After quality monitoring of the amplification products by gel electrophoresis, the target band was obtained by gel electrophoresis and gel excision.
[0033] 5) High-throughput sequencing: Based on the Illumina sequencing platform, high-throughput sequencing was performed on the library-constructed samples. The sequencing read length was 150 bp. Bidirectional sequencing was used to obtain the full-length sequence through paired-end sequencing results, and the number of reads for each peptide was used as the basic data.
[0034] (2) Data processing: 1) Sequencing data alignment: After sequencing, the fastq files are aligned to the reference peptide library using the bowtie2 software (https: / / bowtie-bio.sourceforge.net / bowtie2 / manual.shtml) to obtain the raw read counts for each peptide.
[0035] 2) Data Standardization and Quality Control: First, the raw sequencing data of each sample were normalized to a uniform 1.25 M reads. Then, the normalized reads (NR) of each peptide in each sample were calculated. Further, the sequencing abundance of each peptide in the blank control was combined to calculate the enrichment factor (EF) of each peptide in each sample, which characterizes the binding strength between serum antibodies and peptides. Finally, based on the generalized Poisson distribution model, the statistical significance of each peptide was calculated. P value.
[0036] 3) Identification of positive reactive peptides: Based on preset multiple screening thresholds (normalized reading > 15, enrichment factor > 5, -log10 (p-value) > 36.44), peptides showing a positive reaction in each sample were screened. The -log10 (p-value) threshold was set based on the detection results of the mock IP group. Mock IP refers to an immunoprecipitation control experiment without serum, designed to assess the non-specific binding of peptides to the experimental system and establish a baseline level of background signal.
[0037] 4) Differential analysis: In order to eliminate false positives caused by a single statistical method, this study combined the Wilcoxon test and the edgeR algorithm for differential analysis.
[0038] (3) Experimental results The intersection of the results from the two differential analysis methods yielded a set of 1528 upregulated antigenic peptides. Based on this set, a further screening strategy was implemented: First, using the number of positive reaction differences as the screening index, antigenic peptides with ≥6 positive reaction differences between the colorectal cancer group and the healthy control group were screened, resulting in 44 candidate peptide sequences. A systematic literature review was then conducted on the genes or proteins corresponding to these 44 peptides to assess their correlation with tumor development and progression. The literature review results showed that there were 25 genes / proteins (corresponding to 26 antigenic peptides, with positive reaction details as follows) Figure 1(As shown) Studies have reported their association with tumors. The 26 antigenic peptides and their corresponding gene / protein information are as follows: #509162 (ANO8), #535266 (ARHGEF12), #205985 (ATM), #381773 (BOD1L1), #415727 (CCT6B), #123113 (DNMT1), #181431 (EXOSC10), #117487 (FECH), #402711 (HNRNPLL), #331699 (KTN1), #273623 (LIN9), #403325 (LMO7), #6 40602 (MUC4), #565957 (NOB1), #595736 (PCLO), #314736 (PDZD4), #304675 (PDZRN4), #611799 (PLEKHA1), #524521 (STAU2), #617666 (STAU2), #250520 (TBRG1), #125037 (TMOD1), #436516 (TONSL), #83035 (TP53), #589786 (TRPV2), #244431 (ZNF630).
[0039] Of these, 14 antigenic peptides did not show a positive reaction in the healthy control group, namely #83035 (TP53), #117487 (FECH), #205985 (ATM), #244431 (ZNF630), #250520 (TBRG1), #273623 (LIN9), #381773 (BOD1L1), #403325 (LMO7), #524521 (STAU2), #535266 (ARHGEF12), #565957 (NOB1), #589786 (TRPV2), #611799 (PLEKHA1), and #640602 (MUC4).
[0040] Example 2: Screening for biomarkers for colorectal cancer diagnosis based on peptide microarray chip technology For the 26 antigenic peptides obtained in Example 1, they were first truncated and optimized into short peptides of 20 amino acids in length for easy synthesis based on the shingled method and B-cell antigenic epitope prediction method. Secondly, the optimized antigenic peptides were used to construct a peptide microarray chip for serum detection to further screen biomarkers for colorectal cancer diagnosis.
[0041] 1. Experimental Samples This study collected serum samples from 327 patients with colorectal cancer and 327 healthy controls. All samples were obtained from the First Affiliated Hospital of Zhengzhou University, using the same selection criteria as in Example 1. In the colorectal cancer group, there were 173 males (52.91%) and 154 females (47.09%), with a median age of 62 years and an interquartile range of 53.5–71 years. In the healthy control group, there were 189 males (57.80%) and 138 females (42.20%), with a median age of 60 years and an interquartile range of 52–68.5 years. Among the colorectal cancer patients, 152 (46.48%) were in stage I-II (early stage), and 107 (32.72%) were in stage III-IV.
[0042] Serum collection: 5-10 mL of whole blood was collected from all subjects using red-tipped blood collection tubes. After standing at room temperature for 1 h, the samples were centrifuged at 4℃ and 3500 rpm for 5 min. The supernatant was collected, and each sample was aliquoted, labeled, and stored in a -80℃ freezer to avoid repeated freeze-thaw cycles.
[0043] 2. Sequence truncation and optimization design of candidate antigen peptides For the 26 candidate antigen peptides obtained in Example 1, this study used two strategies for truncation and optimization design, and finally obtained 59 antigen peptides with a length of 20 amino acids.
[0044] First, based on sequence annotation information of protein functional regions in the UniProt database (https: / / www.uniprot.org / ) and the Conserved Domain Database (CDD) (https: / / www.ncbi.nlm.nih.gov / Structure / cdd / wrpsb.cgi), seven candidate polypeptide fragments were identified, located in the known functional regions of their corresponding proteins: #535266 (ARHGEF12), #205985 (ATM), #123113 (DNMT1), #273623 (LIN9), #565957 (NOB1), #595736 (PCLO), and #83035 (TP53). Furthermore, the full-length protein corresponding to candidate antigen polypeptide #640602 (MUC4) has been shown in previous studies to have the potential as a serum diagnostic biomarker for colorectal cancer. Therefore, for the above 8 candidate antigen peptides, a shingled design strategy was adopted, and each peptide was shortened into 4 optimized antigen peptides with a length of 20 amino acids and an overlap of 8 amino acids between adjacent peptide segments.
[0045] For the remaining 18 candidate antigen peptides, this study comprehensively utilized five B-cell epitope prediction algorithms (Table 1) for targeted truncation design. First, sequence information and predicted structural models of each target protein were obtained from UniProt, NCBI (https: / / www.ncbi.nlm.nih.gov / ), and the AlphaFold protein structure database (https: / / alphafold.com / ). Subsequently, the prediction results were compared with the candidate peptide sequences to screen for core regions commonly identified by multiple algorithms. Finally, the core sequences were extended 20 amino acids to both the N-terminus and C-terminus to standardize peptide length.
[0046] In summary, a total of 59 antigenic polypeptides with a length of 20 amino acids were obtained. Their source antigens and sequence ranges are as follows: ANO8 (194-213), ARHGEF12 (804-823, 816-835, 828-847, 840-859), ATM (2685-2704, 2697-2716), BOD1L1 (1-20, 37-56), CCT6B (130-149, 148-167, 166-185), DNMT1 (1429-1448, 1441-1460, 1453-1472, 1465-1484), EXOSC10 (794-813), FECH (55-74), HNRNPLL (258-277, 270-289, 282-301), KTN1 (972-991, 984-1003), LIN9 (209-228,221-240, 233-252, 245-264), LMO7 (1275-1294, 1293-1312), MUC4 (2745-2764, 2757-2776, 2769-2788, 2781-2800), NOB1 (29-48, 41-60, 53-72, 65-84), PCLO (841-860,853-872, 865-884, 877-896), PDZD4 (129-148,140-159), PDZRN4 (483-502), PLEKHA1 (287-306), STAU2 (463-482, 472-491), STAU2 (383-402), TBRG1 (360-379), TMOD1 (228-247), TONSL (1027-1046), TP53 (1-20, 13-32, 25-44, 37-56), TRPV2 (510-529), ZNF630 (197-216, 209-228, 221-240).
[0047] 3. Detection using peptide microarray chips (1) Experimental reagents: 1) 3% BSA blocking solution: Add 7 mL of 1×PBS solution to 3 mL of 10% BSA, mix well, and incubate at 4℃; 2) Serum diluent: Add 9 mL of 1×PBST solution to 1 mL of 10% BSA, mix well, and incubate at 4℃; 3) Washing solution: 1×PBST, stored at room temperature; 4) Secondary antibody incubation solution: Cy3 fluorescently labeled anti-human IgG secondary antibody; 5) 59 antigenic polypeptides: Polypeptides were synthesized using the Fmoc solid-phase synthesis method.
[0048] (2) Preparation of peptide microarray chips: 1) Preparation of antigen peptide samples: Dilute the dissolved antigen peptide with sterile water to the same concentration, and take an equal volume of crystal core. ® Mix protein chip spotting solution B at a 1:1 ratio, transfer an appropriate amount of the mixture to a 96-well polypropylene plate, and store at 4°C for later use.
[0049] 2) Spotting environment control: The Nano-Plotter™ 2.1 multi-channel micro-chip spotter was used for spotting. The humidifier and temperature control equipment were turned on 30 minutes before spotting to balance the environmental conditions. The humidity was set to 60% and the temperature was set to 25℃.
[0050] 3) Spotting program setting: Design the spotting program based on the chip layout and optimize the spotting parameters based on the results of preliminary experiments. Before the actual spotting, calibrate the position parameters using the test chip and sterile water and confirm that the instrument is in normal condition and the spotting points are consistent in shape.
[0051] 4) Formal Spotting and Fixation: Remove the epoxy-based sample from the 4℃ refrigerator, allow it to equilibrate to room temperature for 30 minutes, then open the package and place it sequentially in a dustproof glass cover. Place the 96-well sample plate in the spotting instrument and start the spotting program. After spotting, adjust the temperature to 37℃ and let it stand for 2 hours.
[0052] 5) Applying and storing the sample fences: After the settling period, use the CrystalCore® chip fence application tool to apply 12 sample fences to the sample application surface of the chip, ensuring that the fence positions accurately correspond to the sample application areas. Place the prepared peptide chip in a chip cassette, vacuum it, and store it at -80℃ in the dark.
[0053] (3) Experimental methods 1) Chip preparation: Remove the frozen peptide chip from -80℃ and gradually warm it to -20℃, 4℃ and room temperature while keeping it sealed. Take the CrystalCore® protein chip reaction box, add 100 µL of ultrapure water to maintain humidity inside the box, and place the chip with the sample spotting side facing up in the box.
[0054] 2) Blocking treatment: Using 3% BSA as the blocking solution, slowly add 30 µL to each sample enclosure, tighten the hybridization box lid, and place it on a shaker for 2 h at room temperature.
[0055] 3) Cleaning and drying: Place the chip in a cleaner and clean it 3 times with 1×PBST washing solution for 4 minutes each time, with a cleaning intensity of 5; after cleaning, transfer it to a centrifugal drying chamber for drying.
[0056] 4) Serum incubation: After removing the serum from -80℃, slowly rehydrate it at 4℃, vortex to mix, and centrifuge. Dilute the serum at a ratio of 1:50, and slowly add 25 µL of diluted serum to each sample enclosure. Place the chip in a hybridization instrument and incubate at 37℃ and 10 rpm for 1 h.
[0057] 5) Cleaning and drying: Repeat step 3).
[0058] 6) Secondary antibody incubation: Add 25 µL of Cy3-labeled secondary antibody reaction solution diluted 1:1000 to each sample enclosure under light-protected conditions, and incubate at 37°C and 10 rpm for 1 h in a hybridization apparatus under light-protected conditions.
[0059] 7) Cleaning and drying: Clean and dry as in step 3). Adjust the cleaning conditions as follows: wash twice with 1×PBST solution for 4 min each time, with a strength of 5; wash once with ultrapure water for 2 min, with a strength of 3.
[0060] 8) Fluorescence scanning: Preheat the chip scanner for at least 10 minutes, place the dried chip into the scanner, use LuxScan 3.0 software to extract fluorescence signal values, and save the original image and data files.
[0061] (3) Data processing: 1) Repeatability assessment: The same serum samples are used for microarray detection within the same batch or between different batches. Based on the normality test results of the data distribution, Spearman's or Pearson's correlation coefficient is selected for analysis to examine the repeatability and stability of the detection.
[0062] 2) Array Validity Determination: Validity is determined based on the signal values of the quality control points in each subarray of the chip. The validity criteria are: the difference between the foreground and background values of a positive quality control point is greater than 20,000, and the difference between the foreground and background values of a blank quality control point is less than 100. If any subarray does not meet the above criteria, it is considered invalid, and the corresponding sample needs to be retested.
[0063] 3) Repetition point consistency control: Three technical repetition points are set for each antigenic peptide, and the coefficient of variation of the foreground signal value and the background signal value between the repetition points are calculated. If the coefficient of variation of the foreground value or the background value of any antigenic peptide exceeds 40%, the test result of that sample is deemed invalid and retesting is required.
[0064] 4) In-chip signal correction: To eliminate signal shifts caused by background differences at different sampling locations within the chip, a background normalization method is used to correct the original data. The signal value of each sampling point is expressed as the signal-to-noise ratio (SNR). The arithmetic mean of the SNRs from three technical replicates of the same antigen peptide is then taken as the final detection value for the autoantibody level of that antigen peptide. The formula for calculating SNR is: Signal-to-noise ratio (SNR) = Foreground value / Background value.
[0065] 5) Inter-chip signal standardization: To reduce systematic errors introduced by testing batches or instrument operation between different chips, the signal strength between chips is standardized. The specific steps are as follows: ① Construct a reference dataset: Based on the SNR of each antigenic peptide in the 12th subarray (quality control serum) of each chip, summarize the SNR of the same antigenic peptide in all chips, and take its arithmetic mean as the reference signal intensity of the antigenic peptide to form a reference dataset.
[0066] ② Low signal peptide screening: Antigen peptides with a mean SNR of less than 2.00 in the reference dataset are removed, and the remaining peptides are used for subsequent standardization calculations.
[0067] ③ Outlier handling: Calculate the Pearson correlation coefficient between the 12th subarray of each chip and the reference dataset. If the correlation coefficient is lower than 0.96, outliers that deviate from the diagonal are gradually removed until the correlation coefficient meets the requirements.
[0068] ④ Determination of Standardization Factor: Using the `lm` function in R, a linear regression model passing through the origin is constructed with the signal strength of the reference dataset as the dependent variable and the signal strength of the 12th subarray of each chip as the independent variable. The slope of the model is extracted using the `coef` function and used as the standardization factor for this chip. The calculation formula is as follows: Model Ni=lm (Reference Signal Intensities∼Slide i Block - 1) Factor Ni = coef (Model Ni) Where: Reference_SNR is the SNR value of the reference dataset, and Slide i Block is the SNR value of the 12th subarray in the i-th slide.
[0069] ⑤ Signal value standardization: Multiply the original SNR values of all antigen peptides on each chip by the corresponding standardization factor to obtain the standardized signal values, which are then used for subsequent statistical analysis.
[0070] To screen autoantibodies of antigenic peptides with diagnostic potential for colorectal cancer (especially in the early stages), this study established the following screening criteria: First, the expression level of the autoantibody must be statistically significant between the colorectal cancer group and the healthy control group, as well as between the early colorectal cancer group and the healthy control group; Second, to ensure that the autoantibody has sufficient immunogenicity, its median SNR in the colorectal cancer group must be higher than 1.8.
[0071] (4) Experimental results: After screening, 11 autoantibodies against antigenic peptides with potential early diagnostic value were obtained: anti_PEP01, anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP27, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53. The autoantibodies against these antigenic peptides are respectively: PEP01 antigenic peptide (derived from ANO8, sequence interval 194-213), PEP11 antigenic peptide (derived from CCT6B, sequence interval 148-167), PEP13 antigenic peptide (derived from DNMT1, sequence interval 1429-1448), PEP19 antigenic peptide (derived from HNRNPLL, sequence interval 258-277), PEP21 antigenic peptide (derived from HNRNPLL, sequence interval 282-301), and PEP25 anti_PEP53. Autoantibodies against the original polypeptide (derived from LIN9, sequence range 221-240), PEP27 antigen polypeptide (derived from LIN9, sequence range 245-264), PEP28 antigen polypeptide (derived from LMO7, sequence range 1275-1294), PEP34 antigen polypeptide (derived from NOB1, sequence range 29-48), PEP46 antigen polypeptide (derived from STAU2, sequence range 463-482), and PEP53 antigen polypeptide (derived from TP53, sequence range 13-32) were identified. The expression of these 11 antigen polypeptide autoantibodies in the colorectal cancer group and the healthy control group is shown below. Figure 2 As shown, the expression of autoantibodies against 11 antigenic peptides in the early colorectal cancer group and the healthy control group is as follows: Figure 3 As shown.
[0072] Depend on Figure 2 , Figure 3 It can be seen that the expression levels of autoantibodies of the above 11 antigenic peptides in the serum of the colorectal cancer group / early colorectal cancer group were higher than those in the healthy control group.
[0073] Example 3: ELISA detection of serum autoantibodies against colorectal cancer-associated antigen peptides The expression levels of autoantibodies against the 11 anti-colorectal cancer-related antigen peptides screened in Example 2 were further detected in a large sample of human serum using an indirect enzyme-linked immunosorbent assay (ELISA).
[0074] 1. Experimental Samples The sample for this part of the study is the same as that in Example 2, including 327 colorectal cancer samples and 327 healthy controls, which serve as the validation group.
[0075] Serum collection: 5-10 mL of whole blood was collected from all subjects using red-tipped blood collection tubes. After standing at room temperature for 1 h, the samples were centrifuged at 3500 rpm for 5 min at 4℃. The supernatant was collected, and each sample was aliquoted, labeled, and stored in a -80℃ freezer to avoid repeated freeze-thaw cycles.
[0076] 2. Experimental reagents (1) Blocking solution: 3 mL of 10% BSA, add 7 mL of 1×PBS solution, mix well, and place on ice.
[0077] (2) Serum incubation solution: 1 mL of 10% BSA, add 9 mL of 1×PBST solution, mix well, and place on ice.
[0078] (3) Cleaning solution: 1×PBST, stored at 4℃.
[0079] (4) Secondary antibody incubation solution: horseradish peroxide (HRP) labeled anti-human IgG antibody.
[0080] 3. Experimental methods: (1) Coating: The 11 antigenic polypeptide sequences screened in Example 2 above (chemically synthesized by Sangon Biotech (Shanghai) Co., Ltd.) were coated respectively. The following antigenic polypeptides were coated at a concentration of 2 µg / mL: PEP11, PEP13, PEP25, PEP28, PEP46; the following antigenic polypeptides were coated at a concentration of 1 µg / mL: PEP01, PEP21, PEP27; and the following antigenic polypeptides were coated at a concentration of 0.5 µg / mL: PEP19, PEP34, PEP53; 50 μL / well, overnight at 4℃.
[0081] (2) Blocking: 2% BSA in PBST (PBS, Tween20) solution, 100 μL / well, 37℃ half water bath for 2 h.
[0082] (3) Washing: Wash 3 times with 300 μL / well PBST.
[0083] (4) Primary antibody incubation: The serum to be tested is diluted with PBST containing 1% BSA at a ratio of 1:100, 50 μL / well, and incubated in a half-water bath at 37°C for 1 h.
[0084] (5) Washing: Wash 5 times with 300 μL / well PBST.
[0085] (6) Secondary antibody incubation: HRP-labeled anti-human IgG was diluted with PBST containing 1% BSA at a ratio of 1:5000, 50 μL / well, and incubated in a half-water bath at 37°C for 1 h.
[0086] (7) Washing: Wash 5 times with 300 μL / well PBST.
[0087] (8) Color development: TMB color development system, mix solution A (Solepro, Beijing) and solution B in a 1:1 ratio, 25 μL / well, room temperature and protected from light, to achieve the expected color.
[0088] (9) Termination: Measure absorbance within 10 min after terminating with 25 μL of 10% concentrated sulfuric acid per well.
[0089] (10) Measure absorbance: using OD 450 -OD 620 The relative OD values were calculated, and then the blank control was subtracted. The detection values between different ELISA plates were normalized based on the detection values of the quality control wells, and then subsequent data processing was performed (for details of the data processing method, please refer to the "4. Data Processing" section below).
[0090] 4. Data Processing For the data obtained from the ELISA reader, the OD values at 450 nm wavelength in each ELISA plate were first processed as follows: the OD value at 620 nm wavelength was subtracted, and then the average OD value of the two blank wells in each plate was subtracted to eliminate the influence of background signals. The result was then used as the detection signal value of the sample. Next, quality control and standardization were performed using the following methods: (1) Sample quality control: Two blank control wells are set in each ELISA plate to monitor background noise and non-specific binding. If the OD value of the blank control well is higher than 0.1, the samples in the plate need to be retested to ensure the reliability of the data.
[0091] (2) Standardization: Each ELISA plate has 6 control wells for standardizing OD values between different ELISA plates. For autoantibodies against the same antigenic peptide, the mean, standard deviation, and CV of the OD values of the control samples in each ELISA plate are calculated. If the CV value of a control sample well exceeds 15%, the test result of that well is discarded. For the same detection index, the mean of all control wells of the ELISA plates is used as a reference set. The test values of each plate are standardized based on the reference set to eliminate systematic errors between different plates.
[0092] (3) Statistical Analysis: The diagnostic value of autoantibodies against single antigenic peptides for colorectal cancer was evaluated using ROC curves. The cutoff value was the OD value at which specificity was greater than 90% and the Youden index was maximized. The corresponding AUC, sensitivity, specificity, positive predictive value, negative predictive value, accuracy, precision, and Youden index were calculated and reported. Furthermore, the Delong test was used to compare the differences in AUC values between the two groups. P A value <0.05 indicates a statistically significant difference in AUC between the two groups.
[0093] 5. Experimental Results The expression levels of autoantibodies against 11 antigenic peptides in the colorectal cancer group and the healthy control group were verified. Figure 4 As shown, after ELISA verification, autoantibodies against nine antigenic peptides with statistically significant differences were obtained: anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53.
[0094] Depend on Figure 4 It can be seen that the expression levels of anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53 showed statistical differences.
[0095] Based on the expression levels of autoantibodies against nine antigenic peptides in the colorectal cancer group and the healthy control group, ROC curves were constructed to differentiate colorectal cancer from healthy controls using these autoantibodies. The results are as follows: Figure 5 As shown.
[0096] Depend on Figure 5 The diagnostic value of the autoantibodies anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53 of the above nine antigenic peptides in the validation group samples is shown in Table 2.
[0097] Table 2 shows that the AUC range of individual autoantibodies against the nine antigenic peptides was 0.55-0.73, the sensitivity range was 12.84%-44.95%, and the specificity range was 90.21%-94.19%. Among them, anti-PEP53 had the highest AUC value of 0.73 (95%). CIThe AUC values were 0.69–0.77, with a sensitivity of 44.95%, specificity of 94.19%, positive predictive value of 88.55%, negative predictive value of 63.11%, accuracy of 69.57%, and Youden's index of 0.39. The AUC values for anti-PEP11 and anti-PEP13 were the same, both 0.55 (95%). CI The sensitivities (0.50-0.59) were 12.84% and 17.13%, respectively; the specificities were 93.27% and 92.35%, respectively; the positive predictive values were 65.63% and 69.14%, respectively; the negative predictive values were 51.69% and 52.71%, respectively; and the accuracy was 53.06% and 54.74%, respectively.
[0098] Note: AUC: Area under the receiver operating characteristic curve. 95% CI : 95% confidence interval, Sen: sensitivity, Spe: specificity, Pr+: positive predictive value, Pr-: negative predictive value, Acc: accuracy, YI: Youden index.
[0099] Example 4: Assessment of the ability of autoantibodies against 9 antigenic peptides for the diagnosis of colorectal cancer. This study collected serum samples from 143 patients with colorectal cancer and 143 healthy controls. All samples were obtained from the First Affiliated Hospital of Zhengzhou University, using the same selection criteria as in Example 1. In the colorectal cancer group, there were 79 males (55.24%) and 64 females (44.76%), with a median age of 63 years and an interquartile range of 56–71 years. In the healthy control group, there were 90 males (62.94%) and 53 females (37.06%), with a median age of 64 years and an interquartile range of 55–70 years. Among the colorectal cancer patients, 59 (41.26%) were in stage I-II (early stage), and 52 (36.36%) were in stage III-IV.
[0100] Serum collection: 5–10 mL of whole blood was collected from all subjects using red-tipped blood collection tubes. After being left at room temperature for 2 hours, the serum was collected at 1000 mL / min. g Centrifuge for 15 minutes, collect the supernatant, aliquot each sample into several portions, label them, and store them in a -80℃ freezer to avoid repeated freeze-thaw cycles.
[0101] To further test the diagnostic capabilities of the nine antigenic peptides for colorectal cancer using autoantibodies, serum samples from the test group were analyzed to determine the expression levels of the nine antigenic peptides' autoantibodies. Figure 6 As shown.
[0102] Depend on Figure 6It can be seen that the expression levels of autoantibodies for the nine antigenic peptides all showed statistical differences.
[0103] ROC curves were constructed to differentiate colorectal cancer from healthy controls using autoantibodies containing nine antigenic peptides. The results are as follows: Figure 7 As shown.
[0104] Depend on Figure 7 The diagnostic value of the autoantibodies anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53 of the above nine antigenic peptides in the test group samples is shown in Table 3.
[0105] Table 3. Diagnostic value of autoantibodies against 9 antigenic peptides for colorectal cancer (test group) Note: AUC: Area under the receiver operating characteristic curve. 95% CI : 95% confidence interval, Sen: sensitivity, Spe: specificity, Pr+: positive predictive value, Pr-: negative predictive value, Acc: accuracy, YI: Youden index.
[0106] Table 3 shows that the AUC values of the autoantibodies against the nine antigenic peptides ranged from 0.57 to 0.75, the sensitivity ranged from 9.09% to 50.35%, and the specificity ranged from 90.21% to 95.10%. Among them, anti_PEP53 had the highest AUC value of 0.75 (95%). CI The AUC values were 0.69-0.81, with a sensitivity of 50.35% and a specificity of 94.41%. Anti-PEP13 had the lowest AUC value, at 0.57 (95%). CI 0.50-0.64).
[0107] Example 5: Construction, evaluation, and development of application tools for colorectal cancer diagnostic models After ELISA validation, the autoantibodies (anti_PEP11, anti_PEP13, anti_PEP19, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, anti_PEP53) of the nine antigenic peptides with statistically significant differences obtained in Example 4 were subjected to optimal feature subset screening based on the collinearity feature elimination and recursive feature selection algorithm. A total of seven autoantibodies of antigenic peptides were obtained: anti_PEP13, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53. The autoantibodies against the above antigenic peptides are those of PEP13 antigenic peptide (derived from DNMT1, sequence interval 1429-1448), PEP21 antigenic peptide (derived from HNRNPLL, sequence interval 282-301), PEP25 antigenic peptide (derived from LIN9, sequence interval 221-240), PEP28 antigenic peptide (derived from LMO7, sequence interval 1275-1294), PEP34 antigenic peptide (derived from NOB1, sequence interval 29-48), PEP46 antigenic peptide (derived from STAU2, sequence interval 463-482), and PEP53 antigenic peptide (derived from TP53, sequence interval 13-32).
[0108] The amino acid sequence of the PEP11 antigen polypeptide is: LLDVARTSLQTKVHAELADV The amino acid sequence of the PEP19 antigenic polypeptide (first HNRNPLL antigenic polypeptide) is: DYTKPYLGRRDRGKGRQRQA The amino acid sequence of the PEP13 antigenic polypeptide is: PLAPGSDWRDLPNIEVRLSD (SEQ ID NO.1). The amino acid sequence of the PEP21 antigenic polypeptide (second HNRNPLL antigenic polypeptide) is: HPSSFRHDGYGSHGPLLPLP (SEQ ID NO.2). The amino acid sequence of the PEP25 antigenic polypeptide is: VIGTKVTARLRGVHDGLFTG (SEQ ID NO.3); The amino acid sequence of the PEP28 antigenic polypeptide is: LDEELMVLSSNSMSLTTREP (SEQ ID NO.4); The amino acid sequence of the PEP34 antigen polypeptide is: TIREVVTEIRDKATRRRLAV (SEQ ID NO.5); The amino acid sequence of the PEP46 antigen polypeptide is: TIARELLMNGTSSTAEAIGL (SEQ ID NO.6). The amino acid sequence of the PEP53 antigenic polypeptide is: PLSQETFSDLWKLLPENNVL (SEQ ID NO.7). Subsequently, a multilayer perceptron model was used to construct a model of autoantibodies against seven antigenic peptides, resulting in a colorectal cancer diagnostic model.
[0109] The method for constructing a colorectal cancer diagnostic model is as follows: 1. Training dataset setup: Training and validation set samples: Serum samples from 327 colorectal cancer cases and 327 healthy controls in Example 2 were randomly divided into training and validation sets at a ratio of 7:3. The training set was used to build the model, and the validation set was used for internal validation and parameter tuning of the model.
[0110] Test set samples: Serum samples from 143 colorectal cancer patients and 143 healthy controls in Example 4 were used for external validation of the model.
[0111] 2. Construction of a diagnostic model for colorectal cancer: The detection values of anti_PEP13, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53 autoantibodies in serum samples from the training set were used to construct a colorectal cancer diagnostic model. The construction process is as follows: (1) Construct a dataset of autoantibody expression levels of antigenic peptides of anti_PEP13, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53 for each sample in the training set. (2) Construct a binary label for each sample in the training set (whether or not the patient has colorectal cancer, 0 for healthy and 1 for the disease). (3) The ten-fold cross-validation method was used to construct and tune the model on the training set samples to determine the optimal parameters of the multilayer perceptron model: size = 7, thus obtaining the model for colorectal cancer diagnosis. The input features of the colorectal cancer diagnosis model are the expression levels of autoantibodies anti_PEP13, anti_PEP21, anti_PEP25, anti_PEP28, anti_PEP34, anti_PEP46, and anti_PEP53; the output label is the probability value of whether or not one has colorectal cancer. P value).
[0112] whenP When the value is ≥0.5, the sample is initially identified as a suspected colorectal cancer sample; when P When the value is <0.5, it is preliminarily determined to be a normal sample.
[0113] It should be noted that the detection results in this embodiment are only used as intermediate indicators to assist in diagnosis and should not be used as the sole basis for confirming colorectal cancer. In actual clinical decision-making, a comprehensive evaluation by a physician should be conducted, taking into account the patient's clinical symptoms, imaging findings, and histopathological examinations, to make a final diagnosis. To improve the usability and scalability of the colorectal cancer diagnostic model, an online prediction platform for the colorectal cancer diagnostic model (https: / / cyfan.shinyapps.io / CRCPepAb / ) was built using R software. Based on the shiny (version 1.11.1), shinydashboardPlus (version 2.0.6), and dashboardthemes (version 1.1.6) packages, the MLP_model was deployed to the shinyapps (https: / / www.shinyapps.io / ) online server for access and use. An example prediction diagram of the CRCPepAb tool is shown below. Figure 8 As shown.
[0114] On the single-sample analysis page, after inputting the feature data of a single sample, the tool outputs the probability of the corresponding patient having colorectal cancer. In the multi-sample batch analysis area, it supports uploading batch files containing multiple samples. The tool automatically calculates and returns the probability prediction results for each sample, and supports downloading in CSV format.
[0115] The implementation code for building the diagnostic model and online prediction platform using R software is as follows: library(caret) library(pROC) train_ control<- trainControl(method = "cv", number = 10, search = "grid", classProbs = TRUE, summaryFunction = twoClassSummary, savePredictions = "final") MLP_model<- train(GROUP ~ ., data = train _data, method = "mlp", trControl = train_control, tuneGrid = mlp_grid, metric = "ROC", preProcess = c("center", "scale")) train_predictions<- predict(mlp_model, newdata = train_data, type = "prob")[, 2] train_actual<- train_data[[group_col]] train_roc<- roc(response = train_actual, predictor = train_predictions, levels = levels, direction = "<") val_predictions<- predict(mlp_model, newdata = val_data, type = "prob")[, 2] val_actual<- val_data[[group_col]] val_roc<- roc(response = val_actual, predictor = val_predictions, levels = levels, direction = "<") test_predictions<- predict(mlp_model, newdata = test_data, type = "prob")[, 2] test_actual<- test_data[[group_col]] test_roc<- roc(response = test_actual, predictor = test_predictions, levels = levels, direction = "<") The ROC curves and confusion matrix of the colorectal cancer diagnostic model of this invention on the training set, validation set, and test set are shown below. Figure 9 As shown.
[0116] Depend on Figure 9 It can be seen that the AUC of this model on the training set, validation set, and test set is 0.87 (95%). CI 0.84-0.90), 0.83 (95%) CI 0.77-0.89) and 0.83 (95%) CI The values of 0.79-0.88 indicate that the model can be well used for the diagnosis of colorectal cancer.
Claims
1. The application of reagents for detecting biomarkers in the preparation of products for the diagnosis of colorectal cancer, characterized in that, The biomarker is one or a combination of two or more of the following: autoantibody of CCT6B antigen peptide, autoantibody of first HNRNPLL antigen peptide, autoantibody of DNMT1 antigen peptide, autoantibody of second HNRNPLL antigen peptide, autoantibody of LIN9 antigen peptide, autoantibody of LMO7 antigen peptide, autoantibody of NOB1 antigen peptide, autoantibody of STAU2 antigen peptide, and autoantibody of TP53 antigen peptide.
2. The application as described in claim 1, characterized in that, The reagent is an antigen.
3. The application as described in claim 2, characterized in that, The antigen is one or a combination of two or more of the following: CCT6B antigenic peptide, first HNRNPLL antigenic peptide, DNMT1 antigenic peptide, second HNRNPLL antigenic peptide, LIN9 antigenic peptide, LMO7 antigenic peptide, NOB1 antigenic peptide, STAU2 antigenic peptide, and TP53 antigenic peptide.
4. The application as described in claim 3, characterized in that, The sequence of the CCT6B antigen polypeptide is LLDVARTSLQTKVHAELADV; The sequence of the first HNRNPLL antigen polypeptide is DYTKPYLGRRDRGKGRQRQA; The sequence of the DNMT1 antigen polypeptide is PLAPGSDWRDLPNIEVRLSD; The sequence of the second HNRNPLL antigen peptide is HPSSFRHDGYGSHGPLLPLP. The sequence of the LIN9 antigen peptide is VIGTKVTARLRGVHDGLFTG; The sequence of the LMO7 antigen peptide is LDEELMVLSSNSMSLTTREP; The sequence of the NOB1 antigen polypeptide is TIREVVTEIRDKATRRRLAV; The sequence of the STAU2 antigen polypeptide is TIARELLMNGTSSTAEAIGL; The sequence of the TP53 antigen peptide is PLSQETFSDLWKLLPENNVL.
5. The application as described in claim 4, characterized in that, The biomarker is one or a combination of two or more of the following: autoantibodies against DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide.
6. The application as described in claim 5, characterized in that, The biomarkers are a combination of autoantibodies against DNMT1 antigen peptide, second HNRNPLL antigen peptide, LIN9 antigen peptide, LMO7 antigen peptide, NOB1 antigen peptide, STAU2 antigen peptide, and TP53 antigen peptide.
7. The application as described in claim 4, characterized in that, The product is a protein chip, reagent kit, or formulation.
8. The application as described in claim 7, characterized in that, The kit is an ELISA detection kit, which includes an antigen for detecting biomarkers, the antigen being coated on a solid-phase carrier.
9. The application as described in claim 8, characterized in that, The ELISA test kit further includes any one or more combinations of positive control serum, negative control serum, blocking solution, sample diluent, secondary antibody, secondary antibody diluent, washing solution, coating solution, colorimetric solution, or stop solution; the solid phase carrier is made of polyvinyl chloride, polystyrene, polyacrylamide, or cellulose.
10. The application according to claim 9, characterized in that, The samples are serum, plasma, interstitial fluid, saliva, or urine.