A biomarker for early detection of colorectal cancer and use thereof

Using the mDIP-Seq platform, antibodies against 10 antigenic peptides were screened as biomarkers. A diagnostic model was constructed by combining the random forest algorithm, which solved the problems of high invasiveness, high cost and low sensitivity of existing colorectal cancer detection methods, and achieved efficient and low-cost early diagnosis of colorectal cancer and adenoma.

CN120559235BActive Publication Date: 2026-02-10ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510431527.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-02-10
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Existing colorectal cancer detection methods are highly invasive, costly, and have low sensitivity, making it difficult to diagnose advanced colorectal adenomas and early colorectal cancer. Traditional markers such as CEA show significant individual variability, fecal microbial markers are easily contaminated, and imaging equipment is expensive.

Method used

Plasma antibodies were screened using an immunoprecipitation sequencing platform based on high-resolution mRNA display technology (mDIP-Seq). A multiplex antibody panel was constructed using a random forest algorithm, and antibodies against 10 antigenic peptides were selected as biomarkers for the diagnosis of advanced colorectal adenoma and early colorectal cancer.

Benefits of technology

It significantly improves the diagnostic accuracy of advanced colorectal adenomas and early colorectal cancer, reduces costs, is suitable for large-scale screening, reduces the risk of misdiagnosis and missed diagnosis, and provides support for early detection and intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120559235B_ABST
    Figure CN120559235B_ABST
Patent Text Reader

Abstract

The application provides a biomarker for early colorectal cancer detection and application thereof, uses an immunoprecipitation sequencing platform (mDIP-Seq) based on high-resolution mRNA display technology, screens a plasma antibody diagnostic marker of early colorectal cancer, combines a random forest algorithm to develop a multiple antibody panel suitable for early colorectal cancer diagnosis, constructs a disease prediction model, and improves the accuracy of early colorectal cancer diagnosis. Compared with invasive colonoscopy, the method has the characteristics of low cost, generalization and high throughput, and is suitable for large-scale early colorectal cancer screening in the population.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of parent application 2025101572538 (filed on February 13, 2025). Technical Field

[0002] This invention relates to the field of colorectal cancer diagnosis, and more specifically, to a biomarker for early detection of colorectal cancer and its application. Background Technology

[0003] Colorectal cancer (CRC) has an insidious onset, and early symptoms are often difficult to detect, but it actually has a relatively long window for intervention. Colorectal adenoma (CRA) is a common benign intestinal tumor closely related to the development of CRC and is widely considered a precancerous lesion. Patients diagnosed at the precancerous adenoma stage have a higher 5-year survival rate after radical surgery; therefore, early screening and intervention can significantly improve patient survival rates.

[0004] However, most CRA patients are asymptomatic, with only a few exhibiting clinical manifestations such as abdominal discomfort, changes in bowel habits, or blood in the stool. Furthermore, early symptoms of CRC lack specificity. Therefore, the diagnosis of colorectal tumors relies heavily on invasive colonoscopy biopsy as the most reliable method. However, colonoscopy biopsy is highly invasive, painful, and has low patient compliance, often leading to missed opportunities for optimal treatment. Other existing clinical tests also have many drawbacks in terms of accuracy, sensitivity, and cost. For example, serum carcinoembryonic antigen (CEA) shows significant individual variability and can also yield positive results in other tumors such as gastric, lung, and pancreatic cancer, resulting in low sensitivity and insufficient specificity. Fecal occult blood testing has a high false-positive rate and low sensitivity for early colorectal cancer. Fecal DNA methylation testing is costly. Imaging methods such as MRI and computed tomography (CT) require sophisticated equipment and technology and are expensive, making them difficult to widely use for large-scale screening. Therefore, there is an urgent need to develop a sensitive, reliable, and suitable early detection biomarker for large-scale screening of advanced colorectal adenomas and early colorectal cancer.

[0005] Autoantibodies targeting tumor-associated antigens (TAAs) are increasingly becoming a hot topic in the research of novel biomarkers for tumor plasma diagnosis due to their ability to simultaneously reflect the cancer cell and immune status of the body, their greater stability and ease of detection compared to TAAs, and their advantage of being detectable in the early stages of tumor development and before the appearance of clinical symptoms. For example, anti-p53 antibodies are the most studied autoantibodies related to the diagnosis of colorectal cancer, and can distinguish colorectal cancer from healthy controls or benign diseases. However, due to their low sensitivity, they need to be used in combination with other indicators, and they are difficult to use for early diagnosis of advanced colorectal adenomas.

[0006] Furthermore, numerous studies have demonstrated that the gut microbiome is a crucial pathogenic factor in the occurrence and progression of colorectal cancer, and several fecal microbial biomarkers associated with colorectal tumors have been identified. However, fecal microbial biomarkers suffer from drawbacks such as low biomass and susceptibility to interference from reagents or environmental contaminants. Additionally, the human gut microbiome contains diverse species and expresses a large number of potential antigens; potential transient gut microbiome translocations may trigger durable systemic immune responses. Studies have shown that the temporal stability of antibody epitope libraries is higher than that inferred from microbiome abundance from metagenomic sequencing data of fecal samples; however, no research has yet been conducted on gut microbiome antibodies associated with colorectal tumors.

[0007] Therefore, exploring the plasma antitumor and gut microbial antibody library profiles is of great significance for discovering novel plasma biomarkers that can be used for the early diagnosis of colorectal cancer. Summary of the Invention

[0008] To address the problems existing in the prior art, this invention provides a biomarker for the detection of advanced colorectal adenomas and early colorectal cancer, and its application. Using an immunoprecipitation sequencing platform based on high-resolution mRNA display technology (mDIP-Seq), plasma antibody diagnostic biomarkers for advanced colorectal adenomas and early colorectal cancer were screened. Furthermore, a multiplex antibody panel suitable for the diagnosis of advanced colorectal adenomas and early colorectal cancer was developed using a random forest algorithm, and a disease prediction model was constructed to improve the accuracy of diagnosis. Compared with invasive colonoscopy, this method is low-cost, scalable, and high-throughput, making it suitable for large-scale screening of advanced colorectal adenomas and early colorectal cancer in a population.

[0009] The mRNA display immunoprecipitation sequencing platform (mDIP-Seq) is a novel biomarker screening platform technology constructed for the first time in this invention. By constructing an antigen peptide library, magnetic beads carrying human plasma antibodies are reacted with the antigen peptide library to select antigen peptides that can bind to the magnetic beads. Then, by analyzing the enrichment of antigen peptides in patients with advanced colorectal adenomas, early-stage colorectal cancer, and healthy controls, the differences in antibody profiles between these groups are explored. From this, antigen peptides that can significantly distinguish between patients with advanced colorectal adenomas and healthy individuals, as well as those that can distinguish between early-stage colorectal cancer and healthy individuals, are selected. The antibodies corresponding to these antigen peptides are then biomarkers that can be used to diagnose and predict whether an individual has advanced colorectal adenomas or early-stage colorectal cancer.

[0010] On one hand, the present invention provides the use of a biomarker in the preparation of a reagent for predicting whether an individual has advanced colorectal adenoma. The biomarker includes plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against one or more antigenic peptides as shown in any one or more of the sequence listings Seq ID NO.6 to Seq ID NO.7 and Seq ID NO.9. The anti-gut microbial antibodies include antibodies against one or more antigenic peptides as shown in any one or more of the sequence listings Seq ID NO.1 to Seq ID NO.5, Seq ID NO.8 and Seq ID NO.10.

[0011] Advanced CRA is a precancerous lesion, defined as a villous component of any size ≥1 cm or containing ≥25% villous tissue or exhibiting significant developmental abnormalities. It is highly likely to progress to CRC over time. Identifying tumor biomarkers with early warning potential in the early stages of colorectal cancer development and enabling early diagnosis of CRA is of great significance for improving treatment outcomes and prognosis.

[0012] Based on databases such as UniProt and NCBI, as well as other existing literature, this invention designed and synthesized a human protein oligonucleotide library and an oligonucleotide library covering colorectal tumor-associated gut bacterial antigens in high throughput. Plasma samples were collected from patients with advanced colorectal adenomas and healthy controls. Using a high-resolution immunoprecipitation sequencing platform (mDIP-Seq) based on mRNA display technology, plasma anti-tumor and gut microbial IgG antibody fingerprints were generated during the progression of advanced colorectal adenomas. Finally, antibodies against 10 antigenic peptides that were significantly associated with advanced colorectal adenomas were screened, including 3 plasma autoantibodies and 7 anti-gut microbial antibodies.

[0013] Furthermore, the biomarker includes antibodies against any one or more of the antigenic peptides represented by Seq ID NO.1 to Seq ID NO.8.

[0014] To improve the diagnostic efficacy for advanced colorectal adenomas, three plasma autoantibodies and seven anti-intestinal microbial antibodies were screened together, ultimately yielding antibody biomarkers of eight to ten antigenic peptides. A multiple antibody panel suitable for the diagnosis of advanced colorectal adenomas was developed using a random forest algorithm, and a prediction model for advanced colorectal adenomas was constructed, demonstrating good risk prediction capabilities.

[0015] In some embodiments, the reagent for predicting whether a patient has advanced colorectal adenoma is a detection reagent prepared with the biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies, or other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into standardized reagents or kits suitable for the detection of the biomarker by liquid chromatography-ultraviolet (LC-UV) or liquid chromatography-mass spectrometry (LC-MS).

[0016] Therefore, this invention screened antibodies against 10 antigenic peptides from a large number of antigenic peptides as biomarkers in the test samples, and further screened antibodies against the first 8 antigenic peptides as preferred biomarkers, namely antibodies against antigenic peptides with sequences as described in Seq ID NO.1 to Seq ID NO.8. Because using 8 biomarkers results in fewer biomarkers, the effect can still achieve the same effect as constructing a model with 10 antibody biomarkers. Therefore, it is preferred to use antibody biomarkers against 8 antigenic peptides to construct a diagnostic model, which has good clinical diagnostic value for advanced colorectal adenomas, can significantly improve the diagnostic differentiation and diagnostic efficacy of advanced colorectal adenomas, and realize the risk prediction of advanced colorectal adenomas.

[0017] Furthermore, the reagent is used to detect the content of biomarkers in body fluid samples; the body fluid samples include any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.

[0018] Furthermore, the reagent is used to detect the presence or absence of biomarkers in body fluid samples.

[0019] This invention identifies biomarkers for advanced colorectal adenomas through blood screening. These biomarkers show significant differences in the blood of patients with advanced colorectal adenomas and healthy individuals. By collecting blood samples, the positivity rate of these biomarkers in an individual's blood can be detected to predict or assist in the diagnosis of the individual's likelihood of having advanced colorectal adenomas.

[0020] This invention determines whether an individual has advanced colorectal adenoma by judging the presence of the antibody in a sample, i.e., by assessing its positivity rate. This invention also employs machine learning methods to construct a random forest model for determining whether an individual has advanced colorectal adenoma.

[0021] Furthermore, the detection methods include radiometric methods, immunological methods, fluorescence methods, flow cytometry, latex turbidimetry, biochemical methods, enzymatic methods, hybridization methods, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, chromatography, chemiluminescence methods, magnetoelectric methods, or photoelectric conversion methods.

[0022] The presence or level of biomarkers here is a relative concept. For example, when comparing a diseased group with a non-diseased group, the levels of these specific biomarkers are compared relative to either the diseased or non-diseased group as a baseline. It's possible that the levels of certain biomarkers are higher in the diseased group than in the non-diseased group, and this increase is statistically significant, such as a significant or highly significant increase. Therefore, when judging these biomarkers, if it's a single biomarker, if the probability of a certain risk increases, the biomarker's level will change. This change could be a relative increase or a relative decrease, and this difference could be significant, or even highly significant. Therefore, regardless of the testing method, a pre-defined value (cut-off value) can be used as a standard. A level higher than this value is considered to have changed, and such results can be used for prediction or diagnosis.

[0023] Therefore, in some aspects, the biomarkers described in this invention can be obtained by detecting the content of biomarkers in a sample using any existing known method, such as liquid chromatography, gas chromatography, mass spectrometry, LC-MS, gas chromatography-mass spectrometry (GC-MS), chromatography-mass spectrometry (CC-MS), liquid chromatography-tandem mass spectrometry (LC-MS-MS), nuclear magnetic resonance spectroscopy (NMR), immunochromatographic test strips, immunoreaction chips, capillary electrophoresis, infrared spectroscopy, etc. Any method that can detect the content of protein biomarkers in a sample can be used for the diagnosis of advanced colorectal adenomas. As long as the content of protein biomarkers in a sample can be detected, it can be used to predict or diagnose the probability of a certain disease occurring. It can be understood that the detection here involves testing an individual's sample and comparing it with a pre-set standard. The comparison result is used to judge or predict the disease's state, for example, it can be used to predict the probability of developing advanced colorectal adenomas. This prediction or diagnosis is based on whether it occurs within a certain time period. Of course, such detection can be continuous, using changes in the content of certain substances to infer the disease's progression.

[0024] On the other hand, the present invention provides a kit for predicting whether an individual has advanced colorectal adenoma, the kit comprising detection reagents for biomarkers used as described above.

[0025] In another aspect, the present invention provides a combination of biomarkers for predicting whether an individual has advanced colorectal adenoma, the combination of biomarkers comprising antibodies against any one or more antigenic peptides as shown in the sequence listing Seq ID NO.1 to Seq ID NO.10.

[0026] In another aspect, the present invention provides a system for predicting whether an individual has advanced colorectal adenoma. The system includes a data analysis module for analyzing the detection values ​​of biomarkers, the biomarkers including plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against any one or more antigenic peptides shown in Seq ID NO.6 to Seq ID NO.7 and Seq ID NO.9 of the sequence listing; the anti-gut microbial antibodies include antibodies against any one or more antigenic peptides shown in Seq ID NO.1 to Seq ID NO.5, Seq ID NO.8 and Seq ID NO.10 of the sequence listing.

[0027] Furthermore, the biomarker includes antibodies against any one or more antigenic peptides shown in the sequence listings Seq ID NO.1 to Seq ID NO.8.

[0028] In some implementations, random forest models were constructed using combinations of antibody biomarkers of 2 to 10 antigenic peptides. It was found that antibody models constructed using the top 8 to 10 antigenic peptides all had good diagnostic efficacy. The preferred diagnostic model was constructed using antibodies of 8 antigenic peptides, such as those in the sequence listing Seq ID NO.1 to Seq ID NO.8. Changes in antibody concentrations in these models could be used to distinguish patients with advanced colorectal adenomas from healthy individuals, demonstrating extremely high diagnostic value.

[0029] In summary, the combined diagnostic model of antibodies containing 8 to 10 antigenic peptides constructed in this embodiment has good clinical diagnostic value for advanced colorectal adenomas and can significantly improve the diagnostic differentiation and diagnostic efficacy for advanced colorectal adenomas.

[0030] Furthermore, the data analysis module uses the detection values ​​of known sample markers as the training set, and divides the patients into an advanced colorectal adenoma group and a healthy group based on whether they have advanced colorectal adenoma. The module analyzes the relationship between the detection values ​​of the advanced colorectal adenoma group and the healthy group, and constructs a model.

[0031] Furthermore, the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output the prediction results.

[0032] Furthermore, the detection values ​​are the presence, relative abundance, or concentration of antibodies against the first eight antigenic peptides listed in Seq ID NO.1 to Seq ID NO.10.

[0033] In another aspect, the present invention provides the use of a biomarker in the preparation of a reagent for predicting whether an individual has early colorectal cancer. The biomarker includes plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against any one or more antigenic peptides shown in Seq ID NO.12, Seq ID NO.14 to Seq ID NO.18, and Seq ID NO.20. The anti-gut microbial antibodies include antibodies against any one or more antigenic peptides shown in Seq ID NO.11, Seq ID NO.13, and Seq ID NO.19.

[0034] Based on databases such as UniProt and NCBI, as well as other existing literature, this invention constructed a human protein oligonucleotide library and an oligonucleotide library covering intestinal bacterial antigens. Plasma samples were also collected from patients with early-stage colorectal cancer and healthy controls. Using a high-resolution immunoprecipitation sequencing platform (mDIP-Seq) based on mRNA display technology, plasma anti-tumor and intestinal microbial IgG antibody fingerprints were mapped during the progression of early-stage colorectal cancer. Ultimately, antibodies against 10 antigenic peptides that showed significant association with early-stage colorectal cancer were screened, including 7 plasma autoantibodies and 3 anti-intestinal microbial antibodies.

[0035] Furthermore, the biomarker includes any one or more polypeptide sequences shown in Seq ID NO.11 to Seq ID NO.19.

[0036] To improve the diagnostic efficacy for early colorectal cancer, seven plasma autoantibodies and three anti-gut microbial antibodies were screened together, ultimately yielding antibody biomarkers of nine to ten antigenic peptides. A multiple antibody panel suitable for the diagnosis of early colorectal cancer was developed using a random forest algorithm, and an early colorectal cancer prediction model was constructed, which has good risk prediction capabilities.

[0037] In some embodiments, the reagent for predicting whether a person has early colorectal cancer is a detection reagent prepared with the biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies, or other biological reagents and kits suitable for the detection of the biomarker; it can also be developed into standardized reagents or kits suitable for the detection of the biomarker by liquid chromatography-ultraviolet (LC-UV) or liquid chromatography-mass spectrometry (LC-MS).

[0038] Therefore, this invention screened antibodies against 10 antigenic peptides from a large number of antigenic peptides as biomarkers in the test samples, and further screened antibodies against the first 9 antigenic peptides as preferred biomarkers, namely antibodies against antigenic peptides with sequences as described in Seq ID NO.11 to Seq ID NO.19. Because using 9 biomarkers, the effect can achieve the same effect as constructing a model with 10 antibody biomarkers, it is preferred to use antibody biomarkers against 9 antigenic peptides to construct a diagnostic model. This model has good clinical diagnostic value for early colorectal cancer, can significantly improve the diagnostic and differential diagnostic ability and diagnostic efficacy for early colorectal cancer, and can achieve risk prediction for early colorectal cancer.

[0039] Furthermore, the reagent is used to detect the content of biomarkers in body fluid samples; the body fluid samples include any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.

[0040] Furthermore, the reagent is used to detect the presence or absence of biomarkers in body fluid samples.

[0041] The probability of developing colorectal cancer can be predicted by determining whether antibodies are present in a sample.

[0042] In another aspect, the present invention provides a kit for predicting whether an individual has early-stage colorectal cancer, the kit comprising detection reagents for biomarkers used as described above.

[0043] In another aspect, the present invention provides a combination of biomarkers for predicting whether an individual has early colorectal cancer, the combination of biomarkers comprising antibodies against any one or more antigenic peptides as shown in Seq ID NO.11 to Seq ID NO.20.

[0044] In another aspect, the present invention provides a system for predicting whether an individual has early-stage colorectal cancer. The system includes a data analysis module for analyzing the detection values ​​of biomarkers, the biomarkers including plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against any one or more antigenic peptides shown in Seq ID NO.12, Seq ID NO.14 to Seq ID NO.18, and Seq ID NO.20; the anti-gut microbial antibodies include antibodies against any one or more antigenic peptides shown in Seq ID NO.11, Seq ID NO.13, and Seq ID NO.19.

[0045] Furthermore, the biomarker includes antibodies against any one or more antigenic peptides shown in Seq ID NO.11 to Seq ID NO.19.

[0046] Furthermore, the data analysis module uses the detection values ​​of known samples' markers as the training set, and divides them into an early colorectal cancer group and a healthy group based on whether they have early colorectal cancer. It then analyzes the relationship between the detection values ​​of the early colorectal cancer group and the healthy group to construct a model.

[0047] Furthermore, the system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output the prediction results.

[0048] Furthermore, the detection values ​​are the presence, relative abundance, or concentration of antibodies against the first nine antigenic peptides listed in Seq ID NO.11 to Seq ID NO.19.

[0049] The beneficial effects of this invention are as follows:

[0050] 1. This invention has screened 10 novel antibody biomarkers that can distinguish between patients with advanced colorectal adenomas and healthy individuals. These biomarkers can effectively assess and diagnose patients with advanced colorectal adenomas and healthy individuals, and can more accurately identify patients with advanced colorectal adenomas. Compared with traditional detection methods, this invention reduces the risk of misdiagnosis and missed diagnosis, and provides strong support for the early detection and intervention of the disease.

[0051] 2. A combined differential diagnostic model for advanced colorectal adenomas containing 2 to 10 antibody biomarkers was constructed. This model is convenient, rapid, and the detection results are highly consistent with the clinical gold standard detection results. At the same time, it significantly reduces the cost of diagnosing advanced colorectal adenomas and has good application prospects.

[0052] 3. Ten novel antibody biomarkers were screened that can distinguish between early-stage colorectal cancer patients and healthy individuals, enabling effective assessment and diagnosis of both groups. This allows for more accurate identification of early-stage colorectal cancer patients and reduces the risk of misdiagnosis and missed diagnosis compared to traditional detection methods, providing strong support for early detection and intervention of the disease.

[0053] 4. A combined differential diagnostic model for early colorectal cancer containing 2 to 10 antibody biomarkers was constructed. This model is convenient, rapid, and the detection results are highly consistent with the clinical gold standard detection results. At the same time, it significantly reduces the cost of diagnosing early colorectal cancer and has good application prospects.

[0054] Detailed description

[0055] (1) Diagnosis or testing

[0056] Diagnosis or testing here refers to the detection or analysis of biomarkers in a sample, or the determination of the content of a target biomarker, such as its absolute or relative content. The presence or quantity of the target biomarker then indicates whether the individual providing the sample may have or suffer from a certain disease, or the likelihood of having a certain disease. The meanings of diagnosis and testing are interchangeable here. The results of such testing or diagnosis cannot be directly considered as a direct result of disease; rather, they are intermediate results. If a direct result is obtained, further auxiliary methods such as pathology or anatomy are needed to confirm the presence of a certain disease. For example, this invention provides several novel biomarkers associated with advanced colorectal adenomas and early colorectal cancer. Changes in the content of these biomarkers are directly related to the presence of early colorectal cancer or advanced colorectal adenomas.

[0057] (2) Correlation between markers or biomarkers in patients with early colorectal cancer and advanced colorectal adenoma

[0058] In this invention, "marker" and "biomarker" have the same meaning. Here, "linkage" refers to a direct correlation between the presence or change in the level of a certain biomarker in a sample and a specific disease. For example, a relative increase or decrease in the level indicates a higher likelihood of having the disease compared to healthy individuals.

[0059] If multiple different biomarkers are present simultaneously in a sample, or if their relative levels change, it indicates a higher likelihood of having the disease compared to healthy individuals. In other words, among biomarkers, some have a strong association with the disease, some have a weak association, or some may not be associated with any specific disease. One or more biomarkers with strong associations can be used as diagnostic markers, while biomarkers with weak associations can be combined with strong biomarkers to diagnose a disease, increasing the accuracy of test results.

[0060] The numerous biomarkers in plasma discovered in this invention can be used to differentiate early-stage colorectal cancer, advanced colorectal adenoma, from healthy individuals or those with benign diseases. These biomarkers can be used individually for direct detection or diagnosis, indicating a strong correlation between the relative change in their levels and early-stage or advanced-stage colorectal adenoma. It is understood that one or more biomarkers strongly associated with early-stage or advanced-stage colorectal adenoma can be detected simultaneously. It is generally understood that in some methods, selecting highly correlated biomarkers for detection or diagnosis can achieve a certain level of accuracy, such as 60%, 65%, 70%, 80%, 85%, 90%, or 95%. This indicates that these biomarkers can provide intermediate values ​​for diagnosing a certain disease, but does not necessarily mean a direct diagnosis of that disease.

[0061] (3) Definition of disease terminology

[0062] Colorectal cancer, also known as colon cancer, refers to cancer originating from the colonic epithelium, including colon cancer and rectal cancer. The most common pathological type is adenocarcinoma, with squamous cell carcinoma being rare. Treatment of colorectal cancer should be individualized, selecting appropriate treatment methods based on the patient's age, physical condition, tumor pathological type, and extent of invasion (staging). These methods include radical surgery, chemotherapy, targeted therapy, and radiotherapy. The development of colorectal cancer generally involves a process of normal mucosal hyperplasia, advanced adenoma (malignant), and adenocarcinoma (malignant), typically taking 5-10 years. Therefore, early screening, diagnosis, and treatment are the most effective means of reducing colorectal cancer mortality. In particular, intervention at the polyp adenoma (malignant) stage can effectively prevent the occurrence of colorectal cancer. Advanced adenoma (AA) is a precancerous lesion, defined as a lesion ≥1 cm in size or containing ≥25% villous components of any size or exhibiting significant developmental abnormalities. Over time, it easily develops into colorectal cancer. If tumor biomarkers with certain early warning functions can be found in the early stages of colorectal cancer, and colorectal cancer and advanced adenomas (malignant tumors) can be diagnosed, it will be of great significance for improving the treatment effect and prognosis of patients.

[0063] Early colorectal cancer: Early colorectal cancer refers to colorectal epithelial tumors of any size whose infiltration depth is limited to the mucosa and submucosa, regardless of whether there is lymph node metastasis.

[0064] Advanced adenoma: This refers to tubular villous adenoma, villous adenoma, and / or adenoma with high-grade dysplasia with a diameter >1 cm. Because of its close relationship with the development of colorectal cancer, it is considered a precancerous lesion. Advanced adenomas should be removed promptly to prevent them from developing into colorectal cancer.

[0065] Therefore, the prediction model provided by this invention can quickly distinguish between patients with advanced colorectal adenomas and healthy individuals based on changes in the concentration of biomarkers in body fluid samples.

[0066] (4) The gold standard for the diagnosis of early colorectal cancer: Early colorectal cancer patients often have no symptoms or signs. Diagnosis relies on standardized colonoscopy by qualified physicians and biopsy histopathology as the basis for diagnosis. Attached Figure Description

[0067] Figure 1 This is a schematic diagram of the mDIP-Seq screening process in Example 1;

[0068] Figure 2 This is a schematic diagram of the composition of the plasma antigen peptide and intestinal bacterial antigen peptide libraries constructed in Example 1;

[0069] Figure 3 This is a schematic diagram illustrating the results of verifying the reliability and repeatability of the mDIP-Seq platform in Example 1; where... Figure 3 In this context, A represents the reliability verification result of the mDIP-Seq platform. Figure 3 In this context, B represents the technical repeatability verification result;

[0070] Figure 4 In Example 1, the reported autoantibody markers were detected in the plasma of colorectal cancer patients using the mDIP-Seq platform.

[0071] Figure 5 This refers to the number of antigenic peptides related to HC, CRA, and CRC screened from the human protein antigen library using the mDIP-Seq platform in Example 1.

[0072] Figure 6 This illustrates the relationship between plasma antibody-bound antigen peptides from the HC, CRA, and CRC groups screened from the human protein antigen library in Example 1.

[0073] Figure 7 This refers to the number of antigenic peptides related to HC, CRA, and CRC screened from the intestinal bacterial antibody epitope library using the mDIP-Seq platform in Example 1.

[0074] Figure 8 This illustrates the relationship between plasma antibody-bound antigen peptides from the HC, CRA, and CRC groups screened from the intestinal bacterial antibody epitope library in Example 1.

[0075] Figure 9 This is a graph showing the results of orthogonal partial least squares discriminant analysis of the differences in plasma antibody binding antigen peptides among the HC, CRA, and CRC groups in Example 1;

[0076] Figure 10 This refers to the diagnostic model for advanced colorectal adenoma constructed in Example 2;

[0077] Figure 11 This is the early colorectal cancer diagnostic model constructed in Example 2. Detailed Implementation

[0078] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate understanding of the present invention and are not intended to limit it in any way. The reagents used in this embodiment are all known products and were obtained by purchasing commercially available products.

[0079] Example 1: Screening of antigenic peptide antibody biomarkers for CRA and CRC detection

[0080] Based on databases such as UniProt and NCBI, as well as other existing literature, this invention constructed a human protein oligonucleotide library and an oligonucleotide library covering intestinal bacterial antigens. Plasma samples were collected from patients with early-stage colorectal cancer, patients with advanced-stage colorectal adenomas, and healthy controls. Using a high-resolution immunoprecipitation sequencing platform (mDIP-Seq) based on mRNA display technology, antigenic peptides capable of binding to antibodies in the plasma of patients with colorectal cancer and advanced-stage colorectal adenomas were screened. The flowchart of the screening process using the mDIP-Seq platform (an immunoprecipitation sequencing platform based on high-resolution mRNA display technology) is shown below. Figure 1 Ten antigenic peptides capable of efficiently detecting advanced colorectal adenomas were screened, including three antigenic peptides for detecting plasma autoantibodies and seven antigenic peptides for detecting anti-gut microbial antibodies. Simultaneously, ten antigenic peptides capable of efficiently detecting colorectal cancer were also screened, including seven antigenic peptides for detecting plasma autoantibodies and three antigenic peptides for detecting anti-gut microbial antibodies.

[0081] 1. Research subjects and samples

[0082] This study collected plasma samples from patients in the Department of Colorectal Surgery at the Second Affiliated Hospital of Zhejiang University School of Medicine. The samples included 34 patients with advanced adenomas (CRA), 47 patients with early-stage colorectal cancer (CRC), and 42 healthy controls (HC) (excluding those with intestinal diseases via colonoscopy). Inclusion criteria were: 1) no use of antibiotics, probiotics, traditional Chinese medicine, proton pump inhibitors, or hormones within three months prior to sampling; 2) sampling for adenomas and colorectal cancer patients prior to interventional treatment (including medication and surgery). Exclusion criteria were: vegetarian; patients with diabetes; history of autoimmune diseases or ongoing immunomodulatory therapy; history of malignant diseases or ongoing radiotherapy / chemotherapy for tumors; history of other types of gastrointestinal cancer or colitis-related diseases. Blood samples were collected using EDTK-2K anticoagulant tubes, stored at 4°C, centrifuged at 2,500 rpm for 10 minutes at room temperature, and the supernatant (plasma) was collected, aliquoted into cryovials, and stored at -80°C. All participants voluntarily participated in the study and signed informed consent forms.

[0083] Table 1. Basic Information of Research Subjects

[0084]

[0085] 2. Composition and design of self- and bacterial antigen libraries

[0086] In this embodiment, two antigen libraries were designed and synthesized in high throughput: a human protein antigen library (covering 211,940 antigenic peptides) and an intestinal bacterial antigen library (covering 277,321 antigenic peptides).

[0087] 1) Composition and design of human antigen libraries (Agilent SurePrint platform)

[0088] The entire human proteome was downloaded from the UniProt database and folded at 100% sequence identity (downloaded on December 18, 2018). A peptide library of 66 amino acids with 12 overlapping amino acid sequences was then created, and these peptide sequences were reverse-translated into DNA codons based on human codons. The adapter sequence “GGCTTCCGGAACCTCA” was added to the 5' end of the library, and the sequence “TCTGCCGGTGTAGGCA” was added to the 3' end to produce a set of oligonucleotide sequences of 230 nt in length. Finally, an oligonucleotide library containing 211,940 different sequences covering nearly 20,600 human proteins was synthesized using the Agilent SurePrint platform. The composition of the human antigen library is shown below. Figure 2 The A in the list contains 211,940 antigens that cover human proteins.

[0089] 2) Composition and design of human gut bacterial antigen libraries (Genewiz platform)

[0090] This example incorporates fecal bacterial markers associated with colorectal adenomas and colorectal cancer, based on the work of Wu Y. et al., Thomas AM et al., and Wirbel J et al. Eight bacteria differed in feces between healthy individuals and patients with colorectal adenomas: *Aminipila butyrica*, *Christensenellaceae* R-7 group sp., *Eubacterium coprostanoligenes*, *Eubacterium ruminantium*, *Rothia dentocariosa*, *Ruminiclostridium* 9 sp., *Ruminococcaceae* UCG-005 sp., and *Eillonella parvula*. 24 bacteria that differ between patients with colorectal adenomas and colorectal cancer include: *Clostridium* scindens, *Eubacterium* coprostanoligenes group sp., *Eubacterium* ventriosum group sp., *Ruminococcus* gnavus group sp., *Bacteroides dorei*, *Bacteroides nordii*, *Blautia faecis*, *Blautiasp.*, *Erysipelatoclostridium ramosum*, *Eubacterium ruminantium*, *Hungatella hathewayi* WAL-18680, *Lachnospira pectinoschiza*, *Lachnospiraceae* UCG-010 sp., *Merdibacter massiliensis*, *Parvimonas micra*, *Porphyromonas* sp. 2007b, *Porphyromonas* sp. HMSC077F02, *Roseburia hominis* A2-183, and *Roseburia hominis*. intestinalis, Ruminococcaceae UCG-002sp., Ruminococcus bromii, Streptococcus infantarius, Streptococcus thermophilus TH1435 and Tyzzerella 3sp.Thirteen bacteria that showed differences in feces between healthy individuals and colorectal cancer patients included: Clostridium symbiosum, Faecalibacterium prausnitzii (domaint), Fusobacterium nucleatum, Gemella morbillorum, Hungatella hathewayi, Parvimonas micra, Peptostreptococcus stomatis, Porphyromonas asaccharolytica, Porphyromonas somerae, Porphyromonas uenonis, Prevotella intermedia, Ruminococcus torques, and Solobacterium moorei.

[0091] In addition, this embodiment also includes Prevotella copri, which can elicit an antibody response in more than 95% of individuals, as a positive control; and according to the 2009-2018 Epidemiological Characteristics Report of Acute Diarrhea in China published by the Chinese Center for Disease Control and Prevention, it includes three of the most common intestinal bacterial pathogens in my country, including diarrheagenicescherichia coli, nontyphoidal Salmonella (S. enterica subsp. enterica), and Vibrioparahaemolyticus. Furthermore, this embodiment also includes Enterococcus faecalis V583, Escherichia coli (Strain:NF73-1), and Klebsiella pneumoniae 30660 / NJST258_1, which have been reported to metastasize to extraintestinal organs and tissues. In addition, nine IgA antibody-coated bacteria were added, including Akkermansiamuciniphila ATCC BAA835, Ruminococcus torques ATCC 27756, Dorea formicigenerans ATCC 27755, Ruminococcus bromii, Ruminococcus gnavus AGR2154, Blautia producta ATCC 27340=DSM 2950, ​​Streptococcus lutetiensis O33, Bacteroides fragilis YCH46, and Haemophilus parainfluenzae T3T1.

[0092] Subsequently, the proteome of the aforementioned bacteria was downloaded from Uniprot and NCBI (downloaded on November 19, 2021). GO annotation terms for these proteins were obtained from Uniprot and Interproscan. Proteins with specified membrane localization or flagellar GO terms were selected for further processing. For membrane proteins, to extract extracellular domains and avoid hydrophobic transmembrane domains, the TMHMM (Transmembrane Prediction Using Hidden Markov Models) program was applied to predict transmembrane helices based on Hidden Markov Models. For the selection of secretory proteins, the SignalP 4.0 program was used to predict and remove secretory peptide sequences to obtain mature sequences. Furthermore, secretory proteins belonging to the type I, III, IV, and VI secretory systems of Gram-negative bacteria lacking signal peptides were included based on GO annotations. A polypeptide library of 56 amino acids with 12 overlapping amino acid sequences was then created. The peptide sequences were clustered using the CDHIT program with 75% sequence identity, and cluster representatives were selected for the next step. These peptide sequences were then reverse-translated into DNA codons according to the *E. coli* codons. The adapter sequence “GGGAAACAGCAACGGA” was added to the 5' end of the library, and the sequence “ACATCGTCAGCGTCGA” was added to the 3' end to generate a set of oligonucleotide sequences of 200 nt in length. Finally, an oligonucleotide library containing 277,321 different sequences covering the intestinal bacterial antigens was synthesized using the Genewiz platform. The composition of the human intestinal bacterial antigen library is shown below. Figure 2 The B group contains 277,321 antigens covering the gut microbiota associated with colorectal tumors.

[0093] After obtaining two DNA libraries, PCR amplification of the libraries was performed using the following two rounds of amplification primers. The primers were designed to include sequences encoding the T7 promoter and Kozak at the 5' end, as well as a Flag tag at the 3' end, for in vitro transcription, in vitro translation, and affinity purification of the peptides.

[0094] The first round of PCR primers included the upstream primer (Seq ID NO.41) and downstream primer (Seq ID NO.42) for the human proteome library, and the upstream primer (Seq ID NO.43) and downstream primer (Seq ID NO.44) for the intestinal bacterial antigen library.

[0095] Second-round PCR primers: including upstream primer (Recovery-F, Seq ID NO.45) and downstream primer (Recovery-R, Seq ID NO.46).

[0096] 3. Screening of disease-related antigenic peptides based on the mDIP-Seq platform

[0097] This embodiment is based on the mDIP-Seq platform technology to screen antigenic peptides associated with advanced colorectal adenomas and early colorectal cancer. The specific process is as follows:

[0098] (1) Preparation of diluted input library

[0099] Exon libraries were transcribed using the T7 in vitro transcription kit, and pF30P linkers containing puromycin were ligated to RNA using T4 DNA ligase and Split20 clip DNA fragments. The clip DNA on the RNA was then removed using Lambda exonuclease. The linked RNA was successfully purified using an mRNA purification kit. 0.4 μM of purified RNA was used to construct a 25 μL in vitro translation system using the NEB PURExpress in vitro translation kit. Each 25 μL of translation product was used for screening plasma antibodies from three patients. Subsequently, 500 mM KCl and 60 mM MgCl2 were added and incubated at room temperature for 30 min to promote the formation of the mRNA-peptide fusion product. Then, each 25 μL translation system was diluted with 200 μL of TBK buffer (25 mM Tris, 150 mM KCl, 0.02% Tween-20), and 10 μL of M2 Anti-Flag magnetic beads were added. The mixture was incubated at 4 °C for 2 h, and then washed five times with 1 mL of pre-chilled TBK buffer (25 mM Tris, 150 mM KCl, 0.02% Tween-20) for affinity purification to remove unfused template RNA and sequences containing nonsense mutations. The mRNA-peptide complex was eluted by incubation at room temperature for half an hour with 50 μL of 1X RT reaction solution containing 200 μg / mL Flag peptide. Then, HiscriptIII reverse transcriptase (Vazyme) and RT primers (Lib-rev, Seq ID NO. 47) were added for reverse transcription to generate the mRNA / cDNA-peptide complex (input library). After reverse transcription, a portion of the purified sample was reserved for sequencing to determine the peptide distribution frequency in the input library.

[0100] (2) Enrichment of human plasma IgG

[0101] The procedure was performed in a 96-well plate (Biorad, HSP9601), using 10 μL of Dynabeads containing protein-coupled G. TMProtein G / sample was processed using a 96-well plate magnetic rack. The magnetic beads were washed three times with TBK buffer, then 1 mL of PrG blocking buffer (50% Ultra Block buffer, 20% BSA, 10 μL / mL yeast tRNA, 1 mg / mL sperm DNA) was added. The plate was rotated and blocked at room temperature for 1-2 hours, followed by three washes with TBK. Plasma samples stored at -80℃ were removed and incubated at 4℃ until completely thawed. Patient plasma was diluted 1:50, and 50 μL of diluted serum was added to mix with the PrG-coated magnetic beads. The mixture was incubated overnight at 4℃ to enrich patient plasma IgG.

[0102] (3) Screening of disease-related antigen peptides

[0103] Finally, the magnetic beads were washed three times with approximately 200 μL of TBK buffer, and the input library diluted with Buffer X (20% BSA, 10 μL / mL yeast tRNA, 1 mg / mL sperm DNA, 25 mM Tris-HCl, 150 mM KCl, 0.02% Tween-20) was added and incubated at room temperature for 3 h. After incubation, the samples were washed three times thoroughly with 200 μL of TBK buffer, and then nuclease-free water was added. The mixture was heated at 95 °C for 4 min, and the eluent was collected as the output library. PCR amplification was performed on both the input and output libraries. Sequencing libraries were then constructed from the input and output libraries, with a unique dual index added to each sample for differentiation. Finally, high-throughput sequencing was performed on an Illumina Hiseq PE150 platform.

[0104] PCR amplification primers: including upstream primer (Human-Seq-F, Seq ID NO.48) and downstream primer (Human-Seq-R, Seq ID NO.49) for human proteome libraries; and upstream primer (Bac-Seq-F, Seq ID NO.50) and downstream primer (Bac-Seq-R, Seq ID NO.51) for intestinal bacterial antigen libraries.

[0105] (4) Validation of the sequence screening results

[0106] The screened peptides were further validated using the Split Luciferase Binding Assay (SLBA), which confirmed the presence of antibodies corresponding to the antigenic peptides in plasma.

[0107] The above process is the workflow of the mDIP-Seq platform developed in this embodiment. It combines an in vitro synthesized linear epitope library covering human proteome and colorectal tumor-related gut microbiota antigens. PrG magnetic beads enriched with human plasma IgG antibodies are reacted with antigenic peptides in the library to select antigenic peptides that can bind to antibodies in the plasma of healthy individuals, patients with advanced colorectal adenomas, and patients with colorectal cancer, respectively. The relevant antigenic peptides can be used as biomarkers to predict whether an individual has colorectal cancer or advanced colorectal adenoma.

[0108] In short, the antigen library is amplified and recovered by PCR, and then undergoes a series of steps including in vitro transcription, in vitro translation, flag-tag purification, and reverse transcription to generate an mDIP-Seq input library (peptide-mRNA complex). The input library is then incubated overnight at low temperature with PrG magnetic beads containing human plasma IgG antibodies and blank magnetic beads. After thorough washing, the DNA is eluted by heating at 95°C, and the elution buffer is recovered, which is the output library.

[0109] The abundance distribution of peptides in each sample was fitted using the gamma-poisson model, and the significance p-value of enrichment of each antigenic peptide was calculated to analyze the enriched antigenic peptides in each sample.

[0110] Quality control and validation of the mDIP-Seq platform: The mDIP-Seq platform allows for the full parallel operation of 96 samples, which can relatively reduce any batch or plate-to-plate variability. Figure 3 The composition, feasibility, and reproducibility validation results of two antigen libraries are shown, including membrane proteins, secretory proteins, bacterial flagella, and virulence proteins. Prior to formal screening, this example validated the feasibility of the mDIP-Seq platform using a human antigen library and commercially available anti-TLR4 antibodies containing known antigenic peptides. A single round of immunoprecipitation sequencing revealed that mDIP-Seq successfully enriched the corresponding red-labeled TLR4 antigenic peptides (…). Figure 3 (A) Meanwhile, the peptide enrichment in the output libraries from the three independent experiments showed high reproducibility (Pearson correlation coefficient > 0.95). Figure 3 (B in the text). These results indicate that the mDIP-Seq platform can effectively identify the antigenic peptides corresponding to antibodies, and that the platform can also be applied to antibodies against membrane proteins.

[0111] 4. mDIP-Seq data processing

[0112] Sequencing data was analyzed using Python and bash scripts. Each sample was identified by a 10bp sequence matched on the corresponding Unique Dual Index. To improve data quality, the sequencing company filtered the raw FASTQ files from paired-end sequencing, removing reads with more than 10% N bases or more than 50% low-quality (Q<=5) bases to obtain clean FASTQ files.

[0113] Subsequently, in this embodiment, Cutadapt (version 4.8.0) was used to remove constant regions from each read to generate a new cleanfastq file. Next, the bowtie2-build command was used to construct index files for the human and bacterial reference genome peptide libraries. Bowtie2 (version 2.5.0) was used to align the new cleanfastq file with these index files, using the --very-sensitive parameter to ensure accurate matching of paired-end sequencing data. The alignment results were output in SAM (Sequence Alignment Map) format. Samtools (version 1.16.1) was then used to convert it to BAM (Binary Alignment Map) format, and the Pysam (version 0.22.1) module in Python (version 3.11.5) was used to count the number of peptides in each sample, and the readings were normalized to RPM (read counts per million) files.

[0114] Finally, this embodiment uses the Bash command-line tool and the Phip-stat module (version 0.5.1) in Python (version 3.11.5) to fit the Gamma distribution parameters using the gamma-poisson model and calculate the significance p-value for peptide enrichment in each sample. The code is as follows: "phip gamma-poisson-model -t 99.9 -i RPM_file.tsv -ogamma-poisson-outputfile." The larger the -Log10 p-value of the peptide, the greater the probability of enrichment.

[0115] The study cohort included 123 participants, comprising 34 patients with advanced colorectal cancer (CRA), 47 patients with early-stage colorectal cancer (CRC), and 42 control subjects (HC) who were excluded from intestinal disease via colonoscopy. All patients were diagnosed with the disease and prior to interventional treatment (including medication and surgery) at the time of plasma sample collection. First, this example analyzed the enrichment of colorectal cancer-related autoantibodies reported in the literature. mDIP-Seq detected these autoantibody responses in the CRC population, indirectly validating the feasibility of the mDIP-Seq platform. Figure 4 To reduce the impact of false positive data, a blank PrG control group (Mock IP) without plasma was set up in parallel.

[0116] To reduce the false positive rate for human antigen libraries, this embodiment defines individual-specific enriched self-antigen peptides as peptides that are significantly enriched in patient plasma samples (-Log10 p-value > 2, i.e., p < 0.01) with an enrichment fold of 5 times that of the input library, while not significantly enriched in the blank control group (-Log10 p-value < 1.3, i.e., p > 0.05). The number of specifically enriched peptides in each individual was calculated in each group. Figure 5 Meanwhile, we defined peptides appearing in at least 10% but less than 95% of subjects in the HC, CRA, and CRC cohorts as population-specific enriched antigenic peptides. Therefore, 2,261 autoantibody-binding peptides were selected for case-control analysis, 1,237 for CRA cohort-specific analysis, and 7,296 for early CRC cohort-specific analysis. Figure 6 ).

[0117] Regarding the screening of intestinal bacterial antibody epitope libraries, this embodiment defines peptides that are significantly enriched in plasma samples (-Log10 p-value > 10) with an enrichment fold of 10 times that of the input library, while not significantly enriched in the blank control group (-Log10 p-value < 1.3, i.e., p > 0.05) as individual-specific enriched bacterial antigenic peptides. Figure 7 Similarly, in this embodiment, 2,514 bacterial antibody-binding peptides were selected for case-control analysis, 1,625 bacterial antibody-binding peptides were selected for CRA cohort-specific analysis, and 3,318 bacterial antibody-binding peptides were selected for early CRC cohort-specific analysis. Figure 8 Subsequent orthogonal partial least squares discriminant analysis (OPLS-DA) showed significant differences in plasma IgG antibody enrichment profiles for self- and bacterial antigenic peptides among CRA patients, early CRC patients, and HC patients (p<0.05). Figure 9 ).

[0118] 5. Screening Results

[0119] (1) Biomarkers used to distinguish CRA patients from HC populations

[0120] Finally, through comprehensive screening based on positive rate, importance coefficient, and importance ranking, antibodies against 10 antigenic peptides, as shown in Table 2, were obtained for biomarkers used to distinguish between CRA patients and HC individuals. Detecting the corresponding antibody levels in plasma using these 10 antigenic peptides can efficiently differentiate between patients with advanced colorectal adenomas and healthy individuals. Therefore, the antibodies corresponding to these 10 antigenic peptides are biomarkers that can be used to efficiently predict whether an individual has advanced colorectal adenomas. Among the 10 antigenic peptides, the antibodies against 3 autoantigen peptides are human autoantibodies, and the antibodies against 7 microbial antigen peptides are anti-gut microbial antibodies. The positive rate, importance coefficient, and importance ranking results of the 10 antigenic peptides are shown in Table 3.

[0121] Table 2. Antibodies that can be used to predict antigenic peptides in patients with advanced colorectal adenomas.

[0122]

[0123] Table 3. Positive rates, importance coefficients, and importance rankings of antigenic peptides that can be used to predict patients with advanced colorectal adenomas.

[0124]

[0125]

[0126] The positive rate in Table 3 represents the probability of the antibody for this antigenic peptide appearing positive in the plasma of healthy individuals and patients with advanced colorectal adenomas, respectively. The greater the difference between the two probabilities, the more effective the antibody for this antigenic peptide is in distinguishing between patients with advanced colorectal adenomas and healthy individuals. The importance coefficient refers to the proportion of importance of the antibody for this antigenic peptide in distinguishing between patients with advanced colorectal adenomas and healthy individuals. The larger the importance coefficient, the better the antibody for this antigenic peptide is in distinguishing between patients with advanced colorectal adenomas and healthy individuals, and the greater its importance proportion. The higher the importance ranking, the better the antibody for this antigenic peptide is in distinguishing between patients with advanced colorectal adenomas and healthy individuals. The AUC values, ROC sensitivity, and specificity of the 10 antigenic peptide antibodies obtained through screening for highly effective differentiation between patients with advanced colorectal adenomas and healthy individuals are shown in Table 4.

[0127] Table 4. AUC values, ROC sensitivity, and specificity of 10 antibodies used to distinguish CRA and HC.

[0128] Antibody Antibody antigenic peptide sequence AUC value accuracy Specificity No. 1 Seq ID NO.1 0.7088 0.6812 0.91 No. 2 Seq ID NO.2 0.7681 0.7156 0.70 No. 3 Seq ID NO.3 0.815 0.6402 0.71 No. 4 Seq ID NO.4 0.6531 0.5765 0.89 No. 5 Seq ID NO.5 0.7044 0.6288 0.98 No. 6 Seq ID NO.6 0.7188 0.5309 0.22 No. 7 Seq ID NO.7 0.7144 0.4854 0.21 No. 8 Seq ID NO.8 0.625 0.5271 0.11 No. 9 Seq ID NO.9 0.6875 0.6296 0.41 No. 10 Seq ID NO.10 0.62 0.5198 1.00

[0129] For the antibodies against the 10 selected antigenic peptides, this embodiment also validated their AUC values, ROC sensitivity, and specificity using a validation set. The validation set consisted of patient plasma samples collected from the Department of Colorectal Surgery at the Second Affiliated Hospital of Zhejiang University School of Medicine. The samples included 17 patients with advanced adenomas (CRA) and 21 healthy controls (HC) (intestinal diseases were excluded by colonoscopy). The AUC values, ROC sensitivity, and specificity of the antibodies against these 10 antigenic peptides against the validation set are shown in Table 5.

[0130] Table 5 shows the AUC values, ROC sensitivity, and specificity of 10 antibodies used in the validation set to distinguish between CRA and HC.

[0131]

[0132]

[0133] By detecting the levels of antibodies corresponding to these 10 antigenic peptides in plasma, it is possible to predict whether an individual has advanced colorectal adenoma. The correlation between the antibodies of the 10 antigenic peptides and advanced colorectal adenoma can be distinguished by AUC values, sensitivity, and specificity in Tables 3 and 4, with AUC values ​​being the most intuitive and obvious. The higher the AUC value, the more accurately the biomarker can distinguish between patients with advanced colorectal adenoma and healthy individuals.

[0134] As can be seen from Tables 3 and 4, the levels of antibodies against the 10 antigenic peptides are significantly correlated with whether a patient has advanced colorectal adenoma. Any one of the 10 antigenic peptide antibodies can be used alone to distinguish between patients with advanced colorectal adenoma and healthy individuals.

[0135] (2) Biomarkers for differentiating between early CRC patients and HC populations

[0136] Through comprehensive screening based on positive rate, importance coefficient, and importance ranking, antibodies against 10 antigenic peptides were identified as shown in Table 6 for biomarkers used to differentiate between early-stage CRC patients and healthy individuals. Detecting the corresponding antibody levels in plasma using these 10 antigenic peptides can efficiently distinguish between early-stage colorectal cancer patients and healthy individuals. Therefore, the antibodies corresponding to these 10 antigenic peptides are biomarkers that can be used to efficiently predict whether an individual has early-stage colorectal cancer. Among the 10 antigenic peptides, 7 are human protein antigen autoantibodies, and 3 are anti-gut microbial antibodies. The positive rate, importance coefficient, and importance ranking results for the 10 antigenic peptides are shown in Table 7.

[0137] Table 6. Antibodies that can be used to predict early-stage colorectal cancer patients

[0138]

[0139] Table 7. Positive rates, importance coefficients, and importance rankings of antigenic peptides that can be used to predict early-stage colorectal cancer patients.

[0140]

[0141] The positive rate in Table 7 represents the probability of the antibody for this antigenic peptide appearing positive in the plasma of healthy individuals and patients with early-stage colorectal cancer, respectively. The greater the difference between the two probabilities, the more effective the antibody for this antigenic peptide is in distinguishing between patients with early-stage colorectal cancer and healthy individuals. The importance coefficient refers to the proportion of importance of the antibody for this antigenic peptide in distinguishing between patients with early-stage colorectal cancer and healthy individuals. The larger the importance coefficient, the better the antibody for this antigenic peptide is in distinguishing between patients with early-stage colorectal cancer and healthy individuals, and the greater its proportion of importance. The higher the importance ranking, the better the antibody for this antigenic peptide is in distinguishing between patients with early-stage colorectal cancer and healthy individuals. The detection results of the AUC values, ROC sensitivity, and specificity of the 10 antigenic peptide antibodies obtained from the screening for highly effective differentiation between patients with early-stage colorectal cancer and healthy individuals are shown in Table 8.

[0142] Table 8 shows the AUC values, ROC sensitivity, and specificity of 10 antibodies used to differentiate early-stage CRC and HC.

[0143] Antibody Antibody antigenic peptide sequence AUC value accuracy Specificity No. 11 Seq ID NO.11 0.7101 0.6807 0.82 No. 12 Seq ID NO.12 0.6571 0.6316 0.98 No. 13 Seq ID NO.13 0.6553 0.6203 0.87 No. 14 Seq ID NO.14 0.6857 0.6416 0.99 No. 15 Seq ID NO.15 0.5857 0.6035 1.00 No. 16 Seq ID NO.16 0.6143 0.5827 0.99 No. 17 Seq ID NO.17 0.60 0.5526 0.9528 No. 18 Seq ID NO.18 0.60 0.4959 0.9230 No. 19 Seq ID NO.19 0.5968 0.5837 0.9213 No. 20 Seq ID NO.20 0.6124 0.6124 0.95

[0144] For the antibodies against the 10 selected antigenic peptides, this embodiment also validated their AUC values, ROC sensitivity, and specificity using a validation set. The validation set consisted of patient plasma samples collected from the Department of Colorectal Surgery at the Second Affiliated Hospital of Zhejiang University School of Medicine. The samples included 23 patients with early-stage colorectal cancer (CRC) and 21 healthy controls (HC) (intestinal diseases were excluded by colonoscopy). The AUC values, ROC sensitivity, and specificity of the antibodies against these 10 antigenic peptides against the validation set are shown in Table 9.

[0145] Table 9 shows the AUC values, ROC sensitivity, and specificity of 10 antibodies used in the validation set to distinguish between early CRC and HC.

[0146] Antibody Antibody antigenic peptide sequence AUC value accuracy Specificity No. 11 Seq ID NO.11 0.6591 0.6522 0.8182 No. 12 Seq ID NO.12 0.5417 0.52174 1.00 No. 13 Seq ID NO.13 0.3598 0.3478 0.6364 No. 14 Seq ID NO.14 0.6667 0.6521 1.00 No. 15 Seq ID NO.15 0.5417 0.5217 1.00 No. 16 Seq ID NO.16 0.5416 0.5217 1.00 No. 17 Seq ID NO.17 0.50 0.4783 1.00 No. 18 Seq ID NO.18 0.4545 0.4348 0.9091 No. 19 Seq ID NO.19 0.5455 0.5652 0.9091 No. 20 Seq ID NO.20 0.625 0.6087 1.00

[0147] By detecting the levels of antibodies corresponding to these 10 antigenic peptides in plasma, it is possible to predict whether an individual has early-stage colorectal cancer. The correlation between the antibodies to these 10 antigenic peptides and early-stage colorectal cancer can be distinguished using AUC values, sensitivity, and specificity, as shown in Tables 6 and 7. Among these, the AUC value is the most intuitive and obvious. The higher the AUC value, the more accurately the biomarker can distinguish between patients with early-stage colorectal cancer and healthy individuals.

[0148] As can be seen from Tables 8 and 9, the levels of antibodies against the 10 antigenic peptides are significantly correlated with whether or not a person has early-stage colorectal cancer. Any one of the 10 antigenic peptide antibodies can be used alone to distinguish between patients with early-stage colorectal cancer and healthy individuals.

[0149] Example 2: Construction of the Diagnostic Model

[0150] This embodiment uses 10 antibody biomarkers selected in Example 1 for predicting advanced colorectal adenomas and 10 antibody biomarkers for predicting early colorectal cancer to construct diagnostic models for predicting whether an individual has advanced colorectal adenomas and for predicting whether an individual has early colorectal cancer, respectively. Plasma samples were collected from patients in the Department of Colorectal Surgery at the Second Affiliated Hospital of Zhejiang University School of Medicine. The samples included 56 patients with advanced adenomas (CRA), 48 patients with early colorectal cancer (CRC), and 64 healthy controls (HC) (intestinal diseases were excluded by colonoscopy).

[0151] Next, based on Tables 3 and 7, this embodiment examines the ability of the antibody epitope library to distinguish between classification tasks (CRA vs. HC, early CRC vs. HC) and identifies the antibody with the largest contributing antibody-binding peptide as a diagnostic biomarker. Machine learning analysis is performed using Python code and the scikit-learn library. For the selected antibodies that specifically enrich antigenic peptides, a random forest algorithm is used to evaluate the importance of peptide features and classification performance.

[0152] The data were randomly divided into a training set (including 28 patients with advanced adenomas, 24 patients with early-stage colorectal cancer, and 32 healthy controls) and a validation set (including 28 patients with advanced adenomas, 24 patients with early-stage colorectal cancer, and 32 healthy controls). Antibodies against the top 2, 3, 4, 5, 6, 7, 8, 9, and 10 antigenic peptides and their progressive combinations listed in Tables 2 and 5 were used to detect the presence of antibodies against the antigenic peptides in each sample using the mDIP-Seq platform technology provided in Example 1. A random forest model was constructed using machine learning, and the AUC of the model on the training and validation sets was calculated. The classification accuracy and specificity on the validation set were evaluated by combining the confusion matrix of the prediction results. The parameters of the random forest model are shown in Table 10.

[0153] Table 10. Random Forest Model Parameters

[0154]

[0155] The diagnostic model constructed using the random forest model for predicting patients with advanced colorectal adenomas, and the diagnostic effects of different combinations of antibody numbers are shown in Table 11.

[0156] Table 11. Diagnostic efficacy of random forest models for predicting CRA constructed with different antibody combinations (training set)

[0157] Different combinations AUC accuracy Specificity The first two antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.2) 0.86 0.80 0.70 The first three antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.3) 0.93 0.82 0.80 The first four antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.4) 0.93 0.84 0.80 The first 5 antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.5) 0.95 0.83 0.80 The first 6 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.6) 0.97 0.84 0.90 The first 7 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.7) 0.98 0.84 0.90 The first 8 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.8) 0.99 0.86 0.90 The first 9 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.9) 0.99 0.86 0.90 The first 10 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.10) 0.99 0.86 0.90

[0158] Table 11, compared with Tables 4 and 5, shows that antibodies combining antigenic peptides have a higher ability to distinguish between patients with advanced colorectal adenomas and healthy individuals than antibodies using single antigenic peptides. Random forest models constructed using antibodies containing the first 8, 9, and 10 antigenic peptides all achieved an AUC of 0.99, an accuracy of 86%, and a specificity of 90%, demonstrating excellent discriminatory ability against both patients with advanced colorectal adenomas and healthy individuals. Subsequently, the model performance was evaluated on a validation set, and the evaluation results are shown in Table 12.

[0159] Table 12. Diagnostic efficacy of random forest models for predicting CRA constructed with different antibody combinations (validation set)

[0160] Different combinations AUC accuracy Specificity The first two antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.2) 0.81 0.79 0.70 The first three antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.3) 0.66 0.78 0.80 The first four antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.4) 0.76 0.78 0.80 The first 5 antibody combinations (antibodies against the antigenic peptides of Seq ID NO.1 to Seq ID NO.5) 0.76 0.81 0.80 The first 6 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.6) 0.84 0.78 0.90 The first 7 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.7) 0.80 0.82 0.90 The first 8 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.8) 0.93 0.83 0.90 The first 9 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.9) 0.93 0.82 0.90 The first 10 antibody combinations (antibodies against antigenic peptides from Seq ID NO.1 to Seq ID NO.10) 0.93 0.83 0.90

[0161] The random forest classification model constructed from antibodies against the first eight antigenic peptides is shown below. Figure 10 When the default probability threshold for prediction is 0.5, the AUC value for predicting patients with advanced colorectal adenoma reaches 0.93, with an accuracy of 83% and a specificity of 90%.

[0162] It is evident that random forest models constructed using antibodies with the first 8, 9, or 10 antigenic peptides can all be used to accurately distinguish between patients with advanced colorectal adenomas and healthy individuals. The optimal model is the random forest model constructed using antibodies with the first 8 antigenic peptides, which requires the fewest biomarkers and achieves the best results.

[0163] The diagnostic performance of the random forest model for predicting early colorectal cancer patients using different combinations of antibody numbers is shown in Table 13.

[0164] Table 13. Diagnostic performance of random forest models for predicting early-stage CRC constructed with different antibody combinations (training set)

[0165] Different combinations AUC accuracy Specificity The first two antibody combinations (antibodies against the antigenic peptides of Seq ID NO.11 to Seq ID NO.12) 0.83 0.80 0.70 The first three antibody combinations (antibodies against antigenic peptides of Seq ID NO.11 to Seq ID NO.13) 0.88 0.82 0.80 The first four antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.14) 0.89 0.84 0.80 The first 5 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.15) 0.93 0.83 0.80 The first six antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.16) 0.97 0.84 0.90 The first 7 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.17) 0.97 0.84 0.90 The first 8 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.18) 0.97 0.86 0.90 The first 9 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.19) 0.98 0.86 0.90 The first 10 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.20) 0.98 0.86 0.90

[0166] Table 13, compared with Tables 8 and 9, shows that antibodies combining antigenic peptides have a higher ability to distinguish between early-stage colorectal cancer patients and healthy individuals than antibodies using single antigenic peptides. Random forest models constructed using antibodies containing the first 9 or 10 antigenic peptides achieved AUC values ​​of 0.98, accuracy of 86%, and specificity of 90%, demonstrating excellent discrimination ability against both early-stage colorectal cancer patients and healthy individuals. Subsequently, this embodiment evaluated the model performance on a validation set; the evaluation results are shown in Table 14.

[0167] Table 14. Diagnostic performance of random forest models for predicting early-stage CRC using different antibody combinations (validation set)

[0168] Different combinations AUC accuracy Specificity The first two antibody combinations (antibodies against the antigenic peptides of Seq ID NO.11 to Seq ID NO.12) 0.77 0.75 0.82 The first three antibody combinations (antibodies against antigenic peptides of Seq ID NO.11 to Seq ID NO.13) 0.78 0.75 0.83 The first four antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.14) 0.76 0.76 0.98 The first 5 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.15) 0.77 0.79 0.99 The first six antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.16) 0.78 0.82 0.95 The first 7 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.17) 0.78 0.83 0.91 The first 8 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.18) 0.82 0.84 0.90 The first 9 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.19) 0.88 0.89 1.00 The first 10 antibody combinations (antibodies against antigenic peptides from Seq ID NO.11 to Seq ID NO.20) 0.87 0.88 0.99

[0169] The random forest classification model constructed from antibodies against the first nine antigenic peptides is shown below. Figure 11 When the default probability threshold for prediction is 0.5, the AUC value for predicting patients with early colorectal cancer reaches 0.88, with an accuracy of 89% and a specificity of 100%.

[0170] It is evident that random forest models constructed using antibodies with the first 9 or 10 antigenic peptides can be used to accurately distinguish between early-stage colorectal cancer patients and healthy individuals. The optimal model is the random forest model constructed using antibodies with the first 9 antigenic peptides, which requires the fewest biomarkers and achieves the best results.

[0171] The antibodies containing a total of 20 antigenic peptides provided by this invention have shown good accuracy and specificity in diagnosing CRA patients and early CRC patients, which is of great clinical significance for improving the survival rate of colorectal cancer patients.

[0172] It is understood that the embodiments described in this invention are preferred embodiments and features, and any person skilled in the art can make some modifications and variations based on the spirit of the invention described herein. These modifications and variations are also considered to fall within the scope of this invention and the scope limited by the independent and appended claims.

[0173] sequence list

[0174] Seq ID NO.1

[0175] Antigenic peptide sequence of antibody 1:

[0176] FWADVLGTVAMAADDTNWMYGYDERVGAIELVASIGTIAALINVRTVNRLGMDYSS Seq ID NO.2

[0177] Antigenic peptide sequence of antibody 2:

[0178] PMGEGVQRIAYGDTWSKIRTRDGVEGYVLTTSLSYEMVWQDIDRIVWVDTDSLILR Seq ID NO.3

[0179] Antigenic peptide sequence of antibody 3:

[0180] LTEAVRRDAPYSFAVASDLLLMVQNTTDVHFSSIGILMISAFVEVLHRPGNKLPVQ Seq ID NO.4

[0181] Antigenic peptide sequence of antibody 4:

[0182] DGGDGSDDGSSDNGPSDNGPSDNGSSDNGSSDNGSSGNGSSGNGSSSSVQGSWIKD Seq ID NO.5

[0183] Antigenic peptide sequence of antibody 5:

[0184] PALHKLYGAEQLQQAIANNWHYSVAQWLGVSFSADAGWLSMMTLGSGSGSGSGSGS Seq ID NO.6

[0185] Antigenic peptide sequence of antibody 6:

[0186] TYAGQFVMEGFLRLRWSRFARVLLTRSCAILPTVLVAVFRDLRDLSGLNDLLNVLQSLLLPFAVLPSeq ID NO.7

[0187] Antigenic peptide sequence of antibody 7:

[0188] LLCTARLVGLQLLISCCWAFACHSTESSPDFTLPGDYLLAGLFPLHSGCLQVRHRPEVTLCDRSCSSeq ID NO.8

[0189] Antigenic peptide sequence of antibody 8:

[0190] GLSMLDKYRDAGRLKPYLIFALTDETFSIVCHEEPPKSIPRYWVCCFWISILDQCYG Seq ID NO.9

[0191] Antigenic peptide sequence of antibody 9:

[0192] LLVPALALFLLGYVLSARTWRLLLTGCCSSARASCGSALRGSLVCTQISAAAALAPLTWVAVALLGGSeq ID NO.10

[0193] Antigenic peptide sequence of antibody 10:

[0194] RGNNAYGGALIWSLSKYVPNGCYIQLKRDFTTVNMGNVSASPGMTVINEAFGYGER Seq ID NO.11

[0195] Antigenic peptide sequence of antibody 11:

[0196] GLALPGLGYVELGVLAPLVWEAGLVALAESVRIADGIDGLCCGTAFAAQLGLMLCM Seq ID NO.12

[0197] Antigenic peptide sequence of antibody 12:

[0198] KELERYNKNLEEAKRIGIKKAITANISIGAAFLLIYASYALAFWYGTTLVLSGEYSIGQVLTVFFS

[0199] Seq ID NO.13

[0200] Antigenic peptide sequence of antibody 13:

[0201] NRTLLSLMDAWAGPVVMQLMEAAKPFVRWLTDLCVQLSEVERQIHEIVRAYEWAHH Seq ID NO.14

[0202] Antigenic peptide sequence of antibody 14:

[0203] AEYILSSLISNNGATGTWLYRNESDKVLVQSVCIQIRGQILQKLGMWYEAAELIWASIVGYLALPQSeq ID NO.15

[0204] Antigenic peptide sequence of antibody 15:

[0205] NHNYYLDEFANLLDELLMKINGLSDSLQLPLLEKTSNNTGEARTEESPLVDISSYQAAEPADIKDFSeq ID NO.16

[0206] Antigenic peptide sequence of antibody 16:

[0207] PRLLSKQMAGCLEDCTRQAPESPWEEQLARLLQEAPGKLSLDVEQAPSGQHSQAQLSGQQQRLLAFSeq ID NO.17

[0208] Antigenic peptide sequence of antibody 17:

[0209] GAVQKVSKKLEMHVYSKRLEIMLQDIFGEDCVSVKDDSILSVTVDGKTANLNLETRTVECEEGSEDSeq ID NO.18

[0210] Antigenic peptide sequence of antibody 18:

[0211] IDLRIPVPRVKLIDLEDAVQISGHFHIHHLLGNPEFAAPEVIQGIPVSLGTDIWSIGVLTYVMLSGSeq ID NO.19

[0212] Antigenic peptide sequence of antibody 19:

[0213] IFVIHGGIQSVLDKLYLSYEILPLQADIAFNRALVDFLFSLWDSFIKLMLSFSIPM Seq ID NO.20

[0214] The antigenic peptide sequence of antibody 20:

[0215] RVNLARAVYQDADIYLLDDPLSAVDAEVSRHLFELCICQILHEKITILVTHQLQYLKAASQILILKSeq ID NO.21

[0216] The antigenic peptide encoding nucleotide sequence of antibody 1:

[0217] TTTTGGGCGGATGTGCTGGGCACCGTGGCGATGGCGGCGGATGATAACCAACTGGATGTATGGCTATGATGAACGC

[0218] GTGGGCGCGATTGAACTGGTGGCGAGCATTGGCACCATTGCGGCGCTGATTAACGTGCGCACCGTGAACCGCCTG

[0219] GGCATGGATTATAGCAGCSeq ID NO.22

[0220] The antigenic peptide encoding nucleotide sequence of antibody 2:

[0221] CCGATGGGCGAAGGCGTGCAGCGCATTGCGTATGGCGATACCTGGAGCAAAATTCGCACCCGCGATGGCGTGGAA

[0222] GGCTATGTGCTGACCACCAGCCTGAGCTATGAAATGGTGTGGCAGGATATTGATCGCATTGTGTGGGTGGATAACCG

[0223] ATAGCCTGATTCTGCGCSeq ID NO.23

[0224] The antigenic peptide encoding nucleotide sequence of antibody 3:

[0225] CTGACCGAAGCGGTGCGCCGCGATGCGCCGTATAGCTTTGCGGTGGCGAGCGATCTGCTGCTGATGGTGCAGAAC

[0226] ACCACCGATGTGCATTTTAGCAGCATTGGCATTCTGATGATTAGCGCGTTTGTGGAAGTGCTGCATCGCCCGGGCA

[0227] ACAAACTGCCGGTGCAGSeq ID NO.24

[0228] The antigenic peptide encoding nucleotide sequence of antibody 4:

[0229] GATGGCGGCGATGGCAGCGATGATGGCAGCAGCGATAACGGCCCGAGCGATAACGGCCCGAGCGATAACGGCAG

[0230] CAGCGATAACGGCAGCAGCGATAACGGCAGCAGCGGCAACGGCAGCAGCGGCAACGGCAGCAGCAGCAGCGTGC

[0231] AGGGCAGCTGGATTAAAGATSeq ID NO.25

[0232] The antigenic peptide encoding nucleotide sequence of antibody 5 is as follows:

[0233] CCGGCGCTGCATAAACTGTATGGCGCGGAACAGCTGCAGCAGGCGATTGCGAACAACTGGCATTATAGCGTGGCG

[0234] CAGTGGCTGGGCGTGAGCTTTAGCGCGGATGCGGGCTGGCTGAGCATGATGACCCTGGGCAGCGGCAGCGGCAGC

[0235] GGCAGCGGCAGCGGCAGC

[0236] Seq ID NO.26

[0237] The antigenic peptide encoding nucleotide sequence of antibody 6 is as follows:

[0238] ACCTACGCCGGCCAGTTCGTGATGGAGGGCTTCCTGCGGCTGCGGTGGAGCCGGTTCGCCCGGGTGCTGCTGACCC

[0239] GGAGCTGCGCCATCCTGCCTACCGTGCTGGTGGCCGTGTTCCGGGACCTGCGGGACCTGAGCGGCCTGAACGACCT

[0240] GCTGAACGTGCTGCAGAGCCTGCTGCTGCCTTTCGCCGTGCTGCCTSeq ID NO.27

[0241] The antigenic peptide encoding nucleotide sequence of antibody 7 is as follows:

[0242] CTGCTGTGCACCGCCCGGCTGGTGGGCCTGCAGCTGCTGATCAGCTGCTGCTGGGCCTTCGCCTGCCACAGCACCG

[0243] AGAGCAGCCCTGACTTCACCCTGCCTGGCGACTACCTGCTGGCCGGCCTGTTCCCTCTGCACAGCGGCTGCCTGCA

[0244] GGTGCGGCACCGGCCTGAGGTGACCCTGTGCGACCGGAGCTGCAGCSeq ID NO.28

[0245] The antigenic peptide encoding nucleotide sequence of antibody 8:

[0246] GGCCTGAGCATGCTGGATAAATATCGCGATGCGGGCCGCCTGAAACCGTATCTGATTTTTGCGCTGACCGATGAAA

[0247] CCTTTAGCATTGTGTGCCATGAAGAACCGCCGAAAAGCATTCCGCGCTATTGGGTGTGCTTTTGGATTAGCATTCT

[0248] GGATCAGTGCTATGGCSeq ID NO.29

[0249] The antigenic peptide encoding nucleotide sequence of antibody 9 is as follows:

[0250] CTGCTGGTGCCTGCCCTGGCCCTGTTCCTGCTGGGCTACGTGCTGAGCGCCCGGACCTGGCGGCTGCTGACCGGCT

[0251] GCTGCAGCAGCGCCCGGGCCAGCTGCGGCAGCGCCCTGCGGGGCAGCCTGGTGTGCACCCAGATCAGCGCCGCCG

[0252] CCGCCCTGGCCCCTCTGACCTGGGTGGCCGTGGCCCTGCTGGGCGGCSeq ID NO.30

[0253] The antigenic peptide encoding nucleotide sequence of antibody 10 is as follows:

[0254] CGCGGCAACAACGCGTATGGCCGGCGCCTGATTTGGAGCCTGAGCAAATATGTGCCGAACGGCTGCTATATTCAG

[0255] CTGAAACGCGATTTTACCACCGTGAACATGGGCAACGTGAGCGCGAGCCCGGGCATGACCGTGATTAACGAAGCG

[0256] TTTGGCTATGGCGAACGCSeq ID NO.31

[0257] The nucleotide sequence encoding the antigenic peptide of antibody 11:

[0258] GGCCTGGCGCTGCCGGGCCTGGGCTATGTGGAACTGGGCGTGCTGGCGCCGCTGGTGTGGGAAGCGGGCCTGGTG

[0259] GCGCTGGCGGAAAGCGTGCGCATTGCGGATGGCATTGATGGCCTGTGCTGCGGCACCGCGTTTGCGGCGCAGCTG

[0260] GGCCTGATGCTGTGCATGSeq ID NO.32

[0261] The nucleotide sequence encoding the antigenic peptide of antibody 12:

[0262] AAGGAGCTGGAGCGGTACAACAAGAACCTGGAGGAGGCCAAGCGGATCGGCATCAAGAAGGCCATCACCGCCAA

[0263] CATCAGCATCGGCGCCGCCTTCCTGCTGATCTACGGCCAGCTACGCCCTGGCCTTTGGTACGGCACCACCCTGGTG

[0264] CTGAGCGGCGAGTACAGCATCGGCCAGGTGCTGACCGTGTTCTTCAGCSeq ID NO.33

[0265] The nucleotide sequence encoding the antigenic peptide of antibody 13:

[0266] AACCGCACCCTGCTGAGCCTGATGGATGCGTGGGCGGGCCCGGTGGTGATGCAGCTGATGGAAGCGGCGAAAACCG

[0267] TTTGTGCGCTGGCTGACCGATCTGTGCGTGCAGCTGAGCGAAGTGGAACGCCAGATTCATGAAATTGTGCGCGCGT

[0268] ATGAATGGGCGCATCATSeq ID NO.34

[0269] The nucleotide sequence encoding the antigenic peptide of antibody 14:

[0270] GCCGAGTACATCCTGAGCAGCCTGATCAGCAACAACGGCGCCACCGGCACCTGGCTGTACCGGAACGAGAGCGAC

[0271] AAGGTGCTGTGCAGAGCGTGTGCATCCAGATCCGGGGCCAGATCCTGCAGAAGCTGGGCATGTGGTACGAGGCC

[0272] GCCGAGCTGATCTGGGCCAGCATCGTGGGCTACCTGGCCCTGCCTCAGSeq ID NO.35

[0273] The nucleotide sequence encoding the antigenic peptide of antibody 15:

[0274] AACCACAACTACTACCTGGACGAGTTCGCCAACCTGCTGGACGAGCTGCTGATGAAGATCAACGGCCTGAGCGAC

[0275] AGCCTGCAGCTGCCTCTGCTGGAGAAGACCAGCAACAACACCGGCGAGGCCCGGACCGAGGAGAGCCCTCTGGTG

[0276] GACATCAGCAGCTACCAGGCCGCCGAGCCTGCCGACATCAAGGACTTCSeq ID NO.36

[0277] Coding nucleotide sequence of the antigenic peptide of antibody No. 16:

[0278] CCTCGGCTGCTGAGCAAGCAGATGGCCGGCTGCCTGGAGGACTGCACCCGGCAGGCCCCTGAGAGCCCTTGGGAG

[0279] GAGCAGCTGGCCCGGCTGCTGCAGGAGGCCCCTGGCAAGCTGAGCCTGGACGTGGAGCAGGCCCCTAGCGGCCAG

[0280] CACAGCCAGGCCCAGCTGAGCGGCCAGCAGCAGCGGCTGCTGGCCTTC

[0281] Seq ID NO.37

[0282] Coding nucleotide sequence of the antigenic peptide of antibody No. 17:

[0283] GGCGCCGTGCAGAAGGTGAGCAAGAAGCTGGAGATGCACGTGTACAGCAAGCGGCTGGAGATCATGCTGCAGGA

[0284] CATCTTCGGCGAGGACTGCGTGAGCGTGAAGGACGACAGCATCCTGAGCGTGACCGTGGACGGCAAGACCGCCAA

[0285] CCTGAACCTGGAGACCCGGACCGTGGAGTGCGAGGAGGGCAGCGAGGACSeq ID NO.38

[0286] The nucleotide sequence encoding the antigenic peptide of antibody 18:

[0287] ATCGACCTGCGGATCCCTGTGCCTCGGGTGAAGCTGATCGACCTGGAGGACCGCCGTGCAGATCAGCGGCCACTTCC

[0288] ACATCCACCACCTGCTGGGCAACCCTGAGTTCGCCGCCCCTGAGGTGATCCAGGGCATCCCTGGTGAGCCTGGGCAC

[0289] CGACATCTGGAGCATCGGCGTGCTGACCTACGTGATGCTGAGCGGCSeq ID NO.39

[0290] The nucleotide sequence encoding the antigenic peptide of antibody 19:

[0291] ATTTTTGTGATTCATGGCGGCATTCAGAGCGTGCTGGATAAACTGTATCTGAGCTATGAAATTCTGCCGCTGCAGG

[0292] CGGATATTGCGTTTAACCGCGCGCTGGTGGATTTTCTGTTTAGCCTGTGGGATAGCTTTATTAAACTGATGCTGAGC

[0293] TTTAGCATTCCGATGSeq ID NO.40

[0294] The nucleotide sequence encoding the antigenic peptide of antibody 20:

[0295] CGGGTGAACCTGGCCCGGCCGTGTACCAGGACGCCGACATCTACCTGCTGGACGACCCTCTGAGCGCCGTGGAC

[0296] GCCGAGGTGAGCCGGCACCTGTTCGAGCTGTGCATCTGCCAGATCCTGCACGAGAAGATCACCATCCTGGTGACCC

[0297] ACCAGCTGCAGTACCTGAAGGCCGCCAGCCAGATCCTGATCCTGAAGSeq ID NO.41

[0298] Upstream primer for human proteome library: GGACAATTACAAGGAGGAGCCACCATGGCTTCCGGAACCTCA SeqID NO. 42

[0299] Downstream primers for human proteome library: CTTATCGTCGTCATCCTTGTAATCGCTGCCTACACCGGCAGA SeqID NO. 43

[0300] Upstream primer for intestinal bacterial antigen library: GGACAATTACAAGGAGGAGCCACCATGGGAAACAGCAACGGASeq ID NO.44

[0301] Downstream primers for the intestinal bacterial antigen library: CTTATCGTCGTCATCCTTGTAATCGCTCGACGCTGACGATGTS eq ID NO.45

[0302] Recovery-F: TTCTAATACGACTCACTATAGGGACAATTACAAGGAGGAGCC

[0303] Seq ID NO.46

[0304] Recovery-R:GGAGCCGCTACCCTTATCGTCGTCATCCTTGTA

[0305] Seq ID NO.47

[0306] Lib-rev: GGAGCCGCTACCCTTATCGTCG

[0307] Seq ID NO.48

[0308] Human-Seq-F: ACCATGGCTTCCGGAACCTCA

[0309] Seq ID NO.49

[0310] Human-Seq-R: ATCGCTGCCTACACCGGCAGA

[0311] Seq ID NO.50

[0312] Bac-Seq-F: ACCATGGGAAACAGCAACGGA

[0313] Seq ID NO.51

[0314] Bac-Seq-R:ATCGCTCGACGCTGACGATGT。

Claims

1. The use of a biomarker in the preparation of a reagent for predicting whether an individual has early-stage colorectal cancer, characterized in that, The biomarkers include plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against any one or more antigenic peptides shown in Seq ID NO.12, Seq ID NO.14-Seq ID NO.18, and Seq ID NO.

20. The anti-gut microbial antibodies include antibodies against any one or more antigenic peptides shown in Seq ID NO.11, Seq ID NO.13, and Seq ID NO.

19.

2. The use as described in claim 1, characterized in that, The reagent is used to detect the content of biomarkers in body fluid samples; the body fluid samples include any one or more of saliva, blood, urine, plasma, serum, and cerebrospinal fluid.

3. The use as described in claim 2, characterized in that, The reagent is used to detect the presence or absence of biomarkers in body fluid samples.

4. A kit for predicting whether an individual has early-stage colorectal cancer, characterized in that, The kit includes a detection reagent for biomarkers as described in any one of claims 1 to 3.

5. A system for predicting whether an individual has early-stage colorectal cancer, characterized in that, The system includes a data analysis module for analyzing the detection values ​​of biomarkers, including plasma autoantibodies and anti-gut microbial antibodies. The plasma autoantibodies include autoantibodies against one or more antigenic peptides as shown in Seq ID NO.12, Seq ID NO.14-Seq ID NO.18, and Seq ID NO.

20. The anti-gut microbial antibodies include antibodies against one or more antigenic peptides as shown in Seq ID NO.11, Seq ID NO.13, and Seq ID NO.

19.

6. The system as described in claim 5, characterized in that, The data analysis module uses the detection values ​​of markers from known samples as the training set. Based on whether the sample has early colorectal cancer, the sample is divided into an early colorectal cancer group and a healthy group. The module analyzes the relationship between the detection values ​​of the early colorectal cancer group and the healthy group and constructs a model.

7. The system as described in claim 6, characterized in that, The system also includes a data storage module, a data input interface, and a data output interface; the data storage module is used to store the detection values ​​of biomarkers; the data input interface is used to input the detection values ​​of biomarkers, and the data output interface is used to output the prediction results.

Citation Information

Patent Citations

  • Serological markers for detecting colorectal cancer and their application for inhibiting colorectal cancer cells

    US20130084582A1

  • Mixed protein and autoantibody biomarker panel for diagnosing colorectal cancer

    WO2019048588A1