A marker for distinguishing colorectal cancer from colorectal adenoma and application thereof

By combining exosome surface driving markers and internal core protein markers, and utilizing immunoaffinity magnetic bead method and detection technology, a machine learning model was constructed, which solved the problem of difficulty in distinguishing between colorectal cancer and colorectal adenoma in existing technologies, and achieved efficient early screening and diagnosis.

CN122631891APending Publication Date: 2026-08-25THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611122640.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies are insufficient to efficiently distinguish between colorectal cancer and colorectal adenoma. Traditional screening methods are either highly invasive or have low sensitivity. Current exosome detection technologies cannot resolve heterogeneity, making early diagnosis of lesions difficult.

Method used

We used exosome surface-driven markers consisting of SELL, CD177, ITGAL, OLR1, B2M, and LAG3, and internal core protein markers consisting of TGM3, PSMB4, KRT80, KRT2, DSG1, CPOX, CPN1, SUB1, LSM2, RAB34, PZP, SERPINA5, HYDIN, and PGLYRP1 to separate plasma exosomes using immunoaffinity magnetic beads. We then constructed a machine learning model for risk scoring using PBA and DIA detection technologies.

Benefits of technology

It enables precise differentiation between colorectal cancer and colorectal adenoma, improves the accuracy and sensitivity of early screening, lowers the application threshold of detection, and provides a stable diagnostic target.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122631891A_ABST
    Figure CN122631891A_ABST
Patent Text Reader

Abstract

This invention discloses a biomarker for distinguishing between colorectal cancer and colorectal adenoma and its application, relating to the field of biomedical technology. The biomarker is selected from any one of the following: an exosome surface-driven biomarker composed of SELL, CD177, ITGAL, OLR1, B2M, and LAG3; an internal core protein composed of TGM3, PSMB4, KRT80, KRT2, DSG1, CPOX, CPN1, SUB1, LSM2, RAB34, PZP, SERPINA5, HYDIN, and PGLYRP1; or a single-indicator biomarker, KRT80. This invention enables the detection of colorectal cancer and colorectal adenoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a biomarker for distinguishing between colorectal cancer and colorectal adenoma and its application. Background Technology

[0002] Colorectal cancer (CRC) is one of the most common malignant tumors worldwide. The development of CRC typically involves a gradual progression from normal mucosa to colorectal adenoma (Ad) to early CRC. However, the prognosis for patients in advanced stages is extremely poor. Therefore, early screening is a core means of reducing mortality. Currently used screening methods have significant limitations: colonoscopy, as the gold standard, is difficult to use for large-scale population screening due to its highly invasive nature, cumbersome bowel preparation, and poor patient compliance. Commonly used serum biomarkers such as carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9) have low sensitivity, limited diagnostic value for early lesions, and are prone to non-specific elevations in benign diseases. While SEPTIN9 gene methylation detection provides a new direction for early CRC screening, it still faces practical problems such as limited sensitivity and high testing costs.

[0003] Plasma exosomes (Plasma Exo) are a type of nanoscale membrane vesicle found in plasma. They carry information such as specific proteins and nucleic acids from their source cells, are stable, and easy to obtain, making them ideal liquid biopsy carriers. However, current research on exosomes faces two major bottlenecks: First, traditional batch detection techniques can only reflect the average signal of the exosome population and cannot resolve the high heterogeneity of exosomes, easily masking rare subpopulation characteristics associated with early lesions. Second, single proteomics techniques struggle to balance detection depth and specificity, resulting in poor stability of the screened biomarkers, which cannot effectively distinguish between colorectal adrenal lesions (Ad) and early-stage colorectal cancer (CRC). Ad is the most common precancerous lesion of CRC, with a 25%-40% risk of progression to CRC, making it a key differentiating point for early screening. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a marker for distinguishing colorectal cancer from colorectal adenoma and its application, aiming to solve at least one of the problems in the above-mentioned background art.

[0005] The present invention provides a marker for distinguishing colorectal cancer from colorectal adenoma, selected from any of the following: Exosome surface-driven markers consisting of SELL, CD177, ITGAL, OLR1, B2M, and LAG3; The internal core protein is composed of TGM3, PSMB4, KRT80, KRT2, DSG1, CPOX, CPN1, SUB1, LSM2, RAB34, PZP, SERPINA5, HYDIN, and PGLYRP1. KRT80 single-indicator marker.

[0006] According to one aspect of the above technical solution, the biomarker is derived from plasma exosomes, which are obtained by immunoaffinity magnetic bead separation. The plasma exosomes meet the requirement that the main peak of the particle size is located between 60nm and 120nm, and the positive rates of the three exosome biomarkers CD9, CD63, and CD81 are all ≥90%.

[0007] Another aspect of the present invention is to provide an application of a marker for distinguishing between colorectal cancer and colorectal adenoma, and for preparing a kit for distinguishing between colorectal cancer and colorectal adenoma.

[0008] Furthermore, the kit includes: The marker information acquisition unit is used to acquire the detection information of the marker. The analysis and evaluation unit is used to determine the risk assessment result based on the detection information of the marker.

[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Existing early colorectal cancer screening technologies neglect exosome heterogeneity and struggle to distinguish between stage I / II CRC and precancerous colorectal adrenal gland (AD). This invention addresses these shortcomings by employing a dual-dimensional biomarker approach: exosome surface phenotype and internal core protein. Surface-driven biomarkers can resolve subpopulation heterogeneity at the exosome level, accurately identifying the CRC-specific enriched Cluster 2 and Cluster 8 subpopulations, thus overcoming the technical bottleneck of traditional batch testing masking rare lesion subpopulations. Internal core proteins stably reflect the molecular characteristics of tumor origin. The two types of biomarkers complement each other, offsetting the single-dimensional detection bias, achieving high discriminative power in an internal validation cohort, and providing a novel technical pathway for early risk stratification.

[0010] 2. The biomarkers of this invention underwent rigorous multi-stage screening. Surface-driven biomarkers were validated at the single exosome level, while internal core proteins underwent a four-stage cascade feature screening process: univariate pre-screening, Bootstrap stability selection, elastic network compression, and random forest validation. This eliminated interference from random errors and redundant features. The 14 selected internal core proteins showed highly consistent expression trends in multiple independent cohorts, including blood exosomes, single-cell transcriptomes, and TCGA tumor tissues, demonstrating their broad applicability as diagnostic targets. Furthermore, KRT80 not only serves as a key feature of the model but can also be used as an independent single-indicator detection target, significantly lowering the application threshold. Attached Figure Description

[0011] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a technical roadmap for biomarkers used to distinguish between colorectal cancer and colorectal adenoma in this invention; Figure 2 This is a diagram showing the identification results of exosomes derived from patient plasma in this invention. Figure 2 A in the image: Transmission electron microscopy image (scale bar: 100 nm); Figure 2 B in the figure: Particle size distribution curve of nanoparticle tracking analysis; Figure 2 C: Expression of plasma exosome surface markers CD9, CD63, and CD81 by flow cytometry; Figure 3 This is a data quality control chart for proximity coding detection (PBA) in this invention, wherein, Figure 3 A in the figure: Bar chart of sequencing read counts for each sample, including total reads, tag-matched reads, deduplicated reads, and exosome valid reads; Figure 3 Box plots comparing the Ad group and the CRC group in B: show the number of exosome reads (a), the total number of protein reads (b), and the protein / exosome read ratio (c) between the two groups. Figure 4 This is a graph showing the expression profile and differential analysis of exosomal proteins detected by proximity coding in this invention. Figure 4 A: Cluster heatmap of the top 100 differentially expressed proteins (red: upregulated; blue: downregulated); Figure 4 B in the figure: Principal component analysis plot based on differential protein expression profiles; Figure 5 This is a power validation chart for the exosome biomarker PBA risk score of the present invention (Wilcoxon test p=0.01). Figure 6 This is a graph showing the quality control analysis results at the peptide level in this invention, wherein... Figure 6 A in the diagram represents the peptide mass deviation distribution. Figure 6 B in the figure: Amino acid composition distribution diagram, where red bars (Freq) represent the amino acid frequencies of this type of protein in public databases / previous studies, and blue bars (This study) represent the actual proportion of amino acids in the protein identified in this study; Figure 6 C in the figure represents a violin plot showing the distribution of the coefficient of variation for protein quantification in each group of samples. Figure 7 This is a principal component analysis diagram based on protein expression profiling in this invention; Figure 8 This is a heatmap of differentially expressed proteins clustered in the exosome batch proteome of this invention; Figure 9 This is a graph showing the results of differential protein functional enrichment analysis in this invention; wherein, Figure 9 Figure A: GOBP enrichment analysis of differentially expressed proteins (Gene Ontology Biological Process, top 20); Figure 9 Figure B: Differential protein KEGG pathway enrichment analysis (Kyoto Encyclopedia of Genes and Genomes); Figure 10 This is a diagram showing the identification results of the Weighted Gene Co-expression Network Analysis (WGCNA) module in this invention; Figure 11 This is a heatmap showing the correlation between modules and clinical traits in this invention. Figure 12 This is a bar chart showing the association analysis between the exosome phenotype and the carrier proteome module in this invention. Figure 13 This is a scatter plot of the proximity coding detection phenotypic score (proximity coding detection technology risk score) and the most significantly associated module in this invention; Figure 14 This is a diagram showing the screening and efficacy verification results of 14 internal core proteins in this invention; among them, Figure 14 A in the diagram represents the characteristic frequency map of the stability selection of the bootstrap method. Figure 14 B in the graph represents the importance of stable features in random forests. Figure 14 C: Scoring box plot of 14 internal core protein markers (p < 0.001); Figure 14 D in the figure: Receiver Operating Characteristic (ROC) curve; Figure 15 This is a heatmap for cross-dataset marker validation in this invention; Figure 16 This is a diagram showing the single-cell transcriptome sequencing results and biomarker cell origin analysis of GSE132465 in this invention; wherein, Figure 16 A in the diagram represents the cell origin point of exosome markers. Figure 16 B in the figure: Statistical bar chart of cell origin markers from different sources; Figure 17 This is a dot plot of the single-cell transcriptome of GSE178341 in this invention; Figure 18 This is a bar chart showing the differential expression of marker proteins for the Cancer Genome Atlas-Colon Adenocarcinoma (TCGA-COAD) in this invention. Figure 19 This is a forest plot of the Cox hazard ratio in this invention; Figure 20 This is a bar chart showing the differential expression of marker proteins for colorectal adenocarcinoma (TCGA-READ) in this invention. Figure 21 This is a Kaplan-Meier curve of each marker in the GSE39582 microarray dataset and the total CRC lifetime in this invention; Figure 22 This is a heatmap for cross-queue organization level verification in this invention; Figure 23 This is a box plot of KRT80 expression in TCGA-COAD in this invention; Figure 24 This is a graph showing the gene set enrichment analysis results based on GOBP for the KRT80 high / low expression groups in this invention; Figure 25 This is a graph showing the results of KEGG gene set enrichment analysis of the KRT80 high / low expression groups in this invention; Figure 26 This is a bar chart showing the correlation analysis between KRT80 and neighbor-coding detection genotyping clustering in this invention; In the figure: the significance level of statistical difference is defined as: not significant (ns), p ≥ 0.05; * indicates p < 0.05; ** indicates p < 0.01; *** indicates p < 0.001; **** represents p < 0.0001. Detailed Implementation

[0012] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0013] like Figure 1As shown, the overall technical approach of this invention is as follows: First, blood samples are collected from newly diagnosed hospitalized patients. Adnexal cancer (Ad) and chronic cancer (CRC) patients undergoing colorectal endoscopy are included. After screening according to inclusion and exclusion criteria, exosomes are extracted and separated. The separated exosomes are divided into discovery cohorts, used for data-independent acquisition (DIA) group detection (Ad=25 cases, CRC=25 cases) and proximity barcoding assay (PBA) group analysis of single exosome phenotypes (Ad=5 cases, CRC=5 cases, sample numbers Ad_1 to Ad_5, CRC_1 to CRC_5, respectively). Early cancer biomarker screening analysis is performed based on DIA group data, while single exosome differential subgroup (cluster) and surface biomarker analysis is performed using PBA group data. Further analysis using single-cell external datasets clarifies the cellular origin of biomarker molecules. Based on this, the results of differential analysis, correlation analysis, and regression analysis are integrated to jointly analyze the association between key oncogene biomarkers and key exosome clusters, and a predictive model is constructed using a machine learning model. Finally, multi-dimensional validation was performed using single-cell transcriptome data (GSE132465, GSE178341) and the GSE39582 microarray dataset. In-depth analysis of key differentially expressed proteins was conducted, the value of machine learning models was reviewed, and a systematic screening and evaluation of early diagnostic biomarkers for colorectal cancer was completed.

[0014] Specifically, this involves sample collection and exosome isolation and identification: The samples used in this embodiment were all from subjects admitted to the Second Affiliated Hospital of Nanchang University from 2024 to 2025. The inclusion criteria were: age 18-80 years, pathologically diagnosed with stage I / II CRC or colorectal adrenalectomy, no preoperative anti-tumor treatment, and complete clinical data; the exclusion criteria were: other malignant tumors, severe liver and kidney dysfunction, hemolysis and lipemia, and more than 3 freeze-thaw cycles. Finally, 50 samples were included, including 25 patients with stage I / II CRC and 25 patients with colorectal adrenalectomy. There were no statistically significant differences in baseline characteristics other than age and body mass index between the two groups.

[0015] Samples were collected during the fasting period from 6:00 to 9:00 AM. Peripheral venous blood was collected into EDTA anticoagulant tubes. Plasma was separated within 8 hours by centrifugation at 2000×g for 15 minutes at 4°C. The supernatant was aliquoted into RNase-free / DNase-free cryovials, 1-2 mL per tube, and stored at -80°C. The number of freeze-thaw cycles should not exceed 3.

[0016] Exosome isolation was performed using the immunoaffinity magnetic bead method. The kit used was the Plasma-Serium Exosome Affinity Extraction Kit (Nanjing Yiwei Jianhua Biotechnology Co., Ltd.): 200 μL of plasma was centrifuged at 3000 × g for 10 min at 4°C to remove cell debris. The supernatant was collected and 800 μL of pre-chilled incubation buffer was added. 40 μL of EVLent magnetic beads were added and incubated at room temperature for 1 h. After incubation, the supernatant was discarded after magnetic agitation for 3 min. The cells were then washed once each with 1 mL of incubation buffer II and 1 mL of washing buffer, inverting and mixing 20 times per step. After washing with the washing buffer, the cells were magnetically agitated for another 3 min before discarding the supernatant.

[0017] like Figure 2 As shown, the isolated exosomes must pass the following four quality control standards before they can be used for subsequent testing: (1) Morphological identification: 10 μL of exosome sample was dropped onto a copper grid and allowed to stand at room temperature for 1 min. Excess liquid was then absorbed with filter paper. 10 μL of uranium acetate-hydrogen oxide staining solution was added and stained for 1 min. Excess staining solution was then absorbed with filter paper and allowed to air dry at room temperature. The sample was observed under a transmission electron microscope with an accelerating voltage of 80 kV. Qualified samples showed typical cup-shaped or round vesicle structures with clear boundaries, no impurities, and a diameter of 30 nm-150 nm.

[0018] (2) Particle size and concentration analysis: 10 μL of exosome sample was taken and serially diluted with pre-cooled 1×PBS to the appropriate detection concentration. Nanoparticle tracking analysis (NTA) was used for detection at 25℃, with a single sample detection time of 60 s and three technical replicates for each sample. The acceptable criteria were: the main peak of the particle size was concentrated in the 60 nm-120 nm range, the distribution curve showed a typical single-peak shape, and there were no abnormal peaks; the particle concentration was controlled within (2.35±0.42)×10⁻⁶. 8 Within the particle / mL range, the coefficient of variation among multiple technical replicates is no higher than 15%.

[0019] (3) Biomarker Validation: Take 30 μL of exosome sample, dilute to 90 μL with pre-chilled 1×PBS, add 20 μL of fluorescently labeled CD9, CD63, and CD81 monoclonal antibodies, incubate at 37℃ in the dark for 30 min, then add 1 mL of pre-chilled 1×PBS. Centrifuge at 4℃ and 110000×g for 70 min, discard the supernatant, and repeat centrifugation once; resuspend the sample in 50 μL of pre-chilled 1×PBS. Detect using a nanoflow cytometer, with 3 technical replicates. The pass / fail standard is: the positive rate of all three biomarkers, CD9, CD63, and CD81, is ≥90%.

[0020] (4) Protein integrity test: Take the exosome protein extract, add 5× loading buffer, and boil in a boiling water bath for 5 min to denature the protein; prepare a 12% separating gel and a 5% stacking gel, load 20 μL of sample into each well, and add protein molecular weight standards at the same time; the electrophoresis conditions are 80V for 30 min for the stacking gel stage and 120V for 90 min for the separating gel stage. After electrophoresis, Coomassie brilliant blue staining is used, and the protein bands are observed after destaining. The qualified criteria are: the protein bands are clear and continuous, without obvious tailing, without diffuse bands, and without obvious degradation fragments.

[0021] PBA phenotypic score construction: A subset of exosomes were collected for single-exosome detection using PBA technology: First, exosome capture was mediated by biotinylated cholerae toxin B subunit. The specific binding of the cholerae toxin B subunit to ganglioside GM1 on the exosome surface was utilized to immobilize the exosomes in streptavidin-coated 96-well plates. Then, a detection panel containing antibodies specific to 554 proteins was added, and the mixture was incubated at 37°C for 30 min. After washing, a hybridization template was added, and the mixture was incubated at 42°C for 60 min to complete DNA hybridization. Next, an extension reaction solution containing DNA polymerase was added, and the mixture was incubated at 37°C for 30 min to synthesize complete exosome tag sequences. After washing, end repair, A-tailing, adapter ligation, and polymerase chain reaction amplification were performed to construct a DNA library. After the DNA library passed the Agilent 2100 bioanalyzer test, paired-end 150-base high-throughput sequencing was performed using the DNBSEQ-T7 sequencing platform. At least 90% of the bases had a quality value greater than or equal to 30 to ensure sequencing quality.

[0022] After the sequencing data is standardized, such as Figure 3 As shown in A, the conversion efficiency from original sequencing fragments to tag-matched fragments remained stable for each sample, and the retention rate of reads after deduplication was basically consistent, with no systematic shift observed. Figure 3 As shown in B, the CRC group and the Ad group have a high degree of overlap in the distribution of exosome reads, total protein reads, and protein / exosome read ratio.

[0023] Based on standardized PBA exosome quantitative data and the small sample distribution characteristics, the Mann-Whitney U test was used to compare the differences in protein expression levels between the CRC and Ad groups. The Benjamini-Hochberg false discovery rate (BH) method was used for multiple comparison correction, with corrected p < 0.05 and |log2FC (logarithmic fold change to base 2)| ≥ 0.58 as the screening thresholds for differentially expressed proteins. From 551 stably quantifiable proteins, 18 differentially expressed proteins were screened, of which 8 were upregulated and 10 were downregulated in the CRC group compared to the Ad group. Figure 4As shown, based on the principal component analysis (PCA) results of the above 18 differentially expressed proteins, the CRC group and the Ad group showed a clear separation trend in the principal component space, and the overall expression patterns of exosomal proteins in the two groups were significantly different, suggesting that exosomal proteins have differential enrichment characteristics in different disease stages.

[0024] Using a positive rate >5% and an area under the curve (AUC) ≥0.80 as screening criteria, six plasma exosome surface-driven biomarkers (Exo biomarkers) were identified: SELL, CD177, ITGAL, OLR1, B2M, and LAG3. The counts per million (CPM) for each exosome surface-driven biomarker were calculated, which is calculated as the number of positive exosomes divided by the total number of exosomes multiplied by 10. 6 The receiver operating characteristic (ROC) curves of each exosome surface driving biomarker were weighted and normalized, then summed, and finally standardized using Z-scores to convert the results into a PBA phenotype score in the 0-1 range. Figure 5 As shown, the mean PBA phenotype score in the CRC group was 0.618±0.117, and the mean score in the Ad group was 0.382±0.031. The difference between the two groups was statistically significant (Wilcoxon test, p=0.01), and the intergroup differentiation effect was significant, proving that the PBA phenotype score can effectively reflect the disease-related characteristics of exosomes.

[0025] As an example, not a limitation, a simplified PBA phenotypic score was constructed using LASSO logistic regression (regularization parameter α=1), optimized with leave-one-out cross-validation (LOOCV), and the optimal regularization parameter of 0.2278 was selected using the area under the receiver operating characteristic (ROC) curve as the evaluation metric. The differentiation between colorectal cancer and colorectal adenoma was based on: calculating the linear predictive value for each sample, converting it to a hazard probability using a logistic function (plogis function), and determining the optimal cut-off value (i.e., Youden's index = sensitivity + specificity) using the Youden's index maximization method. 1. The risk probability corresponding to the maximum value is used to make risk judgment and classification, and the risk judgment result is obtained.

[0026] Internal core protein detection and screening: Another portion of exosomes was used for DIA proteomics detection. First, the exosome proteins were removed using a column of the top 14 high-abundance proteins to remove 14 high-abundance plasma proteins such as albumin and immunoglobulin G (IgG), thereby reducing the masking effect of high-abundance proteins on low-abundance exosome proteins. Subsequently, trypsin hydrolysis was performed in parallel using precipitation-assisted enzymatic hydrolysis and filtration-assisted enzymatic hydrolysis. Precipitation-assisted enzymatic hydrolysis involved precipitating proteins with trichloroacetic acid, followed by reconstitution with ethyl ammonium bromide and then enzymatic hydrolysis. Filtration-assisted enzymatic hydrolysis involved two steps of enzymatic hydrolysis after replacing urea with an ultrafiltration tube. The parallel use of both methods ensured an enzymatic hydrolysis efficiency of ≥95%. After desalting and quantification, the enzymatically hydrolyzed peptides were analyzed for DIA using a Vanquish Neo (Thermo Fisher Scientific) nanoliter liquid chromatography-tandem mass spectrometry (LC-MS / MS) system. Chromatographic separation used 0.1% trifluoroacetic acid aqueous solution as mobile phase A and 0.1% trifluoroacetic acid acetonitrile solution as mobile phase B, with nanoliter-level gradient elution. Mass spectrometry was performed in positive ion mode, with a precursor ion scan range of 380 m / z–980 m / z. The primary mass spectrometry resolution was 240,000 at 200 m / z, and the secondary mass spectrometry employed a data-independent acquisition mode with 299 scan windows and a collision energy of 25 eV to ensure accurate qualitative and quantitative analysis of the peptides.

[0027] The raw mass spectrometry data were retrieved from the UniProt Knowledgebase using DIA-NN software. A reverse database was added to control the false discovery rate (FDR) to be less than 1%. Figure 6 As shown in A, the peptide quality error is concentrated around 0 ppm, exhibiting a normal distribution. Figure 6 As shown in B, the peptide length is mainly concentrated in the range of 7 to 16 amino acids, and the proportion of peptides without missing cleavage reaches 83.4%. Figure 6 As shown in C, the protein quantification quality control showed that the normalized protein abundance distribution was concentrated in the Ad group and the CRC group, and the median CV of the quantification within the group was <1.5.

[0028] After sample quality testing and outlier removal, a core cohort of 48 cases was determined (24 cases each in the Ad group and the CRC group). Figure 7 As shown, Principal Component Analysis (PCA) revealed a clear and distinguishable trend in protein expression patterns between the two groups of samples. A total of 4951 proteins were identified, of which 1740 proteins were stably quantified in over 80% of the samples. Technical repeatability showed that the correlation coefficients of protein abundance among the three replicates were all >0.95, indicating good quantitative stability. After quantile standardization of the protein data, the limma algorithm was used to screen for differentially expressed proteins. Using corrected p-values ​​<0.05 and |log2FC (logarithmic fold change to base 2)| ≥0.58 as thresholds, 93 differentially expressed proteins were ultimately selected, as shown below. Figure 8 As shown, cluster analysis reveals a clear distinction between the two groups of expression patterns.

[0029] like Figure 9 As shown in Figure A, the Gene Ontology Biological Process (GOBP) enrichment analysis revealed that differentially expressed proteins were mainly concentrated in biological items such as epidermal development, keratinocyte differentiation, complement activation, and proteolytic regulation. Figure 9 As shown in Figure B, KEGG pathway analysis revealed that differentially expressed proteins showed significant enrichment in pathways such as cell adhesion, proteasome, extracellular matrix (ECM)-receptor interaction, and complement coagulation cascade.

[0030] like Figure 10 As shown, the topological overlap matrix (TOM) distance is calculated based on the protein expression matrix, and hierarchical clustering analysis is performed. The upper part of the figure is a dendrogram of protein hierarchical clustering, and the lower part is a color bar for the modules; different colors represent different co-expression modules. Figure 11 As shown, the module-clinical trait association analysis revealed a significant positive correlation between the MEyellow module (yellow module) and CRC (r=0.31, p). = 0.03).

[0031] PBA phenotypic scores are mapped to proteomics cohorts through cross-platform sample pairing. For example... Figure 12 As shown, five WGCNA modules were identified that were significantly associated with the PBA phenotypic score. Among them, the MEdarkgreen module showed the strongest positive correlation with the PBA phenotypic score (r=0.92, p<0.001), and the MEroyalblue module showed the strongest negative correlation with the PBA phenotypic score. Figure 13 As shown, GOBP enrichment analysis of the MEroyalblue module showed that it was significantly enriched in cytoplasmic transport biological processes (corrected p = 0.002).

[0032] Based on the above results, the following three types of features are included as candidate features for subsequent machine learning models: (1) differentially expressed proteins screened based on DIA proteomics; (2) core proteins of the PBA association module based on WGCNA; and (3) exosome surface driving markers. The steps for constructing the machine learning model include: The first stage is univariate pre-screening, where ROC AUC is calculated for each of the 187 candidate features, and initial features with ROC AUC ≥ 0.65 are selected, resulting in the retention of 72 features.

[0033] The second stage involves bootstrap stability selection. One hundred bootstrap samplings are performed on the 72 features to select stable features with a frequency ≥ 50%, resulting in 34 stable features retained. This effectively avoids model overfitting. For example... Figure 14 As shown in A, the selection frequency of proteins such as TGM3 and PSMB4 all exceeded 90%.

[0034] A dual-algorithm parallel cross-validation strategy was adopted: a third-stage elastic network compression and a fourth-stage random forest validation. In the third stage, the elastic network algorithm (regularization parameter α=0.5) was used to compress 34 stable features using L1 / L2 regularization, selecting 33 key features with non-zero coefficients. In the fourth stage, the random forest algorithm (number of decision trees ntree=500) was used to rank the stable features by their Gini coefficient importance, selecting the top 15 features. The intersection of the two selection results yielded 14 internal core proteins, such as... Figure 14 As shown in B, the random forest feature importance analysis shows that proteins such as KRT80 and TGM3 have the highest contribution in the sample classification process. The 14 core proteins were: TGM3 (transglutaminase 3), PSMB4 (proteasome 20S subunit β4), KRT80 (keratin 80), KRT2 (keratin 2), DSG1 (desmosome core glycoprotein 1), CPOX (coprophyrinogen oxidase), CPN1 (carboxypeptidase N1), SUB1 (RNA polymerase II transcription activator SUB1), LSM2 (U6 small nucleoribonucleoprotein homolog 2), RAB34 (RAS oncogene family member RAB34), PZP (pregnancy zone protein), SERPINA5 (serine protease inhibitor A5), HYDIN (hydrocephalus inducible protein), and PGLYRP1 (peptidoglycan recognition protein 1). Among them, the bootstrap selection frequency of transglutaminase 3 and proteasome 20S subunit β4 was ≥90%, and the bootstrap selection frequency of keratin 80 was 88%. The single index ROC AUC reached 0.816.

[0035] It should be noted that the 14 core internal proteins are widely involved in key pathways of tumorigenesis and development: TGM3 participates in keratinization capsule formation and epithelial-mesenchymal transition, regulating tumor invasion and metastasis; PSMB4 participates in the proteasome pathway, regulating tumor cell protein homeostasis and proliferative activity; KRT80 participates in keratinization capsule formation and cytoskeleton remodeling, associated with tumor cell invasiveness; DSG1 participates in intercellular adhesion junctions, inhibiting tumor cell dissociation and distant metastasis; CPOX participates in heme metabolism, regulating tumor cell energy metabolism and proliferation; RAB34 participates in vesicle transport and exosome secretion, mediating tumor microenvironment regulation; SERPINA5 participates in protease inhibition and immune regulation, associated with tumor invasion and metastasis; PGLYRP1 participates in innate immune response and inflammation regulation, mediating tumor immune surveillance, etc. These functional annotations provide a solid biological basis for the diagnostic efficacy of the model.

[0036] Furthermore, a random forest model (sample size = 48 cases, ntree = 500, 3 candidate features per node) constructed based on 14 internal core protein markers (TGM3, PSMB4, KRT80, KRT2, DSG1, CPOX, CPN1, SUB1, LSM2, RAB34, PZP, SERPINA5, HYDIN, PGLYRP1) demonstrated excellent classification performance in distinguishing between colorectal cancer and colorectal adenoma. Figure 14 As shown in C, the difference in risk scores (marker risk scores) distinguishing CRC from Ad was statistically significant in the core cohort of 48 cases (sample size = 48 cases) (p < 0.05). < 0.001). For example... Figure 14 As shown in D, after stratified 10-fold cross-validation (cross-validation critical value = 0.565), the AUC reached 0.960 (95% confidence interval: 0.915-1.000); after random forest out-of-bag validation, the AUC was 0.950 (95% confidence interval: 0.896-1.000). Both ROC curves are highly concentrated in the upper left corner, and the AUCs both exceed 0.95, fully demonstrating that the random forest model for the biomarkers of these 14 internal core proteins has robust and accurate discriminative ability.

[0037] To avoid the risk of model overfitting caused by small sample data and to verify the generality of the model, such as Figure 15As shown, further independent validation was conducted using two external public blood exosome proteome datasets, mmc4 (iProX database ID: IPX0007566000) and mmc5 (iProX database ID: IPX0009239000). In the mmc4 cohort (20 colorectal cancer patients, 40 healthy controls) and the mmc5 cohort (21 colorectal cancer patients, 16 polyps, 21 healthy controls), the markers of 14 core internal proteins still exhibited excellent diagnostic sensitivity and specificity, and their overall expression trends were highly consistent with the results of the internal cohorts. Among them, markers such as SERPINA5 (serine protease inhibitor A5), B2M (β2-microglobulin), and PZP (pregnancy zone protein) were significantly overexpressed in the colorectal cancer (CRC) group, and the expression levels of some markers showed a gradual upward trend with the progression of disease from healthy controls → colorectal adenoma / polyps → colorectal cancer.

[0038] like Figures 16-17 As shown, in the single-cell transcriptome cohorts GSE132465 and GSE178341, the markers of 14 internal core proteins mainly originated from epithelial cells, which is highly consistent with the epithelial tissue origin of CRC. The expression patterns of the two datasets are highly consistent, proving that the cell origin analysis results are stable.

[0039] like Figures 18-20 As shown, all 32 candidate biomarkers matched in the TCGA-COAD and TCGA-READ cohorts. In the TCGA-COAD cohort, KRT80 was significantly upregulated in colorectal cancer tumor tissues (log2FC=6.38, p<0.001). Cox univariate overall survival (OS) analysis identified 5 biomarkers significantly associated with CRC prognosis (p<0.05). Furthermore, 20 biomarkers remained significantly different in the TCGA-READ cohort after multiple comparison correction.

[0040] like Figure 21 As shown, in the GSE39582 dataset, among 536 CRC patients with survival data, univariate Cox regression analysis showed that 7 biomarkers had significant prognostic value (p<0.05).

[0041] like Figure 22 As shown, most biomarkers exhibit stable differential expression in colorectal cancer tumor tissues, with KRT80, DSG1, and ANXA3 (Annexin A3) being significantly upregulated in tumor tissues.

[0042] like Figure 23As shown, in the TCGA-COAD cohort, KRT80 was significantly overexpressed in colorectal cancer tumor tissues: the average relative expression level in the colorectal cancer group was 1620.9, while that in the healthy control group was only 19.5. It is the gene with the most significant difference among the 14 internal core proteins and has a very strong diagnostic and differential ability for colorectal cancer.

[0043] like Figure 24 As shown, GOBP enrichment analysis revealed that in the KRT80 high-expression group, pathways such as epithelial cell differentiation, keratinization, and epidermal development were significantly activated, while immune-related pathways such as immune defense and adaptive immune response were significantly inhibited. Figure 25 As shown, KEGG pathway enrichment analysis revealed significant enrichment in multiple pathways, including nucleocytoplasmic transport, ubiquitin-mediated proteolysis, and mitophagy.

[0044] like Figure 26 As shown, the association analysis revealed a significant positive correlation between the expression level of KRT80 and the abundance of the Cluster 2 subpopulation obtained by PBA detection (Spearman's rho = 0.67, p = 0.033), demonstrating that KRT80 mainly originates from epithelial-derived CRC-specific exosome subpopulations.

[0045] The above results indicate that KRT80 can not only be included as a key feature in the biomarkers of 14 internal core proteins, but also serve as a standalone in vitro target for early CRC risk detection. In other words, KRT80 as a single biomarker can be used independently as a biomarker for early colorectal cancer risk screening. Its high fold change (log2FC) of 6.38 and single-index ROC AUC of 0.816 in the TCGA-COAD cohort fully validate the clinical application value of this target in rapid screening scenarios at the grassroots level.

[0046] Based on the above results, the KRT80 single-indicator biomarker can be used independently to differentiate between colorectal cancer and colorectal adenoma. This is an example, not a limitation. The detection principle is as follows: the expression level of KRT80 protein is directly compared with a preset threshold to determine the risk assessment result. This preset threshold is the optimal cut-off value determined using the Youden index maximization method. Specifically, the optimal cut-off value is the KRT80 protein expression level corresponding to the maximum sum of sensitivity and specificity minus 1 (Youden index). This preset threshold needs to be recalibrated and determined through receiver operating characteristic (ROC) curve analysis in an independent large-sample validation cohort to ensure its applicability in the target population.

[0047] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0048] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.

Claims

1. A biomarker for differentiating colorectal cancer from colorectal adenoma, characterized in that, Choose from any of the following: Exosome surface-driven markers consisting of SELL, CD177, ITGAL, OLR1, B2M, and LAG3; The internal core protein is composed of TGM3, PSMB4, KRT80, KRT2, DSG1, CPOX, CPN1, SUB1, LSM2, RAB34, PZP, SERPINA5, HYDIN, and PGLYRP1. KRT80 single-indicator marker.

2. The biomarker for distinguishing colorectal cancer from colorectal adenoma according to claim 1, characterized in that, The biomarkers are derived from plasma exosomes, which are obtained by immunoaffinity magnetic bead separation. The plasma exosomes meet the requirement that the main peak particle size is located in the range of 60nm-120nm, and the positive rates of the three exosome biomarkers CD9, CD63, and CD81 are all ≥90%.

3. The application of a biomarker for distinguishing colorectal cancer from colorectal adenoma as described in any one of claims 1-2, characterized in that, This kit is used to prepare a reagent to differentiate between colorectal cancer and colorectal adenoma.

4. The application of the marker for distinguishing colorectal cancer from colorectal adenoma according to claim 3, characterized in that, The kit includes: The marker information acquisition unit is used to acquire the detection information of the marker. The analysis and evaluation unit is used to determine the risk assessment result based on the detection information of the marker.