Serum protein and metabolic marker for diagnosis or auxiliary diagnosis of colorectal cancer or ulcerative colitis and application of serum protein and metabolic marker

By combining serum proteomics and metabolomics analysis methods, machine learning is used to establish a serum biomarker combination model, which solves the problem of early diagnosis of colorectal cancer related to ulcerative colitis, and achieves an efficient and low-trauma diagnostic effect.

CN120142669APending Publication Date: 2025-06-13LONGHUA HOSPITAL SHANGHAI UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510212220.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult to effectively carry out the early diagnosis of ulcerative colitis-related colorectal cancer in the prior art, especially in large-scale population screening and repeated examinations of high-risk populations. Traditional methods have limitations and high false positive rates.

Method used

Using large sample cohort omics data, a method combining serum proteomics and metabolomics analysis is developed to develop a combination of serum biomarker combination models for early diagnosis of ulcerative colitis-associated colorectal cancer, including specific protein markers and metabolic markers.

Benefits of technology

It has achieved high sensitivity and specificity for early diagnosis of ulcerative colitis-related colorectal cancer, reduced the false positive rate, and provided a low-trauma, simple and efficient diagnostic tool suitable for large-scale population screening and monitoring of high-risk populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120142669A_ABST
    Figure CN120142669A_ABST
Patent Text Reader

Abstract

The invention discloses a serum protein and metabolism marker for diagnosis or auxiliary diagnosis of colorectal cancer or ulcerative colitis, the serum protein and metabolism marker comprises 10 serum protein markers or / and 11 serum metabolites, the serum protein markers comprise APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP and CSF1, and the serum metabolism markers comprise LysoPA 18: 2, LysoPA 18: 3, Valproic acid, LysoPA 16: 0, Cis-4-DGTS 12: 0, Choline, Acylcarnitine 19: 3 and Cis-4-. By screening differentially expressed serum proteins and metabolites of patients with colorectal cancer and ulcerative colitis, machine learning is utilized to screen markers and construct a diagnosis model, so that diagnosis and auxiliary diagnosis of colorectal cancer are facilitated. In addition, the invention further discloses application of the serum protein and the metabolic marker, and the serum protein marker and the serum metabolic marker are used independently or in a combined mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical diagnosis, and particularly relates to a diagnostic model for ulcerative colitis-related CRC and its application. In particular, it relates to a serum protein-metabolite biomarker combination model for analyzing the expression levels of serum proteins APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, CSF1 and serum metabolites LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, LysoPA 18:1 separately or in combination, for the early diagnosis of ulcerative colitis-related colorectal cancer. Background Art

[0002] Colorectal cancer (CRC) is a highly prevalent malignant tumor of the digestive system globally. Despite the continuous progress of screening, diagnosis, and treatment technologies for CRC, it remains one of the cancers with the highest mortality rates, seriously threatening human health. [1] In China, the number of CRC patients ranks first in the world, accounting for 9.5% of the causes of cancer deaths, with a serious disease burden. [2] Currently, the most effective way to reduce the burden of CRC is still early diagnosis through population screening. [3] .

[0003] CRC is a typical inflammation-dependent tumor, and "inflammation-cancer transformation" is the key pathological evolution in the occurrence of CRC. The risk of developing CRC in patients with ulcerative colitis (UC) is significantly higher than that in the healthy population, and CRC is the main cause of colectomy and death in this population. [4] Compared with sporadic CRC, the risk of death from UC-related CRC is high, and the 5-year survival rate is poor. A 30-year cohort study from the Netherlands evaluated the reasons for surgical treatment of UC patients from 1991 to 2020. Pathological examinations showed that 16.9% of the patients developed CRC after a median disease course of 11 years, and the proportion of patients undergoing surgery due to CRC increased with the prolongation of the disease course. [5] A meta-analysis of 31,287 UC patients in 44 studies from Asian countries reported that the cumulative risk of CRC increased continuously with the prolongation of the disease course, and the lesion risks at 20 and 30 years were approximately 240 times and 700 times higher than those at 10 years, respectively. [6]In summary, the time window for UC to develop into CRC is usually more than ten years, and the longer the disease course, the higher the cumulative risk of carcinogenesis. Therefore, regular monitoring is required.

[0004] Epidemiological data show that compared with Western industrialized populations, the incidence of UC is increasing more rapidly in the Asia-Pacific region [4] , which leads to an increasing number of high-risk populations for UC-related CRC year by year and requires clinical attention. At present, colonoscopy is the main examination for high-risk populations of clinical CRC [1] . However, the limitations of colonoscopy, such as the dependence on instruments and professional operators, high cost, time-consuming, etc., make it not suitable for large-scale population screening and repeated examinations of high-risk populations; if the patient's bowel preparation is poor, it is even more unfavorable for the discovery of clinical lesions; moreover, mucosal biopsy has limitations in differentiating dysplasia and inflammation; at the same time, the certain damage and complication risks of this invasive examination make the patient compliance poor. In addition, fecal occult blood is a commonly used screening method in clinical practice. Its simple method makes it highly acceptable to patients, but its results are easily interfered with, with a high false positive rate and low sensitivity, and it is difficult to accurately detect early carcinogenesis.

[0005] Serum biomarker detection provides a non-invasive diagnostic approach, which is easier to obtain and less invasive [7] . Blood proteins that appear in the circulation due to cell leakage or active secretion can comprehensively understand human health status and diseases [8] , and serve as the main repository of biomarkers and therapeutic targets, having the most intrinsic predictive potential for diseases [9] . Mass spectrometry-based serum proteomics has been successfully applied to disease diagnosis and the characterization of protein change trajectories in diseases. Many studies have shown that it is an optimal technique for discovering biomarkers in this easily accessible serum sample

[10] . In addition, metabolites are closely related to the occurrence and development of CRC. Serum metabolites have the potential for disease diagnosis and have received increasing attention [7] . Liquid chromatography-mass spectrometry (LC-MS)-based metabolomics is a powerful tool for analyzing metabolite changes and has been used to decipher the metabolic reprogramming of various diseases and as diagnostic markers

[11] . Therefore, the combined analysis of serum proteome and metabolome can reveal disease-related pathways, discover new disease diagnostic markers, and is expected to develop a less invasive, simple, effective and inexpensive diagnostic tool to monitor and diagnose UC, a high-risk population for CRC.

[0006] Although there have been many articles and patents reporting on the discovery of serum CRC markers in recent years, the sample sizes are usually small, and the proteome and metabolome are not organically combined. Moreover, in most cases, the patient population of UC has not been taken into account, and CRC is highly heterogeneous. The detection of a single indicator often lacks specificity, sensitivity, and diagnostic value. A combined detection method needs to be adopted for comprehensive analysis to improve the diagnostic accuracy and have practical application value for the early diagnosis of CRC. Therefore, searching for serum protein-metabolite marker combinations and constructing a CRC diagnosis model have important clinical value.

[0007] Considering the above background, searching for efficient, minimally invasive, and easily implementable early screening methods is the focus and pain point of current clinical research on UC-related CRC.

[0008] References:

[0009] [1] Bretthauer M, M, Wieszczy P, et al. Effect of Colonoscopy Screening on Risks of Colorectal Cancer and Related Death. N Engl J Med. 2022;387(17):1547 - 1556. doi:10.1056 / NEJMoa2208375

[0010] [2] Zhou X, Han J, Zuo A, et al. THBS2+ cancer-associated fibroblasts promote EMT leading to oxaliplatin resistance via COL8A1-mediated PI3K / AKT activation in colorectal cancer. Mol Cancer. 2024;23(1):282. Published 2024 Dec 28. doi:10.1186 / s12943-024-02180-y

[0011] [3]Cervantes A, Adam R, Roselló S, et al. Metastatic colorectal cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol. 2023;34(1):10 - 32. doi:10.1016 / j.annonc.2022.10.003

[0012] [4]Shah SC, Itzkowitz SH. Colorectal Cancer in Inflammatory Bowel Disease: Mechanisms and Management. Gastroenterology. 2022;162(3):715 - 730.e3. doi:10.1053 / j.gastro.2021.10.035

[0013] [5]Heuthorst L, Harbech H, Snijder HJ, et al. Increased Proportion of Colorectal Cancer in Patients With Ulcerative Colitis Undergoing Surgery in the Netherlands. Am J Gastroenterol. 2023;118(5):848 - 854. doi:10.14309 / ajg.0000000000002099

[0014] [6]Bopanna S, Ananthakrishnan AN, Kedia S, Yajnik V, Ahuja V. Risk of colorectal cancer in Asian patients with ulcerative colitis: a systematic review and meta-analysis. Lancet Gastroenterol Hepatol. 2017;2(4):269 - 276. doi:10.1016 / S2468 - 1253(17)30004 - 3

[0015] [7]Yin H,Xie J,Xing S,et al.Machine learning-based analysisidentifies and validates serum exosomal proteomic signatures for thediagnosis of colorectal cancer.Cell Rep Med.2024;5(8):101689.doi:10.1016 / j.xcrm.2024.101689

[0016] [8]Anderson NL,Anderson NG.The human plasma proteome:history,character,and diagnostic prospects[published correction appears in Mol CellProteomics.2003Jan;2(1):50].Mol Cell Proteomics.2002;1(11):845-867.doi:10.1074 / mcp.r200007-mcp200

[0017] [9]Suhre K,McCarthy MI,Schwenk JM.Genetics meets proteomics:perspectives for large population-based studies.Nat Rev Genet.2021;22(1):19-37.doi:10.1038 / s41576-020-0268-2

[0018]

[10] Qu R,Zhang Z,Fu W.Potential of Serum Glycoproteome Profiling inPrediction of Advanced Adenomas and Colorectal Carcinoma:IndividualHeterogeneity Should Be Taken Into Account.Gastroenterology.2024;166(5):946.doi:10.1053 / j.gastro.2023.12.017

[0019]

[11] Shao Y,Li T,Liu Z,et al.Comprehensive metabolic profiling of Parkinson's disease by liquid chromatography-mass spectrometry.Mol Neurodegener.2021;16(1):4.Published 2021 Jan 23.doi:10.1186 / s13024-021-00425-8 Summary of the Invention

[0020] In view of the above deficiencies of the prior art, according to the embodiments of the present invention, it is desired to develop a serum biomarker combination for the diagnosis of UC-related colorectal cancer using large-sample cohort omics data and establish a diagnostic model, and the biomarker combination includes serum proteins and / or serum metabolites. This solves the technical problems in the related art that the UC patient population has not been concerned, the sample size is small, and there is strong heterogeneity in the single-molecule model. In addition, the present invention also hopes to provide a biomarker combination using easily accessible peripheral serum samples, which can be used for the early diagnosis of CRC in healthy populations or UC patients, and can effectively distinguish healthy populations from UC patients, and is suitable for large-scale CRC screening and diagnosis.

[0021] In view of the technical problems in the related art that the UC patient population has not been concerned, the sample size is small, and CRC has high heterogeneity, and the single-molecule model often has insufficient specificity, sensitivity, and diagnostic value. The technical solution adopted by the present invention is as follows:

[0022] First, peripheral serum samples of 101 healthy subjects (Negative control, NC), 99 patients with ulcerative colitis (UC), and 108 patients with colorectal cancer (CRC) were collected for proteomics and metabolomics detection. Through population stratification, they were randomly divided into a discovery cohort (207 cases) and a validation cohort (101 cases). In the discovery cohort (71 NC, 65 UC patients, 71 CRC patients), differentially expressed proteins and metabolites that continuously changed in the NC-UC-CRC progression cohort were identified, serum proteins and metabolites with diagnostic potential for CRC were screened, and then machine learning was used to separately and jointly model the protein and metabolite markers.

[0023] The present invention innovatively proposes a method that combines serum proteomics and metabolomics analysis, and uses machine learning methods to establish potential protein and metabolite biomarker models for the early diagnosis of CRC in UC patients. The finally determined serum protein biomarker combination includes APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, CSF1, and the metabolite biomarker combination includes LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, LysoPA 18:1. The protein and metabolite biomarker combinations show good diagnostic efficacy whether modeled separately or combined, solving the technical problems such as insufficient diagnostic value of a single molecular model in the related art.

[0024] Finally, the diagnostic efficacy of the model for the diagnosis of UC inflammation-cancer transformation and the diagnostic efficacy of sporadic CRC were verified in a validation cohort (30 HC, 34 UC patients, 37 CRC patients).

[0025] The obtained protein biomarker combination model and / or metabolite biomarker combination model can be used to prepare a diagnostic kit for UC-related CRC and sporadic CRC with less trauma, simplicity, effectiveness and low cost.

[0026] Using the serum proteins and / or metabolite biomarkers in the present invention, it is possible to determine the probability of an NC individual being in a state of having UC or CRC, or the probability of a UC patient being in a healthy state or CRC, and it can be used for the early detection or auxiliary detection of colorectal cancer or ulcerative colitis.

[0027] The inventors found that in both the discovery cohort and the validation cohort, all three models showed good CRC diagnostic efficacy. In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) of the protein biomarker model (refer to Appendix Figure 1A ) for NC vs. CRC diagnosis was 1 (95% CI = 1 - 1), 0.997 (95% CI = 0.992 - 1) (refer to Appendix Figure 2A ); the area under the ROC curve (AUC) of the metabolite biomarker model (refer to Appendix Figure 1B ) for NC vs. CRC diagnosis was 1 (95% CI = 1 - 1), 1 (95% CI = 1 - 1) (refer to Appendix Figure 2B); Protein-metabolic biomarker combination model (see Appendix Figure 1C ) The area under the ROC curve (AUC) for the diagnosis of NC vs. CRC was 1 (95% CI = 1 - 1), 1 (95% CI = 1 - 1) (see Appendix Figure 2C ). In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) for the diagnosis of UC vs. CRC by the protein biomarker model was 0.907 (95% CI = 0.855 - 0.96) (see Appendix Figure 2A ), 0.897 (95% CI = 0.825 - 0.969); the area under the ROC curve (AUC) for the diagnosis of UC vs. CRC by the metabolic biomarker model was 0.956 (95% CI = 0.924 - 0.987), 0.988 (95% CI = 0.971 - 1) (see Appendix Figure 2B ); the area under the ROC curve (AUC) for the diagnosis of UC vs. CRC by the protein-metabolic biomarker combination model was 0.971 (95% CI = 0.948 - 0.994), 0.994 (95% CI = 0.985 - 1) (see Appendix Figure 2C ). In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) for the diagnosis of NC vs. UC by the protein biomarker model was 0.934 (95% CI = 0.896 - 0.972), 0.838 (95% CI = 0.742 - 0.934) (see Appendix Figure 2A ); the area under the ROC curve (AUC) for the diagnosis of NC vs. UC by the metabolic biomarker model was 0.864 (95% CI = 0.803 - 0.925), 0.845 (95% CI = 0.751 - 0.938) (see Appendix Figure 2B ); the area under the ROC curve (AUC) for the diagnosis of NC vs. UC by the protein-metabolic biomarker combination model was 0.901 (95% CI = 0.849 - 0.952), 0.89 (95% CI = 0.812 - 0.967) (see Appendix Figure 2C) The above results show that both the protein biomarker model and the metabolite biomarker model have good diagnostic capabilities. The combined model has a stronger diagnostic efficacy. Model selection can be made according to the actual situation, and it can be widely applied to the screening and monitoring of colorectal cancer. The present invention will contribute to promoting the application of minimally invasive serum biomarkers in the long-term monitoring and diagnosis of UC-related CRC.

[0028] The advantages of the present invention are as follows:

[0029] (1) It is a novel serum biomarker, which is more accurate, sensitive, stable and minimally invasive, and is easily accepted by patients. Its successful development contributes to the diagnosis of colorectal cancer.

[0030] (2) For the first time, the protein combinations of APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, CSF1 and the metabolite combinations of LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, LysoPA 18:1 are determined as biomarkers for the diagnosis of colorectal cancer and ulcerative colitis, and a model for the early diagnosis of colorectal cancer in UC patients is established at the same time.

[0031] (3) The diagnostic efficacies of the protein biomarker model, the metabolite biomarker model and the combined model are respectively detected, providing multiple applicable diagnostic model options.

[0032] (4) The model designed based on multiple proteins and metabolites is superior to the model with a single biomarker in terms of sensitivity and specificity, and can reduce the error caused by individual expression differences of a single index, making the results more accurate.

[0033] (5) The present invention has a large sample size, and the diagnostic model is tested through a discovery cohort and a validation cohort, which is more persuasive. Description of the Drawings

[0034] Figure 1A Showing the top 10 serum differential proteins screened according to the MeanDecreaseGini importance ranking of the diagnostic model.

[0035] Figure 1B Showing the top 11 serum differential metabolites screened according to the MeanDecreaseGini importance ranking of the diagnostic model.

[0036] Figure 1CShow the ranking of the MeanDecreaseGini importance of each biomarker in the protein-metabolite combined diagnostic model.

[0037] Figure 2A Show the diagnostic efficacy of the protein biomarker combination model for NC-UC-CRC in the discovery cohort and the validation cohort by ROC curve analysis.

[0038] Figure 2B Show the diagnostic efficacy of the metabolite biomarker combination model for NC-UC-CRC in the discovery cohort and the validation cohort by ROC curve analysis.

[0039] Figure 2C Show the diagnostic efficacy of the protein-metabolite biomarker combination model for NC-UC-CRC in the discovery cohort and the validation cohort by ROC curve analysis. Detailed implementation mode

[0040] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. These embodiments should be understood as only for illustrating the present invention and not for limiting the protection scope of the present invention. After reading the content recorded in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

[0041] The present invention collects peripheral serum samples of 101 healthy subjects (Negative control, NC), 99 patients with ulcerative colitis (UC), and 108 patients with colorectal cancer (CRC), conducts proteomics and metabolomics detection, divides the population into a discovery cohort (207 cases) and a validation cohort (101 cases) by stratification and randomization, identifies differential proteins and metabolites that continuously change in the NC-UC-CRC cohort in the discovery cohort (71 NC, 65 UC patients, 71 CRC patients), screens serum proteins and metabolites with diagnostic potential for CRC, and then separately and jointly models the protein and metabolite biomarkers through machine learning and validates them in the discovery cohort and the validation cohort.

[0042] 1. Experimental materials and methods

[0043] 1.1 Research cohort

[0044] Both the discovery cohort and the validation cohort (a total of 308 cases) are from Longhua Hospital Affiliated to Shanghai University of Traditional Chinese Medicine and have obtained the approval of the Human Research Ethics Committee of Longhua Hospital Affiliated to Shanghai University of Traditional Chinese Medicine (Shanghai, China).

[0045] 1.2 Collection of serum samples

[0046] Collect 1 ml of fasting peripheral blood from the subjects, place it at room temperature for 30 min, and then centrifuge it at 3000 rpm for 10 min to separate the serum. Carefully collect the supernatant, i.e., the serum (do not aspirate the underlying blood clots).

[0047] 1.3 Serum proteome detection

[0048] Normalize the volume of the protein solution according to the BCA quantification results. Resolubilize each sample with 8 M urea denaturant, and add dithiothreitol (DTT) at a final concentration of 5 mM and incubate at 37 °C in a water bath for 1 hour for reduction. Then add iodoacetamide (IAM) at a final concentration of 10 mM and react in the dark for 45 min. Then dilute the sample to a final concentration of 50 mM with 500 mM ammonium bicarbonate (ABC), add trypsin at a ratio of enzyme:protein = 1:35, and digest overnight at 37 °C. Stop the digestion with 10% TFA to make the pH < 2. Desalt the digested sample with MonoSpin C18 and dry it using a vacuum concentrator. Store the sample at -80 °C for mass spectrometry analysis.

[0049] The dried peptide samples were reconstituted with mobile phase A (100% H2O, 0.1% FA). After centrifugation at 20,000 g for 10 min, the supernatant was taken for injection. Separation was performed using nanoElute from Bruker. The sample first entered the trap column for enrichment and desalting, and then was connected in series with a self-packed C18 column (inner diameter 75 μm, particle size of column material 1.8 μm, column length approximately 25 cm). Separation was carried out at a flow rate of 300 nL / min using the following effective gradient: 0 min, 2% mobile phase B (100% ACN, 0.1% FA); 0 - 45 min, mobile phase B linearly increased from 2% to 22%; 45 - 50 min, mobile phase B increased from 22% to 35%; 50 - 55 min, mobile phase B increased from 35% to 80%; 55 - 60 min, 80% mobile phase B. The peptides separated by liquid chromatography were ionized by the CaptiveSpray source and then entered the tandem mass spectrometer timsTOF Pro2 (Bruker Corporation, Billerica, MA) for detection in DIA (Data Independent Acquisition) mode. The main parameter settings were as follows: the ion source voltage was set at 1.7 kV; the ion mobility range was 0.76 - 1.29 V·s / cm2, the first-stage mass spectrometry scan range was 452 - 1152 m / z, and peaks with an intensity higher than 2500 would be detected. The 452 - 1152 m / z range was divided into 4 steps, each step had 7 windows, the Number of Mobility Windows was set to 2, and a total of 56 windows were used for continuous window fragmentation and information collection. The fragmentation mode was CID, the fragmentation energy was 20 - 59 eV, the mass width of each window was 25, and the cycle time for one DIA scan was 1.59 s.

[0050] 1.4 Serum metabolome detection

[0051] Mix 100 μL of plasma sample with 500 μL of extraction solution (methanol to water ratio of 3:1, containing isotope-labeled internal standard). Homogenize the sample mixture by grinding for 4 minutes at 35 Hz and sonication in ice water for 5 minutes, then incubate at 40 °C for 60 minutes to precipitate proteins. Then centrifuge the sample at 12,000 rpm for 15 minutes at 4 °C to collect the supernatant. Another 10 μL of all serum samples from each experimental group were mixed to obtain a mixed serum sample, which was pretreated as described above to obtain a serum quality control sample (QC). Perform untargeted LC-MS / MS to analyze metabolites in the serum. Use a Vanquish Flex UHPLC system (Thermo Fisher Scientific, Waltham, MA) coupled with an ACQUITY UPLC HSS T3 column (2.1 mm × 100 mm, 1.8 μm; Waters, Milford, MA) for chromatographic separation. The mobile phase consists of 5 mmol / L ammonium formate and 5 mmol / L formic acid in water (A) and acetonitrile (B). The autosampler temperature is 4 °C and the injection volume is 2 μL.

[0052] 1.5 Serum proteome analysis

[0053] In terms of proteomic bioinformatics, using the SwissProt human protein sequence as a database, generate a theoretical spectral library using the DIA-NN software [2] with deep learning methods, and integrate the mass spectrometry data collected by DIA for protein identification and quantitative information extraction. Set the digestion parameters as Trypsin / P, the maximum number of missed cleavages as 2, the fixed modification as cysteine reduction alkylation, and the variable modifications as methionine oxidation and protein N-terminal acetylation. The protein identification confidence strategy is Target-decoy, and the protein identification threshold is a spectral false positive <0.01 and a protein false positive <0.01.

[0054] 1.6 Metabolite data analysis and metabolite annotation

[0055] The non-targeted LC-MS / MS data in positive and negative modes were converted to the mzXML format using ProteoWizard and processed by the R package XCMS to generate a data matrix consisting of retention time (RT), mass-to-charge ratio (m / z) values, and peak abundances. Peaks with missing values greater than 80% in the samples and greater than 50% in the QC samples were removed, and the missing values were imputed using the K-nearest neighbor method. The data was normalized using probability quotient normalization (PQN). After normalization, peaks with a relative standard deviation > 30% in the quality control samples were filtered. Metabolite annotation was performed by matching the secondary mass spectra of metabolites in the experiment with the secondary spectra of metabolite standards (self-built library, MASSBANK, LipidBlast, and HMDB). The error in the primary molecular weight was 0.01 Da, and the error in the secondary molecular weight was 0.05 Da. The matching score was for reliable metabolites.

[0056] 2. Model construction and validation

[0057] 2.1 Construction and validation of the protein biomarker diagnostic model

[0058] In the present invention, serum proteins that continuously change in the NC-UC-CRC patient cohort were screened to identify serum proteins with the potential to be diagnostic biomarkers for whether UC patients develop CRC. The results showed that a total of 86 proteins were continuously upregulated / downregulated in the NC-UC-CRC disease progression cohort.

[0059] Random forest modeling was performed using the Random Forest package in R language. The process was as follows: First, the dataset was randomly divided into a training set (2 / 3) and a test set (1 / 3). The data of the control group and the cancer group were taken, and the Random Forest package was used to perform overall modeling on all proteins, and the OBB (accuracy = 1 - OOB), as well as the importance rankings of MeanDecreaseAccuracy and MeanDecreaseGini, were obtained to screen out the top 10 differential genes: APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, CSF1 (see Appendix Figure 1A ). Among them, CRP and CSF1 were continuously upregulated in the NC-UC-CRC disease progression cohort, and APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, and RAB11B were continuously downregulated in the HC-UC-CRC disease progression cohort. The error rate (CV Error) was calculated using the 10-fold 5-time cross-validation method to test and adjust the machine learning model until the optimal model was obtained. Then, the optimal modeling combination was used to continue model construction and confirm the final model. First, in the discovery cohort (see Appendix Figure 2AThe UC group and CRC group of Figure 2A -2) were tested, then the NC group and CRC group were tested, and then the NC and UC groups were tested. The pROC package was used to calculate the AUC value and draw the ROC curve.

[0060] 2.2 Construction and validation of the metabolic biomarker diagnostic model

[0061] To evaluate the diagnostic potential of the selected differential metabolites, a machine learning model (random forest) was established through the R package and 10 5 -fold cross-validation was performed. First, it was tested on the UC group and CRC group of the discovery cohort (see attachment Figure 2B -1), then on the NC group and CRC group, and then on the NC and UC groups. The pROC package was used to calculate the AUC value and draw the ROC curve. The model was used to test the validation cohort (see attachment Figure 2B -2) of the UC group and CRC group, then on the NC group and CRC group, and then on the NC and UC groups. The pROC package was used to calculate the AUC value and draw the ROC curve.

[0062] 2.3 Construction and validation of the protein-metabolism combined model

[0063] The machine learning diagnostic models of proteins and metabolites were integrated and adjusted until the optimal model was obtained. First, it was tested on the UC group and CRC group of the discovery cohort (see attachment Figure 2C -1), then on the NC group and CRC group, and then on the NC and UC groups. The pROC package was used to calculate the AUC value and draw the ROC curve. The model was used to test the validation cohort (see attachment Figure 2C -2) of the UC group and CRC group, then on the NC group and CRC group, and then on the NC and UC groups. The pROC package was used to calculate the AUC value and draw the ROC curve.

[0064] 3. Determination of diagnostic results

[0065] Analyze the levels of APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, CSF1 proteins and metabolites such as LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, and LysoPA18:1 separately or in combination to obtain the probability of a subject being diagnosed with CRC or UC.

[0066] 4. Result Analysis

[0067] In both the discovery cohort and the validation cohort, the three models demonstrated good diagnostic efficacy for CRC and UC. In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) of the protein biomarker model (refer to Appendix Figure 1A ) for the diagnosis of UC vs. CRC was 0.907 (95% CI = 0.855 - 0.96) and 0.897 (95% CI = 0.825 - 0.969) respectively (refer to Appendix Figure 2A ); the area under the ROC curve (AUC) of the metabolite biomarker model (refer to Appendix Figure 1B ) for the diagnosis of UC vs. CRC was 0.956 (95% CI = 0.924 - 0.987) and 0.988 (95% CI = 0.971 - 1) respectively (refer to Appendix Figure 2B ); the area under the ROC curve (AUC) of the protein-metabolite biomarker combination model (refer to Appendix Figure 1C ) for the diagnosis of UC vs. CRC was 0.971 (95% CI = 0.948 - 0.994) and 0.994 (95% CI = 0.985 - 1) respectively (refer to Appendix Figure 2C ). In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) of the protein biomarker model for the diagnosis of NC vs. CRC was 1 (95% CI = 1 - 1) and 0.997 (95% CI = 0.992 - 1) respectively (refer to Appendix Figure 2A) The area under the ROC curve (AUC) of the metabolic biomarker model for the diagnosis of NC vs. CRC was 1 (95% CI = 1 - 1), 1 (95% CI = 1 - 1) (see attached Figure 2B ) The area under the ROC curve (AUC) of the protein-metabolic biomarker combined model for the diagnosis of NC vs. CRC was 1 (95% CI = 1 - 1), 1 (95% CI = 1 - 1) (see attached Figure 2C ). In the discovery cohort and the validation cohort, the area under the ROC curve (AUC) of the protein biomarker model for the diagnosis of NC vs. UC was 0.934 (95% CI = 0.896 - 0.972), 0.838 (95% CI = 0.742 - 0.934) (see attached Figure 2A ) The area under the ROC curve (AUC) of the metabolic biomarker model for the diagnosis of NC vs. UC was 0.864 (95% CI = 0.803 - 0.925), 0.845 (95% CI = 0.751 - 0.938) (see attached Figure 2B ) The area under the ROC curve (AUC) of the protein-metabolic biomarker combined model for the diagnosis of NC vs. UC was 0.901 (95% CI = 0.849 - 0.952), 0.89 (95% CI = 0.812 - 0.967) (see attached Figure 2C ).

[0068] The above results show that both the protein biomarker model and the metabolic biomarker model have good diagnostic capabilities, and the combined model has a stronger diagnostic efficacy. The model can be selected according to the actual situation and widely applied to the screening and monitoring of colorectal cancer and ulcerative colitis.

Claims

1. A serum protein and metabolic marker for diagnosing or assisting in the diagnosis of colorectal cancer or ulcerative colitis, characterized in that: The serum protein and metabolic markers consist of 10 serum protein markers or / and 11 serum metabolites, wherein the serum protein markers are APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, and CSF1; and the serum metabolic markers are LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, and LysoPA 18:

1.

2. Use of the serum protein and metabolic markers described in claim 1 in the preparation of a detection kit for diagnosing or assisting in the diagnosis of colorectal cancer or ulcerative colitis, wherein the serum protein markers and serum metabolic markers are used alone or in combination.

3. The use according to claim 2, characterized in that: The detection kit is used to diagnose or assist in diagnosing whether the subject is colorectal cancer, ulcerative colitis or normal.

4. A detection kit, characterized in that: The detection kit comprises serum protein and metabolic markers and reagents for detecting the expression levels of the serum protein and metabolic markers, wherein the serum protein and metabolic markers consist of 10 serum protein markers or / and 11 serum metabolites, wherein the serum protein markers are APOA4, APOC3, APOF, C4B, CLEC3B, FERMT3, NID2, RAB11B, CRP, and CSF1, and the serum metabolic markers are LysoPA 18:2, LysoPA 18:3, Valproic acid, LysoPA 16:0, cis-4-Decenedioic acid, Trp-Glu, LysoDGTS12:0, Choline, Acylcarnitine 19:3, cis-4-Hydroxy-L-proline, and LysoPA 18:1.

Citation Information

Patent Citations

  • Molecular marker with auxiliary functions of early diagnosis and prognosis monitoring on colorectal cancer, and application thereof

    CN107064508A

  • Plasma metabolism marker combination for diagnosing or monitoring colorectal cancer and application

    CN114705782A

  • Metabolic marker for early auxiliary diagnosis of colorectal cancer and application thereof

    CN116953119A