Marking composition for colorectal cancer detection and application thereof
Relevant metabolites are screened out through LC-MS technology and a Cox regression model is constructed, which solves the problem of insufficient accuracy in early diagnosis of colorectal cancer in the prior art, and realizes accurate evaluation of early diagnosis of colorectal cancer and the identification of high-risk patients.
Patent Information
- Application Number
- CN202510104984.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to provide a low-invasive and accurate method for early diagnosis of colorectal cancer, and the application of metabolic molecular markers in early diagnosis of colorectal cancer has not been fully explored.
A group of relevant metabolites markers were screened through LC-MS technology, including lauric acid, stearic acid, phosphatidylcholine, dehydroepiandrosterone sulfate, etc., and a multi-factor Cox regression model was constructed to predict the early diagnosis of colorectal cancer.
Accurate evaluation of early diagnosis of colorectal cancer has been achieved, and the ability to evaluate and predict early diagnosis of colorectal cancer patients has been improved. It can effectively identify high-risk patients, and monitor and intervene in the early stage in clinical practice to reduce the incidence and mortality of colorectal cancer patients.
Smart Images

Figure CN120161191A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine and relates to a biomarker composition for disease detection. Specifically, it relates to a novel biomarker composition for early diagnosis of colorectal cancer and its applications. This case is a divisional application of the application with the application number CN 202310192260.2 and the invention title "Biomarker Composition for Colorectal Cancer Detection and Its Applications". Background Art
[0002] Colorectal cancer (CRC) has become an increasingly serious challenge worldwide. In 2022, the National Cancer Center released the latest national cancer statistics. The incidence rate of CRC ranks second in China, and the mortality rate ranks fourth in China. Currently, it is believed that most CRCs develop from colorectal adenomas (CRA), among which the carcinogenesis rate of villous adenomas is the highest, and the larger the adenoma, the higher the carcinogenesis rate. Since the survival rate of late-stage radiotherapy and chemotherapy is low, the diagnosis at the CRA stage and the early stage of CRC is considered an effective way to improve the survival rate of colorectal cancer patients.
[0003] Currently, there are various methods for CRC detection, such as non-invasive methods like fecal occult blood test (FOBT) and carcinoembryonic antigen (CEA) test, as well as invasive examinations like colonoscopy. However, due to the low accuracy of non-invasive examinations and the damage caused by invasive examinations, the large-scale application of these methods is restricted. Therefore, a minimally invasive and accurate CRC detection method is needed. Metabolic molecular biomarkers refer to the expression levels of a group of small metabolite molecules, and a mathematical model is established through machine learning for predicting specific clinical targets. In recent years, the detection methods of metabolite molecules have been quite mature, including liquid chromatography-mass spectrometry (LC-MS), gas chromatography-mass spectrometry (GC-MS), and nuclear magnetic resonance technology (NMR), etc.
[0004] The working principle of high performance liquid chromatography (HPLC): The solvent bottle holds the solvent, which is called the mobile phase. The high-pressure pump is used to generate and meter the mobile phase with a specific flow rate, usually several milliliters per minute. The injector can introduce the sample into the continuously flowing mobile phase liquid stream, which carries the sample into the HPLC column. The column contains the chromatographic packing required for separation. This packing is called the stationary phase because it is fixed in place by the column hardware. A detector is needed to view the separated compound bands eluted from the HPLC column. The mobile phase leaves the detector and can be sent to waste or collected as needed. High-pressure tubing and fittings are used to interconnect the pump, injector, column, and detector components to form a pipeline for the mobile phase, sample, and separated compound bands. After the sample enters the HPLC system, the different components in the sample have different affinities (adsorption) for the column. By changing the proportion of the mobile phase, different components are eluted at different times to achieve the separation of the sample.
[0005] The detector is connected to a computer data workstation. Several different types of detectors have been developed to handle different sample compounds with extremely different characteristics. For example, if the compound can absorb ultraviolet light, a UV absorbance detector is used. If the compound emits fluorescence, a fluorescence detector is used. If the compound does not have either of these characteristics, a more general detector type, such as an evaporative light scattering detector, is used. A very effective method is to use multiple detectors in series. For example, UV and / or ELSD detectors can be combined with a mass spectrometer (MS) to analyze the results of chromatographic separation. In this way, more comprehensive analyte information can be obtained with a single injection. The practice of combining a mass spectrometer with an HPLC system is called LC-MS. The purpose of the combination is to perform separation by LC so that the substances entering the mass spectrometer are single substances to avoid mutual interference, which is beneficial for subsequent identification.
[0006] Principle of mass spectrometry (MS). Mass spectrometry is an instrument widely used for qualitative and quantitative analysis of molecules. When using a mass spectrometer to determine the mass of a molecule, the molecule must first be converted into gas-phase ions. To this end, the mass spectrometer transfers charge to the molecule and converts the resulting charged ion current into a proportional current, which is then read by a data system. The data system converts the current into digital information and presents it in the form of a mass spectrum. Ions can be generated in a variety of ways suitable for the target analyte: by laser ablation of a compound dissolved in a matrix and loaded onto a flat surface, such as MALDI; by allowing the compound to interact with high-energy particles or electrons, such as electron impact ionization (EI); by making ionization part of the transport process itself, such as the electrospray ionization (ESI) technique we have learned, which applies a high voltage to the eluent of a liquid chromatograph to generate ions in an aerosol. The mass spectrometer separates, detects, and determines ions based on their mass-to-charge ratio (m / z). A mass spectrum is generated by plotting the relative ion current (signal) against m / z. Small molecules usually carry only one charge: thus m / z is the mass (m) divided by 1. ("1" represents the proton added during the ionization process, denoted as [M+H] + , and [M-H] if the ion is formed by losing a proton - ). At the same time, if the ion is collided, the ion will further fragment to form smaller fragment ions. Even ions with the same mass-to-charge ratio will fragment into different ions depending on their structure, enabling substance identification. In this article, the identification of metabolites is also achieved by comparing the mass-to-charge ratio of ions and their secondary fragments.
[0007] With the development of metabolite molecule detection technology, more and more metabolite molecular markers have been discovered for the early diagnosis of colorectal cancer. Research shows that the application of different metabolite molecular markers can provide directions for the early diagnosis of the occurrence and development of colorectal cancer. However, there are few known studies on how to find a set of metabolite molecular marker combinations for the early diagnosis of colorectal cancer and an optimized mathematical model for prediction with good results. Therefore, selecting a reasonable method to screen out relevant metabolite molecular markers and clarify the potential molecular mechanism of colorectal cancer can provide directions for the early and accurate prevention, diagnosis, and treatment of colorectal cancer, and further improve the clinical diagnosis efficiency and judgment ability of colorectal cancer patients. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a novel model for the early diagnosis of colorectal cancer and its application to overcome the above technical problems existing in the current art.
[0009] The above object of the present invention is achieved by the following technical solutions:
[0010] The first aspect of the present invention provides a combination of related markers for the prognosis prediction and diagnosis of colorectal cancer.
[0011] Furthermore, the metabolic marker combination includes dodecanoic acid, stearic acid, phosphatidylcholine (16:0 / 9:0(CHO)) (PC(16:0 / 9:0(CHO))), dehydroepiandrosterone sulfate, capric acid, uric acid, acylcarnitine (9:1) (ACar(9:1)), pyruvic acid, pseudouridine, L-acetylcarnitine, phosphatidylinositol (18:0 / 22:6) (PI(18:0 / 22:6)), allixin.
[0012] The present invention discloses a marker composition for the detection of colorectal cancer, and the marker composition includes: dodecanoic acid, dehydroepiandrosterone sulfate, capric acid, uric acid.
[0013] Preferably, the marker composition is further selected from one or more of stearic acid, pyruvic acid, pseudouridine, L-acetylcarnitine.
[0014] Preferably, the marker composition includes: dodecanoic acid, stearic acid, dehydroepiandrosterone sulfate, capric acid, uric acid, L-acetylcarnitine.
[0015] Preferably, the marker composition includes: dodecanoic acid, stearic acid, dehydroepiandrosterone sulfate, capric acid, uric acid, pseudouridine.
[0016] Preferably, the biomarker composition comprises Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid.
[0017] Preferably, the biomarker composition comprises: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine, L-Acetylcarnitine.
[0018] Preferably, the biomarker composition comprises: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, L-Acetylcarnitine.
[0019] Preferably, the biomarker composition comprises: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine.
[0020] Preferably, the biomarker composition comprises: Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Stearic acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine.
[0021] Preferably, the biomarker composition consists of Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Stearic acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine.
[0022] The present invention discloses the use of the biomarker composition in the preparation of a reagent for detecting colorectal cancer.
[0023] The present invention discloses the use of the described marker composition in the preparation of a kit for detecting colorectal cancer.
[0024] Preferably, the test sample of the kit is blood.
[0025] The present invention discloses a kit for detecting colorectal cancer, which kit contains the described marker composition.
[0026] Preferably, the kit further includes a standard.
[0027] Another aspect of the present invention provides a method for constructing a risk assessment model for predicting the prognosis of colorectal cancer, and the construction method includes the following steps
[0028] (1) Candidate metabolite analysis: The total metabolites contained in the sample are detected by LC-MS. Then, through P-value screening based on the above 1026 metabolites and correlation analysis with clinical indicators (age, gender, BMI index, etc.), metabolite indicators are screened.
[0029] (2) Model establishment: The total metabolites obtained in step (1) are analyzed by LASSO Cox, and finally the model of early diagnosis metabolites is determined.
[0030] Based on the LC-MS-based metabolomics detection and analysis, the present invention screens and obtains metabolites related to colorectal cancer.
[0031] Through P-value screening and correlation analysis with clinical indicators such as age, gender, and BMI index, the present invention screens and obtains 118 candidate metabolite markers. Combining LASSO model analysis and fully considering the response signal situation of the markers in the sample, the level of background content, and the difficulty of obtaining standards, a combination of multiple metabolite markers is screened.
[0032] The present invention provides a novel model for the early diagnosis of colorectal cancer, and the stable effectiveness of the model is confirmed by multiple test data sets of colorectal cancer patients. The model provides reliable prognostic prediction and diagnostic markers for the evaluation of colorectal cancer patients, improves the evaluation and prediction ability for the early diagnosis of colorectal cancer patients, can effectively identify patients with high-risk colorectal cancer, and can be monitored early and intervened effectively in clinical practice, so as to reduce the incidence and mortality of colorectal cancer patients. The model has broad application prospects in clinical practice. Description of the Drawings
[0033] Figure 1 . Schematic diagram of the screening process of metabolic molecular markers related to colorectal cancer in plasma.
[0034] Figure 2. Schematic diagram of the comparison between the measured metabolite parent ions and daughter ions (upper part) and the parent ions and daughter ions in the database (lower part).
[0035] Figure 3 . ROC curve diagrams of 12 metabolites obtained by screening the training set through the LASSO model, CRA VS NC group.
[0036] Figure 4 . ROC curve diagrams of 12 metabolites obtained by screening the training set through the LASSO model, CRC VS NC group.
[0037] Figure 5 . ROC curve diagrams of 12 metabolites obtained by screening the training set through the LASSO model, CRC VS CRA group.
[0038] Figure 6 . ROC curve diagrams of the combination of 8 metabolites in the training set (CRA group VS NC group).
[0039] Figure 7 . ROC curve diagrams of the combination of 8 metabolites in the training set (CRC group VS NC group).
[0040] Figure 8 . ROC curve diagrams of the combination of 8 metabolites in the training set (CRC group VS CRA group).
[0041] Figure 9 . Comparison diagram of the parent ions and daughter ions of the metabolite Acetylcarnitine standard (left side) and the parent ions and daughter ions of the measured data (right side).
[0042] Figure 10 . Comparison diagram of the parent ions and daughter ions of the metabolite Dodecanoic acid standard (left side) and the parent ions and daughter ions of the measured data (right side).
[0043] Figure 11 . Comparison diagram of the parent ions and daughter ions of the metabolite Dehydroepiandrosterone sulfate standard (left side) and the parent ions and daughter ions of the measured data (right side).
[0044] Figure 12 . Comparison diagram of the parent ions and daughter ions of the metabolite Capric acid standard (left side) and the parent ions and daughter ions of the measured data (right side).
[0045] Figure 13 . Comparison diagram of the parent ions and daughter ions of the metabolite Uric acid standard (left side) and the parent ions and daughter ions of the measured data (right side).
[0046] Figure 14Comparison chart of precursor and product ions of metabolite Pyruvic acid standard (left) and measured data precursor and product ions (right).
[0047] Figure 15 Comparison chart of precursor and product ions of metabolite Pseudouridine standard (left) and measured data precursor and product ions (right).
[0048] Figure 16 Comparison chart of precursor and product ions of metabolite Stearic acid standard (left) and measured data precursor and product ions (right).
[0049] Figure 17 ROC curves of 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, L-Acetylcarnitine) in the training set (CRA group VS NC group).
[0050] Figure 18 ROC curves of 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine, L-Acetylcarnitine) in the training set (CRA group VS NC group).
[0051] Figure 19 ROC curves of 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine) in the training set (CRA group VS NC group).
[0052] Figure 20 ROC curves of 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine) in the training set (CRA group VS NC group).
[0053] Figure 21. ROC curves corresponding to 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, L-Acetylcarnitine) in the training set (CRA group VS NC group).
[0054] Figure 22 . ROC curves corresponding to 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid) in the training set (CRA group VS NC group).
[0055] Figure 23 . ROC curve graphs of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the training set, CRA VS NC group.
[0056] Figure 24 . ROC curve graphs of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the training set, CRC VS NC group.
[0057] Figure 25 . ROC curve graphs of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the training set, CRC VS CRA group.
[0058] Figure 26 . ROC curves corresponding to 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, L-Acetylcarnitine) in the validation set (CRA group VS NC group).
[0059] Figure 27. ROC curves corresponding to 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine, L-Acetylcarnitine) in the validation set (CRA group VS NC group).
[0060] Figure 28 . ROC curves corresponding to 7 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine) in the validation set (CRA group VS NC group).
[0061] Figure 29 . ROC curves corresponding to 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine) in the validation set (CRA group VS NC group).
[0062] Figure 30 . ROC curves corresponding to 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, L-Acetylcarnitine) in the validation set (CRA group VS NC group).
[0063] Figure 31 . ROC curves of 6 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid) in the validation set (CRA group VS NC group).
[0064] Figure 32 . ROC curves of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the validation set, CRA VS NC group.
[0065] Figure 33ROC curves of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the validation set, CRC VS NC group.
[0066] Figure 34 ROC curves of 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the validation set, CRC VS CRA group.
[0067] Figure 35 ROC curves of 8 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine) in the training set corresponding to those in the validation set, CRA VS NC group. Detailed implementation manners
[0068] The following further elaborates on the detailed implementation manners of the present invention with reference to the accompanying drawings. It should be understood that the detailed implementation manners described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.
[0069] Example 1: Screening of metabolic molecular markers related to colorectal cancer and establishment of an evaluation model for early prediction and diagnosis of colorectal cancer
[0070] Plasma is a currently convenient and minimally invasive detection material. Through plasma metabolomics experiments and data analysis on a large-scale population cohort, a group of plasma metabolites were screened and identified in patients with colorectal cancer and adenomas as indicators for early screening of CRC. The detailed steps for screening and identifying a group of plasma metabolites in patients with colorectal cancer and adenomas. The samples selected were human plasma samples, with a total of three groups, namely the normal group (NC group), early adenoma group (CRA group), and colorectal cancer group (CRC group), and the corresponding biological replicates for each group were 77 cases, 81 cases, and 65 cases respectively. LC-MS-based metabolomics detection and analysis were performed on a total of 223 samples.
[0071] The schematic diagram of the main operation process for screening markers is shown in the appendix Figure 1 . The test process is as follows:
[0072] A Sample information
[0073] Human plasma samples, divided into three groups, namely the normal group (NC group), the early adenoma group (CRA group), and the colorectal cancer group (CRC group), with 77, 81, and 65 biological replicates in each group respectively. Based on LC-MS metabolomics detection and analysis were performed on a total of 223 samples.
[0074] Process of sample pretreatment steps for B
[0075] (1) Pipette 100 μL of the sample into an EP tube, add 300 μL of extraction solution (methanol, containing an isotope-labeled internal standard mixture), and vortex for 30 s;
[0076] (2) Sonicate for 10 min (in an ice-water bath);
[0077] (3) Let stand at -40 °C for 1 h;
[0078] (4) Centrifuge the sample at 4 °C, 12000 rpm (centrifugal force 13800 (×g), radius 8.6 cm) for 15 min;
[0079] (5) Take the supernatant and transfer it to an injection vial for on-machine detection. For all samples, take an equal amount of supernatant and mix them into a QC sample for on-machine detection.
[0080] On-machine detection parameters
[0081] The chromatographic parameters are as follows:
[0082] In this project, a Vanquish (Thermo Fisher Scientific) ultra-high performance liquid chromatograph was used. The target compounds were chromatographically separated using a Waters ACQUITY UPLC HSS T3 (2.1 mm × 100 mm, 1.8 μm) liquid chromatography column. The mobile phase A for liquid chromatography was an aqueous phase containing 5 mmol / L ammonium acetate and 5 mmol / L acetic acid, and the mobile phase B was acetonitrile. The sample tray temperature was 4 °C, and the injection volume was 2 μL.
[0083] The mass spectrometry parameters are as follows:
[0084] The Orbitrap Exploris 120 mass spectrometer can perform primary and secondary mass spectrometry data acquisition under the control of control software (Xcalibur, version: 4.4, Thermo). Detailed parameters: Sheath gas flow rate: 50 Arb, Aux gas flow rate: 15 Arb, Capillary temperature: 320 °C, Full ms resolution: 60000, MS / MS resolution: 15000, Collision energy: 10 / 30 / 60 in NCE mode, Spray Voltage: 3.8 kV (positive) or -3.4 kV (negative).
[0085] Data processing and analysis process of D mass spectrometry offline data
[0086] After the raw data is converted into the mzXML format by the ProteoWizard software, the self-written R program package (with the XCMS kernel) is used for peak identification, peak extraction, peak alignment, integration, etc. Then, it is matched with the self-built secondary mass spectrometry database of BiotreeDB (V2.1) for substance annotation, and the Cutoff value of the algorithm score is set to 0.3. 1026 metabolites are identified in all samples through this step.
[0087] Comparison of the detection data with the mass spectrometry database. Figure 2 below shows the matching results of a certain substance, and the other 1025 metabolites are identified in the same way. The upper part of the figure is the parent ion and daughter ion fragments of the measured metabolite, and the lower part is the parent ion and daughter ion fragments measured on the mass spectrometry using high-purity standards in the public database. (See attachment Figure 2 ) If the parent ion and daughter ion are completely matched with the database, it can be identified as the corresponding metabolite.
[0088] E Establishment of an evaluation model for colorectal cancer prognosis prediction and diagnosis
[0089] Based on the above 1026 metabolites, through P-value screening and correlation analysis with clinical indicators (age, gender, BMI index, etc.), 118 metabolite indicators are screened. The names of these 118 metabolites are as follows:
[0090] Dodecanoic acid, FA(16:0), FA(16:1), FA(16:4), FA(17:2), FA(17:3), FA(18:1), FA(18:2), FA(18:4), FA(19:2), FA(20:3), FA(20:4), FA(20:5), FA(21:4), FA(21:5), FA(22:2), FA(22:4), FA(22:5), Glucosyl passiflorate, Quercetin 3-(6-[4-glucosyl-p-coumaryl]glucosyl)(1->2)-rhamnoside, Oleic acid, Isopalmitic acid, Stearic acid, Myristic acid, Eicosadienoic acid, Palmitoleic acid, PI(20:4 / 20:4), SHexCer(d28:1), PC(16:0 / 9:0(CHO)), Undecanoic Acid, Pseudouridine, Palmitic acid, alpha-Linolenic acid, Dehydroepiandrosterone sulfate, Mizolastine, Adrenic acid, Capric acid, 3b, 6a-Dihydroxy-alpha-ionol 9-[apiosyl-(1->6)-glucoside], 9-OxoODE, Eicosapentaenoic acid, Methylmalonic acid, Linoleic acid, Uric Acid, ACar(9:1), Ricinoleic acid, Myzodendrone, Prostaglandin B1, Ruscogenin, 8, 11, 14-Eicosatrienoic acid, Pyruvic acid, Stearidonic acid, Lansiumamide C, Biliverdin, 13-L-Hydroperoxylinoleic acid, Dioscorine, Sphinganine, Dibutyl phthalate, Isoscoparin 2”-(6-(E)-feruloylglucoside)4'-glucoside, FAHFA(22:5 / 22:3), 10E, 12Z-Octadecadienoic acid, 2,6-Di-tert-butyl-4-methylphenol,12-Oxo-2,3-dinor-10,15-phytodienoic acid,2-(5-Methyl-2-furanyl)-3-piperidinol,Hydroxypyruvicacid,FAHFA(16:1 / 22:3),Nordihydrocapsaicin,Thromboxane B3,FAHFA(18:1 / 18:0),5,7-Dihydro-2-methylthieno[3,4-d]pyrimidine,N-Acetyl-2,6-diethylaniline,PC(8:0 / 8:0),Methyl(3x,10R)-dihydroxy-11-dodecene-6,8-diynoate 10-glucoside,FAHFA(16:1 / 18:3),PC(16:2e / 4:0),12-Hydroxydodecanoic acid,5-Methoxytryptamine,SakacinP,Lauroyl diethanolamide,PI(18:1 / 18:1),Sphingosine 1-phosphate,L-Acetylcarnitine,Dulxanthone E,Citric acid,Paradol,Ganoderiol B,MG(18:2(9Z,12Z) / 0:0 / 0:0),PA(8:0 / 8:0),(2R,3R)-2,3-Butanediol,Heptanoic acid,PhenylpyruvicAcid,Hoduloside VI,1,2,3,4,Tetrahydro-1,5,7-trimethylnapthalene,L-AsparticAcid,SQDG(16:4 / 16:4),AcylGlcADG(16:0 / 16:0 / 20:1),PI(16:0 / 18:2),PI(16:0 / 20:4),Pelargonic acid,2-Hydroxybutyric acid,Succinic acid semialdehyde,Resveratrol,FAHFA(16:1 / 16:0),TAG(12:2 / 16:5 / 16:5),FAHFA(20:2 / 18:0),4-(4-Methyl-3-pentenyl)-3-cyclohexene-1-carboxaldehyde,1-Hydroxy-3,6,7-trimethoxy-2-(3-methyl-2-butenyl)-8-(3-hydroxy-3-methyl-1E-butenyl)-xanthone, PI(18:0 / 22:6), HBMP(16:1 / 16:1 / 20:0), alpha-Peroxyachifolide, OxPI(16:0 / 18:1+3O), (2'E,4'Z,7'Z,8E)-Colnelenic acid, PC(16:3 / 26:4), Allixin, Ginsenoyne B, Alpha-dimorphecolicacid, 4-Trimethylammoniobutanoic acid, BL V, (4-Hydroxybenzoyl)choline.,
[0091] Example 2 obtained an early diagnosis model composed of 12 substances through the LASSO model
[0092] Finally, through the LASSO model, 12 metabolites were screened out. The metabolites with the greatest contribution to early diagnosis were screened out from the candidate metabolites (118), and the best model was composed. Finally, an early diagnosis model composed of 12 substances was obtained at the minimum value of the penalty parameter λ of 0.01464075. These 12 metabolites are dodecanoic acid, stearic acid, phosphatidylcholine (16:0 / 9:0(CHO)) (PC(16:0 / 9:0(CHO))), dehydroepiandrosterone sulfate, capric acid, uric acid, acyl carnitine (9:1) (ACar(9:1)), pyruvic acid, pseudouridine, L-acetyl carnitine, phosphatidylinositol (18:0 / 22:6) (PI(18:0 / 22:6)), allixin.
[0093] A multivariate Cox regression model was constructed with these 12 metabolites to obtain the regression coefficient of each substance. Using the regression coefficient and the content of the metabolite, the risk score of each patient was calculated. The specific formula is:
[0094] Y = -0.052 + 0.15 * ACar(9:1) - 0.21 * Stearic acid - 0.049 * Allixin - 0.095 * Capric
[0095] acid - 0.10*Dehydroepiandrosterone sulfate - 0.21*Dodecanoic acid + - 0.19*Pyruvic
[0096] acid + 0.64*Pseudouridine - 0.21*L - Acetylcarnitine + 0.29*PI(18:0 / 22:6) - 0.50*PC(16:0 / 9:0(CHO)) + 0.14*Uric Acid。
[0097] Based on 223 samples in Example 1, through data analysis, the metabolite differences among the CRC group, NC group, and CRA group were compared, and 12 metabolites were selected as prediction indicators. As shown in the figure below, the AUC values of the ROC curves are all greater than 0.8, indicating that the model constructed by the prediction indicators has a high degree of credibility. The ROC curve diagrams of the 12 metabolites are shown in the appendix Figures 3 - 5 。
[0098] Identification of Standard Substances of Metabolites in Example 3
[0099] Based on the obtained 12 metabolic markers, by verifying the original mass spectrometry data, fully considering their response signal conditions in the samples, the levels of background content, and the difficulty of obtaining standard substances. Further, 8 different metabolite combinations were selected. These 8 metabolites are Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine, and L - Acetylcarnitine. Further, the data of these 8 metabolites in the test samples were used to obtain the ROC curve diagram as Figures 6 - 8 shown
[0100] After the preliminary identification of the matching database, the standard substances of the metabolites for the finally established model were identified. That is, high - purity chemicals were purchased, and detection was carried out using the same methods and instruments as those for collecting data, and the LC - MS data of human blood (including retention time, parent ion, and daughter ion mass - to - charge ratio) were compared to achieve accurate identification. The identification spectra are as follows( Figures 9 - 16), where the left side is the standard spectrum (the upper left is the chromatographic peak, representing the retention time; the lower left is the mass spectrometry peak, representing the parent ion and daughter ion fragments), and the right side is the human blood spectrum (the upper right is the chromatographic peak, representing the retention time; the lower right is the mass spectrometry peak, representing the parent ion and daughter ion fragments). After comparison, it is confirmed that 8 metabolites are completely matched.
[0101] Evaluation of different combinations of metabolites as predictors of disease risk in Example 4
[0102] Based on the identified 8 metabolite markers, 7 or 6 of them are selected to build models again. The AUC values of most combinations can still reach above 0.8, indicating that these combinations still have high value for risk assessment of CRC.
[0103] Among them, the measured AUC values of the 7-metabolite combinations for model building are shown in Table 1 below.
[0104] Table 1: Measured AUC of 7-metabolite combination model building in the training set (CRA VS NC group)
[0105]
[0106]
[0107] The ROC curves corresponding to the 7 combinations in Table 1 are shown in Figures 17 - 19 .
[0108] In addition, the measured AUC values of the 6-metabolite combinations for model building are shown in Table 2 below.
[0109] Table 2: Measured AUC of 6-metabolite combination model building in the training set (CRA VS NC group)
[0110]
[0111] The ROC curves corresponding to the 6 combinations in Table 2 are shown in Figures 20 - 22 .
[0112] Further select the 4-substance combination (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) included in both the above 7 groups and 6 groups to build a model and perform ROC curve analysis. The results are shown in Figures 23 - 25 .
[0113] Disease risk assessment of the metabolite combinations obtained from the training set as prediction indicators in the validation set in Example 5
[0114] To verify the metabolite combinations obtained in the training set, human blood samples were recollected and grouped into a normal group (NC group), an early adenoma group (CRA group), and a colorectal cancer group (CRC group), with 52, 54, and 44 biological replicates for each group, respectively, totaling 150 samples.
[0115] Based on the method of Example 1, LC-MS-based metabolomics detection and analysis were performed. ROC verification was carried out using the metabolite combinations obtained in the training set, and the results were consistent with those in the training set, showing a high indicative effect. Among them, the ROC curves corresponding to three combinations of the seven-component combinations in the training set (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, L-Acetylcarnitine), (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine, L-Acetylcarnitine), (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine) in the validation set are shown in Figures 26 - 28 ; The ROC curves corresponding to three combinations of the six-component combinations in the training set (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine), (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, L-Acetylcarnitine), (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid) in the validation set are shown in Figures 29 - 31, the corresponding ROC curves of the 4 compositions (Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid) in the training set in the validation set are shown in Figures 32 - 34 ; the corresponding ROC curves of the 8 compositions (Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine) in the training set in the validation set for the CRA VS NC group are shown in Figure 35 ;.
[0116] The present invention illustrates the process method of the present invention through the above embodiments, but the present invention is not limited to the above process steps, that is, it does not mean that the present invention must rely on the above process steps to be implemented. Those skilled in the art should understand that any improvement of the present invention, the equivalent replacement of the raw materials selected by the present invention, the addition of auxiliary components, the selection of specific methods, etc., all fall within the protection scope and the disclosure scope of the present invention.
Claims
1. A biomarker composition for colorectal cancer detection, characterized in that, The biomarker composition includes: Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid.
2. The biomarker composition according to claim 1, characterized in that, The biomarker composition is further selected from one or more of Stearic acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine.
3. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, L-Acetylcarnitine.
4. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine.
5. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid.
6. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pseudouridine, L-Acetylcarnitine.
7. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, L-Acetylcarnitine.
8. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Stearic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Pyruvic acid, Pseudouridine.
9. The biomarker composition according to claim 2, characterized in that, The biomarker composition includes: Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Stearic acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine.
10. The biomarker composition according to claim 9, characterized in that, The said marker composition consists of Dodecanoic acid, Dehydroepiandrosterone sulfate, Capric acid, Uric Acid, Stearic acid, Pyruvic acid, Pseudouridine, L-Acetylcarnitine.
11. Use of the biomarker composition according to any one of claims 1-10 in the preparation of a reagent for detecting colorectal cancer.