Marker composition for colorectal cancer detection and screening and application thereof
By using a diagnostic model constructed using liquid phase tandem mass spectrometry and a random forest algorithm in colorectal cancer screening, a marker combination was screened out, solving the problems of high invasiveness, high cost, and low sensitivity in existing technologies, and achieving low-invasive, low-cost, and highly sensitive early diagnosis.
Patent Information
- Application Number
- CN202511150034.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing colorectal cancer screening methods are highly invasive, costly, and have poor sensitivity and specificity, making them difficult to implement on a large scale and for early detection.
By collecting human plasma samples, liquid phase tandem mass spectrometry was used to quantitatively detect the concentrations of 33 metabolites, and a diagnostic model was constructed using a random forest machine learning algorithm to screen out a marker combination, including lauric acid, gluconic acid, and L-alanine, for the early diagnosis of colorectal cancer.
It achieves low-invasive, low-cost early diagnosis of colorectal cancer with high sensitivity and specificity, high acceptance among subjects, and is suitable for promotion and application.
Smart Images

Figure CN120652020A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine and relates to a marker composition for disease detection, in particular, to a marker composition for early diagnosis of colorectal cancer and its screening and application. Background Art
[0002] Many colorectal cancer patients are often diagnosed in the late stage when treatment options are limited. Early detection of colorectal adenoma (CRA) through screening can significantly reduce the incidence and impact of colorectal cancer. The 5-year survival rate of early diagnosis can reach over 90%.
[0003] Traditional screening methods for colorectal cancer mainly include colonoscopy, stool examination, and imaging examination. Among them, colonoscopy is the gold standard for colorectal cancer screening. Colonoscopy has high sensitivity and specificity, but it is invasive, has cumbersome intestinal preparation, high cost (especially painless colonoscopy) and low public acceptance. It is difficult to popularize it as a large-scale screening method. Stool examination mainly includes fecal occult blood test (FIT) and fecal DNA test. Although it is non-invasive, it has low sensitivity (only 9.8%-22% sensitivity for adenoma), poor specificity (susceptible to diet / drug interference) and other problems. The ability to detect early lesions is limited. Imaging examinations include CT simulation colonoscopy, MRI, etc. The main problems are that they require radiation exposure, intestinal cleaning preparation and high cost. They are currently only used as auxiliary diagnostic tools.
[0004] Currently, the more recognized markers include serum markers such as CEA and CA19-9, and methylation markers such as SEPTIN9. However, these markers still have problems with poor specificity and sensitivity. For example, serum markers such as CEA and CA19-9 will also increase in other cancers or benign diseases such as breast cancer and lung cancer, and are not sensitive enough to early adenomas / cancer (for example, the sensitivity of CEA to adenomas is only 11%-22%). Although SEPTIN9 has been approved for screening (sensitivity 48%-76%), its sensitivity to adenomas is still less than 20%, and the detection cost is relatively high.
[0005] In recent years, machine learning technology has rapidly developed and become one of the core research areas in artificial intelligence. Using a data-driven approach, it learns patterns from large amounts of historical data and implements predictions or decisions. It is widely used in scenarios such as finance, healthcare, autonomous driving, and recommendation systems. Commonly used algorithms and models include logistic regression, LASSO, SVM, random forests, and neural networks. These algorithms and models are widely used in a variety of application scenarios, including binary probability prediction, high-dimensional sample classification, feature screening of high-dimensional data, complex classification and regression tasks, and nonlinear big data pattern recognition. The main advantages of these algorithms and models in this paper are simplicity and efficiency, strong interpretability, strong generalization, automatic sparse modeling to prevent overfitting, support for parallel computing, and feature self-learning. Summary of the Invention
[0006] Due to the current problems that have prevented the effective and widespread application of CRA and CRC screening, the present invention is committed to resolving these difficulties. By using a small amount of plasma samples, the concentration of marker metabolites is quantitatively detected based on liquid phase tandem mass spectrometry, and the metabolite concentration is input into the diagnostic formula to assess the risk of adenoma or colorectal cancer in subjects. The main advantages are that plasma samples are easier to collect than feces, less invasive than colonoscopy, and more acceptable to subjects. The combined biomarker has a high AUC in the validation cohort, better early specificity and sensitivity, the diagnostic formula is interpretable, and the detection cost is low.
[0007] The above-mentioned purpose of the present invention is achieved through the following technical solutions: The present invention discloses a method for screening a marker composition for colorectal cancer detection, comprising the following steps: 1) Collect human plasma samples and pre-treat them by adding an extract containing isotopic internal standards corresponding to 33 metabolites; 2) Liquid chromatography and mass spectrometry were used to analyze the pre-treated samples, collect data, and perform quantitative analysis on the raw data to obtain the concentration values of 33 metabolites in the samples; 3) The samples were divided into training and validation sets, and the 33 metabolites were modeled using the Random Forest machine learning algorithm. The receiver operating characteristic (ROC) curve of the algorithm model was analyzed. 4) Using a random forest model combined with substance frequency to screen metabolite combinations, modeling the substance combinations using a random forest machine learning algorithm, and validating the metabolite combinations using the receiver operating characteristic (ROC) curve under the algorithm model; The 33 metabolites are lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, L-tryptophan, trimethylamine oxide, L-serine, 1-methylnicotinamide, glycocholic acid, glycochenodeoxycholic acid, L-acetylcarnitine, pseudouridine, stearic acid, dehydroepiandrosterone sulfate, allocholic acid, proline hydroxyproline, L-glycine, L-leucine, glycohyocholic acid sodium salt, glycodeoxycholic acid, linoleic acid, N-acetylglycine, N-acetyl-L-glutamic acid, N-acetylneuraminic acid, and arachidonic acid.
[0008] Preferably, the substance frequency is not less than 800 times.
[0009] Preferably, the substance frequency is not less than 900 times.
[0010] Preferably, the substance frequency is not less than 1000 times.
[0011] Preferably, the AUC value of the ROC curve in step 4) is greater than 0.8.
[0012] The invention discloses a marker composition for colorectal cancer detection obtained by screening the method. The marker composition comprises lauric acid, gluconic acid and L-alanine.
[0013] Preferably, the marker composition further comprises: L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid and hippuric acid.
[0014] Preferably, the marker composition further comprises: 3-indoleacetic acid, L-tyrosine and L-tryptophan.
[0015] Preferably, the marker composition is: lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, and L-tryptophan.
[0016] Preferably, the marker composition is: lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, and L-tryptophan.
[0017] The present invention discloses a kit for detecting colorectal cancer, characterized in that the kit comprises any one of the above-mentioned compositions.
[0018] The present invention discloses a kit for distinguishing colorectal cancer from early adenoma, characterized in that the kit comprises any one of the above-mentioned compositions.
[0019] The present invention discloses a kit for distinguishing early-stage adenomas, characterized in that the kit comprises any one of the above-mentioned compositions.
[0020] The present invention discloses the use of any of the compositions in preparing a reagent and / or a kit for detecting colorectal cancer.
[0021] Preferably, the use is for distinguishing normal people, early adenoma patients and colorectal cancer patients.
[0022] Preferably, the use is for distinguishing normal people from colorectal cancer patients.
[0023] Preferably, the use is for distinguishing patients with early-stage adenoma from patients with colorectal cancer.
[0024] Preferably, the use is for distinguishing normal people from patients with early adenoma.
[0025] The present invention discloses candidate substances for screening marker combinations for colorectal cancer detection. The substances include one or more combinations of lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, L-tryptophan, trimethylamine oxide, L-serine, 1-methylnicotinamide, glycocholic acid, glycochenodeoxycholic acid, L-acetylcarnitine, pseudouridine, stearic acid, dehydroepiandrosterone sulfate, allocholic acid, proline hydroxyproline, L-glycine, L-leucine, glycohyocholic acid sodium salt, glycodeoxycholic acid, linoleic acid, N-acetylglycine, N-acetyl-L-glutamic acid, N-acetylneuraminic acid, and arachidonic acid.
[0026] Preferably, the test sample is plasma.
[0027] Another aspect of the present invention provides a method for constructing a colorectal cancer risk assessment model, the method comprising the following steps: 1) Analysis of candidate substances for marker combination screening: Sample-related metabolite concentrations were obtained by LC-MS / MS testing, and the effectiveness of five machine learning algorithms in screening substance combinations was compared. Specifically, samples from the normal group, CRC group, and CRA group were randomly divided into training and validation sets at a rate of 7 / 3. Then, random sampling was performed at least 100 times to construct individual models. Finally, substances with a frequency of at least 200 times screened by each model were combined, and the ROC curves of the five machine learning algorithm models were analyzed and compared. The machine learning algorithm with the largest AUC was used for subsequent modeling.
[0028] 2) Model building: Utilize the aforementioned preferred machine learning algorithm to model the aforementioned preferred substance combination, and ultimately form the model of early diagnosis metabolites.
[0029] The present invention is based on LC-MS / MS quantitative detection of plasma metabolite concentration analysis to screen and obtain metabolites related to colorectal cancer.
[0030] The present invention uses machine algorithm modeling and screening to perform ROC curve analysis of different models, compare AUC values, and screen out machine learning algorithm models and multiple marker combinations.
[0031] The present invention provides a model for the early diagnosis of colorectal cancer. The model has been validated using real human plasma samples. The validation results show that the model has good sensitivity and specificity in diagnosing CRA and CRC patients, has a certain predictive ability for the early diagnosis of CRA patients, and can relatively accurately identify CRC patients. The detection method is highly accurate, high-throughput, low-cost, and highly accepted by subjects, making it suitable for popularization and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 : Measured ion pair information of 33 metabolites after derivatization.
[0033] Figure 2 : ROC curves of 33 candidate marker combinations modeled using 5 machine learning algorithm training sets (NC group and CRA group).
[0034] Figure 3 : ROC curves of 33 candidate marker combinations modeled using 5 machine learning algorithm training sets (NC group and CRC group).
[0035] Figure 4 : ROC curves of 33 candidate marker combinations modeled using 5 machine learning algorithm training sets (CRA group and CRC group).
[0036] Figure 5 : ROC curves of 33 candidate marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRA group).
[0037] Figure 6 : ROC curves of 33 candidate marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRC group).
[0038] Figure 7 : ROC curves of 33 candidate marker combinations modeled using validation sets of 5 machine learning algorithms (CRA group and CRC group).
[0039] Figure 8: ROC curves of 13 marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRA group).
[0040] Figure 9 : ROC curves of 13 marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRC group).
[0041] Figure 10 : ROC curves of 13 marker combinations modeled using validation sets of 5 machine learning algorithms (CRA group and CRC group).
[0042] Figure 11 : ROC curves of 10 marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRA group).
[0043] Figure 12 : ROC curves of 10 marker combinations modeled using validation sets of 5 machine learning algorithms (NC group and CRC group).
[0044] Figure 13 : ROC curves of 10 marker combinations modeled using validation sets of 5 machine learning algorithms (CRA group and CRC group).
[0045] Figure 14 : ROC curves of the three marker combinations modeled using the validation sets of five machine learning algorithms (NC group and CRA group).
[0046] Figure 15 : ROC curves of the three marker combinations modeled using validation sets of five machine learning algorithms (NC group and CRC group).
[0047] Figure 16 : ROC curves of the three marker combinations modeled using the validation sets of five machine learning algorithms (CRA group and CRC group).
[0048] Figure 17 : A diagnostic formula based on the logistic regression model that can distinguish NC from CRA.
[0049] Figure 18 : Diagnostic confusion matrix of 382 real samples. DETAILED DESCRIPTION
[0050] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.
[0051] Example 1 Screening of metabolites related to colorectal cancer and establishment of an evaluation model for early prediction and diagnosis of colorectal cancer Plasma is currently a readily available and minimally invasive testing material. Through plasma metabolomics experiments and data analysis of a large population cohort, a group of plasma metabolites was screened and identified in patients with colorectal cancer and adenoma as indicators for early CRC screening. Detailed steps for screening and identifying a group of plasma metabolites in patients with colorectal cancer and adenoma. The samples were selected from human plasma samples, with a total of three groups: normal group (NC group), early adenoma group (CRA group), and colorectal cancer group (CRC group). The biological replicates corresponding to each group were 217, 165, and 218 samples, respectively. LC-MS / MS-based quantitative detection and analysis were performed, totaling 600 samples. The samples were all human plasma samples confirmed by colonoscopy gold standard and clinician diagnosis.
[0052] The test process is as follows: A. Sample Information There were three groups of human plasma samples, namely normal group (NC group), early adenoma group (CRA group) and colorectal cancer group (CRC group). The biological replicates of each group were 217, 165 and 218, respectively. LC-MS / MS-based quantitative detection and analysis were performed, totaling 600 samples.
[0053] B. Sample pre-treatment steps (1) Pipette 30 μL of human plasma sample and linear concentration sample into a 96-well plate; (2) Add 100 μL of extraction solution (acetonitrile, containing 33 metabolites corresponding to isotopic internal standards, the 33 metabolites were screened by our laboratory based on non-target platform metabolomics in the early stage), and shake at 1500 rpm for 5 min; (3) Centrifuge at 4000 rpm for 10 min at 4°C; (4) Take 100 μL of the supernatant, load it into a 96-well protein precipitation plate, filter it on a positive pressure device, and collect the filtrate; (5) Add 10 μL each of 200 mM 3-NPH and 100 mM EDC; (6) Add 100 μL of ice-cold 50% methanol water; (7) Place in a -20°C refrigerator for 20 minutes; (8) Take 5 μL of sample for testing.
[0054] C On-machine test parameters The chromatographic parameters are as follows: In this project, the target compounds were separated chromatographically using a Jasper (AB Sciex) high-performance liquid chromatograph on a Kinetex PFP (3.0 mm × 100 mm, 2.6 μm) column. Phase A consisted of 0.1% formic acid in water and phase B consisted of 0.1% formic acid in acetonitrile. A gradient elution process was employed: 10% B (0-0.5 min), 10% to 98% B (0.5-4.5 min), 98% B (4.5-6.0 min), 98% to 10% B (6.0-6.1 min), and 10% B (6.1-8.0 min). The flow rate was 0.35 mL / min. The sample tray temperature was 4°C, and the injection volume was 5 μL.
[0055] The mass spectrometry parameters are as follows: The Citrine™ Triple Quad™ MS / MS mass spectrometer was capable of acquiring precursor and product ion mass spectrometry data under the control of Analyst software (Version 1.6.3, AB). Ion source parameters: Electrospray ionization was used in positive and negative ion scan modes, with curtain gas (CUR) at 20 psi, collision gas (CAD) at 9 psi, ion spray voltage (IS) at 4500 kV, temperature (TEM) at 500°C, ion source gas 1 (GS1) at 40 L / min, and ion source gas 2 (GS2) at 55 L / min. Data were acquired in MRM mode. The mass spectrometry parameters for each compound are shown in the attached figure. Figure 1 .
[0056] D Mass Spectrometry Off-line Data Processing and Analysis Process Raw data were quantitatively analyzed using MultiQuant MD software. The ratio of the analyte signal response in the calibration sample to the corresponding internal standard signal response was fitted to a standard curve equation, along with the labeled concentration of the analyte. The ratio of the peak area of the analyte to the peak area of the internal standard was then substituted into the fitted standard curve equation to calculate the concentration of the sample analyte. Ultimately, the absolute concentrations of 33 metabolites in 600 samples were obtained.
[0057] E. Establishment of an evaluation model for prognosis prediction and diagnosis of colorectal cancer Based on the test concentrations of 33 metabolites in 600 samples, the 600 samples were first randomly divided into a training set and a validation set at a ratio of 7 / 3.
[0058] For the 420 samples in the training set, five machine learning algorithms, namely logistic regression, LASSO, SVM, random forest, and neural network, were used to build models. Each algorithm was randomly sampled 1000 times to build models. The ROC curves of the five algorithm models for distinguishing NC from CRA, NC from CRC, and CRA from CRC were analyzed. See the attached figure. Figure 2-Figure 4 ,Comparing the AUC values, preferably, the random forest algorithm model performed best, with AUC values of 0.871, 0.896 and 0.843 for the NC and CRA groups, NC and CRC groups, and CRA and CRC groups, respectively.
[0059] For the 180 samples in the validation set, similarly, five machine learning algorithms, namely logistic regression, LASSO, SVM, random forest, and neural network, were used for modeling. Each algorithm was randomly sampled 1000 times to model the NC and CRA groups, NC and CRC groups, and CRA and CRC groups. The ROC curves of the five algorithm models were analyzed. See the attached figure. Figure 5-Figure 7 ,Comparing the AUC values, preferably, the random forest algorithm model performed best, with AUC values of 0.877, 0.919 and 0.854 in the NC and CRA groups, NC and CRC groups, and CRA and CRC groups, respectively.
[0060] Example 2: An early diagnosis model consisting of multiple combinations of 13 substances was obtained by screening using a random forest model A Random Forest model screens combinations of substances with a frequency of no less than 800 times Using a random forest model, we screened out 13 substances with a substance frequency of at least 800 (substance frequency is the number of times a substance is consistently screened out after 1,000 runs of modeling with different datasets). These 13 metabolites are lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, and L-tryptophan.
[0061] The 13 metabolite marker combinations were again modeled using 5 machine learning algorithms, and the ROC curves were analyzed. Figures 8-10 , and compared the AUC values of the NC and CRA groups, NC and CRC groups, and CRA and CRC groups. The results showed that random forest still performed best, indicating that the previous random forest modeling was appropriate.
[0062] After random forest modeling was used for the combination of the 13 metabolite markers, the AUC values of the NC and CRA groups were 0.877, the AUC values of the NC and CRC groups were 0.919, and the AUC values of the CRA and CRC groups were 0.854.
[0063] B. Random forest model screening of substance combinations with a substance frequency of not less than 900 times Using the random forest model, we identified 10 substances with a frequency of at least 900. These 10 metabolites are lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, and hippuric acid.
[0064] The 10 metabolite marker combinations were again modeled using 5 machine learning algorithms, and the ROC curves were analyzed. Figure 11-13 , and compared the AUC values of the NC and CRA groups, NC and CRC groups, and CRA and CRC groups. The results showed that random forest still performed best, indicating that the previous random forest modeling was appropriate.
[0065] After random forest modeling was used for the combination of the 10 metabolite markers, the AUC values of the NC and CRA groups were 0.837, the AUC values of the NC and CRC groups were 0.905, and the AUC values of the CRA and CRC groups were 0.848.
[0066] C Random forest model screens combinations of substances with a frequency of no less than 1000 times The random forest model screened three substances with a frequency of at least 1,000. These three metabolites are lauric acid, gluconic acid, and L-alanine.
[0067] The three metabolite marker combinations were again modeled using five machine learning algorithms, and the ROC curves were analyzed. Figure 14-16 , and compared the AUC values of the NC and CRA groups, NC and CRC groups, and CRA and CRC groups. The results showed that random forest still performed best, indicating that the previous random forest modeling was appropriate.
[0068] After random forest modeling was used for the combination of the three metabolite markers, the AUC values of the NC and CRA group were 0.832, the AUC values of the NC and CRC group were 0.811, and the AUC values of the CRA and CRC group were 0.723.
[0069] In summary, compared with the other four methods, the model established using the random forest machine learning algorithm had the largest AUC value and the best model discrimination effect, and the AUCs were all above 0.8, indicating that the model established using the random forest algorithm can effectively distinguish NC from CRA and CRC groups.
[0070] Preferably, the AUC of the 13 marker combination was the largest, reaching 0.877 and 0.919 for the NC group, CRA group, and CRC group, respectively.
[0071] Example 3: Construction of a CRA diagnostic formula using a logistic regression model based on a combination of 13 markers A. Constructing a diagnostic formula A multivariate logistic regression model was constructed using the above 13 substances to obtain the regression coefficient of each substance. The risk score of each patient was calculated using the regression coefficient and the content of metabolites. The specific formula is shown in the attached Figure 17 The calculated cut-off value between NC and CRA is 0.43, that is, when the calculated value is greater than 0.43, it is judged as CRA according to the diagnostic formula; if it is less than or equal to 0.43, it is judged as NC.
[0072] Y = 0.362 + 0.408 gluconic acid - 0.098 L-histidine + 0.719 phenylacetylglutamine - 0.416 choline - 0.596 hippuric acid + 0.105 lactic acid + 0.067 L-tyrosine + 0.072 L-tryptophan - 0.0186 L-alanine + 0.321 DL-3-hydroxybutyric acid + 0.073 indoleacetic acid + 0.378 pyruvic acid + 0.825 lauric acid (concentration unit: ng / mL) B. Prediction Accuracy Verification Based on 382 samples in total, including 217 samples from the NC group and 165 samples from the CRA group in Example 1, the concentration information was input into the diagnostic formula to obtain a confusion matrix for verification accuracy, see Appendix. Figure 18 , the accuracy of the diagnostic formula in predicting NC and CRA is close to 80%, indicating that the diagnostic formula has good diagnostic ability for NC and CRA.
[0073] While the present invention is illustrated by the above-described embodiments, the present invention is not limited to the above-described process steps, and implementation of the present invention is not necessarily dependent on the above-described process steps. Those skilled in the art will appreciate that any improvements to the present invention, equivalent substitutions for the raw materials used, additions of auxiliary components, and selection of specific methods, etc., fall within the scope of protection and disclosure of the present invention.
Claims
1. A method for screening a marker composition for colorectal cancer detection, characterized in that: The steps include: 1) Collect human plasma samples and pre-treat them by adding an extract containing isotopic internal standards corresponding to 33 metabolites; 2) Liquid chromatography and mass spectrometry were used to analyze the pre-treated samples, collect data, and perform quantitative analysis on the raw data to obtain the concentration values of 33 metabolites in the samples; 3) The samples were divided into training and validation sets, and the 33 metabolites were modeled using the Random Forest machine learning algorithm. The receiver operating characteristic (ROC) curve of the algorithm model was analyzed. 4) Using a random forest model combined with substance frequency to screen metabolite combinations, modeling the substance combinations using a random forest machine learning algorithm, and validating the metabolite combinations using the receiver operating characteristic (ROC) curve under the algorithm model; The 33 metabolites are lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, L-tryptophan, trimethylamine oxide, L-serine, 1-methylnicotinamide, glycocholic acid, glycochenodeoxycholic acid, L-acetylcarnitine, pseudouridine, stearic acid, dehydroepiandrosterone sulfate, allocholic acid, proline hydroxyproline, L-glycine, L-leucine, glycohyocholic acid sodium salt, glycodeoxycholic acid, linoleic acid, N-acetylglycine, N-acetyl-L-glutamic acid, N-acetylneuraminic acid, and arachidonic acid.
2. The screening method according to claim 1, wherein The frequency of the substance is not less than 800 times.
3. The screening method according to claim 1, wherein The AUC value of the ROC curve in step 4) is greater than 0.
8.
4. A marker composition for colorectal cancer detection obtained by screening according to any one of claims 1 to 3, characterized in that: The marker composition comprises lauric acid, gluconic acid and L-alanine.
5. The marking composition according to claim 4, characterized in that The marker composition further comprises: L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid and hippuric acid.
6. The marking composition according to claim 5, characterized in that The marker composition further comprises: 3-indoleacetic acid, L-tyrosine and L-tryptophan.
7. The marking composition according to claim 6, characterized in that The marker composition comprises: lauric acid, gluconic acid, L-alanine, L-histidine, choline, phenylacetylglutamine, L-(+) lactic acid, pyruvic acid, DL-3-hydroxybutyric acid, hippuric acid, 3-indoleacetic acid, L-tyrosine, and L-tryptophan.
8. A kit for detecting colorectal cancer, characterized in that: The kit comprises the composition according to any one of claims 4 to 7.
9. Use of the composition according to any one of claims 4 to 7 in the preparation of a reagent and / or a kit for detecting colorectal cancer.
10. The use according to claim 9, characterized in that The use is for distinguishing normal people, early adenoma patients and colorectal cancer patients.