A biomarker for colorectal cancer detection and its application

By screening 26 biomarkers in urine through metabolomics and combining statistical methods and logistic regression equations, a colorectal cancer diagnostic model was constructed, which solved the problem of non-invasive and convenient prediction of colorectal cancer risk in existing technologies and achieved efficient and accurate colorectal cancer risk prediction.

CN115436633BActive Publication Date: 2025-10-03CALIBRA SCIENTIFIC INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211073289.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-10-03
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

Existing technologies lack non-invasive, convenient and efficient methods to predict individual colorectal cancer risk, especially the use of metabolomics technology to screen biomarkers suitable for urine samples for early warning of colorectal cancer.

Method used

The urine of colorectal cancer patients and normal subjects was analyzed using metabolomics methods to screen out metabolites with significant differences. UPLC-MS/MS technology was used to detect 26 biomarkers in urine. A colorectal cancer diagnostic model was constructed by combining random forest, PLS-DA, difference test and SVM statistical methods, and a logistic regression equation was used for prediction.

Benefits of technology

It has achieved non-invasive, convenient and efficient prediction of colorectal cancer risk through urine samples, improved the accuracy of diagnosis, and the AUC value reached 0.957, which has greater advantages and prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115436633B_ABST
    Figure CN115436633B_ABST
Patent Text Reader

Abstract

The present invention provides a biomarker for colorectal cancer detection and its application. By using the metabolomics method, by analyzing the metabolites with significant differences in the urine of colorectal cancer patients and normal people, a series of biomarkers that can early predict the risk of colorectal cancer are screened out, and a group of biomarkers are further screened out from them to construct a diagnostic model for colorectal cancer. This model can be used to conveniently, non-invasively and efficiently predict whether an individual has colorectal cancer, meeting clinical needs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application with application number 202210661330.X, application date June 10, 2022, and invention name “A biomarker for colorectal cancer detection and its application”. Technical Field

[0002] The present invention relates to the field of medicine, and in particular, to the use of metabolomics to screen biomarkers of colorectal cancer and use them for the diagnosis of colorectal cancer, and more particularly to a biomarker for predicting the risk of colorectal cancer by detecting urine samples. Background Art

[0003] Metabolomics is the qualitative and quantitative analysis of small molecule metabolites with a relative molecular weight of less than 1000 in the body. Metabolomics analysis can reveal physiological and pathological conditions and distinguish differences between individuals. With the development of mass spectrometry, liquid chromatography coupled to mass spectrometry (LC-MS) has become the most important research tool in metabolomics research. Currently, metabolomics has been widely applied in clinical diagnosis, primarily to discover metabolic markers relevant to disease diagnosis and treatment.

[0004] Colorectal cancer (CRC) is one of the most common malignant tumors worldwide and in my country. The incidence of CRC is increasing annually in nearly all cancer registries nationwide. Despite progress in the prevention and treatment of CRC through long-term basic research and clinical practice, the overall five-year survival rate remains low, primarily due to the lack of effective biomarkers that can predict the risk of CRC development. Therefore, the key to improving overall survival rates for CRC lies in early detection and treatment.

[0005] At present, the diagnosis of colorectal cancer is still mainly based on colonoscopy and imaging. In the process of research and discovery of cancer biomarkers, various omics technologies based on systems biology also play an important role. Biomarkers discovered based on the results of genomic and proteomics research have been used in cancer research. For example, the genetic diagnostic in vitro diagnostic kit for the detection of KRAS gene mutations and BMP3 / NDRG4 gene methylation in colorectal cancer "KRAS gene mutation and BMP3 / NDRG4 gene methylation and fecal occult blood combined detection kit (PCR fluorescent probe method-colloidal gold method)" was approved for marketing by the National Medical Products Administration on November 9, 2020, and is used for screening high-risk populations for colorectal cancer with poor compliance with colonoscopy.

[0006] In recent years, a wealth of metabolomics research has yielded results that have been increasingly published in various academic journals. In 2014, Cross et al. conducted a serum metabolomics study on 254 patients with colorectal cancer and 254 matched disease-free controls. While they were unable to identify specific serum metabolites directly associated with colorectal cancer risk among the 447 identified serum metabolites, an interesting finding was that the bile acid glycochenodeoxycholate (GCH) was significantly positively correlated with colorectal cancer risk in women. In another metabolomics study focusing on colorectal cancer, Long et al. conducted an untargeted metabolomics study on the serum of 30 CRC patients and 30 healthy controls. These few studies focusing on the early detection and early warning of CRC have theoretically demonstrated the feasibility of discovering CRC-related metabolic biomarkers using metabolomics techniques. However, the sample types required for the metabolic biomarkers for colorectal cancer that have been reported so far are all blood samples, while genetic testing for colorectal cancer risk requires fecal samples, neither of which has advantages in terms of non-invasiveness and simplicity of sample collection.

[0007] Therefore, there is an urgent need to find a biomarker that can be conveniently and quickly sampled non-invasively and can predict early on whether an individual has a risk of colorectal cancer, so as to achieve a more efficient assessment of colorectal cancer risk. Summary of the Invention

[0008] In response to the problems existing in the prior art, the present invention provides a biomarker for colorectal cancer detection. By using the metabolomics method, by analyzing the metabolites with significant differences in the urine of colorectal cancer patients and normal people, a series of biomarkers that can early predict the risk of colorectal cancer (CRC) are screened out, and a group of biomarkers are further screened out to construct a diagnostic model for colorectal cancer. This model can be used to conveniently, non-invasively and efficiently predict whether an individual has colorectal cancer, meeting clinical needs.

[0009] In one aspect, the present invention provides a use of a biomarker in the preparation of an agent for predicting whether an individual has colorectal cancer, wherein the biomarker is selected from one or more of the following: 2-piperidone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvic acid, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidine, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, and tris(hydroxymethyl)aminomethane acetate.

[0010] The present invention conducts non-targeted metabolomics research, using UPLC-MS / MS ultra-high performance liquid chromatography-tandem mass spectrometry to analyze urine samples from a healthy group and a colorectal cancer patient group. Then, four statistical methods, including random forest, PLS-DA, difference test, and SVM, are used to screen for metabolites that are significantly different between the colorectal cancer samples and the control samples. The significantly different metabolites that are screened out by all four statistical analysis methods are selected, ultimately resulting in 26 urine metabolites that serve as biomarkers and can be used to efficiently predict whether an individual has colorectal cancer.

[0011] In some embodiments, the biomarkers that can be used to predict whether an individual has colorectal cancer can be used to prepare detection reagents with the biomarkers as detection targets, such as sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for the detection of the biomarkers; they can also be developed into standardized reagents or kits suitable for LC-UV or LC-MS detection of the biomarkers.

[0012] In some embodiments, the biomarkers of the present invention are obtained by screening urine samples, and are particularly suitable for development into urine detection reagents or kits for predicting colorectal cancer.

[0013] In some embodiments, when the selected biomarker is an amino acid or an amino acid derivative or contains an amino group, such as 4-hydroxyphenylacetylglutamine, N-acetylpentanediamine, N-formylmethionine, N-methyl-4-aminobutyric acid, phenylacetylalanine, phenylacetylglutamate, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, and phenylacetylthreonine, an amino acid analysis method such as PITC method, AQC method, OPA method, or FMOC method can be combined to prepare reagents or kits for detecting these biomarkers suitable for use with an amino acid analyzer or LC-UV.

[0014] Further, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, phenylacetylthreonine, 3-hydroxyaminobenzoic acid, 5-hydroxyindole glucoside, phenylacetylglutamate, phenylacetylhistidine, 2-piperidone, N-formylmethionine, phenylacetylaminoethanesulfonic acid, 3-hydroxyindole sulfate, 6-hydroxyindole sulfate, and trimethylamine-N-oxide.

[0015] By examining the differences in the concentrations of biomarkers in the urine of colorectal cancer patients and normal controls, and sorting them according to the multiples of the differences, we further selected 20 biomarkers with the largest multiples of change between colorectal cancer patients and normal controls from the 26 biomarkers (theoretically, these compounds with large multiples of change would be the most effective markers), which can be used to more effectively differentiate or predict the risk of colorectal cancer, or to construct a diagnostic model for colorectal cancer.

[0016] Furthermore, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.

[0017] By examining the differences in the concentrations of biomarkers in the urine of colorectal cancer patients and normal controls, and sorting them according to the multiples of the differences, we further selected the 10 biomarkers with the largest multiples of change between colorectal cancer patients and normal controls from the 26 biomarkers (theoretically, these compounds with large multiples of change may be the most effective markers), which can be used to more effectively distinguish or predict the risk of colorectal cancer, or to build a diagnostic model for colorectal cancer.

[0018] Furthermore, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvate, N-methyl-4-aminobutyric acid, p-cresol sulfate, phenylacetylmethionine, and phenylacetylthreonine.

[0019] By examining the differences in the concentrations of biomarkers in the urine of colorectal cancer patients and normal controls, and sorting them according to the multiples of the differences, we further selected the five biomarkers with the largest multiples of change between colorectal cancer patients and normal controls from the 26 biomarkers (theoretically, these compounds with large multiples of change may be the most effective markers), which can be used to more effectively distinguish or predict the risk of colorectal cancer, or to build a diagnostic model for colorectal cancer.

[0020] Furthermore, the biomarker is selected from one or more of the following: p-cresol sulfate, phenylacetylthreonine.

[0021] By examining the differences in the concentrations of biomarkers in the urine of colorectal cancer patients and normal controls, and sorting them according to the multiples of the differences, we further selected two biomarkers with the largest multiples of change between colorectal cancer patients and normal controls from the 26 biomarkers (theoretically, these compounds with large multiples of change may be the most effective markers), which can be used to more effectively differentiate or predict the risk of colorectal cancer, or to construct a diagnostic model for colorectal cancer.

[0022] Furthermore, the reagent is used to detect biomarkers in urine.

[0023] The present invention screens urine for biomarkers of colorectal cancer. These biomarkers show significant differences in the urine of patients with colon cancer and those without colon cancer. By collecting urine samples, these biomarkers can be detected in an individual's urine to predict or assist in diagnosing whether the individual has colorectal cancer or the likelihood of developing colorectal cancer. Alternatively, these biomarkers can be detected in the urine of a population, thereby classifying the population into a colorectal cancer group or a non-colorectal cancer group. Compared to blood and feces, urine collection is noninvasive and simple, offering greater advantages and prospects for the use of urine biomarkers in the preparation of diagnostic reagents for colorectal cancer or in the diagnosis of colorectal cancer.

[0024] Furthermore, the detecting of biomarkers in urine is detecting the presence or relative abundance or concentration of biomarkers in an individual's urine sample.

[0025] In some approaches, relative abundance is preferably used as the expression, where the relative abundance is the peak area of ​​the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of ​​a biomarker measured in a control sample (an individual without colon cancer) is 500 and the average peak area measured in a colorectal cancer sample is 3000, then the abundance of the biomarker in the colorectal cancer sample is considered to be 6 times that in the control sample.

[0026] In another aspect, the present invention provides a kit or chip for predicting whether an individual has colorectal cancer, wherein the kit or chip comprises a detection reagent for the biomarker as described above.

[0027] Furthermore, the reagent is used to detect biomarkers in urine.

[0028] On the other hand, the present invention provides a biomarker combination for predicting whether an individual has colorectal cancer, wherein the biomarker combination includes the following biomarkers: 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.

[0029] Furthermore, the biomarker combination includes the following biomarkers: 2-piperidone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvate, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidine, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, and tris(hydroxymethyl)aminomethane acetate.

[0030] In another aspect, the present invention provides a system for predicting whether an individual has colorectal cancer, the system comprising a data analysis module; the data analysis module is configured to analyze the detection values ​​of biomarkers, wherein the biomarkers are one or more selected from the group consisting of: 2-piperidone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvic acid, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidine, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, and tris(hydroxymethyl)aminomethane acetate.

[0031] Further, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, phenylacetylthreonine, 3-hydroxyaminobenzoic acid, 5-hydroxyindole glucoside, phenylacetylglutamate, phenylacetylhistidine, 2-piperidone, N-formylmethionine, phenylacetylaminoethanesulfonic acid, 3-hydroxyindole sulfate, 6-hydroxyindole sulfate, and trimethylamine-N-oxide.

[0032] Furthermore, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.

[0033] Furthermore, the detection value of the biomarker is the detection value of the biomarker in urine.

[0034] Furthermore, the detection value of the biomarker is to detect the presence or relative abundance or concentration of the biomarker in the urine sample of the individual.

[0035] Furthermore, the data analysis module uses random forest or logistic regression equation to construct a model for analysis.

[0036] Furthermore, the data analysis module calculates a prediction value for predicting whether an individual has colorectal cancer by substituting the detection value of the biomarker into a logistic regression equation, thereby evaluating whether the individual has colorectal cancer.

[0037] Furthermore, the logistic regression equation is:

[0038] z=4-hydroxyphenylpyruvate*0.037986+dimethylguanidine*0.4818-N-methyl-4-aminobutyric acid*1.0077-nicotinamide*1.525-p-cresol glucuronate*0.0353-p-cresol sulfate*0.021798-phenylacetylalanine*0.1902+phenylacetylglutamine*0.858-phenylacetylmethionine*0.118805+phenylacetylthreonine*0.59727+0.7486;

[0039]

[0040] Where e is the base of the natural logarithm; p represents the predicted value of whether an individual has colorectal cancer.

[0041] e is the base of natural logarithms, an infinite non-repeating decimal with a value of 2.71828..., and is defined as follows: when n->∞, (1+1 / n) n The limit

[0042] The name of the biomarker represents the relative abundance of the corresponding biomarker in the urine sample, that is, the peak area of ​​the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry.

[0043] Furthermore, when P is greater than 0.5, it is predicted that the individual has a high possibility of colorectal cancer; when p is less than 0.5, it is predicted that the individual has a low possibility of colorectal cancer.

[0044] In another aspect, the present invention provides use of the system as described above for constructing a detection model for predicting the probability value of whether an individual has colorectal cancer.

[0045] The beneficial effects of the present invention are:

[0046] 1. Screened 26 new biomarkers that can early predict the risk of colorectal cancer (CRC);

[0047] 2. We screened 2, 3, 5, 10, 20, and 26 biomarkers to construct a random forest diagnostic model for colorectal cancer. We found that the model constructed using 10 biomarkers was the best.

[0048] 3. Comparison of a random forest model and a logistic regression model constructed using 10 biomarkers revealed that the logistic regression model can further improve detection accuracy and can be used to more efficiently predict whether an individual has colorectal cancer, with an AUC value of 0.957.

[0049] 4. Testing only requires collecting samples through urine, which is non-invasive and more convenient. Compared with testing through serum or fecal samples, it has greater advantages and prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flow chart for screening biomarkers in urine by metabolomics in Example 1;

[0051] Figure 2 is the structural formula of 3-hydroxyindole sulfate in Example 1;

[0052] Figure 3 is the structural formula of 4-hydroxyphenylacetylglutamine in Example 1;

[0053] Figure 4 is the structural formula of 5-hydroxyindole glucoside in Example 1;

[0054] Figure 5 is the structural formula of phenylacetylglutamate in Example 1;

[0055] Figure 6 is the structural formula of phenylacetylhistidine in Example 1;

[0056] Figure 7 is the structural formula of phenylacetylmethionine in Example 1;

[0057] Figure 8 is the structural formula of phenylacetylthreonine in Example 1;

[0058] Figure 9Schematic diagram showing the comparison of the prediction accuracy of constructing a colorectal cancer diagnosis model by selecting 2, 3, 5, 10, 20, and 26 biomarkers from the 26 biomarkers in Example 2;

[0059] Figure 10 : The ROC curve of the random forest model for predicting colorectal cancer constructed in Example 2;

[0060] Figure 11 This is the analysis graph of the random forest model for predicting colorectal cancer in Example 2;

[0061] Figure 12 : The ROC curve of the logistic regression model for predicting colorectal cancer constructed in Example 2;

[0062] Figure 13 This is the analysis graph of the logistic regression model for predicting colorectal cancer in Example 2;

[0063] Figure 14 This is the accuracy evaluation result of the model for predicting whether colorectal cancer is present in Example 3. DETAILED DESCRIPTION

[0064] Below in conjunction with accompanying drawing and embodiment, the present invention is described in further detail, it should be noted that the embodiment described below is intended to facilitate understanding of the present invention, and does not play any limiting role thereto.The reagents used in this embodiment are all known products, obtained by purchasing commercially available products.

[0065] Example 1 Screening for Biomarkers of Colorectal Cancer in Urine Using Metabolomics

[0066] This example first uses non-targeted metabolomics to analyze urine samples from a healthy group and a colorectal cancer patient group using UPLC-MS / MS ultra-high performance liquid chromatography-tandem mass spectrometry. Next, four statistical methods, random forest, PLS-DA, volcano, and support vector machine (SVM), were used to screen for metabolites that were significantly different between colorectal cancer samples and control samples. Metabolites that were significantly different across all four statistical analysis methods were selected, ultimately yielding 26 urine metabolites as biomarkers. The role of these biomarkers in the diagnosis or differentiation of colorectal cancer was then validated (see flowchart for details). Figure 1 ).

[0067] The specific steps are as follows:

[0068] 1. Experimental methods

[0069] ① Sample collection

[0070] Urine samples were collected from 50 patients with colorectal cancer and 50 control subjects (non-colorectal cancer subjects), wherein the colorectal cancer patients were individuals confirmed to have colorectal cancer by colonoscopy.

[0071] ②Sample processing

[0072] Methanol was added to urine samples at a ratio of 1:4, and the mixture was shaken for 3 minutes to mix thoroughly. The samples were then centrifuged at 4000 × g for 10 minutes at 20°C. Four 100 μL aliquots of supernatant were transferred to four sample plates, dried under nitrogen, and reconstituted with the reconstitution solution for subsequent LC-MS / MS analysis.

[0073] ③LC-MS / MS detection and data processing

[0074] The m / z ions were extracted from the raw mass spectrometry data obtained by LC-MS / MS. Metabolites were identified by searching the database. The peak areas of the metabolite chromatographic peaks were checked and normalized, and missing values ​​were filled. The resulting data matrix was then used for subsequent bioinformatics analysis, including four statistical methods: random forest, PLS-DA (partial least squares method), volcano plot, and SVM (support vector machine). A ranked list of differential metabolites that were most effective in grouping colorectal cancer samples and control samples was selected. Finally, metabolites that were identified by all four methods were selected as colorectal cancer biomarkers.

[0075] 2. Experimental results

[0076] Four statistical methods, random forest, PLS-DA, difference test and SVM, screened out 32, 41, 35 and 52 differential metabolites, respectively. Among them, 26 metabolites were screened out by all four data analysis methods, namely 26 biomarkers, as shown in Table 1.

[0077] Table 1. 25 biomarkers for colorectal cancer

[0078]

[0079]

[0080] Example 2: Colorectal cancer prediction model

[0081] This example uses a single biomarker or a combination of multiple biomarkers screened in Example 1 to establish prediction or diagnostic models for colorectal cancer. These models are used to distinguish between colorectal cancer and non-colorectal cancer, or to screen colorectal cancer patients from a population, or to predict whether an individual is a colorectal cancer patient or the likelihood of an individual developing colorectal cancer. Specific models are as follows.

[0082] 1. Single biomarker

[0083] Data were processed using R language software. Urine samples from patients with and without colorectal cancer were divided into two groups. Changes in the concentrations of 26 biomarkers were determined. All test results were subjected to LASSO regression analysis to establish a mathematical model for predicting colorectal cancer. Calibration curves and receiver operating characteristic (ROC) curves were used to evaluate the performance of the regression model.

[0084] The analysis results showed that the 26 biomarkers were significantly correlated with the risk of colorectal cancer. The analysis results are shown in Tables 2 and 3.

[0085] Table 2. Comparison of the correlation test results between 26 biomarkers and the risk of colorectal cancer

[0086]

[0087]

[0088] Table 3. ROC analysis results of single biomarkers

[0089] Serial number Biomarkers AUC value Sensitivity Specificity critical value 1 2-Piperidone 0.7156 0.925 0.68 0.72 2 3-Hydroxyaminobenzoic acid 0.7218 0.53175 0.52 0.82 3 3-Hydroxyindole sulfate 0.7096 0.8711 0.62 0.76 4 4-Hydroxyphenylacetylglutamine 0.7036 1.24985 0.74 0.62 5 4-Hydroxyphenylpyruvic acid 0.7668 0.97835 0.72 0.76 6 5-Hydroxyindole Glucoside 0.7112 0.3662 0.46 0.96 7 6-Hydroxyindole sulfate 0.6864 0.63085 0.48 0.86 8 dimethylguanidine 0.722 0.2471 0.58 0.82 9 N-Acetyl-pentanediamine 0.7796 0.582 0.54 0.9 10 N-Formylmethionine 0.6568 0.20645 0.28 0.98 11 Niacinamide 0.6324 2.27625 0.32 0.98 12 Nicotinamide-N-oxide 0.772 0.1686 0.88 0.58 13 N-methyl-4-aminobutyric acid 0.7444 1.1929 0.62 0.78 14 p-Cresol glucuronate 0.7836 0.86 0.64 0.64 15 p-Cresol Sulfate 0.7348 0.7536 0.64 0.82 16 Phenylacetylalanine 0.7428 1.6654 0.8 0.58 17 Phenylacetylglutamate 0.6988 1.0442 0.68 0.72 18 Phenylacetylglutamine 0.7876 0.5643 0.62 0.84 19 Phenylacetylhistidine 0.7478 0.96145 0.72 0.7 20 Phenylacetylmethionine 0.7768 0.73925 0.7 0.78 21 Phenylacetylserine 0.78 1.116 0.74 0.68 22 Phenylacetamidoethanesulfonic acid 0.6968 0.6231 0.5 0.84 23 Phenylacetylthreonine 0.7352 1.21925 0.72 0.7 24 Trimethylamine-N-oxide 0.6708 0.9524 0.66 0.7 25 Xanthine 0.774 0.8734 0.78 0.68 26 Tris(hydroxymethyl)aminomethane acetate 0.7354 0.72 0.86 0.72

[0090] The correlation between changes in the concentrations of the 26 biomarkers and colorectal cancer can be determined using the OR values ​​and p-values ​​in Table 2, as well as the AUC values ​​in Table 3. The OR and AUC values ​​are the most intuitive and clear. A higher OR value indicates a greater impact on the biomarker in colorectal cancer patients relative to those without the disease, and a more pronounced exposure. A higher AUC value indicates a more accurate distinction between colorectal cancer patients and those without the disease.

[0091] As can be seen from Table 2, the concentration changes of 26 biomarkers are significantly correlated with the risk of colorectal cancer, among which phenylacetylglutamine has the highest correlation, with an OR value of 2.36, followed by phenylacetylthreonine, with an OR value of 1.82.

[0092] As can be seen from Table 3, the concentration changes of any one of the 26 biomarkers used alone to distinguish between colorectal cancer and non-colorectal cancer populations can all achieve an AUC value of over 0.63, indicating high accuracy. Among them, phenylacetylglutamine has the highest AUC value of 0.7876, followed by p-cresol glucuronide with an AUC value of 0.7836.

[0093] 2. Combination of multiple biomarkers

[0094] Although a single biomarker can be used to distinguish colorectal cancer from non-colorectal cancer urine samples or to predict colorectal cancer, generally speaking, a combination of multiple biomarkers can achieve higher accuracy in differentiation or prediction.

[0095] However, a single biomarker with higher accuracy in predicting colorectal cancer may not necessarily play a greater role in the combination after being combined with one or more other biomarkers. It is also not the case that the more biomarkers there are, the higher the predictive accuracy (AUC value) of the combination. Therefore, a large number of verification experiments are still needed.

[0096] Because the AUC and OR values ​​of biomarkers tend to evaluate the relative importance of variables in statistical models and are not suitable for selecting variables to build models, this example preferably uses the 2, 3, 5, 10, 20, and 26 biomarkers with the highest fold change in concentration between colorectal cancer and non-colorectal cancer urine samples to build a diagnostic model for colorectal cancer. The 26 biomarkers are ranked from high to low by their fold change (Fold Change = the mean expression value of the disease sample divided by the mean expression value of the normal sample) in colorectal cancer and non-colorectal cancer urine samples. The results are shown in Table 4.

[0097] Table 4. Ranking of the fold difference in concentration of 26 biomarkers in urine samples of colorectal cancer and non-colorectal cancer

[0098]

[0099]

[0100] Based on the concentration differences of the 26 biomarkers in urine samples of colorectal cancer and non-colorectal cancer provided in Table 4, this example selected 2, 3, 5, 10, 20, and 26 biomarkers from the 26 biomarkers, respectively, and constructed a diagnostic model for colorectal cancer using random forest.

[0101] Among them, the two biomarkers ranked first and second in Table 4 (p-cresol sulfate and phenylacetylthreonine) were selected. In the constructed random forest model, the information gain ratio (GINI coefficient) of p-cresol sulfate was 25.31, and the mean decrease accuracy (MeanDecreaseAccuracy) was 21.17; the GINI coefficient of phenylacetylthreonine was 24.22, and the mean decrease accuracy was 16.71.

[0102] The three biomarkers are the three biomarkers ranked 1 to 3 in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 15.43, and the average reduction precision is 16.37; the GINI coefficient of phenylacetylthreonine is 15.75, and the average reduction precision is 15.04; the GINI coefficient of N-methyl-4-aminobutyric acid is 18.33, and the average reduction precision is 24.42.

[0103] The five biomarkers are ranked 1 to 5 in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 7.86, and the average reduction precision is 10.99; the GINI coefficient of phenylacetylthreonine is 6.39, and the average reduction precision is 5.58; the GINI coefficient of N-methyl-4-aminobutyric acid is 13.73, and the average reduction precision is 25.36; the GINI coefficient of 4-hydroxyphenylpyruvate is 10.43, and the average reduction precision is 45.38; and the GINI coefficient of phenylacetylmethionine is 11.05, and the average reduction precision is 18.74.

[0104] The 10 biomarkers are the ten biomarkers ranked 1 to 10 in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 3.64, and the average reduction precision is 7.56; the GINI coefficient of phenylacetylthreonine is 2.46, and the average reduction precision is 4.80; the GINI coefficient of N-methyl-4-aminobutyric acid is 8.04, and the average reduction precision is 18.60; the GINI coefficient of 4-hydroxyphenylpyruvic acid is 6.25, and the average reduction precision is 12.60; the GINI coefficient of phenylacetylmethionine is 1. The GINI coefficient was 6.26, with an average decrease precision of 12.85; the GINI coefficient of p-cresol glucuronate was 5.20, with an average decrease precision of 11.07; the GINI coefficient of nicotinamide was 6.56, with an average decrease precision of 12.51; the GINI coefficient of phenylacetylalanine was 3.18, with an average decrease precision of 6.30; the GINI coefficient of phenylacetylglutamine was 4.47, with an average decrease precision of 6.83; and the GINI coefficient of dimethylguanidine was 3.43, with an average decrease precision of 9.16.

[0105] The 20 biomarkers are the 20 biomarkers ranked 1 to 20 in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 2.36, with an average reduction accuracy of 6.21; the GINI coefficient of phenylacetylthreonine is 1.73, with an average reduction accuracy of 4.02; the GINI coefficient of N-methyl-4-aminobutyric acid is 5.92, with an average reduction accuracy of 16.23; the GINI coefficient of 4-hydroxyphenylpyruvic acid is 4.10, with an average reduction accuracy of The average reduction accuracy is 9.28; the GINI coefficient of phenylacetylmethionine is 3.79, with an average reduction accuracy of 10.13; the GINI coefficient of p-cresol glucuronate is 3.77, with an average reduction accuracy of 9.49; the GINI coefficient of nicotinamide is 4.67, with an average reduction accuracy of 11.61; the GINI coefficient of phenylacetylalanine is 2.26, with an average reduction accuracy of 5.84; the GINI coefficient of phenylacetylglutamine is 2.67, with an average reduction accuracy of 7.7 1; the GINI coefficient of dimethylguanidine is 2.00, and the average reduction accuracy is 7.77; the GINI coefficient of 3-hydroxyaminobenzoic acid is 2.03, and the average reduction accuracy is 4.32; the GINI coefficient of 5-hydroxyindole glucoside is 2.69, and the average reduction accuracy is 5.66; the GINI coefficient of phenylacetylglutamate is 1.59, and the average reduction accuracy is 4.38; the GINI coefficient of phenylacetylhistidine is 1.62, and the average reduction accuracy is 4.96; 2-piperidin The GINI index of ketones is 1.57, with an average decrease precision of 1.85; the GINI index of N-formylmethionine is 1.45, with an average decrease precision of 2.81; the GINI index of phenylacetamidoethanesulfonic acid is 1.28, with an average decrease precision of 0.79; the GINI index of 3-hydroxyindole sulfate is 1.41, with an average decrease precision of 3.51; the GINI index of 6-hydroxyindole sulfate is 1.57, with an average decrease precision of 1.93; and the GINI index of trimethylamine-N-oxide is 1.02, with an average decrease precision of 2.61.

[0106] The 26 biomarkers are the 26 biomarkers ranked 1 to 26 in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 1.69, and the average reduction precision is 7.04; the GINI coefficient of phenylacetylthreonine is 1.04, and the average reduction precision is 2.80; the GINI coefficient of N-methyl-4-aminobutyric acid is 3.57, and the average reduction precision is 12.93; the GINI coefficient of 4-hydroxyphenylpyruvic acid is 2.45, and the average reduction precision is 5.50; the GINI coefficient of phenylacetylmethionine is 2.68, and the average reduction precision is 7.68; the GINI coefficient of p-cresol glucuronate is 2.61, and the average reduction precision is 8.31; the GINI coefficient of nicotinamide is 2.56, and the average reduction precision is 8.02; the GINI of phenylacetylalanine is 2.84, and the average reduction precision is 13. The GINI coefficient of phenylacetylglutamine was 1.83, with an average decrease in precision of 5.74; the GINI coefficient of dimethylguanidine was 1.34, with an average decrease in precision of 3.76; the GINI coefficient of 3-hydroxyaminobenzoic acid was 1.14, with an average decrease in precision of 4.11; the GINI coefficient of 5-hydroxyindole glucoside was 1.76, with an average decrease in precision of 4.39; the GINI coefficient of phenylacetylglutamic acid was 0.88, with an average decrease in precision of 3.11; the GINI coefficient of phenylacetylhistidine was 1.00, with an average decrease in precision of 4.79; the GINI coefficient of 2-piperidone was 1.20, with an average decrease in precision of 1.80; the GINI coefficient of N-formylmethionine was 0.79, with an average decrease in precision of 2.15; the GINI of phenylacetylaminoethanesulfonic acid was 0.89, with an average decrease in precision of 3.11; The GINI coefficient of the three most common steroids was 0.58, with an average decrease precision of 2.70. The GINI coefficient of 3-hydroxyindole sulfate was 0.96, with an average decrease precision of 3.64. The GINI coefficient of 6-hydroxyindole sulfate was 0.73, with an average decrease precision of 2.70. The GINI coefficient of trimethylamine-N-oxide was 0.74, with an average decrease precision of 2.33. The GINI coefficient of 4-hydroxyphenylacetylglutamine was 0.83, with an average decrease precision of 4.61. The GINI coefficient of N-acetylpentanediamine was 2.22, with an average decrease precision of 7.72. The GINI coefficient of tris(hydroxymethyl)aminomethane acetate was 2.48, with an average decrease precision of 8.06. The GINI coefficient of xanthine was 2.70, with an average decrease precision of 8.67. The GINI coefficient of nicotinamide-N-oxide was 8.21, with an average decrease precision of 16.94. The GINI coefficient of phenylacetylserine was 2.01, with an average decrease precision of 7.16.

[0107] The AUC values ​​and 95% CL confidence spaces of the six random forest diagnostic models constructed with 2, 3, 5, 10, 20, and 26 biomarkers were calculated respectively. The results are shown in Figure 2. Figure 9 shown.

[0108] Depend on Figure 9 As can be seen, when the top two biomarkers from the 26 biomarkers were selected to construct a model, the AUC value only reached 0.922, with a 95% confidence interval of 0.718-0.999. As the number of selected biomarkers increased, the AUC value gradually increased, and the 95% confidence interval gradually narrowed. When 10 biomarkers were selected to construct a colorectal cancer diagnostic model, the AUC value reached 0.935, with a 95% confidence interval of 0.842-0.998. However, when the number of biomarker types further increased to 20 or 26, the room for further AUC increase was very limited, and the confidence interval became larger. In addition, compared with 20 or 26 biomarkers, using 10 biomarkers to construct a model can reduce the number of variables and reduce the complexity of the model. Therefore, it is preferred to use the top 10 biomarkers in Table 4 to construct a colorectal cancer diagnostic model, which not only achieves very good predictive accuracy but also has a simpler and more convenient model.

[0109] A total of 42 clinically known colorectal cancer patients and 42 non-colorectal cancer patients were used as the total data set to detect the biomarker detection values ​​of their urine samples. The random forest model of 10 biomarkers was used for analysis. The analysis graph is as follows: Figure 11 As shown by Figure 11 It can be seen that when the random forest model constructed using 10 biomarkers is used to predict colorectal cancer, there will be certain errors (of course, errors are inevitable). Among 42 colorectal cancer patients, 37 were detected, and among 42 non-colorectal cancer patients, 5 were classified as colorectal cancer patients, with an accuracy rate of 88%. Figure 11 It can be seen that when the predicted value P is greater than 0.5, the probability of predicting that the individual has colorectal cancer is high; when the predicted value p is less than 0.5, the probability of predicting that the individual has colorectal cancer is low.

[0110] Using the top 10 biomarkers ranked by Fold Change, we conducted a multivariate regression analysis and established a logistic regression model to predict whether an individual has colorectal cancer:

[0111] z=4-hydroxyphenylpyruvate*0.037986+dimethylguanidine*0.4818-N-methyl-4-aminobutyric acid*1.0077-nicotinamide*1.525-p-cresol glucuronate*0.0353-p-cresol sulfate*0.021798-phenylacetylalanine*0.1902+phenylacetylglutamine*0.858-phenylacetylmethionine*0.118805+phenylacetylthreonine*0.59727+0.7486;

[0112]

[0113] Where e is the base of the natural logarithm; p represents the predicted value of whether an individual has colorectal cancer; and the name of the biomarker represents the relative abundance of the corresponding biomarker in the urine sample, that is, the peak area of ​​the biomarker in the detection spectrum obtained by high-performance liquid chromatography-tandem mass spectrometry.

[0114] The ROC curve of the logistic regression model for predicting whether an individual has colorectal cancer provided in this embodiment is as follows: Figure 12 As shown in the figure, the AUC value reached 0.957, which was significantly improved compared with the random forest model of 10 biomarkers.

[0115] The logistic regression model for predicting whether an individual has colorectal cancer was used to analyze 50 clinically known colorectal cancer patients and 50 non-colorectal cancer patients as the total data set. The analysis results are as follows: Figure 13 As shown in Table 5,

[0116] Table 5. Analysis results of the model for predicting whether an individual has colorectal cancer

[0117]

[0118] Depend on Figure 13 As can be seen from Table 5, the logistic regression evaluation model constructed using 10 biomarkers to predict whether an individual has colorectal cancer was analyzed. Among 50 colorectal cancer patients, 45 were detected, and among 50 non-colorectal cancer patients, 5 were classified as colorectal cancer patients. The accuracy rate reached more than 90%, and the accuracy was improved.

[0119] Depend on Figure 13 It can also be seen that P of 0.5 can be used as a cutoff point for judgment. When P is greater than 0.5, the probability of predicting that the individual has colorectal cancer is high; when P is less than 0.5, the probability of predicting that the individual has colorectal cancer is low.

[0120] Example 3: Evaluation of a model for predicting colorectal cancer

[0121] This example evaluates the accuracy of the colorectal cancer prediction model constructed in Example 2 for clinical application. The 42 colorectal cancer patients and 42 non-colorectal cancer patients mentioned above are used as the total data set. 8 CRC patients and normal subjects (non-CRC patients) are randomly selected from the data set. Urine samples are taken and the relative abundance of the 10 biomarkers in the model is measured according to the sample processing method in Example 1. The prediction value P is then calculated using the model to predict whether an individual has colorectal cancer. The results are shown in FIG. Figure 14 shown.

[0122] Depend on Figure 14It can be seen that all 8 colorectal cancer patients were detected, and one of the 8 normal people was predicted to have colorectal cancer, with an accuracy rate of 93.75%.

[0123] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. Use of a reagent for detecting a biomarker for preparing a reagent for predicting whether an individual has colorectal cancer, characterized in that: The biomarkers consist of 4-hydroxyphenylpyruvate, dimethylguanidine, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronate, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine and phenylacetylthreonine.

2. The use according to claim 1, characterized in that The reagent is used for detecting biomarkers in urine.

3. The use according to claim 1, characterized in that The detecting of biomarkers in urine is detecting the presence or relative abundance or concentration of biomarkers in a urine sample of an individual.

4. A chip for predicting whether an individual has colorectal cancer, characterized in that: A detection reagent for a biomarker for use according to any one of claims 1 to 3.

5. The chip according to claim 4, wherein: The reagent is used for detecting biomarkers in urine.

Citation Information

Patent Citations

  • Mass spectrometry assay method for detection and quantitation of microbiota-related metabolites

    CN112689755A

  • Mass spectrometry assay method for detection and quantitation of microbiota related metabolites

    US20220050090A1