A colorectal cancer prediction system and applications thereof
By analyzing urine samples using metabolomics, 26 biomarkers were screened, and a random forest diagnostic model was constructed. This solved the problem of non-invasive and convenient prediction of colorectal cancer in existing technologies, and achieved efficient and accurate colorectal cancer risk assessment.
Patent Information
- Application Number
- CN202211073050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-06-10
AI Technical Summary
Current technologies lack non-invasive and convenient methods for early prediction of whether an individual has colorectal cancer, and existing biomarkers mostly rely on blood or stool samples, which are inconvenient to collect.
By analyzing urine samples using metabolomics, 26 significantly different metabolites were screened as biomarkers. A random forest diagnostic model was constructed, and these biomarkers were detected using high performance liquid chromatography-tandem mass spectrometry (UPLC-MS/MS). By combining random forest and logistic regression methods, the individual's risk of colorectal cancer was predicted.
It enables non-invasive and convenient prediction of colorectal cancer risk through urine testing, improving the accuracy and efficiency of diagnosis, with an AUC value of 0.957, and simplifies the sample collection process.
Smart Images

Figure CN115440375B_ABST
Abstract
Description
[0001] This application is a divisional application of application number 202210658811.5, filed on June 10, 2022, entitled "A colorectal cancer prediction system and its application". Technical Field
[0002] This invention relates to the medical field, and more specifically, to the use of metabolomics to screen biomarkers for colorectal cancer and for the diagnosis of colorectal cancer, and particularly to a predictive system for predicting the risk of colorectal cancer by detecting urine samples and its application. Background Technology
[0003] Metabolomics is a discipline that involves the qualitative and quantitative analysis of small molecule metabolites with a relative molecular weight of less than 1000 in the body. Metabolomics analysis can reflect the physiological and pathological conditions of the body and distinguish differences between individuals. With the development of mass spectrometry, liquid chromatography-mass spectrometry (LC-MS) has become the most important research tool in metabolomics research. Currently, metabolomics is widely used in clinical diagnostics, primarily to discover metabolic biomarkers related to disease diagnosis and treatment.
[0004] Colorectal cancer (CRC) is one of the most common malignant tumors globally and in my country. While significant progress has been made in the prevention and treatment of CRC through long-term basic research and clinical practice, the overall five-year survival rate remains low. This is partly due to the lack of effective biomarkers that can predict the early risk of CRC development. Therefore, the key to improving the overall survival rate of colorectal cancer lies in early detection and early treatment.
[0005] Currently, the diagnosis of colorectal cancer mainly relies on colonoscopy and imaging. In the research and discovery of cancer biomarkers, various omics technologies based on systems biology also play an important role. Biomarkers discovered based on genomics and proteomics research results have already been applied in cancer research. For example, the in vitro diagnostic kit for detecting KRAS gene mutations and BMP3 / NDRG4 gene methylation in colorectal cancer, titled "KRAS Gene Mutation and BMP3 / NDRG4 Gene Methylation and Fecal Occult Blood Combined Detection Kit (PCR Fluorescent Probe Method-Colloidal Gold Method)," was approved for marketing by the National Medical Products Administration on November 9, 2020, for screening high-risk individuals with poor colonoscopy adherence to colonoscopy.
[0006] In recent years, a wealth of research findings from metabolomics have been increasingly published in various academic journals. In 2014, Cross et al. conducted a serum metabolomics study on 254 colorectal cancer patients and 254 matched disease-free controls. Of the 447 serum metabolites identified, no specific metabolites were found to be directly associated with rectal cancer risk. However, an interesting finding was that in the female population, the level of glycochenodeoxycholate in bile acids was significantly positively correlated with rectal cancer risk. In another metabolomics study on colorectal cancer, Long et al. first conducted a non-targeted metabolomics study on the serum of 30 CRC patients and 30 healthy controls. These few studies on the early detection and warning of CRC theoretically demonstrate the feasibility of using metabolomics technology to discover CRC-related metabolic biomarkers. However, the reported metabolic biomarkers for colorectal cancer currently require blood samples, while gene testing for colorectal cancer risk requires fecal samples, which do not offer advantages in terms of non-invasiveness and ease of sample collection.
[0007] Therefore, there is an urgent need to find a biomarker that can be conveniently and quickly sampled non-invasively and can predict an individual's risk of colorectal cancer at an early stage, so as to achieve more efficient assessment of colorectal cancer risk. Summary of the Invention
[0008] To address the problems existing in the prior art, this invention provides a biomarker for colorectal cancer detection. Utilizing metabolomics, it analyzes metabolites that show significant differences in urine between colorectal cancer patients and healthy individuals to screen out a series of biomarkers that can predict the early risk of colorectal cancer (CRC). From these, a further set of biomarkers is selected to construct a diagnostic model for colorectal cancer, which can be used to conveniently, non-invasively, and efficiently predict whether an individual has colorectal cancer, meeting clinical needs.
[0009] On one hand, the present invention provides the use of a biomarker in the preparation of a reagent for predicting whether an individual has colorectal cancer, said biomarker being selected from one or more of the following: 2-piperidinone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvic acid, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidinoic acid, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, and tris(hydroxymethyl)aminomethane acetate.
[0010] This invention utilizes non-targeted metabolomics research, employing UPLC-MS / MS (high-performance liquid chromatography-tandem mass spectrometry) to analyze urine samples from two groups: a healthy group and a colorectal cancer patient group. Four statistical methods—random forest, PLS-DA, difference test, and SVM—were then used to screen for metabolites showing significant differences between the colorectal cancer samples and the control samples. Metabolites showing significant differences that were selected in all four statistical analysis methods were chosen, ultimately yielding 26 urinary metabolites that can serve as biomarkers for efficient prediction of colorectal cancer in individuals.
[0011] In some methods, the biomarker that can be used to predict whether an individual has colorectal cancer can be used to prepare detection reagents with the biomarker as the detection target, such as sample pretreatment reagents, antigens or antibodies, and other biological reagents and kits suitable for the detection of the biomarker; or it can be developed into standardized reagents or kits suitable for the LC-UV or LC-MS detection of the biomarker.
[0012] In some embodiments, the biomarkers of the present invention are obtained through screening of urine samples, and are particularly suitable for development into urine test reagents or kits for colorectal cancer prediction.
[0013] In some methods, when the selected biomarkers are amino acids or amino acid derivatives or contain amino groups, such as 4-hydroxyphenylacetylglutamine, N-acetyl-pentanediamine, N-formylmethionine, N-methyl-4-aminobutyric acid, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylethanesulfonic acid, and phenylacetylthreonine, reagents or kits suitable for use with amino acid analyzers or LC-UV for detecting these biomarkers can be prepared in conjunction with amino acid analysis methods such as PITC, AQC, OPA, or FMOC.
[0014] Furthermore, the detection of biomarkers in urine refers to the presence, relative abundance, or concentration of biomarkers in the urine sample of the individual being tested.
[0015] In some methods, relative abundance is preferred, which is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry. For example, if the average peak area of a biomarker measured in a control sample (an individual without colon cancer) is 500, and the average peak area measured in a colorectal cancer sample is 3000, then the abundance of the biomarker in the colorectal cancer sample is considered to be 6 times that in the control sample.
[0016] Further, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvic acid, dimethylguanidinoic acid, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.
[0017] By examining the concentration differences of biomarkers in the urine of colorectal cancer patients and normal individuals, and ranking them according to the fold change, the 10 biomarkers with the largest fold change between colorectal cancer patients and normal controls were further selected from 26 biomarkers (theoretically, these compounds with large fold changes may be the most effective biomarkers). These biomarkers can be used to more effectively distinguish or predict the risk of colorectal cancer, or to construct diagnostic models for colorectal cancer.
[0018] Furthermore, the reagent is used to detect biomarkers in urine.
[0019] This invention identifies biomarkers for colorectal cancer through urine screening. These biomarkers show significant differences in the urine of patients with and without colorectal cancer. By collecting urine samples, the detection of these biomarkers in an individual's urine can predict or aid in the diagnosis of whether that individual has colorectal cancer or the likelihood of having it. Alternatively, these biomarkers can be detected in the urine of a group, thereby classifying that group into a colorectal cancer group or a non-colorectal cancer group. Compared to blood and feces, urine collection is non-invasive and simple, offering greater advantages and prospects for using urine biomarkers in the preparation of diagnostic reagents for colorectal cancer or in the diagnosis of colorectal cancer.
[0020] On the other hand, the present invention provides a kit or chip for predicting whether an individual has colorectal cancer, the kit or chip including detection reagents for biomarkers as described above.
[0021] Furthermore, the reagent is used to detect biomarkers in urine.
[0022] In another aspect, the present invention provides a combination of biomarkers for predicting whether an individual has colorectal cancer, the combination of biomarkers comprising the following biomarkers: 4-hydroxyphenylpyruvic acid, dimethylguanidinoic acid, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.
[0023] Further, the biomarker combination includes the following biomarkers: 2-piperidinone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvic acid, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidinoic acid, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, and tris(hydroxymethyl)aminomethane acetate.
[0024] In another aspect, the present invention provides a system for predicting whether an individual has colorectal cancer, the system including a data analysis module; the data analysis module is used to analyze the detection values of biomarkers, the biomarkers being selected from one or more of the following: 2-piperidinone, 3-hydroxyaminobenzoic acid, 3-hydroxyindole sulfate, 4-hydroxyphenylacetylglutamine, 4-hydroxyphenylpyruvic acid, 5-hydroxyindole glucoside, 6-hydroxyindole sulfate, dimethylguanidinoic acid, N-acetyl-pentanediamine, N-formylmethionine, nicotinamide, nicotinamide-N-oxide, N-methyl-4-aminobutyric acid, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamic acid, phenylacetylglutamine, phenylacetylhistidine, phenylacetylmethionine, phenylacetylserine, phenylacetylaminoethanesulfonic acid, phenylacetylthreonine, trimethylamine-N-oxide, xanthine, tris(hydroxymethyl)aminomethane acetate.
[0025] Further, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvic acid, dimethylguanidinoic acid, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, phenylacetylthreonine, 3-hydroxyaminobenzoic acid, 5-hydroxyindole glucosinolate, phenylacetylglutamic acid, phenylacetylhistidine, 2-piperidinone, N-formylmethionine, phenylacetylaminoethanesulfonic acid, 3-hydroxyindole sulfate, 6-hydroxyindole sulfate, and trimethylamine-N-oxide.
[0026] Further, the biomarker is selected from one or more of the following: 4-hydroxyphenylpyruvic acid, dimethylguanidinoic acid, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.
[0027] Furthermore, the detection value of the biomarker is the detection value of the biomarker in urine.
[0028] Furthermore, the detection value of the biomarker is the presence, relative abundance, or concentration of the biomarker in the urine sample of the individual.
[0029] Furthermore, the data analysis module uses random forest or logistic regression equations to construct models for analysis.
[0030] Furthermore, the data analysis module calculates the predictive value for whether an individual has colorectal cancer by substituting the detection value of the biomarker into the logistic regression equation, thereby assessing whether an individual has colorectal cancer.
[0031] Furthermore, the logistic regression equation is:
[0032] z = 4-hydroxyphenylpyruvic acid * 0.037986 + dimethylguanidinoic acid * 0.4818-N-methyl-4-aminobutyric acid * 1.0077-nicotinamide * 1.525-p-cresol glucuronide * 0.0353-p-cresol sulfate * 0.021798-phenylacetylalanine * 0.1902 + phenylacetylglutamine * 0.858-phenylacetylmethionine * 0.118805 + phenylacetylthreonine * 0.59727 + 0.7486;
[0033]
[0034] Where e is the base of the natural logarithm; p represents the predictive value for whether an individual has colorectal cancer.
[0035] e is the base of the natural logarithm, an infinite non-repeating decimal with a value of 2.71828..., defined as follows: as n approaches infinity, (1 + 1 / n) n The limit ( ).
[0036] The biomarker name represents the relative abundance of the corresponding biomarker in the urine sample, which is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry.
[0037] Furthermore, when p is greater than 0.5, the probability of an individual having colorectal cancer is high; when p is less than 0.5, the probability of an individual having colorectal cancer is low.
[0038] In another aspect, the present invention provides the use of the system described above for constructing a detection model that predicts the probability of whether an individual has colorectal cancer.
[0039] The beneficial effects of this invention are as follows:
[0040] 1. Twenty-six novel biomarkers were identified that can predict the early risk of colorectal cancer (CRC).
[0041] 2. Two, three, five, ten, twenty, and twenty-six biomarkers were selected to construct a random forest diagnostic model for colorectal cancer. It was found that the model using ten biomarkers was the optimal one.
[0042] 3. Comparing the random forest model and the logistic regression model constructed using 10 biomarkers, it was found that the logistic regression model can further improve the detection accuracy and can be used to more efficiently predict whether an individual has colorectal cancer, with an AUC value of 0.957.
[0043] 4. Testing only requires urine samples, which is non-invasive and more convenient, and has greater advantages and prospects compared to testing through serum or fecal samples. Attached Figure Description
[0044] Figure 1 This is a flowchart of the process for screening biomarkers in urine using metabolomics in Example 1;
[0045] Figure 2 The structural formula of 3-hydroxyindole sulfate in Example 1;
[0046] Figure 3 The structural formula of 4-hydroxyphenylacetylglutamine in Example 1;
[0047] Figure 4 The structural formula of 5-hydroxyindoleglucoside in Example 1 is shown.
[0048] Figure 5 The structural formula of phenylacetylglutamic acid in Example 1;
[0049] Figure 6 The structural formula of phenylacetylhistidine in Example 1;
[0050] Figure 7 The structural formula of phenylacetylmethionine in Example 1;
[0051] Figure 8 The structural formula of phenylacetylthreonine in Example 1;
[0052] Figure 9This is a schematic diagram comparing the predictive accuracy of colorectal cancer diagnostic models constructed by selecting 2, 3, 5, 10, 20, and 26 biomarkers from 26 biomarkers in Example 2.
[0053] Figure 10 The ROC curve for the model constructed in Example 2 that predicts whether a colorectal cancer is present;
[0054] Figure 11 This is an analysis graph of the random forest model used in Example 2 to predict colorectal cancer.
[0055] Figure 12 The ROC curve of the logistic regression model for predicting colorectal cancer constructed in Example 2;
[0056] Figure 13 This is an analysis graph of the logistic regression model used in Example 2 to predict colorectal cancer.
[0057] Figure 14 This is the accuracy assessment result of the prediction model for colorectal cancer in Example 3. Detailed Implementation
[0058] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate understanding of the present invention and are not intended to limit it in any way. The reagents used in this embodiment are all known products and were obtained by purchasing commercially available products.
[0059] Example 1: Screening for biomarkers of colorectal cancer in urine using metabolomics
[0060] This embodiment first employs non-targeted metabolomics to analyze urine samples from two groups: a healthy group and a colorectal cancer patient group, using UPLC-MS / MS. Secondly, four statistical methods—random forest, PLS-DA, volcano, and SVM—were used to screen for metabolites showing significant differences between colorectal cancer and control samples. Metabolites showing significant differences in all four statistical analysis methods were selected, ultimately yielding 26 urinary metabolites as biomarkers. The role of these biomarkers in the diagnosis or differentiation of colorectal cancer was then validated (see flowchart). Figure 1 ).
[0061] The specific steps are as follows:
[0062] 1. Experimental Methods
[0063] ① Sample collection
[0064] Urine samples were collected from 50 patients with colorectal cancer and 50 control individuals (those without colorectal cancer). The colorectal cancer patients were those whose colorectal cancer was confirmed by colonoscopy.
[0065] ② Sample processing
[0066] Add methanol to the urine sample at a ratio of 1:4, shake for 3 minutes to mix, and then centrifuge at 20℃ and 4000×g for 10 minutes. Take 100 μL of supernatant from each sample and transfer it to four sample plates. Dry the plates with nitrogen and add reconstitution solution for subsequent LC-MS / MS detection.
[0067] ③LC-MS / MS detection and data processing
[0068] m / z ions were extracted from the raw mass spectrometry data obtained by LC-MS / MS detection. Metabolites were identified by searching databases, and peak areas were obtained by integrating the chromatographic peaks of the metabolites. Data normalization and missing value imputation were performed. The resulting data matrix was then subjected to subsequent bioinformatics analysis, including four statistical methods: random forest, partial least squares analysis (PLS-DA), volcano plot, and support vector machine (SVM). A ranking list of differentially expressed metabolites that were most effective in grouping colorectal cancer samples and control samples was selected for each method. Finally, metabolites selected by all four methods were chosen as biomarkers for colorectal cancer.
[0069] 2. Experimental Results
[0070] Four statistical methods—random forest, PLS-DA, differential test, and SVM—were used to screen for 32, 41, 35, and 52 differentially expressed metabolites, respectively. Among them, 26 metabolites were screened using all four data analysis methods, which are the 26 biomarkers, as shown in Table 1.
[0071] Table 1. 25 biomarkers for colorectal cancer
[0072]
[0073] Example 2: Colorectal Cancer Prediction Model
[0074] This embodiment utilizes single biomarkers or combinations of multiple biomarkers screened in Example 1 to establish predictive or diagnostic models for colorectal cancer. These models are used to distinguish between colorectal cancer and non-colorectal cancer, or to screen for colorectal cancer patients from a population, or to predict whether an individual is a colorectal cancer patient or the likelihood of an individual developing colorectal cancer. The specific models are as follows.
[0075] 1. Single biomarker
[0076] Data was processed using R software. Based on the groups of colorectal cancer patients and non-colorectal cancer individuals, the concentration changes of 26 biomarkers in urine samples from both groups were determined. All test results were subjected to LASSO regression analysis to establish a mathematical model predicting whether an individual has colorectal cancer. Calibration curves and ROC curves were used to evaluate the effectiveness of the regression model.
[0077] The analysis results show that 26 biomarkers are significantly correlated with colorectal cancer. The analysis results are shown in Tables 2 and 3.
[0078] Table 2. Comparison of correlation results between 26 biomarkers and colorectal cancer status.
[0079]
[0080] Table 3. Results of ROC analysis for single biomarkers
[0081]
[0082] The correlation between changes in the concentration of 26 biomarkers and the presence or absence of colorectal cancer can be distinguished using the OR and p-value in Table 2, and the AUC values in Table 3. Among these, the OR and AUC values are the most intuitive and obvious. A higher OR value indicates a greater impact of the biomarker on the individual with colorectal cancer compared to those without, indicating more significant exposure. A higher AUC value indicates that the biomarker can more accurately distinguish between individuals with and without colorectal cancer.
[0083] As shown in Table 2, the concentration changes of 26 biomarkers were significantly associated with the presence or absence of colorectal cancer. Among them, phenylacetylglutamine showed the highest association with an OR of 2.36, followed by phenylacetylthreonine with an OR of 1.82.
[0084] As shown in Table 3, when the concentration change of any one of the 26 biomarkers is used to distinguish between colorectal cancer patients and non-colorectal cancer patients, the AUC value can reach above 0.63, which shows high accuracy. Among them, phenylacetylglutamine has the highest AUC value of 0.7876, followed by p-cresol glucuronide with an AUC value of 0.7836.
[0085] 2. Combination of multiple biomarkers
[0086] While a single biomarker can distinguish between urine samples with colorectal cancer and those without, or to predict colorectal cancer, combining multiple biomarkers generally results in higher accuracy in differentiation or prediction.
[0087] However, a single biomarker that is more accurate in predicting colorectal cancer does not necessarily play a greater role in the combination with one or more other biomarkers. Also, the more biomarkers there are, the higher the predictive accuracy (AUC value) of the combination is not necessarily true. Therefore, a large number of validation experiments are still needed.
[0088] Since the AUC and OR values of biomarkers are more suited to assessing the relative importance of variables in a statistical model and are not suitable for selecting variables to build a model, this embodiment preferably uses 2, 3, 5, 10, 20, and 26 biomarkers with the highest fold change in concentration between colorectal cancer and non-colorectal cancer urine samples to build a diagnostic model for colorectal cancer. The fold change in concentration of the 26 biomarkers between colorectal cancer and non-colorectal cancer urine samples (Fold Change, Fold Change = mean expression of disease samples divided by mean expression of normal samples) is ranked from high to low, and the results are shown in Table 4.
[0089] Table 4. Ranking of the fold differences in concentration of 26 biomarkers between urine samples from colorectal cancer and non-colorectal cancer patients.
[0090]
[0091] Based on the concentration differences of the 26 biomarkers in urine samples from colorectal cancer and non-colorectal cancer provided in Table 4, this embodiment selects 2, 3, 5, 10, 20, and 26 of the 26 biomarkers respectively, and constructs a diagnostic model for colorectal cancer using random forest.
[0092] Among them, the two biomarkers are the first and second ranked biomarkers in Table 4 (p-cresol sulfate and phenylacetylthreonine). In the constructed random forest model, the information gain ratio (GINI coefficient) of p-cresol sulfate is 25.31 and the mean decrease accuracy is 21.17; the GINI coefficient of phenylacetylthreonine is 24.22 and the mean decrease accuracy is 16.71.
[0093] The three biomarkers are the three ranked biomarkers in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate is 15.43, with an average reduction precision of 16.37; the GINI coefficient of phenylacetylthreonine is 15.75, with an average reduction precision of 15.04; and the GINI coefficient of N-methyl-4-aminobutyric acid is 18.33, with an average reduction precision of 24.42.
[0094] The five biomarkers are those ranked 1st to 5th in Table 4. In the constructed random forest model, the GINI coefficient of p-cresol sulfate was 7.86, with an average reduction precision of 10.99; the GINI coefficient of phenylacetylthreonine was 6.39, with an average reduction precision of 5.58; the GINI coefficient of N-methyl-4-aminobutyric acid was 13.73, with an average reduction precision of 25.36; the GINI coefficient of 4-hydroxyphenylpyruvic acid was 10.43, with an average reduction precision of 45.38; and the GINI coefficient of phenylacetylmethionine was 11.05, with an average reduction precision of 18.74.
[0095] The ten biomarkers are those ranked 1st to 10th in Table 4. In the constructed random forest model, the GINI coefficient for p-cresol sulfate was 3.64, with an average reduction precision of 7.56; the GINI coefficient for phenylacetylthreonine was 2.46, with an average reduction precision of 4.80; the GINI coefficient for N-methyl-4-aminobutyric acid was 8.04, with an average reduction precision of 18.60; the GINI coefficient for 4-hydroxyphenylpyruvic acid was 6.25, with an average reduction precision of 12.60; and the GINI coefficient for phenylacetylmethionine was... The GINI coefficient for p-cresol glucuronide was 6.26, with an average reduction precision of 12.85; the GINI coefficient for nicotinamide was 5.20, with an average reduction precision of 11.07; the GINI coefficient for nicotinamide was 6.56, with an average reduction precision of 12.51; the GINI coefficient for phenylacetylalanine was 3.18, with an average reduction precision of 6.30; the GINI coefficient for phenylacetylglutamine was 4.47, with an average reduction precision of 6.83; and the GINI coefficient for dimethylguanidinoic acid was 3.43, with an average reduction precision of 9.16.
[0096] The 20 biomarkers are those ranked 1st to 20th in Table 4. In the constructed random forest model, the GINI coefficient for p-cresol sulfate was 2.36, with an average reduction precision of 6.21; the GINI coefficient for phenylacetylthreonine was 1.73, with an average reduction precision of 4.02; the GINI coefficient for N-methyl-4-aminobutyric acid was 5.92, with an average reduction precision of 16.23; and the GINI coefficient for 4-hydroxyphenylpyruvic acid was 4.10, with an average reduction precision of 9.2. 8; The GINI coefficient of phenylacetylmethionine is 3.79, with an average reduction precision of 10.13; the GINI coefficient of p-cresol glucuronide is 3.77, with an average reduction precision of 9.49; the GINI coefficient of nicotinamide is 4.67, with an average reduction precision of 11.61; the GINI coefficient of phenylacetylalanine is 2.26, with an average reduction precision of 5.84; the GINI coefficient of phenylacetylglutamine is 2.67, with an average reduction precision of 7.71; the GINI coefficient of dimethylguanidinoic acid is... The GINI coefficient for 3-hydroxyaminobenzoic acid was 2.03, with an average reduction precision of 4.32; the GINI coefficient for 5-hydroxyindoleglucoside was 2.69, with an average reduction precision of 5.66; the GINI coefficient for phenylacetylglutamic acid was 1.59, with an average reduction precision of 4.38; the GINI coefficient for phenylacetylhistidine was 1.62, with an average reduction precision of 4.96; and the GINI coefficient for 2-piperidone was 1.57, with an average reduction precision of [missing value]. The GINI coefficient for N-formylmethionine was 1.85, with an average reduction precision of 2.81; the GINI coefficient for phenylacetamidoethanesulfonic acid was 1.28, with an average reduction precision of 0.79; the GINI coefficient for 3-hydroxyindole sulfate was 1.41, with an average reduction precision of 3.51; the GINI coefficient for 6-hydroxyindole sulfate was 1.57, with an average reduction precision of 1.93; and the GINI coefficient for trimethylamine-N-oxide was 1.02, with an average reduction precision of 2.61.
[0097] The 26 biomarkers are those ranked 1st to 26th in Table 4. In the constructed random forest model, the GINI coefficient for p-cresol sulfate was 1.69, with an average reduction precision of 7.04; the GINI coefficient for phenylacetylthreonine was 1.04, with an average reduction precision of 2.80; the GINI coefficient for N-methyl-4-aminobutyric acid was 3.57, with an average reduction precision of 12.93; the GINI coefficient for 4-hydroxyphenylpyruvic acid was 2.45, with an average reduction precision of 5.50; the GINI coefficient for phenylacetylmethionine was 2.68, with an average reduction precision of 7.68; and the GINI coefficient for p-cresol glucuronide was... The GINI coefficient for nicotinamide was 2.61, with an average reduction precision of 8.31; the GINI coefficient for phenylacetylalanine was 2.56, with an average reduction precision of 8.02; the GINI coefficient for phenylacetylglutamine was 1.47, with an average reduction precision of 4.84; the GINI coefficient for phenylacetylglutamine was 1.83, with an average reduction precision of 5.74; the GINI coefficient for dimethylguanidinol was 1.34, with an average reduction precision of 3.76; the GINI coefficient for 3-hydroxyaminobenzoic acid was 1.14, with an average reduction precision of 4.11; the GINI coefficient for 5-hydroxyindoleglucoside was 1.76, with an average reduction precision of 4.39; and the GINI coefficient for phenylacetylglutamic acid was 0. The GINI coefficient for phenylacetylhistidine was 1.00, with an average decrease precision of 4.79; the GINI coefficient for 2-piperidinone was 1.20, with an average decrease precision of 1.80; the GINI coefficient for N-formylmethionine was 0.79, with an average decrease precision of 2.15; the GINI coefficient for phenylacetyl ethanesulfonic acid was 0.58, with an average decrease precision of 2.70; the GINI coefficient for 3-hydroxyindole sulfate was 0.96, with an average decrease precision of 3.64; the GINI coefficient for 6-hydroxyindole sulfate was 0.73, with an average decrease precision of 2.70; and the GINI coefficient for trimethylamine-N-oxide was 0.88, with an average decrease precision of 3.11. The GINI coefficient for 4-hydroxyphenylacetylglutamine was 0.74, with an average reduction precision of 2.33; the GINI coefficient for N-acetyl-pentanediamine was 0.83, with an average reduction precision of 4.61; the GINI coefficient for N-acetyl-pentanediamine was 2.22, with an average reduction precision of 7.72; the GINI coefficient for tris(hydroxymethyl)aminomethane acetate was 2.48, with an average reduction precision of 8.06; the GINI coefficient for xanthine was 2.70, with an average reduction precision of 8.67; the GINI coefficient for nicotinamide-N-oxide was 8.21, with an average reduction precision of 16.94; and the GINI coefficient for phenylacetylserine was 2.01, with an average reduction precision of 7.16.
[0098] The AUC values and 95% CL confidence intervals of the six random forest diagnostic models constructed using 2, 3, 5, 10, 20, and 26 biomarkers, as described above, were calculated respectively. The results are as follows: Figure 9 As shown.
[0099] Depend on Figure 9 As can be seen, selecting the top two biomarkers from 26 biomarkers to construct the model only yields an AUC of 0.922, with a 95% confidence interval (CL) of 0.718–0.999. As the number of biomarkers selected increases, the AUC gradually rises, while the 95% CL gradually narrows. When 10 biomarkers are selected to construct the colorectal cancer diagnostic model, the AUC reaches 0.935, with a 95% CL CL of 0.842–0.998. However, when the number of biomarkers further increases to 20 or 26, the potential for further AUC increases is very limited, and the confidence interval widens. Furthermore, compared to 20 or 26 biomarkers, using 10 biomarkers reduces the number of variables and lowers the model's complexity. Therefore, using the top 10 biomarkers listed in Table 4 to construct the colorectal cancer diagnostic model is preferred, as it not only achieves very good predictive accuracy but also results in a simpler and more convenient model.
[0100] Using 42 clinically known colorectal cancer patients and 42 non-colorectal cancer patients as the total dataset, the biomarker detection values of their urine samples were detected. The data were analyzed using a random forest model with 10 biomarkers, and the analysis graph is shown below. Figure 11 As shown, by Figure 11 It can be seen that the random forest model constructed using 10 biomarkers has a certain error when used to predict colorectal cancer (of course, the error is unavoidable). Among 42 colorectal cancer patients, 37 were detected, and among 42 non-colorectal cancer patients, 5 were classified as colorectal cancer patients, with an accuracy rate of 88%. Figure 11 It can be seen that when the predictive value P is greater than 0.5, the probability of an individual having colorectal cancer is high; when the predictive value p is less than 0.5, the probability of an individual having colorectal cancer is low.
[0101] Using the top 10 Fold Change biomarkers, multivariate regression analysis was performed to establish a logistic regression assessment model for predicting whether an individual will have colorectal cancer.
[0102] z = 4-hydroxyphenylpyruvic acid * 0.037986 + dimethylguanidinoic acid * 0.4818-N-methyl-4-aminobutyric acid * 1.0077-nicotinamide * 1.525-p-cresol glucuronide * 0.0353-p-cresol sulfate * 0.021798-phenylacetylalanine * 0.1902 + phenylacetylglutamine * 0.858-phenylacetylmethionine * 0.118805 + phenylacetylthreonine * 0.59727 + 0.7486;
[0103]
[0104] Where e is the base of the natural logarithm; p represents the predictive value for whether an individual has colorectal cancer; and the biomarker name represents the relative abundance of the corresponding biomarker in the urine sample, which is the peak area of the biomarker in the detection spectrum obtained by high performance liquid chromatography-tandem mass spectrometry.
[0105] The ROC curve of the logistic regression model for predicting whether an individual has colorectal cancer provided in this embodiment is as follows: Figure 12 As shown, the AUC value reached 0.957, which is a significant improvement compared to the random forest model with 10 biomarkers.
[0106] The logistic regression model used to predict whether an individual has colorectal cancer was analyzed using a dataset of 50 clinically known colorectal cancer patients and 50 non-colorectal cancer patients. The results are as follows: Figure 13 As shown in Table 5,
[0107] Table 5. Results of the model for predicting whether an individual will develop colorectal cancer
[0108]
[0109] Depend on Figure 13 As shown in Table 5, the logistic regression assessment model for predicting whether an individual has colorectal cancer, constructed using 10 biomarkers, was used to analyze the results. Among 50 colorectal cancer patients, 45 were detected, and among 50 non-colorectal cancer patients, 5 were classified as colorectal cancer patients, achieving an accuracy rate of over 90%, indicating an improvement in accuracy.
[0110] Depend on Figure 13 It can also be seen that P = 0.5 can be used as a dividing point for judgment. When P is greater than 0.5, the probability of predicting that an individual has colorectal cancer is high; when P is less than 0.5, the probability of predicting that an individual has colorectal cancer is low.
[0111] Example 3: Evaluation of a model for predicting colorectal cancer
[0112] This embodiment evaluates the clinical accuracy of the colorectal cancer prediction model constructed in Example 2. The dataset consists of 42 colorectal cancer patients and 42 non-colorectal cancer patients. Eight CRC patients and healthy individuals (non-CRC patients) are randomly selected from this dataset. Urine samples are collected, and the relative abundance of 10 biomarkers in the model is determined according to the sample processing method in Example 1. The model is then used to calculate the prediction value P to predict whether an individual has colorectal cancer. The results are as follows: Figure 14 As shown.
[0113] Depend on Figure 14It can be seen that all 8 patients with colorectal cancer were detected, and one of the 8 normal individuals was predicted to have colorectal cancer, with an accuracy rate of 93.75%.
[0114] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. The use of a system for constructing a detection model to predict the probability of whether an individual has colorectal cancer, characterized in that, The system includes a data analysis module; the data analysis module is used to analyze the detection values of biomarkers, which are composed of 4-hydroxyphenylpyruvic acid, dimethylguanidinoic acid, N-methyl-4-aminobutyric acid, nicotinamide, p-cresol glucuronide, p-cresol sulfate, phenylacetylalanine, phenylacetylglutamine, phenylacetylmethionine, and phenylacetylthreonine.
2. The use as described in claim 1, characterized in that, The detection value of the biomarker is the detection value of the biomarker in urine.
3. The use as described in claim 2, characterized in that, The detection value of the biomarker is the presence, relative abundance, or concentration of the biomarker in the urine sample of the individual.
4. The use as described in claim 3, characterized in that, The data analysis module uses random forest or logistic regression equations to build models for analysis.
5. The use as described in claim 4, characterized in that, The data analysis module calculates the predictive value for whether an individual has colorectal cancer by substituting the detection value of the biomarker into the logistic regression equation, thereby assessing whether an individual has colorectal cancer.
Citation Information
Patent Citations
Colon cancer diagnostic marker and application thereof
CN109884300A
Method for constructing mathematical model for detecting colorectal cancer in vitro, and application thereof
CN111584008A