Colorectal cancer related methylation related genetic biomarker and application thereof
By identifying the methylation-related genetic biomarker SNP site rs312833 and constructing a risk prediction model, the problem of unsatisfactory prediction performance of traditional models in the Chinese population was solved, achieving highly reliable and applicable early screening and diagnosis of colorectal cancer.
Patent Information
- Application Number
- CN202511853920.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-15
AI Technical Summary
Traditional genetic prediction models for colorectal cancer rely heavily on GWAS loci in European and American populations, resulting in unsatisfactory prediction performance in Chinese populations and a lack of effective biomarkers for early screening and diagnosis.
We identified the methylation-related genetic biomarker SNP rs312833 associated with colorectal cancer, constructed a genetic risk prediction model based on this biomarker, and prepared an auxiliary diagnostic kit for colorectal cancer in the Chinese population, including specific amplification primers and probes, for detecting the genotyping of rs312833 in peripheral blood DNA of the population.
It has enabled highly reliable prediction of colorectal cancer risk in the Chinese population, providing a scientific basis for early screening and individualized prevention in high-risk groups, and improving the accuracy and applicability of diagnosis.
Smart Images

Figure CN122038565A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of genetic engineering and oncology, and relates to methylation-related genetic biomarkers associated with colorectal cancer and their applications. Background Technology
[0002] Colorectal cancer is a common malignant tumor of the digestive system. The severity of colorectal cancer is classified into four stages based on histopathological characteristics. Up to 50% of colorectal cancer patients will develop metastatic disease (stage IV), and their five-year survival rate will drop to 10%. Currently, the commonly used diagnostic methods for colorectal cancer include colonoscopy and stool sample testing; however, there is a lack of effective new biomarkers for early screening and diagnosis.
[0003] The development and progression of colorectal cancer is a complex process involving multiple factors, with genetic factors accounting for 12%–35% of the risk. Individuals with a family history of colorectal cancer have a higher risk of developing the disease. Current genome-wide association studies (GWAS) have identified several single nucleotide polymorphisms (SNPs) closely associated with colorectal cancer risk. Recent research indicates that base changes in genetic variants can influence disease development by regulating m6A or 5mC methylation levels, suggesting the crucial importance of screening for methylation-related genetic variants as novel biomarkers for colorectal cancer screening. The genetic risk score (GRS), as an important method for evaluating risk prediction in epidemiological studies, is used to assess an individual's genetic risk of developing a particular disease. Therefore, identifying methylation-related genetic variants in colorectal cancer and constructing genetic risk prediction models for colorectal cancer is of significant cutting-edge importance for screening high-risk populations and achieving early prevention.
[0004] However, the diversity of genetic backgrounds limits the compatibility of genetic risk prediction models between Asian and European / American populations. Traditional genetic prediction models for colorectal cancer largely rely on GWAS loci in European / American populations, resulting in unsatisfactory predictive performance in Chinese populations and failing to fully realize the potential of genetic prediction models. Summary of the Invention
[0005] To address the issue that traditional colorectal cancer genetic prediction models often rely on GWAS loci in European and American populations, resulting in unsatisfactory prediction performance in Chinese populations, this invention provides a methylation-related genetic biomarker for colorectal cancer and its application. This methylation-related genetic biomarker is suitable for the auxiliary diagnosis of colorectal cancer in the Chinese population.
[0006] In a first aspect, the present invention provides a methylation-related genetic biomarker associated with colorectal cancer, characterized in that the methylation-related genetic biomarker is an SNP site rs312833; the SNP site rs312833 is located at base 75383206 on chromosome 17, and has three genotypes, namely wild-type homozygous CC, heterozygous CT, and homozygous mutant TT.
[0007] Secondly, based on the aforementioned methylation-related genetic biomarkers, this invention constructs a colorectal cancer incidence risk prediction model, with the prediction equation as follows:
[0008] wGRS=0.7463×rs312833 - 0.9171×age - 0.4600×sex
[0009] Among them, the value of rs312833 is "0" for wild homozygous type, "1" for heterozygous type, and "2" for homozygous mutant type; the value of sex is 1 or 2, the value of female is 1, and the value of male is 2; the value of age ranges from 19 to 89, and is determined according to the actual age.
[0010] When wGRS is higher than -0.202, the subjects have a higher risk of colorectal cancer.
[0011] Secondly, the present invention has prepared an auxiliary diagnostic kit for colorectal cancer in the Chinese population based on the methylation-related genetic biomarkers.
[0012] Furthermore, the kit contains specific amplification primers and specific probes for the methylation-related genetic biomarkers, and the kit is used to detect the genotyping of rs312833 in peripheral blood DNA from a population.
[0013] Furthermore, the specific amplification primer sequences for the methylation-related genetic biomarkers include the sequences shown in SEQ ID No. 5, SEQ ID No. 6, SEQ ID No. 7, and SEQ ID No. 8.
[0014] Furthermore, the specific probes for the methylation-related genetic biomarkers include the sequences shown in SEQ ID No. 1, SEQ ID No. 2, SEQ ID No. 3, and SEQ ID No. 4; the sequences shown in SEQ ID No. 1, SEQ ID No. 2, SEQ ID No. 3, and SEQ ID No. 4 are labeled with biotin at both ends.
[0015] Beneficial effects:
[0016] (1) Based on multi-center, large-sample GWAS data of the Chinese population, this invention uses multi-stage validation to identify rs312833, a highly reliable methylation-related genetic biomarker for the auxiliary diagnosis of colorectal cancer in the Chinese population, and further constructs a genetic risk prediction model that can effectively predict the risk of colorectal cancer in the population.
[0017] (2) This invention screened out the genetic variant rs312833, which is significantly associated with the risk of colorectal cancer. The colorectal cancer auxiliary diagnostic kit prepared based on this biomarker helps clinicians quickly grasp the genetic risk of colorectal cancer and has universal applicability to the Chinese population. It provides scientific basis and decision support for screening, individualized prevention and scientific intervention of high-risk groups for colorectal cancer. Attached Figure Description
[0018] Figure 1 The calibration curves and ROC curves of the colorectal cancer genetic risk prediction model constructed based on genetic variation sites in the training and validation sets are examples of the present invention. Detailed Implementation
[0019] The technical solution of the present invention will be described in detail below through embodiments, but the scope of protection of the present invention is not limited to the embodiments described.
[0020] Example 1
[0021] The research method used in this embodiment is as follows:
[0022] Step 1, Sample Collection: Based on the sample inclusion and exclusion criteria, the system collects blood samples from eligible individuals, extracts genomic DNA, and stores it in a biobank.
[0023] 1.1. Determining the study subjects: The case population consisted of patients with pathologically diagnosed primary colorectal adenomas and colorectal cancer, excluding patients with carcinoid tumors, secondary colorectal cancer, and other malignant tumors. The healthy control population consisted of healthy individuals from the same community undergoing physical examinations, and were frequency-matched to the case population based on gender and age.
[0024] This study included 3,920 cases and 4,444 controls from different regions of China. The training set consisted of 1,150 cases and 13,42 controls from Nanjing. The validation set consisted of 2,770 cases and 3,102 controls from East and North China, used to screen for colorectal cancer risk sites. This study was approved by the Ethics Committee of Nanjing Medical University.
[0025] 1.2. Genomic DNA extraction: Peripheral blood genomic DNA was extracted using the adsorption column method, following standard procedures. The OD260 / OD280 ratio was within the range of 1.7-1.9.
[0026] Step 2, Genotyping and Susceptibility Site Screening: Genotyping was performed using whole-genome high-density SNP microarrays and TaqMan probe genotyping. Methylation-related SNP sites that were significantly associated with the risk of colorectal cancer were identified by combining methylation sequencing results and epidemiological data.
[0027] 2.1. Perform genome-wide SNP detection and preliminary screening of methylation-related risk sites on the training set:
[0028] (1) The whole genome SNP site was scanned using the Illumina Human Omni ZhongHua Bead Chips chip, and the genotype was filled using the IMPUTE2 software;
[0029] (2) Combining the m6A-seq results and DNA methylation data, 25 SNPs were screened out near the abnormally enriched m6A peak and DNA methylation peak. The screening results are shown in Table 1.
[0030] (3) SNPs associated with the risk of colorectal cancer were screened in the case group and healthy control group of the training set by logistic regression analysis, and inclusion sites were selected with an association P value <0.05.
[0031] Table 1. Association analysis of 25 SNPs with colorectal cancer (training set)
[0032]
[0033] 2.2. Validation of individual SNP genotyping and risk loci determination in the validation set:
[0034] (1) Design specific amplification primers and specific probes for rs312833. Use the TaqMan MGB probe method to genotype the DNA of the validation set population samples. The specific amplification primers and probes are shown in Table 2.
[0035] The genotyping results are CC, CT, and TT. rs312833 is located at base 75383206 on chromosome 17 and has three genotypes: wild-type homozygous CC, heterozygous CT, and homozygous mutant TT.
[0036] This embodiment uses an additive model to evaluate the relationship between T site and C site and colorectal cancer. The T site significantly increases the risk of colorectal cancer, and its effect value OR is shown in Tables 1 and 2.
[0037] (2) The association strength between candidate SNPs and the risk of colorectal cancer was assessed by logistic regression analysis. The association analysis between candidate SNPs and colorectal cancer (validation set) is shown in Table 2.
[0038] Table 2. Association analysis between candidate SNPs and colorectal cancer (validation set)
[0039]
[0040] Step 3: Construction and Evaluation of the Genetic Risk Prediction Model: Based on the screened risk SNP loci, a genetic score (GRS) model is constructed. The accuracy and reliability of the risk prediction model are evaluated in multi-stage population samples using methods such as calibration curves and receiver operating characteristic (ROC) curves to ensure that the model can effectively predict the risk of colorectal cancer.
[0041] 3.1. Based on the screened risky SNP sites, construct a genetic scoring (GRS) model.
[0042] Based on a genome-wide scanning and single SNP detection strategy, the SNP locus closely associated with the risk of colorectal cancer was identified by fitting a logistic regression model as rs312833. Basic information about the colorectal cancer risk locus rs312833 is shown in Table 3.
[0043] Table 3. Basic information on the colorectal cancer risk locus rs312833
[0044]
[0045] Note: The weighting coefficients are the coefficients of rs312833 obtained from logistic regression, indicated by the superscript in the upper right corner. a The significance is that the weighting coefficients and p-values were adjusted for covariates, with age and sex included as covariates. Population 1 was the Nanjing cohort, with 1150 cases and 1342 controls; Population 2 was the Beijing cohort, with 932 cases and 966 controls; Population 3 was the Nanjing cohort, with 855 cases and 1201 controls; and Population 4 was the Beijing cohort, with 983 cases and 935 controls. All cohorts were independent.
[0046] Using the logistic regression coefficients of rs312833 in the training set as weights, a weighted genetic risk score model (wGRS model) was constructed. Specifically, an additive logistic regression model was used to analyze the association strength between SNPs and the risk of colorectal cancer, while adjusting for confounding variables (such as gender and age). The genetic risk prediction model was fitted by calculating wGRS. The basic steps include:
[0047] (1) Quantitatively score the three genotypes of the SNP, such as wild-type homozygous as "0", heterozygous as "0", and so on.
[0048] "1", homozygous mutant is "2", and the effect scores of all sites are adjusted to be positively associated;
[0049] (2) The weight coefficients of SNP are the regression coefficients obtained from the Logistic regression model, and the equation is constructed as follows:
[0050] wGRS=0.7463×rs312833 - 0.9171×age - 0.4600×sex
[0051] In this equation: rs312833 takes the value of "0" for wild homozygous, "1" for heterozygous, and "2" for homozygous mutant; sex takes the value of 1 or 2, with females taking the value of 1 and males taking the value of 2; age takes the value of 19-89, based on the actual age.
[0052] (3) The population wGRS score was calculated according to the above equation. The range was -1.3771 to 1.4872, the mean was -0.7931, and the standard deviation was 0.5911. Based on the range of mean plus or minus one standard deviation, the sample was divided into high-risk group and low-risk group. When the wGRS was higher than -0.202, it was a high-risk group, and its risk of colorectal cancer was higher.
[0053] 3.2. The wGRS model was used to assess the risk of the validation set population sample, and the ROC curve was used to evaluate the discriminative power of the model in order to test the predictive ability of the identified risk SNPs on the incidence of colorectal cancer in the Chinese population.
[0054] The predictive performance of the GRS model was evaluated using the calibration curve and the area under the ROC curve (AUC). Based on the calibration curve and ROC curve, the model was found to have good calibration and discrimination (AUC training set = 0.628; AUC validation set = 0.654, see...). Figure 1 ).
[0055] Figure 1 The middle A (calibration curve) shows the consistency between the predicted model risk and the actual observed risk. The closer the curve is to the diagonal, the better the calibration. Figure 1 The ROC curve (B-curve) demonstrates the model's ability to distinguish between cases and controls. The AUC on the training set is 0.628, and the AUC on the validation set is 0.654, indicating that the model has stable discriminative power.
[0056] All statistical analyses were performed using PLINK 1.90 and R 4.4.1 software. The statistical significance P-value was 0.05, and all analyses were two-tailed.
[0057] The results from steps one through three show that:
[0058] (1) Based on the logistic regression additive model, with a P value of 0.05 as the evaluation index, the methylation-related colorectal cancer susceptibility site rs312833 was finally screened.
[0059] (2) Based on the screened colorectal cancer genetic loci, a GRS was constructed, and a colorectal cancer incidence risk prediction model was built using the Logistic regression method. Based on the calibration curve and ROC curve, the model was found to have good calibration and discrimination.
[0060] Step 4: Preparation of the Colorectal Cancer Auxiliary Diagnostic Kit: Based on the screened methylation-related genetic biomarkers, an auxiliary diagnostic kit for colorectal cancer in the Chinese population was designed. This kit includes specific amplification primers, specific probes, and commonly used PCR reagents.
[0061] (1) Based on the results of steps one to four, this invention has prepared a colorectal cancer auxiliary diagnostic kit for the Chinese population, containing specific primers for the identified genetic variation risk site rs312833 and other reagents. This kit helps in the early screening of high-risk groups for colorectal cancer and provides strong support for timely intervention.
[0062] A diagnostic kit for colorectal cancer in the Chinese population: This kit is based on the screened methylation-related SNP site rs312833. The diagnostic reagents include specific amplification primers and probes for rs312833, as well as commonly used PCR reagents such as Taq enzyme, dNTP mixture, and deionized water.
[0063] Table 4. Primer and probe sequences for rs312833 specific amplification
[0064]
[0065] Of these, SEQ ID No. 1, SEQ ID No. 2, SEQ ID No. 3, and SEQ ID No. 4 are specific probes (used for hybridization detection), characterized by a biotin label for capture or signal generation during the detection process. Specifically, SEQ ID No. 1 and SEQ ID No. 2 are probe pairs specifically targeting the C allele, used for the specific capture and detection of the C allele. SEQ ID No. 3 and SEQ ID No. 4 are probe pairs specifically targeting the T allele, used for the specific capture and detection of the T allele.
[0066] SEQ ID No. 5, SEQ ID No. 6, SEQ ID No. 7 and SEQ ID No. 8 are specific amplification primers (used for PCR amplification). They are characterized by being biotin-free and being standard PCR primers.
[0067] SEQ ID No. 5 and SEQ ID No. 6 are universal PCR primers for the C allele, used to amplify the DNA region of the C allele. SEQ ID No. 7 and SEQ ID No. 8 are universal PCR primers for the T allele, used to amplify the DNA region of the T allele.
[0068] SEQ ID No. 1-SEQ ID No. 8 together constitute a system for detecting rs312833 genotyping. During detection, the genotype of the sample can be clearly distinguished by judging the signal combination of the C probe and the T probe: only C signal indicates CC type, only T signal indicates TT type, and both indicate CT heterozygous type. This design ensures the accuracy and specificity of genotyping and is the basis for implementing the risk prediction model of this invention.
[0069] As described above, although the invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the invention.
Claims
1. A methylation-related genetic biomarker associated with colorectal cancer, characterized in that, The methylation-related genetic biomarker is the SNP site rs312833; SNP rs312833 is located at base 75383206 on chromosome 17. It has three genotypes: wild-type homozygous CC, heterozygous CT, and homozygous mutant TT.
2. A colorectal cancer incidence risk prediction model constructed based on the methylation-related genetic biomarkers described in claim 1, characterized in that, The prediction equation is: wGRS=0.7463×rs312833 - 0.9171×age - 0.4600×sex; Among them, the value of rs312833 is "0" for wild homozygous type, "1" for heterozygous type, and "2" for homozygous mutant type; sex is 1 or 2, 1 for female and 2 for male; age ranges from 19 to 89, and is determined according to the actual age. When wGRS is higher than -0.202, the subjects have a higher risk of colorectal cancer.
3. A diagnostic kit for colorectal cancer in the Chinese population prepared based on the methylation-related genetic biomarkers described in claim 1.
4. The reagent kit according to claim 3, characterized in that, The kit contains specific amplification primers and specific probes for the methylation-related genetic biomarkers, and the kit shown is used to detect the genotyping of rs312833 in peripheral blood DNA from a population.
5. The reagent kit according to claim 4, characterized in that, The specific amplification primer sequences for the methylation-related genetic biomarkers include the sequences shown in SEQ ID No. 5, SEQ ID No. 6, SEQ ID No. 7, and SEQ ID No.
8.
6. The reagent kit according to claim 4, characterized in that, The specific probes for the methylation-related genetic biomarkers include the sequences shown in SEQ ID No. 1, SEQ ID No. 2, SEQ ID No. 3, and SEQ ID No. 4; The sequences shown in SEQ ID No. 1, SEQ ID No. 2, SEQ ID No. 3 and SEQ ID No. 4 are labeled with biotin.