Application of leukocyte DNA methylation markers in risk prediction of colorectal neoplasms

By targeting bisulfite sequencing to screen for differentially methylated sites in leukocyte DNA, and combining machine learning and lifestyle scoring, a colorectal cancer risk prediction model was constructed. This solved the problem of insufficient sensitivity in the early diagnosis of colorectal cancer and achieved efficient risk assessment and early screening.

CN120485365BActive Publication Date: 2025-11-18CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510567462.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-11-18
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing technologies lack sufficient sensitivity and specificity in the early diagnosis of colorectal cancer, leading to missed opportunities for late diagnosis and treatment, and a lack of effective risk assessment and screening strategies.

Method used

By screening for differentially methylated sites (DMPs) in leukocyte DNA using targeted bisulfite sequencing (TBS), a multi-marker prediction model was constructed using machine learning, and combined with lifestyle scores to identify colorectal cancer risk. The efficacy of the model was validated using an Illumina 935K array and RRBS.

Benefits of technology

It improves the accuracy and sensitivity of colorectal cancer risk prediction, provides an efficient early screening tool, and has high clinical translational potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120485365B_ABST
    Figure CN120485365B_ABST
Patent Text Reader

Abstract

The application discloses application of leukocyte DNA methylation markers in colorectal tumor risk prediction, a series of leukocyte DNA methylation biomarkers related to colorectal tumors are screened and verified, finally, five methylation regions are obtained, and a novel risk prediction model is constructed based on the five methylation regions, the model has high clinical transformation potential, and provides a new risk stratification tool for prevention and early screening of colorectal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedicine, specifically to the application of leukocyte DNA methylation markers in colorectal cancer risk prediction. Background Technology

[0002] Colorectal cancer (CRC) is one of the most common malignant tumors in China. The combined treatment of surgical resection and adjuvant chemotherapy has become the standard of care, significantly improving the overall prognosis of CRC. However, a series of epigenetic and metabolic changes give CRC cancer cells a high capacity for migration and invasion, resulting in a 5-year survival rate of approximately 12% for metastatic CRC. About 30% of CRC patients have distant metastases at diagnosis, and up to 30% die from tumor metastasis and recurrence. Despite advancements in treatment, the 5-year survival rate of CRC remains heavily dependent on early diagnosis. However, early CRC often lacks obvious symptoms, leading to late diagnosis and missed treatment opportunities. Therefore, strengthening early detection and risk assessment is crucial for reducing the disease burden.

[0003] In view of this, the present invention aims to develop a scoring model with high sensitivity and specificity for identifying CRC and precancerous lesions. By integrating lifestyle factors, this study aims to establish a comprehensive risk assessment model to provide a scientific basis for personalized screening strategies. Summary of the Invention

[0004] Given the limitations of existing technologies, this invention utilizes targeted bisulfite sequencing (TBS) to detect colorectal tumors, advanced adenomas, and healthy controls. It screens and identifies differentially methylated sites (DMPs) and regions (DMRs) of leukocyte DNA. Through machine learning, five stable DMRs are selected as methylation biomarkers and incorporated into a multi-marker prediction model. The risk prediction efficacy of the five DMRs and the multi-marker prediction model is independently validated using an Illumina 935K array and RRBS. This model demonstrates excellent recognition performance in detecting colorectal tumors, and combining methylation biomarkers with lifestyle scores can further improve the accuracy of risk prediction.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] The first aspect of this invention provides a biomarker for predicting the risk of colorectal cancer, wherein the biomarker is a methylated biomarker, and the methylated biomarker includes any one of the following methylation regions: chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, chr16: 55866678-55866757.

[0007] Furthermore, the methylation biomarker is a combination of methylation regions chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757.

[0008] Furthermore, chr4: 4859985-4860551 has a low degree of methylation in the patient, chr3: 3170109-3170139 has a low degree of methylation in the patient, chr7: 99818709-99818869 has a high degree of methylation in the patient, chr9: 139582466-139582662 has a high degree of methylation in the patient, and chr16: 55866678-55866757 has a high degree of methylation in the patient.

[0009] As used herein, the terms "patient" or "subject" refer to any animal (e.g., mammal) that has or is suspected of having colorectal cancer, including but not limited to humans, non-human primates, rodents, etc. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Most preferably, the subject is a human.

[0010] Furthermore, the patient has advanced adenoma or is at high risk for colorectal cancer.

[0011] The second aspect of the present invention provides the application of a reagent for detecting biomarkers in a sample in the preparation of a colorectal cancer risk prediction product, wherein the biomarkers are those described in the first aspect of the present invention.

[0012] Furthermore, the reagents include those used in any one or more of the following methylation detection methods, which include: whole-genome bisulfite sequencing (WGBS), pyrosequencing, bisulfite sequencing, methylation-specific polymerase chain reaction (MS-PCR), bisulfite-specific polymerase chain reaction, methylation-sensitive restriction endonuclease-PCR / Southern method, combined bisulfite restriction endonuclease analysis (COBRA), digital polymerase chain reaction, restriction marker genome scanning, CpG island microarray, single nucleotide primer extension (SNUPE), methylation profiling, and methylation microarray detection.

[0013] Furthermore, the product also includes reagents for processing samples.

[0014] Furthermore, the process may include the steps of extracting DNA and converting cytosine into uracil.

[0015] Furthermore, the reagents used in the step of converting cytosine to uracil are most commonly bisulfite reagents.

[0016] Furthermore, the bisulfite reagent includes a bisulfite buffer and a protective buffer.

[0017] Furthermore, the bisulfite is selected from one or more of sodium bisulfite, sodium sulfite, sodium bisulfite, ammonium bisulfite, and ammonium sulfite. DNA treated with the bisulfite reagent will have its unmethylated cytosine nucleotides converted to uracil, while methylated cytosine and other bases remain unchanged. Therefore, it is possible to distinguish between methylated and unmethylated cytosine in, for example, CpG dinucleotide sequences.

[0018] Furthermore, the extraction reagents for extracting DNA may include lysis buffer, binding buffer, washing buffer, and elution buffer.

[0019] Furthermore, the lysis buffer comprises a protein denaturant, a detergent, a pH buffer, and a nuclease inhibitor.

[0020] Furthermore, the binding buffer comprises a protein denaturant and a pH buffer.

[0021] Furthermore, the detergents include, but are not limited to, Tween20, IGEPEAL CA-630, Triton X-100, NP-40, and SDS.

[0022] Furthermore, the pH buffer includes one or more of Tris, boric acid, phosphate, MES, and HEPES.

[0023] Furthermore, the nuclease inhibitor includes one or more of EDTA, EGTA, and DEPC.

[0024] Furthermore, the sample is a sample containing white blood cells.

[0025] Furthermore, the samples containing white blood cells include blood, bone marrow aspiration fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph.

[0026] Furthermore, the samples were derived from individuals at high risk for colorectal cancer.

[0027] Furthermore, the high-risk groups for colorectal cancer include those with unhealthy lifestyles, middle-aged and elderly people, and those with a first-degree relative who has cancer.

[0028] In this invention, unhealthy lifestyle refers to lifestyles or habits that easily lead to various diseases in the human body. Currently, most people in society are in a state of sub-health, experiencing a decline in physical fitness, making them susceptible to illness, and even causing serious diseases such as cancer. Examples include lack of physical exercise, lack of regular checkups, sedentary lifestyles, insufficient sleep, smoking, excessive alcohol consumption, and irregular eating habits. Individuals with unhealthy lifestyles typically maintain these unhealthy habits for a long period and are considered a high-risk group for cancer.

[0029] In this invention, middle-aged and elderly people are defined as those aged 40 and above, according to the "Guidelines for Colorectal Cancer Screening in China".

[0030] In this invention, a family history of cancer in a first-degree relative refers to an individual whose first-degree relatives (parents, siblings, children) have cancer. This usually indicates an increased risk of hereditary cancer in the individual and that they belong to a high-risk group for cancer.

[0031] A third aspect of the present invention provides a colorectal cancer risk prediction product, the product comprising reagents for detecting the biomarkers described in the first aspect of the present invention.

[0032] In this invention, the product may comprise a solid substrate such as a chip, a glass slide, an array, etc., having reagents capable of detecting and / or quantifying one or more blood biomarkers or other sample-derived biomarkers immobilized at predetermined locations on the substrate. As an illustrative example, reagents immobilized at discrete predetermined locations may be provided to the chip for detecting and quantifying the concentration of any amount or any combination of biomarkers in a blood sample.

[0033] Furthermore, the reagents include those used in any one or more of the following methylation detection methods, which include: whole-genome bisulfite sequencing, pyrosequencing, bisulfite sequencing, methylation-specific polymerase chain reaction (PCR), bisulfite-specific PCR, methylation-sensitive restriction endonuclease-PCR / Southern assay, bisulfite-binding restriction endonuclease assay, digital polymerase chain reaction (PCR), restriction marker genome scanning, CpG island microarray, methylation mapping analysis, and methylation chip detection.

[0034] Furthermore, the product also includes reagents for processing samples.

[0035] Furthermore, the sample is a sample containing white blood cells.

[0036] Furthermore, the samples containing white blood cells include blood, bone marrow aspiration fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph.

[0037] Furthermore, the products include reagent kits, methylation panels, chips, systems, devices, and readable media.

[0038] The fourth aspect of the present invention provides a colorectal cancer risk prediction model, wherein the method for constructing the prediction model includes the following steps: obtaining methylation level data of the biomarkers described in the first aspect of the present invention in a sample, wherein the methylation level data of the biomarkers are from people with colorectal cancer and healthy people; and inputting the methylation level data into a machine learning algorithm to construct a prediction model.

[0039] Furthermore, the colorectal tumors include advanced adenomas or colorectal cancer.

[0040] In some embodiments, the methods for constructing the predictive model are known to those skilled in the art, and the steps of associating biomarker methylation levels with a certain probability or risk can be implemented and realized in different ways. Preferably, the biomarker methylation levels are mathematically associated with the fundamental question of whether or not one has colorectal cancer. The determination of biomarker methylation levels can be combined with other clinical characteristics using any suitable existing mathematical methods, and a predictive model can be constructed using machine learning algorithms. These clinical characteristics include age, sex, smoking, and alcohol consumption, etc.

[0041] In some embodiments, the method for constructing the prediction model includes: obtaining the methylation levels of biomarkers and corresponding clinical characteristics of training set samples, including whether the sample has colorectal cancer, age, gender, smoking, and alcohol consumption; constructing a model based on the methylation levels of the biomarkers and the clinical characteristics; obtaining the constructed prediction model using machine learning; and validating the performance of the constructed prediction model. The machine learning includes algorithmic models developed using various development tools; these tools include one or more of TensorFlow, Scikit-Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, Vertex AI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, and Spell. The algorithm model includes one or more of the following: convolutional neural network, autoencoder, deep belief network, linear regression, logistic regression, Lasso regression, Ridge regression, linear discriminant analysis, K-nearest neighbor algorithm, decision tree, perceptron, support vector machine, ensemble learning, correlation analysis, Naive Bayes, AdaBoost, GBDT, XGBoost, LightGBM, CatBoost, or random forest.

[0042] Furthermore, the prediction model obtains the methylation score based on the following formula:

[0043] .

[0044] If the methylation score of the sample being tested is higher than the threshold, the sample is at high risk of having advanced adenoma or colorectal cancer; if the methylation score of the sample being tested is lower than the threshold, the sample is at low risk of not having colorectal cancer or having colorectal cancer.

[0045] In this invention, the term "threshold" refers to a value that is statistically relevant to a particular outcome when compared with the analysis results. In some embodiments, the threshold is determined based on statistical conclusions from analyses of methylation levels of biomarkers in patients or high-risk populations of colorectal cancer and healthy controls. Some such studies are shown in the Examples section of this document, but studies from the literature and the experience of users of the methods described herein can also be used to generate or adjust thresholds.

[0046] The fifth aspect of the present invention provides a colorectal tumor risk prediction system, the system comprising the following modules:

[0047] The data acquisition module is used to acquire methylation level data of the biomarkers of the first aspect of the present invention in the sample to be tested.

[0048] The data analysis module is used to classify and predict the data in the data acquisition module using the prediction model described in the fourth aspect of the present invention, and to obtain the classification result of the test sample as high risk or low risk of colorectal cancer.

[0049] The output module is used to output the classification results.

[0050] The system may be a user's electronic device or a computer system remotely located relative to that electronic device.

[0051] A sixth aspect of the present invention provides a colorectal tumor risk prediction device, the computer device comprising:

[0052] A processor, suitable for implementing various instructions;

[0053] And, a memory adapted to store multiple instructions, said instructions adapted to be loaded by a processor and executed in accordance with the following steps:

[0054] Data is acquired, specifically the methylation level data of the biomarkers described in the first aspect of this invention for the sample to be tested.

[0055] The data is processed by inputting the methylation level data into the prediction model described in the fourth aspect of the present invention for classification prediction, so as to obtain the classification result of the test sample as high risk of colorectal cancer or low risk of colorectal cancer.

[0056] Output the prediction results.

[0057] The device can be a mobile electronic device.

[0058] The processor can also be called a Central Processing Unit (CPU). A processor may be an integrated circuit chip with signal processing capabilities. A processor can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.

[0059] A seventh aspect of the present invention provides a computer-readable medium for colorectal cancer risk prediction, the computer-readable medium storing a plurality of instructions adapted to be loaded by a processor and executed in the following steps:

[0060] Data is acquired, specifically the methylation level data of the biomarkers described in the first aspect of this invention for the sample to be tested.

[0061] The data is processed by inputting the methylation level data into the prediction model described in the fourth aspect of the present invention for classification prediction, so as to obtain the classification result of the test sample as high risk of colorectal cancer or low risk of colorectal cancer.

[0062] Output the prediction results.

[0063] The storage medium of this invention stores program instructions capable of implementing the aforementioned computer-based colorectal cancer risk prediction. These program instructions can be stored in the storage medium as a software product and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.

[0064] It should be understood that the system, device, and readable medium described in this invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings, direct couplings, or communication connections may be indirect couplings or communication connections between devices or units through some interfaces, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0065] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0066] Advantages and benefits of this invention: This invention provides five methylation biomarkers with the most stable predictive performance for colorectal cancer risk, and based on these, a colorectal cancer risk prediction model is constructed. This model demonstrates superior accuracy in predicting colorectal cancer risk compared to traditional lifestyle-based risk prediction models. Furthermore, combining methylation biomarkers with lifestyle risk factors further enhances predictive performance, demonstrating high potential for clinical translation and providing a new risk stratification tool for the prevention and early screening of colorectal cancer. Attached Figure Description

[0067] Figure 1 The results of developing candidate DNA methylation biomarkers specific to colorectal tumors are shown in the figure. The A, B, and C values ​​are chr4: 4859985-4860551, chr16: 55866678-55866757, chr7: 99818709-99818869, chr3: 3170109-3170139, and chr9: 139582466-139582662. The left figure shows the methylation levels of five candidate DMRs in CRC, advanced adenoma, and healthy controls, quantified by targeted bisulfite sequencing (TBS). The right figure shows the ROC curves, which demonstrate the classification performance of the five candidate DMRs and are used to distinguish CRC and advanced adenoma from healthy controls using TBS.DMR (differential methylation regions).

[0068] Figure 2 To validate the results of the selected candidate DNA methylation biomarkers for colorectal tumors, AC represents chr4: 4859985-4860551, chr16: 55866678-55866757, and chr7: 99818709-99818869, respectively. The left figure shows the differences in gene expression levels of the three candidate DMRs in CRC, advanced adenoma, and healthy controls. The right figure shows the ROC curves, demonstrating the performance of the gene expression differences of the three candidate regions in distinguishing between CRC and healthy controls, and between advanced adenoma and healthy controls.

[0069] Figure 3 To validate the results of the selected candidate DNA methylation biomarkers for colorectal tumors, A and B represent chr3: 3170109-3170139 and chr9: 139582466-139582662, respectively. The left figure shows the gene expression differences of the two candidate DMRs in CRC, advanced adenoma, and healthy control groups, while the right figure shows the ROC curves demonstrating the performance of the two candidate regions in distinguishing between CRC and healthy control groups, and between advanced adenoma and healthy control groups.

[0070] Figure 4 To validate the ROC curves of the concentrated methylation model, lifestyle model, and combination model. Detailed Implementation

[0071] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0072] Example

[0073] I. Experimental Methods

[0074] 1. Participants: This study employed a case-control approach, recruiting 325 participants from population-based screening and clinical opportunistic screening cohorts. The screening sample was part of the TARGET-C study (ChiCTR1800015506), a population-based randomized controlled trial of CRC screening conducted from May 2018 to May 2021 in six cities (Taizhou, Lanxi, Changsha, Hefei, Xuzhou, and Kunming) across five provinces (Zhejiang, Hunan, Anhui, Jiangsu, and Yunnan) in China. Permanent residents aged 50–74 years who had resided in the region for more than three years were recruited by trained staff. Individuals with a history of cancer, colorectal surgery, colorectal examination within the past five years, fecal occult blood test within the past year, or a history of serious illness that would cause discomfort during colorectal screening were excluded. Clinical samples were collected from two sources: (1) patients aged 25–84 years who underwent colonoscopy at a hospital endoscopy center, collected by a research team from the School of Public Health, Shenzhen University; and (2) patients aged 45–84 years who underwent colonoscopy in a clinical setting, collected by a research team from Wuhan University. Blood samples were collected from all participants prior to the colonoscopy. The study was approved by the ethics committees of the National Cancer Center / Cancer Hospital, the Chinese Academy of Medical Sciences, Peking Union Medical College, and the respective institutional ethics committees of the participating centers. All participants provided written informed consent, confirming their understanding and capacity to sign the study documents.

[0075] 2. Study Design: In the TARGET-C trial, targeted bisulfite sequencing (TBS) was performed on clinical samples from 48 CRC patients, 50 patients with advanced adenomas, and 49 healthy controls to determine the patterns of DNA methylation specific to advanced adenomas and CRC in leukocytes.

[0076] In addition to the above sample set, whole-genome Illumina 935K microarrays were performed on 56 patients with advanced adenomas and 50 healthy controls. RRBS was also performed on another 22 CRC patients, 20 patients with advanced adenomas, and 30 healthy controls as an independent validation set for risk prediction efficacy.

[0077] All participants were diagnosed via colonoscopy and pathological examination. Differentially methylated sites (DMPs) and regions (DMRs) were identified in each group. Based on the analysis of inter-group differences in the TBS validation set results, DMPs or DMRs that demonstrated good discriminatory ability against CRC and advanced adenomas were identified as candidate biomarkers for constructing a DNA methylation-based colorectal cancer risk prediction model. All candidate biomarkers were included as covariates in the model using TBS data. The model was trained 1000 times using elastic network regression, LASSO regression, and backward stepwise logistic regression. Variables selected from at least 700 of these models were retained to construct the final DNA methylation biomarker panel. Logistic regression was applied to test the model, and DNA methylation scores were calculated. The sensitivity and specificity of colorectal cancer risk prediction models based on DNA methylation scores, lifestyle scores, and their combinations were compared. Receiver operating characteristic (ROC) curves were plotted, and the area under the curve (AUC) was calculated to evaluate model performance. To ensure accuracy, batch effects and cellular heterogeneity were corrected during the preprocessing of the raw Illumina 935K microarray sequencing data, and confounding factors such as sex and age were subsequently adjusted in the differential analysis.

[0078] 3. Blood Processing and Leukocyte Extraction: The morning before the colonoscopy, 5 ml of whole blood was collected from fasting participants in K2-EDTA anticoagulant tubes. The tubes were gently inverted 8-10 times to mix the blood with the anticoagulant, and centrifuged within 4 hours of venipuncture. The collected sample was centrifuged for 12 minutes at 1200 g (±100 g) using a non-cooled horizontal rotor centrifuge. After aspirating the supernatant plasma, the leukocyte layer was extracted, transferred to cryovials, and stored at -80°C. The samples were then transported to the Chinese Academy of Medical Sciences and Peking Union Medical College Hospital for further processing via dry ice or cold chain.

[0079] 4. Detection technology

[0080] 1) Targeted Bisulfite Sequencing (TBS): First, bisulfite polymerase chain reaction (PCR) primers are designed and synthesized for the target region or site. Qualified sample DNA is treated with bisulfite (EZ-DNA Methylation Gold Kit; ZymoResearch). After treatment, unmethylated C becomes U (and T after PCR amplification), while methylated C remains unchanged. The bisulfite-treated template is amplified using a high-fidelity U-resistant DNA polymerase via bisulfite-specific polymerase chain reaction (BSP). BSP amplification products from the same sample are mixed and amplified using primers labeled with Illumina sequencing connectors. Sequencing libraries with different tags are constructed for each sample. The libraries are purified, quantified, mixed, inspected, and sequenced.

[0081] 2) Illumina 935K Microarray: The Illumina 935K microarray (Infinium MethylationEPICv2.0 BeadChip) platform provides high-throughput whole-genome DNA methylation analysis, covering approximately 935,000 CpG sites in the human genome. Genomic DNA is sulfated using the Zymo EZ DNA methylation kit, converting unmethylated cytosine to uracil while retaining methylated cytosine. This modified DNA is amplified, fragmented, and hybridized with the microarray, with fluorescent probes targeting both methylated and unmethylated sites. A scanner measures fluorescence intensity, and the data is imported into GenomeStudio for analysis to determine the methylation level at each site. β values ​​are used to represent DNA methylation levels, ranging from 0 to 1. Microarray detection is based on the principle of randomly distributing samples from different groups in equal proportions on the same chip. This method ensures optimal chip utilization and minimizes detection bias caused by experimental conditions.

[0082] 3) Reducing Representative Bisulfite Sequencing (RRBS): The quality of the raw sequencing data was assessed using FastQC (v0.11.7), and preprocessing was performed using Trimmermatic (v0.36) software. Preprocessing included pruning low-quality reads using a sliding window method (four bases per window, average quality score <15), deleting reads with a quality score <3 or fuzzy bases, pruning adapter sequences, discarding reads smaller than 36 nt, and excluding unpaired reads. Clean reads were aligned to a reference genome using BSMAP (v2.7.3). Methylation levels were calculated at each site, and the accuracy of methylation detection was validated by assessing bisulfite conversion efficiency and enzymatic digestion efficiency. Data with bisulfite conversion efficiency ≥99% and enzymatic digestion efficiency ≥95% were retained.

[0083] 5. Development and Evaluation of the DNA Methylation Scoring Model: Intergroup differential analysis of TBS data identified colorectal tumor-specific methylation sites / regions. Raw TBS data of these differentially expressed sites / regions were extracted from all samples, and a resampling method was used to generate 1000 subsets. In each subset, the differentially expressed sites / regions were used as independent variables, and models were trained using elastic network, LASSO, and stepwise logistic regression. For elastic network and LASSO regression, the frequency of each variable with a non-zero coefficient was calculated in 1000 resampling iterations, while for stepwise logistic regression, the frequency of each variable included in the model was recorded. Variables with a frequency ≥700 in all three methods were retained to construct the DNA methylation score. The formula for constructing the methylation scoring model is as follows:

[0084] This represents the intercept, which is the baseline score when all variables are zero. It reflects the model's baseline predictive power in the absence of any methylation biomarker contribution.

[0085] The range of i is from 1 to n, representing the n independent variables (i.e. DNA methylation biomarkers) included in the model.

[0086] These are coefficients in a logistic regression model, used to quantify the magnitude and direction of the impact of the i-th methylation biomarker on outcomes (such as disease risk).

[0087] This represents the methylation level of the i-th DNA methylation biomarker in a given sample.

[0088] The model performance was evaluated on the development set and further evaluated by splitting the subset into a training set and a test set for 1000 guided resampling iterations to calculate the average AUC value, ensuring robustness and generalizability.

[0089] 6. Development and Evaluation of the Lifestyle Scoring Model: In this study, the lifestyle score was constructed based on a logistic regression model, which primarily referenced the APCS scoring method (Present Risk Stratification Score for Colorectal Cancer). Participant grouping (colorectal cancer or healthy control) was used as the dependent variable (Y), and lifestyle risk factors associated with colorectal cancer (sex, age, smoking, and alcohol consumption) were used as independent variables (X). The regression coefficient (β) for each variable was calculated using logistic regression analysis. The specific logistic regression model is shown below:

[0090] .

[0091] P represents the probability that a participant has colorectal cancer.

[0092] The value is 1 for males and 2 for females.

[0093] : Continuous variable, with the unit of measurement being years.

[0094] The value is 1 for current or past smokers and 0 for non-smokers.

[0095] The value is 1 for current or past drinkers and 0 for non-drinkers.

[0096] The validation method for this model is the same as that for the DNA methylation scoring model mentioned earlier.

[0097] 7. Data Collection, Processing, and Quality Control: Epidemiological survey data of participants in the TARGET-C study were primarily collected through a self-developed web-based data management platform, recorded by trained staff, and monitored by a data monitoring committee. To ensure consistent standards of pathological diagnosis across multiple centers, experienced endoscopists and pathologists from the National Cancer Center centrally reviewed all colonoscopy results and pathology slides following the annual screening program. Basic patient data was collected from hospital medical records using standardized forms.

[0098] 8. Statistical Analysis: Statistical analysis was performed using R 4.3.1 software. The Wilcoxon rank-sum test was used to compare differences between groups. A p-value < 0.05 was considered statistically significant. 935K microarray data were analyzed using the ChAMP package (v2.29.1) and the ChAMPdata package (v2.31.1). Quantile normalization was performed using the `champ.norm` function, while the `champ.refbase` function was used to estimate cell type proportions. Batch effect correction was performed using the `champ.runCombat` function. For sample and probe quality control, the `champ.filter` function was used for quality assessment and visualization to remove low-quality probes and samples. Additional corrections for hematologic heterogeneity and batch effect ensured accurate results. The difference between the case group and the healthy control group was calculated as Δβ = β case -β healthy control ROC curves and AUC were used to evaluate model performance. Sensitivity, specificity, Youden index (Youden index = sensitivity + specificity - 1), and optimal critical value were calculated to assess its discriminative ability and determine the threshold that maximizes both.

[0099] II. Experimental Results

[0100] 1. Study Population Characteristics: Overall, this study included 325 eligible leukocyte DNA samples from individuals who underwent colonoscopy. The development set consisted of 48 CRC patients, 50 patients with advanced adenomas, and 49 healthy controls. The mean age (SD) was 62.4 (8.8), 63.4 (5.7), and 62.1 (9.3) years, respectively. There were no significant differences between the two groups in terms of age (P=0.238), sex (P=0.335), smoking behavior (P=0.243), or alcohol consumption (P=0.121). The validation set consisted of two subsets (Table 1): Subset I included 56 patients with advanced adenomas and 50 healthy controls confirmed by colonoscopy. The mean age (SD) of patients with advanced adenomas was 60.6 (6.7) years, compared to 60.2 (6.5) years in the healthy controls, with no statistically significant difference (P=0.598). There were no significant differences between the two groups in terms of sex (P=0.999), smoking behavior (P=0.811), alcohol consumption (P=0.999), or family history of cancer in first-degree relatives (P=0.262). Subgroup II included 22 CRC patients, 20 patients with advanced adenomas, and 30 healthy controls. The mean ages (SD) were 60.5 (13.1), 50.7 (11.37), and 57.7 (7.4) years, respectively. There were no statistically significant differences between the groups in terms of age (P=0.051), sex (P=0.762), smoking behavior (P=0.089), BMI (P=0.458), or family history of cancer in first-degree relatives (P=0.510).

[0101] Table 1. Basic demographic and clinical characteristics of participants in the validation session

[0102]

[0103] 2. Screening for concentrated DNA methylation biomarkers

[0104] Through a literature review, we identified a multicenter study on DNA methylation in peripheral blood mononuclear cells conducted in a Chinese population. This study, conducted from May 2020 to June 2022, included 1068 individuals aged 23–82 years who underwent colonoscopy at 10 hospitals.

[0105] In the development set, 200 primer pairs were initially designed to amplify regions of candidate DNA methylation markers. However, due to limitations in primer specificity and amplification efficiency, only 177 primer pairs targeting 73 DMPs and 104 DMRs were successfully designed for TBS analysis. Methylation data from CRC, advanced adenomas, and healthy controls were analyzed using the Wilcoxon rank-sum test. Compared to the healthy control group, CRC and advanced adenomas showed statistically significant differences in 3 DMPs (cg05107650, cg27077475, cg00777011) and 11 DMRs (P<0.05).

[0106] 3. Construction and Validation of the Methylation Scoring Model: In 1000 guided training iterations, using elastic network regression, LASSO regression, and logistic regression, five core methylation biomarkers were identified: chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757. Each biomarker was selected more than 700 times by the regression model.

[0107] Gene annotation revealed that no genes were annotated for chr9:139582466-139582662. Other biomarkers included the promoter region of the MSX1 gene (chr4:4859985-4860551), the STAG3 and PVRIG genes (chr7:99818709-99818869), the CES1 gene (chr16:55866678-55866757), and the TRNT1 gene (chr3:3170109-3170139). Hypermethylation of the STAG3, PVRIG, and CES1 genes, and hypomethylation of the MSX1 gene promoter region and the TRNT1 gene may be associated with an increased risk of colorectal cancer.

[0108] For CRC identification, the AUC value of a single biomarker ranges from 0.63 to 0.86 ( Figure 1Among them, the hypomethylated region of chr9:139582466-139582662 showed the highest sensitivity (77.1%) and the highest specificity (91.8%). The hypermethylated region of chr7:99818709-99818869 showed the highest sensitivity (72.9%), while the hypermethylated region of chr4:4859985-4860551 showed the highest specificity (91.8%). For advanced adenomas, the AUC values ​​of individual biomarkers ranged from 0.64 to 0.75. The hypomethylated region of chr3:3170109-3170139 showed the highest sensitivity (90.0%), but its specificity was relatively low (53.1%). The hypomethylated region of chr9:139582466-139582662 showed the highest specificity (73.5%). The hypermethylated region of chr7:99818709-99818869 showed the highest sensitivity (76.0%), while the hypermethylated region of chr16:55866678-55866757 showed the highest specificity (79.6%). In summary, the five methylation biomarkers included in the model demonstrated significant advantages in the detection of colorectal tumors, exhibiting complementary performance in terms of sensitivity and specificity.

[0109] Ultimately, a logistic regression-based scoring model was developed, incorporating five DNA methylation biomarkers:

[0110] f (甲基化评分) =1.025–0.765×β chr4:4859985-4860551 –0.938×β chr3:3170109-3170139 +0.513×β chr7:99818709-99818869 –1.315×β chr9:139582466-139582662 + 0.660×β chr16:55866678-55866757 .

[0111] The risk prediction efficacy of five DMR and multi-marker prediction models was independently validated using an Illumina 935K array and bisulfite sequencing (RRBS). ROC curves were used to evaluate the model's risk prediction performance for colorectal tumors in the validation set. Figure 2-3 For CRC, the model's AUC was 0.881 (95% CI: 0.811–0.950), with a sensitivity of 83.3% and a specificity of 81.6%. For advanced adenomas, the AUC was 0.880 (95% CI: 0.822–0.939), with a sensitivity of 88.0% and a specificity of 87.8%. Its predictive performance for overall colorectal tumors reached an AUC of 0.849 (95% CI: 0.740–0.939). Figure 4 This indicates that it has high accuracy and reliability in distinguishing between colorectal cancer patients and healthy individuals.

[0112] In addition, a lifestyle scoring model based on age, gender, smoking, and drinking was developed, as well as a combined model integrating methylation and lifestyle scores, to compare the discriminative power of the three models. Figure 4 In the uncalibrated analysis, the methylation scoring model demonstrated significantly better discriminative power on the validation set, with an AUC of 0.880, compared to 0.620 for the lifestyle scoring model, highlighting the potential biological relevance of leukocyte DNA methylation biomarkers in CRC detection. The combined model showed the highest AUC (0.887) and improved specificity (82.7%), indicating superior overall performance. However, the calibrated analysis showed that the methylation scoring model maintained the highest predictive performance on the validation set, with a narrower confidence interval than the combined model (AUC = 0.849, 95% CI: 0.741–0.939 vs. AUC = 0.824, 95% CI = 0.716–0.918). Figure 4 ).

[0113] These results demonstrate that the methylation scoring model possesses robust and independent risk prediction capabilities for colorectal cancer. Furthermore, this combined model offers additional benefits by capturing multidimensional risk factors associated with colorectal cancer, highlighting its potential as an innovative and comprehensive screening tool.

[0114] 4. Functional enrichment analysis based on the development of prominent DNA methylation markers: Gene Ontology (GO) enrichment analysis revealed that the identified important DMRs and DMPs were primarily enriched in biological processes related to development and morphogenesis, including embryonic organogenesis, skeletal morphogenesis, limb morphogenesis, and bone development. At the molecular functional level, significantly enhanced DNA-binding transcription activator activity was observed, particularly in functions related to RNA polymerase II-specific transcriptional activation. Enrichment analysis showed that genes corresponding to prominent DNA methylation markers were enriched in several cancer-related pathways, such as those in cancer, and pathways associated with intestinal inflammatory diseases, such as Shigella disease. Furthermore, certain genes (such as FGFR2 and NECTIN1) were enriched in the adhesion linkage pathway. These findings suggest that DNA methylation markers play a crucial role in the occurrence and development of colorectal cancer by influencing developmental regulation, transcriptional activation, and cancer-related signaling pathways.

[0115] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A biomarker for predicting the risk of colorectal cancer, characterized in that, The biomarker is a methylation biomarker, which is a combination of methylation regions chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757; The colorectal tumor is colorectal cancer and / or advanced adenoma.

2. The application of a reagent for detecting the methylation level of biomarkers in a sample in the preparation of colorectal cancer risk prediction products, characterized in that, The biomarker is the biomarker according to claim 1; the colorectal tumor is colorectal cancer and / or advanced adenoma.

3. The application as described in claim 2, characterized in that, The product is a reagent kit, methylation panel, or chip.

4. A colorectal cancer risk prediction product, characterized in that, The product contains a reagent for detecting the methylation level of the biomarker as described in claim 1; The colorectal tumor is colorectal cancer and / or advanced adenoma.

5. The product according to claim 4, characterized in that, The product also includes reagents for processing samples.

6. The product according to claim 5, characterized in that, The sample is a sample containing white blood cells.

7. The product according to claim 6, characterized in that, The samples containing white blood cells include blood, bone marrow aspiration fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, amniotic fluid, or lymph.

8. The product according to claim 4, characterized in that, The products include reagent kits, methylation panels, or chips.

9. A colorectal tumor risk prediction model, characterized in that, The method for constructing the prediction model includes the following steps: obtaining methylation level data of the biomarker of claim 1 from the sample, wherein the methylation level data of the biomarker comes from people with colorectal cancer and healthy people; and inputting the methylation level data into a machine learning algorithm to construct a prediction model. The prediction model obtains the methylation score based on the following formula: f (甲基化评分) =1.025–0.765×β chr4:4859985-4860551 –0.938×β chr3:3170109-3170139 +0.513×β chr7:99818709-99818869 –1.315×β chr9:139582466-139582662 +0.660×β chr16:55866678-55866757 ; The colorectal tumors include advanced adenomas and / or colorectal cancer; If the methylation score of the sample being tested is higher than the threshold, the patient from whom the sample was obtained has a high risk of having advanced adenoma or colorectal cancer; if the methylation score of the sample being tested is lower than the threshold, the patient from whom the sample was obtained has a low risk of having colorectal cancer.

10. A colorectal tumor risk prediction system, characterized in that, The system includes the following modules: A data acquisition module is used to acquire methylation level data of the biomarker of claim 1 in the sample to be tested; The data analysis module is used to classify and predict the data in the data acquisition module using the prediction model described in claim 9, and obtain the classification result of the test sample as high risk of colorectal cancer or low risk of colorectal cancer. The output module is used to output the classification results.

11. A computer device for predicting the risk of colorectal cancer, comprising a memory, a processor, and instructions stored in the memory, characterized in that, The processor executes the instructions to perform the following steps: Acquire data, specifically the methylation level data of the biomarker described in claim 1 of the sample to be tested; The data is processed by inputting the methylation level data into the prediction model of claim 9 for classification prediction, and obtaining the classification result of the test sample as high risk of colorectal cancer or low risk of colorectal cancer. Output the prediction results.

12. A computer-readable storage medium for colorectal tumor risk prediction, having instructions stored thereon, characterized in that, When the instruction is executed by the processor, it performs the following steps: Acquire data, specifically the methylation level data of the biomarker described in claim 1 of the sample to be tested; The data is processed by inputting the methylation level data into the prediction model of claim 9 for classification prediction, and obtaining the classification result of the test sample as high risk of colorectal cancer or low risk of colorectal cancer. Output the prediction results.