Application of leukocyte DNA methylation marker in colorectal tumor risk prediction

Through targeted bisulfite sequencing, screening the differential methylation position of leukocytes with differential methylation positions, and combining machine learning to build a multi-label prediction model, solving the problem of insufficient sensitivity and specificity in early diagnosis of colorectal cancer, achieving higher risk prediction accuracy and early screening effect.

CN120485365AActive Publication Date: 2025-08-15CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510567462.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The lack of high sensitivity and specificity detection methods in the early diagnosis of colorectal cancer results in high advanced diagnosis rates, missed treatment opportunities, and traditional risk assessments rely on lifestyle factors and lack of biomarker support.

Method used

Through targeted bisulfite sequencing (TBS), screening of differential methylation positions (DMP) and regions (DMR) of leukocytes, combining machine learning to build a multi-marker prediction model, combining methylated biomarkers and lifestyle scores, to improve the accuracy of colorectal tumor risk prediction.

Benefits of technology

The constructed predictive model demonstrates excellent identification performance in colorectal tumor risk prediction, improves the accuracy of early screening and clinical translation potential, and provides new risk stratification tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120485365A_ABST
    Figure CN120485365A_ABST
Patent Text Reader

Abstract

The invention discloses an application of leukocyte DNA methylation markers in colorectal tumor risk prediction, a series of leukocyte DNA methylation biomarkers related to colorectal tumors are researched, screened and verified, finally five methylation areas are obtained, a novel risk prediction model is constructed based on the five methylation areas, and the leukocyte DNA methylation markers have high clinical transformation potential and can be used for predicting the risk of colorectal tumors. And a novel risk stratification tool is provided for prevention and early screening of colorectal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedicine, and in particular to the application of leukocyte DNA methylation markers in predicting the risk of colorectal tumors. Background Art

[0002] Colorectal cancer (CRC) is one of the most common malignant tumors in China. The combined treatment of surgical resection and adjuvant chemotherapy has become the standard treatment and has achieved significant improvements in the overall prognosis of colorectal cancer. However, a series of epigenetic and metabolic changes give colorectal cancer cells a high degree of migration and invasion, resulting in a 5-year survival rate of approximately 12% for metastatic colorectal cancer. Approximately 30% of colorectal cancer patients have distant metastases at the time of diagnosis, and up to 30% of patients die from tumor metastasis and recurrence. Despite advances in treatment, the 5-year survival rate of colorectal cancer still relies heavily on early diagnosis. However, early CRC often lacks obvious symptoms, leading to late diagnosis and missed treatment opportunities. Therefore, strengthening early detection and risk assessment is crucial to reducing the burden of the disease.

[0003] In view of this, the present invention aims to develop a scoring model with high sensitivity and specificity for identifying CRC and precancerous lesions. By integrating lifestyle factors, this study aims to establish a comprehensive risk assessment model to provide a scientific basis for personalized screening strategies. Summary of the Invention

[0004] In light of the shortcomings of existing technologies, this study screened and identified differentially methylated sites (DMPs) and regions (DMRs) of leukocyte DNA using targeted bisulfite sequencing (TBS) in samples from colorectal tumors, advanced adenomas, and healthy controls. Using machine learning, five stable DMRs were identified as methylation biomarkers and incorporated into a multi-marker prediction model. The risk prediction performance of the five DMRs and the multi-marker prediction model was independently validated using Illumina 935K arrays and RRBS. The model demonstrated excellent recognition performance in detecting colorectal tumors, and combining methylation biomarkers with lifestyle scores can further improve risk prediction accuracy.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The first aspect of the present invention provides a biomarker for predicting the risk of colorectal tumors, wherein the biomarker is a methylation biomarker, and the methylation biomarker includes any one of the following methylation regions: chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, chr16: 55866678-55866757.

[0007] Furthermore, the methylation biomarker is a combination of methylation regions chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757.

[0008] Furthermore, the methylation level of chr4: 4859985-4860551 is low in the patient, the methylation level of chr3: 3170109-3170139 is low in the patient, the methylation level of chr7: 99818709-99818869 is high in the patient, the methylation level of chr9: 139582466-139582662 is high in the patient, and the methylation level of chr16: 55866678-55866757 is high in the patient.

[0009] As used herein, the term "patient" or "subject" refers to any animal (e.g., mammal) suffering from or suspected of having colorectal cancer, including but not limited to humans, non-human primates, rodents, etc. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to a human subject. Most preferably, the subject is a human.

[0010] Furthermore, the patient suffers from advanced adenoma or is at high risk of colorectal cancer.

[0011] The second aspect of the present invention provides the use of a reagent for detecting a biomarker in a sample in the preparation of a colorectal tumor risk prediction product, wherein the biomarker is the biomarker described in the first aspect of the present invention.

[0012] Furthermore, the reagents include reagents used in any one or more of the following methylation detection methods, including: whole genome bisulfite sequencing (WGBS), pyrosequencing, bisulfite sequencing, methylation-specific polymerase chain reaction (methylation-specific PCR, MS-PCR), bisulfite-specific polymerase chain reaction, methylation-sensitive restriction endonuclease-PCR / Southern method, combined bisulfite restriction endonuclease method (Combined Bisulfite Restriction Analysis, COBRA), digital polymerase chain reaction, restriction landmark genome scanning, CpG island microarray, single nucleotide primer extension SNUPE, methylation profiling, and methylation chip detection.

[0013] Furthermore, the product also includes reagents for processing samples.

[0014] Furthermore, the treatment may include the steps of extracting DNA and converting cytosine into uracil.

[0015] Furthermore, the reagent used in the step of converting cytosine into uracil is most commonly a bisulfite reagent.

[0016] Furthermore, the bisulfite reagent includes a bisulfite buffer and a protective buffer.

[0017] Furthermore, the bisulfite is selected from one or more of sodium bisulfite, sodium sulfite, sodium bisulfite, ammonium bisulfite, and ammonium sulfite. After bisulfite treatment of DNA, unmethylated cytosine nucleotides are converted to uracil, while methylated cytosine and other bases remain unchanged. Thus, for example, methylated and unmethylated cytidines in CpG dinucleotide sequences can be distinguished.

[0018] Furthermore, the extraction reagents for extracting DNA may include a lysis buffer, a binding buffer, a washing buffer and an elution buffer.

[0019] Furthermore, the lysis buffer comprises a protein denaturant, a detergent, a pH buffer and a nuclease inhibitor.

[0020] Furthermore, the binding buffer comprises a protein denaturant and a pH buffer.

[0021] Furthermore, the detergent includes but is not limited to Tween 20, IGEPAL CA-630, Triton X-100, NP-40 and SDS.

[0022] Furthermore, the pH buffer comprises one or more of Tris, boric acid, phosphate, MES and HEPES.

[0023] Furthermore, the nuclease inhibitor includes one or more of EDTA, EGTA and DEPC.

[0024] Furthermore, the sample contains leukocytes.

[0025] Furthermore, the sample containing leukocytes includes blood, bone marrow puncture fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph.

[0026] Furthermore, the samples are from a population at high risk of colorectal cancer.

[0027] Furthermore, the high-risk groups for colorectal cancer include people with unhealthy lifestyles, middle-aged and elderly people, and those with a history of cancer in their first-degree relatives.

[0028] In this context, an unhealthy lifestyle refers to a lifestyle or lifestyle habits that can easily lead to numerous illnesses in the human body. Most people in today's society experience sub-health, a decline in physical fitness, and a greater susceptibility to illness, even leading to serious illnesses such as cancer. Examples include a lack of physical exercise, a lack of proactive medical checkups, prolonged sitting, insufficient sleep, smoking, excessive drinking, and irregular diets. People who maintain these unhealthy lifestyles often engage in these behaviors over a long period of time and are therefore at high risk for cancer.

[0029] In the present invention, middle-aged and elderly people are defined as people aged 40 years and above according to the age recommendation for colorectal cancer screening in the "Guidelines for Colorectal Cancer Screening in China".

[0030] In the present invention, a first-degree relative with a history of cancer refers to the presence of a first-degree relative (parents, siblings, children) of the sample source individual who has cancer, which usually indicates that the sample source individual has an increased risk of hereditary cancer and belongs to a high-risk group for cancer.

[0031] The third aspect of the present invention provides a colorectal tumor risk prediction product, which comprises a reagent for detecting the biomarker described in the first aspect of the present invention.

[0032] In the present invention, the product may comprise a solid substrate such as a chip, a slide, an array, etc., having reagents capable of detecting and / or quantifying one or more blood markers or other sample-derived markers immobilized at predetermined locations on the substrate. As an illustrative example, a chip may be provided with reagents immobilized at discrete predetermined locations for detecting and quantifying the concentration of any number or combination of biomarkers in a blood sample.

[0033] Furthermore, the reagents include reagents used in any one or more of the following methylation detection methods, the methylation detection methods including: whole-genome bisulfite sequencing, pyrophosphate sequencing, bisulfite sequencing, methylation-specific polymerase chain reaction, bisulfite-specific polymerase chain reaction, methylation-sensitive restriction endonuclease-PCR / Southern method, restriction endonuclease method combined with bisulfite, digital polymerase chain reaction, restriction landmark genome scanning, CpG island microarray, methylation profiling, and methylation chip detection.

[0034] Furthermore, the product also includes reagents for processing samples.

[0035] Furthermore, the sample contains leukocytes.

[0036] Furthermore, the sample containing leukocytes includes blood, bone marrow puncture fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph.

[0037] Furthermore, the product includes a kit, a methylation panel, a chip, a system, a device, and a readable medium.

[0038] The fourth aspect of the present invention provides a colorectal tumor risk prediction model, and the method for constructing the prediction model includes the following steps: obtaining the methylation level data of the biomarker described in the first aspect of the present invention in the sample, wherein the methylation level data of the biomarker comes from people with colorectal tumors and healthy people; inputting the methylation level data into a machine learning algorithm to construct a prediction model.

[0039] Furthermore, the colorectal tumor includes advanced adenoma or colorectal cancer.

[0040] In some embodiments, methods for constructing a predictive model are known to those skilled in the art, and the step of associating biomarker methylation levels with a certain likelihood or risk can be implemented and achieved in various ways. Preferably, the biomarker methylation levels are mathematically associated with the underlying question of whether or not a patient has colorectal cancer. The determination of biomarker methylation levels can be combined with other clinical characteristics using any suitable prior art mathematical method, and a predictive model constructed using a machine learning algorithm. Such clinical characteristics include age, gender, smoking, and alcohol consumption.

[0041] In some embodiments, the method for constructing the prediction model includes: obtaining the biomarker methylation levels of training set samples and the corresponding clinical characteristics of the samples, the clinical characteristics including whether the patient has colorectal cancer, age, gender, smoking, and alcohol consumption; constructing a model based on the biomarker methylation levels and the clinical characteristics, using machine learning to obtain the constructed prediction model, and verifying the effectiveness of the constructed prediction model. The machine learning includes algorithm models developed using various development tools; the development tools include one or more of TensorFlow, Scikit Learn, PyTorch, OpenNN, RapidMiner, Azure Machine Learning, Apache Mahout, Shogun, KNIME, Vertex AI, H2Oai, Anaconda, Keras, Tableau, Fast.ai, Catalyst, Amazon ML, MLJAR, and Spell. The algorithm model includes one or more of convolutional neural network, autoencoder, deep belief network, linear regression, logistic regression, Lasso regression, Ridge regression, linear discriminant analysis, K-nearest neighbor algorithm, decision tree, perceptron, support vector machine, ensemble learning, correlation analysis, naive Bayes, AdaBoost, GBDT, XGBoost, LightGBM, CatBoost or random forest.

[0042] Furthermore, the prediction model obtains the methylation score based on the following formula:

[0043]

[0044] If the methylation score of the sample to be tested is higher than the threshold, the sample to be tested has advanced adenoma or a high risk of colorectal cancer; if the methylation score of the sample to be tested is lower than the threshold, the sample to be tested does not have colorectal tumor or has a low risk of colorectal tumor.

[0045] In the present invention, the term "threshold" refers to a value that is statistically correlated with a particular outcome when compared to an analysis result. In some embodiments, the threshold is determined based on statistical conclusions from a population analysis of biomarker methylation levels in patients with colorectal cancer or high-risk populations and healthy controls. Some such studies are shown in the Examples section of this article, but studies from the literature and the experience of users of the methods described herein can also be used to generate or adjust the threshold.

[0046] A fifth aspect of the present invention provides a colorectal cancer risk prediction system, the system comprising the following modules:

[0047] The data acquisition module is used to obtain the methylation level data of the biomarker described in the first aspect of the present invention for the sample to be tested.

[0048] The data analysis module is used to classify and predict the data in the data acquisition module using the prediction model described in the fourth aspect of the present invention, and obtain a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer.

[0049] Output module, used to output classification results.

[0050] The system may be the user's electronic device or a computer system located remotely from the electronic device.

[0051] A sixth aspect of the present invention provides a colorectal tumor risk prediction device, wherein the computer device comprises:

[0052] a processor adapted to implement the instructions;

[0053] and a memory adapted to store a plurality of instructions adapted to be loaded by a processor and executed by the processor:

[0054] Acquire data, and acquire the methylation level data of the biomarker described in the first aspect of the present invention for the sample to be tested.

[0055] Process the data, input the methylation level data into the prediction model described in the fourth aspect of the present invention for classification prediction, and obtain a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer.

[0056] Output the prediction results.

[0057] The device may be a mobile electronic device.

[0058] The processor may also be referred to as a central processing unit (CPU). The processor may be an integrated circuit chip with signal processing capabilities. The processor may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.

[0059] A seventh aspect of the present invention provides a computer-readable medium for predicting colorectal tumor risk, wherein the computer-readable medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the following steps:

[0060] Acquire data, and acquire the methylation level data of the biomarker described in the first aspect of the present invention for the sample to be tested.

[0061] Process the data, input the methylation level data into the prediction model described in the fourth aspect of the present invention for classification prediction, and obtain a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer.

[0062] Output the prediction results.

[0063] The storage medium of the embodiment of the present invention stores program instructions that can implement the above-mentioned computer-implemented colorectal tumor risk prediction, wherein the program instructions can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or a computer device such as a computer, a server, a mobile phone, or a tablet.

[0064] It should be understood that the system, device and readable medium described in the present invention can be implemented in other ways. For example, the device embodiment described above is only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. On the other hand, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0065] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0066] Advantages and benefits of the present invention: This invention provides five methylation biomarkers with the most stable colorectal cancer risk prediction performance, based on which a colorectal cancer risk prediction model is constructed. This model has superior accuracy in predicting colorectal cancer risk compared to traditional lifestyle-based risk prediction models. The combination of methylation biomarkers and lifestyle risk factors further enhances predictive performance, demonstrating high potential for clinical translation and providing a new risk stratification tool for the prevention and early screening of colorectal cancer. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 Figure 1: Result diagram for developing a pooled set of colorectal tumor-specific candidate DNA methylation biomarkers, where AE are chr4: 4859985-4860551, chr16: 55866678-55866757, chr7: 99818709-99818869, chr3: 3170109-3170139, chr9: 139582466-139582662; the left figure shows the methylation levels of the five candidate DMRs in CRC, advanced adenomas and healthy controls, quantified by targeted bisulfite sequencing (TBS); the right figure shows the ROC curve showing the classification performance of the five candidate DMRs, used to distinguish CRC and advanced adenomas from healthy controls using TBS.DMRs (differentially methylated regions).

[0068] Figure 2 Figure 1: Validation of candidate DNA methylation biomarkers specific for colorectal tumors. Figures AE and AE are chr4: 4859985-4860551, chr16: 55866678-55866757, and chr7: 99818709-99818869, respectively. The left figure shows the methylation levels of the three candidate DMRs in CRC, advanced adenomas, and healthy controls. A is the Illumina 935K microarray measurement result, and BC is the RRBS measurement result. The right figure shows the ROC curve showing the classification performance of the three candidate DMRs. DMRs (differentially methylated regions) distinguish CRC from healthy controls, and advanced adenomas from healthy controls.

[0069] Figure 3 To verify the results of the concentrated colorectal tumor-specific candidate DNA methylation biomarkers, AB is chr3: 3170109-3170139, chr9: 139582466-139582662; the left figure shows the methylation levels of the three candidate DMRs in CRC, advanced adenoma and healthy control group, and the right figure is the ROC curve showing the classification performance of the three candidate DMRs. DMRs (differentially methylated regions) distinguish CRC from healthy control group, and advanced adenoma from healthy control group.

[0070] Figure 4 ROC curves of the methylation model, lifestyle model, and combined model in the validation set. DETAILED DESCRIPTION

[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0072] Example

[0073] 1. Experimental Methods

[0074] 1. Participants: This study used a case-control study method and recruited 325 participants from population-based screening and clinical opportunistic screening cohorts. The screening sample was part of the TARGET-C study (ChiCTR1800015506), a population-based randomized controlled trial of CRC screening conducted in six cities (Taizhou, Lanxi, Changsha, Hefei, Xuzhou and Kunming) in five provinces (Zhejiang, Hunan, Anhui, Jiangsu and Yunnan) in China from May 2018 to May 2021. This study recruited permanent residents aged 50-74 years who had lived in the area for more than 3 years and was recruited by trained staff. Individuals with a history of cancer, colorectal surgery, colorectal examination in the past 5 years, fecal occult blood test in the past year, or a history of serious illness that caused discomfort during colorectal screening were excluded. Clinical samples were collected from two sources: (1) patients aged 25–84 years who underwent colonoscopy at the hospital endoscopy center by a research team from the School of Public Health of Shenzhen University, and (2) patients aged 45–84 years who were collected in a clinical setting by a research team from Wuhan University. Blood samples were collected from all participants before colonoscopy. The ethics committees of the National Cancer Center / Cancer Hospital, the Chinese Academy of Medical Sciences, and the Peking Union Medical College, as well as the respective institutional ethics committees of the participating centers, approved the study. All participants provided written informed consent, confirming that they understood and were able to sign the study documents.

[0075] 2. Study design: In the TARGET-C trial, targeted bisulfite sequencing (TBS) was performed on clinical samples from 48 patients with CRC, 50 patients with advanced adenomas, and 49 healthy controls to determine the patterns of advanced adenoma-specific and CRC-specific DNA methylation in leukocytes.

[0076] In addition to the above sample sets, whole-genome Illumina 935K microarray was performed on 56 patients with advanced adenomas and 50 healthy controls, and RRBS was performed on another 22 CRC patients, 20 patients with advanced adenomas, and 30 healthy controls as independent validation sets for risk prediction performance.

[0077] All participants were diagnosed by colonoscopy and pathological examination. Differentially methylated positions (DMPs) and regions (DMRs) were identified in each group. Based on the analysis of intergroup differences in the TBS validation set results, DMPs or DMRs that showed good discrimination between CRC and advanced adenomas were identified as candidate biomarkers for constructing a DNA methylation-based colorectal cancer risk prediction model. All candidate markers were included in the model as covariates of TBS data. The model was trained 1000 times using elastic net regression, LASSO regression, and backward stepwise logistic regression. Variables selected from at least 700 of these models were retained to construct the final DNA methylation biomarker panel. Logistic regression was applied to test the model and calculate the DNA methylation score. The sensitivity and specificity of the colorectal cancer risk prediction model based on the DNA methylation score, lifestyle score, and their combination were compared. Receiver operating characteristic (ROC) curves were plotted, and the area under the curve (AUC) was calculated to evaluate model performance. To ensure accuracy, batch effects and cellular heterogeneity were corrected during the preprocessing of Illumina 935K microarray sequencing raw data, and confounding factors such as gender and age were subsequently adjusted in the differential analysis.

[0078] 3. Blood processing and leukocyte extraction: 5 ml of whole blood was collected from fasting participants in K2-EDTA anticoagulant tubes the morning before colonoscopy. The tube was gently inverted 8-10 times to mix the blood with the anticoagulant and centrifuged within 4 hours after venipuncture. The collected samples were centrifuged at 1200g (±100g) for 12 minutes using a non-cooled horizontal rotor centrifuge. After aspirating the upper plasma, the white blood cell layer was extracted, transferred to a freezing tube, stored at -80°C, and sent to the Chinese Academy of Medical Sciences and Peking Union Medical College Hospital for processing via dry ice or cold chain transportation.

[0079] 4. Detection technology

[0080] 1) Targeted bisulfite sequencing (TBS): First, bisulfite polymerase chain reaction primers are designed and synthesized for the target region or site. Qualified sample DNA is treated with bisulfite (EZ-DNA Methylation Gold Kit; ZymoResearch). After treatment, unmethylated C becomes U (which becomes T after polymerase chain reaction amplification), while methylated C remains unchanged. Bisulfite-specific polymerase chain reaction (BSP) amplification is performed on the bisulfite-treated template using a high-fidelity U-base resistant DNA polymerase. BSP amplification products from the same sample are mixed and amplified using primers labeled with Illumina sequencing connectors. Sequencing libraries with different tags are constructed for each sample. The libraries are purified, quantified, mixed, checked, and sequenced.

[0081] 2) Illumina 935K microarray: The Illumina 935K microarray (Infinium MethylationEPICv2.0 BeadChip) platform provides high-throughput whole-genome DNA methylation analysis, covering approximately 935,000 CpG sites in the human genome. The genomic DNA is sulfated using the Zymo EZ DNA Methylation Kit to convert unmethylated cytosine to uracil while retaining methylated cytosine. This modified DNA is amplified, fragmented, and hybridized to the microarray, with fluorescent probes targeting methylated and unmethylated sites. The scanner measures the fluorescence intensity, and the data is imported into GenomeStudio for analysis to determine the methylation level of each site. The β value is used to represent the DNA methylation level, ranging from 0 to 1. Microarray detection is based on the principle of randomly distributing samples from different groups on the same chip in equal proportions. This approach ensures optimal chip utilization and minimizes detection bias caused by experimental conditions.

[0082] 3) Reduced representation bisulfite sequencing (RRBS): The quality of raw sequencing data was assessed using FastQC (v0.11.7) and preprocessed using Trimmermatic (v0.36) software. Preprocessing included trimming low-quality reads using a sliding window method (four bases per window, average quality score

[0083] We performed a 3-D PCR amplification of the cleaned reads using the BSMAP (v2.7.3) software. We performed a 3-D PCR amplification of the cleaned reads using the BSMAP (v2.7.3) software. We calculated methylation levels at each site, and validated the accuracy of methylation detection by assessing bisulfite conversion efficiency and enzyme digestion efficiency. We retained data with bisulfite conversion efficiency ≥99% and enzyme digestion efficiency ≥95%.

[0084] 5. Development and evaluation of DNA methylation scoring model: Intergroup difference analysis of TBS data identified colorectal tumor-specific methylation sites / regions. The original TBS data of these differential sites / regions were extracted from all samples, and 1000 subsets were generated by the guided resampling method. In each subset, the differential sites / regions were used as independent variables, and the models were trained using elastic net, LASSO, and stepwise logistic regression. For elastic net and LASSO regression, the frequency of each variable with a non-zero coefficient in 1000 resampling iterations was calculated, while for stepwise logistic regression, the frequency of each variable included in the model was recorded. Variables with a frequency ≥ 700 in all three methods were retained to construct the DNA methylation score. The formula for constructing the methylation scoring model is as follows:

[0085] β0 represents the intercept, which is the baseline score when all variables take the value of 0. It reflects the baseline prediction level of the model in the absence of any contribution from methylation biomarkers.

[0086] i ranges from 1 to n, representing the n independent variables (i.e., DNA methylation biomarkers) included in the model.

[0087] β i is the coefficient in the logistic regression model, which quantifies the magnitude and direction of the effect of the ith methylation biomarker on the outcome (e.g., disease risk).

[0088] X i represents the methylation level of the i-th DNA methylation biomarker in a given sample.

[0089] Model performance was evaluated on the development set and further evaluated by partitioning the subset into training and test sets with 1000 bootstrap resampling iterations to calculate the average AUC value to ensure robustness and generalizability.

[0090] 6. Development and evaluation of lifestyle score model: In this study, the lifestyle score was constructed based on a logistic regression model, which was mainly constructed with reference to the APCS score method (established risk stratification score for colorectal cancer). Participant group (colorectal cancer or healthy control) was used as the dependent variable (Y), and lifestyle risk factors associated with colorectal cancer (sex, age, smoking and drinking) were used as independent variables (X). The regression coefficient (β) of each variable was calculated using logistic regression analysis. The specific logistic regression model is shown as follows:

[0091]

[0092] P represents the probability that a participant has colorectal cancer.

[0093] X gender: The value is 1 for male and 2 for female.

[0094] X age : Continuous variable, the unit of variable is year.

[0095] X smoking : For current or former smokers, it takes the value of 1, and for people who have never smoked, it takes the value of 0.

[0096] X alcoholdrinking : Takes a value of 1 for current or past drinkers and 0 for people who never drink.

[0097] The model was validated in the same way as the DNA methylation scoring model mentioned above.

[0098] 7. Data Collection, Processing, and Quality Control: Epidemiological survey data for the TARGET-C study were collected primarily through a self-developed web-based data management platform, recorded by trained staff, and monitored by a data monitoring committee. To ensure uniform standards for multicenter pathological diagnosis, experienced endoscopists and pathologists from the National Cancer Center centrally reviewed all colonoscopy results and pathology slides after the annual screening program. Basic patient data were collected from hospital medical records using standardized forms.

[0099] 8. Statistical analysis: Statistical analysis was performed using R 4.3.1 software. The Wilcoxon rank sum test was used to compare differences between groups. P values < 0.05 were considered statistically significant. The 935K chip data were analyzed using the ChAMP package (v2.29.1) and the ChAMPdata package (v2.31.1). Quantile normalization was performed using the champ.norm function, while cell type proportions were estimated using the champ.refbase function.

[0100] The champ.runCombat function performs batch effect correction. For sample and probe quality control, the champ.QC function is used for quality assessment and visualization, and the champ.filter function is used to remove low-quality probes and samples. Additional corrections for blood cell heterogeneity and batch effects ensure accurate results. The difference between the case group and the healthy control group is calculated as Δβ = β case -β healthy control The performance of the model was evaluated using the ROC curve and AUC. Sensitivity, specificity, Youden index (Youden index = sensitivity + specificity - 1), and the optimal cutoff value were calculated to assess its discrimination ability and determine the threshold that maximized both.

[0101] 2. Experimental Results

[0102] 1. Study Population Characteristics: Overall, 325 qualified leukocyte DNA samples from individuals undergoing colonoscopy were included in this study. The development set included 48 patients with CRC, 50 patients with advanced adenomas, and 49 healthy controls. The mean (SD) age was 62.4 (8.8), 63.4 (5.7), and 62.1 (9.3) years, respectively. There were no significant differences between the two groups with respect to age (P = 0.238), sex (P = 0.335), smoking behavior (P = 0.243), or alcohol consumption (P = 0.121). The validation set was divided into two subsets (Table 1): Subset I included 56 patients with advanced adenomas and 50 healthy controls confirmed by colonoscopy. The mean age (standard deviation, SD) of patients with advanced adenomas was 60.6 (6.7) years; that of healthy controls was 60.2 (6.5) years, a difference that was not statistically significant (P = 0.598). There were no significant differences between the two groups with respect to sex (P = 0.999), smoking behavior (P = 0.811), alcohol consumption (P = 0.999), or a family history of cancer in first-degree relatives (P = 0.262). Subgroup II included 22 patients with CRC, 20 patients with advanced adenomas, and 30 healthy controls. The mean (SD) age was 60.5 (13.1), 50.7 (11.37), and 57.7 (7.4) years, respectively. There were no statistically significant differences between the groups with respect to age (P = 0.051), sex (P = 0.762), smoking behavior (P = 0.089), body mass index (P = 0.458), or a family history of cancer in first-degree relatives (P = 0.510).

[0103] Table 1. Basic demographic and clinical characteristics of the participants in the validation set

[0104]

[0105]

[0106] 2. Discovery of focused DNA methylation biomarker screening

[0107] Through literature review, we identified a multicenter study on DNA methylation in peripheral blood mononuclear cells conducted in a Chinese population. This study was conducted from May 2020 to June 2022 and included 1068 people aged 23-82 years who underwent colonoscopy at 10 hospitals.

[0108] In the development set, 200 primer pairs were initially designed to amplify regions of candidate DNA methylation markers. However, due to limitations in primer specificity and amplification efficiency, only 177 primer pairs targeting 73 DMPs and 104 DMRs were successfully designed for TBS analysis. Methylation data from CRC, advanced adenomas, and healthy controls were analyzed using the Wilcoxon rank sum test. Compared with healthy controls, 3 DMPs (cg05107650, cg27077475, and cg00777011) and 11 DMRs in CRC and advanced adenomas showed statistically significant differences (P < 0.05).

[0109] 3. Construction and Validation of the Methylation Scoring Model: Using elastic net regression, LASSO regression, and logistic regression over 1,000 bootstrapped training iterations, we identified five core methylation biomarkers: chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757. Each marker was selected by the regression model over 700 times.

[0110] Gene annotation revealed that chr9:139582466-139582662 was not annotated to any gene. The remaining markers were located within the promoter region of the MSX1 gene (chr4:4859985-4860551), STAG3 gene, PVRIG gene (chr7:99818709-99818869), CES1 gene (chr16:55866678-55866757), and TRNT1 gene (chr3:3170109-3170139). Hypermethylation of the STAG3, PVRIG, and CES1 genes, and hypomethylation of the MSX1 promoter region and TRNT1 gene may be associated with an increased risk of colorectal cancer.

[0111] For CRC discrimination, the AUC values of individual biomarkers ranged from 0.63 to 0.86 ( Figure 1). Among them, the hypomethylated region of chr9: 139582466-139582662 had the highest sensitivity (77.1%) and the highest specificity (91.8%). The hypermethylated region of chr7: 99818709-99818869 had the highest sensitivity (72.9%), while the hypermethylated region of chr4: 4859985-4860551 had the highest specificity (91.8%). For advanced adenomas, the AUC values of single biomarkers ranged from 0.64 to 0.75. The hypomethylated region of chr3: 3170109-3170139 had the highest sensitivity (90.0%), but its specificity was relatively low (53.1%). The hypomethylated region of chr9: 139582466-139582662 showed the highest specificity (73.5%). The chr7:99818709-99818869 hypermethylation region had the highest sensitivity (76.0%), while the chr16:55866678-55866757 hypermethylation region had the highest specificity (79.6%). In summary, the five methylation biomarkers included in the model showed obvious advantages in detecting colorectal tumors and exhibited complementary performance in terms of sensitivity and specificity.

[0112] Finally, a scoring model based on logistic regression was developed, combining five DNA methylation biomarkers:

[0113] f (甲基化评分) =1.025–0.765×β chr4:4859985-4860551 –0.938×β chr3:3170109-3170139 +0.513×β chr7 : 99818709-99818869 –1.315×β chr9:139582466-139582662 +0.660×β chr16:55866678-55866757 .

[0114] The risk prediction performance of the five DMRs and the multi-marker prediction model was independently validated by Illumina 935K array and bisulfite sequencing (RRBS), and the ROC curve was used to evaluate the risk prediction performance of the model for colorectal cancer in the validation set ( Figure 2-3 For CRC, the model had an AUC of 0.881 (95% CI: 0.811-0.950), a sensitivity of 83.3%, and a specificity of 81.6%. For advanced adenomas, the AUC was 0.880 (95% CI: 0.822-0.939), a sensitivity of 88.0%, and a specificity of 87.8%. The predictive performance for overall colorectal tumors reached an AUC of 0.849 (95% CI: 0.740-0.939) ( Figure 4 ), showing high accuracy and reliability in distinguishing patients with colorectal tumors from healthy individuals.

[0115] In addition, a lifestyle score model based on age, sex, smoking and drinking, as well as a combined model integrating methylation and lifestyle scores were developed to compare the discriminatory abilities of the three models ( Figure 4 ). In the uncalibrated analysis, the methylation score model showed significantly better discrimination in the validation set with an AUC of 0.880, compared to 0.620 for the lifestyle score model, highlighting the potential biological relevance of leukocyte DNA methylation biomarkers in CRC detection. The combined model showed the highest AUC (0.887) and improved specificity (82.7%), indicating superior overall performance. However, the post-calibration analysis showed that the methylation score model maintained the highest predictive performance in the validation set, with a narrower confidence interval than the combined model (AUC = 0.849, 95% CI: 0.741-0.939 vs. AUC = 0.824, 95% CI = 0.716-0.918) ( Figure 4 ).

[0116] These results demonstrate that the methylation score model has robust, independent risk prediction capabilities for colorectal cancer. Furthermore, the combined model offers the additional benefit of capturing the multidimensional risk factors associated with colorectal cancer, highlighting its potential as an innovative and comprehensive screening tool.

[0117] 4. Functional enrichment analysis based on significant DNA methylation marks identified in the development set: Gene ontology (GO) enrichment analysis showed that the identified important DMRs and DMPs were mainly enriched in biological processes related to development and morphogenesis, including embryonic organ development, skeletal system morphogenesis, limb morphogenesis, and bone development. In terms of molecular function, a significant increase in the activity of DNA-binding transcription activators was observed, especially functions related to ribonucleic acid (RNA) polymerase II-specific transcriptional activation. Enrichment analysis showed that genes corresponding to significant DNA methylation marks were enriched in several cancer-related pathways, such as pathways in cancer, and pathways related to intestinal inflammatory diseases such as shigellosis. In addition, certain genes (such as FGFR2 and NECTIN1) were enriched in the adhesion junction pathway. These findings indicate that DNA methylation marks play a key role in the occurrence and development of colorectal tumors by affecting developmental regulation, transcriptional activation, and cancer-related signaling pathways.

[0118] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention can be implemented over a wide range under equivalent parameters, concentrations, and conditions without departing from the spirit and scope of the present invention and without the need for unnecessary experimentation. Although the present invention provides embodiments, it will be understood that further improvements can be made to the present invention. In short, according to the principles of the present invention, this application is intended to include any variations, uses, or improvements to the present invention, including changes made by conventional techniques known in the art that depart from the disclosed scope of this application.

Claims

1. A biomarker for predicting the risk of colorectal cancer, characterized in that: The biomarker is a methylation biomarker, and the methylation biomarker includes any one of the following methylation regions: chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, chr16: 55866678-55866757; Preferably, the methylation biomarker is a combination of methylation regions chr4: 4859985-4860551, chr3: 3170109-3170139, chr7: 99818709-99818869, chr9: 139582466-139582662, and chr16: 55866678-55866757.

2. The biomarker according to claim 1, wherein The methylation level of chr4: 4859985-4860551 is low in the patient, the methylation level of chr3: 3170109-3170139 is low in the patient, the methylation level of chr7: 99818709-99818869 is high in the patient, the methylation level of chr9: 139582466-139582662 is high in the patient, and the methylation level of chr16: 55866678-55866757 is high in the patient; Preferably, the patient suffers from advanced adenoma or is a high-risk group for colorectal cancer.

3. The use of a reagent for detecting biomarkers in a sample in the preparation of a colorectal tumor risk prediction product, characterized in that: The biomarker is the biomarker according to any one of claims 1 to 2.

4. The use according to claim 3, characterized in that The reagents include reagents used in any one or more of the following methylation detection methods, including: whole-genome bisulfite sequencing, pyrosequencing, bisulfite sequencing, methylation-specific polymerase chain reaction, bisulfite-specific polymerase chain reaction, methylation-sensitive restriction endonuclease-PCR / Southern method, bisulfite-coupled restriction endonuclease method, digital polymerase chain reaction, restriction landmark genome scanning, CpG island microarray, methylation profiling, and methylation chip detection; Preferably, the product further comprises reagents for processing the sample; Preferably, the sample is a sample containing leukocytes; Preferably, the sample containing leukocytes includes blood, bone marrow puncture fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph; Preferably, the sample is a sample from a population at high risk of colorectal cancer; Preferably, the high-risk population for colorectal cancer includes people with unhealthy lifestyles, middle-aged and elderly people, and those with a history of cancer in their first-degree relatives.

5. A colorectal tumor risk prediction product, characterized in that: The product comprises a reagent for detecting the biomarker according to any one of claims 1-2; Preferably, the reagents include reagents used in any one or more of the following methylation detection methods, the methylation detection methods including: whole-genome bisulfite sequencing, pyrosequencing, bisulfite sequencing, methylation-specific polymerase chain reaction, bisulfite-specific polymerase chain reaction, methylation-sensitive restriction endonuclease-PCR / Southern method, restriction endonuclease method combined with bisulfite, digital polymerase chain reaction, restriction landmark genome scanning, CpG island microarray, methylation profiling, and methylation chip detection; Preferably, the product further comprises reagents for processing the sample; Preferably, the sample is a sample containing leukocytes; Preferably, the sample containing leukocytes includes blood, bone marrow puncture fluid, urine, pleural effusion, ascites, cerebrospinal fluid, sputum, mucus, amniotic fluid, and lymph; Preferably, the product includes a kit, a methylation panel, a chip, a system, a device, and a readable medium.

6. A colorectal tumor risk prediction model, characterized in that: The method for constructing the prediction model comprises the following steps: obtaining methylation level data of the biomarker according to any one of claims 1 to 2 in a sample, wherein the methylation level data of the biomarker is obtained from a population with colorectal tumors and a healthy population; inputting the methylation level data into a machine learning algorithm to construct a prediction model; Preferably, the colorectal tumor includes advanced adenoma or colorectal cancer.

7. The prediction model according to claim 6, wherein: The prediction model obtains the methylation score based on the following formula: If the methylation score of the sample to be tested is higher than the threshold, the sample to be tested has advanced adenoma or a high risk of colorectal cancer; if the methylation score of the sample to be tested is lower than the threshold, the sample to be tested does not have colorectal tumor or has a low risk of colorectal tumor.

8. A colorectal tumor risk prediction system, characterized in that: The system includes the following modules: A data acquisition module, configured to acquire methylation level data of the biomarker according to any one of claims 1 to 2 of the sample to be tested; A data analysis module, configured to perform classification prediction on the data in the data acquisition module using the prediction model according to any one of claims 6 to 7, and obtain a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer; Output module, used to output classification results.

9. A colorectal tumor risk prediction device, characterized in that: The computer device comprises: a processor adapted to implement the instructions; and a memory adapted to store a plurality of instructions adapted to be loaded by a processor and executed by the processor: Acquiring data, obtaining methylation level data of the biomarker according to any one of claims 1-2 of the sample to be tested; Processing the data, inputting the methylation level data into the prediction model according to any one of claims 6-7 for classification prediction, and obtaining a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer; Output the prediction results.

10. A computer-readable medium for colorectal tumor risk prediction, characterized in that: The computer readable medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the following steps: Acquiring data, obtaining methylation level data of the biomarker according to any one of claims 1-2 of the sample to be tested; Processing the data, inputting the methylation level data into the prediction model according to any one of claims 6-7 for classification prediction, and obtaining a classification result of whether the sample to be tested is at high risk of colorectal cancer or low risk of colorectal cancer; Output the prediction results.

Citation Information

Patent Citations

  • DNA methylation site combination as colorectal tumor marker and application of DNA methylation site combination

    CN117551762A

  • Methylation marker and method for constructing prediction model

    CN117821591A

  • Method for screening colorectal cancer and colorectal polyps or advanced adenomas and application thereof

    WO2023128419A1