Colorectal polyp, adenoma and colorectal cancer methylation marker as well as screening method and application thereof
Characteristic methylation markers were screened through methylated immunoprecipitation sequencing technology, combined with differential region analysis and triclassification model, and the problems of invasiveness and low sensitivity of existing colorectal cancer screening methods were solved, achieving efficient screening of colorectal cancer and its precancerous lesions.
Patent Information
- Application Number
- CN202410136173.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-08
AI Technical Summary
Existing colorectal cancer screening methods such as colonoscopy and blood DNA testing have problems such as strong invasiveness, high false positives and low sensitivity, making it difficult to effectively screen colorectal cancer and its precancerous lesions, especially adenomas and polyps.
The methylated immunoprecipitation sequencing technology was used to screen characteristic methylation markers, and through liquid hybridization capture and high-throughput sequencing, combined with methylation difference region analysis algorithm and random forest model, a triclassification model was constructed to achieve classification prediction of normal and healthy people, colorectal polyps/adenomas and colorectal cancer.
It improves the sensitivity and specificity of colorectal cancer and its precancerous lesions, can detect precancerous lesions early, provides more efficient screening solutions, and reduces resource waste and false positive rates.
Smart Images

Figure CN120442788A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biology and gene detection, and specifically relates to methylation markers for colorectal polyps and adenomas and colorectal cancer, a gene methylation detection method and a prediction model. Background Art
[0002] Colorectal cancer is one of the most common malignancies worldwide, with its incidence and mortality rates increasing year by year. Data from China's National Cancer Center in 2023 showed that colorectal cancer surpassed gastric cancer to become the second most common cancer. Nearly 80% of patients are diagnosed in the late stages of the disease, and nearly half survive less than five years. Therefore, reducing the incidence and mortality of colorectal cancer has become a major public health issue that needs to be addressed urgently in China and globally.
[0003] According to the multi-stage theory of carcinogenesis, colorectal cancer develops morphologically through a phased progression from normal mucosal hyperplasia to polyp and adenoma formation, adenoma carcinomatization, and finally invasion and metastasis. The progression from polyps and adenomas to colorectal cancer takes 10-15 years. Early detection of colorectal cancer has a cure rate of over 90%, while late-stage cases have a cure rate of less than 10%. Screening and intervention are effective measures to reduce colorectal cancer morbidity and mortality.
[0004] Colonoscopy is the gold standard for colorectal cancer screening, but because it is highly invasive and requires cumbersome bowel preparation, the compliance of Chinese people with colonoscopy screening is relatively low, and the number of people undergoing colonoscopy examinations has increased dramatically. Large-scale use of colonoscopy for screening will also result in a huge waste of resources. The traditional screening program uses a two-step screening model that combines a questionnaire survey with two fecal occult blood tests (FIT). Anyone who is positive for any of the items is judged to be initially positive, indicating that they are at high risk and need to undergo a colonoscopy. This screening has problems such as too high false positives and low colorectal cancer detection rate. In addition, this screening model has resulted in insufficient manpower investment in the hospital, resulting in slow project progress.
[0005] For auxiliary diagnosis technology of colorectal cancer based on blood DNA testing, the current products on the market generally have low sensitivity, and the sensitivity to colorectal cancer is less than 85%, which cannot meet clinical needs. According to the public data of products approved for marketing by the National Medical Products Administration in 2023, the blood Septin9 methylation test from Shenzhen Youshengkang can achieve 77.34% sensitivity and 95.95% specificity for colorectal cancer. It also conducted the first clinical study based on blood detection of precancerous lesions, showing a sensitivity of 54.29% for high-grade intraepithelial neoplasia. Clinical practice has shown that the main limitation of blood Septin9 methylation detection is that its sensitivity for identifying colorectal cancer and precancerous lesions (adenomas) is relatively low, and the sensitivity for advanced adenomas is only 7.9%-38.7%. Some studies believe that the diagnostic performance of mSeptin9 testing for precancerous lesions (advanced adenomas and polyps) is poor and the cost is relatively high. Therefore, a lot of research is still needed to improve this situation (TEPUS M, YAU T O. Non-invasive colorectal cancer screening: an overview [J]. Gastrointestinal Tumors, 2020, 7 (3): 62-73.).
[0006] In the actual application scenarios of clinical screening, the detection performance of precancerous lesions has gradually been valued by clinicians. The early detection of colorectal polyps and adenomas, rather than just colorectal cancer, is more in line with the real needs of large-scale screening of colorectal cancer. Currently marketed products or major laboratory protocols use the same markers and the same positive judgment values as those for early diagnosis of colorectal cancer to detect colorectal precancerous lesions (especially high-grade intraepithelial neoplasia). This approach limits the detection of early lesion signals, and therefore the sensitivity to precancerous lesions is generally low. According to the multi-stage theory of the carcinogenesis process, the corresponding methylation profile during the staged evolution of colorectal cancer is likely to change continuously. There is a clear difference between the methylation spectrum in the early stage of colorectal cancer and the methylation spectrum in the late stage of colorectal cancer. Accordingly, the methylation spectrum of colorectal polyps and adenomas is also different from the methylation spectrum of the colorectal cancer stage.
[0007] In exploring marker types for early colorectal cancer screening, ctDNA methylation, compared to ctDNA mutations, not only modifies more sites but also exhibits tissue / cancer specificity, allowing for both signal abundance and signal intensity, offering significant advantages over other indicators. For example, Grail, which published its findings at the 2019 ACSO conference, tested three detection methods: targeted mutation detection based on ultra-deep sequencing, whole-genome methylation detection, and copy number variation detection. Whole-genome methylation detection achieved the best experimental results, and the study found that the scope of methylation could be narrowed to methylation detection of target genes.
[0008] To obtain information on methylation changes associated with specific cancers, the most traditional screening method for methylation markers, the gold standard, is whole-genome bisulfite sequencing (WGBS), which can obtain signals with single-base resolution. There are also reports of using reduced genome methylation sequencing (RRBS) to enrich CpG regions using restriction endonucleases. However, methylated DNA immunoprecipitation sequencing cannot obtain methylation signals with single-base resolution and can only determine whether a region is methylated by enriching peaks. Therefore, it is rarely reported and applied to the screening of cancer-related methylated DNA markers.
[0009] However, the bisulfite treatment process includes denaturation, deamination and desulfonation. The DNA is first denatured into a single strand, and then subjected to high temperature, high salt, acidic and alkaline environments, experiencing the extremes of ice and fire. The resulting converted DNA has the following morphology: mainly single strands, mixed double strands, fragment nicks, gap damage, and uracil state nucleotides. This process generally results in the loss of 90% of the DNA template, and a large amount of methylation information cannot be detected by subsequent processes. At the same time, during the base conversion treatment, there are cases of incomplete sequence conversion or over-conversion, resulting in artificial bias, which will be further amplified by subsequent PCR amplification, resulting in a large amount of data waste. In order to reduce the bias of WGBS, it is necessary to invest more DNA, reduce the number of PCR cycles and optimize the efficiency of PCR amplification enzymes. The above are the main reasons why WGBS sequencing is not efficient in screening ctDNA methylation markers.
[0010] Therefore, in order to carry out early screening for colorectal cancer, there is an urgent need to develop blood samples based on the highest clinical user compliance, explore more efficient and sensitive ctDNA methylation marker screening technologies other than bisulfite treatment methods, and effectively mine new markers that can be used for early screening of colorectal cancer on a large scale, as well as try to distinguish polyps and adenomas from colorectal cancer. Summary of the Invention
[0011] To solve the above technical problems, in a first aspect, the present invention provides a use of a methylation marker in preparing a kit for simultaneously screening colorectal polyps and / or adenomas, and colorectal cancer, wherein the methylation marker comprises multiple of the following 50 chromosomal regions defined by Hg38 coordinates: chr1:3194718-3194839, chr1:10650092-10650200, chr1:21293490-21293543, chr1:24908017-24908181, chr1:170661257-170661455, chr2:5695989-5696187, chr2:477857 88-47785832, chr3:41194800-41194941, chr3:143119154-143119154, ch r3:147413337-147413514, chr4:16898867-16899065, chr4:121380586-1 21380586、chr4:133151609-133151807、chr4:153790249-153790251、chr 4:153790307-153790447, chr4:153792983-153793181, chr5:15503692-1 5503802, chr5:15935273-15935302, chr6:520871-521069, chr6:116125 944-116126142, chr7:711979-712177, chr7:6007924-6008045, chr7:191 45339-19145339, chr7:27094593-27094791, chr7:27157838-27158007, c hr7:97731979-97732177, chr7:97737530-97737644, chr8:41308292-413 08432, chr8:52939654-52939852, chr8:98800641-98800762, chr9:37002 603-37002801, chr9:87663369-87663468, chr10:25176422-25176422, ch r10:100737574-100737772, chr11:44311452-44311645, chr11:11517597 5-115176010、chr12:49903740-49903938、chr12:118981860-118982058、chr13:77918869-77919067, chr14:85530152-85530152, chr16:23836920-23837118, chr16:51149454-51149618, chr16:86578883-86579081, chr17:77467901-77467901, chr17:77467904-77468025, chr18:47251249-47251447, chr19:29525031-29525229, chr19:53255732-53255903, chr20:21713764-21713913, chr21:33025896-33026094. ,
[0012] In some embodiments of this aspect, the methylation markers include at least methylation markers associated with colorectal polyps and adenomas, and the chromosome regions included in the methylation markers are defined by Hg38 coordinates as follows: chr1:170661257-170661455, chr2:5695989-5696187, chr4:133151609-133151807, chr7:711979-712177, chr7:97731979-97732177, chr8:52939654-52939852, chr9:37002603-37002603 02801, chr10:100737574-100737772, chr12:49903740-49903938, chr12:118981860-118982058, chr13:77918869-77919067, chr16:2 3836920-23837118, chr16:86578883-86579081, chr18:47251249-47251447, chr19:29525031-29525229, chr21:33025896-33026094.
[0013] In some embodiments of this aspect, the methylation marker further comprises one or more of the following 15 chromosomal regions defined by Hg38 coordinates: chr1:18631591-18631591, chr2:144517409-144517409, chr2:181457330-181457528, chr2:200585853-200585853, chr5:114363151-114363151, chr5:135535426-135535624, chr10:172 28966-17229164, chr10:103276971-103276971, chr12:102958584-102958782, chr12:132908689-132908887, chr14:6050932 0-60509518, chr14:60509568-60509568, chr14:69548157-69548157, chr16:77434997-77434997, chr18:72867309-72867507.
[0014] In some embodiments of this aspect, the methylation marker further comprises one or more of the following four chromosomal regions defined by Hg38 coordinates: chr10: 17228965-17229164, chr12: 102958583-102958782, chr14: 60509319-60509518, chr18: 72867308-72867507.
[0015] In the second aspect, the present invention provides a method for screening methylation markers for colorectal polyps and / or adenomas, and colorectal cancer. The screening method adopts a methylation differential region analysis algorithm, and the screened methylation markers can realize the classification prediction of normal healthy people, colorectal polyp / adenoma people and colorectal cancer people. The screened methylation markers not only have high sensitivity to colorectal cancer, but also have very high sensitivity to colorectal polyps / adenomas, and can be used to screen precancerous lesions and colorectal cancer.
[0016] The screening method comprises the following steps:
[0017] (1) Treating DNA in the sample to be tested with the methylation immunoprecipitation method, and then obtaining the methylation sequencing results of the sample to be tested by liquid phase hybridization capture and high-throughput sequencing;
[0018] (2) Based on the methylation sequencing results, the reference sequences were compared and the edgeR and DESeq2 algorithms were used to detect the peaks of methylation-enriched regions and screen the differential peaks to identify the differentially methylated regions;
[0019] (3) Use the RPM method to normalize the methylation sequencing results and calculate the RPM value of each sample. The RPM value calculation formula is as follows:
[0020]
[0021] Among them, marker refers to the characteristic gene region, marker_reads_count refers to the number of sequencing fragments (reads) of the marker, CpG_sites refers to the number of CpGs of the marker, and On_target_reads refers to the number of reads aligned to the characteristic gene region;
[0022] (4) Based on the RPM values, the samples to be tested are grouped into two levels and labels are established. In the first level grouping, the labels of colorectal cancer and colorectal polyps / adenomas are set to 1, and the labels of normal health and benign diseases are set to 0. The samples are randomly grouped according to the ratio of training set to test set of 8:2, and RPM matrix 1 is constructed. In the second level grouping, the colorectal cancer label is set to 1, and the colorectal polyps / adenomas label is set to 0. The samples are randomly grouped according to the ratio of training set to test set of 8:2, and RPM matrix 2 is constructed.
[0023] (5) The Feature Selector package was used to rank the importance of differentially methylated regions and screen out multiple characteristic differentially methylated regions as methylation markers.
[0024] The use of the Feature Selector package to rank the importance of the differentially methylated regions preferably includes using the Feature Selector to score and rank the differentially methylated regions, constructing an importance cumulative curve, and screening out multiple significantly different differentially methylated regions as methylation markers.
[0025] In this aspect, the reference sequence is a human gene sequence and / or a gene sequence of a normal sample.
[0026] Preferably, in step (1), the methylation immunoprecipitation method (MeDIP) is a 5-methylated cytosine antibody method.
[0027] Preferably, in step (1), the liquid phase hybridization capture probe is designed to fully cover the target area without gaps or overlaps, and each probe is 120 nt in length; the liquid phase hybridization capture probe is an oligonucleotide modified with 5' biotin.
[0028] Preferably, in step (1), internal standards are added during the liquid phase hybridization capture process, and the internal standards are 7 human endogenous fragments and 2 exogenous fragments pUC19 and λDNA.
[0029] Preferably, in step (1), the liquid phase hybridization capture reaction temperature is 65 degrees Celsius.
[0030] Preferably, in step (1), high-throughput sequencing is performed by a second-generation Illumina sequencing method.
[0031] In a third aspect of the present invention, a method for constructing a three-classification model for simultaneously screening colorectal polyps and / or adenomas, and colorectal cancer, and a three-classification model are provided.
[0032] The three-classification model can be used to distinguish between normal healthy people, people with precancerous lesions and people with cancer. Furthermore, on the basis of distinguishing between people with precancerous lesions and people with cancer, it can also further distinguish between colorectal cancer and colorectal polyps / adenomas, thereby achieving the technical effect of colorectal cancer screening and early screening.
[0033] The method for constructing the three-classification model includes the following steps:
[0034] Step 1: Using any of the methylation markers described in the first aspect of the present invention as candidate markers for constructing a three-classification model; using Feature Selector to score the importance of the candidate markers, constructing a cumulative importance curve, and selecting multiple methylation markers with significant differences;
[0035] Step 2: Use the LazyPredict package to score the Random Forest Classifier, Extra Trees Classifier, and K-Nearest Neighbors Classifier, and select a model based on the scoring results;
[0036] Preferably, when using the lazy predict package for scoring, the indicators considered include AUC value, F1-Score and recall rate; further preferably, the random forest model is selected for the construction of the three-classification model based on the score;
[0037] Step 3: Model the data based on the multiple significantly different methylation markers selected in step 1 and the model selected in step 2, perform 10-fold cross-validation, further screen the methylation markers based on the changes in AUC values, and then verify the modeling results on the test set to establish a three-classification model. Finally, the three-classification model is evaluated on an independent validation set.
[0038] A fourth aspect of the present invention provides a system for simultaneously screening for colorectal polyps and / or adenomas, and colorectal cancer, comprising:
[0039] (1) an acquisition module, configured to acquire measurement data of a methylation marker of a sample to be tested, wherein the methylation marker is selected from any one of the methylation markers described in the first aspect of the present invention;
[0040] (2) A data analysis module, used to input the measurement data of the methylation markers into the three-classification model constructed according to the third aspect of the present invention to obtain a screening result.
[0041] A fifth aspect of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein the computer program comprises:
[0042] (1) a program for performing the method for screening methylation markers according to the second aspect of the present invention; and / or
[0043] (2) A program for executing the method for constructing a three-category model described in the third aspect of the present invention.
[0044] The present invention has the following beneficial effects relative to the prior art:
[0045] 1. The methylation markers screened by the present invention are highly sensitive to colorectal polyps / adenomas and colorectal cancer, and can be used for early screening of colorectal cancer risk and screening of colorectal cancer.
[0046] 2. Using methylation immunoprecipitation sequencing technology to screen methylation markers will help to discover a large number of markers that can be applied to blood DNA testing on a large scale and efficiently, and has higher clinical application value.
[0047] 3. The characteristic methylation difference region analysis algorithm used in the methylation marker screening method provided by the present invention can realize the screening of different methylation markers; the three-classification model constructed by the present invention can realize the classification prediction of normal healthy people, polyp adenoma people, and colorectal cancer. It not only has high sensitivity for colorectal cancer, but also has very high sensitivity for polyps adenoma, which helps to detect precancerous lesions early and also provides a new technical solution for colorectal cancer screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is an IGV visualization of some characteristic difference regions, where Figure 1(A) shows the difference in the SOX11 gene between the colorectal cancer group samples and the polyp adenoma group samples; Figure 1(B) shows the difference in the SFRP1 gene between the colorectal cancer / polyp adenoma group samples and the normal and benign lesion group samples; Figure 1(C) shows the difference in the OLIG2 gene between the colorectal cancer group samples and the polyp adenoma group samples; Figure 1(D) shows the difference in the SIX6 gene between the colorectal cancer / polyp adenoma group samples and the normal and benign lesion group samples.
[0049] Figure 2 is the ROC curve diagram of the training set / test set of the three-classification model (based on 65 differentiated regions) constructed in Example 3, wherein Figure 2 (A) is the ROC curve diagram for distinguishing the colorectal cancer, polyp and / or adenoma groups from the normal and benign lesion groups; Figure 2 (B) is the ROC curve diagram for distinguishing the colorectal cancer group from the polyp and / or adenoma group.
[0050] Figure 3 is the ROC curve diagram of the validation set of the three-classification model (based on 65 differentiated regions) constructed in Example 3, wherein Figure 3 (A) is the ROC curve diagram for distinguishing colorectal cancer, polyp and / or adenoma groups from normal and benign lesion groups; Figure 3 (B) is the ROC curve diagram for distinguishing colorectal cancer group from polyp and / or adenoma group.
[0051] Figure 4 is a ROC curve diagram of the training set / test set of the three-classification model (based on 50 random differential regions) constructed using Example 4, wherein Figure 4 (A) is a ROC curve diagram for distinguishing between colorectal cancer, polyp and / or adenoma groups and normal and benign lesion groups; Figure 4 (B) step two is a ROC curve diagram for distinguishing between colorectal cancer group and polyp and / or adenoma group.
[0052] Figure 5 is a ROC curve diagram of the validation set of the three-classification model (based on 50 random differential regions) constructed using Example 4, wherein Figure 5 (A) is a ROC curve diagram for distinguishing between colorectal cancer, polyp and / or adenoma groups and normal and benign lesion groups; Figure 5 (B) is a ROC curve diagram for distinguishing between colorectal cancer group and polyp and / or adenoma group.
[0053] Figure 6 This is a detection flow chart for specificity and sensitivity analysis of the immunoprecipitation method and bisulfite conversion method using qPCR technology in Comparative Example 1.
[0054] Figure 7 qPCR amplification curves for detecting SALL1, SOX11, IRF4, ALX4, and TAC1 genes after the same sample was treated with the methylation immunoprecipitation method and the bisulfite conversion method in Comparative Example 1, respectively. DETAILED DESCRIPTION
[0055] In order to facilitate understanding by those skilled in the art, some terms appearing in this document are explained and illustrated. It should be noted that these explanations and illustrations are only used to help those skilled in the art understand the present invention and cannot be regarded as limiting the scope of protection of the present invention.
[0056] In the specification and claims of this application, the singular forms "a", "an" and "the" include plural forms unless the context indicates otherwise. Thus, for example, "a reagent" can be understood to include multiple reagent components.
[0057] In the specification and claims of this application, unless otherwise stated, the terms "comprise", "include" or "contain" mean that the listed values, steps or components are included, but other values, steps or components are not excluded.
[0058] In the specification and claims of this application, "subject" or "patient" are used interchangeably herein and refer to a vertebrate, preferably a mammal. The mammal can be a human, non-human primate, mouse, rat, dog, cat, horse or cow, but is not limited to these examples.
[0059] According to an embodiment of the present invention, the marker provided by the present invention refers to a chromosome site or chromosome region that can be used to detect or diagnose whether a subject suffers from a disease, and can be a nucleic acid sequence of a certain length, or can be the nucleotides of a specific site or the nucleotides of two specific sites. "Methylation marker", "methylation region" and "target marker" have the same meaning, and refer to that its methylation level indicates that the subject suffers from a disease, and it should be considered that it can include all transcriptional variants of genes described herein and all promoters and regulatory elements thereof. In addition, it should be understood that the term "marker" should include both the positive strand sequence of the marker or gene and the antisense strand sequence of the marker or gene.
[0060] In some embodiments of the present invention, the target markers of the present invention also include various variants of the above-mentioned genes. Variants include nucleic acid sequences from the same region with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity (i.e., with one or more deletions, insertions, substitutions, reverse sequences, etc.) to the genes or regions described herein. Therefore, the content of this application should be understood as extending to such variants that achieve the same results, despite the fact that the actual nucleic acid sequences between individuals have slight genetic variations. The term "marker" used is broadly interpreted as including both 1) the original markers found in biological samples or genomic DNA (in a specific methylation state), and 2) the processed sequences thereof (e.g., the corresponding regions after immunoprecipitation treatment).
[0061] In some embodiments of the invention, a "normal" sample refers to a sample of the same type isolated from an individual known to be free of the cancer, tumor, polyp, or adenoma.
[0062] The term "adenoma" refers to a benign tumor of glandular origin. Although these growths are benign, they may progress to become malignant over time. The "site" of a polyp, adenoma, cancer, etc., is the tissue, organ, cell type, anatomical region, body part, etc., in a subject's body where the polyp, adenoma, cancer, etc., is located.
[0063] The term "AUC" is an abbreviation for "area under the curve". It specifically refers to the area under the receiver operating characteristic (ROC) curve. The ROC curve is a plot of the true positive rate relative to the false positive rate for different possible cut-off points for a diagnostic test. It shows the balance between sensitivity and specificity (any increase in sensitivity will be accompanied by a decrease in specificity) based on the selected cut-off point. The area under the ROC curve (AUC) is a measure of a diagnostic test (the larger the area, the better; the best is 1; a randomized trial will have an ROC curve with an area on the diagonal of 0.5; Reference: JPEgan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0064] Below, the scheme of the present invention will be explained in conjunction with embodiment.It will be understood by those skilled in the art that the following examples are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention.In the embodiment, if specific technology or conditions are not indicated, the technology or conditions described in the literature in this area or the product instructions are used.The reagents or instruments used are not indicated by the manufacturer, and are all conventional products that can be obtained by commercial purchase.
[0065] In the embodiments of the present invention, the material sources and pretreatment methods are as follows:
[0066] 1. Sample Collection: Blood and tissue samples were collected from subjects with colorectal cancer, colorectal polyps, adenomas, and normal colorectal tissue at multiple clinical hospitals across China. Tissue samples were collected from both adjacent lesions and normal tissue whenever possible. Fresh cancer tissue was preferred, but paraffin-embedded tissue sections were also acceptable.
[0067] 2. Sample pretreatment: When pre-treating cfDNA in blood samples and DNA in tissue samples, common techniques in the field can be used to prepare for subsequent library preparation. For example, a 10 mL blood sample from a patient is centrifuged twice: the first time at 600 x g for 20 minutes at 4°C, and the second time at 16,000 x g for 10 minutes to obtain plasma. 1-2 mL of plasma is extracted using the VAHTS Free-Circulating DNA Maxi Kit (Cat. No. N903-03, Vazyme) to obtain cfDNA. CFDNA concentration is quantitatively determined using Qubit 4.0. The cfDNA yield per sample should not be less than 5 ng. For example, tissue is scraped from a section and tissue genomic DNA is obtained (Tissue Genomic DNA Extraction Kit, Tiangen Biochemical Technology (Beijing) Co., Ltd.). DNA concentration is quantitatively determined using Qubit 4.0. The DNA yield per sample should not be less than 200 ng.
[0068] Example 1 DNA sequencing to obtain methylation sequencing results of the sample to be tested
[0069] (1) Treating DNA in the sample to be tested with the methylation immunoprecipitation method, and then obtaining the methylation sequencing results of the sample to be tested by liquid phase hybridization capture and high-throughput sequencing;
[0070] Blood samples from 1,104 colorectal cancer patients, 335 polyp and adenoma blood samples, and 641 normal colorectal and benign disease blood samples were collected, totaling 2,080 samples. A total of 406 pairs of colorectal cancer and adjacent tissues, 54 polyp and adenoma tissue samples, and 68 normal benign tissue samples were collected, totaling 528 samples.
[0071] 1.1. Construction of a methylated DNA enrichment library
[0072] A DNA library was constructed using a library construction kit for the Illumina high-throughput sequencing platform (e.g., VAHTS Universal ProDNA Library Prep Kit for Illumina, Cat. No. ND608-02, Vazyme), and methylated fragment antibodies were enriched using a methylated DNA enrichment kit (e.g., MagMeDIP kit, Cat. No. C02010021, Diagenode). Fragments were then purified and PCR amplified (the number of cycles was generally 9-11, not more than 15) using a DNA purification kit (e.g., IPure Kit v2, Diagenode) and a library construction kit for the Illumina high-throughput sequencing platform (e.g., VAHTS Universal Pro DNA Library Prep Kit for Illumina, Cat. No. ND608-02, Vazyme). After amplification, the product is purified using an equal volume of magnetic beads (e.g., VAHTS DNA Clean Beads, Cat. No. N411-03, Vazyme) and a MeDIP methylation-enriched library is obtained based on the methylation immunoprecipitation technique (MeDIP).
[0073] The methylation immunoprecipitation method is preferably a 5-methylated cytosine antibody method. Preferably, equal amounts of positive sample libraries and negative sample libraries are added to the same immunoprecipitation reaction, or samples with unknown clinical results can be added to the same immunoprecipitation reaction. The total input amount of DNA library for each immunoprecipitation reaction should be in the range of 10ng to 1000ng. It is possible to perform an immunoprecipitation reaction of a single DNA library sample or to accommodate a DNA library of 100 samples to complete the immunoprecipitation reaction. The preferred DNA library denaturation condition is 95 degrees Celsius for 10 minutes.
[0074] 1.2 Liquid Phase Hybridization Capture
[0075] Using a hybrid capture kit (e.g., Hybrid Capture Reagents, Cat.
[0076] No.REF1005101, Naonda) for liquid phase hybridization capture: 500 ng of methylation enriched library was taken, HumanCot DNA and Nad Nano Blockers, add hybridization reaction solution (containing liquid phase hybridization capture probe), hybridize and capture for 4-16 hours under the hybridization program: 95℃ / 30sec; 65℃ / Hold (100℃ hot cover); then add streptavidin magnetic beads to the hybridization system and incubate for 40 minutes. During this period, vortex and mix every 10 minutes to ensure that the magnetic beads are completely resuspended; wash the bound magnetic beads, and discard the residual liquid at each step; finally, add 20μl of nuclease-free water and gently vortex to mix.
[0077] Among them, the reaction temperature of hybridization capture is 65°C, instead of 63°C for methylation probes designed based on bisulfite conversion. Preferably, the liquid phase hybridization capture probe is designed to fully cover the target area without gaps (gap) and cross-coverage (overlap), and each probe is 120nt in length; the liquid phase hybridization capture probe is an oligonucleotide (Oligo) modified with 5'-end biotin. Preferably, 7 human endogenous fragments and 2 exogenous fragments pUC19 and λDNA are added as internal standards during the liquid phase hybridization capture process. The hybridization capture reaction can be single hybrid or multi-hybrid, and the total amount of MeDIP amplification library input for each hybridization capture reaction should be between 300ng and 8μg. It is preferred to add human placental DNA (Human Cot DNA) and blocking sequences (for example, Nad Nano Blockers, Naonda).
[0078] 1.3 High-throughput sequencing
[0079] The hybridization capture product was PCR amplified and purified using a kit (e.g., VAHTS Universal Pro DNA Library Prep Kit for Illumina, Cat. No. ND608-02, VAHTS DNA Clean Beads, Cat. No. N411-03, Vazyme) to obtain a hybridization capture library. The library concentration was diluted to 4 nM and mixed according to the required data volume. The total data volume should not exceed 120G. After mixing, 5 μl of the library was taken out, 5 μl of 0.2N NaOH was added, pipetting and mixing were performed, and denaturation was performed for 5 minutes. After the end, 990 μl of hybridization reaction solution (HT1 Buffer, REF: 15058251, Illumina) was immediately added. After vortex mixing, 105 μl was taken out and added to 1295 μl of HT1 Buffer. After vortex mixing, the library was prepared for the machine with a concentration of 1.5 pM.
[0080] The sequencer was an Illumina NextSeq 550Dx, and the reagents used were High Output Reagent Cartridge v2 (REF: 15057929, Illumina) (300 cycles), High Output Flow Cell Cartridge v2.5 (REF: 20022408, Illumina), and Buffer Cartridge v2 (REF: 15057941, Illumina). 1300μl of the library was added to the sample position of the High Output Reagent Cartridge v2, and each reagent was added in sequence to begin sequencing. Paired-end sequencing was used, and the total duration was approximately 30 hours.
[0081] Fastp (version 0.22.0) was used to perform quality control on the data and remove low-quality bases. The overall Q20 of the clean data was above 90%, the Q30 was above 85%, and the average sequencing depth was around 300x.
[0082] Through the above operations, the methylation sequencing results of the sample to be tested are obtained.
[0083] Example 2 Screening of methylation markers
[0084] The obtained methylation sequencing result data of the sample to be tested is compared with the human gene sequence for sequence analysis, methylation enrichment region peak detection is performed, and the differential methylation regions of colorectal polyps / adenomas and colorectal cancer are preliminarily screened. The DiffBind tool (version 3.8.4) is used to screen the peaks of difference between colorectal cancer, colorectal polyps / adenomas and normal healthy people. Preferably, the characteristic differential methylation regions of colorectal cancer screened in blood samples and tissue samples are further intersected to obtain a combination of colorectal cancer polyp adenoma and colorectal cancer markers. The specific steps include:
[0085] 1. Group the blood sample data into two groups: colorectal cancer and polyps and adenomas, grouped as the POS group, and normal and benign lesions, grouped as the NEG group.
[0086] 2. Use the DiffBind toolkit to analyze the grouped data, and set the peak length to 200;
[0087] 3. Use the DESeq2 and edgeR algorithms to obtain the differentially methylated regions of blood samples and obtain the intersection;
[0088] 4. Group the tissue sample data and repeat steps 1-3 to obtain the differentially methylated regions of the tissue samples;
[0089] 5. Intersect the results of steps 3 and 4 to obtain differentially methylated regions in blood samples supported by tissue samples;
[0090] 6. Perform CpG site statistics on the differentially methylated regions in step 5;
[0091] 7. Annotate the differentially methylated regions obtained from step 5 to obtain a list of related genes.
[0092] Further detailed steps include:
[0093] (1) Use the edgeR and DESeq2 algorithms of the DiffBind tool to find the common differential peak regions in blood samples (cfDNA) compared with normal blood samples or the human genome (Hg38);
[0094] (2) Use the edgeR and DESeq2 algorithms of the DiffBind tool to find the common differential peak regions in tissue samples compared with normal blood samples or the human genome (Hg38);
[0095] (3) Use the intersect function in the bedtools program to find the common differentially methylated regions between blood samples and tissue samples;
[0096] (4) Calculate the RPM value of each sample using the following formula:
[0097]
[0098] Among them, marker refers to the characteristic gene region, marker_reads_count refers to the number of sequencing fragments (reads) of the marker, CpG_sites refers to the number of CpGs of the marker, and On_target_reads refers to the number of reads aligned to the characteristic gene region;
[0099] (5) Based on the RPM values, the samples to be tested are grouped at two levels and labels are established. In the first level of grouping, the labels of colorectal cancer and colorectal polyps / adenomas are set to 1, and the labels of normal health and benign diseases are set to 0. The samples are randomly grouped according to the ratio of training set to test set of 8:2, and RPM matrix 1 is constructed. In the second level of grouping, the colorectal cancer label is set to 1, and the colorectal polyps / adenomas label is set to 0. The samples are randomly grouped according to the ratio of training set to test set of 8:2, and RPM matrix 2 is constructed.
[0100] (6) The Feature Selector package was used to rank the importance of differentially methylated regions and screen out multiple characteristic differentially methylated regions as methylation markers.
[0101] The method uses the Feature Selector package to rank the importance of the differentially methylated regions, including using Feature Select to score and rank the differentially methylated regions, construct an importance accumulation curve, and screen out multiple significantly different differentially methylated regions as methylation markers.
[0102] The methylation sequencing result data of the test sample obtained in Example 1 was sorted according to the above steps to complete the methylation differential region sorting, and the top 65 were selected as the characteristic methylation differential regions, that is, as methylation markers (there are significant methylation level differences between normal samples and colorectal cancer, colorectal polyps and adenomas). Their human genome database Hg38 coordinates are listed in Table 1. Among them, the IGV visualization results of the SOX11, SFRP1, OLIG2, and SIX6 genes in the characteristic differential regions are shown in Figure 1. Furthermore, the characteristic methylation differential regions between colorectal polyps and adenomas and colorectal cancer are shown in Table 2.
[0103] Table 1: 65 characteristic differentially methylated regions
[0104]
[0105]
[0106]
[0107] Table 2: 20 characteristic differentially methylated regions between polyp and / or adenoma group and colorectal cancer group
[0108]
[0109]
[0110] Example 3 Construction and Validation of a Three-Classification Model (Based on 65 Characteristic Methylation Differential Regions)
[0111] According to actual needs, multiple colorectal cancer and / or colorectal polyp adenoma methylation markers (i.e., the characteristic methylation difference regions screened by the present invention) are selected as markers for constructing a three-classification model. The training set and the test set are split, and the three-classification model is constructed using a random forest model, an extra tree model (Extra Trees Classifier), a K-nearest neighbor model (KneighborsClassifier), etc.
[0112] The construction of the three-classification model includes the following steps:
[0113] Step 1. Using the methylation markers screened by the present invention as candidate markers for constructing a three-classification model. In this example, the 65 methylation marker candidates screened in Example 2 were used. Feature Select was used to perform importance scoring on the 65 candidate markers, construct an importance cumulative curve, and select methylation markers with significant differences. In this example, all 65 candidate markers were selected as methylation markers.
[0114] Step 2. Use the lazy predict package to score the random forest model, the extra tree model, and the K-nearest neighbor model, and select a model based on the score. When using the lazy predict package for scoring, the indicators considered include AUC value, F1-Score, recall rate, etc. In this embodiment, step 2 selects the random forest model for the construction of the three-classification model based on the score.
[0115] Step 3. Use the multiple significantly different methylation markers selected in step 1 and the model selected in step 2 to perform modeling; 10-fold cross validation The methylation markers were further screened based on the changes in the AUC values. In this embodiment, the initial value of AUC reached 0.78. If the AUC value decreased, the candidate marker was deleted; if the AUC value did not decrease, the candidate marker was retained. The model was then validated on the test set, a three-classification model was established, and finally the three-classification model was evaluated on an independent validation set.
[0116] 3.1 Construction of a three-classification model (based on 65 characteristic differentially methylated regions)
[0117] According to the ratio of colorectal cancer: polyps / adenomas: normal and benign lesions = 1:1:1, 600 cfDNA samples were randomly selected, including 200 samples each of colorectal cancer, polyps / adenomas, and normal and benign lesions, to construct a three-classification model.
[0118] The sensitivity and specificity of the three-classification model are shown in Table 3. In the training / test set, the sensitivity of colorectal cancer was 100%, the sensitivity of colorectal polyps / adenomas was 96%, and the specificity of normal health and benign diseases was 99%. The ROC curves are shown in Figure 2. Figure 2(A) is the ROC curve for distinguishing colorectal cancer, colorectal polyps / adenomas, and normal and benign lesions, with an AUC of 0.96. Figure 2(B) is the ROC curve for distinguishing colorectal cancer from colorectal polyps / adenomas, with an AUC of 0.94.
[0119] Table 3: Sensitivity and specificity of the training set and test set of the three-classification model for screening colorectal polyps / adenomas, colorectal cancer, normal and benign (based on 65 characteristic methylation differential regions)
[0120]
[0121]
[0122] Based on two levels of grouping and their corresponding labels, at the first level, colorectal cancer and colorectal polyps / adenomas are labeled 1, while normal health and benign diseases are labeled 0. At the second level, colorectal cancer is labeled 1 and colorectal polyps / adenomas are labeled 0. The model predictions provide corresponding label values, either 0 or 1. In the table, "negative" indicates a first-level model prediction of 0; "positive" indicates a first-level model prediction of 1 and a second-level model prediction of 1; and "polyp / adenoma" indicates a first-level model prediction of 1 and a second-level model prediction of 0. In Tables 4, 6, and 7, "positive" (++), "polyp / adenoma" (+), and "negative" have the same meaning.
[0123] Explanation of specificity results: In Table 3, if polyps / adenomas, normal health and benign diseases are classified as non-colorectal cancer, they are all counted as negative, and the specificity is (195+3) / 200=99%; if only normal health and benign diseases are counted as negative, the specificity is 195 / 200=97.5%.
[0124] 3.2 Validation of the three-classification model (based on 65 characteristic differentially methylated regions)
[0125] In addition to the clinical samples used in Example 3.1, 481 plasma samples were randomly selected, including 172 from the colorectal cancer group, 135 from the colorectal polyp / adenoma group, and 174 from the normal healthy and benign disease group, for independent validation of the three-classification model. Sequencing was performed according to Example 1, or according to methods well known to those skilled in the art.
[0126] The constructed three-classification model was used for analysis, and the results are shown in Table 4 below. The sensitivity for colorectal cancer was 94.19%, the sensitivity for colorectal polyps / adenomas was 74.07%, and the specificity for normal and benign diseases was 97.7%. The ROC values are shown in Figure 3 . Figure 3 (A) is the ROC curve for distinguishing colorectal cancer, colorectal polyps / adenomas, and normal and benign lesions, with an AUC of 0.95. Figure 3 (B) is the ROC curve for distinguishing colorectal cancer from colorectal polyps / adenomas, with an AUC of 0.93.
[0127] Table 4: Sensitivity and specificity of the independent validation set of the three-classification model for screening colorectal polyps / adenomas, colorectal cancer, normal and benign diseases (based on 65 characteristic methylation differential regions)
[0128]
[0129] Explanation of the specificity results: In Table 4, if polyps / adenomas, normal health and benign diseases are classified as non-colorectal cancer and are all counted as negative, the specificity is (145+25) / 174=97.7%; if only normal health and benign diseases are counted as negative, the specificity is 145 / 174=83.3%.
[0130] Example 4 Construction and Validation of a Three-Classification Model (Based on 50 Characteristic Methylation Differential Regions)
[0131] The model construction steps refer to Example 3, except that in step (1), 50 characteristic methylation difference regions among the 65 characteristic methylation difference regions in Table 1 are randomly selected as methylation markers. The Hg38 coordinates of the 50 characteristic methylation difference regions selected in this example are shown in Table 5.
[0132] Table 5: List of 50 characteristic differentially methylated regions
[0133]
[0134]
[0135] 4.1 Construction of a three-classification model (based on 50 characteristic differentially methylated regions)
[0136] The sensitivity and specificity of the three-class classification model (based on 50 characteristic differentially methylated regions) for 600 plasma samples are shown in Table 6. In the training and test sets, the sensitivity for colorectal cancer was 100%, the sensitivity for colorectal polyps / adenomas reached 94.5%, and the specificity for normal health and benign diseases was 99%. The ROC curves are shown in Figure 4. Figure 4(A) shows the ROC curve for distinguishing between colorectal cancer and colorectal polyps / adenomas, and normal and benign lesions. Figure 4(B) shows the ROC curve for distinguishing between colorectal cancer and colorectal polyps / adenomas in step 2.
[0137] Table 6: Sensitivity and specificity of the training set / test set for the three-classification model for screening colorectal polyps / adenomas, colorectal cancer, normal and benign (based on 50 characteristic methylation differential regions)
[0138]
[0139] Explanation of the specificity results: In Table 6, if polyps / adenomas, normal health and benign diseases are classified as non-colorectal cancer, they are all counted as negative, and the specificity is (192+6) / 200=99%; if only normal health and benign diseases are counted as negative, the specificity is 192 / 200=96%.
[0140] 4.2 Validation of the three-classification model (based on 50 characteristic differentially methylated regions)
[0141] The three-class classification model constructed in Example 4.1 was used to analyze the 481 plasma samples selected in Example 3.2. The results are shown in Table 7 below. The sensitivity for colorectal cancer was 94.19%, the sensitivity for polyps / adenomas was 74.07%, and the specificity for normal health and benign diseases was 93.1%. The ROC curves are shown in Figure 5. Figure 5(A) shows the ROC curve for distinguishing between the colorectal cancer, polyp, and / or adenoma groups and the normal and benign disease groups. Figure 5(B) shows the ROC curve for distinguishing between the colorectal cancer group and the polyp and / or adenoma group.
[0142] Table 7: Sensitivity and specificity of the independent validation set for the three-classification model for screening colorectal polyps, adenoma, colorectal cancer, normal and benign (based on 50 characteristic methylation differential regions)
[0143]
[0144] Explanation of the specificity results: In Table 7, if polyps / adenomas, normal health and benign diseases are classified as non-colorectal cancer and are all counted as negative, the specificity is (133+29) / 174=93.1%; if only normal health and benign diseases are counted as negative, the specificity is 133 / 174=76.44%.
[0145] Through Example 4, 50 characteristic methylation difference regions were randomly selected from the 65 screened characteristic methylation difference regions to construct a three-classification model. Both the training / test set and the validation set showed good performance. In the independent validation set of 481 plasma samples, the three-classification model constructed with 65 characteristic methylation difference regions had better specificity, and the independent validation set could reach 97.7%.
[0146] In terms of sensitivity, the three-classification model constructed with 50 characteristic methylation difference regions and the three-classification model constructed with 65 characteristic methylation difference regions both maintained high sensitivity, with a sensitivity of 94.19% for colorectal cancer and 74.07% for polyps / adenomas.
[0147] At the same time, the present invention discovered 20 characteristic methylation difference regions between colorectal polyps / adenomas and colorectal cancer, as shown in Table 2 in Example 2. The present invention provides a possible use of a methylation marker in the preparation of a kit for simultaneous screening of colorectal polyps / adenomas and colorectal cancer, thereby providing strong support for early diagnosis and treatment of the disease.
[0148] Comparative Example 1 Comparison of MeDIP Method and Bisulfite Treatment (BS) Method
[0149] In order to further demonstrate the technical advantages of the MeDIP method used in the present invention compared with the BS method, qPCR technology was used to detect and analyze the specificity and sensitivity of the above two methods. In this comparative example, 48 plasma samples clinically diagnosed as colorectal cancer and 72 plasma control samples with negative colonoscopy results were selected. Among the 48 plasma samples clinically diagnosed as colorectal cancer, there were 6 samples of colorectal cancer stage 0 to I, 10 samples of colorectal cancer stage II, 15 samples of colorectal cancer III, and 17 samples of colorectal cancer IV. The methylation gene qPCR detection process is as follows Figure 6 shown.
[0150] The details are as follows:
[0151] 1) Methylated DNA was treated using the MeDIP method, with the same steps as in 1.1 of Example 1.
[0152] 2) BS method for treating methylated DNA
[0153] Each cfDNA methylation reaction is performed in a separate bisulfite reaction, with approximately 10–100 ng of cfDNA sample input. Methylated DNA treatment based on bisulfite conversion can be performed using commercial kits or custom reagents according to the instructions provided. For example, the ZYMO RESEARCH Biotechnology Company DNA Methylation Kit (EZ DNA Methylation Kit, D5002) can be used for DNA bisulfite modification. The elution volume should be 20–50 μL.
[0154] 3) qPCR detection
[0155] Five target genes were selected: SALL1, SOX11, IRF4, ALX4, and TAC1. The Taqman MGB probe primer pair sequences for the MeDIP method and the BS method are shown in Table 8. The design sequences of the primer probe pairs and the sources of the relevant CpG sites are all part of the characteristic differential methylation regions screened by Example 2. PCR amplification was performed using enriched or bisulfite-converted DNA as a template, and the final concentration of each primer was 10 μM. The PCR reaction system consisted of 2 to 15 μL of enriched template DNA, 2.5 μL of the premix of the above primers, 17.5 μL of PCR reaction solution reagents (e.g., 2xRapid Taq Master Mix, Vazyme), and the total volume was filled with water to 35 μL. The PCR reaction conditions were as follows: 95°C for 5 minutes, 95°C for 15 seconds, 60°C for 40 seconds, and 48 cycles of amplification.
[0156] Table 8: Taqman MGB probe primer pair sequences
[0157]
[0158]
[0159] Result analysis:
[0160] Differences in enrichment effects of target genes between the MeDIP method and the BS method
[0161] Analyze the test data and set the threshold cycle (Ct value) of samples without amplification curve or Undetermined to 48. Set a reference Ct value through ROC curve analysis. If the target gene amplification Ct value of the tested sample is equal to or lower than the set reference Ct value, the sample is judged as a positive sample, otherwise it is judged as a negative sample. Figure 7 The following example shows qPCR amplification curves for five target genes in a sample from a colorectal cancer patient treated with DNA using the MeDIP and BS methods, respectively. The figure shows that the amplification Ct values for the SALL1 gene are 31.28 for the MeDIP method and 37.98 for the BS method; 36.27 for the SOX11 gene and 41.38 for the BS method; 32.6 for the IRF4 gene and 38.46 for the BS method; 34.18 for the ALX4 gene and 38.51 for the BS method; and 34.61 for the TAC1 gene and 37.77 for the BS method. These results indicate that the MeDIP method produces lower amplification Ct values for the target genes compared to the BS method, indicating that the MeDIP method has a better enrichment effect on the target genes in the sample.
[0162] Differences in validation performance of target genes between the MeDIP method and the BS method
[0163] Table 9 shows the performance differences between the MeDIP and BS methods for qPCR target gene validation. For each target gene, the MeDIP method demonstrated higher sensitivity than the BS method for colorectal cancer samples, higher specificity for non-colorectal cancer samples, and higher accuracy. This indicates that the MeDIP method outperforms the BS method for target gene validation.
[0164] In summary, compared with the BS method, the overall performance of the MeDIP method is better.
[0165] Table 9: Performance differences between the MeDIP method and the BS method in qPCR gene validation
[0166]
[0167] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The above implementation description is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. Changes and improvements to the present invention will be possible without exceeding the concept and scope specified in the appended claims. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. Use of a methylation marker in the preparation of a kit for simultaneous screening of colorectal polyps / adenomas and colorectal cancer, characterized in that: The chromosome region locations of the methylation markers include: chr1:3194718-3194839, chr1:10650092-10650200, chr1:21293490-21293543, chr1:24908017-24908181, chr1:170661257-170661455, chr2:5695989-5696187, chr2:47785788-47785832, chr3:4119 4800-41194941, chr3:143119154-143119154, chr3:147413337-147413514, chr4:16898867-16899065, chr4:1 21380586-121380586、chr4:133151609-133151807、chr4:153790249-153790251、chr4:153790307-153790447、 chr4:153792983-153793181, chr5:15503692-15503802, chr5:15935273-15935 302, chr6:520871-521069, chr6:116125944-116126142, chr7:711979-712177, chr7:6007924-6008045, chr7:19145339-19145339, chr7:27094593-27094791, chr7:27157838-27158007, chr7:97731 979-97732177、chr7:97737530-97737644、chr8:41308292-41308432、chr8:52939654-52939852、chr8:98800641-9880 0762, chr9:37002603-37002801, chr9:87663369-87663468, chr10:25176422-25176422, chr10:100737574-100737772 , chr11:44311452-44311645, chr11:115175975-115176010, chr12:49903740-49903938, chr12:118981860-118982058, chr13:77918869-77919067, chr14:85530152-85530152, chr16:23836920-23837118, c hr16:51149454-51149618, chr16:86578883-86579081, chr17:77467901-77467901, ch Multiple of r17:77467904-77468025, chr18:47251249-47251447, chr19:29525031-29525229, chr19:53255732-53255903, chr20:21713764-21713913, chr21:33025896-33026094.
2. The use according to claim 1, characterized in that The chromosome region positions of the methylation markers include at least: chr1:170661257-170661455, chr2:5695989-5696187, chr4:133151609-133151807, chr7:711979-712177, chr7:97731979-97732177, chr8:52939654-52939852, chr9:37002603-37002801, chr10:100737574-1007 37772, chr12:49903740-49903938, chr12:118981860-118982058, chr13:77918869-77919067, chr16:23836920-23837 118. chr16:86578883-86579081, chr18:47251249-47251447, chr19:29525031-29525229, chr21:33025896-33026094.
3. The use according to claim 1 or 2, characterized in that The chromosome region locations of the methylation markers also include: chr1:18631591-18631591, chr2:144517409-144517409, chr2:181457330-181457528, chr2:200585853-200585853, chr5:114363151-114363151, chr5:135535426-135535624, chr10:17228966-17229164, chr10:103276971- 103276971, chr12:102958584-102958782, chr12:132908689-132908887, One or more of chr14:60509320-60509518, chr14:60509568-60509568, chr14:69548157-69548157, chr16:77434997-77434997, chr18:72867309-72867507; and / or one or more of chr10:17228965-17229164, chr12:102958583-102958782, chr14:60509319-60509518, chr18:72867308-72867507.
4. A method for screening methylation markers for colorectal polyps / adenomas and colorectal cancer, the method being used for non-diagnostic purposes, comprising the following steps: (1) Treating DNA in the sample to be tested with the methylation immunoprecipitation method, and then obtaining the methylation sequencing results of the sample to be tested by liquid phase hybridization capture and high-throughput sequencing; (2) Based on the methylation sequencing results, the reference sequences were compared and the edgeR and DESeq2 algorithms were used to detect the peaks of methylation-enriched regions and screen the differential peaks to identify differentially methylated regions; (3) Use the RPM method to normalize the methylation sequencing results and calculate the RPM value of each sample. The RPM value calculation formula is as follows: in, marker refers to the characteristic gene region, marker_reads_count refers to the number of sequencing fragments (reads) of the marker, CpG_sites refers to the number of CpGs of the marker, and On_target_reads refers to the number of reads aligned to the characteristic gene region; (4) Based on the RPM values, the samples to be tested are grouped into two levels and labels are established: in the first level, the labels of the colorectal cancer and colorectal polyp / adenoma groups are set to 1, and the labels of the normal health and benign disease groups are set to 0. The samples are randomly grouped according to the training set: test set ratio of 8:2, and RPM matrix 1 is constructed; in the second level, the labels of the colorectal cancer group are set to 1, and the labels of the colorectal polyp / adenoma group are set to 0. The samples are randomly grouped according to the training set: test set ratio of 8:2, and RPM matrix 2 is constructed; (5) Use the Feature Selector package to rank the importance of the differentially expressed regions and select the characteristic differentially methylated regions as methylation markers; The reference sequence is a human gene sequence and / or a gene sequence of a normal sample.
5. The screening method according to claim 4, wherein In step (1), the methylation immunoprecipitation method is a 5-methylated cytosine antibody method.
6. The screening method according to claim 4, wherein In step (1), the liquid phase hybridization capture probes are designed to fully cover the target region without gaps or overlaps, and each probe is 120 nt in length. The capture probe for liquid phase hybridization is an oligonucleotide modified with biotin at the 5' end.
7. The screening method according to claim 4, wherein In step (1), internal standards are added during the liquid phase hybridization capture process. The internal standards are 7 human endogenous fragments and 2 exogenous fragments pUC19 and λDNA.
8. The screening method according to claim 4, wherein In step (1), the liquid phase hybridization capture reaction temperature is 65 degrees Celsius.
9. The screening method according to claim 4, wherein In step (1), high-throughput sequencing is performed using the second-generation Illumina sequencing method.
10. A method for constructing a three-classification model for simultaneous screening of colorectal polyps / adenomas and colorectal cancer, characterized in that: The construction method comprises the following steps: (1) Using the methylation markers described in any one of claims 1 to 3 as candidate markers for constructing a three-classification model; using Feature Selector to score the importance of the candidate markers, constructing an importance accumulation curve, and selecting multiple methylation markers with significant differences; (2) Use the LazyPredict package to score the Random Forest Classifier, Extra Trees Classifier, and K-nearest neighbor Classifier, and select a model based on the scoring results for the construction of the three-classification model; (3) Modeling is performed using the multiple methylation markers with significant differences selected in step (1) and the model selected in step (2).
11. The construction method according to claim 10, wherein: When using the Lazy predict package for scoring in step (2), the indicators considered include AUC value, F1-Score and recall rate.
12. The construction method according to claim 10, wherein: Step (2) selects the random forest model based on the scoring results for the construction of the three-classification model.
13. The construction method according to claim 10, wherein: When modeling in step (3), 10-fold cross-validation was performed to further screen the methylation markers based on the changes in AUC values. The modeling results were then validated on a test set to establish a three-classification model. Finally, the three-classification model was evaluated on an independent validation set.
14. A computer system for simultaneous screening of colorectal polyps / adenomas and colorectal cancer, characterized in that: include: (1) an acquisition module, configured to acquire measurement data of a methylation marker of a sample to be tested, wherein the methylation marker is selected from the methylation markers described in any one of claims 1 to 3; (2) A data analysis module, which is used to input the measurement data of the methylation markers into the three-classification model obtained according to the construction method described in claim 10 to obtain the screening results.
15. A computer-readable storage medium comprising a stored computer program, characterized in that: The computer program comprises: (1) A program for performing the screening method according to claim 4; and / or (2) A program for executing the construction method according to claim 10.