Methylation prediction markers and screening methods and applications thereof
By constructing methylation gene markers by taking specific bp base sequences during cfDNA comparison and combining them with machine learning methods, the problems of loss and inaccuracy in DNA methylation prediction in existing technologies are solved, and efficient and accurate prediction of methylation levels and cancer risks is achieved.
Patent Information
- Application Number
- CN202411920170.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Existing technologies have problems with large losses and insufficient standardization when predicting DNA methylation levels, and fail to comprehensively analyze the correlation between other base sequences at the ends of cfDNA and methylation, resulting in inaccurate predictions.
Using methylation gene markers, a cancer risk prediction model is constructed by taking a certain bp of base sequence upstream, downstream, or upstream and downstream respectively from the 5' end breakpoint when cfDNA is aligned to the reference genome, and machine learning methods such as the gradient boosting algorithm are used for prediction.
A more comprehensive and accurate prediction of methylation levels and cancer risks was achieved. The model showed good predictive performance in the whole genome, CpG island and CpG shore intervals, with a verified AUC value of up to 0.97-0.91.
Smart Images

Figure CN119351560B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of molecular biotechnology, and in particular, relates to methylation prediction markers and screening methods and applications thereof. Background Art
[0002] DNA methylation plays an important regulatory role in the progression of disease. There have been a large number of studies on the role of cfDNA methylation levels in disease diagnosis and prognosis prediction. However, the gold standard sulfite methylation treatment can cause a large loss of cfDNA, and the enzyme conversion methylation technology also needs further standardization and evaluation.
[0003] cfDNA is not randomly fragmented, and its fragmentation pattern is related to nuclease activity and nucleosome structure. At the same time, studies have shown that the cfDNA fragmentation pattern is related to the epigenetic background. There are significant differences in the fragmentation patterns between methylated and unmethylated cfDNA molecules. Therefore, DNA methylation-related research can be conducted based on the cfDNA fragmentation pattern.
[0004] The study "Epigenetic analysis of cell-free DNA byfragmentomic profiling" published in PANS in 2022 showed that the genome-wide cfDNA 5' CGN / NCG motif ratio was verified to be positively correlated with the methylation level, and the "CGN / NCG motif ratio" can be used to characterize the methylation level; using CGN and NCG (including 8 base sequences, namely CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG), cancer can be predicted.
[0005] It is well known that biological physiological activities are complex, but this method only extracts 8 CpG-related 5'-end base sequences and does not analyze and verify the correlation between other base sequences at the end of cfDNA and methylation.
[0006] Therefore, the field still lacks a more comprehensive, effective and accurate marker for predicting methylation. Summary of the Invention
[0007] The present invention provides a marker that can be used to efficiently and accurately predict methylation.
[0008] In a first aspect of the present invention, there is provided a use of a methylation gene marker or a detection reagent thereof for preparing a detection reagent or a kit for detecting (a) methylation level; and / or (b) cancer risk;
[0009] The methylation gene marker is a set of base sequences obtained by taking n bp upstream, downstream, or upstream and downstream (upstream and downstream) from the 5' end breakpoint when cfDNA is aligned to the reference genome, with an interval of m bp, and taking n bp; m is any integer from 0 to 5, and n is any integer from 1 to 4;
[0010] The methylation gene markers include a set of base sequences coded as A1 to A59 in the following table:
[0011]
[0012] In another preferred embodiment, the methylation gene marker further includes a set of base sequences coded as B1 to B41 in the following table:
[0013]
[0014] In another preferred embodiment, when the direction of the gene marker is upstream and downstream, its base length = 2*n.
[0015] In another preferred embodiment, the cancer is selected from the following group: gastric cancer, breast cancer, colon cancer, bile duct cancer, glioma, sarcoma, epithelial cancer, lymphoma, melanoma, fibroma, meningioma, brain cancer, kidney cancer, thyroid cancer, parathyroid cancer, pituitary tumor, adrenal tumor, osteosarcoma tumor, neuroendocrine system tumor, lung cancer, head and neck cancer, prostate cancer, esophageal cancer, tracheal cancer, liver cancer, bladder cancer, pancreatic cancer, ovarian cancer, uterine cancer, cervical cancer, testicular cancer, skin cancer, or a combination thereof.
[0016] In another preferred embodiment, the cancer is selected from the group consisting of gastric cancer, breast cancer, colon cancer, bile duct cancer, or a combination thereof.
[0017] In a second aspect of the present invention, a method for constructing a cancer risk prediction model is provided, comprising the following steps:
[0018] (s1) providing a first data set for model training and a second data set for model testing, wherein the first data set includes methylation gene marker data of a cancer-positive sample group and methylation gene marker data of a cancer-negative sample group;
[0019] The second data set includes methylation gene marker data of a cancer-positive sample group and methylation gene marker data of a cancer-negative sample group; and
[0020] (s2) constructing a cancer risk prediction model using a machine learning method based on the data in the first dataset, and validating the model on the second dataset, thereby obtaining the cancer risk prediction model.
[0021] In another preferred example, the methylation gene marker data is the frequency of the methylation gene marker.
[0022] In another preferred example, the frequency of the methylation gene marker is the proportion of the cfDNA fragments corresponding to the methylation gene marker in the sample to all the cfDNA fragments.
[0023] In another preferred embodiment, the machine learning method is a gradient boosting algorithm (GBM).
[0024] In another preferred embodiment, the methylation gene markers include a set of base sequences coded as A1 to A59 in the following table:
[0025]
[0026] In another preferred embodiment, the methylation gene marker further includes a set of base sequences coded as B1 to B41 in the following table:
[0027]
[0028] In another preferred embodiment, the methylation gene marker data is obtained by the following steps:
[0029] (i) extracting cfDNA from the sample to be tested and obtaining read data;
[0030] (ii) aligning the read data to a reference genome to obtain the position of the 5' end of the read on the reference genome;
[0031] (iii) based on the position of the 5' end of the read segment on the reference genome, taking n bp upstream, downstream, or both upstream and downstream from the 5' end breakpoint at intervals of m bp to obtain a base sequence, thereby obtaining a set of base sequences of the methylation gene markers in Table A and / or Table B; and
[0032] (iv) Counting the proportion of cfDNA fragments corresponding to each base sequence in the base sequence set in all cfDNA fragments, thereby obtaining the frequency of each base sequence, and thus obtaining the methylation gene marker data.
[0033] In another preferred embodiment, the sample to be tested is a body fluid sample.
[0034] In another preferred embodiment, the sample to be tested is selected from the following group: a plasma sample, a blood sample, a urine sample, an ascites sample, or a combination thereof.
[0035] In a third aspect of the present invention, a cancer risk prediction device is provided, comprising:
[0036] (a) An input module configured to input methylation gene marker data in a sample to be tested; the methylation gene markers include a set of base sequences coded as A1 to A59 in the following table:
[0037]
[0038] (b) a risk assessment module, wherein the risk assessment module is configured to input the input methylation gene marker data into the cancer risk prediction model constructed by the method according to the second aspect of the present invention to perform risk assessment, thereby obtaining a risk assessment result:
[0039] (3) an output module, wherein the output module is configured to output the cancer risk assessment result.
[0040] In another preferred embodiment, the methylation gene marker further includes a set of base sequences coded as B1 to B41 in the following table:
[0041]
[0042] In another preferred embodiment, the risk assessment module performs the following risk assessment:
[0043] When the input methylation marker is an upregulated marker (an upregulated marker in Table A and / or Table B) and is upregulated in the sample to be tested, and when the input methylation marker is a downregulated marker (a downregulated marker in Table A and / or Table B) and is downregulated in the sample to be tested, it indicates that the subject to be tested has a high risk of cancer; otherwise, it indicates that the risk of cancer is not high.
[0044] In another preferred example, the device further includes a sequencing module, which is configured to extract and sequence cfDNA from the sample to be tested to obtain read data.
[0045] In another preferred embodiment, the sample or samples to be tested are body fluid samples of the subject to be tested.
[0046] In another preferred embodiment, the sample to be tested is selected from the following group: a plasma sample, a blood sample, a urine sample, an ascites sample, or a combination thereof.
[0047] In another preferred embodiment, the sequencing is whole genome sequencing.
[0048] In another preferred embodiment, the sequencing is low-depth whole-genome sequencing.
[0049] In another preferred embodiment, the sequencing depth of the sequencing is 1~2X.
[0050] In another preferred example, the device further includes a comparison module, which is configured to compare the read data to a reference genome to obtain the position of the 5' end of the read on the reference genome; and obtain a base sequence obtained by taking n bp upstream, downstream, or upstream and downstream of the 5' end point, with an interval of m bp, thereby obtaining the methylation gene markers (or base sequences) in Table A and / or Table B.
[0051] In another preferred example, the device further includes a statistical module, which is configured to calculate the proportion of cfDNA corresponding to each of the methylation gene markers in the total cfDNA, thereby obtaining methylation gene marker data.
[0052] In a fourth aspect of the present invention, a set of methylation gene markers is provided, wherein the methylation gene marker set is a set of base sequences obtained by taking n bp upstream, downstream, or both upstream and downstream of the 5' breakpoint when cfDNA is aligned to a reference genome, with an interval of m bp;
[0053] The methylation gene marker set includes the set of base sequences coded A1 to A59 in the following table:
[0054]
[0055] In another preferred embodiment, the methylation gene marker set further includes a set of base sequences B1 to B41 in the following table:
[0056]
[0057] In a fifth aspect of the present invention, a computer storage device is provided, wherein the device stores a computer program corresponding to the algorithm of the cancer risk prediction model constructed by the method according to the second aspect of the present invention.
[0058] In a sixth aspect of the present invention, a method for screening methylation gene markers is provided. The method analyzes the differences between the end sequences of fragments in the high-methylation level interval and the low-methylation level interval in healthy human plasma, thereby discovering gene markers with significant differences between the high-methylation level interval and the low-methylation level interval. Specifically, the method comprises the following steps:
[0059] (i) Obtaining hypermethylated and hypomethylated intervals in healthy human chromosomes;
[0060] (ii) extracting cfDNA from healthy human samples and obtaining cfDNA read data;
[0061] (iii) aligning the cfDNA read data to a reference genome to obtain the position of the 5' end of the read on the reference genome; and obtaining cfDNA reads located in the high methylation interval and cfDNA reads located in the low methylation interval, respectively;
[0062] (iv) according to the position of the 5' end of the read segment in step (iii) on the reference genome, starting from the 5' end breakpoint upstream, downstream, and upstream and downstream respectively, at intervals of m bp, and taking n bp to obtain the base sequence; wherein m is any integer from 0 to 5, and n is any integer from 1 to 4;
[0063] (v) Counting the proportion of cfDNA fragments corresponding to each base sequence in the hypermethylated interval to all cfDNA fragments, thereby obtaining the frequency of each base sequence in the hypermethylated interval; counting the proportion of cfDNA fragments corresponding to each base sequence in the hypomethylated interval to all cfDNA fragments, thereby obtaining the frequency of each base sequence in the hypomethylated interval;
[0064] (vi) Calculate the p-value of each base sequence whose frequency in the high-methylation interval is higher than that in the low-methylation interval, perform multiple hypothesis verification, and select base sequences whose Q-value (q-value) is less than a certain threshold; calculate the p-value of each base sequence whose frequency in the high-methylation interval is lower than that in the low-methylation interval, perform multiple hypothesis verification, and select base sequences whose q-value is less than a certain threshold; and
[0065] (vii) Using the Relief method, the base sequences described in step (vi) are filtered and screened, and the base sequences with the largest Relief values are selected to obtain methylation gene markers (i.e., a collection of base sequences with the largest Relief values).
[0066] In another preferred embodiment, the methylation gene marker is used to detect (a) methylation level; and / or (b) cancer risk.
[0067] In another preferred embodiment, the sample is a body fluid sample.
[0068] In another preferred embodiment, the sample is selected from the group consisting of a plasma sample, a blood sample, a urine sample, an ascites sample, or a combination thereof.
[0069] In another preferred embodiment, step (vi) includes the following steps: using the Wilcoxon test to count the pvalue of each base sequence frequency in the high methylation interval that is higher than that in the low methylation interval, using the Benjamini & Hochberg method to perform multiple hypothesis verification, and selecting base sequences with an FDR less than a certain threshold; using the Wilcoxon test to count the pvalue of each base sequence frequency in the high methylation interval that is lower than that in the low methylation interval, using the Benjamini & Hochberg method to perform multiple hypothesis verification, and selecting base sequences with an FDR less than a certain threshold.
[0070] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features described in detail below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be listed here one by one. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 The correlation between the gold standard methylation level and the methylation level calculated based on the tissue (cell) methylation level and the proportion of healthy human cells is shown.
[0072] Figure 2 The correlation between all methylation-related base sequences and methylation levels is shown.
[0073] Figure 3 The correlation between methylation-related base sequences that are not directly related to CGN and NCG and methylation levels is shown.
[0074] Figure 4 The correlation between CGN, NCG and methylation levels is shown.
[0075] Figure 5 The performance of methylation-related base sequences in cancer prediction on the whole genome is shown.
[0076] Figure 6 The performance of methylation-related base sequences in cancer prediction in the CpG island and CpG shore intervals is shown.
[0077] Figure 7 The performance of methylation-related base sequences in cancer prediction in the CpG island interval is shown. DETAILED DESCRIPTION
[0078] After extensive and in-depth research, the inventors have discovered for the first time, through a large number of screenings, base sequences with significant differences between high and low methylation intervals, and the base sequences can be used to judge the level of methylation and / or the risk of cancer. Specifically, the experiments of the present invention show that the model constructed based on the 59 methylation gene markers of the present invention has good cancer risk prediction performance in the whole genome, CpG island and CpG shore intervals, or CpG island intervals, and the AUC values verified are 0.98, 0.96, and 0.91, respectively. On this basis, the present invention has been completed.
[0079] Furthermore, the 59 methylation gene markers of the present invention can be combined with 41 base sequences directly related to CGNs and NCGs to form gene markers for determining methylation levels and / or cancer risk. The model constructed based on these 100 methylation gene markers of the present invention also demonstrated good cancer risk prediction performance across the entire genome, within the CpG island and CpG shore intervals, or within the CpG island interval, with validated AUC values of 0.97, 0.95, and 0.91, respectively.
[0080] the term
[0081] In order to more easily understand the present disclosure, some terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms should have the meaning given below. Other definitions are set forth throughout the application.
[0082] As used herein, the terms "comprising" or "including" may be open, semi-closed, or closed. In other words, the terms also include "consisting essentially of" or "consisting of."
[0083] As used herein, the term "reference value" or "control reference value" refers to a value that is statistically correlated with a particular outcome when compared to the assay results. Research from the literature and user experience with the methods disclosed herein can also be used to generate or adjust reference values. Reference values can also be determined by considering circumstances and outcomes particularly relevant to the patient's ethnic group, medical history, genetics, age, and other factors.
[0084] As used herein, "CpG shore" refers to the interval extending 2 kb upstream and downstream of the CpG island.
[0085] As used herein, "CGNNCG" refers to the 5'-terminal 3 bp bases that start with CG or end with CG.
[0086] As used herein, "taking a base sequence downstream" includes this position.
[0087] As used herein, "taking a sequence upstream" does not include this position.
[0088] CGNNCG characterization of methylation: The study "Epigenetic analysis of cell-free DNA by fragmentomic profiling" published in PANS in 2022 showed that the 5' CGN / NCG motifratio of whole-genome cfDNA was verified to be positively correlated with the methylation level, and the "CGN / NCG motif ratio" can be used to characterize the methylation level; using CGN and NCG (including 8 base sequences, namely CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG), cancer can be predicted.
[0089] This method calculates the differences in the splicing ratios of methylated and unmethylated CpG sites and their ends in a sample, inferring and verifying that the splicing ratio of methylated CG at the CG position is higher than that of unmethylated CG, while the splicing ratio of methylated CG at the position 1 bp before the CG is lower than that of unmethylated CG. Therefore, the 5' end CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG are used to characterize the methylation status. Since the methylation status of a sample is closely related to cancer, these eight base sequences can be used to predict cancer. The Gradient Boosting Machine (GBM) algorithm is a common type of algorithm in machine learning. Its basic principle is to train newly added weak classifiers based on the negative gradient information of the current model loss function, and then cumulatively combine the trained weak classifiers into the existing model to obtain the optimal model. This model has the advantages of good training effect and low overfitting.
[0090] The methylation level in the present invention refers to the degree of DNA methylation at a specific site or region in an organism. Both a significant increase or decrease in the methylation level indicates an abnormal methylation level. Abnormal DNA methylation is closely related to the occurrence, development, and canceration of tumors.
[0091] In tumors, the promoter regions of tumor suppressor genes are often in a hypermethylated state, leading to gene silencing and uncontrolled cell proliferation, thereby increasing the risk of cancer.
[0092] In the present invention, the terms "methylation gene markers of the present invention", "gene markers of the present invention", "markers of the present invention", or "markers shown in Table A and / or Table B" are used interchangeably to refer to any one or more of the methylation gene markers of the present invention.
[0093] In the present invention, the term "methylation gene marker of the present invention" refers to any base sequence or nucleotide sequence shown in Table A and / or Table B.
[0094] In the present invention, the methylation gene marker includes any one of the base sequences shown in A1 to A59 in Table A below, or a combination thereof:
[0095] Table A
[0096]
[0097] In the present invention, the methylation gene marker further includes any one of the base sequences shown in B1 to B41 in Table B below, or a combination thereof:
[0098] Table B
[0099]
[0100] The main advantages of the present invention include:
[0101] (a) Prior art methods for CGN and NCG research focus on CpGs. These methods, by examining only CpGs and their upstream and downstream sequences, reveal that the ratios of methylated and unmethylated CpGs at the 5' end differ. Methylation-related analyses are then performed based on the CGN and NCG at the 5' end of cfDNA fragments. The present method, however, evaluates the differences in the end sequences of multiple cfDNA fragments between hypermethylated and hypomethylated regions to obtain methylation-related base sequences, providing a more comprehensive and effective method.
[0102] The present invention will be further described below in conjunction with specific examples. It should be understood that these examples are intended to illustrate the present invention only and are not intended to limit the scope of the invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise indicated, percentages and parts are by weight.
[0103] Example 1 Correlation between the actual methylation level and the methylation level calculated based on the tissue (cell) methylation level and the proportion of healthy human cells
[0104] Based on the database of methylation levels of 5,820 units in different tissues provided in the existing technical literature ("DNA tissue mapping by genome-wide methylation sequencing for noninvasive prenatal, cancer, and transplantation assessments") and the proportion of tissues (cells) of different sources of cfDNA in healthy human plasma, the methylation level of each unit in healthy human plasma cfDNA was calculated, and units with the same methylation level were merged into the same methylation level range.
[0105] For ease of description, the "methylation level of healthy human cfDNA calculated based on tissue (cell) methylation levels and the proportion of different cells in healthy individuals" will be referred to as the "methylation level" in this article. Whole-genome methylation sequencing data from healthy human plasma samples will also be analyzed. The methylation level within the aforementioned methylation range will be calculated as the "gold standard methylation level," and the correlation between the two will be compared.
[0106] In this example, whole-genome methylation sequencing of healthy human plasma cfDNA samples was performed using bisulfite methylation, resulting in 51.93 GB of sequencing data. Bisulfite treatment maintains methylated Cs in CpGs, while unmethylated Cs are converted to Ts.
[0107] The specific steps are as follows:
[0108] (1) Based on the methylation database information and the proportion of different tissues (cells) of cfDNA in plasma, the methylation level of each unit in healthy people is calculated. The calculation formula is:
[0109]
[0110] in is the methylation level of tissue (cell) 𝑖, is the proportion of tissue (cell) 𝑖 in plasma cfDNA, is the coefficient of tissue (cell) 𝑖. In this example, the coefficients are all 1, and the proportions of neutrophils, B cells, T cells, liver tissue, colon tissue, heart tissue, brain tissue, and lung tissue in the plasma cfDNA of healthy people are set to 0.667, 0.164, 0.175, 0.058, 0.04, 0.035, 0.031, and 0.024 respectively;
[0111] (2) Group the units according to the methylation level. Units with methylation levels <= 10 are merged into the M0_10 interval; units with methylation levels > 10 and <= 20 are merged into the M10_20 interval; units with methylation levels > 20 and <= 30 are merged into the M20_30 interval; units with methylation levels > 30 and <= 40 are merged into the M30_40 interval; units with methylation levels > 40 and <= 50 are merged into the M40_50 interval; units with methylation levels > 50 and <= 60 are merged into the M50_60 interval; units with methylation levels > 60 and <= 70 are merged into the M60_70 interval; units with methylation levels > 70 and <= 80 are merged into the M70_80 interval; units with methylation levels > 80 and <= 90 are merged into the M80_90 interval; units with methylation levels > 90 are merged into the M90 interval;
[0112] (3) Calculate the average methylation level of each methylation level interval as the methylation level of the interval;
[0113] (4) The methylation level of CpG in plasma cfDNA of healthy subjects was calculated using the software Msuit2 (v2.2.0), with the reference sequence hg19 and default parameters selected to obtain the methylation information of each CpG;
[0114] (5) Count the support numbers of reads in which C is detected as C or T in CpG sites within the intervals of M0_10, M10_20, M20_30, M30_40, M40_50, M50_60, M60_70, M70_80, M80_90 and M90 respectively; count the support numbers of reads in which C is detected as C in CpG sites within the intervals of M0_10, M10_20, M20_30, M30_40, M40_50, M50_60, M60_70, M70_80, M80_90 and M90 respectively;
[0115] (6) Calculate the gold standard methylation levels in the intervals M0_10, M10_20, M20_30, M30_40, M40_50, M50_60, M60_70, M70_80, M80_90, and M90 using the following formula:
[0116] Gold standard methylation level = NC / (NC+NT);
[0117] NC is the number of reads supporting that C in the CpG site is detected as C, and NT is the number of reads supporting that C in the CpG site is detected as T.
[0118] (7) Compare the Pearson correlation between the methylation levels described in step (3) and the gold standard methylation described in step (6).
[0119] When calculating the methylation level for each cell, we considered the predominance of white blood cells in cfDNA and the wide variation in their proportion (mean 79%, IQR 74%-88%). To minimize the impact of fluctuations in the proportion of white blood cells in healthy individuals' plasma, the coefficient for white blood cells (neutrophils, B cells, and T cells) was set to 1. When calculating methylation levels in healthy individuals, this coefficient may be greater than 1. The results are shown in Table 1.
[0120] Table 1
[0121]
[0122] The correlation between the gold standard methylation level and the methylation level is 0.998. The methylation level calculated based on the tissue (cell) methylation level and the proportion of healthy human cells can be used to represent the true methylation level, such as Figure 1 shown.
[0123] Example 2 Marker Screening
[0124] To investigate the differences between high-methylation and low-methylation intervals at the end of fragments in healthy human plasma, we first calculated the methylation levels of cfDNA from healthy human plasma in a database of 5,820 units provided in the prior art literature ("DNA tissue mapping by genome-wide methylation sequencing for noninvasive prenatal, cancer, and transplantation assessments"). Units with a methylation level >= 0.8 were considered hypermethylated, while units with a methylation level <= 0.1 were considered hypomethylated. Highly methylated units were then merged into high-methylation intervals, and low-methylated units into low-methylation intervals. Finally, the differences between various base sequences in the high-methylation and low-methylation intervals were calculated, and the top 100 base sequences with significant differences were selected as methylation prediction markers.
[0125] The specific steps are as follows:
[0126] (1) Based on the methylation database information and the proportion of different tissues (cells) of cfDNA in plasma, the methylation level of each unit in healthy people is calculated. The calculation formula is:
[0127]
[0128] in is the methylation level of tissue (cell) 𝑖, is the proportion of tissue (cell) 𝑖 in plasma cfDNA, is the coefficient of tissue (cell) 𝑖. In this example, the coefficients are all 1, and the proportions of neutrophils, B cells, T cells, liver tissue, colon tissue, heart tissue, brain tissue, and lung tissue in healthy human plasma cfDNA are set to 0.667, 0.164, 0.175, 0.058, 0.04, 0.035, 0.031, and 0.024, respectively.
[0129] (2) The unit is defined as a hypermethylated or hypomethylated unit according to the methylation level. When the methylation level is greater than or equal to 0.8, the unit is considered hypermethylated, and when the methylation level is less than or equal to 0.1, the unit is considered hypomethylated.
[0130] (3) Merge all low-methylation units into low-methylation intervals, and merge all high-methylation units into high-methylation intervals, in the format of bed;
[0131] (4) 28 healthy human plasma samples were selected for low-depth whole-genome sequencing data, with a sequencing depth of 1-2X. The obtained fastq was filtered and aligned with the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software picard to obtain the deduplicated alignment files. The alignment information with low alignment quality was removed, and the alignment information on chromosomes X and Y was removed. The position information of the cfDNA molecules aligned to the chromosomes was extracted and converted into bed format;
[0132] (5) Using the software bedtools, the position information of the cfDNA molecules of the samples described in step (4) on the chromosome is compared with the position information of the high methylation interval described in step (3), and the cfDNA fragments of the 28 healthy human samples located in the high methylation interval are obtained; the position information of the cfDNA molecules of the samples described in step (4) on the chromosome is compared with the position information of the low methylation interval described in step (3), and the cfDNA fragments of the 28 healthy human samples located in the low methylation interval are obtained;
[0133] (6) extracting the 5' ends of the cfDNA fragments located in the high methylation interval and the low methylation interval described in step (5) respectively, and finding their positions on the chromosomes;
[0134] (7) According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively taken at intervals of 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the downstream direction (5' to 3' direction); According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively taken at intervals of 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the upstream direction (3' to 5' direction); According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively taken in the upstream and downstream directions;
[0135] (8) calculating the frequencies of the various base sequences described in step (7);
[0136] (9) The Wilcoxon test was used to calculate the p-value of each base sequence whose frequency in the high methylation interval was higher than that in the low methylation interval. The Benjamini & Hochberg method was used to perform multiple hypothesis verification and the base sequences with FDR < 0.05 were selected. The Wilcoxon test was used to calculate the p-value of each base sequence whose frequency in the high methylation interval was lower than that in the low methylation interval. The Benjamini & Hochberg method was used to perform multiple hypothesis verification and the base sequences with FDR < 0.05 were selected.
[0137] (10) Using the Relief method, the base sequences described in step (9) were filtered and screened, and the 100 base sequences with the largest Relief values were selected as methylation-related markers.
[0138] The results show that the length of the hypomethylated interval is 0.543M and the length of the hypermethylated interval is 0.480M. The detailed information of the 100 base sequences is shown in Table 2, of which 41 base fragments are directly related to the 5' end CGN and NCG.
[0139] Table 2 Base sequence information of 100 methylation predictions
[0140]
[0141] Example 3 Correlation between markers and methylation levels
[0142] Another 30 healthy human samples were used to verify the correlation between the markers and methylation levels. The specific steps are as follows:
[0143] (1) Calculate the methylation level of each unit in a healthy person according to step (1) in Example 1;
[0144] (2) Group the units according to the methylation level. Units with methylation levels <= 10 are merged into the M0_10 interval; units with methylation levels > 10 and <= 20 are merged into the M10_20 interval; units with methylation levels > 20 and <= 30 are merged into the M20_30 interval; units with methylation levels > 30 and <= 40 are merged into the M30_40 interval; units with methylation levels > 40 and <= 50 are merged into the M40_50 interval; units with methylation levels > 50 and <= 60 are merged into the M50_60 interval; units with methylation levels > 60 and <= 70 are merged into the M60_70 interval; units with methylation levels > 70 and <= 80 are merged into the M70_80 interval; units with methylation levels > 80 and <= 90 are merged into the M80_90 interval; units with methylation levels > 90 are merged into the M90 interval; methylation level intervals with a length less than 0.2M are filtered out;
[0145] (3) Calculate the average methylation level of each methylation level interval as the methylation level of the interval;
[0146] (4) 30 healthy human plasma samples were selected for low-depth whole-genome sequencing data, with a sequencing depth of 1-2X. The obtained fastq was filtered and aligned with the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software picard to obtain the deduplicated alignment files. The alignment information with low alignment quality was removed, and the alignment information on chromosomes X and Y was removed. The position information of the cfDNA molecules aligned to the chromosomes was extracted and converted into bed format;
[0147] (5) Using the software bedtools, the position information of the cfDNA molecules of the samples described in step (4) on the chromosomes was compared with the position information of the different methylation level intervals described in step (2), and the cfDNA fragments of the samples of 30 healthy people located in different methylation level intervals were obtained;
[0148] (6) extracting the positions of the 5' ends of the cfDNA fragments located in different methylation level intervals on the chromosomes respectively;
[0149] (7) According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the downstream direction (5' end to 3' end); According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the upstream direction (3' end to 5' end); According to the position described in step (6), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated in the upstream and downstream directions;
[0150] (8) calculating the frequencies of the various base sequences described in step (7) for each methylation level interval;
[0151] (9) Extract the frequencies of the 100 base sequences listed in Table 2, add the frequencies of base sequences that are “upregulated” in the hypermethylated interval relative to the hypomethylated interval to obtain the sum of the frequencies of upregulated base sequences (SU), and add the frequencies of base sequences that are “downregulated” in the hypermethylated interval relative to the hypomethylated interval to obtain the sum of the frequencies of downregulated base sequences (SD). Calculate SU / SD, and use the Spearman method to calculate the correlation between SU / SD and methylation levels.
[0152] (10) Extract the frequencies of 51 base sequences that have no direct relationship with CGN and NCG in the base sequences listed in Table 2, add the frequencies of base sequences that are “upregulated” in the high methylation interval relative to the low methylation interval to obtain the sum of the frequencies of upregulated base sequences SU, and add the frequencies of base sequences that are “downregulated” in the high methylation interval relative to the low methylation interval to obtain the sum of the frequencies of downregulated base sequences SD, calculate SU / SD, and use the Spearman method to calculate the correlation between SU / SD and methylation level;
[0153] (11) Extract the frequencies of the eight base sequences at the 5' end: CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG. Add the frequencies of CGA, CGC, CGT, and CGG to obtain the sum of base sequence frequencies (SU). Add the frequencies of ACG, CCG, CGT, and GCG to obtain the sum of base sequence frequencies (SD). Calculate SU / SD. Use the Spearman method to calculate the correlation between SU / SD and methylation levels.
[0154] In the results, M50_60, M60_70, M70_80, and M80_90 were less than 0.2 M in length and were excluded from the analysis. The methylation levels of M0_M10, M10_M20, M20_M30, M30_M40, M40_M50, and M90 were 6.11, 14.86, 24.66, 34.69, 44.84, and 104.51, respectively. Table 3 shows the frequencies of predicted methylation base sequences in sample S1 at different methylation level intervals. The SU, SD, and SU / SD values for different numbers of base sequences were calculated based on Table 3, as shown in Table 4.
[0155] Table 3 Frequency of methylation-predicted base sequences in different methylation level ranges in sample S1
[0156]
[0157] Table 4 Values of SU, SD, and SU / SD in different methylation levels in sample S1
[0158]
[0159] From the results, we can see that the correlation between SU / SD and methylation level of all base fragments of 30 samples is 0.977 ( Figure 1 ), the correlation between SU / SD and methylation level of base fragments not related to CGN and NCG is 0.970 ( Figure 2 ), the correlation between SU / SD of CGNNCG and methylation level was 0.940 ( Figure 3 ). When performing differential analysis between high and low methylation intervals, the chromosome locations of high and low methylation intervals are different.
[0160] The results indicate that: (1) the base sequences with significant frequency differences are caused by methylation rather than by the difference in chromosomal location between high-methylation and low-methylation intervals, and that these base sequences are related to methylation. (2) The correlation between the SU / SD ratio of base segments unrelated to CGN and NCG and methylation levels is higher than that of CGNNCG.
[0161] Example 4 Evaluation of the performance of markers for cancer prediction on a genome-wide basis
[0162] The 58 healthy human plasma samples and 40 cancer plasma samples (10 gastric cancer, 10 breast cancer, 10 colon cancer, and 10 bile duct cancer) from Examples 2 and 3 were used as training sets, and another 87 healthy human plasma samples and 60 cancer plasma samples (15 gastric cancer, 15 breast cancer, 15 colon cancer, and 15 bile duct cancer) were used as test sets to verify the performance of methylation-related base fragments as markers for early cancer screening. The specific steps are as follows:
[0163] (1) Low-depth whole-genome sequencing data was performed on 245 plasma samples, with a sequencing depth of 1-2X. The obtained fastq was filtered and aligned with the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software picard to obtain the deduplicated alignment files. Alignment information with low alignment quality was removed, and alignment information to chromosomes X and Y was removed. The position information of the cfDNA molecules aligned to the chromosomes was extracted and converted into bed format;
[0164] (2) the location of the 5' end of the extracted cfDNA fragment on the chromosome;
[0165] (3) The positions described in step (2) were spaced 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp apart.
[0166] bp downstream (5' to 3' direction) takes 1 bp base, 2 bp base, 3 bp base and 4 bp base; according to the position described in step (2), 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp are respectively spaced upstream (3' to 5' end); according to the position described in step (2), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively taken upstream and downstream;
[0167] (4) Calculate the frequencies of various base sequences in step (3);
[0168] (5) Extract the frequencies of the 100 base sequences in Table 2 from the training set as input values to build the model;
[0169] (6) The gbm algorithm is used to build the model, where the parameters of the trainControl function are:
[0170] "method = "repeatedcv",number = 10, repeats = 10, verboseIter =FALSE, savePredictions=TRUE,classProbs=TRUE,summaryFunction =twoClassSummary";
[0171] The parameters of the train function for training the model are: "method = 'gbm', tuneGrid = data.frame(n.trees=150, interaction.depth=3, shrinkage=0.1, n.minobsinnode=10), preProcess = c("corr", "nzv"))"
[0172] (7) Extract the frequencies of the 100 base sequences in Table 2 from the test set and input them into the model constructed in step (6) to predict health and malignancy;
[0173] (8) Extract the frequencies of 51 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 in the training set as input values and build the model according to step (6);
[0174] (9) Extract the frequencies of 59 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 of the test set and input them into the model constructed in step (8) to predict health and malignancy;
[0175] (10) Extract the frequencies of eight base motifs, CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG, at the 5' end of the training set as input values and construct the model according to step (6);
[0176] (11) Extract the frequencies of the eight base sequences CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG at the 5' end of the test set and input them into the model constructed in step (10) to predict health and malignancy;
[0177] The results showed that when all methylation-related base sequences were used as markers for prediction in the test set, the AUC was 0.97. When 51 base sequences that were not directly related to CGN and NCG were used as markers for prediction, the AUC was 0.98. When 8 base sequences of 5' end CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG were used as markers for prediction, the AUC was 0.90. The performance was significantly improved. Figure 5 shown.
[0178] Example 5 Evaluation of the performance of markers for cancer prediction at CpG island and CpG shore intervals
[0179] The performance of the markers in cancer prediction at the CpG island and CpG shore levels was evaluated using the training and test sets in Example 3. The specific steps were as follows:
[0180] (1) Low-depth whole-genome sequencing data was performed on 245 plasma samples with a sequencing depth of 1-2X. The obtained fastq was filtered and aligned with the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software picard to obtain the deduplicated alignment files. Alignment information with low alignment quality was removed, and alignment information to chromosomes X and Y was removed. The position information of the cfDNA molecules aligned to the chromosomes was extracted and converted into bed format;
[0181] (2) Using the software bedtools, the position information of the sample cfDNA molecules on the chromosome described in step (1) was compared with the position information of the CpG island and CpG shore intervals, and the cfDNA fragments of 245 samples located in the CpG island and CpG shore intervals were obtained;
[0182] (3) The location of the 5' end of the cfDNA fragment extracted in step (2) on the chromosome;
[0183] (4) According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the downstream direction (5' to 3' direction); According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the upstream direction (3' to 5' direction); According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated in the upstream and downstream directions;
[0184] (5) calculating the frequencies of the various base sequences described in step (4);
[0185] (6) Extract the frequencies of the 100 base sequences in Table 2 from the training set as input values and construct the model according to step (6) in Example 4;
[0186] (7) Extract the frequencies of the 100 base sequences in Table 2 from the test set and input them into the model constructed in step (6) to predict health and malignancy;
[0187] (8) Extract the frequencies of 51 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 in the training set as input values, and construct the model according to step (6) in Example 4;
[0188] (9) Extract the frequencies of 51 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 of the test set and input them into the model constructed in step (8) to predict health and malignancy;
[0189] (10) Extract the frequencies of eight base motifs (CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG) at the 5' end of the training set as input values, and construct the model according to step (6) in Example 4;
[0190] (11) The frequencies of eight base sequences, namely CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG, at the 5' end of the test set were extracted and input into the model constructed in step (10) to predict health and malignancy.
[0191] The results showed that when all methylation-related base sequences were used as markers for prediction in the test set, the AUC was 0.95. When 59 base sequences that were not directly related to CGN and NCG were used as markers for prediction, the AUC was 0.96. When 8 base sequences of 5' end CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG were used as markers for prediction, the AUC was 0.77. The performance was significantly improved. Figure 6 shown.
[0192] Example 6 Evaluation of the performance of markers for cancer prediction in the CpG island interval
[0193] The performance of the markers in cancer prediction on the CpG island interval was evaluated using the training and test sets in Example 4. The specific steps were as follows:
[0194] (1) 245 plasma samples were sequenced for low-depth whole-genome sequencing data with a sequencing depth of 1-2X. The obtained fastq was filtered and aligned with the human genome using the software BWA. The alignment files were sorted using the software samtools to obtain the sorted alignment files. The sorted alignment files were deduplicated using the software picard to obtain the deduplicated alignment files. Alignment information with low alignment quality was removed, and alignment information to chromosomes X and Y was removed. The position information of the cfDNA molecules aligned to the chromosomes was extracted and converted into bed format;
[0195] (2) Using the software bedtools, the position information of the cfDNA molecules of the samples described in step (1) on the chromosome was compared with the position information of the CpG island interval, and the cfDNA fragments of 245 samples located in the CpG island interval were obtained;
[0196] (3) The location of the 5' end of the cfDNA fragment extracted in step (2) on the chromosome;
[0197] (4) According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the downstream direction (5' to 3' direction); According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 0 bp, 1 bp, 2 bp, 3 bp, 4 bp and 5 bp in the upstream direction (3' to 5' direction); According to the position described in step (3), 1 bp base, 2 bp base, 3 bp base and 4 bp base are respectively separated by 1 bp, 2 bp base, 3 bp base and 4 bp base in the upstream and downstream direction;
[0198] (5) calculating the frequencies of the various base sequences described in step (4);
[0199] (6) Extract the frequencies of the 100 base sequences in Table 2 from the training set as input values and construct the model according to step (6) in Example 4;
[0200] (7) Extract the frequencies of the 100 base sequences in Table 2 from the test set and input them into the model constructed in step (6) to predict health and malignancy;
[0201] (8) Extract the frequencies of 59 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 in the training set as input values, and construct the model according to step (6) in Example 4;
[0202] (9) Extract the frequencies of 59 base sequences that have no direct relationship with CGN and NCG from the base sequences in Table 2 of the test set and input them into the model constructed in step (8) to predict health and malignancy;
[0203] (10) Extract the frequencies of eight base motifs (CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG) at the 5' end of the training set as input values, and construct the model according to step (6) in Example 4;
[0204] (11) The frequencies of the eight base sequences CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG at the 5' end of the test set were extracted and input into the model constructed in step (10) to predict health and malignancy.
[0205] The results showed that when all methylation-related base sequences were used as markers for prediction in the test set, the AUC was 0.91. When 51 base sequences that were not directly related to CGN and NCG were used as markers for prediction, the AUC was 0.91. When 8 base sequences of 5' end CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG were used as markers for prediction, the AUC was 0.74. The performance was significantly improved. Figure 7 shown.
[0206] From Example 4 to Example 6, it can be seen that the performance of 59 methylation-related base sequences as markers for cancer prediction is significantly improved compared with the existing technology CGNNCG.
[0207] Example 7 Results of marker determination on the whole genome, CpG island, CpG shore and CpG island
[0208] In this example, the present inventors used 59 base sequences of the present invention, 59 base sequences, and 41 base sequences directly related to CGN and NCG (100 base sequences), as well as 8 CGNNCG base sequences, to perform judgments on three randomly selected individual plasma sample data (sample 001, sample 002, and sample 003) from the test set, respectively, on the whole genome, CpG island, CpG shore, and CpG island. When the specificity was 0.95, the judgment results were shown in Tables 5, 6, and 7, respectively.
[0209] Table 5 Judgment results of 59 base sequences
[0210]
[0211] Table 6 Judgment results of 100 base sequences
[0212]
[0213] Table 7 Judgment results of 8 base sequences of CGNNCG
[0214]
[0215] As shown in Tables 5-7, the 59 base sequences of the present invention, the 59 base sequences, and the 41 base sequences directly related to CGN and NCG (100 base sequences) all have very good judgment accuracy, and can give correct judgment results on the whole genome, CpG island, CpG shore, and CpG island. The 8 CGNNCG base sequences have good judgment results on the whole genome, second best judgment results on CpG island and CpG shore, and poor judgment results on CpG island.
[0216] It can be seen that the 59 base sequences of the present invention (sequences coded as A1 to A59) have very good cancer risk judgment effects.
[0217] The existing "CGNNCG modif ratio" method, by statistically analyzing the differences in the splicing ratios of methylated and unmethylated CpG sites and their respective ends in a sample, infers and verifies that the splicing ratio of methylated CG at the CG position is higher than that of unmethylated CG, while the splicing ratio of methylated CG at the position 1 bp before the CG is lower than that of unmethylated CG. Therefore, the 5' end CGA, CGC, CGT, CGG, ACG, CCG, CGT, and GCG are used to characterize the methylation status. Because the methylation status of a sample is closely related to cancer, these eight base sequences can be used to predict cancer. While biological and physiological activities are known to be complex, this method only extracts the eight CpG-related 5' end base sequences and does not analyze and verify the correlation between other base sequences at the end of cfDNA and methylation.
[0218] “Plasma DNA tissue mapping by genome-wide methylation sequencing for noninvasive prenatal, cancer, and transplantation assessments” published in PANS PLUS in 2015 studied the methylation levels of 14 tissues (liver, lung, colon, small intestine, pancreas, adrenal gland, esophagus, adipose tissue, heart, brain, T cells, B cells, neutrophils, and placenta) in the CpG island and CpG shore intervals. The CpG island and CpG shore intervals were divided into non-overlapping 500 bp units, and 5,820 units with significant differences in different tissues (cells) and the methylation levels of the 14 tissues in each unit were obtained. A 2023 study, "The origin of highly elevated cell-free DNA in healthy individuals and patients with pancreatic, colorectal, lung, or ovarian cancer," published on CancerDiscov, found that cfDNA in healthy individuals primarily originates from white blood cells, accounting for an average of 79%. Neutrophils account for 66.7% of this cfDNA, while B cells and T cells contribute 16.4 ± 12.0% and 17.5 ± 4.3%, respectively. The remaining cfDNA originates from the liver, colon, heart, brain, and lung, accounting for 5.8 ± 6.6%, 4.0 ± 7.1%, 3.5 ± 3.4%, 3.1 ± 3.3%, and 2.4 ± 3.4%, respectively. Therefore, we calculated the methylation levels of cfDNA in healthy individuals based on the methylation levels of tissues (cells) and the proportions of different cell types in healthy individuals. We then identified chromosomal regions with varying methylation levels and compared the base sequence differences between low-methylation and high-methylation regions to identify base sequences that can predict methylation levels.
[0219] First, the inventors evaluated the correlation between the gold standard methylation levels of healthy individuals obtained by whole-genome methylation sequencing and the methylation levels calculated based on tissue methylation levels and the percentage of healthy cells. The results showed a Pearson correlation coefficient of 0.998, indicating that the methylation levels calculated based on tissue (cell) methylation levels and the percentage of healthy cells can be used to represent the true methylation levels. The inventors then traversed various base sequences at the end of cfDNA fragments and compared their frequency differences in chromosomal intervals with high methylation levels and low methylation levels in 28 healthy individuals. The top 100 base sequences with significant differences were selected, of which 41 were directly related to 5'-end CGN and NCG. A further 30 healthy human plasma samples were then used to demonstrate a strong correlation between the frequencies of the significantly different base sequences and methylation levels. Sixty cancer plasma samples and 87 healthy human plasma samples were used as a test set for validation.
[0220] The results showed that when the frequencies of the 100 base sequences shown in Table 1 were used as input values to construct the model at the whole genome level, the AUC was 0.97; when the 41 base sequences related to CGN and NCG were removed and the frequencies of the remaining 59 base sequences were used as input values to construct the model, the AUC was 0.98; when the frequencies of CGA, CGC, CGT, CGG, ACG, CCG, CGT and GCG were used as input values to construct the model, the AUC was 0.90.
[0221] All documents mentioned in this application are incorporated herein by reference, just as if each document were incorporated herein by reference individually. It should also be understood that after reading the above teachings of the present invention, those skilled in the art may make various changes or modifications to the present invention, and that such equivalents also fall within the scope of the claims appended hereto.
Claims
1. A use of a detection reagent for a methylation gene marker, characterized in that: For preparing a detection reagent or kit, the detection reagent or kit is used for (a) detecting methylation levels; and / or (b) detecting cancer risk, wherein the detection of cancer risk is to perform health and cancer prediction, and the cancer is selected from the following group: gastric cancer, breast cancer, colon cancer, bile duct cancer, or a combination thereof; The methylation gene marker is a set of base sequences obtained by taking n bp upstream, downstream, or upstream and downstream of the 5' breakpoint when cfDNA is aligned to the reference genome, with an interval of m bp, and taking n bp; m is any integer from 0 to 5, and n is any integer from 1 to 4; The methylation gene markers include a set of base sequences coded as A1 to A59 in the following table: ; When the direction of the methylation gene marker is upstream and downstream, the corresponding base length in the table is 2*n; Among them, the detection reagent or detection kit detects the risk of cancer by detecting the following methylation gene marker data: the frequency of the methylation gene marker, and the frequency of the methylation gene marker is the proportion of the cfDNA fragments corresponding to the methylation gene marker in the sample to all cfDNA fragments.
2. The use according to claim 1, characterized in that The methylation gene markers also include a set of base sequences coded as B1 to B41 in the following table: 。 3. The use according to claim 1, characterized in that The cancers are: gastric cancer, breast cancer, colon cancer and bile duct cancer.
4. A method for constructing a cancer risk prediction model, characterized in that: The steps include: (s1) providing a first data set for model training and a second data set for model testing, wherein the first data set includes methylation gene marker data of a cancer-positive sample group and methylation gene marker data of a cancer-negative sample group; The second data set includes methylation gene marker data of a cancer-positive sample group and methylation gene marker data of a cancer-negative sample group; and (s2) constructing a cancer risk prediction model using a machine learning method based on the data in the first dataset, and validating the model on the second dataset, thereby obtaining the cancer risk prediction model; The methylation gene marker data is the frequency of the methylation gene marker; the frequency of the methylation gene marker is the proportion of the cfDNA fragments corresponding to the methylation gene marker in the sample to all cfDNA fragments; The methylation gene marker data is obtained by the following steps: (i) extracting cfDNA from the sample to be tested and obtaining read data; (ii) aligning the read data to a reference genome to obtain the position of the 5' end of the read on the reference genome; (iii) obtaining a base sequence by selecting n base pairs at intervals of m base pairs upstream, downstream, or both upstream and downstream from the 5' breakpoint, based on the position of the 5' end of the read on the reference genome; m is an integer from 0 to 5, and n is an integer from 1 to 4; thereby obtaining a set of base sequences coded A1 to A59 in the following table: ; and (iv) counting the proportion of cfDNA fragments corresponding to each base sequence in the base sequence set in all cfDNA fragments, thereby obtaining the frequency of each base sequence, thereby obtaining the methylation gene marker data; The cancer risk prediction is health and cancer prediction, and the cancer is selected from the following group: gastric cancer, breast cancer, colon cancer, bile duct cancer, or a combination thereof.
5. The method according to claim 4, wherein The machine learning method is a gradient boosting algorithm.
6. The method according to claim 4, wherein The sample to be tested is selected from the group consisting of a plasma sample, a blood sample, a urine sample, an ascites sample, or a combination thereof.
7. A cancer risk prediction device, characterized in that: The device comprises: (a) An input module configured to input methylation gene marker data in a sample to be tested; the methylation gene markers include a set of base sequences coded as A1 to A59 in the following table: ; (b) a risk assessment module, wherein the risk assessment module is configured to input the input methylation gene marker data into the cancer risk prediction model constructed according to the method of claim 4 to perform risk assessment, thereby obtaining a risk assessment result: (3) an output module, configured to output the cancer risk assessment result; The methylation gene marker data is the frequency of the methylation gene marker, and the frequency of the methylation gene marker is the proportion of the cfDNA fragments corresponding to the methylation gene marker in the total cfDNA fragments in the sample; The cancer risk prediction is health and cancer prediction, and the cancer is selected from the following group: gastric cancer, breast cancer, colon cancer, bile duct cancer, or a combination thereof.
8. A collection of methylation gene markers, characterized in that: The methylation gene marker set is a set of base sequences obtained by taking n bp upstream, downstream, or upstream and downstream of the 5' breakpoint, with an interval of m bp, when cfDNA is aligned to the reference genome; m is any integer from 0 to 5, and n is any integer from 1 to 4; The methylation gene marker set includes the set of base sequences coded A1 to A59 in the following table: 。 9. The set of methylation gene markers according to claim 8, wherein: The methylation gene markers in the collection are obtained by screening through a method comprising the following steps: (i) Obtaining hypermethylated and hypomethylated intervals in healthy human chromosomes; (ii) extracting cfDNA from healthy human samples and obtaining cfDNA read data; (iii) aligning the cfDNA read data to a reference genome to obtain the position of the 5' end of the read on the reference genome; and obtaining cfDNA reads located in the high methylation interval and cfDNA reads located in the low methylation interval, respectively; (iv) according to the position of the 5' end of the read segment in step (iii) on the reference genome, starting from the 5' end breakpoint upstream, downstream, and upstream and downstream respectively, at intervals of m bp, and taking n bp to obtain the base sequence; wherein m is any integer from 0 to 5, and n is any integer from 1 to 4; (v) Counting the proportion of cfDNA fragments corresponding to each base sequence in the hypermethylated interval to all cfDNA fragments, thereby obtaining the frequency of each base sequence in the hypermethylated interval; counting the proportion of cfDNA fragments corresponding to each base sequence in the hypomethylated interval to all cfDNA fragments, thereby obtaining the frequency of each base sequence in the hypomethylated interval; (vi) Calculate the p-value of each base sequence whose frequency in the high-methylation interval is higher than that in the low-methylation interval, perform multiple hypothesis verification, and select base sequences whose Q-value is less than a certain threshold; calculate the p-value of each base sequence whose frequency in the high-methylation interval is lower than that in the low-methylation interval, perform multiple hypothesis verification, and select base sequences whose q-value is less than a certain threshold; and (vii) Using the Relief method, the base sequence described in step (vi) is filtered and screened, and the base sequence with the largest Relief value is selected to obtain a methylation gene marker.
10. A computer storage device, characterized in that: The device stores a computer program corresponding to the algorithm of the cancer risk prediction model constructed according to the method of claim 4; Wherein, the cancer is selected from the group consisting of gastric cancer, breast cancer, colon cancer, bile duct cancer, or a combination thereof.
Citation Information
Patent Citations
Gene marker for determining symptoms of organism, construction method of detection model and detection device
CN118127167A