Methylation marker identification method for free DNA

By identifying differentially methylated regions between abnormal and normal tissues and combining them with a map of hypomethylated regions specific to normal cell types, the problem of low sensitivity in identifying abnormal tissue types in existing technologies has been solved, achieving efficient methylation marker identification and non-invasive disease diagnosis.

CN120913650APending Publication Date: 2025-11-07CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511006649.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing methods have low sensitivity in identifying differentially methylated regions between abnormal and normal tissues, making it difficult to accurately distinguish abnormal tissue types, and they fail to effectively handle the relationship between abnormal-related DMRs and other tissue-type-specific DMRs.

Method used

By using whole-genome methylation sequencing data, differentially methylated regions between abnormal tissues and normal control tissues were identified, region lengths and fragment numbers were calculated, significant DMRs were screened, and combined with normal cell type-specific hypomethylated region maps, three types of methylation markers were identified: DTU-DMR, DU-DMR, and NT-DMR.

Benefits of technology

It effectively captures subtle methylation differences between abnormal and normal tissues, providing rich methylation information to support nanopore sequencing research. Furthermore, by constructing a model using the gradient booster algorithm, it achieves efficient methylation marker recognition and non-invasive disease diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913650A_ABST
    Figure CN120913650A_ABST
Patent Text Reader

Abstract

The invention discloses a methylation marker identification method for free DNA (deoxyribonucleic acid), which comprises the following steps of: identifying three types of differential methylation regions of abnormal tissues relative to normal control tissues on the basis of the fragment methylation level of whole genome methylation sequencing data; based on the abnormal tissue M-DMR set and the normal control tissue-specific low-methylation region set, identifying a region set in which the tissue-specific low-methylation region is not changed in the abnormality generation process; on the basis of the abnormal tissue U-DMR set and other normal tissue specific low methylation region sets, identifying that a non-tissue specific high methylation region is converted into a region set similar to other normal tissue specific low methylation regions in the abnormality generation process; a set of aberrant DMRs, NT-DMRs, that fall within a non-tissue-specific region on the genome are identified. The novel and accurate methylation mark recognition method is provided, the methylation set area recognized through the method is highly related to the abnormal tissue, and tools and data support can be provided for recognition of the abnormal tissue.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of gene recognition and classification, and particularly relates to a methylation marker recognition method for cell-free DNA. BACKGROUND

[0002] Bisulfite sequencing technology can reveal the methylation status of each cytosine site in the genome. Based on bisulfite sequencing data, researchers have proposed various methods for identifying differentially methylated regions (DMRs) between different tissues or cell types. In medical research based on cell-free DNA (cfDNA), tissue-specific and disease-specific DMRs can be used as key feature markers to trace the tissue origin of cfDNA and determine the disease association. Therefore, improving the quality of identified DMRs is of great significance for biomarker development.

[0003] According to the technical process of the existing DMR detection method, it can be roughly divided into two categories: methods based on differentially methylated CpG sites (DMCs) and methods based on candidate regions. In the method based on DMC, researchers usually use classical hypothesis testing methods such as Fisher's exact test, chi-square test, t-test, and analysis of variance to identify DMCs based on raw methylation signals, estimated methylation signals, or smoothed methylation signals, and further combine adjacent DMCs that meet certain criteria to form DMRs.

[0004] The method based on candidate regions takes genomic regions as the analysis target. In some studies, the candidate region is usually a biologically interesting genomic functional region such as a promoter region or a CpG island, or is defined by a fixed size or a sliding window. In other studies, candidate regions can be generated based on adjacent CpG sites with similar methylation levels or similar methylation change directions in different tissues, and then DMRs are identified by comparing the methylation differences of these candidate regions in different tissues.

[0005] Based on the comparison between multiple tissue types or between abnormal and normal tissues, two types of tissue-specific DMRs or abnormal-related DMRs can be identified respectively. However, the methylation difference of abnormal tissues compared with the corresponding normal tissues is often small, and is also affected by inter-individual heterogeneity, sample cell composition confusion, etc. Therefore, the existing DMR identification method has low sensitivity in identifying abnormal-related DMRs. In addition, the existing method has little research on the relationship between abnormal-related DMRs and other tissue-specific DMRs (including normal control tissue-specific DMRs and other normal tissue-specific DMRs), and non-specific regions in the genome. When the identified abnormal-related differential methylation regions overlap with other tissue-specific methylation regions and have consistent methylation patterns, it is difficult to accurately distinguish abnormal tissue types in the identification application based on free DNA methylation.

[0006] Therefore, it is necessary to design a high-sensitivity method to identify abnormal-related DMRs, and on this basis, to identify high-confidence methylation markers based on free DNA identification by combining the relationship between abnormal-related DMRs and other tissue-specific DMRs (including normal control tissue-specific DMRs and other normal tissue-specific DMRs), and non-specific regions in the genome. SUMMARY

[0007] In view of the deficiencies of the prior art, the purpose of the present application is to provide a methylation marker identification method for free DNA, to construct a high-specificity methylation marker map for various known abnormal tissues, and to provide an analysis tool for medical research.

[0008] The present application provides a methylation marker identification method for free DNA, comprising the following steps:

[0009] S1. Based on the fragment methylation level of whole genome methylation sequencing data, three types of differential methylation regions of abnormal tissues relative to normal control tissues are identified, including differential hypomethylation regions (U-DMR), differential hypermethylation regions (M-DMR) and differential mixed methylation regions (X-DMR), and the reliability and significance of each differential methylation region (DMR) are calculated based on the region length and the number of fragments of this methylation type in the abnormal tissue, and the significant DMRs with p value less than 0.05 are screened;

[0010] S2. Based on the abnormal tissue M-DMR set (labeled as MDMR) and the normal control tissue-specific hypomethylation region set (labeled as TUDMR), a set of regions in which the tissue-specific hypomethylation region is not changed during the abnormal generation process (labeled as DTUDMR) is identified;

[0011] S3. Based on the set of aberrant tissue U-DMRs (labeled as UDMR) and the set of other normal tissue specific hypomethylated regions (labeled as TallUDMR), identify the set of regions (DUDMR) that are non-tissue specific hypermethylated regions that are converted to similar regions as other normal tissue specific hypomethylated regions during the process of aberration;

[0012] S4. Identify the set of aberrant DMRs (NTDMR) that fall within the non-tissue specific regions on the genome, which includes U-DMRs, M-DMRs, and X-DMRs that fall within the non-tissue specific regions, complete the identification of aberrant related DMRs.

[0013] Step S1 includes the following steps:

[0014] Quality control and alignment of WGBS data of the aberrant tissue and normal tissue, using fastp tool to remove sequencing adapters and bases with quality less than 30, using bismark tool to align to hg38 genome, extracting the methylation haplotype of each fragment read, obtaining the read methylation information of the aberrant tissue group (G1) and the normal tissue group sample; the methylation haplotype is represented as a 0-1 string, where 0 represents unmethylated state and 1 represents methylated state;

[0015] According to the read methylation information of the aberrant tissue group (G1) and the normal tissue group (G2) sample, generate candidate regions; first identify effective CpG sites, which are defined as CpG sites covered by at least one sample in each group; based on the condition that the distance between adjacent effective CpG sites is less than 500bp, divide the candidate regions, and finally each divided candidate region contains at least 5 CpG sites;

[0016] Collect the methylation information of all reads of each sample in the candidate region, reads covering at least 4 CpG sites, according to the proportion of methylated and non-methylated CpG sites in the reads, divide each group of reads into three categories: U class (mostly unmethylated) fragments, M class (mostly methylated) fragments and X class (mixed type) fragments; U class fragments contain less than or equal to 25% methylated CpG sites, M class fragments contain more than or equal to 75% methylated CpG sites, and X class fragments contain 25% to 75% methylated CpG sites;

[0017] For each group of samples, the set of sub-regions is further divided in each candidate region based on the intersection of each type of fragment (U, M and X) respectively, for each type of fragment formed sub-regions of abnormal tissue group (G1) and normal tissue group (G2), the sub-region difference set of abnormal tissue group (G1) minus normal tissue group (G2) is calculated, and then the sub-region difference set is intersected with the sub-regions formed by other types of fragments to obtain three types of differential methylation region sets: the differential hypomethylation region set U DMR , the differential hypermethylation region set M DMR and the differential mixed methylation region set X DMR , which are expressed by the following formula:

[0018] U DMR =(U set1 -U set2 )∩(M set2 ∪X set2 )

[0019] M DMR =(M set1 -M set2 )∩(U set2 ∪X set2 )

[0020] X DMR =(X set1 -X set2 )∩(U set2 ∪M set2 )

[0021] Wherein, U set1 is the sub-region set corresponding to the U type fragment of G1; M set1 is the sub-region set corresponding to the M type fragment of G1; X set1 is the sub-region set corresponding to the X type fragment of G1; U set2 is the sub-region set corresponding to the U type fragment of G2; M set2 is the sub-region set corresponding to the M type fragment of G2; X set2 is the sub-region set corresponding to the X type fragment of G2;

[0022] The z-score of each DMR is calculated based on the length of the differential methylation region and the number of fragments of this methylation type in the region of abnormal tissue, and then the p value is calculated based on the z-score, which is calculated by the following formula:

[0023]

[0024] Wherein, Length is the length of the current differential methylation region; μ(Length) is the mean of the length of all DMRs on the same chromosome; σ(Length) is the standard deviation of the length of all DMRs on the same chromosome; z1 is the first z-score value; z2 is the second z-score value; N is the number of fragments corresponding to the methylation type in the current region; σ(N) is the standard deviation of the number of fragments of all DMRs on the same chromosome; μ(N) is the mean of the number of fragments of all DMRs on the same chromosome; Φ is the standard normal distribution cumulative function.

[0025] Step S2 comprises the following steps:

[0026] Obtaining a normal cell type-specific hypomethylation map, extracting a plurality of normal cell types corresponding to specific hypomethylation regions contained in a preset number of normal control tissues (G2) as a specific hypomethylation region set TU of the normal control tissues DMR ;

[0027] Based on the abnormal tissue differential hypermethylation region set M obtained in step S1 DMR and the normal control tissue-specific hypomethylation region set TU DMR , identifying a region set DTU DMR unchanged in the tissue-specific hypomethylation region, which is expressed by the following formula:

[0028] DTU DMR = TU DMR -(TU DMR ∩M DMR )

[0029] Wherein, TU DMR is the normal control tissue-specific hypomethylation region TU-DMR set.

[0030] Step S3 comprises the following steps:

[0031] Removing several normal cell types contained in the normal control tissue in the normal cell type-specific hypomethylation map, obtaining specific hypomethylation regions of other normal cell types, and merging into a set TallU DMR ;

[0032] Based on the abnormal tissue differential hypomethylation region U obtained in step S1 DMR and the set TallU DMR , identifying a region set DU DMR converted from a non-tissue-specific hypermethylation region to a region similar to other normal tissue-specific hypomethylation regions, which is expressed by the following formula:

[0033] DU DMR = U DMR∩TallU DMR .

[0034] Step S4 comprises the following steps:

[0035] The other regions of each normal cell type in addition to the pre-set number of specific hypomethylation regions of each normal cell type in the several normal cell type-specific hypomethylation patterns on the human genome are collectively referred to as non-tissue-specific regions NT region ;

[0036] An abnormal tissue DMR set NT falling within the non-tissue-specific regions is identified DMR , which is expressed using the following formula:

[0037] NTM DMR =(M DMR ∩NT region )

[0038] NTU DMR =(U DMR ∩NT region )

[0039] NTX DMR =(X DMR ∩NT region )

[0040] NT DMR =NTM DMR ∪NTU DMR ∪NTX DMR

[0041] wherein M DMR is an abnormal tissue differential hypermethylation region; U DMR is an abnormal tissue differential hypomethylation region; X DMR is an abnormal tissue differential mixed methylation region; NTM DMR is an abnormal tissue differential hypermethylation region falling within the non-tissue-specific regions; NTU DMR is an abnormal tissue differential hypomethylation region falling within the non-tissue-specific regions; and NTX DMR is an abnormal tissue differential mixed methylation region falling within the non-tissue-specific regions.

[0042] Specifically, the three types of methylation markers, DTU-DMR, DU-DMR and NT-DMR, obtained by the steps S2-S4 are the final abnormality-related DMRs identified.

[0043] The application discloses a methylation marker identification method for free DNA, and effectively captures subtle methylation differences between abnormal tissues and normal tissues by using a candidate region generation strategy and an abnormal and normal control group same type fragment sub-region difference set calculation method, and identifies abnormal related differential methylation regions. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 The figure is a flowchart of the method of the application;

[0045] Figure 2 The figure is a differential methylation region between normal liver tissues and abnormal liver tissues identified based on WGBS data of the normal liver tissues and the abnormal liver tissues in the embodiment of the application; wherein Figure 2 (a) is the distribution of the identified abnormal liver tissue related DMR set in the genome; wherein Figure 2 (b) is the distribution of the abnormal liver tissue related differential hypomethylation region U-DMR set in the genomic functional region; Figure 2 (c) is the distribution of the abnormal liver tissue related differential hypermethylation region M-DMR set in the genomic functional region;

[0046] Figure 3 The figure is a KEGG pathway enrichment mode diagram of the abnormal liver tissue U-DMR related genes in the embodiment of the application;

[0047] Figure 4 The figure is the liver cancer noninvasive diagnosis model classification performance of the abnormal liver tissue DTU-DMR, DU-DMR, NTU-DMR and NTM-DMR in the embodiment of the application. DETAILED DESCRIPTION

[0048] The application provides a methylation marker identification method for free DNA, and the flowchart is as shown in the figure Figure 1 The figure shows that the method comprises the following steps:

[0049] S1. Based on the fragment methylation level of whole genome methylation sequencing data, 3 types of differential methylation regions of abnormal tissues relative to normal control tissues are identified, including differential hypomethylation regions (U-DMRs), differential hypermethylation regions (M-DMRs) and differential mixed methylation regions (X-DMRs), and the reliability and significance of each differential methylation region (DMR) are calculated based on the length of the region and the number of fragments of this methylation type in the region of the abnormal tissue, and the significant DMRs with a p value less than 0.05 are screened;

[0050] The WGBS data of the abnormal tissue and the normal tissue are quality controlled and aligned, the sequencing adapters and the bases with a quality lower than 30 are removed by using the fastp tool, the reads are aligned to the hg38 genome by using the bismark tool, and the methylation haplotype of each fragment read is extracted to obtain the read methylation information of the samples in the abnormal tissue group (G1) and the normal tissue group; the methylation haplotype is represented as a 0-1 string, wherein 0 represents an unmethylated state and 1 represents a methylated state;

[0051] According to the read methylation information of the samples in the abnormal tissue group (G1) and the normal tissue group (G2), candidate regions are generated; first, effective CpG sites are identified, which are defined as CpG sites covered by at least one sample in each group; based on the condition that the distance between adjacent effective CpG sites is less than 500 bp, the candidate regions are divided, and finally each divided candidate region contains at least 5 CpG sites;

[0052] The methylation information of all reads of each sample group falling in the candidate regions is collected, the reads covering at least 4 CpG sites are selected, and each group of reads is divided into three categories according to the proportion of methylated and unmethylated CpG sites in the reads: U-type (mostly unmethylated) fragments, M-type (mostly methylated) fragments and X-type (mixed type) fragments; the U-type fragments contain less than or equal to 25% methylated CpG sites, the M-type fragments contain more than or equal to 75% methylated CpG sites, and the X-type fragments contain 25% to 75% methylated CpG sites;

[0053] For each sample group, a set of sub-regions is further divided in each candidate region based on the overlap of each type of fragment (U-type, M-type and X-type), and for each type of fragment of the abnormal tissue group (G1) and the normal tissue group (G2), the difference set of the sub-regions of the abnormal tissue group (G1) minus the normal tissue group (G2) is calculated, and then the intersection operation is performed between the difference set of the sub-regions and the sub-regions formed by other types of fragments to obtain three types of differential methylation region sets: a differential hypomethylation region U-DMR set U DMR , a differential hypermethylation region M-DMR set M DMR and a differential mixed methylation region X-DMR set XDMR , which is expressed using the following equation:

[0054] U DMR = (U set1 - U set2 )∩(M set2 ∪X set2 )

[0055] M DMR = (M set1 - M set2 )∩(U set2 ∪X set2 )

[0056] X DMR = (X set1 - X set2 )∩(U set2 ∪M set2 )

[0057] wherein U set1 is a sub-region set corresponding to the U-type fragments of G1; M set1 is a sub-region set corresponding to the M-type fragments of G1; X set1 is a sub-region set corresponding to the X-type fragments of G1; U set2 is a sub-region set corresponding to the U-type fragments of G2; M set2 is a sub-region set corresponding to the M-type fragments of G2; X set2 is a sub-region set corresponding to the X-type fragments of G2;

[0058] The z-score of each DMR is calculated based on the length of the differential methylation region and the number of fragments of abnormal tissue in the region with this methylation type, and the p-value is calculated based on the z-score, and the following equation is used to calculate:

[0059]

[0060] wherein Length is the length of the current differential methylation region; μ(Length) is the mean of the length of all DMRs on the same chromosome; σ(Length) is the standard deviation of the length of all DMRs on the same chromosome; z1 is the first z-score value; z2 is the second z-score value; N is the number of fragments corresponding to the methylation type in the current region; σ(N) is the standard deviation of the number of fragments of all DMRs on the same chromosome; μ(N) is the mean of the number of fragments of all DMRs on the same chromosome; Φ is the standard normal distribution cumulative function.

[0061] S2. Based on the abnormal tissue M-DMR set (labeled as MDMR) and the normal control tissue-specific hypomethylation region set (labeled as TU DMR), identify the region set (labeled as DTU DMR) in which the tissue-specific hypomethylation region is not changed during the abnormal generation process;

[0062] Step S2 includes the following steps:

[0063] Obtain the normal cell type-specific hypomethylation map, and extract the specific hypomethylation region corresponding to a plurality of normal cell types contained in a preset number of normal control tissues (G2) as the normal control tissue-specific hypomethylation region set TU DMR ;

[0064] Based on the abnormal tissue differential hypomethylation region set M DMR obtained in step S1 and the normal control tissue-specific hypomethylation region set TU DMR , identify the region set DTU DMR in which the tissue-specific hypomethylation region is not changed, which is expressed by the following formula:

[0065] DTU DMR = TU DMR -(TU DMR ∩M DMR )

[0066] Wherein, TU DMR is the normal control tissue-specific hypomethylation region TU-DMR set.

[0067] S3. Based on the abnormal tissue U-DMR set (labeled as UDMR) and the other normal tissue-specific hypomethylation region set (labeled as TallU DMR), identify the region set (DUDMR) in which the non-tissue-specific hypomethylation region is changed into a region similar to the other normal tissue-specific hypomethylation region during the abnormal generation process;

[0068] Step S3 includes the following steps:

[0069] Remove several normal cell types contained in the normal control tissue in the normal cell type-specific hypomethylation map to obtain the specific hypomethylation region of other normal cell types, and combine them into a set TallU DMR ;

[0070] Based on the abnormal tissue differential hypomethylation region U DMR obtained in step S1 and the set TallU DMR , identify the region set DU DMR in which the non-tissue-specific hypomethylation region is changed into a region similar to the other normal tissue-specific hypomethylation region, which is expressed by the following formula:

[0071] DU DMR = U DMR ∩ Tall U DMR .

[0072] S4. Identify the set of abnormal DMRs NTDMR falling in the non-tissue specific region on the genome, which includes U-DMR, M-DMR, and X-DMR falling in the non-tissue specific region, to complete the identification of abnormal related DMRs.

[0073] Step S4 includes the following steps:

[0074] The other regions of each normal cell type on the human genome except for the first pre-set number of specific hypomethylation regions of each normal cell type are collectively referred to as non-tissue specific regions NT region .

[0075] Identify the set of abnormal tissue DMRs NT DMR falling in the non-tissue specific region, which is expressed using the following formula:

[0076] NTM DMR = (M DMR ∩ NT region )

[0077] NTU DMR = (U DMR ∩ NT region )

[0078] NTX DMR = (X DMR ∩ NT region )

[0079] NT DMR = NTM DMR ∪ NTU DMR ∪ NTX DMR

[0080] Where M DMR is an abnormal tissue differential hypermethylation region; U DMR is an abnormal tissue differential hypomethylation region; X DMR is an abnormal tissue differential mixed methylation region; NTM DMR is an abnormal tissue differential hypermethylation region falling in the non-tissue specific region; NTU DMR is an abnormal tissue differential hypomethylation region falling in the non-tissue specific region; NTX DMR is an abnormal tissue differential mixed methylation region falling in the non-tissue specific region.

[0081] Specifically, in actual use, the method can also be applied as follows:

[0082] Collecting existing plasma free DNA (cfDNA) data and marking the source; the source includes abnormal tissues and normal tissues;

[0083] Extracting cfDNA methylation features in different types of methylation marker regions identified by the method;

[0084] Constructing an initial methylation marker identification model based on a gradient boosting machine algorithm;

[0085] Using the obtained cfDNA data sources and cfDNA methylation features as training data sets, training the free DNA methylation marker judgment model to obtain a methylation marker identification model;

[0086] Using the obtained methylation marker identification model to complete the identification of free DNA methylation markers.

[0087] Further, when collecting existing cfDNA data, the Methylation Fragment Level (MFL) is used to calculate the proportion of U-type fragments (methylation CpG site proportion less than or equal to 25%), M-type fragments (methylation CpG site proportion greater than or equal to 75%) and X-type fragments (methylation CpG site proportion between 25% and 75%) in each methylation marker region of the sample, respectively;

[0088] The proportions of U-type, M-type and X-type fragments in the common methylation marker region of all samples are subjected to Z-score standardization processing, and the standardized values are used as feature inputs of the model.

[0089] The following describes the method in combination with an embodiment:

[0090] From the GSE70090 data set, WGBS data of normal liver tissue and liver cancer abnormal tissue is obtained; including 4 normal liver samples and 4 liver cancer abnormal samples.

[0091] The WGBS data of the 4 normal liver samples and 4 liver cancer abnormal samples is analyzed, and based on the method described in step 1, the differential methylation regions between liver cancer abnormal samples and normal liver tissue are identified, and 2150 liver cancer M-DMRs, 8764 liver cancer U-DMRs and 338 liver cancer X-DMRs are identified. Among the identified DMRs, liver cancer U-DMRs account for a large proportion (77.9%), which indicates that during the development of normal liver tissue to liver cancer abnormal tissue, the main type of methylation change is hypomethylation change.

[0092] The distribution characteristics of the liver cancer DMR set in the genome were analyzed, and the number of DMRs on chr19 was significantly higher than that on other chromosomes, suggesting that chr19 might be the main region of methylation change in liver cancer, such as Figure 2 (a); in addition, the liver cancer U-DMR set and the liver cancer M-DMR set had high overlap with the promoter region, as shown in Figure 2 (b) and Figure 2 (c).

[0093] KEGG pathway analysis of the genes overlapped by the liver cancer U-DMR set showed that the liver cancer U-DMR set identified in this embodiment was significantly enriched in a variety of cancer pathways and signal transduction pathways. Figure 3 is a KEGG pathway enrichment pattern diagram of the liver cancer U-DMR related genes provided in this embodiment: specifically, in liver cancer, the U-DMR set was significantly enriched in a variety of cancer-related pathways such as gastric cancer (KEGG ID: hsa05226, FoldEnrichment = 1.84, p.adjust = 1.01e-05), breast cancer (KEGG ID: hsa05224, FoldEnrichment = 1.74, p.adjust = 1.35e-04), basal cell carcinoma (KEGG ID: hsa05217, FoldEnrichment = 2.08, p.adjust = 5.73e-04), hepatocellular carcinoma (KEGG ID: hsa05225, FoldEnrichment = 1.59, p.adjust = 1.27e-03), and colorectal cancer (KEGG ID: hsa05210, FoldEnrichment = 1.76, p.adjust = 4.80e-03); in addition, the identified liver cancer U-DMR was also widely enriched in a variety of key signal transduction pathways, such as the Hippo signal pathway (KEGG ID: hsa04390, FoldEnrichment = 1.89, p.adjust = 1.83x10 -6 ), the Rap1 signal pathway (KEGG ID: hsa04015, FoldEnrichment = 1.75, p.adjust = 1.83x10 -6 ), the calcium signal pathway (KEGG ID: hsa04020, FoldEnrichment = 1.58, p.adjust = 5.08x10 -5 ), and the cAMP signal pathway (KEGG ID: hsa04024, FoldEnrichment = 1.49, p.adjust = 1.82x10 -3 ).

[0094] Based on the liver cancer tissue M-DMR set (MDMR) identified in step 1, and the normal liver tissue-specific hypomethylation region set (TUDMR) in the normal cell type-specific hypomethylation map proposed by Loyfer et al. [1], a set of regions in which the liver tissue-specific hypomethylation region does not change during the development of liver cancer (DTU-DMR) is identified. First, the specific hypomethylation region corresponding to the hepatocyte contained in the normal liver tissue (the first 1000 specific hypomethylation regions) is extracted from the 39 normal cell type-specific hypomethylation maps constructed in the work of Loyfer et al. [1] as the specific hypomethylation region set of the normal liver tissue. Based on the specific hypomethylation region set of the normal liver tissue and the liver cancer-related differential hypermethylation region set MDMR identified in step 1, a set of regions in which the liver tissue-specific hypomethylation region does not change during the development of liver cancer (DTU-DMR) is identified.

[0095] Based on the method, 998 methylation markers of liver cancer DTU-DMR are identified, and such methylation markers are regions that remain tissue-specific hypomethylation during the development of liver cancer. This indicates that normal tissue-specific hypomethylation regions exhibit significant stability during the occurrence of cancer, and only a small part of these regions undergo hypermethylation changes, which is consistent with the conclusion of existing research on the stability of specific methylation patterns between normal tissues.

[0096] The other regions of the first 1000 specific hypomethylation regions of each normal cell type in the 39 normal cell type-specific hypomethylation maps of the human genome are collectively referred to as non-tissue-specific regions, and a set of abnormal tissue DMRs falling within the non-tissue-specific regions is identified.

[0097] Based on the method, 998 methylation markers of liver cancer DTU-DMR are identified, and such methylation markers are regions that remain tissue-specific hypomethylation during the development of liver cancer. This indicates that normal tissue-specific hypomethylation regions exhibit significant stability during the occurrence of cancer, and only a small part of these regions undergo hypermethylation changes, which is consistent with the conclusion of existing research on the stability of specific methylation patterns between normal tissues.

[0098] The plasma free cfDNA data of liver cancer patients and healthy people are collected, the methylation features of the cfDNA samples are extracted in the regions of different types of methylation markers, and a disease non-invasive diagnosis model is constructed based on the gradient boosting machine algorithm to evaluate the effectiveness and reliability of the identified methylation markers in disease non-invasive diagnosis. The liver cancer methylation markers identified in this embodiment exhibit strong classification performance when applied to cancer non-invasive diagnosis. Specifically, as shown in Table 2, the liver cancer NT-DMR, NTU-DMR, NTM-DMR and NTX-DMR methylation markers identified in this embodiment can be used to construct a liver cancer non-invasive diagnosis model, and the model has a high AUC value, which indicates that the methylation markers identified in this embodiment have strong classification performance in the application of liver cancer non-invasive diagnosis. Figure 4In the present embodiment, the WGBS data of plasma cfDNA samples of liver cancer patients were derived from the datasets with the search codes of EGAS00001000566 and EGAS00001003409 in the European Genome-Phenome Archive database (EGA), which included plasma cfDNA data of 58 Hepatocellular Carcinoma (HCC) patients and 40 healthy individuals in total.

[0099] For the methylation sequencing data of each cfDNA sample, we calculated the proportions of U-type fragments (the proportion of methylated CpG sites less than or equal to 25%), M-type fragments (the proportion of methylated CpG sites greater than or equal to 75%) and X-type fragments (the proportion of methylated CpG sites between 25% and 75%) respectively using the Methylation Fragment Level (MFL). Further, we performed Z-score standardization on the proportions of U-type, M-type and X-type fragments in the common methylation marker region of all samples respectively, and used the standardized values as the feature inputs of the liver cancer non-invasive diagnosis model.

[0100] In the model construction part, the present embodiment constructed liver cancer non-invasive diagnosis models based on liver cancer DTUDMR, DUDMR, NTUDMR and NTMDMR respectively, which were used for classifying plasma cfDNA samples of liver cancer patients. We divided the cfDNA sample data into training set and test set using a ratio of 7:3. The training set contained 69 samples, including 41 liver cancer patients and 28 healthy individuals, which were used for training the model. The test set contained 28 samples, including 17 liver cancer patients and 11 healthy individuals, which were used for evaluating the performance of the model. We constructed the cancer detection model using the gradient boosting machine method in the caret R package, and ensured the stability and reliability of the model by 5-fold cross-validation repeated 10 times. Finally, we verified the performance of the liver cancer detection model using the test data set that did not participate in the training, to evaluate its effectiveness in distinguishing HCC patients from healthy individuals.

[0101] In the present embodiment, the classification performance of liver cancer non-invasive diagnosis models based on liver cancer DTUDMR, DUDMR, NTUDMR and NTMDMR was as follows: Figure 4The AUC value of the model based on the DTUDMR is 0.84; the AUC value of the model based on the DUDMR is 0.925, and the classification effect is the best among the four methylation marker training models; the AUC values of the models based on the NTUDMR and the NTMDMR are both more than 0.9. The results verify the potential of the methylation markers identified by the method as non-invasive diagnostic markers of diseases.

Claims

1. A method for methylation marker identification of cell-free DNA, characterized in that, The method comprises the following steps: S1. Based on the fragment methylation level of whole genome methylation sequencing data, identify 3 types of differential methylation regions of abnormal tissues relative to normal control tissues, and calculate the reliability and significance of each differential methylation region based on the length of the region and the number of fragments of this methylation type in the region of the abnormal tissue, and screen the significant DMRs with a p value less than 0.05; S2. Based on the abnormal tissue M-DMR set and the normal control tissue specific hypomethylation region set, identify the region set in which the tissue specific hypomethylation region is not changed during the abnormal generation process; S3. Based on the abnormal tissue U-DMR set and the other normal tissue specific hypomethylation region set, identify the region set in which the non-tissue specific hypermethylation region is converted into a region similar to the other normal tissue specific hypomethylation region during the abnormal generation process; S4. Identify the abnormal DMR set NTDMR falling in the non-tissue specific region on the genome, which includes U-DMR, M-DMR, and X-DMR falling in the non-tissue specific region, and complete the identification of abnormal related DMRs.

2. The method for methylation marker recognition of cell-free DNA according to claim 1, characterized in that, Step S1 comprises the following steps: Quality control and alignment of WGBS data of abnormal tissues and normal tissues, extraction of read methylation information; According to the read methylation information, generate candidate regions; Collect the methylation information of all reads of each sample falling in the candidate region, reads covering at least 4 CpG sites, and classify them; For each group of samples, further divide the sub-region set in each candidate region based on the intersection of each type of fragment; Based on the length of the differential methylation region and the number of fragments of this methylation type in the region of the abnormal tissue, calculate the z-score of each DMR, and then calculate the p value based on the z-score, and screen the significant DMRs with a p value less than 0.

05.

3. The method for methylated marker recognition of cell-free DNA according to claim 2, characterized in that, Step S1 is specifically: Quality control and alignment of WGBS data of abnormal tissues and normal tissues, removal of sequencing adapters and bases with quality lower than 30 using fastp tool, alignment to hg38 genome using bismark tool, extraction of methylation haplotype of each fragment read, and obtaining read methylation information of abnormal tissue group and normal tissue group samples; the methylation haplotype is represented as a 0-1 string, wherein 0 represents unmethylated state and 1 represents methylated state; According to the read methylation information of the abnormal tissue group and the normal tissue group samples, generate candidate regions; first, identify valid CpG sites, which are defined as CpG sites covered by at least one sample in each group; based on the condition that the distance between adjacent valid CpG sites is less than 500 bp, divide the candidate regions, and finally each divided candidate region contains at least 5 CpG sites; The methylation information of all reads of each sample group falling in the candidate region is collected, reads covering at least 4 CpG sites, and each group of reads is divided into three categories according to the proportion of methylated and unmethylated CpG sites in the reads: U-type fragments, M-type fragments and X-type fragments; U-type fragments contain less than or equal to 25% methylated CpG sites, M-type fragments contain more than or equal to 75% methylated CpG sites, and X-type fragments contain 25% to 75% methylated CpG sites; For each group of samples, the set of sub-regions is further divided in each candidate region based on the overlap of each type of fragments, and for each type of fragments formed sub-regions of the abnormal tissue group and the normal tissue group, the sub-region difference set of the abnormal tissue group minus the normal tissue group is calculated, and then the sub-region difference set is intersected with the sub-regions formed by other types of fragments to obtain three types of differential methylation region sets: the differential hypomethylation region set U DMR , the differential hypermethylation region set M DMR , and the differential mixed methylation region set X DMR , which are expressed by the following formulas: U DMR = (U set1 - U set2 )∩(M set2 ∪X set2 ) M DMR = (M set1 - M set2 ) ∩ (U set2 ∪ X set2 ) X DMR = (X set1 - X set2 )∩(U set2 ∪M set2 ) wherein, U set1 is a set of sub-regions corresponding to U-fragments of G1; M set1 is a set of sub-regions corresponding to M-fragments of G1; X set1 is a set of sub-regions corresponding to X-fragments of G1; U set2 is a set of sub-regions corresponding to U-fragments of G2; M set2 is a set of sub-regions corresponding to M-fragments of G2; X set2 is a set of sub-regions corresponding to X-fragments of G2; The z-score of each DMR is calculated based on the length of the differentially methylated region and the number of fragments of abnormal tissue in this region of this methylation type, and the p-value is calculated based on the z-score; Significant DMRs with a p-value less than 0.05 are screened.

4. The method for methylated marker recognition of cell-free DNA according to claim 3, characterized in that, The z-score includes a z-score first value z1 and a z-score second value z2; the following formula is used for calculation: Wherein, Length is the length of the current differentially methylated region; μ(Length) is the mean of the length of all DMRs on the same chromosome; σ(Length) is the standard deviation of the length of all DMRs on the same chromosome; z1 is the z-score first value; z2 is the z-score second value; N is the number of fragments corresponding to the methylation type in the current region; σ(N) is the standard deviation of the number of fragments of all DMRs on the same chromosome; μ(N) is the mean of the number of fragments of all DMRs on the same chromosome.

5. The method for methylated marker recognition of cell-free DNA according to claim 3, wherein, The p-value based on the z-score is calculated using the following formula: Wherein, z1 is the z-score first value; z2 is the z-score second value; Φ is the standard normal distribution cumulative function.

6. The method for methylation marker identification of cell-free DNA according to claim 1, wherein, Step S2 includes the following steps: Obtaining normal cell type specific hypomethylation profiles, extracting a preset number of specific hypomethylation regions corresponding to various normal cell types contained in the normal control tissue as a specific hypomethylation region set TU of the normal control tissue DMR ; Based on the abnormal tissue difference hypomethylation region set M-DMR obtained in step S1 DMR A set of tissue-specific hypomethylation regions TU specific to normal control tissues DMR A set of regions DTU in which the tissue-specific hypomethylation regions are not changed DMR This is expressed using the following equation: DTU DMR = TU DMR - (TU DMR ∩ M DMR ) wherein TU DMR is the set of tissue-specific hypomethylated regions TU-DMR for normal control tissues.

7. The method for methylated marker recognition of cell-free DNA according to claim 1, wherein, Step S3 includes the following steps: Removing several normal cell types contained in normal control tissues in the normal cell type-specific hypomethylation profile, resulting in specific hypomethylation regions of other normal cell types, which are combined into a set TallU DMR ; Based on the abnormal tissue difference hypomethylated region U obtained in step S1 DMR and the set TallU DMR , the non-tissue specific hypermethylated region is identified and converted into a region set DU similar to the other normal tissue specific hypomethylated region DMR , which is expressed using the following formula: DU DMR = U DMR ∩ Tall U DMR .

8. The method for methylated marker recognition of cell-free DNA according to claim 1, wherein, Step S4 includes the following steps: other regions of each normal cell type other than the pre-specified number of specific hypomethylation regions of each normal cell type in several normal cell type-specific hypomethylation patterns on the human genome are collectively referred to as non-tissue-specific regions NT region ; Identifying a set of abnormal tissue DMRs falling within a non-tissue specific region NT DMR is expressed using the following equation: NTM DMR = (M DMR ∩ NT region ) NTU DMR = (U DMR ∩NT region ) NTX DMR = (X DMR ∩NT region ) NT DMR = NTM DMR ∪ NTU DMR ∪ NTX DMR wherein M DMR is an aberrant tissue differential hypermethylated region; U DMR is an aberrant tissue differential hypomethylated region; X DMR is an aberrant tissue differential mixed methylation region; NTM DMR is an aberrant tissue differential hypermethylated region falling within a non-tissue specific region; NTU DMR is an aberrant tissue differential hypomethylated region falling within a non-tissue specific region; NTX DMR is an aberrant tissue differential mixed methylation region falling within a non-tissue specific region.