Genome-wide methylation region splitting method based on the co-methylation status of neighboring CpG sites
Patent Information
- Application Number
- CN202510925931.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-07-07
AI Technical Summary
然而,连锁不平衡概念应用在细胞群体的甲基化上存在一定的局限性:1)计算方法的局限性,该方法会忽略邻近CpG位点甲基化状态完全相同的区域,即下述公式中PA和PB同时等于1(CpG位点甲基化水平均为1的共甲基化)或同时等于0(CpG位点甲基化水平均为0的共非甲基化)时,共甲基化一致性r2会失去作用,造成这些CpG位点区域被忽略;2)拆分方法的局限性:该方法主要在正常组织样本全基因组甲基化测序数据中进行MHB的拆分,这使得在正常组织中超甲基化(Hyper-Methylation)区域而在癌症细胞群中邻近CpG位点完全相同的低甲基化(Hypo-Methylation)区域会被忽视,这些区域往往也为癌症诊断提供有价值的信号
Smart Images

Figure CN120766766B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method for splitting whole-genome methylation regions based on the comethylation status of adjacent CpG sites and its application. Background Technology
[0002] DNA methylation is an epigenetic marker that plays a crucial role in gene regulation, development, and cancer progression. Methylation at CpG sites in mammals is a relatively stable epigenetic modification; it can be transmitted during cell division via DNMT1 and dynamically established or removed by the DNMT3 family and TET proteins. Furthermore, methyltransferases and demethylases exhibit locally consistent activity, resulting in similar methylation states at adjacent CpG sites on DNA molecules, a phenomenon known as co-methylation. Studies have shown that genome-wide hypomethylation and hypermethylation of CpG islands associated with tumor suppressor genes and developmental regulators are key characteristics of cancer cells, and genome-wide CpG site co-methylation status also provides richer information for cancer diagnosis.
[0003] Traditional methylation studies primarily observe the overall methylation status of cell populations by measuring the average methylation levels at CpG sites, differentially methylated promoter regions, and differentially methylated regions. However, tumor heterogeneity is a major characteristic of cancer, making it difficult to observe the diversity of methylation within cell populations caused by this heterogeneity. In next-generation sequencing (NGS) data of cancer samples, methods for measuring cancer signals by co-methylation of adjacent CpG sites in DNA fragments are gaining increasing attention. Multiple studies have proposed various indicators to investigate cancer methylation signals from multiple biological perspectives, such as MHL (Methylated Haplotype Load), MBS (Methylation Block Score), and MHD (Methylated Haplotype Diversity). These methylation signal indicators mainly rely on the co-methylation status of adjacent CpG sites to separate different methylated regions. Currently, the commonly used separation method is mainly based on methylation haplotype blocks (MHBs), which adopts and improves upon the concept of linkage disequilibrium in population genetics to classify MHBs. However, the application of the linkage disequilibrium concept to methylation in cell populations has certain limitations: 1) Limitations of the calculation method: this method ignores regions where adjacent CpG sites have completely identical methylation states, i.e., P in the following formula... A and P BComethylation coherence r is equal to 1 (comethylation where all CpG site methylation levels are 1) or equal to 0 (codemethylation where all CpG site methylation levels are 0). 2 1) It will lose its function, causing these CpG site regions to be ignored; 2) Limitations of the splitting method: This method mainly splits MHB in whole genome methylation sequencing data of normal tissue samples. This makes it possible to ignore hypermethylated regions in normal tissues and hypomethylated regions that are exactly the same as adjacent CpG sites in cancer cell populations. These regions often provide valuable signals for cancer diagnosis.
[0004]
[0005]
[0006] Where D represents the linkage disequilibrium coefficient, which measures the degree of association between two methylation sites; P AB P represents the probability of comethylation at sites A and B; A P represents the probability of methylation at site A; B Indicates the probability of methylation at site B; r 2 This indicates the degree of linkage disequilibrium between two sites, with a value ranging from 0 to 1.
[0007] It is evident that methods for measuring cancer signals by comethylation of adjacent CpG sites in DNA fragments are limited by current methods for delineating genome-wide comethylation regions, resulting in the neglect of signals from numerous comethylated regions across the entire genome. Therefore, this application is proposed. Summary of the Invention
[0008] To address the aforementioned technical issues, this application develops a CoBFP (Co-Methylation / unMethylation Fragment Percent) calculation method, an index for assessing methylation consistency based on the co-methylation status (co-methylation or co-unmethylation) of adjacent CpG sites. Furthermore, it develops a method for segmenting methylated regions of the genome based on CoBFP. This method simultaneously considers both co-methylation and co-unmethylation states, providing a new approach for segmenting methylated regions across the genome and offering more information on methylated regions for the diagnosis of diseases such as cancer.
[0009] Specifically, this application proposes the following technical solution: This application first provides a method for splitting genomic methylation regions based on the comethylation status of adjacent CpG sites, the method comprising the following steps: 1) Methylation information extraction: The sequencing data is compared and analyzed with the human reference genome to obtain the methylation status of CpG sites on each sequencing sequence (fragment) within the genome; 2) CpG site natural distance splitting: Based on the distribution of CpG site distances in the human reference genome, CpG sites within the genome are split according to an appropriate size (preferably 300-800 nt) of natural distance to obtain CpG regions (CpG Bins) that can be completely covered by the sequencing sequence. 3) Candidate Bins Acquisition: Based on the split regions obtained in step 2), filter regions with less than 3 CpG sites to obtain candidate regions; 4) Calculate the proportion of comethylated / unmethylated sequences (CoBFP) within the candidate bins obtained in step 3); 5) Based on the CoBFP calculated in step 4), the genome is divided into comethylated / unmethylated regions according to the threshold.
[0010] Furthermore, step 4) includes the following steps: a. Based on the methylation information of each sequencing sequence obtained in step 1) and the CpG site information in the candidate region obtained in step 3), the methylation status of each CpG site on each sequencing sequence in the candidate region is obtained according to the alignment position information. b. Determine the sequencing sequence type based on the methylation status of adjacent CpG sites on each sequencing sequence; c. Calculate the proportion of methylated / unmethylated sequences (CoBFP).
[0011] Furthermore, in step 4): In step b, the sequencing sequence types include: Type I: Discondant Methylation Fragment (DMF) when the methylation states of the two CpG sites are inconsistent; Type II: Co-unmethylation Fragment (CUF) when both CpG sites are unmethylated; Type III: Co-methylation Fragment (CMF) when both CpG sites are methylated. In step c, the formula for calculating the methylated / unmethylated sequence ratio CoBFP is as follows: (count(CMF)+count(CUF)) / (count(CMF)+count(CUF)+count(DMF)); in: The count(CMF) is the number of type III comethylated sequences defined in b above; The count(CUF) is the number of type II co-unmethylated sequences defined in b above; The count(DMF) is the number of inconsistent methylation states of type I sequences as defined in b above.
[0012] Furthermore, step 5) includes the following steps: a. Use multiple (preferably 2-5) adjacent CpG sites in a window for local sliding windowing, with a sliding window step size of 1 CpG site; b. Within the candidate region obtained in step 3), calculate the co-methylated / unmethylated sequence ratio (CoBFP) of every two CpG sites within the sliding window using CoBFP. Set a CoBFP threshold (preferably 0.9). If the CoBFP calculated based on two CpG sites within the sliding window is greater than or equal to the threshold, continue local sliding windowing. If CoBFP is less than the threshold, split the candidate region at that CpG site and restart the sliding window at the next CpG site. Continue in this manner to split all candidate regions to obtain candidate methylated regions (Candidate Blocks).
[0013] Furthermore, in step 1), the sequencing data is obtained through the following steps: 1.1) DNA was extracted from the test tissue and control samples respectively and subjected to genomic methylation sequencing (WGBS). 1.2) Perform quality control based on the raw sequencing data to remove low-quality sequencing sequences; 1.3) Align the quality-controlled sequencing data to the human reference genome hg19 and remove repetitive sequences.
[0014] Furthermore, the method further includes the following step 6): The above splitting process was performed on the test sample and the control sample respectively. The obtained Candidate Blocks were merged and deduplicated to finally obtain the comethylated region CoBFP-blocks within the genome.
[0015] In some respects, the aforementioned genome is a whole genome.
[0016] This application also provides the application of comethylation states of adjacent CpG sites in the segmentation of genomic methylation regions.
[0017] Furthermore, the application specifically includes the aforementioned method steps.
[0018] This application also provides a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, implement any of the methods described above.
[0019] This application also provides an electronic device, which includes a memory and a processor connected together. The memory is used to store a computer program, and the processor is used to call the computer program to implement any of the methods described above.
[0020] The beneficial technical effects of this application are: (1) This application develops a novel calculation method, CoBFP, for assessing the consistency of methylation status based on adjacent CpG sites; (2) This application constructs a new method for splitting comethylated regions across the entire genome based on the CoBFP index; (3) The comethylation region splitting method constructed in this application takes into account both comethylated and co-nonmethylated regions. Compared with existing comethylation splitting methods, it can obtain more comethylated functional regions and provide more methylation information for the diagnosis of diseases such as cancer. (4) The comethylated regions obtained in this application are called CoBFP-blocks. Compared with traditional differentially methylated regions, they can reduce methylation background noise, significantly improve the signal-to-noise ratio of diseases such as cancer, and improve the detection sensitivity of early cancer screening. Attached Figure Description
[0021] Figure 1 Flowchart of genome-wide methylation region segmentation; Figure 2 A schematic diagram of candidate region acquisition; Figure 3 A schematic diagram of the CoBFP computational model; Figure 4 A schematic diagram of CoBFP-blocks splitting; Figure 5 Genome-wide functional annotation of CoBFP-blocks and MHB; Figure 6 Distribution of primitive differentially methylated regions and methylation levels of MHB and CoBFP-blocks in tumor cell lines and healthy human leukocytes; Figure 7 Comparison of cancer signal detection performance of original differentially methylated regions and MHB and CoBFP-blocks. Detailed Implementation
[0022] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of this application. Where specific conditions are not specified in the examples, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased on the market.
[0023] Some terms are defined herein, unless otherwise defined below. All technical and scientific terms used in the specific embodiments of this application are intended to have the same meaning as commonly understood by those skilled in the art. While it is believed that the following terms will be well understood by those skilled in the art, the following definitions are set forth to better explain this application.
[0024] The term "approximately" in this application refers to an accuracy range that, as would be understood by those skilled in the art, still guarantees the technical effects of the discussed features. This term typically indicates a deviation from the indicated value of ±10%, preferably ±5%.
[0025] As used in this application, the terms “comprising,” “including,” “having,” “containing,” or “involving” are inclusive or open-ended and do not exclude other unlisted elements or method steps. The term “consisting of” is considered a preferred embodiment of the term “comprising.” If a group is defined below as comprising at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists only of those embodiments.
[0026] Furthermore, the terms first, second, third, (a), (b), (c), and similar terms used in the specification and claims are for distinguishing similar elements and are not necessary for the order of description or chronological sequence. It should be understood that such terms are interchangeable in appropriate contexts, and the embodiments described herein can be implemented in a different order than that described or illustrated herein.
[0027] The method for splitting genomic methylation regions based on the comethylation status of adjacent CpG sites according to embodiments of the present invention includes the following steps: 1) Methylation information extraction: The sequencing data is compared and analyzed with the human reference genome to obtain the methylation status of CpG sites on each sequencing sequence (fragment) within the genome; 2) CpG site natural distance splitting: Based on the distribution of CpG site distances in the human reference genome, CpG sites within the genome are split according to an appropriate size of natural distance to obtain CpG regions (CpG Bins) that can be completely covered by the sequencing sequence. 3) Candidate Bins Acquisition: Based on the split regions obtained in step 2), filter regions with less than 3 CpG sites to obtain candidate regions; 4) Calculate the proportion of comethylated / unmethylated sequences (CoBFP) within the candidate bins obtained in step 3). 5) Based on the CoBFP calculated in step 4), perform genome comethylation / unmethylation region separation.
[0028] Step 4) includes the following steps: a. Based on the methylation information of each sequencing sequence obtained in step 1) and the CpG site information in the candidate region obtained in step 3), the methylation status of each CpG site on each sequencing sequence in the candidate region is obtained according to the alignment position information. b. Determine the sequencing sequence type based on the methylation status of adjacent CpG sites on each sequencing sequence; c. Calculate the proportion of methylated / unmethylated sequences (CoBFP).
[0029] In step 4): In step b, the sequencing sequence types include: Type I: Discondant Methylation Fragment (DMF) when the methylation states of the two CpG sites are inconsistent; Type II: Co-unmethylation Fragment (CUF) when both CpG sites are unmethylated; Type III: Co-methylation Fragment (CMF) when both CpG sites are methylated. In step c, the formula for calculating the methylated / unmethylated sequence ratio CoBFP is as follows: (count(CMF)+count(CUF)) / (count(CMF)+count(CUF)+count(DMF)); in, The count(CMF) is the number of type III comethylated sequences; The count(CUF) is the number of type II co-unmethylated sequences; The count(DMF) is the number of sequences with inconsistent methylation states of type I.
[0030] Step 5) includes the following steps: a. Use a window with multiple adjacent CpG sites to perform local sliding windowing, with a sliding window step size of 1 CpG site; b. Within the candidate region obtained in step 3), calculate the co-methylated / unmethylated sequence ratio (CoBFP) of every two CpG sites within the sliding window using CoBFP. Set a CoBFP threshold. If the CoBFP calculated based on two CpG sites within the sliding window is greater than or equal to the threshold, continue local sliding windowing. If CoBFP is less than the threshold, split the candidate region at that CpG site and restart the sliding window at the next CpG site. Continue in this manner to split all candidate regions to obtain candidate methylated regions (Candidate Blocks).
[0031] In step 1), the sequencing data is obtained through the following steps: 1.1) DNA was extracted from the test tissue and control samples, and genomic methylation sequencing was performed. 1.2) Perform quality control based on the raw sequencing data to remove low-quality sequencing sequences; 1.3) Align the quality-controlled sequencing data to the human reference genome hg19 and remove repetitive sequences.
[0032] The genome methylation region splitting method of this invention further includes the following step 6): The above splitting process was performed on the test sample and the control sample respectively. The obtained Candidate Blocks were merged and deduplicated to finally obtain the comethylated region CoBFP-blocks within the genome.
[0033] The genome in question is the whole genome.
[0034] Embodiments of the present invention also include the application of a comethylation state adjacent to a CpG site in the segmentation of genomic methylation regions.
[0035] This invention also includes a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, implement the genome methylation region splitting method of this invention.
[0036] This invention also includes an electronic device comprising a memory and a processor connected together. The memory stores a computer program, and the processor invokes the computer program to implement the genome methylation region splitting method of this invention.
[0037] The present application will now be described in conjunction with specific embodiments.
[0038] Example 1: This application establishes a method system for splitting genomic methylation regions based on the comethylation state of adjacent CpG sites;
[0039] This application utilizes bioinformatics analysis and optimization to establish a set of genome-wide methylation regions, termed CoBFP-blocks. For the disassembly method and detailed steps, please refer to [link to relevant documentation]. Figure 1 Or as follows: DNA was extracted from tumor tissue and normal control samples and whole-genome methylation sequencing (WGBS) was performed.
[0040] Quality control was performed based on the raw WGBS data to remove low-quality sequencing sequences.
[0041] The quality-controlled sequencing data were aligned to the human reference genome hg19, and repetitive sequences were removed.
[0042] Methylation information extraction: Based on the analysis of human reference genome alignment data, the methylation status of CpG sites on each sequencing sequence (fragment) across the whole genome is obtained and tagged with M (methylation) or U (un-methylation); CpG site natural distance splitting: Based on the distribution of CpG site distances in the genome, CpG sites across the entire genome are split according to natural distances of, for example, 500 nt, to obtain CpG regions (CpG Bins) that can be completely covered by the sequencing sequence.
[0043] Candidate Bins Acquisition: (e.g.) Figure 2 As shown in the diagram, based on the split region obtained in step 5), regions with fewer than 3 CpG sites are filtered out to obtain candidate regions.
[0044] Construction of the CoBFP computational model: Based on the methylation information of each sequencing sequence (fragment) of the tissue WGBS obtained in step 4) above and the CpG site information in the candidate region obtained in step 6) above, the methylation status (M and U) of each CpG site on each sequencing sequence (fragment) in the candidate region is obtained according to the alignment position information. The type of fragment can be determined based on the methylation status of adjacent CpG sites on the fragment, such as... Figure 3As shown: when the methylation states of 2 CpG sites are inconsistent (U and M), it is a discordant methylation fragment (DMF, Discordant Methylation Fragment); Type II: when both 2 CpG sites are unmethylated (U), it is a co-unmethylation fragment (CUF, Co-Unmethylation Fragment); Type III: when both 2 CpG sites are methylated (M), it is a co-methylation fragment (CMF, Co-Methylation Fragment); CoBFP calculation: as Figure 3 shown, the Co-Methylation / unMethylation Fragment Percent (CoBFP) of 2 CpG sites represents the ratio of the number of fragments supporting the co-methylation and co-unmethylation states of the 2 CpG sites (count(CMF)+count(CUF)) to the total number of fragments covering the 2 CpG sites (count(CMF)+count(CUF)+count(DMF)).
[0045] Splitting of co-methylation regions (CoBFP-blocks): as Figure 4 shown, for example, a local sliding window with 3 adjacent CpG sites per window (windows 3 CpGs) is used, and the sliding window step size is 1 CpG site; Within the range of Candidate Bins obtained in step 6) above, the CoBFP calculation model constructed in step 7) above is used to calculate the Co-Methylation / unMethylation Fragment Percent (CoBFP) of every two CpG sites in the sliding window; A CoBFP threshold (pcut, percentage cutoff) is set, for example, the threshold is 0.9. If the CoBFP calculated based on two CpG sites in the sliding window ≥ pcut, the local sliding is continued; if CoBFP < pcut, the Candidate Bin is split at this CpG site, and sliding window is restarted at the next CpG site; by this analogy, all Candidate Bins are split to obtain candidate methylation regions (Candidate Blocks).
[0046] Obtaining CoBFP-blocks: The above steps 4) to 8) were performed on tumor tissue and normal control samples respectively. The obtained Candidate Blocks were merged and deduplicated to finally obtain the comethylated regions across the entire genome, which are called CoBFP-blocks.
[0047] Example 2 verifies that this application can obtain more comethylated functional regions; This embodiment used 268 samples of tumor tissue samples from 7 cancer types and normal human leukocyte DNA, as shown in Table 1. WGBS detection was performed. After sequencing data quality control, human reference genome alignment, and deduplication, the whole-genome comethylation region was split using the method of this application for each cancer type and healthy human leukocyte sample. Finally, the methylation regions obtained from the 7 cancer types and healthy human leukocytes were merged and deduplicated to obtain the final comethylation region CoBFP-blocks. These were compared with the MHB splitting results (GUO S, DIEP D, PLONGTHONGKUM N, et al. Identification of methylation haplotype blocks aids in deconvolution of heterogeneous tissue samples and tumor tissue-of-origin mapping from plasma DNA[J / OL]. Nature Genetics, 2017, 49(4): 635-642) to verify that this method can obtain more comethylation functional regions.
[0048] Table 1. Sample distribution of different cancer types Breast cancer 35 Colorectal cancer 31 esophageal cancer 31 liver cancer 36 lung cancer 77 pancreatic cancer 23 Stomach cancer 12 White blood cells in healthy people 23 total 268 The method described in this application yielded 1,055,196 co-methylated CoBFP-blocks across the entire genome, covering a genome size of 120.5M, while the MHB method yielded 147,888 blocks, covering a genome size of 14.4M. This indicates that the number and size of co-methylated regions obtained by the method described in this application are significantly higher than those obtained by MHB, approximately 7-8 times greater. Figure 5As shown, functional annotation of genome-wide methylation regions was performed for CoBFP-blocks and MHB. The results showed that CoBFP-blocks were more widely distributed in intergenic and intronic regions of the genome than MHB. Furthermore, the area of CpG islands covered by CoBFP-blocks was significantly larger than that covered by MHB (22.73 M vs 7.51 M).
[0049] In summary, the results above demonstrate that the method described in this application can obtain more comethylated regions across the entire genome. Furthermore, the functional information of these comethylated regions is more extensive, which can provide more methylation information for the diagnosis of diseases such as cancer.
[0050] Example 3: This application has higher sensitivity and accuracy. This application establishes a comparative example. In this comparative example, DNA from one lung cancer cell line (NCI-H2009) and healthy human leukocytes was used in a mixed dilution experiment to simulate DNA samples with different proportions of tumor cells. The information of the simulated samples is shown in Table 2. Each simulated dilution sample underwent methylation panel targeted capture sequencing at a sequencing depth of 1000X. After sequencing data quality control, alignment with the human reference genome, and deduplication, the methylation level of each simulated sample was calculated within the raw region, MHB, and CoBFP-blocks regions. The methylation levels calculated based on the above three methylation regions were evaluated and compared in the following two aspects:
[0051] 1) The distribution of methylation levels in Raw Region, MHB, and CoBFP-blocks was compared in lung cancer cell line (NCI-H2009, 100% tumor cell percentage) and healthy human leukocytes (WBC, 0% tumor cell percentage), respectively. The main purpose was to assess the differences in the distribution of tumor methylation signal and noise between CoBFP-blocks delineated by the method of this application and conventional differentially methylated regions (Raw Region) and MHB. 2) In simulated samples with different tumor cell percentages (0%, 0.05%, 0.1%, 0.2%, 0.5%), the significantly different percentages of methylated regions, referred to as the Detection Ratio, were used to test the differences in the detection performance of Raw Region, MHB, and CoBFP-blocks for tumor signals, thereby evaluating the lower limit of detection (LoD) of tumor signals under different methylation region segmentation methods.
[0052] Table 2. Information on cell line dilution experiments WBC 0% 5 D2000 0.05% 3 D1000 0.1% 3 D500 0.2% 3 D200 0.5% 3 NCI-H2009 100% 1 like Figure 6 As shown in A and C, comparing the methylation levels of Raw Regions and CoBFP-blocks in lung cancer cell lines and healthy human WBCs revealed that CoBFP-blocks effectively obtained more information on comethylated regions, including hypermethylation and hypomethylation patterns. Figure 6 The C and WBC methylation levels were >0.8 and <0.2, respectively. Meanwhile, CoBFP-blocks exhibited lower background noise compared to the Raw Region. Figure 6 In the figure, A and C represent the WBC methylation level (0.2-0.8), and improve the tumor-specific signal-to-noise ratio. On the other hand, we further compared the effects of CoBFP-blocks and MHB on the distribution of tumor methylation signals in WBCs of lung cancer cell lines and healthy individuals, such as... Figure 6 As shown in B and C, compared to MHB, CoBFP-blocks can retain more tumor-specific low-level comethylation regions ( Figure 6 In B and C, the ordinate is WBC methylation level >0.8).
[0053] Based on the above evaluation results, it can be seen that the CoBFP-blocks separated by the method of this application can effectively improve the tumor-specific methylation signal-to-noise ratio and obtain more co-methylated regions, thus providing more methylation information for the diagnosis of diseases such as cancer.
[0054] Furthermore, the performance of raw differentially methylated regions (Raw Regions), MHBs, and CoBFP-blocks for cancer signal detection was further evaluated. Results are as follows: Figure 7 As shown, the proportion of methylated regions with significant differences is termed the Detection Ratio, which characterizes the strength of the detected tumor signal. The Detection Ratio was compared to that of healthy white blood cell samples (0% tumor cell percentage) to determine whether there was a significant difference. The results showed that CoBFP-blocks achieved a LoD (LoD) of 0.05% for cancer signal detection at the tumor cell percentage level, while MHB's LoD was 0.1% and the LoD for raw regions was 0.2%.
[0055] This demonstrates that the CoBFP-blocks split across the entire genome using the method described in this application can effectively lower the detection limit of cancer methylation signals, exhibiting higher detection sensitivity compared to MHB and Raw Region methods, thus ensuring the accuracy of detecting extremely low tumor methylation signals in applications such as early cancer screening.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for splitting genomic methylation regions based on the co-methylation status of adjacent CpG sites, characterized in that, Includes the following steps: 1) Methylation information extraction: The sequencing data is compared and analyzed with the human reference genome to obtain the methylation status of CpG sites on each sequencing sequence fragment within the genome; 2) CpG site natural distance splitting: Based on the distribution of CpG site distances in the human reference genome, CpG sites within the genome are split according to an appropriate size of natural CpG site distance to obtain CpG regions CpG Bins that can be completely covered by the sequencing sequence. 3) Candidate Bins Acquisition: Based on the split regions obtained in step 2), filter regions with less than 3 CpG sites to obtain candidate regions; 4) Calculate the proportion of comethylated / unmethylated sequences (CoBFP) within the candidate bins obtained in step 3); 5) Based on the CoBFP calculated in step 4), perform genome comethylation / unmethylation region segmentation; Step 4) includes the following steps: a) Based on the methylation information of each sequencing sequence obtained in step 1) and the CpG site information in the candidate region obtained in step 3), the methylation status of each CpG site on each sequencing sequence in the candidate region is obtained according to the alignment position information; b) Based on the methylation status of adjacent CpG sites on each sequencing sequence, the sequencing sequence type is determined. c. Calculate the methylated / unmethylated sequence ratio (CoBFP); In step 4), the sequencing sequence type in step b includes: Type I: DMF (disagreeable methylation state sequence) when the methylation states of the two CpG sites are inconsistent; Type II: CUF (co-unmethylated sequence) when both CpG sites are unmethylated; Type III: CMF (co-methylated sequence) when both CpG sites are methylated. In step c, the formula for calculating the methylated / unmethylated sequence ratio CoBFP is as follows: (count(CMF)+count(CUF)) / (count(CMF)+count(CUF)+count(DMF)); Wherein, count(CMF) is the number of type III comethylated sequences; The count(CUF) is the number of type II co-unmethylated sequences; The count(DMF) is the number of sequences with inconsistent methylation states of type I; Step 5) includes the following steps: A local sliding window is performed using multiple adjacent CpG sites in a single window, with a sliding window step size of one CpG site. Within the candidate region obtained in step 3), the proportion of co-methylated / unmethylated sequences (CoBFP) within each pair of CpG sites in the sliding window is calculated using CoBFP. A CoBFP threshold is set. If the CoBFP calculated based on the two CpG sites within the sliding window is greater than or equal to the threshold, the local sliding window continues. If CoBFP is less than the threshold, the candidate region is split at that CpG site, and the sliding window restarts at the next CpG site. This process is repeated to split all candidate regions to obtain candidate methylated regions (Candidate Blocks).
2. The method for splitting genomic methylation regions according to claim 1, characterized in that, In step 1), the sequencing data is obtained through the following steps: DNA is extracted from the tissue to be tested and control samples and genomic methylation sequencing is performed; quality control is performed based on the raw sequencing data to remove low-quality sequencing sequences; the quality-controlled sequencing data is aligned to the human reference genome hg19 and repetitive sequences are removed.
3. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which includes program instructions that, when executed by a processor, implement the method described in any one of claims 1-2.
4. An electronic device, characterized in that, The electronic device includes a memory and a processor connected together. The memory is used to store a computer program, and the processor is used to call the computer program to implement the method according to any one of claims 1-2.
Citation Information
Patent Citations
Differential methylation region screening method and device
CN114171115A
Screening method of methylation sequencing fragment
CN119649908A