Crowd specific imprinting area detection method based on long-read-long sequencing technology

Through long read and long sequencing technology and multiple statistical models, the problems of low hDMR recognition accuracy and high false positives in the prior art are solved, and high-precision, reusable population-specific imprinting area detection is achieved, which is suitable for large-scale population research and the identification of disease-related epigenetic markers.

CN120472982APending Publication Date: 2025-08-12HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524211.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When identifying population-specific imprinted areas, the limitations of sequencing technology lead to low hDMR recognition accuracy, simple statistical models and lack standardized processes, resulting in many false positive results and making it difficult to reuse in different populations or studies.

Method used

Using long-read and long sequencing technology, single-molecule methylation information is analyzed through the Oxford Nanopore sequencing platform, combined with Clair3, Bcftools, NanoMethPhase and Metilene tools to perform haplotype classification and differential methylation region recognition, and a variety of statistical models are constructed to form a standardized imprint region detection process.

Benefits of technology

It improves the hDMR recognition accuracy, reduces the false positive rate, and forms a reusable detection system. It is suitable for large-scale population research, identifys imprinted areas with significant population specificity, and is used in human epigenetic research and disease susceptibility analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472982A_ABST
    Figure CN120472982A_ABST
Patent Text Reader

Abstract

The invention discloses a crowd specific imprinting region detection method based on a long-read-long sequencing technology, and relates to the fields of bioinformatics and epigenetics, in particular to the crowd specific imprinting region detection method based on the long-read-long sequencing technology. The problems that existing hDMR recognition precision is low, a statistical model is too simple, and a standardized process is lacked are solved. The method comprises the following steps of: 1, acquiring a sequenced BAM file with methylation information of each person in a group of people; 2, performing haplotype typing on the BAM file to obtain a hap1 file and a hap2 file after typing; 3, performing haplotype differential methylation region identification on the hap1 file and the hap2 file to obtain a candidate mark region of the person 1; 4, obtaining a candidate mark area of one crowd; 5, repeatedly executing the steps 1 to 4 to obtain a candidate mark area of another group of crowds; and 6, based on the candidate imprint areas of the two crowds, obtaining specific imprint areas of the two crowds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and epigenetics, and in particular to a method for detecting population-specific imprinted regions based on long-read sequencing technology. Background Art

[0002] Genomic imprinting regions (genomic imprinting regions) are a phenomenon of "uniparental expression" mediated by epigenetic mechanisms, in which the expression of a specific gene is controlled by either the paternal or maternal haplotype, while the other allele is silenced. Imprinting plays a crucial role in mammalian embryonic development, growth regulation, neurodevelopment, and various diseases (such as cancer and metabolic syndrome). The identification of differentially methylated regions (hDMRs) between haplotypes is a prerequisite for identifying imprinted regions. Because imprinted regions often exhibit asymmetry in methylation levels between paternal and maternal haplotypes, accurate detection of hDMRs is a key step in discovering imprinted regions. In human population studies, imprinting patterns are not completely constant but may be influenced by multiple factors such as genetic background, environmental exposures, and population history, resulting in population-specific epigenetic regulatory patterns. Therefore, identifying population-specific imprinting regions (pIRs) has become a key research topic in epigenomics. However, existing technologies and methods still have significant limitations in this regard.

[0003] Current methods for identifying imprinted regions and hDMRs have the following major limitations:

[0004] Limitations of sequencing technology lead to low hDMR identification accuracy: Existing methods mostly use second-generation short-read sequencing (such as WGBS), with a read length of generally around 150bp, which cannot cover multiple continuous methylation sites (CpGs), resulting in poor haplotype typing and thus reducing the accuracy of hDMR identification.

[0005] Unclear thresholds can easily introduce false positives: In actual analyses, statistical models are often simplistic, lacking in-depth modeling of data structure and variational characteristics. Furthermore, some methods oversimplify statistical inference and fail to fully account for potential influences such as noise in methylation data, batch effects, or population structure. This subjective approach can easily introduce a large number of false positives, reducing analytical reliability.

[0006] Lack of standardized processes and difficulty in reusing methods: Existing studies mostly use ad hoc analytical processes without forming a systematic analytical framework, making it difficult to reuse them in different populations or studies, limiting the scalability and efficiency of the research. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems of low hDMR recognition accuracy, overly simple statistical models, and lack of standardized processes, and to propose a population-specific imprinted region detection method based on long-read sequencing technology.

[0008] A method for detecting population-specific imprinted regions based on long-read sequencing technology. The specific process is as follows:

[0009] Step 1: Obtain sorted BAM files with methylation information for each person in a group of people;

[0010] There are N people in each group;

[0011] Each BAM file contains M SNP sites;

[0012] Step 2: Perform haplotype typing on each person's sorted BAM file with methylation information to obtain the typed hap1 file and hap2 file;

[0013] Step 3: Identify haplotype differentially methylated regions on the hap1 and hap2 files obtained in step 2 to obtain candidate imprinted regions (IRs) for one individual.

[0014] Step 4: Based on the candidate imprint regions IRs of one person obtained in step 3, obtain the candidate imprint regions of a group of people;

[0015] Step 5: Repeat steps 1 to 4 to obtain candidate imprint regions for another group of people;

[0016] Step 6: Based on the candidate imprinted regions of the two populations, the specific imprinted regions pIRs of the two populations are obtained.

[0017] Preferably, in step 1, a sorted BAM file with methylation information is obtained for each person in a group of people;

[0018] There are N people in each group;

[0019] Each BAM file contains M SNP sites;

[0020] The specific process is:

[0021] Step 11, obtain the human whole genome Fast5 electrical signal file from the sequencing chip R9 or R10 in the third-generation sequencing platform Oxford Nanopore;

[0022] Step 12: Use Pod5 software to convert the Fast5 electrical signal file into a pod5 format file;

[0023] Step 13: Use Dorado software to perform base calling on the pod5 format file to generate an unaligned BAM file;

[0024] Step 14: Use Dorado software to perform base alignment on the unaligned BAM file generated in step 13 to generate an aligned BAM file;

[0025] Step 15: Perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information;

[0026] The specific process is:

[0027] Use Remora software to perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information;

[0028] Step 16

[0029] Use the Samtools sort command to sort the aligned BAM file with methylation modification information generated in step 15 to generate a sorted BAM file with methylation information;

[0030] At the same time, the Samtools index command was used to index the generated sorted BAM files containing methylation information.

[0031] Preferably, in step 2, haplotype typing is performed on each person's sorted BAM file with methylation information to obtain the typed hap1 file and hap2 file; the specific process is:

[0032] Step 21: Set the parameters of the variant detection tool Clair3, including:

[0033] Set -p to ont;

[0034] Set -m to ont_guppy2;

[0035] Set --enable_phasing to perform phasing operation;

[0036] Input each person's sorted BAM file with methylation information into the variation detection tool Clair3;

[0037] The variant detection tool Clair3 outputs a typing VCF file containing M SNP sites;

[0038] The VCF file of the classification contains 8 columns of information, namely:

[0039] Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column;

[0040] Step 22: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file;

[0041] Step 23: Input each SNP site in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 into the NanoMethPhase tool. The NanoMethPhase tool outputs two typing files, namely haplotype 1 file and haplotype 2 file.

[0042] Preferably, in step 22, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file; the specific process is:

[0043] Keep the rows in the typed VCF file where the "Filter Flag" value is equal to 'PASS';

[0044] Delete the lines in the typing VCF file where the "Filter Flag" value is not equal to 'PASS';

[0045] Finally, the typing VCF file is generated;

[0046] Each SNP site in the typing VCF file is divided into haplotype 1 and haplotype 2.

[0047] Preferably, in step 23, each SNP site in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 are input into the NanoMethPhase tool, and the NanoMethPhase tool outputs two typing files, namely, a haplotype 1 file and a haplotype 2 file;

[0048] The specific process is:

[0049] Step 231: Set Modkit software parameters: pileup, traditional;

[0050] Use Modkit software to convert the sorted BAM file with methylation information generated in step 1 into a tsv format file;

[0051] Step 232: Set the NanoMethPhase software parameters: methyl_call_processer, -tc; -tc value is 0.6;

[0052] Use NanoMethPhase software to generate BED.gz format files from tsv format files;

[0053] Step 233: Set NanoMethPhase software parameters: -of, -mt;

[0054] -of is methylcall, -mt is cpg;

[0055] Each read in the sorted BAM file with methylation information obtained in step 1, the typing VCF file obtained in step 22, and the BED.gz format file obtained in step 232 are input into the NanoMethPhase software. The NanoMethPhase software outputs the typing of each read in the sorted BAM file with methylation information obtained in step 1, which is divided into haplotype 1 file and haplotype 2 file.

[0056] Preferably, in step 3, the haplotype differentially methylated regions are identified on the hap1 file and the hap2 file after typing obtained in step 2 to obtain candidate imprinted regions IRs for one person;

[0057] The specific process is:

[0058] Step 31: Input the haplotype 1 file and haplotype 2 file obtained in step 23 into the differential methylation region identification tool Metilene. The differential methylation region identification tool Metilene outputs an hDMR region file. The hDMR region file contains the chromosome number, chromosome start and end positions of each hDMR region, the methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2;

[0059] Step 32: Merging hDMR regions whose distance is less than 500 bp or overlapping adjacent hDMR regions in the hDMR regions output in step 31 to obtain a merged hDMR region;

[0060] Step 33: Calculate the frequency of the merged hDMR region in the population frequency , retaining hDMR frequency hDMR region >0.5;

[0061] Step 34: searching for the corresponding position of the hDMR region retained in step 33 in the gene annotation database. If the hDMR region retained in step 33 is found, the hDMR region retained in step 33 is annotated and retained; if the hDMR region retained in step 32 is not found, the hDMR region retained in step 32 is deleted.

[0062] Step 35: Sum the lengths of hDMRs whose hDMR regions retained in step 34 fall within the annotated gene functional regions by ΣLength(hDMR);

[0063] Step 36: Calculate the sum of the obtained results in step 35 and the annotated gene length (Gene g ) length ratio LenPro g ,Right now:

[0064]

[0065] Step 37: Keep LenPro g The hDMR regions corresponding to ≥0.1 in step 34 were retained as candidate imprinted regions IRs.

[0066] Preferably, in step 31, the haplotype 1 file and the haplotype 2 file obtained after typing in step 23 are input into the differential methylation region identification tool Metilene, and the differential methylation region identification tool Metilene outputs an hDMR region file, which contains the chromosome number, chromosome start and end positions of each hDMR region, the methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2;

[0067] The specific process is:

[0068] Step 311: inputting the haplotype 1 file and the haplotype 2 file obtained in step 23 into the differentially methylated region identification tool Metilene, setting the parameters of the differentially methylated region identification tool Metilene, and generating an unfiltered hDMR region file by the differentially methylated region identification tool Metilene;

[0069] The specific process is:

[0070] Set the parameters of the differentially methylated region identification tool Metilene, including:

[0071] The maximum CpG spacing parameter (-M) was 500;

[0072] The minimum CpG number parameter (-m) in the region was 5;

[0073] The minimum methylation level difference (-d) was 0.1;

[0074] Input the typing haplotype 1 file and haplotype 2 file obtained in step 23 into the differential methylation region identification tool Metilene, which generates an unfiltered hDMR region file;

[0075] Step 312: Filter the unfiltered hDMR region file generated in step 311 to obtain a filtered hDMR region file. The specific process is as follows:

[0076] The hDMR regions with a length greater than 100 base pairs and a p-value less than 0.05 in the unfiltered hDMR region file generated in step 311 are retained as the filtered hDMR region file.

[0077] Preferably, in step 4, based on the candidate imprint regions IRs of one person obtained in step 3, candidate imprint regions of a group of people are obtained;

[0078] The specific process is:

[0079] Step 41: Construct a population differential methylation background set;

[0080] Step 42: Use Kullback-Leibler divergence to measure the degree of difference KL between the distribution of candidate imprinted region IRs in step 37 and the distribution of the differential methylation background set of a population obtained in step 41; retain the candidate imprinted region IRs in step 37 corresponding to KL>0.2 as the candidate imprinted regions of a population.

[0081] Preferably, in step 41, a differential methylation background set of a population is constructed; the specific steps are as follows:

[0082] Step 411: Calculate the average length L of the merged hDMR regions obtained in step 32, and divide each person's whole genome into N segments of length L according to non-overlapping windows. i ;

[0083] Step 412: Get the hap1 and hap2 files obtained in step 23 and located in segment L. i HAP1 sub-region and HAP2 sub-region within;

[0084] Step 413: Calculate the average CpG methylation level β of the hap1 subregion obtained in step 412 1 ; expressed as:

[0085]

[0086] Where, represents the methylation level of the i-th CpG site in the hap1 subregion obtained in step 412, and n represents the total number of CpG sites in the hap1 subregion obtained in step 412;

[0087] Step 414: Calculate the average CpG methylation level β of the hap2 subregion obtained in step 412 2 ; expressed as:

[0088]

[0089] Step 415: Define the average CpG methylation level β of the hap1 subregion 1 The average CpG methylation level of the HAP2 subregion is β 2 The difference Δβ back =β 1 -β 2 ;

[0090] Step 416: Repeat steps 412 to 415 to calculate each segment L i The corresponding Δβ back , which constitutes an individual’s differential methylation background set;

[0091] Step 417: Repeat steps 411 to 416 to form a differential methylation background set for a population.

[0092] Preferably, in step 6, the specific imprinted regions pIRs of the two populations are obtained based on the candidate imprinted regions of the two populations; the specific process is:

[0093] Step 61: Find the non-overlapping region between the candidate imprint regions of one group of people obtained in step 4 and the candidate imprint regions of another group of people obtained in step 5 as the final candidate imprint region;

[0094] If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 4, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 4;

[0095] If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 5, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 5;

[0096] Step 62: Find the region where the candidate fingerprint regions of one group of people obtained in step 4 overlap with the candidate fingerprint regions of another group of people obtained in step 5;

[0097] Step 63: Use the Mann-Whitney U test to calculate the significance of the overlapping regions, and retain the candidate imprint regions with a p-value less than 0.05;

[0098] Step 64: recalculate the significance of the candidate imprint regions retained in step 63 using multiple testing correction;

[0099] If the p-value of the candidate imprint region retained in step 63 is less than 0.05, the candidate imprint region retained in step 63 belongs to both the candidate imprint regions of the population obtained in step 4 and the candidate imprint regions of the population obtained in step 5;

[0100] If the p-value of the candidate imprint region retained in step 63 is greater than or equal to 0.05, the candidate imprint region retained in step 63 is deleted;

[0101] Step 65:

[0102] The final candidate imprint regions of the population in step 4 retained in step 61 and the final candidate imprint regions of the population in step 4 retained in step 64 are collectively used as the final specific imprint regions pIRs of the population in step 4;

[0103] The final candidate imprint regions belonging to the population in step 5 retained in step 61 and the final candidate imprint regions belonging to the population in step 4 retained in step 64 are collectively taken as the final specific imprint regions pIRs of the population in step 5.

[0104] The beneficial effects of the present invention are:

[0105] This paper proposes a method for detecting population-specific imprinted regions based on long-read sequencing technology. By directly analyzing single-molecule methylation information and reconstructing haplotype methylation patterns, this method can accurately identify hDMRs in individual samples and further statistically screen for population-specific imprinted regions. The entire process does not require family information, making it suitable for practical large-scale population studies.

[0106] Adopting third-generation sequencing technology to improve the accuracy of hDMR identification: using the Oxford Nanopore long-read platform, multiple CpG sites can be analyzed in a single read and their haplotype origins can be retained, greatly improving the accuracy of haplotype methylation identification.

[0107] Build multiple statistical models: By constructing hDMR frequency and hDMR length ratio models, as well as background distribution and density distribution difference models of hDMR in different populations, the false positive rate is effectively reduced and the detection accuracy is improved.

[0108] Standardize the process to form a reusable detection system: This invention provides a complete population-specific imprint region detection process, including data preprocessing, hDMR identification, and population statistical testing. It is suitable for different populations and research tasks and has good portability and scalability.

[0109] This invention has important applications in human epigenetic research, population adaptive evolution, and disease susceptibility analysis. In particular, it can be used to identify potential disease-associated epigenetic markers in population-specific imprinting studies, expanding the research boundaries of precision medicine and population genetics.

[0110] The present invention proposes a method for detecting population-specific imprinted regions based on long-read sequencing, and constructs a complete technical process from haplotype differential methylation region (hDMR) identification, imprinted region (IR) screening, to population-specific imprinted region (pIR) confirmation. This method introduces advanced sequencing technology, statistical modeling and population comparison strategies in multiple key links, significantly improving the accuracy and stability of imprinted region identification. The application of the method of the present invention in large-scale population data analysis can efficiently and accurately identify credible IR regions, and the overlap rate with the existing imprinted gene database reaches 70%, verifying the high reliability of the method. At the same time, the population-specific imprinted regions (pIRs) identified by the present invention are widely distributed in multiple different populations, reflecting good generalization ability and population-specific expression differences, providing strong technical support for subsequent in-depth exploration of epigenetic regulation and population adaptability mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0112] Specific embodiment 1: This embodiment is a method for detecting population-specific imprinted regions based on long-read sequencing technology. The specific process is as follows:

[0113] Step 1: Obtain sorted BAM files with methylation information for each person in a group of people;

[0114] There are N people in each group; 2≤N≤100000;

[0115] Each BAM file contains M SNP sites; 100000≤M≤10000000

[0116] Step 2: Perform haplotype typing on each person's sorted BAM file with methylation information to obtain the typed hap1 file and hap2 file;

[0117] Step 3: Identify haplotype differentially methylated regions (hDMRs) for the hap1 and hap2 files obtained in step 2 to obtain candidate imprinted regions (IRs) for one individual.

[0118] Step 4: Based on the candidate imprint regions IRs of one person obtained in step 3, obtain the candidate imprint regions of a group of people;

[0119] Step 5: Repeat steps 1 to 4 to obtain candidate imprint regions for another group of people;

[0120] Step 6: Based on the candidate imprinted regions of the two populations, the specific imprinted regions pIRs of the two populations are obtained.

[0121] Specific embodiment 2: This embodiment differs from specific embodiment 1 in that: in step 1, a sorted BAM file with methylation information is obtained for each person in a group of people;

[0122] There are N people in each group;

[0123] Each BAM file contains M SNP sites;

[0124] The specific process is:

[0125] Step 11, obtain the human whole genome Fast5 electrical signal file from the sequencing chip R9 or R10 in the third-generation sequencing platform Oxford Nanopore;

[0126] Step 12: Use Pod5 software to convert the Fast5 electrical signal file into a pod5 format file with higher reading efficiency.

[0127] Step 13: Use Dorado software to perform base calling on the pod5 format file to generate an unaligned BAM file;

[0128] Step 14: Use Dorado software to perform base alignment on the unaligned BAM file generated in step 13 to generate an aligned BAM file;

[0129] Step 15: Perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information;

[0130] The specific process is:

[0131] Use Remora software to perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information;

[0132] Step 16

[0133] Use the Samtools sort command to sort the aligned BAM file with methylation modification information generated in step 15 to generate a sorted BAM file with methylation information;

[0134] At the same time, the Samtools index command was used to index the generated sorted BAM files containing methylation information.

[0135] The input of the method of the present invention is sorted BAM files with methylation modification information, which can be generated by a third-generation sequencing platform (such as Oxford Nanopore) and a methylation detection tool (Remora).

[0136] Other steps and parameters are the same as those in the first embodiment.

[0137] Specific embodiment three: This embodiment differs from specific embodiments one or two in that: in step 2, haplotype typing is performed on each person's sorted BAM file with methylation information to obtain the typed hap1 file and hap2 file; the specific process is:

[0138] Step 21: Set the parameters of the variant detection tool Clair3, including:

[0139] Set -p (parameter) to ont;

[0140] Set -m (parameter) to ont_guppy2;

[0141] Set --enable_phasing to perform phasing operation;

[0142] Input each person's sorted BAM file with methylation information into the variation detection tool Clair3;

[0143] The variant detection tool Clair3 outputs a typing VCF file containing M SNP sites;

[0144] The VCF file of the classification contains 8 columns of information, namely:

[0145] Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column;

[0146] Step 22: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file;

[0147] Step 23: Input each high-quality SNP site after typing in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 into the NanoMethPhase tool. The NanoMethPhase tool outputs two typing files, divided into haplotype 1 (hap1) file and haplotype 2 (hap2) file.

[0148] Other steps and parameters are the same as those in the first or second embodiment.

[0149] Specific embodiment 4: This embodiment differs from specific embodiments 1 to 3 in that: in step 22, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file; the specific process is:

[0150] Keep the rows in the VCF files with the value of "Filter Flag" equal to 'PASS' (obtained by clair3 through alignment quality assessment);

[0151] Delete the rows in the VCF file whose "Filter Flag" value is not equal to 'PASS' (delete the values of the rows corresponding to columns 1 to 8);

[0152] Finally, a high-quality typing VCF file is generated;

[0153] Each SNP site in the typing VCF file is divided into haplotype 1 (hap1) and haplotype 2 (hap2);

[0154] The VCF file of the typing has many columns.

[0155] The other steps and parameters are the same as those in the first to third embodiments.

[0156] Specific embodiment 5: This embodiment differs from any one of specific embodiments 1 to 4 in that: in step 23, each high-quality SNP site after typing in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 are input into the NanoMethPhase tool, and the NanoMethPhase tool outputs two typing files, namely, a haplotype 1 (hap1) file and a haplotype 2 (hap2) file;

[0157] The specific process is:

[0158] Step 231: Set Modki software parameters: pileup, traditional;

[0159] Use Modki software to convert the sorted BAM file with methylation information generated in step 1 into a tsv format file;

[0160] Step 232: Set the NanoMethPhase software parameters: methyl_call_processer, -tc; -tc value is 0.6;

[0161] Use NanoMethPhase software to generate BED.gz format files from tsv format files;

[0162] Step 233: Set NanoMethPhase software parameters: -of, -mt;

[0163] -of is methylcall, -mt is cpg;

[0164] Each read in the sorted BAM file with methylation information obtained in step 1, the typing VCF file obtained in step 22, and the BED.gz format file obtained in step 232 are input into the NanoMethPhase software. The NanoMethPhase software outputs the typing of each read in the sorted BAM file with methylation information obtained in step 1, which is divided into haplotype 1 (hap1) file and haplotype 2 (hap2) file.

[0165] Other steps and parameters are the same as those in Specific Embodiments 1 to 4-1.

[0166] Specific embodiment 6: This embodiment differs from any one of specific embodiments 1 to 5 in that: in step 3, the haplotype differentially methylated regions (hDMRs) are identified on the hap1 file and the hap2 file after typing obtained in step 2 to obtain candidate imprinted regions IRs for one person;

[0167] The specific process is:

[0168] Step 31: Input the haplotype 1 (hap1) file and haplotype 2 (hap2) file obtained in step 23 into the differentially methylated region identification tool Metilene. The differentially methylated region identification tool Metilene outputs an hDMR region file. The hDMR region file contains the chromosome number, chromosome start and end positions of each hDMR region, the methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2;

[0169] Step 32: Merging hDMR regions whose distance is less than 500 bp or overlapping adjacent hDMR regions in the hDMR regions output in step 31 to obtain a merged hDMR region;

[0170] Step 33: Calculate the frequency of the merged hDMR region in the crowd (there are many people in a group of people, and the frequency of the merged hDMR region corresponding to each person in the crowd) frequency , retaining hDMR frequency hDMR region > 0.5; at this time we upgrade hDMR from the individual level to the group level.

[0171] Step 34: Search the corresponding location of the hDMR region retained in step 33 in the gene annotation database (GENCODE). If found, the hDMR region retained in step 33 is annotated and retained; if not found, the hDMR region retained in step 32 is deleted;

[0172] Step 35: Sum the lengths of hDMRs whose hDMR regions retained in step 34 fall within the annotated gene functional regions by ΣLength(hDMR);

[0173] Step 36: Calculate the sum of the obtained results in step 35 and the annotated gene length (Gene g ) length ratio LenPro g ,Right now:

[0174]

[0175] Step 37: Keep LenPro g The hDMR regions corresponding to ≥0.1 in step 34 were retained as candidate imprinted regions IRs.

[0176] Other steps and parameters are the same as those in Specific Implementations 1 to 5-1.

[0177] Specific embodiment seven: This embodiment differs from any one of specific embodiments one to six in that: in step 31, the haplotype 1 (hap1) file and the haplotype 2 (hap2) file obtained after typing in step 23 are input into the differential methylation region identification tool Metilene, and the differential methylation region identification tool Metilene outputs an hDMR region file, which contains the chromosome number, chromosome start and end positions, methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2 for each hDMR region;

[0178] The specific process is:

[0179] Step 311: Input the typing haplotype 1 (hap1) file and haplotype 2 (hap2) file obtained in step 23 into the differentially methylated region identification tool Metilene, set the parameters of the differentially methylated region identification tool Metilene, and the differentially methylated region identification tool Metilene generates an unfiltered hDMR region file;

[0180] The specific process is:

[0181] Set the parameters of the differentially methylated region identification tool Metilene, including:

[0182] The maximum CpG spacing parameter (-M) was 500;

[0183] The minimum CpG number parameter (-m) in the region was 5;

[0184] The minimum methylation level difference (-d) was 0.1;

[0185] The haplotype 1 (hap1) file and haplotype 2 (hap2) file obtained in step 23 are input into the differential methylation region identification tool Metilene, which generates an unfiltered hDMR region file.

[0186] Step 312: Filter the unfiltered hDMR region file generated in step 311 to obtain a filtered hDMR region file. The specific process is as follows:

[0187] The hDMR regions in the unfiltered hDMR region file generated in step 311 with a length (end position minus start position) greater than 100 base pairs (bp) and a p-value less than 0.05 are retained as the filtered high-quality hDMR region files.

[0188] The other steps and parameters are the same as those in the first to sixth embodiments.

[0189] Specific embodiment eight: This embodiment differs from any one of specific embodiments one to seven in that: in step 4, based on the candidate imprint regions IRs of one person obtained in step 3, candidate imprint regions of a group of people are obtained;

[0190] The specific process is:

[0191] Step 41: Construct a population differential methylation background set;

[0192] Step 42: Use Kullback-Leibler divergence to measure the degree of difference KL between the distribution of candidate imprinted region IRs in step 37 and the distribution of the differential methylation background set of a population obtained in step 41; retain the candidate imprinted region IRs in step 37 corresponding to KL>0.2 as the candidate imprinted regions of a population.

[0193] Other steps and parameters are the same as those in Specific Embodiments 1 to 7-1.

[0194] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that: in step 41, a differential methylation background set of a population is constructed; the specific steps are as follows:

[0195] Step 411: Calculate the average length L of the merged hDMR region obtained in step 32 (the average length of the merged hDMR regions), and divide each person's whole genome into N segments of length L according to non-overlapping windows. i ;

[0196] Step 412: Get the hap1 and hap2 files obtained in step 23 and located in segment L. i HAP1 sub-region and HAP2 sub-region within;

[0197] Step 413: Calculate the average CpG methylation level β of the hap1 subregion obtained in step 412 1 ; expressed as:

[0198]

[0199] Where, represents the methylation level of the i-th CpG site in the hap1 subregion obtained in step 412, and n represents the total number of CpG sites in the hap1 subregion obtained in step 412;

[0200] Step 414: Calculate the average CpG methylation level β of the hap2 subregion obtained in step 412 2 ; expressed as:

[0201]

[0202] Step 415: Define the average CpG methylation level β of the hap1 subregion 1 The average CpG methylation level of the HAP2 subregion is β 2 The difference Δβ back =β 1 -β 2 ;

[0203] Step 416: Repeat steps 412 to 415 to calculate each segment L i The corresponding Δβ back , which constitutes an individual’s differential methylation background set;

[0204] Step 417: Repeat steps 411 to 416 to form a differential methylation background set for a population.

[0205] Other steps and parameters are the same as those in Specific Implementations 1 to 8-1.

[0206] Specific embodiment 10: This embodiment differs from any one of specific embodiments 1 to 9 in that: in step 6, the specific imprinted regions pIRs of the two populations are obtained based on the candidate imprinted regions of the two populations; the specific process is:

[0207] Step 61: Find the non-overlapping region between the candidate imprint regions of one group of people obtained in step 4 and the candidate imprint regions of another group of people obtained in step 5 as the final candidate imprint region;

[0208] If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 4, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 4;

[0209] If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 5, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 5;

[0210] Follow this step to avoid any overlap;

[0211] Step 62: Find the areas where the candidate fingerprint regions of one group of people obtained in step 4 overlap with the candidate fingerprint regions of another group of people obtained in step 5; this step is repeated for partial overlap and complete overlap;

[0212] Step 63: Use a two-ended Mann-Whitney U test to calculate the significance of the overlapping regions, and retain the candidate imprint regions with a p-value less than 0.05;

[0213] Step 64: recalculate the significance of the candidate imprint regions retained in step 63 using FDR multiple testing correction;

[0214] If the p-value of the candidate imprint region retained in step 63 is less than 0.05, then there is a significant difference in the candidate imprint region between the two populations, and the candidate imprint region retained in step 63 belongs to both the candidate imprint region of the population obtained in step 4 and the candidate imprint region of the population obtained in step 5;

[0215] If the p-value of the candidate imprint region retained in step 63 is greater than or equal to 0.05, then there is no significant difference in the candidate imprint region between the two populations, and the candidate imprint region retained in step 63 is deleted;

[0216] Step 65:

[0217] The final candidate imprint regions of the population in step 4 retained in step 61 and the final candidate imprint regions of the population in step 4 retained in step 64 are collectively used as the final specific imprint regions pIRs of the population in step 4;

[0218] The final candidate imprint regions belonging to the population in step 5 retained in step 61 and the final candidate imprint regions belonging to the population in step 4 retained in step 64 are collectively taken as the final specific imprint regions pIRs of the population in step 5.

[0219] The other steps and parameters are the same as those in the specific implementation modes 1 to 9-1.

[0220] The following examples are used to verify the beneficial effects of the present invention:

[0221] Example 1:

[0222] The results of this invention, as an important derivative work of the project, have further promoted the construction and development of the ChinaMeth resource platform.

[0223] During the implementation of the project, relying on the key methods and core technologies proposed in this invention, representative population-specific imprinting regions (pIRs) in these populations were identified for the first time.

[0224] The study found that although the number of pIRs identified in a particular population was relatively small, the degree of methylation variation was significantly higher than in other populations, demonstrating stronger epigenetic specificity. Furthermore, we further discovered that these pIRs were significantly associated with several key functional genes, such as PLK3, ASPRV1, and MCAT. These genes play an important role in phenotypic traits related to high-altitude adaptation, such as stress response, skin protection, and energy metabolism.

[0225] The present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for detecting population-specific imprinted regions based on long-read sequencing technology, characterized by: The specific process of the method is: Step 1: Obtain sorted BAM files with methylation information for each person in a group of people; There are N people in each group; Each BAM file contains M SNP sites; Step 2: Perform haplotype typing on each person's sorted BAM file with methylation information to obtain the typed hap1 file and hap2 file; Step 3: Identify haplotype differentially methylated regions on the hap1 and hap2 files obtained in step 2 to obtain candidate imprinted regions (IRs) for one individual. Step 4: Based on the candidate imprint regions IRs of one person obtained in step 3, obtain the candidate imprint regions of a group of people; Step 5: Repeat steps 1 to 4 to obtain candidate imprint regions for another group of people; Step 6: Based on the candidate imprinted regions of the two populations, the specific imprinted regions pIRs of the two populations are obtained.

2. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 1, characterized in that: In step 1, a sorted BAM file with methylation information is obtained for each person in a group of people; There are N people in each group; Each BAM file contains M SNP sites; The specific process is: Step 11, obtain the human whole genome Fast5 electrical signal file from the sequencing chip R9 or R10 in the third-generation sequencing platform Oxford Nanopore; Step 12: Use Pod5 software to convert the Fast5 electrical signal file into a pod5 format file; Step 13: Use Dorado software to perform base calling on the pod5 format file to generate an unaligned BAM file; Step 14: Use Dorado software to perform base alignment on the unaligned BAM file generated in step 13 to generate an aligned BAM file; Step 15: Perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information; The specific process is: Use Remora software to perform methylation identification on the aligned BAM file generated in step 14 and the pod5 format file generated in step 12 to generate an aligned BAM file with methylation modification information; Step 16 Use the Samtools sort command to sort the aligned BAM file with methylation modification information generated in step 15 to generate a sorted BAM file with methylation information; At the same time, the Samtools index command was used to index the generated sorted BAM files containing methylation information.

3. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 2, characterized in that: In step 2, haplotype typing is performed on the sorted BAM files containing methylation information of each person to obtain the typed hap1 file and hap2 file; the specific process is: Step 21: Set the parameters of the variant detection tool Clair3, including: Set -p to ont; Set -m to ont_guppy2; Set --enable_phasing to perform phasing operation; Input each person's sorted BAM file with methylation information into the variation detection tool Clair3; The variant detection tool Clair3 outputs a typing VCF file containing M SNP sites; The VCF file of the classification contains 8 columns of information, namely: Chromosome name, SNP position, SNP ID, reference gene, alternative allele, quality value, filter flag, annotation information column; Step 22: Use Bcftools software to filter the typing VCF file to obtain a filtered typing VCF file; Step 23: Input each SNP site in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 into the NanoMethPhase tool. The NanoMethPhase tool outputs two typing files, namely haplotype 1 file and haplotype 2 file.

4. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 3, characterized in that: In step 22, the typing VCF file is filtered using Bcftools software to obtain a filtered typing VCF file; the specific process is as follows: Keep the rows in the VCF file with the "Filter Flag" value equal to 'PASS'; Delete the rows in the typing VCF file where the "Filter Flag" value is not equal to 'PASS'; Finally, the typing VCF file is generated; Each SNP site in the typing VCF file is divided into haplotype 1 and haplotype 2.

5. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 4, characterized in that: In step 23, each SNP site in the filtered typing VCF file obtained in step 22 and each read in the BAM file obtained in step 1 are input into the NanoMethPhase tool, and the NanoMethPhase tool outputs two typing files, namely, a haplotype 1 file and a haplotype 2 file; The specific process is: Step 231: Set Modkit software parameters: pileup, traditional; Use Modkit software to convert the sorted BAM file with methylation information generated in step 1 into a tsv format file; Step 232: Set the NanoMethPhase software parameters: methyl_call_processer, -tc; -tc value is 0.6; Use NanoMethPhase software to generate BED.gz format files from tsv format files; Step 233: Set NanoMethPhase software parameters: -of, -mt; -of is methylcall, -mt is cpg; Each read in the sorted BAM file with methylation information obtained in step 1, the typing VCF file obtained in step 22, and the BED.gz format file obtained in step 232 are input into the NanoMethPhase software. The NanoMethPhase software outputs the typing of each read in the sorted BAM file with methylation information obtained in step 1, which is divided into haplotype 1 file and haplotype 2 file.

6. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 5, characterized in that: In step 3, the haplotype differential methylation regions are identified on the hap1 file and the hap2 file obtained in step 2 to obtain candidate imprinted regions IRs for one individual; The specific process is: Step 31: Input the haplotype 1 file and haplotype 2 file obtained in step 23 into the differential methylation region identification tool Metilene. The differential methylation region identification tool Metilene outputs an hDMR region file. The hDMR region file contains the chromosome number, chromosome start and end positions of each hDMR region, the methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2; Step 32: Merging hDMR regions whose distance is less than 500 bp or overlapping adjacent hDMR regions in the hDMR regions output in step 31 to obtain a merged hDMR region; Step 33: Calculate the frequency of the merged hDMR region in the population frequency , retaining hDMR frequency hDMR region >0.5; Step 34: searching for the corresponding position of the hDMR region retained in step 33 in the gene annotation database. If the hDMR region retained in step 33 is found, the hDMR region retained in step 33 is annotated and retained; if the hDMR region retained in step 32 is not found, the hDMR region retained in step 32 is deleted. Step 35: Sum the lengths of hDMRs whose hDMR regions retained in step 34 fall within the annotated gene functional regions by ΣLength(hDMR); Step 36: Calculate the sum of the obtained results in step 35 and the annotated gene length (Gene g ) length ratio LenPro g ,Right now: Step 37: Keep LenPro g The hDMR regions corresponding to ≥0.1 in step 34 were retained as candidate imprinted regions IRs.

7. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 6, characterized in that: In step 31, the haplotype 1 file and the haplotype 2 file obtained in step 23 are input into the differential methylation region identification tool Metilene, which outputs an hDMR region file. The hDMR region file includes the chromosome number, chromosome start and end positions, methylation levels on hap1 and hap2, and the difference Δβ between the methylation levels of hap1 and hap2 for each hDMR region. The specific process is: Step 311: inputting the haplotype 1 file and the haplotype 2 file obtained in step 23 into the differentially methylated region identification tool Metilene, setting the parameters of the differentially methylated region identification tool Metilene, and generating an unfiltered hDMR region file by the differentially methylated region identification tool Metilene; The specific process is: Set the parameters of the differentially methylated region identification tool Metilene, including: The maximum CpG spacing parameter (-M) was 500; The minimum CpG number parameter (-m) in the region was 5; The minimum methylation level difference (-d) was 0.1; Input the typing haplotype 1 file and haplotype 2 file obtained in step 23 into the differential methylation region identification tool Metilene, which generates an unfiltered hDMR region file; Step 312: Filter the unfiltered hDMR region file generated in step 311 to obtain a filtered hDMR region file. The specific process is as follows: The hDMR regions with a length greater than 100 base pairs and a p-value less than 0.05 in the unfiltered hDMR region file generated in step 311 are retained as the filtered hDMR region file.

8. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 7, characterized in that: In step 4, based on the candidate imprint regions IRs of one person obtained in step 3, a candidate imprint region of a group of people is obtained; The specific process is: Step 41: Construct a population differential methylation background set; Step 42: Use Kullback-Leibler divergence to measure the degree of difference KL between the distribution of candidate imprinted region IRs in step 37 and the distribution of the differential methylation background set of a population obtained in step 41; retain the candidate imprinted region IRs in step 37 corresponding to KL>0.2 as the candidate imprinted regions of a population.

9. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 8, characterized in that: In step 41, a differential methylation background set of a population is constructed; the specific steps are as follows: Step 411: Calculate the average length L of the merged hDMR regions obtained in step 32, and divide each person's whole genome into N segments of length L according to non-overlapping windows. i ; Step 412: Get the hap1 and hap2 files obtained in step 23 and located in segment L. i HAP1 sub-region and HAP2 sub-region within; Step 413: Calculate the average CpG methylation level β of the hap1 subregion obtained in step 412 1 ; expressed as: Where, represents the methylation level of the i-th CpG site in the hap1 subregion obtained in step 412, and n represents the total number of CpG sites in the hap1 subregion obtained in step 412; Step 414: Calculate the average CpG methylation level β of the hap2 subregion obtained in step 412 2 ; expressed as: Step 415: Define the average CpG methylation level β of the hap1 subregion 1 The average CpG methylation level of the HAP2 subregion is β 2 The difference Δβ back =β 1 -β 2 ; Step 416: Repeat steps 412 to 415 to calculate each segment L i The corresponding Δβ back , which constitutes an individual’s differential methylation background set; Step 417: Repeat steps 411 to 416 to form a differential methylation background set for a population.

10. The method for detecting population-specific imprinted regions based on long-read sequencing technology according to claim 9, characterized in that: In step 6, specific imprinted regions pIRs of the two populations are obtained based on the candidate imprinted regions of the two populations; The specific process is: Step 61: Find the non-overlapping region between the candidate imprint regions of one group of people obtained in step 4 and the candidate imprint regions of another group of people obtained in step 5 as the final candidate imprint region; If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 4, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 4; If the final candidate imprint region only appears in the candidate imprint regions of the population obtained in step 5, the final candidate imprint region is included in the candidate imprint regions of the population obtained in step 5; Step 62: Find the region where the candidate fingerprint regions of one group of people obtained in step 4 overlap with the candidate fingerprint regions of another group of people obtained in step 5; Step 63: Use the Mann-Whitney U test to calculate the significance of the overlapping regions, and retain the candidate imprint regions with a p-value less than 0.05; Step 64: recalculate the significance of the candidate imprint regions retained in step 63 using multiple testing correction; If the p-value of the candidate imprint region retained in step 63 is less than 0.05, the candidate imprint region retained in step 63 belongs to both the candidate imprint regions of the population obtained in step 4 and the candidate imprint regions of the population obtained in step 5; If the p-value of the candidate imprint region retained in step 63 is greater than or equal to 0.05, the candidate imprint region retained in step 63 is deleted; Step 65: The final candidate imprint regions of the population in step 4 retained in step 61 and the final candidate imprint regions of the population in step 4 retained in step 64 are collectively used as the final specific imprint regions pIRs of the population in step 4; The final candidate imprint regions belonging to the population in step 5 retained in step 61 and the final candidate imprint regions belonging to the population in step 4 retained in step 64 are collectively taken as the final specific imprint regions pIRs of the population in step 5.