A method and electronic device for methylation detection

By utilizing double-stranded DNA information and UMI information from sequencing data to filter methylation sites, the methylation detection process is optimized, solving the problem of mutations affecting the accuracy of methylation identification in existing technologies. This achieves efficient, low-cost, and high-quality methylation site identification, improving the accuracy of methylation modeling.

CN117352054BActive Publication Date: 2026-06-02BOE TECHNOLOGY GROUP CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2023-09-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methylation detection procedures fail to effectively consider the impact of mutations on the accuracy of methylation identification, leading to false positives and false negatives, which affects the accuracy of subsequent methylation modeling.

Method used

By utilizing double-stranded DNA information from sequencing data and combining it with UMI information to filter methylation sites, the methylation detection process is optimized, high-quality methylation sites are identified, and systematic mutations in the PCR and sequencing processes are eliminated.

Benefits of technology

It improves the accuracy of methylation site identification, reduces time and cost, can efficiently identify high-quality methylation sites, reduces false positives and false negatives, and improves the accuracy of subsequent methylation modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117352054B_ABST
    Figure CN117352054B_ABST
Patent Text Reader

Abstract

This invention discloses a method and electronic device for methylation detection, used to improve the accuracy of mutation identification and enhance the accuracy of existing methylation site identification, providing high-quality methylation sites for subsequent methylation modeling and methylation model identification and classification. The method includes: acquiring sequencing data; determining the reference genome corresponding to the sequencing data; comparing the sequencing data with the corresponding reference genome to obtain the position of the sequencing data in the corresponding reference genome; identifying the methylation sites of the sequencing data based on their position in the corresponding reference genome, and using the identified methylation sites as candidate methylation sites; identifying the double-stranded DNA corresponding to the sequencing data based on the UMI information in the sequencing data; filtering the candidate methylation sites using the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a method and electronic device for methylation detection. Background Technology

[0002] cfDNA (circulating free DNA) is cell-free nucleic acid in blood plasma, consisting of DNA fragments released during cell apoptosis and necrosis. Currently, cfDNA methylation patterns are considered the preferred technique for detecting tumor methylation biomarkers, with key features including high-efficiency transformation experiments, identification of high-quality methylation sites, and highly accurate methylation models.

[0003] High-cycle-number PCR (polymerase chain reaction) induced by methylation can lead to various mutations, and mutations are also easily introduced into the DNA of organisms during replication. These mutations pose challenges to the subsequent identification of true methylation sites. Current methylation detection procedures do not consider the impact of mutations on the accuracy of methylation identification. Summary of the Invention

[0004] This invention provides a method and electronic device for methylation detection, which can improve the accuracy of mutation identification, enhance the accuracy of existing methylation site identification, and provide high-quality methylation sites for subsequent methylation modeling and methylation model identification and classification.

[0005] In a first aspect, embodiments of the present invention provide a method for methylation detection, the method comprising:

[0006] Acquire sequencing data, determine the reference genome corresponding to the sequencing data, compare the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome;

[0007] Based on the location of the sequencing data in the corresponding reference genome, the methylation sites of the sequencing data are identified, and the identified methylation sites are used as candidate methylation sites.

[0008] Based on the UMI information in the sequencing data, the double-stranded DNA corresponding to the sequencing data is identified. The candidate methylation sites are filtered by the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.

[0009] As an optional implementation, before comparing the sequencing data with the corresponding reference genome, the following steps are also included:

[0010] The sequencing data is preprocessed to obtain preprocessed sequencing data.

[0011] As an optional implementation, sequencing data can be preprocessed in any one or more ways:

[0012] Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0013] Remove bases in the sequencing data whose base quality is below the second threshold;

[0014] Remove base sequences from sequencing data whose length is below the third threshold.

[0015] As an optional implementation, after obtaining the position of the sequencing data in the corresponding reference genome, the method further includes:

[0016] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0017] As an optional implementation, the step of identifying methylation sites in the sequencing data based on their positions in the corresponding reference genome, and using the identified methylation sites as candidate methylation sites, includes:

[0018] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0019] As an optional implementation, before identifying the methylation sites of the sequencing data based on their positions in the corresponding reference genome, the method further includes:

[0020] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0021] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0022] As an optional implementation, the step of filtering candidate methylation sites using the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site includes:

[0023] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0024] As an optional implementation, the step of filtering candidate methylation sites using the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site includes:

[0025] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0026] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0027] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0028] As an optional implementation, if the candidate methylation sites at corresponding positions of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical, the method further includes:

[0029] By comparing the candidate methylation sites at corresponding positions in the first double-stranded DNA and the corresponding reference genome of the sequencing data, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained; or,

[0030] Candidate methylation sites at corresponding positions are compared between the second double-stranded DNA and the corresponding reference genome of the sequencing data. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0031] Secondly, an electronic device provided by an embodiment of the present invention includes a processor and a memory, wherein the memory is used to store a program executable by the processor, and the processor is used to read the program in the memory and perform the following steps:

[0032] Acquire sequencing data, determine the reference genome corresponding to the sequencing data, compare the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome;

[0033] Based on the location of the sequencing data in the corresponding reference genome, the methylation sites of the sequencing data are identified, and the identified methylation sites are used as candidate methylation sites.

[0034] Based on the UMI information in the sequencing data, the double-stranded DNA corresponding to the sequencing data is identified. The candidate methylation sites are filtered by the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.

[0035] As an optional implementation, before comparing the sequencing data with the corresponding reference genome, the processor is further configured to perform:

[0036] The sequencing data is preprocessed to obtain preprocessed sequencing data.

[0037] As an optional implementation, the processor is specifically configured to preprocess the sequencing data in one or more ways:

[0038] Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0039] Remove bases in the sequencing data whose base quality is below the second threshold;

[0040] Remove base sequences from sequencing data whose length is below the third threshold.

[0041] As an optional implementation, after obtaining the location of the sequencing data in the corresponding reference genome, the processor is further configured to execute:

[0042] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0043] As an optional implementation, the processor is specifically configured to execute:

[0044] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0045] As an optional implementation, before identifying the methylation sites of the sequencing data based on their positions in the corresponding reference genome, the processor is further configured to perform:

[0046] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0047] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0048] As an optional implementation, the processor is specifically configured to execute:

[0049] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0050] As an optional implementation, the processor is specifically configured to execute:

[0051] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0052] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0053] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0054] As an optional implementation, if the candidate methylation sites at corresponding positions of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical, the processor is further configured to execute:

[0055] By comparing the candidate methylation sites at corresponding positions in the first double-stranded DNA and the corresponding reference genome of the sequencing data, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained; or,

[0056] Candidate methylation sites at corresponding positions are compared between the second double-stranded DNA and the corresponding reference genome of the sequencing data. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0057] Thirdly, embodiments of the present invention also provide an apparatus for methylation detection, the apparatus comprising:

[0058] The gene alignment module is used to acquire sequencing data, determine the reference genome corresponding to the sequencing data, align the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome.

[0059] The site identification module is used to identify methylation sites in the sequencing data based on their location in the corresponding reference genome, and to use the identified methylation sites as candidate methylation sites.

[0060] The site filtering module is used to identify the double-stranded DNA corresponding to the sequencing data based on the UMI information in the sequencing data, and to filter candidate methylation sites by the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.

[0061] As an optional implementation, before comparing the sequencing data with the corresponding reference genome, a preprocessing module is further included, specifically for:

[0062] The sequencing data is preprocessed to obtain preprocessed sequencing data.

[0063] As an optional implementation, the preprocessing module is specifically used to preprocess the sequencing data in any one or more ways:

[0064] Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0065] Remove bases in the sequencing data whose base quality is below the second threshold;

[0066] Remove base sequences from sequencing data whose length is below the third threshold.

[0067] As an optional implementation, after obtaining the position of the sequencing data in the corresponding reference genome, a filtering module is further included, specifically for:

[0068] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0069] As an optional implementation, the site identification module is specifically used for:

[0070] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0071] As an optional implementation, before identifying the methylation sites of the sequencing data based on their positions in the corresponding reference genome, a site conversion module is further included, specifically for:

[0072] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0073] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0074] As an optional implementation, the site filtering module is specifically used for:

[0075] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0076] As an optional implementation, the site filtering module is specifically used for:

[0077] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0078] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0079] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0080] As an optional implementation, if the candidate methylation sites at corresponding positions of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical, the site filtering module is further configured to:

[0081] By comparing the candidate methylation sites at corresponding positions in the first double-stranded DNA and the corresponding reference genome of the sequencing data, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained; or,

[0082] Candidate methylation sites at corresponding positions are compared between the second double-stranded DNA and the corresponding reference genome of the sequencing data. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0083] Fourthly, embodiments of the present invention also provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the steps of the method described in the first aspect above.

[0084] These or other aspects of this application will become more apparent in the following description of embodiments. Attached Figure Description

[0085] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0086] Figure 1 This is a flowchart illustrating an embodiment of a methylation detection method provided by the present invention.

[0087] Figure 2 A schematic diagram illustrating the principle of methylation identification provided in an embodiment of the present invention;

[0088] Figures 3A-3B An example diagram of sequence characteristics of single-stranded sequencing data provided in an embodiment of the present invention;

[0089] Figure 4 A schematic diagram of a cfDNA conversion process provided in an embodiment of the present invention;

[0090] Figure 5 A schematic diagram illustrating the mapping and identification of double-stranded DNA according to an embodiment of the present invention;

[0091] Figures 6A-6C An example diagram of sequence characteristics of double-stranded sequencing data provided in an embodiment of the present invention;

[0092] Figure 7 This is an overall implementation flowchart of methylation detection provided by an embodiment of the present invention;

[0093] Figure 8 A schematic diagram of an electronic device provided in an embodiment of the present invention;

[0094] Figure 9 This is a schematic diagram of a methylation detection device provided in an embodiment of the present invention. Detailed Implementation

[0095] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0096] In this embodiment of the invention, the term "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0097] The application scenarios described in the embodiments of this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems. In the description of this invention, unless otherwise stated, "multiple" means two or more.

[0098] Before introducing the methylation detection method provided in the embodiments of this application, the technical background of the embodiments of this application will be described in detail below for ease of understanding.

[0099] Cell-free DNA (cfDNA) is cell-free nucleic acid found in blood plasma, consisting of DNA fragments released during cell apoptosis and necrosis. In cancer patients, cfDNA also contains circulating tumor DNA (ctDNA), which forms the molecular basis for liquid biopsies of tumors. Currently, cfDNA methylation patterns are considered the preferred technique for detecting tumor methylation biomarkers and are applied in early cancer screening, cancer diagnosis, personalized medicine, real-time monitoring, and prognostic surveillance.

[0100] The core of tumor methylation biomarker detection technology lies in high-efficiency transformation experiments, identification of high-quality methylation sites, and highly accurate methylation models. However, high-cycle-number PCR (polymerase chain reaction) for methylation can lead to various mutations, and biological DNA (deoxyribonucleic acid) is prone to mutations during replication, especially C>T or C>A mutations. These mutations pose challenges to the subsequent identification of true methylation sites. In the subsequent methylation modeling process, the number and location of methylations used for modeling vary significantly depending on the project requirements; however, false positives or false negatives from a few or dozens of methylation sites can lead to modeling errors or subsequent model identification errors.

[0101] Current methylation detection procedures do not consider the impact of mutations on the accuracy of methylation identification. Identifying mutations through resequencing increases costs; using existing mutation identification methods to identify methylation data involves a data conversion process, which itself contains errors and may lead to false positives and false negatives.

[0102] The methylation detection method provided in this application is based on the core idea of ​​improving the accuracy of mutation identification by utilizing double-stranded DNA information from sequencing data. It optimizes double-stranded DNA technology for methylation detection to identify high-quality methylation sites, thus solving the problem of inaccurate mutation methylation identification in previous methylation site identification schemes. It eliminates the need for additional sequencing and data conversion processes, resulting in lower time costs and efficient identification of high-quality methylation sites.

[0103] like Figure 1 As shown in the figure, the implementation flow of a methylation detection method provided in this embodiment is as follows:

[0104] Step 100: Obtain sequencing data, determine the reference genome corresponding to the sequencing data, compare the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome;

[0105] During implementation, a reference genome for the corresponding species is determined based on the sequencing data. The reference genome for the corresponding species includes multiple versions, and one version can be selected according to the needs.

[0106] The sequencing data in this embodiment includes one or more sequences, which can be single-stranded DNA sequences or double-stranded DNA sequences. This embodiment does not impose too many limitations on this.

[0107] Optionally, the sequencing data in this embodiment includes a UMI, and the length of the UMI is known, meaning that after acquiring the sequencing data, a UMI of the corresponding length can be extracted from the sequencing data. The UMI (Unique Molecular Identifier) ​​is a unique tag sequence added to each fragment of the original sample's genome after fragmentation. It is used to distinguish thousands of different fragments within the same sample, providing error correction and increasing accuracy during sequencing. The principle of UMI is that each DNA fragment of the same sample carries a unique tag sequence, which is processed together with the target sequence through library construction, PCR (polymerase chain reaction) amplification, and then sequenced. In the final sequenced sequences, sequences with different tags represent that they come from different original DNA fragment molecules; when the sequence characteristics are consistent and the UMI molecular tags are consistent, it means that these sequences are all amplified from the same original DNA fragment. Consistent sequence characteristics specifically refer to sequence bases; in this text, it can refer to sequences aligned to the same genomic location with consistent sequence characteristics. Since errors in PCR and sequencing occur randomly, these molecular tags can be used to eliminate systematic mutations introduced during PCR and sequencing processes in a redundancy removal process.

[0108] In practice, UMIs are extracted from sequencing data based on their known lengths, typically ranging from 4 to 8 bp. For example, based on the set UMI length, the UMIs are extracted from the sequencing data and stored in the sequencing file, such as the header line of a FastQ file. An example of a header line is shown below:

[0109] A00311:228:HLGMGDSX3:1:1264:23078:15522_1:N:0:ACTCTCGA+CTGT ACCA:CCTTCG.

[0110] In this sequence, “A00311:228:HLGMGDSX3:1:1264:23078:15522_1:N:0:” represents the sequence identifier for sequencing data, which includes the sequencing instrument ID, sequencing run number, and flowcell ID. The sequencing run number is used to identify a single sequencing run and distinguish different sequencing processes. The flowcell ID is a unique identifier for the sequencer chip and is used to distinguish different chips. “ACTCTCGA+CTGTACCA” represents the paired-end index information of the sequencing data. “CCTTCG” is the UMI sequence extracted from the sequencing data.

[0111] In some embodiments, before comparing the sequencing data with the corresponding reference genome, the following steps may also be performed:

[0112] The sequencing data is preprocessed to obtain preprocessed sequencing data; wherein the data preprocessing is used to evaluate the sequencing data in order to obtain high-quality data.

[0113] In some examples, sequencing data are preprocessed in one or more ways:

[0114] Method 1) Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0115] Method 2) Remove bases from the sequencing data whose base quality is below the second threshold;

[0116] Method 3) Remove base sequences in the sequencing data whose sequence length is lower than the third threshold.

[0117] In practice, adapter sequences in the sequencing data can be removed. The removal criterion is that when the length of the adapter sequence at the end of the sequencing data is greater than or equal to a first threshold (e.g., 3 bp), the corresponding adapter sequence is removed. Low-quality bases can also be filtered out. Low-quality bases are defined as those with a quality value below a second threshold. Bases with a quality value below the second threshold (e.g., 25) are removed / filtered. The second threshold can be adjusted according to actual needs, ranging from 20 to 30. After removing adapter sequences and low-quality bases, the sequencing data continues to be filtered by length. When the sequence length in the sequencing data is lower than a third threshold (e.g., 50), the sequence is removed. For example, if the sequence length of Read1 or Read2 in the sequencing data is less than 50, it is filtered. The size of the third threshold can be adjusted according to specific circumstances. If more sequences are needed, the third threshold can be lowered, such as to 20. If longer sequences are needed, the third threshold can be increased to 100 (e.g., PE150, PE250, PE300, etc.). The third threshold should not exceed the difference between the Read length and the UMI length of the sequencing data.

[0118] In some embodiments, before aligning the sequencing data with the corresponding reference genome, the reference genome in the database can be formatted (i.e., a reference genome index can be constructed). The database corresponding to the reference genome can be selected based on the sequencing species corresponding to the sequencing data. For example, if the sequencing data is human, versions such as hg19, hg38, and hs1 can be used as reference genomes. The reference genome is formatted using alignment software (such as bwa, bowtie2, hisat2, minimap2, etc.) (this process can also be referred to as constructing a reference genome index), and the selection is based on the formatted index supporting subsequent alignment methods.

[0119] Then, methods such as bismark, bwa, and bowtie2 can be used to align the methylated sequencing data to the constructed reference genome index, determining the corresponding reference genome and the position of the sequencing data within it. The bam format files, which contain the corresponding reference genome and its position information, are then sorted and indexed.

[0120] Typically, because methylated sequencing data undergoes C>T and G>A conversions, it cannot be directly aligned to the reference genome. It is generally necessary to perform base conversions at the corresponding methylation sites on the reference genome, such as C>T and G>A conversions. At the same time, when aligning sequencing data and the reference genome, alignment software takes into account the multiple conversion information of methylation sites. For example, when C>T and G>A conversions are performed, there are 2×2 alignment methods for the two sequences in the sequencing data. By combining the alignment of paired sequences, the optimal alignment result can be obtained.

[0121] In some embodiments, before identifying the methylation sites of the sequencing data based on their location in the corresponding reference genome, the following steps may also be performed:

[0122] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0123] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0124] It should be noted that the corresponding position in this embodiment is understood as the alignment position when the sequencing data and the reference genome are compared to methylation sites.

[0125] Step 101: Identify the methylation sites in the sequencing data based on their positions in the corresponding reference genome, and use the identified methylation sites as candidate methylation sites.

[0126] In practice, the purpose of identifying methylation sites in sequencing data is to find out which sites are methylation sites.

[0127] In some embodiments, after obtaining the location of the sequencing data in the corresponding reference genome, the following steps may also be performed:

[0128] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0129] It should be noted that repetitive sequences in the sequencing data originate from PCR over-amplification. The definition criteria are: alignment to the same chromosome, consistent alignment start position, and consistent UMI. Repetitive sequences are filtered from the sequencing data based on these criteria. The filtered BAM format files are then sorted and indexed.

[0130] In some embodiments, methylation sites in sequencing data are identified in the following manner, and the identified methylation sites are used as candidate methylation sites:

[0131] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0132] For example, the site information corresponding to the same base information can be identified as a candidate methylation site.

[0133] In practice, methylation extraction can be performed on sequencing data, and the methylation sites at corresponding positions can be compared between the sequencing data and the reference genome. The principle is that the methylated C sites in the sequencing data are retained as C sites when compared to the reference genome, while the unmethylated C sites in the sequencing data are converted to T sites when compared to the reference genome.

[0134] like Figure 2As shown in the diagram, this embodiment provides a schematic diagram of methylation identification. In Example A, the sequencing data includes methylated sequence reads, specifically: 5'…TTGGCATGTTTAAACGTT…3'. The specific information of the reference genome sequence corresponding to this sequencing data (methylated sequence reads) is: 5'…CCGGCATGTTTAAACGCT…3'. Here, "me" represents methylated C. The methylated C site in the sequencing data is still a C site when corresponding to the reference genome sequence, while the unmethylated C site is a T site when corresponding to the reference genome sequence. Based on the alignment results, the methylated C sites and unmethylated C sites in the sequencing data can be obtained. Optionally, the methylated C sites can be used as candidate methylation sites. In Example B, the specific information of the sequencing data (methylated sequence reads) is: 5'…TTGGCATGTTTAAACGTT…3', and the specific information of the corresponding reference genome sequence is: 5'…CCGGCATGTTTAAACGCT…3'. The sequencing data contains mutated T sites. Since there may be mutated sites among the candidate methylation sites, it is necessary to filter the candidate methylation sites to obtain high-quality methylation sites.

[0135] Step 102: Based on the UMI information in the sequencing data, identify the double-stranded DNA corresponding to the sequencing data, and filter the candidate methylation sites by the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.

[0136] In this implementation, the sequencing data is single-stranded DNA. The site information of the double-stranded DNA corresponding to the sequencing data includes positional information and base information.

[0137] In some embodiments, the target methylation site is obtained in the following manner:

[0138] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0139] like Figures 3A-3B As shown in the figure, this embodiment provides an example diagram of the sequence characteristics of single-stranded sequencing data, wherein, Figure 3A In this process, the C sites in the sequencing data are first converted, and then the double-stranded DNA corresponding to the sequencing data is identified based on UMI. The corresponding position of the C site in the sequencing data is the T site. Based on the principle that the corresponding position of the unmethylated C site is the T site after conversion, it is shown that the C site in the sequencing data is the unmethylated C site. Figure 3BIn this process, the C sites in the sequencing data are first converted, and then the double-stranded DNA corresponding to the sequencing data is identified based on UMI. The position of the C site in the sequencing data corresponding to the double-stranded DNA is the C site. Based on the principle that the C site obtained after the conversion of the methylated C site is the C site, it is explained that the C site in the sequencing data is the methylated C site.

[0140] In some embodiments, the target methylation site is obtained in the following manner:

[0141] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0142] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0143] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0144] In practice, the double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand complementary to that strand. It is also necessary to identify whether the first and second double-stranded DNA sequences correspond to the same DNA template. For example... Figure 4 As shown in the diagram, this embodiment provides a schematic diagram of the cfDNA conversion process, where the sequencing data is cfDNA. First, the cfDNA is converted into one strand (e.g., the sense strand) top strand and another complementary strand (the antisense strand) bottom strand. The sense strand and the antisense strand must meet the following conditions:

[0145] Condition 1) The UMI sequences of the two strands are completely inversely complementary.

[0146] Condition 2) When the two strands are aligned to the same reference genome (i.e., the same chromosome), the differences on the left and right sides at the alignment position are m and n, respectively. The total positional difference between the two strands is G, where G = m + n; Figure 5 As shown in the figure, this embodiment provides a schematic diagram of mapping and identifying double-stranded DNA. Due to the existence of operations such as filtering, PCR errors, and mutations, the two strands can be designed with a certain degree of error tolerance by setting the G value.

[0147] 1) The smaller G is, the higher the accuracy. This value can be adjusted according to specific needs. The strictest standard is 0, that is, m=0, n=0;

[0148] 2) The maximum value of G should be less than 6, and m>=G / 3, n>=G / 3.

[0149] Optionally, the length (L) of the mapping region after the positive and negative strands are overlapped can be determined when the positive and negative strands are aligned to the same position in the reference genome and the UMI is consistent. m This is used to further determine the double-stranded DNA template in the sequencing data.

[0150] In some embodiments, if the candidate methylation sites at corresponding positions of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical, then any of the following is also performed:

[0151] (1) Compare the candidate methylation sites at the corresponding positions of the first double-stranded DNA and the corresponding reference genome of the sequencing data. Filter the candidate methylation sites according to the comparison results to obtain the target methylation sites.

[0152] (2) Compare the candidate methylation sites at the corresponding positions of the second double-stranded DNA and the corresponding reference genome of the sequencing data. Filter the candidate methylation sites according to the comparison results to obtain the target methylation sites.

[0153] like Figures 6A-6C As shown in the figure, this embodiment provides an example diagram of the sequence characteristics of double-stranded sequencing data, wherein, Figure 6A In this process, sequencing data is first converted into one strand of a double-stranded DNA template and the other strand complementary to that strand. The first double-stranded DNA sequence corresponding to the first strand and the second double-stranded DNA sequence corresponding to the other strand are then determined. When two strands of the same DNA template have been identified according to the above steps, the candidate methylation sites at corresponding positions of the first and second double-stranded DNA sequences are compared. If the Read1 and Read2 bases are identical at the same position on both strands, and this position is a C site (C base), then this position is considered a real methylation site (C). That is, firstly, the two strands must be paired DNA template strands, and secondly, the methylation sites at the same positions on both strands must have identical bases.

[0154] Figure 6BIn this process, sequencing data is first converted into one strand of a double-stranded DNA template and another strand complementary to that strand. The first double-stranded DNA sequence corresponding to the first strand and the second double-stranded DNA sequence corresponding to the other strand are then determined. When two strands of the same DNA template have been identified according to the above steps, the candidate methylation sites at corresponding positions in the first and second double-stranded DNA sequences are compared. If the Read1 and Read2 bases at the same position on both strands are different, a mutation is considered to exist at that position, and the methylation status cannot be determined. If only single-strand information is used, this position may be directly identified as methylated or unmethylated. In this embodiment, by double-strand pairing, the difference in bases at the same position on both strands is determined to be T / A or C / G, respectively. This position may be due to a methylated C mutating to T during PCR (as shown in the example in the figure), or it may be a C>A mutation or a G>T mutation. This position may also be due to other bases, A or T, mutating to C or G during PCR. The figure only illustrates one case.

[0155] Figure 6C In this process, sequencing data is first converted into one strand of a double-stranded DNA template and another strand complementary to that strand. The first double-stranded DNA sequence corresponding to the first strand and the second double-stranded DNA sequence corresponding to the other strand are then determined. When two strands of the same DNA template have been identified according to the above steps, the candidate methylation sites at corresponding positions in the first and second double-stranded DNA sequences are compared. If the two strands are at the same position with identical Read1 and Read2 bases, and the reference genome identifies this position as a C site, but subsequent identification reveals it to be a T site, then this position should not be identified as an unmethylated site. Because a specific set of sites will be selected for subsequent modeling and identification, the identification criteria are: whether the position is methylated or unmethylated. When the genetic mutation accounts for a certain proportion in the sample population (threshold setting), this position should not be used as a feature set for subsequent methylation sites. The figure only shows one case; other cases include T>C and C>A mutations in the reference genome.

[0156] like Figure 7 As shown in the figure, this embodiment also provides an overall implementation process for methylation detection, as detailed below:

[0157] Step 700: Obtain sequencing data and extract the UMI of the sequencing data;

[0158] Step 701: Filter the adapter sequences with a length greater than or equal to the first threshold and the bases with a quality lower than the second threshold in the sequencing data. Filter the sequence length of the filtered sequencing data again for base sequences with a length lower than the third threshold.

[0159] Step 702: Determine the database corresponding to the sequencing data, format the database, compare the sequencing data with the reference genome in the database, and determine the position of the sequencing data in the corresponding reference genome;

[0160] Step 703: Based on the UMI of the sequencing data, filter out repetitive sequences contained in the sequencing data;

[0161] Step 704: Identify the methylation sites in the sequencing data based on their positions in the corresponding reference genome, and use the identified methylation sites as candidate methylation sites.

[0162] Step 705: Based on the UMI information in the sequencing data, identify the double-stranded DNA corresponding to the sequencing data, filter the candidate methylation sites using the site information of the double-stranded DNA corresponding to the sequencing data, and obtain the target methylation site.

[0163] The following provides specific experimental data for methylation identification using the methylation detection method provided in this embodiment. In this implementation, methylation sequencing was performed on 18 samples using the PE250 sequencing strategy. After sequencing, UMI data extraction and data preprocessing were performed. The processing results are shown in the table below, where cleanReads represents the sequencing data and readTrimed represents the number of remaining paired sequences after this operation.

[0164] Table 1: Methylation sequencing data volume and filtered data volume

[0165] Sample name cleanReads readTrimed Lib1 67,904,400 67,904,400 Lib2 71,188,500 71,188,500 Lib3 68,592,900 68,592,900 Lib4 65,237,700 65,237,700 Lib5 60,938,700 60,938,700 Lib6 92,699,400 92,699,400 Lib7 75,907,800 75,907,800 Lib8 116,725,500 116,725,500 Lib9 71,881,500 71,881,500 Lib10 96,522,600 96,522,600 Lib11 93,974,700 93,974,700 Lib12 93,222,430 93,222,430 Lib13 55,484,072 55,484,072 Lib14 104,099,522 104,099,522 Lib15 178,425,060 178,425,060 Lib16 55,983,812 55,983,812 Lib17 58,175,583 58,175,583 Lib18 50,644,350 50,644,350

[0166] In this embodiment, hg38 is used as the reference genome, and an index is constructed based on it. Based on the UMI of the sequencing data, repetitive sequences contained in the sequencing data are filtered. The filtered sequencing data is compared with the reference genome and the results are statistically analyzed. Here, mapRatio represents the alignment rate; dedupRatio represents the proportion of repetitive sequences filtered; CpgRatio represents the identified methylation rate; and Paired represents the proportion of double-stranded DNA identified.

[0167] Table 2: Statistics on Methylation Identification Information

[0168] Sample name mapRatio dedupRatio CpgRatio Paired (%) Lib1 68.30% 80.22% 56.10% 0.04 Lib2 82.10% 76.09% 51.60% 0.05 Lib3 83.80% 71.60% 57.10% 0.05 Lib4 82.00% 77.57% 51.00% 0.05 Lib5 82.80% 76.88% 53.30% 0.05 Lib6 81.70% 74.90% 44.80% 0.05 Lib7 76.90% 68.15% 45.30% 0.05 Lib8 78.90% 41.44% 48.60% 0.05 Lib9 79.10% 78.51% 46.40% 0.06 Lib10 80.30% 69.43% 37.00% 0.05 Lib11 62.70% 70.53% 51.20% 0.06 Lib12 62.50% 14.94% 43.00% 0.07 Lib13 82.90% 74.26% 41.80% 0.03 Lib14 69.80% 49.68% 32.30% 0.16 Lib15 68.40% 53.67% 38.70% 0.14 Lib16 43.80% 2.81% 26.50% 0.14 Lib17 84.10% 80.07% 48.30% 0.05 Lib18 29.10% 9.37% 9.70% 0.09

[0169] This embodiment performs methylation screening on the identified double-stranded DNA template sequence to identify various possible mutations, facilitating the selection of higher-quality methylation sites (i.e., target methylation sites) from candidate methylation sites. Here, realCpgRatio represents the proportion of identified real methylation sites; PCRmutationRatio represents the proportion of identified PCR-mutated methylation sites; germlineMutationRatio represents the proportion of identified germline-mutated methylation sites; and othersRatio represents the proportion of other types of methylation.

[0170] Table 3: Statistics of various mutations and unidentified types in methylation sites

[0171]

[0172] This embodiment utilizes double-stranded UMI technology to improve the accuracy of mutation identification by combining UMI with double-stranded information. UMI extraction and data preprocessing are performed on methylation sequencing data, followed by database formatting, and the sequencing data is aligned to the formatted database. Repetitive sequence filtering is performed based on UMI, and methylation sites are identified to obtain candidate methylation sites. Mutations are identified based on double-stranded UMI, and methylation sites are filtered to obtain high-quality methylation sites. This embodiment can detect the problem of methylation sites being identified as unmethylated sites due to genetic mutations, and can detect accuracy issues caused by PCR mutations, providing high-quality methylation data for subsequent methylation model establishment and model identification. No additional sequencing or data conversion processes are required, resulting in lower time costs and more efficient detection.

[0173] Based on the same inventive concept, this embodiment of the invention also provides an electronic device. Since this electronic device is the same as the electronic device in the method of this embodiment of the invention, and the principle of solving the problem by this electronic device is similar to that of this method, the implementation of this electronic device can refer to the implementation of the method, and the repeated parts will not be described again.

[0174] like Figure 8 As shown, the electronic device includes a processor 800 and a memory 801. The memory 801 stores programs executable by the processor 800. The processor 800 reads the programs from the memory 801 and performs the following steps:

[0175] Acquire sequencing data, compare the sequencing data with the reference genome in the database, determine the reference genome corresponding to the sequencing data, and the position of the sequencing data in the corresponding reference genome;

[0176] Based on the position of the sequencing data in the corresponding reference genome, the methylation sites at the corresponding positions in the sequencing data and the reference genome are compared, and candidate methylation sites in the sequencing data are determined according to the comparison results.

[0177] Based on the sequencing data, PCR amplification was performed to obtain each methylation site in the double-stranded DNA. Candidate methylation sites in the sequencing data were filtered to obtain the target methylation sites.

[0178] As an optional implementation, before aligning the sequencing data with the corresponding reference genome, the processor 800 is further configured to perform:

[0179] The sequencing data is preprocessed to obtain preprocessed sequencing data.

[0180] As an optional implementation, the processor 800 is specifically configured to preprocess the sequencing data in one or more ways:

[0181] Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0182] Remove bases in the sequencing data whose base quality is below the second threshold;

[0183] Remove base sequences from sequencing data whose length is below the third threshold.

[0184] As an optional implementation, after obtaining the location of the sequencing data in the corresponding reference genome, the processor 800 is further configured to execute:

[0185] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0186] As an optional implementation, the processor 800 is specifically configured to perform:

[0187] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0188] As an optional implementation, before identifying the methylation sites of the sequencing data based on their positions in the corresponding reference genome, the processor 800 is further configured to perform:

[0189] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0190] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0191] As an optional implementation, the processor 800 is specifically configured to perform:

[0192] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0193] As an optional implementation, the processor 800 is specifically configured to perform:

[0194] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0195] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0196] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0197] As an optional implementation, if the candidate methylation sites of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical at corresponding positions, then the processor 800 is further configured to execute:

[0198] By comparing the candidate methylation sites at corresponding positions in the first double-stranded DNA and the corresponding reference genome of the sequencing data, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained; or,

[0199] Candidate methylation sites at corresponding positions are compared between the second double-stranded DNA and the corresponding reference genome of the sequencing data. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0200] Based on the same inventive concept, this embodiment of the invention also provides a methylation detection device. Since this device is the same as the device in the method of this embodiment of the invention, and the principle of the device in solving the problem is similar to that of the method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0201] like Figure 9 As shown, the device includes:

[0202] The gene alignment module 900 is used to acquire sequencing data, determine the reference genome corresponding to the sequencing data, align the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome.

[0203] The site identification module 901 is used to identify methylation sites in the sequencing data based on their positions in the corresponding reference genome, and to use the identified methylation sites as candidate methylation sites.

[0204] The site filtering module 902 is used to identify the double-stranded DNA corresponding to the sequencing data based on the UMI information in the sequencing data, and filter the candidate methylation sites through the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site.

[0205] As an optional implementation, before comparing the sequencing data with the corresponding reference genome, a preprocessing module is further included, specifically for:

[0206] The sequencing data is preprocessed to obtain preprocessed sequencing data.

[0207] As an optional implementation, the preprocessing module is specifically used to preprocess the sequencing data in any one or more ways:

[0208] Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold;

[0209] Remove bases in the sequencing data whose base quality is below the second threshold;

[0210] Remove base sequences from sequencing data whose length is below the third threshold.

[0211] As an optional implementation, after obtaining the position of the sequencing data in the corresponding reference genome, a filtering module is further included, specifically for:

[0212] Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

[0213] As an optional implementation, the site identification module 901 is specifically used for:

[0214] If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

[0215] As an optional implementation, before identifying the methylation sites of the sequencing data based on their positions in the corresponding reference genome, a site conversion module is further included, specifically for:

[0216] The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and...

[0217] The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

[0218] As an optional implementation, the site filtering module 902 is specifically used for:

[0219] By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

[0220] As an optional implementation, the site filtering module 902 is specifically used for:

[0221] The double-stranded DNA corresponding to the sequencing data is converted into one strand of the double-stranded DNA template and another strand that is complementary to the first strand.

[0222] Determine the first double-stranded DNA sequence corresponding to one strand and the second double-stranded DNA sequence corresponding to the other strand;

[0223] Candidate methylation sites at corresponding positions are compared between the first and second double-stranded DNA sequences. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0224] As an optional implementation, if the candidate methylation sites of the first double-stranded DNA sequence and the second double-stranded DNA sequence are identical at corresponding positions, then the site filtering module 902 is further configured to:

[0225] By comparing the candidate methylation sites at corresponding positions in the first double-stranded DNA and the corresponding reference genome of the sequencing data, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained; or,

[0226] Candidate methylation sites at corresponding positions are compared between the second double-stranded DNA and the corresponding reference genome of the sequencing data. Based on the comparison results, the candidate methylation sites are filtered to obtain the target methylation sites.

[0227] Based on the same inventive concept, this disclosure provides a computer storage medium comprising: computer program code, which, when executed on a computer, causes the computer to perform any of the methylation detection methods described above. Since the principle by which the computer storage medium solves the problem is similar to that of the methylation detection method, the implementation of the computer storage medium can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0228] In specific implementation, computer storage media can include: Universal Serial Bus Flash Drive (USB), portable hard drive, Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disk, and other storage media that can store program code.

[0229] Based on the same inventive concept, this disclosure also provides a computer program product, which includes computer program code that, when executed on a computer, causes the computer to perform any of the methylation detection methods described above. Since the principle by which the above-described computer program product solves the problem is similar to that of the methylation detection method, the implementation of the above-described computer program product can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0230] Computer program products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0231] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0232] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 Devices that specify the functions in one or more boxes.

[0233] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction device, which is implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0234] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0235] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method of methylation detection, characterized in that, The method includes: Acquire sequencing data, determine the reference genome corresponding to the sequencing data, compare the sequencing data with the corresponding reference genome, and obtain the position of the sequencing data in the corresponding reference genome; Based on the location of the sequencing data in the corresponding reference genome, the methylation sites of the sequencing data are identified, and the identified methylation sites are used as candidate methylation sites. Based on the UMI information in the sequencing data, the double-stranded DNA corresponding to the sequencing data is identified. Candidate methylation sites are filtered using the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site. Specifically, the double-stranded DNA corresponding to the sequencing data is converted into one strand of a double-stranded DNA template and another strand complementary to that strand. The first double-stranded DNA sequence corresponding to the first strand and the second double-stranded DNA sequence corresponding to the other strand are determined. Candidate methylation sites at corresponding positions of the first and second double-stranded DNA sequences are compared, and the candidate methylation sites are filtered based on the alignment results to obtain the target methylation site. If one strand is the sense strand and the other strand is the antisense strand, the UMI sequences of the sense and antisense strands are completely inversely complementary. The differences on the left and right sides of the alignment position between the sense and antisense strands and the same reference genome are m and n, respectively, and the total positional difference between the two strands is G, where G = m + n / 2. n; Based on the alignment of the sense and antisense strands to the same position in the reference genome and the consistency of the UMI, the length of the overlapping region of the sense and antisense strands is used to determine the double-stranded DNA template in the sequencing data; if the candidate methylation sites of the first double-stranded DNA sequence and the second double-stranded DNA sequence are consistent at the corresponding positions, the process further includes: aligning the candidate methylation sites of the first double-stranded DNA and the corresponding reference genome of the sequencing data at the corresponding positions, filtering the candidate methylation sites according to the alignment results, and obtaining the target methylation site; or, aligning the candidate methylation sites of the second double-stranded DNA and the corresponding reference genome of the sequencing data at the corresponding positions, filtering the candidate methylation sites according to the alignment results, and obtaining the target methylation site.

2. The method of claim 1, wherein, Before comparing the sequencing data with the corresponding reference genome, the process also includes: The sequencing data is preprocessed to obtain preprocessed sequencing data.

3. The method of claim 2, wherein, Sequencing data can be preprocessed using one or more methods: Remove adapter sequences in the sequencing data that are longer than or equal to the first threshold; Remove bases in the sequencing data whose base quality is below the second threshold; Remove base sequences from sequencing data whose length is below the third threshold.

4. The method of claim 1, wherein, After obtaining the position of the sequencing data in the corresponding reference genome, the method further includes: Based on the UMI information of the sequencing data, repetitive sequences contained in the sequencing data are filtered out, which is used to identify the methylation sites in the filtered sequencing data.

5. The method of claim 1, wherein, The process of identifying methylation sites in the sequencing data based on their location in the corresponding reference genome, and using the identified methylation sites as candidate methylation sites, includes: If the sequencing data and the corresponding reference genome have the same base information at the same position, then the candidate methylation sites in the sequencing data are determined based on the site information corresponding to the same base information.

6. The method of claim 1, wherein, Before identifying the methylation sites of the sequencing data based on their location in the corresponding reference genome, the process further includes: The C sites in the reference genome are converted to T sites to obtain the first reference genome, which is used to identify methylation sites in the sequencing data; and... The G sites in the reference genome are converted to A sites to obtain the second reference genome, which is used to identify methylation sites in the sequencing data.

7. The method of claim 1, wherein, The process of filtering candidate methylation sites using the site information of the double-stranded DNA corresponding to the sequencing data to obtain the target methylation site includes: By comparing sequencing data with corresponding double-stranded DNA at the corresponding candidate methylation sites, and filtering the candidate methylation sites based on the comparison results, the target methylation sites are obtained.

8. An electronic device, comprising: The device includes a processor and a memory for storing a program executable by the processor, and the processor for reading the program in the memory and performing the steps of the method according to any one of claims 1 to 7.

9. A computer storage medium having stored thereon a computer program, characterized in that When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 7.