Method, device, medium and program for detecting mutations on basis of methylation data

By using a methylation-based method, 3-base genome data is reduced to 4-base genome data. By combining the base complementarity between the positive and negative strands, candidate mutation sites are identified, which solves the problem of low accuracy in detecting low-abundance mutations in traditional methods and achieves accurate detection at methylation sequencing depth.

WO2026044955A1PCT designated stage Publication Date: 2026-03-05GUANGZHOU BURNING ROCK DX CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133822
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-26
Filing Date
2024-11-22
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Traditional mutation detection methods based on methylation data struggle to improve the detection accuracy for low-abundance mutations, especially in early tumor detection and minimal residual disease scenarios, where the frequency of tumor-derived mutations is as low as one in a thousand or even less, resulting in low detection accuracy.

Method used

By acquiring the methylation data of the sample to be tested, and using the base composition data of the 3-base genome after chemical or enzymatic transformation, combined with the complementary base pairing relationship between the positive and negative strands and the genotype correspondence before and after transformation, the base composition data of the 4-base genome before transformation is determined, thereby identifying candidate mutation sites and filtering them to generate mutation detection results.

Benefits of technology

Accurate detection of low-abundance SNV mutation sites was achieved at methylation sequencing depth, improving the detection accuracy of low-abundance mutations. It can simultaneously detect methylation signals and gene mutations in one reaction, and the two detection modes do not interfere with each other.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133822_05032026_PF_FP_ABST
    Figure CN2024133822_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A method, device, medium and program for detecting mutations on the basis of methylation data. The method comprises obtaining methylation data of a sample to be tested; on the basis of the methylation data, obtaining base composition data of a 3-base genome after chemical or enzymatic conversion; on the basis of the base composition data of the 3-base genome after chemical or enzymatic conversion, determining base composition data of a 4-base genome before chemical or enzymatic conversion according to the complementary base-pairing relationship between a positive strand and a negative strand, and / or the corresponding relationship between genotypes before and after chemical or enzymatic conversion; on the basis of the base composition data of the 4-base genome, identifying candidate mutation sites; and filtering the candidate mutation sites to generate detection results relating to single nucleotide variation mutation sites. According to the method, the single nucleotide variations can be detected from whole genome methylation data, and the accuracy of detection for low abundance mutations can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, equipment, media, and procedures for detecting mutations based on methylation data.

[0001] This application claims priority to Chinese Patent Application No. 202411184079.8, filed on August 26, 2024, entitled "Method, apparatus, medium and procedure for detecting mutations based on methylation data", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This invention relates generally to the processing of biological information, and more specifically to methods, computing devices, computer storage media, and computer program products for detecting mutations based on methylation data. Background Technology

[0003] Traditional methods for determining DNA methylation levels, such as bisulfite sequencing, utilize sulfite treatment to measure DNA methylation. Currently, commonly used DNA methylation sequencing technologies involve chemical (e.g., bisulfite conversion, BC) or enzymatic (EC) conversions, i.e., BC or EC conversions. Specifically, bisulfite or enzymatic methods are used to convert cytosine (C) in DNA to uracil (U), while methylated C sites do not undergo this conversion. Methylation generally occurs at CpG sites (where the two bases C and G are adjacent). The proportion of C successfully converted to U at a given CpG site indicates the methylation level at different DNA sites. In methylation data, because most C in CpG sites is converted to U (corresponding to T in DNA), traditional mutation detection software cannot directly distinguish between C→T and G→A mutations.

[0004] Traditional mutation detection methods based on methylation data include analysis software such as BC, EC-SNPer, and Bis-SNP, which detect single nucleotide variants from whole-genome methylation data. These methods all employ Bayesian methods to estimate mutation types. However, these methods are designed for high-frequency germline mutations (SNPs). In scenarios such as early tumor detection and minimal residual disease (MRD), the frequency of tumor-derived mutations can be as low as one in a thousand or even lower, making them comparable to background noise. This results in low accuracy when detecting low-abundance mutations.

[0005] In summary, the traditional mutation detection methods based on methylation data have the following drawbacks: they are difficult to improve the accuracy of detecting low-abundance mutations. Summary of the Invention

[0006] This invention provides a method, computing device, computer storage medium, and computer program product for detecting mutations based on methylation data. It can not only detect single nucleotide variations from whole-genome methylation data, but also significantly improve the accuracy of detecting low-abundance mutations.

[0007] According to a first aspect of the present invention, a method for detecting mutations based on methylation data is provided. The method includes: acquiring methylation data of a sample to be tested; obtaining, based on the methylation data, the base composition data of a 3-base genome after chemical or enzymatic (i.e., BC or EC) transformation; determining, based on the base composition data of the 3-base genome, the base composition data of a 4-base genome before BC or EC transformation according to the complementary base pairing relationship between the positive and negative strands and / or the correspondence between genotypes before and after BC or EC transformation; identifying candidate mutation sites based on the base composition data of the 4-base genome; optionally, the method further includes: filtering the candidate mutation sites to generate detection results for mutation sites related to single nucleotide variations.

[0008] According to a second aspect of the invention, a computing device is also provided, the device comprising: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute one or more programs to cause the device to perform the method of the first aspect of the invention.

[0009] According to a third aspect of the invention, a non-transient computer-readable storage medium is also provided. This non-transient computer-readable storage medium stores machine-executable instructions that, when executed, cause a machine to perform the method of the first aspect of the invention.

[0010] According to a fourth aspect of the invention, a computer program product is also provided. The computer program product includes instructions that, when executed by a machine, implement the method of the first aspect of the invention.

[0011] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the invention, nor is it intended to limit the scope of the invention.

[0012] The present invention achieves the following technical advantages over the prior art: The method of the present invention only requires data analysis and processing of the methylation data obtained from methylation detection to obtain gene mutation detection at the same time as obtaining methylation detection results. That is, the present invention can detect methylation signals and gene mutations simultaneously in one reaction, and the two detection modes do not interfere with each other. Furthermore, it can achieve accurate detection of low abundance (AF as low as 0.4%) SNV mutation sites at methylation sequencing depth (e.g., 10000X). Attached Figure Description

[0013] Figure 1 shows a schematic diagram of a system for implementing a method for detecting mutations based on methylation data according to an embodiment of the present invention.

[0014] Figure 2 shows a flowchart of a method for detecting mutations based on methylation data according to an embodiment of the present invention.

[0015] Figure 3 shows a schematic diagram of base changes at the same site according to an embodiment of the present invention during BC or EC transformation and real mutation.

[0016] Figure 4 illustrates a method for determining a 3-base genome after BC or EC conversion during methylation and a 4-base genome before BC or EC conversion according to embodiments of the present invention.

[0017] Figure 5 shows a flowchart of a method for calculating the number of G bases in the negative chain that are complementary to the T bases in the positive chain, which are converted from C bases via BC or EC.

[0018] Figure 6 shows a flowchart of a method for calculating the number of bases in the positive chain that are converted from C bases to T bases via BC or EC according to an embodiment of the present invention.

[0019] Figure 7 shows a flowchart of a method for determining the number of T bases in the positive chain before BC or EC conversion according to an embodiment of the present invention.

[0020] Figure 8 schematically illustrates a block diagram of an electronic device suitable for implementing embodiments of the present invention.

[0021] Figure 9 schematically illustrates the detection sensitivity for different cancer types and stages according to an embodiment of the present invention.

[0022] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0023] Preferred embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0024] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects.

[0025] As described above, traditional mutation detection methods based on methylation sequencing data typically employ Bayesian methods to estimate mutation types. These methods are primarily designed for high-frequency germline mutations (SNPs), assuming, for example, that the mutation being tested is diploid and relying on the distribution of mutation frequency in the population. These premises and assumptions are not applicable to scenarios such as early tumor detection and MRD (Malignant Transformation Disease). This is because, in blood samples from early-stage cancer patients, the frequency of tumor-derived mutations may be as low as one in a thousand, orders of magnitude lower than some background noise terms. Therefore, the accuracy of traditional methods for detecting low-abundance mutations is relatively low, requiring appropriate methods to distinguish between background noise and low-frequency mutation signals. Consequently, there is a limitation in improving the accuracy of low-abundance mutation detection.

[0026] In order to at least partially address one or more of the above-mentioned problems and other potential problems, exemplary embodiments of the present invention propose a scheme for mutation detection based on methylation data.

[0027] In this scheme, the relevant terms are defined as follows:

[0028] Terminology definition:

[0029] Methylation:

[0030] In this article, methylation and DNA methylation have the same meaning, referring to the methylation process that occurs at the 5th carbon atom of cytosine in CpG dinucleotides. As a relatively stable modification state, it can be inherited by newly generated daughter DNA during DNA replication under the action of DNA methyltransferases. It is an important epigenetic mechanism. When DNA is methylated, methylation of the gene promoter region can lead to transcriptional silencing of tumor suppressor genes, so it is closely related to the occurrence of tumors.

[0031] Methylation detection methods:

[0032] This invention utilizes various existing techniques to detect whether a site is methylated. Common methods include chemical or enzymatic BC or EC conversion, where one of the unmethylated cytosine BC or EC bases is converted to uracil (U) or a base that is substantially equivalent to uracil in base pairing (e.g., dihydrouracil, DHU). In subsequent amplification, the corresponding uracil pairs as thymine (T) with adenine (A), resulting in the unmethylated cytosine appearing as thymine in the detection results (e.g., sequencing results). By comparing with a reference sequence, it is possible to determine whether cytosine in a DNA molecule or fragment is methylated.

[0033] BC or EC conversion:

[0034] In this article, "BC or EC conversion" refers to the conversion of unmethylated cytosine bases (BC or EC) in the DNA of the sample to uracil methyl groups, so that the methylated or methylated cytosine at the methylation site is represented as thymine in the sequencing results. This includes the chemical or enzymatic BC or EC conversions mentioned above. For example, a commonly used BC or EC conversion to BS (Bisulfite) conversion occurs when the original DNA sequence is treated with sulfite during methylation sequencing. Unmethylated C becomes U (which becomes T during subsequent DNA replication), while methylated C (including 5mC and 5hmC) remains unchanged.

[0035] Sample to be tested:

[0036] The sample to be tested in this article can be a DNA sample from any source and of any length, such as genomic DNA, DNA fragments, or extracellular cell-free DNA, such as the subject's tissue, plasma, or blood.

[0037] Methylation sequencing data:

[0038] The methylation detection data of the test sample is obtained using existing technologies, such as, but not limited to, the following steps: First, DNA extraction and purification of the test sample is performed; then, BC or EC transformation is performed on the DNA of the extracted and purified test sample; whole-genome library construction is performed (e.g., but not limited to, whole-genome library construction based on ELSA-seq); then, specific target regions are co-enriched in the same hybridization capture reaction using customized cancer methylation profiling RNA probes and mutation detection RNA probes; the target library is quantified by real-time PCR; and sequencing is performed using a sequencer to generate methylation detection data of the test sample.

[0039] Methylation data:

[0040] In this paper, methylation data can be obtained using any method known in the art, such as methylation sequencing data obtained through methylation sequencing technology. The methylation data can also be data obtained through further screening of methylation sequencing data, such as data quality preprocessing, using trimmomatic to remove low-quality sequences and artificially added bases in the sequence, evaluation (using FASTP software), genome alignment (using Bismark software), or removal of duplicate data resulting from sample / experimental techniques.

[0041] 3-base genome base composition data:

[0042] In methylation detection, sequencing results where all C bases are converted to T bases are generally referred to as a 3-base genome. In this article, "3-base genome base composition data" refers to the proportion of 3 bases (A, G, T) at each site in the methylation data.

[0043] 4. Base composition data of the genome:

[0044] In this article, "4-base genome base composition data" refers to the composition ratio data of the original base composition data (A, G, C, T) of each site in the sample DNA that has not undergone BC or EC transformation during the methylation detection process.

[0045] genotype:

[0046] In this article, "genotype" refers to the base composition at a specific site on a DNA sample.

[0047] Gene frequency:

[0048] In this article, "gene frequency" refers to the proportion of mutated genotypes at a certain site in a single DNA sample to non-mutated genotypes.

[0049] Specifically, the present invention provides a mutation detection scheme based on methylation data, comprising: acquiring methylation data of a sample to be tested; obtaining base composition data of a 3-base genome after BC or EC transformation based on the methylation data; determining the base composition data of a 4-base genome before BC or EC transformation based on the base composition data of the 3-base genome after BC or EC transformation, according to the complementary base pairing relationship between the positive and negative strands and / or the correspondence between genotypes before and after BC or EC transformation; identifying candidate mutation sites based on the base composition data of the 4-base genome; optionally, the method further comprises: filtering the candidate mutation sites to generate detection results for mutation sites related to single nucleotide variations.

[0050] This invention restores the 3-base genome after BC or EC conversion to the 4-base genome before BC or EC conversion, and then detects mutations based on the 4-base genome algorithm. Therefore, this invention can detect methylation signals and gene mutations simultaneously in one reaction, and the two detection modes do not interfere with each other. It can also achieve accurate detection of low abundance (AF as low as 0.4%) SNV mutation sites at methylation sequencing depth (e.g., 10000X).

[0051] Figure 1 illustrates a schematic diagram of a system 100 for implementing a method for detecting mutations based on methylation data according to an embodiment of the present invention. As shown in Figure 1, the system 100 includes a computing device 110 and a sequencing device 130. In some embodiments, the computing device 110 and the sequencing device 130 interact with each other via a network (not shown).

[0052] Regarding sequencing device 130, it is used, for example, for sample library preparation and sequencing of target libraries to generate methylation detection data for the test sample or multiple tissue samples or blood samples. Various methods can be used for sample library preparation, and in some embodiments, such as, but not limited to, ELSA-seq (Burning Rock Biotech, Guangzhou, China), sample library preparation is performed. The sample library preparation method includes, for example, the following steps: first, DNA extraction and purification; then, methylation treatment; subsequently, construction of a whole-genome library using ELSA-seq; co-enrichment of specific target regions using customized cancer methylation profiling RNA probes and mutation detection RNA probes in the same or separate hybridization capture reactions; and quantification of the target library by real-time PCR. Finally, targeted sequencing is performed using a next-generation sequencing instrument, but not limited to, Illumina's NovaSeq 6000. It should be understood that the above sample library preparation and target library sequencing methods are merely exemplary, and other methods can also be used for sample library preparation and target library sequencing.

[0053] Regarding computing device 110, it can be used to detect mutations based on methylation sequencing data. In some embodiments, computing device 110 may have one or more processing units, including dedicated processing units such as GPUs, FPGAs, and ASICs, and general-purpose processing units such as CPUs. One or more virtual machines may also run on each computing device. Computing device 110 includes, for example, a methylation data acquisition unit 112, a base composition data acquisition unit 114 for a chemically or enzymatically transformed 3-base genome, a base composition data determination unit 116 for a chemically or enzymatically transformed 4-base genome, a candidate mutation site identification unit 118, and a detection result generation unit 120. The methylation data acquisition unit 112, the base composition data acquisition unit 114 for a chemically or enzymatically transformed 3-base genome, the base composition data determination unit 116 for a chemically or enzymatically transformed 4-base genome, the candidate mutation site identification unit 118, and the detection result generation unit 120 may be configured on one or more computing devices 110.

[0054] Regarding the methylation data acquisition unit 112, it is used to acquire the methylation data of the sample to be tested.

[0055] Unit 114 for obtaining base composition data of a chemically or enzymatically transformed 3-base genome, which is used to obtain base composition data of a chemically or enzymatically transformed 3-base genome based on methylation data.

[0056] The base composition data determination unit 116 for the 4-base genome before chemical or enzymatic transformation is used to determine the base composition data of the 4-base genome before chemical or enzymatic transformation based on the base composition data of the 3-base genome after BC or EC transformation, according to the complementary base pairing relationship between the positive and negative strands and / or the correspondence between the genotypes before and after chemical or enzymatic transformation.

[0057] Regarding the candidate mutation site identification unit 118, it is used to identify candidate mutation sites based on the base composition data of the 4-base genome.

[0058] The detection result generation unit 120 is used to filter candidate mutation sites to generate detection results for mutation sites of single nucleotide variations.

[0059] The following description, in conjunction with Figures 2, 3, and 4, describes a mutation detection method based on methylation data according to an embodiment of the present invention. Figure 2 shows a flowchart of a mutation detection method 200 based on methylation data according to an embodiment of the present invention. It should be understood that method 200 can be performed, for example, at the electronic device 800 described in Figure 8. It can also be performed at the computing device 110 described in Figure 1. It should be understood that method 200 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.

[0060] In step 202, computing device 110 acquires methylation data of the sample to be tested.

[0061] Regarding the methylation data of the sample to be tested, it is acquired, for example, by computing device 110 from sequencing device 130. In some embodiments, the methylation data of the sample to be tested is generated, for example, by the following steps: first, DNA extraction and purification of the sample to be tested is performed; then, enzymatic BC or EC transformation is performed on the DNA of the extracted and purified sample to be tested; whole-genome library construction is performed (e.g., but not limited to whole-genome library construction based on ELSA-seq method (Burning Rock Biotech, Guangzhou, China)); subsequently, specific target regions are co-enriched in the same hybridization capture reaction using customized cancer methylation profiling RNA or DNA probes and mutation detection RNA or DNA probes; the target library is quantified by real-time PCR; and sequencing is performed by a sequencer to generate methylation data of the sample to be tested.

[0062] For example, computing device 110 uses the Illumina platform data filtering tool trimmomatic to remove low-quality sequences and artificially added bases from the methylation data of the sample to be tested; then, it uses bismark software to align the raw methylation data BC or EC-seq to the human reference genome hg19, and then removes polymerase chain reaction amplification (PCR duplication) based on the sequence alignment results and UMI.

[0063] At step 204, computing device 110 obtains base composition data of the chemically or enzymatically transformed 3-base genome based on methylation data. For example, based on methylation data, it obtains the number of bases (i.e., the total number of reads for each base) of the OT and OB strands after chemically or enzymatically (or “BC or EC”) transformation at the same mutation site.

[0064] Figure 4 illustrates a schematic diagram of the method for determining a 3-base genome after BC or EC conversion during methylation and a 4-base genome before BC or EC conversion according to embodiments of the present invention. As shown in Figure 4, marker 430 indicates a site on the OT chain that is converted from the original 4-base genotype 410 (BC or EC) to the 3-base genotype indicated by marker 430 via methylation process 420. The C base 412 in the original genotype 410 is converted to the T base 432 in the 3-base genotype 430 via methylated BC or EC conversion. Therefore, the number of all T bases in the 3-base genotype 430 of the OT chain, converted via BC or EC (i.e., ...) The observed value is equal to the number of all bases in the OT chain that are T bases before and after BC or EC conversion (i.e., The number of bases in the OT chain that are converted from C bases to T bases via BC or EC (i.e., The sum of ) should be noted. It should be noted that due to the methylation process, only one chain is involved. For example, in Figure 4, the 4-base genotype 410 of the OT chain is converted to the 3-base genotype 430 via BC or EC, while the 440 of the OB chain remains a 4-base genotype.

[0065] In step 206, the computing device 110, based on the base composition data of the 3-base genome after chemical or enzymatic transformation, determines the base composition data of the 4-base genome before chemical or enzymatic transformation according to the complementary base pairing relationship between the positive and negative strands and / or the correspondence between genotypes before and after chemical or enzymatic transformation. For example, for the same mutation site, the number of bases (i.e., the total number of reads for each base) of the OT and OB strands before BC or EC transformation is restored. By employing the above methods, the present invention can restore the base change from C bases to T bases caused by methylation BC or EC transformation.

[0066] For example, as shown in Figure 4, the computing device 110 uses the complementary base pairing relationship between the OT and OB chains to reduce the base change from C to T caused by methylation of BC or EC transformation 450, so as to convert the observed value of the 3-base genotype 430 after methylation of BC or EC transformation into the true value of the 4-base genotype 460 before BC or EC transformation. Thus, based on the true value of the 4-base genotype 460, mutations can be accurately identified.

[0067] Methods for determining the base composition data of a 4-base genome before BC or EC transformation include, for example: determining the distribution data of the number of A, T, G, and C bases (corresponding to the number of reads) of the positive and negative strands after BC or EC transformation based on the base composition data of the 3-base genome after BC or EC transformation; and determining the genotype and gene frequency of the 4-base genome before BC or EC transformation based on the complementary base pairing relationship between the positive and negative strands, the correspondence between genotypes before and after BC or EC transformation, and the distribution data of the number of A, T, G, and C bases of the positive and negative strands after BC or EC transformation.

[0068] It should be understood that during methylation, generally only one strand (positive or negative) of the chain is altered, while the other strand usually remains unchanged. That is, for the same site, only one strand (original top strand, OT chain) or the original bottom strand (OB chain) will be affected by BC or EC conversion. In contrast, during mutation, the bases of both the positive and negative strands generally change. As shown in Figure 4, the OT chain undergoes BC or EC conversion (C base to T base), while the OB chain does not undergo BC or EC conversion.

[0069] Figure 3 illustrates the base changes at the same site during BC or EC conversion and true mutation according to an embodiment of the present invention. For example, as shown in Figure 3, for the same site, as indicated by label 310, the original OT strand has a C base and the OB strand has a G base. If a true C→T base mutation occurs, as indicated by label 314, while the C base in the positive strand converts to a T base, the corresponding G base in the negative strand also converts to an A base. However, if a BC or EC conversion occurs at the same site, as indicated by label 312, the C base in the positive strand converts to a T base, while the corresponding G base in the negative strand remains a G base and does not undergo a conversion.

[0070] Therefore, the above characteristics can be used to distinguish whether the base changes originate from methylation or mutation. Table 1 below schematically illustrates the correspondence of base changes before and after BC or EC transformation via methylation. According to Table 1, for sites with an A or G base genotype before BC or EC transformation, after BC or EC transformation via methylation, a T base will be observed in the OB strand, and an A base and a G base will be observed in the OT strand, respectively. For sites with a T or C base genotype before BC or EC transformation, after BC or EC transformation via methylation, the corresponding observed values ​​in the OT strand are all T bases, which cannot be distinguished; the corresponding observed values ​​in the OB strand are A bases and G bases, which can be correctly distinguished. Therefore, based on the correspondence of genotypes before and after BC or EC transformation, combined with the base data observed in the OT and OB strands before BC or EC transformation, the base composition data of the 4-base genome before BC or EC transformation can be determined.

[0071] Table 1

[0072] In Table 1 above, T C The BC or EC bases, representing methylation, are initially C bases before conversion and are then converted to T bases via methylation. In other words, a C base is converted to a T base via BC or EC. T T The BC or EC bases representing methylation are both T bases before and after the transformation. Assuming the current site serving as the template is a C base, then if the observed value of the current site is T... T If the observed value at the current site is T, it indicates a mutation from a C base to a T base; if the observed value at the current site is T... C The bases indicate that the base change at the current site is not due to a mutation, but rather to the conversion of methylated BC or EC.

[0073] It should be understood that the OT and OB chains form corresponding double helices, and the relative proportions of complementary base pairs in the OT and OB chains are equal. Specifically, the ratio of the number of A, T, C, and G bases in the OT chain is equal to the ratio of the number of T, A, G, and C bases in the OB chain.

[0074] In high-throughput sequencing, the OT and OB strands are sequenced separately. While the direct correspondence between the bases observed in the OT and OB strands is difficult to determine, the number of A, T, G, and C bases observed at a specific locus in both strands (i.e., the number of reads for that corresponding base) can be obtained. Therefore, by utilizing the principle of complementary base pairing between the positive and negative strands of DNA, the correspondence between genotypes before and after BC or EC transformation, and the distribution data of the number of A, T, G, and C bases in the positive and negative strands, the genotypes and gene frequencies of the four-base genome before BC or EC transformation can be determined.

[0075] For example, based on the principle of complementary base pairing between the positive and negative strands of DNA, it can be assumed that the frequencies of complementary base pairings in the OT and OB strands are approximately the same. The following formula (1) schematically illustrates the frequency relationship between the bases in the OT and OB strands.

[0076] In the above formula (1), This represents the total number of A bases in the OT chain. This represents the total number of A bases in the OB chain. This represents the total number of C bases in the OT chain. This represents the total number of C bases in the OB chain. This represents the number of bases in the OT chain that are T bases before and after BC or EC conversion. This represents the number of bases in the OB chain that are T bases before and after the BC or EC conversion. This represents the number of bases in the OT chain that are converted from C bases to T bases via BC or EC. This represents the number of bases in the OB chain that are converted from C bases to T bases via BC or EC. Representative and The number of complementary G bases, that is, the number of G bases in the OB chain that are complementary to all the C bases in the OT chain that are converted to T bases via BC or EC. Representative and The number of complementary G bases, that is, the number of G bases in the OT chain that are complementary to all the C bases in the OB chain that are converted to T bases via BC or EC.

[0077] It should be understood that, based on the complementary base pairing relationship between the positive and negative strands and the correspondence between the genotypes before and after BC or EC transformation, the following formulas (2) to (3) can be derived.

[0078] According to formulas (2) and (3), the number of G bases in the OB chain that are complementary to all C bases in the OT chain is determined. The number of G bases in the OB chain that are complementary to all C bases converted to T bases via BC or EC in the OT chain. ratio Equal to the number of C bases in the OT chain The number of bases in the OT chain that are converted from C bases to T bases via BC or EC. ratio And the number of G bases in the OT chain that are complementary to all C bases in the OB chain. The number of G bases in the OT chain that are complementary to all C bases converted to T bases via BC or EC in the OB chain. ratio Equal to the number of C bases in the OB chain The number of bases in the OB chain that are converted from C bases to T bases via BC or EC. ratio This is used to determine the genotype and gene frequency of the 4-base genome before BC or EC transformation.

[0079] The following, in conjunction with formulas (4) to (7), further explains the method for determining the genotype and gene frequency of the 4-base genome before BC or EC transformation.

[0080] As shown in formulas (4) and (5), the number of all T bases in the OT chain is [missing information]. This equals the number of bases in the OT chain that are T bases before and after the BC or EC conversion. The number of bases in the OT chain that are converted from C bases to T bases via BC or EC. sum And the number of bases that make all T bases in the OB chain This equals the number of bases in the OB chain that are T bases before and after the BC or EC conversion. The number of bases in the OB chain that are converted from C bases to T bases via BC or EC. sum This is used to determine the genotype and gene frequency of the 4-base genome before BC or EC transformation.

[0081] As shown in formulas (6) and (7), the number of all G bases in the OT chain is [missing information]. This equals the number of G bases in the OT chain that are complementary to all the T bases in the OB chain that are converted from C bases via BC or EC. The number of G bases in the OT chain that are complementary to all C bases in the OB chain. sum And the number of bases that make all G bases in the OB chain This equals the number of G bases in the OB chain that are complementary to all the T bases in the OT chain that are converted from C bases via BC or EC. The number of G bases in the OB chain that are complementary to all C bases in the OT chain. sum

[0082] For example, based on the above formulas (1) to (7), the number of T bases, A bases, G bases, and C bases in the OT and OB chains before BC or EC conversion are finally obtained, that is, the base composition data of the 4-base genome before BC or EC conversion.

[0083] The following example, using formula (8), illustrates the method for calculating the number of T bases in the OT chain before BC or EC conversion. For example, the number of T bases in the OT chain before BC or EC conversion ( That is, the true number of T bases is equal to the number of T bases in the OT chain before and after BC or EC conversion. The following will explain this in conjunction with Figure 7 and Method 700. The calculation algorithm will not be elaborated here.

[0084] The following example, using formula (9), illustrates the method for calculating the number of A bases in the OT chain before BC or EC conversion. For example, the number of A bases in the OT chain before BC or EC conversion ( That is, the true number of A bases) is equal to the total number of A bases in the OT chain (i.e., ).

[0085] The following example, using formula (10), illustrates the method for calculating the number of G bases in the OT chain before BC or EC conversion. For example, the number of G bases in the OT chain before BC or EC conversion ( That is, the true number of G bases) is equal to the total number of G bases in the OT chain (i.e., ).

[0086] Formula (11) exemplarily illustrates a method for calculating the number of C bases in the OT chain before BC or EC conversion. For example, the number of C bases in the OT chain before BC or EC conversion ( That is, the actual number of C bases is equal to the total number of C bases in the OT chain. In addition to the number of bases in the OT chain that are converted from C bases to T bases via BC or EC about The calculation algorithm will be explained in conjunction with Figure 6 and Method 600 below, and will not be repeated here.

[0087] It should be understood that the same method can also be used to estimate the number of T bases, A bases, G bases, and C bases in the real OB chain before BC or EC conversion, which will not be elaborated here.

[0088] In step 208, computing device 110 identifies candidate mutation sites based on the base composition data of the 4-base genome.

[0089] For example, the computing device 110 adds up the number of bases (i.e., the total number of reads for each base) of the OT and OB strands before BC or EC transformation at the same mutation site to obtain the number of bases of each base before BC or EC transformation at the same mutation site, which is used to identify candidate mutation sites based on the number of bases of each base before BC or EC transformation. In some embodiments, the computing device 110 identifies a first candidate mutation site based on the number of bases (i.e., the total number of reads for each base) of the OT strand before BC or EC transformation at the same mutation site; identifies a second candidate mutation site based on the number of bases (i.e., the total number of reads for each base) of the OB strand before BC or EC transformation at the same mutation site; and cross-checks the first and second candidate mutation sites to determine the selected mutation site.

[0090] Specifically, for example, computing device 110 identifies mutation sites based on the base composition data of the pre-BC or EC 4-base genome estimated via step 206, using conventional mutation detection methods, such as, but not limited to, the VarScan2 mutation detection software. The VarScan2 mutation detection software provides three commands for mutation site identification: mpileup2snp / pileup2snp (used to call single nucleotide polymorphism (SNP) variants individually), mpileup2indel / pileup2indel (used to call insertion or deletion (indel) variants individually), and mpileup2cns / pileup2cns (used to call both SNP and indel variants simultaneously). In some embodiments, computing device 110 extracts a set of candidate point mutation sites whose allele frequencies exceed a set threshold based on the determined pre-BC or EC 4-base genome base composition data.

[0091] It should be understood that by employing the above-mentioned means, the method of the present invention only needs to perform data analysis and processing on the methylation data obtained from methylation detection to obtain gene mutation detection at the same time as obtaining methylation detection results.

[0092] In some embodiments, method 200 further includes step 210.

[0093] In step 210, the computing device 110 filters candidate mutation sites to generate detection results for mutation sites of single nucleotide variants.

[0094] A method for generating detection results for single nucleotide variant (SNV) mutation sites includes, for example, filtering candidate mutation sites to generate detection results for SNV mutation sites, including: filtering out candidate mutation sites that meet a first predetermined condition; and filtering the remaining candidate mutation sites based on the presence of insertion and deletion variations near the candidate mutation sites and alignment quality, in order to generate detection results for SNV mutation sites. The first predetermined condition includes at least one of the following: coverage less than a predetermined coverage threshold; mutation support count less than a predetermined support count threshold; and mutation frequency less than a predetermined frequency threshold.

[0095] It should be understood that sequencing errors are higher in low-coverage regions, affecting the reliability of test results. Therefore, it is necessary to filter out candidate mutation sites with coverage less than a predetermined coverage threshold.

[0096] A method for filtering out candidate mutation sites that meet a first predetermined condition may include, for example, determining a predetermined frequency threshold based on whether the relevant mutation of the candidate mutation site belongs to a known mutation and the mutation type. Specifically, for example, the computing device 110 determines whether the relevant mutation of the candidate mutation site belongs to a known mutation; if it is determined that the relevant mutation belongs to a known mutation, it determines the predetermined frequency threshold as a first value; if it is determined that the relevant mutation does not belong to a known mutation, it determines the predetermined frequency threshold as a second value, the second value being different from the first value. For the determined predetermined frequency threshold, the computing device 110 filters out candidate mutation sites whose mutation frequency is less than the predetermined frequency threshold.

[0097] In some embodiments, the method for determining a predetermined frequency threshold includes, for example, setting different predetermined frequency thresholds for different mutation types. For example, since the mutation type G>A is more prone to false positives, the predetermined frequency threshold corresponding to the mutation type G>A is higher than the corresponding predetermined frequency thresholds for other mutation types.

[0098] As mentioned above, the reason for filtering remaining candidate mutation sites based on the presence of insertions and deletions near them is that, since methylated data is a 3-base genome, its data complexity is lower than that of a 4-base genome, making it more prone to generating insertions and deletions during sequence alignment. Therefore, insertions and deletions near the target site can easily lead to the target site being incorrectly identified as a mutation site. Thus, it is necessary to filter for insertions and deletions near the target site, and reads with low alignment quality are also filtered out.

[0099] The present invention further provides a method for implementing mutation detection based on methylation data (e.g., the MMcall algorithm). This method includes, for example, data preparation, algorithm construction, and preferred filtering steps. The following description, in conjunction with specific embodiments, exemplifies each step of the above method.

[0100] The data preparation steps are similar to the example given earlier. First, the sample library is prepared using... This method (Burning Rock Biotech, Guangzhou, China) mainly includes the following steps: 1) DNA extraction and purification; 2) enzymatic BC or EC transformation; 3) through... The method involved: 4) constructing a whole-genome library; 5) using customized cancer methylation profiling RNA probes and mutation detection RNA probes to co-enrich specific target regions in the same hybridization capture reaction; and 6) quantifying the target library using real-time polymerase chain reaction (PCR). Sequencing was finally performed using an Illumina NovaSeq 6000.

[0101] Subsequently, for example, the Illumina platform data filtering tool Trimmomatic is used to remove low-quality sequences and artificially added bases from the methylation data of the test samples. Then, the raw methylation data is aligned to the human reference genome hg19 using bismark software using BS-seq. Based on the sequence alignment results and UMI, polymerase chain reaction (PCR) duplication is removed. Thus, the methylation data is obtained.

[0102] Regarding the algorithm construction steps, it should be understood that in methylation data, unmethylated C bases will be converted to T bases via BC or EC conversion. This BC or EC conversion occurs in both the OT and OB chains. The principle can be summarized as the correspondence shown in Table 1 above. Table 1 specifically indicates that for sites where the genotype before BC or EC conversion is A or G, a T base will be observed in the OB chain, but both A and G bases can be observed in the OT chain. Similarly, if the genotype before BC or EC conversion (or the "true genotype") is T or C, the BC or EC conversion results in T bases in the OT chain (or "observed values"), making them indistinguishable. However, the BC or EC conversion results in A and G bases in the OB chain, which can be correctly distinguished. Based on the above principles, this invention can determine the genotype of the 4-base genome before BC or EC conversion based on 3-genome data, using different observed values ​​from the OT and OB chains.

[0103] It should be understood that in high-throughput sequencing, the OT and OB strands are sequenced separately, and we do not know their correspondence; we can only know the total number of A, T, G, and C bases observed in the OT and OB strands at a certain locus, but we do not know which A bases observed in the OT strand correspond to which T bases observed in the OB strand. To accurately estimate allele frequency (AF), this invention utilizes the principle of complementary pairing between the two DNA strands, calculating the true genotype and gene frequency by analyzing the distribution of the total number of A, T, G, and C bases in the OT and OB strands.

[0104] Specifically, in order to achieve this goal, we set that the frequencies of each base in the OT chain and the OB chain are approximately the same, as shown in formula (1) above.

[0105] To correspond with this, we further subdivide the total number of G bases, which leads to the formula (2) above. That is, the number of G bases in the OB chain that are complementary to all C bases in the OT chain. The number of G bases in the OB chain that are complementary to all C bases converted to T bases via BC or EC in the OT chain. ratio Equal to the number of C bases in the OT chain The number of bases in the OT chain that are converted from C bases to T bases via BC or EC. ratio

[0106] In addition, this makes the number of all T bases in the OT chain... This equals the number of bases in the OT chain that are T bases before and after the BC or EC conversion. The number of bases in the OT chain that are converted from C bases to T bases via BC or EC. sum And the number of bases that make all T bases in the OB chain This equals the number of bases in the OB chain that are T bases before and after the BC or EC conversion. The number of bases in the OB chain that are converted from C bases to T bases via BC or EC. sum As indicated by formulas (4) and (5) above.

[0107] Taking the OT chain as an example, under the condition that the above formulas (2), (4) and (5) are satisfied, the present invention, based on the base composition data of the 3-base genome after BC or EC transformation, and according to the base complementary pairing relationship between the positive and negative strands and / or the correspondence between the genotypes before and after BC or EC transformation, derives the proportional relationship indicated by formulas (12) and (9) below.

[0108] From (12) and (9) above, we can further derive the proportional relationship indicated by formula (13) below.

[0109] This leads to the equation (14) below indicating the following information about... (It represents the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are converted from C bases via BC or EC).

[0110] In the above formula (14), the formula is indicated for calculating... The algorithm is as follows: That is, count the number of all A bases in the OB chain. Number of bases with G base Add them together to obtain the sum of the first dibase numbers of all A and G bases in the OB chain. The number of bases of all C bases in the OT chain Divide by the sum of the total number of T and C bases in the OT chain. In order to obtain the first proportion of C bases in the OT chain And the number of all G bases in the OB chain. In the middle, subtract the first proportion of C bases in the OT chain. The sum of the first dibases of all A and G bases in the OB chain The product of these numbers is used to obtain the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are transformed from C bases via BC or EC.

[0111] Similarly, the present invention can use the algorithms indicated by the following formulas (15) and (16) to calculate the number of bases in the OT chain that are converted from C bases to T bases via BC or EC. And the number of bases in the OT chain that are T bases before and after BC or EC conversion.

[0112] Therefore, we can obtain: the number of T bases in the OT chain before BC or EC conversion ( That is, the actual number of T bases), and the number of A bases in the OT chain before BC or EC conversion ( That is, the actual number of bases in A, and the number of bases in C before BC or EC conversion in the OT chain. That is, the actual number of C bases), and the number of G bases in the OT chain before BC or EC conversion ( That is, the algorithm for calculating the true number of bases in G bases. For example, as shown in formulas (8) to (11) below.

[0113] It should be understood that the OB chain can also use the same method to calculate the true number of bases for the four bases. The algorithm construction process will not be elaborated here.

[0114] In some embodiments, the present invention may further include a filtering step, namely, identifying candidate mutation sites based on the calculated base composition data of the 4-base genome before BC or EC conversion, and filtering the candidate mutation sites (filtering out possible false positives) to generate detection results for mutation sites of single nucleotide variants.

[0115] Regarding the filtering step, in the above embodiments, after the preceding BC or EC transformation and calculation, we calculated the base composition data of the 4-base genome before BC or EC transformation (e.g., the true number of bases in the four bases). Then, the present invention uses a mutation detection method (e.g., but not limited to, Varscan2 mutation detection software) to identify candidate mutation sites.

[0116] Subsequently, the present invention uses, for example, the following filtering rules to filter candidate mutation sites to generate detection results for mutation sites of single nucleotide variants. The specific filtering rules include the following:

[0117] 1. Overall coverage of the sites.

[0118] It should be understood that sequencing errors are higher in low-coverage regions, affecting the reliability of test results. Therefore, it is necessary to filter out candidate mutation sites with coverage less than a predetermined coverage threshold.

[0119] 2. The number of reads that support low-frequency mutations (supporting read length, i.e., supporting reads).

[0120] Unlike traditional 4-base genomes, only the bases actually detected and not estimated using the reverse method are counted. For example, when the alternative allele is T, the T bases of the OT strand are not counted; only the A bases of the OB strand are counted.

[0121] 3. Mutation frequency.

[0122] Different mutation frequency thresholds are used for known and unknown mutations. Different thresholds are also assigned to different mutation types; for example, G>A is more likely to produce false positives, so a higher threshold is assigned.

[0123] 4. SNP filtering.

[0124] In the absence of paired white blood cells, this invention uses information such as population frequency to filter out possible germline mutations.

[0125] 5. Establish a white blood cell baseline and filter low-frequency mutations.

[0126] This filtering rule is also intended to filter clonal hematopoiesis.

[0127] 6. Establish read-level filtering.

[0128] Because methylated data is a 3-base genome, its data complexity is lower than that of a 4-base genome, making it more prone to insertion and deletion variations during sequence alignment. Insertion and deletion variations near the target site can easily lead to the target site being incorrectly identified as a mutation. Therefore, reads with insertion and deletion variations near the target site or with low alignment quality will also be filtered out.

[0129] To further verify the present invention, a set of hybrid cell lines was designed. Table 2 below shows a list of six cancer-related mutations included in the hybrid cell lines, which were diluted to a low frequency (e.g., 0.4% and 0.2%).

[0130] Table 2

[0131] Using the aforementioned mutations as target sites, this invention designs probes with a length of 120 bp and a distance of 24 bp between each probe. The probes are divided into two types: fully methylated and fully non-methylated. In fully methylated probes, only C in non-CpG DNA is converted to T, while C in CpG DNA remains C. This type of probe mainly captures DNA sequences with high methylation levels. Fully non-methylated probes convert all C to T, mainly capturing DNA sequences with low methylation levels. Since the sequences of the OT and OB chains are not complementary after BC or EC transformation, separate probes are also required for the OT and OB chains. In summary, this invention designs four probes for each region: a fully methylated OT probe, a fully non-methylated OT probe, a fully methylated OB probe, and a fully non-methylated OB probe.

[0132] For the aforementioned mixed cell lines, this invention simultaneously employed three methods for mutation detection. The three methods are: 35000x conventional mutation sequencing with UMI (Method 1), 15000x conventional mutation detection without UMI (Method 2) (for related methods, see Liang, N., et al., Ultrasensitive detection of circulating tumor DNA via deep methylation sequencing aided by machine learning. Nat Biomed Eng, 2021.5(6):p.586-599.), and Method 3, which utilizes methylation sequencing data from embodiments of this invention to detect mutations, for example, detecting mutations based on approximately 10000x methylation sequencing data of the target region. No false positives were observed with any of the three methods. Table 3 below shows the detected mutation frequencies for a mixed sample target mutation abundance (AF) of 0.4% using the three methods.

[0133] Table 3

[0134] Table 4 below shows the mutation frequencies detected by the three methods for the target AF0.4% in mixed samples.

[0135] Table 4

[0136] Based on the mutation detection results, all target mutations were detected normally by conventional sequencing. In both sets of data, method 3 successfully detected 4 mutations, and the frequency of detected mutations was not significantly different from that of conventional sequencing using methods 1 and 2.

[0137] Methylation sequencing data is typically less complex than normal sequencing data and has higher background noise. However, the mutation detection method based on methylation data in this invention can achieve mutation detection capabilities close to those of conventional sequencing by using only methylation data for data analysis (e.g., but not limited to, the MMcall algorithm) without introducing additional false positives. A single detection reaction can simultaneously obtain both methylation signals and gene mutation results.

[0138] To further verify the value of this invention in clinical testing, the performance of the mutation detection method based on methylation sequencing data (e.g., the MMcall algorithm) of this invention was tested in real clinical samples, and the results were compared with methylation models (for related content on methylation models, see Liang, N., et al., Ultrasensitive detection of circulating tumor DNA via deep methylation sequencing aided by machine learning. Nat Biomed Eng, 2021.5(6):p.586-599.). To cover common tumor mutation sites, this invention designed a panel containing 513 mutation sites. The probe design method of this panel is consistent with the performance verification, and the DNA input is 5 ng.

[0139] This experiment included samples from 322 age-matched healthy individuals and samples from patients with multiple different cancer types. Specific sample information is shown in Table 5. The methylation model technique successfully detected 14 cases of lung cancer without false positives, while missing 7 cases. Using the mutation detection method based on methylation sequencing data (e.g., the MMcall algorithm) of this invention, with a specificity of 98.8%, one additional case of stage III lung cancer was detected among the samples missed by methylation sequencing. Table 5 below shows a comparison of the results of detecting positive samples using the mutation detection method based on methylation sequencing data (e.g., the MMcall algorithm) and the methylation model method of this invention for clinical samples, with an overall pan-cancer detection rate of 32.9%. The detection sensitivity across different cancer types and stages is shown in Figure 9.

[0140] Table 5

[0141] In the above scheme, the present invention can convert the 3-base genome after BC or EC conversion into the 4-base genome before BC or EC conversion, and then detect mutations based on the 4-base genome algorithm and filtering criteria. The method of the present invention only needs to perform data analysis and processing on the methylation data obtained by methylation detection to obtain gene mutation detection at the same time as obtaining methylation detection results. That is, the present invention can detect methylation signals and gene mutations simultaneously in one reaction, and the two detection modes do not interfere with each other. Furthermore, it can achieve accurate detection of low abundance (AF as low as 0.4%) SNV mutation sites at methylation sequencing depth (e.g., 10000X).

[0142] Figure 5 illustrates a flowchart of a method 500 according to an embodiment of the present invention for calculating the number of G bases in the negative strand that are complementary to the T bases in the positive strand, which are converted from C bases via BC or EC. It should be understood that method 500 can be performed, for example, at the electronic device 800 described in Figure 8. It can also be performed at the computing device 110 described in Figure 1. It should be understood that method 500 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.

[0143] At step 502, the computing device 110 divides the number of all C bases in the positive chain by the sum of the number of all T bases and C bases in the positive chain to obtain a first ratio.

[0144] For example, computing device 110 will count the number of all C bases in the OT chain. Divide by the sum of the number of all T and C bases in the OT chain. In order to obtain the first proportion

[0145] At step 504, the computing device 110 calculates the number of G bases in the negative chain that are complementary to the T bases in the positive chain, which are converted from C bases to T bases via chemical or enzymatic conversion, based on the number of all G bases in the negative chain, the first ratio, and the sum of the number of G bases and A bases in the negative chain.

[0146] For example, computing device 110 is based on the number of all G bases in the OB chain. First proportion And the sum of the number of G bases and A bases in the OB chain. Calculate the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are transformed from C bases via BC or EC.

[0147] Specifically, in some embodiments, the number of all A bases in the OB chain is... Number of bases with G base Add them together to obtain the sum of the first dibase numbers of all A and G bases in the OB chain. The number of bases of all C bases in the OT chain Divide by the sum of the total number of T and C bases in the OT chain. In order to obtain the first proportion of C bases in the OT chain And the number of all G bases in the OB chain. In the middle, subtract the first proportion of C bases in the OT chain. The sum of the first dibases of all A and G bases in the OB chain The product of these numbers is used to obtain the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are transformed from C bases via BC or EC.

[0148] The following uses formulas (12) to (14) to explain the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are transformed from C bases via BC or EC. algorithm.

[0149] In the above formulas (12) to (14), It represents the sum of the first dibases of all A and G bases in the OB chain. This represents the highest percentage of C bases in the OT chain. This represents the number of G bases in the OB chain. This represents the number of G bases in the OB chain that are complementary to the T bases in the OT chain, which are converted from C bases via BC or EC.

[0150] Figure 6 illustrates a flowchart of a method 600 for calculating the number of bases in the positive chain that are converted from a C base to a T base via BC or EC, according to an embodiment of the present invention. It should be understood that method 600 can be performed, for example, at the electronic device 800 described in Figure 8. It can also be performed at the computing device 110 described in Figure 1. It should be understood that method 600 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.

[0151] At step 602, the computing device 110 adds the number of A bases and the number of G bases in the negative chain to obtain the sum of the first dibases of all A bases and G bases in the negative chain.

[0152] For example, computing device 110 will count the number of all A bases in the OB chain. Number of bases with G base Add them together to obtain the sum of the first dibase numbers of all A and G bases in the OB chain.

[0153] At step 604, the computing device 110 divides the number of all A bases in the negative chain by the sum of the first dibases to obtain a second proportion of A bases in the negative chain.

[0154] For example, computing device 110 will count the number of all A bases in the OB chain. Divide by the sum of the first dibases In order to obtain the second proportion of A bases in the OB chain

[0155] At step 606, the computing device 110 adds the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion to the number of bases in the positive chain that are T bases before and after chemical or enzymatic conversion, in order to obtain the sum of T bases in the positive chain.

[0156] For example, computing device 110 calculates the number of bases in the OT chain that are converted from C bases to T bases via BC or EC. The number of bases in the OT chain that are T bases before and after BC or EC conversion. Add them together to obtain the sum of the T bases in the OT chain.

[0157] At step 608, the computing device 110 makes the product of the sum of the T bases in the positive chain and the second proportion of the A bases in the negative chain equal to the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion, in order to calculate the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion.

[0158] For example, computing device 110 causes the sum of the bases of the T bases in the OT chain to be... The second proportion of A bases in the OB chain The product equals the number of bases in the OT chain that are converted from C bases to T bases via BC or EC. Used to calculate the number of bases in the OT chain that are converted from C bases to T bases via BC or EC.

[0159] The following uses formula (15) to explain the algorithm for determining the number of bases in the OT chain that are converted from C bases to T bases via BC or EC.

[0160] In the above formula (15), It represents the sum of the first dibases of all A and G bases in the OB chain. This represents the second highest percentage of A bases in the OB chain.

[0161] Figure 7 shows a flowchart of a method 700 for determining the number of T bases in the positive chain before BC or EC conversion according to an embodiment of the present invention. It should be understood that method 700 can be performed, for example, at the electronic device 800 described in Figure 8. It can also be performed at the computing device 110 described in Figure 1. It should be understood that method 700 may also include additional actions not shown and / or the actions shown may be omitted, and the scope of the invention is not limited in this respect.

[0162] At step 702, the computing device 110 adds the number of A bases and the number of G bases in the negative chain to obtain the sum of the first dibases of all A bases and G bases in the negative chain.

[0163] For example, computing device 110 will count the number of all A bases in the OB chain. Number of bases with G base Add them together to obtain the sum of the first dibase numbers of all A and G bases in the OB chain.

[0164] At step 704, the computing device 110 divides the number of G bases in the negative chain that are complementary to the T bases in the positive chain by chemical or enzymatic conversion from C bases to T bases by the sum of the first dibase numbers in order to obtain the third proportion of G bases in the negative chain.

[0165] For example, computing device 110 calculates the number of bases in the OB chain that are complementary to the G bases in the OT chain, which are converted from C bases to T bases via BC or EC. Divide by the sum of the first dibases In order to obtain the third proportion of G bases in the OB chain

[0166] In step 706, the computing device 110 adds the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion to the number of bases in the positive chain that are T bases before and after chemical or enzymatic conversion, in order to obtain the sum of T bases in the positive chain.

[0167] For example, computing device 110 calculates the number of bases in the OT chain that are converted from C bases to T bases via BC or EC. The number of bases in the OT chain that are T bases before and after BC or EC conversion. Add them together to obtain the sum of the T bases in the OT chain.

[0168] At step 708, the computing device 110 makes the product of the sum of T bases in the positive chain and the third proportion of G bases in the negative chain equal to the number of T bases in the positive chain before and after BC or EC conversion, so as to calculate the number of T bases in the positive chain before and after BC or EC conversion.

[0169] For example, computing device 110 causes the sum of the bases of the T bases in the OT chain to be... The third proportion of G bases in the OB chain The product equals the number of bases in the OT chain that are T bases before and after the BC or EC conversion. This is used to calculate the number of bases in the OT chain that are T bases before and after BC or EC conversion. The following uses formula (16) to explain the algorithm for determining the number of T bases in the OT chain before BC or EC conversion.

[0170] In step 710, the computing device 110 determines the number of T bases in the positive chain before chemical or enzymatic transformation based on the calculated number of T bases before and after chemical or enzymatic transformation.

[0171] For example, computing device 110 is based on the calculated number of bases in the OT chain that are T bases before and after BC or EC conversion. Determine the number of T bases in the OT chain before BC or EC conversion. As shown in formula (8) above.

[0172] Figure 8 schematically illustrates a block diagram of an electronic device 800 suitable for implementing embodiments of the present invention. The electronic device 800 may be used to implement methods 200, 500 to 700 shown in Figures 2, 5 to 7. As shown in Figure 8, the electronic device 800 includes a central processing unit (i.e., CPU 801), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (i.e., ROM 802) or loaded from a storage unit 808 into a random access memory (i.e., RAM 803). Various programs and data required for the operation of the electronic device 800 may also be stored in the RAM 803. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output interface (i.e., I / O interface 805) is also connected to the bus 804.

[0173] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, and storage unit 808. CPU 801 executes the various methods and processes described above, such as methods 200, 500 to 700. For example, in some embodiments, methods 200, 500 to 700 may be implemented as computer software programs stored in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more operations of methods 200, 500 to 700 described above may be performed. Alternatively, in other embodiments, CPU 801 may be configured to execute one or more actions of methods 200, 500 to 700 by any other suitable means (e.g., by means of firmware).

[0174] It should be further noted that the present invention can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the present invention.

[0175] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0176] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0177] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0178] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0179] These computer-readable program instructions can be provided to a processor in a voice interaction device, a general-purpose computer, a special-purpose computer, or a processing unit of another programmable data processing device, thereby producing a machine such that, when executed by the processing unit of the computer or other programmable data processing device, these instructions create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0180] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0182] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0183] The above are merely optional embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for detecting mutations based on methylation data, characterized in that, include: Obtain methylation data of the sample to be tested; Based on methylation data, obtain the base composition data of the 3-base genome after chemical or enzymatic transformation; Based on the base composition data of the 3-base genome, the base composition data of the 4-base genome before chemical or enzymatic transformation is determined according to the complementary base pairing relationship between the positive and negative strands and / or the correspondence between genotypes before and after chemical or enzymatic transformation. Based on the base composition data of the 4-base genome, candidate mutation sites are identified; The method further includes: The candidate mutation sites are filtered to generate detection results for mutation sites of single nucleotide variants.

2. The method according to claim 1, characterized in that, Data determining the 4-base genome composition before chemical or enzymatic transformation includes: Based on the base composition data of the aforementioned 3-base genome, the distribution data of the number of A, T, G, and C bases in the positive and negative strands after chemical or enzymatic transformation were determined, respectively; and Based on the complementary base pairing relationship between the positive and negative strands, the correspondence between genotypes before and after chemical or enzymatic transformation, and the distribution data of the number of A, T, G, and C bases in the positive and negative strands after chemical or enzymatic transformation, the genotypes and gene frequencies of the four-base genome are determined.

3. The method according to claim 2, characterized in that, Determining the genotype and gene frequency of the 4-base genome includes: The ratio of the number of G bases in the negative strand that are complementary to all C bases in the positive strand to the number of G bases in the negative strand that are complementary to all C bases in the positive strand through chemical or enzymatic conversion to T bases is equal to the ratio of the number of all C bases in the positive strand to the number of all C bases in the positive strand through chemical or enzymatic conversion to T bases; and The ratio of the number of G bases in the positive strand that are complementary to all C bases in the negative strand to the number of G bases in the positive strand that are complementary to all C bases in the negative strand through chemical or enzymatic conversion to T bases is equal to the ratio of the number of all C bases in the negative strand to the number of all C bases in the negative strand through chemical or enzymatic conversion to T bases, in order to determine the genotype and gene frequency of the 4-base genome.

4. The method according to claim 2, characterized in that, Determining the genotype and gene frequency of the 4-base genome includes: This ensures that the total number of T bases in the positive strand is equal to the sum of the number of T bases in the positive strand before and after chemical or enzymatic transformation, and the number of bases in the positive strand that are converted from C bases to T bases through chemical or enzymatic transformation; and The number of all T bases in the negative strand is equal to the sum of the number of all T bases in the negative strand before and after chemical or enzymatic transformation and the number of C bases in the negative strand that are chemically or enzymatically transformed into T bases, in order to determine the genotype and gene frequency of the 4-base genome.

5. The method according to claim 2, characterized in that, Determining the genotype and gene frequency of the 4-base genome includes: This makes the total number of G bases in the positive strand equal to the sum of the number of G bases in the positive strand that are complementary to all C bases in the negative strand through chemical or enzymatic conversion to T bases, and the number of G bases in the positive strand that are complementary to all C bases in the negative strand; and This makes the total number of G bases in the negative chain equal to the sum of the number of G bases in the negative chain that are complementary to all C bases in the positive chain that have been chemically or enzymatically converted to T bases, and the number of G bases in the negative chain that are complementary to all C bases in the positive chain.

6. The method according to claim 2, characterized in that, Determining the genotype and gene frequency of the 4-base genome includes: Divide the number of all C bases in the positive strand by the sum of the number of all T bases and C bases in the positive strand to obtain the first ratio; and Based on the number of all G bases in the negative chain, the first ratio, and the sum of the number of all G bases and A bases in the negative chain, calculate the number of G bases in the negative chain that are complementary to the T bases in the positive chain, which are converted from C bases to T bases via chemical or enzymatic transformation.

7. The method according to claim 2 or 3, characterized in that, Determining the genotype and gene frequency of the 4-base genome also includes: The number of T bases in the positive strand before chemical or enzymatic transformation is calculated based on the number of T bases in the positive strand that have not undergone chemical or enzymatic transformation, the number of T bases in the positive strand that have undergone chemical or enzymatic transformation from C bases to T bases, the number of G bases in the negative strand that are complementary to the T bases in the positive strand that have undergone chemical or enzymatic transformation from C bases to T bases, and the total number of A bases in the negative strand.

8. The method according to claim 6, characterized in that, The number of G bases in the negative chain that are complementary to the T bases in the positive chain through chemical or enzymatic conversion from C bases is obtained through the following methods: Add the number of A bases and the number of G bases in the negative chain to obtain the sum of the first dibases of all A bases and G bases in the negative chain; Divide the total number of C bases in the positive chain by the sum of the total number of T bases and C bases in the positive chain to obtain the first proportion of C bases in the positive chain. as well as Subtract the product of the first proportion of C bases in the positive chain and the sum of the first dibases of all A and G bases in the negative chain from the total number of G bases in the negative chain. This will give the number of G bases in the negative chain that are complementary to the T bases in the positive chain, which are converted from C bases to T bases via chemical or enzymatic transformation.

9. The method according to claim 4, characterized in that, Determining the genotype and gene frequency of the 4-base genome includes: Add the number of C bases in the positive strand to the number of T bases that were converted from C bases to T bases via chemical or enzymatic conversion, in order to obtain the number of C bases in the positive strand before the chemical or enzymatic conversion.

10. The method according to claim 9, characterized in that, The number of bases in the positive chain that are converted from C bases to T bases via chemical or enzymatic conversion is obtained through the following methods: Add the number of A bases and the number of G bases in the negative chain to obtain the sum of the first dibases of all A bases and G bases in the negative chain; Divide the total number of A bases in the negative chain by the sum of the first dibases to obtain the second percentage of A bases in the negative chain; The number of bases in the positive chain that are converted from C bases to T bases through chemical or enzymatic transformation is added to the number of bases in the positive chain that are T bases before and after chemical or enzymatic transformation, so as to obtain the sum of T bases in the positive chain. as well as The product of the sum of T bases in the positive chain and the second proportion of A bases in the negative chain is equal to the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion, in order to calculate the number of bases in the positive chain that have been converted from C bases to T bases via chemical or enzymatic conversion.

11. The method according to claim 7, characterized in that, Calculating the number of T bases in the OT chain before chemical or enzymatic transformation includes: Add the number of A bases and the number of G bases in the negative chain to obtain the sum of the first dibases of all A bases and G bases in the negative chain; Divide the number of G bases in the negative chain that are complementary to the T bases in the positive chain by chemical or enzymatic conversion from C bases to T bases by the sum of the first dibase numbers to obtain the third proportion of G bases in the negative chain. The number of bases in the positive chain that are converted from C bases to T bases through chemical or enzymatic transformation is added to the number of bases in the positive chain that are T bases before and after chemical or enzymatic transformation, so as to obtain the sum of T bases in the positive chain. The product of the sum of T bases in the positive chain and the third proportion of G bases in the negative chain is equal to the number of T bases in the positive chain before and after chemical or enzymatic transformation, for use in calculating the number of T bases in the positive chain before and after chemical or enzymatic transformation; and Based on the calculated number of T bases in the positive chain before and after chemical or enzymatic transformation, determine the number of T bases in the positive chain before chemical or enzymatic transformation.

12. The method according to claim 1, characterized in that, Filtering candidate mutation sites to generate detection results for single nucleotide variants includes: Candidate mutation sites that meet a first predetermined condition are filtered out. The first predetermined condition includes at least one of the following: The coverage is less than the predetermined coverage threshold; The number of mutation supports is less than a predetermined support threshold; and The mutation frequency is less than a predetermined frequency threshold; and For the remaining candidate mutation sites, filtering is performed based on the presence of insertion and deletion variations near the candidate mutation sites and the alignment quality, in order to generate detection results for mutation sites with single nucleotide variations.

13. The method according to claim 12, characterized in that, Candidate mutation sites that meet the first predetermined criteria were filtered out, including: A predetermined frequency threshold is determined based on whether the relevant mutations at candidate mutation sites belong to known mutations and the mutation type.

14. The method according to claim 13, characterized in that, Determining the predetermined frequency threshold includes: Determine whether the relevant mutation at the candidate mutation site belongs to a known mutation; If the relevant mutation is determined to be a known mutation, a predetermined frequency threshold is set as a first value; and If it is determined that the relevant mutation does not belong to the known mutations, a predetermined frequency threshold is determined as a second value, which is different from the first value.

15. The method according to claim 13, characterized in that, Determining the predetermined frequency thresholds includes setting different predetermined frequency thresholds for different mutation types.

16. A computing device, characterized in that, include: At least one processing unit; At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to perform the steps of the method according to any one of claims 1 to 15.

17. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a machine, implements the method according to any one of claims 1 to 15.

18. A computer program product, characterized in that, The computer program product includes instructions that, when executed by a machine, implement the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Next generation sequencing-based point mutation detection filtering method and apparatus, and storage medium

    CN107944223A

  • Method for detecting mutation, electronic equipment and computer storage medium

    CN111292802A

  • Method and device for simultaneously detecting DNA methylation and genome variation and storage medium

    CN112634984A

  • Method and device for variation detection based on methylation sequencing data

    CN113674802A

  • Method, apparatus, medium and program for detecting mutations based on methylation data

    CN118969083A