Variant calling with methylation-level estimation

HK40138092APending Publication Date: 2026-09-25ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
HK62026125753
Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2026-07-06
Publication Date
2026-09-25
Estimated Expiration
2044-06-25

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure describes methods, non-transitory computer-readable media, and systems that can simultaneously determine cytosine bases of a target genomic sample and an estimated methylation level value of genotype detection. The disclosed system may utilize a Bayesian method on nucleotide read data of a target genomic sample to generate an estimated methylation level value indicative of genomic coordinates for the target genomic sample to contain a reference cytosine base, or a nucleobase, which may be referred to as cytosine. The disclosed system may estimate a methylation level value based on a prior genotype probability and observed nucleobases at genomic coordinates from a read stack of the target genomic sample. Based on the estimated methylation level value and a base detection quality metric, the disclosed system may generate a posterior genotype probability of the genomic sample at the genomic coordinates. Based on the posterior genotype probability, the disclosed system can generate a genotype detection of the target genomic sample.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480041051.0 (22) Application Date 2024.06.26 (30) Priority Data 63 / 510603 2023.06.27 US (85) PCT International Application Entering National Phase Date 2025.12.18 (86) PCT International Application Application Data PCT / US2024 / 035562 2024.06.26 (87) PCT International Application Publication Data WO2025 / 006565 EN 2025.01.02 (71) Applicant Inmena Ltd. Address California, USA (72) Inventors J.B. D. Andrews K.H. Schaeffler (74) Patent Agency Beijing Panhua Weiye Intellectual Property Agency Co., Ltd. 11280 Patent Attorney Wang Bo (51) Int.Cl. G16B 20 / 20 (2006.01) G16B 40 / 30 (2006.01) (54) Invention Title Variant Detection Using Methylation Level Estimation (57) Abstract This disclosure describes a method, non-transitory computer-readable medium, and system for simultaneously determining estimated methylation level values ​​for cytosine bases and genotype detection of a target genome sample. The disclosed system utilizes a Bayesian method to generate estimated methylation level values ​​from nucleotide read data of a target genome sample, which indicate the genomic coordinates where the target genome sample contains a reference cytosine base, or a nucleobase that may be referred to as cytosine. The disclosed system can estimate the methylation level value based on a prior genotype probability and observed nucleobases at genomic coordinates of read stacks from the target genome sample. Based on the estimated methylation level value and a base detection quality metric, the disclosed system can generate a posterior genotype probability of the genome sample at those genomic coordinates. Based on this posterior genotype probability, the system disclosed by the institute can generate genotype detection for the target genome sample.Claims 4 pages, Description 36 pages, Drawings 21 pages, CN 121359206 A 2026.01.16 CN 1 21 35 92 06 A 1. A method comprising: identifying, for a target genome sample, a nucleotide read containing one or more nucleobases transformed by methylation sequencing; determining an estimated methylation level of cytosine bases at the genomic coordinates based on a prior genotype probability of the target genome sample at genomic coordinates and observed nucleobases at the genomic coordinates within the nucleotide read; generating a posterior genotype probability of the target genome sample at the genomic coordinates based on the estimated methylation level of the observed nucleobases and a base detection quality metric using a variant detection model; and generating a genotype detection of a predicted combination of nucleobases at the genomic coordinates contained in the target genome sample based on the posterior genotype probability. 2. The method of claim 1, further comprising generating a refined methylation level value of the cytosine base at the genomic coordinates based on the posterior genotype probability and the observed nucleotides. 3. The method of claim 2, further comprising generating the refined methylation level value by: determining a genotype-specific methylation level value corresponding to each possible genotype at the genomic coordinates based on the observed nucleotides; and weighting the genotype-specific methylation level value based on the posterior genotype probability. 4. The method of claim 1, wherein generating the genotype detection of the target genome sample is further based on sequencing metrics corresponding to the nucleotide reads. 5. The method of claim 1, wherein the posterior genotype probability comprises a subset of the posterior genotype probabilities of cytosine bases in the positive or negative strand. 6. The method of claim 1, further comprising determining the estimated methylation level value by: determining the probability of each nucleotide at the genomic coordinate on the positive and negative strands; determining the probability of the observed nucleotide at the genomic coordinate based on the number of observed nucleotides, the probability of each nucleotide at the genomic coordinate, and the prior genotype probability; and generating an estimated positive-strand methylation level value for the nucleotide at the genomic coordinate on the positive strand and an estimated negative-strand methylation level value for the nucleotide at the genomic coordinate on the negative strand by performing a Bayesian inversion on the probabilities of the observed nucleotides. 7. The method of claim 6, further comprising determining the probability of each nucleotide by: determining that the probability of a given thymine base on the positive strand is approximately the estimated methylation level value; and determining that the probability of a given cytosine base on the positive strand is approximately 1 minus the estimated methylation level value.8. The method of claim 6, further comprising determining the probability of each nucleobase by: determining the probability of a given adenine base on the negative strand to be approximately equal to the estimated methylation level value; and determining the probability of a given guanine base on the negative strand to be approximately equal to 1 minus the estimated methylation level value. 9. The method of claim 1, wherein the variant detection model comprises a Hidden Markov Model (HMM) modified to receive input of a base detection quality metric based on the estimated methylation level value and the corresponding observed nucleobase. 10. The method of claim 1, further comprising generating the genotype detection based on a predicted combination of nucleobases corresponding to the highest posterior genotype probability. 11. The method of claim 1, wherein identifying the nucleotide read comprising the one or more nucleobases converted by the methylation sequencing comprises identifying the nucleotide read comprising a thymine base or a uracil base converted from a cytosine base by the methylation sequencing. 12. A system comprising: at least one processor; and a nontransitory computer-readable medium including instructions that, when executed by the at least one processor, cause the system to: identify, for a target genome sample, a nucleotide read comprising one or more nucleosides converted by methylation sequencing; determine an estimated methylation level of cytosine bases at the genomic coordinates based on a prior genotype probability of the target genome sample at genomic coordinates and observed nucleosides at the genomic coordinates within the nucleotide reads; generate a posterior genotype probability of the target genome sample at the genomic coordinates based on the estimated methylation level of the observed nucleosides and a base detection quality metric using a variant detection model; and generate a genotype detection of a predicted combination of nucleosides at the genomic coordinates contained in the target genome sample based on the posterior genotype probability. 13. The system of claim 12, further comprising instructions, which, when executed by the at least one processor, cause the system to generate a refined methylation level value of the cytosine base at the genomic coordinates based on the posterior genotype probability and the observed nucleotides. 14. The system of claim 13, further comprising instructions, which, when executed by the at least one processor, cause the system to generate the refined methylation level value by: determining a genotype-specific methylation level value corresponding to each possible genotype at the genomic coordinates based on the observed nucleotides; and weighting the genotype-specific methylation level value based on the posterior genotype probability.15. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to generate the genotype detection of the target genome sample based on sequencing metrics corresponding to the nucleotide reads. 16. The system of claim 12, wherein the posterior genotype probability comprises a subset of the posterior genotype probabilities of cytosine bases in the positive or negative strand. 17. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to determine the estimated methylation level value by: determining the probability of each nucleobase at the genomic coordinates on the positive and negative strands; determining the probability of the observed nucleobase at the genomic coordinates based on the number of each observed nucleobase, the probability of each nucleobase at the genomic coordinates, and the prior genotype probability; and generating an estimated positive-strand methylation level value for the nucleobase at the genomic coordinates on the positive strand and an estimated negative-strand methylation level value for the nucleobase at the genomic coordinates on the negative strand by performing a Bayesian inversion on the probability of the observed nucleobase. 18. The system of claim 17, further comprising instructions, when executed by the at least one processor, causing the system to determine the probability of each nucleobase by: determining the probability of a given thymine base on the positive strand to be approximately the estimated methylation level value; and determining the probability of a given cytosine base on the positive strand to be approximately 1 minus the estimated methylation level value. 19. The system of claim 17, further comprising instructions, when executed by the at least one processor, causing the system to determine the probability of each nucleobase by: determining the probability of a given adenine base on the negative strand to be approximately the estimated methylation level value; and determining the probability of a given guanine base on the negative strand to be approximately 1 minus the estimated methylation level value. 20. The system of claim 12, wherein the variant detection model comprises a Hidden Markov Model (HMM), the HMM being modified to receive input of a base detection quality metric based on the estimated methylation level value and the corresponding observed nucleobase. 21. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to generate the genotype detection based on a predicted combination of nucleobases corresponding to the highest posterior genotype probability.22. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to identify a nucleotide read containing one or more nucleobases converted by the methylation sequencing by identifying the nucleotide read containing a thymine base or uracil base converted from a cytosine base by the methylation sequencing. 23. A non-transitory computer-readable medium storing instructions, which, when executed by at least one processor, cause a computing device to: identify, for a target genome sample, a nucleotide read comprising one or more nucleosides converted by methylation sequencing; determine an estimated methylation level of cytosine bases at the genomic coordinates based on a prior genotype probability of the target genome sample at genomic coordinates and observed nucleosides at the genomic coordinates within the nucleotide reads; generate a posterior genotype probability of the target genome sample at the genomic coordinates based on the estimated methylation level of the observed nucleosides and a base detection quality metric using a variant detection model; and generate a genotype detection of a predicted combination of nucleosides at the genomic coordinates contained in the target genome sample based on the posterior genotype probability. 24. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the estimated methylation level value by: determining the probability of each nucleobase at the genomic coordinates on the positive and negative strands; determining the probability of the observed nucleobase at the genomic coordinates based on the number of each observed nucleobase, the probability of each nucleobase at the genomic coordinates, and the prior genotype probability; and generating an estimated positive-strand methylation level value for the nucleobase at the genomic coordinates on the positive strand and an estimated negative-strand methylation level value for the nucleobase at the genomic coordinates on the negative strand by performing a Bayesian inversion on the probabilities of the observed nucleobases. Claims 3 / 4, page 4, CN 121359206 A 25. The non-transitory computer-readable medium of claim 23, wherein the variant detection model comprises a Hidden Markov Model (HMM), the HMM being modified to receive input of a base detection quality metric based on the estimated methylation level value and the corresponding observed nucleobases. 26. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the genotype detection based on a predicted combination of nucleobases corresponding to the highest posterior genotype probability.27. The non-transitory computer-readable medium of claim 23, wherein identifying the nucleotide read comprising one or more nucleobases converted by said methylation sequencing comprises identifying the nucleotide read comprising a thymine base or a uracil base converted from a cytosine base by said methylation sequencing. Claims 4 / 4 Page 5 CN 121359206 A Variant Detection Using Methylation Level Estimation Cross-Reference to Related Applications

[0001] This application claims priority and benefit to U.S. Provisional Application No. 63 / 510,603, filed June 27, 2023, entitled “VARIANT CALLING WITH METHYLATION-LEVEL ESTIMATION,” the entire contents of which are incorporated herein by reference. Background Art

[0002] In recent years, biotechnology companies and research institutions have improved the hardware and software used for nucleotide sequencing of genomic samples and for determining nucleobase detection. For example, some existing sequencers and sequencing data analysis software (collectively referred to as "existing sequencing systems") predict individual nucleotide bases within a sequence using conventional Sanger sequencing or sequencing-by-synthesis (SBS) methods. When using SBS, existing sequencing systems can monitor thousands to millions of oligonucleotides synthesized in parallel from a template to predict the nucleotide base detection of an ever-growing number of nucleotide reads. In many existing sequencing systems, a camera captures images of irradiated fluorescent tags incorporated into the oligonucleotides. After capturing such images, some existing sequencing systems determine the nucleotide base detection of the nucleotide reads corresponding to the oligonucleotides and send the base detection data to a computing device with sequencing data analysis software that aligns the nucleotide reads to a reference genome. Based on the differences between the aligned nucleotide reads and the reference genome, existing systems can further utilize variant detectors to identify variants in the genome sample, such as single nucleotide variants (SNVs), insertions or deletions (indels), or other variants within the genome sample.

[0003] In addition to improved genome sequencing, biotechnology companies and research institutions have also improved methods for detecting methylation of cytosine bases at specific genomic regions (e.g., regions encoding or promoting genes) and for detecting methylation of larger nucleotide fragments or the entire genome in a sample. For example, some existing sequencing systems can use sequencing equipment and corresponding sequencing data analysis software to identify when a methyl or hydroxymethyl group has been added to the cytosine bases of the sample's deoxyribonucleic acid (DNA)—where methylated cytosine bases are often part of a cytosine-guanine dinucleotide pair in the 5'-C phosphate-G-3' (CpG) conformation in mammals.For example, existing sequencing systems can detect methylated cytosine by: (i) enzymatically converting methylated or unmethylated cytosine bases at CpG or other cytosine sites in a sample nucleotide fragment to uracil bases (e.g., dihydrouracil); (ii) using a sequencing device to determine the base detection of nucleotide reads in the sample, wherein the sequencing device detects the uracil bases as thymine bases during polymerase chain reaction (PCR) amplification; and (iii) comparing the base detections from these nucleotide reads with a reference genome or non-enzymatically converted nucleotide reads from the sample. Based on the comparison of nucleotide reads from the sample with the reference genome or the non-enzymatically converted nucleotide reads, existing sequencing systems can identify thymine bases from these nucleotide reads that do not match the cytosine bases at CpG or other cytosine sites in the reference genome or these non-enzymatically converted nucleotide reads, thereby detecting methylated cytosine bases in the sample nucleotide fragment.

[0004] Despite these recent advances, existing sequencing and methylation detection systems face several technical drawbacks. Different types of methylation assays detect methylated cytosine by converting methylated or unmethylated cytosine bases to uracil bases, and subsequently, in some cases, to thymine bases. Since oligonucleotides extracted from the genome sample are replicated as part of the methylation sequencing assay, the complementary strand reflects the cytosine-to-thymine substitution region by replacing guanine with adenine. While these conversions facilitate methylation detection, they can also negatively impact the performance and accuracy of existing sequencing systems.

[0005] For example, when performing C>T conversion-based sequencing, existing sequencing and methylation detection systems fail to consistently generate accurate genotype detections when simultaneously detecting genotype and methylation levels. Partly due to C>T conversions in methylation assays, existing sequencing systems tend to produce biased genotype detections and biased methylation level estimates. Transformed methylated or unmethylated cytosine bases often introduce noise into sequence data, which in turn hinders accurate variant detection. Due to such transformations and noise in methylation assays, existing methylation detection systems often overestimate the methylation levels of C / A, C / T, G / A, and G / T genotypes. Furthermore, existing sequencing systems frequently determine inaccurate base detections of genomic regions containing transformed methylated or unmethylated cytosine bases. For example, due to noise in the sequence data, existing sequencing systems often incorrectly classify methylation events as C / T or G / A genotypes.

[0006] As noted above, in addition to biased genotype detection, existing sequencing systems generate biased methylation level estimates when simultaneously determining genotypes.Existing sequencing systems often rely on overly simplistic methods or algorithms to estimate methylation levels. For example, some existing systems attempt to improve the accuracy of simultaneous methylation level estimation and genotype detection by correcting data, ignoring data, or creating statistical models that account for difficult-to-detect genomic regions. Consequently, existing systems often suffer from coding errors and yield limited benefits in terms of accuracy.

[0007] In addition to accuracy challenges, some existing sequencing and methylation detection systems inefficiently consume excessive processing materials, time, and computational resources. Existing sequencing and methylation detection systems often require multiple assays and samples to accurately determine methylation levels and variant detection. For example, in addition to using a separate variant detector to accurately determine both methylation level values ​​and genotype detection, existing systems often require separate methylation assays. Therefore, some existing systems require multiple samples from a single organism, performing both sequencing and methylation assays on these multiple samples in separate computational analyses. Replication of genomic samples often requires replication of computer processing, computer storage, software programs, and other resources to sequence and determine the methylation levels of the same genomic sequence. Therefore, existing systems often consume excessive genome samples, significant time and computer processing resources to sequence and determine the methylation level of a single genome sequence.

[0008] These problems and challenges, along with additional problems and challenges, exist in existing sequencing and methylation detection systems. Summary of the Invention

[0009] This disclosure describes one or more embodiments of systems, methods, and nontransitory computer-readable storage media that address one or more of the above-described problems or provide other advantages over the prior art. For example, the disclosed system uses a Bayesian method to accurately and simultaneously determine the estimated methylation level of cytosine bases and genotype detection of a target genome sample using nucleotide read data of that sample. Such estimated methylation level values ​​may include genomic coordinates at which the target genome sample contains a reference cytosine base or a nucleobase that may be referred to as cytosine (e.g., a cytosine base on the negative strand). Specifically, the disclosed system may estimate methylation level values ​​based on prior genotype probabilities and observed nucleobases at genomic coordinates where reads from the target genome sample are stacked. Based on the estimated methylation level and base detection quality metric, the disclosed system generates a posterior genotype probability of the genome sample at the specified genomic coordinates. Based on this posterior genotype probability, the disclosed system generates a genotype detection for the target genome sample. In some embodiments, the disclosed system further refines the estimated methylation level based on the posterior genotype probability and observed nucleobases from read stacking to generate a more accurate refined methylation level for the target genome sample.Specification 2 / 36 pages 7 CN 121359206 A

[0010] Additional features and advantages of one or more embodiments of this disclosure will be set forth in the following description and will be apparent in part from the description, or may be learned by practice of such exemplary embodiments. Brief Description of the Drawings

[0011] This patent or patent application document contains at least one color drawing. A copy of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the necessary fees.

[0012] Detailed description refers to the drawings, which are briefly described below.

[0013] Figure 1 illustrates an environment in which a methylation genotype detection system can operate according to one or more embodiments of this disclosure.

[0014] Figures 2A and 2B illustrate an overview of a methylation genotype detection system according to one or more embodiments of this disclosure, which simultaneously determines the genotype detection and methylation level value of observed bases from read stacking.

[0015] Figure 3 illustrates various types of methylation sequencing schemes that can be utilized by a methylation genotype detection system according to one or more embodiments of the present disclosure.

[0016] Figure 4 illustrates a methylation genotype detection system according to one or more embodiments of the present disclosure that accurately estimates genotype and methylation level values ​​based on observed nucleobases.

[0017] Figure 5 illustrates a model used by existing sequencing and methylation detection systems that inaccurately estimate methylation level values.

[0018] Figure 6 illustrates a graph showing methylation level values ​​estimated by existing sequencing systems compared to the true methylation level values ​​of a set of genotypes.

[0019] Figures 7A and 7B illustrate a series of actions by which the methylation genotype detection system determines estimated methylation level values ​​of cytosine bases according to one or more embodiments of the present disclosure.

[0020] Figure 8 illustrates an overview of a methylation genotype detection system that generates posterior genotype probabilities according to one or more embodiments of the present disclosure.

[0021] Figure 9 illustrates a methylation genotype detection system that, according to one or more embodiments of the present disclosure, determines the probability of observed bases and the probability of posterior genotypes in part based on estimated methylation level values.

[0022] Figure 10 illustrates a methylation genotype detection system that generates refined methylation level values ​​according to one or more embodiments of the present disclosure.

[0023] Figures 11A and 11B illustrate diagrams demonstrating the improvements made by the methylation genotype detection system in accurately predicting methylation level values ​​in both germline and somatic alleles according to one or more embodiments of the present disclosure.

[0024] Figures 12A and 12B illustrate a series of graphs according to one or more embodiments of the present disclosure, demonstrating that the methylation genotype detection system detects single nucleotide polymorphisms (SNPs) more accurately.

[0025] Figures 13A and 13C illustrate a series of graphs according to one or more embodiments of the present disclosure, demonstrating that the methylation genotype detection system accurately detects genotypes at genomic coordinates.

[0026] Figure 14 illustrates a flowchart of a series of actions for determining estimated methylation levels and generating genotype detections according to one or more embodiments of the present disclosure.

[0027] Figure 15 illustrates a block diagram of an example computing device according to one or more embodiments of the present disclosure. Specification 3 / 36 pages 8 CN 121359206 A Detailed Description

[0028] The present disclosure describes one or more embodiments of a methylation genotype detection system that can accurately determine the genotype and methylation level of a genomic sample from nucleotide-read data of the sample. For example, a methylation genotype detection system can access nucleotide reads containing nucleotides that have been converted by methylation sequencing for a target genome sample. This methylation genotype detection system can further estimate the methylation level of candidate cytosine bases at genomic coordinates based on a prior genotype probability and the nucleotides observed at genomic coordinates within the nucleotide reads. The methylation genotype detection system further utilizes a variant detection model to generate a posterior genotype probability of the target genome sample at genomic coordinates based on the estimated methylation level of the observed nucleotides and a base detection quality metric. Based on this posterior genotype probability, the methylation genotype detection system can predict the genotype detection of the target genome sample at genomic coordinates. In some embodiments, the methylation genotype detection system further refines and increases the accuracy of the estimated methylation level values ​​of the target genome sample based on the posterior genotype probability.

[0029] As just noted, the methylation genotype detection system can identify nucleotide reads containing one or more nucleotides converted by methylation sequencing for a target genome sample. As part of the identification of transformed nucleobases, this methylation genotyping system identifies genomic coordinates from identified nucleotide reads containing candidate methylable cytosine bases within the target genome sample. In some implementations, the methylation genotyping system identifies genomic coordinates of alternative haplotypes containing cytosine bases. This methylation genotyping system can identify candidate cytosine bases within the target genome sample, even if they do not match a reference genome.More specifically, the methylation genotype detection system identifies genomic coordinates of a reference or target genome sample containing a cytosine base, including but not limited to (i) genomic coordinates where the reference base is a cytosine base and the detected genotype contains a cytosine base; or (ii) genomic coordinates where the reference base in the target genome sample is not a cytosine base (e.g., ref=A) but the detected genotype contains a cytosine base.

[0030] In cases where genomic coordinates containing alternative haplotypes have been identified, such as (i) and (ii) genomic coordinates containing candidate cytosine bases, the methylation genotype detection system can determine the estimated methylation level of the candidate cytosine base at the genomic coordinates based on the prior genotype probability of the target genome sample and based on the nucleobases observed within the nucleotide reads. For example, the methylation genotype detection system can generate estimated methylation level values ​​for (a) genomic coordinates where the reference base is a cytosine base and the prior genotype probability indicates the cytosine base, and (b) genomic coordinates where the target genome sample contains a cytosine base as an alternative haplotype. Specifically, the methylation genotype detection system determines the probability of a given observed read stack based on the prior genotype probability from the observed nucleobase and the probability of each nucleobase at the genomic coordinates within the read stack on both the positive and negative strands.

[0031] For the prior genotype probability, the methylation genotype detection system assumes that the different prior genotype probabilities of each nucleobase serve as the basis for determining the estimated methylation level value. For example, the methylation genotype detection system can determine that a haplotype that does not match cytosine is not methylated. As further described below, in some cases, this methylation genotype detection system assumes that (i) the prior probability of a thymine base on the positive strand is approximately equal to a β value, where β represents the position-specific fraction of a cytosine base that is methylated into a thymine base; (ii) the prior probability of a cytosine base on the positive strand is approximately equal to Equation 1 minus the β value; and (iii) the prior probability of an adenine or guanine base on the positive strand is determined by a base detection error probability (e.g., a base detection error greater than 3). In contrast, the prior probability of an adenine, guanine, or thymine base on the negative strand is also determined by a base detection error probability. Furthermore, the prior probability of a cytosine on the positive strand is approximately equal to Equation 1 minus the base detection error probability. This methylation genotype detection system can then generate an estimated methylation level value for the cytosine base by performing a Bayesian inversion on the prior probabilities exhibited by the observed nucleobases within a given read stack.

[0032] Based on the estimated methylation level values ​​of observed nucleosides and base detection quality metrics (e.g., Q score) and / or other sequencing metrics, in some embodiments, the methylation genotype detection system utilizes a variant detection model to generate a posterior genotype probability of a target genome sample at genomic coordinates. More specifically, the methylation genotype detection system determines inputs representing the estimated methylation level values ​​and base detection quality metrics, and feeds these inputs into a variant detection model modified to receive inputs from such methylation level sources. The methylation genotype detection system uses the variant detection model to generate a posterior genotype probability, and further determines the highest posterior genotype probability as the genotype detection.

[0033] In some embodiments, the methylation genotype detection system further utilizes the posterior genotype probability to generate a refined methylation level value of nucleosides at genomic coordinates. This refined methylation level value may represent the percentage of cytosine methylation at genomic coordinates. More specifically, the methylation genotype detection system can determine (a) the genomic coordinates of a reference base that is a cytosine base and a prior genotype probability indicating the cytosine base, and (b) the genomic coordinates of a target genome sample containing a cytosine base as an alternative haplotype. This methylation genotype detection system can determine the refined methylation level based on the posterior genotype probability, the number of reads in the read stack, and the number of methylated nucleobases in the read stack. Because the methylation genotype detection system determines the refined methylation level based on the posterior genotype probability, the refined methylation level can be more accurate than an estimated methylation level.

[0034] As indicated above, the methylation genotype detection system offers several technical advantages over existing sequencing systems by, for example, improving the accuracy and computational efficiency of methylation and genotype detection compared to existing sequencing systems. As mentioned, the methylation genotype detection system improves the accuracy of methylation level values ​​and genotype detection compared to existing sequencing systems. Methylation genotyping systems more accurately identify potential methylation sites in target genome samples. As mentioned, existing systems often fail to estimate methylation at genomic coordinates of target genome samples that do not align with cytosine bases in the reference genome. In contrast, methylation genotyping systems estimate the methylation level of cytosine bases identified within the target genome sample, rather than just the methylation level of locations in the target genome sample containing cytosine bases that match the reference genome. As indicated below, methylation genotyping systems exhibit precision and recall for germline and somatic SNV detection with accuracy approaching that of unmethylated whole-genome sequencing.

[0035] Methylation genotyping systems also more accurately estimate the methylation level of cytosine bases within the target genome sample.Because the methylation genotype detection system considers methylated bases (e.g., cytosine and thymine) on both the positive and negative strands of the target genome sample, it generates more accurate estimated methylation level values ​​for such bases. The methylation genotype detection system can utilize these estimated methylation level values ​​to generate more accurate genotype detections than existing variant detectors based on methylation reads. Furthermore, in some embodiments, the methylation genotype detection system can generate refined methylation level values ​​that improve accuracy relative to initially determined methylation level values ​​by utilizing posterior genotype probabilities used to determine genotype detection.

[0036] In addition to improved genotype and methylation detection accuracy, in some embodiments, the methylation genotype detection system also improves processing and physical resource efficiency compared to existing systems. As described above, because the accuracy of existing genotype detection and methylation assays may not be suitable for clinical benchmarks, some existing systems perform (i) separate methylation sequencing assays to chemically or enzymatically transform nucleotide reads from a genomic sample and determine methylation levels, and (ii) separate DNA sequencing runs that use non-chemically or non-enzymatically transformed nucleotide reads from a genomic sample to determine variant detection. Such separate methylation sequencing assays and DNA sequencing consume and replicate computer processing, memory storage, physical space and reagents for nucleotide sample slides (e.g., flow cells), and software programs (e.g., separate methylation analysis and variant detection software). In contrast to these bifurcation approaches, in some embodiments, methylation genotype detection systems can simultaneously determine methylation level values ​​indicating the methylation level of cytosine bases in a target genomic sample and generate variant detections of the genomic sample with improved accuracy. Therefore, methylation genotype detection systems can efficiently generate epigenetic and genetic sequencing data from a single genomic sample. By generating both methylation level values ​​and variant detection from the same genomic sample, the methylation genotyping system further reduces the amount of computer processing, computer storage, software programs, space on nucleotide sample slides in sequencing equipment, and other resources required to generate accurate sequencing and methylation data.

[0037] As discussed above, this disclosure uses various terms to describe the features and advantages of the methylation genotyping system. As used herein, for example, the term "methylation sequencing assay" refers to an assay that detects, measures, or quantifies the methylation of cytosine from oligonucleotides or other nucleotide sequences. In some cases, methylation sequencing assays detect or quantify the methylation of cytosine at a specific target genomic region or in a specific cell type. Some methylation sequencing assays quantify methylation based on methylation level values.

[0038] Accordingly, the term “methylation level value” refers to a numerical value indicating the amount, percentage, ratio, or quantity of cytosine with added or bonded methyl or hydroxymethyl groups. For example, a methylation level value includes a score (e.g., ranging from 0 to 1) that indicates the percentage or ratio of cytosine bases (e.g., at CpG or other cytosine sites) with added methyl groups at a specific genomic coordinate or genomic region. In some cases, the methylation level value is expressed as a β value ( ) or an M value. For illustration, a β value can be used to estimate the methylation level using the ratio of the signal intensity between methylated alleles and unmethylated alleles at the same genomic coordinate, where 0 represents complete unmethylation and 1 represents complete methylation. In another example, a β value can include a genomic coordinate-specific fraction of cytosine bases that have been methylated to thymine bases. Conversely, an M value can represent the log2 ratio of the signal intensities of methylated and unmethylated probes corresponding to cytosine bases. As described below, the disclosed methylation genotype detection system can determine chain-specific methylation level values, such as a first methylation level value for nucleotides at genomic coordinates on the positive strand and a second methylation level value for nucleotides at genomic coordinates on the negative strand.

[0039] As used herein, the term "refined methylation level value" refers to a modified or updated methylation level value based on new or previously unavailable data (e.g., posterior genotype probabilities). For example, a refined methylation level value includes a score (e.g., ranging from 0 to 1) that indicates the percentage or ratio of cytosine bases (e.g., at CpG or other cytosine sites) at a specific genomic coordinate or genomic region where methyl groups have been added. In some cases, the refined methylation level value is expressed as a β value ( ) or an M value. Furthermore, in some embodiments, the methylation genotype detection system generates a refined methylation level value that is more accurate than an initial estimated methylation level value. For example, the methylation genotype detection system may use posterior genotype probabilities and observed nucleotides to generate a refined methylation level value.

[0040] As used herein, the term "variant detection model" (or simply "variant detector") refers to a probabilistic model for generating rapid sequencing data from nucleotide reads of a sample nucleotide sequence, which includes variant detection and associated metrics. For example, in some cases, a variant detection model refers to a Bayesian probabilistic model for generating variant detection based on nucleotide reads of a sample nucleotide sequence. Such models can handle or analyze sequencing metrics corresponding to read stacking (e.g., multiple nucleotide reads corresponding to a single genome coordinate), including mapping quality, base quality, and various hypotheses, including foreign reads, missing reads, co-detection, etc.Variant detection models may also include multiple components, including but not limited to different software applications or components for mapping and alignment, sorting, repeat tagging, calculating read stacking depth, and variant detection. In some cases, a variant detection model refers to the ILLUMINA DRAGEN model used for variant detection functionality as well as mapping and alignment functionality.

[0041] As further used herein, the term "nucleotide read" (or simply "read") refers to one or more nucleobases (or nucleobase pairs) inferred from all or part of a sample nucleotide sequence (e.g., a sample genomic sequence, complementary DNA). Specifically, a nucleotide read includes a determined or predicted sequence from the nucleobase detection of a nucleotide sequence (or monoclonal nucleotide sequence set) of a sample library fragment corresponding to a genomic sample. For example, in some cases, sequencing devices determine nucleotide reads by generating nucleobase detection of nucleobases passing through nanopores in a nucleotide sample slide, by fluorescent tagging, or by clusters in a flow cell.

[0042] As used herein, the term "nucleobase" refers to a nitrogenous base. Specifically, a nucleobase comprises a component of a nucleotide. For example, a nucleobase may be adenine (A), cytosine (C), guanine (G), or thymine (T).

[0043] As used herein, the term "observed nucleobase" refers to a nucleobase that has been identified or predicted for a nucleotide read. In particular, observed nucleobases include nucleobases detected by sequencing equipment for a nucleotide read. For example, observed nucleobases may include detected nucleobases that, after the nucleotide read has been mapped and aligned, are aligned to genomic coordinates corresponding to cytosine bases in the same reference genome. In some embodiments, a methylation genotype detection system may identify observed nucleobases at a given genomic coordinate from one or more nucleotide reads.

[0044] Also, as used herein, the term "target genome sample" refers to a target genome or a portion of a genome that has undergone sequencing. For example, a genome sample comprises a nucleotide sequence (or a copy of such an isolated or extracted sequence) isolated or extracted from a sample organism. Specifically, a genome sample comprises a whole genome isolated or extracted (in whole or in part) from a sample organism and consisting of nitrogen-containing heterocyclic bases. A genome sample may include fragments or molecules of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other aggregated forms of nucleic acids or chimeric or hybrid forms of nucleic acids as described below. In some cases, the genome sample is present in a sample prepared or isolated by a kit and received by a sequencing device.

[0045] As further used herein, the term “genomic coordinates” (or sometimes simply “coordinates”) refers to a specific location or orientation of a nucleobase within a genome (e.g., the genome of an organism or a reference genome).In some cases, genomic coordinates include an identifier of a specific chromosome in the genome and an identifier of the location of a specific nucleus base on that chromosome. For example, one or more genomic coordinates may include a chromosome number, name, or other identifier (e.g., chr1 or chrX) and one or more specific locations, such as a numbered position following the chromosome identifier (e.g., chr1:1234570 or chr1:1234570-1234870). Furthermore, in some specific implementations, genomic coordinates refer to the source of a reference genome (e.g., mt of a mitochondrial DNA reference genome or SARS-CoV-2 of a SARS-CoV-2 virus reference genome) and the location of the source nucleus base in the reference genome (e.g., mt:16568 or SARS-CoV-2:29001). In contrast, in some cases, genomic coordinates refer to the location of a reference nucleus base without relating to a chromosome or source (e.g., 29727).

[0046] As used herein, the term “impute / imputation” refers to the statistical inference or estimation of the genotype of genomic coordinates or genomic regions. More specifically, imputation may include statistically inferring the genotypes of one or more alleles corresponding to haplotypes of genomic regions of a genome sample. For example, imputation may refer to using marker variants surrounding a genomic region to determine the genotype probability of an allele corresponding to a haplotype of a genomic region. In one or more embodiments, the methylation genotype detection system uses a reference set from a haplotype database and a genotype imputation model (e.g., a hidden Markov model) to imput genotype detection probabilities as the basis for genotype detection.

[0047] As used herein, the term “genotype probability” refers to the likelihood, probability, or score that a genome sample contains a particular genotype at genomic coordinates or a genomic region. Specifically, genotype probability may include a numerical score or measurement indicating the likelihood of a particular genotype. For example, genotype probability may include a numerical score between 0 and 1, where a higher score corresponds to a greater likelihood of a given genotype. Thus, in some cases, genotype probability includes the likelihood of a homozygous reference genotype at one or more genomic coordinates being between 0 and 1, the likelihood of a heterozygous variant genotype, or the likelihood of a homozygous variant genotype. Relatedly, the term "prior genotype probability" refers to the estimated genotype probability before collecting and / or analyzing new data or information (e.g., newly identified metrics or newly observed events). For example, prior genotype probability may refer to the estimated genotype probability before and / or before considering estimated methylation level values. The term "posterior genotype probability" refers to the genotype probability that takes into account or reflects new data or information (e.g., newly identified metrics or newly observed events).For example, the posterior genotype probability can refer to the estimated genotype probability due to imputation and / or consideration of estimated methylation level values ​​or other measures.

[0048] As used herein, the term "base detection quality metric" refers to a specific score or other measurement indicating the accuracy of nucleotide base detection. Specifically, a base detection quality metric includes a value indicating the probability that one or more predicted nucleotide base detections at genomic coordinates contain errors. For example, in some specific embodiments, a base detection quality metric may include a Q score (e.g., a Phred quality score) that predicts the probability of error in the detection of any given nucleotide base. For example, the quality score (or Q score) may indicate the probability of incorrect nucleobase detection at genomic coordinates, with a score equal to 1:100 for Q20, 1:1,000 for Q30, 1:10,000 for Q40, and so on.

[0049] Also as used herein, the term "reference genome" refers to a digital nucleic acid sequence assembled as a representative example (or multiple representative examples) of genes and other genetic sequences of an organism. Regardless of sequence length, in some cases, a reference genome is defined as an exemplary set of genes or a set of nucleic acid sequences in a digital nucleic acid sequence representing an organism. For example, a linear human reference genome may be GRCh38 (or another version of the reference genome) from the Genome Reference Consortium. GRCh38 may include alternating continuous sequences representing alternating haplotypes or alternating nucleobases, such as SNPs and small insertions and deletions (e.g., 10 or fewer base pairs, 50 or fewer base pairs).

[0050] As further used herein, the term “genotype detection” refers to identifying or predicting a specific genotype of a genome sample at a genomic locus. Specifically, genotype detection may include predicting a specific genotype of a genome sample relative to a reference genome or reference sequence at genomic coordinates or genomic regions. For example, in some cases, genotype detection includes identifying or predicting that a genome sample includes both nucleobases and complementary nucleobases at genomic coordinates that are homozygous or heterozygous with respect to a reference base or variant (e.g., a homozygous reference base is represented as 0|0, or a heterozygous variant on a particular strand is represented as 0|1). Genotype detection is often determined against genomic coordinates or genomic regions where single nucleotide variants (SNVs) or other variants have been identified for an organism population. In this disclosure, among other genotype detections, methylation genotype detection systems predict the genotype detection of SNV regions within a genomic sample.

[0051] As used herein, the term "plus strand / positive strand" refers to a DNA strand in which the sequence directly corresponds to the sequence of an RNA transcript that has been translated or can be translated into an amino acid sequence.Relatedly, the term "minus strand / negative strand" refers to a single DNA strand complementary to the positive strand.

[0052] The following paragraphs describe a methylation genotype detection system with reference to exemplary drawings depicting example embodiments and specific implementations. For example, Figure 1 illustrates a schematic diagram of a computing system 100 in which a methylation genotype detection system 106 according to one or more embodiments of the present disclosure operates. As shown, the computing system 100 includes a server device 102, a sequencing device 114, and a user client device 110 connected via a network 118. While Figure 1 illustrates an embodiment of the methylation genotype detection system 106, the present disclosure describes the following alternative embodiments and configurations. As shown in Figure 1, the sequencing device 114, the server device 102, and the user client device 110 can communicate with each other via the network 118. The network 118 includes any suitable network through which the computing devices can communicate. An example network is discussed in more detail below with reference to Figure 15.

[0053] As indicated in FIG1, sequencing device 114 includes sequencing device system 116 for sequencing genomic samples or other nucleic acid polymers, such as when sequencing oligonucleotides extracted from genomic samples as part of a methylation sequencing assay. In some embodiments, by performing sequencing device system 116, sequencing device 114 analyzes nucleotide sequences or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data directly or indirectly on sequencing device 114 using computer-implemented methods and systems (described herein). More specifically, sequencing device 114 receives a nucleotide sample slide (e.g., a flow cell) comprising nucleotide sequences extracted from a sample, and subsequently copies and determines the nucleobase sequence of such extracted nucleotide sequences. As part of a methylation sequencing assay, for example, sequencing device 114 may determine the nucleobase detection of nucleotide reads containing CpG sites or other cytosine sites.

[0054] As described above, by executing sequencing device system 116, sequencing device 114 can run one or more sequencing cycles as part of a sequencing run. For example, by executing methylation genotype detection system 106, sequencing device 114 can (i) sequence certain uracil bases converted from methylated or unmethylated cytosine bases and which are part of a nucleotide read, and (ii) determine the thymine nucleobase detection for such uracil bases as part of a methylation sequencing assay. In one or more embodiments, sequencing device 114 utilizes sequencing by synthesis (SBS) to sequence nucleic acid polymers into nucleotide reads.

[0055] In some cases, server device 102 is located at or near the same physical location as sequencing device 114 or away from sequencing device 114.In practice, in some implementations, server device 102 and sequencing device 114 are integrated into the same computing device. Server device 102 may run sequencing system 104 and / or methylation genotype detection system 106 to generate, receive, analyze, store, and transmit digital data, such as by receiving base detection data, methylation assay data, and / or generating genotype detections.

[0056] As further presented in FIG1, sequencing device 114 may send (and server device 102 may receive) base detection data generated during sequencing operations of sequencing device 114. By executing software in the form of sequencing system 104 or methylation genotype detection system 106, server device 102 may analyze read data and / or detection data, such as sequencing metrics received from sequencing device 114, and may determine the nucleobase sequence of nucleotide reads. In some implementations, sequencing system 104 of server device 102 determines the sequence of nucleosides in DNA and / or RNA segments or oligonucleotides. In addition to processing and determining the sequence of nucleic acid polymers, sequencing system 104 also generates variant detection files indicating one or more genotype detections and / or variant detections for one or more genomic coordinates. Server device 102 may also communicate with user client device 110. Specifically, server device 102 may send data to user client device 110, including variant detection files (VCFs) or other information indicating nucleobase detections, methylation level values, sequencing metrics, error data, or other metrics.

[0057] In some embodiments, server device 102 includes a distributed set of servers, wherein server device 102 includes a number of server devices distributed across network 118 and located in the same or different physical locations. Further, server device 102 may include a content server, application server, communication server, web hosting server, or another type of server.

[0058] As further shown and indicated in FIG1, user client device 110 may generate, store, receive, and transmit digital data. Specifically, user client device 110 may receive variant detection, methylation level values, and corresponding sequencing metrics from server device 102, or base detection data (e.g., BCL or FASTQ) and corresponding sequencing metrics from sequencing device 114. Furthermore, user client device 110 may communicate with server device 102 to receive VCF, which includes nucleobase detection and / or other metrics, such as base detection quality metrics or filter-based metrics. User client device 110 may accordingly present or display information related to variant detection or other nucleobase detection to the user associated with it within a graphical user interface. Specifically, user client device 110 may present results from methylation sequencing assays or graphs indicating methylation level values ​​of target cytosine bases.

[0059] Although FIG1 depicts user client device 110 as a desktop or laptop computer, user client device 110 may include various types of client devices. For example, in some embodiments, user client device 110 includes non-mobile devices such as desktop computers or servers, or other types of client devices. In yet other embodiments, user client device 110 includes mobile devices such as laptops, tablets, mobile phones, or smartphones. Additional details regarding user client device 110 are discussed below in conjunction with FIG15. Specification 9 / 36 pages 14 CN 121359206 A

[0060] As further shown in FIG1, user client device 110 includes sequencing application 112. Sequencing application 112 may be a web application or native application (e.g., a mobile application, a desktop application) stored and executed on user client device 110. The sequencing application 112 may include instructions (when executed) that cause the user client device 110 to receive data from the methylation genotype detection system 106 and present (e.g., base detection data from BCL, data from VCF, or data from methylation sequencing assays) for display at the user client device 110.

[0061] As further illustrated in FIG1, one version of the methylation genotype detection system 106 may be located on the user client device 110 as part of the sequencing application 112, or on the sequencing device 114 as part of the sequencing device system 116. In some embodiments, the methylation genotype detection system 106 is implemented (e.g., wholly or partially located) on the user client device 110. In other embodiments, the methylation genotype detection system 106 is implemented by one or more other components of the computing system 100, such as the sequencing device 114. Specifically, the methylation genotype detection system 106 may be implemented in a variety of different ways across the sequencing device 114, the user client device 110, and the server device 102. As illustrated in Figure 1, the methylation genotype detection system 106 is implemented by a sequencing system 104 (e.g., fully or partially), which is implemented by a server device 102. In at least one example, the methylation genotype detection system 106 may be downloaded from the server device 102 to the sequencing device 114 and / or the user client device 110, wherein all or part of the functionality of the methylation genotype detection system 106 is performed at each corresponding device within the computing system.

[0062] As further illustrated in Figure 1, the methylation genotype detection system 106 may implement a variant detection model 120 and a methylation assay system 122. By executing the variant detection model 120, the methylation genotype detection system 106 can align nucleotide reads with a reference genome and determine variant detection based on the aligned nucleotide reads.In some implementations, the methylation genotype detection system 106 analyzes each nucleotide of each read to determine the “fit” position of the nucleotide read relative to a reference sequence—for example, the position where a base within the read aligns with a base in the reference genome (or receives information indicating that position). In some cases, the methylation genotype detection system 106 aligns many reads at a single genomic coordinate, resulting in read stacking. In some implementations, the methylation assay system 122 is also implemented by the methylation genotype detection system 106. The methylation assay system 122 determines the methylation level value at CpG sites or other cytosine sites. As previously mentioned, the variant detection model 120 and / or the methylation assay system 122 may be implemented by the methylation genotype detection system 106.

[0063] As previously mentioned, the methylation genotype detection system 106 may generate both genotype detection and methylation level values ​​based on nucleotide read data. Figures 2A and 2B illustrate an overview of a methylation genotype detection system 106 according to one or more embodiments of the present disclosure, which detects genotypes from read stacks and generates methylation level values. Figures 2A and 2B illustrate a series of actions 200, including action 202 of identifying nucleotide reads of a genomic sample, action 204 of determining an estimated methylation level value of nucleosides at genomic coordinates, action 206 of generating a posterior genotype probability of the genomic sample at genomic coordinates, action 208 of generating genotype detection of the genomic sample at genomic coordinates, and action 210 of generating a refined methylation level value of nucleosides at genomic coordinates.

[0064] As shown in Figure 2A, the methylation genotype detection system 106 performs action 202 of identifying nucleotide reads. Specifically, the methylation genotype detection system 106 performs action of identifying nucleotide reads containing one or more nucleosides converted by methylation sequencing for a target genomic sample. In some cases, for example, the methylation genotype detection system 106 receives data representing nucleotide reads 228 of a genome sample that has been sequenced by a sequencing device. Such data for nucleotide reads includes sequences of nucleobase detections determined by the sequencing device. In some specific embodiments, after receiving the read data, the methylation genotype detection system 106 aligns the nucleotide reads 228 with a reference genome 230. Based on the aligned nucleotide reads, the methylation genotype detection system 106 can determine the genomic coordinates of the target genome sample relative to the reference genome 230 and one or more nucleobase detections of the genomic region. Specification 10 / 36 pages 15 CN 121359206 A

[0065] As further illustrated in FIG2A, the methylation genotype detection system 106 performs an action 204 of determining an estimated methylation level value.Specifically, the methylation genotype detection system 106 determines an estimated methylation level 216 for a cytosine base (or other nucleobase) at a genomic coordinate based on a prior genotype probability 212 of the observed nucleobase 214 at the genomic coordinate within the nucleotide read. As mentioned above, such estimated methylation level values ​​can be determined for genomic coordinates of a genomic sample containing a reference cytosine base or a nucleobase that may be referred to as cytosine (e.g., a cytosine base on the negative strand). In some specific embodiments, the methylation genotype detection system 106 utilizes a Bayesian method to determine the estimated methylation level value. Figures 7A and 7B illustrate the methylation genotype detection system 106 for determining the estimated methylation level value according to one or more embodiments of the present disclosure.

[0066] Figure 2B illustrates the methylation genotype detection system 106 performing the action 206 of generating a posterior genotype probability. In some embodiments, the methylation genotype detection system 106 uses a variant detection model 222 to generate a posterior genotype probability 224 of the target genome sample at genomic coordinates based on an estimated methylation level value 216 of observed nucleobases and a base detection quality metric 220. In some embodiments, the methylation genotype detection system 106 (in addition to the base detection quality metric 220) also uses a sequencing metric 218 as input to the variant detection model 222 as part of generating the posterior genotype probability 224. Figures 9 and 10 and the corresponding paragraphs provide additional details on how the methylation genotype detection system 106 generates the posterior genotype probability according to one or more embodiments of the present disclosure.

[0067] As further illustrated in Figure 2B, the methylation genotype detection system 106 performs an action 208 to generate a genotype detection. In some examples, based on the posterior genotype probability 224, the methylation genotype detection system 106 generates a genotype detection of the target genome sample at the genomic coordinates corresponding to the observed nucleotide 214. For example, the genotype detection may indicate that the target genome sample contains a predicted combination of nucleotides at the genomic coordinates. In one or more embodiments, the methylation genotype detection system 106 generates a genotype detection of the target genome sample by determining a predicted combination of nucleotides corresponding to the highest posterior genotype probability.

[0068] As further shown in FIG2B, in some embodiments, the methylation genotype detection system 106 performs an action 210 of generating a refined methylation level value. Specifically, the methylation genotype detection system 106 may generate a refined methylation level value of cytosine nucleotides at the genomic coordinates based on the posterior genotype probability 224 and the observed nucleotide 226. Similar to estimating methylation level values ​​216, such refined methylation level values ​​can be determined for genomic coordinates of a genome sample containing a reference cytosine base or a nucleobase that may be called cytosine (e.g., a cytosine base on the negative strand).The refined methylation level value generated as part of action 210 can be more accurate than the estimated methylation level value 216. More specifically, in some cases, the methylation genotype detection system 106 determines the refined methylation level value based on modified or updated input data—that is, the posterior genotype probability 224, the number of reads in the read stack, and the number of methylated bases in the read stack. Figure 10 and the corresponding discussion further detail how the methylation genotype detection system 106 generates refined methylation level values ​​according to one or more embodiments of the present disclosure.

[0069] As mentioned, the methylation genotype detection system 106 identifies nucleotide reads containing one or more nucleotides converted by methylation sequencing. Figure 3 illustrates that the methylation genotype detection system 106 according to one or more embodiments of the present disclosure can utilize various types of methylation sequencing protocols (or receive data from these methylation sequencing protocols) as part of identifying nucleotide reads containing converted nucleotides. As an overview, different methylation sequencing protocols utilize different conversions. For example, some methylation sequencing protocols convert unmethylated cytosine bases to thymine bases (i.e., C-to-T conversion). Other methylation sequencing protocols convert methylated cytosine bases to thymine bases (i.e., 5mC, 5hmC-to-T conversion). The methylation genotype detection system 106 can utilize methylation data from the specification of any type of methylation sequencing protocol, page 11 / 36, 16 CN 121359206 A. The following paragraphs further detail the various effects of C-to-T conversion protocol 302 and 5mC, 5hmC-to-T conversion protocol 304, as well as other conversion protocols.

[0070] Figure 3 illustrates C-to-T conversion protocol 302. In C-to-T conversion protocol 302, unmethylated bases are converted to thymine bases. For example, C-to-T conversion protocol 302 includes bisulfite and enzymatic methylation sequencing (EM-seq) for methylation sequencing assays. EM-seq can be performed, for example, as described by Romualdas Vaisvila et al., Enzymatic Methyl Sequencing Detects DNA Methylation at Single-Base Resolution from Picograms of DNA, 30 Genome Research 1280-1289 (2021), the entire text of which is incorporated herein by reference.

[0071] Figure 3 depicts an example target genome sample containing unmethylated cytosine base 308 and methylated cytosine base 306. As shown in Figure 3, unmethylated cytosine base 308 is converted to thymine base 310, while methylated cytosine base 306 remains unconverted.As shown, 98% of the cytosine bases are converted to thymine bases, while approximately 2% of the methylated cytosine bases remain cytosine bases.

[0072] Figure 3 further illustrates 5mC, 5hmC to T conversion scheme 304. In these conversion schemes, the methylated bases are converted to thymine bases. For example, Tet-assisted pyridineborane sequencing (TAPS) uses the deca-eutectoid translocation (TET) enzyme for methylation sequencing determination, as described by Yibin Liu et al., “Bisulfite-free Direct Detection of 5-Methylcystosine and 5-Hydroxymethylcystosine at Base Resolution,” 36 Nature Biotechnology 424-29 (2019). In some TET-dependent assays, methylation sequencing assays use TET enzymes to convert 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) into oxidation products, which are then deaminated using apolipoprotein B mRNA editing enzyme, catalytic peptide (APOBEC) 3A, or another APOBEC protein by converting the unmodified cytosine to uracil bases. In some cases, such converted uracil bases are detected as thymine bases during sequencing.

[0073] As further shown in Figure 3, an example target genome sample contains unmethylated cytosine base 314 and methylated cytosine base 316. As part of the 5mC, 5hmC to T conversion scheme 304, the methylated cytosine base 316 is converted to thymine base 318, and the unmethylated cytosine base 314 remains unchanged. Unlike C-to-T conversion scheme 302 (in which most cytosine bases are converted to thymine bases), in the 5mC, 5hmC-to-T conversion scheme 304, only 2% of the cytosine bases are converted to thymine bases. Subsequently, 98% of the cytosine bases remain cytosine bases. The methylation genotype detection system 106 can accurately and simultaneously generate methylation level values ​​and genotype detection for genome samples that have undergone either C-to-T conversion scheme 302 or 5mC, 5hmC-to-T conversion scheme 304.

[0074] As previously mentioned, the methylation genotype detection system 106 estimates genotype and methylation level values ​​based on observed nucleobases. Figure 4 illustrates the goal of the methylation genotype detection system 106 according to one or more embodiments of the present disclosure to accurately estimate genotype and methylation level values ​​based on observed nucleobases. Generally, Figure 4 illustrates a side-by-side comparison of two example genome samples. More specifically, Figure 4 includes (i) a genome sample 416 containing a fully unmethylated CC genotype and (ii) a genome sample 418 containing a fully methylated CC genotype.Genome samples 416 and 418 each comprise two homologous strands—that is, each sample contains nucleotides in both the positive and negative strands at genomic coordinates. As shown in Figure 4, the goal of the methylation genotype detection system 106 is to accurately estimate the genotype and methylation level values ​​at genomic coordinates.

[0075] As shown in Figure 4, the methylation genotype detection system 106 aims to accurately estimate genotype and methylation level values ​​based on observed nucleotides. Generally, the methylation genotype detection system 106 accesses nucleotide read data to identify nucleotides observed at specific genomic locations. Observed nucleotides covering or overlapping a single genomic coordinate or region can also be referred to as read stacking. For illustration, observed nucleotide 406 represents multiple nucleotide reads corresponding to genomic coordinate 402. Observed nucleobase 406 comprises observed nucleobases from four nucleotide reads from the positive strand 410 and observed nucleobases from four nucleotide reads from the negative strand 412. Similarly, observed nucleobase 408 comprises observed nucleobases from the positive strand 420 and observed nucleobases from the negative strand 414 at genomic coordinate 404. Figure 4 depicts such observed nucleobases only at a single genomic coordinate, but does not depict the remaining nucleobases from the corresponding nucleotide reads.

[0076] Figure 4 illustrates a genomic sample 416 having cytosine bases on two haplotypes of the same strand (e.g., the + strand) at genomic coordinate 402. The genotype of genomic sample 416 at genomic coordinate 402 is CC. More specifically, the CC genotype represents a portion of the two haplotypes of genomic sample 416 at genomic coordinate 402, thus illustrating that genomic sample 416 is diploid. As further shown in Figure 4, the cytosine base at genomic coordinate 402 is unmethylated (i.e., ). As further described below, the methylation genotype detection system 106 includes instructions or algorithms designed to accurately infer the CC genotype and zero methylation based on observed nucleobases 406 from nucleotide reads covering genomic coordinate 402.

[0077] Figure 4 also illustrates genomic samples 418 with cytosine bases on two haplotypes of the same strand (e.g., + strand) at genomic coordinate 404. The genotype at genomic coordinate 404 is CC. As shown, each cytosine base in the cytosine base at genomic coordinate 404 is methylated (i.e., ). As further described below, the methylation genotype detection system 106 includes instructions or algorithms designed to accurately infer the CC genotype and full methylation based on observed nucleobases 408 at genomic coordinate 404.

[0078] As previously mentioned, existing sequencing and methylation detection systems tend to generate inaccurate methylation level values.Figure 5 illustrates the models used by existing sequencing and methylation detection systems that inaccurately estimate methylation levels. More specifically, Figure 5 illustrates how one or more existing sequencing and methylation detection systems attempt to estimate methylation levels (for genomic coordinates based on observed nucleotides). For example, existing sequencing and methylation detection systems seek to estimate methylation levels representing the proportion of methylated conversions in read stacks. In some examples, methylation levels indicate the proportion of cytosine bases converted to thymine bases on the positive strand or the proportion of guanine bases converted to adenine bases on the negative strand.

[0079] Figure 5 illustrates two examples of observed nucleotides from nucleotide reads covering two genomic coordinates. As shown in Figure 5, genotype 506 CC of the first genomic sample has been detected at the first genomic coordinate. The cytosine bases at the first genomic coordinate are completely unmethylated ( ). As shown in the figure, the observed nucleobase 502 corresponding to the first genomic coordinate contains three cytosine bases and one thymine base on the positive strand, and four cytosine bases on the negative strand.

[0080] Figure 5 also illustrates genotype 508 of the second genomic sample that has been detected at the second genomic coordinate. As shown in Figure 5, the observed nucleobase 504 corresponds to genotype 508. The positive strand of the observed nucleobase 504 contains two cytosine bases and two thymine bases. Similarly, the negative strand of the observed nucleobase 504 contains two cytosine bases and two thymine bases. As further illustrated in Figure 5, genotype 508 of CT is completely unmethylated.

[0081] As shown in Figure 5, given the detected genotypes 506 and 508 and the observed nucleobases 502 and 504, existing sequencing and methylation detection systems often inaccurately estimate methylation levels. In some examples, existing sequencing and methylation detection systems (pages 13 / 36, CN 121359206 A) estimate methylation levels using the observed number of each nucleobase. For example, and as shown in Figure 5, existing sequencing and methylation detection systems may rely on equations such as Equation 510 to estimate methylation levels. Figure 5 illustrates Equation 512 and Equation 514, which demonstrate specific implementations of Equation 510 using different observed nucleobases. More specifically, Equation 512 shows the implementation of Equation 510 for observed nucleobase 502, and Equation 514 shows the implementation of Equation 510 for observed nucleobase 504. As shown in Equation 512, for a first genome sample at the first genomic coordinate corresponding to genotype 506, existing sequencing and methylation detection systems may estimate a methylation level of 0.25 based on observed nucleobase 502.However, the estimated methylation level value of 0.25 deviates from the true methylation level value of 0.

[0082] In some cases, existing sequencing and methylation detection systems estimate methylation level values ​​that deviate significantly from the true methylation level values. For example, and as illustrated in Figure 5, existing sequencing and methylation detection systems may use Equation Specific 514 to estimate the methylation level value of cytosine bases in a second genome sample located at the second genome coordinates corresponding to genotype 508. When compared to the true methylation level value of 0, existing sequencing and methylation detection systems significantly overestimate the methylation level value by 0.5. As described below, the methylation genotype detection system 106 implements a model to avoid the inaccurate methylation level value estimates generated by Equation 510 used by some existing sequencing and methylation detection systems.

[0083] Technical deficiencies in existing sequencing and methylation detection systems often lead to them overestimating the methylation level value of a particular genotype. Figure 6 and the corresponding paragraph illustrate the overestimated methylation level values ​​generated by existing sequencing and methylation detection systems. Figure 6 illustrates a schema showing the methylation level values ​​estimated by existing sequencing and methylation detection systems compared to the true methylation level values ​​for a specific genotype. As an overview, Figure 6 includes schema 602 for genotype CA, schema 604 for genotype CC, schema 606 for genotype CG, and schema 608 for genotype CT. The schema illustrated in Figure 6 also includes reference lines 610a, 610b, 610c, and 610d, which indicate estimated methylation level values ​​that will be exactly equal to the true methylation level values. Although Figure 6 illustrates a schema generated from simulated data, existing sequencing and methylation detection systems using Equation 510 will necessarily use simulated observed nucleotide reads to generate the estimated methylation level values ​​in the schema of Figure 6.

[0084] As just noted, the schema illustrated in Figure 6 depicts results from simulated data, where observed nucleotide reads are simulated. For example, read stacking was simulated for each real genotype (e.g., CA, CC, CG, and CT) and its real methylation level value. Existing sequencing and methylation detection systems were used to analyze the simulated nucleotide reads and estimate the methylation level values.

[0085] As illustrated in Figure 6, existing sequencing and methylation detection systems tend to overestimate the methylation level values ​​of genotypes containing adenine and thymine bases. More specifically, existing sequencing and methylation detection systems tend to assume the presence of adenine and thymine bases due to methylation conversion. However, at the genomic coordinates depicted in Figure 6 for real CA or CT genotypes, not all adenine or thymine bases represent methylation conversion. As shown in Figure 6, existing sequencing and methylation detection systems tend to overestimate the methylation level values ​​in the case of CA and CT genotypes.For example, Figures 602 and 608 depict how estimated methylation levels can be higher than the actual methylation levels.

[0086] As further shown in Figure 6, existing sequencing and methylation detection systems can overestimate and underestimate the methylation levels of the CG genotype. For the CG genotype, methylation can occur on both the positive and negative strands. For example, C on the positive strand can be methylated, and C on the negative strand corresponding to G can also be methylated. Methylation of cytosine bases on both the positive and negative strands further corrupts the data, causing existing sequencing and methylation detection systems to overestimate and underestimate the methylation levels of the CG genotype. Also depicted in Figure 6, 606, existing sequencing and methylation detection systems tend to overestimate the low-range methylation levels of the CG genotype and underestimate the high-range methylation levels of the CG genotype. For example, diagram 606 indicates that existing sequencing and methylation detection systems overestimate the estimated methylation level on the left side of diagram 606, while underestimating the methylation level relative to reference line 610c.

[0087] As further shown in Figure 6, existing sequencing and methylation detection systems estimate the methylation level of the CC genotype more accurately than other genotypes. As illustrated in Figure 6, the estimated methylation level shown in diagram 604 roughly tracks reference line 610b. However, this accuracy is limited to the true CC genotype.

[0088] As mentioned, the methylation genotype detection system 106 generates a more accurate methylation level estimate, in part, by using prior genotype probabilities to determine the estimated methylation level of cytosine bases at genomic coordinates while detecting the genotype of the genome sample. Figures 7A and 7B illustrate a series of actions according to one or more embodiments of the present disclosure, through which the methylation genotype detection system 106 determines an estimated methylation level value for a cytosine base (or other candidate nucleobase). Generally, the methylation genotype detection system 106 utilizes a Bayesian method to determine an estimated methylation level value for a cytosine base, or a nucleobase that may be called cytosine (e.g., a cytosine base on the negative strand). As a summary, Figures 7A and 7B illustrate a series of actions including action 702 of identifying observed nucleobases, action 704 of determining the prior probability of each nucleobase at genomic coordinates, action 706 of determining the probability of observed nucleobases for each possible genotype, and action 708 of performing a Bayesian inversion.

[0089] As shown in Figure 7A, the methylation genotype detection system 106 performs action 702 of identifying observed nucleobases. Generally, the methylation genotype detection system 106 considers genotypes that can be methylated. In some specific implementations, the methylation genotype detection system 106 performs action 704 by first determining the genomic coordinates of cytosine bases in the target genome sample.Some existing sequencing and methylation detection systems determine the genomic coordinates of cytosine bases by referring to a reference genome. However, these existing systems may fail to assess cytosine bases in the target genome sample that do not align with cytosine bases in the reference genome. Therefore, in some implementations, a methylation genotype detection system 106 identifies genomic coordinates within the target genome sample that can be detected as containing cytosine. Thus, the methylation genotype detection system 106 predicts the methylation level of genomic coordinates containing cytosine alleles. More specifically, the methylation genotype detection system 106 identifies and considers (i) genomic coordinates where the reference base is a cytosine base and the detected genotype contains a cytosine base, or (ii) genomic coordinates where the reference base in the target genome sample is not a cytosine base (e.g., ref=A) but the detected genotype contains a cytosine base. In some examples, the methylation genotype detection system 106 identifies genomic coordinates in the target genome sample that can be detected as cytosine bases based on prior genotype probability.

[0090] Furthermore, and as part of the action 702 of performing the identification of observed nucleosides, the methylation genotype detection system 106 identifies nucleotide reads corresponding to the genomic coordinates of cytosine bases in the target genome sample. As previously described, the methylation genotype detection system 106 identifies genomic coordinates covering cytosine bases in the reference genome or nucleotide reads aligned to those coordinates. Alternatively or additionally, the methylation genotype detection system 106 identifies genomic coordinates covering nucleosides in the target genome sample that may be referred to as cytosine or nucleotide reads aligned to those coordinates. The methylation genotype detection system 106 may compile read stacks containing observed nucleosides 714 at genomic coordinates within the nucleotide reads. For example, observed nucleosides 714 are aligned to genomic coordinates containing cytosine bases in the reference genome or target genome sample.

[0091] As further shown in FIG7A, the methylation genotype detection system 106 may determine observed nucleosides corresponding to the positive and negative strands of the target genome sample. For example, and as illustrated in Figure 7A, the methylation genotype detection system 106 identifies observed nucleotides compared to genomic coordinates with a true genotype of CC. As shown as part of action 702, the methylation genotype detection system 106 identifies two positive-strand cytosine bases and two positive-strand thymine bases. The methylation genotype detection system 106 identifies four negative-strand cytosine bases. Positive-strand nucleotides are included in the nucleobase prediction of sample nucleotides 716 at genomic coordinates.

[0092] As indicated in Figure 7A, the combination of observed nucleotides 714 illustrated in Figure 7A provides evidence of either a CT genotype or a methylated CC genotype.Specifically, positive-strand thymine bases can be evidence of transformed methylated (or unmethylated) positive-strand cytosine bases. Therefore, the methylation genotype detection system 106 considers both such genotypes when determining prior probabilities.

[0093] Figure 7A further illustrates the methylation genotype detection system 106 performing the action 704 of determining the prior probability of each nucleobase at a genomic coordinate. As previously mentioned, the methylation genotype detection system 106 determines the prior probability of each nucleobase at a genomic coordinate containing at least one cytosine base (e.g., CC, CT, CA, or CG). Generally, the methylation genotype detection system 106 assumes that genomic coordinates with observed nucleobases supporting a genotype or haplotype without cytosine bases are not subject to methylation transformation and therefore does not utilize the Bayesian transformation depicted in Figures 7A-7B as part of the genotype detection for determining such genomic coordinates. However, for genomic coordinates with observed nucleobases supporting a genotype containing cytosine bases, the methylation genotype detection system 106 is performed as depicted in Figures 7A and 7B, and predicts the prior probability of the nucleobases, and subsequently predicts the posterior genotype probability of such genomic coordinates by incorporating methylation level values.

[0094] As indicated by graph 710 in Figure 7A, in some embodiments, the methylation genotype detection system 106 accesses or identifies the prior probability of observed nucleobases 714 on the positive strand. As an overview of Figure 710, the methylation genotype detection system 106 determines (i) the prior probability of a thymine base on the positive strand is approximately equal to a β value, where the β value represents the position-specific fraction of a cytosine base converted to a thymine base by methylation, (ii) the prior probability of a cytosine base on the positive strand is approximately equal to Equation 1 minus the β value, and (iii) the prior probability of an adenine or guanine base on the positive strand is determined by the base detection error probability (e.g., a base detection error greater than 3).

[0095] As just noted and more specifically shown in Figure 7A, Figure 710 depicts the prior probability of each nucleobase at a genomic coordinate on the positive strand. As shown, the methylation genotype detection system 106 assumes that the prior probabilities of positive-strand adenine and positive-strand guanine bases are largely unaffected by methylation assay conversion. Therefore, in some specific implementations, the methylation genotype detection system 106 determines the prior probability of a positive-strand adenine base and the prior probability of a positive-strand guanine base to be equal to the prior probability representing a base detection error.

[0096] As mentioned, a positive-strand thymine base can be evidence of a converted methylated (or unmethylated) positive-strand cytosine base. More specifically, the prior probability of positive-strand thymine reflects the likelihood that a cytosine base has been converted to a thymine base.Therefore, the prior probability of a positive-strand thymine base is approximately equal to the positive-strand methylation level value. Thus, the prior probability of a positive-strand thymine base at genomic coordinates can be expressed by the equation. In some specific implementations, the positive-strand methylation level value represents the percentage of cytosine methylation on the positive strand.

[0097] As further illustrated in Figure 7A, the methylation genotype detection system 106 considers that a positive-strand cytosine base can be methylated into a thymine base. Therefore, the methylation genotype detection system 106 determines the prior probability of a positive-strand cytosine base not only based on base detection errors but also on whether the cytosine base has been methylated into thymine. Therefore, the methylation genotype detection system 106 can determine the prior probability of a positive-strand cytosine base at genomic coordinates and the positive-strand methylation level value. Specification 16 / 36 pages 21 CN 121359206 A

[0098] In some embodiments, and as shown in FIG7A, the methylation genotype detection system 106 does not incorporate the prior probability of a base detection error into the probability of a positive-strand cytosine base or the probability of a positive-strand cytosine base. More specifically, it is approximate, and in many cases, the prior probability of a base detection error is negligible compared to the positive-strand methylation level value. In some embodiments, the methylation genotype detection system 106 incorporates the prior probability of a base detection error into or into. For example, in some embodiments, the methylation genotype detection system 106 will approximate as equal to.

[0099] The effect of methylation conversion on the prior probability of a positive-strand nucleotide base has a similar case, involving complementary bases on the negative strand. As an overview of Figure 712, the prior probability of adenine, guanine, or thymine bases on the negative strand determined by the methylation genotype detection system 106 is also determined by the base detection error probability. Further, the prior probability of cytosine on the positive strand is approximately equal to Equation 1 minus the base detection error probability. Specifically, Figure 712 depicts the probability of each nucleobase at genomic coordinates on the negative strand. In some embodiments, the methylation genotype detection system 106 determines the prior probability of a negative strand adenine base, the prior probability of a negative strand guanine base, and the prior probability of a negative strand thymine base to be equal to the prior probability representing a base detection error. In some examples, the methylation genotype detection system 106 divides by 3 because it assumes that an erroneous base detection can be detected as any of a negative strand adenine base, a negative strand guanine base, or a negative strand thymine base. The methylation genotype detection system 106 does not record the most likely error and assumes that all three have equal probability.The methylation genotype detection system 106 applies similar reasoning to determine and (in some specific implementations) , and utilizes the same values ​​when determining the impact of base detection errors on those probabilities.

[0100] As further illustrated by Graph 712 in Figure 7A, the methylation genotype detection system 106 determines the prior probability of a negative-strand cytosine base equal to the probability where represents a base detection error. This reflects the assumption of the methylation genotype detection system 106 that, in the case of a potential CC genotype, the methylation genotype detection system 106 will observe cytosine bases with a higher probability. Under this assumption, the methylation genotype detection system 106 will predict a lower probability of observing adenine, guanine, and thymine bases proportional to the base detection error. Furthermore, since negative-strand cytosine bases cannot be methylated, the prior probability of negative-strand cytosine bases is generally unaffected by methylation. corresponds to G on the original DNA template.

[0101] As shown in FIG7B, the methylation genotype detection system 106 performs the action 706 of determining the probability of observed nucleosides for each possible genotype. In some embodiments, the methylation genotype detection system 106 determines the prior probability of observed nucleosides at genomic coordinates based on the number of observed nucleosides for each nucleotide read from a target genome sample at genomic coordinates, the probability of each nucleobase at genomic coordinates, and the corresponding prior genotype probability. For example, in some specific embodiments, the methylation genotype detection system 106 uses the following equation to determine the prior probability of observed nucleosides at genomic coordinates on the positive strand:

[0102]

[0103] where represents the number of observed positive-strand adenine bases, represents the number of observed positive-strand cytosine bases, represents the number of observed positive-strand guanine bases, and

[0104] represents the number of observed positive-strand thymine bases at genomic coordinates. In the above equation, Table 17 / 36, page 22, CN 121359206 A, represents the positive-strand methylation level value and the total number of observed nucleosides on the positive strand. For example, the methylation genotype detection system 106 determines the probability of observed nucleosides on the positive strand given the assumption that a given genotype is a true potential genotype. For example, and as shown in Figure 7B, the methylation genotype detection system 106 determines the probability of observed nucleosides given a true genotype of CC. In some embodiments, the methylation genotype detection system 106 determines the probability of observed nucleosides at genomic coordinates given each possible true genotype (e.g., AA, AC, AG, AT, CC, CG, CT, GG, GT, TT).

[0105] In some embodiments, as part of action 706, the methylation genotype detection system 106 uses the following equation to determine the probability of an observed nucleobase at a genomic coordinate on the negative strand:

[0106]

[0107] where represents the number of observed negative strand adenine bases, represents the number of observed negative strand cytosine bases, represents the number of observed negative strand guanine bases, and

[0108] represents the number of observed negative strand thymine bases at a genomic coordinate. In the above equation, represents the negative strand methylation level value (or the percentage of cytosine methylation on the negative strand), and represents the total number of observed nucleobases on the negative strand. In some embodiments,

[0109] the methylation genotype detection system 106 determines the probability of an observed nucleobase on the negative strand given the assumption that a given genotype is a true potential genotype. For example, and as shown in FIG7B, the methylation genotype detection system 106 determines the probability of observed nuclei given a true genotype of CC. The methylation genotype detection system 106 determines the probability of observed nuclei for each possible potential genotype.

[0110] As described above, the methylation genotype detection system 106 determines the probability of observed nuclei at the genomic coordinates of the positive and negative strands. As further shown in FIG7B, the methylation genotype detection system 106 compresses the probabilities of the positive and negative strands into the probability of observed nuclei at the entire stack of observed nuclei at a given genomic coordinate. In some specific embodiments, the methylation genotype detection system 106 uses the following equation to determine the probability of observed nuclei at a given genomic coordinate:

[0111]

[0112] where represents the observed nuclei at the genomic coordinates of the positive strand, and represents the observed nuclei at the genomic coordinates of the negative strand. In the above equation, represents the methylation value or cytosine methylation percentage at the genomic coordinates, represents the positive strand methylation level, and represents the negative strand methylation level. represents the total number of observed nucleotides at the genomic coordinates, and is equal to the sum of the total number of observed nucleotides on the positive strand and the total number of observed nucleotides on the negative strand. As mentioned above, represents the hypothetical true genotype, and the methylation genotype detection system 106 determines the probability of observed nucleotides for each possible genotype at the genomic coordinates. The methylation genotype detection system 106 enumerates across all possible genotypes.

[0113] As further shown in Figure 7B, the methylation genotype detection system 106 performs an action 708 of performing a Bayesian inversion. More specifically, the methylation genotype detection system 106 generates an estimated methylation level value as represented by by performing a Bayesian inversion on the probability of observed nucleotides determined in action 706.Bayesian inversion can be represented by the following equation: Specification 18 / 36 Page 23 CN 121359206 A

[0114]

[0115] Wherein the methylation genotype detection system 106 determines the prior genotype probability of each potential genotype given the data. More specifically, the methylation genotype detection system 106 determines the prior genotype probability of a given observed nucleobase at the genomic coordinates and the total number of observed nucleobases at the genomic coordinates.

[0116] As further illustrated in Figure 7B, the methylation genotype detection system 106 further performs Bayesian inversion to solve for the positive and negative methylation levels. As shown, the methylation genotype detection system 106 generates estimated methylation levels for both the positive and negative strands. In one example, at the CG heterozygous position, the methylation genotype detection system 106 generates a positive-strand methylation level value for the positive strand of haplotype C and a negative-strand methylation level value for the negative strand of haplotype G. For example, the methylation genotype detection system 106 can determine the positive-strand methylation level value using the following equation:

[0117]

[0118] where the methylation genotype detection system 106 estimates the expected value of the positive-strand methylation level value given the data. In some specific implementations, the methylation genotype detection system 106 estimates by integrating all possible values ​​from 0 to 1. Generally, the methylation genotype detection system 106 analytically solves the integral to obtain the above expression, which contains a single fraction of the sum of different terms—each term being a genotype-specific term. As shown in the above equation, represents the prior genotype probability and the probability of the observed nucleobase at the genomic coordinates. The above fraction also includes the probability of the positive-strand methylation level value and the probability of the negative-strand methylation level value. These probabilities can be tuned using a β prior function expressed as follows

[0119]

[0120] where represents the tunable parameter. When and equals 1, the distribution collapses to a uniform distribution. In some embodiments, the methylation genotype detection system 106 is tuned to find the β prior that produces the highest accuracy.

[0121] In some embodiments, the methylation genotype detection system 106 generates estimated methylation level values ​​of genomic coordinates by combining positive and negative methylation level values. The methylation genotype detection system 106 can utilize the estimated methylation level values ​​to improve the accuracy of genotype detection. Furthermore, the methylation genotype detection system 106 can refine the estimated methylation level values ​​to generate refined methylation level values. Figure 10 and the corresponding paragraph illustrate embodiments of the methylation genotype detection system 106 for generating refined methylation level values ​​according to one or more embodiments.

[0122] As mentioned, in some embodiments, the methylation genotype detection system 106 utilizes a variant detection model to generate a posterior genotype probability based on estimated methylation levels of observed nucleobases and a base detection quality metric. In some cases, the variant detection model includes an imputation model and other components. Figure 8 illustrates an overview of the methylation genotype detection system 106 for generating posterior genotype probabilities according to one or more embodiments of this disclosure, page 19 / 36, 24 CN 121359206 A. In some examples, the methylation genotype detection system 106 applies a genotype imputation model (such as a hidden Markov model (HMM)-based genotype imputation model) to nucleotide reads corresponding to genomic coordinates. By applying the genotype imputation model, the methylation genotype detection system 106 can determine the posterior genotype probability 816 and haplotype detection 818 for genomic regions. Figure 8 illustrates that the methylation genotype detection system 106 uses the Genotype Probability Imputation and Phased Approach (GLIMPSE) as a genotype imputation model to determine the posterior genotype probability 816 at genomic coordinates.

[0123] Through imputation, the methylation genotype detection system 106 imputs one or more genotype detections of a target genome sample. As shown in Figure 8, for example, the methylation genotype detection system 106 determines the probability of an observed base 804 at genomic coordinates 800 from a target genome sample. The probability of an observed base 804 is expressed as where represents a given observed nucleobase and represents a base of a candidate haplotype under consideration. For example, the methylation genotype detection system 106 considers the observed nucleobases and tests these observed nucleobases against identified candidate haplotypes. The methylation genotype detection system 106 attempts to align the observed nucleobases with haplotypes in various ways to identify possible combinations. For example, if nucleotide read 802 contains only adenine or cytosine bases at genomic coordinate 800, the methylation genotype detection system 106 performs a test to determine whether the potential haplotype is an adenine or cytosine base.

[0124] In some specific embodiments, the methylation genotype detection system 106 determines the probability of the observed base 804 (and base detection quality metric 822) based on an estimated methylation level value 820. Each base detection quality metric in the base detection quality metric 822 indicates the probability that the observed base in the nucleotide read is an error. More specifically, the methylation genotype detection system 106 determines the observed base 804 based on a base detection quality metric (e.g., BASEQ score) of the observed base. The estimated methylation level value 820 indicates the probability that the observed base has been converted by methylation sequencing. Figure 9 and the corresponding paragraph provide additional details on how the methylation genotype detection system 106 determines the probability of the observed base 804 according to one or more embodiments.

[0125] As shown in FIG8, the methylation genotype detection system 106 generates a priori genotype probability 824 based on observed bases 804. Specifically, the methylation genotype detection system 106 may aggregate the probabilities of observed bases for all observed stacked bases to generate the probabilities of observed bases for candidate genotypes. The probabilities of observed bases for candidate genotypes can be expressed as , where represents a given observed nucleobase, and represents the bases of the candidate genotype under consideration. In some embodiments, the methylation genotype detection system 106 generates a priori genotype probability 824 based on observed bases for candidate genotypes. For example, in some embodiments, the methylation genotype detection system 106 generates the priori genotype probability 824 by multiplying the probabilities of observed bases to determine the corresponding candidate genotype.

[0126] As further indicated in FIG8, genomic coordinates 800 correspond to variable positions (or variable genomic coordinates) of haplotype reference set 806. In some cases, the methylation genotype detection system 106 further deconvolves the vector of probabilities of the observed base 804 into two independent vectors of haplotype allele probabilities (or, simply put, haplotype probabilities), where each vector corresponds to one of two complementary haplotypes.

[0127] Based on the haplotype probabilities from the independent vectors, in some specific implementations, the methylation genotype detection system 106 uses the haplotype version of the HMM to interpolate two target haplotypes as haplotype detection 818 during the iteration process. As shown in Figure 8, for example, the methylation genotype detection system 106 selects haplotype 810 based on the estimated haplotype reference set 806 and target haplotype 808 for each genome sample. After selecting haplotypes for a given genome sample, the methylation genotype detection system 106 stores the reference version and target version of the selected haplotype as a positional Brauth-Whiler transform (PBWT) 812. Instruction manual, pages 20 / 36, 25 CN 121359206 A

[0128] As further shown in Figure 8, in some embodiments, the methylation genotype detection system 106 performs a linear time sampling algorithm based on a haplotype interpolation version of the HMM developed by Na Li and Matthew Stephens in “Modeling Linkage Disequilibrium and Identifying Recombination Hotspots Using Single-Nucleotide Polymorphism Data,” 165 Genetics 2213-2233 (2003), sampling haplotype 814 in PBWT 812 format, the entire text of which is incorporated herein by reference.By performing a linear-time sampling algorithm as part of the sampler iteration, the methylation genotype detection system 106 further determines (and updates) the phase of two imputed haplotypes at genomic coordinates 800 of the target genome sample.

[0129] As further shown in FIG8, based on the imputed and phased haplotypes, the methylation genotype detection system 106 determines the posterior genotype probability 816 of genomic coordinates 800 of the genome sample exhibiting a specific genotype (e.g., a reference allele or a substitute allele). The methylation genotype detection system 106 further determines the genotype detection for genomic regions of each genome sample in the genome sample. For example, in some specific implementations, the methylation genotype detection system 106 generates the genotype detection based on determining a predicted combination of nucleobases corresponding to the highest posterior genotype probability. As indicated above, in some embodiments, the methylation genotype detection system 106 uses a modified version of GLIMPSE, developed by Rubinacci, S., Ribeiro, D.M., Hofmeister, R.J., et al., “Efficient phasing and imputation of low-coverage sequencing data using large reference panels,” Nat Genet 53, 120–126 (2021) (hereinafter referred to as Rubinacci), https: / / doi.org / 10.1038 / s41588-020-00756-0, which is incorporated herein by reference in its entirety. In other embodiments, the methylation genotype detection system 106 utilizes a modified HMM model to receive the estimated methylation level value 820 as input. For example, the methylation genotype detection system 106 may utilize a modified version of GLIPOSE as described by Rubinacci.

[0130] As indicated above, in some embodiments, the methylation genotype detection system 106 utilizes prior genotype probabilities as input to an HMM model. As further described below, the methylation genotype detection system 106 may (i) determine the probability of observed bases at genomic coordinates based on estimated methylation level values ​​and base detection quality metrics, and (ii) further determine prior genotype probabilities (as input to the HMM model) based on the corresponding probabilities of observed bases. Figure 9 illustrates a methylation genotype detection system 106 according to one or more embodiments of the present disclosure, which determines the probability of observed bases at genomic coordinates and the posterior genotype probability of a genomic sample at genomic coordinates in part based on estimated methylation level values ​​of observed bases.

[0131] As shown in FIG9, the methylation genotype detection system 106 determines the probability of a base observed given a haplotype base for each observed nucleobase at a given genomic coordinate. The probability of a base observed given a haplotype base can be expressed as , where represents the observed nucleobase and represents the haplotype base.

[0132] As shown in FIG9, the methylation genotype detection system 106 uses various equations to determine the probability of a base observed given a haplotype based on different observed base-haplotype combinations, as follows: Specification 21 / 36 pages 26 CN 121359206 A

[0133]

[0134] where represents the estimated positive strand methylation level value of the observed base on the positive strand at the genomic coordinate, represents the estimated negative strand methylation level value based on the observed base on the negative strand at the genomic coordinate, and represents the probability that the observed base is an error.

[0135] As further shown in Figure 9, in some embodiments, for observed base-haplotype base combinations not explicitly listed, the methylation genotype detection system 106 uses the following equation to determine the probability of the observed base:

[0136]

[0137] where represents the probability that the observed base is erroneous. In some embodiments, the methylation genotype detection system 106 determines this based on a base detection quality metric (e.g., BASEQ score). In one example, is equal to .

[0138] As previously mentioned, the methylation genotype detection system 106 determines the probability represented by for each observed nucleobase at a given genomic coordinate. Thus, the methylation genotype detection system 106 can determine several values ​​for the entire read stack. The methylation genotype detection system 106 further aggregates the values ​​into the probability of the observed base in the case of a genotype at a given genomic coordinate. For example, the probability of the observed base in the case of a given genotype can be expressed as , where represents the observed nucleobase, and represents the genotype at the genomic coordinate.

[0139] As discussed above with respect to Figure 8, the methylation genotype detection system 106 can input the probability represented by into the variant detector. In some specific embodiments, and as shown in Figure 9, the variant detector collapses the probability into the probability of the observed read stack, represented by . The methylation genotype detection system 106 can further utilize the variant detector to reverse the probability of the observed read stack to generate a posterior genotype probability, where D represents the data or the observed read stack.

[0140] The methylation genotype detection system 106 can utilize the posterior genotype probability in at least several ways.As indicated above, the methylation genotype detection system 106 can detect the genotype for a given genome coordinate by determining the highest posterior genotype probability from the posterior genotype probabilities. As described in step 27 CN 121359206 A on pages 22 / 36 of the specification below, such genotype detection outperforms existing sequencing and methylation detection systems and approaches the accuracy of unmethylated whole-genome sequencing. In addition to genotype detection, the methylation genotype detection system 106 can also determine a refined (and sometimes more accurate) methylation level value of candidate cytosine bases at the target genome coordinate based on the posterior genotype probability.

[0141] As just indicated, in some specific embodiments, the methylation genotype detection system 106 generates a refined methylation level value of cytosine bases at the genome coordinate based on the posterior genotype probability and the observed nucleobases at the genome coordinate. Figure 10 illustrates a methylation genotype detection system 106 for generating refined methylation level values ​​according to one or more embodiments of the present disclosure. In some embodiments, the methylation genotype detection system 106 can more accurately predict the methylation level values ​​of cytosine bases at genomic coordinates based on more accurate prediction of genotypes.

[0142] As shown in Figure 10, the methylation genotype detection system 106 can generate refined methylation level values ​​using the following equation ( ):

[0143]

[0144] where represents the count of methylated nucleobases at genomic coordinates, and represents the total number of nucleobases at genomic coordinates. In the above equation, represents the set of potential genotypes (i.e., CC, CA, CG, and CT) at genomic coordinates. In some specific embodiments, represents the estimated methylation level value representing data from read stacking at genomic coordinates.

[0145] As further illustrated in Figure 10, the methylation genotype detection system 106 uses different equations to generate refined methylation level values ​​based on observed or predicted genotypes. In some implementations, for example, for the predicted genotypes CC, CA, and CG, the methylation genotype detection system 106 generates a refined methylation level value using the following equation:

[0146]

[0147] where represents the count of methylated nucleobases at genomic coordinates and represents the total number of nucleobases observed at genomic coordinates.In the above equation, and represent the tunable parameters used in the following β prior function for estimating methylation level values:

[0148]

[0149] In contrast, for predicted genotype CT, the methylation genotype detection system 106 can generate refined methylation level values ​​using the following equation:

[0150]

[0151] where represents the count of methylated nuclei at genomic coordinates, and represents the total number of observed nuclei at genomic coordinates. In the above equation, and represent the tunable parameters used in the above β prior function. Additionally, and as shown in FIG10, represents the hypergeometric function. Specification 23 / 36 pages 28 CN 121359206 A

[0152] As mentioned, in some examples, the refined methylation level value ( ) is more accurate than the estimated methylation level value ( ). Specifically, the methylation genotype detection system 106 improves the accuracy of refined methylation level values ​​by calculating estimates of different methylation levels and weighting these estimates based on posterior genotype probabilities. For example, for a given genomic coordinate, only one candidate genotype in a subset of candidate genotypes (CC, CA, CG, or CT) is accurate. Based on the posterior genotype probability indicating a strong likelihood of the indicator genotype CC, for example, the methylation genotype detection system 106 may use collapsed whole expression and reweight the final refined methylation level values. For example, in some implementations, the methylation genotype detection system 106 may determine the refined methylation level value based on an equation specific to that given genotype, based on a posterior genotype probability indicating a given genotype is above a threshold. If the genotype is uncertain based on the posterior genotype probability, the methylation genotype detection system 106 may also allow mixed weighting of expressions.

[0153] As mentioned, the methylation genotype detection system 106 improves the accuracy of both methylation level estimation and genotype detection compared to existing sequencing and methylation detection systems. Figures 11A to 13C illustrate various diagrams and graphs depicting such accuracy improvements made by the methylation genotype detection system 106.

[0154] Figures 11A to 11B illustrate diagrams demonstrating the improvements made by the methylation genotype detection system 106 in accurately predicting methylation level values. More specifically, Figure 11A illustrates how the methylation genotype detection system 106 accurately estimates the methylation level values ​​of cytosine bases on the positive and negative strands of phylogenetic alleles. Figure 11B illustrates how the methylation genotype detection system 106 accurately estimates the methylation level values ​​of cytosine bases on the positive strand with a 10% C allele frequency.Although the diagrams in Figures 11A and 11B depict results from simulated data, where observed nucleotide reads are simulated for a given genotype, the methylation genotype detection system 106 will necessarily generate more accurate estimated methylation level values ​​in the diagrams of Figures 11A and 11B using the equations described above to determine refined methylation level values. Therefore, the type of estimated methylation level values ​​depicted in the diagrams of Figures 11A and 11B is refined methylation level values.

[0155] Specifically, Figure 11A illustrates diagram 1102 depicting estimated methylation level values ​​generated by the methylation genotype detection system 106 compared to true methylation level values. More specifically, diagram 1102 reflects the accuracy of the estimated methylation level values ​​generated by the methylation genotype detection system 106 for cytosine bases at genomic coordinates on both the positive and negative strands of germline alleles. As shown in Figure 11A, for the depicted genotypes CA, CC, CG, and CT on the positive strand and CG, GA, GG, and GT on the negative strand, the estimated methylation level values ​​closely track the true methylation level values ​​illustrated by the dashed lines. For example, as discussed above with respect to Figure 6, this contrasts with existing systems that often inaccurately estimate the methylation level values ​​for the CA, CT, and CG genotypes. However, the methylation genotype detection system 106 accurately estimates the methylation level values ​​for these same genotypes, as shown in Figure 11A.

[0156] Figure 11B illustrates a schematic diagram 1104 depicting the estimated methylation level values ​​generated by the methylation genotype detection system 106 compared to the true methylation level values ​​from somatic alleles. More specifically, schematic diagram 1104 depicts the accuracy of the methylation level values ​​generated by the methylation genotype detection system 106 for genomic samples with somatic variants having a 10% variant allele frequency (VAF) of cytosine alleles. As shown, the estimated methylation level values ​​generated by the methylation genotype detection system 106 closely track the actual methylation level values ​​indicated by the dashed lines. Although the diagrams illustrated in Figure 11B corresponding to the CA, CT, and CG genotypes depict a wider range of variations in the results, the estimated methylation level values ​​generated by the methylation genotype detection system 106 are more accurate than those predicted by existing sequencing and methylation detection systems described above. In fact, many existing sequencing and methylation detection systems that rely on SBS technology are completely unable to predict methylation levels from somatic alleles. Specification 24 / 36 pages 29 CN 121359206 A

[0157] In addition to generating more accurate methylation level values, the methylation genotype detection system 106 also detects single nucleotide polymorphisms (SNPs) more accurately.Figures 12A and 12B illustrate a series of curves according to one or more embodiments of the present disclosure, demonstrating that the methylation genotype detection system 106 detects SNPs more accurately than existing sequencing and methylation detection systems.

[0158] Figures 12A and 12B illustrate curves mapping the SNP precision and SNP recall of various variant detector (VC) algorithms. For example, the curves map results from the current VC algorithm, the methylation-aware VC algorithm (methylation-aware), and the methylation-masking VC algorithm (methylation-masking). The current VC algorithm data represents the basic variant detector utilized by existing methylation and detection systems, such as the model used in Figure 5. The methylation-aware VC algorithm data represents the results generated by the methylation genotype detection system 106 according to one or more embodiments of the present disclosure. The methylation-masking VC algorithm data represents the results generated by a variant detector that is more complex and accurate than the "current" VC. For example, the methylation-masking VC algorithm can improve SNP precision and SNP recall relative to the current VC algorithm by taking into account methylation transformation. More specifically, the methylation-masking VC algorithm reduces the prior genotype probability of the variants detected corresponding to the transformed nucleobases determined by methylation sequencing. For example, the methylation-masking VC algorithm assigns lower quality scores (e.g., Q scores of 0) to thymine bases on the positive strand or adenine bases on the negative strand.

[0159] As shown in Figure 12A, the methylation genotype detection system 106 improves SNP precision and recall compared to existing sequencing and methylation detection systems. More specifically, Figure 12A illustrates two schemas and the enriched regions of these schemas. These schemas show the results of the EM-Seq protocol and the whole-genome sequencing (WGS) protocol. The results of the methylation genotype detection system 106 are shown by lines 1218a to 1218b on the curves illustrated in Figure 12A. Lines 1206a to 1206b reflect the performance of the current VC algorithm; lines 1216a to 1216b reflect the performance of the methylation genotype detection system 106, which can also be referred to as the methylation-sensing VC algorithm; and lines 1218a to 1218b reflect the performance of the methylation-masking VC algorithm.

[0160] For the EM-Seq scheme, and as shown by line 1206a, the current VC algorithm corresponds to a decrease in both precision and recall. Both the methylation genotype detection system 106 and the methylation-masking VC algorithm show improved performance relative to existing sequencing and methylation detection systems in the EM-Seq scheme. For example, as shown by lines 1218a to 1218b, the methylation genotype detection system 106 performs similarly to the WGS algorithm in terms of both SNP precision and recall.

[0161] The methylation genotype detection system 106 can also accurately predict single nucleotide variants (SNVs) in somatic variant detection.Figure 12B illustrates the accurate detection of SNVs by the methylation genotype detection system 106 in somatic variant detection. As illustrated in Figure 12B, the methylation genotype detection system 106 is one of the first known systems for the accurate detection of somatic variants from short-read methylation sequencing data. Lines 1222a and 1222b represent the somatic variant detection performance of the methylation genotype detection system 106 (methylsensing) on ​​different samples NA12877 and NA12878, where 10% of NA12877 was added to NA12878 at TruSight One (TSO) enriched genomic regions with 300x coverage. In contrast, lines 1224a and 1224b represent the detection performance of the current VC algorithm for somatic variants of NA12877 and NA12878 in different samples, where 10% of NA12877 were also added to the TSO-enriched genomic regions of NA12878 with a 300x coverage. Further, lines 1220a and 1220b represent the detection performance of the methylation-masked VC algorithm for somatic variants of NA12877 and NA12878 in different samples, where 10% of NA12877 were also added to the TSO-enriched genomic regions of NA12878 with a 300x coverage. As shown in both Figures 12A and 12B, the methylation genotype detection system 106 performs with higher SNV precision and recall compared to the methylation masking method. The methylation genotype detection system 106 also matches the current VC method used in the WGS protocol.

[0162] As previously discussed, the methylation genotype detection system 106 can detect genotypes more accurately than existing sequencing and methylation detection systems that perform C>T conversion-based sequencing. Figures 13A to 13C illustrate a series of graphs according to one or more embodiments of the present disclosure, demonstrating that the methylation genotype detection system 106 accurately detects the genotype at the coordinates of Group A on pages 25 / 36 of the gene specification 30 CN 121359206. Figures 13A to 13C illustrate the accuracy of genotype detection by the methylation genotype detection system 106 (methylation sensing) under different conditions relative to the methylation-masked VC scheme and the current VC scheme. For example, Figure 13A illustrates the effect of high coverage on a genomic region with 100x read coverage, Figure 13B illustrates the effect of low coverage on a genomic region with 10x read coverage, and Figure 13C illustrates the effect of low data quality.

[0163] Figure 13A shows the effect of high coverage (100x) on verification. As just suggested, coverage describes the average number of reads that align to or cover a known reference base.Figure 1302 shows the validation results of the current VC protocol, Figure 1304 shows the validation results of the methylation genotype detection system 106, and Figure 1306 shows the validation results of the methylation-masked VC protocol. As shown in Figure 13A, both the methylation genotype detection system 106 and the methylation-masked VC protocol detect genotypes more accurately than the current VC protocol. The current VC protocol refers to existing sequencing and methylation detection systems. For example, the current VC protocol often misdetects genotypes where the true genotype is CA, CC, or CG.

[0164] Figure 13B shows the impact of low coverage (10x) on validation. Figure 1308 shows the validation results of the current VC protocol, Figure 1310 shows the validation results of the methylation genotype detection system 106, and Figure 1312 shows the validation results of the methylation-masked VC protocol. As shown in Figure 13B, the current VC protocol is negatively affected and misdetects more genotypes with low coverage. The methylation genotype detection system 106 detects the CA and CG genotypes more accurately than the methylation-masked VC protocol, as shown in boxes 1320 and 1322, respectively. Furthermore, as shown in box 1324, although the methylation genotype detection system 106 may falsely detect the CT genotype, the number of false bases detected is less than those falsely detected by the methylation-masked VC protocol. Therefore, the methylation genotype detection system 106 outperforms existing systems in validation, even at low coverage.

[0165] Figure 13C illustrates the impact of low-quality data on validation. Graph 1314 shows the validation results of the current VC protocol, Graph 1316 shows the validation results of the methylation genotype detection system 106, and Graph 1318 shows the validation results of the methylation-masked VC protocol. While Figures 13A and 13B show results corresponding to data with a BASEQ score of 40, the data in Figure 13C has a BASEQ score of 10. For low-quality data, and as highlighted in 1326 and 1328, the methylation genotype detection system 106 detects the CA and CG genotypes more accurately than the methylation-masked VC scheme. As highlighted in 1330, the methylation genotype detection system 106 tends to misdetect the true CT genotype as CC. However, overall, the methylation genotype detection system 106 detects most genotypes more accurately than current VC systems and methylation-masked VC systems.

[0166] Figures 1 through 13C, the corresponding text, and examples provide a number of different methods, systems, apparatuses, and nontransitory computer-readable media for the methylation genotype detection system 106. In addition to the foregoing, one or more specific embodiments may be described according to flowcharts including actions for achieving a particular result, as shown in Figure 14. Figure 14 illustrates a flowchart of a series of actions 1400 for generating genotype detection and estimating methylation level values ​​according to one or more embodiments of the present disclosure.While Figure 14 illustrates actions according to one embodiment, alternative embodiments may omit, add, reorder, and / or modify any of the actions shown in Figure 14. The actions of Figure 14 may be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device or system to perform the actions depicted in Figure 14. In yet another embodiment, the system includes at least one processor and a non-transitory computer-readable medium that includes instructions that, when executed by one or more processors, cause the system to perform the actions of Figure 14.

[0167] As shown in Figure 14, a series of actions 1400 includes actions 1402 for identifying nucleotide reads, actions 1404 for determining estimated methylation level values, actions 1406 for generating posterior genotype probabilities, and actions 1408 for generating genotype detection. For example, a series of actions 1400 may include actions for performing any of the operations described in the following clauses:

[0168] Clause 1. A method comprising: Specification 26 / 36 Page 31 CN 121359206 A

[0169] For a target genome sample, identifying a nucleotide read containing one or more nucleobases converted by methylation sequencing;

[0170] Determining an estimated methylation level value of cytosine bases at the genomic coordinates based on a prior genotype probability of the target genome sample at genomic coordinates and observed nucleobases at the genomic coordinates within the nucleotide read;

[0171] Generating a posterior genotype probability of the target genome sample at the genomic coordinates based on the estimated methylation level value of the observed nucleobases and a base detection quality metric using a variant detection model; and

[0172] Genotype detection of the target genome sample containing a predicted combination of nucleobases at the genomic coordinates based on the posterior genotype probability.

[0173] Clause 2. The method of Clause 1, further comprising generating a refined methylation level value for the cytosine base at the genomic coordinate based on the posterior genotype probability and the observed nucleotide.

[0174] Clause 3. The method of Clause 2, further comprising generating the refined methylation level value by:

[0175] determining a genotype-specific methylation level value corresponding to each possible genotype at the genomic coordinate based on the observed nucleotide; and

[0176] weighting the genotype-specific methylation level value based on the posterior genotype probability.

[0177] Clause 4. The method of Clause 1, wherein generating the genotype detection of the target genome sample is further based on sequencing metrics corresponding to the nucleotide read.

[0178] Clause 5. The method of Clause 1, wherein the posterior genotype probability comprises a subset of the posterior genotype probabilities of cytosine bases in the positive or negative strand.

[0179] Clause 6. The method of Clause 1, further comprising determining the estimated methylation level value by:

[0180] determining the probability of each nucleobase at the genomic coordinate on the positive and negative strands;

[0181] determining the probability of the observed nucleobase at the genomic coordinate based on the number of each observed nucleobase, the probability of each nucleobase at the genomic coordinate, and the prior genotype probability; and

[0182] generating an estimated positive-strand methylation level value for the nucleobase at the genomic coordinate on the positive strand and an estimated negative-strand methylation level value for the nucleobase at the genomic coordinate on the negative strand by performing a Bayesian inversion on the probability of the observed nucleobase.

[0183] Clause 7. The method of Clause 6, further comprising determining the probability of each nucleobase by:

[0184] determining the probability of a given thymine base on the positive strand approximately equal to the estimated methylation level value; and

[0185] determining the probability of a given cytosine base on the positive strand approximately equal to 1 minus the estimated methylation level value.

[0186] Clause 8. The method of Clause 6, further comprising determining the probability of each nucleobase by:

[0187] determining the probability of a given adenine base on the negative strand approximately equal to the estimated methylation level value; and

[0188] determining the probability of a given guanine base on the negative strand approximately equal to 1 minus the estimated methylation level value.

[0189] Clause 9. The method of Clause 1, wherein the variant detection model comprises a Hidden Markov Model (HMM), the HMM being modified to receive input of a base detection quality metric based on the estimated methylation level value and the corresponding observed nucleobase. Specification 27 / 36 pages 32 CN 121359206 A

[0190] Clause 10. The method of Clause 1, further comprising generating the genotype detection based on a predicted combination of nucleobases corresponding to the highest posterior genotype probability.

[0191] Clause 11. The method of Clause 1, wherein identifying the nucleotide read comprising the one or more nucleobases converted by the methylation sequencing comprises identifying the nucleotide read comprising a thymine base or a uracil base converted from a cytosine base by the methylation sequencing.

[0192] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. The particularly suitable techniques are those in which nucleic acids are attached to fixed positions in the array so that their relative positions do not change and in which the array is repeatedly imaged.Implementations that obtain images in different color channels (e.g., matching different markers used to distinguish one nucleobase type from another nucleobase type) are particularly suitable. In some implementations, the process of determining the nucleotide sequence of the target nucleic acid (i.e., nucleic acid polymer) can be automated. Preferred implementations include sequencing-by-synthesis (SBS) technology.

[0193] SBS technology typically involves the enzymatic extension of a nascent nucleic acid chain by repeatedly adding nucleotides to the template chain. In conventional SBS methods, a single nucleotide monomer can be provided to the target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to the target nucleic acid in the presence of a polymerase in delivery.

[0194] SBS can utilize nucleotide monomers having a terminator motif or nucleotide monomers lacking any terminator motif. Methods utilizing nucleotide monomers lacking a terminator include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in further detail below. In methods using nucleotide monomers lacking a terminator, the number of nucleotides added in each cycle is typically variable and depends on the template sequence and the manner of nucleotide delivery. For SBS technology utilizing nucleotide monomers with a terminator motif, the terminator may be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or the terminator may be reversible, as in sequencing methods developed by Solexa (now Illumina, Inc.).

[0195] SBS technology can utilize nucleotide monomers with or without a labeled motif. Therefore, incorporation events can be detected based on: the characteristics of the label, such as the fluorescence of the label; the characteristics of the nucleotide monomer, such as molecular weight or charge; byproducts of the incorporated nucleotide, such as the release of pyrophosphate; and so on. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or alternatively, the two or more different labels may be indistinguishable under the detection technology used. For example, the different nucleotides present in the sequencing reagent may have different labels, and they may be distinguishable using appropriate optics, as exemplified by sequencing methods developed by Solexa (now Illumina, Inc.).

[0196] Preferred embodiments include pyrosequencing technology.Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001), “Pyrosequencing sheds light on DNA sequencing.”, Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998), “A sequencing method based on real-time pyrophosphate.”, Science 281(5375), 363; US Patent No. 6,210,891; US ​​Patent No. 6, The disclosures of U.S. Patent Nos. 258,568 and 6,274,320, the entire contents of which are incorporated herein by reference. In pyrosequencing, the released PPi can be detected by being immediately converted to ATP by adenosine triphosphate (ATP) sulfatase, and the level of ATP produced can be detected by photons generated by luciferase. The nucleic acid to be sequenced can be attached to a feature in an array, and the array can be imaged to capture the chemiluminescent signal generated due to the incorporation of nucleotides at the feature in the array. Images can be obtained after processing the array with a specific type of nucleotide (e.g., A, T, C, or G). The images obtained after adding each type of nucleotide will differ in which feature in the array is detected. These differences in the images reflect the different sequence contents of the feature on the array. However, the relative position of each feature will remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein. For example, the images obtained after processing the array with each different nucleotide type can be processed in the same way as illustrated in this paper for images obtained from different detection channels used in reversible terminator-based sequencing methods.

[0197] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, cleavable or photobleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method was commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, the disclosures of each of which are incorporated herein by reference. The availability of fluorescently labeled terminators (where the termination can be reversible and the fluorescent label can be cleaved) facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.

[0198] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection markers can be removable, for example, by cleavage or degradation. Images can be captured after the markers are incorporated into the arrayed nucleic acid signatures. In a specific embodiment, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally distinct marker. Four images are then obtained, each using a detection channel selective for one of the four different markers. Alternatively, different nucleotide types can be added sequentially, and images of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid signature with the specific type of nucleotide incorporated. Different signatures may or may not be present in different images due to the different sequence contents of each signature. However, the relative positions of the signatures will remain unchanged in the images. Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the markers can be removed, and the reversible terminator portion can be removed for subsequent cycles of nucleotide addition and detection. Removing these markers after they have been detected in a particular cycle and before subsequent cycles provides the advantage of reducing background signal and crosstalk between cycles. Examples of available marker and removal methods are illustrated below.

[0199] In a particular embodiment, some or all of the nucleotide monomers may include a reversible terminator. In such an embodiment, the reversible terminator / cleavable fluorophore may include a fluorophore linked to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15: 1767-1776 (2005), which is incorporated herein by reference).Other methods have separated terminator chemistry from the cleavage of the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), the full text of which is incorporated herein by reference). Ruparel et al. described the development of reversible terminators that use small 3'-allyl groups to block elongation, but can be easily deblocked by short-time treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker that can be easily cleaved by exposure to long-wavelength ultraviolet light for 30 seconds. Therefore, disulfide reduction or photocleavage can be used as the cleavable linker. Another method of reversible termination is the use of natural termination, which occurs after the placement of a bulk dye on a dNTP. The presence of a charged bulk dye on the dNTP can act as an efficient terminator due to steric and / or electrostatic hindrance. The presence of an incorporation event prevents further incorporation unless the dye is removed. The cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entirety.

[0200] Additional exemplary SBS systems and methods that may be used in conjunction with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entirety.

[0201] Some embodiments may use fewer than four different labels to detect four different nucleotides. For example, the methods and systems described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232 may be used to perform SBS. As a first example, a pair of nucleotide types may be detected at the same wavelength, but distinguished based on the intensity difference of one member relative to the other member, or based on a change in one member of the pair that results in a significant appearance or disappearance of a signal compared to the detected signal of the other member of the pair (e.g., by chemical modification, photochemical modification, or physical modification).As a second example, three of the four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type lacks a marker that is detectable under those conditions or is minimally detectable under those conditions (e.g., minimal detection due to background fluorescence, etc.). The incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their respective signals, and the incorporation of the fourth nucleotide type can be determined based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type may include a marker detected in two different channels, while other nucleotide types are detected in no more than one channel. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations. An exemplary implementation combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type lacking a label or minimally detected in either channel (e.g., unlabeled dGTP).

[0202] Additionally, as described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, the first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.

[0203] Some embodiments may utilize ligation-based sequencing technology. Such technologies utilize DNA ligases to incorporate oligonucleotides and recognize the incorporation of such oligonucleotides. Oligonucleotides typically have distinct labels associated with the identity of specific nucleotides in the sequence to which they hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show nucleic acid features that have been incorporated with a specific type of label. Due to the different sequence contents of each feature, different features may or may not be present in different images, but the relative positioning of the features will remain unchanged within the images. Images obtained by ligation-based sequencing methods can be stored, processed, and analyzed as described herein.Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.

[0204] Some embodiments may utilize nanopore sequencing (Deamer, DW and Akeson, M., “Nanopores and nucleic acids: prospects for ultrarapid sequencing.”, Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis.”, Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin and J. A. Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope”, Nat. Mater., 2:611-615 (2003), the full text of which is incorporated herein by reference). In such embodiments, the target nucleic acid passes through a nanopore. Nanopores can be synthetic pores or biomembrane proteins, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the pore's conductivity.(The full text of the publications of these documents are incorporated herein by reference: U.S. Patent Nos. 7,001,792; Soni, GV, and Meller, “A. Progress toward ultrafast DNA sequencing using solid-state nanopores.”, Clin. Chem. 53, 1996–2001 (2007); Healy, K., “Nanopore-based single-molecule DNA analysis.”, Nanomed., 2, 459–481 (2007); Cockroft, SL, Chu, J., Amorin, M., and Ghadiri, M. R., “A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.”, J. Am. Chem. Soc. 130, 818–820 (2008).) Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed like images, based on the exemplary processing of optical images and other images described herein.

[0205] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interaction between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414 (each of which is incorporated herein by reference), or by using a zero-mode waveguide, as described, for example, in U.S. Patent No. 7,315,019 (which is incorporated herein by reference), and by using fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference).Illumination can be limited to a volume of approximately 1 / 2000 that surrounds the surface-tethered polymerase, allowing for observation of the incorporation of fluorescently labeled nucleotides against a low background (Levene, M. J. et al., “Zero-mode vegetatives for single-molecule analysis at high concentrations.”, Science 299, 682-686 (2003); Lundquist, PM et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures.”, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008)). (The full text of these publications is incorporated herein by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.

[0206] Some SBS implementations include the detection of protons released during nucleotide incorporation into the extension product. For example, sequencing based on the detection of released protons can use electrical detectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies) or the sequencing methods and systems described in US 2009 / 0026082 A1, US 2009 / 0127589 A1, US 2010 / 0137143 A1, or US 2010 / 0282617 A1, each of which is incorporated herein by reference. The methods described herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates for proton detection. More specifically, the methods described herein can be used to generate a population of amplicon clones for proton detection.

[0207] The SBS method described above can be advantageously performed in a variety of formats, enabling the simultaneous manipulation of multiple different target nucleic acids. In specific embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a specific substrate.This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in multiple ways, as described on pages 31 / 36 of the specification (CN 121359206 A). In embodiments using surface-bound target nucleic acids, the target nucleic acids may be in an array format. In an array format, the target nucleic acid can typically bind to the surface in a spatially distinguishable manner. The target nucleic acid can bind by direct covalent attachment, attachment to beads or other particles, or binding to polymerases or other molecules attached to the surface. The array may include a single copy of the target nucleic acid at each site (also referred to as a feature region), or multiple copies having the same sequence may be present at each site or feature region. Multiple copies may be generated by amplification methods such as, for example, bridging amplification or emulsion PCR, which are described in further detail below.

[0208] The methods described herein may use arrays having feature portions at any of a variety of densities, including, for example, at least about 10 feature portions / cm², 100 feature portions / cm², 500 feature portions / cm², 1,000 feature portions / cm², 5,000 feature portions / cm², 10,000 feature portions / cm², 50,000 feature portions / cm², 100,000 feature portions / cm², 1,000,000 feature portions / cm², 5,000,000 feature portions / cm², or higher.

[0209] An advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Therefore, this disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Therefore, the integrated system of this disclosure may include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc. A flow cell in the integrated system may be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference. As illustrated with respect to a flow cell, one or more fluid components of the integrated system may be used for amplification and detection methods. Taking a nucleic acid sequencing implementation as an example, one or more fluid components of the integrated system may be used for the amplification methods set forth herein and for delivering sequencing reagents in sequencing methods (such as those illustrated above). Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of generating amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Serial No. 13 / 273,666, which is incorporated herein by reference.

[0210] The sequencing system described above sequences nucleic acid polymers present in a sample received by the sequencing equipment. As defined herein, “sample” and its derivatives are used in their broadest sense to include any specimen, culture, etc., suspected of containing a target. In some embodiments, a sample includes nucleic acids in the form of DNA, RNA, PNA, LNA, chimeric or hybrid forms. A sample may include any biological, clinical, surgical, agricultural, atmospheric or aquatic plant or animal specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh frozen or formalin-fixed paraffin-embedded nucleic acid specimens. It is also contemplated that a sample may be derived from: a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched to) a single individual (such as tumor samples and normal tissue samples), or a sample from a single source containing two different forms of genetic material (such as maternal DNA and fetal DNA obtained from a maternal subject), or a sample containing contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of nucleic acid material may include nucleic acids obtained from newborns, such as nucleic acids commonly used for newborn screening.

[0211] Nucleic acid samples may include high molecular weight substances, such as genomic DNA (gDNA). Samples may include low molecular weight substances, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another embodiment, the low molecular weight substance includes enzymatically fragmented or mechanically fragmented DNA. Samples may include cell-free circulating DNA. In some embodiments, samples may include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microscopy, surgical resection, and other clinical or laboratory samples. In some embodiments, samples may be epidemiological samples, agricultural samples, forensic samples, or pathogenic samples. In some embodiments, samples may include nucleic acid molecules obtained from animals (such as human or mammalian sources). In another embodiment, samples may include nucleic acid molecules obtained from non-mammalian sources (such as plants, bacteria, viruses, or fungi). In some embodiments, the source of nucleic acid molecules may be archived or extinct samples or species.

[0212] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples containing low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or may include forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel.Nucleic acid samples can be purified samples or lysed products containing crude DNA, such as those derived from oral swabs, paper, fabric, or other substrates that can be impregnated with saliva, blood, or other bodily fluids. Therefore, in some embodiments, the nucleic acid sample may include a small amount of DNA (such as genomic DNA) or fragmented portions of DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from a victim's hair, skin, tissue samples, autopsy, or remains. In some embodiments, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence may include nucleic acids obtained from non-human DNA (such as microbial, plant, or insect DNA). In some embodiments, the target sequence or amplified target sequence is directed for human identification purposes. In some embodiments, this disclosure relates in its entirety to methods for identifying the characteristics of forensic samples. In some embodiments, this disclosure relates in its entirety to methods for human identification using one or more target-specific primers disclosed herein or designed with one or more target-specific primers designed using the primer design standards outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence may be amplified using any one or more target-specific primers disclosed herein or using the primer standards outlined herein.

[0213] Components of the methylation genotype detection system 106 may include software, hardware, or both. For example, components of the methylation genotype detection system 106 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., user client device 110). When executed by one or more processors, the computer-executable instructions of the methylation genotype detection system 106 may cause the computing device to perform the bubble detection method described herein. Alternatively, components of the methylation genotype detection system 106 may include hardware, such as a dedicated processing device for performing a particular function or set of functions. Additionally or alternatively, components of the methylation genotype detection system 106 may include a combination of computer-executable instructions and hardware.

[0214] Furthermore, components of the methylation genotype detection system 106, which perform the functions described herein with respect to the methylation genotype detection system 106, may be implemented, for example, as part of a standalone application, as a module of an application, as a plugin of an application, as one or more library functions that can be detected by other applications, and / or as a cloud computing model. Thus, components of the methylation genotype detection system 106 may be implemented as part of a standalone application on a personal computing device or mobile device.Additionally or alternatively, components of the methylation genotype detection system 106 may be implemented in any application providing sequencing services, including but not limited to Illumina BaseSpace, Illumina DRAGEN, Illumina NextSeq, Illumina TruSeq, or Illumina TruSight software. "Illumina", "BaseSpace", "DRAGEN", "NextSeq", "TruSeq", and "TruSight" are registered trademarks or trademarks of Illumina Corporation in the U.S. and / or other countries.

[0215] As discussed in more detail below, embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware such as, for example, one or more processors and system memory. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specific Description 33 / 36 pages 38 CN 121359206 A One or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any media content access device described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more processes described herein.

[0216] A computer-readable medium may be any available medium accessible by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of this disclosure may include at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0217] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSDs) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that is accessible by a general-purpose or special-purpose computer.

[0218] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices.When information is transferred or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately considers that connection as a transmission medium. The transmission medium may include a network and / or a data link that can be used to carry desired program code means in the form of computer-executable instructions or data structures, and which is accessible by a general-purpose or special-purpose computer. The combinations described above should also be included within the scope of computer-readable media.

[0219] Furthermore, upon arrival at various computer system components, program code means in the form of computer-executable instructions or data structures may be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link may be buffered in RAM within a network interface module (e.g., a NIC) and then ultimately transferred to computer system RAM and / or to a less transient computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) may be included in computer system components that also (or even primarily) utilize the transmission medium.

[0220] Computer-executable instructions include, for example, instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a function or group of functions. In some embodiments, the computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary numbers, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or methodological actions, it should be understood that the subject matter as defined in the appended claims is not necessarily limited to the described features or actions. Rather, the described features and actions are disclosed as exemplary forms of implementing the claims.

[0221] Those skilled in the art will understand that this disclosure can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, portable computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. This disclosure can also be practiced in a distributed system environment, wherein both local and remote computer systems perform tasks via network links (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may reside in both local and remote memory storage devices.

[0222] Embodiments of this disclosure can also be implemented in a cloud computing environment.In this specification, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then thus scaled.

[0223] The cloud computing model can consist of various features such as, for example, on-demand self-service, widespread network access, resource pooling, rapid elasticity, metered services, etc. The cloud computing model can also exhibit various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). The cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and in the claims, “cloud computing environment” is an environment in which cloud computing is employed.

[0224] Figure 15 illustrates a block diagram of a computing device 1500 that can be configured to perform one or more of the processes described above. It will be understood that one or more computing devices, such as computing device 1500, can implement the methylation genotype detection system 106. As shown in FIG15, computing device 1500 may include a processor 1502, memory 1504, storage device 1506, I / O interface 1508, and communication interface 1510, which are communicatively coupled via communication infrastructure 1512. In some embodiments, computing device 1500 may include fewer or more components than those shown in FIG15. The following paragraphs describe the components of computing device 1500 shown in FIG15 in more detail.

[0225] In one or more embodiments, processor 1502 includes hardware for executing instructions, such as those constituting a computer program. By way of example and not limitation, in order to execute instructions for dynamically modifying workflows, processor 1502 may retrieve (or obtain) instructions from internal registers, internal caches, memory 1504, or storage device 1506, and decode and execute these instructions. Memory 1504 may be volatile or non-volatile memory for storing data, metadata, and programs for execution by a processor. Storage device 1506 includes storage means for storing data or instructions for performing the methods described herein, such as a hard disk, flash drive, or other digital storage device.

[0226] I / O interface 1508 allows a user to provide input to computing device 1500, receive output from computing device, and otherwise transfer data to and receive data from computing device. I / O interface 1508 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O interfaces.I / O interface 1508 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O interface 1508 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be servicing a particular implementation.

[0227] Communication interface 1510 may include hardware, software, or both. In any event, communication interface 1510 may provide one or more interfaces for communication (such as, for example, packet-based communication) between computing device 1500 and one or more other computing devices or networks. By way of example and not limitation, communication interface 1510 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (such as Wi-Fi).

[0228] Additionally, communication interface 1510 may facilitate communication with various types of wired or wireless networks. Communication interface 1510 can also facilitate communication using various communication protocols. Communication infrastructure 1512 may also include hardware, software, or both that couple components of computing device 1500 to each other. For example, communication interface 1510 may use one or more networks and / or protocols (see page 35 / 36 of CN 121359206 A) to enable multiple computing devices connected through a particular infrastructure to communicate with each other to perform one or more aspects of the process described herein. For illustration, the sequencing process may allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.

[0229] In the foregoing description, this disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of this disclosure have been described with reference to the details discussed herein, and various embodiments are illustrated in the accompanying drawings. The above description and figures are illustrative of this disclosure and should not be construed as limiting the disclosure. Numerous specific details have been described to provide a thorough understanding of various embodiments of this disclosure.

[0230] This disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The embodiments described herein should be considered exemplary in all respects and not limiting. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other or in parallel with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description.All changes within the equivalent meaning and scope of the claims shall be included within its scope. Instruction manual page 36 / 36, page 41, CN 121359206 A, Figure 1; Instruction manual figure 1 / 21, page 42, CN 121359206 A, Figure 2A; Instruction manual figure 2 / 21, page 43, CN 121359206 A, Figure 2B; Instruction manual figure 3 / 21, page 44, CN 121359206 A, Figure 3; Instruction manual figure 4 / 21, page 45, CN 121359206 A, Figure 4; Instruction manual figure 5 / 21, page 46, CN 121359206 A, Figure 5; Instruction manual figure 6 / 21, page 47, CN 121359206 A, Figure 6; Instruction manual figure 7 / 21, page 48, CN 121359206 A, Figure 7A; Instruction manual figure 8 / 21, page 49, CN 121359206 A, Figure 7B; Instruction manual figure 9 / 21, page 50, CN 121359206 A, Figure 8. Instruction manual illustrations, page 10 / 21, figure 51, CN 121359206 A, figure 9; page 11 / 21, figure 52, CN 121359206 A, figure 10; page 12 / 21, figure 53, CN 121359206 A, figure 11A; page 13 / 21, figure 54, CN 121359206 A, figure 11B; page 14 / 21, figure 55, CN 121359206 A, figure 12A; page 15 / 21, figure 56, CN 121359206 A, figure 12B; page 16 / 21, figure 57, CN 121359206 A, figure 13A; page 17 / 21, figure 58, CN 121359206 A, figure 13B; page 18 / 21, figure 59, CN 121359206 A, figure 13C. Figure 14 of the instruction manual, page 19 / 21, 60 CN 121359206 A; Figure 15 of the instruction manual, page 20 / 21, 61 CN 121359206 A; Figure 16 of the instruction manual, page 21 / 21, 62 CN 121359206 A.

Claims

1. A method comprising: identifying, for a target genomic sample, nucleotide reads comprising one or more nucleobases transformed by a methylation sequencing assay; determining, based on a prior genotype probability of the target genomic sample at a genomic coordinate and an observed nucleobase at the genomic coordinate within the nucleotide reads, an estimated methylation level value of a cytosine base at the genomic coordinate; generating, using a variant call model, a posterior genotype probability of the target genomic sample at the genomic coordinate based on the estimated methylation level value of the observed nucleobase and a base call quality metric; and generating, based on the posterior genotype probability, a genotype call of the target genomic sample comprising a predicted combination of nucleobases at the genomic coordinate.

2. The method of claim 1, further comprising generating a refined methylation level value of the cytosine base at the genomic coordinate based on the posterior genotype probability and the observed nucleobase.

3. The method of claim 2, further comprising generating the refined methylation level value by: determining genotype-specific methylation level values corresponding to each possible genotype at the genomic coordinate based on the observed nucleobase; and weighting the genotype-specific methylation level values based on the posterior genotype probability.

4. The method of claim 1, wherein generating the genotype call of the target genomic sample is further based on a sequencing metric corresponding to the nucleotide reads.

5. The method of claim 1, wherein the posterior genotype probability comprises a subset of posterior genotype probabilities of cytosine bases in a positive strand or a negative strand.

6. The method of claim 1, further comprising determining the estimated methylation level value by: determining a probability of each nucleobase at the genomic coordinate on a positive strand and a negative strand; determining a probability of the observed nucleobases at the genomic coordinate based on a number of each observed nucleobase, the probability of each nucleobase at the genomic coordinate, and the prior genotype probability; and generating an estimated positive strand methylation level value of a nucleobase at the genomic coordinate on the positive strand and an estimated negative strand methylation level value of a nucleobase at the genomic coordinate on the negative strand by performing a Bayesian inversion on the probability of the observed nucleobases.

7. The method of claim 6, further comprising determining the probability of each nucleobase by: determining that a probability of a given thymine base on the positive strand approximates the estimated methylation level value; and determining that a probability of a given cytosine base on the positive strand approximates 1 minus the estimated methylation level value.

8. The method of claim 6, further comprising determining the probability of each nucleobase by: determining that a probability of a given adenine base on the negative strand approximates the estimated methylation level value; and determining that a probability of a given guanine base on the negative strand approximates 1 minus the estimated methylation level value. determining that a probability of a given guanine base on the negative strand approximates to 1 minus the estimated methylation level value.

9. The method of claim 1, wherein the variant call model comprises a hidden Markov model (HMM) that is modified to receive input based on the estimated methylation level value and a base call quality metric for a corresponding observed nucleobase.

10. The method of claim 1, further comprising generating the genotype call based on a predicted combination of nucleobases corresponding to a highest posterior genotype probability.

11. The method of claim 1, wherein identifying the nucleotide reads comprising the one or more nucleobases that are converted by the methylation sequencing assay comprises identifying the nucleotide reads comprising a thymine base or a uracil base that is converted from a cytosine base by the methylation sequencing assay.

12. A system comprising: at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: identify, for a target genomic sample, nucleotide reads comprising one or more nucleobases that are converted by a methylation sequencing assay; determine, based on a prior genotype probability of the target genomic sample at a genomic coordinate and an observed nucleobase at the genomic coordinate within the nucleotide reads, an estimated methylation level value for a cytosine base at the genomic coordinate; generate, using a variant call model, a posterior genotype probability for the target genomic sample at the genomic coordinate based on the estimated methylation level value and a base call quality metric for the observed nucleobase; and generate, based on the posterior genotype probability, a genotype call for the target genomic sample comprising a predicted combination of nucleobases at the genomic coordinate.

13. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to generate a refined methylation level value for the cytosine base at the genomic coordinate based on the posterior genotype probability and the observed nucleobase.

14. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to generate the refined methylation level value by: determining, based on the observed nucleobase, genotype-specific methylation level values corresponding to each possible genotype at the genomic coordinate; and weighting the genotype-specific methylation level values based on the posterior genotype probability.

15. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to generate the genotype call for the target genomic sample based on sequencing metrics corresponding to the nucleotide reads.

16. The system of claim 12, wherein the posterior genotype probabilities comprise a subset of posterior genotype probabilities for cytosine bases in the positive strand or the negative strand.

17. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to determine the estimated methylation level values by: determining a probability of each nucleobase at the genomic coordinate on the positive strand and the negative strand; determining a probability of the observed nucleobases at the genomic coordinate based on a number of each observed nucleobase, the probability of each nucleobase at the genomic coordinate, and the prior genotype probabilities; and generating an estimated positive strand methylation level value for nucleobases at the genomic coordinate on the positive strand and an estimated negative strand methylation level value for nucleobases at the genomic coordinate on the negative strand by performing a Bayesian inversion on the probabilities of the observed nucleobases.

18. The system of claim 17, further comprising instructions that, when executed by the at least one processor, cause the system to determine the probability of each nucleobase by: determining that a probability of a given thymine base on the positive strand approximates the estimated methylation level value; and determining that a probability of a given cytosine base on the positive strand approximates 1 minus the estimated methylation level value.

19. The system of claim 17, further comprising instructions that, when executed by the at least one processor, cause the system to determine the probability of each nucleobase by: determining that a probability of a given adenine base on the negative strand approximates the estimated methylation level value; and determining that a probability of a given guanine base on the negative strand approximates 1 minus the estimated methylation level value.

20. The system of claim 12, wherein the variant call model comprises a hidden Markov model (HMM) that is modified to receive inputs based on the estimated methylation level values and a base call quality metric for corresponding observed nucleobases.

21. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to generate the genotype calls based on determining a predicted combination of nucleobases corresponding to highest posterior genotype probabilities.

22. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to identify the nucleotide reads comprising the one or more nucleobases transformed by the methylation sequencing assay by identifying the nucleotide reads comprising thymine bases or uracil bases transformed by the methylation sequencing assay from cytosine bases.

23. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to: identifying, for a target genomic sample, nucleotide reads comprising one or more nucleobases converted by a methylation sequencing assay; determining, based on a prior genotype probability of the target genomic sample at a genomic coordinate and an observed nucleobase at the genomic coordinate within the nucleotide read, an estimated methylation level value of a cytosine base at the genomic coordinate; generating, using a variant call model, a posterior genotype probability of the target genomic sample at the genomic coordinate based on the estimated methylation level value of the observed nucleobase and a base call quality metric; and generating, based on the posterior genotype probability, a genotype call of the target genomic sample comprising a predicted combination of nucleobases at the genomic coordinate.

24. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the estimated methylation level value by: determining a probability of each nucleobase at the genomic coordinate on a plus strand and a minus strand; determining a probability of the observed nucleobase at the genomic coordinate based on a number of each observed nucleobase, the probability of each nucleobase at the genomic coordinate, and the prior genotype probability; and generating an estimated plus strand methylation level value of a nucleobase at the genomic coordinate on the plus strand and an estimated minus strand methylation level value of a nucleobase at the genomic coordinate on the minus strand by performing a Bayesian inversion on the probability of the observed nucleobase.

25. The non-transitory computer-readable medium of claim 23, wherein the variant call model comprises a hidden Markov model (HMM) modified to receive input based on the estimated methylation level value and a base call quality metric of a corresponding observed nucleobase.

26. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the computing device to generate the genotype call based on determining a predicted combination of nucleobases corresponding to a highest posterior genotype probability.

27. The non-transitory computer-readable medium of claim 23, wherein identifying the nucleotide reads comprising the one or more nucleobases converted by the methylation sequencing assay comprises identifying the nucleotide reads comprising a thymine base or a uracil base converted from a cytosine base by the methylation sequencing assay.