Genotype calling from sequencing data representing spontaneous deamination and / or methylation

The deamination-aware variant calling system addresses biased genotype calls and resource inefficiencies by estimating deamination events and integrating methylation analysis, improving accuracy and efficiency in sequencing systems.

WO2026117717A1PCT designated stage Publication Date: 2026-06-04ILLUMINA INC
View PDF 38 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-11-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing sequencing systems discard mismatched base calls from duplex reads, leading to biased genotype calls, reduced read data, and inefficient use of resources, particularly affecting low-quality samples and methylation detection.

Method used

A deamination-aware variant calling system that estimates probabilities of spontaneous deamination events and accounts for converted haplotype intermediate bases to generate accurate genotype calls from duplex consensus reads, integrating methylation-level determination in a single analysis.

Benefits of technology

Improves genotype call accuracy, reduces sequencing depth and resource consumption, and enhances methylation-level estimation by analyzing duplex consensus read data, particularly beneficial for high-throughput and clinical laboratories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025057329_04062026_PF_FP_ABST
    Figure US2025057329_04062026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure describes methods, non-transitory computer readable media, and systems that account for spontaneous deamination, sequencing errors, and / or methylation of nucleobases to more accurately predict genotype from reads of a sample. The disclosed systems can estimate a first set of probabilities that a sample exhibits observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases. The disclosed system can also determine a second set of probabilities that the sample exhibits converted haplotype intermediate bases given an underlying haplotype. Using a variant caller, the disclosed systems can determine a third set of probabilities of the sample exhibiting the observed duplex consensus bases at the genomic coordinate given candidate haplotype bases based on the first and second sets of probabilities. The disclosed systems may further use the third set of probabilities to determine posterior genotype probabilities and generate a genotype call.
Need to check novelty before this filing date? Find Prior Art

Description

GENOTYPE CALLING FROM SEQUENCING DATA REPRESENTING SPONTANEOUS DEAMINATION AND / OR METHYLATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 726,005, filed on November 27, 2024, entitled “GENOTYPE CALLING FROM SEQUENCING DATA REPRESENTING SPONTANEOUS DEAMINATION AND / OR METHYLATION,” (IP-2816-PRV), which is incorporated herein by reference in its entirety.BACKGROUND

[0002] In recent years, biotechnology firms and research institutions have improved hardware and software for sequencing nucleotides and determining nucleobase calls for genomic samples. For instance, some existing sequencing machines and sequencing-data-analysis software (together “existing sequencing systems”) predict individual nucleobases within sequences by using conventional Sanger sequencing or sequencing-by-synthesis (SBS) methods. When using SBS, existing sequencing systems can monitor many millions to billions of oligonucleotides being synthesized in parallel from templates to predict nucleobase calls for growing nucleotide reads. In many existing sequencing systems, a camera captures images of irradiated fluorescent tags incorporated into oligonucleotides. After capturing such images, some existing sequencing systems determine nucleobase calls for nucleotide reads corresponding to the oligonucleotides and send base-call data to a computing device with sequencing-data-analysis software, which aligns nucleotide reads with a reference genome. Based on differences between the aligned nucleotide reads and the reference genome, existing systems can further utilize a variant caller to identify variants of a genomic sample, such as single nucleotide variants (SNVs), insertions or deletions (indels), or other variants within the genomic sample.

[0003] In addition to improved genomic sequencing, biotechnology firms and research institutions have also improved methods of detecting variants at particular genomic regions (e.g., regions encoding or promoting genes) and detecting variants in larger nucleotide fragments or whole genomes of a sample. For instance, some existing systems can use duplex sequencing to improve the accuracy of genomic sequencing. In contrast to traditional sequencing methods, which often sequence only one strand of deoxyribonucleic acid (DNA) or treat both strands independently, duplex sequencing can compare the sequence of both the plus (e.g., forward) and minus (e.g., reverse) strands of a DNA duplex. The use of duplex sequencing often facilitates the identification of more accurate genetic variants while reducing sequencing errors. For example, duplex sequencing can result in lower error rates relative to existing sequencing systems because errors in one strand can be identified by checking the corresponding base on the complement strand.Attorney Docket No. IP-2816-PCT 1 Patent Application

[0004] Despite these recent advances, existing sequencing systems face several technical shortcomings. In particular, existing sequencing systems often discard data resulting in less- accurate genotype calls. While existing sequencing systems can use duplex sequencing, some existing systems require that complementary base calls be present on both complementary strands of duplex reads to determine a base call. Discrepancies between two complementary strands of a duplex read are often flagged as an error or assigned a zero-quality score and excluded from downstream processing, which ensures that some existing sequencing systems use only consistent bases for genotype calling. While discarding mismatched base calls in duplex sequencing can filter out errors and improve downstream analysis, discarding such mismatched bases can also negatively impact the accuracy of existing sequencing systems. Discarded mismatched bases from strands of duplex reads may indeed represent real genetic variants or biological phenomena rather than errors. By discarding mismatched bases from strands of duplex reads, existing sequencing systems can reduce the overall read data — including valuable data in difficult-to-call regions — available for analysis in variant calling. By automatically discarding mismatched base calls, existing sequencing systems can become biased toward high-confidence calls (e.g., inflated quality scores for variant calls) while overlooking real but less frequent variants or other causes of mismatched base calls.

[0005] In addition to overly confident genotype calls, discarded base call data can disproportionately impact low-quality samples. By discarding mismatched bases in strands of duplex reads, existing sequencing systems discard sequence data necessary for read coverage of a sample’s genomic regions and sensitivity for variant calls, especially for low-quality samples. To illustrate, low-quality samples may include degraded DNA that gives rise to more frequent mismatched base calls. By simply discarding mismatched base call data, existing sequencing systems may face difficulties retrieving meaningful data from the nucleotide reads of low-quality samples, thereby resulting in no reference or variant calls for certain genomic regions of the corresponding samples.

[0006] Beyond failing to account for mismatched base calls, some existing systems often cannot account for processes that may give rise to such a mismatch. To illustrate, some existing sequencing systems can use sequencing devices and corresponding sequencing-data-analysis software to identify when a methyl or hydroxymethyl group has been added to a cytosine base of a sample’s DNA. For example, existing sequencing systems can detect methylated cytosines by (i) enzymatically converting methylated or unmethylated cytosine bases at 5’-C-phosphate-G-3’ (CpG) or other cytosine sites from a sample nucleotide fragment into uracil bases (e.g., dihydrouracil); (ii) determining base calls of nucleotide reads for the sample using a sequencing device, where the sequencing device detects the uracil bases as thymine bases during polymerase chain reaction (PCR) amplification; and (iii) comparing the base calls from the nucleotide reads toAttorney Docket No. IP-2816-PCT 2 Patent Applicationa reference genome or non-enzymatically converted nucleotide reads from the sample. As indicated above, however, some existing sequencing systems may discard mismatched base calls in which the mismatch originates from methylation assay conversion protocols.

[0007] In addition to accuracy and methylation-error challenges, some existing systems inefficiently consume an inordinate amount of processing materials, time, and computing resources. Existing sequencing and methylation detection systems often require multiple assays and samples to accurately determine methylation levels and variant calls. For example, existing systems often require a separate methylation assay in addition to the utilization of a separate variant caller to accurately determine both methylation-level values and genotype calls. Accordingly, some existing systems require multiple samples from a single organism on which to perform both sequencing and methylation assays in separate computational analyses. The duplication of genomic samples often necessitates a duplication of computer processing, computer storage, software programs, and other resources to sequence and determine methylation levels for the same genomic sequence. Thus, existing systems often consume excessive genomic samples, significant time, and computer processing resources to both sequence and determine methylation levels for a single genomic sequence.

[0008] These, along with additional problems and issues exist in existing sequencing systems.SUMMARY

[0009] This disclosure describes one or more embodiments of systems, methods, and non- transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can account for spontaneous deamination that occurs to nucleobases during storage, library preparation, or other deamination events to more accurately generate genotype calls from nucleotide reads subject to a methylation assay. For example, based on target genomic sample’s duplex consensus reads for a genomic region, the disclosed system can estimate a first set of probabilities of the target genomic sample exhibiting an observed duplex consensus base given converted haplotype intermediate bases. Such converted haplotype intermediate bases may comprise a haplotype base that has been converted from one base class (e.g., cytosine) to another base class (e.g., adenine, guanine, or thymine) prior to sequencing the target genomic sample. The disclosed systems can further determine a second set of probabilities of the target genomic sample exhibiting the converted intermediate bases given various candidate haplotype bases. Based on the first and second set of probabilities, the disclosed system can determine a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases given the candidate haplotype bases. Such sets of probabilities further feed genotype calling. Based on posterior genotype probabilities derivedAttorney Docket No. IP-2816-PCT 3 Patent Applicationfrom the first, second, and third sets of probabilities, the disclosed systems generate a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic region.

[0010] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The detailed description refers to the drawings briefly described below.

[0012] FIG. 1 illustrates an environment in which a deamination-aware variant calling system can operate in accordance with one or more embodiments of the present disclosure.

[0013] FIG. 2 illustrates an overview of duplex sequencing to form duplex consensus reads in accordance with one or more embodiments of the present disclosure.

[0014] FIGS. 3A-3B illustrate an overview diagram of the deamination-aware variant calling system generating genotype calls by determining first, second, and third sets of probabilities — basecall probabilities, pre-sequencing intermediate base probabilities, and candidate-haplotype- derived consensus base probabilities, respectively — in accordance with one or more implementations of the present disclosure.

[0015] FIG. 4 illustrates the deamination-aware variant calling system determining the first set of probabilities of a target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases in accordance with one or more embodiments of the present disclosure.

[0016] FIG. 5 illustrates the deamination-aware variant calling system determining a second set of probabilities of a target genomic sample exhibiting converted haplotype intermediate bases given candidate haplotype bases in accordance with one or more implementations of the present disclosure.

[0017] FIG. 6 illustrates an overview of the deamination-aware variant calling system generating a third set of probabilities of a target genomic sample exhibiting observed duplex consensus bases given candidate haplotype bases and generating a genotype call based on the third set of probabilities in accordance with one or more implementations of the present disclosure.

[0018] FIGS. 7A-7B illustrate improvements in accuracy by the deamination-aware variant calling system calling C-to-T somatic variants relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.Attorney Docket No. IP-2816-PCT 4 Patent Application

[0019] FIGS. 8A-8B illustrate improvements in accuracy by the deamination-aware variant calling system calling T-to-C somatic variants relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.

[0020] FIG. 9 illustrates a flowchart of a series of acts for generating genotype calls based on duplex consensus base data in accordance with one or more embodiments of the present disclosure.

[0021] FIG. 10 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION

[0022] This disclosure describes embodiments of a deamination-aware variant calling system that can account for spontaneous deamination that occurs to nucleobases during storage or library preparation in determining genotype calls from duplex consensus reads of a genomic sample subject to a methylation assay. In particular, the deamination-aware variant calling system identifies, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads, where the duplex consensus bases comprise methylated or deaminated cytosine bases on a single strand. By analyzing error probabilities from the collapsed single-stranded reads, the deamination-aware variant calling system determines a first set of probabilities that the target genomic sample exhibits observed duplex consensus at a genomic coordinate given converted haplotype intermediate bases. The deamination-aware variant calling system further determines a second set of probabilities that the target genomic sample exhibits converted haplotype intermediate bases given underlying haplotype bases at the genomic coordinate. The deamination-aware variant calling system can further determine a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases based on the underlying haplotype bases and the first and second sets of probabilities. By deriving posterior genotype probabilities from the third set of probabilities, the deamination-aware variant calling system can generate a genotype call for the target genomic sample at the genomic coordinate.

[0023] As indicated above, in some embodiments, the deamination-aware variant calling system uses duplex-consensus-read data to estimate spontaneous deamination events for a genomic sample. To estimate such deamination, the deamination-aware variant calling system considers converted haplotype intermediate bases that are unobserved but may arise from artifacts during library preparation, storage, or processes such as conversion protocols in methylation assays. In particular, the deamination-aware variant calling system estimates converted haplotype intermediate bases that show mismatched collapsed nucleobases on different strands of a duplex consensus read (e.g., T+C or G+A'). The deamination-aware variant calling system can leverage such mismatched converted haplotype intermediate bases to further estimate a probability of aAttorney Docket No. IP-2816-PCT 5 Patent Applicationtarget genomic sample having an observed duplex consensus base given candidate haplotype bases (e.g., A, C, T, or G).

[0024] To estimate or account for spontaneous deamination events, the deamination-aware variant calling system recognizes certain observed duplex consensus bases. For instance, in contrast to existing sequencing systems that typically discard mismatching observed duplex consensus bases, such as T+C~ or G+A“on plus and minus strands, respectively, the deamination-aware variant calling system recognizes certain observed duplex consensus bases that arise from conversion protocols in methylation assays, library preparation, or other sources. Such observed duplex consensus bases can include a methylated cytosine base on a plus strand (e.g.,mC+:= T C' ) or a methylated cytosine base on a minus strand (e.g.,mC" := G+A'). By determining the first, second, and third probabilities above using such observed duplex consensus bases, the deamination-aware variant calling system leverages unique duplex consensus bases representing mismatching bases on complementary strands to estimate the effects of intermediate base conversions on genotype calling.

[0025] Based on duplex consensus reads comprising such duplex consensus bases for strandspecific methylated cytosines, as mentioned above, the deamination-aware variant calling system determines a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases (e.g., shown as Ri below) at a genomic coordinate given converted haplotype intermediate bases (e.g., shown as H'j below) that represent a haplotype base converted from one base class to another base class. For ease of reference, this disclosure will sometimes and interchangeably refer to “basecall probabilities” as an abbreviated term for such a first set of probabilities of a target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases. As explained below, the deamination- aware variant calling system can estimate the basecall probabilities based on strand-specific error probabilities. More specifically, the deamination-aware variant calling system can determine the likelihood that a collapsed base on a plus single-stranded read or on a minus single-stranded read is an error. Based on the plus-strand error probability and the minus-strand error probability, the deamination-aware variant calling system determines the basecall probabilities. In some implementations, the plus-strand error probability and the minus-strand error probabilities represent sequencing errors by a sequencing device during a sequencing run. The deamination- aware variant calling system can use a combination of the plus-strand and minus-strand error probabilities to estimate whether an observed duplex consensus base at a genomic coordinate arises from a deaminated cytosine base, a methylated cytosine base, or a sequencing error.

[0026] In addition to determining the first set of probabilities, the deamination-aware variant calling system can determine a second set of probabilities of the target genomic sample exhibitingAttorney Docket No. IP-2816-PCT 6 Patent Applicationthe converted haplotype intermediate bases (e.g., shown as H'j below) given candidate haplotype bases (e.g., shown as Hj below) at the relevant genomic coordinate. For ease of reference, this disclosure will sometimes and interchangeably refer to “pre-sequencing intermediate base probabilities” as an abbreviated term for such a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases. To determine the pre-sequencing intermediate base probabilities, for example, the deamination-aware variant calling system determines duplex-consensus base-conversion probabilities that approximate a likelihood that a haplotype base converts from one base class to another base class. In some cases, the deamination-aware variant calling system determines duplex-consensus baseconversion probabilities on a genomic-sample-wide basis and stores the duplex-consensus baseconversion probabilities in a lookup table. In some examples, the deamination-aware variant calling system further uses estimated methylation-level values to predict the pre-sequencing intermediate base probabilities. Generally, the estimated methylation-level values indicate a likelihood that a haplotype base will be converted to a converted haplotype intermediate base during a conversion protocol of a methylation assay.

[0027] Based on the first set of probabilities and the second set of probabilities — also called basecall probabilities and pre-sequencing intermediate base probabilities, respectively — the deamination-aware variant calling system can generate a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases (e.g., shown as Ri below) at the genomic coordinate given the candidate haplotype bases (e.g., shown as Hj below). For ease of reference, this disclosure will sometimes and interchangeably refer to “candidate-haplotype- derived consensus base probabilities” as an abbreviated term for such a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases. To generate such basecall probabilities, for instance, the deamination-aware variant calling system can determine a probability that the target genomic sample exhibits each observed duplex consensus base (e.g., A, T, C, G,mC+, ormC ) given each candidate haplotype base (e.g., A, T, C, G) at the genomic coordinate. By using the basecall probabilities and pre-sequencing intermediate base probabilities, the deamination-aware variant calling system introduces a model that considers converted haplotype intermediate bases not considered by existing systems to generate more accurate basecall probabilities.

[0028] The third set of probabilities — also called the basecall probabilities — form the basis for posterior genotype probabilities and corresponding genotype calls. Based on the basecall probabilities, in some implementations, the deamination-aware variant calling system generates posterior genotype probabilities. The deamination-aware variant calling system may use such posterior genotype probabilities to generate a genotype call that the target genomic sampleAttorney Docket No. IP-2816-PCT 7 Patent Applicationcomprises a predicted combination of nucleobases at the genomic coordinate. For example, the deamination-aware variant calling system may select a genotype call with a highest posterior genotype probability as the genotype call.

[0029] In addition to genotype calls, the deamination-aware variant calling system can also estimate deamination or methylation levels for a target genomic sample. Additionally, the deamination-aware variant calling system can estimate sequencing errors by the sequencing device during a sequencing run. In some implementations, the deamination-aware variant calling system further utilizes the posterior genotype probabilities to generate a refined methylation-level value for a nucleobase at a genomic coordinate. The refined methylation-level value can represent a cytosine methylation percentage at the genomic coordinate. More specifically, the methylationgenotype-calling system can determine refined methylation-level values (a) for genomic coordinates at which a reference base is a cytosine base and a prior genotype probability indicates a cytosine base and (b) for genomic coordinates at which the target genomic sample comprises cytosine bases as alternative haplotypes. In addition to relying on the posterior genotype probabilities, the deamination-aware variant calling system may determine the refined methylationlevel value further based on a number of reads in the read pileup and a number of methylated nucleobases in the read pileup. Because the deamination-aware variant calling system determines the refined methylation-level value based on posterior genotype probabilities, the refined methylation-level value may be more accurate than estimated methylation-level values.

[0030] As indicated above, the deamination-aware variant calling system provides several technical advantages relative to existing sequencing systems by, for example, improving genotypecalling accuracy, methylation-level estimations, and computational efficiency relative to existing sequencing systems. In particular, the deamination-aware variant calling system can more accurately generate genotype calls by preserving and analyzing data from mismatched bases on strands of a duplex consensus read and converted haplotype intermediate bases. The deamination- aware variant calling system represents and accounts for certain converted haplotype intermediate bases (e.g.,mC+:= T C' ormC := G+A') that are often disregarded or assigned zero-value quality scores by existing sequencing systems. In some cases, the newly represented converted haplotype intermediate bases and observed duplex consensus bases can arise from conversion protocols for methylation assays, spontaneous deamination events, or sequencing errors. By analyzing data from duplex consensus reads that represent these mismatched bases — and determining sets of probabilities of a target genomic sample exhibiting observed duplex consensus bases (e.g., shown as Ri below) and converted haplotype intermediate bases (e.g., shown as H'j below) — the deamination-aware variant calling system can generate more accurate genotype calls. Unlike existing sequencing systems, the genotype calls generated by the deamination-aware variant callingAttorney Docket No. IP-2816-PCT 8 Patent Applicationsystem account for duplex-consensus-read data representing mismatched bases that can help distinguish between sequencing errors and true biological variations, especially in complex genomic regions or low-frequency variants. By generating more accurate genotype calls from duplex consensus reads, the deamination-aware variant calling system can lower sequencing depths and the amount of sequencing data required per sample, which is particularly advantageous in high- throughput sequencing laboratories and / or clinical laboratories that process significant amounts of sequencing data. As described below, FIGS. 7A-8B demonstrate improvements in variant calling accuracy achieved by the deamination-aware variant calling system when retaining and analyzing such duplex-consensus-read data.

[0031] As part of improving such variant calls, the deamination-aware variant calling system analyzes and accounts for more accurate and reliable read data — which accounts for the physical state and deamination events of a sample’s templates for nucleotide reads — that would otherwise be discarded as noise. Consequently, the deamination-aware variant calling system can (i) use more DNA extracted from a sample to determine genotype calls or (ii) in some cases or genomic regions, require less DNA from a sample relative to existing sequencing systems to determine genotype calls from reads subject to a methylation sequencing assay. As mentioned, existing sequencing systems cannot differentiate between mismatching bases on strands of a duplex consensus read arising from sequencing errors, on the one hand, and mismatches caused by deamination events or other biological processes, on the other hand. Instead of discarding mismatching bases on a duplex consensus read as noise, the deamination-aware variant calling system accounts for and recognizes the physical state of a target genomic sample’s reads collapsed into a duplex consensus read, thereby identifying true variants or biological phenomena that is often obscured by existing sequencing systems. By analyzing duplex-consensus-read data that represent these mismatched bases — and determining sets of probabilities of a target genomic sample exhibiting observed duplex consensus bases (e.g., shown as Ri below) and converted haplotype intermediate bases (e.g., shown as H'j below) — the deamination-aware variant calling system constructs a first-of-its-kind model that can account for deamination events in duplex consensus reads to generate more accurate genotype calling based on relatively less of a sample’s DNA relative to existing sequencing systems. By uniquely analyzing duplex-consensus-read data and generating more accurate genotype calls from duplex consensus reads, the resulting reduction in the required DNA from a sample is particularly advantageous in high-throughput sequencing laboratories and / or clinical laboratories that process a significant number of samples and amounts of sequencing data. The deamination-aware variant calling system thus leverages converted haplotype intermediate bases data (e.g.,mC+:= TC’ ormC" := G A ) to improve accuracy of genotype calls and, in some cases, deamination or methylation level values based on DNA extracted from a sample. As describedAttorney Docket No. IP-2816-PCT 9 Patent Applicationfurther below, the deamination-aware variant calling system can also use such data to generate more accurate estimated methylation-level values, error probabilities, and base-call-quality scores.

[0032] Beyond improved genotype-call accuracy and methylation-level estimates, in some embodiments, the deamination-aware variant calling system improves computing efficiency in processing and physical resources relative to existing systems. As noted above, because state-of- the-art genotype-calling and methylation detection accuracy can be unfit for clinical benchmarks, some existing systems execute (i) a separate methylation sequencing assay to chemically or enzymatically convert nucleotide reads from a genomic sample and determine deamination or methylation levels and (ii) a separate DNA sequencing run with non-chemically or non- enzymatically converted nucleotide reads from the genomic sample to determine variant calls. Such separate methylation sequencing assays and DNA sequencing can consume and duplicate computer processing, memory storage, physical space and reagents for a nucleotide-sample slide (e.g., flow cell), and software programs (e.g., separate methylation analysis and variant calling software). In contrast to such a bifurcated approach, in some embodiments, the deamination-aware variant calling system can simultaneously determine methylation-level values indicating levels of methylation of a target genomic sample’s cytosine bases and generate variant calls for the genomic sample with improved accuracy. Thus, the deamination-aware variant calling system can efficiently generate epigenetic and genetic sequencing data from a single genomic sample. By generating both methylation-level values and variant calls from the same genomic sample, the deamination-aware variant calling system further reduces the amount of computer processing, computer storage, software programs, space used on a nucleotide-sample slide in a sequencing device, and other resources to generate accurate sequencing and methylation data.

[0033] As suggested by the foregoing discussion, this disclosure utilizes a variety of terms to describe features and benefits of the deamination-aware variant calling system. Additional detail is hereafter provided regarding the meaning of these terms as used in this disclosure. As used in this disclosure, for instance, the term “target genomic sample” (or simply “sample”) refers to a specimen, culture, or the like that is suspected of including a target nucleic acid. In some embodiments, the sample comprises DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acids as targets. The sample can likewise include any biological, clinical, surgical, agricultural-atmospheric, or aquatic-based specimen containing one or more nucleic acids. A target genomic sample also includes any isolated or extracted nucleic acid sample from an organism, such a genomic DNA, fresh-frozen, or formalin-fixed paraffin-embedded nucleic acid specimen. In some cases, accordingly, a target genomic sample can include a full genome or partial genome that is isolated or extracted (e.g., in whole or in part by a kit) from an organism and that is prepared to undergo sequencing or an assayAttorney Docket No. IP-2816-PCT 10 Patent Applicationin a sequencing device. A target genomic sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.

[0034] The target genomic sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source, such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0035] As further used herein, the term “nucleotide read” (or simply “read”) refers to an inferred sequence of one or more nucleobases (or nucleobase pairs) from all or part of a sample nucleotide sequence (e.g., a sample genomic sequence, complementary DNA). In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a nucleotide sequence (or group of monoclonal nucleotide sequences) from a sample library fragment corresponding to a genomic sample. For example, in some cases, a sequencing device determines a nucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a cluster in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g., assembled nucleotide reads) satisfying a threshold number ofAttorney Docket No. IP-2816-PCT 11 Patent Applicationnucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.

[0036] As used herein, the term “single strand” refers to one strand of DNA or a linear chain of nucleotides that forms one half or a part of double-stranded nucleotide sequences (e.g., DNA). In particular, a single strand can refer to one nucleotide read in a pair of paired-end reads, such as one read from a pair of collapsed single-stranded reads. In some cases, a single strand comprises one of two complementary, antiparallel strands of genomic material. Single strands can be identified based on the direction of sequencing relative to a reference sequence in a reference genome. For example, a single strand refers to a plus (e.g., Rl) or minus (e.g., R2) nucleotide read. In some cases, a single strand refers to a read from single-end sequencing.

[0037] As used herein, the term “plus strand” (or “positive strand”) refers to a first read from a pair of paired-end reads, such as a read designated by convention as a first read from a pair of collapsed single-stranded reads. In particular, the plus strand can include a first read from one end of a nucleotide fragment and be in a forward orientation relative to a reference sequence of a reference genome. For example, a plus strand includes a forward 1st, reverse 2nd (F1R2) read of a nucleotide fragment in a pair of paired-end reads (or alternatively referred to as Rl).

[0038] Relatedly, the term “minus strand” (or “negative strand”) refers to a second read from a pair of paired-end reads, such as a read designated by convention as a second read from a pair of collapsed single-stranded reads. In particular, the minus strand can include a second read from an opposite end of a nucleotide fragment and be in a reverse orientation relative to a reference sequence of a reference genome. For example, a minus strand includes a forward 2nd, reverse 1st (F2R1) read of a nucleotide fragment in a pair of paired-end reads (or alternatively referred to as R2).

[0039] As used herein, the term “collapsed single-stranded read” refers to a sequence of nucleobases that is representative of nucleotide reads aligning to the same genomic region or positions on a single strand (e.g., plus strand or minus strand). In particular, a collapsed singlestranded read includes a data string for a consolidated sequence of nucleobases that represent a sample’s nucleotide reads that align to either a plus or minus single strand. For example, during sequencing, multiple nucleotide reads may be generated for the same genomic coordinate(s) or nucleotide fragment due to sequencing coverage or repetitive sequencing. In some examples, nucleotide reads that align to the same genomic coordinate(s) on the plus strand or minus strand can be very similar or identical in nucleobase content. The aligned reads can be collapsed into a collapsed single-stranded read that is representative of nucleotide reads aligning to the same genomic coordinate(s) of the same strand. For instance, a nucleobase on a collapsed single-stranded read at a genomic coordinate may be denoted by C+, which represents a cytosine on a plus strand.Attorney Docket No. IP-2816-PCT 12 Patent ApplicationIn some examples, the bases on the collapsed single-stranded read at the genomic coordinate on the complimentary strand (or the minus strand) may be complemented to match the notation on the plus strand. To illustrate, the complemented base on the collapsed single-stranded read at the genomic coordinate may be denoted by C-, which complements that C+ at the genomic coordinate on the plus strand.

[0040] As used herein, the term “duplex consensus read” refers to a sequence of duplex consensus bases that is representative of nucleotide reads aligning to the same genomic position on complementary strands of DNA. In particular, a duplex consensus read can comprise a combination of collapsed single-stranded reads for complementary strands of DNA. For example, a duplex consensus read may include duplex consensus bases at a genomic coordinate that represent both a nucleobase on a plus collapsed single-stranded read at a genomic coordinate and a nucleobase on a complementary minus collapsed single-stranded at the genomic coordinate.

[0041] As used herein, the term “duplex consensus base” refers to a nucleobase that makes up duplex consensus reads. In particular, a duplex consensus base comprises a consensus nucleobase representing a nucleobase on one or both of a collapsed plus single-stranded read and the complementary collapsed minus single-stranded read at a genomic coordinate within a duplex consensus read. For example, a duplex consensus base at a genomic coordinate may be denoted by A+A-, which represents an adenine on the plus collapsed single-stranded read and, by inference, a thymine on the minus collapsed single-stranded read. Even though the minus collapsed singlestranded read includes a thymine, in some cases, the minus strand base notation is complemented to match the plus strand base notation. For instance, the notation A+A- can indicate an adenine base on the minus collapsed single-stranded read, emphasizing alignment rather than focusing on the specific complementary base. As indicated by the foregoing example, the deamination-aware variant calling system can use an identical notation logic to represent an actual “A+T-,” where an adenine base on a plus collapsed single-stranded read actual has a complementary thymine base on a minus collapsed single-stranded read, instead as “A+A-” to facilitate alignment of plus and minus collapsed single-stranded reads. As a further example, a duplex consensus base at a genomic coordinate may be denoted by T+C-, which represents a thymine on the plus single-stranded read and a complemented minus strand base, that is, C-, which represents what is actually a G- on the minus single-stranded read. As explained further below, such duplex consensus bases can indicate a deamination event and be represented as a converted haplotype intermediate base asmC+:= T C' ormC := G A'.

[0042] As used herein, the term “deaminated cytosine base” refers to a cytosine base that has lost an amino group. In particular, a deaminated cytosine base may comprise a cytosine molecule that has lost its amino group (-NEE). Deamination can occur spontaneously or due to externalAttorney Docket No. IP-2816-PCT 13 Patent Applicationprocesses. For example, a cytosine base may be deaminated as part of a methylation assay. To illustrate, the term “methylated cytosine base” refers to a cytosine base to which a methyl group or hydroxymethyl group has been added or bonded. In particular, a methylated cytosine base refers to a cytosine base on a single strand to which a methyl group (CH3) is added to the carbon-5 position of the cytosine ring. As part of some methylation assays, a methylated or unmethylated cytosine base can be enzymatically converted to a uracil base or thymine base. Additionally, a deaminated cytosine base can arise from artifacts during library preparation or storage of a genomic sample.

[0043] As used herein, the term “observed duplex consensus base” refers to a duplex consensus base that has been determined or predicted at a genomic coordinate within a duplex consensus read. In particular, an observed duplex consensus base includes nucleobases called by a sequencing device for nucleotide reads that have been collapsed into collapsed single-stranded reads and, subsequently, collapsed into duplex consensus reads. Accordingly, an observed duplex consensus base may comprise nucleobase call by a sequencing device for a collapsed plus single-stranded read and a collapsed minus single-stranded read. In some implementations, the deamination-aware variant calling system can identify observed duplex consensus bases at a given genomic coordinate for duplex consensus reads.

[0044] As used herein, the term “genomic coordinate” (or sometimes simply “coordinate”) refers to a particular location or position of a nucleobase within a genome (e.g., an organism’s genome or a reference genome). Accordingly, a genomic coordinate can refer to a particular location or position at which a reference base occurs, a variant different from the reference base occurs, a variant base within an insertion has been inserted, or a deletion occurs. For instance, a genomic coordinate can refer to a position of a nucleobase within an insertion. In some examples, a genomic coordinate can be the coordinate of an anchor nucleobase that is positioned upstream of an insertion of a sample relative to nucleobases within a reference genome or that is otherwise positioned offset from an insertion for a variant call of the sample. In some cases, a genomic coordinate includes an identifier for a particular chromosome of a genome and an identifier for a position of a nucleobase within the particular chromosome. For instance, a genomic coordinate or coordinates may include a number, name, or other identifier for a chromosome (e.g., chrl or chrX) and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl: 1234570 or chrl: 1234570-1234870). In certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mitochondrial DNA reference genome or SARS-CoV-2 for a reference genome for the SARS-CoV-2 virus) and a position of a nucleobase within the source for the reference genome (e.g., mt: 16568 or SARS-CoV- 2:29001). By contrast, in certain cases, a genomic coordinate refers to a position of a nucleobase within a reference genome without reference to a chromosome or source (e.g., 29727).Attorney Docket No. IP-2816-PCT 14 Patent Application

[0045] As used herein, the term “nucleobase” refers to a nitrogenous base. In particular, nucleobases comprise components of nucleotides. For example, a nucleobase may be an adenine (A), cytosine (C), guanine (G), or thymine (T).

[0046] As used herein, the term “base class” refers to a particular type or class of nitrogenous base. For instance, a genome or nucleotide sequence may include five different base classes, including adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U).

[0047] As also used herein, the term “haplotype” refers to nucleotide sequences that are present in an organism (or present in organisms from a population) and inherited from one or more ancestors. In particular, a haplotype can include alleles or other nucleotide sequences present in organisms of a population and inherited together by such organisms respectively from a single parent. In one or more embodiments, haplotypes include a set of SNPs on the same chromosome that tend to be inherited together. In some cases, data representing a haplotype or a set of different haplotypes are stored or otherwise accessible on a haplotype database.

[0048] As used herein, the term “haplotype base” refers to a specific nucleobase from a haplotype at a genomic coordinate. For example, a haplotype base can be an SNP, an insertion, or a deletion that is inherited in linkage with other genetic variants. Relatedly, the term “candidate haplotype base” refers to a specific base class for a haplotype base. For example, a candidate haplotype base can be any of the five different base classes, including adenine (A), cytosine (C), guanine (G), or thymine (T) / uracil (U) at a genomic coordinate.

[0049] As used herein, the term “converted haplotype intermediate base” refers to a haplotype base that has been converted from one base class to another base class but the conversion was unobserved or not directly detected. In some cases, for example, the converted haplotype intermediate base comprises a hidden variable that is not observed. In particular, the converted haplotype intermediate base may occur prior to an observed duplex consensus base. For example, a converted haplotype intermediate base may occur because of methylation or artifacts during preparation of a template derived from the genomic sample. More specifically, converted haplotype intermediate bases arise from artifacts that alter the haplotype base that occur upstream of sequencing.

[0050] As used herein, the term “variant call model” (or simply “variant caller”) refers to a probabilistic model that generates sequencing data from nucleotide reads of a sample nucleotide sequence, including variant calls and associated metrics. For example, in some cases, a variant call model refers to a Bayesian probability model that generates variant calls based on nucleotide reads of a sample nucleotide sequence. Such a model can process or analyze sequencing metrics corresponding to read pileups (e.g., multiple nucleotide reads corresponding to a single genomic coordinate), including mapping quality, base quality, and various hypotheses including foreignAttorney Docket No. IP-2816-PCT 15 Patent Applicationreads, missing reads, joint detection, and more. A variant call model may likewise include multiple components, including, but not limited to, different software applications or components for mapping and aligning, sorting, duplicate marking, computing read pileup depths, and variant calling. In some cases, the variant call model includes the ILLUMINA DRAGEN model for variant calling functions and mapping and alignment functions.

[0051] As used herein, the term “genotype call” refers to a determination or prediction of a particular genotype of a genomic sample at a genomic locus. In particular, a genotype call can include a prediction of a particular genotype of a genomic sample with respect to a reference genome or a reference sequence at a genomic coordinate or a genomic region. For instance, in some cases, a genotype call includes a determination or prediction that a genomic sample comprises both a nucleobase and a complementary nucleobase at a genomic coordinate that is either homozygous or heterozygous for a reference base or a variant (e.g., homozygous reference bases represented as 0|0 or heterozygous for a variant on a particular strand represented as 0| 1). A genotype call is often determined for a genomic coordinate or genomic region at which a single nucleotide variant (SNV) or other variant has been identified for a population of organisms. In this disclosure, among other genotype calls, the methylation-genotype-calling system predicts genotype calls for SNV regions within a genomic sample.

[0052] The following paragraphs describe the deamination-aware variant calling system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which a deamination- aware variant calling system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server device(s) 110, and a client device 114. As shown in FIG. 1, the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114 can communicate with each other via a network 118. The network 118 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 10. While FIG. 1 shows an embodiment of the deamination-aware variant calling system 106, this disclosure describes alternative embodiments and configurations below.

[0053] As indicated by FIG. 1, the sequencing device 102 comprises a computing device and a sequencing device system 104 for sequencing a genomic sample or other nucleic-acid polymer. In some embodiments, by executing the sequencing device system 104 using a processor, the sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 102. More particularly, theAttorney Docket No. IP-2816-PCT 16 Patent Applicationsequencing device 102 receives nucleotide-sample slides (e.g., flow cells) comprising nucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments.

[0054] In one or more embodiments, the sequencing device 102 utilizes sequencing-by- synthesis (SBS) techniques to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition, or in the alternative to communicating across the network 118, in some embodiments, the sequencing device 102 bypasses the network 118 and communicates directly with the local device 108, and / or the client device 114. By executing the sequencing device system 104, the sequencing device 102 can further store the nucleobase calls as part of a base-call data fde that is formatted as a binary base call (BCL) file and / or a FASTQ file and send the BCL file and / or FASTQ file to the local device 108, and / or the server device(s) 110.

[0055] As further indicated by FIG. 1 , the local device 108 is located at or near a same physical location of the sequencing device 102. Indeed, in some embodiments, the local device 108 and the sequencing device 102 are integrated into a same computing device. The local device 108 may run the sequencing device system 104 and / or the deamination-aware variant calling system 106 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As shown in FIG. 1, the sequencing device 102 may send (and the local device 108 may receive) base-call data generated during a sequencing run of the sequencing device 102. The local device 108 may also communicate with the client device 114. In particular, the local device 108 can send data to the client device 114, including a binary alignment map (BAM) fde, a variant call format (VCF) fde, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.

[0056] As further indicated by FIG. 1, the server device(s) 110 are located remotely from the local device 108 and the sequencing device 102. Like the local device 108, in some embodiments, the server device(s) 110 include a version of (or are otherwise able to access or implement) the deamination-aware variant calling system 106. For example, the server device(s) 110 can implement the deamination-aware variant calling system 106 as part of a sequencing system 112. Accordingly, the server device(s) 110 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such basecall data. As indicated above, the sequencing device 102 may send (and the server device(s) 110 may receive) base-call data from the sequencing device 102. The server device(s) 110 may also communicate with the client device 114. In particular, the server device(s) 110 can send data to the client device 114, including BAM fries, VCF fries, or other sequencing related information.

[0057] In some embodiments, the server device(s) 110 comprise a distributed collection of servers where the server device(s) 110 include a number of server devices distributed across theAttorney Docket No. IP-2816-PCT 17 Patent Applicationnetwork 118 and located in the same or different physical locations. Further, the server device(s) 110 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0058] As further illustrated and indicated in FIG. 1, by executing a sequencing application 116, the client device 114 can generate, store, receive, and send digital data. In particular, the client device 114 can receive sequencing data from the local device 108 or receive base-call data files (e.g., BCL and FASTQ) and sequencing metrics from the sequencing device 102. Furthermore, the client device 114 may communicate with the local device 108 or the server device(s) 110 to receive a VCF comprising genotype or variant calls and / or other metrics, such as base-call-quality metrics or pass-filter metrics. The client device 114 can accordingly present or display information pertaining to variant calls or other genotype calls within a graphical user interface of the sequencing application 116 to a user associated with the client device 114. For example, the client device 114 can present nucleobase calls, genotype calls, variant calls, and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of the sequencing application 116.

[0059] Although FIG. 1 depicts the client device 114 as a desktop or laptop computer, the client device 114 may comprise various types of client devices. For example, in some embodiments, the client device 114 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 114 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 114 are discussed below with respect to FIG. 10.

[0060] As further illustrated in FIG. 1, the client device 114 includes the sequencing application 116. The sequencing application 116 may be a web application or a native application stored and executed on the client device 114 (e.g., a mobile application, desktop application). The sequencing application 116 can include instructions that (when executed) cause the client device 114 to receive data from the deamination-aware variant calling system 106 and present, for display at the client device 114, base-call data or data from an alignment data fde or VCF. Furthermore, the sequencing application 116 can instruct the client device 114 to display summaries for multiple sequencing runs.

[0061] As further illustrated in FIG. 1, a version of the deamination-aware variant calling system 106 may be located and / or implemented (e.g., entirely or in part) on the client device 114 or the sequencing device 102. In yet other embodiments, the deamination-aware variant calling system 106 is implemented by one or more other components of the computing system 100, such as the local device 108. In particular, the deamination-aware variant calling system 106 can be implemented in a variety of different ways across the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114. For example, the deamination-aware variantAttorney Docket No. IP-2816-PCT 18 Patent Applicationcalling system 106 can be downloaded from the server device(s) 110 to the client device 114 and / or the local device 108 where all or part of the functionality of the deamination-aware variant calling system 106 is performed at each respective device within the computing system 100.

[0062] As mentioned previously, in some embodiments, sequencing systems use duplex sequencing and duplex consensus reads to achieve more accurate sequencing results. FIG. 2 illustrates an overview of duplex sequencing in accordance with one or more embodiments of the present disclosure. By way of overview, FIG. 2 illustrates a series of acts comprising an act 202 of sequencing single-stranded plus and single-stranded minus reads where the deamination-aware variant calling system 106 sequences individual nucleotide reads. The deamination-aware variant calling system 106 performs an act 204 of collapsing the single-stranded plus reads and singlestranded minus reads into collapsed single-stranded reads. The deamination-aware variant calling system 106 further performs an act 206 of combining the collapsed single-stranded reads into duplex consensus reads.

[0063] As shown in FIG. 2, the deamination-aware variant calling system 106 performs the act 202 of sequencing single-stranded plus reads and single-stranded minus reads. By performing sequencing cycles on the sequencing device 102, for example, the deamination-aware variant calling system 106 determines nucleotide sequences of complementary strands of DNA in the form of nucleotide reads. FIG. 2 illustrates the deamination-aware variant calling system 106 determining nucleotide sequences of an original DNA molecule having a methylated cytosine on one strand and a guanine on the complementary strand at a genomic coordinate. As shown in FIG. 2, the deamination-aware variant calling system 106 sequences or determines nucleotide reads 210. The nucleotide reads 210 comprise individual nucleotide reads that map to a reference genome 208. The nucleotide reads 210 comprise a proportion of all nucleotide reads that map to a genomic region. In some embodiments, the deamination-aware variant calling system 106 can determine that the nucleotide reads 210 originate from the same original molecule from a genomic sample. By contrast, the deamination-aware variant calling system 106 can also determine that the nucleotide reads 210 without respect to a molecule from a genomic sample. Further, in some examples, the deamination-aware variant calling system 106 can determine which of the nucleotide reads comprise single-stranded plus reads and single-stranded minus reads that correspond to the plus strand and minus strand, respectively, of the original molecule from a genomic sample.

[0064] As further shown in FIG. 2, the deamination-aware variant calling system 106 performs the act 204 of collapsing the reads into collapsed single-stranded reads. The deamination-aware variant calling system 106 consolidates overlapping single-stranded reads from the same genomic region into a single continuous sequence or what is called a collapsed single-stranded read. As shown in FIG. 2, the deamination-aware variant calling system 106 collapses single-stranded plusAttorney Docket No. IP-2816-PCT 19 Patent Applicationreads 212 into a collapsed plus single-stranded read 216. As further shown in FIG. 2, the collapsed plus single-stranded read 216 includes a thymine base at the genomic coordinate indicating a C-to- T conversion as part of a methylation assay or a deamination event. Conversely, the deamination- aware variant calling system 106 collapses single-stranded minus reads 214 into a collapsed minus single-stranded read 218. The collapsed plus single-stranded read 216 and the collapsed minus single-stranded read 218 reflect the most likely base at each genomic coordinate of a plus strand and a minus strand, respectively. As illustrated in FIG. 2, the collapsed minus single-stranded read 218 includes a guanine base at the genomic coordinate.

[0065] In some implementations, and as part of performing the act 204 of collapsing the singlestranded plus and minus reads into collapsed single-stranded reads, the deamination-aware variant calling system 106 adjusts base-call quality scores of each base call. Generally, the deamination- aware variant calling system 106 can adjust the base-call quality scores (e.g., Q-score) of individual base calls to reflect increased confidence in the base calls. In some implementations, the deamination-aware variant calling system 106 increases the base-call quality scores of base calls that occur in a higher proportion of single-stranded plus and single-stranded minus reads. Where there is disagreement in certain base calls among single-stranded reads, the deamination-aware variant calling system 106 can lower the base-call quality score. Furthermore, the deamination- aware variant calling system 106 can increase the base-call quality scores for base calls corresponding with higher read depth.

[0066] As further shown in FIG. 2, the deamination-aware variant calling system 106 performs the act 206 of combining collapsed single-stranded reads into duplex consensus reads. For example, and as illustrated in FIG. 2, the deamination-aware variant calling system 106 combines the collapsed plus single-stranded read 216 with the collapsed minus single-stranded read 218 to form a duplex consensus read. As part of the act 204, the deamination-aware variant calling system 106 aligns complementary collapsed plus single-stranded reads and collapsed minus single-stranded reads.

[0067] In some embodiments, the deamination-aware variant calling system 106 can compare collapsed bases at each genomic coordinate along the duplex consensus read to adjust the base-call quality score of each base call. Matching collapsed bases from the collapsed plus single-stranded read 216 and the collapsed minus single-stranded read 218 are designated as duplex consensus bases in the duplex consensus read. As mentioned previously, existing sequencing systems often compare collapsed bases along a single-stranded plus read and a single-stranded minus read. If the collapsed bases at a genomic coordinate match, existing sequencing systems often increase the base-call quality score of the base call at the genomic coordinate. If the collapsed bases at a genomic coordinate do not match, however, existing sequencing systems may assign a placeholder baseAttorney Docket No. IP-2816-PCT 20 Patent Application(e.g., N base) and / or assign a base-call quality score of zero (e.g., Q-score = 0). For example, as illustrated in FIG. 2, the collapsed minus single-stranded read 218 and the collapsed plus singlestranded read 216 contain mismatched collapsed bases at a genomic coordinate. In contrast to existing sequencing systems that simply discard base calls or assign a zero base-call quality score corresponding to mismatched collapsed bases between the collapsed plus single-stranded read 216 and the collapsed minus single-stranded read 218, the deamination-aware variant calling system 106 recognizes new duplex consensus bases to differentiate mismatched collapsed bases arising from spontaneous deamination or other processes.

[0068] As mentioned previously, in some embodiments, the deamination-aware variant calling system 106 accounts for spontaneous deamination of nucleobases to more accurately predict genotype and methylation level from nucleotide reads. More particularly, the deamination-aware variant calling system 106 can estimate spontaneous deamination by determining various sets of probabilities of bases converting from one base class to another. FIGS. 3A-3B depict an overview diagram of the deamination-aware variant calling system 106 estimating the effects of spontaneous deamination by generating genotype calls based on determined sets of probabilities — including basecall probabilities, pre-sequencing intermediate base probabilities, and candidate-haplotype- derived consensus base probabilities — in accordance with one or more implementations of the present disclosure.

[0069] As shown in FIG. 3 A, the deamination-aware variant calling system 106 performs an act 302 of identifying duplex consensus reads. In particular, the deamination-aware variant calling system 106 identifies, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads. In some embodiments, the duplex consensus bases comprise at least one methylated cytosine base or deaminated cytosine base on a single strand. FIG. 3A illustrates a collapsed plus single-stranded read 326 and a collapsed minus singlestranded read 328. Together, the collapsed plus single-stranded read 326 and the collapsed minus single-stranded read 328 form a duplex consensus read. As further shown in FIG. 3A, the duplex consensus read comprises a duplex consensus base 312 and a duplex consensus base 314. The duplex consensus base 312 comprises T+C~, which is indicative of a methylated cytosine base or a deaminated cytosine base on a plus single strand. For example, the deaminated cytosine base can indicate a methylated cytosine base on the plus single strand (mC+) or an unmethylated cytosine base on the plus single strand (UC+) that has been converted to a thymine. The duplex consensus base 314 comprises G+A~, which is indicative of a deaminated cytosine base on a minus single strand. For example, the deaminated cytosine base on the minus single strand can indicate a methylated cytosine base on the minus single strand (mC~) or an unmethylated cytosine base on the minus single strand (UC~) that has been converted to a thymine.Attorney Docket No. IP-2816-PCT 21 Patent Application

[0070] By way of overview, different methylation sequencing protocols utilize different conversions. For example, some methylation sequencing protocols convert unmethylated cytosine bases to thymine bases (which are sometimes represented as C-to-T conversions). Other methylation sequencing protocols may convert methylated cytosine bases to thymine bases (which are sometimes represented as 5mC, 5hmC-to-T conversions). The deamination-aware variant calling system 106 may utilize methylation data stemming from either type of methylation sequencing protocol. The following paragraphs further detail various effects of C-to-T conversion protocols and 5mC, 5hmC-to-T conversion protocols, among other conversion protocols.

[0071] In the C-to-T conversion protocols, unmethylated bases are converted into thymine bases. For instance, the C-to-T conversion protocols include Bisulfite and Enzymatic Methyl-seq (EM-seq) used for methylation sequencing assays. EM-seq can be performed, for instance, as described by Romualdas Vaisvila et al., Enzymatic Methyl Sequencing Detects DNA Methylation at Single-Base Resolution from Picograms of DNA, 30 Genome Research 1280-1289 (2021), which is hereby incorporated by reference in its entirety.

[0072] In 5mC, 5hmC-to-T conversion protocols, methylated bases are converted into thymine bases. For instance, Tet-assisted pyridine borane sequencing (TAPS) uses a ten-eleven translocation (TET) enzyme for a methylation sequencing assay, as described by Yibin Liu et al., “Bisulfite-free Direct Detection of 5-Methylcystosine and 5-Hydroxymethylcystosine at Base Resolution,” 36 Nature Biotechnology 424-29 (2019). In some assays that rely on a TET enzyme, a methylation sequencing assay converts 5-Methylcystosine (5mC) and 5-Hydroxymethylcystosine (5hmC) into oxidized products using a TET enzyme and then uses an Apolipoprotein B mRNA Editing Enzyme, Catalytic Polypeptide (APOBEC) 3 A or another APOBEC protein to deaminate unmodified cytosines by converting them to uracil bases. In some cases, such converted uracil bases are detected as thymine bases during sequencing.

[0073] In some embodiments presented herein, the deamination-aware variant calling system 106 uses nucleotide reads (e.g., single-stranded reads) from bisulfite converted samples. In bisulfite sequencing, DNA is chemically treated with sodium bisulfite, which results in the conversion of unmethylated cytosine bases to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, the modified cytosine bases, 5mC and 5hmC, are resistant to bisulfite conversion, and are sequenced as cytosine bases. It will be understood by one of ordinary skill in the art that other methylation conversion methods can also be used to generate nucleotide reads for use in the methods and embodiments presented herein. For example, an alternative to bisulfite conversion is Enzymatic Methyl-seq (EM-seq). EM-seq has been described in neb.com / - / media / nebus / files / manuals / manuale7120.pdf, the content of which is incorporated herein by reference in its entirety. Briefly, EM-seq conversion uses an enzymatic method that results in theAttorney Docket No. IP-2816-PCT 22 Patent Applicationconversion of unmethylated cytosine bases to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, the modified cytosine bases, 5mC and 5hmC, are resistant to the enzymatic conversion, and are sequenced as cytosine bases. Since only a very small fraction of cytosines in a typical sample are methylated, in some embodiments, the result of either bisulfite or EM-seq conversion methods is a set of sequence reads that is substantially made up of only 3 bases: A, T, and G.

[0074] Another example of an alternative method to bisulfite conversion is TET-assisted pyridine borane sequencing (TAPS). TAPS has been described in Liu, et al. Bisulfite-free direct detection of 5 -methylcytosine and 5-hydroxymethylcytosine at base resolution. Nat Biotechnol. 2019 Apr;37(4):424-429. doi: 10.1038 / s41587-019-0041-2, the content of which is incorporated herein by reference in its entirety. TAPS results in conversion of 5mC and 5hmC to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, unmethylated cytosine bases are resistant to the TAPS conversion, and are sequenced as cytosine bases. One of ordinary skill in the art will recognize that, where TAPS conversion is utilized to generate sequence reads, the reference sequences described herein can be modified accordingly to reflect methylated targets where methylated cytosine bases are converted and sequenced as thymine bases (C-to-T conversion) and unmethylated cytosine bases are not converted and sequenced as cytosine bases. An enzymatic alternative to TAPS involves use of a modified cytidine deaminase enzyme, engineered to selectively deaminate only 5mC and 5hmC, while unmethylated cytosine bases are not converted. Similar to TAPS, this modified cytidine deaminase method results in conversion of 5mC and 5hmC to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, unmethylated cytosine bases are resistant to the TAPS conversion, and are sequenced as cytosine bases. Use of modified cytidine deaminase enzymes, such as APOBEC3A, has been described in International Application No. PCT / US2023 / 17846, filed April 7, 2023, and titled “Altered Cytidine Deaminases and Methods of Use,” the contents of which is incorporated herein by reference in its entirety. One of ordinary skill in the art will recognize that since only the methylated cytosines in a typical sample are converted using these alternative conversion methods, in some embodiments, the result of either TAPS or modified cytidine deaminase conversion methods is a set of sequence reads that is substantially made up of 4 bases: A, T, C and G.

[0075] As shown in FIG. 3 A, the deamination-aware variant calling system 106 can use protocols that convert methylated (or unmethylated) cytosine bases to thymine bases. Such conversions can be expressed in duplex consensus reads as T+C~, which is indicative of a methylated cytosine base on a plus single strand (mC+). In examples where unmethylated cytosine bases are converted thymine bases, the duplex consensus reads T+C~ indicate an unmethylated cytosine base on the plus single strand (UC+). In either instance, the T+C~ duplex consensus baseAttorney Docket No. IP-2816-PCT 23 Patent Applicationindicates that the cytosine base on the plus strand has been converted to a thymine base, which is indicative of a deaminated cytosine on the plus single strand. The duplex consensus base 314 comprises G+A~, which is indicative of a methylated cytosine base on a minus single strand (mC~) or an unmethylated cytosine base on the minus single strand (UC~). The G+A~ duplex consensus base indicates that a guanine base on a minus single strand has been converted to an adenine base, which signifies a deaminated cytosine base on the minus single strand.

[0076] As further illustrated in FIG. 3 A, the deamination-aware variant calling system 106 further performs an act 304 of determining a first set of probabilities exhibiting observed duplex consensus bases given converted haplotype intermediate bases. In particular, the deamination- aware variant calling system 106 determines basecall probabilities — that is, a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases that represent a haplotype base converted from one base class to another base class and that comprise one or more deaminated cytosine bases. As shown in FIG. 3A, the first set of probabilities can be expressed aswhere Rtrepresents observed duplex consensus bases and H'j represents a converted haplotype intermediate base.

[0077] FIG. 3A illustrates an example of how Rtrepresents the observed duplex consensus bases. As mentioned previously, the observed duplex consensus bases comprise nucleobases called by a sequencing device for a nucleotide read. For example, and as illustrated in FIG. 3A, the observed duplex consensus bases 320 comprise T+C~ or a thymine base on the collapsed plus single-stranded read and a cytosine base on the collapsed minus single-stranded read.

[0078] As further shown in FIG. 3A and as implied by the abbreviated term basecall probabilities, the deamination-aware variant calling system 106 determines the first set of probabilities of the observed duplex consensus bases given converted haplotype intermediate bases. The converted haplotype intermediate basesgenerally comprises converted haplotype bases represented on a duplex consensus read. The converted haplotype intermediate base includes the descriptor “intermediate” because it comprises a type of unobserved hidden variable. More specifically, converted haplotype intermediate bases may arise from methylation or artifacts during preparation or storage of a genomic sample. For example, FIG. 3A illustrates an original molecule 316 of a genomic sample. The original molecule 316 may have a haplotype base at a genomic coordinate 322. The haplotype base comprises the original base within the original molecule 316 before conversions from methylation assays or other sources.

[0079] The haplotype base can be converted to a converted haplotype intermediate base 318 by several causes. Namely, the converted haplotype intermediate base 318 may arise from spontaneous deamination of a base at the genomic coordinate 322. For example, the base at theAttorney Docket No. IP-2816-PCT 24 Patent Applicationgenomic coordinate 322 may be deaminated as part of C-to-T conversions from a methylation assay. Additionally, the converted haplotype intermediate base 318 can further arise from artifacts during preparation or storage of the genomic sample. A converted haplotype intermediate base cannot be observed because sequencing the genomic sample may introduce sequencing error before outputting the observed duplex consensus bases 320. FIG. 4 and the corresponding paragraphs further detail how the deamination-aware variant calling system 106 determines the basecall probabilities in accordance with one or more embodiments of the present disclosure.

[0080] As further shown in FIG. 3 A, the deamination-aware variant calling system 106 performs an act 306 of determining a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate base given candidate haplotype bases. In particular, the deamination-aware variant calling system 106 determines basecall probabilities — that is, a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate base given candidate haplotype bases at the genomic coordinate. As shown in FIG. 3 A, the second set of probabilities can be expressed as P(HJ |Hy), where H- represents a converted haplotype intermediate base and Hj represents a candidate haplotype base. As previously mentioned, the converted haplotype intermediate base (H- ) comprises converted bases represented on duplex consensus reads.

[0081] FIG. 3 A illustrates an example of how Hj represents a candidate haplotype base. For instance, and as illustrated in FIG. 3A, a candidate haplotype base 324 comprises candidate base classes at the genomic coordinate. More specifically, the candidate haplotype base represents the base of a candidate haplotype being considered. For example, the deamination-aware variant calling system 106 considers converted haplotype intermediate bases and tests them against candidate haplotype bases that are being considered. FIG. 5 and the corresponding discussion further detail how the deamination-aware variant calling system 106 determines the second set of probabilities in accordance with one or more embodiments of the present disclosure.

[0082] Based on the first and second sets of probabilities — and as illustrated in the following figure labelled as FIG. 3B — the deamination-aware variant calling system 106 performs an act 308 of generating a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases given the candidate haplotype bases. In particular, based on the basecall probabilities and the pre-sequencing intermediate base probabilities, the deamination-aware variant calling system 106 generates, utilizing a variant call model, candidate-haplotype-derived consensus base probabilities — that is, a third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases. The third set of probabilities can be expressed asrepresents the observedAttorney Docket No. IP-2816-PCT 25 Patent Applicationduplex consensus bases and Hj represents candidate haplotype bases. In particular, the deamination-aware variant calling system 106 considers the observed duplex consensus bases from a sequencing device and tests them against the candidate haplotype bases.

[0083] As further illustrated in FIG. 3B, the deamination-aware variant calling system 106 further performs an act 310 of generating a genotype call. In particular, the deamination-aware variant calling system 106 generates, based on posterior genotype probabilities derived from the third set of probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate. For instance, the genotype call may indicate that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate(s). In one or more implementations, the deamination-aware variant calling system 106 generates the genotype call for the genomic sample by determining a predicted combination of nucleobases corresponding to a highest priority genotype probability.

[0084] As mentioned, the deamination-aware variant calling system 106 can determine basecall probabilities. FIG. 4 illustrates the deamination-aware variant calling system 106 determining a first set of probabilities of a target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases in accordance with one or more implementations of the present disclosure.

[0085] As illustrated in FIG. 4, the deamination-aware variant calling system 106 performs an act 402 of determining observed duplex consensus bases at a genomic coordinate. As described previously, the deamination-aware variant calling system 106 can determine an observed duplex consensus base using duplex sequencing. For example, and as illustrated in FIG. 4, the deamination-aware variant calling system 106 identifies nucleobases 408 at a genomic coordinate for individual nucleotide reads. More specifically, the nucleobases 408 can originate from plus single strands or minus single strands (e.g., a single-stranded plus read or a single-stranded minus read). The deamination-aware variant calling system 106 collapses the nucleobases 408 from plus single strands into a collapsed base 406a for a single-stranded plus read and a collapsed base 406b for a single-stranded minus read. The resulting observed duplex consensus base thus comprises T+C~.

[0086] As mentioned previously, the deamination-aware variant calling system 106 recognizes new duplex consensus bases. FIG. 4 illustrates duplex consensus bases recognized by the deamination-aware variant calling system 106. As shown in FIG. 4, the deamination-aware variant calling system 106 may associate certain notations with duplex consensus bases. For example, the deamination-aware variant calling system 106 associates an A notation for an A+A~ duplex consensus base, a C notation for a C+C~ duplex consensus base, a G notation for a G+G~ duplex consensus base, and a T notation for a T+T~ duplex consensus base. The deamination-awareAttorney Docket No. IP-2816-PCT 26 Patent Applicationvariant calling system 106 associates new notationsmC+andmC~ indicating a methylated cytosine base on a plus single strand and a minus single strand, respectively. More specifically, the deamination-aware variant calling system 106 associates themC+notation with a T+C~ duplex consensus base. The deamination-aware variant calling system 106 associates themC~ notation with a G+A~ duplex consensus base.

[0087] FIG. 4 further illustrates the deamination-aware variant calling system 106 performing an act 404 of determining a first set of probabilities of a target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases. Because the converted haplotype intermediate bases are unobserved and unknown, the deamination-aware variant calling system 106 tests different converted haplotype intermediate bases and determines the basecall probabilities based on estimated error probabilities (e).

[0088] As shown in FIG. 4, the deamination-aware variant calling system 106 uses various equations to determine the basecall probabilitieswhere Rtrepresents observed duplex consensus bases and H'j represents converted haplotype intermediate bases.In function (1), e+represents a plus-strand error probability that a collapsed base on a collapsed plus single-stranded read is an error, e~ represents a minus-strand error probability that a collapsed base on a collapsed minus single-stranded read is an error. In particular, e+and e~ account for base calling errors by a sequencing device during a sequencing run. The deamination-aware variant calling system 106 can use the combination of e+and e~ to estimate an error probability that an observed duplex consensus base is an error. Such an estimate of error probability is a type of base- call-quality score. An observed duplex consensus base can accordingly have two base-call-quality scores. In some implementations, the deamination-aware variant calling system 106 determines e+and e~ based on base-call quality metrics (e.g., BASEQ score). In one example, e+and e~ equal 1O_BASEQ / IO £or t e conapSecipjussingle-stranded read or the collapsed minus single-stranded read, respectively.Attorney Docket No. IP-2816-PCT 27 Patent Application

[0089] As further shown in function (1), when the observed duplex consensus base (R equals the converted haplotype intermediate base (Hy) — and when the converted haplotype intermediate base comprises A or C — the deamination-aware variant calling system 106 determines the probability of the observed duplex consensus bases using the following function: l-y-4yy. But when the observed duplex consensus base (R equals the converted haplotype intermediate base (Hy) — and when the converted haplotype intermediate base comprises G or T — the deamination- aware variant calling system 106 determines the probability of the observed duplex consensus bases using the following function:

[0090] As indicated above and as relevant to function (1), in some implementations, the deamination-aware variant calling system 106 recognizes or represents two new observed duplex consensus bases (R;),mC+andmC~ . More specifically, the observed duplex consensus basemC+represents a methylated cytosine base on a plus strand. As mentioned previously, the observed duplex consensus basemC+comprises duplex consensus base T+C~. The observed duplex consensus basemC~ represents a methylated cytosine base on a minus strand and comprises the duplex consensus base G+A~. The deamination-aware variant calling system 106 uses function (1) to estimate different sources for the observed duplex consensus bases (R;),mC+andmC~. More specifically, the deamination-aware variant calling system 106 estimates that the duplex consensus bases (R;),mC+andmC~ came (or were derived) from a methylated cytosine base, a deaminated cytosine base, or a sequencing error. More specifically, the deamination-aware variant calling system 106 can use e+and e~ values to estimate sequencing errors and, accordingly, infer effects of cytosine base methylation and / or cytosine base deamination.

[0091] As further shown by function (1), when the observed duplex consensus base (R equals the converted haplotype intermediate base (Hy) — and when the converted haplotype intermediate base comprisesmC+or —the deamination-aware variant calling system 106 determines the probability of observed duplex consensus bases using the following function:y y Insome implementations, the deamination-aware variant calling system 106 may mark the new R; bases as C or G and use a methylation tag to distinguishmC+from a C ormC~ from a G in a duplex consensus read. The deamination-aware variant calling system 106 may further add secondary simplex base-call-quality scores (e.g., Q scores) as a non-standard tag in a duplex consensus read.

[0092] Function (1) further shows how the deamination-aware variant calling system 106 can determine the probability of observed duplex consensus bases based on specific combinations of the observed duplex consensus base (R and the converted haplotype intermediate baseTo illustrate, if E {(A,mC“), (mC“, A), (C,mC+), (mC+, C)}, the deamination-aware variantAttorney Docket No. IP-2816-PCT 28 Patent Applicationcalling system 106 determines the probability of observed duplex bases using the following function: y(l-e“). Additionally, ifG {(G,mC“), (mC“, G), (T,mC+), (mC+, T)}, the deamination-aware variant calling system 106 determines the probability of observed duplex bases using the following function: y (l-e+). As further shown by function (1), if none of the prior conditions are satisfied, the deamination-aware variant calling system 106 determines the probability of observed duplex bases using the following function: y y.

[0093] As previously described, the deamination-aware variant calling system 106 can further determine a second set of probabilities, which this disclosure also calls pre-sequencing intermediate base probabilities. FIG. 5 illustrates the deamination-aware variant calling system 106 determining a second set of probabilities of a target genomic sample exhibiting converted haplotype intermediate bases given candidate haplotype bases in accordance with one or more implementations.

[0094] As illustrated in FIG. 5, the deamination-aware variant calling system 106 considers converted haplotype intermediate bases and determines probabilities that such intermediate bases are present at a genomic coordinate in an approach that tests such bases against candidate haplotype bases. The deamination-aware variant calling system 106 attempts to align converted haplotype intermediate bases with the haplotype bases in various ways to identify possible combinations. FIG. 5 illustrates a number of equations for determining a probability of a target genomic sample exhibiting a converted haplotype intermediate base given various candidate haplotype bases.

[0095] As illustrated for embodiments shown in FIG. 5, the deamination-aware variant calling system 106 determines duplex-consensus base-conversion probabilities that a haplotype duplex consensus base converts to a converted haplotype intermediate duplex consensus base. The deamination-aware variant calling system 106 can determine a duplex-consensus base-conversion probability based on a collapsed-plus-strand base conversion probability approximating a likelihood of a haplotype base converting from one base class to another base class on a collapsed plus single-stranded read and a collapsed minus-strand base-conversion probability approximating a likelihood of a haplotype base converting from one base class to another base class on a collapsed minus single-stranded read.

[0096] For example, the deamination-aware variant calling system 106 can determine a probability of a converted haplotype intermediate base given a candidate haplotype base A using the following equation:Attorney Docket No. IP-2816-PCT 29 Patent ApplicationIn function (2), H' represents a converted haplotype intermediate base, H represents a candidate haplotype base of adenine (A), and a represents base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class. To illustrate, ^dupiexsignjfies aduplex-consensus base-conversion probability that approximates the likelihood of a haplotype adenine base (A+A“) stays as an adenine base (A+A“). The deamination-aware variant calling system 106 can determine that a^Aexis some value close to 1 as having a high likelihood. By contrast, the duplex-consensus base-conversion probability aA^lex, which approximates the likelihood of a haplotype base converting from an adenine base to a cytosine converted haplotype intermediate base, is a much lower value as the haplotype base converting from an A to a C would require conversions or errors on both the plus strand and the minus strand. More specifically, a^Gexcomprises the likelihood of a haplotype base A+A~ converting to a converted haplotype intermediate base C+C~. In some embodiments, the deamination-aware variant calling system 106 determines that aA^exand aA^lexare, likewise, small values.

[0097] As mentioned previously, relative to existing sequencing systems, the deamination- aware variant calling system 106 recognizes or encodes for at least two new converted haplotype intermediate bases:mC+■= T+C~ andmC~ ■= G+A~ . The converted haplotype intermediate basemC+represents a methylated cytosine base on a plus strand. The converted haplotype intermediate basemC~ represents a methylated cytosine base on a minus strand. Instead of assigning a base- call-quality score of zero to these new converted haplotype intermediate bases, the deamination- aware variant calling system 106 increases the size of the lookup table. As further shown in FIG. 5, the deamination-aware variant calling system 106 also considers duplex-consensus baseconversion probabilities that approximate a likelihood of a haplotype base converting to a methylated cytosine on a collapsed plus single-stranded read or a collapsed minus single-stranded read. For instance, the deamination-aware variant calling system 106 can determine baseconversion probabilities using the following equations: As further shown in function (2),represents a duplex-consensus base-conversion probability approximating a likelihood of a haplotype base A (or the duplex consensus base A+A“) converting to the converted haplotypeAttorney Docket No. IP-2816-PCT 30 Patent Applicationintermediate basemC+(or the duplex consensus base T+C“). The term simplex in FIG. 5 and the equation above denotes a collapsed single-stranded read. By contrast, a^^+xrepresents a collapsed-plus-strand base-conversion probability of a haplotype base converting from an A+to a T+on a collapsed plus single-stranded read; and a^™pc-xrepresents a collapsed-minus-strand base-conversion probability of a haplotype base converting from an A~ to a C~ on a collapsed minus single-stranded read.

[0098] In addition to examples of collapsed-plus-strand and collapsed-minus-strand baseconversion probabilities, the equation above also includes some examples of duplex-consensus base-conversion probabilities and how the latter relate to such collapsed-strand base-conversion probabilities. In particular,represents a duplex-consensus base-conversion probability approximating a likelihood of a haplotype base A (or the duplex consensus base A+A“) converting to the converted haplotype intermediate basemC~ (or the duplex consensus base G+A-)- As shown, this conversion requires only the change from the collapsed base A+to G+on the collapsed plus single-stranded read. Accordingly, a™^rG+xrepresents a collapsed-plus-strand base-conversion probability of a haplotype base converting from an A+to G+on the collapsed plus single-stranded read.

[0099] In some implementations, the deamination-aware variant calling system 106 determines the duplex-consensus base-conversion probabilities, the collapsed-plus-strand baseconversion probabilities, and the collapsed-minus-strand base-conversion probabilities for a target genomic sample. In particular, the deamination-aware variant calling system 106 predicts a on a genomic-sample-wide basis. The deamination-aware variant calling system 106 can predetermine a values for a genomic sample and store the duplex-consensus base-conversion probability (cr) values in a lookup table. More specifically, the deamination-aware variant calling system 106 can determine duplex-consensus base-conversion probabilities specific to a target genomic sample.

[0100] In some implementations, the deamination-aware variant calling system 106 can use the genomic-sample-wide duplex-consensus base-conversion probability to determine an estimated methylation-level value ( ). For instance, the deamination-aware variant calling system 106 can estimate the methylation-level value (0) based on both the genomic-sample- wide duplex-consensus base-conversion probability (cr) and genomic-coordinate-specific parameters. To illustrate, the deamination-aware variant calling system 106 can estimate the methylation-level value (0) based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, a probability of each nucleobase at the genomic coordinate on collapsed plus single-stranded reads or minus single-stranded reads, and genomic-sample-wide duplex-consensus base-conversionAttorney Docket No. IP-2816-PCT 31 Patent Applicationprobabilities (cr). Accordingly, the methylation-level value ( ) can refer to a genomic coordinate and not an entire nucleotide read.

[0101] The paragraphs above describe how the deamination-aware variant calling system 106 determines the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases when the candidate haplotype base is adenine. The deamination-aware variant calling system 106 can estimate additional sets of probabilities for candidate haplotype bases guanine, cytosine, and thymine using the following respective equations:In functions (3), (4), and (5), / ?+represents an estimated methylation-level value for a collapsed plus single-stranded read, and f represents an estimated methylation-level value for a collapsed minus single-stranded read. In contrast to the base-conversion probabilities (cr), which the deamination-aware variant calling system 106 estimates on a genomic-sample- wide basis, theAttorney Docket No. IP-2816-PCT 32 Patent Applicationdeamination-aware variant calling system 106 determines methylation-level values specific to genomic position. In particular, the deamination-aware variant calling system 106 estimates methylation-level values f>+for each cytosine coordinate within the target genomic sample.

[0102] As further shown in FIG. 5, the deamination-aware variant calling system 106 determines the collapsed-plus-strand base-conversion probability that a cytosine base has converted to a thymine base (a^^,+x) based on an estimated methylation-level value f / ?+) for a collapsed plus single-stranded read. In another example, the deamination-aware variant calling system 106 determines a collapsed-minus-strand base-conversion probability that a guanine base has converted to an adenine basebased on an estimated methylation-level value for a collapsed minus single-stranded read ( / ?_).

[0103] Additionally, or alternatively, in some implementations, the deamination-aware variant calling system 106 determines an empirical duplex-consensus base-conversion probability for a genomic sample. In particular, instead of determining duplex-consensus base-conversion probabilities based on plus-strand and minus-strand error probabilities, in some implementations, the deamination-aware variant calling system 106 empirically counts the number of occurrences ofmC+(T+C~) andmC~ (G+A~) duplex consensus bases. For instance, the deamination-aware variant calling system 106 determines a first set of duplex consensus bases comprising a thymine on the collapsed plus single-stranded read and a cytosine on the collapsed minus single-stranded read. The deamination-aware variant calling system 106 further determines a second set of duplex consensus bases comprising a guanine on the collapsed plus single-stranded read and a cytosine on the collapsed minus single-stranded read. Based on the observed number ofmC+(T+C~) andmC~ (G+A~) duplex consensus bases, the deamination-aware variant calling system 106 estimates an empirical duplex-consensus base-conversion probability.

[0104] In some embodiments, the deamination-aware variant calling system 106 generates the estimated methylation-level value for the collapsed plus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed plus single-stranded reads. Additionally, the deamination-aware variant calling system 106 generates the estimated methylation-level value for the collapsed minus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed minus single-stranded reads. In some implementations, the deamination- aware variant calling system 106 determines the estimated methylation-level values in accordance with methods described in International Application No. PCT / US2024 / 035562, entitled “VariantAttorney Docket No. IP-2816-PCT 33 Patent ApplicationCalling with Methylation-Level Estimation,” filed June 26,2024 (IP-2578-PCT), the disclosure of which is incorporated herein by reference in its entirety.

[0105] Additionally, and as mentioned, in some implementations, the deamination-aware variant calling system 106 generates the estimated methylation-level value based on a combination of genomic-coordinate-specific parameters and genomic-sample-wide duplex consensus baseconversion probabilities (cr). More specifically, the deamination-aware variant calling system 106 can determine a genomic-sample- wide plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylation-level value for a collapsed plus single-stranded read and determine a genomic-sample-wide minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylationlevel value for a collapsed minus single-stranded read. The deamination-aware variant calling system 106 generates the estimated methylation-level value for the collapsed plus single-stranded read based on a combination of the genomic-sample- wide plus-strand base-conversion probability, numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed plus single-stranded reads. The deamination-aware variant calling system 106 can further generate the estimated methylation-level value for the collapsed minus single-stranded read based on a combination of the genomic-sample- wide minus-strand base-conversion probability, numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed minus single-stranded reads. The deamination-aware variant calling system 106 can determine the methylation-level values ( ) by using a Bayesian computation and utilizing the genomic-sample-wide duplex-consensus base-conversion probabilities as priors.

[0106] As mentioned previously, the deamination-aware variant calling system 106 can generate a genotype call based on posterior genotype probabilities derived from candidate- haplotype-derived consensus base probabilities. FIG. 6 illustrates an overview of the deamination- aware variant calling system 106 generating a genotype call in accordance with one or more implementations of the present disclosure. By way of overview, FIG. 6 illustrates a series of acts 600 by which the deamination-aware variant calling system 106 generates candidate-haplotype- derived consensus base probabilities, generates posterior genotype probabilities, and, based thereon, generates a genotype call.

[0107] As shown in FIG. 6, the deamination-aware variant calling system 106 performing an act 602 of generating candidate-haplotype-derived consensus base probabilities — that is, a third set of probabilities of a target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given candidate haplotype bases. More specifically, the deamination-aware variant calling system 106 uses a variant call model to generate the candidate-haplotype-derivedAttorney Docket No. IP-2816-PCT 34 Patent Applicationconsensus base probabilities based on the basecall probabilities and the pre-sequencing intermediate base probabilities. In some implementations, the deamination-aware variant calling system 106 uses the following equation to generate such candidate-haplotype-derived consensus base probabilities or third set of probabilities of the target genomic sample exhibiting observed duplex consensus bases given the candidate haplotype bases P( / ?£| F / ) :In function (6),represents the first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases,represents the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases, and H' represents converted haplotype intermediate bases.

[0108] FIG. 6 further illustrates the deamination-aware variant calling system 106 performing an act 604 of determining a probability of the observed duplex consensus bases given a genotype at the genomic coordinate. The deamination-aware variant calling system 106 can determine severalvalues for a read pileup at a given genomic coordinate or genomic region. In some implementations, the deamination-aware variant calling system 106 aggregates the P( / ? H / ) into a probability of an observed base given a genotype at the genomic coordinate. For example, the probability of an observed base given a genotype can be expressed as P(Ri \9k where Rtrepresents an observed nucleobase and Qkrepresents the genotype at the genomic coordinate.

[0109] As further shown in FIG. 6, the deamination-aware variant calling system 106 performs an act 606 of generating posterior genotype probabilities. In some embodiments, the deamination- aware variant calling system 106 uses a variant call model to generate posterior genotype probabilities (P(^fcl^)) for the target genomic sample at the genomic coordinate based on the candidate-haplotype-derived consensus base probabilities (or third set of probabilities). In some embodiments, the deamination-aware variant calling system 106 further inputs base-call quality metrics and sequencing metrics into the variant call model to generate the posterior genotype probabilities.

[0110] In one or more embodiments, as part of performing the act 606, the deamination-aware variant calling system 106 inputs the probabilities represented byor P(Ri \ k) intoavariant caller. In some embodiments, the variant caller collapses the probabilities into a probability of the observed read pileup represented as P(D | < / k). The deamination-aware variant calling system 106 can further utilize the variant caller to invert the probability of the observed read pileup toAttorney Docket No. IP-2816-PCT 35 Patent Applicationgenerate posterior genotype probability P(Qk\D), where D represents the data or the observed read pileup.[OHl] FIG. 6 also illustrates the deamination-aware variant calling system 106 performing an act 608 of generating a genotype call. In particular, the deamination-aware variant calling system 106 generates a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate. For instance, the genotype call may indicate that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate(s). In one or more implementations, the deamination-aware variant calling system 106 generates the genotype call for the target genomic sample by determining a predicted combination of nucleobases corresponding to a highest posterior genotype probability.

[0112] The deamination-aware variant calling system 106 can leverage the posterior genotype probabilities to generate the genotype call according to P(< / (|O) in at least a couple of ways. As indicated above, the deamination-aware variant calling system 106 can determine, from among the posterior genotype probabilities according to P(< / (|O) for a given genomic coordinate, a highest posterior genotype probability as the genotype call for the given genomic coordinate. Such genotype calls outperform the accuracy of existing sequencing and methylation detection systems and approaches the accuracy of non-methylated whole genome sequencing. In addition to genotype calls, the deamination-aware variant calling system 106 can leverage the posterior genotype probabilities according to P(< / (|O) to determined refined (and sometimes more accurate) methylation-level values for candidate cytosine bases at target genomic coordinates.

[0113] As mentioned, in contrast to existing systems that either rely exclusively on data for single-stranded reads or assign a zero quality score (or otherwise ignore) mismatched bases on duplex consensus reads, the deamination-aware variant calling system 106 can improve accuracy by (i) analyzing duplex consensus bases and (ii) uniquely encoding converted haplotype intermediate bases in such duplex consensus bases to estimate effects of deamination events. More specifically, the deamination-aware variant calling system 106 reduces false negatives for somatic variants. FIGS. 7A-8B illustrate box plots reflecting simulated data illustrating how the deamination-aware variant calling system 106 improves accuracy relative to existing systems. In particular, FIGS. 7A and 7B illustrate improvements in accuracy by the deamination-aware variant calling system 106 calling C-to-T somatic variants relative to existing sequencing systems in accordance with one or more implementations of the present disclosure. FIGS. 8 A and 8B illustrate improvements in accuracy by the deamination-aware variant calling system 106 calling T-to-C somatic variants relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.Attorney Docket No. IP-2816-PCT 36 Patent Application

[0114] The box plots illustrated in FIGS. 7A-8B portray varying levels of variant allele frequency (VAF) for respective somatic variants within simulated data. The simulated data portrayed within the box plots illustrated in FIGS. 7A-8B are at 2000x coverage. As shown, all box plots illustrated in FIGS. 7A-8B have an x-axis correlating with duplex rate and ay-axis correlating with variant call accuracy. The x-axis duplex rate indicates a proportion of reads in a given pileup that are duplex consensus reads. To illustrate, a duplex rate of 0.0, which is shown at the left end of all box plots in FIGS. 7A-8B, indicates that only collapsed single-stranded plus reads or collapsed single-stranded minus reads have been captured. The box plots at the duplex rate of 0.0 are representative of the performance of existing systems that process only collapsed singlestranded reads and not relationships between duplex consensus bases like the deamination-aware variant calling system 106.

[0115] To further illustrate the axes of FIGS. 7A-8B, on the opposite end of the x-axis, a duplex rate of 1.0 indicates that 100% of the read pileup comprises duplex consensus reads capturing both collapsed plus single-stranded reads and their complement collapsed minus single-stranded reads. Box plots at the duplex rate of 1.0 indicate the performance of deamination-aware variant calling system 106 which is able to accurately analyze and use data from duplex consensus bases to estimate spontaneous deamination events. More specifically, and as shown in FIGS. 7A-8B, the deamination-aware variant calling system 106 improves variant call accuracy as the duplex rate increases because the deamination-aware variant calling system 106 can better process and analyze duplex consensus bases by encoding converted haplotype intermediate bases in such duplex consensus bases and assigning corresponding base-call-quality scores.

[0116] Additionally, the box plots in FIGS. 7A-8B compare the performance of the deamination-aware variant calling system 106 with a baseline variant caller. The baseline variant caller illustrated in FIGS. 7A-8B analyzes simple chemistry that does not include methylation of nucleotide reads by a methylation sequencing assay. The baseline variant caller accounts for duplex bases by (i) representing certain intermediate bases and (ii) using one set of errors for collapsed bases on collapsed single-stranded reads and another set of errors for duplex consensus bases on duplex nucleotide reads. But the baseline variant caller does not recognize (or allow for) T+C- duplex consensus bases or model T+C- intermediate bases. Instead, when faced with T+C- duplex consensus bases, the baseline variant caller collapses T+C- duplex consensus bases on duplex nucleotide reads to an N-base and assigns a zero-quality score. Because the baseline variant caller does not deal with the complications of calling variants with methylation, the baseline variant caller illustrated in FIGS. 7A-8B is relatively accurate. In some examples, existing sequencing systems fail to accurately predict both methylation-level values and genotype calls. FIGS. 7A-8B illustrateAttorney Docket No. IP-2816-PCT 37 Patent Applicationthat the deamination-aware variant calling system 106 can achieve competitive performance with a simple and accurate non-methylated pipeline baseline variant caller.

[0117] As mentioned, FIGS. 7A-7B illustrate improvements in accuracy by the deamination- aware variant calling system 106 calling C-to-T somatic variants. Box plots 700 in FIG. 7A portray the accuracy of variant calls generated by the deamination-aware variant calling system 106 for a genomic sample with C-to-T somatic variants at 0.5% VAF. Box plots 704 in FIG. 7B portray the accuracy of variant calls generated by the deamination-aware variant calling system 106 for a genomic sample with C-to-T somatic variants at 0.2% VAF. The box plots 700 and box plots 704 illustrated in FIGS. 7A-7B, respectively, also show increasing beta_plus_truth values or true methylation-level values ( / ?+) for collapsed plus single-stranded reads. FIGS. 7A-7B illustrate true methylation-level values of 0.0, 0.05, 0.5, and 0.95. The box plots 700 illustrate baseline variant caller performance boxes 702. While FIG. 7A illustrates the baseline variant caller performance boxes 702 in only the first box plot, the baseline variant caller performance boxes 702 can be carried over to all the box plots 700. Similarly, baseline variant caller performance boxes 706 in FIG. 7B can be carried over to all the box plots 704.

[0118] As mentioned, the box plots 700 and the box plots 704 in FIGS. 7A-7B, respectively, show the performance of existing sequencing systems that analyze only collapsed single-stranded reads (duplex rate = 0.0) as compared to the performance of the deamination-aware variant calling system 106 that analyzes data from duplex consensus reads (duplex rate = 1.0) by determining the basecall probabilities, pre-sequencing intermediate base probabilities, and candidate-haplotype- derived consensus base probabilities described above. As shown in FIGS. 7A-7B, the deamination- aware variant calling system 106 improves the variant calling accuracy of C-to-T somatic variants to be competitive with a simple methylation unaware baseline variant caller with diminishing returns at high methylation-level values.

[0119] FIGS. 8A-8B illustrate improvements in accuracy by the deamination-aware variant calling system 106 calling T-to-C somatic variants. Box plots 800 in FIG. 8A portray the accuracy of variant calls generated by the deamination-aware variant calling system 106 for a genomic sample with T-to-C somatic variants at 0.5% VAF. Box plots 804 in FIG. 8B portray the accuracy of variant calls generated by the deamination-aware variant calling system 106 for a genomic sample with T-to-C somatic variants at 0.2% VAF. The box plots 800 and box plots 804 illustrated in FIGS. 8A-8B, respectively, also show increasing beta_plus_truth values or true methylationlevel values ( / ?+) for collapsed plus single-stranded reads. FIGS. 8A-8B illustrate true methylationlevel values of 0.0, 0.05, 0.5, and 0.95. The box plots 800 illustrate baseline variant caller performance boxes 802. While FIG. 8A illustrates the baseline variant caller performance boxes 802 in only the first box plot, the baseline variant caller performance boxes 802 can be carried overAttorney Docket No. IP-2816-PCT 38 Patent Applicationto all the box plots 800. Similarly, baseline variant caller performance boxes 806 in FIG. 8B can be carried over to all the box plots 804.

[0120] As mentioned, the box plots 800 and the box plots 804 in FIGS. 8A-8B, respectively, show the performance of existing systems that analyze only collapsed single-stranded reads (duplex rate = 0.0) or fail to recognize T+C- and G+A- duplex consensus bases on duplex nucleotide reads as compared to the performance of the deamination-aware variant calling system 106 that analyzes data from duplex consensus reads (duplex rate = 1.0) by determining the basecall probabilities, presequencing intermediate base probabilities, and candidate-haplotype-derived consensus base probabilities described above. As shown in FIGS. 8A-8B, the deamination-aware variant calling system 106 improves the variant calling accuracy of C-to-T somatic variants to be competitive with a simple methylation unaware baseline variant caller with diminishing returns at high methylationlevel values.

[0121] Turning now to FIG. 9, this figure illustrates an example flowchart of a series of acts for generating a genotype call based on duplex consensus base data in accordance with one or more embodiments of the present disclosure. While FIG. 9 illustrates acts according to particular embodiments, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 9. The acts of FIG. 9 can be performed as part of a method. Alternatively, a non- transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device to perform the acts depicted in FIG. 9. In still further embodiments, a system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 9.

[0122] As shown in FIG. 9, the series of acts 900 includes an act 902 of identifying duplex consensus reads, an act 904 of determining a first set of probabilities, an act 906 of determining a second set of probabilities, an act 908 of generating, utilizing a variant caller and based on the first set of probabilities and the second set of probabilities, a third set of probabilities, and an act 910 of generating a genotype call.

[0123] For example, the series of acts 900 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A computer-implemented method comprising: identifying, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads, wherein the duplex consensus bases comprise at least one methylated cytosine base or deaminated cytosine base on a single strand;Attorney Docket No. IP-2816-PCT 39 Patent Applicationdetermining a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases that represent a haplotype base converted from one base class to another base class and that comprise the at least one methylated cytosine base or deaminated cytosine base; determining a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases at the genomic coordinate; generating, utilizing a variant call model and based on the first set of probabilities and the second set of probabilities, a third set of probabilities of the target genomic sample exhibiting the observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases; and generating, based on posterior genotype probabilities derived from the third set of probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate.CLAUSE 2. The computer-implemented method of clause 1, wherein the converted haplotype intermediate bases comprise a haplotype intermediate base representing a methylated cytosine base on a plus strand or a haplotype intermediate base representing a methylated cytosine base on a minus strand.CLAUSE 3. The computer-implemented method of clause 1 or 2, wherein an observed duplex consensus base of the observed duplex consensus bases comprises an observed duplex consensus base representing a methylated cytosine base on a plus strand or an observed duplex consensus base representing a methylated cytosine base on a minus strand.CLAUSE 4. The computer-implemented method of any of clauses 1-3, further comprising determining the first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases by: determining a plus-strand error probability that a collapsed base on a collapsed plus singlestranded read is an error; determining a minus-strand error probability that a collapsed base on a collapsed minus single-stranded read is an error; and determining the first set of probabilities of the target genomic sample based on the plusstrand error probability and the minus-strand error probability.Attorney Docket No. IP-2816-PCT 40 Patent ApplicationCLAUSE 5. The computer-implemented method of any of clauses 1-4, further comprising determining the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining, for the target genomic sample, a first set of duplex consensus bases comprising a thymine on a collapsed plus single-stranded read and a cytosine on a collapsed minus single-stranded read; determining, for the target genomic sample, a second set of duplex consensus bases comprising a guanine on the collapsed plus single-stranded read and a cytosine on the collapsed minus single-stranded read; and generating an empirical duplex-consensus base-conversion probability based on the first set of duplex consensus bases and the second set of duplex consensus bases.CLAUSE 6. The computer-implemented method of any of clauses 1-5, further comprising determining the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining collapsed-plus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed plus single-stranded reads; determining collapsed-minus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed minus single-stranded reads; generating duplex-consensus base-conversion probabilities based on the collapsed-plus- strand base-conversion probabilities and the collapsed-minus-strand base-conversion probabilities; and determining the second set of probabilities based on the duplex-consensus base-conversion probabilities.CLAUSE 7. The computer-implemented method of any of clauses 1-6, wherein the duplexconsensus base-conversion probabilities are specific to the target genomic sample.CLAUSE 8. The computer-implemented method of any of clauses 1-6, further comprising: determining, at the genomic coordinate, a collapsed-plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylationlevel value for a collapsed plus single-stranded read; andAttorney Docket No. IP-2816-PCT 41 Patent Applicationdetermining, at the genomic coordinate, a collapsed-minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylationlevel value for a collapsed minus single-stranded read.CLAUSE 9. The computer-implemented method of any of clauses 1-8, further comprising: generating the estimated methylation-level value for the collapsed plus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed plus singlestranded reads; and generating the estimated methylation-level value for the collapsed minus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed minus singlestranded reads.CLAUSE 10. The computer-implemented method of any of clauses 1-9, further comprising: determining a genomic-sample-wide plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylation-level value for a collapsed plus single-stranded read; determining a genomic-sample-wide minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylation-level value for a collapsed minus single-stranded read; generating the estimated methylation-level value for the collapsed plus single-stranded read further based on the genomic-sample-wide plus-strand base-conversion probability; and generating the estimated methylation-level value for the collapsed minus single-stranded read further based on the genomic-sample-wide minus-strand base-conversion probability.

[0124] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.Attorney Docket No. IP-2816-PCT 42 Patent Application

[0125] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.

[0126] SBS can utilize nucleotide monomers that have a terminator moiety or those that lack any terminator moieties. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using y-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is generally variable and dependent upon the template sequence and the mode of nucleotide delivery. For SBS techniques that utilize nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditions used as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator can be reversible as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).

[0127] SBS techniques can utilize nucleotide monomers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; or the like. In embodiments where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods developed by Solexa (now Illumina, Inc.).

[0128] Preferred embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on realtime pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted toAttorney Docket No. IP-2816-PCT 43 Patent Applicationadenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C or G). Images obtained after addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed and analyzed using the methods set forth herein. For example, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different detection channels for reversible terminatorbased sequencing methods.

[0129] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, in WO 04 / 018497 and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference. This approach is being commercialized by Solexa (now Illumina Inc.), and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators in which both the termination can be reversed, and the fluorescent label cleaved facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from these modified nucleotides.

[0130] Preferably in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured following incorporation of labels into arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show nucleic acid features that have incorporated nucleotides of a particular type. Different features are present or absent in the different images due the different sequence content of each feature. However, the relative position of the features will remain unchanged in the images. Images obtained from such reversible terminator- SBS methods can be stored, processed and analyzed as set forth herein. Following the image capture step, labels can be removed, and reversible terminator moieties can be removed for subsequent cycles of nucleotideAttorney Docket No. IP-2816-PCT 44 Patent Applicationaddition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.

[0131] In particular embodiments some or all of the nucleotide monomers can include reversible terminators. In such embodiments, reversible terminators / cleavable fluors can include fluor linked to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3' allyl group to block extension but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30 second exposure to long wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of natural termination that ensues after placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.

[0132] Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.

[0133] Some embodiments can utilize detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detectedAttorney Docket No. IP-2816-PCT 45 Patent Applicationfor the other member of the pair. As a second example, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples, is a fluorescent-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).

[0134] Further, as described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.

[0135] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained following treatment of an array of nucleic acid features with the labeled sequencing reagents. Each image will show nucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based sequencing methods can be stored, processed and analyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No.Attorney Docket No. IP-2816-PCT 46 Patent Application6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.

[0136] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147- 151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed and analyzed as set forth herein. In particular, the data can be treated as an image in accordance with the exemplary treatment of optical images and other images that is set forth herein.

[0137] Some embodiments can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). The illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-modeAttorney Docket No. IP-2816-PCT 47 Patent Applicationwaveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed and analyzed as set forth herein.

[0138] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.

[0139] The above SBS methods can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows convenient delivery of sequencing reagents, removal of unreacted reagents and detection of incorporation events in a multiplex manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as, bridge amplification or emulsion PCR as described in further detail below.

[0140] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.

[0141] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprisingAttorney Docket No. IP-2816-PCT 48 Patent Applicationcomponents such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeqTM platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference. The sequencing system described above sequences nucleic acid polymers present in samples received by a sequencing device, as described further above.

[0142] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one orAttorney Docket No. IP-2816-PCT 49 Patent Applicationmore target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.

[0143] The components of the deamination-aware variant calling system 106 can include software, hardware, or both. For example, the components of the deamination-aware variant calling system 106 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the client device 114, the local device 108, or the server device(s) 110). When executed by the one or more processors, the computer-executable instructions of the deamination-aware variant calling system 106 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the deamination-aware variant calling system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the deamination-aware variant calling system 106 can include a combination of computer-executable instructions and hardware.

[0144] Furthermore, the components of the deamination-aware variant calling system 106 performing the functions described herein with respect to the deamination-aware variant calling system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the deamination- aware variant calling system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the deamination-aware variant calling system 106 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.

[0145] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non- transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., aAttorney Docket No. IP-2816-PCT 50 Patent Applicationmemory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0146] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0147] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0148] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.

[0149] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0150] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, orAttorney Docket No. IP-2816-PCT 51 Patent Applicationspecial purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0151] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0152] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0153] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.Attorney Docket No. IP-2816-PCT 52 Patent Application

[0154] FIG. 10 illustrates a block diagram of a computing device 1000 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1000 may implement the deamination-aware variant calling system 106 and the sequencing device system 104. As shown by FIG. 10, the computing device 1000 can comprise a processor 1002, a memory 1004, a storage device 1006, an I / O interface 1008, and a communication interface 1010, which may be communicatively coupled by way of a communication infrastructure 1012. In certain embodiments, the computing device 1000 can include fewer or more components than those shown in FIG. 10. The following paragraphs describe components of the computing device 1000 shown in FIG. 10 in additional detail.

[0155] In one or more embodiments, the processor 1002 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1002 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1004, or the storage device 1006 and decode and execute them. The memory 1004 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1006 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0156] The I / O interface 1008 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1000. The I / O interface 1008 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1008 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1008 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.

[0157] The communication interface 1010 can include hardware, software, or both. In any event, the communication interface 1010 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1000 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1010 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.Attorney Docket No. IP-2816-PCT 53 Patent Application

[0158] Additionally, the communication interface 1010 may facilitate communications with various types of wired or wireless networks. The communication interface 1010 may also facilitate communications using various communication protocols. The communication infrastructure 1012 may also include hardware, software, or both that couples components of the computing device 1000 to each other. For example, the communication interface 1010 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.

[0159] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.

[0160] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Attorney Docket No. IP-2816-PCT 54 Patent Application

Claims

1. CLAIMSWe claim:

1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: identify, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads, wherein the duplex consensus bases comprise at least one methylated cytosine base or deaminated cytosine base on a single strand; determine a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases that represent a haplotype base converted from one base class to another base class and that comprise the at least one methylated cytosine base or deaminated cytosine base; determine a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases at the genomic coordinate; generate, utilizing a variant call model and based on the first set of probabilities and the second set of probabilities, a third set of probabilities of the target genomic sample exhibiting the observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases; and generate, based on posterior genotype probabilities derived from the third set of probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate.

2. The system of claim 1, wherein the converted haplotype intermediate bases comprise a haplotype intermediate base representing a methylated cytosine base on a plus strand or a haplotype intermediate base representing a methylated cytosine base on a minus strand.

3. The system of claim 1 , wherein an observed duplex consensus base of the observed duplex consensus bases comprises an observed duplex consensus base representing a methylated cytosine base on a plus strand or an observed duplex consensus base representing a methylated cytosine base on a minus strand.Attorney Docket No. IP-2816-PCT 55 Patent Application4. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases by: determining a plus-strand error probability that a collapsed base on a collapsed plus singlestranded read is an error; determining a minus-strand error probability that a collapsed base on a collapsed minus single-stranded read is an error; and determining the first set of probabilities of the target genomic sample based on the plusstrand error probability and the minus-strand error probability.

5. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining, for the target genomic sample, a first set of duplex consensus bases comprising a thymine on a collapsed plus single-stranded read and a cytosine on a collapsed minus single-stranded read; determining, for the target genomic sample, a second set of duplex consensus bases comprising a guanine on the collapsed plus single-stranded read and a cytosine on the collapsed minus single-stranded read; and generating an empirical duplex-consensus base-conversion probability based on the first set of duplex consensus bases and the second set of duplex consensus bases.

6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining collapsed-plus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed plus single-stranded reads; determining collapsed-minus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed minus single-stranded reads;Attorney Docket No. IP-2816-PCT 56 Patent Applicationgenerating duplex-consensus base-conversion probabilities based on the collapsed-plus- strand base-conversion probabilities and the collapsed-minus-strand base-conversion probabilities; and determining the second set of probabilities based on the duplex-consensus base-conversion probabilities.

7. The system of claim 6, wherein the duplex-consensus base-conversion probabilities are specific to the target genomic sample.

8. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, at the genomic coordinate, a collapsed-plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylation-level value for a collapsed plus single-stranded read; and determine, at the genomic coordinate, a collapsed-minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylationlevel value for a collapsed minus single-stranded read.

9. The system of claim 8, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the estimated methylation-level value for the collapsed plus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed plus singlestranded reads; and generate the estimated methylation-level value for the collapsed minus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed minus singlestranded reads.

10. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to: determine a genomic-sample-wide plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylation-level value for a collapsed plus single-stranded read;Attorney Docket No. IP-2816-PCT 57 Patent Applicationdetermine a genomic-sample- wide minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylation-level value for a collapsed minus single-stranded read; generate, for the genomic coordinate, the estimated methylation-level value for the collapsed plus single-stranded read further based on the genomic-sample-wide plus-strand baseconversion probability; and generate, for the genomic coordinate, the estimated methylation-level value for the collapsed minus single-stranded read further based on the genomic-sample-wide minus-strand base-conversion probability.

11. A computer-implemented method comprising: identifying, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads, wherein the duplex consensus bases comprise at least one methylated cytosine base or deaminated cytosine base on a single strand; determining a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases that represent a haplotype base converted from one base class to another base class and that comprise the at least one methylated cytosine base or deaminated cytosine base; determining a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases at the genomic coordinate; generating, utilizing a variant call model and based on the first set of probabilities and the second set of probabilities, a third set of probabilities of the target genomic sample exhibiting the observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases; and generating, based on posterior genotype probabilities derived from the third set of probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate.

12. The computer-implemented method of claim 11, wherein, the converted haplotype intermediate bases comprise a haplotype intermediate base representing a methylated cytosine base on a plus strand or a haplotype intermediate base representing a methylated cytosine base on a minus strand.

13. The computer-implemented method of claim 11, wherein an observed duplex consensus base of the observed duplex consensus bases comprises an observed duplex consensusAttorney Docket No. IP-2816-PCT 58 Patent Applicationbase representing a methylated cytosine base on a plus strand or an observed duplex consensus base representing a methylated cytosine base on a minus strand.

14. The computer-implemented method of claim 11, further comprising determining the first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases by: determining a plus-strand error probability that a collapsed base on a collapsed plus singlestranded read is an error; determining a minus-strand error probability that a collapsed base on a collapsed minus single-stranded read is an error; and determining the first set of probabilities of the target genomic sample based on the plusstrand error probability and the minus-strand error probability.

15. The computer-implemented method of claim 13, further comprising determining the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining, for the target genomic sample, a first set of duplex consensus bases comprising a thymine on a collapsed plus single-stranded read and a cytosine on a collapsed minus single-stranded read; determining, for the target genomic sample, a second set of duplex consensus bases comprising a guanine on the collapsed plus single-stranded read and a cytosine on the collapsed minus single-stranded read; and generating an empirical duplex-consensus base-conversion probability based on the first set of duplex consensus bases and the second set of duplex consensus bases.

16. The computer-implemented method of claim 11, further comprising determining the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining collapsed-plus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed plus single-stranded reads; determining collapsed-minus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed minus single-stranded reads;Attorney Docket No. IP-2816-PCT 59 Patent Applicationgenerating duplex-consensus base-conversion probabilities based on the collapsed-plus- strand base-conversion probabilities and the collapsed-minus-strand base-conversion probabilities; and determining the second set of probabilities based on the duplex-consensus base-conversion probabilities.

17. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to: identify, for a target genomic sample, duplex consensus reads comprising duplex consensus bases from collapsed single-stranded reads, wherein the duplex consensus bases comprise at least one methylated cytosine base or deaminated cytosine base on a single strand; determine a first set of probabilities of the target genomic sample exhibiting observed duplex consensus bases at a genomic coordinate given converted haplotype intermediate bases that represent a haplotype base converted from one base class to another base class and that comprise the at least one methylated cytosine base or deaminated cytosine base; determine a second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given candidate haplotype bases at the genomic coordinate; generate, utilizing a variant call model and based on the first set of probabilities and the second set of probabilities, a third set of probabilities of the target genomic sample exhibiting the observed duplex consensus bases at the genomic coordinate given the candidate haplotype bases; and generate, based on posterior genotype probabilities derived from the third set of probabilities, a genotype call that the target genomic sample comprises a predicted combination of nucleobases at the genomic coordinate.

18. The non-transitory computer-readable medium of claim 17, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the second set of probabilities of the target genomic sample exhibiting the converted haplotype intermediate bases given the candidate haplotype bases by: determining collapsed-plus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed plus single-stranded reads; determining collapsed-minus-strand base-conversion probabilities that approximate a likelihood of a haplotype base converting from one base class to another base class on one or more collapsed minus single-stranded reads;Attorney Docket No. IP-2816-PCT 60 Patent Applicationgenerating duplex-consensus base-conversion probabilities based on the collapsed-plus- strand base-conversion probabilities and the collapsed-minus-strand base-conversion probabilities; and determining the second set of probabilities based on the duplex-consensus base-conversion probabilities.

19. The non-transitory computer-readable medium of claim 18, wherein the duplexconsensus base-conversion probabilities are specific to the target genomic sample.

20. The non-transitory computer-readable medium of claim 18, further comprising instructions that, when executed by the at least one processor, cause the computing device to: determine, at the genomic coordinate, a collapsed-plus-strand base-conversion probability that a cytosine base has converted to a thymine base based on an estimated methylation-level value for a collapsed plus single-stranded read; and determine, at the genomic coordinate, a collapsed-minus-strand base-conversion probability that a guanine base has converted to an adenine base based on an estimated methylationlevel value for a collapsed minus single-stranded read.

21. The non-transitory computer-readable medium of claim 20, further comprising instructions that, when executed by the at least one processor, cause the computing device to: generate the estimated methylation-level value for the collapsed plus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed plus singlestranded reads; and generate the estimated methylation-level value for the collapsed minus single-stranded read based on numbers of each observed base, a probability of each nucleobase at the genomic coordinate, and prior genotype probabilities at the genomic coordinate on collapsed minus singlestranded reads.Attorney Docket No. IP-2816-PCT 61 Patent Application