Mapping and alignment of methylation sequencing reads
Patent Information
- Application Number
- PCT/US2026/013341
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-06-26
- Filing Date
- 2026-01-30
- Publication Date
- 2026-09-03
Smart Images

Figure US2026013341_03092026_PF_FP_ABST
Abstract
Description
MAPPING AND ALIGNMENT OF METHYLATION SEQUENCING READSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 830,785, entitled, “MAPPING AND ALIGNMENT OF METHYLATION SEQUENCING READS,” filed on June 26, 2025 (IP-2950-PRV3); U.S. Provisional Patent Application No. 63 / 800,233, entitled, “MAPPING AND ALIGNMENT OF METHYLATION SEQUENCING READS,” filed on May 5, 2025 (IP-2950-PRV2); and U.S. Provisional Patent Application No. 63 / 763,181, entitled, “METHYLATION MAPPING AND ALIGNMENT,” filed on February 25, 2025 (IP-2950-PRV). The aforementioned applications are hereby incorporated by reference in their entirety.BACKGROUND
[0002] In recent years, biotechnology firms and research institutions have improved hardware and software for sequencing nucleotides and determining nucleobase calls for genomic samples. For instance, some existing sequencing machines and sequencing-data-analysis software (together “existing sequencing systems”) predict individual nucleobases within sequences by using conventional Sanger sequencing or sequencing-by-synthesis (SBS) methods. When using SBS, existing sequencing systems can monitor millions to billions of oligonucleotides being synthesized in parallel from templates to predict nucleobase calls for growing nucleotide reads. In many existing sequencing systems, a camera captures images of irradiated fluorescent tags incorporated into oligonucleotides. After capturing such images, some existing sequencing systems determine nucleobase calls for nucleotide reads corresponding to the oligonucleotides and send base-call data to a computing device with sequencing-data-analysis software, which aligns nucleotide reads with a reference genome. Based on differences between the aligned nucleotide reads and the reference genome, existing sequencing systems can further utilize a variant caller to identify variants of a genomic sample, such as single nucleotide variants (SNVs), insertions or deletions (indels), or other variants within the genomic sample.
[0003] In addition to improved genomic sequencing, biotechnology firms and research institutions have also improved methods of detecting methylation of cytosine bases at particular genomic regions (e.g., regions encoding or promoting genes) and detecting methylation of larger nucleotide fragments or whole genomes of a sample. Existing sequencing systems that can detect methylation are referred to herein as “existing methylation sequencing systems.” For instance, some existing methylation sequencing systems can use sequencing devices and corresponding sequencing-data-analysis software to identify when a methyl or hydroxymethyl group has beenAttorney Docket No. IP-2950-PCT 1 Patent Applicationadded to a cytosine base of a sample’s deoxyribonucleic acid (DNA) — where the methylated cytosine base is often part of a cytosine-guanine-dinucleotide pair in a 5’ — C — phosphate — G — 3’ (CpG) configuration in mammals.
[0004] Such existing methylation sequencing systems can detect methylated cytosines by (i) enzymatically or chemically converting methylated or unmethylated cytosine bases at CpG or other cytosine sites from a sample nucleotide fragment into uracil bases (e.g., dihydrouracil); (ii) determining base calls of nucleotide reads for the sample using a sequencing device, where the sequencing device detects the uracil bases as thymine bases during polymerase chain reaction (PCR) amplification; and (iii) comparing the base calls from the nucleotide reads to a reference genome or non-enzymatically converted nucleotide reads from the sample. Based on the comparison of nucleotide reads from the sample to a reference genome or the non-enzymatically converted nucleotide reads, existing sequencing systems can identify thymine bases from the nucleotide reads that do not match cytosine bases at CpG or other cytosine sites within the reference genome or the non-enzymatically converted nucleotide reads and, thereby, detect methylated cytosine bases in a sample nucleotide fragment. Because about 98% of cytosines are uC, while only about 2% are mC, the reads produced by these different library preparation methods differ greatly in the amount of cytosine that is converted to T. Also, in mammalian genomes, mC is concentrated in cytosine guanine (CG) dinucleotide contexts.
[0005] In addition to detecting methylated cytosine, some existing methylation sequencing systems can also map and align nucleotide reads to a reference genome and determine genotype calls based on the mapped-and-aligned read. But such existing systems consume unnecessary memory to support genotype calling. For instance, certain methylation sequencing systems use a copy of an entire reference genome with all cytosines converted to thymine, and another referencegenome copy with all guanines converted to adenine. Such all-cytosine-converted reference genomes are sometimes referred to as “3 -base reference genomes” because each of the resulting copies will include only 3 bases. However, the doubled reference-genome size increases compute and memory burden. To compensate for the increased compute and / or memory burden, some existing methylation sequencing systems reduce the depth of read subsequences (e.g., k-mers) for mapping. By reducing the depth of read subsequences, however, such existing systems decrease mapping and / or variant-calling accuracy.
[0006] To facilitate read alignment, some existing methylation sequencing systems use an alignment process that reduces mapping quality of methylation sequencing reads. In particular, some such systems score alignments of methylation sequencing reads to '3 -base reference genomes to account for nucleobases in the methylation sequencing reads that have been converted by the methylation sequencing assay. However, alignment scoring against a 3 -base reference genomeAttorney Docket No. IP-2950-PCT 2 Patent Applicationleads to lower mapping quality (e.g., MAPQ) scores. Because 3-base reference genomes include only 3 bases, the simplified reference genomes have reduced sequence complexity, and the likelihood that a methylation sequencing read could map and / or align ambiguously to multiple locations on the reference genome is increased.
[0007] In the alternative to mapping and aligning reads to a 3-base reference genome, one potential approach would be to map and align methylation sequencing reads using a sequencing system that does not account for methylation (hereinafter, “existing non-methylation sequencing system”). For example, an existing non-methylation sequencing system could map and / or align reads from library preps that convert mC to T, with a conventional, unconverted reference genome that does not account for methylation conversions. But, even with only about 2% of C converted to T, such converted-read-to-unconverted-reference mapping introduces mismatches to mapping and aligning and leads to inaccuracy. Because nucleobases in the methylation sequencing reads converted by the methylation sequencing assay will not match the corresponding nucleobase in the unconverted reference genome, these existing non-methylation sequencing systems overly penalize correct read alignments during alignment scoring.
[0008] Further compounding the foregoing alignment challenges, some existing methylation sequencing systems cannot effectively mask (e.g., soft-clip) nucleobases from a methylation sequencing read as part of a read alignment process. Nucleobases on an end of a methylation sequencing read may not accurately represent a sample due to sequencing error or because the nucleobase is actually from an adapter sequence. When existing methylation sequencing systems cannot generate or score candidate alignments with methylation sequencing reads having one or more masked nucleobases, the systems are more likely to inaccurately align such reads and compromise downstream analyses. For example, an inaccurate nucleobase, when not masked and scored accordingly, can decrease a likelihood of selecting a candidate alignment because the unmasked nucleobase does not match a reference base in the corresponding reference sequence. Furthermore, the methylation sequencing system may inaccurately determine an unmasked nucleobase to be a variant or represent a conversion caused by the methylation sequencing assay. Thus, these existing methylation sequencing systems may less accurately align reads and more likely compromise downstream genotype calling and methylation level estimation.
[0009] Due to decreased accuracy of such mapping-and-alignment, some existing methylation sequencing systems bifurcate and employ multiple assays and samples to accurately determine methylation levels and variant calls. But such a bifurcated approach inefficiently consumes an inordinate amount of processing materials, time, and computing resources. For example, existing systems often require a separate methylation assay in addition to the utilization of a separate variant caller to accurately determine both methylation-level values and genotype calls. Accordingly, someAttorney Docket No. IP-2950-PCT 3 Patent Applicationexisting methylation sequencing systems require multiple samples from a single organism on which to perform both sequencing and methylation assays in separate computational analyses. The duplication of genomic samples often necessitates a duplication of computer processing, computer storage, software programs, and other resources to sequence and determine methylation levels for the same genomic sequence. Thus, existing methylation sequencing systems often consume excessive genomic samples, significant time, and computer processing resources to both sequence and determine methylation levels for a single genomic sample.
[0010] These, along with additional problems and issues exist in existing methylation sequencing systems and existing non-methylation sequencing systems.SUMMARY
[0011] This disclosure describes one or more embodiments of systems, methods, and non-transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can (a) convert nucleobases at candidate methylation sites in a sample’s read subsequences to facilitate mapping methylation sequencing reads for the sample to similarly converted reference subsequences and (b) align the sample’s methylation sequencing reads with reference sequences. For example, the disclosed systems can access read subsequences of methylation sequencing reads of a target genomic sample. The methylation sequencing reads can include nucleobases that have been modified by a methylation sequencing assay. The disclosed systems can further identify a candidate methylation site within the read subsequences and convert at least one nucleobase at the candidate methylation site from one base type (e.g., A, C, T, G) to another and map the converted read subsequences to converted reference sequences taken from one or more reference sequences. Based on the converted read-to-reference-subsequence mapping, the disclosed systems can generate candidate alignments of the methylation sequencing reads to the reference sequences.
[0012] Independent of such mapping and alignment of methylation sequencing reads, the disclosed systems can generate a tolerance-adjusted alignment score for candidate read alignments using a tolerance metric. In determining an alignment score, the disclosed systems apply a tolerance metric that increases a likelihood of selecting the candidate read alignment with a read-reference-base mismatch within the context of a candidate methylation site (e.g., within a CG context) relative to another candidate read alignment with a read-reference-base mismatch outside of the context of the candidate methylation site. Based on tolerance-adjusted alignment scores, the disclosed systems can select a candidate read alignment for improved downstream analyses, such as determining a genotype, sequence variant, or methylation-level value.Attorney Docket No. IP-2950-PCT 4 Patent Application
[0013] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The detailed description refers to the drawings briefly described below.
[0015] FIG. 1 illustrates an environment in which a methylation sequencing map / align system can operate in accordance with one or more embodiments of the present disclosure.
[0016] FIG. 2 illustrates an overview of the methylation sequencing map / align system mapping and aligning methylation sequencing reads for genotype calling in accordance with one or more embodiments of the present disclosure.
[0017] FIG. 3 A illustrates an overview of the methylation sequencing map / align system obtaining sequences from a variant site and a methylated site in the context of a methylation sequencing assay and converting one or more nucleobases at a candidate methylation site in accordance with one or more embodiments of the present disclosure.
[0018] FIG. 3B illustrates the methylation sequencing map / align system mapping converted read subsequences to converted reference subsequences, in accordance with one or more embodiments of the present disclosure.
[0019] FIG. 3C illustrates the methylation sequencing map / align system aligning methylation sequencing reads to one or more reference sequences, in accordance with one or more embodiments of the present disclosure.
[0020] FIG. 4A illustrates the methylation sequencing map / align system determining a read-reference-base mismatch within a CG context of a candidate alignment and alignment scoring as a result of the read-reference-base mismatch, in accordance with one or more embodiments of the present disclosure.
[0021] FIG. 4B illustrates the methylation sequencing map / align system generating alignment scores based on read-reference comparisons between methylation sequencing reads and one or more reference sequences in accordance with one or more embodiments of the present disclosure.
[0022] FIG. 5 illustrates improvements in accuracy by the methylation sequencing map / align system in determining a methylation-level value relative to existing methylation sequencing systems in accordance with one or more implementations of the present disclosure.
[0023] FIG. 6 A illustrates improvements in accuracy by the methylation sequencing map / align system in calling single nucleotide polymorphisms (SNPs) relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.Attorney Docket No. IP-2950-PCT 5 Patent Application
[0024] FIG. 6B illustrates improvements in accuracy by the methylation sequencing map / align system in calling single nucleotide polymorphisms (SNPs) and insertions / deletions (INDELs) relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.
[0025] FIG. 6C illustrates improvements in accuracy by the methylation sequencing map / align system in calling single nucleotide variants (SNVs) and insertions / deletions (INDELs) relative to existing sequencing systems in accordance with one or more implementations of the present disclosure.
[0026] FIG. 7 illustrates a flowchart of a series of acts for mapping and aligning methylation sequencing reads in accordance with one or more embodiments of the present disclosure.
[0027] FIG. 8 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION
[0028] This disclosure describes embodiments of methods, non-transitory computer-readable media, and systems that can (a) convert nucleobases at candidate methylation sites in a sample’s read subsequences (e.g., k-mers from reads) to support mapping methylation sequencing reads from which the read subsequences derive to methylation-site converted reference subsequences and further (b) align unconverted versions of the methylation sequencing reads with one or more reference sequences from a reference genome. In some embodiments, for example, the disclosed systems use methylation sequencing techniques that convert methylated cytosine (mC) to thymine (T) — or other nucleobase conversions noted below — in methylation sequencing reads that have been processed by a methylation sequencing assay. As described below, the disclosed systems can improve memory efficiency and variant-calling accuracy, among other things, by utilizing the unique conversion-based mapping and unconverted alignment described herein.
[0029] For example, the disclosed systems can access read subsequences of methylation sequencing reads that have been obtained from a target genomic sample. The methylation sequencing reads can include nucleobases that have been modified by a methylation sequencing assay performed on the target genomic sample. The disclosed systems can further identify a candidate methylation site within the read subsequences and convert at least one nucleobase at the candidate methylation site from one base type (e.g., A, C, T, G) to another. The disclosed systems can further map the converted read subsequences to converted reference sequences, which are taken from one or more reference sequences and selectively converted at a candidate methylation site. Based on the mapping of the converted read subsequences to the converted reference subsequences,Attorney Docket No. IP-2950-PCT 6 Patent Applicationthe disclosed systems can generate candidate alignments of the original, unconverted methylation sequencing reads to the original, unconverted one or more reference sequences.
[0030] To access read subsequences, in some embodiments, the disclosed systems extract read subsequences (e.g., k-mers) from methylation sequencing reads. Such read subsequences can include nucleobases that, as part of the methylation sequencing read, has been modified by a methylation sequencing assay. For example, the methylation sequencing assay may have previously converted (e.g., chemically or enzymatically) methylated cytosines to thymine in the methylation sequencing reads.
[0031] Having accessed or extracted read subsequences, the disclosed systems can convert the read subsequences by converting, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences. When converting read subsequences, in some embodiments, the candidate methylation site comprises a cytosine guanine (CG) dinucleotide. To convert nucleobases within such a CG dinucleotide, for example, the disclosed systems convert one or more nucleobases based on the type of modifications made by the methylation sequencing library preparation methods. In some embodiments, for example, the disclosed systems convert any CG dinucleotides in the read subsequences to TG (thymine guanine) if the read subsequence is from a sense strand, and to CA (cytosine adenine) if the read subsequence is from an antisense strand.
[0032] In addition to accessing and converting read subsequences, the disclosed systems can also access and convert reference subsequences taken from one or more reference sequences. The disclosed systems can convert reference subsequences by converting nucleobases from one base type to another base type at candidate methylation sites within reference subsequences. For example, when the candidate methylation site comprises a CG dinucleotide, the disclosed systems can generate a first set of reference subsequences with all CG dinucleotides converted to TG, and a second set of reference subsequences with all CG dinucleotides converted to CA. The converted read subsequences, along with unconverted read subsequences that do not include a candidate methylation site, can be stored in a data structure, such as a hash table.
[0033] After accessing converted read subsequences and reference subsequences, the disclosed systems can map the converted read subsequences to converted reference subsequences. From the results of the mapping, the disclosed systems can further store genomic coordinates. For example, the disclosed systems store genomic coordinates from the one or more reference sequences that correspond to where the converted read subsequences map to the converted reference subsequences.
[0034] Having mapped converted read subsequences to reference subsequences, the disclosed systems can subsequently generate candidate alignments of the (unconverted) methylationAttorney Docket No. IP-2950-PCT 7 Patent Applicationsequencing reads to (unconverted) reference sequences at genomic coordinates corresponding to the mapped converted read subsequences. In some embodiments, for instance, the disclosed systems can determine a mismatch-tolerant-adjusted alignment score for candidate alignments such that read-reference-base mismatches that are in the context of a candidate methylation site (e.g., in the context of a CG dinucleotide) are penalized less than a read-reference-base mismatch that is not in the context of a candidate methylation site. Thus, the likelihood of selecting a candidate alignment that includes a read-reference-base mismatch in the context of a candidate methylation site is increased compared to another candidate alignment that has a read-reference-base mismatch that is not in the context of a candidate methylation site. The disclosed systems can select a candidate alignment with a highest determined alignment score for use in downstream processes, including genotyping and / or variant calling.
[0035] To perform genotype and / or variant calling, among other things, the disclosed systems can (i) determine genomic coordinates or other candidate sites comprising methylated nucleobases within the target genomic sample based on comparing selected alignments of methylation sequencing reads to a reference sequence and (ii) determine variant calls or reference calls at such coordinates or sites based on the selected read alignments. Such variant calls can include, for example, single nucleotide variants (SNVs) and / or insertions and / or deletions (INDELs), or structural variants (SVs) based on analyses comparing selected alignments of methylation sequencing reads to one or more reference sequences. In addition to such genotype or variant calling, the disclosed systems can also estimate a methylation level for a genomic coordinate or genomic region of the target genomic sample.
[0036] The present disclosure relates to several improvements to solve technical problems associated with existing methylation sequencing systems and existing non-methylation sequencing systems. For example, the disclosed systems save memory and improve computational efficiency relative to existing methylation sequencing systems by converting the nucleobases of read subsequences and mapping the converted read subsequences to converted reference subsequences. As indicated above, to map methylation sequencing reads, some existing methylation sequencing systems map reads to two entire copies of a reference genome, where one reference-genome copy has all cytosines converted to thymine, and another reference-genome copy has all guanines converted to adenine. This dual-copy technique doubles the size of reference sequences from, for example, a human genome of 3.2 billion bases stored in about 750 megabytes (MBs) to two reference-genome copies consuming about 1,500 MBs, thereby increasing compute burden unless accuracy tradeoffs (such as reduced subsequence depth) are made. In contrast, the disclosed systems conserve memory and improve efficiency by mapping the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences —Attorney Docket No. IP-2950-PCT 8 Patent Applicationrequiring a single reference-genome copy (e.g., stored in about 750 MBs). The smaller reference sequence also improves efficiency for downstream steps, such as generating subsequences and mapping read subsequences to reference subsequences. The converted reference subsequences can comprise one or more nucleobases at a candidate methylation site that have been selectively converted from one base type to another base type. Thus, the disclosed systems can tailor read mapping to a given candidate methylation site (e.g., CG site) by selectively generating converted reference subsequences comprising such a candidate methylation site. The disclosed systems, therefore, save memory and computer processing relative to existing methylation sequencing systems, without having to make the additional alignment or variant-calling accuracy tradeoffs made by some existing methylation sequencing systems. Such memory and computer-processing savings provide a concrete technological improvement to the field of methylation sequencing in part because the disclosed systems facilitate broader adoption of sensitive genetic screening in resource-limited or high-volume laboratory environments.
[0037] In addition to memory savings, in some cases, the disclosed systems improve mapping-and-alignment accuracy by aligning methylation sequencing reads to one or more (unconverted) reference sequences. As discussed above, some existing methylation sequencing systems align methylation sequencing reads to a reference genome with all cytosines converted to thymine, or with all guanines converted to adenine, such that the reference genome(s) only has 3 nucleobases. By generating candidate alignments of the methylation sequencing reads to (unconverted) reference sequences of the one or more reference sequences, the disclosed systems improve read-alignment accuracy. Because alignment is based on original, unconverted reference sequences with 4 nucleobases, the disclosed systems’ approach to generating candidate alignments of the methylation sequencing reads both preserves sequence complexity and avoids needlessly reducing MAPQ, unlike existing methylation sequencing systems that score alignments based on a 3 -base reference genome. Because alignment with a 4-base reference sequence is less ambiguous than alignment with a 3 -base reference sequence, the disclosed systems can identify candidate read alignments of higher alignment and mapping scores. Thus, alignments will have higher MAPQ scores and greater accuracy because the disclosed systems can leverage such scores to determine where a methylation sequencing read most accurately aligns.
[0038] Beyond improved mapping-and-alignment accuracy, the disclosed systems improve alignment accuracy by scoring read-reference-base mismatches that occur in the context of a candidate methylation site with a tolerance metric. As discussed above, some existing nonmethylation sequencing systems would align methylation sequencing reads to an unconverted reference sequence. But such existing systems overly penalize read-reference-base mismatches due to conversions from the methylation sequencing assay and score the correct alignment lower thanAttorney Docket No. IP-2950-PCT 9 Patent Applicationincorrect alignments, thereby resulting in the correct alignment not being selected. In contrast, the disclosed systems can (i) determine, for a candidate alignment of a methylation sequencing read, a read-reference-base mismatch within a context of a candidate methylation site (e.g., within a CG context) between a target nucleobase of the methylation sequencing read and a corresponding reference nucleobase and (ii) generate an alignment score for the candidate alignment based on a tolerance metric that increases a likelihood of selecting the candidate alignment relative to another candidate alignment mismatch outside of the context of the candidate methylation site (e.g., outside of the CG context). With the disclosed tolerance metric, read-reference-base mismatches in the context of a candidate methylation site (e.g., a CG dinucleotide) will not be overly penalized in the alignment scoring because these mismatches may be due a nucleobase being chemically or enzymatically converted by the methylation conversion assay. The disclosed systems thus improve alignment accuracy compared to approaches that would equally penalize read-reference-base mismatches that are in the context of a candidate methylation site and such mismatches that are not.
[0039] By improving read-alignment accuracy, the disclosed systems also improve downstream accuracy in genotype calling. Indeed, by converting nucleobases at candidate methylation sites in a sample’s read subsequences, mapping corresponding methylation sequencing reads from which the read subsequences derive to methylation-site converted reference subsequences, and generating candidate read alignments of the methylation sequencing reads to one or more reference sequences, the disclosed systems select candidate alignments that result in more accurate variant calls relative to existing methylation sequencing systems. As further detailed below and depicted in FIGS. 6 A - 6C, the disclosed systems demonstrate improved precision and sensitivity in variant calling relative to existing methylation sequencing systems (e.g., EM-Seq) and existing non-methylation sequencing systems, and comparable accuracy to Whole Genome Sequencing (WGS) by ILLUMINA DRAGEN.
[0040] Independent of such improved alignment and variant-calling accuracy, the disclosed systems improve accuracy by masking nucleobases that would ordinarily be scored for alignment during the alignment process. As discussed above, some existing methylation sequencing systems are not compatible with masking nucleobases in a candidate alignment, including nucleobases in a methylation sequencing read that may be inaccurately base-called or may be left over from an adaptor sequence, resulting in inaccurate alignment and downstream analyses. In contrast, the disclosed systems can mask (e.g., soft-clip) one or more nucleobases from a methylation sequencing read and generate candidate alignment with the methylation sequencing read having one or more masked nucleobases and a corresponding reference sequence. Indeed, the disclosed systems introduce a first-of-its-kind model for masking bases in methylation sequencing reads thatAttorney Docket No. IP-2950-PCT 10 Patent Applicationexisting methylation sequencing systems cannot perform. The disclosed systems can thus increase a likelihood that a correct candidate alignment is selected and decrease false positives in variant calling and methylation level determination. By enabling masking (e.g., soft-clipping), the disclosed systems improve the accuracy of alignment, genotype calling, and methylation-level value determination.
[0041] Beyond improved genotype-call accuracy and / or improved computing and memory, in some embodiments, the disclosed systems improve efficiency in processing and physical resources relative to existing methylation sequencing systems. Because state-of-the-art genotype-calling and methylation detection accuracy can be unfit for clinical benchmarks, some existing methylation sequencing systems execute (i) a separate methylation sequencing assay to chemically or enzymatically convert nucleobases for nucleotide reads from a genomic sample and determine methylation levels and (ii) a separate DNA sequencing run with non-chemically or non-enzymatically converted nucleotide reads from the genomic sample to determine variant calls. Such separate methylation sequencing assays and DNA sequencing can consume and duplicate computer processing, memory storage, physical space and reagents for a nucleotide-sample slide (e.g., flow cell), and software programs (e.g., separate methylation analysis and variant calling software). In contrast to such a bifurcated approach, in some embodiments, the disclosed systems can concurrently (or in parallel) determine methylation-level values indicating levels of methylation of a target genomic sample’s cytosine bases and generate variant calls for the genomic sample with improved accuracy relative to existing methylation sequencing systems. Thus, the disclosed systems can efficiently generate epigenetic and genetic sequencing data from a single genomic sample. By performing both genotype calling and methylation-level-value determinations from methylation sequencing reads of a single sample, the disclosed systems also decrease consumption of sequencing reagents and materials, which are major contributors to the overall cost of next-generation physical resources required in sequencing workflows of high-throughput sequencing laboratories and / or clinical laboratories that process a significant number of samples and amounts of sequencing data.
[0042] As suggested by the foregoing discussion, this disclosure utilizes a variety of terms to describe features and benefits of the methylation sequencing map / align system. Additional detail is hereafter provided regarding the meaning of these terms as used in this disclosure. As used in this disclosure, for instance, the term “target genomic sample” (or simply “sample”) refers to a specimen, culture, or the like that is suspected of including a target nucleic acid. In some embodiments, the sample comprises DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acids as targets. The sample can likewise include any biological, clinical, surgical, agricultural-atmospheric, or aquatic-basedAttorney Docket No. IP-2950-PCT 11 Patent Applicationspecimen containing one or more nucleic acids. A target genomic sample also includes any isolated or extracted nucleic acid sample from an organism, such a genomic DNA, fresh-frozen, or formalin-fixed paraffin-embedded nucleic acid specimen. In some cases, accordingly, a target genomic sample can include a full genome or partial genome that is isolated or extracted (e.g., in whole or in part by a kit) from an organism and that is prepared to undergo sequencing or an assay in a sequencing device. In the latter such cases of a partial genome, the term “sample nucleotide sequence” refers to a sequence of nucleotides isolated or extracted from a sample organism (or a copy of such an isolated or extracted sequence). A target genomic sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.
[0043] The target genomic sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source, such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0044] As further used herein, the term “methylation sequencing assay” refers to an assay that detects, measures, or quantifies methylation of cytosine from an oligonucleotide or other nucleotide sequence. In some cases, a methylation sequencing assay detects or quantifies methylation of cytosine at particular target genomic regions or in particular cell types. Some methylation sequencing assays quantify methylation in terms of methylation-level values.
[0045] Relatedly, as used herein, the term “directional methylation sequencing assay” refers to a methylation sequencing assay that can distinguish between sense and antisense strands of DNAAttorney Docket No. IP-2950-PCT 12 Patent Applicationin a target genomic sample. There are several kinds of directional sequencing methods known to those of ordinary skill in the art. For example, some methods use strand-specific adapters that are joined to the end of a sense strand or antisense strand of DNA from the target genomic sample. The DNA is then denatured and methylated cytosines are converted to thymine. After sequencing, the sequence from the directional adapter can be used to identify whether the methylation sequencing read is from the sense strand or the antisense strand.
[0046] As used herein, the term “nucleobase” or “base” refers to a nitrogenous base. In particular, nucleobases comprise components of nucleotides. For example, a nucleobase may be an adenine (A), cytosine (C), guanine (G), or thymine (T).
[0047] As used herein, the term “base type” refers to a particular type or class of nitrogenous base. For instance, a genome or nucleotide sequence may include five different base types, including adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U).
[0048] As further used herein, the term “nucleobase call” (or simply “base call”) refers to a determination or prediction of a particular nucleobase (or nucleobase pair) for an oligonucleotide (e.g., nucleotide read) during a sequencing cycle or for a genomic coordinate of a genomic sample. In particular, a nucleobase call can indicate a determination or prediction of the type of nucleobase that has been incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., read-based nucleobase calls). In some cases, for a nucleotide read, a nucleobase call includes a determination or a prediction of a nucleobase based on intensity values resulting from fluorescent-tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., in a cluster of a flow cell). As suggested above, a single nucleobase call can be an adenine (A) call, a cytosine (C) call, a guanine (G) call, a thymine (T) call, or an uracil (U) call. Note that the terms nucleobase and nucleotide base are interchangeable.
[0049] As further used herein, the term “nucleotide read” (or simply “read”) refers to an inferred sequence of one or more nucleobases (or nucleobase pairs) from all or part of a sample nucleotide sequence (e.g., a sample genomic sequence, complementary DNA). In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a nucleotide sequence (or group of monoclonal nucleotide sequences) from a sample library fragment corresponding to a genomic sample. For example, in some cases, a sequencing device determines a nucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a cluster in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form aAttorney Docket No. IP-2950-PCT 13 Patent Applicationcontiguous sequence (e.g., assembled nucleotide reads) satisfying a threshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.
[0050] As used herein, the term “methylation sequencing read” refers to a nucleotide read comprising nucleobase calls (or corresponding data) for a nucleobase that has been modified or otherwise affected by a methylation sequencing assay. In particular, a methylation sequencing read can include one or more nucleobase calls for one or more corresponding nucleobases that have been chemically or enzymatically converted from one base type to another base type during the methylation sequencing assay. For example, where a target genomic sample had a CG dinucleotide with a methylated cytosine, a methylation sequencing read may have nucleobase calls for a TG dinucleotide or a CA dinucleotide due to the methylation sequencing assay.
[0051] As used herein, the term “read subsequence” refers to a subsequence of a methylation sequencing read or other nucleotide read. In particular, a read subsequence can include a sequence of nucleobases that is a portion or part of nucleobase sequence from a methylation sequencing read. For example, a read subsequence may be a nucleotide k-mer comprising a subset of nucleobases of length k (e.g., as measured by a number of nucleobases) extracted from a nucleotide read. Accordingly, one methylation sequencing read can provide a plurality of read subsequences representing different portions of the methylation sequencing read.
[0052] As used herein, the term “k-mer” refers to a subsequence of nucleobases of a predetermined length k, where k represents an integer number of nucleobases. One of ordinary skill will be able to determine an appropriate value for k, for example 1, 10, 20, 30, 40, 50, 60 or any integer therebetween. In some embodiments, k is between 10 and 30 bases. A k-mer can encompass different types of k-mers, such as minimizers, spaced k-mers, syncmers, strobemers, gapped k-mers, super-k-mers, and unitigs.
[0053] As used herein, the term “converted read subsequence” refers to a read subsequence comprising one or more nucleobase calls for (or other data representing) one or more nucleobases that have been converted at one or more candidate methylation sites through a bioinformatic process. In particular, a converted read subsequence can comprise one or more nucleobase calls for one or more nucleobases that have been changed by a bioinformatic (or data-manipulation) from one base type to another base type at one or more candidate methylation sites within the read subsequence. For example, after a bioinformatic conversion, a converted read subsequence may include “TG” or “CA,” where the original methylation read subsequence had a “CG.” While the methylation sequencing map / align system converts nucleobase calls or other data representing nucleobases in read subsequences as part of a bioinformatic or data-manipulation process — and does not enzymatically or chemically convert actual nucleobases in read subsequences — thisAttorney Docket No. IP-2950-PCT 14 Patent Applicationdisclosure often describes and / or claims converting nucleobases in a read subsequence as a form of shorthand for brevity.
[0054] As used herein, the term “candidate methylation site” refers to a predetermined motif of nucleobases that can appear in a sequence or subsequence of nucleobases. A candidate methylation site can be predetermined based on a particular motif of nucleobases that will sometimes (but not necessarily always) include a methylated nucleobase in a target genomic sample. For example, a candidate methylation site can be a cytosine guanine (CG) dinucleotide on either a sense strand or a reverse complement of an antisense strand.
[0055] As used herein, the term “reference subsequence” refers to a subsequence of a reference sequence. In particular, a reference subsequence can include a sequence of nucleobases that is a portion or part of nucleobase sequence from a reference sequence. For example, a reference subsequence may be a nucleotide k-mer comprising a subset of nucleobases of length k (e.g., as measured by a number of nucleobases) extracted from a reference sequence. Accordingly, one reference sequence can provide a plurality of reference subsequences representing different portions of the reference sequence.
[0056] As used herein, the term “converted reference subsequence” refers to a reference subsequence comprising one or more nucleobase calls for (or other data representing) one or more nucleobases that have been selectively converted at one or more candidate methylation sites through a bioinformatic process. In particular, a converted reference subsequence can comprise one or more nucleobase calls for one or more nucleobases that have selectively been changed by a bioinformatic (or data-manipulation) process from one base type to another base type at one or more candidate methylation sites within the reference subsequence. Accordingly, the disclosed systems can identify a candidate methylation site within a reference subsequence and replace a nucleobase of the candidate methylation site with another nucleobase of another base type to generate a converted reference subsequence. For example, a converted reference subsequence may include “TG” or “CA,” where the original reference subsequence had a “CG.” While the methylation sequencing map / align system converts nucleobase calls or other data representing nucleobases in reference subsequences as part of a bioinformatic or data-manipulation process — and does not enzymatically or chemically convert actual nucleobases in reference subsequences — this disclosure often describes and / or claims converting nucleobases in a reference subsequence as a form of shorthand for brevity.
[0057] As used herein, the term “single strand” refers to one strand of DNA or a linear chain of nucleotides that forms one half or a part of double-stranded nucleotide sequences (e.g., DNA). In particular, a single strand can refer to one nucleotide read in a pair of paired-end reads. In some cases, a single strand comprises one of two complementary, antiparallel strands of genomicAttorney Docket No. IP-2950-PCT 15 Patent Applicationmaterial. Single strands can be identified based on the direction of sequencing relative to a reference sequence in a reference genome. For example, a single strand refers to a plus (e.g., Rl) or minus (e.g., R2) nucleotide read. In some cases, a single strand refers to a read from single-end sequencing.
[0058] As used herein, the term “sense strand” (or “plus strand, forward strand, positive strand, or Watson strand”) refers to a first read from a pair of paired-end reads, such as a read designated by convention as a first read from a pair of single-stranded reads. In particular, the plus strand can include a first read from one end of a nucleotide fragment and be in a forward orientation relative to a reference sequence of a reference genome. For example, a plus strand includes a forward 1st, reverse 2nd (F1R2) read of a nucleotide fragment in a pair of paired-end reads (or alternatively referred to as Rl).
[0059] Relatedly, the term “antisense strand” (or “minus strand, reverse strand, negative strand, or Crick strand”) refers to a second read from a pair of paired-end reads, such as a read designated by convention as a second read from a pair of single-stranded reads. In particular, the minus strand can include a second read from an opposite end of a nucleotide fragment and be in a reverse orientation relative to a reference sequence of a reference genome. For example, a minus strand includes a forward 2nd, reverse 1st (F2R1) read of a nucleotide fragment in a pair of paired-end reads (or alternatively referred to as R2).
[0060] As further used herein, the term “genomic coordinate” (or sometimes simply “coordinate”) refers to a particular location or position of a nucleobase within a genome (e.g., an organism’s genome or a reference genome). In some cases, a genomic coordinate includes an identifier for a particular chromosome of a genome and an identifier for a position of a nucleobase within the particular chromosome. For instance, a genomic coordinate or coordinates may include a number, name, or other identifier for a somatic or sex chromosome (e.g., chrl or chrX) and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl: 1234570 or chrl: 1234570-1234870). In some cases, a genomic coordinate refers to a genomic coordinate on a sex chromosome (e.g., chrX or chrY). Consequently, the [insert name] system can determine genotype probabilities for a genotype call (e.g., a variant call) for a genomic coordinate on a sex chromosome. Further, in certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mitochondrial DNA reference genome or SARS-CoV-2 for a reference genome for the SARS-CoV-2 virus) and a position of a nucleobase within the source for the reference genome (e.g., mt: 16568 or SARS-CoV-2:29001). By contrast, in certain cases, a genomic coordinate refers to a position of a nucleobase within a reference genome without reference to a chromosome or source (e.g., 29727).Attorney Docket No. IP-2950-PCT 16 Patent Application
[0061] As used herein, a “genomic region” refers to a range of genomic coordinates. Like genomic coordinates, in certain implementations, a genomic region may be identified by an identifier for a chromosome and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl: 1234570-1234870). In various implementations, a genomic coordinate includes a position within a reference genome. In some cases, a genomic coordinate is specific to a particular reference genome. Relatedly, as used herein, the term “reference span” refers to a span of nucleobase positions within a linear reference genome. In other words, a reference span includes a span of nucleobases between two respective genomic coordinates of the linear reference genome.
[0062] As noted above, a genomic coordinate includes a position within a reference genome. Such a position may be within a particular reference genome. As used herein, the term “reference genome” refers to a digital nucleic acid sequence assembled as a representative example (or representative examples) of genes and other genetic sequences of an organism. Regardless of the sequence length, in some cases, a reference genome represents an example set of genes or a set of nucleic acid sequences in a digital nucleic acid sequenced determined by scientists as representative of an organism of a particular species. For example, a linear human reference genome may be GRCh38 or other versions of reference genomes from the Genome Reference Consortium. As noted above, in some cases, a reference genome includes multi-base codes. As a further example, a reference genome may include a graph reference genome that includes both a linear reference genome and paths representing nucleic acid sequences from ancestral haplotypes, such as Illumina DRAGEN Graph Reference Genome hgl9.
[0063] Relatedly, as further used herein, the term “reference sequence” refers to a digital nucleic acid sequence used as a representative example (or representative examples) of at least a portion of a gene or other genetic sequence of an organism. Accordingly, a reference sequence can represent a part of a reference genome. For example, a reference sequence can include a primary contiguous sequence or an alternate contiguous sequence extracted or otherwise isolated from a reference genome. A reference sequence may be one of a plurality of digital nucleic acid sequences included in a reference genome.
[0064] As used herein, the term “mapping-quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In some embodiments, for example, a mapping-quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong}, rounded to the nearest integer. In the alternative to a mean or median mapping quality, inAttorney Docket No. IP-2950-PCT 17 Patent Applicationsome implementations, a mapping-quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.
[0065] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith-Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith- Waterman scoring.
[0066] As used herein, the term “candidate alignment” (or “candidate read alignment”) refers to a proposed alignment between a nucleotide read and a reference sequence. In particular, a candidate alignment can include an alignment between a methylation sequencing read and a reference sequence at a particular location (e.g., at one or more genomic coordinates or a corresponding genomic region).
[0067] As used herein, the term “conversion mismatch” refers to a mismatch between a nucleobase in one or more methylation sequencing reads and a corresponding nucleobase in one or more reference sequences, where the mismatch or corresponding nucleobase occurs within a context of a candidate methylation site. A mismatch may be determined to be within a context of a candidate methylation site based on one or more nucleobases in a reference sequence and / or based on one or more nucleobases in a methylation sequencing read.
[0068] Relatedly, as used further herein, the term “CG-conversion mismatch” refers to a mismatch between a nucleobase in one or more methylation sequencing reads and a corresponding nucleobase in one or more reference sequences, where the mismatch or corresponding nucleobase occurs within a CG context. Thus, a CG-conversion mismatch is a type of conversion mismatch when the candidate methylation site comprises a CG dinucleotide.
[0069] As used herein, the term “CG context” (or “cytosine guanine context”) refers to a region within a nucleotide read or a reference sequence comprising a dinucleotide, such as CG, TG, or CA, known to produce a mismatch due to the effects of a methylation sequencing assay. For example, a CG context may include: i) a CG dinucleotide in the one or more reference sequences, and / or ii) a sequence of nucleobases in the methylation sequencing reads that have been converted from a methylated CG dinucleotide by a methylation sequencing assay (e.g., a thymine guanine (TG) dinucleotide or a cytosine adenine (CA) dinucleotide).Attorney Docket No. IP-2950-PCT 18 Patent Application
[0070] As used herein, the term “tolerance metric” refers to a value, score, or other quantity that accounts for a mismatch between one nucleobase and another nucleobase within the context of a candidate methylation site. Accordingly, a tolerance metric can be a component of or applied to an alignment score. In particular, a tolerance metric can be a value applied to (or be part of) an alignment score for a candidate alignment comprising a mismatch within a CG context to increase a likelihood of selecting the candidate alignment relative to another candidate alignment comprising a mismatch outside of a CG context. Accordingly, a tolerance metric may penalize a conversion mismatch (e.g., CG-conversion mismatch) less than a mismatch that is not within the context of a candidate methylation site.
[0071] As used herein, the term “methylation-level value” refers to a numeric value indicating an amount, percentage, ratio, or quantity of cytosine to which a methyl group or hydroxymethyl group has been added or bonded. For instance, a methylation-level value includes a score (e.g., ranging from 0 to 1) that indicates a percentage or ratio of cytosine bases (e.g., at CpG or other cytosine sites) for particular genomic coordinates or genomic regions to which a methyl group has been added. In some cases, a methylation-level value is expressed as a beta value ( / ?) or an M value. To illustrate, a beta value may estimate a methylation level using a ratio of signal intensities between methylated alleles corresponding to a genomic coordinate and unmethylated alleles corresponding to the genomic coordinate, where 0 represents completely unmethylated and 1 represents completely methylated. In another example, a beta value may comprise a genomic-coordinate-specific fraction of cytosine bases that are methyl-converted to thymine bases. By contrast, an M value may represent a log2 ratio of signal intensities of a methylated probe and an unmethylated probe corresponding to a cytosine base.
[0072] As used herein, the term “variant” refers to a nucleobase or multiple nucleobases that do not align with, differs from, or varies from a corresponding nucleobase (or nucleobases) in a reference sequence or a reference genome. For example, a variant includes a SNP, an indel, or a structural variant that indicates nucleobases in a sample nucleotide sequence that differ from nucleobases in corresponding genomic coordinates of a reference sequence or a reference genome.
[0073] Along these lines, a “variant call” (or “variant nucleobase call”) refers to a nucleobase call comprising a mutation or a variant at a particular genomic coordinate or genomic region with respect to a reference. In particular, a variant call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate or region that differs from a reference nucleobase (or sequence of reference nucleobases) at the same genomic coordinate or region within a reference genome. Conversely, a “reference call” (or “non-variant nucleobase call” or “non-variant call”) refers to a nucleobase call comprising a non-variant or a reference nucleobase at a genomic coordinate or a genomic regionAttorney Docket No. IP-2950-PCT 19 Patent Applicationwith respect to a reference. In particular, a non-variant or reference call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate or region that matches a reference nucleobase (or sequence of reference nucleobases) at the same genomic coordinate or region within a reference genome.
[0074] As used herein, the term “genotype call” refers to a determination or prediction of a particular genotype of a genomic sample at a genomic locus. In particular, a genotype call can include a prediction of a particular genotype of a genomic sample with respect to a reference genome or a reference sequence at a genomic coordinate or a genomic region. For instance, in some cases, a genotype call includes a determination or prediction that a genomic sample comprises both a nucleobase and a complementary nucleobase at a genomic coordinate that is either homozygous or heterozygous for a reference base or a variant (e.g., homozygous reference bases represented as 0|0 or heterozygous for a variant on a particular strand represented as 0|l). A genotype call is often determined for a genomic coordinate or genomic region at which a single nucleotide variant (SNV) or other variant has been identified for a population of organisms. In this disclosure, among other genotype calls, the methylation-genotype-calling system predicts genotype calls for SNV regions within a genomic sample.
[0075] The following paragraphs describe the methylation sequencing map / align system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which a methylation sequencing map / align system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server device(s) 110, and a client device 114. As shown in FIG. 1, the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114 can communicate with each other via a network 118. The network 118 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 8. While FIG. 1 shows an embodiment of the methylation sequencing map / align system 106, this disclosure describes alternative embodiments and configurations below.
[0076] As indicated by FIG. 1, the sequencing device 102 comprises a computing device and a sequencing device system 104 for sequencing a genomic sample or other nucleic-acid polymer. In some embodiments, by executing the sequencing device system 104 using a processor, the sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 102. More particularly, the sequencing device 102 receives nucleotide-sample slides (e.g., flow cells) comprising nucleotideAttorney Docket No. IP-2950-PCT 20 Patent Applicationfragments extracted from samples and further copies and determines the nucleobase sequence of such extracted nucleotide fragments.
[0077] In one or more embodiments, the sequencing device 102 utilizes sequencing-by-synthesis (SBS) techniques to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition, or in the alternative to communicating across the network 118, in some embodiments, the sequencing device 102 bypasses the network 118 and communicates directly with the local device 108, and / or the client device 114. By executing the sequencing device system 104, the sequencing device 102 can further store the nucleobase calls as part of a base-call data fde that is formatted as a binary base call (BCL) file and / or a FASTQ file and send the BCL file and / or FASTQ file to the local device 108, and / or the server device(s) 110. As described further below, in some embodiments, the sequencing device 102 and the computing system 100 can generate and analyze methylation sequencing reads as part of a methylation sequencing assay.
[0078] As further indicated by FIG. 1 , the local device 108 is located at or near a same physical location of the sequencing device 102. Indeed, in some embodiments, the local device 108 and the sequencing device 102 are integrated into a same computing device. The local device 108 may run the sequencing device system 104 and / or the methylation sequencing map / align system 106 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As shown in FIG. 1, the sequencing device 102 may send (and the local device 108 may receive) base-call data generated during a sequencing run of the sequencing device 102. The local device 108 may also communicate with the client device 114. In particular, the local device 108 can send data to the client device 114, including a binary alignment map (BAM) fde, a variant call format (VCF) fde, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0079] As further indicated by FIG. 1, the server device(s) 110 are located remotely from the local device 108 and the sequencing device 102. Like the local device 108, in some embodiments, the server device(s) 110 include a version of (or are otherwise able to access or implement) the methylation sequencing map / align system 106. For example, the server device(s) 110 can implement the methylation sequencing map / align system 106 as part of a sequencing system 112. Accordingly, the server device(s) 110 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such basecall data. As indicated above, the sequencing device 102 may send (and the server device(s) 110 may receive) base-call data from the sequencing device 102. The server device(s) 110 may also communicate with the client device 114. In particular, the server device(s) 110 can send data to the client device 114, including BAM fdes, VCF files, or other sequencing related information.Attorney Docket No. IP-2950-PCT 21 Patent Application
[0080] In some embodiments, the server device(s) 110 comprise a distributed collection of servers where the server device(s) 110 include a number of server devices distributed across the network 118 and located in the same or different physical locations. Further, the server device(s) 110 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0081] As further illustrated and indicated in FIG. 1, by executing a sequencing application 116, the client device 114 can generate, store, receive, and send digital data. In particular, the client device 114 can receive sequencing data from the local device 108 or receive base-call data files (e.g., BCL and FASTQ) and sequencing metrics from the sequencing device 102. Furthermore, the client device 114 may communicate with the local device 108 or the server device(s) 110 to receive a VCF comprising genotype or variant calls and / or other metrics, such as base-call-quality metrics or pass-filter metrics. The client device 114 can accordingly present or display information pertaining to variant calls or other genotype calls within a graphical user interface of the sequencing application 116 to a user associated with the client device 114. For example, the client device 114 can present nucleobase calls, genotype calls, variant calls, and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of the sequencing application 116.
[0082] Although FIG. 1 depicts the client device 114 as a desktop or laptop computer, the client device 114 may comprise various types of client devices. For example, in some embodiments, the client device 114 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 114 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 114 are discussed below with respect to FIG. 8.
[0083] As further illustrated in FIG. 1, the client device 114 includes the sequencing application 116. The sequencing application 116 may be a web application or a native application stored and executed on the client device 114 (e.g., a mobile application, desktop application). The sequencing application 116 can include instructions that (when executed) cause the client device 114 to receive data from the methylation sequencing map / align system 106 and present, for display at the client device 114, base-call data or data from an alignment data file or VCF. Furthermore, the sequencing application 116 can instruct the client device 114 to display summaries for multiple sequencing runs.
[0084] As further illustrated in FIG. 1, a version of the methylation sequencing map / align system 106 may be located and / or implemented (e.g., entirely or in part) on the client device 114 or the sequencing device 102. In yet other embodiments, the methylation sequencing map / align system 106 is implemented by one or more other components of the computing system 100, such as the local device 108. In particular, the methylation sequencing map / align system 106 can beAttorney Docket No. IP-2950-PCT 22 Patent Applicationimplemented in a variety of different ways across the sequencing device 102, the local device 108, the server device(s) 110, and the client device 114. For example, the methylation sequencing map / align system 106 can be downloaded from the server device(s) 110 to the client device 114 and / or the local device 108 where all or part of the functionality of the methylation sequencing map / align system 106 is performed at each respective device within the computing system 100.
[0085] FIG. 2 illustrates an overview of the methylation sequencing map / align system 106 mapping and aligning methylation sequencing reads for genotype calling in conjunction with a methylation sequencing assay and in accordance with one or more embodiments of the present disclosure. By way of overview, FIG. 2 illustrates the methylation sequencing map / align system 106 performing a series of acts comprising an act 210 of accessing read subsequences (e.g., k-mers) from methylation sequencing reads. The methylation sequencing map / align system 106 further performs an act 212 of converting one or more nucleobases within the read subsequences from one base type to another base type, The methylation sequencing map / align system 106 further performs an act 226 of mapping the converted read subsequences to converted reference subsequences, and an act 230 of aligning methylation sequencing reads with one or more reference sequences. Having mapped and aligned methylation sequencing reads, as further depicted in FIG. 2, the methylation sequencing map / align system 106 can perform an act 240 of genotype calling and — optionally and in parallel — an act 250 of determining methylation-level value(s) for nucleobase(s) within a sample nucleotide sequence 204.
[0086] As shown in FIG. 2, for instance, the methylation sequencing map / align system 106 inputs or runs the sample nucleotide sequence 204 through a methylation sequencing assay 202. As described below, the methylation sequencing assay 202 detects or quantifies methylation of cytosine bases at particular genomic regions and / or in particular cell types — thereby determining the methylation-level value(s). As part of quantifying methylation, the methylation sequencing assay 202 converts mC to uracil (U) or thymine (T). But the methylation sequencing assay 202 can perform other types of conversions, as explained in the paragraphs following this description of FIG. 2.
[0087] As indicated above, the sample nucleotide sequence 204 comprises cytosine bases 200. As depicted in FIG. 2, for instance, the sample nucleotide sequence 204 includes fifteen of the cytosine bases 200, including both methylated and unmethylated cytosine bases. As depicted by FIG. 2, the open circles of the sample nucleotide sequence 204 represent twelve unmethylated cytosine bases, and the dark-filled circles of the sample nucleotide sequence 204 represent three methylated cytosine bases.
[0088] As further part of the methylation sequencing assay 202, in some embodiments, the methylation sequencing map / align system 106 amplifies and determines nucleobase calls for theAttorney Docket No. IP-2950-PCT 23 Patent Applicationsample nucleotide sequence 204 and complementary strands using the sequencing device 102. The sequencing device 102 can accordingly generate a BCL fde, FASTQ fde, or other base-call-data file comprising methylation sequencing reads 206. In some such cases, the methylation sequencing map / align system 106 uses SBS to determine nucleobase calls for the sample nucleotide sequence 204 when sequencing or amplifying a methylation sequencing read of the methylation sequencing reads 206, including thymine nucleobase calls for one or more of the cytosine bases 200 that have been converted into uracil bases or thymine bases. Along with other determined nucleotide reads, in some cases, the sequencing device 102 sends the base-call data to the server device(s) 110.
[0089] As further shown in FIG. 2, the methylation sequencing map / align system 106 performs the act 210 of accessing read subsequences. For example, the read subsequences can be subsequences extracted or otherwise taken from methylation sequencing reads generated by the methylation sequencing assay 202 that has been performed on a target genomic sample. The read subsequences can comprise k-mers. In some embodiments, the read subsequences comprise singlestranded read subsequences. For illustrative purposes, FIG. 2 depicts an example read subsequence as “ — CG — TG — ” with capital letters representing nucleobases in candidate methylation sites.
[0090] The read subsequences accessed at act 210 can include one or more nucleobases that have been modified by the methylation sequencing assay 202. For example, in embodiments where the methylation sequencing assay 202 converts methylated cytosines (e.g., 5mC, 5hmC) to thymines, the read subsequences can comprise a thymine guanine (TG) dinucleotide where, in the original nucleic acid in the target genomic sample, there was a cytosine guanine (CG) dinucleotide with a methylated cytosine. In some embodiments that include a directional methylation sequencing assay, read subsequences can be from sense strands, or from reverse complements of antisense strands.
[0091] After accessing the read subsequences, as shown in FIG. 2, the methylation sequencing map / align system 106 performs the act 212 of converting nucleobases within the read subsequences. Unlike the enzymatic or chemical conversions of nucleobases performed by the methylation sequencing assay 202, however, the methylation sequencing map / align system 106 converts nucleobase calls (e.g., “C” or “T”) or other data representing nucleobases in the read subsequences as part of a bioinformatic process. For example, during the act 212, the methylation sequencing map / align system 106 can identify one or more candidate methylation sites (e.g., CG dinucleotides) within the read subsequences, and convert one or more nucleobases at the one or more candidate methylation sites from one base type to another base type (e.g., from C to T, or from Gto A, as further described herein). By performing the act 212, the methylation sequencing map / align system 106 can generate converted read subsequences. The converted read subsequences can comprise, for example, single-stranded converted read subsequences. For illustrative purposes,Attorney Docket No. IP-2950-PCT 24 Patent ApplicationFIG. 2 depicts an example converted read subsequence as “ — TG — TG — ” with capital letters representing nucleobases in candidate methylation sites.
[0092] In some embodiments, the act 212 includes the methylation sequencing map / align system 106 selectively converting one or more nucleobases at the one or more candidate methylation sites. In embodiments where the candidate methylation site comprises a cytosine guanine (CG) dinucleotide, for instance, the methylation sequencing map / align system 106 selectively converts one or more nucleobases of a CG dinucleotide within a read subsequence from a sense strand or within a read subsequence from a reverse complement of an antisense strand. For example, in some embodiments, converting the one or more nucleobases does not replace nucleobases outside of CG dinucleotides (e.g., does not replace a C that is not part of a CG dinucleotide). In other words, in some embodiments, the methylation sequencing map / align system 106 performs the act 212 by converting nucleobases only if they are within the context of a candidate methylation site, such as CG dinucleotides. Thus, in some embodiments, the methylation sequencing map / align system 106 identifies CG dinucleotides within the read subsequences and selectively replaces one or more nucleobases within the CG dinucleotides.
[0093] In some embodiments, as part of the act 212, the methylation sequencing map / align system 106 replaces a nucleobase of the candidate methylation site (e.g., CG dinucleotide) with one base type instead of another base type. For example, in some embodiments, the methylation sequencing map / align system 106 extracts subsequences (e.g., k-mers) from a target genomic sample’s methylation sequencing reads and converts all remaining CG sites to either TG (if the read is from the sense strand) or CA (if the read is from the antisense strand).
[0094] For example, as shown in FIG. 2, an example read subsequence shown at act 210 includes a CG dinucleotide that was not modified by the methylation sequencing assay 202, for example, because the cytosine base in the exemplary target genomic sample was not methylated. During the act 212, the methylation sequencing map / align system 106 has converted the cytosine base of the CG dinucleotide to a thymine base, resulting in a TG dinucleotide.
[0095] As indicated above, in certain implementations, the methylation sequencing map / align system 106 converts nucleobases as part of a bioinformatic process only if the nucleobases are part of a candidate methylation site. But such bioinformatic-based or data-based conversion can be selective. When a nucleobase is positioned at the end of a methylation sequencing read such that it is unclear whether the nucleobase is part of a CG dinucleotide (e.g., the methylation sequencing read begins with a G or ends with a C), in some embodiments, the methylation sequencing map / align system 106 does not convert that nucleobase.
[0096] In addition to accessing and converting read subsequences, as further shown in FIG. 2, the methylation sequencing map / align system 106 performs the act 222 of accessing referenceAttorney Docket No. IP-2950-PCT 25 Patent Applicationsubsequences from the one or more reference sequences 220. Accordingly, the methylation sequencing map / align system 106 generates reference subsequences from the one or more reference sequences 220. Similar to the read subsequences discussed above with reference to the act 210, the reference subsequences in the act 222 can comprise k-mers. For illustrative purposes, FIG. 2 depicts an example reference subsequence as “ — CG — CG — ” with capital letters representing nucleobases in candidate methylation sites.
[0097] Having accessed the reference subsequences, as further shown in FIG. 2, the methylation sequencing map / align system 106 optionally performs the act 224 of converting reference subsequences. Again, unlike the enzymatic or chemical conversions of nucleobases performed by the methylation sequencing assay 202, the methylation sequencing map / align system 106 converts nucleobase calls (e.g., “C” or “T”) or other data representing nucleobases in the reference subsequences as part of a bioinformatic or data-manipulation process. In some embodiments, the methylation sequencing map / align system 106 converts, at a candidate methylation site within the reference subsequences, one or more nucleobases from one base type to another base type to generate converted reference subsequences. For example, the methylation sequencing map / align system 106 can identify candidate methylation sites (e.g., CG dinucleotides) within the reference subsequences and replace a nucleobase within the candidate methylation site with another nucleobase of another base type. Thus, in some embodiments, the converted reference subsequences comprise a replaced nucleobase of a CG dinucleotide differing from a reference nucleobase of the one or more reference sequences. For illustrative purposes, FIG. 2 depicts an example converted reference subsequence as “ — TG — TG — ” with capital letters representing nucleobases in candidate methylation sites.
[0098] In some embodiments, the methylation sequencing map / align system 106 generates two sets of converted reference subsequences, where a first set of converted reference subsequences comprise a first replaced nucleobase and a second set of converted reference subsequences comprise a second, different replaced nucleobase. This double-set approach can be advantageous based on the methylation sequencing assay 202 when, for example, the methylation sequencing assay 202 generates two different base conversions among the methylation sequencing reads. Accordingly, the methylation sequencing map / align system 106 can generate converted reference subsequences to accord with possible conversions from the methylation sequencing assay 202 among the methylation sequencing reads 206. In some embodiments, the methylation sequencing map / align system 106 generates a first set of converted reference subsequences by replacing cytosine nucleobases of CG dinucleotides with thymine nucleobases; and generates a second set of converted reference subsequences by replacing guanine nucleobases of CG dinucleotides with adenine nucleobases.Attorney Docket No. IP-2950-PCT 26 Patent Application
[0099] After generating or accessing converted reference subsequences, in some embodiments, the methylation sequencing map / align system 106 stores the converted reference subsequences in a data structure, such as a hash table. In some embodiments, the data structure further comprises unconverted reference subsequences that do not include a given candidate methylation site. While the methylation sequencing map / align system 106 can construct the data structure (e.g., hash table) including converted reference subsequences, in other embodiments, this data structure has already been generated, and the methylation sequencing map / align system 106 merely accesses the data structure comprising converted reference subsequences stored in memory.
[0100] After generating or accessing both converted read subsequences and reference subsequences, as shown in FIG. 2, the methylation sequencing map / align system 106 performs the act 226 of mapping the converted read subsequences to the converted reference subsequences. In some embodiments, the methylation sequencing map / align system 106 maps converted read subsequences to converted reference subsequences stored in a hash table. As indicated above, in some embodiments, the converted reference subsequences include a first set of converted reference subsequences in which all CG sites are converted to TG and a second set of converted reference subsequences in which all CG sites are converted to CA.
[0101] The methylation sequencing map / align system 106 can map the converted read subsequences to the converted reference subsequences by any method. For example, the methylation sequencing map / align system 106 can map the converted read subsequences to the converted reference subsequences by seed mapping and chaining. Such seed mapping and chaining can be performed by, for instance, matching converted read subsequences to converted reference sequences and chaining the matches with additional matches. For illustrative purposes, FIG. 2 depicts an example of mapping a converted read subsequence to a converted reference subsequence as “ — TG — TG — ” above “ — TG — TG — ” with capital letters representing nucleobases in candidate methylation sites.
[0102] In some embodiments, the methylation sequencing map / align system 106 also maps unconverted read subsequences that do not include a given candidate methylation site to one or more unconverted reference subsequences that do not include a given candidate methylation site. Thus, the mapping at act 226 does not necessarily have to be limited to converted-to-converted subsequences and can integrate mapping between converted read subsequences and unconverted read subsequences to converted reference subsequences and unconverted reference subsequences.
[0103] Having mapped read subsequences extracted from the methylation sequencing reads 206, the methylation sequencing map / align system 106 can identify genomic coordinates of matches from the mapping process for use in downstream alignment steps detailed below. ForAttorney Docket No. IP-2950-PCT 27 Patent Applicationexample, based on the mapped genomic coordinates, the methylation sequencing map / align system 106 can identify a genomic region having genomic coordinates in the reference sequence.
[0104] After mapping the read subsequences, as further shown in FIG. 2, the methylation sequencing map / align system 106 performs the act 230 of aligning methylation sequencing reads. For example, the methylation sequencing map / align system 106 can generate candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences based on genomic coordinates determined during the mapping process of act 226. For illustrative purposes, FIG. 2 depicts an example of aligning a methylation sequencing read with a reference sequence as “ - CG — TG - ” abovewith capital letters representing nucleobases in candidate methylation sites.
[0105] To support more accurate read alignments, the methylation sequencing map / align system 106 can generate candidate alignments between the target genomic sample’s original methylation sequencing reads and original reference sequences. Accordingly, the methylation sequencing map / align system 106 can generate candidate alignments of the methylation sequencing reads 206 to the one or more reference sequences 220, which have not undergone the process of bioinformatical or data-based conversion at act 212 or act 224. Thus, in some embodiments, when the methylation sequencing map / align system 106 generates candidate alignments, the methylation sequencing reads 206 will only comprise conversions that have been physically (e.g., chemically or enzymatically) caused by the methylation sequencing assay 202. Accordingly, in some embodiments, the one or more reference sequences 220 does not include any converted nucleobases. Thus, in some embodiments, generating the candidate alignments comprises generating one or more candidate alignments of an unconverted methylation sequencing read to an unconverted reference sequence.
[0106] For example, in the example depicted in FIG. 2, the methylation sequencing map / align system 106 generates a candidate alignment of (a) a read subsequence including a CG dinucleotide (e.g., that has not been bioinformatically converted to a TG dinucleotide) and a TG dinucleotide (e.g., that may have been physically converted from a CG dinucleotide by the methylation sequencing assay 202), to a (b) reference sequence that includes two CG dinucleotides (e.g., that have not been bioinformatically converted to a TG dinucleotide).
[0107] As part of generating candidate alignments for the act 230, the methylation sequencing map / align system 106 can determine alignment scores for the candidate alignments. In some embodiments, for instance, the methylation sequencing map / align system 106 scores the candidate alignments based on matches and mismatches between nucleobases in the methylation sequencing reads 206 and the one or more reference sequences 220. As further described below, the methylation sequencing map / align system 106 can determine a read-reference-base mismatchAttorney Docket No. IP-2950-PCT 28 Patent Applicationbetween a nucleobase of a methylation sequencing read (e.g., of the methylation sequencing reads 206) and a corresponding nucleobase of the one or more reference sequences 220, that occurs in the context of a candidate methylation site (e.g., in a CG context). This disclosure describes an embodiment of the methylation sequencing map / align system 106 determining read-reference-base mismatches, mapping read subsequences, and determining alignment scores below with respect to FIGS. 4 A and 4B.
[0108] As indicated above, the methylation sequencing map / align system 106 can determine a mismatch-tolerant-adjusted alignment score for candidate alignments. In some embodiments, the methylation sequencing map / align system 106 scores a mismatch differently if the mismatch is within the context of a candidate methylation site (e.g., in a CG context) than if the mismatch is not within the context of the candidate methylation site, as further described with respect to FIG.4B. In some embodiments, the methylation sequencing map / align system 106 scores candidate alignments nucleobase-by-nucleobase using a tolerance metric (e.g., CG-conversion mismatch metric) such that nucleobase mismatches that are in a CG context are not penalized to the same degree as mismatches that are not in a CG context. For example, the methylation sequencing map / align system 106 can give match a score of 1, a CG-conversion mismatch a score of 0, a masked (e.g., soft-clipped) nucleobase a score of 0, and a (non-CG) mismatch a score of -5.
[0109] The methylation sequencing map / align system 106 can select a candidate alignment based on alignment scoring. For example, the methylation sequencing map / align system 106 can select a candidate alignment with a highest score. Because of the mismatch-tolerant-adjusted alignment scoring, a candidate alignment comprising a mismatch in a CG context may be the highest scoring candidate alignment.
[0110] Based on the selected candidate alignments, as further shown in FIG. 2, the methylation sequencing map / align system 106 optionally performs the act 240 of genotype calling. In some embodiments, the methylation sequencing map / align system 106 determines a genotype call for the target genomic sample at the candidate methylation site based on a selected alignment of the candidate alignments of the methylation sequencing reads to the reference sequences.
[0111] For example, the methylation sequencing map / align system 106 can determine genomic coordinates or other candidate sites comprising methylated nucleobases within the target genomic sample based on comparing selected alignments of methylation sequencing reads 206 to the one or more reference sequences 220. As part of the genotype calling, the methylation sequencing map / align system 106 can determine variant calls or reference calls at genomic coordinates or other candidate sites based on comparing selected alignments of methylation sequencing reads 206 to the one or more reference sequences 220. Such variant calls can include, for example, single nucleotide variants (SNVs) and / or insertions and / or deletions (INDELs), or structural variantsAttorney Docket No. IP-2950-PCT 29 Patent Application(SVs) based on analyses comparing selected alignments of methylation sequencing reads to one or more reference sequences. In some embodiments, the methylation sequencing map / align system 106 determines a genotype call in accordance with methods described in as described in U.S. Provisional Pat. App. No. 63 / 726,005, entitled “Genotype Calling from Sequencing Data Representing Spontaneous Deamination and / or Methylation,” (IP-2816-PRV) filed November 27, 2024; and International Application No. PCT / US2025 / 057329, entitled “Genotype Calling from Sequencing Data Representing Spontaneous Deamination and / or Methylation,” (IP-2816-PCT) filed November 26, 2025. Bot of the disclosures of which are incorporated herein by reference in their entirety.
[0112] In addition to genotype calling, as further shown in FIG. 2, the methylation sequencing map / align system 106 optionally performs the act 250 of determining methylation-level value(s). In some embodiments, the methylation sequencing map / align system 106 determines a methylation-level based on a comparison of the methylation sequencing reads 206 to the one or more reference sequences 220 or to other nucleotide reads that have not undergone chemical or enzymatic methylation conversion (not illustrated), or based on other analyses of the methylation sequencing reads 206. In some embodiments, the methylation sequencing map / align system 106 determines estimated methylation-level value(s) in accordance with methods described in International Application No. PCT / US2023 / 076472, entitled “Detecting and Correcting Methylation Values from Methylation Sequencing Assays,” fded October 10, 2023 (IP-2455 -PCT), and / or in accordance with methods described in International Application No. PCT / US2024 / 035562, entitled “Variant Calling with Methylation-Level Estimation,” filed June 26, 2024 (IP-2578-PCT), the disclosures of which are incorporated herein by reference in their entirety.
[0113] As noted above, the methylation sequencing assay 202 can perform various types of methylation sequencing assays and perform different types of conversions. By way of overview, a methylation sequencing assay can use different methylation sequencing protocols that utilize different conversions. For example, some methylation sequencing protocols convert unmethylated cytosine bases to thymine bases (which are sometimes represented as C-to-T conversions). Other methylation sequencing protocols may convert methylated cytosine bases to thymine bases (which are sometimes represented as 5mC, 5hmC-to-T conversions). The methylation sequencing map / align system 106 may utilize methylation data stemming from either type of methylation sequencing protocol. The following paragraphs further detail various effects of C-to-T conversion protocols and 5mC, 5hmC-to-T conversion protocols, among other conversion protocols.
[0114] In the C-to-T conversion protocols, unmethylated bases are converted into thymine bases. For instance, the C-to-T conversion protocols include Bisulfite and Enzymatic Methyl-seqAttorney Docket No. IP-2950-PCT 30 Patent Application(EM-seq) used for methylation sequencing assays. EM-seq can be performed, for instance, as described by Romualdas Vaisvila et al., Enzymatic Methyl Sequencing Detects DNA Methylation at Single-Base Resolution from Picograms of DNA, 30 Genome Research 1280-1289 (2021), which is hereby incorporated by reference in its entirety.
[0115] In 5mC, 5hmC-to-T conversion protocols, methylated bases are converted into thymine bases. For instance, Tet-assisted pyridine borane sequencing (TAPS) uses a ten-eleven translocation (TET) enzyme for a methylation sequencing assay, as described by Yibin Liu et al., “Bisulfite-free Direct Detection of 5-Methylcystosine and 5-Hydroxymethylcystosine at Base Resolution,” 36 Nature Biotechnology 424-29 (2019). In some assays that rely on a TET enzyme, a methylation sequencing assay converts 5-Methylcystosine (5mC) and 5-Hydroxymethylcystosine (5hmC) into oxidized products using a TET enzyme and then uses an Apolipoprotein B mRNA Editing Enzyme, Catalytic Polypeptide (APOBEC) 3 A or another APOBEC protein to deaminate unmodified cytosines by converting them to uracil bases. In some cases, such converted uracil bases are detected as thymine bases during sequencing.
[0116] In some embodiments presented herein, the methylation sequencing map / align system 106 uses methylation sequencing reads from bisulfite converted samples. In bisulfite sequencing, DNA is chemically treated with sodium bisulfite, which results in the conversion of unmethylated cytosine bases to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, the modified cytosine bases, 5mC and 5hmC, are resistant to bisulfite conversion, and are sequenced as cytosine bases. It will be understood by one of ordinary skill in the art that other methylation conversion methods can also be used to generate nucleotide reads (e.g., methylation sequencing reads) for use in the methods and embodiments presented herein. For example, an alternative to bisulfite conversion is Enzymatic Methyl-seq (EM-seq). EM-seq has been described in neb.com / - / media / nebus / files / manuals / manuale7120.pdf, the content of which is incorporated herein by reference in its entirety. Briefly, EM-seq conversion uses an enzymatic method that results in the conversion of unmethylated cytosine bases to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, the modified cytosine bases, 5mC and 5hmC, are resistant to the enzymatic conversion, and are sequenced as cytosine bases. Since only a very small fraction of cytosines in a typical sample are methylated, in some embodiments, the result of either bisulfite or EM-seq conversion methods is a set of sequence reads that is substantially made up of only 3 bases: A, T, and G.
[0117] Another example of an alternative method to bisulfite conversion is TET-assisted pyridine borane sequencing (TAPS). TAPS has been described in Liu, et al. Bisulfite-free direct detection of 5 -methylcytosine and 5-hydroxymethylcytosine at base resolution. Nat Biotechnol.2019 Apr;37(4):424-429. doi:10.1038 / s41587-019-0041-2, the content of which is incorporatedAttorney Docket No. IP-2950-PCT 31 Patent Applicationherein by reference in its entirety. TAPS results in conversion of 5mC and 5hmC to uracil bases, and the resulting uracil bases are ultimately sequenced as thymine bases. In contrast, unmethylated cytosine bases are resistant to the TAPS conversion, and are sequenced as cytosine bases. One of ordinary skill in the art will recognize that, where TAPS conversion is utilized to generate nucleotide reads (e.g., methylation sequencing reads), the reference sequences described herein can be modified accordingly to reflect methylated targets where methylated cytosine bases are converted and sequenced as thymine bases (C-to-T conversion) and unmethylated cytosine bases are not converted and sequenced as cytosine bases. An enzymatic alternative to TAPS involves use of a modified cytidine deaminase enzyme, engineered to selectively deaminate only 5mC and 5hmC, while unmethylated cytosine bases are not converted. Similar to TAPS, this modified cytidine deaminase method results in conversion of 5mC and 5hmC to uracil bases or of 5mC and 5hmC to thymine bases or modified thymine analogs that are sequenced as thymine bases. In contrast, unmethylated cytosine bases are resistant to the TAPS conversion, and are sequenced as cytosine bases. Use of modified cytidine deaminase enzymes, such as AP0BEC3A, has been described in International Application No. PCT / US2023 / 17846, filed April 7, 2023, and titled “Altered Cytidine Deaminases and Methods of Use,” the contents of which is incorporated herein by reference in its entirety. One of ordinary skill in the art will recognize that, where this modified cytidine deaminase conversion method is utilized to generate nucleotide reads (e.g., methylation sequencing reads), the reference sequences described herein can be modified accordingly to reflect methylated targets where methylated cytosine bases are converted and sequenced as thymine bases (C-to-T conversion) and unmethylated cytosine bases are not converted and sequenced as cytosine bases. One of ordinary skill in the art will recognize that since only the methylated cytosines in a typical sample are converted using these alternative conversion methods, in some embodiments, the result of either TAPS or modified cytidine deaminase conversion methods is a set of sequence reads that is substantially made up of 4 bases: A, T, C and G.
[0118] As indicated above, in accordance with one or more embodiments, FIG. 3A depicts the methylation sequencing map / align system 106 accessing sequences from a variant site and a methylated site in the context of a methylation sequencing assay and converting nucleobases at a candidate methylation site. In particular, FIG. 3A compares the characteristics of methylation sequencing reads that cover a variant site with methylation sequencing reads that cover a methylated site. FIG. 3 A also illustrates exemplary strand-specific conversions that the methylation sequencing map / align system 106 generates at a candidate methylation site. Accordingly, both variant site and methylated site depicted in FIG. 3A represent examples of candidate methylation sites.Attorney Docket No. IP-2950-PCT 32 Patent Application
[0119] Panel 300 in FIG. 3 A, for instance, schematically illustrates nucleotide sequences generated at a site with a heterozygous C to T variant in the context of a methylation sequencing assay that converts methylated cytosine to thymine. In the example of FIG. 3A, in panel 300, a nucleic acid from one haplotype in the target genomic sample includes a thymine (T) at a site 301a followed by a guanine (G). The target genomic sample undergoes a directional methylation sequencing assay that converts methylated cytosine to thymine, which does not affect the T at a site 301b. As shown in panel 300, after sequencing and alignment (with the sequence from the antisense strand being reverse-complemented), both the sense and antisense strand show a T at the site 301b. The methylation sequencing reads 306 are aligned to the reference sequence, which has a C at a site 301c in an alignment pile-up 305. As shown in the alignment pile-up 305, about half of the methylation sequencing reads 306 that cover the site 301c have a T, because target genomic sample is heterozygous for the sequence variant, but the T is found in both the sense and the antisense strand.
[0120] By contrast, panel 320 in FIG. 3A schematically illustrates nucleotide sequences generated at a methylated site with a methylation sequencing assay that converts methylated cytosine to thymine. For example, as illustrated in FIG. 3A, panel 320, some types of methylation sequencing assays result in converted methylated CG dinucleotides being read as TG in methylation sequencing reads from a sense strand, and as CA in methylation sequencing reads from the antisense strand. Accordingly, the methylation sequencing map / align system 106 generates converted read subsequences to accord with possible conversions from the methylation sequencing assay.
[0121] For example, in the embodiment illustrated in panel 320, a nucleic acid from the target genomic sample includes a methylated cytosine (C) followed by a guanine (G) at a candidate methylation site 307a. In particular, the target genomic sample undergoes a directional methylation sequencing assay that converts methylated cytosine to thymine, which converts the methylated cytosine at the candidate methylation site 307a to thymine. As shown in panel 320, after sequencing and alignment (with the sequence from the antisense strand being reverse-complemented), the sense strand sequence shows a TG dinucleotide at a candidate methylation site 307b, and the antisense strand sequence has a CA dinucleotide at the candidate methylation site 307b. The methylation sequencing reads 309 are aligned to the reference sequence which has a CG at a candidate methylation site 307c in an alignment pile-up 308. As seen in the alignment pile-up 308, the methylation sequencing reads 309 from the sense strand have a TG dinucleotide (e.g., a conversion from C to T) at the candidate methylation site 307c, while the methylation sequencing reads 309 from the antisense strand have a CA dinucleotide (e.g., a conversion from Gto A) at the candidate methylation site 307c.Attorney Docket No. IP-2950-PCT 33 Patent Application
[0122] In some embodiments, alignments of methylation sequencing reads including a conversion of a methylated nucleobase at a candidate methylation site are different from alignments of methylation sequencing reads with a sequence variant at the candidate methylation site. In some embodiments, this can enable the methylation sequencing map / align system 106 (or the sequencing system 112 encompassing the methylation sequencing map / align system 106) to distinguish between nucleobase differences due to sequence variants in the target genomic sample versus nucleobase differences due to conversion of methylated nucleobases, e.g., by the methylation sequencing assay 202. For example, the alignment pile-up 305 and alignment pile-up 308 are distinguishable from each other. The alignment pile-up 305 includes the C to T variant in methylation sequencing reads 306 from both sense and antisense strands, while the alignment pileup 308 includes a T instead of C only on the methylation sequencing reads 309 from the sense strand and an A instead of a G only on the methylation sequencing reads 309 from the antisense strand. In some embodiments, the methylation sequencing map / align system 106 (or the sequencing system 112 encompassing the methylation sequencing map / align system 106) can distinguish between sequence variant loci and methylation loci based on such patterns.
[0123] In FIG. 3A, panel 330 schematically illustrates bioinformatic conversion of a CG dinucleotide to a TG dinucleotide in a subsequence from a sense strand, and to a CA dinucleotide in a subsequence from a reverse complement of an antisense strand. Based on the methylation sequencing assay conversion pattern shown in panel 320, in some embodiments, the methylation sequencing map / align system 106 replaces a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand and replaces a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand. Similarly, in some embodiments, the methylation sequencing map / align system 106 generates a first set of converted reference subsequences by replacing a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase, and a second set of converted reference subsequences by replacing a guanine nucleobase of a CG dinucleotide with an adenine nucleobase.
[0124] As illustrated in panel 330 in FIG. 3A, during a bioinformatic conversion process, the methylation sequencing map / align system 106 can replace a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand (e.g., CG to TC), and can replace a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand (e.g., CGto CA).
[0125] While in the context of FIG. 3A, the methylation sequencing assay selectively converts methylated cytosines to thymine (including the methylated cytosine of the candidate methylation site 307a). As indicative above, in alternative embodiments, a site like that of the candidateAttorney Docket No. IP-2950-PCT 34 Patent Applicationmethylation site 307a includes an unmethylated cytosine, and a methylation sequencing assay selectively converts unmethylated cytosines to uracil or thymine (e.g., either mOT or uOT), leading to the same pattern of nucleobase conversions in methylation sequencing reads as the pattern depicted in FIG. 3A for methylation sequencing reads in the alignment pile-up 308.
[0126] After converting nucleobases within read subsequences, as depicted in FIGS. 3B and 3C, the methylation sequencing map / align system 106 can (i) map converted read subsequences to converted reference subsequences and (ii) align methylation sequencing reads from which the read subsequences derive to reference sequences. For instance, FIG. 3B depicts the methylation sequencing map / align system 106 mapping converted read subsequences to converted reference subsequences, in accordance with one or more embodiments of the present disclosure. In the example shown in FIG. 3B, the methylation sequencing map / align system 106 accesses a methylation sequencing read 310. The methylation sequencing read 310 includes a T, which has been enzymatically or chemically converted from a methylated cytosine by the methylation sequencing assay 202, followed by a G. The methylation sequencing read 310 further includes a C, which was not enzymatically or chemically converted by the methylation sequencing assay 202 because it was not methylated, followed by a G. The methylation sequencing read 310 further includes a read subsequence that does not include a candidate methylation site (e.g., does not include a CG).
[0127] As indicated above, in some embodiments, the methylation sequencing map / align system 106 extracts converted read subsequences from the methylation sequencing read 310. Because the methylation sequencing read 310 in FIG. 3B is from a sense strand, the methylation sequencing map / align system 106 replaces all CG dinucleotides with TG as part of converting a read subsequence extracted from the methylation sequencing read 310.
[0128] In addition to extracting and converting read subsequences from methylation sequencing reads, in certain implementations, the methylation sequencing map / align system 106 extracts and converts reference subsequences from references sequences (e.g., reference sequences of a reference genome). As shown in FIG. 3B, accordingly, the methylation sequencing map / align system 106 further performs an act 312 of converting subsequences taken from a reference sequence data fde 325, which can include one or more reference sequences from a reference genome. As depicted in FIG. 3B, the methylation sequencing map / align system 106 generates a first converted reference subsequence with all CG dinucleotides converted to TG, and a second converted reference subsequence with all CG dinucleotides converted to CA.
[0129] After converting the extracted reference subsequences, as depicted in FIG. 3B, the methylation sequencing map / align system 106 further performs an act 314 of mapping the converted and unconverted read subsequences to converted and unconverted reference sequences.Attorney Docket No. IP-2950-PCT 35 Patent ApplicationFor example, the methylation sequencing map / align system 106 compares nucleobases in the converted read subsequences to nucleobases in the converted reference sequences. As part of the act 314, the methylation sequencing map / align system 106 can also map unconverted read subsequences (e.g., that did not have a candidate methylation site) to unconverted reference subsequences (e.g., that did not have a candidate methylation site) in a similar fashion. In the example of FIG. 3B, if the mapping of a read subsequence to a reference subsequence includes no mismatches, the methylation sequencing map / align system 106 stores a genomic coordinate from the mapped reference sequence as a candidate genomic coordinate for generating candidate alignments.
[0130] After mapping read subsequences to reference subsequences, as depicted in FIG. 3C and in accordance with one or more embodiments, the methylation sequencing map / align system 106 aligns methylation sequencing reads to one or more reference sequences. To facilitate generating candidate alignments, the methylation sequencing map / align system 106 adds genomic coordinates from the mapping at act 314 as candidate alignments of corresponding methylation sequencing reads to one or more reference sequences.
[0131] During act 316 shown in FIG. 3C, the methylation sequencing map / align system 106 determines alignment scores for a candidate alignment of the methylation sequencing read 310 to one or more reference sequences. As shown here, the methylation sequencing read 310 represents the original and non-data-converted read from a methylation sequencing assay (e.g., the methylation sequencing assay 202) and includes converted nucleobases from the methylation sequencing assay, but does not include any nucleobases converted bioinformatically. Similarly, the reference sequence shown as part of the act 316 represents the original reference sequence from the reference sequence datafile.
[0132] As part of the example alignment score for the candidate alignment shown in FIG. 3C, the methylation sequencing map / align system 106 determines a match metric of 1 for a match between nucleobases, a mismatch metric of -5 for a mismatched pair of nucleobases outside of a CG context, a tolerance metric of 0 for a CG-conversion mismatch, and a soft-clip metric of 0 for a soft-clipped (e.g., masked) nucleobase. In the example shown in FIG. 3C, for the unconverted subsequence, there are 9 bases that are all matches with a match metric of 1, providing a match metric of 9 * 1 = 9 for the unconverted subsequence. For the converted subsequence, there is 1 CG-converted mismatch with a tolerance metric of 0 and 8 matches and with a match metric of 1, therefore, the combined metric is 1 *0 + 8* 1 = 8 for the converted subsequence. The total alignment score for the candidate alignment shown in FIG. 3C is, therefore, 9+8=17. In this summed-metric approach to alignment scoring, the methylation sequencing map / align system 106 thereforeAttorney Docket No. IP-2950-PCT 36 Patent Applicationdetermines an alignment score of 17 for the candidate alignment in the act 316 in the exemplary embodiment shown in FIG. 3C.
[0133] In the embodiment shown in FIG. 3C, accordingly, the tolerance metric of 0 neither increases nor decreases the alignment score. But those of ordinary skill in the art will recognize that the tolerance metric can be adapted based on the alignment scoring scheme and the other alignment score components.
[0134] As indicated above and to provide more detail on alignment scoring, FIGS. 4A-4B depict the methylation sequencing map / align system 106 determining read-reference-base mismatches and determining alignment scores for candidate alignments. For example, FIG. 4A illustrates an embodiment of the methylation sequencing map / align system 106 determining a read-reference-base mismatch within a CG context, also referred to herein as a CG-conversion mismatch, in accordance with one or more embodiments of the present disclosure.
[0135] As part of evaluating candidate alignments, the methylation sequencing map / align system 106 can determine a read-reference-base mismatch between a target nucleobase in one or more methylation sequencing reads and the corresponding nucleobase in the one or more reference sequences that occurs within the context of a candidate methylation site. As described in more detail above, such a read-reference-base mismatch is an example of a conversion mismatch. As a specific example of a read-reference-base mismatch within a candidate methylation site, in some embodiments, the methylation sequencing map / align system 106 can determine a read-reference-base mismatch that occurs within a CG context as the candidate methylation site, also referred to a CG-conversion mismatch. Below are further details regarding determining a read-reference-base mismatch is within a CG context. Similar approaches can be used for other types of conversion mismatches, e.g., mismatches that are within the context of another type of candidate methylation site.
[0136] The methylation sequencing map / align system 106 can use one approach or a combination of several approaches to determine a read-reference-base mismatch that is within a CG context of a candidate alignment. For example, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on a nucleotide sequence in the one or more reference sequences at the site of the read-reference-base mismatch. In some such embodiments, the methylation sequencing map / align system 106 determines the read-reference-base mismatch is within a CG context by determining that the mismatch between target nucleobase of the methylation sequencing read and the corresponding reference nucleobase is at a cytosine guanine (CG) dinucleotide within the one or more reference sequences. In this example, the methylation sequencing map / align system 106 would determine whether the one or more reference sequences has a CG dinucleotide at the site of the mismatch. IfAttorney Docket No. IP-2950-PCT 37 Patent Applicationso, the methylation sequencing map / align system 106 can count the mismatch as a CG-conversion mismatch for purposes of alignment scoring as further described with respect to FIG. 4B. The following paragraphs describe other approaches that the methylation sequencing map / align system 106 can utilize (in addition or in the alternative to the approach described in this paragraph) to determine that a read-reference-base mismatch is within a CG context.
[0137] Accordingly, additionally or alternatively, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on a nucleotide sequence in a methylation sequencing read at the site of the read-reference-base mismatch. In some embodiments, for example, the methylation sequencing map / align system 106 determines that the mismatch between the target nucleobase of the methylation sequencing read and the corresponding reference nucleobase is at a thymine guanine (TG) dinucleotide or a cytosine adenine (CA) dinucleotide within the methylation sequencing read. In this example, the methylation sequencing map / align system 106 determines whether the methylation sequencing read has a TG dinucleotide or a CA dinucleotide at the site of the read-reference-base mismatch. If the methylation sequencing read includes either of these dinucleotides, the methylation sequencing map / align system 106 can count the read-reference-base mismatch as a CG-conversion mismatch for purposes of alignment scoring, including determining a tolerance metric, as further described with respect to FIG. 4B.
[0138] Additionally or alternatively, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on the base type (e.g., A, C, T, or G) of the target nucleobase in the read-reference-base mismatch. Additionally or alternatively, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on the strand (sense or antisense) of the methylation sequencing read. In some embodiments, the methylation sequencing map / align system 106 determines if the read-reference-base mismatch is between a thymine nucleobase of a TG dinucleotide on a methylation sequencing read from a sense strand, and a corresponding reference nucleobase; or if the mismatch is between an adenine nucleobase of a CA dinucleotide on methylation sequencing read from a reverse complement of an antisense strand, and a corresponding reference nucleobase. If the candidate alignment includes of these read-reference-base mismatches, the methylation sequencing map / align system 106 can count the mismatch as a CG-conversion mismatch for purposes of alignment scoring, as further described with respect to FIG. 4B.
[0139] Additionally or alternatively, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on the base type of the target nucleobase in the mismatch and the base type of the corresponding reference nucleobase.Attorney Docket No. IP-2950-PCT 38 Patent ApplicationFor example, the methylation sequencing map / align system 106 can determine whether the target nucleobase is a thymine (T) and the corresponding reference nucleobase is a cytosine (C). As another example, the methylation sequencing map / align system 106 can determine whether the target nucleobase is an adenine (A) and the corresponding reference nucleobase is a guanine (G).
[0140] As yet another example of CG-context determination, additionally or alternatively, the methylation sequencing map / align system 106 can determine that a read-reference-base mismatch is within a CG context based on a context base (e.g., a nucleobase immediately upstream or downstream) immediately adjacent to the target nucleobase in the methylation sequencing read. The paragraphs below refer to the relevant target nucleobase and context base on a methylation sequencing read as a target read base and a context read base, respectively. In some embodiments, if the target read base is located at the end of a methylation sequencing read such that the context read base is not present, the methylation sequencing map / align system 106 can assume that the read-reference-base mismatch is within a CG context. In alternative embodiments, if the target read base is located at the end of a methylation sequencing read such that the context read base is not present, the methylation sequencing map / align system 106 assumes the read-reference-base mismatch is not within a CG context.
[0141] As shown in FIG. 4A, for example, the methylation sequencing map / align system 106 determines a read-reference-base mismatch within a CG context based on the target nucleobase in the methylation sequencing read, the corresponding reference nucleobase, and a context nucleobase adjacent to the target nucleobase in the methylation sequencing read. In FIG. 4A, a sample nucleotide sequence 401 from a target genomic sample includes a CG site 402a with a methylated cytosine nucleobase. As depicted in FIG. 4A, the sample nucleotide sequence 401 represents the original DNA sequence of the target genomic sample (e.g., in vivo or in vitro) upon which a template can be adapted for a methylation sequencing assay or for constructing methylation sequencing reads with adapters. After a methylation sequencing assay converts methylated cytosines in extracted counterparts of the sample nucleotide sequence 401 to guanines — and both sequencing and mapping of counterpart methylation sequencing reads are performed — a CG site 402b within a methylation sequencing read includes a TG dinucleotide in the sense read sequence 403a and a CA dinucleotide in the reverse complement of the antisense strand sequence 405a.
[0142] As depicted in FIG. 4A, the methylation sequencing map / align system 106 aligns the methylation sequencing reads — including the sense read sequence 403 a and the antisense strand sequence 405a — to a reference sequence 407a. Both the reference sequences 407a and 407b include the original CG dinucleotide at CG sites 402c and 402d, respectively. Thus, the candidate alignments in FIG. 4A exhibit a read-reference-base mismatch at the CG sites 402c and 402dAttorney Docket No. IP-2950-PCT 39 Patent Applicationbecause the methylation sequencing reads have had one or more nucleobases chemically or enzymatically converted by a methylation sequencing assay.
[0143] In FIG. 4A, the methylation sequencing map / align system 106 determines a read-reference-base mismatch is within a CG context for the sense strand sequence 403b if the target read base 420a (TargetReadBase) is thymine (T), the corresponding reference nucleobase (RefBase) is cytosine (C), and the context read base 421a (ContextReadBase) downstream is guanine (G). The methylation sequencing map / align system 106 determines a read-reference-base mismatch is within a CG context for the reverse complement of the antisense strand sequence 405b if the target read base 420b (TargetReadBase) is adenine (A), the corresponding reference nucleobase (RefBase) is guanine (G), and the context read base 421b (ContextReadBase) upstream is cytosine (C). If either of these are true, the methylation sequencing map / align system 106 can score the mismatch as a CG-conversion mismatch as further described with respect to FIG. 4B.
[0144] FIG. 4B illustrates the methylation sequencing map / align system generating alignment scores based on read-reference comparisons of a methylation sequencing read and one or more reference sequences in accordance with one or more embodiments of the present disclosure. As shown in FIG. 4B, for example, the methylation sequencing map / align system 106 generates one or more alignment score(s) 414a, 414b, through 414n for each of the respective candidate alignments 406a, 406b, through 406n based on comparing nucleobases within a subset of the methylation sequencing reads 404a, 404b, through 404n with nucleobases of reference sequences 410a, 410b, through 41 On at the genomic regions of the respective candidate alignments 406a, 406b, through 406n.
[0145] In particular, as illustrated in FIG. 4B, the methylation sequencing map / align system 106 identifies read-reference comparisons 412a, 412b through 412n based on the comparison of nucleobases from candidate alignments 406a, 406b through 406n. As depicted in FIG. 4B, the candidate alignments 406a, 406b through 406n include alignments between the methylation sequencing reads 404a, 404b through 404n and reference sequences 410a, 410b through 41 On, respectively. Based on the read-reference comparisons 412a-412n from the candidate alignments 406a-406n, the methylation sequencing map / align system 106 determines alignment scores 414a-414n for each respective candidate alignment 406a-406n. For example, the read-reference comparisons 412a-412n can include read-reference match(es), read-reference mismatch(es), CG-conversion mismatch(es), and soft clip(s).
[0146] For example, for the candidate alignment 406a, the methylation sequencing map / align system 106 identifies read-reference comparisons 412a. From the read-reference comparisons 412a, the methylation sequencing map / align system 106 determines components or metrics that increase or decrease alignment score(s) 414a. In particular, in some embodiments, the methylationAttorney Docket No. IP-2950-PCT 40 Patent Applicationsequencing map / align system 106 increases the alignment score(s) 414a for each match between a nucleobase of the methylation sequencing read 404a and a nucleobase of the reference sequence 410a, as represented by the read-reference comparisons 412a. Further, the methylation sequencing map / align system 106 decreases the alignment score(s) 414a for each mismatch (e.g., outside of a CG context) between a nucleobase of the methylation sequencing read 404a and a nucleobase of the reference sequence 410a, as represented by the read-reference comparisons 412a, by applying a mismatch-penalty metric.
[0147] Further, as shown in FIG. 4B, the methylation sequencing map / align system 106 does not increase or decrease the alignment score(s) 414a for each CG-conversion mismatch between a nucleobase of the methylation sequencing read 404a and a nucleobase of the reference sequence 410a, as represented by the read-reference comparisons 412a, by applying a tolerance metric. Similarly, the methylation sequencing map / align system 106 does not increase or decrease the alignment score(s) 414a between a soft-clipped (e.g., masked) nucleobase of the methylation sequencing read 404a and a nucleobase of the reference sequence 410a, as represented by the readreference comparisons 412a.
[0148] Accordingly, as shown in FIG. 4B, the methylation sequencing map / align system 106 generates the alignment score(s) 414a based on the read-reference comparisons 412a in the respective genomic region of the candidate alignment 406a. Moreover, the methylation sequencing map / align system 106 performs similar steps to determine one or more alignment score(s) 414b-414n for the remaining candidate alignments 406b-406n.
[0149] As indicated above and in FIGS. 3C and 4B, in some embodiments, the methylation sequencing map / align system 106 generates an alignment score based on a tolerance metric for a read-reference-base mismatch within a CG context. For example, the tolerance metric can comprise an alignment score component that is less of a penalty than a mismatching-penalty metric applied to a read-reference-base mismatch outside of a CG context. In some embodiments, the tolerance metric increases a likelihood of selecting the candidate alignment relative to another candidate alignment comprising a read-reference-base mismatch outside of a CG context.
[0150] Accordingly, the methylation sequencing map / align system 106 can apply different values when determining alignment scores for candidate alignments with read-reference-base mismatches within a CG context and candidate alignments with read-reference-base mismatches outside of a CG context. For example, the methylation sequencing map / align system 106 can further determine an additional read-reference-base mismatch not within the CG context between an additional nucleobase of the methylation sequencing read and an additional corresponding reference nucleobase. The methylation sequencing map / align system 106 can accordingly generate an alignment score based on a mismatching-penalty metric that decreases the value of or penalizesAttorney Docket No. IP-2950-PCT 41 Patent Applicationa candidate alignment with a read-reference-base mismatch outside a CG context. Consequently, in some embodiments, the mismatching-penalty metric decreases a likelihood of selecting the candidate alignment relative to another candidate alignment because of the mismatch not within the CG context.
[0151] As indicated above, in some embodiments, the methylation sequencing map / align system 106 can mask (e.g., soft clip) one or more nucleobases within a methylation sequencing read of the methylation sequencing reads. For example, unlike existing methylation sequencing systems, the methylation sequencing map / align system 106 can mask one or more nucleobases on or near an end of a methylation sequencing read (e.g., within a threshold number of nucleobases from an end of a methylation sequencing read). In some embodiments, accordingly, the methylation sequencing map / align system 106 generates a candidate alignment of the methylation sequencing read comprising the one or more masked nucleobases and a corresponding reference sequence. In some instances, the methylation sequencing map / align system 106 masks nucleobases at or near an end of a methylation sequencing read that may be base-called less accurately (e.g., have a lower base calling quality(Q)-score), or that may be left over from an adapter sequence and not representative of the target genomic sample.
[0152] As mentioned, FIG. 5 illustrates improvements in accuracy by the methylation sequencing map / align system 106 in determining a methylation-level value. Plots 500 in FIG. 5 portrays the observed average percentage of nucleobases that are methylated (y-axis) per read cycle number (x-axis), referred to as M-bias. In the example of FIG. 5, the target genomic sample is unmethylated cell free DNA, and the ground truth methylation-level value is 0% methylation.
[0153] The plots 500 in FIG. 5 show the performance of the methylation sequencing map / align system 106 as compared to the performance of an existing methylation sequencing system. The methylation sequencing map / align system 106 can mask one or more nucleobases (e.g., soft-clipping), for example, at the end of methylation sequencing reads, and generate alignment scores for candidate alignments of methylation sequencing reads with masked nucleobases. As shown in FIG. 5, the existing methylation sequencing system does not support masking (e.g., soft-clipping) and instead uses another process to trim adapter sequence nucleobases from the methylation sequencing reads.
[0154] As further shown in the plots 500 of FIG. 5, the existing methylation sequencing system determined inaccurate methylation-level values that are significantly greater than the ground truth methylation-level value of 0%, particularly in later sequencing cycles. The existing methylation sequencing system generates inflated and inaccurate methylation-level values likely because existing systems cannot trim adapter nucleobases correctly from a methylation sequencing read and / or cannot mask nucleobases that are sequencing errors or adapter sequence nucleobases.Attorney Docket No. IP-2950-PCT 42 Patent ApplicationExisting methylation sequencing systems often interpret untrimmed / unmasked nucleobases as C to T methylation conversions and count them toward the methylation-level value, which raises the methylation-level value above the ground truth zero. In contrast, as shown in FIG. 5, the methylation sequencing map / align system 106 determines the methylation-level value to be closer to the expected 0%, including in the later sequencing cycles. The methylation sequencing map / align system 106 generates more accurate methylation-level values likely because it can mask nucleobases which would otherwise contribute towards an inaccurate methylation-level value. Thus, the methylation sequencing map / align system 106 improves accuracy in determining a methylation-level value as compared to the existing methylation sequencing system.
[0155] As mentioned, FIG. 6A illustrates improvements in accuracy by the methylation sequencing map / align system 106 in calling single nucleotide polymorphisms (SNPs). In FIG. 6A, bar graph 600 depicts the Fl score (based on precision and recall) in calling SNPs using three existing sequencing systems, and the methylation sequencing map / align system 106. To generate the data in the bar graph 600, methylation sequencing reads were taken from reference sample HG001, which underwent a methylation sequencing assay that converted methylated cytosines. In the context of bar graph 600, the bar for “Germline” refers to Fl scores for an existing nonmethylation sequencing system, which uses an unconverted reference sequence. The bar for “Methylation” refers to Fl scores for an existing methylation sequencing system that uses two copies of a 3 -base reference sequence for mapping and alignment. The bar for “Germline Graph” refers to Fl scores for an existing non-methylation sequencing system like the “Germline” system, but that performs mapping with a graph reference genome with multiple reference sequences.
[0156] As depicted in the bar graph 600, the methylation sequencing map / align system 106 improves accuracy in calling SNPs compared to the three existing sequencing systems. Indeed, the methylation sequencing map / align system 106 exhibits Fl scores of nearly 1 full value greater than existing sequencing systems. In the context of the bar graph 600, where even small Fl -score increments matter, the Fl-score improvements demonstrate improved and measurable performance.
[0157] FIG. 6B illustrates further improvements in accuracy by the methylation sequencing map / align system 106 in calling single nucleotide polymorphisms (SNPs) and insertions / deletions (INDELs) from methylation sequencing reads from reference sample HG001. In FIG. 6B, table 620 illustrates improvements in accuracy in calling SNPs and INDELs as compared to existing sequencing systems by comparing the reduction in SNP false positives (FP) and false negatives (FN) and INDEL false positives (FP) and false negatives (FN). In table 620 of FIG. 6B, the column for “Methylation” refers to FN or FP counts for the same existing methylation sequencing system as the “Methylation” system of FIG. 6A, and the column for “Germline Graph” refers to FN or FPAttorney Docket No. IP-2950-PCT 43 Patent Applicationcounts for the same existing non-methylation sequencing system as the “Germline Graph” system of FIG 6A.
[0158] As shown in the table 620, the methylation sequencing map / align system 106 significantly reduces the number of both false positives and false negatives for SNPs and INDELs. Although the “Germline Graph” system reduces false positives and false negatives in the data presented in FIG. 6B, other data (not presented here) showed that the “Germline Graph” mapping and alignment were inaccurate in highly methylated regions. In contrast, the methylation sequencing map / align system 106 did not have such inaccuracy in mapping and alignment in highly methylated regions.
[0159] FIG. 6C illustrates improvements in accuracy by the methylation sequencing map / align system 106 in calling single nucleotide variants (SNVs) and insertions / deletions (INDELs). In FIG.6C, the graphs 640 show precision and sensitivity in calling SNVs and INDELs for three existing sequencing systems: EM-seq, the methylation sequencing map / align system 106, and whole genome sequencing (WGS). EM-seq converts unmethylated cytosines and variant calling was performed with public tools bis-SNP and EpiDiverse. The methylation sequencing map / align system 106 used methylation sequencing reads from a methylation sequencing assay that converts methylated cytosines. WGS used normal, unconverted sequencing reads prepared from PCR and PCR-Free methods and existing sequencing systems. The samples used were a mix of reference samples HG001 and HG002 and all data is at about 3 OX coverage.
[0160] As shown in the graphs 640, the methylation sequencing map / align system 106 achieved over 99.5% accuracy for SNV detection and 97.5% accuracy for INDEL calling. The accuracy of the methylation sequencing map / align system 106 is significantly greater than that of EM-seq. The accuracy of the methylation sequencing map / align system 106 approaches the performance offered by WGS, meaning the methylation sequencing map / align system 106 could be used as a single system for both methylation analyses and variant / genotype calling.
[0161] Turning now to FIG. 7, this figure illustrates an example flowchart of a series of acts for mapping and aligning methylation sequencing reads in accordance with one or more embodiments of the present disclosure. While FIG. 7 illustrates acts according to particular embodiments, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 7. The acts of FIG. 7 can be performed as part of a method. Alternatively, a non-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device to perform the acts depicted in FIG. 7. In still further embodiments, a system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 7.Attorney Docket No. IP-2950-PCT 44 Patent Application
[0162] As shown in FIG. 7, the series of acts 700 includes an act 702 of accessing read subsequences from methylation sequencing reads, an act 704 of converting one or more nucleobases from one base type to another base type, an act 706 of mapping the converted read subsequences to converted reference subsequences, and an act 708 of generating candidate alignments of the methylation sequencing reads to reference sequences.
[0163] For example, the series of acts 700 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A computer-implemented method comprising:accessing, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;converting, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences; mapping the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andgenerating candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.CLAUSE 2. The computer-implemented method of clause 1, wherein converting, at the candidate methylation site within the read subsequences, comprises selectively converting one or more cytosine guanine (CG) dinucleotides within a read subsequence from a sense strand or within a read subsequence from a reverse complement of an antisense strand.CLAUSE 3. The computer-implemented method of clause 1 or 2, wherein: converting the one or more nucleobases comprises converting a target nucleobase at the candidate methylation site within a single-stranded read subsequence from a single-stranded methylation sequencing read to generate a converted single-stranded read subsequence;mapping the converted read subsequences comprises mapping the converted singlestranded read subsequence to a converted reference subsequence; andgenerating the candidate alignments comprises generating a candidate alignment of the single-stranded methylation sequencing read to a reference sequence of the one or more reference sequences at a genomic coordinate corresponding to the mapped and converted single-stranded read subsequence.CLAUSE 4. The computer-implemented method of clause 3, wherein converting the target nucleobase comprises:Attorney Docket No. IP-2950-PCT 45 Patent Applicationreplacing a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand; orreplacing a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand.CLAUSE 5. The computer-implemented method of any of clauses 1-4, further comprising:determining, for a candidate alignment of a methylation sequencing read, a mismatch within a cytosine guanine (CG) context between a target nucleobase of the methylation sequencing read and a corresponding reference nucleobase; andgenerating an alignment score for the candidate alignment of the methylation sequencing read and a corresponding reference sequence based on a tolerance metric that increases a likelihood of selecting the candidate alignment relative to another candidate alignment mismatch outside of the CG context.CLAUSE 6. The computer-implemented method of clause 5, wherein determining the mismatch within the CG context comprises:determining the mismatch between target nucleobase of the methylation sequencing read and the corresponding reference nucleobase at a cytosine guanine (CG) dinucleotide within the one or more reference sequences; ordetermining the mismatch between the target nucleobase of the methylation sequencing read and the corresponding reference nucleobase at a thymine guanine (TG) dinucleotide or a cytosine adenine (CA) dinucleotide within the methylation sequencing read.CLAUSE 7. The computer-implemented method of clause 6, wherein determining the mismatch between the target nucleobase of the methylation sequencing read and the corresponding reference nucleobase comprises:determining the mismatch between a thymine nucleobase of a TG dinucleotide on a methylation sequencing read from a sense strand and the corresponding reference nucleobase; or determining the mismatch between an adenine nucleobase of a CA dinucleotide on methylation sequencing read from a reverse complement of an antisense strand and the corresponding reference nucleobase.CLAUSE 8. The computer-implemented method of any of clauses 5-7, further comprising:determining an additional mismatch not within the CG context between an additional nucleobase of the methylation sequencing read and an additional corresponding reference nucleobase; andAttorney Docket No. IP-2950-PCT 46 Patent Applicationgenerating the alignment score for the candidate alignment of the methylation sequencing read and a corresponding reference sequence based further on a mismatching-penalty metric that decreases a likelihood of selecting the candidate alignment relative to another candidate alignment because of the mismatch not within the CG context.CLAUSE 9. The computer-implemented method of any of clauses 1-8, further comprising:masking one or more nucleobases within a methylation sequencing read of the methylation sequencing reads; andgenerating a candidate alignment of the methylation sequencing read comprising the one or more masked nucleobases and a corresponding reference sequence.CLAUSE 10. The computer-implemented method of any of clauses 1-9, wherein the methylation sequencing assay chemically or enzymatically converts methylated cytosine nucleobases to thymine nucleobases.CLAUSE 11. The computer-implemented method of any of clauses 1-10, wherein converting the one or more nucleobases at the candidate methylation site within the read subsequences comprises:identifying CG dinucleotides within the read subsequences; andreplacing a nucleobase of the CG dinucleotides with one base type instead of another base type.CLAUSE 12. The computer-implemented method of clause 11, wherein replacing the nucleobase of CG dinucleotides within the read subsequences comprises:replacing a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand; orreplacing a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand.CLAUSE 13. The computer-implemented method of any of clauses 1-12, wherein the converted reference subsequences comprise a replaced nucleobase of a CG dinucleotide as compared to a reference nucleobase of the one or more reference sequences.CLAUSE 14. The computer-implemented method of any of clauses 1-13, wherein generating the candidate alignments comprises generating one or more candidate alignments of an unconverted methylation sequencing read to an unconverted reference sequence.CLAUSE 15. The computer-implemented method of any of clauses 1-13, further comprising:generating reference subsequences from the one or more reference sequences; andAttorney Docket No. IP-2950-PCT 47 Patent Applicationselectively converting, at a candidate methylation site within the reference subsequences, one or more nucleobases from one base type to another base type to generate converted reference subsequences.CLAUSE 16. The computer-implemented method of clause 15, wherein selectively converting, at the candidate methylation site within the reference subsequences, one or more nucleobases from one base type to another base type comprises:identifying CG dinucleotides within the reference subsequences;generating a first set of converted reference subsequences by replacing cytosine nucleobases of CG dinucleotides with thymine nucleobases; andgenerating a second set of converted reference subsequences by replacing guanine nucleobases of CG dinucleotides with adenine nucleobases.CLAUSE 17. The computer-implemented method of any of clauses 1-16, wherein the converted reference subsequences are stored in a hash table that further comprises unconverted reference subsequences that do not include a given candidate methylation site.CLAUSE 18. The computer-implemented method of any of clauses 1-17, further comprising mapping unconverted read subsequences that do not include a given candidate methylation site to a set of unconverted reference subsequences that do not include a given candidate methylation site.CLAUSE 19. The computer-implemented method of clause 1 , wherein converting the one or more nucleobases does not replace nucleobases outside of CG dinucleotides.CLAUSE 20. The computer-implemented method of any of clauses 1-19, further comprising determining a genotype call for the target genomic sample at the candidate methylation site based on a selected alignment of the candidate alignments of the methylation sequencing reads to the reference sequences.CLAUSE 21. A system comprising:at least one processor; anda non-transitory computer readable medium storing instructions that, when executed by the at least one processor, cause the system to:access, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;convert, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences;map the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andAttorney Docket No. IP-2950-PCT 48 Patent Applicationgenerate candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.CLAUSE 22. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:access, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;convert, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences; map the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andgenerate candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.
[0164] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (e.g., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.
[0165] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0166] SBS can utilize nucleotide monomers that have a terminator moiety or those that lack any terminator moieties. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using y-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is generally variable and dependent upon the template sequence and the mode of nucleotide delivery. For SBS techniques that utilize nucleotide monomers having a terminator moiety, the terminator can be effectively irreversible under the sequencing conditionsAttorney Docket No. IP-2950-PCT 49 Patent Applicationused as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator can be reversible as is the case for sequencing methods developed by Solexa (now Illumina, Inc.).
[0167] SBS techniques can utilize nucleotide monomers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be detected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleotide monomer such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; or the like. In embodiments where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniques being used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequencing methods developed by Solexa (now Illumina, Inc.).
[0168] Preferred embodiments include pyrosequencing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res.11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) “A sequencing method based on realtime pyrophosphate.” Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by reference in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-produced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C or G). Images obtained after addition of each nucleotide type will differ with regard to which features in the array are detected. These differences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed and analyzed using the methods set forth herein. For example, images obtained after treatment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different detection channels for reversible terminatorbased sequencing methods.Attorney Docket No. IP-2950-PCT 50 Patent Application
[0169] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, in WO 04 / 018497 and U.S. Pat. No.7,057,026, the disclosures of which are incorporated herein by reference. This approach is being commercialized by Solexa (now Illumina Inc.), and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators in which both the termination can be reversed, and the fluorescent label cleaved facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from these modified nucleotides.
[0170] Preferably in reversible terminator-based sequencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. However, the detection labels can be removable, for example, by cleavage or degradation. Images can be captured following incorporation of labels into arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image will show nucleic acid features that have incorporated nucleotides of a particular type. Different features are present or absent in the different images due the different sequence content of each feature. However, the relative position of the features will remain unchanged in the images. Images obtained from such reversible terminator- SBS methods can be stored, processed and analyzed as set forth herein. Following the image capture step, labels can be removed, and reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.
[0171] In particular embodiments some or all of the nucleotide monomers can include reversible terminators. In such embodiments, reversible terminators / cleavable fluors can include fluor linked to the ribose moiety via a 3' ester linkage (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3' allyl group to block extension but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30 secondAttorney Docket No. IP-2950-PCT 51 Patent Applicationexposure to long wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of natural termination that ensues after placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.
[0172] Additional exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No.2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No.2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No.2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No.2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.
[0173] Some embodiments can utilize detection of four different nucleotides using fewer than four different labels. For example, SBS can be performed utilizing methods and systems described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair. As a second example, three of four different nucleotide types can be detected under particular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleotide type into the nucleic acid can be determined based on absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary configurations are not considered mutually exclusive and can be used in various combinations. An exemplary embodiment that combines all three examples, is a fluorescent-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g. dATP having a label that is detected in the first channel when excited by a first excitationAttorney Docket No. IP-2950-PCT 52 Patent Applicationwavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, or minimally, detected in either channel (e.g. dGTP having no label).
[0174] Further, as described in the incorporated materials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0175] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained following treatment of an array of nucleic acid features with the labeled sequencing reagents. Each image will show nucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based sequencing methods can be stored, processed and analyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No.6,769,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.
[0176] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological membrane protein, such as a-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G. V. & Meller, "A. ProgressAttorney Docket No. IP-2950-PCT 53 Patent Applicationtoward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Data obtained from nanopore sequencing can be stored, processed and analyzed as set forth herein. In particular, the data can be treated as an image in accordance with the exemplary treatment of optical images and other images that is set forth herein.
[0177] Some embodiments can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate-labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No.7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero-mode waveguides as described, for example, in U.S. Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). The illumination can be restricted to a zeptoliter-scale volume around a surface-tethered polymerase such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M. J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science 299, 682-686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 826-828 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed and analyzed as set forth herein.
[0178] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.Attorney Docket No. IP-2950-PCT 54 Patent Application
[0179] The above SBS methods can be advantageously carried out in multiplex formats such that multiple different target nucleic acids are manipulated simultaneously. In particular embodiments, different target nucleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows convenient delivery of sequencing reagents, removal of unreacted reagents and detection of incorporation events in a multiplex manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or feature. Multiple copies can be produced by amplification methods such as, bridge amplification or emulsion PCR as described in further detail below.
[0180] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or higher.
[0181] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprising components such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeqTM platform (Illumina, Inc., San Diego,Attorney Docket No. IP-2950-PCT 55 Patent ApplicationCA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference. The sequencing system described above sequences nucleic acid polymers present in samples received by a sequencing device, as described further above.
[0182] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0183] The components of the methylation sequencing map / align system 106 can include software, hardware, or both. For example, the components of the methylation sequencing map / align system 106 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the client device 114, the local device 108, or the server device(s) 110). When executed by the one or more processors, the computer-executable instructions of the methylation sequencing map / align system 106 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the methylation sequencing map / align system 106 can comprise hardware, suchAttorney Docket No. IP-2950-PCT 56 Patent Applicationas special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the methylation sequencing map / align system 106 can include a combination of computer-executable instructions and hardware.
[0184] Furthermore, the components of the methylation sequencing map / align system 106 performing the functions described herein with respect to the methylation sequencing map / align system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the methylation sequencing map / align system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the methylation sequencing map / align system 106 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, Illumina 5-Base DNA Prep, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” “5-Based DNA Prep,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0185] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0186] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0187] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phase-Attorney Docket No. IP-2950-PCT 57 Patent Applicationchange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0188] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0189] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0190] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0191] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personalAttorney Docket No. IP-2950-PCT 58 Patent Applicationcomputers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0192] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0193] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.
[0194] FIG. 8 illustrates a block diagram of a computing device 800 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 800 may implement the methylation sequencing map / align system 106 and the sequencing device system 104. As shown by FIG. 8, the computing device 800 can comprise a processor 802, a memory 804, a storage device 806, an I / O interface 808, and a communication interface 810, which may be communicatively coupled by way of a communication infrastructure 812. In certain embodiments, the computing device 800 can include fewer or more components than those shown in FIG. 8. The following paragraphs describe components of the computing device 800 shown in FIG. 8 in additional detail.
[0195] In one or more embodiments, the processor 802 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 802 mayAttorney Docket No. IP-2950-PCT 59 Patent Applicationretrieve (or fetch) the instructions from an internal register, an internal cache, the memory 804, or the storage device 806 and decode and execute them. The memory 804 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 806 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0196] The I / O interface 808 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 800. The I / O interface 808 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 808 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 808 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0197] The communication interface 810 can include hardware, software, or both. In any event, the communication interface 810 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 800 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 810 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.
[0198] Additionally, the communication interface 810 may facilitate communications with various types of wired or wireless networks. The communication interface 810 may also facilitate communications using various communication protocols. The communication infrastructure 812 may also include hardware, software, or both that couples components of the computing device 800 to each other. For example, the communication interface 810 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.
[0199] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanyingAttorney Docket No. IP-2950-PCT 60 Patent Applicationdrawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0200] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Attorney Docket No. IP-2950-PCT 61 Patent Application
Claims
CLAIMSWe claim:
1. A computer-implemented method comprising:accessing, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;converting, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences; mapping the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andgenerating candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.
2. The computer-implemented method of claim 1, wherein converting, at the candidate methylation site within the read subsequences, comprises selectively converting one or more cytosine guanine (CG) dinucleotides within a read subsequence from a sense strand or within a read subsequence from a reverse complement of an antisense strand.
3. The computer-implemented method of claim 1, wherein:converting the one or more nucleobases comprises converting a target nucleobase at the candidate methylation site within a single-stranded read subsequence from a single-stranded methylation sequencing read to generate a converted single-stranded read subsequence;mapping the converted read subsequences comprises mapping the converted singlestranded read subsequence to a converted reference subsequence; andgenerating the candidate alignments comprises generating a candidate alignment of the single-stranded methylation sequencing read to a reference sequence of the one or more reference sequences at a genomic coordinate corresponding to the mapped and converted single-stranded read subsequence.
4. The computer-implemented method of claim 3, wherein converting the target nucleobase comprises:replacing a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand; orreplacing a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand.Attorney Docket No. IP-2950-PCT 62 Patent Application5. The computer-implemented method of claim 1, further comprising: determining, for a candidate alignment of a methylation sequencing read, a mismatch within a cytosine guanine (CG) context between a target nucleobase of the methylation sequencing read and a corresponding reference nucleobase; andgenerating an alignment score for the candidate alignment of the methylation sequencing read and a corresponding reference sequence based on a tolerance metric that increases a likelihood of selecting the candidate alignment relative to another candidate alignment comprising a mismatch outside of the CG context.
6. The computer-implemented method of claim 5, wherein determining the mismatch within the CG context comprises:determining the mismatch between a target nucleobase of the methylation sequencing read and the corresponding reference nucleobase at a cytosine guanine (CG) dinucleotide within the one or more reference sequences; ordetermining the mismatch between the target nucleobase of the methylation sequencing read and the corresponding reference nucleobase at a thymine guanine (TG) dinucleotide or a cytosine adenine (CA) dinucleotide within the methylation sequencing read.
7. The computer-implemented method of claim 6, wherein determining the mismatch between the target nucleobase of the methylation sequencing read and the corresponding reference nucleobase comprises:determining the mismatch between a thymine nucleobase of a TG dinucleotide on a methylation sequencing read from a sense strand and the corresponding reference nucleobase; or determining the mismatch between an adenine nucleobase of a CA dinucleotide on methylation sequencing read from a reverse complement of an antisense strand and the corresponding reference nucleobase.
8. The computer-implemented method of claim 5, further comprising: determining an additional mismatch not within the CG context between an additional nucleobase of the methylation sequencing read and an additional corresponding reference nucleobase; andgenerating the alignment score for the candidate alignment of the methylation sequencing read and a corresponding reference sequence based further on a mismatching-penalty metric thatAttorney Docket No. IP-2950-PCT 63 Patent Applicationdecreases a likelihood of selecting the candidate alignment relative to another candidate alignment because of the mismatch not within the CG context.
9. The computer-implemented method of claim 1, further comprising: masking one or more nucleobases within a methylation sequencing read of the methylation sequencing reads; andgenerating a candidate alignment of the methylation sequencing read comprising the one or more masked nucleobases and a corresponding reference sequence.
10. The computer-implemented method of claim 1, wherein the methylation sequencing assay chemically or enzymatically converts methylated cytosine nucleobases to thymine nucleobases.
11. The computer-implemented method of claim 1 , wherein converting the one or more nucleobases at the candidate methylation site within the read subsequences comprises:identifying CG dinucleotides within the read subsequences; andreplacing a nucleobase of the CG dinucleotides with one base type instead of another base type.
12. The computer-implemented method of claim 11, wherein replacing the nucleobase of CG dinucleotides within the read subsequences comprises:replacing a cytosine nucleobase of a CG dinucleotide with a thymine nucleobase within a read subsequence from a sense strand; orreplacing a guanine nucleobase of a CG dinucleotide with an adenine nucleobase within a read subsequence from a reverse complement of an antisense strand.
13. The computer-implemented method of claim 1, wherein the converted reference subsequences comprise a replaced nucleobase of a CG dinucleotide as compared to a reference nucleobase of the one or more reference sequences.
14. The computer-implemented method of claim 1, wherein generating the candidate alignments comprises generating one or more candidate alignments of an unconverted methylation sequencing read to an unconverted reference sequence.
15. The computer-implemented method of claim 1, further comprising:Attorney Docket No. IP-2950-PCT 64 Patent Applicationgenerating reference subsequences from the one or more reference sequences; and selectively converting, at a candidate methylation site within the reference subsequences, one or more nucleobases from one base type to another base type to generate converted reference subsequences.
16. The computer-implemented method of claim 15, wherein selectively converting, at the candidate methylation site within the reference subsequences, one or more nucleobases from one base type to another base type comprises:identifying CG dinucleotides within the reference subsequences;generating a first set of converted reference subsequences by replacing cytosine nucleobases of CG dinucleotides with thymine nucleobases; andgenerating a second set of converted reference subsequences by replacing guanine nucleobases of CG dinucleotides with adenine nucleobases.
17. The computer-implemented method of claim 1, wherein the converted reference subsequences are stored in a hash table that further comprises unconverted reference subsequences that do not include a given candidate methylation site.
18. The computer-implemented method of claim 1, further comprising mapping unconverted read subsequences that do not include a given candidate methylation site to a set of unconverted reference subsequences that do not include a given candidate methylation site.
19. The computer-implemented method of claim 1 , wherein converting the one or more nucleobases does not replace nucleobases outside of CG dinucleotides.
20. The computer-implemented method of claim 1, further comprising determining a genotype call for the target genomic sample at the candidate methylation site based on a selected alignment of the candidate alignments of the methylation sequencing reads to the reference sequences.
21. A system comprising:at least one processor; anda non-transitory computer readable medium storing instructions that, when executed by the at least one processor, cause the system to:Attorney Docket No. IP-2950-PCT 65 Patent Applicationaccess, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;convert, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences;map the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andgenerate candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.
22. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:access, for a target genomic sample, read subsequences from methylation sequencing reads comprising nucleobases modified by a methylation sequencing assay;convert, at a candidate methylation site within the read subsequences, one or more nucleobases from one base type to another base type to generate converted read subsequences; map the converted read subsequences to converted reference subsequences that are converted from one or more reference sequences; andgenerate candidate alignments of the methylation sequencing reads to reference sequences of the one or more reference sequences at genomic coordinates corresponding to the mapped converted read subsequences.Attorney Docket No. IP-2950-PCT 66 Patent Application