Methods and systems for phasing sequence reads

By determining linkage and mapping sequence reads to a reference genome using geographic location and haplotype data, the methods enhance the accuracy of sequence assembly and variant detection by correctly phasing reads into haplotypes, addressing the limitations of traditional sequencing methods.

WO2026096262A1PCT designated stage Publication Date: 2026-05-07ILLUMINA INC
View PDF 20 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-10-22
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Traditional nucleic acid sequencing methods lose information about the original position and connectivity of sequence fragments, making it difficult to phase reads into haplotypes, which is crucial for improving sequence assembly and variant detection, especially in diploid organisms like humans.

Method used

Methods and systems that determine linkage information between sequence reads based on their geographic location on a flow cell, map them to a reference genome using candidate haplotypes, analyze the likelihoods of haplotype correspondence, and phase the reads accordingly, utilizing a forward and backward approach to enhance accuracy.

Benefits of technology

Improves the accuracy of sequence assembly and variant detection by correctly assigning reads to parental haplotypes, reducing false positives and negatives, and enhancing the utility of sequencing systems for diploid genomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052049_07052026_PF_FP_ABST
    Figure US2025052049_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for phasing sequence reads. In some embodiments, the methods includes steps of generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. Further disclosed herein are methods and systems for detecting a phased sequence variant, for determining a haplotype nucleotide sequence, and for detecting a structural variant.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR PHASING SEQUENCE READSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U. S. Provisional Application No.63 / 713,986, filed October 30, 2024, the content of which is incorporated by reference in its entirety.BACKGROUNDField

[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, this disclosure relates to systems and methods for phasing sequence reads.Description

[0003] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods, including Sequencing by Synthesis (SBS), use a shotgun approach to sequence large genomic DNA fragments, sometimes called template genomic sequences. During SBS sequencing the template genomic sequences are first fragmented into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been sequenced, knowledge of their original position in the genome, and their connectivity and proximity to each other in the original template genomic sequence is lost.

[0004] In diploid organisms such as humans, two haplotypes exist for each chromosome: one for the paternal homolog and one for the maternal homolog. Phasing is the process of assigning sequence reads into one of two or more haplotypes. In humans, this would include assigning each sequence read to either the maternal or the paternal homolog. Phasing information can be useful for improving downstream sequence assembly and for detecting variants. For example, it is clinically relevant for some variants whether two or more variants are inherited in the same haplotype (cis) or different haplotypes (trans).SUMMARY

[0005] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of theclaims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0006] Disclosed herein are methods for phasing sequence reads from a genomic DNA sample. In some embodiments, the methods include generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0007] In some embodiments, the method comprises dividing an alignment of the sequence reads to the reference genome into bins, and phasing sequence reads by bin. In some embodiments, the method comprises phasing the sequence reads in a sequence of bins, wherein the sequence of bins proceeds in a first direction along the reference genome. In some embodiments, the method further includes determining a most-probable first parental haplotype and a most-probable second parental haplotype from the set of candidate haplotypes for each bin and a most-probable phasing for each sequence read in the bm. In some embodiments, the method comprises analyzing sequence reads in each bin m a sequence of bins to determine: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype. In some embodiments, determining the distribution of candidate haplotype likelihoods for a first parental haplotype, the distribution of candidate haplotype likelihoods for a second parental haplotype, and the probability for each sequence read that the sequenceread originated from the first parental haplotype, is based on a determination from a previous bin in the sequence of bins. In some embodiments, the probability for each sequence read that the sequence read originated from the first parental haplotype is determined based on linkage information with sequence reads in a previous bin in the sequence of bins. In some embodiments, determining the distribution of candidate haplotype likelihoods for a first parental haplotype, the distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability' for each sequence read that the sequence read originated from the first parental haplotype comprises analyzing sequence similarity between sequence reads in each bin and each haplotype of the set of candidate haplotypes.

[0008] In some embodiments, the method further comprises phasing sequence reads sequentially by bin in a second direction along the reference genome. In some embodiments, phasing sequence reads sequentially by bin in a second direction comprises analyzing sequence reads in each bin in a sequence of bins to determine: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype, based on a determination from a previous bin in a sequence of bins, and based on a determination from a bin from phasing sequence reads sequentially by bin in the first direction.

[0009] In some embodiments, the method comprises separating an alignment of the sequence reads to the reference genome into overlapping sections, separating each overlapping section into bins, and phasing sequence reads by bin. In some embodiments, each overlapping section is about 1 Mbp in size, wherein each bin is about 4 kbp in size, and wherein each overlapping section overlaps with another overlapping section by about 100 kbp. In some embodiments, the method comprises phasing sequence reads within each overlapping section independently. In some embodiments, the method comprises analyzing sequence read phasing within an overlap of two consecutive overlapping sections and updating phasing of one or more sequence reads based on the analysis of sequence read phasing within the overlap.

[0010] In some embodiments, the linkage information is determined based on location information of clusters of the fragments on the flow cell. In some embodiments, the linkage information comprises information linking a first sequence read from a first cluster toa second sequence read from a second cluster based on a distance between the first and second clusters m the flow cell.

[0011] In another aspect, disclosed herein are methods for detecting a phased sequence variant. In some embodiments the methods include phasing sequence reads according to the methods described herein, and detecting one or more phased sequence variants.

[0012] According to a further aspect, disclosed herein are computer-implemented methods for phasing sequence reads from a genomic DNA sample. In some embodiments, the methods include receiving sequence reads generated from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0013] According to a further aspect, disclosed herein are methods of determining a haplotype nucleotide sequence. In some embodiments, the methods include phasing sequence reads according to the methods described herein; and assembling phased sequence reads into a first parental haplotype, thereby determining a haplotype nucleotide sequence.

[0014] Further disclosed herein are methods of detecting a structural variant in a genomic DNA sample. In some embodiments, the method includes phasing sequence reads according to the methods described herein, generating one or more sequence assemblies based on the phased sequence reads; and detecting a structural variant based on the one or more sequence assemblies.

[0015] In some embodiments, the methods include storing a set of sequence reads and associated phasing information in an electronic file. In some embodiments, the electronic file comprises a BAM file, and wherein the BAM file comprises a tag that indicates whether a sequence read is phased to a first parental haplotype or a second parental haplotype, or a probability of phasing to a first parental haplotype or a second parental haplotype. In some embodiments, the methods further include comprising storing imputed genotypes in an electronic file. In some embodiments, the electronic file comprises a VCF file.

[0016] Further disclosed herein are systems. In some embodiments, the systems are for phasing sequence reads from a genomic DNA sample. In some embodiments, the system comprises one or more processors having instructions that when executed perform a method comprising: receiving sequence reads generated from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0017] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium comprises a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads generated from fragments of the genomic DN A sample bound to a flow cell; determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; map the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyze the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phase the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, m which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the followingdescriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.

[0019] FIG. 1 is a flow diagram that schematically illustrates an exemplary method for phasing sequence reads from a genomic DNA sample.

[0020] FIG. 2 is a flow diagram that schematically illustrates a process of analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes, that may take place within the method of FIG. 1.

[0021] FIG. 3 A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.

[0022] FIG. 3B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 3A.

[0023] FIG. 4 is a schematic representation of an exemplary method for phasing sequence reads.

[0024] FIG. 5 is a schematic representation of an exemplary method based on a forward- backward HMM algorithm.

[0025] FIG. 6 is a schematic representation of exemplary analysis steps within a forward pass of an exemplary bin h.

[0026] FIG. 7 is a graphic showing phased and unphased alignments in an example genomic region that does not include a structural variant.

[0027] FIG. 8 is a graphic showing phased and unphased alignments in an example genomic region that includes a large deletion.

[0028] FIGs. 9A and 9B illustrate how the phasing methods described herein can be applied m the context of variant calling for cancer according to some embodiments of the disclosed technology.DETAILED DESCRIPTION

[0029] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to oneembodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.

[0030] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0031] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.Overview

[0032] In recent years, biotechnology firms and research institutions have improved hardware and software for sequencing nucleotides and determining variant calls for genomic samples. For instance, some existing sequencing platforms determine individual nucleobases within sequences from genomic samples’ cells by using conventional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existingplatfomis can monitor millions to billions of nucleic acid polymers being synthesized in parallel to predict nucleobase calls from a larger base call dataset. For instance, a camera in many SBS platforms captures images of irradiated fluorescent tags incorporated into oligonucleotides for determining the nucleobase calls. After capturing such images, existing sequencing platforms send base call data (or image-based data) to a computing device to apply sequencing data analysis software that determines a nucleobase sequence for a genomic sample or other nucleic acid polymer. For instance, such software maps and aligns nucleotide reads determined by the sequencing platform for a genomic sample with a reference genome. Based on differences between the aligned nucleotide reads and the reference genome, existing data analysis software can further utilize a variant caller to identify genotype and / or variants within a genomic sample, such as single nucleotide polymorphisms (SNPs), insertions or deletions (in dels), or structural variants,

[0033] Despite these recent advances, existing nucleobase sequencing platforms and sequencing data analysis software (together and hereinafter, “existing sequencing systems”) often utilize reference genomes that misrepresent certain populations and foment inaccurate read mapping and alignment and mistaken variant calling. For example, some existing sequencing systems use a linear reference genome that purportedly represents a consensus or example of genes and other nucleotide sequences of an organism. But about 93% of the primary assembly for the most common linear human reference genome, GRCh38 from the Genome Reference Consortium, is based on libraries from only 11 individuals, with 70% of the linear human reference genome coming from 1 individual. Accordingly, many existing systems use a linear reference genome that does not represent certain populations, common variants, or common population haplotypes.

[0034] To address this lack of genetic representation in linear reference genomes, some sequencing systems generate or use a graph reference genome. For example, some graph reference genomes include both a linear reference genome and graph augmentations with multi-nucleobase codes representing SNPs and / or indels and alternate contiguous sequences representing various alternative population haplotypes at given genomic regions. In some cases, such graph reference genomes stack and index numerous alternate contiguous sequences that can respectively stretch relatively long nucleobase distances (e.g., hundreds to thousands of base pairs in length) and, consequently, include redundant reference nucleobasesoverlapping a same region. The one-size-fits-all approach to graph reference genomes can accordingly consume excessive amounts of memory encoding alternate contiguous sequences.

[0035] While such graph reference genomes better account for some populations’ genetics, the expanded representation of existing graph reference genomes often includes an exorbitant number of alternative paths for alleles that are similar to other genomic regions and paths in the graph reference genome. Consequently, existing sequencing systems can significantly increase the difficulty of mapping and aligning sequence reads because sequence reads may map to multiple similar alternative paths in a genomic region, increasing confusion between multiple look-alike genomic regions. Indeed, the one-size-fits-all approach to graph reference genomes can decrease quality scores for mapping reads and undermine the expected utility of a more diverse set of alternate contiguous sequence to which reads can be mapped.

[0036] Indeed, these generic graph reference genomes — with an excessive number of alternative paths representing alternative contiguous sequences — frequently cause existing sequencing systems to misalign, incorrectly match, or miscall variants for a large number of samples as well as increase the chances of mismatched alignments with reads from a genomic sample. Due to having multiple look-alike population haplotypes that cover a given genomic region of a primary contiguous sequence, diminishing mapping quality (e.g., MAPQ 0) as such population haplotypes increase m number for the given genomic region, existing sequencing systems have often failed to scale up candidate population haplotypes in a graph reference genome without detrimental effects to the accuracy of mapping and alignment and consequent reductions to variant-calling accuracy.

[0037] Furthermore, although the human genome is largely diploid, the human reference genome assemblies utilized by many existing sequencing systems comprise a haploid representation of different reference genomic samples. By using a haploid reference genome for mapping and alignment of diploid genomic samples, existing sequencing systems frequently align nucleotide reads and determine variant calls that are negatively influenced by a reference bias in favor of alignment of such reads with alleles represented by the reference genome and to the detriment of alternative alleles — despite allele-variant differences between the sample diplotype and the haploid reference genome. Such a reference bias often leads to false positives (FPs) and false negatives (FNs) in variant calls, thus reducing the accuracy of existing sequencing systems m generating variant calls for a diploid genomic sample. These,along with additional problems and issues exist m existing sequencing systems. Methods and systems for creating a personalized haplotype database for improved mapping and alignment of nucleotide reads and improved genotype calling are disclosed in U. S. Provisional App. No.63 / 558,754, filed on February 28, 2024, which is hereby incorporated by reference in its entirety..

[0038] During massively parallel sequencing processes, such as those found in Next Generation Sequencing (NGS) systems, relatively short reads of 100-500 nucleotides are typically generated. During this process, sample nucleic acids may be fragmented by tagmentation of the sample directly on the flow cell, in some embodiments. In an example process, transposons are linked to the flow cell, and when contacted by sample nucleic acids may fragment the sample nucleic acids resulting in many fragments from the same nucleic acid molecule being generated. This example process may result in fragments from the same nucleic acid molecule being bound to the flow cell in geographically adjacent or nearby locations. These fragments then give rise to nucleic acid clusters at nearby locations on the flow cell, and these nucleic acid clusters will participate in sequencing reactions to generate reads or read pairs. Therefore, reads or read pairs that are generated by sequencing these fragments which are located nearby each other on the flow cell may be determined by the disclosed systems or methods to be “connected” or “linked”, or that there are “links” among the reads or read pairs since they may have derived from the same originating nucleic acid molecule. This type of linking information or connectivity information may enable reconstruction of a nucleic acid molecule by way of grouping together linked read pairs coming from nearby locations.

[0039] One way to improve accuracy of sequence read assembly and variant detection is to determine which of two parental haplotypes a sequence read corresponds to and assign each sequence read to one of the two parental haplotypes. Disclosed herein are methods and systems for phasing sequence reads from a genomic DNA sample. In some embodiments, sequence reads are generated from fragments of the genomic DNA sample which are bound to the flow cell. For example, long (e.g., greater than 500 bp) DNA fragments may be flowed across a sequencing flow cell and may attach to transposome complexes that are embedded on the surface of the flow cell. The transposome complexes can fragment the long DNA fragments into shorter fragments (e.g., of less than 500 bp), and attach sequencing adapters. Each shorterfragment can be amplified to form clusters, and the sequence of each shorter fragment can be determined using the clusters. The physical location of each cluster can be recorded in addition to sequence reads.

[0040] Next, in some embodiments, linkage information can be determined between the sequence reads based on the geographic location of each fragment on the flow cell and the mapping location on a reference genome. For example, given the physical location of two clusters on a flow cell and the locations the sequence reads map to on a reference genome, sequence reads may be linked in that a probability can be determined related to whether the sequence reads derive from the same original long DNA fragment, because sequence read clusters that are physically closer together on the flow cell are more likely to have derived from the same original long DNA fragment.

[0041] Next, in some embodiments, the sequence reads are mapped to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes. The linkage information can be used to map sequence reads because sequence reads which came from physically proximate clusters on the flow cell may be more likely to have derived from the same original long DNA fragment and may thus be more likely to be genomically proximate to one another. The set of candidate haplotypes can be a database of previously-determined haplotypes from different individuals in a population. The set of candidate haplotypes can include alternate sequences so that sequence reads which map to the reference genome sequence alone with low quality may match the alternate sequences from the set of candidate haplotypes, to thereby determine a mapping location on the reference genome.

[0042] Next, in some embodiments, the sequence reads are analyzed to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes, and the sequence reads are phased based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. For example, in some embodiments, the sequence reads are analyzed and phased by dividing the alignment to the reference genome into smaller subsections called “bins”, and analyzing sequence reads in each bin, one bin after another in a sequence.

[0043] In each bin, sequence reads can be compared to the set of candidate haplotypes to determine which of the candidate haplotypes are a likely first parental haplotype,and which are a likely second parental haplotype, for the bin. Furthermore, each sequence read can be analyzed to determine the likelihood that the sequence read is from a first parental haplotype or a second parental haplotype. The analysis and phasing process can be based on determinations of likely first and second parental haplotypes from the previous bin in the sequence of bins, and on links with sequence reads in previous bins. This is because one bin is likely to have the same likely first and second parental haplotypes as the previous bin unless a recombination event has occurred between bins, and because sequence reads which derive from the same original long DNA fragment would be in phase with one another because they would be from the same parental haplotype.

[0044] In some embodiments, once the methods and systems analyze all of the bins in a sequence in a first direction (e.g., “forward”), the methods and systems repeat the process of analyzing and phasing sequence reads in a second direction (e.g., “backward”). During the pass in the second direction, determinations of likely first and second parental haplotypes and read phasing from the pass in the first direction can be used in the analysis, in addition to determinations from previous bins in the pass in the second direction.

[0045] Based on the analysis, the methods and systems described herein can phase sequence reads by assigning each sequence read to a first parental haplotype or a second parental haplotype. The methods and systems can also determine, for each bin, the most likely first parental haplotype and second parental haplotype from the set of candidate haplotypes. This phasing and haplotype information can be used for various downstream purposes, for example, for sequence assembly and for determining phased variant calls.Definitions

[0046] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.

[0047] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.

[0048] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.

[0049] Language of degree used herein, such as the terms “approximately,” “about,” “generally,” and “substantially” represent a value, amount, or characteristic close to the stated value, amount, or characteristic that still performs a desired function or achieves a desired result. In some embodiments these terms encompass minor variations (for example, up to + / - 10%) from the stated value or amount.

[0050] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”),

[0051] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.

[0052] As used herein, the term “consists essentially of” (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.

[0053] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5 -methyl- CTP, 5-methyl-dCTP, HP, dlTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2 -thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0054] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents, A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.

[0055] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the senes.

[0056] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments,tlie reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. Regardless of the sequence length, in some cases, a reference genome represents an example set of genes or a set of nucleic acid sequences in a digital nucleic acid sequenced determined by scientists as representative of an organism of a particular species. For example, a linear human reference genome may be GRCh38 or other versions of reference genomes from the Genome Reference Consortium, In some cases, a reference genome includes inulti-base codes. As a further example, a reference genome may include a graph reference genome that includes both a linear reference genome and paths representing nucleic acid sequences from ancestral haplotypes, such as Illumina DRAGEN Graph Reference Genome hgl9. Graph genomes are data structures that can be stored in a computer memory and represent a genome as a type of graph. In general, a genome graph may have a structure with nodes representing DNA sequences, and edges connecting nodes which are adjacent to one another in the genome. The use of a graph genome and its related data structure may allow a computer to determine match sequence reads to different possible reference sequences more efficiently and accurately than analyses using linear genomic data.

[0057] The term “nucleic acid sample” herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence. In certain embodiments the nucleic acid sample comprises at least one nucleic acid sequence. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from ahuman subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein, A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.

[0058] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. Forexample, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis (SBS), sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0059] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 300 bp, and includes paired end sequence reads.

[0060] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more. A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed by assembling multiple short sequence reads.

[0061] As used herein, a “file” includes electronic files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.

[0062] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, m certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood of the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to m the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.

[0063] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).

[0064] Alignment may be performed by modifications and / or combinations of methods such as Burrows- Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BL. AT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, EL, AND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoalign & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0065] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.

[0066] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.

[0067] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith-Waterman score or a variation or version of a Smith- Waterman score for local alignment, suchas various settings or configurations used by DRAGEN by Illumina, Inc. for Smith- Waterman scoring.

[0068] Relatedly, as used herein, the term “mapping quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In some embodiments, for example, a mapping quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong}, rounded to the nearest integer. In the alternative to a mean or median mapping quality, in some implementations, a mapping quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.

[0069] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquid. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g. due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides etc.), nylon, ceramics, resms, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solid supports for some embodiments are located within a flow cell apparatus. Exemplary flow cells are set forth in further detail below.

[0070] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.

[0071] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1 mm2, 500 pm2, 100 pm2, 25 pm2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.

[0072] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially flat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.

[0073] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.

[0074] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element- containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. App, Pub, No. 2010 / 0120098, which is incorporated herein by reference. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.

[0075] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-type or mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transpositionreaction. For example, the transposon end can comprise DNA, RNA, modified bases, nonnatural bases, modified backbone, and can comprise nicks in one or both strands.

[0076] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads in ambiguous positions.

[0077] In one or more embodiments, the system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.

[0078] Moreover, in one or more embodiments, one or more sequencing data files in which the system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map ( CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.

[0079] Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping ofoligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a nonpatterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow’ cell.

[0080] The term “genotype call” refers to a determination or prediction of a particular genotype of a genomic sample or a sample nucleotide sequence at a genomic locus. In particul r, a genotype call for a nucleic acid can include a prediction of a particular genotype of the nucleic acids in a genomic sample with respect to a reference genome or a reference sequence at a genomic coordinate or a genomic region. For instance, in some cases, a genotype call includes a determination or prediction that a genomic sample comprises both a nucleobase and a complementary nucleobase at a genomic coordinate that is either homozygous or heterozygous for a reference base or a variant (e.g., homozygous reference bases represented as 0|0 or heterozygous for a variant on a particular strand represented as 0| 1 ). Accordingly, a genotype call can include a prediction of a variant or reference base for one or more alleles of a genomic sample and indicate zygosity with respect to a variant or reference base. A genotype call is often determined for a genomic coordinate or genomic region of a nucleic acid at which an SNP, insertion, deletion, or other variant has been identified for a population of organisms.Methods for Phasing Sequence Reads

[0081] In an aspect, disclosed herein are methods for phasing sequence reads from a genomic DNA sample.

[0082] In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. In some embodiments, generating sequence reads from fragments of the genomic DNA sample bound to a flow cell comprises: providing transposome complexes, wherein the transposome complexes comprise a transposase and a first polynucleotide comprising an end sequence and a first tag; contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides; and amplifying the fragmented target polynucleotides to form aplurality of nucleic acid clusters on a substrate. In some embodiments, the transposome complexes are attached to a substrate of the flow cell. In some embodiments, the sequence reads comprise paired end sequence reads.

[0083] FIG. 1 is a flow diagram that schematically illustrates an exemplary method 100 for phasing sequence reads from a genomic DNA sample. In some embodiments, the method 100 is implemented on a processor. The method 100 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system, for example the computing system 3000. When the method 100 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device, for example the server device 3102,

[0084] As shown in FIG. 1, the method 100 for phasing sequence reads from a genomic DNA sample may start from start block 110. The method 100 may proceed to block 120, wherein sequence reads sequenced from fragments of a genomic DNA sample bound to a flow cell are received by a system. In one embodiment the sequence reads are stored in data files on a computer system, for example the computer system 3000, and may be, for example, retrieved from a computer or electronic storage device, for example the data store 390. The sequence reads may be generated as further described above. The sequence reads may be stored, retrieved and manipulated in any electronic file format.

[0085] Receiving the sequence reads may be performed by a nucleic acid sequencing system itself, or alternatively, the sequence reads in some embodiments may be received by a cloud-based electronic system or memory which then carries out the functions mentioned herein.Determining linkage information

[0086] Once the sequence reads are received, the method 100 may proceed to block 130, wherein linkage information between the sequence reads based on the geographic location of each fragment on the flow cell is determined.

[0087] Embodiments of the present disclosure relate to methods and systems which use “links” or “linkage information” between sequence reads. The “link” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule.In some next generation sequencing (NGS) systems, fragments of DNA, such as genomic DNA, from a biological source are sheared to create shorter fragments which can be sequenced in a single read. The shearing process can create these shorter fragments which land on the flow cell and the spatial location of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments which come from the same nucleic acid molecule have been found to bind closer together on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are close together spatially and also close together on the genome, the clusters are more likely to have come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty7in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability that two reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.

[0088] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by7well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). As described above, according to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in spatially nearby positions as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In regular library preparations with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell, flow cell proximity information is retained. This proximity information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.

[0089] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. In some embodiments, the complexed strand and transposome is in solution, and is then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.

[0090] Once the fragments have been bound to substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. The location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster.

[0091] In some embodiments, linking information is determined by analyzing (for example, statistically analyzing with a model) the genomic distance between two reads and the proximity between the two reads on the flow cell. In some embodiments, the methods and systems determine whether the genomic distance and / or flow cell proximity is below a threshold. In some embodiments, the methods and systems determine the presence or absence of a link between the two sequence reads. In some embodiments, the methods and systems determine a linking quality score between the two sequence reads. In some embodiments, the methods and systems analyze genomic distance and flow cell proximity for a plurality of pairs of two sequence reads (for example, each possible pair) of two sequence reads in a dataset. Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U. S. Provisional PatentApplication Nos. 63 / 600460, 63 / 614066, 63 / 700,049, and 63 / 700,262 the disclosure of each of which is incorporated herein by reference in its entirety.

[0092] In some embodiments, sequence reads comprise a barcode. As used herein “barcode” refers to a short, unique nucleic acid sequence used to tag or label different nucleic acid samples or different nucleic acid molecules. In some embodiments, long DNA template molecules are fragmented, and a barcode sequence is added to each fragment, with a unique barcode sequence for each long DNA template. The fragments are sequenced to produce short sequence reads. In some embodiments, the sequence reads are analyzed, and the barcodes are used to identify which sequence reads are part of the same long DNA template molecule. Thus, barcodes can help retain connectivity7information for short sequence reads. Such embodiments can be considered an alternative form of linking and can be used in the methods described herein.Mapping sequence reads to a reference genome

[0093] The method 100 may proceed to block 140, wherein sequence reads are mapped to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes.

[0094] In some embodiments, for example, at block 140, the sequence reads are initially mapped and aligned using a set of candidate haplotypes. In one embodiment, the candidate haplotypes are stored in a database of population haplotypes. These databases of haplotypes may include panels of any number of known haplotypes, for example a panel of 128, 256, 512 or any other number of known haplotypes. This mapping of the sequence reads to the known haplotypes is performed to generate an initial alignment data file. This initial alignment datafile can be any type of data file, including for example a rescored binary alignment map (BAM) file.

[0095] In some embodiments, the reference genome comprises a “graph” reference genome comprising a primary contiguous genomic sequence augmented by a plurality of alternate genomic contiguous genomic sequences representing candidate haplotypes associated with the reference genome.

[0096] In some embodiments, linkage information is used to map sequence reads to the reference genome. For example, once the nucleotide sequence of each cluster has beendetermined, the method can start to map those reads to determine the original target polynucleotide from which the read originated. In some embodiments, the mapping process takes into account the spatial location of each cluster, such that clusters which are closer to each other on the flow cell are more likely to have originated from the same target polynucleotide.

[0097] For example, in some embodiments, the methods and systems align an initial set of reads to the reference genome (e.g., a subset of all of the sequence reads) in order to define linking model parameters based on flow' cell proximity and mapping location. In some embodiments, once linking model parameters for the current sample are defined, the full link-informed alignment process begins. For example, when reads are being aligned to reference, various alignment candidates are evaluated for each read, and potential links for each candidate alignment position of a read are evaluated and a boost to the mapping quality is calculated from the linking information in each possible candidate position. For example, in some embodiments, all candidate alignments are obtained and a pairwise comparison between potentially linked reads is performed. In some embodiments, a link bonus for each candidate alignment is determined based on the linking model parameters. Alignment can thus proceed iteratively using linking information to align sequence reads to the reference genome.

[0098] Furthermore, by mapping the sequenced fragments to target polynucleotides using the flow cell proximity information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the spatial location of each cluster into account during the mapping process. Therefore, spatial information that includes relative distances between various clusters on a flow' cell may be leveraged to adjust mapping information, thereby increasing the read quality and alignment accuracy of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. Increasing the read quality of these previously discarded reads, by improving the confidence of read pair’s alignment based on linking information wdtli a high link quality score, may improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling. Processing DNA samples suitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments.

[0099] In some embodiments, once a linking model is obtained, linking information is used to align sequence reads to the reference genome. In some embodiments, linking information is updated based on the alignments. In some embodiments, the alignments can be updated based on the updated linking information. In some embodiments, the process continues to iteratively update alignments and linking information.

[0100] In some embodiments, at block 140, initial alignments of the set of sequence reads with respective genomic regions of the reference genome are determined. In some embodiments, subsets of sequence reads are identified that, according to the initial alignments, align to respective regions of the reference genome. In some embodiments, an initial set of haplotype likelihoods are determined for each respective region of the reference genome by comparing the sequence reads with the set of candidate haplotypes. Moreover, in some embodiments, the initial alignments are determined by (1) identifying, as indicated within the set of candidate haplotypes, allele-variant differences between candidate haplotypes and a primary contiguous sequence at respective genomic regions of the reference genome, and (2) rescoring candidate reference alignments of the set of sequence reads with the primary contiguous sequence according to the identified allele- variant differences.

[0101] Having determined, via the aforementioned alignment method, or by another method, initial alignments for a set of sequence reads, in one or more embodiments, an alignment data file is generated including information for analyzing sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haploptypes. For example, such an alignment data file (e.g., a personalized Binary Alignment Map (BAM) file) can include the initial alignments of the set of sequence reads, alignment scores corresponding to the initial alignments, and one or more candidate haplotypes identified for each sequence read of the set of sequence reads. The alignment data file can include a set of initial haplotype likelihoods. The alignment data file can also include linkage information determined earlier in the workflow, for example, a list for each sequence read of the all the sequence reads to which it is linked.Analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes

[0102] The method 100 may proceed to process block 150, wherein the sequence reads are analyzed to determine the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. The specific features of the process block 150 are explained in more detail with reference to FIG. 2, which is a block diagram that illustrates additional details on the process taking place within the process block 150.Phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes

[0103] The method 100 may proceed to block 160, wherein the sequence reads are phased based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0104] For example, at block 160, sequence reads can be assigned to one of two haplotypes based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. Sequence reads can be assigned to the most probable haplotype based on the analysis at block 160.

[0105] The method 100 may end at end block 170.Process of analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes

[0106] FIG. 2 is a block diagram that illustrates additional details on the specific methods taking place within the process block 150, wherein the sequence reads are analyzed to determine the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. As shown in FIG. 2, the process 150 for analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes may include the following steps.Divide alignment of the sequence reads to the reference genome into overlapping sections

[0107] First, at block 210, the alignment of the sequence reads to the reference genome is divided into overlapping sections. In some embodiments, the overlapping sections are about 1 Mbp in length, and have an overlap of about 100 kbp. However, one of ordinaryskill in the art will recognize that various sizes of overlapping sections could be used. In some embodiments, each overlapping section is processed independently. In some embodiments, each overlapping section is processed in parallel.Divide each overlapping section into bins

[0108] At block 220, each overlapping section is divided into bins, which refers to a subsection of the reference genome that is smaller than the overlapping section. In some embodiments, each bin is of approximately (e.g., within 10%) the same size as other bins. In some embodiments, each bin is about 4 kbp in size. One of ordinary skill in the art will recognize that other sizes may be used that are less than or greater than 4 kbp, such as, but not limited to, 250 bp, 500 bp, 1 kbp, 1500 bp, 5 kbp, 10 kbp, and so forth. It should be realized that bin sizes can be adjusted based on a balance of computational efficiency and accuracy. For example, smaller bin sizes may improve accuracy by providing a more granular determination of sequencing depth along the reference genome, but may be computationally more intense to process. A desired bin size can be determined based on these considerations.First direction: Analyze sequence reads

[0109] At block 230, the methods and systems can proceed in a first direction, bin by bin, along the sequence of bins to analyze sequence reads.

[0110] At block 230, the methods and systems can determine, at each bin: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype.

[0111] In some embodiments, the methods and systems determine a most-probable first parental haplotype and a most- probable second parental haplotype from the set of candidate haplotypes for each bin and a most-probable phasing for each sequence read in the bin.

[0112] In some embodiments, the analysis in each bin is based on a determination of a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and / or a probability for each sequence read that the sequence read originated from the first parental haplotype, from a previous bin in the sequence of bins. In some embodiments, the methods and systems analyzethe sequence reads in a given bin b based on sequence information and linking information to update a determination of a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype from the previous bin in the sequence, bin b-1.

[0113] In some embodiments, the methods and systems analyze linking information between sequence reads in bin b with sequence reads in the previous bin in the sequence of bins, bin b-1. In some embodiments, the methods and systems analyze sequence similarity between sequence reads in each bin and each haplotype of the set of candidate haplotypes.

[0114] In some embodiments, the methods and systems analyze sequence reads using an imputation model. In some embodiments, the imputation model is a hidden Markov model (HMM) algorithm. In some embodiments, the first direction is a forward pass through each bin in a sequence of bins along the reference genome,

[0115] For example, in some embodiments, the methods and systems utilize an HMM algorithm to impute haplotype set posterior probabilities for sets of candidate haplotypes across adjacent bins of a sequence of bins of a reference genome associated with a genomic sample. Based on the haplotype set posterior probabilities and on likely phasing of linked reads in a previous bin in the sequence of bins, the methods and systems can determine a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype.

[0116] Furthermore, in various embodiments, the methods and systems can utilize a variety of methods of scoring sets of candidate haplotypes. In some embodiments, for example, the methods and systems utilize a Variational Bayesian model implementing an iterative algorithm to generate haplotype set scores for the respective sets of candidate haplotypes. Also, in one or more embodiments, the methods and systems categorize sequence reads as inherited from respective first and second parents, generate individual haplotype scores for the categorized reads, and generate haplotype set scores for pairs of candidate population haplotypes by combining individual haplotype scores from sequence readsinherited from the first parent with individual haplotype scores from sequence reads inherited from the second parent.Second direction: Analyze sequence reads

[0117] At block 240, the methods and systems can proceed in a second direction along the sequence of bins to analyze sequence reads to determine: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype. In some embodiments, the second direction is a backwards pass through each bin in a sequence of bins along the reference genome.

[0118] The methods and systems can analyze the sequence reads similarly to what is described for block 240. For example, in some embodiments, the methods and systems utilize an HMM algorithm to impute haplotype set posterior probabilities for sets of candidate haplotypes across adjacent bins of a set of bins of a reference genome associated with a genomic sample.

[0119] At block 240, the methods and systems can also analyze the sequence reads based on a determination from the analysis in the first direction -- such as a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype. Thus, in some embodiments, based on the haplotype set posterior probabilities and on likely phasing of linked reads in a previous bin in the sequence of bins and from the same bin in a forward pass of the bin, the methods and systems can determine a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype.Joining overlapping sections together

[0120] At block 250, the methods and systems can compare sequence read phasing within the overlap of two consecutive overlapping sections and join overlapping sections together.

[0121] For example, if in a first overlapping section a first set of sequence reads are assigned to haplotype Hl and a second set of sequence reads are assigned to haplotype H2, but in the overlap with a second overlapping section, the same first set of sequence reads is assigned to H2 and the second set of sequence reads is assigned to Hl, phasing can be updated in one of the overlapping sections to concord with the other. Thus, the methods and systems can join together the phasing determinations made in the overlapping sections.

[0122] Thus, in some embodiments, the methods and systems analyze sequence read phasing within an overlap of two consecutive overlapping sections and update phasing of one or more sequence reads based on the analysis of sequence read phasing within the overlap. The methods and systems can adjust chromosome-wide or genome-wide phasing of the sequence reads based on the overlaps.Methods of Determining a Haplotype Nucleotide Sequence

[0123] Further disclosed herein are methods of determining a haplotype nucleotide sequence. In some embodiments, the method includes phasing sequence reads according to any of the methods described herein; and assembling phased sequence reads into a first parental haplotype, thereby determining a haplotype nucleotide sequence. In some embodiments, the method further includes assembling phased sequence reads into a second parental haplotype, thereby determining a second haplotype nucleotide sequence.Methods of Detecting Phased Sequence Variants

[0124] Other aspects include methods of detecting a phased sequence variant.

[0125] In conventional variant-calling methods, variants are first called without read phasing information and are then phased downstream of calling. In some embodiments of the present disclosure, the phasing methods described herein allow the methods and systems to phase and call sequence variants at the same time. By phasing and calling sequence variants simultaneously, the methods and systems described herein increase both variant phasing accuracy and variant calling accuracy. Furthermore, in the context of tumor-only somatic variant calling, read phasing information can advantageously help distinguish somatic variants from gemiline variants.

[0126] In some embodiments, the method includes phasing sequence reads according to any of the methods disclosed herein, and detecting one or more phased sequence variants. In some embodiments, the sequence variant is a structural variant (for example, an insertion, deletion, or translocation of more than 35bp, 50 bp, 75bp, or 100 bp or more). In some embodiments, the sequence variant is a small variant (less than 35bp, 50 bp or less).

[0127] Also disclosed herein are methods of detecting a structural variant in a genomic DNA sample. In some embodiments, the method includes phasing sequence reads according to any of the methods described herein; generating one or more sequence assemblies based on the phased sequence reads; and detecting a structural variant based on the one or more sequence assemblies.

[0128] In some embodiments, the methods and systems determine one or more genotype calls based on the phased sequence reads. For example, the methods and systems can determine whether a genomic sample is homozygous or heterozygous for one or more variants.Systems for Phasing Sequence Reads

[0129] Additional aspects include electronic systems. In some embodiments, the systems are for phasing sequence reads from a genomic DNA sample. In some embodiments, the system includes a processor configured to perform any of the methods described herein. In some embodiments, the method comprises: receiving sequence reads generated from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0130] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads generated from fragments of the genomic DNA sample bound to aflow cell; determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; map the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyze the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phase the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes. The non-transitory computer-readable media can include a plurality of instructions, which when executed by at least one processor, cause the at least one processor to perform any of the methods described herein.

[0131] FIG. 3A illustrates a diagram of an environment in which a system for detecting phasing sequence reads from a genomic DNA sample can operate in accordance with one or more implementations. The following paragraphs describe the phasing system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 3A illustrates a schematic diagram of a computing system 3000 in which a phasing application 3106 operates in accordance with one or more implementations. As illustrated, the computing system 3000 includes one or more server device(s) 3102 connected to a user client device 3108, a local device 3118, and a sequencing device 3114 via a network 3112. The network 3112 can comprise any suitable network over which computing devices can communicate.

[0132] As shown m FIG. 3 A, the computing system 3000 includes the server device(s) 3102. In various implementations, the server device(s) 3102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic-acid polymers. In some implementations, the server device(s) 3102 receive various data from the sequencing device 3114, such as data from a sample genome and / or sequence reads. The server device(s) 3102 may also communicate with the user client device 3108. In particular, the server device(s) 3102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 3108.

[0133] As shown, the server device(s) 3102 includes a sequencing application 3110. In general, the sequencing application 3110 analyzes the data (such as call data) received from the sequencing device 3114 or elsewhere to determine nucleobase sequences for nucleic-acid polymers. For example, the sequencing application 3110 can receive raw data from thesequencing device 3114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 3110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0134] As also shown, the sequencing application 3110 includes the phasing application 3106. As described below, in some embodiments, the phasing application 3106 can phase sequence reads from a genomic DNA sample. For example, in some embodiments, the phasing application 3106 receives sequence reads generated from fragments of the genomic DNA sample bound to a flow cell; determines linkage information between the sequence reads based on the geographic location of each fragment on the flow cell; maps the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes; analyzes the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; and phases the sequence reads based on the l inkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

[0135] While the sequencing application 3110 has been described as including the phasing application 3106, other systems or methods may be included within the sequencing application 3110, such as an application to detect sequence variants or to assemble sequence reads (not illustrated).

[0136] Moreover, while the phasing application 3106 is described being implemented on the server device(s ) 3102, as part of the sequencing application 3110, in some implementations, the phasing application 3106 is implemented by (such as located entirely or in part) on the user client device 3108, the sequencing device 3114, and / or the local device 3118. As mentioned, in some implementations, phasing application 3106 is implemented by one or more other components of the computing system 3000, such as the sequencing device 3114. In particular, the phasing application 3106 can be implemented in a variety of different ways across the server device(s) 3102, the network 3112, the user client device 3108, the local device 3118, and the sequencing device 3114.

[0137] As further shown in FIG. 3A, the computing system 3000 includes the user client device 3108. In various implementations, the user client device 3108 can generate, store, receive, and send digital data. In particular, the user client device 3108 can receive the data from the sequencing device 3114. As further illustrated, the user client device 3108 includes asequencing application 3110. The sequencing application 3110 may be a web application or a native application stored and executed on the user client device 3108 (for example, a mobile application, desktop application, or web application). The sequencing application 3110 can receive data from the sequencing application 3110 and / or phasing application 3106. For example, the user client device 3108 can receive variant call files and / or alignment files from the sequencing application 3110.

[0138] The sequencing application 3110 can also include instructions that (when executed) cause the user client device 3108 to receive data from the phasing application 3106 and present data from the sequencing device 3114 and / or the server device(s) 3102, Furthermore, the sequencing application 3110 can instruct the user client device 3108 to display data for variant calls, such as nucleobase calls or an indication of a sequence variant or genotype. Indeed, the user client device 3108 can display nucleobase call results for a genome sample and / or an indication of a predicted variant or genotype.

[0139] As further shown in FIG. 3A, the computing system 3000 includes the sequencing device 3114. In various implementations, the sequencing device 3114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 3114 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 3114. More particularly, the sequencing device 3114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 3114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 3112, m some implementations, the sequencing device 3114 bypasses the network 3112 and communicates directly with the user client device 3108.

[0140] As further depicted in FIG. 3A, in some implementations, the server device(s) 3102 includes a distributed collection of servers, where the server device(s) 3102 include several server devices distributed across the network 3112 and located in the same or different physical locations. For instance, the server device(s) 3102 can be implemented, in whole or in part, on the local device 3118. To illustrate, the local device 3118 may implement the sequencing application 3110 and / or the phasing application 3106. Further, the serverdevice(s) 3102 and / or the local device 3118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0141] The user client device 3108 illustrated in FIG. 3 A can include various types of client devices. For example, in some implementations, the user client device 3108 includes non-niobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 3108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0142] Though FIG. 3A illustrates the components of the computing system 3000 communicating via the network 3112, in certain implementations, the components of computing system 3000 can also communicate directly with each other, bypassing the network 3112, For instance, in some implementations, the user client device 3108 communicates directly with the sequencing device 3114. Additionally, in some implementations, the user client device 3108 communicates directly with the phasing application 3106 and / or the server device(s) 3102. In some implementations, the user client device 3108 communicates directly with the local device 3118. Moreover, the phasing application 3106 can access one or more databases housed on or accessed by the server device(s) 3102 or elsewhere in the computing system 3000.

[0143] FIG. 3B is a block diagram of an exemplary server device 3102 that may be used in connection with the computing system 3000 of FIG. 3 A. The server device 3102 may be configured to phase sequence reads from a genomic DNA sample. The general architecture of the server device 3102 depicted in FIG 3B includes an arrangement of computer hardware and software components. The server device 3102 may include many more (or fewer) elements than those shown in FIG. 3B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 3102 includes a processing unit 310, a network interface 320, a computer readable medium drive 330, an input / output device interface 340, a display 350, and an input device 360, all of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity to one or more networks or computing systems. The processing unit 310 may thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350via the input / output device interface 340. The input / output device interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0144] The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the server device 3102, The memory 370 may store a reference genome 373, such as for use by the sequencing application 3110. The memory 370 may further include computer program instructions and other information for implementing aspects of the present di sclosure,

[0145] For example, in one embodiment, the memory 370 includes a sequencing application 3110, which may include a phasing application 3106. The phasing application 3106 can perform the methods disclosed herein. In addition, memory 370 may include or communicate with the data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference genomes.

[0146] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0147] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD-ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0148] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0149] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0150] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such asPDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0151] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, m a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0152] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resourcestorage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0153] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0154] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0155] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework ( University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Additional Embodiments of Methods for Phasing Sequence Reads

[0156] FIG. 4 is a schematic representation of an exemplary method for phasing sequence reads. As shown in FIG. 4, a database of candidate haplotypes (HapDB 401), sequence reads 402, and a linking model 403 which determines linkage information, are input to a graph mapper module 410. The graph mapper module 410 is a graph reference genomethat includes both a linear reference genome and paths representing nucleic acid sequences from ancestral haplotypes. In some embodiments, the graph mapper module 410 is Illumina DRAGEN Graph Reference Genome hg19. The graph mapper module 410 maps the sequence reads 402 to a reference genome using a graph reference genome with a primary sequence, where the candidate haplotypes of the HapDB 401 are alternate sequences. The graph mapper 410 then outputs a BAM file 420 which contains haplotype likelihoods and link information for each of the sequence reads.

[0157] The BAM file 420 and the HapDB 401 are then fed into a personalization process 430, which determines a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype. This process is described in more detail in FIGS. 5-6. The personalization process 430 outputs the most likely first parental haplotype and most likely second parental haplotype for each bin (selected haplotypes 440), a personalized imputed VCF 450 which includes imputed genotype calls for the sample, and a phased BAM file 460 where sequence reads are phased to a first parental haplotype or a second parental haplotype.

[0158] FIG. 5 illustrates an exemplary data flow diagram based on a forwardbackward hidden Markov Model (HMM) process. For each overlapping section 500 of the alignment (e.g., about 1 Mbp in size), the HMM proceeds in a forward pass 510 along a sequence of bins 520 in overlapping section 500. Figure 6 below describes the analysis which occurs during the forward pass. The HMM receives sequence read information from a database 530 of sequence reads for each bin, as well as determinations from the previous bin m the sequence of bins. After the forward pass is complete, the HMM proceeds to perform a backward pass 540 through the sequence of bins. During the backward pass, the HMM receives sequence read information from the database 530 of sequence reads for each bin, as well as determinations from the previous bin in the sequence of bins and from the same bin during the forward pass.

[0159] FIG. 6 is a schematic representation of exemplary' analysis steps within a forward pass of an exemplary bin b.

[0160] During the analysis, the sequence reads in bin b are analyzed to determine: 1) the likely haplotypes for parental haplotype 1 (Hl);2) the likely haplotypes for parental haplotype 2 (H2); and3) the probability, for each sequence read, that the read originated from Hl.

[0161] Because the analysis is based on prior determinations of 1) - 3) above from previously processed bins m the sequence of bins, the analysis in bin b can be considered an update of previous determinations from the immediately previous bin ( >-l). For the first bin in the sequence, uniform priors may be assumed, as it may be assumed that each candidate haplotype is equally likely, and phasing of sequence reads to one parental haplotype or the other is equally likely. The update uses information from the bin b, including the read data X, to update the prior determination of 1) - 3) above.

[0162] To update 1 ) - 2), the system takes the posterior probabilities corresponding to likely Hl and H2 haplotypes determined from bm b-\, adds a constant to represent the probability of a recombination event between bin b-l and b, and renormalizes the posterior probabilities to become the prior probabilities for bm b. The sequence reads in bin b are compared to each candidate haplotype, and the probability of each candidate haplotype can be adjusted based on sequence similarity with the sequence reads. For example, if the sequence reads match one candidate haplotype, the probability for that candidate haplotype can be increased during the update, and if the sequence reads do not match another candidate haplotype, the probability for that candidate haplotype can be decreased during the update.

[0163] To update 3), the system determines a priori read phasing probabilities by analyzing links from reads in bin b to other reads which have been phased in previous bins. If there are multiple links to sequence reads in previous bins, the system determines which link has the highest quality score, and the phasing of that sequence read is used to provide the phasing probability. Sequence reads which overlap a heterozygous variant can be phased based on which allele they include at the site of the heterozygous variant.

[0164] The bin data Ais a matrix that records how well each sequence read matches each candidate haplotype m the database. It is expected that all sequence reads should match at least one parental haplotype from the database. In the case of sequence reads that overlap a heterozygous variant, roughly half of the reads are expected to match the first parental haplotype and half match the second parental haplotype.

[0165] A model is used to determine the likelihood of Hl, H2 and the read phasings for the bin data X. In theory, this model and the priors described above are sufficient to useBayes rule to get updated posterior probabilities; however, this is intractable due to the large number of ways the reads may be collectively phased. To overcome this limitation, approximate inference with a variational Bayes algorithm is used to determine a most-probable first parental haplotype (Hl) and a most-probable second parental haplotype (H2) for bm b from the set of candidate haplotypes, and a most-probable phasing for each sequence read in the bin b.Electronic Files

[0166] In some embodiments, the methods and systems store a set of sequence reads and associated phasing information in an electronic file. In some embodiments, the electronic file comprises a BAM file. In some embodiments, the BAM file comprises a tag that indicates whether a sequence read is phased to a first parental haplotype or a second parental haplotype, or a probability of phasing to a first parental haplotype or a second parental haplotype. In some embodiments, the methods and systems store a set of genotype calls in a VCF file.

[0167] The methods and systems described herein can include a step of storing information generated by the methods and systems, in an electronic file, or retrieving information from an electronic file. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.

[0168] In some embodiments, the electronic file includes information related to a co-proximity property, where reads that are nearby within an individual’s genome are also nearby on the flow cell surface. This property enables improved mapping, variant calling, and phasing, among other enhancements. The following is a list of terms that are used to describe data in Tables 1 and 2 below.

[0169] Template molecule: A template molecule is a long DNA molecule from either a standard or high molecular weight extraction that is loaded into the sequencing cartridge for on-flow cell tagmentation.

[0170] Proximal reads: Proximal reads are reads from the same template molecule that are nearby on the flow cell surface.

[0171] Proximity group: A proximity group is a set of reads that have been derived from the same template molecule.

[0172] Genomic proximity: Genomic proximity is the genomic distance between reads within a proximity group.

[0173] Template length: Genomic distance spanned by a proximity group from the same original template.

[0174] Co-proximity rate: Percentage of ail reads that are co-proximal with at least one other read.

[0175] Co-proximity quality: A Phred-scaled quality score that estimates the probability that two reads derive from the same original template molecule,

[0176] In some embodiments, the method includes a step of storing sequence reads or other sequence information in a BAM file or a VCF file. In some embodiments, the BAM file has tags and and / or VCF file has fields as described below. In some embodiments, metrics for these tags are stored in a separate CSV file with the definitions provided.

[0177] Table I describes new tags for BAM files. In some embodiments, the BAM file also includes other tags, for example debug tags (xn, xp, xf, xa, xb, ld, ad, bd, ws, wln, and pq)- Table I: New BAM tags

[0178] In some embodiments, the VCF file includes fields related to variant calling and haplotype information (for example, Whatshap phase).

[0179] In some embodiments, the method includes creating an electronic file that includes one or more of the following metrics:Table 2EXAMPLES

[0180] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1

[0181] In the following example, the phasing method 100 described herein was tested on sequence reads. The phased sequence reads were used to impute small variants.

[0182] Sequence reads were obtained from human genomic sample HG002 using an on-flowcell method. Briefly, DNA fragments were flowed across a flow cell with transposome complexes embedded on the surface. The transposome complexes further fragmented the DNA fragments and attached sequencing adapters. Sequence reads were obtained by cluster amplification, and the flow cell location of the clusters was recorded. The sequence reads were processed through a prototype pipeline to give a BAM file of corrected mappings with the xl and hl tags shown in Tables 1 -2, The BAM file was processed with a prototype personalization / read-phasing pipeline shown in Figs. 4-6, which output personalized (imputed) phased VCF files. These phased variant calls were compared to a phased VCF file truthset for the HG002 genome. The imputed recall and precision were calculated for a subset of the truth VCF file, and those variants that were also present in the HapDB. The software Whatshap was used to compare and calculate the switch error between the imputed variants and the truthset.

[0183] The results are shown in Table 1 below, which demonstrates that the switch error rate is the fraction of adjacent heterozygous variants that are incorrectly phased (with respect to a truthset). Recall and precision are with respect to the truthset in confident regions intersected with HapDB variants. Sequence reads were phased to haplotype Hl or H2 withsome probability value. In Table 1, Q10 and Q20 refer to Phred-scale quantities. The “% reads phased at > Q10” is the fraction of sequence reads phased to Hl with a probability that is less than 10% or greater than 90%, and the “% reads phased at > Q20” is the fraction of sequence reads phased to Hl with a probability that is less than 1% or greater than 99%.Table 1Example 2

[0184] In the following example, the exemplary phased alignments are described, where the sequence reads were phased using the method 100 described herein. The phased reads were aligned to reference human genome hg38 in two haplotypes, Hl and H2, in a region that does not include a structural variant (FIG. 7) and in a region that includes a deletion (FIG.8). The remaining unphased reads were also aligned to the reference genome in the same regions.

[0185] FIG. 7 is a graphic showing a phased and unphased alignments m an example genomic region that does not include a structural variant. As shown in FIG. 7, most sequence reads overlapping a heterozygous variant are phased. Some sequence reads which do not overlap a heterozygous SNP have also been phased based on linking information,

[0186] FIG, 8 is a graphic showing a phased and unphased alignments in an example genomic region that includes a large deletion. Many fewer reads are phased to the haplotype with the deletion.Example 3

[0187] In the following example, the phasing method 100 described herein was applied in the context of variant calling for cancer. Method 100 includes determining linkage information between sequence reads and using such linkage information in mapping and phasing the sequence reads, resulting in sequence read phasing information. The sequence read phasing information obtained from method 100 can advantageously help distinguish between a variant that would be associated with cancer (herein referred to as the “somatic variant”) and a variant that would occur in the germ line genome (herein referred to as the “germ line variant”), because the sequence read phasing information can be used to separate sequence reads corresponding to different haplotypes and this increases the dynamic range of detection.

[0188] As illustrated in FIG. 9A, for a sample that only contains normal cells ha ving the same germline genome, if there is a mutation (indicated by “X”) in one haplotype (“phase 2”) of the genome, the mutation is expected to occur in all of the genetic material in the sample associated with this haplotype. Therefore, the mutation should be detected in nearly 100% of the sequence reads associated with this haplotype, subject to sequencing errors. However, if the sequence reads are not phased, then in the example of FIG. 9A, it would be determined that about 50% of all sequence reads have the mutation, since the sequence reads are not assigned to particular haplotypes.

[0189] A tumor sample obtained from a cancer patient is usually heterogeneous or impure, containing both normal cells and cancerous cells. As illustrated in FIG. 9B, for a heterogeneous / impure tumor sample, if the cancer causes a mutation (indicated by “X”) in one haplotype (“phase 2”) of the genome, the mutation is expected to be found in the genetic material associated with this haplotype from the cancerous cells, but not from the normal cells. Therefore, the mutation should be detected in a fraction of the sequence reads associated with this haplotype, for instance, about 60% in the example of FIG. 9B may have the mutation. However, if the sequence reads are not phased, then it would be determined that about 30% of all sequence reads have the mutation in the example of FIG. 9B, since the sequence reads are not assigned to particular haplotypes.

[0190] Without the sequence read phasing information, it is difficult to distinguish a somatic variant from a germline variant, because in view of sequencing errors, the percentage of sequence reads having a variant may not be sufficiently different for the two cases to besuggestive. For instance, 30% in the example of FIG. 9B versus 50% in the example of FIG.9 A may not be sufficiently different from one another in view of sequencing errors.

[0191] Using the sequence read phasing information obtained from method 100 to separate sequence reads corresponding to different haplotypes increases the dynamic range of detection, such that the percentage of sequence reads having a variant can be a useful indicator for distinguishing a somatic variant from a germline variant. For instance, 60% in the example of FIG. 9B versus 100% in the example of FIG. 9A can be sufficiently different despite sequencing errors. This allows the disclosed technology to efficiently distinguish a somatic variant from a germline variant when sequencing a tumor sample alone. This differs from existing technologies that can only determine whether a mutation or variant is somatic or germline by requiring that a control sample also be sequenced and comparing results from the tumor sample to those from the control sample. Therefore, compared to existing technologies, the disclosed technology, which can determine somatic variations from germline variations without reference to a control, allows for a more efficient workflow and can save sequencing costs.Other Considerations

[0192] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0193] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusivesense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0194] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0195] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.

[0196] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0197] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

[0198] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.

Claims

WHAT IS CLAIMED IS:

1. A method for phasing sequence reads from a genomic DNA sample, the method comprising:generating sequence reads from fragments of the genomic DNA sample bound to a flow cell;determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell;mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes;analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; andphasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

2. The method of claim 1, wherein the method comprises dividing an alignment of the sequence reads to the reference genome into bins, and phasing sequence reads by bin.

3. The method of claim 1 or claim 2, wherein the method comprises phasing the sequence reads in a sequence of bins, wherein the sequence of bins proceeds in a first direction along the reference genome.

4. The method of claim 3, further comprising determining a most-probable first parental haplotype and a most-probable second parental haplotype from the set of candidate haplotypes for each bin and a most-probable phasing for each sequence read in the bin.

5. The method of any one of claims 1-4, wherein the method comprises analyzing sequence reads in each bin in a sequence of bins to determine: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype.

6. The method of claim 5, wherein determining the distribution of candidate haplotype likelihoods for a first parental haplotype, the distribution of candidate haplotype likelihoods for a second parental haplotype, and the probability for each sequence read that thesequence read originated from the first parental haplotype, is based on a determination from a previous bin in the sequence of bins.

7. The method of any one of claims 5-6, wherein the probability for each sequence read that the sequence read originated from the first parental haplotype is determined based on linkage information with sequence reads in a previous bin in the sequence of bins.

8. The method of any one of claims 5-7, wherein determining the distribution of candidate haplotype likelihoods for a first parental haplotype, the distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype comprises analyzing sequence similarity between sequence reads in each bin and each haplotype of the set of candidate haplotypes.

9. The method of any one of claims 3-8, wherein the method further comprises phasing sequence reads sequentially by bin in a second direction along the reference genome.

10. The method of claim 9, wherein phasing sequence reads sequentially by bin in a second direction comprises analyzing sequence reads in each bin in a sequence of bins to determine: a distribution of candidate haplotype likelihoods for a first parental haplotype, a distribution of candidate haplotype likelihoods for a second parental haplotype, and a probability for each sequence read that the sequence read originated from the first parental haplotype, based on a determination from a previous bin in a sequence of bins, and based on a determination from a bin from phasing sequence reads sequentially by bin in the first direction.

11. The method of any one of claims 1-10, wherein the method comprises separating an alignment of the sequence reads to the reference genome into overlapping sections, separating each overlapping section into bins, and phasing sequence reads by bin.

12. The method of claim 11, wherein each overlapping section is about 1 Mbp in size, wherein each bin is about 4 kbp in size, and wherein each overlapping section overlaps with another overlapping section by about 100 kbp.

13. The method of claim 11 or claim 12, w’herein the method comprises phasing sequence reads within each overlapping section independently.

14. The method of any one of claims 11-13, wherein the method comprises analyzing sequence read phasing within an overlap of two consecutive overlapping sectionsand updating phasing of one or more sequence reads based on the analysis of sequence read phasing within the overlap.

15. The method of any one of claims 1-14, wherein the linkage information is determined based on location information of clusters of the fragments on the flow cell.

16. The method of claim 15, wherein the linkage information comprises information linking a first sequence read from a first cluster to a second sequence read from a second cluster based on a distance between the first and second clusters in the flow cell.

17. A method for detecting a phased sequence variant, comprising:phasing sequence reads according to the method of any one of claims 1-16; and detecting one or more phased sequence variants.

18. A computer-implemented method for phasing sequence reads from a genomic DNA sample, the method comprising:receiving sequence reads generated from fragments of the genomic DNA sample bound to a flow cell;determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell;mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes;analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; andphasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

19. A method of determining a haplotype nucleotide sequence, the method comprising:phasing sequence reads according to the method of any of claims 1-18; and assembling phased sequence reads into a first parental haplotype, thereby determining a haplotype nucleotide sequence.

20. A method of detecting a structural variant in a genomic DNA sample, the method comprising:phasing sequence reads according to the method of any of claims 1-18;generating one or more sequence assemblies based on the phased sequence reads; anddetecting a structural variant based on the one or more sequence assemblies.

21. The method of any of claims 1 -20, further comprising storing a set of sequence reads and associated phasing information in an electronic file.

22. The method of claim 21, wherein the electronic file comprises a BAM file, and wherein the BAM file comprises a tag that indicates whether a sequence read is phased to a first parental haplotype or a second parental haplotype, or a probability of phasing to a first parental haplotype or a second parental haplotype.

23. The method of any of claims 1-22, further comprising storing imputed genotypes in an electronic file,24. The method of claim 23, wherein the electronic file comprises a VCF file.

25. A system for phasing sequence reads from a genomic DNA sample, comprising one or more processors having instructions that when executed perform a method comprising:receiving sequence reads generated from fragments of the genomic DNA sample bound to a flow cell;determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell;mapping the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes;analyzing the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; andphasing the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

26. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to:receive sequence reads generated from fragments of the genomic DNA sample bound to a flow cell;determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell;map the sequence reads to a reference genome using the linkage information and nucleic acid sequence data from a set of candidate haplotypes;analyze the sequence reads to determine likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes; andphase the sequence reads based on the linkage information and on the likelihoods that the sequence reads correspond to a haplotype from the set of candidate haplotypes.

Citation Information

Patent Citations

  • Polymerase enzymes and reagents for enhanced nucleic acid sequencing

    US20080108082A1

  • Transposon end compositions and methods for modifying nucleic acids

    US20100120098A1

  • Determining a mudweight of drilling fluids for drilling through naturally fractured formations

    US62636004P0

  • Interactive techniques for speeding up homomorphic linear equation solving on encrypted data

    US62637000P0

  • Method And System For Touchless Gesture Detection And Hover And Touch Detection

    US62637002P0