Read mapping enhanced by spatial information
By leveraging spatial information from flowcell proximity, the method improves read mapping accuracy and reduces errors in nucleic acid sequencing, addressing the challenge of fragmented genomic sequence connectivity loss.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-02
AI Technical Summary
Existing nucleic acid sequencing methods struggle with accurate read mapping due to loss of connectivity and proximity information of fragmented genomic sequences, leading to incorrect alignments and errors in variant calling and gene expression studies.
Utilize spatial (flowcell proximity) information to enhance read mapping by determining a linking score bonus based on the proximity of polynucleotide clusters on a flow cell and genomic distance, updating mapping to improve accuracy.
Enhances mapping accuracy, reduces errors, and increases the number of correctly mapped reads, particularly in regions with genetic repeats and low-complexity sequences, improving variant calling and gene expression analysis.
Smart Images

Figure US2025047172_02042026_PF_FP_ABST
Abstract
Description
ILLINC.834WO / IP-2777-PCT PATENTREAD MAPPING ENHANCED BY SPATIAL INFORMATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 700,262, filed September 27, 2024, the content of which is incorporated by reference in its entirety.BACKGROUNDField of the Invention
[0002] This disclosure relates generally to sequencing nucleic acids, and more specifically to methods of using spatial (flowcell proximity) information to produce more accurate read mapping.Description of the Related Art
[0003] In recent years, biotechnology firms and research institutions have improved hardware and software for sequencing nucleotides and determining variant calls for genomic samples. For instance, some existing nucleobase sequencing platforms determine individual nucleobases within sequences from a genomic samples’ cells by using conventional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms can monitor millions to billions of nucleic acid polymers being synthesized in parallel to predict nucleobase calls from a larger base call dataset. For instance, a camera in many SBS platforms captures images of excited fluorescent tags incorporated into oligonucleotides for determining the nucleobase calls. After capturing such images, existing sequencing platforms send base call data (or image-based data) to a computing device to apply sequencing data analysis software that determines a nucleobase sequence for a genomic sample or other nucleic acid polymer. In some cases, such software maps and aligns nucleotide reads determined by the sequencing platform with a reference genome comprising at least a primary contiguous sequence. An example reference genome is the genome assembly GRCh38 from the National Center for Biotechnology Information at the National Library of Medicine. Basedon differences between the aligned nucleotide reads and the reference genome, existing data analysis software can further utilize a variant caller to identify genotype and / or variants within a genomic sample. These variants can include single nucleotide polymorphisms (SNPs), insertions or deletions (indels), or structural variants, among other genetic variabilities.
[0004] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods, use a shotgun approach to sequence large genomic DNA fragments, called template genomic sequences. Specifically, template genomic sequences are first fragmented in solution into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been read, knowledge of their connectivity and proximity to each other in the original template genomic sequence is lost. The process of ordering the sequence fragments to arrive at the sequence of the original template genomic sequence is generally referred to as “assembly.” Assembly processes can be computationally intensive and time-consuming. In addition, sequence and assembly errors can become a problem depending upon the sequencing methodology used and the quality of genomic DNA samples under evaluation.
[0005] In existing read mapping methods, mapping a read from the template nucleic acid molecule may use an alignment process that identifies the best match between the sequence read and the reference genome based on sequence similarity and homology between the two sequences. However, when nucleic acid samples are fragmented and sequenced in typical shotgun methods, information pertaining to the identity of which fragments came from which originating nucleic acid molecules can be lost. This origin information can be difficult or impossible to reconstruct using typical shotgun methods. For example, existing read mapping methods often struggle in regions where a sequence read may align with multiple sites on the reference genome due to genetic repeats in the reference genome or low-complexity sequences. Incorrect alignments can introduce errors into downstream analyses and interpretations, such as variant calling, structural variant detection, and gene expression studies. Therefore, there is a need for improved read mapping systems and methods.SUMMARY
[0006] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.
[0007] In one aspect, the disclosed technology provides a method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The method may comprise step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. The method may further comprise step (c) updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value. In some embodiments, the method further comprises updating the mapping of the second read pair. In some embodiments, updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the second candidate locations. In some embodiments, updating the mapping comprises updating the alignment scores associated with the second candidate locations. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair based on the updated mapping. In some embodiments, the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score. In some embodiments, updating themapping of the first read pair comprises realigning the first read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the first candidate locations. In some embodiments, updating the mapping comprises updating the alignment scores associated with the first candidate locations. In some embodiments, the method further comprises determining a select location on the reference sequence for the first read pair based on the updated mapping. In some embodiments, the select location on the reference sequence for the first read pair is associated with a maximal updated alignment score. In some embodiments, the linking score bonus is a Phred scaled score or is determined based on a linking quality statistical model. In some embodiments, combining the linking score bonus with the alignment scores associated with the first candidate locations comprises adding the linking score bonus to the alignment scores associated with the first candidate locations. In some embodiments, the method further comprises adding the updated mapping to a BAM file which stores the read pairs.
[0008] In another aspect, the disclosed technology provides a system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. The method may further comprise step (c) updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value. In some embodiments, the method further comprises updating the mapping of the second read pair. In some embodiments, updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the second candidate locations. In some embodiments, updating the mapping comprises updating the alignment scores associatedwith the second candidate locations. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair based on the updated mapping. In some embodiments, the selected location on the reference sequence for the second read pair is associated with a maximal updated alignment score. In some embodiments, updating the mapping of the first read pair comprises realigning the first read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the first candidate locations. In some embodiments, updating the mapping comprises updating the alignment scores associated with the first candidate locations.
[0009] In another aspect, the disclosed technology provides a computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising instructions that, when executed by a processor, cause the processor to perform a method comprising step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. The method may further comprise step (c) updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value. In some embodiments, the method further comprises updating the mapping of the second read pair. In some embodiments, updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the second candidate locations. In some embodiments, updating the mapping comprises updating the alignment scores associated with the second candidate locations. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair based on the updated mapping.
[0010] In another aspect, the disclosed technology provides a method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The method may comprise step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. The method may further comprise step (c) determining a select location on the reference sequence for the first read pair based on the linking score bonus. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair. In some embodiments, the selected location on the reference sequence for the second read pair is associated with a maximal updated alignment score. In some embodiments, the select location on the reference sequence for the second read pair is determined based on a function of the linking score bonus, the alignment scores associated with the first candidate locations, and the alignment scores associated with the second candidate locations. In some embodiments, the linking score bonus is a Phred-scaled score. In some embodiments, the linking score bonus is determined based on a linking quality statistical model. In some embodiments, the selected location on the reference sequence for the first read pair is associated with a maximal updated alignment score.
[0011] In another aspect, the disclosed technology provides a system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locationsand each of the second candidate locations. The method may further comprise step (c) determining a select location on the reference sequence for the first read pair based on the linking score bonus. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair. In some embodiments, the selected location on the reference sequence for the second read pair is associated with a maximal updated alignment score. In some embodiments, the select location on the reference sequence for the second read pair is determined based on a function of the linking score bonus, the alignment scores associated with the first candidate locations, and the alignment scores associated with the second candidate locations. In some embodiments, the linking score bonus is a Phred scaled score. In some embodiments, the linking score bonus is determined based on a linking quality statistical model.
[0012] In another aspect, the disclosed technology provides a computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising instructions that, when executed by a processor, cause the processor to perform a method comprising step (a) mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. The method may further comprise step (b) determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. The method may further comprise step (c) determining a select location on the reference sequence for the first read pair based on the linking score bonus. In some embodiments, the method further comprises determining a select location on the reference sequence for the second read pair. In some embodiments, the selected location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
[0013] Features which are described in the context of separate aspects and embodiments of the invention may be used together and / or be interchangeable. Similarly, features described in the context of a single embodiment may also be provided separately or in any suitable sub-combination.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein).
[0015] FIG. 1 is a flowchart illustrating a method for enhanced mapping according to some embodiments of the disclosed technology.
[0016] FIG. 2 is a flowchart illustrating an alternative method for enhanced mapping according to some embodiments of the disclosed technology.
[0017] FIG. 3 A is a schematic diagram that illustrates the spatial relationship of a flow cell and target polynucleotides on the flow cell.
[0018] FIG. 3B is a schematic diagram that illustrates the spatial relationship of the flow cell and the polynucleotide clusters derived from the target polynucleotides.
[0019] FIG. 3C is a schematic diagram that illustrates the mapping of read pairs on a reference genome.
[0020] FIG. 3D illustrates part of an exemplary process of computing the linking score bonus.
[0021] FIG. 4 illustrates an environment in which a mapping system can operate in accordance with one or more embodiments of the present disclosure.
[0022] FIG. 5 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure.DETAILED DESCRIPTION
[0023] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.Overview
[0024] One embodiment is a system and method for determining the strength of a link between two or more read pairs on a flow cell. The “link” as discussed herein represents the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long DNA, such as genomic DNA, from a biological source are sheared to create shorter fragments which can be sequenced in a single read. The shearing process can create these shorter fragments which bind to the flow cell. In some embodiments, a read pair is obtained from each polynucleotide cluster (which may contain thousands of polynucleotide molecules that are generated from one fragment, such as by bridge amplification, and thus have the same sequence) on the flow-cell. A sequence read generated by a short-read NGS system may be about 100 to 500 nucleotides long. The spatial location of each fragment on the flowcell has been found to relate to the original nucleic acid molecule from which the fragment was derived. Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U.S. Provisional Patent Application Nos. 63 / 600460 and 63 / 614066, the disclosure of each of which is incorporated herein by reference in its entirety.
[0025] For example, fragments which come from the same nucleic acid molecule have been found to be located in closer proximity to each other on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of polynucleotides on a flow cell are close together spatially, meaning they are in close proximity, and also their corresponding sequencing results (e.g., read-pairs) are found to map to locations close together on the genome, the seeds for these clusters (i.e., nucleic acid fragments from the original sample that are amplified to form these clusters) are more likely to have physically come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell in proximity to each other, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors maychange based on a variety of experimental conditions. Embodiments of the disclosure provide methods for improving mapping utilizing linking information.
[0026] Embodiments of the invention relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). As described above, according to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in spatially nearby positions as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In regular library preps with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell, spatial information is retained. This spatial (flowcell proximity) information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.
[0027] For example, one embodiment is a method for sequencing nucleic acid molecules by fragmenting them on a flow cell using bound transposome complexes. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences, such that the transposome can be used to fragment the nucleic acid molecules and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the flow cell. The method can include contacting the transposome complexes with the nucleic acid molecules under conditions to fragment them and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. (San Diego, CA).
[0028] Once the fragments have been bound to a substrate, the bound fragments can be amplified to form a plurality of polynucleotide clusters on the substrate. The location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster. Once the nucleotide sequence of each cluster has been read, the systemor method can map those reads to a reference sequence (e.g., reference genome sequence) to determine the original target polynucleotide from which the read originated. In some embodiments, the mapping process may consider the spatial location of each cluster on the flow cell, such that clusters which are closer in proximity to each other on the flow cell are more likely to have originated from the same target polynucleotide. Clusters that are in a close proximity to one another may also have a small relative displacement with respect to one another. By considering spatial (flowcell proximity or relative displacement) information associated with each read or read pair when mapping the sequenced fragments to the reference sequence, the method may perform more accurate mapping operations as compared to methods that do not take the spatial location of each cluster into account during the mapping process. Therefore, spatial information that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the mapping quality of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. However, improving the confidence of read pair’s alignment based on linking information may allow these previously discarded reads to be retained and disambiguated as well as improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling.
[0029] Understanding that fragments located spatially near each other on a flow cell are more likely to have derived from the same original nucleic acid molecule allows flow cell location to be used to inform the determination of whether two fragments can be assigned to a particular location on the genome. If the mapping of the reads on the reference genome is clear, links can be evaluated as a combination of genomic proximity and flow cell proximity. Conversely, knowledge of links between neighboring fragments can also clarify ambiguous read mapping. For example, when a first fragment aligns uniquely to region A and a second, linked fragment aligns ambiguously to regions A or B, it is more likely that the second fragment aligns to region A since it was linked to a fragment known to derive from region A. The knowledge of links between fragments on the flow cell can assist in mapping of fragment reads to the genome, but links themselves can also be discovered / evaluated based in part on mapping. Therefore, embodiments relate to systems and methods for enhancing the mapping capabilities of such linked reads. In some embodiments, the link between different clusters of polynucleotide fragments on the flow cell is given a linking score that quantifies the strengthof the association between the fragments. For example, the linking score may comprise a Phred scaled score (e.g., lOLogio) of the link probability, which is the likelihood that two read pairs’ associated clusters are nearby on the flow cell due to linking instead of by chance, and the link probability can be determined from a linking quality statistical model that has been fitted to experimental data. Additional details regarding the construction and fitting of the linking quality statistical model can be found in U.S. Provisional Patent Application No. 63 / 700,049 filed on September 27, 2024, the disclosure of which is incorporated herein by reference in its entirety. Fragments which are more likely to be derived from the same original nucleic acid molecule are given a higher linking score and fragments which are less likely to be from the same original nucleic acid molecule are given a lower linking score. In some embodiments, a “linking score bonus” may be determined based on the genomic distance on the original nucleic acid molecule between two fragments and the relative proximity on the flow cell of the two polynucleotide clusters for the two fragments. The true genomic distance between the two fragments may be unknown, but an estimated value (or several estimated values) based on preliminary mapping may be used. Thus, for example, the knowledge of links may allow wrong alignments to be corrected, correct alignments to have their MAPQ increased, and questionable / uncertain alignments to have their MAPQ improved by adding the linking score bonus to their preliminary alignment scores.
[0030] One embodiment is a method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. In this embodiment, a first read pair and a second read pair are preliminarily mapped to a reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. Once this preliminary mapping information has been determined, the system may determine a linking score bonus between the first read pair and the second read pair (for each combination of candidate alignment locations of the first read pair and the second read pair) based on the flow cell proximity on the flow cell between the polynucleotide cluster for the first read pair and the polynucleotide cluster for the second read pair along with an estimated genomic distance between each of the first candidate locations and each of the second candidate locations. With this linking information determined, the system may update themapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value.
[0031] Another embodiment is a system or method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. In this embodiment, the method includes mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. Once these first candidate and second candidate locations are known, the system or method determines a linking score bonus between the first read pair and the second read pair based on the flow cell proximity between the location of the clusters for the first read pair and the second read pair on the flow cell and the estimated genomic distance between each of the first candidate locations and each of the second candidate locations. Then, a select location on the reference sequence for the first read pair may be determined based in part on the linking score bonus.
[0032] These embodiments increase mapping accuracy, can assist in the discovery / evaluation of links, and may be compute-efficient by not greatly increasing the cost of a standard read mapping processes. The disclosed technology provides such an enhanced, accurate mapping method that is also useful for link discovery / evaluation. The disclosed technology may also greatly reduce the accuracy gap between short-read and long-read sequencing. In addition, embodiments may help increase the number of reads which can be processed by taking reads which would otherwise have had a low or no ability to be mapped to a reference genome and determine the proper mapping for such reads.Definitions
[0033] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. The use of the term “including” as well as other forms, such as “include”, “includes,” and “included,” is not limiting. The use of the term “having” as well as other forms, such as “have”, “has,” and “had,” is not limiting. As used in this specification, whether in a transitional phrase or in the body of the claim, the terms “comprise(s)” and “comprising” are to be interpreted as having an open-ended meaning. That is, the above terms are to be interpreted synonymously with thephrases “having at least” or “including at least.” For example, when used in the context of a process, the term “comprising” means that the process includes at least the recited steps, but may include additional steps. When used in the context of a compound, composition, or device, the term “comprising” means that the compound, composition, or device includes at least the recited features or components, but may also include additional features or components.
[0034] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.
[0035] As used herein, a “nucleotide” includes a nitrogen containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is a ribose, and in deoxyribonucleotides (DNA), the sugar is a deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2' position in ribose. The nitrogen containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-l atom of deoxyribose is bonded to N-l of a pyrimidine or N-9 of a purine. The phosphate groups may be in the mono- , di-, or tri-phosphate form. These nucleotides may be natural nucleotides, but it is to be further understood that non-natural nucleotides, modified nucleotides or analogs of the aforementioned nucleotides can also be used.
[0036] As used herein, “nucleobase” is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can be naturally occurring or synthetic. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted at the 8 position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6- diaminopurine, N6-ethano-2,6-diaminopurine, 5- methylcytosine, 5-(C3-C6)- alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510 and PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (“Practical Handbook of Biochemistry and Molecular Biology”, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all herein incorporated by reference in their entireties.
[0037] The terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecules” are used interchangeably herein and refer to a covalently linked sequence of nucleotides of any length (i.e., ribonucleotides for RNA, deoxyribonucleotides for DNA, analogs thereof, or mixtures thereof) in which the 3 ’ position of the pentose of one nucleotide is joined by a phosphodiester group to the 5’ position of the pentose of the next. The terms should be understood to include, as equivalents, analogs of either DNA, RNA, cDNA, or antibody-oligo conjugates made from nucleotide analogs and to be applicable to single stranded (such as sense or antisense) and double stranded polynucleotides. The term as used herein also encompasses cDNA, which is complementary or copy DNA produced from an RNA template, for example by the action of reverse transcriptase. This term refers only to the primary structure of the molecule. Thus, the term includes, without limitation, triple-, double- and single-stranded deoxyribonucleic acid (“DNA”), as well as triple-, double- and singlestranded ribonucleic acid (“RNA”). The nucleotides include sequences of any form of nucleic acid.
[0038] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locusinherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.
[0039] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series.
[0040] The term “primer,” as used herein refers to an isolated oligonucleotide that is capable of acting as a point of initiation of synthesis when placed under conditions inductive to synthesis of an extension product (e.g., the conditions include nucleotides, an inducing agent such as DNA polymerase, and a suitable temperature and pH). The primer is preferably single stranded for maximum efficiency in amplification, but may alternatively be double stranded. If double stranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to prime the synthesis of extension products in the presence of the inducing agent. The exact lengths of the primers will depend on many factors, including temperature, source of primer, use of the method, and the parameters used for primer design.
[0041] As used herein the term “chromosome” refers to the heredity-bearing gene carrier of a living cell, which is derived from chromatin strands comprising DNA and proteincomponents (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein.
[0042] A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences.
[0043] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. In some embodiments, the term “reference genome” refers to a digital nucleic acid sequence assembled as a representative example (or representative examples) of genes and other genetic sequences of an organism. Regardless of the sequence length, in some cases, a reference genome represents an example set of genes or a set of nucleic acid sequences in a digital nucleic acid sequenced determined by scientists as representative of an organism of a particular species. For example, a linear human reference genome may be GRCh38 or other versions of reference genomes from the Genome Reference Consortium. As noted above, in some cases, a reference genome includes multi-base codes. As a further example, a reference genome may include a graph reference genome that includes both a linear reference genome and paths representing nucleic acid sequences from ancestral haplotypes, such as Illumina DRAGEN Graph Reference Genome hgl9.
[0044] For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. For example, the reference sequence can be a reference human genome sequence, such as hgl9 or hg38. In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hgl9. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.
[0045] The term “nucleic acid sample” herein refers to a sample, typically derived from a biological fluid, cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments the nucleic acid sample comprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (e.g., surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (e.g., patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (e.g., namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein.
[0046] As used herein, the term “cluster” or “clump” refers to a group of molecules, e.g., a group of DNA, or a group of signals. In some embodiments, the signals of a cluster are derived from different features. In some embodiments, a signal clump represents a physical region covered by one amplified oligonucleotide. Each signal clump could be ideally observed as several signals. Accordingly, duplicate signals could be detected from the same clump of signals. In some embodiments, a cluster or clump of signals can comprise one or more signalsor spots that correspond to a particular feature. When used in connection with microarray devices or other molecular analytical devices, a cluster can comprise one or more signals that together occupy the physical region occupied by an amplified oligonucleotide (or other polynucleotide or polypeptide with a same or similar sequence). For example, where a feature is an amplified oligonucleotide, a cluster can be the physical region covered by one amplified oligonucleotide. In other embodiments, a cluster or clump of signals need not strictly correspond to a feature. For example, spurious noise signals may be included in a signal cluster but not necessarily be within the feature area. For example, a cluster of signals from four cycles of a sequencing reaction could comprise at least four signals. For example, the term “cluster of oligonucleotides” (or “oligonucleotide cluster”, “cluster of polynucleotides”, “polynucleotide cluster”, “cluster of nucleic acids”, or “nucleic acid cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-patterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.
[0047] The term “next generation sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing-by- synthesis using reversible dye terminators, and sequencing-by-ligation.
[0048] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriateto determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region, e.g., that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (e.g., 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. A sequence read may be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0049] In some embodiments, the term “read” refers to an inferred or predicted sequence of one or more nucleotide bases (or nucleobase pairs) from all or part of a sample genomic sequence (e.g., a sample genomic sequence, complementary DNA). Such a sample nucleotide sequence may take the form of a sample genomic sequence from genomic DNA (gDNA), a transcriptomic sequence from complementary DNA (cDNA), a transcriptomic sequence from RNA, or other nucleotide sequence. In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a polynucleotide fragment (or group of monoclonal polynucleotide fragments) from a sequencing library corresponding to a genomic sample. For example, in some embodiments, a sequencing device determines a nucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a well in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g., assembled nucleotide reads) satisfying athreshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.
[0050] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood that the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood that the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.
[0051] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0052] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP,SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.
[0053] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.
[0054] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.
[0055] As further used herein, the term “genomic coordinate” (or sometimes simply “coordinate”) refers to a particular location or position of a nucleobase within a genome (e.g., an organism’s genome or a reference genome). In some cases, a genomic coordinate includes an identifier for a particular chromosome of a genome and an identifier for a position of a nucleobase within the particular chromosome. For instance, a genomic coordinate or coordinates may include a number, name, or other identifier for a somatic or sex chromosome (e.g., chrl or chrX) and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl : 1234570 or chrl : 1234570-1234870). In some cases, a genomic coordinate refers to a genomic coordinate on a sex chromosome (e.g., chrX or chrY). Consequently, the mapping system can determine genotype probabilities for a genotype call (e.g., a variant call) for a genomic coordinate on a sex chromosome. Further, in certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mitochondrial DNA reference genome or SARS-CoV-2 for a reference genome for the SARS- CoV-2 virus) and a position of a nucleobase within the source for the reference genome (e.g., mt: 16568 or SARS-CoV-2:29001). By contrast, in certain cases, a genomic coordinate refers to a position of a nucleobase within a reference genome without reference to a chromosome or source (e.g., 29727).
[0056] As used herein, a “genomic region” refers to a range of genomic coordinates. Like genomic coordinates, in certain implementations, a genomic region may be identified by an identifier for a chromosome and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl: 1234570-1234870). In various implementations, a genomic coordinate includes a position within a referencegenome. In some cases, a genomic coordinate is specific to a particular reference genome. Relatedly, as used herein, the term “reference span” refers to a span of nucleobase positions within a linear reference genome. In other words, a reference span includes a span of nucleobases between two respective genomic coordinates of the linear reference genome.
[0057] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith- Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.
[0058] Relatedly, as used herein, the term “mapping quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In some embodiments, for example, a mapping-quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong}, rounded to the nearest integer. In the alternative to a mean or median mapping quality, in some implementations, a mapping quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.
[0059] As further used herein, the term “genotype call” refers to a determination or prediction of a particular genotype of a genomic sample or a sample nucleotide sequence at a genomic locus. In particular, a genotype call can include a prediction of a particular genotype of a genomic sample with respect to a reference genome or a reference sequence at a genomic coordinate or a genomic region. For instance, in some cases, a genotype call includes a determination or prediction that a genomic sample comprises both a nucleobase and a complementary nucleobase at a genomic coordinate that is either homozygous or heterozygous for a reference base or a variant (e.g., homozygous reference bases represented as 0|0 orheterozygous for a variant on a particular strand represented as 0| 1). Accordingly, a genotype call can include a prediction of a variant or reference base for one or more alleles of a genomic sample and indicate zygosity with respect to a variant or reference base. A genotype call is often determined for a genomic coordinate or genomic region at which an SNP, insertion, deletion.
[0060] As used herein, the term “variant” refers to a nucleobase or multiple nucleobases that do not align with, differ from, or vary from a corresponding nucleobase (or nucleobases) in a reference sequence or a reference genome. For example, a variant includes a SNP, an indel, or a structural variant that indicates nucleobases in a sample nucleotide sequence that differ from nucleobases in corresponding genomic coordinates of a reference sequence.
[0061] Along these lines, a “variant call” (or “variant nucleobase call”) refers to a nucleobase call comprising a mutation or a variant at a particular genomic coordinate or genomic region with respect to a reference. In particular, a variant call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate or region that differs from a reference nucleobase (or sequence of reference nucleobases) at the same genomic coordinate or region within a reference genome. Conversely, a “non-variant call” (or “non-variant nucleobase call” or “reference call”) refers to a nucleobase call comprising a non-variant or a reference nucleobase at a genomic coordinate or a genomic region with respect to a reference. In particular, a non- variant or reference call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate or region that matches a reference nucleobase (or sequence of reference nucleobases) at the same genomic coordinate or region within a reference genome.
[0062] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquid. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g. due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics,polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solid supports for some embodiments are located within a flow cell apparatus. Exemplary flow cells are set forth in further detail below.
[0063] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.
[0064] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1mm2, 500 pm2, 100 pm2, 25 pm2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.
[0065] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially flat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.
[0066] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.
[0067] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element- containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, amixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.
[0068] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-type or mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, nonnatural bases, modified backbone, and can comprise nicks in one or both strands.
[0069] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads in ambiguous positions.
[0070] In one or more embodiments, the mapping system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads,alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.
[0071] Moreover, in one or more embodiments, one or more sequencing data files in which the mapping system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.Embodiments of an Enhanced Mapping System
[0072] Aspects of the disclosed systems and methods are useful for enhancing the accuracy and efficiency of mapping reads on a flow cell to their correct position in a genomic sequence. These read mapping systems incorporate spatial (flowcell proximity) information associated with each read or read pair to produce more accurate alignments with a reference genome and more accurate measures of mapping quality of each read (“MAPQs”). In some embodiments, the disclosed enhanced mapping systems and methods may map a batch of reads from a flow-cell region (e.g., a rectangular region, such as a tile) using a standard mapping process (without using any flowcell proximity information or link information) on a first pass, and then the mapping system may perform a second pass to search for possible spatial links among read pairs in the batch. The system may then determine linking score bonuses for candidate alignments between the reads and the reference genome that are consistent with these links. This process of using a linking score bonus was found to result in the remapping and correction of approximately 10% of the read pairs in each batch (for example, the remapping is only triggered for those read pairs where the MAPQ is low enough such that adding a linking score bonus to increase the mapping quality would improve the overall mapping, or where competing alignment scores are close enough to that of the most likely alignments such that the linking score bonus could cause a change in the winning alignment). Thus, it was discovered that applying the linking score bonuses resulted in a change to the mapping resultswhich improved the accuracy of mapping reads within the batch to their proper genomic locations.
[0073] In this aspect, the enhanced mapping system may first process the “source” and “destination” read pairs (i.e., the first read- pair and the second read-pair in a pair of readpairs) from a rectangular flow cell area as a batch in a first mapping pass to determine any links between those source and destination read pairs. Read pairs within a nested set of flow cell rectangles will have link information determined for all read pairs within the batch. Thus, each link between each read pair will be known. However, read pairs near the outer edges of the rectangular flow cell area may be processed in another overlapping batch process. After the first mapping pass, a defined flow cell neighborhood, or Euclidean distance, around each “source” read pair may be searched for “destination” read pairs with possible links suggested by their genomic proximity. Next, the genomic positions of two spatially proximate read pairs may be compared. When comparing the genomic positions of two spatially proximate reads, not only are the read pair candidates associated with the highest alignment scores of these read pairs considered, but all plausible alignment candidates for each read pair may be considered.
[0074] The disclosed methods determine a possible link to be when a “source read pair” alignment candidate and a “destination read pair” alignment candidate are found to be in the same chromosome, and when the distance between their positions in the chromosome is no more than a predetermined maximal genomic link distance. For each possible link between a source read pair alignment candidate and a destination read pair alignment candidate, a linking score bonus applied to the source read pair alignment candidate may be determined. This determination may be based on priors and likelihoods for same-molecule and different- molecule explanations, as mentioned below. For example, the linking score bonus may comprise the link probability LinkProb(A)(d) measured in the Phred scale (e.g., lOLogio), where the link probability LinkProb(A)(d) is the likelihood that two particular read pairs are nearby due to linking (i.e., coming from the same nucleic acid molecule) instead of by chance (i.e., coming from different nucleic acid molecules). The link probability LinkProb(A)(d) may be a function of A, which is the flow cell proximity of the two read pairs on the flow cell, and d, which is the genomic distance of the two read pairs on the originating nucleic acid molecule. The genomic distance d may be obtained by preliminarily determining the alignments or mapping of the two read pairs to a reference genome. For example, the genomic distance dmay be estimated as the distance between the alignment candidates for the two read pairs in the reference genome. The link probability LinkProb(A)(d) may be described by a linking quality statistical model and fitted to experimental data. In some embodiments, the determination may further consider how much worse any destination read pair alignment candidate may have scored than the destination read pair’s winning alignment (e.g., the alignment candidate associated with the highest alignment score) and may use this difference to adjust the linking score bonus accordingly. The result of this determination may include positive linking score bonuses which become applied to a small subset of all the alignment candidates for each read pair in the batch. Once this information is determined, a determination may be made of which read pairs require a second mapping pass to incorporate the determined linking score bonuses. With the use of a linking score bonus, only a small subset of read pairs (approximately 10% of the read pairs in each batch) are remapped, thereby increasing the mapping accuracy while preserving the computational efficiency for the entire sequencing run.
[0075] In some cases, adding a linking score bonus to the already-determined alignment scores may only trigger a second mapping pass if the mapping quality is below a threshold (e.g., 3, 10, 33, or any value therebetween) such that increasing the mapping quality would be helpful to improve the overall mapping. In some cases, adding a linking score bonus to competing alignment candidates (e.g., the alignment candidate associated with sufficiently high alignment scores) may only trigger a second mapping pass if the competing alignment scores are close enough to that of the most likely alignments, such that the linking score bonus could cause a relevant MAPQ improvement or an outright change in the winning alignment (e.g., the alignment candidate associated with the highest alignment score). In some cases, a second mapping pass may be triggered when the winning alignment is clipped (i.e., having nucleobases that are in the sequencing read but not in the reference sequence), and another clipped alignment candidate that could plausibly form a split-read alignment with the winning alignment gets a linking score bonus. In some embodiments, only about 10% of read pairs within any particular batch may trigger a second mapping pass.
[0076] In some embodiments, each read pair which is subject to a second mapping pass gets returned to the top of a map / align pipeline (which may be part of the mapping system 206 of FIG. 4), along with a list of linking score bonuses applied to specific alignment candidates. In some embodiments, seed mapping and alignment scoring may be repeated,generating the same set of alignment candidates and associated alignment scores as those in the first pass, to which the bonuses are applied. Because of these linking score bonuses, different winning alignments may be selected, or different MAPQs may be determined. In some embodiments, second-pass mapping results may be returned to the mapping system, and may replace the batch’s first-pass mapping results for the affected read pairs. In some embodiments, when the second mapping pass completes, updated batch mapping results may be considered final and used in a downstream analysis (e.g., variant calling). In alternative embodiments, seed mapping and alignment scoring may not be repeated in a second pass, but rather first-pass alignment scoring results may be stored before link analysis and recalled afterward, and then the linking score bonuses may be applied to these first-pass alignment scoring results.
[0077] In one aspect, the disclosed technology provides an enhanced method 400 of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, as shown in FIG. 1. The method 400 involves a remapping step, i.e., a second mapping pass. The method 400 may start at a start step and then move to a step 401, which involves mapping a first read pair and a second read pair taken from a flow cell to a reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair. In some embodiments, mapping the first read pair and the second read pair to the reference sequence comprises determining the alignments of the first read pair and the second read pair to the reference sequence using a local alignment method.
[0078] After step 401, the method 400 may move to step 402, where a linking score bonus between the first read pair and the second read pair is determined based on the flow cell proximity between the clusters on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. In some embodiments, the linking score bonus is a Phred scaled score. In some embodiments, the linking score bonus is determined based on a linking quality statistical model. The linking quality statistical model may be constructed to describe the link probability, which is the probability that, for two read pairs derived from the same nucleic acid fragment (e.g., a genomic fragment) at two regions separated by a given genomic distance on that nucleic acid fragment, two clusters on the flow cell associated with the two read pairs would be separated by a flow cell proximity (e.g., avector separation) on the flow cell. In some embodiments, the linking quality statistical model is fitted (i.e., the parameters in the model are quantified) to data generated from each sequencing run. In some embodiments, determining the linking score bonus comprises determining that the flow cell proximity between the clusters for the first read pair and the second read pair on the flow cell is below a predetermined proximity value. In some embodiments, determining the flow cell proximity between the first read pair cluster and the second read pair cluster on the flow cell comprises determining a radius around the cluster for the first read pair. In some embodiments, determining the flow cell proximity between the clusters for the first read pair and the second read pair on the flow cell comprises determining a sliding window around the cluster for the first read pair. In some embodiments, a linking quality is used to determine the shape of the sliding window around the cluster for the first read pair. For example, when the linking quality is below a certain threshold, a radius may be used to determine whether the first and second clusters are linked. However, when the linking quality is above a certain threshold the method may use a sliding window having other geometric shapes (such as a square, a rectangle, or a hexagon) to determine when the fragments at the first and second clusters are linked. In some embodiments, the use of sliding window may result in redundant processing of overlapped flow cell areas, but the overall mapping speed may only be reduced by about 10% to 20%. In some embodiments, the flow cell proximity is measured in flow cell units. In different sequencing instruments, one flow cell unit may correspond to about 1 micron, 20 micron, 40 micron, 60 micron, 80 micron, 100 micron, 200 micron, or any value therebetween, or may be defined as the distance between sequencing wells on the flow cell. In some embodiments, the analyses related to linking score bonus may be performed on a field programmable gate array (FPGA), so that significant CPU costs and slowdown may be avoided.
[0079] After step 402, the method 400 may move to step 403, wherein the mapping of the first read pair is updated (i.e., the first read pair is remapped), if a probability that altering the mapping will increase the mapping accuracy is above a predetermined value (e.g., 50%, 70%, 90%, 95%, 99%, or any value therebetween). The probability can be determined based on the determined linking score bonus. In some embodiments, step 403 may include a sub-step 404 that involves updating the mapping of the second read pair (i.e., the second read pair is remapped). In some embodiments, updating the mapping of the second read pair comprisesrealigning the second read pair to the reference sequence using a local alignment method (such as a method based on the Smith-Waterman algorithm). In some embodiments, updating the mapping comprises updating the second candidate locations in the computer memory (e.g., in the BAM file). In some embodiments, updating the mapping comprises updating the alignment scores associated with the second candidate locations in the computer memory (e.g., in the BAM file). In some embodiments, updating the mapping of the first read pair comprises realigning the first read pair to the reference sequence using a local alignment method. In some embodiments, updating the mapping comprises updating the first candidate locations in the computer memory (e.g., in the BAM file). In some embodiments, updating the mapping comprises updating the alignment scores associated with the first candidate locations in the computer memory (e.g., in the BAM file). In some embodiments, the mapping is only updated if a mapping quality score for the first read pair is lower than a predetermined threshold. In some embodiments, the probability is determined based on a combination of the linking score bonus and the alignment scores associated with the first candidate locations. In some embodiments, combining the linking score bonus with the alignment scores associated with the first candidate locations comprises adding the linking score bonus to the alignment scores associated with the first candidate locations. In some embodiments, the probability is further determined based on the alignment scores associated with the second candidate locations. In some embodiments, the overall mapping speed may only be reduced by about 10% to 20% even with the incorporation of the remapping steps 403 and / or 404. In some embodiments, even though the method 400 utilizes remapping steps 403 and / or 404 internally, it acts similar to a single-pass mapping process when observed from the outside.
[0080] After step 403, the method 400 may move to step 405, which involves determining a select location on the reference sequence for the first read pair based on the updated mapping. In some embodiments, step 405 may include a sub-step 406 that involves determining a select location on the reference sequence for the second read pair based on the updated mapping. In some embodiments, the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score. In some embodiments, the select location on the reference sequence for the first read pair is associated with a maximal updated alignment score.
[0081] After step 405, the method 400 may move to step 408, which involves adding the updated mapping to a BAM file which stores the read pairs, or to step 410, which involves adding the linking score bonus and / or information regarding each of the possible links between the first read pair and the second read pair (for example, there may be a linking score bonus between each candidate alignment of the first read pair and each candidate alignment of the second read pair) to a BAM file which stores the read pairs. The method 400 then ends at an end step.
[0082] In another aspect, the disclosed technology provides a method of improving accuracy of split-alignments of read pairs to a reference sequence. Linking information can improve the sensitivity of split-alignment of reads by providing clearer evidence that one portion of the pair-end read came from a particular genomic region, which is distant from another genomic region that the rest of the pair-end read came from. In some embodiments, a split-aligned read pair is defined as a read pair that comprises different segments that align to disjoint regions in the reference sequence. Thus, to improve the alignment accuracy for the split-aligned read pair, the method may provide an updated mapping of a first segment of the read pair using the method 400 described above. In some embodiments, split-alignments are selected by scoring various combinations of candidate partial alignments of a read, such as by summing local alignment scores of the candidate partial alignments participating in a candidate split alignment and subtracting break penalties for genomic jumps between the partial alignments. A linking score bonus which is found to apply to a candidate partial alignment, especially a template-outermost partial alignment such as a 5’ partial alignment, may then contribute to the score sum for the split alignment. In another aspect, the disclosed technology provides a system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The system may include a memory storing instructions and a processor that, when executing the instructions, is configured to perform the method 400 described above. In another aspect, the disclosed technology provides a computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The computer program may include instructions that, when executed by a processor, cause the processor to perform the method 400 described above.
[0083] In another aspect, the disclosed technology provides an enhanced method 500 of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a referencesequence, as shown in FIG. 2. The method 500 does not involve a remapping step, i.e., a second mapping pass. The method 500 may start at a start step and then move to a step 501, wherein a first read pair and a second read pair is mapped to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for a first read pair and a plurality of second candidate locations and associated alignment scores for a second read pair. In some embodiments, mapping the first read pair and the second read pair to the reference sequence comprises determining the alignments of the first read pair and the second read pair to the reference sequence using a local alignment method.
[0084] After step 501, the method 500 may move to step 502, wherein a linking score bonus is determined between the first read pair and the second read pair based on the flow cell proximity between the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations. Similar to the process mentioned in Fig. 1, the linking score bonus may be a Phred scaled score, determined based on a linking quality statistical model. The linking quality statistical model may be constructed to describe the link probability, which is the probability that, for two read pairs derived from the same nucleic acid fragment (e.g., a genomic fragment) at two regions separated by a given genomic distance on that nucleic acid fragment, two clusters on the flow cell associated with the two read pairs would be separated by a flow cell proximity (e.g., a vector separation) on the flow cell. In some embodiments, the linking quality statistical model is fitted (i.e., the parameters in the model are quantified) to data generated from each sequencing run. In some embodiments, computing the linking score bonus comprises determining that the flow cell proximity between the clusters for the first read pair and the second read pair on the flow cell is below a predetermined proximity value. In some embodiments, the analyses related to linking score bonus may be performed on a field programmable gate array (FPGA), so that significant CPU costs and slowdown may be avoided.
[0085] After step 502, the method 500 may move to step 503, wherein a select location on the reference sequence is determined for the first read pair based on the linking score bonus. In some embodiments, step 503 may include a sub-step 504 that involves determining a select location on the reference sequence for the second read pair. In some embodiments, the select location on the reference sequence for the second read pair isassociated with a maximal updated alignment score. In some embodiments, the select location on the reference sequence for the second read pair is determined based on a function of the linking score bonus, the alignment scores associated with the first candidate locations, and the alignment scores associated with the second candidate locations. In some embodiments, the select location on the reference sequence for the first read pair is associated with a maximal updated alignment score. In some embodiments, the select location on the reference sequence for the first read pair is determined based on a combination of the linking score bonus and the alignment scores associated with the first candidate locations. In some embodiments, combining the linking score bonus with the alignment scores associated with the first candidate locations comprises adding the linking score bonus to the alignment scores associated with the first candidate locations. In some embodiments, the select location on the reference sequence for the first read pair is further determined based on the alignment scores associated with the second candidate locations.
[0086] After step 503, the method 500 may move to step 505, wherein the select location on the reference sequence for the first read pair and / or the select location on the reference sequence for the second read pair is added to a BAM file which stores the read pairs. Alternatively, the method 500 may move to a step 506, wherein the linking score bonus and / or information regarding each of the possible links between the first read pair and the second read pair (for example, there may be a linking score bonus between each candidate alignment of the first read pair and each candidate alignment of the second read pair) is added to a BAM file which stores the read pairs.
[0087] In another aspect, the disclosed technology provides a method of improving accuracy of split-alignments of read pairs to a reference sequence. Linking information can improve the sensitivity of split-alignment of reads by providing clearer evidence that one portion of the pair-end read came from a particular genomic region, which is distant from another genomic region that the rest of the pair-end read came from. In some embodiments, a split-aligned read pair is defined as a read pair that comprises different segments that align to disjoint regions in the reference sequence. Thus, to improve the alignment accuracy for the split-aligned read pair, the method may determine a select location on the reference sequence for a first segment of the read pair using the method 500 described above. In some embodiments, split-alignments are selected by scoring various combinations of candidatepartial alignments of a read, such as by summing local alignment scores of the candidate partial alignments participating in a candidate split alignment and subtracting break penalties for genomic jumps between the partial alignments. A linking score bonus which is found to apply to a candidate partial alignment, especially a template-outermost partial alignment such as a 5’ partial alignment, may then contribute to the score sum for the split alignment. In another aspect, the disclosed technology provides a system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The system may include a memory storing instructions and a processor that, when executing the instructions, is configured to perform the method 500 described above. In another aspect, the disclosed technology provides a computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence. The computer program may include instructions that, when executed by a processor, cause the processor to perform the method 500 described above.
[0088] In some embodiments, the link information can improve mapping accuracy dramatically, such as when one read pair aligns uniquely to region A, while a second read pair nearby on the flow cell aligns ambiguously to either region A or B, the most likely hypothesis consistent with these observables is that these two read pairs both came from the same nucleic acid molecule, i.e., from genomic region A. Therefore, the second read pair can be determined to map to region A, i.e., the chance is smaller that separate nucleic acid molecules from regions A and B randomly appeared so close together on the flow cell. Thus, link information can disambiguate mapping among several possible alignment positions, such as in the case of segmental duplicate copies.
[0089] In some embodiments, a sequencing system may be used to determine the sequence and location of read pairs in each cluster on the flow cell. The system may create a BAM file entry for each such read pair which includes the sequence of the read pair and identification of one or more possible target polynucleotides. For any BAM file which includes more than one possible target polynucleotide, the system may attempt to disambiguate the read pairs by referring to the spatial location (flowcell proximity) information contained with each read pair. As mentioned above, if some read pairs are very near one another on the flow cell (for example, within 10 flow cell units, 50 flow cell units, 100 flow cell units, 500 flow cell units, or any value therebetween, where one flow cell unit may correspond to about 1 micron,20 micron, 40 micron, 60 micron, 80 micron, 100 micron, 200 micron, or any value therebetween, in different sequencing instruments), then the mapping system 206 as shown in FIG. 4 may determine that one read pair is more likely to have originated from a particular target polynucleotide than another read pair. If a particular read pair can be disambiguated, then the BAM file entry for that read pair is updated to indicate only the correct target polynucleotide.
[0090] In some embodiments, in order to manage the spatial location of each cluster during the mapping process, each flow cell may be divided into swaths and tiles. For example, each swath is a longitudinal stripe of the flow cell, and each tile is an area within each stripe. In some embodiments, each tile on the flow cell may be given a unique tile number, which includes the corresponding swath where the tile is located. When the nucleotide sequence of a cluster is determined, it may be stored using a filename which includes not only the sequence information but also the identification of the swath and tile. A non-limiting example of storing such information includes storing the information in a specific file format, such as a FASTQ file format. Then, during the assembly process, if the process determines that a particular read pair maps to more than one possible target polynucleotide, the process may read the spatial location of the cluster on the flow cell and use that spatial location to determine if one of the mapping assignments is more likely than the other based on the spatial location of each read pair on the flow cell. Read pairs which are within a relatively short distance with one another (for example, within 10 flow cell units, 50 flow cell units, 100 flow cell units, 500 flow cell units, or any value therebetween, where one flow cell unit may correspond to about 1 micron, 20 micron, 40 micron, 60 micron, 80 micron, 100 micron, 200 micron, or any value therebetween, in different sequencing instruments) are more likely to have derived from the same target polynucleotide, so the process may assign a read pair which is mapped to multiple target polynucleotides to a particular one based on its flow cell proximitys from other read pairs’ clusters on the flow cell. Therefore, Euclidean distances between clusters may be used to improve the speed and quality of mapping read pairs to their target polynucleotides.
[0091] In some embodiments, for two read pairs A and B with any given flow cell vector separation (also referred to as flow cell proximity) A = (B.x, B.y) - (A.x, A.y) (in other words, the flow cell proximity between the two clusters for the two read pairs on the flowcell can be described with respect to their longitudinal and latitudinal locations) and genomicdistance d on the originating nucleic acid molecule, the linking score bonus, LinkPhred(A)(d), may comprise a Phred scaled score 10Logio(LinkProb(A)(d)). In some embodiments, the link probability LinkProb(A)(d) may be a product of LinkPrior(A) and LinkDistPDF(A)(d), where the link probability LinkProb(A)(d) is the likelihood that these two read pairs are nearby due to linking instead of by chance; LinkPrior(A) is the prior probability of having a real link between A and B, meaning they actually came from the same nucleic acid molecule due to their proximity on the flowcell surface; and LinkDistPDF(A)(d) is the link distribution function of genomic distances d > 0 given that there is a real link between A and B. In one embodiment, both the LinkPrior(A) and LinkDistPDF(A)(d) functions are estimated (fitted to experimental data) for each nucleic acid sample processed, such as by mapping a small portion of a flow cell without using link information, and then measuring the increased rates of observing read pairs nearby on the flow cell and also nearby genomically. In another embodiment, these functions may be estimated (fitted to experimental data) once, such as for a single sequencing assay, and used for many applicable samples.
[0092] FIG. 3A and FIG. 3B illustrate the definitions of the aforementioned genomic distance and flow cell proximity. For example, as shown in FIG. 3 A, a flow cell 100 includes tile 1001, tile 1002, tile 1003, and tile 1004. A target polynucleotide 151 is above tile 1001, a target polynucleotide 152 is above tile 1002, and a target polynucleotide 150 is above tile 1004. Target polynucleotide 150 includes a region 1501 and a region 1502, which are separated by a genomic distance d. The genomic distance d is unknown but can be estimated by preliminary alignment as explained below.
[0093] FIG. 3B shows that after the target polynucleotides are fragmented on the flow cell using surface-immobilized transposome complexes, bound fragments of the target polynucleotides are amplified to form a plurality of polynucleotide clusters on the flow cell 100. Fragments that were derived from the same template genomic sequence are more likely to bind to the flow cell in spatially nearby positions. For example, a plurality of polynucleotide clusters 170, which includes cluster 1701 and cluster 1702, is derived from the target polynucleotide 151 in tile 1001. A plurality of polynucleotide clusters 180, which includes cluster 1801 and cluster 1802, is derived from the target polynucleotide 152 in tile 1002. A plurality of polynucleotide clusters 190, which includes cluster 1901 and cluster 1902, is derived from the target polynucleotide 150 in tile 1004. The clusters 1901 and 1902 are derivedfrom the regions 1501 and 1502, respectively, as shown in FIG. 3A. The clusters 1901 and 1902 are separated by a flow cell vector separation (also referred to as flow cell proximity) A as shown in FIG. 3B. For example, clusters 1901 and 1902 may be located at (xl, yl) and (x2, y2) on the flow cell, respectively. Thus, clusters 1901 and 1902 are separated by A = (|xl-x2|, |yl-y2|).
[0094] After the clusters shown in FIG. 3B are sequenced to obtain read pairs, they may be mapped to a reference sequence. For example, FIG. 3C illustrates the mapping of read pairs obtained from clusters 1901 and 1902 on a reference genome 111. After a preliminary alignment with a standard alignment process, a read pair’s position on the reference genome may not be uniquely determined. For example, the read pair obtained from cluster 1901 may align to candidate positions 1401a, 1401b and 1401c (e.g., chrl:100,000, chr2:500,000, and chr2:550,000) on the reference genome 111 with different alignment scores. Similarly, the read pair obtained from cluster 1902 may align to candidate positions 1402a and 1402b (e.g., chrl:90,000 and chr2:600,000) on the reference genome 111 with different alignment scores. However, the results of this preliminary alignment process may be used to provide a few estimations for the unknown genomic distance between the two read pairs obtained from clusters 1901 and 1902. For example, the estimated genomic distance is d’ between candidate positions 1401a and 1402a, and the estimated genomic distance is d” between candidate positions 1401c and 1402b. In some embodiments, among the candidate positions for each read pair, the candidate position having the highest alignment score is chosen to be the estimated position for the read pair on the reference genome. In some embodiments, when estimating the genomic distance between two read pairs, the estimated position for each read pair is used. A possible link may exist between the same-chromosome alignment candidates. For example, between chrl: 100,000 and chr 1 :90, 000, the estimated genomic distance is 10,000 bp and the determined linking score bonus may be 35; between chr2:500,000 and chr2:600,000, the estimated genomic distance is 100,000 bp and the determined linking score bonus may be 5; and between chr2: 550,000 and chr2: 600,000, the estimated genomic distance is 50,000 bp and the determined linking score bonus may be 15.
[0095] In some embodiments, one linking score bonus is applied to a source read pair alignment candidate. For each source read pair alignment candidate, there may be manydestination read pairs, and for each destination read pair, there may be many destination read pair alignment candidates. Thus, when considering multiple destination alignment candidates for a particular destination read pair, an exemplary method is to use the greatest linking score bonus among the multiple destination alignment candidates of the particular destination read pair (for the source alignment candidate) as the linking score bonus from that destination read pair (for the source alignment candidate). When considering multiple destination read pairs, an exemplary method is to use the sum (in a log-probability space such as Phred scale) of the linking score bonuses from the multiple destination read pairs as the overall linking score bonus for the source alignment candidate. Alternatively, maximization may be used instead of summation, or other methods of combining the linking score bonuses from the multiple destination read pairs may be used.
[0096] In some embodiments, computing the linking score bonus may include a process 300 as illustrated in FIG. 3D. The process 300 may include a sub-process 301 to “loop over” (i.e., consider) all batches (i.e., flow cell regions) in a flow cell. The sub-process 301 may include a sub-process 302 to loop over all source read pairs in each batch. The sub-process 302 may include a sub-process 303 to loop over all candidate alignments X.A of each source read pair X (where the candidate alignments X.A may be obtained by preliminarily aligning the source read pair X to the reference genome sequence). The sub-process 303 may include a sub-process 304 to loop over all destination read pairs Y nearby each source read pair X. The sub-process 304 may include a sub-process 305 to loop over all candidate alignments Y.A of each destination read pair Y (where the candidate alignments Y.A may be obtained by preliminarily aligning the destination read pair Y to the reference genome sequence), and to determine the link probability LinkProb(A)(d) or linking score bonus LinkPhred(A)(d) between the particular candidates X.A and Y.A. In some embodiments, the sub-process 305 includes finding the maximum of the linking score bonuses (“MaxPhredBonus”) for one source candidate X.A from all the candidates Y.A of the destination read pair Y. In some embodiments, the sub-process 304 includes summing the MaxPhredBonus values from all destination read pairs Y, yielding an overall linking score bonus for the source alignment candidate X.A.
[0097] FIG. 4 illustrates a schematic diagram of a computing system 200 in which a mapping system 206 operates in accordance with one or more embodiments. As illustrated,the computing system 200 includes a sequencing device 202 connected to a local device 108 (e.g., a local server device), one or more server device(s) 210, and a client device 214. As shown in FIG. 4, the sequencing device 202, the local device 208, the server device(s) 210, and the client device 214 can communicate with each other via a network 218. The network 218 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 5. While FIG. 4 shows an embodiment of the mapping system 206, this disclosure describes alternative embodiments and configurations below.
[0098] As indicated by FIG. 4, the sequencing device 202 comprises a computing device and a sequencing device system 204 for sequencing a genomic sample or other nucleic- acid polymer. In some embodiments, by executing the sequencing device system 204 using a processor, the sequencing device 202 analyzes polynucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 202. More particularly, the sequencing device 202 receives nucleotide-sample slides (e.g., flow cells) comprising polynucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted polynucleotide fragments.
[0099] In one or more embodiments, the sequencing device 202 utilizes sequencing-by-synthesis (SBS) techniques to sequence polynucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition or in the alternative to communicating across the network 218, in some embodiments, the sequencing device 202 bypasses the network 218 and communicates directly with the local device 208 or the client device 214. By executing the sequencing device system 204, the sequencing device 202 can further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the local device 208 and / or the server device(s) 210.
[0100] As further indicated by FIG. 4, the local device 208 is located at or near a same physical location of the sequencing device 202. Indeed, in some embodiments, the local device 208 and the sequencing device 202 are integrated into a same computing device. The local device 208 may run the mapping system 206 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based onanalyzing such base-call data. As shown in FIG. 4, the sequencing device 202 may send (and the local device 208 may receive) base-call data generated during a sequencing run of the sequencing device 202. By executing software in the form of the mapping system 206, the local device 208 may align nucleotide reads with a reference genome utilizing a BAM file 212 and determine genetic variants based on the aligned nucleotide reads. The local device 208 may also communicate with the client device 214. In particular, the local device 208 can send data to the client device 214, including a binary alignment map (BAM) file, a variant call format (VCF) file, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.
[0101] As further indicated by FIG. 4, the server device(s) 210 are located remotely from the local device 208 and the sequencing device 202. Similar to the local device 208, in some embodiments, the server device(s) 210 include a version of (or are otherwise able to access or implement) the mapping system 206. Accordingly, the server device(s) 210 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As indicated above, the sequencing device 202 may send (and the server device(s) 210 may receive) base-call data from the sequencing device 202. The server device(s) 210 may also communicate with the client device 214. In particular, the server device(s) 210 can send data to the client device 214, including BAM files, VCF files, or other sequencing related information.
[0102] In some embodiments, the server device(s) 210 comprise a distributed collection of servers where the server device(s) 210 include a number of server devices distributed across the network 218 and located in the same or different physical locations. Further, the server device(s) 210 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0103] As indicated above, as part of the server device(s) 210 or the local device 208, the mapping system 206 can generate, encode, and / or implement the BAM file 212 to determine alignments of nucleotide reads from a genomic sample with a reference genome. For instance, the mapping system 206 can identify candidate alignments of one or more nucleotide reads with a primary contiguous sequence, generate primary alignment scores for the candidate alignments, and adjust the alignment scores based on population variant dataindicated in the BAM file 212, as described in greater detail below in relation to the subsequent figures.
[0104] As further illustrated and indicated in FIG. 4, by executing a sequencing application 216, the client device 214 can generate, store, receive, and send digital data. In particular, the client device 214 can receive sequencing data from the local device 208 or receive call files (e.g., BCL) and sequencing metrics from the sequencing device 202. Furthermore, the client device 214 may communicate with the local device 208 or the server device(s) 210 to receive a VCF comprising genotype or variant calls and / or other metrics, such as a base-call-quality metrics or pass-filter metrics. The client device 214 can accordingly present or display information pertaining to variant calls or other genotype calls within a graphical user interface of the sequencing application 216 to a user associated with the client device 214. For example, the client device 214 can present genotype calls, variant calls, and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of the sequencing application 216.
[0105] Although FIG. 4 depicts the client device 214 as a desktop or laptop computer, the client device 214 may comprise various types of client devices. For example, in some embodiments, the client device 214 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 214 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 214 are discussed below with respect to FIG. 5.
[0106] As further illustrated in FIG. 4, the client device 214 includes the sequencing application 216. The sequencing application 216 may be a web application or a native application stored and executed on the client device 214 (e.g., a mobile application, desktop application). The sequencing application 216 can include instructions that (when executed) cause the client device 214 to receive data from the structural-variant-aware mapping system 206 and present, for display at the client device 214, base-call data or data from a VCF. Furthermore, the sequencing application 216 can instruct the client device 214 to display summaries for multiple sequencing runs.
[0107] As further illustrated in FIG. 4, a version of the mapping system 206 may be located and / or implemented (e.g., entirely or in part) on the client device 214 or thesequencing device 202. In yet other embodiments, the mapping system 206 is implemented by one or more other components of the computing system 200, such as the local device 208. In particular, the mapping system 206 can be implemented in a variety of different ways across the sequencing device 202, the local device 208, the server device(s) 210, and the client device 214. For example, the mapping system 206 can be downloaded from the server device(s) 210 to the mapping system 206 and / or the local device 208 where all or part of the functionality of the mapping system 206 is performed at each respective device within the computing system 200.
[0108] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.
[0109] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.
[0110] SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. PatentApplication Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.
[0111] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system comprising components such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference.
[0112] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic or pathogenic sample. In some embodiments,the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another embodiment, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus or fungus. In some embodiments, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0113] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.
[0114] The components of the mapping system 206 can include software, hardware, or both. For example, the components of the mapping system 206 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the client device 208). When executed by the one or more processors, the computer-executable instructions of the mapping system 206 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the mapping system 206 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the mapping system 206 can include a combination of computer-executable instructions and hardware.
[0115] Furthermore, the components of the mapping system 206 performing the functions described herein with respect to the mapping system 206 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud- computing model. Thus, components of the mapping system 206 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the mapping system 206 may be implemented in any application that provides sequencing services including, but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0116] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitorycomputer- readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0117] Embodiments of the disclosure methods can be implemented or performed by a machine, such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor or group of processors for performing the methods described herein may be of various types, including programmable devices (e.g., CPLDs and FPGAs) and non-programmable devices such as gate array ASICs or general-purpose microprocessors.
[0118] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are non-transitory computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0119] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0120] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carrydesired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0121] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0122] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general- purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0123] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked(either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0124] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0125] A cloud- computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.
[0126] FIG. 5 illustrates a block diagram of a computing device 3700 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 3700 may implement the mapping system 206 and the sequencing system 204. As shown by FIG. 5, the computing device 3700 can comprise a processor 3702, a memory 3704, a storage device 3706, an I / O interface 3708, and a communication interface 3710, which may be communicatively coupled by way of a communication infrastructure 3712. In certain embodiments, the computing device 3700 can include fewer or more components than those shown in FIG. 5. The following paragraphs describe components of the computing device 3700 shown in FIG. 5 in additional detail.
[0127] In one or more embodiments, the processor 3702 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 3702 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 3704, or the storage device 3706 and decode and execute them. The memory 3704 may be a volatile or non-volatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 3706 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0128] The I / O interface 3708 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 3700. The I / O interface 3708 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 3708 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 3708 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0129] The communication interface 3710 can include hardware, software, or both. In any event, the communication interface 3710 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 3700 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 3710 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.
[0130] Additionally, the communication interface 3710 may facilitate communications with various types of wired or wireless networks. The communication interface 3710 may also facilitate communications using various communication protocols.The communication infrastructure 3712 may also include hardware, software, or both that couples components of the computing device 3700 to each other. For example, the communication interface 3710 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.
[0131] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0132] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Additional Notes
[0133] Various embodiments of the present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or mediums) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0134] For example, the functionality described herein may be performed as software instructions are executed by, and / or in response to software instructions being executed by, one or more hardware processors and / or any other suitable computing devices. The software instructions and / or other executable code may be read from a computer readable storage medium (or mediums). Computer readable storage mediums may also be referred to herein as computer readable storage or computer readable storage devices.
[0135] The computer readable storage medium can be a tangible device that can retain and store data and / or instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device (including any volatile and / or non-volatile electronic storage devices), a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0136] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions forstorage in a computer readable storage medium within the respective computing / processing device.
[0137] Computer readable program instructions (as also referred to herein as, for example, “code,” “instructions,” “module,” “application,” “software application,” and / or the like) for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the "C" programming language or similar programming languages. Computer readable program instructions may be callable from other instructions or from itself, and / or may be invoked in response to detected events or interrupts. Computer readable program instructions configured for execution on computing devices may be provided on a computer readable storage medium, and / or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression or decryption prior to execution) that may then be stored on a computer readable storage medium. Such computer readable program instructions may be stored, partially or fully, on a memory device (e.g., a computer readable storage medium) of the executing computing device, for execution by the computing device. The computer readable program instructions may execute entirely on a user's computer (e.g., the executing computing device), partly on the user’ s computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0138] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0139] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart(s) and / or block diagram(s) block or blocks.
[0140] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer may load the instructions and / or modules into its dynamic memory and send the instructions over a telephone, cable, or optical line using a modem. A modem local to a server computing system may receive the data on the telephone / cable / optical line and use a converter device including the appropriate circuitry to place the data on a bus. The bus may carry the data to a memory, from which a processor may retrieve and execute the instructions. The instructions received by the memory may optionally be stored on a storage device (e.g., a solid-state drive) either before or after execution by the computer processor.
[0141] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a service, module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. In addition, certain blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states relating thereto can be performed in other sequences that are appropriate.
[0142] It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. For example, any of the processes, methods, algorithms, elements, blocks, applications, or other functionality (or portions of functionality) described in the preceding sections may be embodied in, and / or fully or partially automated via, electronic hardware such application-specific processors (e.g., application-specific integrated circuits (ASICs)), programmable processors (e.g., field programmable gate arrays (FPGAs)), application-specific circuitry, and / or the like (any of which may also combine custom hard-wired logic, logic circuits, ASICs, FPGAs, etc. with custom programming / execution of software instructions to accomplish the techniques).
[0143] Any of the above-mentioned processors, and / or devices incorporating any of the above-mentioned processors, may be referred to herein as, for example, “computers,” “computer devices,” “computing devices,” “hardware computing devices,” “hardware processors,” “processing units,” and / or the like. Computing devices of the above-embodiments may generally (but not necessarily) be controlled and / or coordinated by operating system software, such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.),Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating systems. In other embodiments, the computing devices may be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality, such as a graphical user interface (“GUI”), among other things.
[0144] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.
[0145] It is to be understood that the ranges provided herein include the stated range and any value or sub-range within the stated range, as if such value or sub-range were explicitly recited. For example, a range from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly recited limits of from about 2 kbp to about 20 kbp, but also to include individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc. Furthermore, when “about” and / or “substantially” are / is utilized to describe a value, this is meant to encompass minor variations (up to + / - 10%) from the stated value.
[0146] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.
[0147] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the methods described herein may be made without departing from the spirit of the disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.
[0148] Features, materials, characteristics, or groups described in conjunction with a particular aspect, or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not restricted to the details of any foregoing examples. The protection extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.
[0149] Furthermore, certain features that are described in this disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations, one or more features from a claimed combination can, in some cases, be excised from the combination, and the combination may be claimed as a sub-combination or variation of a sub-combination.
[0150] Moreover, while operations may be depicted in the drawings or described in the specification in a particular order, such operations need not be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated in the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some examples, the actual steps taken in the processes illustrated and / or disclosed may differ from those shown in the figures. Depending on the example, certain of the steps described above may be removed or others may be added. Furthermore, the features and attributes of the specific examples disclosed above may be combined in different ways to form additional examples, all of which fall within the scope of the present disclosure.
[0151] For purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not necessarily all such advantages may be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.
[0152] Conditional language, such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more examples or that one or more examples necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, and / or steps are included or are to be performed in any particular example.
[0153] Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z. Thus, such conjunctive language is not generally intended to imply that certain examples require the presence of at least one of X, at least one of Y, and at least one of Z.
[0154] Language of degree used herein, such as the terms “approximately,” “about,” “generally,” and “substantially” represent a value, amount, or characteristic close to the stated value, amount, or characteristic that still performs a desired function or achieves a desired result.
[0155] The scope of the present disclosure is not intended to be limited by the specific disclosures of preferred examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as non-exclusive.
Claims
WHATIS CLAIMED IS:
1. A method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, the method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations; and c. updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value.
2. The method of claim 1 , further comprising updating the mapping of the second read pair.
3. The method of claim 1 or 2, wherein updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method.
4. The method of claim 2, wherein updating the mapping comprises updating the second candidate locations.
5. The method of claim 2, wherein updating the mapping comprises updating the alignment scores associated with the second candidate locations.
6. The method of claim 2, further comprising determining a select location on the reference sequence for the second read pair based on the updated mapping.
7. The method of claim 6, wherein the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
8. The method of any of claims 1-7, wherein updating the mapping of the first read pair comprises realigning the first read pair to the reference sequence using a local alignment method.
9. The method of any of claims 1-8, wherein updating the mapping comprises updating the first candidate locations.
10. The method of any of claims 1-9, wherein updating the mapping comprises updating the alignment scores associated with the first candidate locations.
11. The method of any of claims 1-10, further comprising determining a select location on the reference sequence for the first read pair based on the updated mapping.
12. The method of claim 11, wherein the select location on the reference sequence for the first read pair is associated with a maximal updated alignment score.
13. The method of any of claims 1-12, wherein the linking score bonus is a Phred scaled score or is determined based on a linking quality statistical model.
14. The method of claim 13, wherein combining the linking score bonus with the alignment scores associated with the first candidate locations comprises adding the linking score bonus to the alignment scores associated with the first candidate locations.
15. The method of any of claims 1-14, further comprising adding the updated mapping to a BAM file which stores the read pairs.
16. A system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations; and c. updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value.
17. The system of claim 16, wherein the method further comprises updating the mapping of the second read pair.
18. The system of claim 17, wherein updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method.
19. The system of claim 17, wherein updating the mapping comprises updating the second candidate locations.
20. The system of claim 17, wherein updating the mapping comprises updating the alignment scores associated with the second candidate locations.
21. The system of claim 17, wherein the method further comprises determining a select location on the reference sequence for the second read pair based on the updated mapping.
22. The system of claim 21, wherein the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
23. The system of any of claims 16-22, wherein updating the mapping of the first read pair comprises realigning the first read pair to the reference sequence using a local alignment method.
24. The system of any of claims 16-22, wherein updating the mapping comprises updating the first candidate locations.
25. The system of any of claims 16-22, wherein updating the mapping comprises updating the alignment scores associated with the first candidate locations.
26. A computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising instructions that, when executed by a processor, cause the processor to perform a method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance betweeneach of the first candidate locations and each of the second candidate locations; and c. updating the mapping of the first read pair if a probability, determined based on the linking score bonus, that altering the mapping will increase the mapping accuracy is above a predetermined value.
27. The computer program of claim 26, wherein the method further comprises updating the mapping of the second read pair.
28. The computer program of claim 27, wherein updating the mapping of the second read pair comprises realigning the second read pair to the reference sequence using a local alignment method.
29. The computer program of claim 27, wherein updating the mapping comprises updating the second candidate locations.
30. The computer program of claim 27, wherein updating the mapping comprises updating the alignment scores associated with the second candidate locations.
31. The computer program of claim 27, wherein the method further comprises determining a select location on the reference sequence for the second read pair based on the updated mapping.
32. A method of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, the method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations; and c. determining a select location on the reference sequence for the first read pair based on the linking score bonus.
33. The method of claim 32, further comprising determining a select location on the reference sequence for the second read pair.
34. The method of claim 33, wherein the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
35. The method of claim 33, wherein the select location on the reference sequence for the second read pair is determined based on a function of the linking score bonus, the alignment scores associated with the first candidate locations, and the alignment scores associated with the second candidate locations.
36. The method of any of claims 32-35, wherein the linking score bonus is a Phred scaled score.
37. The method oof any of claims 32-36, wherein the linking score bonus is determined based on a linking quality statistical model.
38. The method of any of claims 32-37, wherein the select location on the reference sequence for the first read pair is associated with a maximal updated alignment score.
39. A system of mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations; and c. determining a select location on the reference sequence for the first read pair based on the linking score bonus.
40. The system of claim 39, wherein the method further comprises determining a select location on the reference sequence for the second read pair.
41. The system of claim 40, wherein the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
42. The system of claim 40, wherein the select location on the reference sequence for the second read pair is determined based on a function of the linking score bonus, the alignment scores associated with the first candidate locations, and the alignment scores associated with the second candidate locations.
43. The system of any of claims 39-42, wherein the linking score bonus is a Phred scaled score.
44. The system of any of claims 39-43, wherein the linking score bonus is determined based on a linking quality statistical model.
45. A computer program for mapping read pairs sequenced from polynucleotide clusters on a flow cell to a reference sequence, comprising instructions that, when executed by a processor, cause the processor to perform a method comprising: a. mapping a first read pair and a second read pair to the reference sequence to obtain a plurality of first candidate locations and associated alignment scores for the first read pair and a plurality of second candidate locations and associated alignment scores for the second read pair; b. determining a linking score bonus between the first read pair and the second read pair based on the flow cell proximity of the clusters for the first read pair and the second read pair on the flow cell and the genomic distance between each of the first candidate locations and each of the second candidate locations; and c. determining a select location on the reference sequence for the first read pair based on the linking score bonus.
46. The computer program of claim 45, wherein the method further comprises determining a select location on the reference sequence for the second read pair.
47. The computer program of claim 46, wherein the select location on the reference sequence for the second read pair is associated with a maximal updated alignment score.
Citation Information
Patent Citations
Method of nucleic acid amplification
US20050100900A1
Labelled nucleotides
US20060188901A1
Modified polymerases for improved incorporation of nucleotide analogues
US20060240439A1
Polymerases
US20060281109A1
Modified nucleotides
US20070166705A1