Systems and methods for quantifying sequence read connectivity

Statistical modeling of read pairs in nucleic acid sequencing improves alignment accuracy and phasing by quantifying linking quality, addressing the loss of connectivity in fragmented genomic sequences and mixed samples.

WO2026072436A1PCT designated stage Publication Date: 2026-04-02ILLUMINA INC
View PDF 32 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing methods, particularly shotgun sequencing, lose connectivity and proximity information of fragmented genomic sequences, leading to difficulties in reconstructing phasing information for diploid and polyploid genomes and mixed samples, and fail to accurately quantify linking quality between read pairs.

Method used

A method involving statistical modeling to determine a linking quality score between read pairs by fitting a model to sequencing data, using a Gauss-Newton or Levenberg-Marquardt optimizer, and applying it to BAM files to improve alignment accuracy and phasing information.

Benefits of technology

Enhances the accuracy of read pair alignments and phasing by quantifying the likelihood that read pairs originated from the same nucleic acid fragment, improving computational efficiency and reducing errors in genomic reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046976_02042026_PF_FP_ABST
    Figure US2025046976_02042026_PF_FP_ABST
Patent Text Reader

Abstract

In some embodiments, the disclosed technology relates to a method of determining a linking quality score between two read pairs. In some embodiments, the disclosed technology relates to a method for assigning read pairs to target nucleic acid polymers. In some embodiments, the disclosed technology relates to a method of producing a mapping of read pairs to a reference sequence. In some embodiments, the disclosed technology relates to a method of producing a mapping of read pairs to a reference sequence. In some embodiments, the disclosed technology relates to a method of improving accuracy of split-alignments of read pairs to a reference sequence.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR QUANTIFYING SEQUENCE READ CONNECTIVITYCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No, 63 / 700,049, filed September 27, 2024, the content of which is incorporated by reference in its entirety.BACKGROUNDField of the Invention

[0002] This disclosure relates generally to sequencing nucleic acids, and more specifically to determining probabilistic models that quantify the quality of links between read pairs obtained from a sequencing flow cell and using one or more models to determine a linking quality for the links found between two read pairs obtained from a sequencing flow cell.Description of the Related Art

[0003] In recent years, biotechnology firms and research institutions have improved hardware and software for sequencing nucleotides and determining variant calls for genomic samples. For instance, some existing nucleobase sequencing platforms determine individual nucleobases within sequences from a genomic samples’ cells by using conventional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms can monitor millions to billions of nucleic acid polymers being synthesized in parallel to predict nucleobase calls from a larger base call dataset. For instance, a camera in many SBS platforms captures images of excited fluorescent tags incorporated into oligonucleotides for determining the nucleobase calls. After capturing such images, existing sequencing platforms send base call data (or image-based data) to a computing device to apply sequencing data analysis software that determines a nucleobase sequence for a genomic sample or other nucleic acid polymer. In some cases, such software maps and aligns nucleotide reads determined by the sequencing platform with a reference genome comprising at least a primaryreference genome, existing data analysis software can further utilize a variant caller to identify genotype and / or variants within a genomic sample. These variants can include single nucleotide polymorphisms (SNPs), insertions or deletions (indels), or structural variants among other genetic variabilities.

[0004] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods, use a shotgun approach to sequence large genomic DNA fragments, called template genomic sequences. Specifically, template genomic sequences are first fragmented in solution into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been read, knowledge of their connectivity and proximity to each other in the original template genomic sequence is lost. The process of ordering the sequence fragments to arrive at the sequence of the original template genomic sequence may be achieved by alignment to a reference genome sequence or by “assembly.” Assembly processes can be computationally intensive and timeconsuming. In addition, sequence and assembly errors can become a problem depending upon the sequencing methodology used and the qualify of genomic DNA samples under evaluation,

[0005] Moreover, many genomes of interest contain more than one version of each chromosome. For example, the human genome is diploid, having two sets of chromosomes — one set inherited from each parent. Some organisms have polyploid genomes with more than two sets of chromosomes. Examples of polyploid organisms include animals, such as salmon, and many plant species such as wheat, apple, oat and sugar cane. When diploid and polyploid genomes are fragmented and sequenced in typical shotgun methods, phasing information, pertaining to the identity of which fragments came from a particular chromosome, is lost. This phasing information can be difficult or impossible to reconstruct using typical shotgun methods.

[0006] Somewhat similar, yet often more complex difficulties, can arise when mixed samples are evaluated. Mixed samples can contain nucleic acid molecules, such as chromosomes, mRNA transcripts, plasmids etc., from two or more organisms. Mixed samples derived from multiple organisms are often referred to as metagenomic samples. Other examples of mixed samples include different cell or tissue samples that although being derivedtissues made up from a mixture of healthy cells and cancerous cells, tissues having both pre- cancerous cells and cancerous cells, and tissue samples having two or more different types of cancerous cells. Indeed, there may be a variety of different types of cancer cells in one sample as is the case for cancer samples that have mosaicity. Another example of different cells derived from a single organism are mixtures of maternal and fetal cells obtained from a pregnant female (e.g. from the blood or from tissues). When mixed nucleic acid samples are fragmented and sequenced in typical shotgun methods information pertaining to the identity’ of which fragments came from which cell, organism or other source can be lost. This origin information can be difficult or impossible to reconstruct using typical shotgun methods.SUMMARY

[0007] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0008] In one aspect, the disclosed technology provides a method of determining a linking quality score between two read pairs. The method may comprise step (a) obtaining a first set of read pairs from a flow cell comprising clusters of polynucleotides bound to locations on the flow cell. The method may further comprise step (b) fitting a linking quality statistical model to data generated from the first set of read pairs. In some embodiments, such a statistical model may estimate the probability that, for two read pairs derived from the same nucleic acid fragment at two regions separated by a given genomic distance, two clusters on the flow cell associated with the two read pairs would be separated by a relative displacement. The method may further comprise step (c) determining the linking quality score between two read pairs in a second set of read pairs from the flow cell based on the statistical model, the genomic distanceassociated with the two read pairs.

[0009] In some embodiments, the first set of read pairs may be sequenced in a sequencing run using a first sample. In some embodiments, the second set of read pairs may be sequenced in the same sequencing run using a second sample. In some embodiments, the first set of read pairs each has a mapping quality score (MAPQ) of greater than 60. In some embodiments, the first set of read pairs are located within at least 200 kb of one another on the same nucleic acid fragment and their corresponding clusters are located within 1000 flow cell units of one another on the flow cell.

[0010] In some embodiments, fitting the linking quality’ statistical model to data generated from the first set of read pairs may comprise constructing a two-dimensional histogram of the count of the relative displacement between two clusters on the flow cell associated with two read pairs in the first set of read pairs. In some cases, the two read pairs were derived from two regions on the same nucleic acid fragment, wherein the two regions are separated by a given genomic distance.

[0011] In some embodiments, the contour lines of the histogram may comprise an oval-like shape. In some embodiments, the major axis of the oval-like shape may align with the direction of flow' of sequencing reaction reagents on the flow cell. In some embodiments, the oval-like shape may center at the origin. In some embodiments, the contour lines of the histogram may comprise a circle. In some embodiments, the circle may center at the origin. In some embodiments, the contour lines of the histogram may comprise a pair of blobs. In some embodiments, the pair of blobs may be symmetric with respect to the x-axis and the y-axis of the histogram. For example, the y-axis of the histogram may align with the direction of flow of sequencing reaction reagents on the flow' cell. In some embodiments, each blob of the pair of blobs may be a distance away from the x-axis of the histogram.

[0012] In some embodiments, fitting the statistical model may comprise using an optimizer for weighted non-linear least squares to obtain model parameters for the statistical model. In some embodiments, the optimizer is based on the Gauss-Newton or Levenberg- Marquardt algorithm. In some embodiments, the first and second sets of read pairs may be obtained from a BAM file, and wherein the determined linking quality score is added to the BAM file. In some embodiments, the nucleic acid fragment may be a genomic fragment.read pairs to target nucleic acid polymers, comprising. The method may comprise step (a) fragmenting the target nucleic acid polymers in a flow cell to generate polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common target nucleic acid polymer are preferentially captured at proximal locations on the flow cell. The method may further comprise step (b) sequencing the polynucleotide fragments to generate read pairs. The method may further comprise step (c) aligning the read pairs to a reference sequence to determine a position on the reference sequence for each read pair. The method may further comprise step (d) assigning two read pairs to the same target nucleic acid polymer based on the linking quality score between the two read pairs. In some embodiments, the linking quality score between the first read pair and the second read pair may be determined according to the method of determining a linking quality' score between two read pairs as described above. In alternative embodiments, however, the linking quality score may be directly calculated by inputting the results of steps (b) and / or (c) (e.g., genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexistmg / stored / received linking quality statistical model.

[0014] In some embodiments, step (a) may comprise adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. In some embodiments, the inserts comprise cleavage sites. In some embodiments, the inserts are cleaved at the cleavage sites. In some embodiments, inserts are added into the target nucleic acid polymers by transposases. In some embodiments, each of the inserts comprises a first transposon element and a second transposon element that is contiguous with the first transposon element. In some embodiments, the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, the modified nucleic acid polymers are attached to the flow cell prior to producing the fragments. In some embodiments, the transposases are removed from the modified nucleic acid polymers prior to attaching of the modified nucleic acid polymer to the flow cell. In some embodiments, the transposases are removed from the modified nucleic acid polymer after attaching the modified nucleic acid polymers to the flow cell. In some embodiments, the polynucleotide fragments passivelyare actively transported to the locations on the flow cell. In some embodiments, the polynucleotide fragments are amplified at the locations to produce amplified polynucleotide fragments. In some embodiments, step (b) may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. In some embodiments, the amplifying comprises extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. In some embodiments, the amplifying comprises extending at least two primer species in a bridge amplification technique. In some embodiments, step (a) may comprise stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer.

[0015] In another aspect, the disclosed technology provides a method of producing a mapping of read pairs to a reference sequence. The method may comprise step (a) fragmenting target nucleic acid polymers in a flow cell to generate polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell. The method may further comprise step (b) sequencing the polynucleotide fragments to generate read pairs. The method may further comprise step (c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. The method may further comprise step (d) producing the mapping, comprising selecting a first position among the plurality of candidate positions for a first read pair based on the linking quality score between the first read pair and a second read pair, wherein the second read pair may optionally have an alignment score associated with a second position that is above a predetermined value. In some embodiments, the linking quality score between the first read pair and the second read pair may be determined according to the method of determining a linking quality score between two read pairs as described above. In alternative embodiments, however, the linking quality score may be directly calculated by inputting the results of steps (b) and / or (c) (e.g., genomic distance between two read pairs and relative displacement between two clusters on the flow' cellmodel.

[0016] In some embodiments, step (a) may comprise adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. In some embodiments, the inserts comprise cleavage sites. In some embodiments, the inserts are cleaved at the cleavage sites. In some embodiments, inserts are added into the target nucleic acid polymers by transposases. In some embodiments, each of the inserts comprises a first transposon element and a second transposon element that is contiguous with the first transposon element. In some embodiments, the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, the modified nucleic acid polymers are attached to the flow cell prior to producing the fragments. In some embodiments, the transposases are removed from the modified nucleic acid polymers prior to attaching of the modified nucleic acid polymer to the flow cell. In some embodiments, the transposases are removed from the modified nucleic acid polymer after attaching the modified nucleic acid polymers to the flow cell. In some embodiments, the polynucleotide fragments passively diffuse to the locations on the flow' cell. In some embodiments, the polynucleotide fragments are actively transported to the locations on the flow cell. In some embodiments, the polynucleotide fragments are amplified at the locations to produce amplified polynucleotide fragments. In some embodiments, step (b) may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. In some embodiments, the amplifying comprises extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. In some embodiments, the amplifying comprises extending at least two primer species in a bridge amplification technique. In some embodiments, step (a) may comprise stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer.

[0017] In another aspect, the disclosed technology provides a method of producing a mapping of read pairs to a reference sequence. The method may comprise step (a) fragmenting target nucleic acid polymers in a flow cell to generate polynucleotide fragmentsfrom a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell. The method may further comprise step (b) sequencing the polynucleotide fragments to generate read pairs. The method may further comprise step (c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. The method may further comprise step (d) producing the mapping, comprising selecting a first position among the plurality of candidate positions for a first read pair based on a function of optionally (i) the alignment score associated with the first position, optionally (li) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and optionally (ni) the linking quality score between the first read pair and the second read pair. In some embodiments, the linking quality score between the first read pair and the second read pair may be determined according to the method of determining a linking quality score between two read pairs as described above. In alternative embodiments, however, the linking quality score may be directly calculated by inputting the results of steps (b) and / or (c) (e.g., genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexist! ng / stored / received linking quality statistical model.

[0018] In some embodiments, step (a) may comprise adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. In some embodiments, the inserts comprise cleavage sites. In some embodiments, the inserts are cleaved at the cleavage sites. In some embodiments, inserts are added into the target nucleic acid polymers by transposases. In some embodiments, each of the inserts comprise a first transposon element and a second transposon element that is contiguous with the first transposon element. In some embodiments, the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, the modified nucleic acid polymers are attached to the flow cell prior to producing the fragments. In some embodiments, the transposases are removed from the modified nucleic acid polymers prior to attaching the modified nucleic acid polymer to the flow cell. In some embodiments, the transposases are removed from the modified nucleic acid polymer after attaching the modified nucleic aciddiffuse to the locations on the flow cell. In some embodiments, the polynucleotide fragments are actively transported to the locations on the flow cell. In some embodiments, the polynucleotide fragments are amplified at the locations to produce amplified polynucleotide fragments. In some embodiments, step (b) may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. In some embodiments, the amplifying comprises extending at least one primer species atached to the locations to produce the amplified polynucleotide fragments atached to the locations via the primer species. In some embodiments, the amplifying comprises extending at least two primer species in a bridge amplification technique. In some embodiments, step (a) may comprise stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer.

[0019] In another aspect, the disclosed technology provides a method of improving the accuracy of split-alignments of read pairs to a reference sequence. The method may comprise step (a) fragmenting target nucleic acid polymers in a flow cell to generate polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell. The method may further comprise step (b) sequencing the polynucleotide fragments to generate read pairs. The method may further comprise step (c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. The method may further comprise step (d) for a first read pair comprising different, segments that align to disjoint regions in the reference sequence, selecting a first position among the plurality of candidate positions for a first, segment of the first read pair based on the Unking quality score between the first segment, and a second read pair, wherein the second read pair may optionally have an alignment score associated with a second position that is above a predetermined value. In some embodiments, the linking quality score between the first read pair and the second read pair may be determined according to the method of determining a linking quality score between two read pairs as described above. In alternative embodiments, however, the linking quality score may be directly calculated by inputting the results of steps (b) and / or (c) (e.g., genomiccell associated with the two read pairs) to a preexisting / stored / received linking quality statistical model.

[0020] In some embodiments, step (a) may comprise adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. In some embodiments, the inserts comprise cleavage sites. In some embodiments, the inserts are cleaved at the cleavage sites. In some embodiments, inserts are added into the target nucleic acid polymers by transposases. In some embodiments, each of the inserts comprises a first transposon element and a second transposon element that is contiguous with the first transposon element. In some embodiments, the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, the modified nucleic acid polymers are attached to the flow cell prior to producing the fragments. In some embodiments, the transposases are removed from the modified nucleic acid polymers prior to attaching of the modified nucleic acid polymer to the flow cell. In some embodiments, the transposases are removed from the modified nucleic acid polymer after attaching the modified nucleic acid polymers to the flow' cell. In some embodiments, the polynucleotide fragments passively diffuse to the locations on the flow cell. In some embodiments, the polynucleotide fragments are actively transported to the locations on the flow cell. In some embodiments, the polynucleotide fragments are amplified at the locations to produce amplified polynucleotide fragments. In some embodiments, step (b) may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. In some embodiments, the amplifying comprises extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. In some embodiments, the amplifying comprises extending at least two primer species in a bridge amplification technique. In some embodiments, step (a) may comprise stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer.

[0021] In another aspect, the disclosed technology provides a method of improving accuracy of split-alignments of read pairs to a reference sequence. The method may comprisefragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell. The method may further comprise step (b) sequencing the polynucleotide fragments to generate read pairs. The method may further comprise step (c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. The method may- further comprise step (d) for a first read pair comprising different segments that align to disjoint regions in the reference sequence, selecting a first position among the plurality of candidate positions for a first segment of the first read pair based on a function of optionally (i) the alignment score associated with the first segment, optionally (ii) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and optionally (iii) the linking quality score between the first segment and the second read pair. In some embodiments, the linking quality score between the first read pair and the second read pair may be determined according to the method of determining a linking quality score between two read pairs as described above. In alternative embodiments, however, the linking quality score may be directly calculated by inputting the results of steps (b) and / or (c) (e.g., genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexistmg / stored / received linking quality statistical model.

[0022] In some embodiments, step (a) may comprise adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. In some embodiments, the inserts comprise cleavage sites. In some embodiments, the inserts are cleaved at the cleavage sites. In some embodiments, inserts are added into the target nucleic acid polymers by transposases. In some embodiments, each of the inserts comprises a first transposon element and a second transposon element that is contiguous with the first transposon element. In some embodiments, the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, the modified nucleic acid polymers are attached to the flow cell prior to producing of the fragments. In some embodiments, the transposases are removed from the modified nucleic acid polymers prior to attaching theremoved from the modified nucleic acid polymer after attaching the modified nucleic acid polymers to the flow cell. In some embodiments, the polynucleotide fragments passively diffuse to the locations on the flow cell. In some embodiments, the polynucleotide fragments are actively transported to the locations on the flow cell. In some embodiments, the polynucleotide fragments are amplified at the locations to produce amplified polynucleotide fragments. In some embodiments, step (b) may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. In some embodiments, the amplifying comprises extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. In some embodiments, the amplifying comprises extending at least two primer species in a bridge amplification technique. In some embodiments, step (a) may comprise stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer,

[0023] Features which are described in the context of separate aspects and embodiments of the invention may be used together and / or be interchangeable. Similarly, features described in the context of a single embodiment may also be provided separately or in any suitable sub-combination.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG,” herein).

[0025] FIG. 1A is a schematic diagram that illustrates the spatial relationship of a flow cell and target polynucleotides on the flow cell.

[0026] FIG. I B is a schematic diagram that illustrates the spatial relationship of the flow cell and the nucleic acid clusters derived from the target polynucleotides.a reference genome.

[0028] FIG. 2 A is a flowchart illustrating a method for determining a linking quality score according to some embodiments of the disclosed technology.

[0029] FIG. 2B is a flowchart illustrating a method for assigning read pairs to target nucleic acid polymers according to some embodiments of the disclosed technology.

[0030] FIG. 2C is a flowchart illustrating a method for producing a mapping of read pairs to a reference sequence according to some embodiments of the disclosed technology.

[0031] FIG. 2D is a flowchart illustrating a method for improving accuracy of splitalignments of read pairs to a reference sequence according to some embodiments of the disclosed technology.

[0032] FIG. 3A shows experimental data obtained from a sequencing run in a first Illumina sequencing platform.

[0033] FIG. 3B shows experimental data obtained from a sequencing run in the first Illumina sequencing platform at a higher resolution.

[0034] FIG. 4 shows experimental data obtained from a sequencing run in a second Illumina sequencing platform.

[0035] FIG. 5 shows the distribution of the hotspot component according to some embodiments of the disclosed technology.

[0036] FIG. 6 shows the distribution of the diffusion component according to some embodiments of the disclosed technology.

[0037] FIG. 7 shows one of the distributions of the y-axis component according to some embodiments of the disclosed technology.

[0038] FIG. 8 is an experimentally obtained two-dimensional histogram.

[0039] FIG. 9 A shows the count of the genomic distance GDIST between two linked read pairs from an experiment. FIG 9B shows a “background model” where the read pairs are not linked and compares it with the data shown in FIG. 9A.

[0040] FIG. 10 illustrates a computer environment in which a linking quality system can operate in accordance with one or more embodiments of the present disclosure.

[0041] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.Overview

[0042] During massively parallel sequencing processes, such as those found in Next Generation Sequencing (NGS) systems, relatively short reads of 100-500 nucleotides are typically generated. During this process, sample nucleic acids may be fragmented by tagmentation of the sample directly on the flow cell, in some embodiments. In an example process, transposons are linked to the flow' cell, and when contacted by sample nucleic acids may fragment the sample nucleic acids resulting in many fragments from the same nucleic acid molecule being generated. This example process may result in fragments from the same nucleic acid molecule being bound to the flow' cell in geographically adjacent or nearby locations. These fragments then give rise to nucleic acid clusters at nearby locations on the flow cell, and these nucleic acid clusters will participate in sequencing reactions to generate reads or read pairs. Therefore, reads or read pairs that are generated by sequencing these fragments which are located nearby each other on the flow cell may be determined by the disclosed systems or methods to be “connected” or “linked”, or that there are “links” among the reads or read pairs since they may have derived from the same originating nucleic acid molecule.

[0043] This type of linking information or connectivity information may enable reconstruction of a nucleic acid molecule by way of grouping together linked read pairs coming from nearby locations. However, in some cases, two reads or read pairs that are nearby on the flow' cell may not have originated from the same nucleic acid molecule; they may be nearby by chance alone. Therefore, in order to leverage the linking information of read pairs in processes such as reconstruction of the originating nucleic acid molecule or mapping of the sequence reads to a reference sequence, there is a need for methods that can accurately quantify the relative likelihood that two read pairs obtained from nearby locations on a flow celllikelihood is referred to as a “linking quality” and the measure of linking quality between the two read pairs is a “linking quality score.” In previous approaches to quantify the linking quality, linking quality metrics may be limited to static ranges, which cannot account for variations in sequencing systems (such as variations in imaging components, fluidics systems, excitation optics, etc.). Previous approaches may not account for the experimental observation that two links having the same genomic distance and flow cell distance can differ in linking quality due to different orientations of the relative displacement on the flow cell. Moreover, previous approaches may use large lookup tables to compute linking quality metrics, which is not computationally efficient.

[0044] To quantify the linking quality more accurately and more efficiently, the disclosed technology provides statistical models that describe, for two linked read pairs at a certain genomic distance on the originating nucleic acid molecule, the probability' that the two read pairs’ clusters are at a given relative displacement on the flow cell. Each statistical model is fited to sequencing data in which preliminary alignments of the read pairs to a reference sequence have been made. Once a model has been chosen and fited to data to obtain the model parameters, the disclosed technology can further provide prescribed methods for determining a linking quality score between read pairs obtained from the flow cell and then applying the determined linking quality scores in phasing, mapping, split alignment, etc. In one embodiment, the linking quality score may comprise a Phred scaled score (e.g., lOLogio) of the link probability, which is the likelihood that two read pairs’ associated clusters are nearby on the flow cell due to linking instead of by chance, and the link probability can be calculated from the fitted statistical model.

[0045] In some embodiments, a sequencing system may determine the sequence and location of read pairs in each cluster on the flow cell. The system may create a BAM file for each such read pair which includes the sequence of the read pair and identification of one or more possible target polynucleotides. For any BAM file which includes more than one possible target polynucleotide, the system may attempt to disambiguate the read pairs byreferring to the spatial location information contained with each read pair. As mentioned above, if some read pairs are very near one another on the flow cell, then the system may determine that one read pair is more likely to have originated from a particular target polynucleotide thanthat read pair is updated to indicate only the correct target polynucleotide.

[0046] In some embodiments, in order to manage the spatial location of each cluster, each flow cell may be divided into swaths and tiles. For example, each swath is a longitudinal stripe of the flow cell, and each tile is an area within each stripe. In some embodiments, each tile on the flow'' cell may be given a unique tile number, which includes the corresponding swath where the tile is located. When the nucleotide sequence of a cluster is determined, it may be stored using a filename which includes not only the sequence information but also the identification of the swath and tile. A non-limiting example of storing such information includes storing the information in a specific file format, such as a FASTQ file format. Then, during the assembly process, if the process determines that a particular fragment maps to more than one possible target polynucleotide, the process may read the spatial location of the cluster and use that spatial location to determine if one of the mapping assignments is more likely than the other based on the spatial location of each read pair on the flow cell. Read pairs which are within a relatively short distance with one another are more likely to have derived from the same target polynucleotide, so the process may assign a read pair which is mapped to multiple target polynucleotides to a particular one based on their relative spatial locations on the flow' cell.

[0047] FIG. 1A and FIG IB illustrate the definitions of the aforementioned genomic distance and relative displacement. For example, as shown in FIG. 1 A, a flow cell 100 includes tile 1001, tile 1002, tile 1003, and tile 1004. A target polynucleotide 151 is above tile 1001, a target polynucleotide 152 is above tile 1002, and a target, polynucleotide 150 is above tile 1004. Target, polynucleotide 150 includes a region 1501 and a region 1502, which are separated by a genomic distance d. The genomic distance d is unknown but can be estimated by preliminary' alignment as explained below.

[0048] FIG. IB shows that after the target, polynucleotides are fragmented on the flow cell using surface-immobilized transposome complexes, bound fragments of the target polynucleotides are amplified to form a plurality of nucleic acid clusters on the flow' cell 100. Fragments that were derived from the same template genomic sequence are more likely to bind to the flow cell in spatially nearby positions. For example, a plurality of nucleic acid clusters 170, which includes cluster 1701 and cluster 1702, is derived from the target polynucleotidecluster 1802, is derived from the target polynucleotide 152 in tile 1002. A plurality of nucleic acid clusters 190, which includes cluster 1901 and cluster 1902, is derived from the target polynucleotide 150 in tile 1004. The clusters 1901 and 1902 are derived from the regions 1501 and 1502, respectively, as shown in FIG. 1A. The clusters 1901 and 1902 are separated by a flow cell vector separation (also referred to as relative displacement) A as shown in FIG. IB.

[0049] After the clusters shown in FIG. IB are sequenced to obtain read pairs, they may be mapped to a reference sequence. For example, FIG. 1C illustrates the mapping of read pairs obtained from clusters 1901 and 1902 on a reference genome 111. After a preliminary alignment with a standard alignment process, a read pair’s position on the reference genome may not be uniquely determined. For example, the read pair obtained from cluster 1901 may align to candidate positions 1401a, 1401 b and 1401c on the reference genome 111 with different alignment scores. Similarly, the read pair obtained from cluster 1902 may align to candidate positions 1402a and 1402b on the reference genome 111 with different alignment scores. However, the results of this preliminary alignment process may be used to provide a few estimations for the unknown genomic distance between the two read pairs obtained from clusters 1901 and 1902, For example, the estimated genomic distance is d’ between candidate positions 1401a and 1402a, and the estimated genomic distance is d” between candidate positions 1401c and 1402b, In some embodiments, among the candidate positions for each read pair, the candidate position having the highest alignment score is chosen to be the estimated position for the read pair on the reference genome. In some embodiments, when calculating / estimating the genomic distance between two read pairs, the estimated position for each read pair is used.

[0050] In some embodiments, for two read pairs A and B with any given flow cell vector separation (also referred to as relative displacement) A:::(B.x, B.y) - (A.x, A.y)::::(Ax, Ay) and genomic distance g; on the originating nucleic acid molecule, the linking quality score, LinkPhred(A)($j), may comprise a Phred scaled scorereads linked)), wherein the link probability P(AX, Ay, reads linked) is the likelihood that A and B are nearby on the flow cell due to linking instead of by chance, which may be described by a linking quality statistical model and fited to experimental data. The true genomic distance gtis unknown, but an estimated value based on preliminary al ignment may be used.

[0051] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art. The use of the term “including” as well as other forms, such as “include”, “includes,” and “included,” is not limiting. The use of the term “having” as well as other forms, such as “have”, “has,” and “had,” is not limiting. As used in this specification, whether in a transitional phrase or in the body of the claim, the terms “comprise(s)” and “comprising” are to be interpreted as having an open-ended meaning. That is, the above terms are to be interpreted synonymously with the phrases “having at least” or “including at least.” For example, when used in the context of a process, the term “comprising” means that the process includes at least the recited steps, but may include additional steps. When used in the context of a compound, composition, or device, the term “comprising” means that the compound, composition, or device includes at least the recited features or components, but may also include additional features or components.

[0052] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary' skill in the art to which the present disclosure belongs. See, e.g. Singleton et. al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al,, Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.

[0053] As used herein, a “nucleotide” includes a nitrogen containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is a. ribose, and in deoxyribonucleotides (DNA), the sugar is a deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2’ position in ribose. The nitrogen containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-l atom of deoxyribose is bonded to N-l of a pyrimidine or N-9 of a purine. The phosphate groups may be in the mono- , di-, or tri-phosphate form. These nucleotides may be natural nucleotides, but it is to be furtheraforementioned nucleotides can also be used.

[0054] As used herein, “nucleobase” is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can be naturally occurring or synthetic. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purines substituted at the 8 position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7- deaza-adenme, N4-ethanocytosme, 2,6- diaminopurine, N6-ethano-2,6-diammopurine, 5- methylcytosine, 5-(C3-C6)- alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-mdole, ethenoadenine and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510 and PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (“Practical Handbook of Biochemistry' and Molecular Biology”, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all herein incorporated by reference in their entireties,

[0055] The terms “polynucleotide,” “oligonucleotide,” “nucleic acid” and “nucleic acid molecules” are used interchangeably herein and refer to a covalently linked sequence of nucleotides of any length (i.e., ribonucleotides for RNA, deoxyribonucleotides for DNA, analogs thereof, or mixtures thereof) in which the 3’ position of the pentose of one nucleotide is joined by a phosphodiester group to the 5’ position of the pentose of the next. The terms should be understood to include, as equivalents, analogs of either DNA, RNA, cDNA, or antibody-oligo conjugates made from nucleotide analogs and to be applicable to single stranded (such as sense or antisense) and double stranded polynucleotides. The term as used herein also encompasses cDNA, that is complementary or copy DNA produced from an RNA template, for example by the action of reverse transcriptase. This term refers only to the primary structure of the molecule. Thus, the term includes, without limitation, triple-, double- and single-stranded deoxyribonucleic acid (“DNA”), as well as triple-, double- and single- stranded ribonucleic acid (“RN A”). The nucleotides include sequences of any form of nucleic acid.nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.

[0057] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety' of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, senes of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at. lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A senes of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series.

[0058] The term “primer,” as used herein refers to an isolated oligonucleotide that is capable of acting as a point of initiation of synthesis when placed under conditions inductive to synthesis of an extension product (e.g., the conditions include nucleotides, an inducing agent such as DN A polymerase, and a suitable temperature and pH). The primer is preferably single stranded for maximum efficiency in amplification, but may alternatively be double stranded.extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to prime the syn thesis of extension products in the presence of the inducing agent. The exact lengths of the primers will depend on many factors, including temperature, source of primer, use of the method, and the parameters used for primer design.

[0059] As used herein the term “chromosome” refers to the heredity-bearing gene carrier of a living cell, which is derived from chromatin strands comprising DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein.

[0060] A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences.

[0061] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. In some embodiments, the term “reference genome” refers to a digital nucleic acid sequence assembled as a representative example (or representati ve examples) of genes and other genetic sequences of an organism. Regardless of the sequence length, in some cases, a reference genome represents an example set of genes or a set of nucleic acid sequences in a digital nucleic acid sequenced determined by scientists as representative of an organism of a particular species. For example, a linear human reference genome may be GRCh38 or other versions of reference genomes from the Genome Reference Consortium. As noted above, in some cases, a reference genome includes multi-base codes. As a further example, a reference genome may include a graph reference genome that includes both a linear reference genome and paths representing nucleic acid sequences from ancestral haplotypes, such as Illumina DRAGEN Graph Reference Genome hg!9.

[0062] For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10’’ times larger, or at least about 106times larger, or at least about 107times larger. In one example, thegenomic reference sequences. For example, the reference sequence can be a reference human genome sequence, such as hg!9 or hg38. In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hg!9. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub- chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.

[0063] The term “nucleic acid sample” herein refers to a sample, typically derived from a biological fluid, cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments the nucleic acid sample comprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (e.g., surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (e.g., patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employ ed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (e.g., namely, a sample that is not subjected to any suchbiological “test” samples with respect to the methods described herein.

[0064] As used herein, the term “cluster” or “clump” refers to a group of molecules, e.g., a group of DNA, or a group of signals. In some embodiments, the signals of a cluster are derived from different features. In some embodiments, a signal clump represents a physical region covered by one amplified oligonucleotide. Each signal clump could be ideally observed as several signals. Accordingly, duplicate signals could be detected from the same clump of signals. In some embodiments, a cluster or clump of signals can comprise one or more signals or spots that correspond to a particular feature. When used in connection with microarray devices or other molecular analytical devices, a cluster can comprise one or more signals that together occupy the physical region occupied by an amplified oligonucleotide (or other polynucleotide or polypeptide with a same or similar sequence). For example, where a feature is an amplified oligonucleotide, a cluster can be the physical region covered by one amplified oligonucleotide. In other embodiments, a cluster or clump of signals need not strictly correspond to a feature. For example, spurious noise signals may be included in a signal cluster but not necessarily be within the feature area. For example, a cluster of signals from four cycles of a sequencing reaction could comprise at least four signals. For example, the term “cluster of oligonucleotides” (or “oligonucleotide cluster”, “cluster of polynucleotides”, “polynucleotide cluster”, “cluster of nucleic acids”, or “nucleic acid cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-patterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.

[0065] The term “next generation sequencing (NGS)” herein refers to sequencing methods that allow' for massively parallel sequencing of clonally amplified molecules and ofsynthesis using reversible dye terminators, and sequencing-by-ligation.

[0066] The term “read” or “sequence read” (or sequencing reads) refer to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region, e.g., that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (e.g., 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. A sequence read may be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g,, in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTS EQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0067] In some embodiments, the term “read” refers to an inferred or predicted sequence of one or more nucleotide bases (or nucleobase pairs) from all or part of a sample genomic sequence (e.g., a sample genomic sequence, complementary DNA). Such a. sample nucleotide sequence may take the form of a sample genomic sequence from genomic DNA (gDNA), a transcriptomic sequence from complementary DNA (cDNA), a transcriptomic sequence from RNA, or other nucleotide sequence. In particular, a nucleotide read includes a determined or predicted sequence of nucleobase calls for a polynucleotide fragment (or group of monoclonal polynucleotide fragments) from a sequencing library corresponding to anucleotide read by generating nucleobase calls for nucleobases passed through a nanopore of a nucleotide-sample slide, determined via fluorescent tagging, or determined from a well in a flow cell. In some cases, a nucleotide read can refer to a particular type of read, such as a nucleotide read synthesized from sample library fragments that are shorter than a threshold number of nucleobases (e.g., SBS reads). In these or other cases, another type of nucleotide read can refer to (i) assembled nucleotide reads that have been assembled from shorter nucleotide reads to form a contiguous sequence (e.g., assembled nucleotide reads) satisfying a threshold number of nucleobases, (ii) circular consensus sequencing (CCS) reads satisfying the threshold number of nucleobases, or (iii) nanopore long reads satisfying the threshold number of nucleobases.

[0068] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood of the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.

[0069] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).methods such as Burrows- Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BEAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0071] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.

[0072] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.

[0073] As further used herein, the term “genomic coordinate” (or sometimes simply “coordinate”) refers to a particular location or position of a nucleobase within a genome (e.g,, an organism’s genome or a reference genome). In some cases, a genomic coordinate includes an identifier for a particular chromosome of a genome and an identifier for a position of a nucleobase within the particular chromosome. For instance, a genomic coordinate or coordinates may include a number, name, or other identifier for a somatic or sex chromosome (e.g., chrl or chrX) and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl : 1234570 or chrl : 1234570-1234870). In some cases, a genomic coordinate refers to a genomic coordinate on a sex chromosome (e.g., chrX or chrY). Consequently, the mapping system can determine genotype probabilities for a genotype call (e.g., a variant call) for a genomic coordinate on a sex chromosome. Further, in certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mitochondrial DNA reference genome or S ARS-Co V-2 for a reference genome for the SARS- CoV-2 virus) and a position of a nucleobase within the source for the reference genome (e.g., mt: 16568 or SARS-CoV-2:29001). By contrast, in certain cases, a genomic coordinate referssource (e.g., 29727).

[0074] As used herein, a “genomic region” refers to a range of genomic coordinates. Like genomic coordinates, in certain implementations, a genomic region may be identified by an identifier for a chromosome and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl: 1234570-1234870). In various implementations, a genomic coordinate includes a position within a reference genome. In some cases, a genomic coordinate is specific to a particular reference genome. Reiatedly, as used herein, the term “reference span” refers to a span of nucleobase positions within a linear reference genome. In other words, a reference span includes a span of nucieobases between two respective genomic coordinates of the linear reference genome.

[0075] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucieobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith-Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.

[0076] Reiatedly, as used herein, the term “mapping quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In some embodiments, for example, a mapping-quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong] , rounded to the nearest integer. In the alternative to a mean or median mapping quality, in some implementations, a mapping quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.prediction of a particular genotype of a genomic sample or a sample nucleotide sequence at a genomic locus. In particular, a genotype call can include a prediction of a particular genotype of a genomic sample with respect to a reference genome or a reference sequence at a genomic coordinate or a genomic region. For instance, in some cases, a genotype call includes a determination or prediction that a genomic sample comprises both a nucleobase and a complementary nucleobase at a genomic coordinate that is either homozygous or heterozygous for a reference base or a variant (e.g., homozygous reference bases represented as 0|0 or heterozygous for a variant on a particular strand represented as Oil). Accordingly, a genotype call can include a prediction of a variant or reference base for one or more alleles of a genomic sample and indicate zygosity with respect to a variant or reference base. A genotype call is often determined for a genomic coordinate or genomic region at which an SNP, insertion, deletion.

[0078] As used herein, the term “variant” refers to a nucleobase or multiple nucleobases that do not align with, differ from, or vary? from a corresponding nucleobase (or nucleobases) in a reference sequence or a reference genome. For example, a variant includes a SNP, an indel, or a structural variant that indicates nucleobases in a sample nucleotide sequence that differ from nucleobases in corresponding genomic coordinates of a reference sequence.

[0079] Along these lines, a “variant call” (or “variant nucleobase call”) refers to a nucleobase call comprising a mutation or a variant at a particular genomic coordinate or genomic region with respect to a reference. In particular, a variant call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate or region that differs from a reference nucleobase (or sequence of reference nucleobases) at the same genomic coordinate or region within a reference genome. Conversely, a “non-variant call” (or “non-variant nucleobase call” or “reference call”) refers to a nucleobase call comprising a non-variant or a reference nucleobase at a genomic coordinate or a genomic region with respect to a reference. In particular, a non- variant or reference call includes a determination or prediction that a genomic sample comprises a particular nucleobase (or sequence of nucleobases) at a genomic coordinate orgenomic coordinate or region within a reference genome.

[0080] As used herein, the term "solid support" refers to a rigid substrate that is insolubie in aqueous liquid. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g. due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, poly imides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solid supports for some embodiments are located within a flow cell apparatus. Exemplar}' flow cells are set forth in further detail below.

[0081] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow' of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026, WO 91 / 06678; WO 07 / 123744, US 7,329,492; US 7,211 ,414, US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.

[0082] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays wall have a variable or random spatial pattern of features. Alternatively , a solid support used m a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns arepatterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about lnim2, 500 gm2, 100 gm2, 25 gm2, 10 gm2, 5 gm2, 1 gm2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nni2, 1 gm2, 2.5 pm2, 5 pm2, 10 gm2, 100 gm2, or 500 gni2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.

[0083] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially fiat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.

[0084] As an example, in some embodiments, fragments derived from a long nucleic acid molecule captured at the surface of a flow cell occur in a line across the surface of the flow cell (e.g. if the nucleic acid was stretched out prior to fragmentation or amplification) or in a cloud on the surface. Further, a physical map of the immobilized nucleic acid can then be generated. The physical map thus correlates the physical relationship of clusters after the immobilized nucleic acid is amplified. Specifically, the physical map is used to calculate the probability that sequence data, obtained from any two clusters are linked, as described m the incorporated materials of WO 2012 / 025250. Alternatively or additionally, the physical map can be indicative of the genome of a particular organism in a. metagenomic sample. In this latter case the physical map can indicate the order of sequence fragments in the organism's genome; however, the order need not be specified and instead the mere presence of two or more fragments in a common organism (or other source or origin) can be sufficient basis for a physical map that characterizes a mixed sample and one or more organisms therein.

[0085] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like.examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.

[0086] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element-containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent. complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.

[0087] As used herein, the term "transposon element." is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-type or mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of USelements can comprise any nucleic acid or nucleic acid analog suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DMA, RNA, modified bases, non-natural bases, modified backbone, and can comprise nicks in one or both strands.

[0088] A standard NGS sequencing run y ields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library’ preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally’ be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (e.g., X and Y coordinates obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads in ambiguous positions.

[0089] In one or more embodiments, the mapping system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.

[0090] Moreover, in one or more embodiments, one or more sequencing data files in which the mapping system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.

[0091] The disclosed systems and methods may process data resulting from sequencing polynucleotides using massively parallel or next-generation sequencing processes to characterize a linking quality model, which can be used to determine a linking quality metric or score between reads or read pairs. This linking quality score may be used to quantify the linking quality between read pairs in various applications. Additional details regarding how the linking information in fragmented nucleic acid samples is preserved are described in U.S. Patent Number 10,246,746, the disclosure of which is incorporated herein by reference.

[0092] After or during a sequencing process, once the nucleotide sequence of each cluster has been determined, the disclosed methods may obtain and map a batch of sequence reads, or read pairs, from a first region on the flow cell using a standard mapping process (which does not use any linking information or connectivity' information), such as a process based on the Needleman-Wunsch algorithm or the Smith-Waterman algorithm. For example, the system may obtain all of the read pairs from clusters within a specific region of the flow cell, (e.g., a rectangular area) and then map those read pairs to a reference genome. Following this standard mapping process, the mapped location of the read pair to the reference genome may temporarily be used to estimate the “true” position of the read pair on the reference genome.

[0093] For example, some read pairs may map to the reference genome with a high mapping quality (e.g., having a high MAPQ) and other read pairs may be unable to be aligned at all. For example, a read pair which corresponds to a highly repeated region of the genome may not be easily mapped to the reference genome because the nucleotide sequence of the read pair is found in multiple locations within the reference genome. By knowing that the cluster for the un-mapped read pair was located on the flow cell adjacent to a second cluster for a second read pair that was mapped with a high quality to a specific repeated region in the reference genome may allow the system to determine that the specific repeated region is the most likely correct location of the un-mapped read pair on the reference genome, as adjacent clusters on the flow' cell may be more likely to come from nearby locations on the genome, for example, under sequencing conditions in which polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell.pair on a reference genome, the system may then determine a statistical model that can be used to create a score of the linking quality between read pairs obtained from the flow cell. For example, determining the statistical model may involve fitting the statistical model to the data from the first round of sequence mapping. The fitted, fully characterized statistical model may- then be used to improve the accuracy of assigning each read pair to the reference genome, in some embodiments.

[0095] After determining the appropriate statistical model, and then fitting the statistical model to the data from the first round of sequence mapping, the disclosed methods can then determine a linking quality score between two read pairs that were sequenced in the same sequencing run from a second region on the same flow cell, utilizing the fitted linking quality statistical model. The linking quality score can be calculated as a function of flow' cell proximity and genomic distance, quantifying how- likely that these read pairs came from the same target nucleic acid molecule. The linking quality score may be utilized or applied in analyses such as phasing, mapping, split alignment, etc.

[0096] A technical problem in previous approaches for quantifying linking quality is that the linking quality behave differently not only from sequencer to sequencer, but also from sequencing run to sequencing run, which may impact secondary analyses as well as tertiary analyses. Details regarding secondary' and tertiary' analyses can be found in “Pereira, R., Oliveira, J., & Sousa, M. (2020). Bioinformatics and computational tools for nextgeneration sequencing analysis in clinical genetics. Journal of clinical medicine, 9(1), 132”, the disclosure of which is incorporated herein by reference. Therefore, in some embodiments, the Unking quality statistical model is fitted (i.e., the parameters in the model are quantified) for each sequencing run. Fitting the model for each sequencing run can help evaluate run-to- run variability of the sequencing process and improve the sequencing reactions and conditions so that the linking information among fragments is better preserved (i.e., having higher link qualities among reads or read pairs).

[0097] In one aspect, the disclosed technology provides systems and methods for determining linking quality between two read pairs. Specifically, an example method 300 that can determine a linking quality score is shown in FIG. 2A. The method 300 may start at a start step and move to step 301 to obtain a first set of read pairs from a flow cell, which mayexample, the first set of read pairs may have been sequenced in a sequencing run using a first nucleic acid sample. The process 300 then may move to a step 302 wherein an initial mapping and mapping quality of each read pair on the flow cell may be determined by mapping the read pairs with a reference sequence. In some embodiments, mapping the read pairs with a reference sequence may be performed by the ILLUMINA DRAGEN® data analysis tools. In some embodiments, the first set of read pairs are considered to be properly mapped if the mapping quality score (MAPQ) of the read pair is determined to be greater than 30, 40 or 60. Of course, other MAPQ scores, such as greater than 30, 40, 50 or more may also be used as a cutoff for determining whether a particular read pair has been properly mapped to the reference genome. In some embodiments, the process also determines the distance between clusters of nucleic acids for their associated read pairs on the flow cell. For example, the process may identify each set of read pairs whose clusters are located within a specific distance from one other on the flow cell and are also within a particular genomic distance on the mapped nucleic acid. As one example, the system may determine clusters associated with read pairs whose corresponding clusters are 1000 flow cell units of one another on the flow cell and which are within at least 200 kb of one another on the same nucleic acid fragment as read pairs which may be linked to one another. In different sequencing instruments, one flow cell unit may correspond to about 1 micron, 20 microns, 40 microns, 60 microns, 80 microns, 100 microns, 200 microns, or any value therebetween, or may be defined as the distance between sequencing wells on the flow cell. Of course, other distances may also be contemplated as thresholds for determining whether two read pairs are linked to one another. For example, clusters associated with read pairs which are within a threshold of 100, 200, 300, 400, 500, 600, 700, 800, 900 or more flow cell units from each other may be considered to be linked. Similarly, sets of read pairs which are located at least 50kb, 75kb, lOOkb, 150kb, 200kb, 250kb or more base pairs apart from each other may be considered to be linked.

[0098] The method 300 may then move to step 303 to fit a linking quality statistical model to the data. Data used to fit the model may have been generated from the first set of read pairs. In particular, the statistical model may estimate the link probabilityreads linked), which is the probability that, for two read pairs derived from the same nucleic acid fragment (e.g., a genomic fragment) at two regions separated by a giventhe two read pairs would be separated by a relative displacement (e.g., a vector separation) on the flow cell. The genomic distance between the two read pairs can be obtained or estimated by preliminarily aligning or mapping the two read pairs to a reference sequence as mentioned above.

[0099] The procedure for fitting the statistical model may include using an optimizer for weighted non-linear least squares to obtain model parameters for the statistical model. In some embodiments, the optimizer is based on the Gauss-Newton or Levenberg- Marquardt algorithm. For example, the optimizer may be MINPACK. The result of step 303 (i.e., the fitted statistical model) may be a probability as a function of (i) the genomic distance of two read pairs and (ii) the relative displacement between two clusters on the flow cell associated with the two read pairs, with model parameters that have been fully quantified.

[0100] Prior to fitting the linking quality statistical model to data generated from the first set of read pairs, the data may first be processed and represented in the form of a two- dimensional histogram. Specifically, the two-dimensional histogram may plot the count of the relative displacement between two clusters on the flow cell associated with two read pairs in the first set of read pairs. In particular, the two read pairs being ploted may have been derived from two regions on the same nucleic acid fragment, and the two regions were separated by a given genomic distance on that nucleic acid fragment. For example, the read pairs being ploted may be read pairs whose clusters are 1000 flow cell units of one another on the flow cell and which are within at least 200 kb of one another on the same nucleic acid fragment. In some cases, the contour lines of the histogram comprise an oval-like shape (e.g., see the y-axis component discussed in connection with FIG. 3B). Moreover, without being bound by theory, the major axis of the oval-like shape may align with the direction of flow of sequencing reaction reagents on the flow cell. Additionally or alternatively, the oval-like shape may center at the origin of the histogram. In some cases, the contour lines of the histogram comprise a circle (e.g., see the hotspot or diffusion component discussed in connection with FIG. 3B), and the circle may center at the origin in some examples. In some cases, the contour lines of the histogram comprise a pair of “blobs” (e.g., see the poles component discussed in connection with FIG. 3B). For example, the pair of blobs at the poles may be symmetric with respect to the x-axis and the y -axis of the histogram, and without being bound by theory, the y-axis ofcell. Additionally or alternatively, each blob of the pair of blobs may be a distance away from the x-axis of the histogram by a predetermined amount.

[0101] The method 300 may then move to step 304 to determine the linking quality score between two read pairs in a second set of read pairs from the flow' cell by applying the statistical model to the second set of read pairs. The determination may be based on the statistical model obtained from step 303, the genomic distance of the two read pairs in the second set, and the relative displacement between two clusters on the flow cell associated with the two read pairs in the second set. For example, the fitted statistical model, or the link probability p(&x, Ay, g,- 1 reads linked), obtained at step 303 (based on data generated from the first set of read pairs) is also applied to the case of the second set of read pairs (which may be distinct from the first set of read pairs). The second set of read pairs may have been sequenced in the same sequencing run using a second sample. In particular, the determination may simply involve inserting (i) the genomic distance of the two read pairs in the second set and (ii) the relative displacement between two clusters on the flow cell associated with the two read pairs in the second set into the fitted statistical model, or the link probability' Ay, g,- 1 reads linked), obtained at step 303 to determine the linking quality between the second set of read pairs. The genomic distance between the two read pairs can be obtained or estimated by preliminarily aligning or mapping the two read pairs to the reference sequence. In some embodiments, the first and second sets of read pairs were obtained from a BAM file, and the determined linking quality score by applying the statistical model may be added to the BAM file. The method 300 then ends at an end step.

[0102] The results (i.e., the determined linking quality scores) of the method 300 may be optionally used in various applications, including phasing, mapping, and split alignment, winch will be described in more detail in connection with FIG. 2B, FIG, 2C, and FIC. 2D below, respectively. Specifically, at the decision block, the disclosed systems and methods may determine which application (method 400, method 500, or method 600) will be selected.

[0103] In another aspect, the disclosed technology provides a method 400 for assigning read pairs to target nucleic acid polymers (phasing), as shown in FIG. 2B. The method 400 may start at a start step and move to step 401 to fragment the target nucleic acidmay then become captured on the flow cell under conditions wherein polynucleotide fragments from a common target nucleic acid polymer are preferentially captured at proximal locations on the flow cell. In some cases, the polynucleotide fragments may passively diffuse to the locations on the flow cell. In other cases, the polynucleotide fragments may be actively transported to the locations on the flow cell.

[0104] In some cases, step 401 comprises stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer. Additionally or alternatively, step 401 may include adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. Specifically, the inserts may include cleavage sites, so that step 401 may further include cleaving the inserts at the cleavage sites. The inserts may be added into the target nucleic acid polymers by transposases. For example, each of the inserts may comprise a first transposon element and a second transposon element that is contiguous with the first transposon element. The transposases may comprise ligands and the flow cell’s surface may comprise receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, step 401 further comprises attaching the modified nucleic acid polymers to the flow cell prior to producing the fragments. Moreover, the transposases may either be removed from the modifi ed nucleic acid polymers prior to or after the attaching of the modified nucleic acid polymer to the flow cell.

[0105] The method 400 may then move to step 402 to sequence the polynucleotide fragments to generate read pairs. In some embodiments, method 400 further comprises amplifying the polynucleotide fragments at the locations to produce amplified polynucleotide fragments, and step 402 may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. Specifically, the amplification may include extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. Additionally or alternatively, the amplification may include extending at least two primer species in a bridge amplification technique.reference sequence to determine a position on the reference sequence for each read pair. A standard alignment process can be used at tins stage. This alignment can give rise to estimated genomic distances between two read pairs on the reference sequence.

[0107] The method 400 may then move to step 404 to assign two read pairs to the same target nucleic acid polymer based on the linking quality score between the two read pairs determined according to the method 300 described above. In alternative embodiments, however, the linking quality score used in step 404 may not necessarily be determined according to method 300. For example, the linking quality score may be directly calculated by inputting the results of steps 402 and 403 (e.g., estimated genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexisting / stored / received linking quality’ statistical model (e.g., steps 301 and 302 may be omitted in the method 303). The method 400 then ends at an end step.

[0108] In another aspect, the disclosed technology provides a method 500 of producing a mapping of read pairs to a reference sequence, as shown in FIG. 2C. The method 500 may start at a start step and move to step 501 to fragment the target nucleic acid polymers in a flow cell to generate polynucleotide fragments. These polynucleotide fragments may then become captured on the flow cell under conditions wherein polynucleotide fragments from a common target nucleic acid polymer are preferentially captured at proximal locations on the flow cell. In some cases, the polynucleotide fragments may passively diffuse to the locations on the flow cell. In other cases, the polynucleotide fragments may be actively transported to the locations on the flow cell.

[0109] In some cases, step 501 comprises stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer. Additionally or alternatively, step 501 may include adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. Specifically, the inserts may include cleavage sites, so that step 501 may further include cleaving the inserts at the cleavage sites. The inserts may be added into the target nucleic acid polymers by transposases. For example, each of the inserts may comprise a first transposon element and a second transposon element that is contiguous with the first transposon element. The transposases may comprise ligands and the flow cell’s surface may comprise receptors that bind to the ligands to attachembodiments, step 501 further comprises ataching the modified nucleic acid polymers to the flow cell prior to producing the fragments. Moreover, the transposases may either be removed from the modified nucleic acid polymers prior to or after the attaching of the modified nucleic acid polymer to the flow cell.

[0110] The method 500 may then move to step 502 to sequence the polynucleotide fragments to generate read pairs. In some embodiments, method 500 further comprises amplifying the polynucleotide fragments at the locations to produce amplified polynucleotide fragments, and step 502 may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. Specifically, the amplification may include extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. Additionally or alternatively, the amplification may include extending at least two primer species in a bridge amplification technique.

[0111] The method 500 may then move to step 503 to align the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. A standard alignment process can be used at this stage. This alignment can give rise to estimated genomic distances between two read pairs on the reference sequence.

[0112] The method 500 may then move to step 504 to produce the mapping utilizing the linking quality score. In one embodiment, step 504 may involve selecting a first position among the plurality of candidate positions on the reference sequence for a first read pair based on the linking quality score between the first read pair and a second read pair determined according to the method 300 described above. For example, the selected first position may correspond to the highest linking quality score. Moreover, the second read pair may have an alignment score associated with a second position on the reference sequence that is above a predetermined value (for example, 75% of nucleobases matching, 80% of nucleobases matching, 85% of nucleobases matching, 90% of nucleobases matching, 95% of nucleobases matching, 99% of nucleobases matching, or any value therebetween). Alternatively, step 504 may involve selecting a first position among the plurality of candidatefirst position, (ii) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and (lii) the linking quality score between the first read pair and the second read pair determined according to the method 300 described above. In some embodiments, the function may be a weighted sum and the weights for (i), (ii), and (hi) may be in the range of about 0.2--0.4, 0.2 0.4. and 0.3---0.5, respectively. For example, the selected first position may correspond to the maximum of that function. In alternative embodiments, however, the linking quality score used in step 504 may not necessarily be determined according to method 300. For example, the linking quality score may be directly calculated by inputting the results of steps 502 and 503 (e.g., estimated genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexisting / stored / received linking quality statistical model (e.g., steps 301 and 302 may be omitted in the method 303). The method 500 then ends at an end step.

[0113] In another aspect, the disclosed technology provides a method 600 of improving accuracy of split-alignments of read pairs to a reference sequence, as shown in FIG. 2D. The method 600 may start at a start step and move to step 601 to fragment the target nucleic acid polymers in a flow cell to generate polynucleotide fragments. These polynucleotide fragments may then become captured on the flow cell under conditions wherein polynucleotide fragments from a common target nucleic acid polymer are preferentially captured at proximal locations on the flow cell. In some cases, the polynucleotide fragments may passively diffuse to the locations on the flow cell. In other cases, the polynucleotide fragments may be actively transported to the locations on the flow cell.

[0114] In some cases, step 601 comprises stretching the target nucleic acid polymer along the flow cell prior to producing the polynucleotide fragments of the target nucleic acid polymer. Additionally or alternatively, step 601 may include adding inserts into the target nucleic acid polymers to form modified nucleic acid polymers. Specifically, the inserts may include cleavage sites, so that step 601 may further include cleaving the inserts at the cleavage sites. The inserts may be added into the target nucleic acid polymers by transposases. For example, each of the inserts may comprise a first transposon element and a second transposon element that is contiguous with the first transposon element. The transposases may comprisethe modified nucleic acid polymers to the surface prior to producing the fragments. In some embodiments, step 601 further comprises attaching the modified nucleic acid polymers to the flow cell prior to producing the fragments. Moreover, the transposases may either be removed from the modified nucleic acid polymers prior to or after the attaching of the modified nucleic acid polymer to the flow cell.

[0115] The method 600 may then move to step 602 to sequence the polynucleotide fragments to generate read pairs. In some embodiments, method 600 further comprises amplifying the polynucleotide fragments at the locations to produce amplified polynucleotide fragments, and step 602 may comprise determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations. Specifically, the amplification may include extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species. Additionally or alternatively, the amplification may include extending at least two primer species in a bridge amplification techni que.

[0116] The method 600 may then move to step 603 to align the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair. A standard alignment process can be used at this stage. This alignment can give rise to estimated genomic distances between two read pairs on the reference sequence.

[0117] The method 600 may then move to step 604 to select a “winning” alignment utilizing the linking quality score for a first read pair that, is split-aligned (e.g., the first read pair comprising different segments that align to disjoint regions in the reference sequence). In one embodiment, step 604 may involve selecting a first position among the plurality of candidate positions on the reference sequence for a first segment, of the first read pair, based on the linking quality score between the first segment and a second read pair determined according to the method 300 as described above. For example, the selected first position may correspond to the highest linking quality score. Moreover, the second read pair may have an alignment score associated with a second position on the reference sequence that is above a predetermined value (for example, 75% of nucleobases matching, 80% of nucleobasesmatching, 99% of nucleobases matching, or any value therebetween). Alternatively, step 604 may involve selecting a first position among the plurality of candidate positions on the reference sequence for a first segment of the first read pair based on a function of (i) the alignment score associated with the first segment, (ii) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and (lii) the linking quality score between the first segment and the second read pair determined according to the method 300 as described above. In some embodiments, the function may be a weighted sum and the weights for (i), (ii), and (iii) may be in the range of about 0.2-0.4, 0.2-0.4, and 0.3-0.5, respectively. For example, the selected first position may correspond to the maximum of that function. In alternative embodiments, however, the linking quality score used in step 604 may not necessarily be determined according to method 300. For example, the linking quality' score may be directly calculated by inputting the results of steps 602 and 603 (e.g., estimated genomic distance between two read pairs and relative displacement between two clusters on the flow cell associated with the two read pairs) to a preexisting / stored / received linking quality statistical model (e.g., steps 301 and 302 may be omitted in the method 303). The method 600 then ends at an end step.Experimental Observations

[0118] FIG. 3A show's experimental data obtained from a sequencing run in a first HXUMINA® sequencing platform. Each sub-figure is a two-dimensional histogram of the count (n) of the relative displacement (XDIST, YDIST) between two clusters on the flow cell, where the two clusters are associated with two read pairs that either were derived from the same originating nucleic acid molecule or landed nearby on the flow? cell by chance, and that are found to be at a given genomic separation distance. The data included in the histogram come from a random selection of read-pairs from the result of the sequencing run (for example, 1000 '24 read-pairs per tile are randomly selected), and for each selected read-pair, the clusters within a 2000 flow cell unit “search disc’1(i.e., the radius being considered) around the cluster associated with the selected read-pair are considered. The genomic separation distance (3000) is labeled on the top panel of the sub-figure, and the subfigures run from a genomic separation distance of 5000 base pairs to a genomic separation distance of 95000 base pairs. The genomicof the read pairs to a reference genome sequence (e.g., similar to step 302 of process 300). As shown in the figure, different shades are used to represent different ranges of the count (n). The relative displacement between clusters (XDIST, YDIST) is measured in flow cell units. The high counts near the center of the histograms (for example, as indicated by reference numeral 3001 ) show that the read-pairs in these data are linked, as compared to scenarios where the data are not linked and the histograms would be generally flat.

[0119] FIG. 3B again show's experimental data obtained from a sequencing run in the Illumina sequencing platform, but at a higher resolution. FIG. 3B shows two sub-figures, with a first sub-figure showing a genomic separation distance of 5000 base pairs and a second sub-figure showing a genomic separation distance of 40,000 base pairs. Each sub-figure is a two-dimensional histogram of the count (n) of the relative displacement (XDIST, YDIST) between two clusters on the flow cell, where the two clusters are associated with two read pairs that either were derived from the same originating nucleic acid molecule or landed nearby on the flow cell by chance, and that are found to be at a given genomic separation distance. The data included in the histogram come from a random selection of read-pairs from the result of the sequencing run, and for each selected read-pair, the clusters within a 2000 flow cell unit search disc around the cluster associated with the selected read-pair are considered. The genomic separation distance (3000) is labeled on the top panel of the sub-figure. The genomic distance is measured in base pairs and may be determined by preliminary alignment of the read pairs to a reference genome sequence (e.g., similar to step 302 of process 300). Different shades are used to represent different ranges of the count (n). The relative displacement (XDIST, YDIST) is measured in flow cell units. Each histogram can be viewed as a sum-up of four components as indicated in the figure.

[0120] The first component (3501) of the histogram, denoted as “hotspot”, is characterized by small flow cell displacements and circular symmetry, and its boundary appears to grow with the genomic distance. The hotspot component constitutes about 50-100% of the total counts. The second component (3502), denoted “y-axis”, is characterized by having small x-axis displacements and a large range of y-axis displacements. The y-axis component constituted about 0-25% of the total counts. The third component (3503), named “diffusion”, has large flow cell displacements and circular symmetry. The diffusion component constitutesaxis displacements while its y-axis displacements are near the maximal extent. The poles component constitutes about 0-25% of the total counts. The percentage of each component may change from instrument to instrument and / or run to run.

[0121] FIG. 4 shows experimental data obtained from a sequencing run in a second Illumina sequencing platform. Each sub-figure is a two-dimensional histogram of the count (n) of the relative displacement (XDIST, YDIST) between two clusters on the flow cell, where the two clusters are associated with two read pairs that either were derived from the same originating nucleic acid molecule or landed nearby on the flow' cell by chance, and that are found to be at a given genomic separation distance. The data included in the histogram come from a random selection of read-pairs from the result of the sequencing run, and for each selected read-pair, the clusters within a 2000 flow cell unit search disc around the cluster associated with the selected read-pair are considered. The genomic separation distance (4000) is labeled on the top panel of the sub-figure. The genomic distance is measured in base pairs and may be determined by preliminary alignment of the read pairs to a reference genome sequence (e.g., similar to step 302 of process 300). As shown in the figure, different shades are used to represent different ranges of the count (n). The relative displacement (XDIST, YDIST) is measured in flow' cell units. The high counts near the center of the histograms (4001) show that the read-pairs in these data are linked, as compared to scenarios where the data are not linked and the histograms would be generally flat. There appeared to be many more counts in the diffusion component (4503) compared to data obtained from the first Illumina platform. The y-axis component is not obvious, and there appears to be an additional small x-axis component (4505). Without being bound by theory, the differences in the results may be due to differences in the sample preparations used in the different sequencing platforms, such as the direction and magnitude of reagent flows, etc.

[0122] Without being bound by theory, these experimental data shown in FIGs. 3 A, 3B and 4 are consistent with sequencing conditions in which I ) the target nucleic acid molecule can be in configurations that range from linear to tightly packed; 2) a y-axis aligned flo w exists when reagents are applied to the flow cell; and / or 3) the target nucleic acids can diffuse. The disclosed linking quality statistical models can accurately capture the features observed from these experimental data across various sequencing platforms. Moreover, it has been discovereddistance can differ in linking quality due to different orientations of the relative displacement on the flow cell, and the disclosed models can accurately capture this feature as well.Example Linking Quality Statistical Model

[0123] Given a set of candidate genomic distances between two read pairs {,g;} (which may be obtained from the alignment or assembly of the read pairs), the statistical model for linking quality7describes the link distribution functionLinkDisplacementPDF(Ag,-, reads linked ) for all i,

[0124] where (Ax, Ay) or (XDIST, YDIST) is the flow-cell relative displacement between the two clusters for the two read pairs and is measured in flow cell units in FIGs. 5-8 described below.

[0125] Based on the above experimental observations, an example model was constructed to include in the link distribution function LinkDisplacementPDF the four different components mentioned in connection with FIG. 3B. Thus, the link distribution function is a weighted sum of a hotspot component, a diffusion component, a y-axis component, and a poles component. The weights are model parameters that were quantified by fitting the model to the experimental data. The weight for the hotspot component, n hotspot, may be interpreted as the number of proximal read pairs in the search disc on the flow cell attributed to the hotspot component. The weight for the diffusion component, n diffusion, may be interpreted as the number of proximal read pairs in the search disc on the flow cell attributed to the diffusion component. The weight for the y-axis component, n y-axis, may be interpreted as the number of proximal read pairs in the search disc on the flow7cell attributed to the y-axis component. The weight for the poles component, n poles, may be interpreted as the number of proximal read pairs in the search disc on the flow7cell attributed to the poles component.

[0126] The hotspot component is modelled byd(Ay, AvJ ™ - - - where N is the two-dimensional normal distribution, d is the flow cell distance from (0, 0) to (Ax, Ay), and the hotspot varianceincreases with increasing genomic distance g, in which awspot and hotspot are constant model parameters. The variance intercept ahotspot is the constant, intercept parameter in the abovegenomic distance coefficient footspot is rate at which the hotspot variance increases with genomic distance. Each sub-figure of FIG. 5 shows the distribution p of the hotspot component as a function of Axwhen Ay= 0 for a given genomic distance g, along with experimental data.The genomic distance (5000) is labeled on the top panel of each sub-figure.

[0127] The diffusion component is modelledwhere d is the flow cell distance from (0, 0) to (Ax, Ay), and J? is a constant model parameter. Each sub-figure of FIG. 6 shows the distribution p of the diffusion component as a function of Ax when Ay= 0 for a given genomic distance g, along with experimental data. The genomic distance (6000) is labeled on the top panel of each sub-figure. Experimental observations confirm that the diffusion component is independent of the genomic distance.

[0128] The y-axis component is modelled by two independent probability 1 sigmoidi~&(,g)OA.J -!•(#))) , . x > distributions= . ■.■■■■— ... ■■■■■■ ■■■ . "■■■■;■ . ■■■■■■ . and p(&rn) ~ M0, a ) that is 2( iAy|+a) log(rG)+«)-log (a) 'v vassumed to be independent of Ax, where r(g) ~ eg (where the constant of proportionality E is a model parameter, also called Y limit genomic distance coefficient, that characterizes the rate at which the maximal extent of the y-axis component along the y-axis increases with genomic distance) and g is the genomic distance between two read pairs, a is a positive, constant model parameter, also called Y offset parameter, that governs how sharp the y-axis component is along the y-axis. b(g~) = mg + £, where the Y limit drop-off intercept { governs how rapidly the y-axis component goes to zero near the maximal extent along the y-axis, and the Y limit drop-off intercept genomic distance coefficient m governs how the rapidity of dropoff of the y-axis component near the maximal extent increases with genomic distance. N is the normal distribution, and u2(g) ~ eg, where the X variance genomic distance coefficient c is the rate at which the variance er2of the x-coordinate of linked read pairs attributed to the y- axis component increases with genomic distance g. Without being bound by theory, this component may give rise from polynucleotide fragments that are derived from a somewhat linear originating nucleic acid molecule and are aligned along the y-axis due to flows of sequencing reagents. FIG. 7 shows the distributionthe y-axis component as a function of Aywhen the genomic distance is 50000 bp, along with experimental data. The sy mbols used in FIG. 7 are the same as those used in FIG. 6. As shown in FIG. 7, most of thedistribution decays similar to 1 / | Ay|.

[0129] The poles component is modelled by two independent probability distributionsTfmean = r(p), sd = s($)) , F is the Gamma distribution, r($) = pg (where the constant of proportionality p, also called Xydist mean genomic distance coefficient, is a model parameter that characterizes the rate at which the mean distance from the origin of the poles component increases with genomic distance), s2(g) = ^g (where the constant of proportionality also called Xydist variance genomic distance coefficient, is a model parameter that characterizes the rate at which the variance of the distance from the origin of the poles component increases with genomic distance), and g is the genomic distance, d and 0 are the radius and angle of the flow cell vector separation (Ax, Ay). p(6\g) is assumed to be angle independent and modelled with the Beta , 7 and having a standard deviation = t (alsocalled Angle concentration), which is a model parameter that governs how much of a full circle is filled by the poles component. Without being bound by theory, this component may give rise from polynucleotide fragments that are derived from a fully linear target nucleic acid molecule that is aligned along the y-axis due to reagents flows, but with ends free to swing. FIG. 8 shows the above parameters, r (reference numeral 8001), 5 (reference numeral 8002), and r (reference numeral 8003), indicated in an experimentally obtained two-dimensional histogram, of the count (n) of the relative displacement between two clusters on the flow cell associated with two read pairs that were derived from the same originating nucleic acid molecule at a genomic separation distance of 30000 bp.

[0130] To obtain the full marginal model for the link probability'g,- 1 reads linked), the link distribution functionLinkDisplacementPDF(Ax,reads linked ), expressed as a weighted sum of the four components, is multiplied with P(g J reads linked). The link probabilityg J reads linked) is then fitted to an experimentally obtained three-dimensional histogram, which is a collection of two-dimensional histograms, each representing the count (n) of the relative di splacement (XDIST, YDIST) between two clusters on the flow cell, where the two clusters are associated with two read pairs that either were derived from the samefound to be at a genomic separation distance g, within an 1000 bp bin. The data included in this three-dimensional histogram come from a random selection of read-pairs from the result of a sequencing run, and for each selected read-pair, the clusters within a 2000 flow cell unit search disc around the cluster associated with the selected read-pair are considered. P{gt | reads linked), the frequency of occurrence of the genomic distance g between two linked read pairs, can be modelled by a Gamma distribution that is parameterized by K shape, which is the shape parameter in the Gamma distribution, and Kjscale, which is the scale parameter in the Gamma distribution. FIG. 9A shows the experimentally obtained count of the genomic distance GDIST between two linked read pairs, and is fitted to a Gamma distribution. FIG. 9B shows a “background model” (9000) where the read pairs are not linked (i.e., the sequencing condition does not enable polynucleotide fragments from a common nucleic acid polymer to be preferentially captured at proximal locations on the flow cell) and compares it with the data shown in FIG. 9A. The distribution of the background model is generally flat and can be characterized by a model parameter Q, which is the un- linked background rate, i.e., the number of proximal but un-linked read pairs per 1000 bp genomic distance in the search disc.

[0131] Overall, this model used less than 20 parameters, and once these parameters have been estimated for a sequencing run, the linking quality’ score can be calculated with a closed-form expression.

[0132] In some embodiments, the library’ preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment required for the systems. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The sample preparation steps can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. In some embodiments, the complexed strand and transposome is in solution, and is then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposomecomplexes in solution become immobilized to the substrate.

[0133] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids are attached at fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.

[0134] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.

[0135] SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251 , U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.

[0136] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly, the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techniques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering amplificationcomprising components such as pumps, valves, reservoirs, fluidic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated sy stem can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification method set forth herein and for the deliver}' of sequencing reagents in a sequencing method such as those exemplified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to cany out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference.

[0137] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory-obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic or pathogenic sample. In some embodiments, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another embodiment, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus or fungus. In some embodiments, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0138] Further, the methods and compositions disclosed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degradedsamples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or more military services or any such personnel. The nucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be impregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue samples, autopsy or remains of a victim. In some embodiments, nucleic acids including one or more target sequences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target sequences are directed to purposes of human identification. In some embodiments, the disclosure relates generally to methods for identifying characteristics of a forensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target-specific primers disclosed herein or one or more targetspecific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.

[0139] FIG. 10 illustrates a schematic diagram of a computing system 200 in which a linking quality system 206 operates in accordance with one or more embodiments of the disclosed systems and methods. As illustrated, the computing system 200 includes a sequencing device 202 connected to a local device 108 (e.g., a local server device), one or more server device(s) 210, and a client device 214. As shown in FIG. 10, the sequencing device 202, the local device 208, the server device(s) 210, and the client device 214 can communicate "with each other via a network 218. The network 218 comprises any suitable network over which computing devices can communicate. While FIG. 10 show's an embodiment of theconfigurations below.

[0140] As indicated by FIG. 10, the sequencing device 202 comprises a computing device and a sequencing device system 204 for sequencing a genomic sample or other nucleic- acid polymer. In some embodiments, by executing the sequencing device system 204 using a processor, the sequencing device 202 analyzes polynucleotide fragments or oligonucleotides extracted from genomic samples to generate nucleotide reads or other data utilizing computer implemented methods and systems either directly or indirectly on the sequencing device 202. More particularly, the sequencing device 202 receives nucleotide-sample slides (e.g., flow cells) comprising polynucleotide fragments extracted from samples and further copies and determines the nucleobase sequence of such extracted polynucleotide fragments.

[0141] In one or more embodiments, the sequencing device 202 utilizes sequencing-by-synthesis (SBS) techniques to sequence polynucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. In addition or in the alternative to communicating across the network 218, in some embodiments, the sequencing device 202 bypasses the network 218 and communicates directly with the local device 208 or the client device 214. By executing the sequencing device system 204, the sequencing device 202 can further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send the BCL file to the local device 208 and / or the server device(s) 210.

[0142] As further indicated by FIG. 10, the local device 208 is located at or near a same physical location of the sequencing device 202. Indeed, in some embodiments, the local device 208 and the sequencing device 202 are integrated into a same computing device. The local device 208 may run the linking quality system 206 to generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As shown in FIG. 10, the sequencing device 202 may send (and the local device 208 may receive) base-call data generated during a sequencing run of the sequencing device 202. By executing software in the form of the linking quality system 206, the local device 208 may align nucleotide reads with a reference genome utilizing a BAM file 212 and determine genetic variants based on the aligned nucleotide reads. The local device 208 may also communicate with the client device 214. In particular, the local device 208 can sendformat (VCF) file, or other information indicating nucleobase calls, sequencing metrics, error data, or other metrics.

[0143] As further indicated by FIG. 10, the server device(s) 210 are located remotely from the local device 208 and the sequencing device 202. Similar to the local device 208, in some embodiments, the server device(s) 210 include a version of (or are otherwise able to access or implement) the linking quality system 206. Accordingly, the server device(s) 210 may generate, receive, analyze, store, and transmit digital data, such as by receiving base-call data or determining variant calls based on analyzing such base-call data. As indicated above, the sequencing device 202 may send (and the server device(s) 210 may receive) base-call data from the sequencing device 202. The server device(s) 210 may also communicate with the client device 214. In particular, the server device(s) 210 can send data to the client device 214, including BAM files, VCF files, or other sequencing-related information.

[0144] In some embodiments, the server device(s) 210 comprise a distributed collection of servers where the server device(s) 210 include a number of server devices distributed across the network 218 and located in the same or different physical locations. Further, the server device(s) 210 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0145] As indicated above, as part of the server device(s) 210 or the local device 208, the linking quality system 206 can generate, encode, and / or implement the BAM file 212 to determine alignments of nucleotide reads from a genomic sample with a reference genome. For instance, the linking quality system 206 can identify candidate alignments of one or more nucleotide reads with a primary contiguous sequence, generate primary alignment scores for the candidate alignments, and adjust the alignment scores based on population variant data indicated in the BAM file 212, as described in greater detail below in relation to the subsequent figures.

[0146] As further illustrated and indicated in FIG. 10, by executing a sequencing application 216, the client device 214 can generate, store, receive, and send digital data. In particular, the client device 214 can receive sequencing data from the local device 208 or receive call files (e.g., BCL) and sequencing metrics from the sequencing device 202. Furthermore, the client device 214 may communicate with the local device 208 or the serveras a base-call-quality metrics or pass-filter metrics. The client device 214 can accordingly present or display information pertaining to variant calls or other genotype calls within a graphical user interface of the sequencing application 216 to a user associated with the client device 214. For example, the client device 214 can present genotype calls, variant calls, and / or sequencing metrics for a sequenced genomic sample within a graphical user interface of the sequencing application 216.

[0147] Although FIG. 10 depicts the client device 214 as a desktop or laptop computer, the client device 214 may comprise various types of client devices. For example, in some embodiments, the client device 214 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 214 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0148] As further illustrated in FIG. 10, the client device 214 includes the sequencing application 216. The sequencing application 216 may be a web application or a native application stored and executed on the client device 214 (e.g., a mobile application, desktop application). The sequencing application 216 can include instructions that (when executed) cause the client device 214 to receive data from the structural -variant-aware linking quality system 206 and present, for display at the client device 214, base-call data or data from a VCF. Furthermore, the sequencing application 216 can instruct the client device 214 to display summaries for multiple sequencing runs.

[0149] As further illustrated in FIG. 10, a version of the linking quality system 206 may be located and / or implemented (e.g., entirely or in part) on the client device 214 or the sequencing device 202. In yet other embodiments, the linking quality system 206 is implemented by one or more other components of the computing system 200, such as the local device 208. In particular, the linking quality system 206 can be implemented in a variety of different ways across the sequencing device 202, the local device 208, the server device(s) 210, and the client device 214. For example, the linking quality system 206 can be downloaded from the server device(s) 210 to the linking quality system 206 and / or the local device 208 where all or part of the functionality of the linking quality system 206 is performed at each respective device within the computing system 200.purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer- readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0151] Embodiments of the disclosure methods can be implemented or performed by a machine, such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor or group of processors for performing the methods described herein may be of various types, including programmable devices (e.g., CPLDs and FPGAs) and non-programmable devices such as gate array ASICs or general-purpose microprocessors.

[0152] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are non-transitory computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least tw'O distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0153] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magneticdesired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0154] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly view's the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to cany desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0155] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0156] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general- purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above.claims.

[0157] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wareless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0158] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0159] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.

[0160] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspectsaccompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.

[0161] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Additional Notes

[0162] Various embodiments of the present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or mediums) having computer readable program instructions thereon for causing a processor to carry’ out aspects of the present disclosure.

[0163] For example, the functionality’ described herein may be performed as software instructions are executed by, and / or in response to software instructions being executed by, one or more hardware processors and / or any other suitable computing devices. The software instructions and / or other executable code may be read from a computer readable storage medium (or mediums). Computer readable storage mediums may also be referred to herein as computer readable storage or computer readable storage devices.

[0164] The computer readable storage medium can be a tangible device that can retain and store data and / or instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device (including any volatile and / or non-volatile electronic storage devices), asemiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0165] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0166] Computer readable program instructions (as also referred to herein as, for example, “code,” “instructions,’1“module,’1“application,” “software application,” and / or the like) for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the "C"instructions may be callable from other instructions or from itself, and / or may be invoked in response to detected events or interrupts. Computer readable program instructions configured for execution on computing devices may be provided on a computer readable storage medium, and / or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression or decryption prior to execution) that may then be stored on a computer readable storage medium. Such computer readable program instructions may be stored, partially or fully, on a memory’ device (e.g., a computer readable storage medium) of the executing computing device, for execution by the computing device. The computer readable program instructions may execute entirely on a user's computer (e.g., the executing computing device), partly on the user’s computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0167] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It wall be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0168] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block orreadable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart] s) and / or block diagram(s) block or blocks.

[0169] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer may load the instructions and / or modules into its dynamic memory and send the instructions over a telephone, cable, or optical line using a modem. A modem local to a server computing system may receive the data on the telephone / cable / optical line and use a converter device including the appropriate circuitry to place the data on a bus. The bus may carry' the data to a memory, from which a processor may retrieve and execute the instructions. The instructions received by the memory may optionally be stored on a storage device (e.g., a solid-state drive) either before or after execution by the computer processor.

[0170] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a sendee, module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. In addition, certain blocks may be omitted in some implementations. The methods and processes describedcan be performed in other sequences that are appropriate.

[0171] It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware- based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. For example, any of the processes, methods, algorithms, elements, blocks, applications, or other functionality (or portions of functionality) described in the preceding sections may be embodied in, and / or fully or partially automated via, electronic hardware such application-specific processors (e.g., application-specific integrated circuits (ASICs)), programmable processors (e.g., field programmable gate arrays (FPGAs)), application-specific circuitry’, and / or the like (any’ of which may also combine custom hard-wired logic, logic circuits, ASICs, FPGAs, etc. with custom programming / execution of software instructions to accomplish the techniques).

[0172] Any of the above-mentioned processors, and / or devices incorporating any of the above-mentioned processors, may be referred to herein as, for example, “computers,” “computer devices,” “computing devices,” “hardware computing devices,” “hardware processors,” “processing units,” and / or the like. Computing devices of the above-embodiments may generally (but not necessarily) be controlled and / or coordinated by operating system software, such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating systems. In other embodiments, the computing devices may be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality, such as a graphical user interface (“GUI”), among other things.

[0173] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it ismanner in the various examples unless the context clearly dictates otherwise.

[0174] It is to be understood that the ranges provided herein include the stated range and any value or sub-range within the stated range, as if such value or sub-range were explicitly recited. For example, a range from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly recited limits of from about 2 kbp to about 20 kbp, but also to include individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc. Furthermore, when “about” and / or “substantially” are / is utilized to describe a value, this is meant to encompass minor variations (up to + / - 10%) from the stated value.

[0175] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.

[0176] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the methods described herein may be made without departing from the spirit of the disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.

[0177] Features, materials, characteristics, or groups described in conjunction with a particular aspect, or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not restricted to the details of any foregoing examples. The protection extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations, one or more features from a claimed combination can, in some cases, be excised from the combination, and the combination may be claimed as a sub-combination or variation of a sub-combination.

[0179] Moreover, while operations may be depicted in the drawings or described in the specification in a particular order, such operations need not be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated in the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some examples, the actual steps taken in the processes illustrated and / or disclosed may differ from those shown in the figures. Depending on the example, certain of the steps described above may be removed or others may be added. Furthermore, the features and attributes of the specific examples disclosed above may be combined in different ways to form additional examples, all of which fall within the scope of the present disclosure.

[0180] For purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not necessarily all such advantages may be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.

[0181] Conditional language, such as “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more examplesor prompting, whether these features, elements, and / or steps are included or are to be performed in any particular example.

[0182] Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to convey that an item, term, etc. may be either X, Y, or Z. Thus, such conjunctive language is not generally intended to imply that certain examples require the presence of at least one of X, at least one of Y, and at least one of Z.

[0183] Language of degree used herein, such as the terms “approximately,” “about,” “generally,” and “substantially” represent a value, amount, or characteristic close to the stated value, amount, or characteristic that still performs a desired function or achieves a desired result.

[0184] The scope of the present disclosure is not intended to be limited by the specific disclosures of preferred examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as non-exclusive.

Claims

1. A method of determining a linking quality score between two read pairs, comprising:(a) obtaining a first set of read pairs from a flow cell comprising clusters of polynucleotides bound to locations on the flow cell;(b) fitting a linking quality statistical model to data generated from the first set of read pairs, wherein the statistical model estimates the probability that, for two read pairs derived from the same nucleic acid fragment at two regions separated by a given genomic distance, two clusters on the flow cell associated with the two read pairs would be separated by a relative displacement; and(c) determining the linking quality score between two read pairs in a second set of read pairs from the flow cell based on the statistical model, the genomic distance of the two read pairs, and the relative displacement between two clusters on the flow cell associated with the two read pairs.

2. The method of claim 1, wherein the first set of read pairs is sequenced in a sequencing run using a first sample.

3. The method of claim 2, wherein the second set of read pairs is sequenced in the same sequencing run using a second sample.

4. The method of claim 1 , wherein the first set of read pairs each has a mapping quality score (MAPQ) of greater than 60.

5. The method of claim 1, wherein the first, set of read pairs is located within at. least 200 kb of one another on the same nucleic acid fragment and their corresponding clusters are located within 1000 flow cell units of one another on the flow cell.

6. The method of claim 1 , wherein fiting the linking quality statistical model to data generated from the first set. of read pairs comprises constructing a two-dimensional histogram of the count of the relative displacement between two clusters on the flow cell associated with two read pairs in the first set of read pairs, wherein the two read pairs were derived from two regions on the same nucleic acid fragment, the two regions separated by a given genomic distance.

7. The method of claim 6, wherein the contour lines of the histogram comprise an ovallike shape.direction of flow of sequencing reaction reagents on the flow cell.

9. The method of claim 7, wherein the oval-like shape centers at the origin.

10. The method of claim 6, wherein the contour lines of the histogram comprise a circle.

11. The method of claim 10, wherein the circle centers at the origin.

12. The method of claim 6, wherein the contour lines of the histogram comprise a pair of blobs.

13. The method of claim 12, wherein the pair of blobs are symmetric with respect to the x- axis and the y-axis of the histogram, wherein the y-axis of the histogram aligns with the direction of flow' of sequencing reaction reagents on the flow cell.

14. The method of claim 13, wherein each blob of the pair of blobs is a distance away from the x-axis of the histogram.

15. The method of claim 1, wherein fitting the statistical model comprises using an optimizer for weighted non-linear least squares to obtain model parameters for the statistical model.

16. The method of claim 15, wherein the optimizer is based on the Gauss-Newton or Levenberg-Marquardt algorithm.

17. The method of claim 1 , wherein the first and second sets of read pairs are obtained from a BAM file, and wherein the determined linking quality score is added to the BAM file.

18. The method of claim 1, wherein the nucleic acid fragment is a genomic fragment.

19. A method for assigning read pairs to target nucleic acid polymers, comprising:(a) fragmenting the target nucleic acid polymers in a flow' cell to generate polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common target nucleic acid polymer are preferentially captured at proximal locations on the flow cell;(b) sequencing the polynucleotide fragments to generate read pairs;(c) aligning the read pairs to a reference sequence to determine a position on the reference sequence for each read pair; and(d) assigning two read pairs to the same target nucleic acid polymer based on the linking quality score between the two read pairs as determined according to claim(a) fragmenting target nucieic acid polymers in a flow ceil to generate polynucleotide fragments winch become captured on the flow' cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow' cell;(b) sequencing the polynucleotide fragments to generate read pairs;(c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair; and(d) producing the mapping, comprising selecting a first position among the plurality' of candidate positions for a first read pair based on the linking quality score between the first read pair and a second read pair determined according to claim 1, wherein the second read pair has an alignment score associated with a second position that is above a predetermined value.

21. A method of producing a mapping of read pairs to a reference sequence, comprising:(a) fragmenting target nucleic acid polymers in a flow cell to generate polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell;(b) sequencing the polynucleotide fragments to generate read pairs;(c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair; and(d) producing the mapping, comprising selecting a first position among the plurality of candidate positions for a first read pair based on a function of (i) the alignment score associated with the first position, (li) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and (ni) the linking quality score between the first read pair and the second read pair determined according to claim 1.

22. A method of improving the accuracy of split-alignments of read pairs to a reference sequence, comprising:polynucleotide fragments which become captured on the flow cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell;(b) sequencing the polynucleotide fragments to generate read pairs;(c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair; and(d) for a first read pair comprising different segments that align to disjoint regions in the reference sequence, selecting a first position among the plurality of candidate positions for a first segment of the first read pair based on the linking quality score between the first segment and a second read pair determined according to claim 1, wherein the second read pair has an alignment score associated with a second position that is above a predetermined value.

23. A method of improving the accuracy of split-alignments of read pairs to a reference sequence, comprising:(a) fragmenting target nucleic acid polymers in a flow cell to generate polynucleotide fragments which become captured on the flow' cell under conditions wherein polynucleotide fragments from a common nucleic acid polymer are preferentially captured at proximal locations on the flow cell;(b) sequencing the polynucleotide fragments to generate read pairs,(c) aligning the read pairs to the reference sequence to generate a plurality of candidate positions on the reference sequence and associated alignment scores for each read pair; and(d) for a first read pair comprising different segments that align to disjoint regions in the reference sequence, selecting a first position among the plurality of candidate positions for a first segment of the first read pair based on a function of (i) the alignment score associated with the first segment, (li) the alignment score associated with a second position among the plurality of candidate positions for a second read pair, and (ni ) the linking quality score between the first segment and the second read pair determined according to claims 1-18.nucleic acid polymers to form modified nucleic acid polymers.

25. The method of claim 24, wherein the inserts comprise cleavage sites.

26. The method of claim 25, further comprising cleaving the inserts at the cleavage sites.

27. The method of claim 24, wherein the inserts are added into the target nucleic acid polymers by transposases.

28. The method of claim 27, wherein each of the inserts comprises a first transposon element and a second transposon element that is contiguous with the first transposon element.

29. The method of claim 27, wherein the transposases comprise ligands and a surface of the flow cell comprises receptors that bind to the ligands to attach the modified nucleic acid polymers to the surface prior to the producing of the fragments.

30. The method of claim 29, further comprising attaching the modified nucleic acid polymers to the flow cell prior to the producing of the fragments.

31. The method of claim 30, wherein the transposases are removed from the modified nucleic acid polymers prior to the ataching of the modified nucleic acid polymer to the flow cell,32. The method of claim 30, wherein the transposases are removed from the modified nucleic acid polymer after the attaching of the modified nucleic acid polymers to the flow cell.

33. The method of any of claims 19-23, wherein the polynucleotide fragments passively diffuse to the locations on the flow cell.

34. The method of any of claims 19-23, wherein the polynucleotide fragments are actively transported to the locations on the flow cell.

35. The method of any of claims 19-23, further comprising amplifying the polynucleotide fragments at the locations to produce amplified polynucleotide fragments, wherein (b) comprises determining nucleotide sequences from the polynucleotide fragments by extension of primers hybridized to priming sites of the amplified polynucleotide fragments at the locations.

36. The method of claim 35, wherein the amplifying comprises extending at least one primer species attached to the locations to produce the amplified polynucleotide fragments attached to the locations via the primer species.primer species in a bridge amplification technique.

38. The method of any of claims 19-23, wherein (a) comprises stretching the target nucleic acid polymer along the flow cell prior to the producing of the polynucleotide fragments of the target nucleic acid polymer.

Citation Information

Patent Citations

  • Preserving genomic connectivity information in fragmented genomic DNA samples

    US10246746B2

  • Method of nucleic acid amplification

    US20050100900A1

  • Labelled nucleotides

    US20060188901A1

  • Modified polymerases for improved incorporation of nucleotide analogues

    US20060240439A1

  • Polymerases

    US20060281109A1