System and method for determining interlocking of sequence reads on flow cell

By analyzing the spatial location and genomic distance of sequence reads in a flow cell, calculating the link quality score, and optimizing the library preparation process, the problem of read linkability information loss was solved, the accuracy and efficiency of genome sequencing were improved, and resource consumption was reduced.

CN121368802APending Publication Date: 2026-01-20ILLUMINA INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202480039861.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2024-09-10
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

In the process of genomic DNA sequencing, existing technologies lose link and proximity information between reads, resulting in uneven coverage and difficulty in accurately mapping genomic regions. Furthermore, traditional methods require increasing read depth to improve coverage, which increases resource consumption.

Method used

By analyzing the spatial location and genomic distance of sequence reads in the flow cell, the link quality score is calculated, high-quality read links are screened to improve the linking rate in the flow cell, and the library preparation process is optimized by using low-concentration loading, multiple wash loading, and high molecular weight DNA extraction techniques.

Benefits of technology

It improves the link rate between reads, reduces false positives and false negatives, reduces data redundancy, improves the accuracy and efficiency of genome analysis, and reduces computational and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121368802A_ABST
    Figure CN121368802A_ABST
Patent Text Reader

Abstract

DNA sequencing systems and methods are described. Systems and methods may calculate a genomic zero distribution of read pairs on a flow cell having approximately zero probabilities linked to each other, and calculate a spatial zero distribution of read pairs having approximately zero probabilities linked spatially. The system may generate a link mass score representing a probability of incorrectly assigning a link between the two polynucleotide reads based on the calculated genomic zero distribution and the calculated spatial zero distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 582,092, filed September 12, 2023, and U.S. Provisional Application No. 63 / 565,255, filed March 14, 2024, the contents of which are incorporated by reference in their entirety. BACKGROUND TECHNICAL FIELD

[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, the present disclosure relates to systems and methods for determining and improving the number of links, percentage of links, and quality of links between read pairs on a sequencing flow cell. BACKGROUND

[0003] Genomic DNA is often too long to be directly sequenced using modern sequencing technologies. Library preparation is a step performed prior to genomic sequencing to facilitate the sequencing process and ensure accurate and efficient analysis of the genomic DNA. Library preparation involves fragmenting the DNA into smaller, manageable pieces. This fragmentation can be achieved through physical or enzymatic methods. Fragmented DNA allows for more efficient sequencing and enables the original genome to be reconstructed during data analysis. Library preparation also involves attaching adapter sequences to the fragmented DNA. Adapters contain specific sequences that are recognized by the sequencing platform and used to sequence the DNA fragments. These adapters provide priming sites and identification tags for the sequencing process.

[0004] Traditional nucleic acid sequencing methods and several types of next-generation sequencing methods use shotgun sequencing to sequence large genomic DNA fragments, referred to as template nucleic acid molecules. Specifically, the template nucleic acid molecules are first fragmented in solution into smaller pieces that are suitable for next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that the knowledge of the connectivity and proximity of the smaller sequence fragments from the template genomic sequence to each other is lost when the fragments are read. To accurately map the genomic regions, traditional nucleic acid sequencing methods often sequence a particular region (or genome) multiple times during the sequencing process.

[0005] Next generation sequencing methods will sequence each base in a sample several times. This is for several reasons. First, each base has multiple observations that produce a more reliable determination of the correct sequence of the nucleic acid. In addition, some reads are not evenly distributed across the genome, simply because the reads will sample the genome in a random and non-dependent manner. Thus, many bases will be covered by fewer reads than the average coverage, while other bases will be covered by more reads than the average coverage. This coverage is represented by the coverage metric, which is the number of times a sample genome has been sequenced, the "depth" of the sequencing coverage. For applications that aim to sequence only a defined subset of the entire genome, such as targeted resequencing or RNA sequencing, the coverage means the number of times the subset is sequenced. For example, for targeted resequencing, the coverage means the number of times the targeted subset of the genome is sequenced.

[0006] The depth of the sequencing coverage generally determines whether a variant can be discovered at a particular base position with a certain confidence. The sequencing coverage requirement varies by application, but generally, at higher coverage depths, each base is covered by a larger number of aligned sequence reads, and thus, the base can be identified with higher confidence. To achieve high coverage in genomic sequencing, researchers often increase the read depth, which involves generating a larger number of sequencing reads for a given sample. This can be achieved by sequencing more DNA molecules or by sequencing the target nucleic acid additional times to determine the correct sequence information. Additionally, the library preparation process is often optimized to minimize DNA loss and bias, ensuring that a larger proportion of the original DNA is successfully sequenced. SUMMARY

[0007] One embodiment is a system for sequencing polynucleotides, comprising at least one processor; and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to: retrieve data comprising polynucleotide sequence reads and their spatial locations on a flow cell; wherein the flow cell comprises clusters of amplified nucleic acid fragments of polynucleotides, and wherein the probability of clusters being positioned close to each other on the flow cell is related to the genomic distance between nucleic acid fragments of the polynucleotides, and wherein each cluster provides one or more of the sequence reads; determine links between sequence reads that originated from the same polynucleotide molecule; and store information about the links between sequence reads in computer memory, wherein at least one of the following: (a) linked sequence reads have a link quality score of at least 20 on the Phred scale, and linked sequence coverage has an average depth of at least 30 times (30X) coverage of a reference genome; (b) the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 45%, and the average sequence coverage is at least 70X coverage for all sequence reads; (c) the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 35%, and the average sequence coverage is at least 90X coverage for all sequence reads; and (d) the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 25%, and the average sequence coverage is at least 120X coverage for all sequence reads.

[0008] Another embodiment is a method for sequencing polynucleotides, comprising obtaining sequence reads from a flow cell, wherein the flow cell comprises clusters of amplified nucleic acid fragments of polynucleotides, and wherein the probability of clusters being positioned close to each other on the flow cell is related to the genomic distance between nucleic acid fragments of the polynucleotides, and wherein each cluster provides one or more of the sequence reads; and determining links between sequence reads that originated from the same polynucleotide molecule, wherein at least one of the following: linked sequence reads have a link quality score of at least 20 on the Phred scale, and linked sequence coverage has an average depth of at least 30 times (30X) coverage of a reference genome; (b) the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 45%, and the average sequence coverage is at least 70X coverage for all sequence reads; (c) the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 35%, and the average sequence coverage is at least 90X coverage for all sequence reads; and the percentage of sequence reads on the flow cell that have a determined link to at least one other sequence read is greater than 25%, and the average sequence coverage is at least 120X coverage for all sequence reads.

[0009] In some aspects, the technology described herein relates to a method of generating a linkage quality score, the method comprising calculating a genomic zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other on a flow cell; calculating a spatial zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other in space; and generating a linkage quality score representing a probability of an incorrect assignment of linkage between two polynucleotide reads based on the calculated genomic zero distribution and the calculated spatial zero distribution.

[0010] In some embodiments, a method for generating a linkage quality score is provided, the method comprising calculating a genomic zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other on a flow cell; calculating a spatial zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other in space; and generating a linkage quality score representing a probability of an incorrect assignment of linkage between two polynucleotide reads based on the calculated genomic zero distribution and the calculated spatial zero distribution.

[0011] In some embodiments, the method comprises calculating the genomic zero distribution of a plurality of read pairs comprises determining read pairs from clusters on a flow cell. In some embodiments, calculating the genomic zero distribution of a plurality of read pairs comprises determining read pairs from clusters that are spatially distant from each other on the same or opposite surface of a flow cell. In some embodiments, the clusters used to calculate the zero distribution are on the same surface of a flow cell but on opposite ends. In some embodiments, the clusters used to calculate the zero distribution are opposite surfaces of a flow cell.

[0012] In some embodiments, the method comprises calculating the spatial zero distribution of a plurality of read pairs comprises determining read pairs from clusters on a flow cell. In some embodiments, the technology described herein relates to a method wherein a first cluster has nucleic acids derived from a first chromosome and a second cluster has nucleic acids from a different second chromosome.

[0013] In some embodiments, the method can further comprise applying a threshold level of spatial distance and genomic distance to a plurality of read pairs on a flow cell.

[0014] In some embodiments, the method can further comprise using the quality of linkage to determine at least one threshold level such that read pairs from a cluster are derived from the same DNA molecule with a predetermined probability of spatial distance and genomic distance. In some embodiments, the method can further comprise estimating a linkage rate for total linkages between read pairs on a flow cell for a threshold level.

[0015] In some embodiments, the method can further include selecting a threshold level based on the link quality score. In some embodiments, the method can further include adjusting an assay parameter based on the link quality score. In some embodiments, the assay parameter includes at least one of: a flow cell design, a coverage, a DNA loading density, a DNA loading method, a DNA loading strategy, an intensity of an electric field applied to the flow cell, a linearity of flow over the flow cell, a 3D structure of the DNA, a solution applied to the flow cell, and an index. Similarly, in some embodiments, the assay parameter of the flow cell design includes at least one of a pitch size, a well depth, and a flow cell size.

[0016] In some embodiments, the techniques described herein relate to a method, wherein generating the link quality score includes generating a false positive rate. In some embodiments, the false positive rate is proportional to a product of a percentage of genomic zero links and a percentage of spatial zero links. In some embodiments, generating the link quality score includes generating an expected percentage of genomic zero links and spatial zero links for links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance.

[0017] In some embodiments, the expected percentage of genomic zero links and spatial zero links is proportional to a product of the false positive rate and a number of links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance. In some embodiments, generating the link quality score includes generating a false discovery rate. In some embodiments, the false discovery rate is proportional to an estimated number of false links divided by a total number of links in the plurality of read pairs on the flow cell. In some embodiments, the link quality score is expressed using a Phred scale.

[0018] In some embodiments, the method involves computing a genomic zero distribution, and performing a spatial zero distribution of read pairs on a subset of the plurality of read pairs on the flow cell. In some embodiments, the read pairs represent total read pairs on the flow cell. In some embodiments, a quality of a link between two polynucleotides represents a quality of a link of the plurality of reads on the flow cell being below a threshold level of spatial distance and genomic distance. In some embodiments, the quality of the link between the two polynucleotides is applied to individual read pairs within the plurality of reads on the flow cell.

[0019] In some embodiments, the method can further include assigning a candidate link between two polynucleotide reads using spatial position data. In some embodiments, assigning the candidate link between the two polynucleotide reads using the spatial position data includes determining a link quality score that a first cluster and a second cluster on a sequencing flow cell originate from a same target polynucleotide. In some embodiments, generating the link quality score includes a statistical test including at least one of a t-test, a chi-squared test, a permutation test, and a deconvolution analysis. Attached Figure Description

[0020] The exemplary features of this disclosure will become apparent from the following detailed description and accompanying drawings, in which similar reference numerals correspond to similar but potentially different components. For brevity, reference numerals or features having the functions previously described may be described in conjunction with or without conjunction with other drawings in which they appear. While this disclosure has been illustrated and described in detail in the accompanying drawings and the foregoing description, such illustration and description should be considered illustrative or exemplary rather than restrictive. This disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will understand and implement variations of the disclosed embodiments in practice with respect to the claimed disclosure.

[0021] Figure 1A Non-limiting examples of solid carriers that can be used to perform the disclosed sequencing techniques are illustrated schematically.

[0022] Figure 1B A non-limiting example of template molecules on a flow cell is illustrated schematically.

[0023] Figure 2 A flowchart is shown as an example method for ordering reads when there is a high link rate between sequential reads.

[0024] Figure 3 A flowchart is shown for an example method used to generate link quality scores.

[0025] Figure 4 An overview of the methods used to estimate percentage link rate is shown.

[0026] Figure 5A A bar chart is shown illustrating the difference in linkage rates under the processing conditions of the sequencing system at two different concentrations of input DNA. Figure 5B A line graph depicting the relationship between the read coverage of deduplicated links and the combined false positive (FP) plus false negative (FN) rate of phylogenetic detection is shown.

[0027] Figure 6 Screens A and B show bar graphs illustrating the results of an experiment reporting linker rates performed using crude samples of patient blood. In Screen A, the white blood cell count (WBC) for each blood sample is shown as dots. In Screen B, the nanograms of genomic DNA in 25 µl of blood are shown as dots.

[0028] Figure 7 Screen A and Screen B show two line graphs of the link quality Q score calculated using the false positive rate for two different experimental systems.

[0029] Figure 8 Figures A-F show schematic diagrams of a random flow cell (panel A), a patterned flow cell (panel B), center-to-center spacing of flow cells (panel C), reduced center-to-center spacing of flow cells based on feature size (panel D), an electron micrograph image of a cross-sectional side view of a flow cell well (panel E), and a top-down electron micrograph image of a flow cell well (panel F).

[0030] Figure 9 Exemplary post-sequencing spatial positions of sequencing reads on a solid support are illustrated in accordance with some embodiments.

[0031] Figure 10 Non-limiting examples of metrics related to spatial distance information from a cluster are illustrated in accordance with some embodiments.

[0032] Figure 11 A flowchart showing an example method for generating a linkage quality score is shown.

[0033] Figure 12 A system including a memory including a sequencing module for achieving high linkage rates is schematically illustrated. DETAILED DESCRIPTION

[0034] All patents, applications, published applications and other publications referred to herein are incorporated by reference herein in their entirety. In the event of inconsistency between the terminology, usage or definitions set forth in the patents, applications, published applications and other publications incorporated by reference herein and the terminology, usage or definitions set forth herein, the terminology, usage or definitions set forth herein control.

[0035] The present disclosure describes sequencing methods performed using improved assays that result in a relatively high rate of linking between sequence reads. Linking between two sequence reads on a flow cell indicates that the two reads were found on the original template nucleic acid molecule from which the reads originated to be in close proximity to one another prior to fragmentation. Thus, reads that are "linked" to one another will map in close proximity to one another on the original template nucleic acid molecule. Determining links involves analyzing sequence reads that are spatially close to one another on a flow cell. This involves the discovery that when template nucleic acid molecules, such as genomic DNA, are fragmented on a flow cell, sequences that are positioned spatially close to one another on the flow cell are more likely to have originated from the same template nucleic acid molecule. Where a first sequence and a second sequence are found to be in close proximity to one another on the original template, the first sequence positioned in close proximity to the second sequence can be assigned a link between the two sequences, which is useful for proper alignment and assembly of the sequenced nucleic acid molecule. Since there is a random chance that two unrelated sequences will end up in close proximity to one another on a flow cell, a link quality score can be assigned to the link in order to determine and potentially filter true links between sequence reads from false links. The link quality score can also be used to determine whether a sequencing system and method is generating more or less true links between sequences.

[0036] The present disclosure describes methods and systems that can improve the rate at which links are generated between sequence reads, and also provides systems and methods that increase the rate at which links are created between sequences by performing various assays to increase the linking rate. Various techniques for increasing the number of sequence read links on a flow cell can include library preparation techniques that result in a flow cell with an increased number of linked reads compared to a flow cell with sequence reads that have not been subjected to the library preparation techniques. For example, some techniques discovered for increasing the number of links between sequence reads can include, by way of non-limiting example, low concentration loading, multiple wash loading methods, use of a crude lysate, use of high molecular weight DNA extraction techniques, use of additives, and optimization of flow rates. These techniques will be discussed in more detail in the following sections.

[0037] The present disclosure provides systems and methods for sequencing nucleic acids using improved assays, such as polynucleotides, where the rate of linking between sequence reads on the surface of a flow cell is increased. Typical procedures used to prepare libraries for sequencing short reads are not necessarily the best procedures for generating linked reads. When the goal is to produce linked reads, the methods traditionally used to prepare libraries and cluster on instruments that are used for short read sequencing can not be the most efficient. Linked reads, which provide a more comprehensive genomic insight by preserving information about longer DNA fragments, can use a consistent sequencing process, while the preparation stage and clustering techniques can be adapted to meet the specific goals of linked read sequencing. Various assay parameters have been adjusted to increase the linking rate. For example, it has been found that adding a lower concentration of DNA on the flow cell increases the number of true links on the flow cell. This result can be interpreted as a lower concentration of DNA added to the flow cell results in less crowding of unrelated DNA on the flow cell prior to fragmentation, which in turn will result in less overlap of unrelated fragments on the flow cell. A lower false linking rate allows for a higher genomic and spatial distance threshold, which is a criterion for establishing links with high quality. Thus, the total amount of DNA on the flow cell can be reduced, but the linking rate between sequences will be increased. It was found that unextracted crude lysates and samples from simple lysis as from blood or other sources produce higher linking rates compared to samples where DNA is extracted in a standard manner. This enhanced performance is attributed to the high molecular weight of the DNA in these crude lysates. Similarly, high molecular weight extraction kits or gentle sample processing can help generate conditions for better and more spatial linking overall or in percentage.

[0038] The benefits of linked reads can be partially quantified by determining the number of false positives and false negatives when analyzing a ground truth set. Based on empirical analysis, increasing the linking rate in sequencing data results in diminishing returns of performance gains for identifying false positives and false negatives once a depth of approximately 30x linking coverage can be achieved. Thus, each base should be covered on average by 30 different reads on the flow cell. Experiments were conducted to determine the effect of coverage depth on false positive (FP) and false negative (FN) reduction in phylogenetic variant identification. It was observed that the combined FP+FN decreases as the depth of de-duplicated linked read coverage increases, which indicates that higher coverage depth generally results in fewer errors in variant identification. However, the performance gains begin to plateau once approximately 30x coverage depth is achieved. Similarly, the percentage reduction in FP+FN as coverage depth increases from 8x to 16x, 16x to 24x, 24x to 32x, and 32x to 40x shows the highest reduction observed when coverage depth increases from 8x to 16x.

[0039] Some embodiments are directed to providing a linkage quality of at least Phred score of 20. Phred quality score is a common metric used in sequencing to indicate the quality of nucleotide calls. The score is logarithmically linked to the probability of error. For example, Phred score of 20 corresponds to a probability of error of 1%, score of 30 to 0.1%, and so on. In the context of the metric of "linkage quality" on the Phred scale, determining a lower threshold depends on the specific requirements of the sequencing project. However, as a general guideline: Phred score of 20 is traditionally considered the minimum of reliable data, equivalent to a confidence level of 99%. In some embodiments, at a depth of coverage of 70x, the linkage rate of reads with linkage quality score of 20 is about 45% to 50%. This is equivalent to about 30x depth coverage of linked data.

[0040] The present disclosure also provides systems and methods for establishing linkage quality between two or more pairs of reads on a flow cell. As discussed herein, "linkage" is the probability that two pairs of reads on a sequencing flow cell originated from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long DNA from a biological source, such as genomic DNA, are fragmented to produce shorter fragments that can be sequenced in a single read. The fragmentation process will produce these shorter fragments that land on a flow cell, and the spatial location of each fragment can be correlated to the original nucleic acid molecule from which the fragment originated. For example, it has been found that fragments from the same nucleic acid molecule bind closer together on a flow cell than fragments from different original nucleic acid molecules. Thus, if two clusters of reads on a flow cell are spatially close together and also close together on the individual's genome, the clusters are more likely to be from the same nucleic acid molecule. However, unrelated fragments can also bind close to each other to a flow cell, leading to uncertainty in the probability that adjacent clusters originated from the same molecule. Many factors can influence the probability that unrelated clusters will land in similar regions, and these factors can change based on various experimental conditions, as described herein. Embodiments of the present disclosure provide statistical methods for calculating the probability that two reads are linked, such that on a flow cell, the two reads originated from the same nucleic acid molecule.

[0041] Embodiments of the present disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acids and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind to capture primers, which are then used to generate clusters by well-known techniques, such as those provided by Illumina, Inc. (San Diego, CA). As described above, according to the methods of the present disclosure, fragments derived from the same template genomic sequence are more likely to bind to the flow cell in spatially proximate locations compared to fragments from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In conventional library preparation where fragmentation occurs prior to loading, fragments can land anywhere in the flow cell, regardless of whether they are from the same molecule. However, when the fragmentation is performed directly on the flow cell, spatial information is preserved. This spatial information can be used to help guide assembly of the original template genomic sequence and variant identification, as will be described in more detail below.

[0042] For example, one embodiment is a method for assigning nucleic acid sequence reads to a target polynucleotide, the method comprising providing a transposome complex. In some embodiments, the transposome complex comprises a transposase and a first polynucleotide having end sequences that can be used to fragment the target polynucleotide and insert an end sequence or tag into each fragment that can be used to bind to a capture probe located on a substrate. The method can comprise contacting the transposome complex bound to a solid substrate with the target polynucleotide under conditions to fragment the target polynucleotide, and a capture sequence can be added to the end of each fragment.

[0043] Once the fragments have been bound to the substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. The location of each cluster on the flow cell can then be determined prior to, during, or after performing a sequencing-by-synthesis reaction (SBS) to obtain the nucleotide sequence of each fragment located in each cluster. Once the nucleotide sequence of each cluster has been determined, the method can begin mapping those reads to determine the original target polynucleotides from which the reads originated. In some embodiments, the mapping process takes into account the spatial location of each cluster such that clusters that are closer to each other on the flow cell are more likely to originate from the same target polynucleotide. In some embodiments, the library preparation steps are performed on the flow cell, which can reduce the complexity and number of equipment associated with the system. Furthermore, by using the spatial information that accompanies each cluster to map the sequenced fragments to the target polynucleotides, the method performs a more accurate mapping operation compared to methods that do not take into account the spatial location of each cluster during the mapping process. Thus, the spatial information including the relative distance between various clusters on the flow cell can be utilized to adjust the mapping information, thereby improving the read quality of previously identified multiply-mapped reads. In the past, identified multiply-mapped reads can have been discarded. By improving the read quality of these previously discarded reads by increasing the confidence of the read alignment based on linkage information having a high linkage quality score, the alignment information and information quality used in certain genomic analysis applications, including but not limited to variant calling, can be improved.

[0044] Processing DNA samples that are amenable to high-throughput sequencing that preserve information about the original configuration of the DNA sample provides useful information about co-localized fragments. However, co-localized fragments will not always originate from the same molecule, which creates some difficulty in defining the confidence that any two co-localized fragments originate from the same molecule. The ambiguity between true and false links can limit the usefulness of spatial information or for integrating linked reads into mapping or alignment processes. Therefore, assessing and quantifying the relationships of fragments on a flow cell can have a significant impact on mapping, alignment, and variant identification. The ability to filter out false links between reads and to quantify high-quality links and resulting read alignments impacts processing requirements. By improving the quality of linked reads, the methods of the present disclosure also serve to reduce the memory footprint of the sequencing process. High-quality linked reads provide more accurate and reliable genomic information. This means that fewer reads are needed to confidently cover and assemble specific regions of a genome. In contrast, in lower-quality sequencing, more reads (and thus more data) are typically needed to achieve the same level of confidence, resulting in data redundancy. By reducing the number of redundant reads, the overall amount of data that needs to be stored and processed is reduced, thereby reducing memory requirements. With high-quality linked reads, the computational processes involved in sequencing, such as alignment and assembly, can be more accurate when it is difficult to map genomic regions. High-quality reads align to the reference genome more accurately and with less computational effort. This means that these processes require less computational power and memory. Additionally, high-quality data can reduce the need for extensive error correction and data cleaning steps, which are resource-intensive tasks.

[0045] High-quality linked reads reduce the need for extensive data cleaning and filtering processes, which are necessary to deal with errors and ambiguities in lower-quality data. These processes not only consume computational resources, but also increase memory footprint by generating additional intermediate data files. As read quality improves, polynucleotide sequences in difficult regions of the genome can be more accurately mapped and aligned.

[0046] Described herein are systems and methods for establishing a link quality score for links determined between pairs of reads based on spatial information obtained on a flow cell. The spatial information may, for example, be Cartesian coordinates of a cluster containing a particular read on a flow cell. In one embodiment, the spatial information can include the location of a well on a substrate. To establish these spatial links between two reads, in one embodiment, two thresholds are used. The first threshold is a spatial distance threshold, which represents the physical distance between two reads on a flow cell.

[0047] In some embodiments, spatial distances can be measured in units of nanometers. In some embodiments, spatial distances can be measured in units relative to the length of a flow cell. For example, flow cell units can be relative to the size and / or spacing of the patterned clusters on the flow cell. In some embodiments, two differently patterned flow cells can have different absolute length units due to different densities of clusters on the surface. In some embodiments, spatial distances can be in absolute length units, or any other length units consistent with the present disclosure. In some embodiments, spatial distances can be included in a FASTQ file, which is typically a text file containing sequence data from clusters that pass through a filter on a flow cell. The FASTQ file can be used as sequence input for alignment and other secondary analysis software.

[0048] The second threshold is a genomic distance threshold, which represents the distance between two reads on the genome after mapping to a reference genome. Empirical methods for establishing thresholds will vary widely between experimental conditions. The present disclosure provides methods of attaching a linkage quality score as a factor of the spatial and genomic distance between two potentially linked reads to a linkage. As described in more detail below, one method of determining the quality of a linkage between two reads is to estimate the null distribution of pairs of read pairs. This null distribution can provide the basis for calculating a "false discovery rate," which can then be used as a proxy for the linkage quality score of a linkage.

[0049] The linkage quality score is defined as a numerical representation that quantifies the reliability of a linkage between two read pairs. Multiple metrics that affect the quality of a linkage can be used to calculate this score, and the linkage quality score can serve as a composite measure that simplifies complex relationships into a single value.

[0050] Link quality scores can provide a basis for comparison or decision making. For example, a high link quality score between two pairs of reads can indicate that the two reads are likely to have originated from the same DNA fragment, and thus should be paired for further analysis, and conditions used to generate the link can be adjusted and evaluated based on the score. The formula used to calculate the link quality score can vary, and can be determined based on a false discovery rate, a measure quantifying type II error (false negative error), a statistical or mathematical model, a weighted average of different contribution measures, and / or a machine learning model trained to predict link quality based on multiple features. As a non-limiting example, the following measures can be helpful in evaluating the quality score of a linked read on a flow cell. In some embodiments, read length and quality can be useful, as longer reads with high base quality scores can be more likely to produce accurate linking information, which in turn benefits genome assembly and variant identification. In some embodiments, coverage depth and uniformity can also be helpful in ensuring that genomic regions are adequately sequenced, and have uniform coverage for regions that are difficult to sequence. In some embodiments, phasing efficiency, which measures how accurately linked reads are assigned to their respective haplotypes, e.g., in a training set, can also impact the precision of haplotype analysis. In these cases, the link quality score is intended to encapsulate a variety of different considerations into a single number representing the overall "quality" of the link, facilitating quantitative analysis.

[0051] A relevant aspect to consider is the relationship between the surface area in a flow cell and the likelihood that two read fragments will land close to each other by chance. Due to the finite surface area that accommodates reads, a small surface area in a flow cell reduces the probability that two read fragments will land in close proximity. Conversely, a larger surface area in a flow cell increases the likelihood that the chance occurrence of read fragments, including those from different chromosomes, will land in close proximity. Thus, establishing spatial links with a large threshold distance for read fragments in a flow cell, in turn, leads to an increased likelihood of considering false links. The larger the area considered, the potential for information gain from the sequencing process is potentially reduced, especially if a minimum link quality threshold is maintained. Link quality scores convey whether a link is false or not, and downstream methods can determine which quality scores to consider. While links can still exist within a larger surface area, distinguishing true links from false links presents additional challenges.

[0052] Reference is now made to Figure 1AIn some embodiments, the flow cell 100 that provides spatial information of read pairs can include a plurality of lanes 110. Each lane 110 includes a plurality of surfaces. As shown, in some embodiments of the flow cell 100, a lane includes a top surface 112 and a bottom surface 114. If reads being compared are on opposite surfaces, the distance between them is considered infinite by the assumption that they cannot be linked. In some embodiments, each surface is subdivided into a plurality of tiles 120. As shown, a cluster 130 can be located on a tile 120 designated 1201. This designation is used as an illustrative example only and is not limited to the alphanumeric characters shown in the figure. In some embodiments, the tiles 120 include two-dimensional X-Y coordinates as shown to provide spatial information between clusters. In some embodiments, the X-Y coordinates can be derived from information stored in a FASTQ file. In some embodiments, the X-Y coordinates can be stored in or derived from a BCL (base call) file, which is a binary file format commonly associated with next generation sequencing (NGS) platforms.

[0053] In some embodiments, the subdivision of the surface into tiles 120 is artificial separation, such that the surface of the flow cell is not separated into physical tiles, but rather the image captured by the camera can be segmented into tiles. As shown, the tiles 120 are subdivided into strips, which generally correspond to the pixel width of the camera used to capture the image of the flow cell. In some embodiments, the tiles 120 represent the size of the image that can be captured by the camera. In some embodiments, the X-Y coordinates are pixel values. In some embodiments, 1 unit of the tile 120 can be approximated as 1 / 10 of a pixel. Physical separation is contemplated in some embodiments, where the tiles can have physical barriers, wells, and other structures that separate one portion of the flow cell from another portion of the flow cell. In some embodiments, the spatial information of a cluster such as the cluster 130 is obtained by a camera that processes pixel values of a digital image, including the X-Y coordinates.

[0054] An example of a determination that can be provided regarding spatial information of sequenced reads and / or associated library preparation can be a determination performed on a substrate having transposome complexes immobilized thereon. In some embodiments, the transposome complexes can include, for example, a transposase and a first polynucleotide including a terminal sequence and a first tag. The determination can be performed by contacting the transposome complexes with a target polynucleotide under conditions that fragment the target polynucleotide. The fragmented target polynucleotide can then be amplified to form a plurality of nucleic acid clusters on the substrate. The plurality of nucleic acid clusters on the substrate are microscopically observable, and their positional data can be recorded. After the positional information has been obtained, nucleic acid sequence reads of the fragmented nucleic acids can then be sequenced, and corresponding positional data can be stored.

[0055] In some embodiments, spatial information can be used to link reads together, where the linkage between reads can be physical or non-physical. In some embodiments, spatial information is used to link reads together to form longer linked reads with one or more pairs of reads. Linked pairs of reads have expected properties, such as but not limited to expected length distribution, distance between pairs, and number of pairs. These properties can be leveraged in genomic analysis. Reference is now made to Figure 1B In one non-limiting example, flow cell 148 contains a set of clusters, where each cluster is composed of nucleic acid fragments. Fragment 1 is located in cluster 151 and is derived from segment 152 on original nucleic acid 150. Similarly, fragment 2 is located in adjacent cluster 155 and is derived from segment 156 on original nucleic acid 150. Fragments 1 and 2 can be linked using spatial information that confirms that fragments 1 and 2 are from the same original nucleic acid 150 or any other polynucleotide consistent with the present disclosure. The length of the linked read construct that connects these fragments can be represented as the length of fragment 1 and fragment 2 plus, for example, 2 units. Thus, in this example, the distance between fragment 1 and fragment 2 is 2 units, such as 2 units of length on the flow cell. Note that any arbitrary unit can be used, such as X-Y coordinates or flow cell coordinates, as shown by distance 190, which is 4 flow cell units of distance from seat. In the case where clusters 151 and 155 are in close proximity in a packed array, a transformation can be used to transform flow cell units to X-Y coordinates, and vice versa. In another non-limiting example, flow cell 160 can contain fragment 3 and fragment 4, which can be falsely labeled as linked. In this example, fragment 3 can be derived from original nucleic acid 170, while fragment 4 can be derived from a different original nucleic acid 180. In another non-limiting example, fragments 3, 4, and 5 can be linked together using spatial information to confirm that fragments 3, 4, and 5 are derived from the same polynucleotide molecule. In this example, the distance between fragment 3 and fragment 4 can be 5 units or less, and the distance between fragment 4 and fragment 5 can be 5 units or less.

[0056] Figure 2 is a flowchart of an example method 200 for sequencing polynucleotides that achieve a desired linkage rate between reads. At step 210, an assay is performed to prepare nucleic acids, such as DNA or RNA, for sequencing. This preparation typically involves procedures such as extraction, quantification, and fragmentation of the nucleic acids, ensuring that the nucleic acids are in proper form and concentration for subsequent analysis. In the methods described herein, fragmentation can occur after the DNA or RNA is first bound to a flow cell.

[0057] After preparation, the workflow proceeds to step 220, where sequence reads are obtained from the flow cell. As described herein, the flow cell can include clusters of amplified nucleic acid fragments of polynucleotides, and when performing the methods of the present disclosure, the probability of clusters being positioned close to each other on the flow cell correlates to the genomic distance between nucleic acid fragments of the polynucleotides. In some embodiments, each cluster provides one of the sequence reads.

[0058] At step 230, the method identifies the spatial location of the sequence reads on the flow cell. The exact identification of where each read is located on the flow cell can be embedded in a taken image of the flow cell surface, where clusters can correspond to specific pixels in the image. Methods of converting from pixels to flow cell units or other distance measures are further described herein.

[0059] Next, at step 240, the method can proceed by determining links between the sequence reads obtained in the previous step, where a link will indicate that sequences originated from the same polynucleotide molecule. In line with the present disclosure, various parameters can be adjusted in order to consistently achieve high link rates between sequence reads. However, there are some trade-offs with sequence depth, coverage, and link rate. For example, as previously described, in order to increase coverage, researchers can typically increase the loading of sample onto the flow cell to generate more reads, but this strategy can result in lower link rates. The following non-limiting parameters demonstrate high link rates and high coverage, and can be achieved following the methods provided by the present disclosure. However, it is noted that in some cases, increasing from 0.5x to lx coverage can increase link rates, as additional coverage can begin to drop when approaching full occupancy of the flow cell. As an additional example, gross-fragmentation has been shown to achieve both high coverage and high link rates under a range of sample inputs.

[0060] In order to optimize link rates in sequencing, several strategies and parameters can be adjusted: reducing the concentration of sample, as has been demonstrated in the data sets described herein; utilizing gross-fragmentation, which can affect link rates depending on its level of processing; performing high molecular weight (HMW) DNA extraction to preserve larger DNA fragments; implementing a multiple flush technique, which can involve multiple sample streams to optimize the sample environment; adding specific additives to stabilize DNA or improve reaction conditions; optimizing flow rates within the sequencing device to affect the movement and capture of DNA; ensuring gentle sample processing to maintain DNA integrity; employing molecular combing to stretch DNA molecules, which can enhance the clarity of spatial relationships between segments; and designing active flow cells and microchannels to facilitate efficient link capture. Additionally, multiple sample streams can also enable the repeated addition of low concentrations of template to maintain high link rates, but also maintain high coverage. These methods, along with considerations for flow cell design, DNA loading density, methods, and strategies, can be used to enhance link rates, and thus improve the quality of sequencing data.

[0061] Then, at step 250, the workflow evaluates the quality of the links by generating a link quality score. In some embodiments, given the flow cell coordinates (e.g., Cartesian coordinates) and the genomic coordinates of two reads, the link quality score can be the Phred-scaled log-likelihood of the two reads not being linked. In some embodiments, the score is related to the probability that the link between the two polynucleotide reads is correctly linked. The process reaches a decision point at step 260, where it is determined whether the target coverage has been reached. In some embodiments, if the coverage is insufficient, the workflow can optionally loop back to step 270, where assay optimization is optionally performed to further improve the sequencing output.

[0062] Once it is confirmed that the target coverage has been reached, the sequence of steps ends, which indicates the end of the experimental workflow and that the data is ready for downstream analysis. In some embodiments, the linked sequence reads have a link quality score of at least 20, and the linked sequence coverage has an average of at least 30X coverage of the linked sequence reads. In some embodiments, the linked sequence reads have a link quality score of at least 25, and the linked sequence coverage has an average of at least 30X coverage of the linked sequence reads. In some embodiments, the percentage of sequence reads on the flow cell that have a non-spurious link with at least one other sequence read is greater than 45%, and the average sequence coverage is at least 70X coverage for all sequence reads. In some embodiments, the percentage of sequence reads on the flow cell that have a non-spurious link with at least one other sequence read is greater than 35%, and the average sequence coverage is at least 90X coverage for all sequence reads. In some embodiments, the percentage of sequence reads on the flow cell that have a non-spurious link with at least one other sequence read is greater than 25%, and the average sequence coverage is at least 120X coverage for all sequence reads.

[0063] Method of calculating quality score

[0064] In some embodiments, a method for generating a link quality score is provided, the method comprising calculating a genomic zero distribution of a plurality of pairs of reads on a flow cell that have approximately zero probability of being linked to each other; calculating a spatial zero distribution of a plurality of pairs of reads that have approximately zero probability of being linked to each other in space; and generating a link quality score representing a probability of incorrectly assigning a link between two polynucleotide reads based on the calculated genomic zero distribution and the calculated spatial zero distribution.

[0065] The present disclosure provides methods for generating quality scores that facilitate, but do not necessarily perform, the methods of the present disclosure. For example, the methods of the present disclosure can generate a large (or small) number of links, each of which can have a high quality score. In such cases, all links can be used without filtering or otherwise assessing the quality of the links. However, it has been found that link quality scores are a useful measure to distinguish true links from false links. Thus, Figure 3 A flowchart of a method 300 for generating link quality scores by computing a spatial zero distribution 310 and a genomic zero distribution 320 is provided. At step 310, the method can proceed by computing a genomic empirical zero distribution of pairs of reads on a flow cell that have approximately zero probability of being linked to each other. For example, pairs of reads on opposite surfaces of a flow cell cannot truly be linked to each other because template molecules cannot flow through the space separating the flow cell surfaces. In another example, pairs of reads that are also chosen to be far enough apart that fragments from the original DNA molecule cannot separate that distance on the flow cell can also be used to derive a zero distribution because these pairs of reads should also not be linked. In this context, "far" can refer to a predetermined threshold distance that would provide a high probability that the pairs of reads are not from the same origin. The sufficient distance to conclude that two pairs of reads are not from the same origin can depend on the underlying statistical distribution of the data and the particular context. Generally, a sufficient distance will be a distance significantly greater than the average separation between pairs of reads that would be expected based on the underlying assumptions about the distribution of fragmented DNA sequences on the flow cell surface, which can indicate that two objects are not from the same origin. Thus, the identification of links that are too far apart on the flow cell surface to have originated from the same nucleic acid molecule facilitates the estimation of the genomic zero distribution for the entire data set. These long distance links (which are non-candidate links due to their spatial separation) provide valuable data for understanding the background noise within the sequencing process. These estimates can be readily collected by selecting clusters on opposite surfaces or marked with spatial data located at the far opposite ends of the flow cell (e.g., the beginning and end regions of the flow cell). Thus, this subset of the total sequencing data can be used to generate an estimate of the genomic zero distribution. In some embodiments, a full zero distribution is not necessary for generating meaningful link quality scores.

[0066] In some embodiments, read pairs used to derive the zero distribution (genomic or spatial) have a near zero probability of being truly linked to each other (i.e., originating from the same molecule). As described above, cross-surface read pairs can be used to derive genomic distance zero values because selecting these read pairs allows the method to select read pairs that are not from the same molecule without having to additionally set constraints on genomic distance. Thus, some embodiments can avoid biased zero distributions, where bias can be introduced by arbitrarily selecting, for example, large genomic distances that can introduce a non-representative sample of read lengths with bias. Similarly, to compute spatial distance zero values, the disclosed methods can identify a set of reads that are not from the same molecule without setting any constraints on the spatial distance between them in order to not bias my spatial zero distribution. As described above, cross-chromosome reads can be used to derive zero distributions without biasing the population used for estimation.

[0067] The genomic zero distribution allows for determining the expected level of noise or random variation in the data, thus the genomic zero distribution provides the basis for statistical tests. In one formulation, the probability (p-value) of observing a given genomic linkage can be computed under the default expectation of no linkage. If the observed result is unlikely to occur by chance (e.g., if the p-value is below a predetermined threshold), the result provides evidence for rejecting the default expectation and indicates a true genomic linkage.

[0068] The genomic zero distribution can be compared to the observed data to identify significant genomic signals. For example, a difference between the observed data and the genomic / spatial distance zero distribution will be due to the presence of true linkages. In some embodiments, a significant signal means that there is strong evidence that there is an effect in the data beyond random variability. This determination is typically made using statistical tests that compare the observed data to a zero distribution (random or expected distribution) and compute a p-value. If the p-value is below a predetermined threshold (typically denoted by a, such as 0.05), the result is considered significant. As described above, the genomic signal of interest will be the true linkage between read pairs that flow on the pool from the same polynucleotide sequence. Similarly, the quality of experimental procedures and data processing methods can be evaluated by comparing the observed data to the genomic zero distribution. A difference between the observed data and the zero distribution can indicate problems such as batch effects, sample contamination, or technical bias.

[0069] After aligning reads to a reference genome, a genomic null distribution of read pairs on the flow cell can be performed, assuming a linkage probability of zero. The alignment process maps short DNA sequences (reads) obtained from the flow cell to their corresponding genomic locations in the reference genome. After alignment, reads can be processed to identify read pairs that are properly aligned and in the expected orientation. Additionally, after aligning read pairs to the reference genome, the process of assigning read pairs to a particular chromosome can include analyzing the alignment information and determining the genomic location of each read pair. The mapping process can provide this information by indicating the mapping location of each read pair to the reference genome. For example, to assign a read pair to a chromosome, the mapping location is checked against each chromosome in the reference genome, which can be represented by a range of genomic coordinates. If the mapping location of a read pair falls within the range of a particular chromosome, then the read pair is assigned to that chromosome. This assignment can be performed using a variety of methods, such as comparing the mapping location to a pre-defined template chromosome coordinate file, or utilizing a reference genome annotation that provides chromosome information for each genomic location. Additionally, particular alignment tools or bioinformatics pipelines can provide built-in functionality for assigning reads to chromosomes based on alignment results.

[0070] Method 300 can proceed to step 320 by computing a spatial null distribution of a plurality of read pairs having approximately zero probability of being linked spatially to one another. The read pairs can be aligned to a reference genome, where the read pairs can be given a genomic location, e.g., a genomic location on a chromosome. When two read pairs are on different chromosomes, it is less likely that the two read pairs are linked because the reads are from different chromosomes (i.e., different molecules) in the reference genome. Note that in some cases, such as structural variants (SVs) of translocation or fusion events, these reads can be involved. In such cases, the reads, although appearing on different chromosomes, can indeed be linked, challenging the initial assumption of their independence. Considering structural variants (SVs) in sequencing data to challenge the null assumption is generally not a significant factor for several reasons. First, the occurrence of SVs in a genome is relatively rare. In the vast landscape of genomic sequences, instances of read pairs from different chromosomes being linked due to SVs, such as translocations or fusions, are uncommon. As a result, the majority of read pairs on different chromosomes are indeed unlinked, originating from separate chromosomal molecules. This rarity implies that the proportion of reads potentially linked by SVs is small when compared to the total number of reads, thereby minimizing their impact on the overall data interpretation.

[0071] Additionally, by carefully selecting regions without structural variants, the likelihood of encountering reads linked by SVs is reduced, which effectively reduces the potential impact of SVs on the null hypothesis. Another consideration is the statistical significance of SVs in large-scale genomic data. The sheer volume of unlinked reads in a typical sequencing project tends to dilute the impact of the relatively few reads linked by SVs. This statistical perspective ensures that the reads linked by SVs have minimal overall impact on the dataset, thereby preserving the integrity of the data analysis. Thus, the estimate of the spatial null distribution can be generated from this subset of read pairs of the total sequencing data that indicate reads on different pairs of chromosomes.

[0072] Once the null distribution is calculated, the method can proceed to step 325 by calculating a false positive rate proportional to the product of the percentage of spatial and genomic null links below a threshold spatial and genomic distance. It should be recognized that alternative methods for calculating a link quality score can not employ a false positive rate as recited in step 325, and thus it can be optional. For example, other positive and negative predictive values can be used to test true positive and true negative results. Link quality scores can be derived from the null distribution to estimate the reliability or confidence of the observed data. These scores can be based on metrics such as signal-to-noise ratio, deviation from the null distribution, or the likelihood of observing the data given the default expectation. Higher link quality scores indicate greater reliability and confidence of the observed results. Statistical tests of interest can be performed on the observed data to obtain each of the expected test statistics or p-values (false positive, false negative, true positive, true negative). Test statistics can be calculated on the sampled data, where the p-value represents the probability of obtaining a test statistic as extreme or more extreme than the observed value under the default expectation. In other examples, the false positive rate can be used as an intermediate measure to estimate the false discovery rate, or any other type of Type I error measure.

[0073] The method can then proceed to step 330 by generating a link quality score representing the probability of incorrectly assigning a link between two polynucleotide reads. The link quality score can be generated based on the calculated genomic null distribution and the calculated spatial null distribution. Calculating a false discovery rate (FDR) from the null distribution can involve determining the proportion of false positives among total positive discoveries when there is no true effect. The null distribution can be generated by simulating data under the assumption that there is no true effect or association. The present disclosure provides methods that can avoid the step of fitting a null distribution or simulating a null distribution by randomly shuffling the observed data. Alternatively, the present disclosure provides methods that utilize complementary information on the genomic sequence (spatial link data), and a subset of this data can be selected (which will not represent realistic or possible scenarios) to estimate the null distribution and generate the link quality score. Thus, the systems and methods of the present disclosure can advantageously improve the quality and efficiency of data processing associated with polynucleotide sequencing. In particular, practical applications of quantifiable linked reads can include improvements in mapping and variant identification.

[0074] At step 335, the method calculates a linkage quality score for different threshold combinations. In some embodiments, step 335 can include calculating a linkage quality score by determining a false discovery rate for each threshold. In some embodiments, step 335 can include calculating a linkage quality score by determining a previous zero distribution for a new threshold. In some embodiments, step 335 can include generating a linkage quality score based on a false discovery rate. By utilizing multiple or different threshold combinations, different linkage quality scores can be achieved, allowing for more informed decisions or more accurate placement of linked read pairs. To calculate a linkage quality score for different threshold combinations, various parameters and metrics can be considered, such as the number of linked read pairs, sequence depth, read overlap, sequence quality, and alignment score. In some embodiments, the alignment score can be output from a Smith-Waterman alignment. For each combination of these metrics, a threshold can be established, above or below which a linkage is considered to meet that particular quality criterion. A linkage quality score can then be calculated based on how many thresholds are met and to what extent those thresholds are met. This can be achieved through a weighted average, or a composite linkage quality score can be generated. The composite linkage quality score can be designed to represent individual scores as a comprehensive metric or a series of metrics that encapsulate the overall quality of linkages in a sequencing experiment. Alternatively or additionally, outputs from different thresholds can be used in a hierarchical manner. High quality connections determined by applying the most stringent threshold can be used or analyzed first. This can be beneficial in applications where computational resources are limited or there is some benefit to quickly identifying the most reliable linkages. Lower quality connections that meet less stringent thresholds can be considered later or as a supplement to high quality linkages.

[0075] Once linkage quality scores have been calculated at step 335, process 300 can move to decision step 340 to determine whether a target threshold or optimal threshold has been reached. If the target threshold has been reached, process 300 can terminate at an end step. If the target threshold has not been reached, process 300 can return to step 325, so the calculations can be repeated in order to determine a target threshold or optimal threshold. A predetermined target threshold for identifying true connections between read pairs can be set based on prior knowledge from previous verifications. For example, a target threshold can be based on setting a minimum number of linked read pairs or a maximum allowed distance between linked read pairs on a flow cell. An optimal threshold can be determined through post-analysis of previous target thresholds. Thresholds can be customized to maximize particular metrics that can be important for a given task, such as sensitivity, specificity, or a balanced combination of both. Methods such as cross-validation can be used to fine-tune the target threshold in order to minimize error, including false positives and false negatives of linkages between read pairs.

[0076] Optionally, at step 350, various assay parameters can be adjusted, including a flow cell design, a coverage, a DNA loading density, a DNA loading method, a DNA loading strategy, a strength of an electric field applied to the flow cell, a linearity of flow over the flow cell, a 3D structure of the DNA, a solution applied to the flow cell, and non-limiting examples of indexing. The flow cell design includes a pitch size, a well depth, and a flow cell size.

[0077] In some embodiments, generating the linkage quality score includes generating a false positive rate. Note that there are several statistical tests that can be used to generate the linkage quality score that can or can not be directly associated with the false positive rate or the false discovery rate. Generally, the false positive rate is proportional to the product of the percentage of genomic zero links and the percentage of spatial zero links. Thus, calculating the linkage quality score can include generating an expected percentage of genomic zero links and spatial zero links for links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance. The resulting expected percentage of genomic zero links and spatial zero links will also be proportional to the product of the false positive rate and the number of links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance. In some embodiments, generating the linkage quality score includes generating a false discovery rate. The false discovery rate can be proportional to the estimated number of false links divided by the total number of links in the plurality of read pairs on the flow cell. The false discovery rate can be expressed using a Phred scale.

[0078] In some embodiments, the linkage quality score includes a score representing the likelihood that a read from a particular cluster links to another cluster originating from the same original target polynucleotide. In some embodiments, the linkage quality score for a pair of reads from two clusters is 0. In some embodiments, the linkage quality score for a pair of reads from two clusters is higher than 30. In some embodiments, the positional information includes a first spatial coordinate and a second spatial coordinate in a 2D coordinate system. In some embodiments, the method further includes classifying the plurality of reads determined from the nucleic acid cluster by the spatial coordinates of the nucleic acid cluster. This linkage quality score can be analogous to a mapping quality (MAPQ) score, e.g., the likelihood that a read’s mapped position to a particular target polynucleotide is incorrect. The linkage quality score provides a new complementary type of information. For example, the linkage quality score reflects proximity information that is different from other techniques for contiguity or origination information. The proximity information tells how close a pair of reads are to each other within some threshold, but does not require that the pair of reads be directly adjacent (or contiguous with the linked reads). The proximity information for polynucleotide fragments bound to a flow cell refers to the spatial relationship and relative positioning of these fragments on the surface of the flow cell during sequencing. Thus, these methods are adapted to this complementary information, and generally, more linkage information can be determined at a lower cost relative to other methods.

[0079] In some embodiments, the linkage quality score can be used in combination with a mapping quality (MAPQ) score to generate more high quality mappings and assign nucleic acid sequence reads to a target nucleotide. For example, a first read pair with a low quality MAPQ score can be linked to a second read pair with a high quality MAPQ score, such that the first read pair can be aligned with high confidence. The linkage quality score provides a measure of quantifying high confidence.

[0080] Figure 4 Comprised of three different panels, the figure shows an overview of the method for estimating percent linkage rate. The linkage rate estimate can be obtained by dividing the number of linked reads by all observed reads, and is the proportion of observed linkages (or connections) between sequence reads relative to the total sequence reads. The linkage quality score can be used to quantify the improvement in linkage rate and subsequent coverage. By evaluating the linkage quality across a threshold and comparing results from different experiments, researchers can make informed decisions about the relevance and accuracy of their genomic studies.

[0081] The figure is divided into panels A, B, and C. The first panel A, labeled "Linkage Quality," presents a set of curves, each curve representing a different spatial threshold (40, 50, 70, 100, 300, and 500). The x-axis of the plot is the genomic threshold, spanning 10^3 to 10^7 base pairs (bp). The y-axis represents the linkage quality, which ranges from 0 to 40. The curves show how the linkage quality changes across various genomic thresholds for the specified spatial thresholds. As the genomic threshold increases, the linkage quality appears to first slightly increase, then decrease, forming a downward trend. The region of the plot with a light green background highlights the region where the linkage quality is highest (and is desirable for later applications) for various spatial thresholds.

[0082] The second panel, labeled "Reference Genome," shows an array of shaded shapes (402, 403), each representing a sequence read that covers the reference genome, with the depth indicated by the vertical stacking of the sequence reads. Some of the sequence reads are shaded with the same color (402), indicating that those reads have linked read pairs. The panel also shows unlinked read pairs (401) in light gray, which are expected to form a percentage of the total number of sequence reads. The panel serves as a visual representation of the genomic region under consideration in the analysis.

[0083] The third plot presents a histogram showing linkage rates from two different hypothetical experiments. The y-axis represents linkage rates, in percentages from 0% to 80%. The histogram provides two orange bars of different heights. The difference in their heights indicates a change in linkage rates between the two experiments and / or two different conditions or treatments. The histogram serves as an example tool to monitor and adjust sequencing experiments, sample preparation, and methods to optimize the balance between linked reads and overall coverage to achieve, for example, maximum linked read coverage. Variables that significantly contribute to enhancing the number of successfully linked reads can be identified by comparing linkage rates between different conditions. For example, in scenarios where samples are scarce or DNA quality is compromised, adjustments can be made to the sequencing protocol, such as changing DNA input, library preparation methods, or sequencing procedures and additives, to improve linkage quality and achieve desired results. The ability to make informed decisions about these trade-offs via linkage quality scores can enhance the efficiency of experiments and the quality of sequencing data. This flexibility and adaptability highlight the innovative aspects of the technology, providing a method to dynamically monitor and refine sequencing efforts.

[0084] Various assay parameters can be adjusted, including, by way of non-limiting examples, flow cell design, coverage, DNA loading density, DNA loading method, DNA loading strategy, strength of electric field applied to the flow cell, linearity of flow on the flow cell, 3D structure of DNA, solution applied to the flow cell, and indexing. Flow cell design includes pitch size, well depth, and flow cell size.

[0085] Figure 5AA bar graph showing one example demonstrating the impact of processing conditions on a sequencing system and the respective linkage rate difference for two absolute amounts (or concentrations) of DNA input. In the experiment, extraction of genomic DNA, typically yielding DNA fragments larger than 60 kb, was performed for sequencing at 500 ng and 50 ng loading inputs. Different inputs resulted in differences in linkage rates. The observed variation in linkage rates indicates that the amount of DNA loaded onto a flow cell can have a significant impact on the success of linkage across reads. This relationship between loading input and linkage rate is an important consideration for optimizing sequencing experiments. Higher DNA input can enhance the chances of reads being linked due to increased opportunities for overlapping reads; however, this must be balanced against the possibility of flow cell overcrowding, which can have adverse effects on sequencing quality. These differences can be further adjusted and would be expected due to differences in fluidics between systems caused by different formulations and flow cell geometries. Variability in linkage rates can also be attributed to differences in the sequencing system itself, including the flow cell's fluid dynamics and physical design. These systems can differ in their handling of samples and reagents and the distribution pattern of DNA fragments during sequencing. Different sequencing systems can require individualized optimization of DNA input amounts to achieve efficient linkage and coverage. Understanding these nuances via linkage quality score allows for fine-tuning of experiments to suit the potentially unique characteristics of a sequencing system.

[0086] Figure 5B A line graph showing the relationship between de-duplicated linked read coverage and combined false positive (FP) plus false negative (FN) rates for germline detection. There are three lines, each representing a different linkage quality score threshold: Q25 at 70x coverage, Q25 at 140x coverage, and Q30 at 70x coverage. The x-axis, representing de-duplicated linked read coverage, ranges from 0 to 45. The y-axis represents the combined FP+FN and ranges from 0 to 12,000. The lines on the graph show that as de-duplicated linked read coverage increases, the combined FP+FN decreases, indicating an improvement in accuracy. However, the rate of improvement in accuracy seems to slow as coverage exceeds approximately 30x, indicating diminishing returns with further increases in coverage. Each line on the graph flattens as it moves to the right, reinforcing the idea of a plateau effect in performance gains. The line representing Q25 at 70x appears to have the highest FP+FN rate across all coverage levels compared to the other two lines, while Q30 at 70x generally has the lowest FP+FN rate. The line for Q25 at 140x shows intermediate performance between the lines for Q25 and Q30 at 70x.

[0087] It is expected and observed that flushing during loading of the sample results in higher linkage coverage. Consider loading a sample with a low concentration and performing a single flush to remove excess reagents and unbound sample. This single flush of low concentration is expected to result in reduced occupancy of the binding sites of the flow cell, and because of the lower chance of crowding resulting in false linkages as described above, the false positive rate is also lower. It is expected that this single flush using a relatively low concentration will result in fewer occupied binding sites on the flow cell. However, as mentioned previously, because of the reduced likelihood of crowding that can result in false linkages, it also contributes to a lower false positive rate. Thus, this lower false positive rate contributes to both higher linkage quality and higher linkage rate. It is important to note that the term "lower concentration" is relative to the initial concentration used; if this lower concentration represents the optimal level, the overall linkage coverage of a single experiment can not necessarily be lower, but can be at the optimal level balancing linkage quality and coverage. Conversely, if the sample is prepared with a single flush at a higher concentration, there can be increased occupancy and increased false positives, resulting in a reduced linkage rate and also suboptimal linkage coverage.

[0088] Thus, one method to increase linkage rate is to perform multiple flushes at low concentration. The first flush at low concentration is expected to result in reduced occupancy of the binding sites of the flow cell, and a low false positive rate. The second flush at low concentration will then occupy binding sites that were not occupied during the first flush. Fragmentation can optionally be performed between flush steps. Thus, after the second flush, fragments from the second flush will not be overmixed with the bound fragments of the first flush. The flushes can be performed two times, three times, four times, or more. This method results in increased occupancy of the flow cell, resulting in overall more sequences. The low false positive rate results in a higher linkage rate, which in turn results in overall increased linkage coverage of the sequencing run. In some embodiments, a single flush can result in a higher linkage rate. Nonetheless, it is still feasible to retain the multiple flush method as an option, particularly when tagging is employed between flushes. This technique has been observed to maintain high linkage rates even in the case of multiple flushes in NSX runs, which indicates that the effectiveness of this strategy can vary.

[0089] Improved linkage rate method

[0090] Methods according to the present disclosure can be adapted to different sequencers, increasing linkage rates and leading to overall increased linkage coverage of sequencing runs. The amount of improvement in linkage rates for each system can vary depending on, for example, the density of nanopores on a flow cell or the flushing steps during a sequencing experiment. Accordingly, the present disclosure provides methods that achieve increased linkage rates of at least a threshold performance. Additionally, the present disclosure provides methods that perform sequencing after linkage rates have been optimized, without the need for repeated optimization steps during the sequencing process. In some embodiments, methods can achieve a relative performance gain over normal operating methods, such as methods performed without the need to determine spatial information of reads on a flow cell surface.

[0091] In some embodiments, the improved methods can be quantified based on linkage rates between linked sequence reads having a high mapping quality score. In some embodiments, the linked sequence reads have a mapping quality score (MAPQ) of at least 25. In some embodiments, the linked sequence reads have a MAPQ of at least any one of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, and 35. In any one of these embodiments, the linked sequence coverage can be 30X coverage. In some embodiments, the linked sequence coverage can be at least any one of 10X, 15X, 20X, 25X, 35X, 40X, 45X, 50X, 60X, and 70X. In some embodiments, the linked sequence reads have a mapping quality score (MAPQ) of at least 25, and the linked sequence coverage has an average value of at least 30X coverage of the linked sequence reads.

[0092] In some embodiments, the improved methods can be quantified based on linkage quality scores. As described herein, linkage quality scores can be quantified based on a phred scale. Accordingly, in some embodiments, the linked sequence reads have a linkage quality score of at least 25. In some embodiments, the linked sequence reads have a linkage quality score of at least any one of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, and 35. In any one of these embodiments, the linked sequence coverage can be 30X coverage. In some embodiments, the linked sequence coverage can be at least any one of 10X, 15X, 20X, 25X, 35X, 40X, 45X, 50X, 60X, and 70X. In some embodiments, the linked sequence reads have a linkage quality score of 25, and the linked sequence coverage has an average value of at least 30X coverage of the linked sequence reads.

[0093] In some embodiments, the improved method can be quantified based on the percentage of sequence reads on a flow cell that have non-spurious links with at least one other sequence read. In some embodiments, the percentage of sequence reads on a flow cell that have non-spurious links is greater than 45%. In some embodiments, the percentage of sequence reads on a flow cell that have non-spurious links is greater than 45% and the average sequence coverage is at least 70X coverage for all sequence reads. As described above, the percentage of links that have non-spurious links can be increased by underloading the sample, resulting in lower overall coverage. Thus, the method according to the present disclosure can be quantified based on both linkage rate and coverage. In some embodiments, the coverage can be the coverage of linked sequence reads. In some embodiments, the coverage can be the coverage of all sequence reads. In some embodiments, the percentage of sequence reads on a flow cell that have non-spurious links with at least one other sequence read is greater than 35% and the average sequence coverage is at least 90X coverage for all sequence reads. In some embodiments, the percentage of sequence reads on a flow cell that have non-spurious links with at least one other sequence read is greater than 25% and the average sequence coverage is at least 120X coverage for all sequence reads.

[0094] Types of errors that can be used in linkage quality score

[0095] The present disclosure provides methods of determining a confidence level of linkage data. In some embodiments, the confidence level of linkage data can be related to the probability of the presence of false links in a set of links filtered by a threshold level. The probability of false links in a set of links is a Type I error, also known as a false positive error. A Type I error occurs when the null expectation of being true (no link) is incorrectly rejected in favor of the alternative expectation (a link) when it is not. Several measures can be used to estimate and quantify the rate of Type I error, including the false positive rate (FPR) and the false discovery rate (FDR). Various measures can be used to determine the confidence level of linkage data.

[0096] For example, the false positive rate (FPR) represents the proportion of instances or cases in which the null expectation is true but is incorrectly rejected as a significant negative instance. The FPR measures the rate at which an experiment incorrectly identifies an effect that is not present. The FPR can be calculated as the ratio of false positives to the total number of false positives plus true negatives. Thus, the false positive rate (FPR) can be calculated using the following equation: FPR = FP / (FP + TN). Here, FP represents the number of false positives (instances that are incorrectly classified as positives), and TN represents the number of true negatives (instances that are correctly classified as negatives).

[0097] The above equation represents the proportion of instances that are actually negative examples but are classified as positive examples by the prospective test (or binary classification model). The false positive rate can be used to estimate the performance of a classification model or the expected error rate in a test, indicating the frequency with which the test or model makes Type I errors (incorrectly rejecting a true null hypothesis of expectation).

[0098] The false discovery rate (FDR) is a related but different measure that is primarily used in the context of multiple prospective tests. The FDR can be used to measure the proportion of positive results (significant discoveries) among all significant results that are false positives. In other words, the FDR addresses the question of the proportion of incorrect rejections among all rejections. The FDR can be used with multiple statistical tests that are performed simultaneously. Various methods can be used to control the FDR, such as the Benjamini-Hochberg procedure, for example, to limit the number of false discoveries. The FDR addresses the challenge that arises when a large number of statistical tests are performed simultaneously, increasing the likelihood that a statistically significant result is obtained by chance alone.

[0099] The false discovery rate (FDR) reflects the proportion of results that are identified as significant that are actually false positives. The false discovery rate (FDR) can be calculated using the following equation: FDR = FP / (FP + TP). Here, (FP + TP) represents the total number of data points, where FP represents the number of false positives (instances that are incorrectly classified as positive examples) and TP represents the number of true positives (instances that are correctly classified as positive examples).

[0100] The equation quantifies the proportion of instances that are classified as positive examples but are actually negative examples in a binary classification problem. The false discovery rate can be used to control the proportion of false positive discoveries among all results that are deemed significant in a multiple prospective test. By quantifying the proportion of false positives within a selected population, the FDR can balance the number of true discoveries against the risk of making false positive claims.

[0101] In some embodiments, FDR provides a practical way to balance between identifying meaningful associations in population health studies and minimizing the risk of false discoveries. By determining the FDR at a pre-defined threshold, the value of the FDR can be used to distinguish between truly meaningful associations and false discoveries. In some embodiments, the FDR can be regulated using FDR control procedures, which are designed to control the FDR, which is the expected proportion of "discoveries" (zero-valued hypotheses that are rejected) that are false (incorrect rejections of zero values). Equivalently, the FDR is the expected ratio of the number of false positive classifications (false discoveries) to the total number of positive classifications (rejections of zeros). The total number of rejections of zero values includes both the number of false positives (FPs) and true positives (TPs). In contrast to familywise error rate (FWER) control procedures, such as the Bonferroni correction, which control the probability of at least one Type I error, FDR control procedures provide less stringent control over Type I errors.

[0102] Some embodiments can also involve estimation of the number of true links that were not used because the true link was above a threshold level. Type II metrics can be used to determine whether there is more data that can be salvaged. Type II error metrics assess the performance of a statistical test in failing to detect true effects, i.e., they quantify the proportion of actual positives that were incorrectly classified as negatives or correctly identified as positives. Metrics such as the false negative rate (FNR) and sensitivity quantify the proportion of actual positives that were incorrectly classified as negatives or correctly identified as positives, respectively. The power and 1 - β represent the probability of correctly rejecting a false null expectation, indicating the ability of the model to detect true effects. The miss rate highlights the proportion of actual positive instances that the model missed. Specificity assesses the accuracy of the model in correctly identifying negative instances, while the fall-out assesses the tendency of the model to incorrectly classify negatives as positives. The Type II error rate ( β ) represents the probability of failing to reject a false null expectation when the false null expectation is actually false. These metrics collectively provide a comprehensive view of the ability of a model or test to avoid Type II errors and accurately identify true effects or outcomes.

[0103] In some embodiments, Type II metrics can be preferably adapted to a commonly used size or scale, or can be provided with a scale that provides intuitive meaning. For example, in genomics, Phred-scaled quality scores have become a standardized way of assessing the accuracy of DNA sequencing data, helping researchers make informed decisions about the reliability of identified genetic variations. The logarithmic nature of this scale facilitates efficient data representation, sensitivity enhancement, and cross-platform comparability.

[0104] When metrics are subject to a Phred scale, such as quality scores in sequencing data, the logarithmic nature of the Phred scale helps to amplify differences between values. Small differences in raw data can result in large differences on the Phred scale. This amplification of change can help to identify subtle distinctions or anomalies that can be missed when inspecting raw data directly. Furthermore, the logarithmic transformation of the Phred scale can facilitate comparisons across different experiments, instruments, or laboratories. Since the scale is standardized and logarithmic, it reduces the impact of variations between platforms or processes, making it easier to compare results consistently. The magnitude of the value of the false positive rate can result in more favorable / interpretable quality metrics. In some embodiments, FDR can be more preferable than FPR. In some embodiments, Type I error metrics can be preferable to Type II error metrics. (More preferable FDR than FPR).

[0105] To calculate FDR, the observed test statistics or p-values are compared to a null distribution. The proportion of false positive discoveries from the null distribution above a specified threshold represents the expected number of false positives. Here, FPR is the proportion of false positives among all pairs tested (thus false links among the magnitude of zero values). FDR is then calculated by dividing the number of false positives by the total number of positive discoveries (both true and false positives) obtained from the observed data. By comparing the observed test statistics or p-values to a null distribution, FDR provides an estimate of the proportion of false positives among significant discoveries. This quantifies the rate at which false positives would be expected to occur when the default expectation is true.

[0106] In some embodiments, linking nucleic acid sequence reads includes determining a distance between each cluster and using the determined distance to assign reads to a particular target polynucleotide. In some embodiments, assigning nucleic acid sequence reads includes determining a linking quality score that sequence reads from at least a first cluster and a second cluster on a substrate originate from the same target polynucleotide. In some embodiments, determining that sequence reads from more than two clusters originate from the same target polynucleotide is based on a linking quality score. In some embodiments, the method further includes increasing a linking quality score of sequence reads from at least a first cluster when a spatial distance and a genomic distance between reads originating from the first cluster and the second cluster are below a threshold. In some embodiments, assigning nucleic acid sequence reads to a particular target polynucleotide includes determining a linking quality score that a plurality of sequence reads from a cluster on a substrate originate from the same target polynucleotide. In some embodiments, the method further includes increasing a linking quality score of sequence reads from a cluster when a spatial distance and a genomic distance between the cluster and one or more other clusters are below a threshold.

[0107] Linked reads can be performed in a genome reference dependent or independent manner. For genome independent linking, only spatial information is used to form links between reads. In a non-limiting example, reads are linked when they are separated by 50 flow cell distance units. In this example, reference based information, including but not limited to genomic information, etc., is not considered when linking these reads. For genome dependent linking, spatial information and genomic information are used to form links between reads. In a non-limiting example of genome dependent linking, when reads are aligned to a reference genome, a pair of reads or multiple reads are linked when they are separated by 50 flow cell distance units and within 10 kbp distance.

[0108] In some embodiments, alignment data can be analyzed to create a genome null distribution by considering the expected genomic properties of pairs of reads. For example, if reads are randomly sampled from a genome without any particular linkage, the distribution of distances between mapped pairs should follow a random distribution pattern. In some embodiments, a genome null distribution can be generated by computing the observed distribution of distances of correctly aligned pairs of reads and comparing it to the expected distribution under the assumption of no genomic linkage. This can be achieved by various statistical methods. For example, one method can involve computing the empirical distribution of distances between correctly aligned pairs of reads and comparing it to a theoretical null distribution. The null distribution can be generated by randomly permuting the mapped positions of pairs of reads while maintaining their original genomic positions. By repeating this permutation process multiple times and generating null distributions, a statistical baseline can be established, which can be compared to the observed distribution of distances.

[0109] In some embodiments, computing a genome null distribution for a plurality of pairs of reads includes determining pairs of reads from clusters on a flow cell. In some embodiments, the clusters on the flow cell can be in a random flow cell or in a patterned flow cell. In some embodiments, the position data of the pairs of reads can be the positions of the clusters on the flow cell. In some embodiments, the step of determining pairs of reads from clusters on a flow cell can be performed before, during, or after the data analysis of the null distribution. For the clusters of pairs of reads that can be used to compute the null distribution, there can be multiple markers. For example, the clusters for the null distribution can be spatially distanced from each other on the same or opposite surfaces of the flow cell. In some embodiments, the clusters can be on the same surface of the flow cell but on opposite ends. In some embodiments, the clusters can be on opposite surfaces of the flow cell.

[0110] In some embodiments, a spatial zero distribution of a plurality of read pairs can be computed by determining read pairs from clusters on a flow cell. In some embodiments, computing a zero distribution includes estimating a zero distribution based on a subsample of total data generated in a sequencing run. In some embodiments, computing a zero distribution can be performed during the first few cycles of a sequencing run and used to determine whether an experiment should continue. In some cases, the zero distribution can be reused for subsequent experiments.

[0111] In some embodiments, a cluster can be used to compute a zero distribution, where a first cluster has read pairs originating from a first chromosome and a second cluster has read pairs from a different second chromosome. In some embodiments, computing a zero distribution can include applying a threshold level of spatial distance and genomic distance to the plurality of read pairs on a flow cell. In some embodiments, a method can include using quality of links to determine at least one threshold level such that read pairs from a cluster originate from the same DNA molecule with a predetermined probability of spatial distance and genomic distance. For example, a method can include estimating a link rate of total links between read pairs on a flow cell for a threshold level. Accordingly, these methods can be used to select a threshold level based on a link quality score. Based on the quality level, different assay parameters can be explored and adjusted to assess the impact on the number and quality of links generated.

[0112] In some embodiments, a method includes computing a genomic zero distribution of a plurality of read pairs includes determining read pairs from clusters on a flow cell. In some embodiments, computing a genomic zero distribution of a plurality of read pairs includes determining read pairs from clusters that are spatially distant from each other on the same or opposite surface of a flow cell. In some embodiments, clusters used to compute a zero distribution are on the same surface of a flow cell but on opposite ends. In some embodiments, clusters used to compute a zero distribution are opposite surfaces of a flow cell.

[0113] In some embodiments, a method includes computing a spatial zero distribution of a plurality of read pairs includes determining read pairs from clusters on a flow cell. In some embodiments, a method described herein involves a method where a first cluster has nucleic acids originating from a first chromosome and a second cluster has nucleic acids from a different second chromosome.

[0114] In some embodiments, a method can further include applying a threshold level of spatial distance and genomic distance to the plurality of read pairs on a flow cell.

[0115] In some embodiments, a method can further include using quality of links to determine at least one threshold level such that read pairs from a cluster originate from the same DNA molecule with a predetermined probability of spatial distance and genomic distance. In some embodiments, a method can further include estimating a link rate of total links between read pairs on a flow cell for a threshold level.

[0116] In some embodiments, the method can further comprise selecting a threshold level based on the link quality score. In some embodiments, the method can further comprise adjusting an assay parameter based on the link quality score. In some embodiments, the assay parameter comprises at least one of: a flow cell design, a coverage, a DNA loading density, a DNA loading method, a DNA loading strategy, an intensity of an electric field applied to the flow cell, a linearity of flow on the flow cell, a 3D structure of DNA, a solution applied to the flow cell, and an index. Similarly, in some embodiments, the assay parameter of the flow cell design comprises at least one of a pitch size, a well depth, and a flow cell size.

[0117] In some embodiments, the techniques described herein relate to a method, wherein generating the link quality score comprises generating a false positive rate. In some embodiments, the false positive rate is proportional to a product of a percentage of genomic zero links and a percentage of spatial zero links. In some embodiments, generating the link quality score comprises generating an expected percentage of genomic zero links and spatial zero links for links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance.

[0118] In some embodiments, the expected percentage of genomic zero links and spatial zero links is proportional to a product of a false positive rate and a number of links from the plurality of read pairs on the flow cell that are below a threshold level of spatial distance and genomic distance. In some embodiments, generating the link quality score comprises generating a false discovery rate. In some embodiments, the false discovery rate is proportional to an estimated number of false links divided by a total number of links in the plurality of read pairs on the flow cell. In some embodiments, the link quality score is expressed using a Phred scale.

[0119] In some embodiments, the method involves computing a genomic zero distribution, and performing a spatial zero distribution of read pairs on a subset of the plurality of read pairs on the flow cell. In some embodiments, the read pairs represent total read pairs on the flow cell. In some embodiments, a quality of a link between two polynucleotides represents a quality of links of the plurality of reads on the flow cell being below a threshold level of spatial distance and genomic distance. In some embodiments, the quality of the link between the two polynucleotides is applied to individual read pairs within the plurality of reads on the flow cell.

[0120] In some embodiments, the method can further include assigning a candidate linkage between two polynucleotide reads using spatial position data. In some embodiments, assigning a candidate linkage between two polynucleotide reads using spatial position data includes determining a linkage quality score that a first cluster and a second cluster on a sequencing flow cell originate from the same target polynucleotide. In some embodiments, generating a linkage quality score includes a statistical test including at least one of a t-test, a chi-squared test, a permutation test, and a deconvolution analysis.

[0121] In some embodiments, the systems and methods can determine an appropriate threshold to remove pairs of reads that are located near each other on a flow cell but are not derived from the same polynucleotide molecule. Different datasets can use different thresholds, and the selection of a threshold can involve a tradeoff between sensitivity (the ability to detect true linkages) and specificity (the ability to exclude false linkages). Setting a strict threshold can increase specificity by removing more false linkages, but can also result in excluding true genomic associations. Conversely, setting a lenient threshold can increase sensitivity, but can result in including false positive linkages. Regardless, without having a metric to quantify the quality of the data after applying a threshold, such as a linkage quality score, the confidence assigned to a linkage after establishing a threshold is unclear.

[0122] Examples of parameters that affect linkage quality

[0123] Linkage quality scores associated with data obtained by sequencing a biological sample on a flow cell can be subject to a variety of factors. Such factors include the accuracy of base calling, the choice of sequencing chemistry, the range of read lengths, and the intensity of signal derived from fluorescent labels, all of which can have an impact on the assignment of linkage quality scores. Additional considerations including the density of clusters on a flow cell, the number of cycles of various phases of sequencing, the quality of sample material being sequenced, such as the molecular weight of polynucleotides, and the specific GC content of a sequence can also result in fluctuations in linkage quality scores.

[0124] In some embodiments, methods according to the present disclosure can use a concentration of less than 5 picomoles in a loaded sample on a flow cell. In some embodiments, methods according to the present disclosure can use a concentration of less than any one of 1 picomole, 2 picomoles, 3 picomoles, 4 picomoles, 5 picomoles, 6 picomoles, 7 picomoles, 8 picomoles, 9 picomoles, and 10 picomoles in a loaded sample on a flow cell. In some embodiments, methods according to the present disclosure can use less than 75 nanograms of a loaded sample. In some embodiments, methods according to the present disclosure can use less than any one of 25 nanograms, 50 nanograms, 75 nanograms, and 100 nanograms of a loaded sample. In some embodiments, methods according to the present disclosure can use less than any one of 250 nanograms, 500 nanograms, 750 nanograms, and 1000 nanograms of a loaded sample.

[0125] In some embodiments, methods according to the present disclosure can use gentle processing techniques. In some embodiments, gentle processing techniques can be quantified by the procedures used during sample preparation. In some embodiments, methods according to the present disclosure can use fewer than three purification steps, such as two, one, or no purification steps, which are performed on the loaded sample prior to loading the sample. In some embodiments, the length of polynucleotides determined as part of a quality control (QC) process can be used to quantify the effectiveness of sample preparation procedures as described. In some embodiments, the length of polynucleotides can be known based on the typical output of the assay design. In some embodiments, methods can use polynucleotides having at least 60,000 base pairs in the loaded sample. In some embodiments, methods can use polynucleotides having at least any one of 10 kbp, 20 kbp, 30 kbp, 40 kbp, 50 kbp, 60 kbp, 70 kbp, 80 kbp, 90 kbp, and 100 kbp in the loaded sample.

[0126] In some embodiments, methods according to the present disclosure can use a flow rate of less than 250 µL / min for the flow of solution across the flow cell during sample loading. In some embodiments, methods can use a flow rate that is reduced by any one of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90%. In some embodiments, the flow rate can be maintained at less than 250 µL / min, with particular examples employing rates such as 50 µL / min, 100 µL / min, and other similar scales. In some embodiments, methods according to the present disclosure can use a flow rate of greater than 250 µL / min for the flow of solution across the flow cell during sample loading. In some embodiments, methods can use a flow rate that is increased by any one of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90%.

[0127] In some embodiments, methods according to the present disclosure can use a crude lysate loading sample. Utilizing a crude lysate as a loading sample is an innovative approach mentioned in the present disclosure to potentially increase linkage rates in sequencing methods. This strategy involves bypassing the conventional DNA extraction and purification steps and using the lysate, which is a mixture of cellular components obtained after cell lysis, directly in the sequencing process. The underlying principle behind this approach is that crude lysate generally contains high molecular weight DNA, which is more likely to maintain physical linkage between different parts of the genome. When these high molecular weight DNA fragments are used for sequencing, the likelihood of capturing longer stretches of physically linked genomes increases. This can result in enhanced linkage rates, as more extensive regions of the genome can be sequenced in a single read or closely associated reads. Furthermore, by skipping the DNA extraction process, which can sometimes shear the DNA into smaller fragments, the integrity of these long DNA molecules is better preserved. This preservation can be used to maintain the likelihood of physical linkage between reads, which in turn leads to high linkage rates. In some embodiments, obtaining sequence reads can comprise loading a crude lysate loading sample onto a flow cell, wherein the crude lysate sample has been lysed but not extracted. In some embodiments, a crude lysate can comprise an unpurified mixture obtained by lysing a biological sample through mechanical, chemical, or thermal means to release cellular contents. In some embodiments, a crude lysate can be used directly in the sequencing process without further purification.

[0128] In some embodiments, methods according to the present disclosure can sequence polynucleotides after two or more loading cycles on a flow cell. In some embodiments outlined in the present disclosure, methods present a method of sequencing polynucleotides after two or more wash loading cycles on a flow cell. This technique involves repeatedly introducing a sample into a flow cell, followed by washing it out, and then repeating the process multiple times. The underlying principle behind this approach is that multiple wash loading cycles can potentially enhance the distribution and occupancy of polynucleotides on the surface of a flow cell. Each wash cycle has the potential to deposit a different set of polynucleotide fragments onto the flow cell. By repeating the process, the method can maximize the coverage of genomic regions being sequenced, while also potentially reducing the amount of sample loaded during each cycle. This increased coverage is crucial for ensuring comprehensive genomic analysis, particularly in studies where detecting low-abundance sequences is relevant.

[0129] According to the present disclosure, it has been observed that various sequencing methods achieve different linkage rates, and these rates can be used as a metric for optimization. By measuring the linkage rate, the sequencing protocol can be adjusted to optimize this rate, which can involve changing the cluster density on the flow cell, the number of cycles, the quality of the sample material in terms of molecular weight. In some embodiments, the percentage of sequence reads that have a linkage with a second sequence read is greater than 45%. In order to achieve such a high linkage rate, the sequencing method should be adjusted and optimized. The methods according to the present disclosure consistently achieve linkage rates that meet or exceed this threshold.

[0130] In some embodiments, sample loading includes the introduction of one or more additives. In some embodiments, sample loading includes the introduction of one or more additives selected from the group consisting of magnesium, detergent, betaine, DMSO, chelating agents, reducing agents, and PCR enhancers and stabilizers. These additives can be selected based on their ability to enhance the linkage rate of the sequencing process, but are often used in typical sequencing methods in different combinations or concentrations.

[0131] Figure 6 Two panels are shown displaying a bar graph labeled“coarse lysate blood samples” reporting the linkage read rate for 11 patients (numbered 1 through 11). The bar graph shows two types of data: one represented by orange bars and labeled“NS6K-Crude” which corresponds to the Q25 linkage rate from the coarse sample of a crude lysate blood sample sequenced on a NovaSeq 6000s. These percentages align with the left vertical axis. In panel A, the second type of data is represented by dots and labeled“WBC” which stands for white blood cells, the counts are shown on the right vertical axis. The bars show some variation in linkage rate between patients, with percentages ranging from below 20% (for patient 9) to above 80%. The WBC counts are plotted on the same horizontal axis as the bars, with the highest count near 40 (in thousands per µL) and the lowest count around 5. As observed in the right axis of the graph, patient 9 has an abnormally high white blood cell count, resulting in likely overloading of the sample on the flow cell and a corresponding decrease in the Q25 linkage rate. Panel B shows the same linkage rate data as panel A, but the dots represent nanograms of genomic data found in 25 µl of blood.

[0132] The following experimental section details non-limiting examples of procedures for processing various biological samples and extracting high molecular weight (HMW) DNA. Blood samples can be stored at 4°C and processed within 3 days, using 25 μΐ, per reaction and collected in K2EDTA. Bone marrow aspirate (BMA) samples have similar storage and processing requirements, but use 15 μΐ, per reaction and are also collected in K2EDTA. In some examples, a range of 15 μΐ, to 30 μΐ, can be used. Tissue culture cells and / or fresh frozen tissue are harvested by centrifugation or trypsin digestion, followed by cell counting. After harvesting, any liquid should be removed as much as possible, and the remaining cell pellet should be processed immediately or frozen at -80°C.

[0133] For HMW DNA extraction, the following general considerations can help maintain DNA integrity: avoid high temperatures, extreme pH, intercalating dyes, UV light, and vortexing. DNA should be pipetted with wide-bore tips, remove all RNA, and do not use guanidinium thiocyanate or phenol / chloroform to minimize damage. Extracted DNA should then be resuspended in low-salt buffer or water. Several specific DNA extraction kits are listed that are available from various suppliers, with the general recommendation to follow the manufacturer's blood, bodily fluid, and tissue culture cell protocols. If no specific protocol is available for tissue culture cells, previous experimental results will suggest using the same procedure as for blood.

[0134] For quality control (QC) and storage, the consistency of DNA concentration should be checked on both Nanodrop and Qubit, and purity should be checked on Nanodrop at the specified optimal optical density (OD) ratio. DNA size should be estimated using DNA genomic screening tape or a pulsed field gel electrophoresis system, with the goal of a DNA integrity number (DIN) greater than 9. Finally, DNA should be stored at -20°C for long-term storage, with the recommendation to avoid more than ten freeze-thaw cycles and to aliquot DNA if necessary.

[0135] The following experimental procedure provides an example of a method for preparing blood and bone marrow aspirate (BMA) samples, and involves several reagents and a two-step heat incubation process.

[0136] First, the liquid is removed from the cell pellets, which should be about 25 μΐ^ in volume. These pellets can be processed immediately or stored at -80 °C for later use. For the lysis step, a 10x lysis buffer containing 2% sodium dodecyl sulfate (SDS), 500 mM Tris HC1 pH 8.0, 10 mM ethylenediaminetetraacetic acid (EDTA), and 5% Tween is used. This buffer is stored at room temperature. Also, a solution of non-heat-resistant proteinase K stored at -20 °C is used, as well as a 10x tagging buffer containing 50 mM magnesium acetate (MgAc) and 100 mM Tris acetate (TrisAc) pH 7.6, stored at room temperature.

[0137] For the lysis mixture, 25 μΐ^ of lysis buffer, 30 μΐ^ of proteinase K, and 170 μΐ^ of water are mixed for blood or cell samples, while 180 μΐ^ of water is used for BMA samples. Then, 25 μΐ^ of blood or cell sample volume and 15 μΐ^ of BMA sample volume are added to the lysis mixture. The mixture is mixed with wide-bore pipette tips and placed in a heating block set at 37 °C for 15 minutes, then transferred to a second heating block at 55 °C for 10 minutes. After the heat treatment, the samples are cooled to room temperature. At the same time, an inhibition mixture can be prepared with 40% a-cyclodextrin and tagging buffer. Once the lysate has cooled, 50 μΐ^ of this inhibition mixture is added to each sample, after which it is mixed 10 times with wide-bore pipette tips. Then, the samples are ready to be loaded onto the flow cell for further analysis.

[0138] As an alternative workflow, a different type of proteinase K can be used, such as proteinase K from Tritirachium album, which can be inactivated at a higher temperature of 75-80 °C for 10 minutes, or by using an inhibitor such as tetrapeptyl chloromethyl ketone (TCK) at a concentration of 0.05-0.08 mg / mL, which can be combined with the inhibition mixture.

[0139] Figure 7Two plots showing the calculated linkage quality scores of an experimental system by using the estimated zero distribution are shown. The Y-axis represents the Q-value, which is a log scale used to convey accuracy, and where a higher Q-value indicates better quality links. This value can be directly derived from the FDR. The X-axis represents the genomic threshold, while different colors represent various spatial thresholds. As the genomic threshold increases, the linkage quality generally improves until a peak is reached. Beyond this point, increasing the genomic threshold further results in a decrease in the linkage quality score. Each point on the plot corresponds to all links created at a particular genomic threshold and spatial threshold. This plot enables the selection of a line such as Q20, Q25, or Q30, and assigns the corresponding linkage quality score to all data points above this quality threshold. The highest linkage quality score threshold can be balanced with the number of links obtained from this threshold. This approach allows the quantification of a desired threshold or thresholds based on a reproducible linkage quality score. Increasing the threshold, such as from Q20 to Q30, results in a lower linkage rate, as the selection of links becomes more conservative. If the threshold is too strict, a certain percentage of reads with links can be eliminated. Therefore, this approach can be used to balance between a desired Q-score and the number of links obtained. Additionally, these thresholds can be used as a direct input for downstream analysis, such as error correction or variant identification. Additionally, the approach can iteratively eliminate lower linkage rates, which efficiently balances the sequencing system for downstream analysis. In this iterative process, one would start with a lower quality threshold, such as Q20, which includes more links, but potentially with lower confidence. As the threshold is increased, for example to Q30, the criteria for inclusion of links becomes more stringent, resulting in fewer but higher confidence links. This selective process eliminates lower quality links, resulting in a compact dataset optimized for both linkage quality and number. By quantifying and setting a desired threshold, one can reproducibly achieve a balance that suits the specific needs of downstream analysis.

[0140] Similar variations can be observed in NovaSeq 6K runs. Certain runs exhibit lower linkage quality compared to other runs. Factors such as the type of flowcell, experimental conditions such as underloaded DNA or normal DNA amounts, distance between wells, well depth, and flow rate on the flowcell all affect linkage quality. The distinction between linkage formation and linkage detection in genomic sequencing can be used to explain the usefulness of certain methods of the present disclosure. The process of forming linkages is affected by factors such as the use of high molecular weight (HMW) DNA or pulsed loading, which affect how DNA fragments physically associate on the flowcell. On the other hand, setting a quality threshold is about determining the ability to detect these linkages. It is a measure of the confidence of the linkages identified, not a measure of all linkages present, including those that are undetectable. The utility of setting a quality threshold is to balance the number of reads that form detectable linkages. Each threshold value representing a different spatial distance allows for quantification of the quality of an experiment at a given threshold level. However, this approach only quantifies those linkages that are detectable, not all linkages that are actually formed. In some embodiments, more linkages can occur, but can not be detectable. Based on the assigned threshold values and the method of characterizing true linkages and false linkages, linkages occurring at combinations of genome and flowcell distance that are indistinguishable from the zero distribution rate will still exist, but are undetectable. For example, while an improvement in linkage rate can also affect the zero distribution of unlinked reads, a sample with a higher linkage rate for very distant reads can not be detected because these distant linkages will exceed the threshold physical distance, in this case most reads are unlinked. However, in general, a higher percentage (e.g., 40% or 50%) indicates better performance. This approach is applicable to NextSeq 2K runs with higher DNA loading as well as NextSeq 2K runs with underloaded DNA. Finally, compounding (e.g., combining all threshold values) of multiple threshold values results in a significantly improved linkage rate.

[0141] Figure 8 A schematic of a patterned flowcell is shown, as well as variable parameters of the patterned flowcell. As Figure 7 As explained, experimental conditions significantly affect the linkage quality score obtained from linkages in a particular run. Figure 8Examples of pitch and feature size for patterned flow cells and random flow cells are shown. At the top, there are two images labeled respectively. The "random flow cell" image shows a plurality of colored spots dispersed in a non-uniform distribution, representing a flow cell with randomly organized capture sites. The "patterned flow cell" image depicts a more ordered array of colored spots, exemplifying a flow cell with systemically arranged capture sites. Below these images, there are diagrams explaining the concepts of "pitch" and "feature size" and the effects of "reduced pitch" and "reduced feature size." "Pitch" is defined as the center-to-center distance between adjacent features, represented by two orange circles, with a line indicating the measurement between their centers. To the right, the concept of "reduced pitch" is illustrated by moving the features closer together, as indicated by the arrow. Similarly, "feature size" is depicted by the diameter of the orange circles, and "reduced feature size" is represented by the smaller circles with dashed outlines, indicating a reduction in feature size. At the bottom of the figure are two micrographs. On the left is a cross-sectional image of a flow cell with several wells or pits to capture DNA molecules. On the right is a top-down micrograph of a flow cell surface, showing a high density of uniform wells, which are the features in which DNA molecules are captured during the sequencing process. This figure illustrates the structural differences between random flow cells and patterned flow cells, and highlights the concepts of pitch and feature size in flow cell design, which are relevant parameters that affect the methods of the present disclosure. The comparison shows that adjustments in pitch and feature size can affect the distribution and occupancy of DNA fragments on the flow cell, which in turn can affect sequencing output.

[0142] Variable parameters of the flow cell, such as pitch and feature size, can determine several aspects of the sequencing process, including the density with which sequencing reads or fragments can be packed and the extent to which the optical system can distinguish between individual clusters. Changes in these variables can result in differences in sequencing quality and / or linkage quality scores between pairs of reads. For example, a smaller pitch can allow for more DNA clusters, but can also result in overlapping signals between wells due to links originating from different proximal but unrelated wells, which would include both more true links and false links, making it challenging to assign linkage quality scores to reads in a sequence. Conversely, a larger pitch can provide a clearer resolution between true links and false links, resulting in a less significant false discovery rate, but at the cost of reduced throughput or the number of sequences of parallel reads. A larger pitch can also increase the minimum size required for a single polynucleotide molecule to span two nanopores. Experimental conditions including temperature, reagent quality, and flow rate can affect linkage quality scores in combination with the design parameters exemplified in Figure 8 For example, poor experimental conditions can introduce noise into the data and alter the biochemical processes by which a biological sample is fragmented, making it difficult to establish reliable links between pairs of reads, resulting in lower linkage quality scores.

[0143] Figure 9 It is shown how block X-Y coordinates can be converted to flow cell X-Y coordinates, where flow cell X coordinate = block X coordinate x block number; and flow cell Y coordinate = block Y coordinate x strip number. Flow cells often have different coordinate systems depending on the granularity of the data or components of the system. In sequencing technology, flow cells are often divided into blocks, each of which can contain many DNA clusters or spots. Block X-Y coordinates are a way of locating a particular spot within that smaller unit (block). Figure 9 Equations are provided for converting block X-Y coordinates to flow cell X-Y coordinates. Specifically, block X coordinates can be multiplied by the block number to obtain flow cell X coordinates. Similarly, block Y coordinates can be multiplied by the strip number to obtain flow cell Y coordinates. The block number indicates the particular block of interest, while the strip number can indicate the block of a particular row or column on the flow cell. Multiplying the local coordinates by these values converts them to global coordinates, which can be used to locate a spot anywhere on the flow cell. Similar processes can be used to convert spatial coordinates to units of interest, such as units comparable to the length of a DNA molecule or insert size.

[0144] Figure 10 A non-limiting example 1000 of flow cell grid geometry 1005 and spatial unit 1010 is provided. As shown at grid 1005, the nanowells of a flow cell are organized into a grid. The nanowell grid 1005 is overlaid with a pixel grid (not shown) that is derived from an image of the flow cell grid 1005 by a CCD camera. In this example, 1 pixel in the captured image represents a pixel size of 10 FQUs, and each pixel size is 345 nm. Various cameras of different magnification levels will have different pixel sizes, and Figure 10 The pixel size disclosed in the example is provided as an example only. In this example, with the pixel size and the pixel to FQU ratio known, the physical FQU size is determined to be 34.5 nm. Other metrics obtained with this information include nanowell pitch, diameter, and gap distance. In this example, the nanowell grid is organized into a hexagonal pattern.

[0145] In one embodiment, Figure 11A flowchart of a method 1100 for generating a linkage quality score by computing a spatial zero distribution and a genomic zero distribution is provided. At step 1110, the method can proceed by computing a genomic zero distribution of a plurality of read pairs that have approximately zero probability of being linked on a flow cell. The genomic zero distribution of read pairs can be estimated by selecting read pairs that have approximately zero probability of being truly linked to each other. For example, read pairs on opposite surfaces of a flow cell cannot be truly linked to each other because the same template molecule cannot bind to both surfaces of a flow cell at the same time because the two surfaces appear on opposite sides of the flow cell. In another example, read pairs that are far enough apart from each other such that fragments from the original DNA molecule cannot separate that distance on a flow cell can also be selected to derive a zero distribution because these read pairs should also not be linked. In this context, "far" can refer to a predetermined threshold distance that would provide a high probability that the read pairs are not from the same origin. The sufficient distance to conclude that two read pairs are not from the same origin can depend on the underlying statistical distribution of the data and the particular context. Generally, a sufficient distance would be a distance that is significantly greater than the average separation between read pairs that would be expected based on the underlying assumptions about the distribution of fragmented DNA sequences on the surface of a flow cell, which can indicate that two objects are not from the same origin. Thus, identifying candidate links from links that are physically unable to originate from the same nucleic acid molecule provides an estimate of the genomic zero distribution of the entire dataset. These estimates can be readily collected by selecting clusters on opposite surfaces or with spatial data markers located at a far distance opposite ends of the flow cell, e.g., the beginning and end regions of the flow cell. Thus, this subset of the total sequencing data can be used to generate an estimate of the genomic zero distribution. In some embodiments, a full zero distribution is not necessary for generating a meaningful linkage quality score.

[0146] The method 1100 can proceed to step 1120 by computing a spatial zero distribution of a plurality of read pairs that have approximately zero probability of being linked to each other spatially. The read pairs can be aligned to a reference genome, where the read pairs can be given a genomic location, e.g., a genomic location on a chromosome. When two read pairs are on different chromosomes, these two read pairs are less likely to be linked because the reads are from different chromosomes (i.e., different molecules). Thus, an estimate of the spatial zero distribution can be generated from this subset of read pairs of the total sequencing data that indicate reads on different pairs of chromosomes.

[0147] Once the null distribution is calculated, the method can proceed to step 1125 by calculating a false positive rate proportional to the product of the percentage of space and genome null links that are below a threshold space and genome distance. It should be recognized that alternative methods for calculating link quality scores can not employ a false positive rate as recited in step 1125, and thus it can be optional. For example, other positive and negative predictive values can be used to test true positive and true negative results. Link quality scores can be derived from the null distribution to estimate the reliability or confidence of the observed data. These scores can be based on metrics such as signal to noise, deviation from the null distribution, or the likelihood of observing the data given a default expectation. Higher link quality scores indicate greater reliability and confidence of the observed results. Statistical tests of interest can be performed on the observed data to obtain each of the expected test statistics or p-values (false positive, false negative, true positive, true negative). Test statistics can be calculated on the sampled data, where the p-value represents the probability of obtaining a test statistic as extreme or more extreme than the observed value under the default expectation. In other examples, the false positive rate can be used as an intermediate measure to estimate the false discovery rate, or any other type of Type I error measure.

[0148] The method can then proceed to step 1130 by generating a link quality score representing the probability of incorrectly assigning a link between two polynucleotide reads. The link quality score can be generated based on the calculated genome null distribution and the calculated space null distribution. Calculating the false discovery rate (FDR) from the null distribution can involve determining the proportion of false positives among total positive discoveries when there is no true effect. The null distribution can be generated by simulating data under the assumption that there is no true effect or association. The present disclosure provides methods that can avoid the step of fitting a null distribution or simulating a null distribution by randomly shuffling the observed data. As an alternative, the present disclosure provides methods that utilize complementary information on the genomic sequence (spatial link data), and a subset of this data can be selected that would not represent realistic or possible scenarios to estimate the null distribution and generate link quality scores. Thus, the systems and methods of the present disclosure can advantageously improve the quality and efficiency of data processing associated with polynucleotide sequencing. In particular, the practical application of quantifiable linked reads can include improvements in mapping, alignment, and variant identification.

[0149] At step 1135, the method calculates a linkage quality score for different threshold combinations. In some embodiments, step 1135 can include calculating a linkage quality score by determining a zero distribution for each threshold. In some embodiments, step 1135 can include calculating a linkage quality score by determining a false discovery rate for each threshold. In some embodiments, step 1135 can include calculating a linkage quality score by determining a previous zero distribution for a new threshold. In some embodiments, step 1135 can include generating a linkage quality score based on a false discovery rate. By utilizing multiple or different threshold combinations, different linkage quality scores can be achieved, allowing for more informed decisions or more accurate placement of linked read pairs. To calculate a linkage quality score for different threshold combinations, various parameters and metrics can be considered, such as the number of linked read pairs, sequence depth, read overlap, sequence quality, and alignment score. For each combination of these metrics, a threshold can be established, above or below which a linkage is considered to meet that particular quality standard. A linkage quality score can then be calculated based on how many thresholds are met and to what extent those thresholds are met. This can be achieved through a weighted average, or a composite linkage quality score can be generated. The composite linkage quality score can be designed to represent individual scores as a comprehensive metric or a series of metrics that encapsulate the overall quality of linkages in a sequencing experiment. Alternatively or additionally, outputs from different thresholds can be used in a hierarchical manner. High quality connections determined by applying the most stringent threshold can be used or analyzed first. This can be beneficial in applications where computational resources are limited or there is some benefit to quickly identifying the most reliable linkages. Lower quality connections that meet less stringent thresholds can be considered later or as a supplement to high quality linkages.

[0150] Once linkage quality scores have been calculated at step 1135, process 1100 can move to decision step 1140 to determine whether a target threshold or optimal threshold has been reached. If the target threshold has been reached, process 1100 can terminate at an end step. If the target threshold has not been reached, process 1100 can return to step 1125, so the calculations can be repeated in order to determine a target threshold or optimal threshold. A predetermined target threshold for identifying true connections between read pairs can be set based on prior knowledge from previous verifications. For example, a target threshold can be based on setting a minimum number of linked read pairs or a maximum allowed distance between linked read pairs on a flow cell. An optimal threshold can be determined through post-analysis of previous target thresholds. Thresholds can be tailored to maximize particular metrics that can be important for a given task, such as sensitivity, specificity, or a balanced combination of both. The target threshold can be fine-tuned using methods such as cross-validation in order to minimize error, including false positives and false negatives of linkages between read pairs.

[0151] Embodiments of the present disclosure also include systems for analyzing and assembling polynucleotide sequences.Figure 12 is a diagram of an example computing system 1200 that can be used in conjunction with the example sequencing systems. The computing system 1200 can be configured to determine DNA sequences by using the sequencing and assembly methods disclosed herein. The general architecture of the computing system 1200 includes an arrangement of computer hardware and software components. The computing system 1200 can include more (or fewer) elements. However, in order to provide the disclosure that can be implemented, it is not necessary to show all of these generally conventional elements.

[0152] As illustrated, the computing system 1200 includes a processing unit 1210, a network interface 1220, a computer-readable media drive 1230, an input / output device interface 1240, a display 1250, and an input device 1260, all of which can communicate with one another by way of a communication bus. The network interface 1270 can provide connectivity to one or more networks or computing systems. Thus, the processing unit 1210 can receive information and instructions from other computing systems or services via a network. The processing unit 1210 can also communicate with the memory 1270 via the input / output device interface 1240 and further provide output information to the optional display 1250. The input / output device interface 1240 can also accept input from the optional input device 1260, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, game pad, accelerometer, gyroscope, or other input device.

[0153] The memory 1270 can contain computer program instructions (grouped in modules or components in some embodiments) that the processing unit 1210 executes in order to implement one or more embodiments. The memory 1270 generally includes RAM, ROM, and / or other persistent, auxiliary, or non-transitory computer-readable media. The memory 1270 can store an operating system 1272 that provides computer program instructions for use by the processing unit 1210 in the general management and operation of the computing device 1200. The memory 1270 can also include computer program instructions and other information for implementing aspects of the present disclosure.

[0154] For example, in one embodiment, the memory 1270 includes an alignment module 1274 for analyzing and assembling polynucleotide sequences. The module 1274 can perform the methods disclosed herein, including the methods described with respect to the flowcharts of FIGS. 1-3, for example. In addition, the memory 1270 can include and / or be in communication with a data store 1290 that stores one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining DNA sequences and providing assembly processes in accordance with the present disclosure. Figure 2

[0155] System and instrument ​

[0156] Aspects of the present disclosure relate to methods for identifying links across an entire genome. In some embodiments, a method can include receiving a BAM file comprising spatial information. The method can proceed by splitting the BAM file into surfaces and chromosomes. For each subset of the BAM file (surface / chromosome), a "KD-tree" can be constructed, which is a data structure for querying m-dimensional ranges, where m > 1. The method can then proceed for each point p in each KD-tree t. The KD-tree t can then be queried to find all points p_neighbors within a spatial distance threshold of p. Finally, the KD-tree t can be queried to find each point p2 in p_neighbors. If p and p2 are within a genomic distance threshold, the method can determine a link and then record (p, p2) as a link.

[0157] In some embodiments, a method of finding links between pairs of reads on a flow cell can include the step of providing sequencing data for pairs of reads from clusters on a flow cell. The method can also include filtering clusters that are spatially distant from each other, and / or filtering clusters that are distant from each other on a genome. The method can include selecting adjacent clusters that are within a spatial distance threshold as adjacent clusters; and assigning a link to two pairs of reads in the adjacent clusters when the adjacent clusters are within a genomic distance threshold. In some embodiments, assigning a link to two pairs of reads can occur when the genomic distance threshold is a preset threshold. In some embodiments, the method can generate a first subset of sequencing data for pairs of reads by selecting clusters that are spatially distant from each other. In some embodiments, the method can include selecting a first cluster having nucleic acids derived from a first chromosome and a second cluster having nucleic acids from a different second chromosome.

[0158] In some embodiments, the clusters are on the same surface of the flow cell but on opposite ends. In some embodiments, the clusters are opposite surfaces of the flow cell. In some embodiments, calculating a spatial zero-distribution of a plurality of pairs of reads includes determining pairs of reads from clusters on a flow cell. In some embodiments, the first cluster has nucleic acids derived from a first chromosome and the second cluster has nucleic acids from a different second chromosome.

[0159] Various embodiments of the present disclosure can be a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0160] For example, the functionality described herein can be performed as software instructions executed by one or more hardware processors and / or in response to software instructions being executed by one or more hardware processors and / or any other suitable computing device. The software instructions and / or other executable code can be read from computer-readable storage medium (or media). Computer-readable storage medium can also be referred to herein as computer-readable storage device or computer-readable storage devices.

[0161] A computer-readable storage medium can be a tangible device that can retain and store data and / or instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or raised structures in

[0162] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions into the computing / processing device for storage in a computer readable storage medium within the respective computing / processing device.

[0163] Computer readable program instructions for carrying out operations of the present disclosure (also referred to herein as, for example, "code", "instructions", "modules", "applications", "software applications", etc.) can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and a procedural programming language such as the "C" programming language or similar programming languages. The computer readable program instructions can be invoked from other instructions or from themselves, and / or can be invoked in response to detected events or interrupts. Computer readable program instructions configured for execution on a computing device can be provided on a computer readable storage medium and / or as a digital download (and can be initially stored in a compressed or installable format requiring installation, decompression or decryption prior to execution), which can then be stored on a computer readable storage medium. Such computer readable program instructions can be partially or entirely stored on a memory device (e.g., computer readable storage medium) of the computing device executing the instructions, for execution by the computing device. The computer readable program instructions can execute entirely on the user's computer (e.g., computing device of execution), partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, an electronic circuit, including for example a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuit, in order to perform aspects of the present disclosure.

[0164] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0165] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including

[0166] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions can initially be borne on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions and / or modules into its dynamic memory and send the instructions via a modem over a telephone, cable, or optical line. A modem local to the server computing system can receive the data on the telephone / cable / optical line and use a converter device to place the data on the bus. The bus can carry the data to memory, from which a processor can retrieve and execute the instructions. The instructions received by the memory can optionally be stored on a storage device, such as a solid state drive, either before or after execution by the computer processor.

[0167] The flow diagrams and the block diagrams in the drawings are meant only to illustrate possible architectures, functions, and operations for systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical functions (‘’or a part of an instruction” can be implemented by different combinations of hardware circuits and processors that execute software). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. Also, a number of these blocks can be optional. The methods and processes described herein are also not limited to any particular sequence of steps, and the blocks or states relating thereto can be performed in other suitable orders.

[0168] It should also be noted that each of the elements of the block and / or flow diagrams, and combinations thereof, can be implemented by specialized hardware-based systems that perform the specified functions or actions, or combinations of special-purpose hardware and computer instructions. For example, any of the processes, methods, elements, blocks, applications, or other functions (or portions thereof) described in the foregoing detailed description can be embodied in electronic hardware (including both application specific hardware and / or programmable logic), software, or combinations thereof. Any of the foregoing embodiments can be fully automated or partially automated via the electronic hardware.

[0169] Any of the processors described above and / or devices incorporating any of the processors described above can be referred to herein as, for example, a “computer,” “computer device,” “computing device,” “hardware computing device,” “hardware processor,” “processing unit,” or the like. The computing devices of the embodiments described above are generally (but not necessarily) controlled and / or coordinated by operating system software such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating system. In other embodiments, the computing devices can be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide user interface functionality, such as a graphical user interface (“GUI”), and the like.

[0170] References throughout this specification to “one example,” “another example,” “an example,” and so on, mean that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and can or can not be present in other examples. In addition, it should be understood that the described elements for any example can be combined in any suitable manner in the various examples, unless the context clearly dictates otherwise.

[0171] It should be understood that the ranges provided herein include the specified ranges and any values or sub-ranges within the specified ranges, as if such values or sub-ranges were explicitly listed. For example, a range of from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly recited limits of about 2 kbp to about 20 kbp, but also individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc. Furthermore, when using “about” and / or “substantially” to describe a value, it means including minor variations (up to + / - 10%) of the stated value.

[0172] In some embodiments, the methods can be written in any of a variety of suitable programming languages, for example, compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages can be scripting languages such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java, or Python. In some embodiments, the methods can be standalone applications with data input and data display modules. Alternatively, the methods can be computer software products and can include such classes in which distributed objects include applications containing the computational methods as described herein.

[0173] In some embodiments, the methods can be incorporated into existing data analysis software, such as that found on sequencing instruments. Software including the methods implemented by the computers as described herein are installed directly onto a computer system, or are held indirectly on a computer readable medium and loaded onto a computer system as needed. Further, the methods can be located on a computer remote from where the data is generated, such as software found on a server held in another location relative to where the data is generated (such as provided by a third party service provider).

[0174] The assay instrument, desktop computer, laptop computer, or server can include a processor in operative communication with accessible memory containing instructions for implementing the systems and methods. In some embodiments, the desktop computer or laptop computer is in operative communication with one or more computer-readable storage media or devices and / or output devices. The assay instrument, desktop computer, and laptop computer can operate under many different computer-based operating languages, such as those used by Apple-based computer systems or PC-based computer systems. The assay instrument, desktop computer, and / or laptop computer and / or server system can also provide a computer interface for creating or modifying experiment definitions and / or conditions, viewing data results, and monitoring experiment progress. In some embodiments, the output device can be a graphical user interface, such as a computer monitor or computer screen, a printer, a handheld device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (e.g., iPAD), a hard drive, a server, a memory stick, a flash drive, or the like.

[0175] The computer-readable storage device or medium can be any device such as a server, mainframe, supercomputer, tape system, or the like. In some embodiments, the storage device can be located on-site in close proximity to the location of the assay instrument, for example, adjacent or in close proximity to the assay instrument. For example, relative to the assay instrument, the storage device can be located in the same room, in the same building, in an adjacent building, on the same floor in one building, on a different floor in one building, or the like. In some embodiments, the storage device can be located off-site or remote from the assay instrument. For example, relative to the assay instrument, the storage device can be located in a different part of a city, a different city, a different state, a different country, or the like. In embodiments where the storage device is located remote from the assay instrument, communication between the assay instrument and one or more of the desktop computer, laptop computer, or server is typically through an internet connection (wirelessly or with a network cable through an access point). In some embodiments, the storage device can be maintained and managed by the individual or entity directly associated with the assay instrument, while in other embodiments, the storage device can be maintained and managed by a third party, typically remote from the location of the individual or entity associated with the assay instrument. In embodiments as described herein, the output device can be any device for visualizing data.

[0176] The assay instrument, desktop computer, laptop computer, and / or server system itself can be used to store and / or retrieve computer-implemented software programs including computer code for performing and implementing the computational methods as described herein, data for use in implementing the computational methods, and the like. One or more of the assay instrument, desktop computer, laptop computer, and / or server can include one or more computer-readable storage media for storing and / or retrieving computer-implemented software programs including computer code for performing and implementing the computational methods as described herein, data for use in implementing the computational methods, and the like. The computer-readable storage media can include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a magnetic tape, a flash memory stick or card, and the like. In addition, a network, including the Internet, can be a computer-readable storage medium. In some embodiments, the computer-readable storage media refers to computing resources storage that is accessible by a computer network through the Internet or a service provider’s network, rather than, for example, a local desktop computer or laptop computer at a location remote from the assay instrument.

[0177] In some embodiments, the computer-readable storage media for storing and / or retrieving computer-implemented software programs including computer code for performing and implementing the computational methods as described herein, data for use in implementing the computational methods, and the like, is operated and maintained by a service provider in operative communication with the assay instrument, desktop computer, laptop computer, and / or server system through an Internet connection or a network connection.

[0178] In some embodiments, the hardware platform for providing the computing environment includes a processor (i.e., CPU), where processor time and memory layout, such as random access memory (i.e., RAM), are system considerations. For example, smaller computer systems provide inexpensive, fast processors, as well as large memory and storage capabilities. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, the hardware platform for performing the computational methods as described herein includes one or more computer systems having one or more processors. In some embodiments, smaller computer clusters are brought together to create a supercomputer network.

[0179] In some embodiments, the computational methods as described herein are executed on a collection of interconnected or internally connected computer systems (i.e., grid technology) that can run various operating systems in a coordinated fashion. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are examples of coordinating multiple independent computer systems for the purpose of processing large amounts of data. These systems can provide a Perl interface to submit, monitor, and manage large sequence analysis jobs on a cluster in serial or parallel configuration. One aspect of the present disclosure relates to workflow modules that can be integrated into existing workflows. In some embodiments, the workflow module can be a two-lane sequencing module and can be integrated into NGS sequence analysis platforms, such as DRAGEN from Illumina ™ Bio-ID platform.

[0180] Definitions

[0181] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0182] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The use of the terms "including" as well as "comprising" are not limiting. The use of the term "have" is not limiting. As used in this specification, the terms "comprises", "comprising", "includes", "including", "has", "having" or variants thereof are each to be construed as open-ended, i.e., to mean including, but not limited to. As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having" or variants thereof are each to be construed as open-ended, i.e., to mean including, but not limited to. For example, when used in the context of a process, the term "comprising" means that the process includes at least the recited steps, but also can include additional steps. When used in the context of a compound, composition, or device, the term "comprising" means that the compound, composition, or device includes at least the recited features or components, but also can include additional features or components.

[0183] The terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" are used interchangeably herein and refer to a sequence of nucleotides (i.e., ribonucleotides, if of RNA; deoxyribonucleotides, if of DNA; their analogs, or mixtures thereof) of any length covalently linked, wherein the 3' position of the pentose of one nucleotide is joined by a phosphodiester bond to the 5' position of the pentose of the next nucleotide. The terms shall be understood to include analogs of DNA, RNA, cDNA or antibody-oligonucleotide conjugates made from nucleotide analogs as equivalents. The terms apply to single (such as sense or antisense) and double-stranded polynucleotide. As used herein, the term also encompasses cDNA, which is complementary DNA or copy DNA produced from an RNA template, for example, by the action of a reverse transcriptase. The term refers only to the primary structure of the molecule. Thus, the term includes, but is not limited to, triple-, double- and single-stranded deoxyribonucleic acid ("DNA"), and triple-, double- and single-stranded ribonucleic acid ("RNA"). Nucleotides include sequences of any form of nucleic acid. As will be evident from the examples below and elsewhere herein, a nucleic acid can have a naturally occurring nucleic acid structure or a non-naturally occurring nucleic acid analog structure. A nucleic acid can comprise phosphodiester bonds; however, in some embodiments, a nucleic acid can have other types of backbones, including, for example, phosphoramidate, phosphorothioate, phosphodiester, O-methyl phosphoroamidate, and peptide nucleic acid backbones and linkages. A nucleic acid can have a positively charged backbone; a non-ionic backbone and a non-ribosomal backbone. A nucleic acid can also comprise one or more carbocyclic sugars. As specified, a nucleic acid used in a method or composition herein can be single-stranded, or alternatively, double-stranded. In some embodiments, a nucleic acid can comprise portions of both double-stranded and single-stranded sequences, for example, as evidenced by a forked adapter. A nucleic acid can comprise any combination of deoxyribonucleotides and ribonucleotides; and any combination of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, isoguanine; and base analogs, such as nitropyrrole (including 3-nitropyrrole) and nitroindole (including 5-nitroindole), and the like. In some embodiments, a nucleic acid can include at least one wobble base. A wobble base can base pair with more than one different type of base, and can be used, for example, when included in an oligonucleotide primer or insert for random hybridization in a complex nucleic acid sample, such as a genomic DNA sample. Examples of wobble bases include inosine, which can pair with adenine, thymine, or cytosine. Other examples include hypoxanthine, 5-nitroindole, acyclic 5-nitroindole, 4-nitropyrrole, 4-nitroimidazole, and 3-nitropyrrole. Wobble bases that can base pair with at least two, three, four, or more types of bases can be used.

[0184] As used herein, a probability-related phrase and / or term refers to the probability that fragments of a polynucleotide are in close proximity to each other on a flow cell. Relatedly, the probability can be related to the distance between fragments of a polynucleotide in a polynucleotide. In the context of sequencing technologies, such as those used in next generation sequencing (NGS), a polynucleotide is a long chain of nucleotides, which can be DNA or RNA. The process typically involves fragmenting these long molecules into smaller pieces in order to more easily sequence. When these fragments are placed onto a flow cell, which is a glass slide that is used as a solid surface for sequencing reactions, they are typically distributed across the flow cell surface in a uniform and random distribution in space.

[0185] The probability that fragments of a polynucleotide are in close proximity to each other on a flow cell is related to the distance between fragments in the actual polynucleotide chain refers to the likelihood that segments that are closer together along the chain of a polynucleotide will end up positioned closer to each other on a flow cell during sequencing. This correlation can be a result of the fragmentation process or the method used to affix the fragments to the flow cell.

[0186] For example, if a fragmentation method only randomly cuts a polynucleotide, then fragments from the same region of the original polynucleotide will be in close proximity on the flow cell. However, if the fragments are attached to the flow cell in a random manner, then the initial proximity in the polynucleotide can not be preserved. On the other hand, certain methods, such as those according to the present disclosure, can preserve some spatial information, resulting in a higher probability that sequences in a polynucleotide that are close in location remain close on the flow cell. The relevance of this correlation depends on the sequencing technology used. In some technologies, preserving the relative position of fragments can be beneficial to reconstructing the original sequence of a polynucleotide, as it can aid the mapping and assembly process. Here, in the context of genomic fragments, “spatially co-localized” refers to the presence of two or more DNA fragments in close physical proximity to each other in a given space. The term is used to describe how fragments are positioned relative to each other on a flow cell during sequencing.

[0187] As used herein, the term "fragment" when used in reference to a first nucleic acid is intended to mean a second nucleic acid having a portion or part of the sequence of the first nucleic acid. Typically, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physically removing from a larger nucleic acid, by copying or amplifying a region of a larger nucleic acid, by degrading other portions of a larger nucleic acid, combinations thereof, and the like. The term can be similarly used to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus that an individual inherits from one of its parents. A haplotype can include two or more loci from all or a portion of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions, and the like. The term "phasing alleles" refers to the distribution of specific alleles from a particular chromosome or portion thereof. Thus, the "phase" of two alleles can refer to a characterization or representation of the relative position of two or more alleles on one or more chromosomes.

[0188] As used herein, the term "active region" or "region of interest" refers to a segment of the genome that is specifically targeted for sequencing or is currently being analyzed during a step of a sequencing method. These regions can be individual regions or windows that cover multiple sequence reads at a time. When referring to methods of assembly or structural variant detection, the active region is typically the focus of applying advanced sequencing techniques to obtain highly accurate sequences. In the context of structural variant detection, the active region can be scrutinized using specialized techniques that can detect larger scale genomic alterations such as inversions, translocations, or large indels. These variants can not be apparent in typical sequencing methods and often require methods such as paired-end or long read sequencing to span the entire region of interest. This is also relevant for de novo assembly of genomes, where the active region can be targeted for sequencing with higher coverage depth or longer reads for each step of the sequencing process in order to ensure that these important parts of the genome are assembled correctly.

[0189] As used herein, the term "anchor read" refers to a read that can be mapped to a unique location in the genome with high confidence or definitively. Anchor reads are used as reliable reference points in the mapping process, providing high confidence alignments between sequence reads and a reference genome. These anchor reads are often characterized by a high degree of similarity to known sequences in the reference genome, often facilitated by a process of assigning a high quality alignment score based on the number of matches, mismatches, gaps, and other criteria.

[0190] As used herein, the term "flanking" genomic sequencing refers to the stretching of a DNA or RNA segment located at a distance from a particular region of interest, such as an anchor read, a gene, a mutation site, or a repetitive element. These regions can be used as reference points and can not necessarily be directly adjacent to the region of interest. The distance between the flanking region and the target can vary widely, from just a few base pairs to thousands of bases away, depending on the genome and the method used to link reads to anchor reads. For example, some methods of the present disclosure are able to link reads from thousands of bases away and can be even more sensitive to structural variants that are thousands of bases long.

[0191] In the context of anchor sequence reads, as described above, flanking regions are used as reference points for alignment, but do not need to be immediately adjacent to the sequence of interest. Anchor reads can include sequences that are hundreds or even thousands of base pairs away from the flanking regions. These non-adjacent flanking regions are particularly useful when the anchor read includes a repetitive sequence that occurs frequently in the genome, or in identifying structural variants. By identifying unique flanking sequences at a distance, methods according to the present disclosure can still map the anchor read to the correct location on the genome.

[0192] Using distal flanking regions is a useful strategy of the present disclosure for genomic sequencing to enable accurate mapping. It allows for unambiguous alignment of reads that would otherwise be difficult to place due to the presence of repeats or complex sequences. By considering a range of distances for potential flanking regions, various tools can effectively "anchor" reads to their correct location in the genome, which is useful for reliable genome assembly and accurate identification of genetic variants.

[0193] As used herein, in the context of genomic sequencing, the term "unambiguous mapping" refers to the process of correctly and uniquely assigning a sequenced DNA segment to a single location in a reference genome. This means that the sequence of the segment is so unique that it matches one and only one region in the reference genome with high confidence. As an example, because genomes often contain repetitive sequences, challenges in mapping can arise. If a segment is from a repetitive region, the segment can map to multiple locations, resulting in ambiguous mapping. Ambiguity in mapping can complicate genetic analysis and can lead to incorrect conclusions. Therefore, the goal is to achieve unambiguous mapping as much as possible, which is more likely for longer reads, longer synthetic reads, linked reads of long sequences, or segments that include unique sequences flanking a repetitive region.

[0194] As used herein, the term "ambiguous mapping" or "ambiguously mapped" refers to the scenario when a DNA or RNA fragment (nucleotide sequence) aligns to two or more positions in a target polynucleotide sequence, where the confidence of the two or more positions is low and / or the confidence levels are similar. When a genome is sequenced, individual fragments are generated and then typically matched to a reference genome to determine their original location. This process is called mapping. If a read comes from a unique sequence in the genome, the read can be unambiguously mapped. However, if the read originates from a sequence that is repeated, for example, in the genome, the mapping process can find multiple potential origins for the read. These multiple matching positions make it unclear where the read actually came from, hence the term "ambiguous mapping."

[0195] As used herein, the term "alignment field" refers to the data categories within an alignment record that specifically detail the relationship between a sequence read and a reference sequence. These alignment records are typically stored in a standard format, such as the Sequence Alignment / Mapping (SAM) file, which is widely used to store sequence alignment data. The SAM format organizes the alignment information into several pre-defined fields, each of which represents a particular aspect of the alignment. For example, fields such as QNAME (query name), FLAG (alignment properties), RNAME (reference sequence name), POS (alignment position) are standard components of an alignment record. Additional fields include MAPQ (mapping quality), which indicates the confidence of the alignment, and CIGAR (compact idiosyncratic gap alignment report), which succinctly characterizes how the read aligns to the reference, encompassing matches, mismatches, insertions, and deletions.

[0196] As described herein, the alignment fields can be used to interpret the quality and accuracy of the alignment. These fields contain information such as the exact (or approximate) starting position of the alignment on the reference sequence, the sequence of the read itself, the quality scores of each base in the read, and details about the pairing of the read in paired-end sequencing. For example, the CIGAR string can be used to identify mismatches and gaps that can imply changes between the read and the reference.

[0197] As described herein, the alignment fields can also indicate ambiguous alignments if, for example, the MAPQ score is low, which indicates that the alignment of the read to multiple positions in the reference genome is equally good. Another indication of ambiguity can be inferred from the FLAG field, which can indicate whether the read is mapped in the correct pair. An improperly paired read typically results from one read in a pair being confidently mapped to one location while its pair is mapped to another location or not at all. In cases where the reference genome contains repetitive sequences, reads originating from such regions can be mapped to several positions with similar scores, resulting in an ambiguous alignment. Reads that are ambiguously aligned can be flagged and optionally excluded from further analysis.

[0198] As used herein, the term "baseline scenario" (particularly when it relates to the use of ground truth datasets) refers to a set of sequence data that has been validated and used as a comparative standard to evaluate the quality of an assessment. The size of the sequence data can vary from short sequences to long sequences up to the size of the reference genome. A baseline scenario can be generated for a portion of a sequencing dataset and used as a comparison for the remainder of the same sequencing dataset. For example, a portion of sequencing data can be evaluated for a certain metric, such as sequence depth, and used to determine whether the remainder of the sequencing data (or a portion thereof) is abnormal and indicative of a certain genomic variant.

[0199] Ground truth datasets can include sequences with known variants, including single nucleotide polymorphisms (SNPs), insertions, deletions, and other genetic features that have been rigorously tested and deemed highly accurate. These ground truth sets can be employed as a benchmark to evaluate the extent to which a new sequencing run can identify and replicate known genetic variations. They provide a point of comparison to determine the error rate of a new sequencing process by highlighting the differences between the newly sequenced data and the validated sequences.

[0200] As used herein, the term "putative" generally refers to "generally believed or thought to be," which implies a hypothesis based on some evidence but not solid evidence. In the context of genomics, when referring to "putative structural variants," the term indicates that these are structural changes in the genome, such as deletions, duplications, insertions, inversions, or translocations, that have been identified as possible or likely variations from the reference genome but have not yet been fully validated. Putative structural variants are typically identified through computational analysis of genomic data as described herein. Methods according to the present disclosure can predict these variants by analyzing patterns in the sequencing data that indicate deviations from the expected alignment to the reference genome. For example, reads that span the breakpoint junction of an inversion or multiple sets of linked reads, or clusters of reads indicative of a duplication can lead to the identification of putative structural variants. However, these predictions can require further investigation to determine their validity.

[0201] As used herein, in the context of identifying structural variants in polynucleotides, the term "threshold distance" refers to a predefined maximum / minimum distance that sequence reads must fall within relative to anchor sequence reads to be considered relevant, such as, for example, as part of the same structural variant event. The use of a threshold distance can be used to filter out less relevant reads when analyzing high-throughput sequencing data to detect genomic rearrangements, such as deletions, insertions, duplications, inversions, or translocations.

[0202] As described above, anchor sequence reads are those that can be aligned with high confidence to a known location on the reference genome. In the vicinity of these anchor reads, other reads that do not align with a high MAPQ (e.g., MAPQ = 60) can still provide information for variant detection if they are within some proximity (threshold distance). The threshold distance can range depending on the type of structural variant being investigated and the sequencing technology used. For example, for small indel (insertion / deletion), the threshold distance can be very small, typically in the range of a few bases to 50 bases, as the variation is relatively close to the anchor reads. For larger structural variants, the threshold distance can be set to a few hundred bases to a few thousand bases. The larger the expected variant, the larger the distance that can be considered. When a portion of a chromosome has been significantly rearranged, the threshold distance can be very large, spanning tens of thousands of bases to hundreds of thousands of bases, as the reads indicating the breakpoints of such events can be far away from the anchors in the linear genomic sequence.

[0203] These threshold distances can or can not be arbitrary. The threshold can be determined based on empirical evidence and statistical models that take into account the distribution of reads and the expected frequency of sequencing errors or natural genomic variations. By setting appropriate threshold distances, researchers can minimize false positives (incorrectly identifying a variant where there is none) and false negatives (failing to detect an actual variant). The threshold distances as disclosed herein are useful parameters in the bioinformatics pipeline for structural variant detection, balancing sensitivity (detecting true variants) and specificity (not identifying false variants).

[0204] Note that in the context of spatially linked reads, distance can refer to genomic distance or physical distance in the flow cell. As used in this disclosure, the term distance can refer to both (e.g., the threshold distance is applied to both genomic distance and physical distance) and / or can be understood in context to refer to one or the other type of distance.

[0205] As used herein, genomic distance refers to the number of base pairs between two points on a sequence within a genome. Genomic distance is a linear measure that only considers the length of the sequence, independent of the three-dimensional structure of the polynucleotide. For example, if one gene starts at position 100,000 on a chromosome and another gene starts at position 300,000, the genomic distance between them is 200,000 base pairs. As described herein, in the context of identifying structural variants, a threshold genomic distance can be set to determine the distance at which two reads can still be considered potentially relevant to the same structural variant. If two reads are within this threshold genomic distance, they can be analyzed together to identify potential deletions, insertions, or other variants.

[0206] Similarly, the term“physical distance” refers to the actual space between two fragments of a polynucleotide in a flow cell. This distance can reflect the way the DNA was fragmented on the flow cell. When applying a threshold to physical distance, researchers often look at the interaction between DNA segments in three-dimensional space, such as in chromosome conformation capture experiments (e.g., Hi-C). A threshold for physical distance can be used to determine whether two DNA fragments are close enough to one another to have originated from the same original polynucleotide sequence.

[0207] Thresholds for both genomic distance and physical distance can be used to interpret complex genomic data. For genomic distance, thresholds can be applied in the sequence alignment and variant calling process as described herein to decide whether reads should be considered together for variant detection. For example, in paired-end sequencing, if the distance between two reads exceeds the expected genomic distance based on insert size, this can indicate a potential deletion or insertion.

[0208] For physical distance, thresholds are used to analyze the linkage between fragments of a polynucleotide. Here, thresholds can help identify fragments that are spatially co-localized (such as within a physical distance threshold) at a frequency higher or lower than expected chance versus random chance.

[0209] As used herein, the phrase“spatially close” refers to the proximity of sequence fragments relative to one another or within a given space. In a broad sense, this means that the fragments are close to one another in physical distance, which can be measured in units such as nanometers or distance units on a flow cell. It is context-dependent to define what is considered“close.” Close can be defined by a threshold distance that sets a cutoff for how close two points should be considered to be spatially close. Close can also refer to a distance, such as a degree to which two fragments are close to one another, and does not necessarily imply close proximity.

[0210] As used herein, in the context of genomic sequencing, the phrase“spatially linked pair of reads” refers to a pair of DNA sequence reads that originated from the same polynucleotide sequence and are expected to be separated by a distance based on, for example, the size of the fragments. These pairs of reads are considered“linked” because they were physically connected in the genome during, for example, library preparation for sequencing, before the DNA was fragmented.

[0211] Spatially linked pairs of reads are very useful when determining which sequence reads are linked to other reads, such as anchor sequence reads. As described above, anchor sequence reads are reads that have been confidently mapped to a particular location on a reference genome. By looking at spatially linked pairs of reads, researchers can infer where the other fragment should be mapped to the genome. If the second read in the pair is not mapped to the expected location (based on the known length of the DNA fragment), this can indicate a structural variant between the two reads.

[0212] Detecting structural variants including insertions, deletions, inversions, and translocations often presents a challenge because they inherently involve larger, more complex alterations to the genome than single nucleotide polymorphisms (SNPs) or small insertions-deletions. In this context, high-confidence anchor reads become particularly important. When reads are mapped to a reference genome, some reads can align perfectly or nearly perfectly, serving as anchor reads, while other reads can not align well or can align to multiple locations. These less reliably mapped reads can in fact indicate a structural variant, and their accurate mapping often relies on the context provided by anchor reads.

[0213] For example, in the case of a deletion in a sample genome relative to a reference genome, an anchor read can align well on one end but have a "dangling" end that does not align anywhere nearby. The presence of a high-confidence anchor read can provide the context needed to identify that the dangling end is not a sequencing error or artifact but can be part of a structural variant. Likewise, for insertions, translocations, or inversions, an anchor read can provide a stable framework within which unusual or less confidently mapped reads can be understood.

[0214] In paired-end sequencing, one read in the pair can serve as an anchor read, while the other read spans the structural variant. The anchor read ensures that the pair exists in a particular region, giving the bioinformatician confidence to explore the other read in the pair for what it can reveal about structural changes in the genome. Tools specifically for detecting structural variants often use these anchor reads as a starting point for "walking" along the genome to find the boundaries of the structural variant.

[0215] As used herein, the term "nucleotide sequence" is intended to refer to the order and type of nucleotide polymers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats, including, for example, depictions, images, electronic media, symbol series, numeric series, alphabetic series, color series, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g., indicating molecular structure of nucleotide subunits), or at lower resolution (e.g., indicating chromosomal regions, such as haplotype steps). The series of "A," "T," "G," and "C" letters is a well-known representation of a sequence of DNA, which can be related to the actual sequence of a DNA molecule at single nucleotide resolution. A similar representation is used for RNA, except that "U" is used in place of "T" in the series.

[0216] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquids. The substrate can be non-porous or porous. The substrate can optionally be capable of absorbing liquid (e.g., due to porosity), but will generally be sufficiently rigid such that the substrate does not significantly expand upon absorption of liquid, and does not substantially shrink upon removal of liquid by drying. Non-porous solid supports are generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylic, polystyrene, and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon ™ , cyclic olefins, polyimides, etc.), nylon, ceramic, resin, Zeonor, silica or silica-based materials (including silicon and modified silicon), carbon, metal, inorganic glass, optical fiber bundles, and polymers. Solid supports that are particularly useful for some embodiments are located within a flow cell device. Exemplary flow cells are set forth in further detail below.

[0217] As used herein, the term "flow cell" is intended to mean a chamber having a surface through which one or more fluid reagents can flow. Typically, a flow cell will have an inflow inlet and an outflow outlet to facilitate the flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells that can be readily used in the methods of the present disclosure, as well as related fluidic systems and detection platforms, are described, for example, in Bentley et al., Nature, 456:53-59 (2008); WO 04 / 018497, US 7,057,026, WO 91 / 06678, WO 07 / 123744, US 7,429,492, US 7,311,414, US 7,415,019, US 7,405,381, and US 2008 / 0108082, each of which is incorporated herein by reference.

[0218] In many embodiments, the solid support to which the nucleic acids are attached in the methods set forth herein will have a continuous or unitary surface. Thus, fragments can be attached at spatially random locations, with the distance between nearest neighbor fragments (or nearest neighbor clusters derived from fragments) being variable. The resulting array will have a spatial pattern of variable or random features. Alternatively, the solid support used in the methods set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide locations at which the modified nucleic acid polymers or fragments thereof can be attached. Particularly useful repeating patterns are hexagonal patterns, linear patterns, grid patterns, patterns having reflection symmetry, patterns having rotational symmetry, and the like. The features (e.g., wells) to which the modified nucleic acid polymers or fragments thereof are attached can each have a dimension of less than about 1 mm 2 , 500 µm 2 , 100 µm 2 , 25 µm2 10 pm 2 5 pm 2 1 pm 2 500 nm 2 or 100 nm 2 . Alternatively or additionally, each feature can have an area greater than about 100 nm 2 350 nm 2 500 nm 2 1 pm 2 2.5 pm 2 5 pm 2 10 pm 2 100 pm 2 or 500 pm 2 . Nucleic acid clusters or colonies produced by amplification of fragments on an array, whether patterned or spatially random, can similarly have areas within a range above or between upper and lower limits selected from those exemplified above.

[0219] For embodiments of arrays comprising features on a surface, the features can be discrete, separated by interstitial regions. Alternatively, some or all of the features on the surface can be contiguous (i.e., not separated by interstitial regions). Whether the features are discrete or contiguous, the average size of the features and / or the average distance between features can vary, such that the array can be high density, medium density, or lower density. High density arrays are characterized by features having an average spacing of less than about 15 pm. Medium density arrays have an average feature spacing of about 15 pm to 30 pm, while lower density arrays have an average feature spacing of greater than 30 pm. Arrays for use in the present application can have, for example, a feature spacing of less than 100 pm, 50 pm, 10 pm, 5 pm, 1 pm, or 0.5 pm. Alternatively or additionally, the feature spacing can be, for example, greater than 0.1 pm, 0.5 pm, 1 pm, 5 pm, 10 pm, 50 pm, or 100 pm.

[0220] As used herein, the term "origin" is intended to include the source of a nucleic acid molecule, such as a tissue, cell, organelle, compartment, or organism. The term can be used to identify or distinguish the origin of a particular nucleic acid in a mixture that includes the origin of several other nucleic acids. The origin can be a particular organism in a metagenomic sample of organisms having several different species. In some embodiments, the origin will be identified as an individual origin (e.g., an individual cell or organism). Alternatively, the origin can be identified as a species that encompasses several individuals of the same type in a sample (e.g., a species of bacteria or other organism in a metagenomic sample having several individual members of that species as well as members of other species).

[0221] As used herein, the term "surface" when used in reference to a material is intended to mean the outer portion or layer of the material. The surface can be in contact with another material, such as a gas, a liquid, a gel, a polymer, an organic polymer, a second surface of a similar or different material, a metal, or a coating. The surface or regions thereof can be substantially flat. The surface can have surface features, such as pores, pits, channels, ridges, raised regions, pegs, rods, and the like. The material can be, for example, a solid support, a gel, and the like.

[0222] As an example, in some embodiments, fragments derived from long nucleic acid molecules captured at the surface of a flow cell appear in lines across the surface of the flow cell (e.g., if the nucleic acids were stretched prior to fragmentation or amplification) or in clouds on the surface. Further, a physical map of the immobilized nucleic acids can then be generated. Thus, after amplification of the immobilized nucleic acids, the physical map is related to the physical relationship of the clusters. Specifically, the physical map is used to calculate the probability that sequence data obtained from any two clusters are linked, as described in the incorporated material of WO 2012 / 025250. Alternatively or additionally, the physical map can indicate the genome of a particular organism in a metagenomic sample. In the latter case, the physical map can indicate the order of sequence fragments in the genome of the organism; however, it is not necessary to specify the order, and instead, it can be sufficient that two or more fragments are present in a common organism (or other source or origin) to be a sufficient basis for a physical map characterizing the mixed sample and one or more organisms therein. In some embodiments, the polynucleotides can be provided by a sample comprising a plurality of polynucleotides from at least two organisms. In some embodiments, the method can comprise determining linked sequence reads associated with an individual organism. In some embodiments, the method can comprise performing metagenomic analysis on the plurality of polynucleotides.

[0223] In some embodiments, the physical map is generated by imaging the solid support to establish the location of the immobilized nucleic acid molecules across the surface of the solid. In some embodiments, the immobilized nucleic acids are imaged by adding an imaging agent to the solid support and detecting a signal from the imaging agent. In some embodiments, the imaging agent is a detectable label. Suitable detectable labels include, but are not limited to, protons, haptens, radionuclides, enzymes, fluorescent labels, chemiluminescent labels, and / or chromogenic agents. For example, in some embodiments, the imaging agent is an intercalating dye or a non-intercalating DNA binding agent. Any suitable intercalating dye or non-intercalating DNA binding agent can be used as known in the art, including but not limited to those set forth in U.S. 2012 / 0282617, which is incorporated herein by reference.

[0224] In certain embodiments, a plurality of modified nucleic acid molecules is flowed onto a flow cell comprising a plurality of nanochannels. As used herein, the term nanochannel refers to a narrow channel into which a long linear nucleic acid molecule is stretched. In some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 300, 400, 400, 500, 600, 700, 1400, 900, or no more than 500 individual long strands of nucleic acid are stretched across each nanochannel. In some embodiments, individual nanochannels are separated by a physical barrier that prevents individual long strands of target nucleic acid from interacting with multiple nanochannels. In some embodiments, a solid support comprises at least 10, 50, 100, 300, 500, 500, 3000, 5000, 10000, 30000, 50000, 80000, or at least 100000 nanochannels.

[0225] As used herein, the term "target" when used in reference to a nucleic acid polymer is intended to linguistically distinguish the nucleic acid, e.g., from other nucleic acids, modified forms of nucleic acids, fragments of nucleic acids, and the like. Any of the various nucleic acids set forth herein can be identified as a target nucleic acid, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complementary DNA (cDNA), and derivatives or analogs of these nucleic acids.

[0226] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon-element containing composition (e.g., a transposon, a transposon end, a transposon end composition), and catalyzing the insertion or transposition of the transposon-element containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes, and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, which is incorporated by reference herein. While many of the embodiments described herein relate to Tn5 transposases and / or hyperactive Tn5 transposases, it will be appreciated that transposon elements that are capable of being inserted with sufficient efficiency to label a target nucleic acid of use can be used. In particular embodiments, preferred transposition systems are capable of inserting transposon elements to label a target nucleic acid in a random or nearly random fashion. As used herein, the term "transposome" is intended to mean a transposase bound to a nucleic acid. Typically, the nucleic acid is double stranded. For example, a complex can be the product of incubating a transposase with double stranded transposon DNA under conditions that support non-covalent complex formation. The transposon DNA can include, but is not limited to, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions, or other nucleic acids capable of interacting with a transposase, such as a hyperactive Tn5 transposase.

[0227] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule or portion thereof that includes a nucleotide sequence that forms a transposome with a transposase or integrase. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, the transposon element is capable of forming a functional complex with a transposase in a transposition reaction. As non-limiting examples, the transposon element can include a 19-bp outer end ("OE") transposon end, an inner end ("IE") transposon end, or a "mixed end" ("ME") transposon end recognized by wild-type or mutant Tn5 transposases, or Rl and R2 transposon ends, as set forth in the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, which is incorporated by reference herein. The transposon element can comprise any nucleic acid or nucleic acid analog suitable for forming a functional complex with a transposase or integrase in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, non-natural bases, modified backbones, and can comprise a nick in one or both strands.

[0228] A typical NGS sequencing run produces millions of short sequences that are ultimately mapped to a reference genome. Due to ambiguous genomic locations, a certain percentage of good quality reads (1-5%) are discarded. The need to resolve such reads that would otherwise be discarded can be addressed by implementing increased read length (2x500 or long read sequencing), designing specialized processes to map reads to specific regions of the genome (targeted identifiers), using expensive and time consuming library preparation (Illumina CLR), or a combination thereof. However, such methods are expensive, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from the surface of a solid support can be leveraged to identify fragments generated from a single long input fragment and subsequently used to improve mapping of reads in ambiguous locations.

[0229] Various implementations of the disclosure can be implementations of systems, methods, and / or computer program products at any possible degree of integration. Computer program products can include computer readable storage media (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0230] For example, the functionality described herein can be performed as software instructions executed by one or more hardware processors and / or any other suitable computing devices. The software instructions and / or other executable code can be read from computer-readable storage media (or media). Computer-readable storage media can also be referred to herein as computer-readable storage, computer-readable storage device, or computer-readable storage devices.

[0231] A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, semiconductor, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or raised structures in grooves of a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted via a wire cable, as long as the signals have a practical effect such as triggering a device, a system, or a human.

[0232] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0233] Computer-readable program instructions (also referred to herein as, for example, "code," "instructions," "modules," "applications," "software applications," etc.) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and procedural programming languages ​​such as the "C" programming language or similar programming languages. Computer-readable program instructions may be invoked from other instructions or from themselves, and / or may be invoked in response to a detected event or interrupt. Computer-readable program instructions configured to execute on a computing device may be provided on a computer-readable storage medium, and / or as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption prior to execution), which may then be stored on a computer-readable storage medium. Such computer-readable program instructions may be stored, partially or entirely, on a memory device (e.g., a computer-readable storage medium) for execution by the computing device. Computer-readable program instructions may execute entirely on a user's computer (e.g., an execution computing device), partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry in order to perform aspects of this disclosure.

[0234] This document describes aspects of the disclosure with reference to flowchart illustrations and / or step diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each step in the flowchart illustrations and / or step diagrams, and combinations of steps in the flowchart illustrations and / or step diagrams, can be implemented by computer-readable program instructions.

[0235] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including

[0236] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions can initially be borne on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions and / or modules into its dynamic memory and send the instructions via a modem over a telephone, cable, or optical line. A modem local to the server computing system can receive the data on the telephone / cable / optical line and use a converter device to place the data on the bus. The bus can carry the data to memory, from which a processor can retrieve and execute the instructions. The instructions received by the memory can optionally be stored on a storage device, such as a solid state drive, either before or after execution by the computer processor.

[0237] The flow and step diagrams in the drawings illustrate the architecture, functionality, and operations of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each step in the flow or step diagrams can represent a portion of a service, module, segment, or instruction, including one or more executable instructions for implementing the specified logical functions (s). In some alternative implementations, the functions noted in the steps can occur out of the order noted in the figures. For example, two steps shown in succession can in fact be executed substantially concurrently or the steps can sometimes be executed in the reverse order, depending upon the functionality involved. Also, certain steps can be omitted in some implementations. The methods and processes described herein are also not limited to any particular order or sequence, and the steps or states related thereto can be performed in other suitable orders or sequences.

[0238] It should also be noted that each of the steps in the step diagrams and / or flowchart illustrations, and combinations of steps in the step diagrams and / or flowchart illustrations, can be implemented by a special-purpose based hardware-based system that performs specified functions or actions, or a combination of custom hardware and computer instructions. For example, any of the processes, methods, elements, steps, applications or other functionalities (or portions thereof) described in the foregoing detailed sections can be embodied in electronic hardware such as application specific processors (e.g., application specific integrated circuits (ASICs)), programmable processors (e.g., field programmable gate arrays (FPGAs)), dedicated circuits, and / or the like (any of which can also be in combination with custom hardwired logic, logic circuits, ASICs, FPGAs, etc. to implement these technologies, with custom programming / software executing thereon).

[0239] Any of the above-described processors and / or devices incorporating any of the above-described processors can be referred to herein, e.g., as a “computer,” “computer device,” “computing device,” “hardware computing device,” “hardware processor,” “processing unit,” or the like. The computing devices of the above-described embodiments are generally (but not always) controlled and / or coordinated by operating system software such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating system. In other embodiments, the computing devices can be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality such as a graphical user interface (“GUI”), among other things.

[0240] References throughout this specification to “one example,” “another example,” “an example,” and so on, mean that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and can or can not be present in other examples. In addition, it should be understood that the described elements for any example can be combined in a wide variety of ways without departing from the example, unless explicitly stated otherwise.

[0241] It should be understood that the ranges provided herein include the specified ranges and any values or sub-ranges within the specified ranges, as if such values or sub-ranges were explicitly listed. For example, a range of from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly listed limits of from about 2 kbp to about 20 kbp, but also individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc. Furthermore, when using “about” and / or “substantially” to describe a value, this means including minor variations (up to + / - 10%) of the stated value.

[0242] While several examples have been described in detail, it will be appreciated that modifications can be made to the disclosed examples. Therefore, the above description should not be construed as limiting, but merely as exemplification.

[0243] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods described herein can be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the form of the methods described herein can be made without departing from the spirit of the disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.

[0244] Features, materials, characteristics or groups described in conjunction with a particular aspect or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All features and / or steps disclosed in the specification (including any accompanying claims, abstract and drawings) can be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Protection is not sought for any of the examples in their individual parts but for combinations only. Protection is sought solely for the features of the disclosure as disclosed in the specification (including any accompanying claims, abstract and drawings) or any novel combination of such features or any novel step of any method or process so disclosed.

[0245] Also, certain features described in the context of separate implementations in the disclosure can also be implemented in combinations with each other. Conversely, various features described in the context of a single implementation can also be implemented separately or in any appropriate subcombination. Moreover, although features can be described above as acting in certain combinations and / or initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the combination can be claimed as a new combination in its own right, or in a subcombination.

[0246] Moreover, while operations can be depicted in the drawings or described in the specification in a particular order, this should not be understood as requiring or implying that such operations be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated into the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the described operations. Further, the illustrated or other operations can be rearranged or reordered in other implementations. Those skilled in the art will understand that the actual steps taken in the processes shown and / or disclosed can differ from those shown in the figures. Depending on the example, certain of the steps described above can be removed, others can be added, and some steps can be modified.

[0247] For purposes of the present disclosure, certain aspects, advantages, and novel features are described herein. It is to be understood that not necessarily all such advantages can be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the examples described herein can have other merits, and / or can be adapted by one of ordinary skill in the art for use with other technologies and / or under a variety of other conditions.

[0248] Conditional language used herein, such as, among others, "can," "could," "might," "may," "e.g.," and the like, unless specifically stated otherwise, generally are intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language generally is not intended to imply that one or more examples are required to include a feature, element, and / or step in order to implement the example.

[0249] Conjunctive language such as the phrase "at least one of X, Y, and Z," is generally understood to convey the meaning that the item, term, etc. can be either X, Y, or Z, or a combination thereof. Thus, such conjunctive language is generally not intended to imply that certain examples require at least one of X, at least one of Y, and at least one of Z.

[0250] As used herein, degree language such as the terms "about," "approximately," "generally," and "substantially" mean values, amounts, or characteristics that are close to that which is stated, indicated, or claimed, in that the stated, indicated, or claimed values, amounts, or characteristics are within acceptable limits or tolerances when compared to the stated, indicated, or claimed values, amounts, or characteristics.

[0251] The scope of the disclosure is not intended to be limited to the particular examples in this section or elsewhere in this specification, and can be defined by the claims as presented in this section or elsewhere in this specification or in the future. The language of the claims should be construed based on the language employed by the claimant in the claims, and not limited to examples in the specification or described during the prosecution of the application, which should be understood as non-exclusive.

Claims

1. A method of generating a linkage quality score, the method comprising: computing a genomic zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other on a flow cell; computing a spatial zero distribution of a plurality of read pairs having approximately zero probability of linkage to each other; and generating a linkage quality score representing a probability of an incorrect assignment of a linkage between two polynucleotide reads in the plurality of read pairs based on the computed genomic zero distribution and the computed spatial zero distribution.

2. The method of claim 1, wherein computing the genomic zero distribution of the plurality of read pairs comprises determining read pairs from clusters of nucleic acid sequences on the flow cell.

3. The method of claim 2, wherein the clusters are spatially distant from each other on the same or opposite surfaces of a flow cell.

4. The method of claim 3, wherein the clusters are on the same surface of the flow cell but on opposite ends.

5. The method of claim 3, wherein the clusters are on opposite surfaces of the flow cell.

6. The method of claim 1, wherein computing the spatial zero distribution of the plurality of read pairs comprises determining the nucleic acid sequences of read pairs from clusters on the flow cell.

7. The method of claim 6, wherein a first cluster has nucleic acid sequences derived from a first chromosome and a second cluster has nucleic acid sequences from a different second chromosome.

8. The method of claim 1, further comprising applying a threshold level limiting a spatial distance and a genomic distance between the plurality of read pairs on the flow cell.

9. The method of claim 1, further comprising using the quality of the linkage to determine at least one threshold level such that read pairs are derived from the same DNA molecule with a predetermined probability of a spatial distance and a genomic distance.

10. The method of claim 8 or 9, further comprising estimating a linkage rate of total linkages between read pairs on the flow cell for the threshold level.

11. The method of claim 1, further comprising selecting a threshold level based on the linkage quality score.

12. A system for sequencing polynucleotides, the system comprising: at least one processor; and the processor configured to perform a method comprising: retrieving data comprising polynucleotide sequence reads and their spatial locations on a flow cell, wherein the flow cell comprises clusters of amplified nucleic acid fragments of polynucleotides, and wherein a probability of the clusters being located close to each other on the flow cell is related to a genomic distance between the nucleic acid fragments of the polynucleotides, and wherein each cluster provides one or more of the sequence reads; determining linkages between sequence reads originating from the same polynucleotide, wherein a linkage between sequence reads is based on at least one of: (a) linked sequence reads have a linkage quality score of at least 20 on a Phred scale, and linked sequence coverage has an average depth of at least 30 times (30X) of a reference genome; (b) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 45% and the average sequence coverage is at least 70X coverage for all sequence reads; (c) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 35% and the average sequence coverage is at least 90X coverage for all sequence reads; and (d) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 25% and the average sequence coverage is at least 120X coverage for all sequence reads.

13. A method for sequencing a polynucleotide, the method comprising: obtaining sequence reads from a flow cell, wherein the flow cell comprises clusters of amplified nucleic acid fragments of the polynucleotide, and wherein the probability that the clusters are positioned close to each other on the flow cell is related to the genomic distance between the nucleic acid fragments of the polynucleotide, and wherein each cluster provides one or more of the sequence reads; and determining linkages between sequence reads that originated from the same polynucleotide molecule, wherein at least one of the following is true: (a) linked sequence reads have a linkage quality score of at least 20 on the Phred scale and linked sequence coverage has an average depth of at least 30-fold (30X) coverage of a reference genome; (b) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 45% and the average sequence coverage is at least 70X coverage for all sequence reads; (c) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 35% and the average sequence coverage is at least 90X coverage for all sequence reads; and (d) the percentage of sequence reads on the flow cell that have a determined linkage to at least one other sequence read is greater than 25% and the average sequence coverage is at least 120X coverage for all sequence reads.

14. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises using a polynucleotide concentration of less than 10 picomoles in a loading sample on the flow cell.

15. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises loading less than 1 microgram of polynucleotide in a loading sample on the flow cell.

16. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises loading a polynucleotide having at least 60,000 base pairs in a loading sample on the flow cell.

17. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises using a flow rate of less than 250 µL / min for solution flow across the flow cell during sample loading on the flow cell.

18. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises using less than three purification steps on a loaded sample prior to performing sample loading onto the flow cell.

19. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises loading a crude lysate loaded sample onto the flow cell, wherein the crude lysate sample has been lysed but not extracted.

20. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises sequencing the polynucleotides after two or more loading cycles on the flow cell.

21. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises performing more than two wash cycles during sample loading onto the flow cell.

22. The method of claim 13, wherein the percentage of sequence reads from the flow cell that have a link to a second sequence read is greater than 45%.

23. The method of claim 13, wherein obtaining sequence reads from the flow cell comprises introducing one or more additives onto the flow cell, the one or more additives selected from the group consisting of: magnesium, detergent, betaine, DMSO, chelating agents, reducing agents, and PCR enhancers and stabilizers.

Citation Information

Patent Citations

  • Polymerase enzymes and reagents for enhanced nucleic acid sequencing

    US20080108082A1

  • Transposon end compositions and methods for modifying nucleic acids

    US20100120098A1

  • Detection using a dye and a dye modifier

    US20120282617A1

  • Labelled nucleotides

    US7057026B2

  • Ornamental lamp assembly

    US7311414B2