Methods and systems for determining a short tandem repeat genotype for a nucleic acid sample
Patent Information
- Application Number
- PCT/US2026/013888
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2026-02-04
- Publication Date
- 2026-08-27
Smart Images

Figure US2026013888_27082026_PF_FP_ABST
Abstract
Description
ILLINC.866WO / IP-2917-PCT PATENT METHODS AND SYSTEMS FOR DETERMINING A SHORT TANDEM REPEAT GENOTYPE FOR A NUCLEIC ACID SAMPLECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U. S. Provisional Application No.63 / 762,225, filed February’ 24, 2025, the content of which is incorporated by reference in its entirety.BACKGROUNDField
[0002] The present disclosure relates to nucleic acid sequencing systems and methods. In particular, this disclosure relates to systems and methods for determining a short tandem repeat (STR) genotype for a nucleic acid sample and for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence.Description
[0003] STRs are abundant throughout the human genome, occurring every 6 to 10 kbp on average. Common forms of STRs include dinucleotide repeats, trinucleotide repeats, and tetranucleotide repeats wherein these sequences of nucleotide bases are repeated consecutively at a given locus in a genome. STRs are also referred to as microsatellites or simple sequence repeats (SSRs).
[0004] Changes to the patterns or number of STRs has been correlated with a range of neurological and neuromuscular diseases. The expansion of STRs has been found to have a role in many human phenotypes, including the myotonic dystrophies, Fragile X syndrome, Huntington’s disease, hereditary cerebellar ataxias, amyotrophic lateral sclerosis and frontotemporal dementia. In some cases, the size of the expansion of an SI R can be associated with the severity of the disease. Bird TD (1993), Myotonic dystrophy type 1, GeneReviews. University of Washington, Seattle.
[0005] One way to determine the presence of expanded STRs in a patient is through traditional nucleic acid sequencing methods. However, several types of next-generation sequencing methods, including Sequencing by Synthesis (SBS), use a shotgun approach to sequence large genomic DNA fragments, sometimes called template genomic sequences.During SBS sequencing the template genomic sequences are first fragmented into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been sequenced, knowledge of their original position in the genome, and their connectivity and proximity to each other in the original template genomic sequence is lost. This can make it difficult to determine the correct sequence, and original position of certain repeated fragments, since the smaller sequence fragments generated during SBS sequencing may align to many different regions of a template genomic sequence due to the repeated nucleic acid sequences found in STRs.SUMMARY
[0006] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.
[0007] An aspect of the disclosure is directed to a method for determining a short tandem repeat (STR) genotype for a nucleic acid sample, including: obtaining flow cell data including 1) sequence reads from a flow cell including clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample includes STR regions and flanking regions; identifying STR sequence reads by determining sequence reads comprising a repeated STR motif; aligning flaking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region; determining the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads; re-mapping STR sequence readsto the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determining an STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0008] In some embodiments, the repeated STR motif includes a repeated nucleotide sequence of 1-20 base pairs (bp). In some embodiments, the flanking regions include regions of the reference genome extending from a boundary of a STR region to up to about 350 kbp from the boundary of the STR region. In some embodiments, at least a portion of the STR sequence reads have a nucl eotide sequence that is entirely a repeated STR motif.
[0009] In some embodiments, the STR genotype may be determined based on sequence reads which map entirely within an STR region, and sequence reads which map partially to an STR region and partially to a flanking region.
[0010] In some embodiments, identifying STR sequence reads includes comparing sequence reads to STR motif sequences and identifying sequence reads which have an alignment score with a confidence above a predetermined threshold. In some embodiments, identifying STR sequence reads includes identifying sequence reads which map to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences include a repeated STR motif, with an alignment score above a predetermined threshold. In some embodiments, STR sequence reads can be identified from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read.
[0011] In some embodiments, the STR genotype includes at least one estimated STR region length. Further, estimating the STR region length can include determining a linking rate by determining a proportion of the sequence reads across the nucleic acid sample which have a link to another sequence read based on flow cell proximity and genomic proximity; counting in-repeat reads which were re-mapped to the STR region based on proximity with at least one flanking region sequence read; estimating a total number of inrepeat reads based on the linking rate and on the count of in-repeat reads which were re-mapped to the STR region based on proximity to at least one flanking region sequence read; and estimating a STR region length based on the estimated total number of in-repeat reads.
[0012] In some embodiments, the nucleic acid sample includes two haplotypes. The STR genotype may include an estimated STR region length for each haplotype. In someembodiments, sequence reads may be assigned to a haplotype based on flow cell proximity information.
[0013] In some embodiments, the method comprises determining probabilities that sequence reads mapped to an STR region belong to each of two parental haplotypes based on flow cell proximity to at least one sequence read mapped to a flanking region. In some embodiments, an STR region length for each haplotype can be based on STR sequence reads assigned to each haplotype. In some embodiments, estimating an STR region length may be based on sequence reads which map entirely within an STR region, and sequence reads which map partially to an STR region and partially to a flanking region.
[0014] In some embodiments, the alignment of sequence reads to the STR region may be stored in an electronic file.
[0015] In a further aspect, disclosed herein are systems for determining a short tandem repeat (STR) genotype for a nucleic acid sample. In some embodiments the system comprises a memory storing instructions and a processor configured to perform a method comprising: obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions; identifying STR sequence reads by determining sequence reads comprising a repeated STR motif; aligning flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region; determining the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads; re-mapping STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determining a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0016] In a further aspect, disclosed herein is a non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one process, cause the processor to: obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flowcell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions; identify STR sequence reads by determining sequence reads comprising a repeated STR motif; align flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region; determine the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads; re-map STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determine a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0017] In a further aspect, disclosed herein are methods for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence. In some embodiments, the method includes obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments; aligning sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest; identifying sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region; and re-mapping the identified sequence reads to the region of interest on the reference sequence.
[0018] In some embodiments, the flanking region covers a region of the reference genome extending from a boundary of the region of interest up to about 350 kbp from the boundary of the region of interest.
[0019] In some embodiments, the method includes comparing flow cell proximity with sequence read clusters corresponding to at least one sequence read which aligns to the flanking region, and selecting a new mapping location on the reference sequence for the identified sequence reads based on flow cell proximity.
[0020] In some embodiments, the region of interest includes an STR region, and the one or more decoy contiguous sequences includes a repeated STR motif. In some embodiments, the repeated STR motif includes a repeated nucleotide sequence of 1 -20 bp. The one or more decoy contiguous sequences can include a nucleotide sequence that is entirely a repeated STR motif.
[0021] In some embodiments, the method includes aligning the sequence reads to the reference sequence and a plurality of decoy contiguous sequences, wherein the plurality of decoy contiguous sequences includes a plurality of STR sequence motifs. In some embodiments, the method includes re-mapping the identified sequence reads to one of a plurality of STR regions having different STR sequence motifs.
[0022] In some embodiments, the alignment of sequence reads to the region of interest may be stored in an electronic file.
[0023] In a further aspect, disclosed herein are methods of determining an STR genotype for a nucleic acid sample. In some embodiments, the method includes re-mapping sequence reads taken from a nucleic acid sample to a region of interest of a reference genome according to the method of claim 18, wherein the region of interest comprises an STR region; and determining an STR genotype based on the sequence reads re-mapped to the STR region.
[0024] In a further aspect, disclosed herein are systems for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising: obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments; aligning sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest; identifying sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence reads which aligns to the flanking region; and re-mapping the identified sequence reads to the region of interest on the reference sequence.
[0025] In another aspect, disclosed herein is a non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments; align sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest; identify sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region; and re-map the identified sequence reads to the region of interest on the reference sequence.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.
[0027] FIG. 1 A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.
[0028] FIG. IB is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 1A.
[0029] FIG. 2 is a flow diagram that schematically illustrates steps for obtaining sequence reads and flow cell locations from a nucleic acid sample placed on a flow cell.
[0030] FIG. 3A is a flow diagram that schematically illustrates a method for determining a short tandem repeat (STR) genotype for a nucleic acid sample.
[0031] FIG. 3B is a flowchart further illustrating a process of applying a correction factor based on linking rate, that can take place as part of the method shown in FIG. 3 A.
[0032] FIG. 4 is a flow diagram that schematically illustrates a method for re¬ mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence.
[0033] FIG. 5 is a bar graph showing counts of realigned sequence reads for two samples in the FXN gene and a third sample in the DMPK gene and compare the results of STR detection from standard whole genome sequencing (WGS) to sequencing using proximity information according to some embodiments.DETAILED DESCRIPTION
[0034] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.
[0035] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entirety for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.
[0036] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.Overview
[0037] While STR regions are common throughout the genome, it has historically been difficult to determine the sequence of an STR region and / or determine an STR genotype of an individual for a number of reasons. One reason is that because there is variation in the length of a given STR region m different individuals, the length of an STR region in a particular nucleic acid sample may not match the length of the same STR region in the reference sequence used for alignment. This can lead to misalignment issues. In addition, STR regions may also be longer than sequence reads from next generation sequencing (NGS) methods, meaning that NGS sequence reads may not cover the entire length of an STR region. Thus, NGS sequence reads may, in some embodiments, only span one boundary of an STR region or may fall entirely within an STR region and comprise only the repeated STR motif.
[0038] Because the nucleic acid sequence in an STR is repetitive, it can be difficult to determine the length of an STR region based on sequence reads which are not long enough to span an entire STR region. Furthermore, two STR regions in distinct locations in the genome may have identical repeated motifs and thus produce sequence reads which can only ambiguously map to either STR region. Finally, because individuals are diploid, i.e., having two complete sets of chromosomes, there can also be variation between each haplotype in a person’s nucleic acid sample. Thus, the length of a given STR region may be different in each allele, which can also make it difficult to determine allele-specific STR information when there is a lack of phasing information for each STR region.
[0039] In current SBS methods, nucleic acid sequences are fragmented and bound to flow cells. In embodiments of the present disclosure, relatively long nucleic acid fragments, approximately 10 kb or more, are instead flowed across a SBS flow cell and then are bound to the flow cell and fragmented using bound transposons. Clusters of identical bound amplified fragments are then formed on the flow cell for each bound fragment. It has been discovered that the proximity of each fragment cluster on the flow cell may be correlated with the fragment’s location from the original nucleic acid sample. Fragments bound to the flow cell in closer proximity to one another have a higher likelihood of being from the same original nucleic acid fragment than fragments bound more distal to one another on the flow cell. This flow cell proximity information can therefore be used to more accurately determine the location and length of an STR region in a nucleic acid sample.
[0040] In some embodiments, the methods and systems described herein can be used to accurately align sequence fragments taken from SBS sequencing systems and containing STR regions to their original location on a target nucleic acid using flow cell proximity information. The systems and methods can obtain flow cell data including 1) sequence reads from the flow cell, which comprises clusters of nucleic acids from the nucleic acid sample, and 2) geographic locations on the flow cell of the clusters of the nucleic acids. The methods and systems can determine sequence reads that include a repeated STR motif to identify STR sequence reads needing an alignment. For example, the system may identify sequence reads which contain STR sequences and also unique sequences, or spanning reads, which are not part of the STR indicating that the particular read may span a region on the target nucleic acid which flanks the 3’ or 5’ end of the SIR. The methods and systems can align sequence reads to a reference sequence to determine any flanking region sequence reads that map to the regions flanking a known STR region. Once a flanking sequence read has been determined, the methods and systems can determine other sequence reads derived from flow cell clusters which are in proximity to clusters containing the flanking sequence read and also include the same STR repeated nucleotides. The identified sequence reads which are in proximity to the flanking sequence read and contain similar or the same repeated nucleic acids would have a higher probability' of mapping to the same target nucleic acid than sequence reads having the STR sequences, but located more distal to the flanking sequence read. By analyzing the proximity’ of each flanking region to other STR-containing sequence read, the methods andsystems can map (or re-map) STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read. Finally, the methods and systems can determine a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0041] In some embodiments, the methods and systems can perform further steps to further improve the accuracy of STR genotyping. For example, in some embodiments, the methods and systems may apply a correction factor based on the overall linking rate across the nucleic acid sample. In some embodiments, the methods and systems determine a linking rate by determining the proportion of sequence reads that have a link to another sequence read based on flow cell proximity and genomic proximity. The methods and systems can use the linking rate to adjust an estimate of an STR region length,
[0042] As another example, the methods and systems can also use phasing information, in some embodiments, to more accurately determine an STR genotype. In some embodiments, assigning sequence reads to a haplotype advantageously allows the methods and systems to estimate a STR region length for each haplotype. Phasing information can be especially useful when the STR region length is greater than the sequence read length in both haplotypes, and the sequence reads include in-repeat reads (IRRs). In previous methods, because in-repeat reads could come from either allele, it was difficult or not possible assign these reads to either haplotype, and previous methods and systems may have been less accurate in distinguishing STR length between haplotypes. Embodiments of the methods and systems described herein can advantageously use flow cell proximity information to assign STR sequence reads to a particular haplotype, improving accuracy and recall in the detection of STR length variants. Data received and leveraged from the proximity information resolves previously known issues described above, enabling more accurate detection from short read sequencing.
[0043] Embodiments further relate to methods and systems for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence. These methods can be particularly useful when attempting to determine a mapping location for sequence reads that are difficult to map to the reference genome because of sequence similarity with multiple locations on the reference genome. The region of interest can be an STR region,which due to the repeated nucleotides within the STR makes it difficult to map to a particular genomic region.
[0044] To re-map sequence reads, the methods and systems disclosed herein can obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments. The methods and systems can align sequence reads to a reference sequence and to one or more decoy contiguous sequences. The one or more decoy contiguous sequences can be nucleic acid sequences that are separate from, or in addition to, a reference sequence as a possible mapping / alignment location for sequence reads. Each decoy contiguous sequence can comprise a nucleic acid sequence of interest. For example, if the region of interest is an STR region, the one or more decoy contiguous sequences can each comprise tandem repeats of an STR motif. The reference sequence can include a region of interest and a flanking region which is adjacent to the region of interest. The methods and systems can identify sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold, and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region. Then, the methods and systems can re-map the identified sequence reads to the region of interest on the reference sequence.
[0045] In other methods, decoy contiguous sequences may be used to capture and filter or discard “problematic” sequence reads that are difficult to map to a reference sequence, to keep those sequence reads from being used in downstream steps such as sequence assembly or other analyses of a location or region of interest, for example, in variant calling or STR genotyping. This was because the difficult-to-map reads were less likely to have been mapped accurately to the reference genome and would thus complicate those downstream analyses. In contrast, the methods and systems disclosed herein can, in some embodiments, advantageously use one or more decoy contiguous sequences to capture sequence reads and use flow cell proximity information to re-map the sequence reads to the reference sequence, making those sequence reads available for downstream analyses. Thus, the re-mapping methods and systems disclosed herein can further improve the accuracy of such downstream steps because a greater proportion of the sequence reads from the sample are able to be leveraged.Definitions
[0046] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0047] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.
[0048] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art,
[0049] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.
[0050] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).
[0051] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.
[0052] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.
[0053] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to,ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0054] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.
[0055] “Fragmentation” as described herein refers to the shearing or fragmenting of nucleic acid into shorter lengths. Fragmentation methods include enzymatic, physical (including sonication, nebulization, needle shearing, microwave, etc.), and chemical (including depurination, hydrolysis, oxidation, etc.). The terms “fragmenting enzymes” or “enzymebased fragmentation” or “enzyme fragmentation” as used herein refers to enzymes that fragment nucleic acid. The enzymes can be a single enzyme or two or more enzymes that work together to fragment the nucleic acid. Some enzymes work on single stranded nucleic acid whereas others work on double stranded nucleic acid and yet others work on one strand of a double stranded nucleic acid. Fragmenting enzymes can cut randomly or specifically. Non¬ limiting examples of fragmenting enzymes include transposase, restriction enzymes,Argonaute, CRISPR-associated nuclease (Cas), endonucleases, exonuclease, topoisomerase, FragmentaseTM(New England Biolabs, Ipswich, MA). Preferred fragmentation embodiments include methods that fragment while retaining proximity information of the fragments.
[0056] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of " A," " T," " G," and " C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that " T" is replaced with " U" in the series.
[0057] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 900 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. Examples of reference genomes include GRCh38 from the Genome Reference Consortium.
[0058] The term “nucleic acid sample” herein may refer to a sample, typically derived from any organism, including but not limited to animals, plants, fungi, and microbes. For example, such samples may be derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. Alternatively, the sample may be microbial such as bacteria, viral, or fungal. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.
[0059] The sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, andother clinical or laboratory-obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0060] As further used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample (e.g., nucleic acid sample). In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device (including an imaging device, such as a CCD or CMOS) that incorporate nucleobases into growing oligonucleotides to determine nucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell or other nucleotide-sample slide. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more nucleic acid samples seeded in clusters throughout a flow cell. Upon completing a sequencing run, a sequencing device can generate base-call data in a file, such as a binary base call (BCL) sequence file or a fast-all quality (FASTQ) file.
[0061] Relatedly, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample’s sequence (e.g., a genomic or transcriptomic sequence from a sample) or a corresponding adapter sequence. In some cases, a sequencing cycle includes an iteration of both incorporating nucleobases into clusters of oligonucleotides using sequencing chemistry and capturing images of such clusters attached to a nucleotide-sample slide (e.g., a flow cell). Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer (e.g., a sample genomic sequence). For example, in one or more embodiments, each sequencing cycle involves incorporating nucleobases into either a single nucleotide read in which DNA or RNA strands are read in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends but in different cycles. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining aparticular nucleobase added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleobases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within an SBS run. A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For instance, one cluster of oligonucleotides or a set of clusters of oligonucleotides may be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated and another cluster of oligonucleotides or another set of clusters of oligonucleotides may be concurrently undergoing an indexing cycle in which nucleobases corresponding to an indexing sequence for a nucleotide read are incorporated.
[0062] Further, as used herein, the term “nucleotide-sample slide” (or “nucleotide-sample substrate”) refers to a plate or substrate, such as a flow cell, comprising oligonucleotides for sequencing nucleotide sequences from nucleic acid samples or other sample nucleic-acid polymers. In particular, a nucleotide-sample slide can refer to a substrate containing fluidic channels through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, a flow cell (e.g., a patterned flow cell or non-patterned flow cell) may comprise small fluidic channels and oligonucleotide samples that can be bound to adapter sequences on the substrate. In other implementations, a nucleotide-sample slide can be an open substrate with one or more regions for oligonucleotide samples to be analyzed and the oligonucleotide samples may be positioned using charged pads or other means. In yet another implementation, the nucleotide-sample slide can be a membrane having a nanopore through which one or more oligonucleotide samples may pass.
[0063] Relatedly, as used herein, the term “region of a nucleotide-sample slide” (or “nucleotide-sample slide region”) refers to an area that is part of a nucleotide-sample slide. In particular, a region of a nucleotide-sample slide can refer to a discrete portion of a nucleotide- sample slide that differs from other portions of the nucleotide-sample slide. For instance, a region of a nucleotide-sample slide can include a subsection of patterned flow cell comprising one or more wells (e.g., a nano-wells) or a discrete subsection of a non-pattered flow cell (e.g., a subsection corresponding to one or more clusters). In some cases, a region (e.g., section) ofa nucleotide-sample slide includes a tile or a sub-tile of a flow cell having clusters of oligonucleotides growing in parallel.
[0064] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for exampl e, that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques.
[0065] Embodiments described herein can be used with any suitable sequencing chemistry, such as sequencing by synthesis (SBS), sequencing by binding, sequencing by ligation, or nanopore sequencing.
[0066] SBS can be with or without the use of reversible terminators. For example, SBS can be initiated by contacting the target nucleic acids with one or more nucleotides (e.g., labelled, synthetic, modified, or a combination thereof), DNA polymerase, etc. Those features where a primer is extended using the target nucleic acid as a template will incorporate a labeled nucleotide that can be detected. The incorporation time used in a sequencing run can be significantly reduced using the altered polymerases described herein. Optionally, the labeled nucleotides can further include a reversible termination property that terminates further primer extension once a nucleotide has been added to a primer. For example, a nucleotide analoghaving a reversible terminator moiety can be added to a primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can then be repeated n times to extend the primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems, and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al.. Nature 456:53-59 (2008); WO 04 / 018497; WO 91 / 06678; WO 07 / 123744; U. S. Pat. Nos. 7,057,026 B2, 7,329,492 B2, 7,211,414 B2, 7,315,019 B2, 7,405,281 B2, and 8,343,746 B2. Sequence reads can be generated using instruments such as MiniSeq™, MiSeq™, NextSeq™, HiSeq™ and NovaSeq™ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0067] One example of SBS is termed sequencing by binding. One implementation of sequencing by binding includes cycles of initiating sequencing of a template with a reversible blocker on the 3’ end to prevent additional bases from incorporating, interrogating the template by flooding the flow cell with fluorescently tagged bases that do not include a blocker and measuring an emitted signal of bound bases, activating the 3’ end via removal of the reversible blocker, and incorporating the complementary base from unlabeled, blocked nucleotides. Reads using sequencing by binding can be generated from using instruments such as OnsoTMsequencing instruments from Pacific Biosciences of California, Inc. (Menlo Park, CA). Another implementation of sequencing by binding could be sequencing by avidity. In sequencing by avidity, fluorescent dye-labeled cores, termed “avidites” are used. One potential cycle of sequencing by avidity includes providing a reagent of polymerase and reversibly terminated nucleotides to templates immobilized on a solid surface, de-blocking the incorporated nucleotides, flowing a set of four types of avidites, washing away unbound avidites, detecting the incorporated bases / nucleotides, and removing the bound avidites. The steps in the cycle of sequencing by avidity may be performed in other orders. Sequencing by avidity is described in Arslan, S., Garcia, F. J., Guo, M. et al. Sequencing by avidity enables high accuracy with low reagent consumption. Nat Biotechnol 42, 132–138 (2024). doi.org / 10.1038 / s41587-023-01750-7, which is incorporated by reference in its entirety.Reads using sequencing by avidity can be generated using instruments such as AvitiIMsequencing instruments from Element Biosciences (San Diego).
[0068] One example of SBS using an open flow cell and without using reversible terminators is disclosed in Almogy, G.(2022) “Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform” doi.org / 10.1101 / 2022.05.29.493900, which is incorporation by reference in its entirety. Sequence reads using an open flow cell can be generated using instruments such as UG 100TM Sequencer from Ultima Genomics, Inc, (Fremont, CA)
[0069] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are described in U. S. Pat. Nos. 8,262,900 B2, 7,948,015 B2, 8,349,167 B2, and U. S. Pat. Pub.2010 / 0137143 Al, which are incorporated by reference in its entirety,
[0070] Sequence reads can be generated using instruments such as DNBSEQ™ sequencing instruments from MGI Tech Co., Ltd. (Shenzhen, China) and as SURFSeq™, FASTASeq™, and GenoLab™ sequencing instruments from GeneMind Biosciences Co., Ltd. (Shenzhen, China).
[0071] Some embodiments can use methods involving the real-time monitoring of DN / X polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate- labeled nucleotides, or with zeromode waveguides. Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), which are incorporated by reference in its entirety. Techniques for sequencing using zeromode waveguides is described in U. S. Pat. No.6,917,726 B2, which is incorporated by reference in its entirety.
[0072] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood that the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignmentof a read to the reference sequence for human chromosome 13 will tell the likelihood that the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference sequence (e g., chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence,
[0073] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0074] Alignment may be performed by modifications and / or combinations of methods such as Burrows- Wheeler Aligner (BWA), ISAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASFLX, Cloudburst, CUDA-EC, CUSFIAW, CUSHAW2, CUSI IAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRIMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.
[0075] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference sequence, by alignment.
[0076] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.
[0077] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference sequence. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference sequence. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith-Waterman score for local alignments, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.
[0078] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 100 bp, and includes paired end sequence reads,
[0079] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more, A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed by assembling multiple short sequence reads.
[0080] As used herein, a “file” includes electronic files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.
[0081] The terms “solid support,” “solid surface,” and other grammatical equivalents herein refer to any substrate that is appropriate for or can be modified to be appropriate for the attachment of enzymes, nucleic acids, and complexes thereof. As will be appreciated by those in the art, the number of possible substrates is very large. Possible substrates include, but are not limited to, glass and modified or functionalized glass, polymers (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, TeflonIM, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, plastics, optical fiber bundles, quartz, metal oxides, inorganic oxides, other suitable transparent materials, other suitable non-transparent materials,other suitable translucent materials, and combinations thereof. The composition and geometry' of the solid support can vary with its use.
[0082] In some embodiments, the solid support or solid surface is a planar structure, such as a flowcell, slide, chip, microchip, array, microarray, wafer, panel, charge pad, and / or web. The planar structure can be a single surface structure having a single surface of sample / reaction sites. The planar structure can be a dual surface structure. One example of a dual surface structure includes a top substrate having a top surface of sample / reactions sites, a bottom substrate having a bottom surface of sample / reactions sites, and a spacer layer separating the top substrate and the bottom substrate. The solid support or solid surface can be open to direct application of a fluid. One example of an open solid support or open solid surface is an open flow cell having a single surface structure without an inlet port. In some embodiments, the solid support is not necessarily planar, such as, for example, the surface of a well, tube, or other vessel. Nonlimiting examples include the surface of a microcentrifuge tube, a well of a multiwell plate, and the like,
[0083] In some embodiments, the solid support comprises one or more surfaces of a flowcell or flow cell. The term “flowcell” or “flow cell” as used herein refers to a solid surface across which one or more fluid reagents can be flowed. Examples of flowcells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U. S. 7,057,026 B2; WO 91 / 06678; WO 07 / 123744; U. S. 7,329,492 B2; U. S.7,211,414 B2; U. S. 7,315,019 B2; U. S. 7,405,281 B2, and U. S. Pat. Pub. 2008 / 0108082 Al, each of which is incorporated herein by reference in its entirety. In some embodiments, the flowcells can be one or more flow lanes. For flow cells having a plurality of flow lanes, each of the flow lanes can be independently accessed or two or more flow lanes can be accessed as a group.
[0084] In some embodiments, the solid support or solid surface is a non-planar structure, such beads, microspheres, and / or inner and / or outer surface of a tube or vessel. The terms “beads”, “microspheres,” or “particles” or grammatical equivalents herein is refer to small discrete particles. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, acrylamide, paramagnetic materials, tlioria sol, carbon graphite, titanium dioxide, latex, polysaccharide (e.g. DextranfM, SepharoseIM, cellulose, agarose), nylon, cross-linked micelles, TeflonIM, as well as any other materials outlined herein for solid supports may all be used. “Microsphere Detection Guide” from Bangs Laboratories, Fishers Ind. is a helpful guide. In certain embodiments, the microspheres are magnetic microspheres or beads. The beads need not be spherical; irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead sizes range from nanometers, e.g., 100 nm, to millimeters, e.g., 1 mm, with beads from about 0.2 micron to about 200 microns being preferred, and from about 0.5 to about 5 micron being particularly preferred, although in some embodiments smaller or larger beads may be used.
[0085] In some embodiments, the solid support comprises a patterned surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in an ordered patern. A “patterned surface” refers to an arrangement of different regions in or on an exposed layer of a solid support. The features can be separated by interstitial regions that contribute to the pattern. In some embodiments, the interstitial regions can be a different height, creating wells or raised platform patterns. In other embodiments, the interstitial regions can have a different surface charges or surface energies. In yet other embodiments, the interstitial regions can have a different attachment moieties. In some embodiments, the pattern can be any suitable pattern, such as a grid patterns, radial patterns, and combinations thereof. In some embodiments, a patterned surface can contain pre¬ determined locations of features but the features are not arrayed in a repetitive patern. Examples of grid patterns include rectangular patterns, hexagonal patterns, triangular, and other suitable grid patterns. The regions for immobilization of molecules may be depressed regions, elevated regions, or planar regions relative to the interstitial regions. The regions may be fabricated as is generally known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, microetching techniques, and combinations thereof. As will be appreciated by those in the art, the technique used will depend on the composition and shape of the regions. For example, the regions for immobilization of molecules of a patterned surface may be wells, pits, channels, posts, pillars, ridges, stripes, swirls, lines, and other suitable topographies. For example, the wells may have any opening in any shape, such as circular, oval, polygonal (e.g., hexagonal, octagonal, square, rectangular, elliptical, etc.). Exemplary patterned surfaces that can be used in the methodsand compositions set forth herein are described in U. S. Pat. No. 8,778,849 B2, which is incorporated herein by reference in its entirety.
[0086] In some embodiments, the solid support comprises a surface suitable for immobilization of molecules, such as enzymes, nucleic acids, and complexes thereof, in a random distribution over the solid support. Exemplary random distribution over a solid support is described in U. S. Pat. No. 8,241,573 B2, which is incorporated herein by reference in its entirety.
[0087] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow' of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,129,492; US 7,211,414; US 7,115,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.
[0088] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1mm2, 500 um2, 100 pm2, 25 jim2, 10 pm2, 5 pm2, 1 m2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random)can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.
[0089] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially flat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.
[0090] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.
[0091] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element-containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U. S. Pat. App. Pub. 2010 / 0120098, which is incorporated herein by reference in its entirety. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include,without limitation, Tn5 DNA, a portion of Tn5 DNA, a fusion of Tn5 or a portion of Tn5 with one or more auxiliary' proteins, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.
[0092] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end (" OE") transposon end, inner end (" IE") transposon end, or "mosaic end" (" ME") transposon end recognized by a wild-type or mutant Tn5 transposa se, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App, Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, non¬ natural bases, modified backbone, and can comprise nicks in one or both strands.
[0093] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference sequence. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches can be costly, laborious, and time intensive. Spatial information (e.g., X and Y coordinates) obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads m ambiguous positions.
[0094] In one or more embodiments, the system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Suchsequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.
[0095] Moreover, in one or more embodiments, one or more sequencing data files in which the system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.
[0096] Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster” or “colony”) refers to a localized group or collection of DNA or RNA on a nucleotide-sample support, such as a flow cell, particle, polymer scaffold, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, the cluster can comprise one or more concatemers, such as, for example, a polony or a nanoball. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a non-patterned flow cell. In typical embodiments, a cluster is the product of an amplification reaction. A cluster of oligonucleotides can be imaged utilizing one or more light signals, changes in pH, changes in conductance, and other signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent labeled nucleotides incorporated into oligonucleotides, fluorescent labeled nucleotides bound but not incorporated into oligonucleotides, and other fluorescent labeled complexes associated with incorporated or bound nucleotides from one or more clusters on a flow cell. Examples of other sequencing procedures are set forth herein. In some embodiments, a cluster can be monoclonal or polyclonal.
[0097] The term “immobilized”, “affixed” and “attached” are used interchangeably herein and both terms are intended to encompass direct or indirect, covalent or non-covalent attachment unless indicated otherwise, either explicitly or by context.
[0098] Exemplary covalent attachment includes, for example, those that result from the use of click chemistry techniques. Exemplary non-covalent attachment includes, but are not limited to, non-specific interactions (e.g. hydrogen bonding, ionic bonding, van der Waals interactions etc.) or specific interactions (e.g. affinity interactions, receptor-ligand interactions, antibody-epitope interactions, avidin-biotin interactions, streptavidin -biotin interactions, lectin-carbohydrate interactions, etc.). Exemplary attachments are set forth in U. S, Pat. Nos. 6,737,236 Bl; 7,259,258 B2; 7,375,234 B2 and 7,427,678 B2; and U. S. Pat. Pub.2011 / 0059865 Al, each of which is incorporated herein by reference in its entirety.
[0099] In certain embodiments, the molecules (e.g. nucleic acids, enzymes) remain immobilized or attached to the solid support under the conditions in which it is intended to use the solid support, for example in applications requiring nucleic acid amplification and / or sequencing. In other embodiments, the molecules are reversibly immobilized and can be removed from the solid support through the use of cleavable sites, linkers, and the like.
[0100] Some embodiments further comprise amplifying and / or replicating one or more nucleic acid templates, including fragments thereof. The amplifying and / or replicating comprises use of one or more of a bridge amplification reaction, an isothermal bridge amplification reaction, a rolling circle amplification (RCA) reaction, a modified rolling circle multiple displacement amplification, a helicase-dependent amplification reaction, a recombinase-dependent amplification reaction, a single-stranded DNA binding (SSB) protein mediated Isothermal amplification, a PCR reaction, a strand-displacement reaction, a ligase chain reaction, a transcription-mediated reaction, a loop-mediated amplification reaction, other suitable reactions, and combinations thereof. Amplification can occur on the sequencing instrument or separately from the sequencing instrument.
[0101] Some embodiments further comprise rolling circle amplification / replication used to form polonies. The term “polony” or “polonies” used herein refers to a nucleic acid library molecule clonally amplified in-solution or on-support to generate an amplicon that can serve as a template molecule for sequencing. In some aspects, a linear library molecule can be circularized to generate a circularized library molecule, and the circularized library moleculecan be clonally amplified in-solution or on-support to generate a concatemer. In some aspects, the concatemer can serve as a nucleic acid template molecule which can be sequenced. The concatemer is sometimes referred to as a polony. In some aspects, a polony includes nucleotide strands.
[0102] Some embodiments further comprise rolling circle amplification / replication used to form nucleic acid nanoballs. The term “nucleic acid nanoball” may be a concatemer comprising multiple copies of a target nucleic acid molecule. These nucleic acid copies may be arranged one after another in a continuous linear strand of nucleotides. These nucleic acid copies may result in a nanoball folding configuration. The multiple copies of a target nucleic acid molecule in a nucleic acid nanoball may each contain an adaptor sequence of known sequence to facilitate amplification or sequencing. The adaptor sequence of each target nucleic acid molecule may be the same or different. The nucleic acid nanoball can be loaded on the surface of solid support. The nanoball can be attached to the surface of solid support by any suitable method. Non-limiting examples of such methods include nucleic acid hybridization, biotin streptavidin binding, thiol binding, photoactive binding, covalent binding, antibody¬ antigen, physical constraints via hydrogels or other porous polymers, etc., or combinations thereof. In some cases, the nanoball can be digested with an enzyme (nuclease, etc.) to produce a smaller nanoball or a fragment from the nanoball.
[0103] In some embodiments, sequence reads comprise a barcode. As used herein “barcode” refers to a short, unique nucleic acid sequence used to tag or label different nucleic acid samples or different nucleic acid molecules. In some embodiments, long DNA template molecules are fragmented, and a barcode sequence is added to each fragment, with a unique barcode sequence for each long DNA template. The fragments are sequenced to produce short sequence reads. In some embodiments, the sequence reads are analyzed, and the barcodes are used to identify which sequence reads are part of the same long DNA template molecule. Thus, barcodes can help retain connectivity information for short sequence reads. Such embodiments can be considered an alternative form of retaining connectivity information between DNA fragments and can be used in the methods described herein.
[0104] In any of the embodiments summarized herein, the analytes are obtained from a population of cells, a single cell, a population of cell nuclei, or a cell nucleus. In any of the embodiments summarized herein, the analytes are analyzed using various analyses,depending on what the analyte is. For example, analysis may include DNA analysis, RNA analysis, protein analysis, tagmentation, nucleic acid amplification, nucleic acid sequencing, nucleic acid library preparation, assay for transposase accessible chromatic using sequencing (ATAC-seq), contiguity-preserving transposition (CPT-seq), single cell combinatorial indexed sequencing (SCI-seq), or single cell genome amplification, or any combination thereof.
[0105] In some embodiments, a sample includes a single cell, and the single cell is fixed. In some embodiments, the cells can be fixed with a fixative. As used herein, a fixative generally refers to an agent that can fix cells. For example, fixed cells can stabilize protein complexes, nucleic acid complexes, or protein-nucleic acid complexes in the cell. Suitable fixatives and cross-linkers can include, alcohol or aldehyde based fixatives, formaldehyde, glutaraldehyde, ethanol-based fixatives, methanol-based fixatives, acetone, acetic acid, osmium tetraoxide, potassium dichromate, chromic acid, potassium permanganate, mercurials, picrates, formalin, paraformaldehyde, amine-reactive NHS-ester crosslinkers such as bis[sulfosuccinimidyl] suberate (BS3), 3,3'-dithiobis sulfosuccinimidylpropionate] (DTSSP), ethylene glycol bis[sulfosuccinimidylsuccmate] (sulfo-EGS), disuccinimidyl glutarate (DSG), dithiobis[succinimidyl propionate] (DSP), disuccinimidyl suberate (DSS), ethylene glycol bis[succinimidylsuccinate] (EGS), NHS-ester / diazirine crosslinkers such as NHS- diazirine, NHS-LC-diazirine, NHS-SS-diazirine, sulfo-NHS-diazirine, sulfo-NHS-LC-diazirine, and sulfo-NHS-SS-diazirine. In some embodiments, fixing a cell preserves the internal state of the cell thereby preventing modification of the cell during subsequent analysis or during performance of an assay.
[0106] In some embodiments, the sample includes a nucleic acid source, such as a single cell, a single nucleus, or a population of cells or population of nuclei, and the single cell, single nucleus, population of cells, or population of nuclei is encapsulated within a droplet. In some embodiments, the cell is fixed prior to encapsulation. As used herein, a droplet may include a hydrogel bead, which is a bead for encapsulating a single cell, and composed of a hydrogel composition. In some embodiments, the droplet is a homogeneous droplet of hydrogel material or is a hollow droplet having a polymer hydrogel shell. Whether homogenous or hollow, a droplet may be capable of encapsulating a single cell. As used herein, the term “hydrogel” refers to a substance formed when an organic polymer (natural or synthetic) is cross-linked via covalent, ionic, or hydrogen bonds to create a three-dimensionalopen-lattice structure that entraps water molecules to form a gel. In some embodiments, the hydrogel may be a biocompatible hydrogel. As used herein, the term “biocompatible hydrogel” refers to a polymer that forms a gel that is not toxic to living cells and allows sufficient diffusion of oxygen and nutrients to entrapped cells to maintain viability. In some embodiments, the hydrogel material includes alginate, acrylamide, or poly-ethylene glycol (PEG), PEG-acrylate, PEG-amine, PEG-carboxylate, PEG-dithiol, PEG-epoxide, PEG-isocyanate, PEG-maleimide, polyacrylic acid (PAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), polystyrene sulfonate (PSS), polyvinylpyrrolidone (PVPON), N, N’-bis(acryloyl)cystamine, polypropylene oxide (PPO), poly(hydroxyethyl methacrylate) (PHEMA), poly(N-isopropylacrylamide) (PNIPAAm), poly(lactic acid) (PLA), poly(lactic-co-glycolic acid) (PLGA), poly caprolactone (PCL), poly(vinylsulfonic acid) (PVSA), poly(L-aspartic acid), poly(L-glutamic acid), polylysine, agar, agarose, heparin, alginate sulfate, dextran sulfate, hyaluronan, pectin, carrageenan, gelatin, chitosan, cellulose, collagen, bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethyleneglycol diallyl ether, ethyleneglycol diacrylate, polymethyleneglycol diacrylate, polyethyleneglycol diacrylate, trimethylopropoane trimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetracrylate, or combinations or mixtures thereof. In some embodiments, the hydrogel is an alginate, acrylamide, or PEG based material. In some embodiments, the hydrogel is a PEG based material with acrylate-dithiol, epoxide-amine reaction chemistries. In some embodiments, the hydrogel forms a polymer shell that includes PEG-maleimide / dithiol oil, PEG-epoxide / amine oil, PEG-epoxide / PEG-amine, or PEG-dithiol / PEG-acrylate. In some embodiments, the hydrogel material is selected in order to avoid generation of free radicals that have the potential to damage intracellular biomolecules. In some embodiments, the hydrogel polymer includes 60-90% fluid, such as water, and 10-30% polymer. In certain embodiments, the water content of hydrogel is about 70-80%. As used herein, the term “about” or “approximately”, when modifying a numerical value, refers to variations that can occur in the numerical value. For example, variations can occur through differences in the manufacture of a particular substrate or component. In one embodiment, the term “about” means within 1%, 5%, or up to 10% of the recited numerical value.
[0107] Hydrogels may be prepared by cross-linking hydrophilic biopolymers or synthetic polymers. Thus, m some embodiments, the hydrogel may include a crosslinker. Asused herein, the term “crosslinker” refers to a molecule that can form a three-dimensional network when reacted with the appropriate base monomers. Examples of the hydrogel polymers, which may include one or more crosslinkers, include but are not limited to, hyaluronans, chitosans, agar, heparin, sulfate, cellulose, alginates (including alginate sulfate), collagen, dextrans (including dextran sulfate), pectin, carrageenan, poly lysine, gelatins (including gelatin type A), agarose, (meth)acrylate-oligolactide-PEO-oligolactide-(meth)acrylate, PEO-PPO-PEO copolymers (Pluronics), poly(phosphazene), poly(methacrylates), poly(N-vinylpyrrolidone), PL(G)A-PEO-PL(G)A copolymers, poly(ethylene imine), polyethylene glycol (PEG)-thiol, PEG-acrylate, acrylamide, N, N’-bis(acryloyl)cystamine, PEG, polypropylene oxide (PPO), polyacrylic acid, poly(hydroxy ethyl methacrylate) (PHEMA), poly(methyl methacrylate) (PMMA), poly(N-isopropylacrylamide) (PNIPAAm), poly(lactic acid) (PLA), poly(lactic-co-gly colic acid) (PLGA), poly caprolactone (PCI.,), poly(vinylsulfonic acid) (PVSA), poly(L-aspartic acid), poly(L-glutamic acid), bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethyleneglycol diallyl ether, ethyleneglycol diacrylate, polymethyleneglycol diacrylate, polyethyleneglycol diacrylate, trimethylopropoane trimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetracrylate, or combinations thereof. Thus, for example, a combination may include a polymer and a crosslinker, for example polyethylene glycol (PEG)-thiol / PEG-acrylate, acrylamide / N, N’-bis(acryloyl)cystamine (B / XCy), or PEG / polypropylene oxide (PPO). In some embodiments, the polymer shell includes a four-arm polyethylene glycol (PEG). In some embodiments, the four-arm polyethylene glycol (PEG) is selected from the group consisting of PEG-acrylate, PEG-amine, PEG-carboxylate, PEG-dithiol, PEG-epoxide, PEG-isocyanate, and PEG-maleimide
[0108] In some embodiments, the crosslinker is an instantaneous crosslinker or a slow crosslinker. An instantaneous crosslinker is a crosslinker that instantly crosslinks the hydrogel polymer, and is referred to herein as click chemistry. Instantaneous crosslinkers may include dithiol oil + PEG-maleimide or PEG epoxide + amine oil. A slow crosslinker is a crosslinker that slowly crosslinks the hydrogel polymer, and may include PEG-epoxide + PEG-aniine or PEG-dithiol + PEG-acrylate. A slow crosslinker may take more than several hours to crosslink, for example more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 hours to crosslink. In some embodiments provided herein, droplets are formulated by an instantaneous crosslinker, andtliereby preserve the cell state better compared to a slow crosslinker. Without wishing to be bound by theory, cells may possibly undergo physiological changes by intracellular signaling mechanisms during longer crosslinking times.
[0109] In some embodiments, a crosslinker forms a disulfide bond in the hydrogel polymer, thereby linking hydrogel polymers. In some embodiments, the hydrogel polymers form a hydrogel matrix having pores (for example, a porous hydrogel matrix). These pores are capable of retaining sufficiently large particles, such as a single cell or nucleic acids extracted therefrom within the droplet, but allow other materials, such as reagents, to pass through the pores, thereby passing in and out of the droplets. In some embodiments, the pore size of the droplets is finely tuned by varying the ratio of the concentration of polymer to the concentration of crosslinker. In some embodiments, the ratio of polymer to crosslinker is 30:1, 25:1, 20:1, 19:1, 18:1, 17:1, 16:1, 15:1, 14:1, 13:1, 12:1, 11:1, 10:1, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:15, 1:20, or 1:30, or a ratio within a range defined by any two of the aforementioned ratios. In some embodiments, additional functions such as DNA primer, or charged chemical groups can be grafted to polymer matrix to meet the requirements of different applications.
[0110] As used herein, the term “porosity” means the fractional volume (dimension-less) of a hydrogel that is composed of open space, for example, pores or other openings. Therefore, porosity measures void spaces in a material and is a fraction of volume of voids over the total volume, as a percentage between 0 and 100% (or between 0 and 1). Porosity of the hydrogel may range from 0.5 to 0.99, from about 0.75 to about 0.99, or from about 0.8 to about 0.95.
[0111] In some embodiments, the droplet can have any pore size that allows for sufficient diffusion of reagents while concomitantly retaining the single cell or nucleic acids extracted therefrom. As used herein, the term “pore size” refers to a diameter or an effective diameter of a cross-section of the pores. The term “pore size” can also refer to an average diameter or an average effective diameter of a cross-section of the pores, based on the measurements of a plurality of pores. The effective diameter of a cross-section that is not circular equals the diameter of a circular cross-section that has the same cross-sectional area as that of the non-circular cross-section. In some embodiments, the hydrogel can be swollen when the hydrogel is hydrated. The sizes of the pores size can then change depending on thewater content in the hydrogel. In some embodiments, the pores of the hydrogel can have a pore of sufficient size to retain the encapsulated cell within the hydrogel but allow reagents to pass through. In some embodiments, the interior of the droplet is an aqueous environment. In some embodiments, the single cell disposed within the droplet is free from interaction with the polymer shell of the droplet and / or is not in contact with the polymer shell. In some embodiments, a polymer shell is formed around a cell, and the cell is in contact with the polymer shell due to the polymer shell being brought to the cell surface due to passive adsorption or in a targeted manner, such as by being attached to an antibody or other specific binding molecule,
[0112] In some embodiments, the droplet is of a sufficient size to encapsulate a single cell. In some embodiments, the droplet has a diameter of about 20 gm to about 200 gm, such as 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 gm, or a diameter within a range defined by any two of the aforementioned values. The size of the droplet may change due to environmental factors. In some embodiments, the droplets expand when they are separated from the continuous oil phase and are immersed in an aqueous phase. In some embodiments, expansion of the droplet increases the efficiency of performing assays on the genetic material inside the encapsulated cells. In some embodiments, expansion of the droplet creates a larger environment for indexed inserts to be amplified during PCR, which may otherwise be restricted in current cell based assays.
[0113] In some embodiments, a droplet is prepared by dynamic means, such as by vortex assisted emulsion, microfluidic droplet generation, or valve based microfluidics. In some embodiments, the droplets are formulated in a uniform size distribution. In some embodiments, the size of the droplets is finely tuned by adjusting the size of the microfluidic device, the size of the one or more channels, or the flow rate through the microfluidic channels. In some embodiments, the resulting droplet has a diameter ranging from 20 to 200 gm, for example, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 gm, or a diameter within a range defined by any two of the aforementioned values.
[0114] In some embodiments, analyzing one or more analytes may include various analyses, depending on what the analyte is. For example, analyzing may include DNA analysis, RNA analysis, protein analysis, tagmentation, nucleic acid amplification, nucleic acidsequencing, nucleic acid library preparation, assay for transposase accessible chromatic using sequencing (ATAC-seq), contiguity-preserving transposition (CPT-seq), single cell combinatorial indexed sequencing (SCI-seq), or single cell genome amplification, or any combination thereof.
[0115] DNA analysis refers to any technique used to amplify, sequence, or otherwise analyze DNA. DNA amplification can be accomplished using PCR techniques or pyrosequencing. DNA analysis may also comprise non-targeted, non-PCR based DNA sequencing (e.g., metagenomics) techniques. As a non-limiting example, DNA analysis may include sequencing the hyper-variable region of the 16S rDNA (ribosomal DNA) and using the sequencing for species identification via DNA.
[0116] RNA analysis refers to any technique used to amplify, sequence, or otherwise analyze RNA, The same techniques used to analyze DNA can be used to amplify and sequence RNA. RNA, which is less stable than DNA is the translation of DNA in response to a stimuli. Therefore, RNA analysis may provide a more accurate picture of the metabolically active members of the community and may be used to provide information about the community function of organisms in a sample. Further, simultaneous analysis of both DNA and RNA may be beneficial to efficiently determination of both DNA and RNA related interrogations. Nucleic acid sequencing refers to use of sequencing to determine the order of nucleotides in a sequence of a nucleic acid molecule, such as DNA or RNA.
[0117] As used herein, the term “reagent” describes an agent or a mixture of two or more agents useful for reacting with, interacting with, diluting, or adding to a sample, and may include agents used in assays described herein, including agents for lysis, nucleic acid analysis, nucleic acid amplification reactions, protein analysis, tagmentation reactions, ATAC-seq, CP T-seq, or SCI-seq reactions, or other assays. Thus, reagents may include, for example, buffers, chemicals, enzymes, polymerase, primers having a size of less than 50 base pairs, template nucleic acids, nucleotides, labels, dyes, or nucleases. In some embodiments, the reagent includes lysozyme, proteinase K, random hexamers, polymerase (for example, Φ29 DNA polymerase, Taq polymerase, Bsu polymerase), transposase (for example, Tn5), primers (for example, P5 and P7 adaptor sequences), ligase, catalyzing enzyme, deoxy nucleotide triphosphates, buffers, or divalent cations.Systems
[0118] Additional aspects include electronic systems for determining a short tandem repeat (STR) genotype for a nucleic acid sample. In some embodiments, the system includes a processor configured to perform any of the methods described herein. In some embodiments, the system is configured to obtain data from a SBS sequencing system which has determined the sequences reads on a flow cell. The system may determine the nucleotide sequence of each sequence read and the geographic location on the flow cell of the clusters of nucleic acids which were used to sequence the particular read. The system may also determine which of the sequence reads contain a repeated STR motif. The system may also determine sequence reads which align to flanking regions, which are a genomic region adjacent to a known STR region. The system may then determine the flow cell proximity of sequence reads derived from flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads. The system may map and / or re-map STR sequence reads based on flow cell proximity with one or more of the flanking region sequence reads. The system may then determine a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0119] Additional aspects include systems for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence. In some embodiments, the system is configured to obtain data from a SBS sequencing system which has determined the sequences reads on a flow cell. The system may determine the nucleotide sequence of each sequence read and the geographic location on the flow cell of the clusters of nucleic acids which were used to sequence the particular read. The system may align the sequence reads to the reference sequence and to one or more decoy contiguous sequences. The one or more decoy contiguous sequences can be a set of nucleic acid sequences that are separate from or in addition to the reference sequence. The one or more decoy contiguous sequences can each include a nucleic acid sequence of interest. For example, in the case of re-mapping STR sequence reads, the one or more decoy contiguous sequences can include a plurality of decoy contiguous sequences, each with a different repeated STR motif. The reference sequence can include a region of interest and a flanking region which is adjacent to the region of interest. For example, the region of interest can be an STR region. The system can proceed to identify sequence reads which map to the one or more decoy contiguous sequences with an alignmentscore above a threshold, for example a mapping quality (MAPQ) score a threshold. The system can store this set of sequence reads in a data store. The system can further identify which of that set of sequence reads are proximate on the flow cell to at least one sequence reads which aligns to the flanking region. For example, the system can query the set of sequence reads to identify sequence reads which have flow cell distance within a threshold to at least one flanking region sequence read. The system can proceed to re-map the identified sequence reads to the region of interest on the reference sequence.
[0120] FIG. 1A illustrates a diagram of an environment in which a system for determining a STR genotype for a nucleic acid sample can operate in accordance with one or more implementations. The following paragraphs describe the STR genotyping system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 1 A illustrates a schematic diagram of a computing system 1000 in which a STR genotyping application 1106 operates in accordance with one or more implementations. As illustrated, the computing system 1000 includes one or more server device(s) 1102 connected to a user client device 1108, a local device 1118, and a sequencing device 1114 via a network 1112. The network 1112 can comprise any suitable network over which computing devices can communicate.
[0121] As shown in FIG. 1A, the computing system 1000 includes the server device(s) 1102. In various implementations, the server device(s) 1102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 1102 receive various data from the sequencing device 1114, such as data from a sample genome, flow cell cluster locations, and / or sequence reads. The server device(s) 1102 may also communicate with the user client device 1108. In particular, the server device(s) 1102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 1108.
[0122] As shown, the server device(s) 1102 includes a sequencing application 1110. In general, the sequencing application 1110 analyzes the data (such as call data) received from the sequencing device 1114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 1110 can receive raw data from the sequencing device 1114 and determine a nucleobase sequence for a sample genome or anucleic-acid segment. In some implementations, the sequencing application 1110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.
[0123] As also shown, the sequencing application 1110 includes the STR genotyping application 1106. As described below, in some embodiments, the STR genotyping application 1106 can determine a STR genotype for a nucleic acid sample. For example, in some embodiments, the STR genotyping application 1106 obtains flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions; identifies STR sequence reads by determining sequence reads comprising a repeated STR motif; align flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region; determines the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads; re-maps STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determines a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0124] While the sequencing application 1110 has been described as including the STR genotyping application 1106, other systems or methods may be included within the sequencing application 1110, such as an application to re-map sequence reads, detect sequence variants or to assemble sequence reads (not illustrated).
[0125] Moreover, while the STR genotyping application 1106 is described being implemented on the server device(s ) 1102, as part of the sequencing application 1110, in some implementations, the S TR genotyping application 1106 is implemented by (such as located entirely or in part) on the user client device 1108, the sequencing device 1114, and / or the local device 1118. As mentioned, in some implementations, STR genotyping application 1106 is implemented by one or more other components of the computing system 1000, such as the sequencing device 1114. In particular, the STR genotyping application 1106 can be implemented in a variety of different ways across the server device(s) 1102, the network 1112, the user client device 1108, the local device 1118, and the sequencing device 1114.
[0126] As further shown in FIG. 1A, the computing system 1000 includes the user client device 1108. In various implementations, the user client device 1108 can generate, store, receive, and send digital data. In particular, the user client device 1108 can receive the data from the sequencing device 1114. As further illustrated, the user client device 1108 includes a sequencing application 1110. The sequencing application 1110 may be a web application or a native application stored and executed on the user client device 1108 (for example, a mobile application, desktop application, or web application). The client device 1108 can receive data from the sequencing application 1110 and / or STR genotyping application 1106. For example, the user client device 1108 can receive variant call files and / or alignment files from the sequencing application 1110.
[0127] The sequencing application 1110 can also include instructions that (when executed) cause the user client device 1108 to receive data from the STR genotyping application 1106 and present data from the sequencing device 1114 and / or the server device(s) 1102. Furthermore, the sequencing application 1110 can instruct the user client device 1108 to display data for variant calls, such as nucleobase calls or an indication of a sequence variant or genotype. Indeed, the user client device 1108 can display nucleobase call results for a genome sample and / or an indication of a predicted variant or genotype.
[0128] As further shown in FIG. 1A, the computing system 1000 includes the sequencing device 1114. In various implementations, the sequencing device 1114 can sequence a nucleic acid sample or other nucleic-acid polymer. For example, the sequencing device 1114 analyzes nucleic-acid segments or oligonucleotides extracted from nucleic acid samples to generate data either directly or indirectly on the sequencing device 1114. More particularly, the sequencing device 1114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from nucleic acid samples. In one or more implementations, the sequencing device 1114 utilizes sequencing by synthesis (SBS) to sequence a nucleic acid sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 1112, in some implementations, the sequencing device 1114 bypasses the network 1112 and communicates directly with the user client device 1108.
[0129] As further depicted in FIG. 1A, in some implementations, the server device(s) 1102 includes a distributed collection of servers, where the server device(s) 1102include several server devices distributed across the network 1112 and located in the same or different physical locations. For instance, the server device(s) 1102 can be implemented, in whole or in part, on the local device 1118. To illustrate, the local device 1118 may implement the sequencing application 1110 and / or the STR genotyping application 1106. Further, the server device(s) 1102 and / or the local device 1118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0130] The user client device 1108 is illustrated in FIG. 1A can include various types of client devices. For example, in some implementations, the user client device 1108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 1108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.
[0131] Though FIG. 1A illustrates the components of the computing system 1000 communicating via the network 1112, in certain implementations, the components of computing system 1000 can also communicate directly with each other, bypassing the network 1112. For instance, in some implementations, the user client device 1108 communicates directly with the sequencing device 1114. Additionally, in some implementations, the user client device 1108 communicates directly with the STR genotyping application 1106 and / or the server device(s) 1102. In some implementations, the user client device 1108 communicates directly with the local device 1118. Moreover, the STR genotyping application 1106 can access one or more databases housed on or accessed by the server device(s) 1102 or elsewhere in the computing system 1000.
[0132] FIG. IB is a block diagram of an exemplary server device 1102 that may be used in connection with the computing system 1000 of FIG. 1A. The server device 1102 may be configured to determine a STR genotype for a nucleic acid sample. The general architecture of the server device 1102 depicted in FIG. IB includes an arrangement of computer hardware and software components. The server device 1102 may include many more (or fewer) elements than those shown in FIG. IB. It is not necessary', however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 1102 includes a processing unit 110, a network interface 120, a computer readable medium drive 130, an input / output device interface 140, a display 150, and an input device 160, all of which may communicate with one another by way of a communication bus.The network interface 120 may provide connectivity to one or more networks or computing systems. The processing unit 110 may thus receive information and instructions from other computing systems or services via a network. The processing unit 110 may also communicate to and from memory 170 and further provide output information for an optional display 150 via the input / output device interface 140. The input / output device interface 140 may also accept input from the optional input device 160, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0133] The memory 170 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 110 executes in order to implement one or more embodiments. The memory 170 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 170 may store an operating system 172 that provides computer program instructions for use by the processing unit 110 in the general administration and operation of the server device 1102, The memory 170 may store a reference sequence 173, such as for use by the sequencing application 1110. The memory 170 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0134] For example, in one embodiment, the memory 170 includes a sequencing application 1110, which may include a STR genotyping application 1106. The STR genotyping application 1106 can perform the methods disclosed herein. In addition, memory 170 may include or communicate with the data store 190 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference sequences.
[0135] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification orannotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.
[0136] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.
[0137] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.
[0138] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third-party service provider.
[0139] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer-based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server systemmay further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.
[0140] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.
[0141] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readablestorage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a sendee provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0142] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a sendee provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.
[0143] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.
[0144] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework ( University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Methods
[0145] FIG. 2 is a flow diagram that schematically illustrates an exemplary method 200 for steps for obtaining sequence reads and flow cell locations from a nucleic acid sample placed on a flow cell. In some embodiments, the nucleic acid sample comprises genomic nucleic acids. In some embodiments, the nucleic acid sample comprises cDNA. In some embodiments, the nucleic acid sample comprises genomic DNA. For example, a DNA sample 201 may be placed on flow cell 205. In some embodiments, the DNA sample 201 comprises DNA fragments that are greater than 500 bp in length, greater than 1 kbp in length, greater than 5 kbp in length, greater than 10 kbp in length, greater than 100 kbp in length, greater than 200 kbp in length, greater than 250 kbp in length, greater than 300 kbp in length, greater than 500 kbp in length, or more. In some embodiments the DNA fragments are high molecular weight DNA of about 250 kbp in length or more.Obtaining Sequence Reads
[0146] In some embodiments, the method 200 includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. The method 200 may proceed to block 210, wherein sequence reads are obtained by SBS or other methods from each of the clusters generated on the flow cell.
[0147] Sequence reads can be obtained from processing sample nucleic acids through a sequencing instrument (e.g., FIG. 2). For example, in some embodiments, the flow cell comprises transposome complexes bound to the flow cell. The transposome complexes include a transposase and a first polynucleotide comprising an end sequence and a first tag. A library preparation of the genomic sequences is then contacted with the flow cell and transposomes in order to contact the transposome complexes with the target genomic DNA sample under conditions to cause the transposome to fragment the genomic DN A sample. Because the transposomes are bound to the flow cell, following cleavage of the genomic DNA sample, the resulting fragments become bound to the flow cell. The process may then include amplifying the fragmented genomic DNA to form a plurality of nucleic acid clusters on a flow cell. A sequencing by synthesis process may then be started to sequence the nucleic acids in each cluster on the flow cell to generate sequence reads. In some embodiments, the sequencereads comprise paired end sequence reads where each nucleic acid is sequenced by two primers bound in opposite directions to one another on the bound fragment.
[0148] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). According to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in close physical proximity as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In some embodiments, the library preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment associated with the systems. In some library preparations with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell proximity information is retained. This flow cell proximity information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.
[0149] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. In some embodiments, the complexed strand and transposome is in solution, and is then brought toward a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.
[0150] While embodiments involving transposome complexes have been described above, a variety of sequencing library preparation methods and sequencing techniques may be used to capture flow cell and genomic proximity information for sequence reads. For example, in some embodiments, the methods herein include rolling circle amplification or DNA nanoball sequencing. For example, in some embodiments, any library preparation method which facilitates the recording and storing of flow cell proximity and genomic proximity information. For example, the methods can include library preparation by a multitude of different methods to prepare the nucleic acid fragments on the flow cell.Obtaining Geographic Location Information
[0151] The method 200 may proceed to block 220, wherein geographic location information for each of the sequence reads is obtained. Geographic location information can include locations of sequence read clusters on the flow cell 205, where each cluster corresponds to a nucleic acid fragment that is sequenced at block 210. For example, flow cell locations of the sequence read clusters on the flow cell can be obtained. For example, in some embodiments, the geographic location information comprises coordinates in a cartesian coordinate system. For example, the dimensions of the flow cell may be mapped with a cartesian coordinate system, e.g., with x and y dimensions, and locations of clusters on a flow cell may be assigned to a coordinate using this system. The coordinates can be used to determine the flow cell proximity between two or more clusters on the flow cell.
[0152] For example, once the fragments have been bound to substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. While block 220 is described as taking place after block 210 in the example of method 200, the location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster.Determining STR Genotype Using Flow Cell Locations
[0153] The method 200 may proceed to block 230, wherein an STR genotype is determined using flow cell proximity. This may be accomplished as described below with reference to FIG. 3A and FIG. 3B.
[0154] In alternative embodiments, the method 200 proceeds to a step of re¬ mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence, as described below with reference to FIG. 4.Methods of determining a short tandem repeat ( STR) genotype for a nucleic acid sample
[0155] FIG. 3A is a flow diagram that schematically illustrates an exemplary method 300 for determining a short tandem repeat (STR) genotype for a nucleic acid sample. In some embodiments, the nucleic acid sample comprises two haplotypes. In some embodiments, the STR genotype comprises an estimated STR region length for each haplotype. The method 300 can begin from a start block.Obtaining flow cell data
[0156] The method 300 can proceed to block 310, wherein flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions, is obtained.
[0157] This may be accomplished using the methods described with respect for FIG. 2. For example, a nucleic acid sample can be applied to a flow cell. For example, in some embodiments, the flow cell comprises transposome complexes bound to the flow cell. Because the transposomes are bound to the flow cell, following cleavage of the genomic DNA sample, the resulting fragments become bound to the flow cell. The process may then include amplifying the fragmented genomic DNA to form a plurality of nucleic acid clusters on a flow cell. A sequencing by synthesis process may then be started to sequence the nucleic acids in each cluster on the flow cell to generate sequence reads. Sequence reads can thus be obtained from sample nucleic acids. Furthermore, the location on the flow cell of the clusters of the nucleic acid segments can be obtained. An electronic system can receive the flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions.
[0158] In some embodiments, the nucleic acid sample comprises STR regions and flanking regions. In some embodiments, an STR region comprises a region of a genome thatcomprises a repeated STR motif. In some embodiments, a flanking region comprises a region of a genome that is adjacent to an STR region, upstream or downstream.Identify STR sequence reads
[0159] The method 300 may proceed to block 320, wherein STR sequence reads are identified by determining sequence reads comprising a repeated STR motif.
[0160] In some embodiments, the repeated STR motif comprises a repeated nucleotide sequence of 1-20 bp. The repeated STR motif can be, for example, a dinucleotide repeat, a trinucleotide repeat, or a tetranucleotide repeat,
[0161] In some embodiments, at least a portion of the STR sequence reads have a nucleotide sequence that is entirely a repeated STR motif. These sequence reads may be referred to as in-repeat reads (IRRs).
[0162] STR sequence reads can also include sequence reads which map partially to an STR region and partially to a flanking region. Sequence reads which span an STR repeat sequence and a flanking region may be referred to as spanning reads.
[0163] Sequence reads comprising a repeated STR motif can be determined using any method. Exemplary approaches for this step are described further below.
[0164] In some embodiments, identifying STR sequence reads comprises comparing sequence reads to STR motif sequences and identifying sequence reads which have an alignment score with a confidence above a predetermined threshold. For example, the methods and systems can query the sequence reads to determine if they match (e.g., align with a confidence score above a predetermined threshold) an STR motif sequence.
[0165] In some embodiments, identifying STR sequence reads comprises identifying sequence reads which map to one or more decoy contiguous sequences with an alignment score above a predetermined threshold. In some embodiments, the one or more decoy contiguous sequences comprise a repeated STR motif. In some embodiments, the methods and systems make use of a plurality of decoy contiguous sequences, which each comprise a different STR motif sequence. For example, in some embodiments, the plurality of decoy contiguous sequences includes a decoy contiguous sequence for each possible 1-20 bp STR motif.
[0166] In previous sequencing methods, decoy contiguous sequences may be used to identify and discard sequence reads which complicate the process of determining a nucleic acid sequence. According to the present disclosure, decoy contiguous sequences may advantageously be used to identify and retain difficult-to-map sequences such as STR sequence reads, as further described below in the method 300. This approach is also described further below with reference to FIG. 4 and the method 400.
[0167] In some embodiments, the methods and systems can store the set of identified STR sequence reads in an electronic file, for use in later steps of the method 300.Aligning flanking region sequence reads to a reference sequence
[0168] The method 300 may proceed to block 330, wherein flanking region sequence reads are aligned to a reference sequence to obtain a genomic location of the STR sequence reads. In some embodiments, the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region. Alignment may proceed by any method, for example, by a local alignment method.
[0169] Thus, the methods and systems can identify flanking region sequence reads which align to the reference sequence in a genomic region adjacent to a known STR region. The size and location of the flanking region in the reference sequence may be predetermined based on prior knowledge about the location of STR regions.
[0170] In some embodiments, the flanking regions comprise regions of the reference genome extending from a boundary of a STR region to up about 350 kbp from the boundary of the STR region. In some embodiments, the flanking regions comprise regions of the reference genome extending from a boundary of a STR region to up about 10 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 250 kbp, about 300 kbp, about 350 kbp, about 400 kbp, about 450 kbp, about 500 kbp, or more from the boundary of the STR region. For example, some embodiments using on-flow cell tagmentation may advantageously retain connectivity information between sequence reads that are up to about 350 kbp away from each other, meaning that flanking regions can include up to 350 kbp from the boundary of an STR region, increasing the number of flanking region sequence reads which can be leveraged to re¬ map STR sequence reads later in the method 300 as described below.
[0171] In some embodiments, the methods and systems can store a file comprising flanking region sequence reads, such as for use in later steps of the method 300.
[0172] While the block 330 has been shown in the flow chart of FIG. 3A as taking place after the block 320, for avoidance of doubt, it should be understood that block 320 and block 330 may take place at a different time, block 320 and block 330 may take place at the same time (in parallel), or block 330 may take place before block 320. For example, in some embodiments, the methods and system align all sequence reads against a reference sequence and one or more decoy contiguous sequences, and then identify STR sequence reads and flanking region sequence reads.Determining the proximity of flow cell clusters
[0173] The method 300 can proceed to block 340, wherein the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads is determined.
[0174] For example, the methods and systems can use the locations on the flow cell of the clusters of the nucleic acids corresponding to each STR sequence read and each flanking region sequence read to determine flow cell proximities. In some embodiments, the flow cell proximity is the relative displacement between clusters on the flow cell. For example, in embodiments where the flow cell location is stored as an x,y coordinate on the flow cell, the coordinates can be used to determine the flow cell proximity between two or more clusters on the flow cell.
[0175] In some embodiments, the methods and systems identify STR sequence reads that are from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read. In some embodiments, the methods and systems identify S TR sequence reads that are from clusters that have a relative displacement within a predetermined threshold, to at least one cluster corresponding to a flanking region sequence read. For example, the predetermined threshold may comprise a predetermined threshold for relative distance on the flow cell between clusters. In some embodiments, the methods and systems store an electronic file comprising the set of STR sequence reads that are from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read.
[0176] In some embodiments, the methods and systems store an electronic file comprising flow cell proximity information between STR sequence reads and flanking sequence reads.Re-mapping STR sequence reads based on flow cell proximity
[0177] The method 300 can proceed to block 350, wherein STR sequence reads are re-mapped to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read.
[0178] For example, if a STR sequence read comes from a cluster that is proximate on the flow cell to a cluster corresponding to a flanking region sequence read, and the flanking region sequence read is mapped to the flanking region of a particular STR, the STR sequence read is re-mapped to that particular STR. Because clusters that are proximate on the flow cell are more likely to have derived from the same original long nucleic acid from the nucleic acid sample, the sequence reads from those clusters can be mapped near one another. Thus, flow cell proximity can be used to determine the mapping location for STR sequence reads with accuracies not easily achievable by conventional methods.
[0179] By mapping the sequenced fragments to the reference sequence using the flow cell proximity information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the flow cell proximity of each cluster into account during the mapping process. Therefore, flow cell proximity that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the read quality of previously identified multi¬ mapped reads. In the past, identified multi-mapped reads may have been discarded. The ability to map or re-map sequence reads based on flow cell proximity to other sequence reads mapped to a reference sequence, can improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling and STR genotyping. For example, processing DNA samples suitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments, as further described herein.Assigning sequence reads to a haplotype
[0180] The method 300 may proceed to block 360, wherein sequence reads are assigned to a haplotype. This process is also referred to as phasing sequence reads.
[0181] In some embodiments, the methods and systems can initially phase a subset of the sequence reads based on sequence information. For example, in the flanking regions, sequence reads which include a first allele at a particular locus are assigned to a first haplotype, and sequence reads which include a second allele at the locus are assigned to a second haplotype.
[0182] In some embodiments, the method includes assigning sequence reads to a haplotype based on flow cell proximity information. In some embodiments, flow cell proximity information is used to extend phasing information to other sequence reads which cannot be phased because there is no difference in sequence between haplotypes in that region. For example, sequence reads which include an STR motif, and particularly in-repeat reads which are entirely STR motif sequences, may not be able to be phased based on sequence information alone, because the STR motif sequence is the same between haplotypes. In some embodiments, flow cell proximity to a phased sequence read is used to phase the in-repeat reads.
[0183] In some embodiments, the method comprises determining probabilities that sequence reads mapped to an STR region belong to each of two parental haplotypes based on flow cell proximity to at least one sequence read mapped to a flanking region and assigned to a haplotype. For example, the methods and systems can analyze flow cell proximities with sequence reads that are assigned to a haplotype, for example, flanking region sequence reads that are assigned to a haplotype, and assign unphased sequence reads to a haplotype based on linking information. In some embodiments, the linking information may include spatial metrics, for example, X- and Y-axis displacement of the unphased sequence reads on the flow cell. In some embodiments, the linking information may also include the genomic distance between two reads. Linking information may be used to calculate a Phred quality score to determine if two reads can be considered “linked.” In some embodiments, two reads may be considered “linked” if the Phred score is greater than 10.
[0184] As further described herein, in some embodiments, assigning sequence reads to a haplotype using flow cell proximity can advantageously allow the methods and systems to determine an estimated STR region length for each haplotype. Thus, in someembodiments, downstream analyses to estimate STR region length are performed separately for each haplotype.Applying correction factor based on linking rate
[0185] The method 300 may optionally proceed to block 370, wherein a correction factor is applied based on a linking rate. This step is further described below with respect to FIG. 3B.Determining an STR genotype for the nucleic acid sample
[0186] The method 300 may proceed to block 380, wherein a STR genotype is determined for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
[0187] In some embodiments, the STR genotype is determined based on sequence reads which map entirely within an STR region, e.g., in-repeat reads, and sequence reads which map partially to an STR region and partially to a flanking region, e.g., spanning reads.
[0188] In some embodiments, the STR genotype comprises an STR region length. In some embodiments, the methods and systems estimate a STR region length based on sequence reads which map entirely within an STR region, and sequence reads which map partially to an STR region and partially to a flanking region. In some embodiments, the sequence reads are classified into separate sets as a flanking sequence read, a spanning read, or an in-repeat read. In some embodiments, the methods and systems analyze one or more of the sets of sequence reads to estimate an STR region length.
[0189] In some embodiments, the S TR region length is estimated based on the estimated total number of in-repeat reads after application of the linking rate correction factor, as described further below with respect to FIG. 3B. For avoidance of doubt, the STR region length can further be based on other STR sequence reads, including spanning reads. Thus, in some embodiments, the STR genotype is based on the estimated total number of in-repeat reads and based on spanning reads.
[0190] As further described herein, in embodiments that include assigning sequence reads to a haplotype, the methods and systems can determine an estimated STR region length for each haplotype. For example, the methods and systems can analyze sequencereads assigned to a first haplotype Hl to determine an estimated STR region length for Hl, and separately analyze sequence reads assigned to a second haplotype H2 to determine an estimated STR region length for H2.Linking Rate Correction Factor
[0191] FIG. 3B is a flowchart that illustrates the process taking place within block 370, wherein a correction factor based on linking rate is applied.
[0192] The process 370 can begin from a start block. The process 370 can proceed to block 3710, wherein a linking rate is determined. In some embodiments, the linking rate is determined by determining a proportion of the sequence reads which have a link to another sequence read across the nucleic acid sample, wherein the link is based on flow cell proximity and genomic proximity.
[0193] Embodiments of the present disclosure relate to methods and systems which use “links” or “linking information” between sequence reads. The “link” or “linking information” as discussed herein refers, in some embodiments, to the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long nucleic acids, such as genomic DNA, from a sample are sheared to create shorter fragments which can be sequenced in a single read. The shearing process can create these shorter fragments which land on the flow cell and the flow cell proximity of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments which come from the same nucleic acid molecule have been found to bind closer together on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are in close proximity and also close together on the genome, the clusters are more likely to have come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability thattwo reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.
[0194] In some embodiments, linking information is determined by analyzing, for example, statistically analyzing with a model, the genomic distance between two reads and flow cell proximity between the two clusters on the flow cell. In some embodiments, the methods and systems determine whether the genomic distance and / or flow cell proximity is below a threshold. In some embodiments, the methods and systems determine the presence or absence of a link between the two sequence reads. In some embodiments, the methods and systems determine a linking quality score between the two sequence reads. In some embodiments, the methods and systems analyze genomic distance and flow cell proximity for a plurality of pairs of two sequence reads (for example, each possible pair) of two sequence reads in a dataset. Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U. S. Provisional Patent Application Nos. 63 / 600460, 63 / 614066, 63 / 800,049, and 63 / 800,262, the disclosure of each of which is incorporated herein by reference in its entirety.
[0195] In some embodiments, the methods and systems thus determine a linking rate by determining the proportion of sequence reads that have a link to another sequence read based on flow cell proximity and genomic proximity. In some embodiments, the linking rate is determined across the data set of the sequence reads taken from the nucleic acid sample, including regions beyond the S TR region and flanking region.
[0196] The process 370 can proceed to block 3720, wherein the methods and systems count in-repeat reads which were re-mapped to the STR region based on proximity with at least one flanking region sequence read. For example, the re-mapped sequence reads can come from the output of block 350 of the method 300.
[0197] The process 370 can proceed to block 3730, wherein the methods and systems estimate a total number of in-repeat reads based on the linking rate and based on the count of in-repeat reads which were re-mapped to the STR region based on proximity to at least one flanking region sequence read. In some embodiments, the methods and systems use the output of block 3710, the linking rate over the entire nucleic acid sample, and the output ofblock 3720, the count of re-mapped in- repeat reads, to estimate a total number of in-repeat sequence reads. Thus, in some embodiments, the methods and systems use the linking rate as a scaling factor to estimate the total number of in-repeat reads that could be recovered if there were a 100% linking rate. In some embodiments, the methods and systems divide the count of re-mapped in-repeat reads by the linking rate to estimate the total number of in-repeat reads.
[0198] The process 370 can then end, and the method 300 can proceed to block 380 as further described above.Methods of re- mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence
[0199] FIG. 4 is a flow diagram that schematically illustrates an exemplary method 400 for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence. The method 400 can begin at a start block.Obtaining flow cell data
[0200] The method 400 can proceed to block 410, wherein the methods and systems obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments. This can be accomplished using any of the methods described at block 310 of the method 300.
[0201] For example, nucleic acid sample can be applied to a flow cell. For example, in some embodiments, the flow cell comprises transposome complexes bound to the flow cell. Because the transposomes are bound to the flow cell, following cleavage of the genomic DNA sample, the resulting fragments become bound to the flow cell. The process may then include amplifying the fragmented genomic DNA to form a plurality of nucleic acid clusters on a flow cell. A sequencing by synthesis process may then be started to sequence the nucleic acids in each cluster on the flow cell to generate sequence reads. Sequence reads can thus be obtained from sample nucleic acids. Furthermore, the location on the flow cell of the clusters of the nucleic acid segments can be obtained. An electronic system can receive the flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from thenucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions.Aligning sequence reads to the reference sequence and to one or more decoy contiguous sequences
[0202] The method 400 can proceed to block 420, wherein sequence reads are aligned to the reference sequence and to one or more decoy contiguous sequences. Alignment may proceed by any method, for example, by a local alignment method.
[0203] In some embodiments, the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest. In some embodiments, the reference sequence comprises a region of interest and two flanking regions which are adjacent to the region of interest on either side (upstream and downstream).
[0204] In some embodiments, the reference sequence comprises a plurality of regions of interest, and a plurality of flanking regions adjacent to at least one regions of interest of the plurality of regions of interest.
[0205] In some embodiments, the flanking region comprises a region of the reference genome extending from a boundary of a region of interest to up to about 350 kbp from the boundary of the region of interest. In some embodiments, the flanking region comprises a region of the reference genome extending from a boundary of a region of interest to up about 10 kbp, about 50 kbp, about 100 kbp, about 200 kbp, about 250 kbp, about 300 kbp, about 350 kbp, about 400 kbp, about 450 kbp, about 500 kbp, or more from the boundary of the region of interest. For example, some embodiments using on-flow cell tagmentation may advantageously retain connectivity information between sequence reads that are up to about 350 kbp away from each other, meaning that a flanking region can, in some embodiments, include up to 350 kbp from the boundary of a region of interest, increasing the number of flanking region sequence reads which can be leveraged to re-map sequence reads later in the method 400 as described below.
[0206] In some embodiments, the methods and systems can store a file comprising flanking region sequence reads, such as for use in later steps of the method 400.
[0207] In some embodiments, the region of interest comprises a STR region, and wherein the one or more decoy contiguous sequences comprise a repeated STR motif. In someembodiments, the repeated STR motif comprises a repeated nucleotide sequence of 1-20 bp. In some embodiments, the one or more decoy contiguous sequences comprise a nucleotide sequence that is entirely a repeated STR motif
[0208] In some embodiments, the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest. In some embodiments, a plurality of decoy contiguous sequences are used. For example, in some embodiments, the method comprises aligning the sequence reads to the reference sequence and a plurality of decoy contiguous sequences, wherein the plurality of decoy contiguous sequences comprises a plurality of STR sequence motifs. In some embodiments, the method comprises using decoy sequence reads to enable the alignment of sequences which previously were unable to be aligned to a reference sequence. In this embodiment, the system or method may attempt to re-map sequence reads which could not be previously mapped, by aligning those sequence reads to decoy contiguous sequences from a reference sequence. In some embodiments, re-mapping sequence reads using decoy contiguous sequences and flow cell proximity information increases the proportion of sequence reads available for downstream analyses.Identify sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region
[0209] The method 400 can proceed to block 430, wherein sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region are identified.
[0210] For example, the methods and systems can identify sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold by comparing and / or aligning sequence reads to one or more decoy contiguous sequences. In some embodiments, an alignment score takes the form of a Smith-Waterman score or a variation or version of a Smith- Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring. In some embodiments, the threshold is predetermined and sets a minimum alignment score for mapping to one or more decoy contiguous sequences.
[0211] The methods and systems can further identify, from the sequence reads which map to a decoy contiguous sequence, sequence reads which are proximate on the flow cell to at least one sequence read which aligns to the flanking region. For example, the methods and systems can use the locations on the flow cell of the clusters of the nucleic acids corresponding to each sequence read mapped to the one or more decoy contiguous sequences with an alignment score above a threshold, and each flanking region sequence read, to determine flow cell proximities. In some embodiments, the flow cell proximity is the relative displacement between clusters on the flow cell. For example, in embodiments where the flow cell location is stored as an x,y coordinate on the flow cell, the coordinates can be used to determine the flow cell proximity between two or more clusters on the flow cell.
[0212] In some embodiments, the methods and systems identify sequence reads mapped to the one or more decoy contiguous sequences with an alignment score above a threshold that are from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read. In some embodiments, the methods and systems identify sequence reads mapped to the one or more decoy contiguous sequences with an alignment score above a threshold, that are from clusters that have a relative displacement within a predetermined threshold, to at least one cluster corresponding to a flanking region sequence read. For example, the predetermined threshold may comprise a predetermined threshold for relative distance on the flow cell between clusters.
[0213] In some embodiments, the methods and systems store an electronic file comprising the set of sequence reads that are from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read, and that map to the one or more decoy contiguous sequences with an alignment score above a threshold. In some embodiments, the methods and systems store an electronic file comprising flow cell proximity information between sequence reads that mapped to the one or more decoy contiguous sequences with an alignment score above a threshold, and flanking sequence reads.Re-map the identified sequence reads to the region of interest on the reference sequence
[0214] The method 400 can proceed to block 440, wherein the identified sequence reads are re-mapped to the region of interest on the reference sequence.
[0215] In some embodiments, the method comprises comparing flow cell proximity with sequence read clusters corresponding to at least one sequence read which aligns to the flanking region, and selecting a new mapping location on the reference sequence for the identified sequence reads based on flow cell proximity.
[0216] For example, the sequence reads identified in step 430 can be re-mapped on the basis of the flanking sequence reads to which the identified sequence reads were proximate.
[0217] In embodiments w’herein the region of interest comprises an STR region, the method can comprise re-mapping the identified sequence reads to the STR region. In embodiments wherein the reference sequence comprises a plurality of regions of interest, and the plurality of regions of interest comprise a plurality of STR regions, the method can comprise re-mapping the identified sequence reads to one of a plurality of STR regions having different STR sequence motifs.
[0218] In some embodiments, the method comprises storing the alignment of sequence reads to the region of interest, such as the updated alignment of sequence reads based on the re-mapping in block 440, in an electronic file.
[0219] The method 400 can end at an end block.Method of Determining STR Genotype Using Decoy Re-Mapping
[0220] It will be appreciated that the methods of re-mapping using decoy contiguous sequences can be used in methods of determining an STR genotype. For example, further disclosed herein are method of determining an SIR genotype for a nucleic acid sample. In some embodiments, the method comprises re-mapping sequence reads taken from a nucleic acid sample to a region of interest of a reference genome according to the methods described herein, wherein the region of interest comprises an STR region; and determining an STR genotype based on the sequence reads re-mapped to the SIR region, for example by any of the steps of the methods described herein.EXAMPLES
[0221] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, w’hich are not in any way intended to limit the scope of thepresent disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1
[0222] In the following example, methods of re-mapping sequence reads to an STR region based on flow cell proximity information were evaluated on three reference cell lines in genes which include expanded STR regions, NA15850 (FXN gene), NA16207 (FXN gene) and NA04648 (DMPK gene). Sample NA15850 has an expansion of the FXN locus (GAA repeat) with a size of 650 / 1030, wherein 650 is the true size and 1030 is the length of FXN expansion in each allele. Sample NA16207 also has an expansion of the FXN locus (GAA repeat) with a size of 280 / 830. Sample NA04648 has an expansion of DMPK (CAG repeat) with a size of 5 / 1008,
[0223] Sequence reads were generated from a nucleic acid (in this case, DNA) sample from each of the reference cell lines. Flow cell locations were recorded. Sequence reads were aligned to a reference genome to map flanking sequence reads. STR sequence reads (including in-repeat reads (IRR) and spanning sequence reads) were identified based on including an STR motif. In-repeat reads, spanning sequence reads, and flanking sequence reads were re-mapped based on flow cell proximity information.
[0224] FIG. 5 shows counts of realigned sequence reads reference cell lines NA158850 FXN gene), NA16207 (FXN gene), and NA04648 (DMPK gene). The number on the top of each plotted bar indicates the estimated size using standard WGS (whole genome sequencing) compared to the methods for determining STR described herein which use proximity information to improve S TR length determinations.
[0225] As shown in FIG. 5, each of the samples showed successful re-mapping of IRRs and improved estimation of IRR numbers and expansion size compared to standard WGS. Performing additional steps such as correction / adjustment based on linking rate and inclusion of phasing information may improve the repeat size estimate.Other Considerations
[0226] Various embodiments of the present disclosure may be a system, a method, and / or a computer program product at any possible technical detail level of integration. Thecomputer program product may include a computer readable storage medium (or mediums) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.
[0227] For example, the functionality described herein may be performed as software instructions are executed by, and / or in response to software instructions being executed by, one or more hardware processors and / or any other suitable computing devices. The software instructions and / or other executable code may be read from a computer readable storage medium (or mediums). Computer readable storage mediums may also be referred to herein as computer readable storage or computer readable storage devices,
[0228] The computer readable storage medium can be a tangible device that can retain and store data and / or instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device (including any volatile and / or non-volatile electronic storage devices), a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a solid state drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (D VD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0229] Computer readable program instructions described herein can be downloaded to respective computmg / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission,routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0230] Computer readable program instructions (as also referred to herein as, for example, “code,” “instructions,” “module,” “application,” “software application,” and / or the like) for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, and procedural programming languages, such as the " C" programming language or similar programming languages. Computer readable program instructions may be callable from other instructions or from itself, and / or may be invoked in response to detected events or interrupts. Computer readable program instructions configured for execution on computing devices may be provided on a computer readable storage medium, and / or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression or decryption prior to execution) that may then be stored on a computer readable storage medium. Such computer readable program instructions may be stored, partially or fully, on a memory device (e.g., a computer readable storage medium) of the executing computing device, for execution by the computing device. The computer readable program instructions may execute entirely on a user's computer (e.g., the executing computing device), partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry’, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable programinstructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0231] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0232] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart(s) and / or block diagram(s) block or blocks.
[0233] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer may load the instructions and / or modules into its dynamic memory and send the instructions over a telephone, cable, or optical line using a modem. A modem local to a server computing system may receive the data on the telephone / cable / optical line and use a converter device including the appropriate circuitry to place the data on a bus. The bus may cany the data to a memory, from which a processor may retrieve and execute the instructions.The instructions received by the memory may optionally be stored on a storage device (e.g., a solid-state drive) either before or after execution by the computer processor.
[0234] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a service, module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. In addition, certain blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states relating thereto can be performed in other sequences that are appropriate.
[0235] It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. For example, any of the processes, methods, algorithms, elements, blocks, applications, or other functionality (or portions of functionality) described in the preceding sections may be embodied in, and / or fully or partially automated via, electronic hardware such application-specific processors (e.g., application-specific integrated circuits (ASICs)), programmable processors (e.g., field programmable gate arrays (FPGAs)), application-specific circuitry, and / or the like (any of which may also combine custom hard-wired logic, logic circuits, ASICs, FPGAs, etc. with custom programming / execution of software instructions to accomplish the techniques).
[0236] Any of the above-mentioned processors, and / or devices incorporating any of the above-mentioned processors, may be referred to herein as, for example, “computers,” “computer devices,” “computing devices,” “hardware computing devices,” “hardware processors,” “processing units,” and / or the like. Computing devices of the above-embodiments may generally (but not necessarily) be controlled and / or coordinated by operating systemsoftware, such as Mac OS, iOS, Android, Chrome OS, Windows OS (e.g., Windows XP, Windows Vista, Windows 7, Windows 8, Windows 10, Windows 11, Windows Server, etc.), Windows CE, Unix, Linux, SunOS, Solaris, Blackberry OS, VxWorks, or other suitable operating systems. In other embodiments, the computing devices may be controlled by a proprietary operating system. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file system, networking, I / O services, and provide a user interface functionality, such as a graphical user interface (“GUI”), among other things,
[0237] Reference throughout the specification to “one example”, “another example”, “an example”, and so forth, means that a particular element (e.g., feature, structure, and / or characteristic) described in connection with the example is included in at least one example described herein, and may or may not be present in other examples. In addition, it is to be understood that the described elements for any example may be combined in any suitable manner in the various examples unless the context clearly dictates otherwise.
[0238] It is to be understood that the ranges provided herein include the stated range and any value or sub-range within the stated range, as if such value or sub-range were explicitly recited. For example, a range from about 2 kbp to about 20 kbp should be interpreted to include not only the explicitly recited limits of from about 2 kbp to about 20 kbp, but also to include individual values, such as about 3.5 kbp, about 8 kbp, about 18.2 kbp, etc., and sub-ranges, such as from about 5 kbp to about 10 kbp, etc. Furthermore, when “about” and / or “substantially” are / is utilized to describe a value, this is meant to encompass minor variations (e.g., up to + / - 10%) from the stated value.
[0239] While several examples have been described in detail, it is to be understood that the disclosed examples may be modified. Therefore, the foregoing description is to be considered non-limiting.
[0240] While certain examples have been described, these examples have been presented by way of example only, and are not intended to limit the scope of the disclosure. Indeed, the novel methods described herein may be embodied in a variety of other forms. Furthermore, various omissions, substitutions and changes in the methods described herein may be made without departing from the spirit of the disclosure. The accompanying claimsand their equivalents are intended to cover such forms or modifications as would fall within the scope and spirit of the disclosure.
[0241] Features, materials, characteristics, or groups described in conjunction with a particular aspect, or example are to be understood to be applicable to any other aspect or example described in this section or elsewhere in this specification unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The protection is not restricted to the details of any foregoing examples. The protection extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.
[0242] Furthermore, certain features that are described in this disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combmation. Moreover, although features may be described above as acting in certain combinations, one or more features from a claimed combination can, in some cases, be excised from the combination, and the combination may be claimed as a sub-combination or variation of a sub-combination.
[0243] Moreover, while operations may be depicted in the drawings or described in the specification in a particular order, such operations need not be performed in the particular order shown or in sequential order, or that all operations be performed, to achieve desirable results. Other operations that are not depicted or described can be incorporated in the example methods and processes. For example, one or more additional operations can be performed before, after, simultaneously, or between any of the described operations. Further, the operations may be rearranged or reordered in other implementations. Those skilled in the art will appreciate that in some examples, the actual steps taken in the processes illustrated and / or disclosed may differ from those shown in the figures. Depending on the example, certain of the steps described above may be removed or others may be added. Furthermore, the featuresand attributes of the specific examples disclosed above may be combined in different ways to form additional examples, all of which fall within the scope of the present disclosure.
[0244] For purposes of this disclosure, certain aspects, advantages, and novel features are described herein. Not necessarily all such advantages may be achieved in accordance with any particular example. Thus, for example, those skilled in the art will recognize that the disclosure may be embodied or carried out in a manner that achieves one advantage or a group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein.
[0245] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0246] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.
[0247] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.
[0248] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.
[0249] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes w’hich come within the meaning and range of equivalency of the claims are to be embraced within their scope,
[0250] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.
[0251] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.
Claims
WHAT IS CLAIMED IS:
1. A method for determining a short tandem repeat (STR) genotype for a nucleic acid sample, the method comprising:obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions;identifying STR sequence reads by determining sequence reads comprising a repeated STR motif;aligning flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region;determining the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads;re-mapping STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determining a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
2. The method of claim 1, wherein the repeated STR motif comprises a repeated nucleotide sequence of 1-20 bp.
3. The method of claim 1, wherein the flanking regions comprise regions of the reference genome extending from a boundaiy of a STR region to up to about 350 kbp from the boundary of the STR region.
4. The method of claim 1, wherein at least a portion of the STR sequence reads have a nucleotide sequence that is entirely a repeated STR motif.
5. The method of claim 1, wherein the STR genotype is determined based on sequence reads which map entirely within an STR region, and sequence reads which map partially to an STR region and partially to a flanking region.
6. The method of claim 1, wherein identifying STR sequence reads comprises comparing sequence reads to STR motif sequences and identifying sequence reads which have an alignment score with a confidence above a predetermined threshold.
7. The method of claim 1, wherein identifying STR sequence reads comprises identifying sequence reads which map to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a repeated STR motif, with an alignment score above a predetermined threshold.
8. The method of claim 1, wherein the method comprises identifying STR sequence reads that are from clusters that are proximate on the flow cell to at least one cluster corresponding to a flanking region sequence read.
9. The method of claim 1, wherein the STR genotype comprises at least one estimated STR region length, and wherein the method comprises:determining a linking rate by determining a proportion of the sequence reads across the nucleic acid sample which have a link to another sequence read based on flow cell proximity and genomic proximity;counting in-repeat reads which were re-mapped to the STR region based on proximity with at least one flanking region sequence read;estimating a total number of in-repeat reads based on the linking rate and on the count of in-repeat reads which were re-mapped to the STR region based on proximity to at least one flanking region sequence read; andestimating a STR region length based on the estimated total number of in-repeat reads.
10. The method of claim 1, wherein the nucleic acid sample comprises two haplotypes, and wherein the STR genotype comprises an estimated STR region length for each haplotype.
11. The method of claim 1, further comprising assigning sequence reads to a haplotype based on flow cell proximity information.
12. The method of claim 11, wherein the method comprises determining probabilities that sequence reads mapped to an STR region belong to each of two parental haplotypes based on flow cell proximity to at least one sequence read mapped to a flanking region.
13. The method of claim 11, wherein the method comprises estimating a STR region length for each haplotype based on STR sequence reads assigned to each haplotype.
14. The method of claim 1, wherein the method comprises estimating a STR region length based on sequence reads which map entirely within an STR region, and sequence reads which map partially to an STR region and partially to a flanking region.
15. The method of claim 1, further comprising storing the alignment of sequence reads to the STR region in an electronic file.
16. A system for determining a short tandem repeat (STR) genotype for a nucleic acid sample, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising:obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions;identifying STR sequence reads by determining sequence reads comprising a repeated STR motif;aligning flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region;determining the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads;re-mapping STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determining a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
17. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to:obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cellof the clusters of the nucleic acids, wherein the nucleic acid sample comprises STR regions and flanking regions;identify STR sequence reads by determining sequence reads comprising a repeated STR motif;align flanking region sequence reads to a reference sequence to obtain a genomic location of the flanking region sequence reads, wherein the flanking region sequence reads comprise sequence reads that align to the reference sequence in a genomic region adjacent to a known STR region;determine the proximity of flow cell clusters corresponding to STR sequence reads and flow cell clusters corresponding to flanking region sequence reads;re-map STR sequence reads to the STR regions of the reference sequence based on flow cell proximity with at least one flanking region sequence read; and determine a STR genotype for the nucleic acid sample based on the STR sequence reads mapped to the STR regions of the reference sequence.
18. A method for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence, the method comprising:obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments;aligning sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest;identifying sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region; andre-mapping the identified sequence reads to the region of interest on the reference sequence.
19. The method of claim 18, wherein the flanking region covers a region of the reference genome extending from a boundary of the region of interest up to about 350 kbp from the boundary of the region of interest.
20. The method of claim 18, wherein the method comprises comparing flow cell proximity with sequence read clusters corresponding to at least one sequence read which aligns to the flanking region, and selecting a new mapping location on the reference sequence for the identified sequence reads based on flow cell proximity.
21. The method of claim 18, wherein the region of interest comprises a STR region, and wherein the one or more decoy contiguous sequences comprise a repeated STR motif.
22. The method of claim 21, wherein the repeated STR motif comprises a repeated nucleotide sequence of 1-20 bp.
23. The method of claim 21, wherein the one or more decoy contiguous sequences comprises a nucleotide sequence that is entirely a repeated STR motif24. The method of claim 21, wherein the method comprises aligning the sequence reads to the reference sequence and a plurality of decoy contiguous sequences, wherein the plurality of decoy contiguous sequences comprises a plurality of STR sequence motifs.
25. The method of claim 24, wherein the method comprises re-mapping the identified sequence reads to one of a plurality of STR regions having different STR sequence motifs.
26. The method of claim 18, further comprising storing the alignment of sequence reads to the region of interest in an electronic file.
27. A method of determining an STR genotype for a nucleic acid sample, the method comprising:re-mapping sequence reads taken from a nucleic acid sample to a region of interest of a reference genome according to the method of claim 18, wherein the region of interest comprises an S TR region; anddetermining an S TR genotype based on the sequence reads re-mapped to the STR region.
28. A system for re-mapping sequence reads taken from a nucleic acid sample to a region of interest on a reference sequence, comprising a memory storing instructions and a processor that, when executing the instructions, is configured to perform a method comprising:obtaining flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments;aligning sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest;identifying sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence reads which aligns to the flanking region; and re-mapping the identified sequence reads to the region of interest on the reference sequence.
29. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to:obtain flow cell data comprising 1) sequence reads from a flow cell comprising clusters of nucleic acids from the nucleic acid sample and 2) locations on the flow cell of the clusters of the nucleic acid segments;align sequence reads to the reference sequence and to one or more decoy contiguous sequences, wherein the one or more decoy contiguous sequences comprise a nucleic acid sequence of interest, wherein the reference sequence comprises a region of interest and a flanking region which is adjacent to the region of interest;identify sequence reads which map to the one or more decoy contiguous sequences with an alignment score above a threshold and which are proximate on the flow cell to at least one sequence read which aligns to the flanking region; andre-map the identified sequence reads to the region of interest on the reference sequence.