Sequence alignment system and method for identifying short motifs in high-error single molecule reads
A multi-stage alignment method optimizes sequence alignment for nanopore sequencers, addressing inefficiencies in identifying short motifs from long, high-error reads, reducing computational time and costs in NIPT assays.
Patent Information
- Application Number
- JP2025181509
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-28
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-10
AI Technical Summary
Existing sequence alignment tools are inefficient in identifying short motifs from long sequences with high error rates, particularly in noninvasive prenatal testing (NIPT) assays, due to computational resource constraints and sensitivity/specificity requirements.
A multi-stage alignment method using Burrows-Wheeler transformation with varying seed lengths and sensitivity parameters, iteratively masking and aligning sequence reads to enhance sensitivity and reduce computational load, optimized for nanopore sequencers.
The method reduces computational time and resource requirements, enabling rapid analysis of nanopore sequencer data within one hour, and significantly cuts hardware costs by up to 50%, while maintaining sensitivity and specificity.
Smart Images

Figure 2026021415000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS none
[0002] Incorporation by Reference All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
[0003] Field FIELD OF THE INVENTION Embodiments of the present invention relate generally to aligning sequences, and more specifically to efficiently identifying short motifs in long sequences with a high error rate. [Background technology]
[0004] background Commercially available sequencing systems generally generate either short reads with low error rates (i.e., Illumina sequencers) or long reads with high error rates (i.e., Pacific Biosciences sequencers). As a result, most sequence alignment tools have been developed and optimized for both of these use cases: (1) identifying short motifs from short reads with low error rates, or (2) identifying long motifs from long reads with high error rates. However, in certain assays, it is desirable to be able to identify short motifs from long sequences with high error rates. Summary of the Invention
[0005] Disclosure Overview The present invention relates generally to sequence alignment, and more particularly to the efficient identification of short motifs in long sequences with a high error rate.
[0006] In some embodiments, a method for aligning sequence reads to a reference sequence is provided, the method can include: aligning a first set of sequence reads from an entire population of sequence reads to a reference sequence by Burrows-Wheeler transformation using a first seed length, where the first seed length is selected based on an error rate of the sequence reads; masking the first set of sequence reads so that the entire population of sequence reads includes a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by Burrows-Wheeler transformation using a second seed length to achieve greater sensitivity of the second alignment step, where the second seed length is smaller or shorter than the first seed length; and determining alignments of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0007] In some embodiments, the method further includes iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a larger seed length, and determining alignment of the sequence reads with the additional sets of sequence reads. nothing.
[0008] In some embodiments, the first seed length is less than 10 bases. In some embodiments, the first seed length is less than 5 bases. In some embodiments, the first seed length is 4 bases.
[0009] In some embodiments, the error rate of the sequence reads is at least 5%. In some embodiments, the error rate of the sequence reads is at least 10%. In some embodiments, the error rate of the sequence reads is at least 15%.
[0010] In some embodiments, the sequence reads are sequenced from multiple concatemers, each concatemer formed from linked oligonucleotide sequences, and the oligonucleotide sequences correspond to multiple loci from a set of chromosomes. In some embodiments, the set of chromosomes includes chromosomes 13, 18, 22, X, and Y. In some embodiments, the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y. In some embodiments, the method further includes calculating the frequency at which each locus is found in the sequence reads.
[0011] In some embodiments, a method for aligning sequence reads to a reference sequence is provided, the method including: aligning a first set of sequence reads from an entire population of sequence reads to the reference sequence by Burrows-Wheeler transform using a first sensitivity parameter set, where the first sensitivity parameter set is selected based on an error rate of the sequence reads; masking the first set of sequence reads such that the entire population of sequence reads includes a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by Burrows-Wheeler transform using a second sensitivity parameter set, where the second sensitivity parameter set provides a higher sensitivity than the first sensitivity parameter set; and determining alignments of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0012] In some embodiments, the method further includes iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a sensitivity parameter set that results in higher sensitivity, and determining the alignment of the sequence reads with the additional sets of sequence reads.
[0013] In some embodiments, the sensitivity parameter is selected from the group consisting of seed generation, chaining and filtering, and thresholding.
[0014] In some embodiments, a computer product includes a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any of the methods described above.
[0015] In some embodiments, a system includes the computer product described above and one or more processors for executing instructions stored on a computer-readable medium.
[0016] In some embodiments, the system comprises means for performing any of the above methods.
[0017] In some embodiments, the system includes one or more processors configured to perform any of the above methods.
[0018] In some embodiments, the system includes modules for performing each of the steps of any of the above methods. [Brief explanation of the drawings]
[0019] The novel features of the invention are set forth with particularity in the following claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0020] [Figure 1] FIG. 1 is a block diagram illustrating one embodiment of a computer system configured to implement one or more aspects of the present invention. [Figure 2]In one embodiment, the sample reads in the NIPT secondary analysis can be concatemers of short NIPT indexes corresponding to unique loci on the chromosome of interest (i.e., 13, 18, 22, X, Y), and these NIPT indexes in the sample reads are aligned to an index database to determine the frequency of the NIPT indexes in the sample reads. [Figure 3] 1 shows a flowchart of an embodiment of a multi-stage alignment process. [Figure 4] It compares key cost / performance metrics between a conventional BWA aligner (blue) and an embodiment of the optimized aligner described herein (orange). DETAILED DESCRIPTION OF THE INVENTION
[0021] Detailed Description Nanopore-based sequencing systems can generate vast amounts of long-read sequencing data, depending on the number of nanopore-based sensors fabricated within the sensing chip. In some embodiments, each sensing chip can have millions of cells, each with a nanopore sensor. One advantage of using nanopore sequencers is their ability to generate long reads. However, the current raw read accuracy of nanopore sequencers is generally lower (i.e., between 80 and 95%) than established short-read technologies (i.e., greater than 99% accuracy).
[0022] To take advantage of the long read capability of nanopore sequencers, relatively short target sequences can be concatenated into longer linear sequences that can be efficiently sequenced by nanopore sequencers. The concatemers can be sequenced in parallel by the nanopore sequencer to generate a consensus sequence with the desired accuracy (i.e., greater than 99%, 99.9%, or 99.99%). To generate a consensus sequence, the sequence fragments must be aligned.
[0023] Although the alignment methods and systems are described in the context of a nanopore sequencer, other sequences generated by other types of sequencing devices may also be aligned using the alignment methods and systems described herein. For example, non-limiting examples of sequence assays suitable for use in the methods disclosed herein include nanopore sequencing (U.S. Patent Application Publication Nos. 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycle sequencing (Sears et al., Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol., 3:39-42 (1992)), matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS), and the like. ; Fu et al., Nature Biotech., 16:381-384 (1998)), sequencing by hybridization (Drmanac et al., Nature Biotech., 16:54-58 (1998), NGS methods including, but not limited to, sequencing by synthesis (e.g., HiSeq™, MiSeq™, or Genome Analyzer, each available from Illumina), sequencing by ligation (e.g., SOLiD™, Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent™, Life Technologies), and SMRT™ sequencing (e.g., Pacific Biosciences).
[0024] Commercially available sequencing technologies include the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, CA), the sequencing-by-synthesis platforms from Illumina / Solexa (San Diego, CA) and Helicos Biosciences (Cambridge, MA), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, CA). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisher Scientific), and nanopore sequencing (Genia Technology, Roche Sequencing Solutions, Santa Clara, CA) and Oxford Nanopore Technologies (Oxford, UK).
[0025] While alignment analysis is not a new topic, certain sequencing applications present challenges. For example, noninvasive prenatal testing (NIPT) assays present an underdeveloped case in the bioinformatics world. Identifying short motifs—e.g., 40 base pairs in one embodiment—NIPT indexes from long and noisy read sequences (in one embodiment, concatemeric nanopore reads with 1,000–5,000 base pairs and 10–20% error) presents new considerations for algorithm design, as it must satisfy both computational time / resource constraints and the assay's strict sensitivity and specificity requirements. Most existing alignment tools and methods are designed for short motifs sequenced with low error rates (e.g., Illumina sequencers) and long motifs with high error rates (e.g., Pacific Bioscience sequencers). In fact, the use cases of short motifs and high error rates for NIPT assays are precisely the weaknesses of these approaches, resulting in disadvantages in terms of sensitivity, computational resources, or both.
[0026] As a specific consideration, genome aligners are designed for cases where only a small number of very long reference sequences (e.g., chromosomes of millions of base pairs) are used as alignment targets, and each read sequence is expected to uniquely align to a single target locus. In contrast, in one embodiment, the NIPT assay probes biological samples with artificial short sequence motifs (NIPT indexes), cleaves them with a restriction enzyme (i.e., Pstl), ligates them to long templates (i.e., greater than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 1000 bases) containing non-genomic short oligos, and sequences the templates as long, high-error reads (i.e., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20%). The NIPT assay can have over 5000 very short reference sequences (40 base pair NIPT index), and the read sequence of each concatemer is a non-overlapping sequence, as shown in Figure 2. The database aligns to multiple targets as a set of concatemeric reads with a maximum error rate of 20%. The goal of the secondary analysis is to identify those indexes in the concatemeric reads, i.e., to find the frequency of each database index in a given patient sample. The index frequencies are then sent to a tertiary analysis module (i.e., "Forte"), which builds statistical models from the chromosome frequencies and establishes screening test results (probabilities), recommendations to clinicians, etc.
[0027] Similar reasons, as well as poor noise tolerance, are applicable to analytical tools covering other classes of alignment problems, which may not be compatible or optimal for the types of alignment problems found in certain types of sequencing-based assays. For example, short motifs and a well-defined ground truth (e.g., a set of 5000+ target NIPT indices without possible allelic variants and / or other mismatches typical of genomic analysis) make this type of problem seemingly suitable for alignment-free approaches, such as using sequence k-mer seeding / hashing. However, the presence of high noise in the reads would require very short k-mer lengths to maintain high sensitivity. For example, internal data showed that changing the k-mer length from 4 base pairs to 5 base pairs can actually result in a loss of sensitivity of 10% or more, depending on the read error rate. Short (shorter) k-mers also create additional computational burden (typically with a complexity of O(n 2 ), where n is the number of k-mers and is inversely proportional to the length of the k-mer.
[0028] The inventors have conceived a novel registration method that utilizes multi-stage secondary analysis, with each stage gradually reducing the amount of data analyzed in the next stage but increasing the comprehensiveness of the search on the remaining data received from the previous stage. In this way, less noisy registrations can be quickly identified from an initial large pool of data in early stages, while very noisy registrations can be similarly quickly identified from smaller pools of data in later stages of the calculation, thus maintaining the target sensitivity while reducing the overall calculation time.
[0029] We found that the alignment method requires a non-trivial computational implementation, as we showed that a naive, non-optimized implementation of the alignment method results in additional auxiliary data / computational costs (negating the benefits of the alignment method) and potential incompatibilities with downstream processing.
[0030] To address all of this, we implemented our alignment method as a highly optimized and concise extension to the industry-standard BWA-MEM, a well-established open-source genome alignment tool available under the Apache license. We had to modify and tune its core algorithms, data structures, and context-dependent optimizations to fit the unique nature of the alignment problem at hand for NIPT assays. We also had to modify its overall infrastructure to accommodate multi-stage analysis functionality. More specifically, the following parts of BWA-MEM were modified or tuned: core algorithms (seed generation, chaining and filtering, thresholding, etc.), data structures (reads, references, alignment information, SAM records, reduced I / O), memory usage and access patterns (I / O buffering, array creation and propagation), and the overall infrastructure and logic of the execution flow (to accommodate multi-stage analysis functionality, we added a new iterative alignment step to BWA, and we added iteration-specific parameter creation, checkpointing, etc.). A flowchart of the modified and optimized workflow of BWA-MEM is shown in Figure 3.
[0031] For example, in some embodiments, the seed parameters are 19 bases or base pairs kme The kmer is reduced from r (i.e., the default in BWA-MEM) to a 4-base or base-pair kmer. The reduced kmer size is based on the error rate. If the error rate of the sequence reads decreases, the seed length can be increased to reduce computation time. The increased seed size of the first pass compared to the seed size of the second pass allows the first pass to quickly identify oligo sequences that match the NIPT index (however, with lower sensitivity), which allows these sequences to be removed from the unaligned sequence pool, which can then be processed using higher sensitivity parameters in subsequent passes, such as smaller seed sizes (seed generation, linking and filtering, thresholding, etc.). The seeding process involves converting the kmer to an integer value (i.e., each base type has a specific value, and the kmer is the sum of the base values). These integer values can be compared much more easily and quickly with the corresponding kmer integer value from the reference NIPT index than by attempting to match the kmer string itself.
[0032] Along with seed generation, linking can be adjusted to improve performance. Linking is the process of finding groups of seeds, called chains, which are overlapping seeds or seeds that are collinear and close to each other. Chains can be identified by various grouping algorithms, such as pigeonholing or greedy clustering algorithms.
[0033] This project was exciting because we developed a novel use for perhaps today's most popular bioinformatics tool. The significance of this project lies in the fact that it now allows nanopore sequencing-based NIPT assay secondary analysis to be completed rapidly (i.e., in less than 1, 2, 3, 4, 5, or 6 hours). It also enables analysis on virtually any type of commodity hardware, in stark contrast to the initial project estimates of very expensive high-end computing hardware.
[0034] The work in this project also forms the basis for similar algorithmic approaches in other areas (i.e., sequencing-based assays) that can utilize long reads with short and noisy motifs but lack highly optimized tools and algorithms, e.g., inference of splice variant events in RNA-seq isoforms, metatranscriptomics, etc.
[0035] One impact of the optimized alignment method is a direct cost reduction in the final product. Initial estimates of the hardware configuration required for secondary analysis of NIPT assays, prior to optimization, predicted a computational workstation with 256–512 GB of RAM and 32–40 CPU cores. The optimized alignment method reduced computation time by a factor of two (reduced CPU cores) and reduced expected RAM memory usage by 98%. While the exact cost-saving impact is difficult to estimate, it is tens of thousands of dollars for each installed configuration, potentially saving up to 50% of computing hardware costs.
[0036] The second impact was a reduction in computational time. By reducing computational time, it was also possible to fit the analysis per run cycle into a one-hour timeframe, so that a secondary run analysis of the previous run cycle could be performed while the instrument was sequencing the next run cycle, effectively producing secondary run analysis results within an hour of completing the sequencing experiment. This important product requirement was not feasible and was at risk with previous, unoptimized approaches.
[0037] The alignment methods described herein are particularly suited to identifying short motifs from long, high error reads. In some embodiments, the short motifs are about 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15 ...5000, 15000, 15000, 15000, 15000, 15000, In some embodiments, high-error reads are raw reads with an error rate of more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20%. In some embodiments, long reads have at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 bases or base pairs. In some embodiments, the algorithm can be applied to shorter high-error reads depending on the output of the sequencer. For example, shorter reads can be less than 1000, 900, 800, 700, 600, or 500 bases.
[0038] The alignment methods and algorithms described herein may be implemented on a computer system. For example, Figure 1 is a block diagram illustrating one embodiment of a computer system 100 configured to implement one or more aspects of the present invention. As shown, computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104 coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. Memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, which in turn is coupled to a switch 116.
[0039] In operation, I / O bridge 107 is configured to receive user input information from input device(s) 108 (e.g., keyboard, mouse, video / image capture device, etc.) and forward the input information to CPU 102 for processing via communication path 106 and memory bridge 105. In some embodiments, the input information is a live feed from a camera / image capture device or video data stored on a digital storage medium on which object detection operations are performed. Switch 116 is configured to provide connections between I / O bridge 107 and other components of computer system 100, such as network adapter 118 and various add-in cards 120 and 121.
[0040] As also shown, I / O bridge 107 is coupled to system disk 114, which may be configured to store content, applications, and data used by CPU 102 and parallel processing subsystem 112. As a general matter, system disk 114 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (compact disc-read only memory), DVD-ROMs (digital versatile disc-ROMs), Blu-ray discs, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components, such as universal serial bus or other port connections, compact disc drives, digital versatile disc drives, film recording devices, etc., may also be connected to I / O bridge 107.
[0041] In various embodiments, memory bridge 105 may be a northbridge chip and I / O bridge 107 may be a southbridge chip. Additionally, communication paths 106 and 113, as well as other communication paths within computer system 100, may be implemented using any technically suitable protocol, including, but not limited to, Accelerated Graphics Port (AGP), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0042] In some embodiments, the parallel processing subsystem 112 is connected to a display device 110, which may be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. The parallel processing subsystem 112 includes a graphics subsystem that supplies pixels. In such embodiments, the parallel processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be integrated across one or more parallel processing units (PPUs) included within the parallel processing subsystem 112. In other embodiments, the parallel processing subsystem 112 incorporates circuitry optimized for general-purpose and / or computational processing. Similarly, such circuitry may be integrated across one or more PPUs included within the parallel processing subsystem 112 configured to perform such general-purpose and / or computational operations. In yet other embodiments, one or more PPUs included within the parallel processing subsystem 112 may be configured to perform graphics processing, general-purpose processing, and computational processing operations. The system memory 104 includes at least one device driver 103 configured to manage the processing operations of one or more PPUs within the parallel processing subsystem 112. The system memory 104 also includes a software application 125 that executes on the CPU 102 and can issue commands to control the operation of the PPUs.
[0043] In various embodiments, parallel processing subsystem 112 may be integrated with one or more other elements of Figure 1 to form a single system. For example, parallel processing subsystem 112 may be integrated with CPU 102 and other connecting circuitry on a single chip to form a system-on-chip (SoC).
[0044] It will be understood that the systems shown herein are illustrative and that variations and modifications are possible. The connection topology, including the number and placement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112, may be modified as needed. For example, in some embodiments, the system memory 104 may connect directly to the CPU 102 without going through the memory bridge 105, with other devices communicating with the system memory 104 through the memory bridge 105 and the CPU 102. In other alternative topologies, the parallel processing subsystem 112 may connect directly to the I / O bridge 107 or to the CPU 102 rather than the memory bridge 105. In still other embodiments, the I / O bridge 107 and the memory bridge 105 may be integrated into a single chip instead of existing as one or more separate devices. Finally, in certain embodiments, one or more components shown in FIG. 1 may not be present. For example, the switch 116 may be eliminated, and the network adapter 118 and add-in cards 120, 121 connect directly to the I / O bridge 107.
[0045] Figure 4 compares key cost / performance metrics between the traditional BWA (blue) and the v0.4 multi-stage aligner (orange) running on the vhmem hardware of the RSS-SC1 cluster to process the San Jose test case 20180503_uzuki_000001_WSK18R04C8_L03. Cost (first five bars) is significantly reduced, while completeness (last four bars) is comparable.
[0046] When a feature or element is referred to herein as being "on" another feature or element, it can be directly on the other feature or element, or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being "directly" on another feature or element, there are no intervening features or elements present. When a feature or element is referred to as being "connected," "attached," or "coupled" to another feature or element, it will be understood that it can also be directly connected, attached, or coupled to the other feature or element, or that intervening features or elements may be present. In contrast, when a feature or element is referred to as being "directly connected," "directly connected," or "directly coupled" to another feature or element, there are no intervening features or elements present. Although described or illustrated with respect to one embodiment, features and elements so described or illustrated may be applicable to other embodiments. Features and elements located "adjacent" to another feature Those skilled in the art will also understand that references to structures or features may have overlapping or underlying portions of adjacent features.
[0047] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. For example, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. As used herein, it is further understood that the terms "comprises" and / or "comprising" specify the presence of stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ."
[0048] Spatially relative terms such as "under," "below," "lower," "over," "upper," etc. may be used herein to describe the relationship of one element or feature to another element or feature shown in the figures for ease of description. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation shown in the figures. For example, if the device in the figures is turned over, an element described as "under" or "beneath" another element or feature would then be "over" the other element or feature. Thus, the exemplary term "under" can encompass both an orientation of above and below. The device may be oriented in other ways (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly. Similarly, terms such as "upwardly," "downwardly," "vertical," "horizontal," etc. are used herein for descriptive purposes only, unless otherwise specified.
[0049] The terms "first" and "second" may be used herein to describe various features / elements (including steps), but these features / elements should not be limited by these terms unless the context dictates otherwise. These terms may be used to distinguish one feature / element from another. Thus, a first feature / element described below could be referred to as a second feature / element, and similarly, a second feature / element described below could be referred to as a first feature / element without departing from the teachings of the present invention.
[0050] Throughout this specification and the claims that follow, unless the context dictates otherwise, the word "comprise," and variations such as "comprises" and "comprising," mean that they can be used in conjunction with methods and articles (e.g., compositions and apparatuses, including methods). For example, the term "comprising" is understood to imply the inclusion of any element or step described therein, but does not include the exclusion of any other element or step.
[0051] As used herein in the specification and claims, including those used in the examples, unless expressly stated otherwise, all numbers can be read as if they begin with the word "about" or "approximately," even if that term is not expressly stated. The phrase "about" or "approximately" is used when describing a size and / or location to indicate that the described value and / or location is within a reasonably expected range of value and / or location. It can be indicated that a value is within a certain range. For example, a numerical value can have a value of + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Numerical values given herein should also be understood to include about or approximately that value, unless the context dictates otherwise. For example, if the value "10" is disclosed, "about 10" is also disclosed. Any numerical range described herein is intended to include all subranges contained therein. It is also understood that when a value is disclosed as being "less than or equal to," "greater than or equal to" and possible ranges between values are also disclosed, as one of ordinary skill in the art will appreciate. For example, if a value "X" is disclosed, "less than or equal to X" as well as "greater than or equal to X" (e.g., X is a numerical value) are also disclosed. It is also understood that throughout this patent application, data is provided in many different formats, and that this data represents endpoints and starting points, and ranges for any combination of the data points. For example, if a specific data point "10" and a specific data point "15" are disclosed, it is understood that greater than, greater than, less than, less than, less than, and equal to 10 and 15 are considered to be disclosed, as well as between 10 and 15. It is also understood that each unit between two specified units is also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
[0052] While various exemplary embodiments have been described above, several modifications can be made to the various embodiments without departing from the scope of the present invention, as set forth in the claims. For example, the order in which the various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments, one or more method steps may be skipped entirely. Optional features of the various apparatus and system embodiments may be included in some embodiments and not in other embodiments. Therefore, the foregoing description has been provided primarily for illustrative purposes and should not be construed as limiting the scope of the present invention, as set forth in the claims.
[0053] The examples and figures included herein illustrate, by way of illustration and not limitation, specific embodiments in which the subject matter may be practiced. As noted above, other embodiments may be utilized and derived therefrom, resulting in structural and logical substitutions and changes being made without departing from the scope of the present disclosure. Such embodiments of the inventive subject matter may be individually or collectively referred to herein by the term "invention," merely for convenience when more than one is actually disclosed, and without any intention to intentionally limit the scope of this patent application to any single invention or inventive concept. Thus, while specific embodiments have been illustrated and described herein, any configuration calculated to achieve the same purpose may be substituted for the specific embodiment shown. The present disclosure is intended to encompass any and all adaptations or modifications of the various embodiments. Combinations of the above-described embodiments, and other embodiments not specifically described herein, will be apparent to those skilled in the art upon reviewing the above description.
Claims
1. 1. A method for aligning sequence reads to a reference sequence, comprising: aligning a first set of sequence reads from an entire population of sequence reads to a reference sequence by Burrows-Wheeler transformation using a first seed length, wherein the first seed length is selected based on an error rate of the sequence reads; masking the first set of sequence reads such that the entire population of sequence reads comprises a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by the Burrows-Wheeler transformation using a second seed length, wherein the second seed length is shorter than the first seed length; and determining an alignment of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads; The method comprising:
2. 2. The method of claim 1, further comprising iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a shorter seed length, and determining alignments of the sequence reads with the additional sets of sequence reads.
3. 2. The method of claim 1, wherein the first seed length is less than 10 bases.
4. 2. The method of claim 1, wherein the first seed length is less than 5 bases.
5. 2. The method of claim 1, wherein the first seed length is four bases.
6. 2. The method of claim 1, wherein the error rate of the sequence reads is at least 5%.
7. 2. The method of claim 1, wherein the error rate of the sequence reads is at least 10%.
8. 2. The method of claim 1, wherein the error rate of the sequence reads is at least 15%.
9. 2. The method of claim 1, wherein the sequence reads are sequenced from multiple concatemers, each concatemer formed from oligonucleotide sequences linked together, and the oligonucleotide sequences correspond to multiple loci from a set of chromosomes.
10. 10. The method of claim 9, wherein the set of chromosomes comprises chromosomes 13, 18, 22, X, and Y.
11. 10. The method of claim 9, wherein the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y.
12. 10. The method of claim 9, further comprising calculating the frequency with which each locus is found in the sequence reads.
13. 1. A method for aligning sequence reads to a reference sequence, comprising: Aligning a first set of sequence reads from an entire population of sequence reads to a reference sequence by Burrows-Wheeler transformation using a first set of sensitivity parameters, wherein the first set of sensitivity parameters is selected based on an error rate of the sequence reads. It is something that masking the first set of sequence reads such that the entire population of sequence reads comprises a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by a Burrows-Wheeler transformation using a second sensitivity parameter set, wherein the second sensitivity parameter set provides a higher sensitivity than the first sensitivity parameter set; and determining an alignment of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads; The method comprising:
14. 14. The method of claim 13, further comprising iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a sensitivity parameter set that results in higher sensitivity, and determining alignment of the sequence reads with the additional sets of sequence reads.
15. The method of claim 13 , wherein the sensitivity parameters are selected from the group consisting of seed generation, chaining and filtering, and thresholding.
16. A computer product comprising a computer readable medium storing a plurality of instructions for controlling a computer system to perform the operations of the method of any one of claims 1 to 15.
17. A computer product according to claim 16; one or more processors for executing instructions stored on the computer-readable medium.
18. A system comprising means for carrying out the method according to any one of claims 1 to 15.
19. A system comprising one or more processors configured to carry out the method of any one of claims 1 to 15.
20. A system comprising modules for respectively carrying out the steps of the method according to any one of claims 1 to 15.