Sequence alignment systems and methods for identifying short motifs in high error single molecule reads
Through multi-stage secondary analysis and BWA-MEM algorithm optimization, the problem of existing tools identifying short motifs in long sequences with high error rate is solved, and efficient and low-cost short motif recognition is achieved, which is suitable for NIPT determination.
Patent Information
- Application Number
- CN202510528608.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-28
- Filing Date
- 2021-05-26
- Publication Date
- 2025-08-08
AI Technical Summary
Existing sequence alignment tools are difficult to effectively identify short motifs in long sequences with high error rates, especially in non-invasive prenatal detection (NIPT) assays. The existing methods have limitations in computing time and resources, and cannot meet stringent measurement sensitivity and specificity requirements.
A multi-stage secondary analysis method is adopted, and the seed length and sensitivity parameters are adjusted, and the Burrows-Wheeler transformation is used for comparison, gradually reducing the data volume and improving search detail. Combined with the optimization of the BWA-MEM algorithm, it is adapted to long reads with high error rate to identify short motifs.
It realizes efficient identification of short motifs in a short time, reduces computing cost and hardware requirements, is suitable for various commercial hardware, reduces computing time and improves sensitivity, and is suitable for NIPT assays based on nanopore sequencing.
Smart Images

Figure CN120442767A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application 202180038644.8, filed on May 26, 2021, “Sequence alignment system and method for identifying short motifs in high-error single-molecule reads.”
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] none.
[0004] Incorporated by reference
[0005] All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. Technical Field
[0006] Embodiments of the present invention relate generally to aligning sequences, and more particularly to efficiently identifying short motifs in long sequences with high error rates. Background Art
[0007] Commercial sequencing systems generally produce short reads with low error rates (i.e., Illumina sequencers) or long reads with high error rates (i.e., Pacific Biosciences sequencers). Consequently, most sequence alignment tools have been developed and optimized for two use cases: (1) identifying short motifs from short reads with low error rates, or (2) identifying long motifs from long reads with high error rates. However, in some assays, it is desirable to be able to identify short motifs from long sequences with high error rates. Summary of the Invention
[0008] The present invention relates generally to aligning sequences, and more particularly to efficiently identifying short motifs in long sequences with high error rates.
[0009] In some embodiments, a method for aligning sequence reads to a reference sequence is provided. The method may include aligning a first set of sequence reads from a whole population of sequence reads to a reference sequence using a Burrows-Wheeler transform using a first seed length, wherein the first seed length is selected based on an error rate of the sequence reads; masking the first set of sequence reads so that the whole population of sequence reads includes a subset of masked sequence reads and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence using a Burrows-Wheeler transform using a second seed length, wherein the second seed length is less than or shorter than the first seed length to achieve higher sensitivity for the second alignment step; and determining an alignment of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0010] In some embodiments, the method further comprises iteratively masking and aligning additional sets of sequence reads with each subsequent set of reads having a greater seed length, and determining an alignment of the sequence read with the additional sets of sequence reads.
[0011] In some embodiments, the first seed is less than 10 bases in length. In some embodiments, the first seed is less than 5 bases in length. In some embodiments, the first seed is 4 bases in length.
[0012] In some embodiments, the error rate of the sequence reads is at least 5%. In some embodiments, the error rate of the sequence reads is at least 10%. In some embodiments, the error rate of the sequence reads is at least 15%.
[0013] In some embodiments, the sequence reads are sequenced from a plurality of concatemers, wherein each concatemer is formed from oligonucleotide sequences that have been linked together, wherein the oligonucleotide sequences correspond to a plurality of loci from a set of chromosomes. In some embodiments, the set of chromosomes comprises chromosomes 13, 18, 22, X, and Y. In some embodiments, the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y. In some embodiments, the method further comprises calculating the frequency with which each locus is found in the sequence reads.
[0014] In some embodiments, a method for aligning sequence reads to a reference sequence is provided. The method may include aligning a first set of sequence reads from a whole group of sequence reads to a reference sequence using a Burrows-Wheeler transform using a first set of sensitivity parameters, wherein the first set of sensitivity parameters is selected based on an error rate of the sequence reads; masking the first set of sequence reads so that the whole group of sequence reads includes a subset of masked sequence reads and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence using a Burrows-Wheeler transform using a second set of sensitivity parameters, wherein the second set of sensitivity parameters results in a higher sensitivity than the first set of sensitivity parameters; and determining an alignment of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0015] In some embodiments, the method further comprises iteratively masking and aligning additional sets of sequence reads with each subsequent set of reads having a set of sensitivity parameters that results in a higher sensitivity, and determining the alignment of the sequence reads with the additional sets of sequence reads.
[0016] In some embodiments, the sensitivity parameter is selected from the group consisting of seed generation, linking and filtering, and thresholding.
[0017] In some embodiments, a computer product includes a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any of the above methods.
[0018] In some embodiments, a system includes the above computer product and one or more processors configured to execute instructions stored on a computer-readable medium.
[0019] In some embodiments, a system includes means for performing any of the above methods.
[0020] In some embodiments, a system includes one or more processors configured to perform any of the above methods.
[0021] In some embodiments, the system includes modules that respectively perform the steps of any of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The novel features of the present invention are particularly set forth in the appended claims. The features and advantages of the present invention may be better understood by reference to the following detailed description taken in conjunction with the accompanying drawings, which set forth illustrative embodiments in which the principles of the present invention are utilized. In the accompanying drawings:
[0023] Figure 1 is a block diagram illustrating one embodiment of a computer system configured to implement one or more aspects of the present invention.
[0024] Figure 2 It is shown that in one embodiment, the sample reads in the NIPT secondary analysis can be concatemers of short NIPT indices corresponding to unique loci on the target chromosomes (i.e., 13, 18, 22, X, Y), and these NIPT indices in the sample reads are aligned with the index database to determine the frequency of the NIPT indices in the sample reads.
[0025] Figure 3 A flow chart illustrating an embodiment of a multi-stage alignment process.
[0026] Figure 4 Comparison of key cost / performance metrics between a conventional BWA aligner (blue) and an embodiment of the optimized aligner described herein (orange). DETAILED DESCRIPTION
[0027] Nanopore-based sequencing systems can generate a large amount of long-read sequencing data, depending on the number of nanopore-based sensors manufactured in the sensor chip. In some embodiments, each sensor chip may have millions of cells, each cell having a nanopore sensor. One advantage of using a nanopore sequencer is the ability to generate long reads. However, the raw read accuracy of current nanopore sequencers is generally lower (i.e., between 80% and 95%) than established short-read technologies (i.e., accuracy greater than 99%).
[0028] In order to take advantage of the long read capability of the nanopore sequencer, relatively short target sequences can be concatenated into longer linear sequences that can be efficiently sequenced in the nanopore sequencer. The concatemers can be sequenced in parallel by the nanopore sequencer to generate a consensus sequence with the desired accuracy (ie, for example, greater than 99%, 99.9% or 99.99%). In order to generate a consensus sequence, the sequence fragments need to be aligned.
[0029] Although the alignment methods and systems are described in the context of nanopore sequencers, other sequences generated by other types of sequencing equipment can also be aligned using the alignment methods and systems described herein. For example, non-limiting examples of sequence determination suitable for use with the methods disclosed herein include nanopore sequencing (U.S. Patent Publication Nos. 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycle sequencing (Sears et al., Biotechniques 13:626-633 (1992)), solid phase sequencing (Zimmerman et al., Methods Mol. Cell Biol 3:39-42 (1992)), sequencing with mass spectrometry such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nature Biotechnology 3:41-43), and sequencing with ionization time-of-flight mass spectrometry (TFSMS). Biotech 16:381-384 (1998)), sequencing by hybridization (Drmanac et al., Nature Biotech 16:54-58 (1998)), and NGS methods, including but not limited to sequencing by synthesis (e.g., HiSeq TM , MiSeq TM or Genome Analyzer, both available from Illumina), sequencing by ligation (e.g., SOLiD TM , Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent TM , Life Technologies) and Sequencing (eg, Pacific Biosciences).
[0030] Commercially available sequencing technologies include the sequencing-by-hybridization platform of Affymetrix Inc. (Sunnyvale, California), the sequencing-by-synthesis platforms of Illumina / Solexa (San Diego, California) and Helicos Biosciences (Cambridge, Massachusetts), and the sequencing-by-ligation platform of Applied Biosystems (Foster City, California). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisher Scientific) and nanopore sequencing (Genia Technology, Roche Sequencing Solutions, Santa Clara, California); and Oxford Nanopore Technologies (Oxford, United Kingdom).
[0031] While alignment analysis is not a new topic, certain sequencing applications present challenges. For example, non-invasive prenatal testing (NIPT) assays present a case that is not well developed in the field of bioinformatics. Identifying short motifs, such as a 40 base pair NIPT index in one embodiment, from long and noisy read sequences (in one embodiment, concatenated nanopore reads with 1000-5000 base pairs and 10%-20% errors) is a new consideration for algorithm design because of the need to meet both computational time / resource constraints and to meet stringent assay sensitivity and specificity requirements. Most existing alignment tools and methods are designed for short motifs sequenced with low error rates (e.g., Illumina sequencers) and for long motifs with high error rates (e.g., Pacific Bioscience sequencers): in fact, the use case of short motifs and high error rates for NIPT assays is precisely the weakness of these methods, which suffer in terms of sensitivity, computational resources, or both.
[0032] As a specific consideration, genome aligners are designed for situations where only a few very long reference sequences (e.g., chromosomes of millions of base pairs) are used as alignment targets and each read sequence is expected to uniquely align to a single target locus. In contrast, in one embodiment, a NIPT assay probes a biological sample with artificial short sequence motifs (NIPT indexes), cuts them with a restriction enzyme (i.e., Pst1) and ligates them into long templates (i.e., greater than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 1000 bases) comprising non-genomic short oligonucleotides, and sequences the templates into long, high-error reads (i.e., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20%). NIPT assays can have over 5000 very short reference sequences (NIPT index of 40 base pairs), and each concatemer read sequence is mapped to multiple targets as a set of non-overlapping segments with an error rate of up to 20%, e.g. Figure 2 The goal of the secondary analysis is to identify those indices that are in the concatenated reads, i.e., to find the frequency of each database index in a given patient sample. The index frequencies are then sent to a tertiary analysis module (i.e., "Forte"), which builds a statistical model from the chromosome frequencies and establishes screening test results (probabilities), clinician recommendations, etc.
[0033] Similar reasons, along with low noise resilience, apply to analytical tools covering other classes of alignment problems, rendering them inapplicable or unsuitable for the types of alignment problems found in certain types of sequencing-based assays. For example, short motifs and well-defined true values (e.g., a set of 5000+ target NIPT indices with no possible allelic polymorphisms and / or other inconsistencies typical for genomic analysis) would make such problems seem amenable to alignment-free approaches such as using sequence k-mer seeding / hashing. However, the presence of high noise in the reads will require very short k-mer lengths to maintain high sensitivity. For example, internal data have shown in practice that changing the k-mer length from 4 to 5 base pairs can result in a loss of sensitivity of 10% or more, depending on the read error rate. (Shorter) k-mers also impose an additional computational burden (typically of complexity O(n)). 2 ), where n is the number of k-mers, which is inversely proportional to the k-mer length).
[0034] We conceived a novel alignment approach that utilizes multiple stages of secondary analysis, with each stage progressively reducing the amount of data to be analyzed in the next stage or stages, but increasing the exhaustiveness of the search over the remaining data received from the previous stage or stages. In this way, alignments with low noise levels can be identified rapidly from an initially large data pool in one or more early stages, while very noisy alignments can be identified equally rapidly from a smaller data pool in one or more later computational stages, thereby maintaining target sensitivity while reducing overall computational time.
[0035] We find that alignment methods require a significant computational implementation, as we show that initial, non-optimized implementations of alignment methods result in additional auxiliary data / computational costs (which offset the benefits of the alignment methods) and potential incompatibilities with downstream processing.
[0036] To address all these issues, we implemented our alignment approach as a highly optimized and concise extension to the industry standard BWA-MEM, a well-established open source genome alignment tool available under the Apache license. We had to change and adapt its core algorithms, data structures, and context-dependent optimizations to suit the unique nature of our alignment problem at hand for NIPT assays. We also had to change its overall infrastructure to accommodate our multi-stage analysis capabilities. More specifically, the following parts of BWA-MEM were changed or adapted: the core algorithms (seed generation, linking and filtering, thresholding, etc.), data structures (reads, references, alignment information, SAM records, reduced I / O), memory usage and access patterns (I / O buffering, array creation and propagation), and the overall infrastructure and logic of the execution pipeline (to accommodate the multi-stage analysis capabilities, we added a new iterative alignment process to BWA, added iteration-specific parameter creation, checkpoints, etc.). The flowchart of the modified and optimized workflow of BWA-MEM is shown in Figure 2. Figure 3 shown.
[0037] For example, in some embodiments, the seed parameter is reduced from 19 bases or base pairs kmer (i.e., the BWA-MEM default value) to 4 bases or base pairs kmer. The size of the reduced kmer is based on the error rate. If the error rate of the sequence reads is reduced, the seed length can be increased to reduce the calculation time. Compared to the second pass seed size, the increased first pass seed size allows the first pass to quickly identify oligonucleotide sequences that match the NIPT index (but with lower sensitivity), which allows these sequences to be removed from the unaligned sequence pool and can be processed in the next pass using higher sensitivity parameters, such as smaller seed size (seed generation, linking and filtering, thresholds, etc.). The inoculation process involves converting kmer to integer values (i.e., each basic type has a specific value and kmer is the sum of basic values). Comparing these integer values with the corresponding kmer integer values from the reference NIPT index is much easier and faster than trying to match the kmer string itself.
[0038] As seeds are generated, chaining can be adjusted to improve performance. Chaining is the process of finding groups of seeds, called chains, that are overlapping seeds or seeds that are collinear and close to each other. Chains can be identified using various grouping algorithms, such as pigeonholes or greedy clustering.
[0039] This project is exciting because we have developed a new use for arguably the most popular bioinformatics tool available today. The importance of this project lies in the fact that it enables rapid secondary analysis of nanopore sequencing-based NIPT assays (i.e., less than 1, 2, 3, 4, 5, or 6 hours). It also enables analysis on virtually any type of commodity hardware, in stark contrast to the initial project estimates of very expensive high-end computing hardware.
[0040] The work in this project also lays the foundation for similar algorithmic approaches in other fields (i.e., sequencing-based assays) that can exploit long reads carrying short, noisy motifs but lack highly optimized tools and algorithms, for example, inferring splice variant events in RNA-seq isoforms, metatranscriptomics, etc.
[0041] One of the impacts of the optimized alignment method is a direct reduction in the cost of the final product. Initial estimates prior to optimization of the hardware configuration required for secondary analysis of NIPT assays predicted a computational workstation with 256-512GB of RAM and 32-40 CPU cores. The optimized alignment method reduced computation time by a factor of 2 (reduced CPU cores) and is expected to reduce RAM memory usage by 98%. While it is difficult to estimate the exact cost reduction impact, it is fair to say that each installation, which would cost tens of thousands of dollars, could potentially save up to 50% in computational hardware costs.
[0042] The second impact is a reduction in computational time. This reduction allows analysis of each run to be accommodated within a one-hour timeframe. This allows secondary analysis of the previous run to be performed while the instrument is sequencing the next run, effectively making secondary analysis results available within an hour of sequencing completion. This significant product requirement was previously unfeasible and risky with unoptimized methods.
[0043] The alignment methods described herein are particularly suitable for identifying short motifs from long, high-error reads. In some embodiments, the short motif is less than about 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bases or base pairs. In some embodiments, high-error reads are raw reads with an error rate greater than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20%. In some embodiments, the long reads are at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 bases or base pairs. In some embodiments, the algorithm can be applied to shorter length high error reads, depending on the output of the sequencer. For example, the shorter reads can be less than 1000, 900, 800, 700, 600, or 500 bases.
[0044] The alignment method algorithms described herein can be implemented on a computer system. For example, Figure 1 1 is a block diagram illustrating one embodiment of a computer system 100 configured to implement one or more aspects of the present invention. As shown, computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104, which is coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. Memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, and I / O bridge 107 is in turn coupled to a switch 116.
[0045] In operation, I / O bridge 107 is configured to receive user input information from input device 108 (e.g., keyboard, mouse, video / image capture device, etc.) and forward the input information to CPU 102 for processing via communication path 106 and memory bridge 105. In some embodiments, the input information is a real-time feed from a camera / image capture device or video data stored on a digital storage medium on which object detection operations are performed. Switch 116 is configured to provide connections between I / O bridge 107 and other components of computer system 100, such as a network adapter 118 and various expansion cards 120 and 121.
[0046] As also shown, I / O bridge 107 is coupled to a system disk 114, which can be configured to store content and applications, as well as data for use by CPU 102 and parallel processing subsystem 112. Generally, system disk 114 provides non-volatile storage for applications and data and can include fixed or removable hard drives, flash memory devices, and CD-ROMs (Compact Disc Read Only Memory), DVD-ROMs (Digital Versatile Disc-ROMs), Blu-ray Discs, HD-DVDs (High Definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components such as a universal serial bus or other port connector, an optical drive, a digital versatile disc drive, a video recorder, and the like can also be connected to I / O bridge 107.
[0047] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths within computer system 100, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0048] In some embodiments, parallel processing subsystem 112 includes a graphics subsystem that delivers pixels to display device 110, which can be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, or the like. In such embodiments, parallel processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry can be incorporated across one or more parallel processing units (PPUs) included in parallel processing subsystem 112. In other embodiments, parallel processing subsystem 112 incorporates circuitry optimized for general-purpose and / or computational processing. Similarly, such circuitry can be incorporated across one or more PPUs included in parallel processing subsystem 112, which are configured to perform such general-purpose and / or computational operations. In still other embodiments, the one or more PPUs included in parallel processing subsystem 112 can be configured to perform graphics processing, general-purpose processing, and computational processing operations. System memory 104 includes at least one device driver 103, which is configured to manage processing operations of one or more PPUs within parallel processing subsystem 112. System memory 104 also includes software applications 125 that execute on CPU 102 and can issue commands that control the operation of the PPU.
[0049] In various embodiments, the parallel processing subsystem 112 may be coupled with Figure 1 For example, parallel processing subsystem 112 may be integrated with CPU 102 and other connected circuits on a single chip to form a system on a chip (SoC).
[0050] It should be understood that the system shown herein is illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112, may be modified as desired. For example, in some embodiments, system memory 104 may be connected to CPU 102 directly rather than through memory bridge 105, and other devices will communicate with system memory 104 via memory bridge 105 and CPU 102. In other alternative topologies, parallel processing subsystem 112 may be connected to I / O bridge 107 or directly to CPU 102 rather than being connected to memory bridge 105. In still other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip rather than existing as one or more discrete devices. Finally, in some embodiments, Figure 1 One or more of the components shown in FIG may not exist. For example, the switch 116 may be removed, and the network adapter 118 and expansion cards 120 , 121 may be connected directly to the I / O bridge 107 .
[0051] Figure 4 Comparison of key cost / performance metrics between traditional BWA (blue) and the v0.4 multi-stage aligner (orange) executed on the vhmem hardware of the RSS-SC1 cluster to process the San Jose test case 20180503_uzuki_000001_WSK18R04C8_L03. Cost (first 5 bars) is significantly reduced, while completeness (last 4 bars) is comparable.
[0052] The present invention also relates to the following embodiments:
[0053] 1. A method for aligning sequence reads to a reference sequence, the method comprising:
[0054] aligning a first set of sequence reads from the entire population of sequence reads to a reference sequence using a Burrows-Wheeler transform using a first seed length, wherein the first seed length is selected based on an error rate for the sequence reads;
[0055] masking the first set of sequence reads such that the entire population of sequence reads includes masked sequence reads and a subset of unmasked sequence reads;
[0056] aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence using the Burrow-Wheeler transform using a second seed length, wherein the second seed length is less than the first seed length; and
[0057] An alignment of the sequence reads to the reference sequence is determined based on the first set of sequence reads and the second set of sequence reads.
[0058] 2. The method of embodiment 1, further comprising iteratively masking and aligning additional sets of sequence reads with each subsequent set of reads having a smaller seed length, and
[0059] An alignment of the sequence read to another set of sequence reads is determined.
[0060] 3. The method of embodiment 1, wherein the first seed is less than 10 bases in length.
[0061] 4. The method of embodiment 1, wherein the first seed is less than 5 bases in length.
[0062] 5. The method of embodiment 1, wherein the first seed is 4 bases in length.
[0063] 6. The method of embodiment 1, wherein the error rate of the sequence reads is at least 5%.
[0064] 7. The method of embodiment 1, wherein the error rate of the sequence reads is at least 10%.
[0065] 8. The method of embodiment 1, wherein the error rate of the sequence reads is at least 15%.
[0066] 9. The method of embodiment 1, wherein the sequence reads are sequenced from a plurality of concatemers, wherein each concatemer is formed by oligonucleotide sequences that have been ligated together, wherein
[0067] The oligonucleotide sequences correspond to multiple loci from a set of chromosomes.
[0068] 10. The method of embodiment 9, wherein the set of chromosomes comprises chromosomes 13, 18, 22, X, and Y.
[0069] 11. The method of embodiment 9, wherein the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y.
[0070] 12. The method of embodiment 9, further comprising calculating the frequency with which each locus is found in the sequence reads.
[0071] 13. A method for aligning sequence reads to a reference sequence, the method comprising:
[0072] aligning a first set of sequence reads from the entire population of sequence reads to a reference sequence using a Burrows-Wheeler transform using a first set of sensitivity parameters, wherein the first set of sensitivity parameters is selected based on an error rate of the sequence reads;
[0073] masking the first set of sequence reads such that the entire population of sequence reads includes masked sequence reads and a subset of unmasked sequence reads;
[0074] aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence using the Burrow-Wheeler transform using a second set of sensitivity parameters, wherein the second set of sensitivity parameters results in higher sensitivity than the first set of sensitivity parameters; and
[0075] An alignment of the sequence reads to the reference sequence is determined based on the first set of sequence reads and the second set of sequence reads.
[0076] 14. The method of embodiment 13, further comprising iteratively masking and aligning an additional set of sequence reads with each subsequent set of reads having a set of sensitivity parameters that results in higher sensitivity, and determining the alignment of the sequence reads with the additional set of sequence reads.
[0077] 15. The method of embodiment 13, wherein the sensitivity parameter is selected from the group consisting of seed generation, concatenation and filtering, and thresholding.
[0078] 16. A computer product comprising a computer readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any one of the above methods.
[0079] 17. A system comprising:
[0080] The computer product of embodiment 16; and
[0081] One or more processors for executing instructions stored on the computer-readable medium.
[0082] 18. A system comprising means for performing any of the above methods.
[0083] 19. A system comprising one or more processors configured to perform any of the above methods.
[0084] 20. A system comprising modules that respectively perform the steps of any one of the above methods.
[0085] When a feature or element is referred to as being "on" another feature or element in this article, it can be directly on the other feature or element, or there can be intermediate features and / or elements. On the contrary, when a feature or element is referred to as being "directly on" another feature or element, there can be no intermediate features or elements. It will also be understood that when a feature or element is referred to as being "connected," "attached" or "coupled" to another feature or element, it can be directly connected, attached or coupled to the other feature or element, or there can be intermediate features or elements. On the contrary, when a feature or element is referred to as being "directly connected," "directly attached" or "directly coupled" to another feature or element, there can be no intermediate features or elements. Although described or illustrated for one embodiment, the features and elements described or illustrated in this manner can be applied to other embodiments. Those skilled in the art will also recognize that the structure or feature mentioned as being "adjacent" to another feature can have a portion that overlaps or is located below the adjacent feature.
[0086] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. For example, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be further understood that when the terms "include" and / or "comprise" are used in this specification, they specify the presence of the specified features, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components and / or their groups. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and can be abbreviated as " / ".
[0087] For ease of description, spatially relative terms such as "below", "below", "below", "above", "above", etc. may be used herein to describe the relationship between an element or feature and another element or feature, as shown in the accompanying drawings. It should be understood that, in addition to the orientations depicted in the accompanying drawings, spatially relative terms are also intended to cover different orientations of the device in use or operation. For example, if the device in the accompanying drawings is inverted, the elements described as being "below" or "below" other elements or features will then be oriented to be "above" other elements or features. Therefore, the exemplary term "below" can cover both "above" and "below" orientations. The device can be oriented in other ways (rotated 90 degrees or in other orientations), and the spatially relative descriptors used herein are interpreted accordingly. Similarly, unless specifically noted otherwise, terms such as "upward", "downward", "vertical", "horizontal" are used herein only for the purpose of explanation.
[0088] Although the terms "first" and "second" may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms unless the context indicates otherwise. These terms can be used to distinguish one feature / element from another. Thus, without departing from the teachings of the present invention, the first feature / element discussed below may be referred to as the second feature / element, and similarly, the second feature / element discussed below may be referred to as the first feature / element.
[0089] Throughout this specification and the claims that follow, unless the context requires otherwise, the word "comprise" and variations such as "include" and "comprising" mean that various components can be employed together in methods and articles (e.g., compositions and apparatus including devices and methods). For example, the term "comprising" will be understood to imply the inclusion of any stated elements or steps but not the exclusion of any other elements or steps.
[0090] As used herein in the specification and claims, including as used in the examples, and unless otherwise expressly specified, all numbers can be interpreted as if preceded by the word "about" or "approximately", even if the term does not explicitly appear. When describing amplitude and / or position, the phrase "about" or "approximately" can be used to indicate that the value and / or position described are within the reasonable expected range of value and / or position. For example, a numerical value can have a value of + / -0.1% of a specified value (or range of values), + / -1% of a specified value (or range of values), + / -2% of a specified value (or range of values), + / -5% of a specified value (or range of values), + / -10% of a specified value (or range of values), etc. Unless the context indicates otherwise, any numerical value given herein should also be understood to include about or approximately that value. For example, if the value "10" is disclosed, "about 10" is also disclosed. Any numerical range recited herein is intended to include all subranges contained therein. It should also be understood that when a value is disclosed, "less than or equal to" that value, "greater than or equal to" that value, and possible ranges between values are also disclosed, as appropriately understood by those skilled in the art. For example, if a value "X" is disclosed, "less than or equal to X" and "greater than or equal to X" are also disclosed (e.g., where X is a numerical value). It should also be understood that throughout this application, data is provided in a variety of different formats, and that the data represents endpoints and starting points, as well as ranges for any combination of data points. For example, if a specific data point "10" and a specific data point "15" are disclosed, it should be understood that values greater than, greater than or equal to, less than, less than or equal to, equal to 10 and 15, and between 10 and 15 are considered disclosed. It should also be understood that every unit between two specific units is also disclosed. For example, if 10 and 15 are disclosed, 11, 12, 13, and 14 are also disclosed.
[0091] Although various illustrative embodiments have been described above, any of a variety of changes may be made to the various embodiments without departing from the scope of the invention as described in the claims. For example, in alternative embodiments, the order in which the various method steps described are performed may often be changed, while in other alternative embodiments, one or more method steps may be skipped entirely. In some embodiments, optional features of the various device and system embodiments may be included, while in other embodiments, they may not be included. Accordingly, the foregoing description is provided primarily for illustrative purposes and should not be construed as limiting the scope of the invention as set forth in the claims.
[0092] The examples and illustrations included herein show, by way of illustration and not limitation, specific embodiments in which this theme can be put into practice. As mentioned, other embodiments can be utilized and derived therefrom so that structural and logical replacements and changes can be made without departing from the scope of this disclosure. These embodiments of the subject matter of the present invention can be referred to herein individually or collectively by the term "the present invention", which is merely for convenience, rather than for the purpose of actively limiting the scope of this application to any single inventive concept when more than one inventive concept is actually disclosed. Therefore, although specific embodiments have been shown and described herein, any arrangement calculated for achieving the same purpose can replace the specific embodiments shown. This disclosure is intended to cover any and all modifications or variations of the various embodiments. After reviewing the above description, the combination of the above embodiments and other embodiments not explicitly described herein will be apparent to those skilled in the art.
Claims
1. A method for aligning sequence reads to a reference sequence, the method comprising: aligning a first set of sequence reads from the entire population of sequence reads to a reference sequence using a Burrows-Wheeler transform using a first set of sensitivity parameters, wherein the first set of sensitivity parameters is selected based on an error rate of the sequence reads; masking the first set of sequence reads such that the entire population of sequence reads includes masked sequence reads and a subset of unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence using the Burrow-Wheeler transform using a second set of sensitivity parameters, wherein the second set of sensitivity parameters results in higher sensitivity than the first set of sensitivity parameters; and An alignment of the sequence reads to the reference sequence is determined based on the first set of sequence reads and the second set of sequence reads.
2. The method of claim 1 , further comprising iteratively masking and aligning an additional set of sequence reads with each subsequent set of reads having a set of sensitivity parameters that results in a higher sensitivity, and determining the alignment of the sequence reads with the additional set of sequence reads.
3. The method of claim 1, wherein the sensitivity parameter is selected from the group consisting of seed generation, concatenation and filtering, and thresholding.
4. A computer product comprising a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any one of the above methods.
5. A system comprising: The computer product according to claim 4; and One or more processors for executing instructions stored on the computer-readable medium.
6. A system comprising means for performing any of the above methods.
7. A system comprising one or more processors configured to perform any of the above methods.
8. A system comprising modules respectively performing the steps of any one of the above methods.
Citation Information
Patent Citations
Nanopore Based Molecular Detection and Sequencing
US20130244340A1
DNA sequencing by synthesis using modified nucleotides and nanopore detection
US20130264207A1
Nucleic acid sequencing using tags
US20140134616A1
Nucleic acid sequencing by nanopore detection of tag molecules
US20150119259A1
Methods for creating bilayers for use with nanopore sensors
US20150337366A1