Sequence alignment systems and methods for identifying short motifs in high-error single-molecule reads
A multi-stage alignment method optimizes sequence alignment for nanopore sequencers by varying seed lengths and sensitivity parameters, addressing inefficiencies in identifying short motifs with high error rates in NIPT assays, reducing computational time and hardware needs.
Patent Information
- Application Number
- JP2022572640
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-28
- Filing Date
- 2021-05-26
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2041-05-26
AI Technical Summary
Existing sequence alignment tools are inefficient for identifying short motifs in long sequences with high error rates, particularly in noninvasive prenatal testing (NIPT) assays, due to computational resource constraints and sensitivity/specificity requirements.
A multi-stage alignment method using Burrows-Wheeler transformation with varying seed lengths and sensitivity parameters, iteratively masking and aligning sequence reads to improve sensitivity and reduce computational load, optimized for nanopore sequencers.
The method reduces computational time and hardware requirements, enabling rapid secondary analysis of NIPT assays on commodity hardware, while maintaining sensitivity and specificity, and achieving consensus sequences with high accuracy.
Smart Images

Figure 0007767320000001 
Figure 0007767320000002 
Figure 0007767320000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS none
[0002] Incorporation by Reference All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
[0003] Field FIELD OF THE INVENTION Embodiments of the present invention relate generally to aligning sequences, and more specifically to efficiently identifying short motifs in long sequences with a high error rate. [Background technology]
[0004] background Commercially available sequencing systems generally generate either short reads with low error rates (i.e., Illumina sequencers) or long reads with high error rates (i.e., Pacific Biosciences sequencers). As a result, most sequence alignment tools have been developed and optimized for both of these use cases: (1) identifying short motifs from short reads with low error rates, or (2) identifying long motifs from long reads with high error rates. However, in certain assays, it is desirable to be able to identify short motifs from long sequences with high error rates. Summary of the Invention
[0005] Disclosure Overview The present invention relates generally to sequence alignment, and more particularly to the efficient identification of short motifs in long sequences with a high error rate.
[0006] In some embodiments, a method for aligning sequence reads to a reference sequence is provided, the method can include: aligning a first set of sequence reads from an entire population of sequence reads to a reference sequence by Burrows-Wheeler transformation using a first seed length, where the first seed length is selected based on an error rate of the sequence reads; masking the first set of sequence reads so that the entire population of sequence reads includes a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by Burrows-Wheeler transformation using a second seed length to achieve greater sensitivity of the second alignment step, where the second seed length is smaller or shorter than the first seed length; and determining alignments of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0007] In some embodiments, the method further includes iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a larger seed length, and determining alignment of the sequence reads with the additional sets of sequence reads.
[0008] In some embodiments, the first seed length is less than 10 bases. In some embodiments, the first seed length is less than 5 bases. In some embodiments, the first seed length is 4 bases.
[0009] In some embodiments, the error rate of the sequence reads is at least 5%. In some embodiments, the error rate of the sequence reads is at least 10%. In some embodiments, the error rate of the sequence reads is at least 15%.
[0010] In some embodiments, the sequence reads are sequenced from multiple concatemers, each concatemer formed from linked oligonucleotide sequences, and the oligonucleotide sequences correspond to multiple loci from a set of chromosomes. In some embodiments, the set of chromosomes includes chromosomes 13, 18, 22, X, and Y. In some embodiments, the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y. In some embodiments, the method further includes calculating the frequency at which each locus is found in the sequence reads.
[0011] In some embodiments, a method for aligning sequence reads to a reference sequence is provided, the method including: aligning a first set of sequence reads from an entire population of sequence reads to the reference sequence by Burrows-Wheeler transform using a first sensitivity parameter set, where the first sensitivity parameter set is selected based on an error rate of the sequence reads; masking the first set of sequence reads such that the entire population of sequence reads includes a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by Burrows-Wheeler transform using a second sensitivity parameter set, where the second sensitivity parameter set provides a higher sensitivity than the first sensitivity parameter set; and determining alignments of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads.
[0012] In some embodiments, the method further includes iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a sensitivity parameter set that results in higher sensitivity, and determining the alignment of the sequence reads with the additional sets of sequence reads.
[0013] In some embodiments, the sensitivity parameter is selected from the group consisting of seed generation, chaining and filtering, and thresholding.
[0014] In some embodiments, a computer product includes a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any of the methods described above.
[0015] In some embodiments, a system includes the computer product described above and one or more processors for executing instructions stored on a computer-readable medium.
[0016] In some embodiments, the system comprises means for performing any of the above methods.
[0017] In some embodiments, the system includes one or more processors configured to perform any of the above methods.
[0018] In some embodiments, the system includes modules for performing each of the steps of any of the above methods. [Brief explanation of the drawings]
[0019] The novel features of the invention are set forth with particularity in the following claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0020] [Figure 1] FIG. 1 is a block diagram illustrating one embodiment of a computer system configured to implement one or more aspects of the present invention. [Figure 2]In one embodiment, the sample reads in the NIPT secondary analysis can be concatemers of short NIPT indexes corresponding to unique loci on the chromosome of interest (i.e., 13, 18, 22, X, Y), and these NIPT indexes in the sample reads are aligned to an index database to determine the frequency of the NIPT indexes in the sample reads. [Figure 3] 1 shows a flowchart of an embodiment of a multi-stage alignment process. [Figure 4] It compares key cost / performance metrics between a conventional BWA aligner (blue) and an embodiment of the optimized aligner described herein (orange). DETAILED DESCRIPTION OF THE INVENTION
[0021] Detailed Description Nanopore-based sequencing systems can generate vast amounts of long-read sequencing data, depending on the number of nanopore-based sensors fabricated within the sensing chip. In some embodiments, each sensing chip can have millions of cells, each with a nanopore sensor. One advantage of using nanopore sequencers is their ability to generate long reads. However, the current raw read accuracy of nanopore sequencers is generally lower (i.e., between 80 and 95%) than established short-read technologies (i.e., greater than 99% accuracy).
[0022] To take advantage of the long read capability of nanopore sequencers, relatively short target sequences can be concatenated into longer linear sequences that can be efficiently sequenced by nanopore sequencers. The concatemers can be sequenced in parallel by the nanopore sequencer to generate a consensus sequence with the desired accuracy (i.e., greater than 99%, 99.9%, or 99.99%). To generate a consensus sequence, the sequence fragments must be aligned.
[0023] Although the alignment methods and systems are described in the context of a nanopore sequencer, other sequences generated by other types of sequencing devices may also be aligned using the alignment methods and systems described herein. For example, non-limiting examples of sequence assays suitable for use in the methods disclosed herein include nanopore sequencing (U.S. Patent Application Publication Nos. 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycle sequencing (Sears et al., Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol., 3:39-42 (1992)), matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nature 106:101-102 (1992)), and ELISA. Biotech., 16:381-384 (1998)), sequencing by hybridization (Drmanac et al., Nature Biotech., 16:54-58 (1998)), NGS methods including, but not limited to, sequencing by synthesis (e.g., HiSeq™, MiSeq™, or Genome Analyzer, each available from Illumina), sequencing by ligation (e.g., SOLiD™, Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent™, Life Technologies), and SMRT™ sequencing (e.g., Pacific Biosciences).
[0024] Commercially available sequencing technologies include the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, CA), the sequencing-by-synthesis platforms from Illumina / Solexa (San Diego, CA) and Helicos Biosciences (Cambridge, MA), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, CA). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisher Scientific), and nanopore sequencing (Genia Technology, Roche Sequencing Solutions, Santa Clara, CA) and Oxford Nanopore Technologies (Oxford, UK).
[0025] While alignment analysis is not a new topic, certain sequencing applications present challenges. For example, noninvasive prenatal testing (NIPT) assays present an underdeveloped case in the bioinformatics world. Identifying short motifs—e.g., 40 base pairs in one embodiment—NIPT indexes from long and noisy read sequences (in one embodiment, concatemeric nanopore reads with 1,000–5,000 base pairs and 10–20% error) presents new considerations for algorithm design, as it must satisfy both computational time / resource constraints and the assay's strict sensitivity and specificity requirements. Most existing alignment tools and methods are designed for short motifs sequenced with low error rates (e.g., Illumina sequencers) and long motifs with high error rates (e.g., Pacific Bioscience sequencers). In fact, the use cases of short motifs and high error rates for NIPT assays are precisely the weaknesses of these approaches, resulting in disadvantages in terms of sensitivity, computational resources, or both.
[0026] As a specific consideration, genome aligners are designed for cases where only a small number of very long reference sequences (e.g., chromosomes of millions of base pairs) are used as alignment targets, and each read sequence is expected to uniquely align to a single target locus. In contrast, in one embodiment, the NIPT assay probes biological samples with artificial short sequence motifs (NIPT indexes), cleaves them with a restriction enzyme (i.e., Pstl), ligates them to long templates (i.e., greater than 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 1000 bases) containing non-genomic short oligos, and sequences the templates as long, high-error reads (i.e., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20%). An NIPT assay can have over 5,000 very short reference sequences (40 base pair NIPT indices), and each concatemeric read sequence aligns to multiple targets as a set of non-overlapping segments with a maximum error rate of 20%, as shown in Figure 2. The goal of secondary analysis is to identify those indices in the concatemeric reads, i.e., to find the frequency of each database index in a given patient sample. The index frequencies are then sent to a tertiary analysis module (i.e., "Forte"), which builds statistical models from the chromosome frequencies and establishes screening test results (probabilities), recommendations to clinicians, etc.
[0027] Similar reasons, as well as poor noise tolerance, are applicable to analytical tools covering other classes of alignment problems, which may not be compatible or optimal for the types of alignment problems found in certain types of sequencing-based assays. For example, short motifs and a well-defined ground truth (e.g., a set of 5000+ target NIPT indices without possible allelic variants and / or other mismatches typical of genomic analysis) make this type of problem seemingly suitable for alignment-free approaches, such as using sequence k-mer seeding / hashing. However, the presence of high noise in the reads would require very short k-mer lengths to maintain high sensitivity. For example, internal data showed that changing the k-mer length from 4 base pairs to 5 base pairs can actually result in a loss of sensitivity of 10% or more, depending on the read error rate. Short (shorter) k-mers also create additional computational burden (typically with a complexity of O(n 2 ), where n is the number of k-mers and is inversely proportional to the length of the k-mer.
[0028] The inventors have conceived a novel registration method that utilizes multi-stage secondary analysis, with each stage gradually reducing the amount of data analyzed in the next stage but increasing the comprehensiveness of the search on the remaining data received from the previous stage. In this way, less noisy registrations can be quickly identified from an initial large pool of data in early stages, while very noisy registrations can be similarly quickly identified from smaller pools of data in later stages of the calculation, thus maintaining the target sensitivity while reducing the overall calculation time.
[0029] We found that the alignment method requires a non-trivial computational implementation, as we showed that a naive, non-optimized implementation of the alignment method results in additional auxiliary data / computational costs (negating the benefits of the alignment method) and potential incompatibilities with downstream processing.
[0030] To address all of this, we implemented our alignment method as a highly optimized and concise extension to the industry-standard BWA-MEM, a well-established open-source genome alignment tool available under the Apache license. We had to modify and tune its core algorithms, data structures, and context-dependent optimizations to fit the unique nature of the alignment problem at hand for NIPT assays. We also had to modify its overall infrastructure to accommodate multi-stage analysis functionality. More specifically, the following parts of BWA-MEM were modified or tuned: core algorithms (seed generation, chaining and filtering, thresholding, etc.), data structures (reads, references, alignment information, SAM records, reduced I / O), memory usage and access patterns (I / O buffering, array creation and propagation), and the overall infrastructure and logic of the execution flow (to accommodate multi-stage analysis functionality, we added a new iterative alignment step to BWA, and we added iteration-specific parameter creation, checkpointing, etc.). A flowchart of the modified and optimized workflow of BWA-MEM is shown in Figure 3.
[0031] For example, in some embodiments, the seed parameter is reduced from a 19-base or base-pair kmer (i.e., the default for BWA-MEM) to a 4-base or base-pair kmer. The reduced kmer size is based on the error rate. If the error rate of the sequence reads decreases, the seed length can be increased to reduce computation time. The increased seed size for the first pass compared to the seed size for the second pass allows the first pass to quickly identify oligo sequences that match the NIPT index (however, with lower sensitivity), which allows these sequences to be removed from the unaligned sequence pool, which can then be processed using higher sensitivity parameters in subsequent passes, such as smaller seed sizes (seed generation, linking and filtering, thresholding, etc.). The seeding process involves converting kmers to integer values (i.e., each base type has a specific value, and the kmer is the sum of the base values). These integer values can be compared much more easily and quickly with corresponding kmer integer values from a reference NIPT index than by attempting to match the kmer string itself.
[0032] Along with seed generation, linking can be adjusted to improve performance. Linking is the process of finding groups of seeds, called chains, which are overlapping seeds or seeds that are collinear and close to each other. Chains can be identified by various grouping algorithms, such as pigeonholing or greedy clustering algorithms.
[0033] This project was exciting because we developed a novel use for perhaps today's most popular bioinformatics tool. The significance of this project lies in the fact that it now allows nanopore sequencing-based NIPT assay secondary analysis to be completed rapidly (i.e., in less than 1, 2, 3, 4, 5, or 6 hours). It also enables analysis on virtually any type of commodity hardware, in stark contrast to the initial project estimates of very expensive high-end computing hardware.
[0034] The work in this project also forms the basis for similar algorithmic approaches in other areas (i.e., sequencing-based assays) that can utilize long reads with short and noisy motifs but lack highly optimized tools and algorithms, e.g., inference of splice variant events in RNA-seq isoforms, metatranscriptomics, etc.
[0035] One impact of the optimized alignment method is a direct cost reduction in the final product. Initial estimates of the hardware configuration required for secondary analysis of NIPT assays, prior to optimization, predicted a computational workstation with 256–512 GB of RAM and 32–40 CPU cores. The optimized alignment method reduced computation time by a factor of two (reduced CPU cores) and reduced expected RAM memory usage by 98%. While the exact cost-saving impact is difficult to estimate, it is tens of thousands of dollars for each installed configuration, potentially saving up to 50% of computing hardware costs.
[0036] The second impact was a reduction in computational time. By reducing computational time, it was also possible to fit the analysis per run cycle into a one-hour timeframe, so that a secondary run analysis of the previous run cycle could be performed while the instrument was sequencing the next run cycle, effectively producing secondary run analysis results within an hour of completing the sequencing experiment. This important product requirement was not feasible and was at risk with previous, unoptimized approaches.
[0037] The alignment method described herein is particularly suitable for identifying short motifs from long high error readings.In some embodiments, the short motif is less than about 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 bases or base pairs.In some embodiments, high error readings are raw readings that have an error rate of more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20%. In some embodiments, the long read has at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000 bases or base pairs. In some embodiments, the algorithm can be applied to shorter high-error reads depending on the output of the sequencer. For example, the shorter read can be less than 1000, 900, 800, 700, 600, or 500 bases.
[0038] The alignment methods and algorithms described herein may be implemented on a computer system. For example, Figure 1 is a block diagram illustrating one embodiment of a computer system 100 configured to implement one or more aspects of the present invention. As shown, computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104 coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. Memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, which in turn is coupled to a switch 116.
[0039] In operation, I / O bridge 107 is configured to receive user input information from input device(s) 108 (e.g., keyboard, mouse, video / image capture device, etc.) and forward the input information to CPU 102 for processing via communication path 106 and memory bridge 105. In some embodiments, the input information is a live feed from a camera / image capture device or video data stored on a digital storage medium on which object detection operations are performed. Switch 116 is configured to provide connections between I / O bridge 107 and other components of computer system 100, such as network adapter 118 and various add-in cards 120 and 121.
[0040] As also shown, I / O bridge 107 is coupled to system disk 114, which may be configured to store content, applications, and data used by CPU 102 and parallel processing subsystem 112. As a general matter, system disk 114 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (compact disc-read only memory), DVD-ROMs (digital versatile disc-ROMs), Blu-ray discs, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components, such as universal serial bus or other port connections, compact disc drives, digital versatile disc drives, film recording devices, etc., may also be connected to I / O bridge 107.
[0041] In various embodiments, memory bridge 105 may be a northbridge chip and I / O bridge 107 may be a southbridge chip. Additionally, communication paths 106 and 113, as well as other communication paths within computer system 100, may be implemented using any technically suitable protocol, including, but not limited to, Accelerated Graphics Port (AGP), HyperTransport, or any other bus or point-to-point communication protocol known in the art.
[0042] In some embodiments, parallel processing subsystem 112 includes a graphics subsystem that provides pixels to display device 110, which may be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. In such embodiments, parallel processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be integrated across one or more parallel processing units (PPUs) included within parallel processing subsystem 112. In other embodiments, parallel processing subsystem 112 incorporates circuitry optimized for general-purpose and / or computational processing. Similarly, such circuitry may be integrated across one or more PPUs included within parallel processing subsystem 112 configured to perform such general-purpose and / or computational operations. In yet other embodiments, one or more PPUs included within parallel processing subsystem 112 may be configured to perform graphics processing, general-purpose processing, and computational processing operations. System memory 104 includes at least one device driver 103 configured to manage the processing operations of one or more PPUs in parallel processing subsystem 112. System memory 104 also includes software applications 125 that execute on CPU 102 and can issue commands to control the operation of the PPUs.
[0043] In various embodiments, parallel processing subsystem 112 may be integrated with one or more other elements of Figure 1 to form a single system. For example, parallel processing subsystem 112 may be integrated with CPU 102 and other connecting circuitry on a single chip to form a system-on-chip (SoC).
[0044] It will be understood that the systems shown herein are illustrative and that variations and modifications are possible. The connection topology, including the number and placement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112, may be modified as needed. For example, in some embodiments, the system memory 104 may connect directly to the CPU 102 without going through the memory bridge 105, with other devices communicating with the system memory 104 through the memory bridge 105 and the CPU 102. In other alternative topologies, the parallel processing subsystem 112 may connect directly to the I / O bridge 107 or to the CPU 102 rather than the memory bridge 105. In still other embodiments, the I / O bridge 107 and the memory bridge 105 may be integrated into a single chip instead of existing as one or more separate devices. Finally, in certain embodiments, one or more components shown in FIG. 1 may not be present. For example, the switch 116 may be eliminated, and the network adapter 118 and add-in cards 120, 121 connect directly to the I / O bridge 107.
[0045] Figure 4 compares key cost / performance metrics between the traditional BWA (blue) and the v0.4 multi-stage aligner (orange) running on the vhmem hardware of the RSS-SC1 cluster to process the San Jose test case 20180503_uzuki_000001_WSK18R04C8_L03. Cost (first five bars) is significantly reduced, while completeness (last four bars) is comparable.
[0046] When a feature or element is referred to herein as being "on" another feature or element, it can be directly on the other feature or element, or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being "directly" to another feature or element, there are no intervening features or elements present. When a feature or element is referred to as being "connected," "attached," or "coupled" to another feature or element, it will be understood that it can also be directly connected, attached, or coupled to the other feature or element, or that intervening features or elements may be present. In contrast, when a feature or element is referred to as being "directly connected," "directly connected," or "directly coupled" to another feature or element, there are no intervening features or elements present. Although described or illustrated with respect to one embodiment, features and elements so described or illustrated can be applied to other embodiments. It will also be understood by those skilled in the art that a reference to a structure or feature located "adjacent" to another feature can have portions that overlap or underlie the adjacent feature.
[0047] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. For example, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. As used herein, it is further understood that the terms "comprises" and / or "comprising" specify the presence of stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ."
[0048] Spatially relative terms such as "under," "below," "lower," "over," "upper," etc. may be used herein to describe the relationship of one element or feature to another element or feature shown in the figures for ease of description. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation shown in the figures. For example, if the device in the figures is turned over, an element described as "under" or "beneath" another element or feature would then be "over" the other element or feature. Thus, the exemplary term "under" can encompass both an orientation of above and below. The device may be oriented in other ways (e.g., rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly. Similarly, terms such as "upwardly," "downwardly," "vertical," "horizontal," etc. are used herein for descriptive purposes only, unless otherwise specified.
[0049] The terms "first" and "second" may be used herein to describe various features / elements (including steps), but these features / elements should not be limited by these terms unless the context dictates otherwise. These terms may be used to distinguish one feature / element from another. Thus, a first feature / element described below could be referred to as a second feature / element, and similarly, a second feature / element described below could be referred to as a first feature / element without departing from the teachings of the present invention.
[0050] Throughout this specification and the claims that follow, unless the context dictates otherwise, the word "comprise," and variations such as "comprises" and "comprising," mean that they can be used in conjunction with methods and articles (e.g., compositions and apparatuses, including methods). For example, the term "comprising" is understood to imply the inclusion of any element or step described therein, but does not include the exclusion of any other element or step.
[0051] As used herein in the specification and claims, including those used in the examples, unless expressly specified otherwise, all numbers can be read as if preceded by the word "about" or "approximately," even if that term is not explicitly indicated. The phrase "about" or "approximately" can be used when describing a size and / or location to indicate that the described value and / or location is within a reasonably expected range of values and / or locations. For example, a numerical value can have a value of + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Numeric values given herein should also be understood to include about or approximately that value, unless the context dictates otherwise. For example, if the value "10" is disclosed, "about 10" is also disclosed. Any numerical range described herein is intended to include all subranges contained therein. It is also understood that when a value is disclosed as being "less than or equal to," "greater than or equal to" and possible ranges between values are also disclosed, as would be understood by one of ordinary skill in the art. For example, if a value "X" is disclosed, "less than or equal to X" as well as "greater than or equal to X" (e.g., X is a numeric value) are also disclosed. It is also understood that throughout this patent application, data is provided in many different formats, and this data represents endpoints and starting points, and ranges for any combination of the data points. For example, if a specific data point "10" and a specific data point "15" are disclosed, it is understood that greater than, greater than, less than, less than, less than, and equal to 10 and 15 are considered to be disclosed, as well as between 10 and 15. It is also understood that each unit between two specified units is also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.
[0052] While various exemplary embodiments have been described above, several modifications can be made to the various embodiments without departing from the scope of the present invention, as set forth in the claims. For example, the order in which the various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments, one or more method steps may be skipped entirely. Optional features of the various apparatus and system embodiments may be included in some embodiments and not in other embodiments. Therefore, the foregoing description has been provided primarily for illustrative purposes and should not be construed as limiting the scope of the present invention, as set forth in the claims.
[0053] The examples and figures included herein illustrate, by way of illustration and not limitation, specific embodiments in which the subject matter may be practiced. As noted above, other embodiments may be utilized and derived therefrom, resulting in structural and logical substitutions and changes being made without departing from the scope of the present disclosure. Such embodiments of the inventive subject matter may be individually or collectively referred to herein by the term "invention," merely for convenience when more than one is actually disclosed, and without any intention to intentionally limit the scope of this patent application to any single invention or inventive concept. Thus, while specific embodiments have been illustrated and described herein, any configuration calculated to achieve the same purpose may be substituted for the specific embodiment shown. The present disclosure is intended to encompass any and all adaptations or modifications of the various embodiments. Combinations of the above-described embodiments, and other embodiments not specifically described herein, will be apparent to those skilled in the art upon reviewing the above description.
Claims
1. 1. A method for aligning sequence reads to a reference sequence, comprising: aligning a first set of sequence reads from an entire population of sequence reads to a reference sequence by Burrows-Wheeler transformation using a first seed length, wherein the first seed length is selected based on an error rate of the sequence reads; masking the first set of sequence reads such that the entire population of sequence reads comprises a subset of masked and unmasked sequence reads; aligning a second set of sequence reads from the unmasked sequence reads to the reference sequence by the Burrows-Wheeler transformation using a second seed length, wherein the second seed length is shorter than the first seed length; and determining an alignment of the sequence reads to the reference sequence based on the first set of sequence reads and the second set of sequence reads; Including, iteratively masking and aligning additional sets of sequence reads with each subsequent read set having a shorter seed length, and determining alignments of the sequence reads with the additional sets of sequence reads; The method further comprises:
2. 2. The method of claim 1, wherein the first seed length is less than 10 bases.
3. 2. The method of claim 1, wherein the first seed length is less than 5 bases.
4. 2. The method of claim 1, wherein the first seed length is four bases.
5. 2. The method of claim 1, wherein the error rate of the sequence reads is at least 5%.
6. 2. The method of claim 1, wherein the error rate of the sequence reads is at least 10%.
7. 2. The method of claim 1, wherein the sequence reads are sequenced from multiple concatemers, each concatemer formed from oligonucleotide sequences linked together, and the oligonucleotide sequences correspond to multiple loci from a set of chromosomes.
8. 8. The method of claim 7, wherein the set of chromosomes comprises chromosomes 13, 18, 22, X, and Y.
9. 8. The method of claim 7, wherein the set of chromosomes is selected from the group consisting of chromosomes 13, 18, 22, X, and Y.
10. 8. The method of claim 7, further comprising calculating the frequency with which each locus is found in the sequence reads.
11. A computer product comprising a computer readable medium storing a plurality of instructions for controlling a computer system to perform the operations of the method of any one of claims 1 to 10.
12. A computer product according to claim 11; one or more processors for executing instructions stored on the computer-readable medium.
13. A system comprising means for carrying out the method according to any one of claims 1 to 10.
14. A system comprising one or more processors configured to perform the method of any one of claims 1 to 10.
15. A system comprising modules for respectively carrying out the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Genomics infrastructure for on-site or cloud-based DNA and RNA processing and analysis
JP2019510323A
Genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis
US20170124254A1
Genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis
WO2017123664A1