Multi-region alignment processing

The method addresses inefficiencies in conventional read alignment by using a global alignment score threshold to efficiently align sequences across multiple candidate regions, improving both speed and accuracy in genomic read mapping.

WO2025122571A1PCT designated stage expired Publication Date: 2025-06-12ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058393
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Conventional read alignment methods, such as the WFA algorithm, face inefficiencies when dealing with multiple candidate alignment regions and split-read alignments, particularly in handling low-scoring 'noise' matches and requiring extensive exploration of alignment matrices.

Method used

A computer-implemented method for aligning a first sequence of nucleobases to a second sequence using a global alignment score threshold, which allows for efficient alignment processing across multiple candidate regions by iteratively adjusting the score threshold and aggregating partial alignments from different regions.

Benefits of technology

This approach significantly improves the efficiency of read alignment by reducing computational expense in low-scoring regions and enabling effective split-read alignment support, thereby enhancing the accuracy and speed of genomic read mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058393_12062025_PF_FP_ABST
    Figure US2024058393_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A process for aligning a first sequence of nucleobases to a second sequence of nucleobases includes identifying a collection of candidate alignment regions of the second sequence. The process additionally includes performing alignment processing for the candidate alignment regions using a global alignment score threshold, the alignment processing including, attempting, for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold. The process then determines whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, and performs processing based on the determining.
Need to check novelty before this filing date? Find Prior Art

Description

MULTI-REGION ALIGNMENT PROCESSINGBACKGROUND

[0001] Two approaches for local sequence alignment are the Needleman-Wunsch and Smith-Waterman algorithms. Both align one sequence of nucleobases to another sequence of nucleobases by exploring the space of alignment possibilities, attempting to match bases of one sequence to the other sequence - possibly with insertions and deletions - to produce an alignment and a corresponding alignment score. Various scoring mechanisms could be employed to score the alignments. In some examples, score bonuses are awarded when the two bases are corresponding positions of the two sequences match. In other examples, score penalties are taken for mismatches, opening a gap (insertion or deletion, or ‘indel’) in the alignment, and / or for lengthening a gap. Other score mechanisms are possible.

[0002] Often the alignment processing to identify alignments is facilitated using a two-dimensional matrix view in which one sequence extends horizontally across the top of the matrix, the other sequence extends vertically down the left size of the matrix, and a process iteratively explores alignment possibilities via horizontal, vertical, and diagonal movement through the matrix. Under this arrangement, horizontal and vertical movement corresponds to indels, diagonal movement corresponds to single-base matches or mismatches, and the process attempts to create a trace from the top left corner of the matrix to the bottom right corner of the matrix to define an alignment between the two sequences. Different alignment paths through the matrix correspond to different alignments and will have corresponding alignment scores depending on the bonuses / penalties taken to complete those paths. The path(s) with the ‘best’ scores are often taken to be the best alignment(s) between the two sequences.

[0003] Exploration of the entire alignment space represented by the matrix may be inefficient. Some approaches therefore limit the space being explored to a subset of the entire matrix. For instance, one approach might explore only a band of matrix cells that are within some range of the matrix diagonal that extends from the top left to bottom right of the matrix. Another approach sets a maximum score loss and the process exploresIP-2682-PCTalignment paths with a score loss up to, but not exceeding, that allowable maximum score loss. The process will stop exploring any given partial alignment path if continuing on that path would result in exceeding the maximum score loss. If no alignment results at the current maximum score loss, the maximum score loss is increased to open the possibility of an alignment at that new maximum score loss. This can be iterated until an alignment is found. For example, the processing might initially tolerate zero score loss and check whether any path can be formed without taking a loss. If no such path can be formed, the process can increase the maximum score loss to increase the score loss tolerance, and continue exploring the alignment space using that new maximum score loss. Increasing the maximum allowed score loss at each iteration allows the process to stray further from the matrix diagonal (representing a perfect alignment), which correlates to exploration of more gapped alignments. Eventually the maximum score loss will increase enough that one or more alignment(s) will be found.SUMMARY

[0004] Some deficiencies exist in the conventional approaches for practical read alignment. Read alignment is often performed between a read sequence and a reference genome in which there are multiple reference genome regions that are candidates for alignment of the read sequence to the reference sequence. Seed mapping, for example, might identify several candidate alignment regions of the reference where the read sequence could be aligned, and it may be desired to explore all of these candidate regions. While one or more of these regions might yield good alignment(s) with comparably high / good score(s), there could also be large number of low-scoring ‘noise’ matches discovered by the seed mapping. This could be problematic for an alignment algorithm that performs high-scoring alignments efficiently but that struggles from an efficiency standpoint with low-scoring regions. The wavefront alignment (WFA) algorithm is one such example. The WFA algorithm would execute independently for each candidate alignment region to find the best alignment score(s) in each region. Some of those alignment scores might be relatively good, while some might be extremely poor (low scoring). Low-scoring alignments are expensive in WFA because it must iterate longer and explore large regions of the alignment matrices to reach large score losses.IP-2682-PCT

[0005] In addition, much secondary analysis, particularly structural variant calling, requires split-read alignment support. In a scenario of a 10 kilobase pair (Kbp) deletion, for example, a read spanning that deletion should be able to align a first portion on its left flank and a second portion on its right flank. There are appropriate scoring models for these structural variant jumps or breaks, with far sub-linear score penalties, e.g., a break penalty linear with the logarithm of jump distance, for instance. One approach is to produce many partially-aligned fragments with local alignment scoring, then combine them in various combinations to achieve higher-scoring split alignments. However, WFA, by way of example, becomes much more expensive when producing substantially clipped local alignments, and does not suggest any method of natively producing splitread alignments.

[0006] Shortcomings of the prior art are overcome and additional advantages are provided through the provision of a computer-implemented method for aligning a first sequence of nucleobases to a second sequence of nucleobases. The method identifies a collection of candidate alignment regions of the second sequence. The method also performs alignment processing for the candidate alignment regions using a global alignment score threshold, the alignment processing including attempting, for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold. The method determines whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, and performs processing based on that determination.

[0007] Additional aspects of the present disclosure are directed to systems and computer program products configured to perform the methods described above and herein. The present summary is not intended to illustrate each aspect of, every implementation of, and / or every embodiment of the present disclosure. Additional features and advantages are realized through the concepts described herein.IP-2682-PCTBRIEF DESCRIPTION OF THE DRAWINGS

[0008] Aspects described herein are particularly pointed out and distinctly claimed as examples in the claims at the conclusion of the specification. The foregoing and other objects, features, and advantages of the disclosure are apparent from the following detailed description taken in conjunction with the accompanying drawings in which:

[0009] FIG. 1 depicts a conceptual diagram of aligning a first sequence of nucleobases to a second sequence of nucleobases at plurality of candidate alignment regions, in accordance with aspects described herein;

[0010] FIG. 2 depicts an example process for aligning a first sequence of nucleobases to a second sequence of nucleobases, in accordance with aspects described herein;

[0011] FIG. 3 depicts an example of aggregating partial alignments from multiple alignment regions to produce a full alignment, in accordance with aspects described herein;

[0012] FIGS. 4A-4G collectively depict a time-series conceptual representation of the aggregation of partial alignments, in accordance with aspects described herein; and

[0013] FIG. 5 depicts one example of such a computer system and associated devices to incorporate and / or use aspects described herein.DETAILED DESCRIPTION

[0014] Described herein are approaches to efficiently perform alignment processing to align a first sequence to a second sequence, even in situations where there are multiple candidate alignment regions to process and / or when considering split alignments. The first and second sequences could be any nucleobase sequences. By way of example and not limitation, the first sequence may be a read sequence determined by a sequencing instrument and the second sequence may be a reference sequence or portion thereof.

[0015] In accordance with one aspect, alignment scoring is performed for each of multiple candidate alignment regions of the second sequence, e.g., the reference using aIP-2682-PCTglobal alignment score threshold, for instance a global score-loss counter, that is maintained and applies to all of the candidate alignment regions. In examples, the alignment scoring is performed for each of the candidate alignment regions in parallel (e.g., simultaneously, contemporaneously) and jointly. Under this approach, the global alignment score threshold is used as a threshold for the alignment processing being performed for each of those candidate alignment regions, the global alignment score threshold is iteratively changed as alignment processing progresses, and changing the global alignment score threshold changes it for the processing at each region. At each iteration, the alignment processing for the candidate alignment regions uses the current global alignment score threshold and attempts, for each of those candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with an alignment score that is within that global alignment score threshold.

[0016] By way of example, FIG. 1 depicts a conceptual diagram of aligning a first sequence of nucleobases to a second sequence of nucleobases at plurality of candidate alignment regions, and FIG. 2 depicts an example process for aligning a first sequence of nucleobases to a second sequence of nucleobases, in accordance with aspects described herein. The sequence lengths, number and length of candidate alignment regions, alignment scores, score losses, thresholds and other example numbers presented in this description are used by way of example only and not limitation. For instance, in practice the sequences may be hundreds, thousands, or millions of bases long and the number of candidate alignment regions could be tens or hundreds, as examples.

[0017] Referring initially to FIG. 1, 102 represents a reference sequence, i.e., the second sequence of nucleobases. Seed mapping for identifying candidate alignment regions of the reference sequence 102 has identified regions 104, 106, and 108 of the reference sequence as candidate alignment regions. Details of each of the three regions are shown below the reference sequence. The alignment processing is to attempt to align the read sequence 110 to each of these candidates. The read sequence 110 in this example has a length of 10 nucleobases: T-C-G-T-A-C-G-T-C-C.IP-2682-PCT

[0018] Referring to FIG. 2, the process identifies (202) the candidate alignment regions (e.g., 104, 106, 108 of FIG. 1) of the second sequence, then sets (204) a global alignment score threshold, which represents a mismatch tolerance for aligning the read sequence to the reference sequence. In a particular example, this could be a maximum score loss that is tolerated for aligning the read sequence to the reference sequence. As explained further below, this global alignment score threshold may be set again when iterating aspects of FIG. 2.

[0019] Assume that the global alignment score threshold is initially set to a score loss of 0 (SL=0). The process will then attempt (206) to align the read sequence to the reference sequence in each candidate alignment region first at that score loss. If an alignment at SL=0 is found, this means that a perfect alignment of the read sequence to the reference sequence was found at a candidate alignment region, or possibly more than one such region. In any case, a found alignment at the current SL could be output at that point in the process.

[0020] Referring back to the example of FIG. 1, the process will attempt to align the read sequence 110 to each of the three candidate alignment regions 104, 106, 108 at the current global alignment score threshold, SL=0 in this example. With respect to region 104, the first five bases, T-C-G-T-A, match but a two-base insertion is reflected. Therefore, a score loss would need to be taken when exploring the alignment space of the read sequence 110 to this candidate alignment region 104, and therefore the alignment processing with respect to this region could halt / break after finding that no alignment in this region is possible at no score loss. Similarly, the alignment processing with respect to candidate alignment region 106 will be unsuccessful at zero score loss because of a base mismatch / point mutation (T vs. C) at position 6, and the alignment processing with respect to candidate alignment region 108 will be unsuccessful at zero score loss because of a two-base deletion. Therefore, a score loss would also need to be taken when exploring the alignment spaces of the read sequence 110 to each of the candidate alignment regions 104, 106, 108, and so no alignment to any of the candidates is possible at SL=0.IP-2682-PCT

[0021] Returning to FIG. 2, the processing proceeds by determining (208) whether to continue to a next iteration. The determination could be based on any of various parameters, including the result of 206. Continuing with a practical example in which at least some amount of alignment mismatch is tolerated and therefore exploration at higher score loss is desired, the process determines to continue to a next iteration (208, Y) and returns to 204 to change the global alignment score threshold. The change will increase the mismatch tolerance of successful alignment, meaning will change the global alignment score threshold to tolerate a greater degree of mismatch between the two sequences. In the example of a score loss, the process might set the global alignment score threshold to reflect a score loss of ‘ 1’ (SL=1), meaning 1 point of score loss is tolerable for a successful full alignment. The process continues by attempting (206) alignment of the read sequence to the reference sequence to explore alignment paths in the candidate alignment regions at that changed global alignment score threshold, i.e., SL=1.

[0022] Referring again to FIG. 1, assume the applicable score loss scheme is such that a base mismatch results in a score loss of 1, opening an indel / gap results in a score loss of 2, and an indel of base length n results in an additional score loss of n. In that case, neither the processing of candidate alignment region 104 nor the processing of candidate alignment 108 will result in an alignment with a score loss within the global alignment score threshold of 1, since the two-base indel to successfully align to each region 104, 108 would require a score loss of 2 for opening the gap plus 2 for the two- base indel, for a total of 4 score loss. However, alignment of the read sequence 110 to the reference sequence at candidate alignment region 106 is possible at that score loss. That is, a full alignment of the read sequence 110 to the reference sequence at region 106 is possible at a score loss of 1 (due to the single base mismatch). The aligning attempted at 206 on this iteration would result in an alignment in region 106 with a score within the global alignment score threshold.

[0023] If at any point an alignment at the current global alignment score threshold is found, the alignment can be output, retained in memory, or handled in any other way desired. If desired, the alignment processing can halt on the first instance of finding anIP-2682-PCTend-to-end alignment of the full read in any candidate alignment region. This may be considered the globally best scoring alignment. Alternatively, since processing of each candidate alignment region may not proceed at the same rate, it may be desired in these situations to let the processing at that iteration continue to discover whether any other alignments at the global alignment score threshold result. Furthermore, after finding on or more alignments at a given threshold (at SL=1 in this example), the processing could continue with possibly higher SL values until at least one second-best alignment candidate is found, say at SL=2 or higher. Using the example discussed above, the process could continue by (i) determining to continue at 208, (ii) setting the global alignment score threshold (at 204) to SL=2, and (iii) attempting (at 206) to align the candidates at SL=2. In the example of FIG. 1, this will be unsuccessful for the other candidates 104, 108, so the process could iterate again, this time to SL=3, which again would result in unsuccessful alignment for the other candidates 104, 108. Finally, a full alignment will be found for both candidate regions 104, 1408 after setting the global alignment score threshold to SL=4. One or both alignments could be output at that point.

[0024] As noted above, the determination at 208 about whether to continue alignment processing for a next iteration could based on any of various parameters. One goal may be to find at least some minimum number of candidate alignments regardless how high the mismatch tolerance must be to find those. Another might be to find all alignments that result within some maximum global alignment score threshold, say SL=7.

[0025] Yet another approach can consider an alignment confidence metric, for instance a MAPQ calculation for mapping quality. This metric can be determined as a function of (at least) a difference between two global alignment scores, for instance between (i) the global alignment score at which a best alignment was found (SL=x) and (ii) another global alignment score (SL=y), such as one at which a next-best alignment was found or the current global alignment score threshold in the iterating of the alignment processing. The iterating could halt once, for example, the score loss of the second global alignment score threshold (SL=y) exceeds the score loss of the first global alignment score threshold (SL=x) by enough that MAPQ would have reached some threshold maximum, for instance MAPQ=60. In other words, MAPQ can be estimated inIP-2682-PCTproportion to the difference between two score losses (SLy - SLx), and there is some score loss y’ at which SLy’-SLx corresponds to that maximum MAPQ, meaning any further score loss increase above y’ would result in an undesirably large MAPQ. Thus, once reaching y’ in the alignment processing of FIG. 2, it may not be worth exploring greater score losses because MAPQ would be worse than the maximum of interest, and therefore the iterating could halt after that y’ is reached and if no next-best alignment has already been found.

[0026] To help illustrate, assume that the maximum desired MAPQ is set at 60 and that this corresponds to a score loss difference (SL2-SL1) of 20. If the best alignment is found at SL1=7O, then processing might continue to search for the next-best alignment at scores between 70 and 90 (20 more than 70) until either a next-best alignment is found or the process has iterated to the point where the score loss threshold (global alignment score threshold) has reached 90 and no successful alignment in any region was produced at that threshold. Similarly, if instead SL1 was 80, then processing could stop after SL2 reaches 100 (20 more than 80), or earlier.

[0027] Using the example of FIG. 1, the best alignment was found at a global alignment score threshold corresponding to SL=1, and the second best was found at a global alignment score threshold corresponding to SL=4. This corresponds to a score difference of 3. If a threshold alignment confidence metric was set such that a score loss difference of 3 is too high, meaning SL=3 is the maximum, then the iterating might have stopped after exploring at SL=3, prior to finding either of the second-best alignments at SL=4.

[0028] It is noted that the alignment confidence approach could be taken with respect to an alignment that is not necessarily the best. For instance, one approach might find the best alignment, then the second-best alignment, and then all other alignments within some maximum score loss or alignment confidence (e.g., MAPQ) relative to that second- best alignment. Various other approaches are possible.

[0029] In some embodiments, the collection of a candidate alignment regions could be refined as alignments are found. For instance, a candidate alignment region in whichIP-2682-PCTan alignment is found at a given global alignment score threshold may be removed from consideration on subsequent iterations of the alignment processing (higher global alignment score thresholds), if desired.

[0030] In one example of alignment processing and working through the iterations, the aligning could proceed from an initial anchor / seed position that is not at one end or the other of the read sequence, extend first in one direction, for instance leftward, from the anchor position, and then switch to extension in the other direction from the anchor position. As the aligning proceeds in a given iteration, any necessary score loss would be taken and accumulated, not to exceed the global alignment score threshold on that iteration. If on a given iteration the aligning in the one direction reaches the beginning of the read sequence and the accumulated score loss from progressing in that direction does not exceed the then-current global alignment threshold, then the aligning switches to extension in the other direction from the anchor position, continuing taking any necessary additional score loss, with the total not to exceed the global alignment score threshold on that iteration. Further examples of score loss accumulation and directional progression are discussed below.

[0031] Split-read alignments may be supported in the context above by exploring potential aggregation of multiple partial alignments from different candidate alignment regions. In one particular approach, alignment processing proceeds as above based on seed mapping (or other candidate identification approach) and iteratively works through global alignment score thresholds, e.g., via a score loss counter for SL=0, 1, 2, etc. In this approach, the alignment of each regional alignment fragment, also referred to as a partial alignment, within a given candidate alignment region could extend initially in one direction in the read - leftward, for instance - exploring alignment paths within the global alignment score threshold and trying to reach the beginning of the read. As part of this, it could explore taking a score penalty to bridge further in that direction to another partial alignment, for instance one in another candidate alignment region upstream in the reference. If a partial alignment reaches the beginning of the read, the alignment processing could use further score loss increments to extend in the other direction - rightward in this example - trying to reach the end of the read, while exploring bridgingIP-2682-PCTfurther in that other direction to another partial alignment, for instance one in another candidate alignment region downstream in the reference. In this manner, at any point in the alignment processing two partial alignments in a corresponding two candidate alignment regions that have aligned to adjacent positions in the read from opposite directions may potentially bridge between each other with a structural variant jump by paying an appropriate break penalty. The break penalty could be a function of distance and orientation, for instance. A partial alignment extending leftward may thus reach further left by jumping to another partial alignment extending further left. If this leads to a path reaching the beginning of the read, this establishes a completed leftward alignment path. Similarly, rightward extension of that fragment may consider a jump further right to another partial alignment. It is seen that a complete / full alignment could be built as an aggregate of two or more partial alignments, and could be considered a valid full alignment within the current global alignment score threshold as long as the accumulated score loss, which includes the respective loss within each partial alignment plus any score penalties associated with the bridging, does not exceed that global alignment score loss threshold.

[0032] In a particular approach, the read sequence is anchored to each candidate alignment region based on seed matching and the matrix is extended by some number of bases, say 50 to 100, on each end of the read so that the read sequence is anchored somewhere in the diagonal away from the top left and bottom right comers of the matrix. This enables exploration of the space left and right of the anchor position.

[0033] In a situation of a gapped alignment that spans two candidate regions, the read is anchored somewhere in each region and the alignment processing can begin exploring in one direction (say leftward / upstream) in each region. Assume the first region is located upstream from the second region. The bases in the first region might match relatively well working leftward, while the second region may not match well on its left side because of the gap between the two regions. Similarly, after working leftward in the first region and reaching the left end of the read, the processing of that region can start working rightward / down stream but may encounter significant mismatching because of the gap. However, the processing of the first region can consider downstream jumpsIP-2682-PCTrightward to the second region with relatively little score loss compared to the score loss taken to finish the rightward alignment in the first region, and similarly the processing of the second region can consider upstream jumps leftward to the first region with relatively little score loss compared to the score loss taken to finish the leftward alignment in the second region.

[0034] FIG. 3 depicts an example of aggregating partial alignments from multiple alignment regions to produce a full alignment, in accordance with aspects described herein. The processing in this example identified three candidate alignment regions, 304, 306, 308, in reference sequence 302. The read sequence is given by 310.

[0035] With respect to candidate alignment region 304, assume the read is anchored by bases C-G in positions 2 and 3. The alignment processing for this region might begin by working leftward to match base T in position 1 and reach the leftmost portion of the read. At this point, the processing begins working rightward from position 4 but finds the beginning of what would eventually be found to be a significantly long chain of mismatched bases. The processing therefore might consider whether there are any partial alignments downstream, e.g., in regions 306 or 308, to which the partial alignment T-C-G of region 304 might bridge.

[0036] With respect to candidate alignment region 306, assume the read is anchored by bases T-A in positions 4 and 5. The alignment processing for this region might begin by working leftward but immediately encounter a mismatch at position 3. This could prompt the alignment processing of 306 to consider whether there are any partial alignments in another region, say region 304, to which the partial alignment of region 304 might bridge. In this case, the alignment processing might identify that it could bridge the partial alignment T-C-G of region 304 to the partial alignment T-A of region 306 to result in a partial alignment T-C-G-T-A for positions 1 through 5 of the read sequence that includes a gap (gap 1) between them. At that point, the read is aligned in the leftward direction to the beginning of the read. The alignment processing for region 306 might therefore begin working rightward at that point and find the beginning of a long chain of mismatched bases beginning at position 7. The processing therefore might consider whether there are any partial alignments downstream, e.g., in region 308, to which theIP-2682-PCTpartial alignment of region 306 (which itself might be bridged on its left end) might bridge.

[0037] Similarly, with respect to candidate alignment region 308, the alignment processing for this region might begin by working leftward from the anchor position, encounter mismatches beginning at position 6, but identify a possibility of bridging to the partial alignment of region 306 at its position 6. In this case, the alignment processing might identify that it could bridge the partial alignment of region 308 to the partial alignment of region 306 via a gap (gap 2). The partial alignment of region 306 is itself bridged to the partial alignment of region 304 which reaches a front end of the read sequence, and the processing identifies that the read has been aligned fully in the left direction from the anchor position in region 308. If there are remaining bases of the read to the right of the anchor position in region 310, the alignment can proceed rightward to the end of the read sequence. The processing will recognize that aggregation of these three partial alignments results in a full alignment of the read sequence to the reference sequence. If gap penalties and any other penalties taken in the partial alignments (here there are none because there are no base mismatches or indels within any partial alignment) do not aggregate to exceed the global alignment score threshold, then this full alignment satisfies that threshold to provide a useful alignment of the read sequence to the reference sequence.

[0038] In a split-read situation where there are two partial alignments to bridge, it might make sense to determine that one partial alignment fragment has already reached one end of the read to make it eligible for another fragment to bridge to. If trying to support aggregation of three (or more) partial alignments, such read-end anchoring might only make sense to enforce at one (globally chosen) end.

[0039] Additionally, a given bridged alignment (two or more partial alignments that have been bridged together) is not necessarily be a full alignment, meaning it may not reach both ends of the read. In other words, one of the components being bridged together with another component may itself already be a bridged partial alignment. Using a left-end anchoring requirement, for example, which was depicted and described with reference to FIG. 3, the process could first bridge a left-anchored partial alignment A to aIP-2682-PCTmiddle partial alignment B, forming a left-anchored aggregate partial alignment AB. The process could then bridge this left-anchored aggregate partial alignment AB to a right partial alignment C to form a left-anchored aggregate alignment ABC. Even a triplefragment alignment may not be a full alignment yet, requiring further iteration, potentially though additional global alignment score threshold(s) with higher mismatch tolerance, before the ‘C’ portion of ABC reaches the right end of the read, either within the regions of C or by bridging to yet another partial alignment in another region.

[0040] FIGS. 4A-4G collectively depict a time-series conceptual representation of the aggregation of partial alignments, in accordance with aspects described herein. This example aggregates three partial alignments to produce a full alignment of the read sequence to the reference sequence, but the principles discussed herein could be extended to aggregate any number of partial alignments. For simplicity, the depictions reflect (using the # sign) the portions of the read sequence that are partially or fully aligned and omit the reference sequence and positioning thereof. It is noted that each partial alignment depicted by a string of # characters may be a perfect or imperfect partial alignment, meaning one with no base mismatches or one with mismatches, indels, or the like. That is, a partial alignment need not be perfect with respect to the bases that it covers.

[0041] In FIG. 4A, alignment processing has proceeded to build three partial alignments labeled A, B, and C, in a respective three different candidate alignment regions by extending the partial alignments leftward (indicated by ‘<’ at the left side of each partial alignment string denoted by # characters). In FIG. 4B, partial alignment A reaches the left end of the read and the alignment processing switches to exploring rightward to extending the partial alignment A in that direction. Meanwhile, the alignment processing has extended partial alignments B and C leftward relative to FIG. 4A.

[0042] At FIG. 4C, partial alignments A and B have reached adjacent positions (at 402) in the read and become bridged (FIG. 4D) into aggregate partial alignment AB. Recognizing that the partial alignment AB is fully left anchored having extended to the beginning of the read sequence, the alignment processing will switch to extending theIP-2682-PCTpartial alignment rightward as indicated by ‘>’ at the right end of the partial alignment B (which has been aggregated with A). At FIG. 4E, aggregate fragment AB and partial alignment C have reached adjacent positions (at 404) in the read. At FIG. 4F, aggregate partial alignment AB and partial alignment C become bridged into aggregate partial alignment ABC. The alignment processing will switch to extending the partial alignment C (now bridged to B and C) rightward as indicated by ‘>’ at the right end of the partial alignment C. At FIG. 4G, the aggregate partial alignment ABC reaches the right end of the read, and becomes a full alignment.

[0043] In terms of implementation, one approach is to consider an aggregate partial alignment, such as AB of FIGS. 4D and 4E, as an evolution of its rightmost fragment, e g., B. When A is left-anchored and A and B reach adjacent positions and become bridged, fragment B thereby becomes left-anchored by virtue of its bridged alignment path through A. Then being left-anchored, fragment B switches to rightward extension. Fragment C bridges to the now left-anchored fragment B (which has a leftward path through A), and fragment C becomes left-anchored by virtue of its bridged alignment path through B and A. Then being left-anchored, fragment C switches to rightward extension, and eventually reaches the right end of the read sequence such that the aggregate becomes a complete alignment.

[0044] This implementation approach permits more complex split-alignment evolutions. For instance, although partial alignment B is left-anchored after bridging to A and switches to rightward extension, the alignment processing could continue rightward extension of partial alignment A and might plausibly align the right end of the read sequence on its own, i.e., before partial alignment B aligns the right end of the read sequence. It is also possible that another partial alignment, other than partial alignment B, might bridge to partial alignment A, building its own split-read alignment in competition with that of partial alignment B.

[0045] In the examples above, bridging occurs between adjacent partial alignments, however this need not be the case. Opportunities for bridging can be explored for partial alignments whether upstream or downstream from the current position, across chromosomes, etc. The score penalties can vary depending on the size of the gap, forIP-2682-PCTexample. Additionally, aspects can account for differences in orientation. For instance, continued alignment after a jump need not necessarily proceed in the direction of the jump - for instance a jump downstream in the reference sequence might be followed by subsequent alignment in the upstream direction of the reference sequence as the read sequence is aligned in its downstream direction (e.g., moving rightward in the read sequence but leftward in the reference sequence), for instance in the case of an inversion.

[0046] The split alignment processing can occur within the context of the alignment processing described above with reference to FIGS. 1 and 2, i.e., iteratively working through global alignment score thresholds to find alignments. The scenario of FIG. 4 might therefore proceed during the course of this alignment processing through potentially several iterations. Partial alignments may be built gradually during the iteration through these global alignment score thresholds. The alignment processing can therefore maintain alignment status and results, for instance alignment positions, partial alignments, and scoring, for both current and past iterations of alignment processing, and for each of the regions. This can facilitate the exploration and identification of partial alignments to possibly bridge together.

[0047] While the WFA algorithm mentioned previously can offer a dramatic speedup of a core bioinformatics algorithm, its advantage to practical genomic read mapping is limited as currently defined. In addition to the drawbacks noted previously, there are other ways to avoid excessive Smith-Waterman-type computation in read mapping. Some secondary analysis mappers, for example, compute much cheaper (from a resource consumption standpoint) gapless alignments of all candidates first, and heuristically determine which candidates should proceed to Smith- Waterman-type gapped alignment, with only about 2% of candidates undergoing the Smith-Waterman alignment algorithm. This heuristic effort limitation results in small accuracy losses, where additional Smith- Waterman-type work would have found better alignments. However, this pertains to short-read mapping; long reads tend to require gapped (and split-read) alignment for most or all candidates. Thus, simply replacing Smith-Waterman-type calculations in secondary analysis mappers with the WFA algorithm may be unlikely to improve performance, as it might replace only a minority (30-40%) of current total compute consumption. Further,IP-2682-PCTthis might not result in significant speed improvements because some Smith-Waterman operations are for lower-scoring candidates where WFA remains expensive. In contrast, aspects described herein provide advantages. By comprehensively replacing candidate alignment including split-read analysis, it can solve a large majority of the total compute and improve accuracy somewhat by eliminating heuristic tradeoffs Additionally, the multi-region scoring using a single, global alignment score threshold applicable at any given time to each candidate alignment region could perform alignment processing cheaper than current methods because the alignment processing would not be left running to expensive completion in poor-scoring candidate alignment regions, i.e., it would halt based on halting parameter(s) before expending such resources. This could provide a better path forward for long-read mapping, or for non-FPGA accelerated software installations of secondary analysis software, as examples.

[0048] Processing described herein could be performed by any appropriate system, for instance by a genomic sequencing instrument that sequences a sample to produce the read sequence, a computer system that is in communication with such a sequencing instruction, or any other computer system, for instance. Example computer systems could be existent in a cloud environment and / or an on-site environment, for instance at a same site as a sequencing instrument. Therefore, and referring back to the process of FIG. 2, the process may be executed, in one or more examples, by a processor or processing circuitry of one or more computer s / computer systems, such as those described herein. Code or instructions implementing the process(es) of FIG. 2 may be part of one or more modules and / or sub-module(s).

[0049] Additional examples of processing of FIG. 2 are provided. The process identifies (202) a collection of candidate alignment regions of the second sequence, sets (204) a global alignment score threshold, and performs alignment processing for the candidate alignment regions using the global alignment score threshold. The alignment processing includes attempting (206), for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold.IP-2682-PCT

[0050] In examples, the attempting (206) to align the first sequence to the second sequence for a candidate alignment region of the collection of candidate alignment regions includes first aligning a portion of the first sequence to the second sequence in a first direction for one or more nucleobases until reaching a first end of the first sequence, and next aligning the portion of the first sequence to the second sequence in a second direction for one or more nucleobases until reaching a second end of the first sequence.

[0051] The alignment processing may or may not result in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold. The process determines whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, and performs processing based on that determination. For instance, based on determining that the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, this can include outputting the at least one of those alignment(s).

[0052] In addition, the process includes determining (at 208) whether to continue / halt this process by returning to 204 and iterating through 204, 206, 208. At each such iteration, there may be such a determination whether to halt or continue. Thus, if determined that the alignment processing results in no alignments of the first sequence to the second sequence with a score within the global alignment score threshold on the current iteration, the performing processing could include determining to continue to a next iteration (208, Y) and iterating, one or more times, the alignment processing and the determining. At each iteration of the iterating, the process changes the global alignment score threshold (i.e., at 204) to increase a mismatch tolerance of successful alignment of the first sequence to the second sequence, the alignment processing uses the changed global alignment score threshold (at 206), and the process determines whether the alignment processing using the changed global alignment score threshold results in one or more alignments of the first sequence to the second sequence with a score within the changed global alignment score threshold. Based on determining, at any iteration of the iterating, that the alignment processing performed at the iteration results in an alignmentIP-2682-PCTof the first sequence to the second sequence with a score within the global alignment score threshold at the iteration, the process can output the alignment. The process can determine to continue the iterating to search for a next-best alignment, if desired.

[0053] The determining whether to halt the iterating at a current iteration may be based on (i) the global alignment score threshold at the current iteration and (ii) the global alignment score threshold at a prior iteration, of the iterating, that resulted in an alignment of the first sequence to the second sequence. For instance, the determining whether to halt can include determining whether an alignment confidence metric, determined as a function of a difference between the global alignment score threshold at the current iteration and the global alignment score threshold at a prior iteration (for instance one at which the best-scoring alignment was found), corresponds to a threshold alignment confidence.

[0054] As another example, the process can determine to halt the iterating (208, N) based on the mismatch tolerance corresponding to the global alignment score threshold increasing to a threshold mismatch tolerance as a result of changing the global alignment score threshold. For example, the processing might halt once the alignment processing at a maximum tolerable global alignment score threshold has been performed.

[0055] In some examples, the attempting to align (206) includes attempting to aggregate a plurality of partial alignments of the first sequence to the second sequence to produce a full alignment of the first sequence to the second sequence within the global alignment score threshold, where each partial alignment of the plurality of partial alignments has the first sequence anchored at a respective candidate alignment region of the collection of candidate alignment regions and partially aligning to the second sequence. The attempting to aggregate can include, for example, for each partial alignment of the plurality of partial alignments: 1. aligning a respective portion of the first sequence to the second sequence in a first direction until reaching either (i) a position, in the first sequence, at which to bridge to a next partial alignment of the partial alignments in the first direction or (ii) a first end of the first sequence, and 2. aligning the respective portion of the first sequence to the second sequence in a second direction until reaching either (i) a position, in the first sequence, at which to bridge to a next partialIP-2682-PCTalignment in the second direction or (ii) a second end of the first sequence. In this manner, each partial alignment can be aligned in one direction until it is anchored in that direction by reaching one end of the read sequence either within its region or by bridging to another partial alignment that is so anchored in that one direction, and aligned in another direction until it is anchored in that other direction by reaching the other end of the read sequence either within its region or by bridging to another partial alignment that is so anchored in that other direction.

[0056] Processes described herein may be performed singly or collectively by one or more computer systems. FIG. 5 depicts one example of such a computer system and associated devices to incorporate and / or use aspects described herein. A computer system may also be referred to herein as a data processing device / system, computing device / system / node, or simply a computer. The computer system may be based on one or more of various system architectures and / or instruction set architectures, such as those offered by International Business Machines Corporation (Armonk, New York, USA) or Intel Corporation (Santa Clara, California, USA), as examples.

[0057] FIG. 5 shows a computer system 500 in communication with external device(s) 512. Computer system 500 includes one or more processor(s) 502, for instance central processing unit(s) (CPUs). A processor can include functional components used in the execution of instructions, such as functional components to fetch program instructions from locations such as cache or main memory, decode program instructions, and execute program instructions, access memory for instruction execution, and write results of the executed instructions. A processor 502 can also include register(s) to be used by one or more of the functional components. Computer system 500 also includes memory 504, input / output (I / O) devices 508, and VO interfaces 510, which may be coupled to processor(s) 502 and each other via one or more buses and / or other connections. Bus connections represent one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include the Industry Standard Architecture (ISA), the Micro Channel Architecture (MCA), the Enhanced ISA (EISA), the Video ElectronicsIP-2682-PCTStandards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI).

[0058] Memory 504 can be or include main or system memory (e.g. Random Access Memory) used in the execution of program instructions, storage device(s) such as hard drive(s), flash media, or optical media as examples, and / or cache memory, as examples. Memory 504 can include, for instance, a cache, such as a shared cache, which may be coupled to local caches (examples include LI cache, L2 cache, etc.) of processor(s) 502. Additionally, memory 504 may be or include at least one computer program product having a set (e.g., at least one) of program modules, instructions, code or the like that is / are configured to carry out functions of embodiments described herein when executed by one or more processors.

[0059] Memory 504 can store an operating system 505 and other computer programs 506, such as one or more computer programs / applications that execute to perform aspects described herein. Specifically, programs / applications can include computer readable program instructions that may be configured to carry out functions of embodiments of aspects described herein.

[0060] Examples of EO devices 508 include but are not limited to microphones, speakers, Global Positioning System (GPS) devices, cameras, lights, accelerometers, gyroscopes, magnetometers, sensor devices configured to sense light, proximity, heart rate, body and / or ambient temperature, blood pressure, and / or skin resistance, and activity monitors. An EO device may be incorporated into the computer system as shown, though in some embodiments an EO device may be regarded as an external device (512) coupled to the computer system through one or more EO interfaces 510.

[0061] Computer system 500 may communicate with one or more external devices 512 via one or more EO interfaces 510. Example external devices include a keyboard, a pointing device, a display, and / or any other devices that enable a user to interact with computer system 500. Other example external devices include any device that enables computer system 500 to communicate with one or more other computing systems or peripheral devices such as a printer. A network interface / adapter is an example EOIP-2682-PCTinterface that enables computer system 500 to communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), providing communication with other computing devices or systems, storage devices, or the like. Ethernet-based (such as Wi-Fi) interfaces and Bluetooth® adapters are just examples of the currently available types of network adapters used in computer systems (BLUETOOTH is a registered trademark of Bluetooth SIG, Inc., Kirkland, Washington, U.S.A.).

[0062] The communication between I / O interfaces 510 and external devices 512 can occur across wired and / or wireless communications link(s) 511, such as Ethernet-based wired or wireless connections. Example wireless connections include cellular, Wi-Fi, Bluetooth®, proximity -based, near-field, or other types of wireless connections. More generally, communications link(s) 511 may be any appropriate wireless and / or wired communication link(s) for communicating data.

[0063] Particular external device(s) 512 may include one or more data storage devices, which may store one or more programs, one or more computer readable program instructions, and / or data, etc. Computer system 500 may include and / or be coupled to and in communication with (e.g. as an external device of the computer system) removable / non-removable, volatile / non-volatile computer system storage media. For example, it may include and / or be coupled to a non-removable, non-volatile magnetic media (typically called a "hard drive"), a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and / or an optical disk drive for reading from or writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM or other optical media.

[0064] Computer system 500 may be operational with numerous other general purpose or special purpose computing system environments or configurations. Computer system 500 may take any of various forms, well-known examples of which include, but are not limited to, personal computer (PC) system(s), server computer system(s), such as messaging server(s), thin client(s), thick client(s), workstation(s), laptop(s), handheld device(s), mobile device(s) / computer(s) such as smartphone(s), tablet(s), and wearable device(s), multiprocessor system(s), microprocessor-based system(s), telephonyIP-2682-PCTdevice(s), network appliance(s) (such as edge appliance(s)), virtualization device(s), storage controller(s), set top box(es), programmable consumer electronic(s), network PC(s), minicomputer system(s), mainframe computer system(s), and distributed cloud computing environment(s) that include any of the above systems or devices, and the like.

[0065] Aspects described herein may be a system, a method, and / or a computer program product, any of which may be configured to perform or facilitate aspects described herein.

[0066] In some embodiments, aspects of the present invention may take the form of a computer program product, which may be embodied as computer readable medium(s). A computer readable medium may be a tangible storage device / medium having computer readable program code / instructions stored thereon. Example computer readable medium(s) include, but are not limited to, electronic, magnetic, optical, or semiconductor storage devices or systems, or any combination of the foregoing. Example embodiments of a computer readable medium include a hard drive or other mass-storage device, an electrical connection having wires, random access memory (RAM), read-only memory (ROM), erasable-programmable read-only memory such as EPROM or flash memory, an optical fiber, a portable computer disk / diskette, such as a compact disc read-only memory (CD-ROM) or Digital Versatile Disc (DVD), an optical storage device, a magnetic storage device, or any combination of the foregoing. The computer readable medium may be readable by a processor, processing unit, or the like, to obtain data (e.g. instructions) from the medium for execution. In a particular example, a computer program product is or includes one or more computer readable media that includes / stores computer readable program code to provide and facilitate one or more aspects described herein.

[0067] As noted, program instruction contained or stored in / on a computer readable medium can be obtained and executed by any of various suitable components such as a processor of a computer system to cause the computer system to behave and function in a particular manner. Such program instructions for carrying out operations to perform, achieve, or facilitate aspects described herein may be written in, or compiled from code written in, any desired programming language. In some embodiments, such programmingIP-2682-PCTlanguage includes object-oriented and / or procedural programming languages such as C, C++, C#, Java, etc.

[0068] Program code can include one or more program instructions obtained for execution by one or more processors. Computer program instructions may be provided to one or more processors of, e.g., one or more computer systems, to produce a machine, such that the program instructions, when executed by the one or more processors, perform, achieve, or facilitate aspects of the present invention, such as actions or functions described in flowcharts and / or block diagrams described herein. Thus, each block, or combinations of blocks, of the flowchart illustrations and / or block diagrams depicted and described herein can be implemented, in some embodiments, by computer program instructions.

[0069] Although various embodiments are described above, these are only examples.

[0070] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising”, when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0071] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below, if any, are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of one or more embodiments has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiment was chosen and described in order to best explain various aspects and the practical application, and to enable others of ordinaryIP-2682-PCTskill in the art to understand various embodiments with various modifications as are suited to the particular use contemplated.IP-2682-PCT

Claims

CLAIMSWhat is claimed is:

1. A method of aligning a first sequence of nucleobases to a second sequence of nucleobases, the method comprising: identifying a collection of candidate alignment regions of the second sequence; performing alignment processing for the candidate alignment regions using a global alignment score threshold, the alignment processing comprising: attempting, for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold; determining whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold; and performing processing based on the determining.

2. The method of claim 1, wherein based on determining that the alignment processing results in no alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises iterating, one or more times, the alignment processing and the determining, wherein at each iteration of the iterating, the method changes the global alignment score threshold to increase a mismatch tolerance of successful alignment of the first sequence to the second sequence, the alignment processing uses the changed global alignment score threshold, and the determining determines whether the alignment processing using the changed global alignment score threshold results in one or more alignments of the first sequence to the second sequence with a score within the changed global alignment score threshold.IP-2682-PCT3. The method of claim 2, wherein based on determining, at an iteration of the iterating, that the alignment processing performed at the iteration results in an alignment of the first sequence to the second sequence with a score within the global alignment score threshold at the iteration, the method further comprises outputting the alignment.

4. The method of claim 3, further comprising continuing the iterating to search for a next-best alignment of the first sequence to the second sequence.

5. The method of claim 4, further comprising determining whether to halt the iterating at a current iteration, wherein the determining whether to halt is based on (i) the global alignment score threshold at the current iteration and (ii) the global alignment score threshold at a prior iteration, of the iterating, that resulted in an alignment of the first sequence to the second sequence.

6. The method of claim 5, wherein the determining whether to halt comprises determining whether an alignment confidence metric, determined as a function of a difference between the global alignment score threshold at the current iteration and the global alignment score threshold at a prior iteration, corresponds to a threshold alignment confidence.

7. The method of claim 2, further comprising halting the iterating based on the mismatch tolerance increasing to a threshold mismatch tolerance as a result of changing the global alignment score threshold.

8. The method of claim 1, wherein based on determining that the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises outputting at least one alignment of the one or more alignments.

9. The method of claim 1, wherein the attempting to align comprises attempting to aggregate a plurality of partial alignments of the first sequence to the second sequence to produce a full alignment of the first sequence to the second sequence within the global alignment score threshold, each partial alignment of the plurality ofIP-2682-PCTpartial alignments having the first sequence anchored at a respective candidate alignment region of the collection of candidate alignment regions and partially aligning to the second sequence.

10. The method of claim 9, wherein the attempting to aggregate comprises, for each partial alignment of the plurality of partial alignments: aligning a respective portion of the first sequence to the second sequence in a first direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment, of the partial alignments, in the first direction, or (ii) a first end of the first sequence; and aligning the respective portion of the first sequence to the second sequence in a second direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment in the second direction, or (ii) a second end of the first sequence.

11. The method of claim 1, wherein the attempting to align the first sequence to the second sequence for a candidate alignment region of the collection of candidate alignment regions comprises: first aligning a portion of the first sequence to the second sequence in a first direction for one or more nucleobases until reaching a first end of the first sequence; and next aligning the portion of the first sequence to the second sequence in a second direction for one or more nucleobases until reaching a second end of the first sequence.

12. A computer system comprising: a memory; and a processing circuit in communication with the memory, wherein the computer system is configured to perform a method comprising:IP-2682-PCTidentifying a collection of candidate alignment regions of the second sequence; performing alignment processing for the candidate alignment regions using a global alignment score threshold, the alignment processing comprising: attempting, for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold; determining whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold; and performing processing based on the determining.

13. The computer system of claim 12, wherein based on determining that the alignment processing results in no alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises iterating, one or more times, the alignment processing and the determining, wherein at each iteration of the iterating, the method changes the global alignment score threshold to increase a mismatch tolerance of successful alignment of the first sequence to the second sequence, the alignment processing uses the changed global alignment score threshold, and the determining determines whether the alignment processing using the changed global alignment score threshold results in one or more alignments of the first sequence to the second sequence with a score within the changed global alignment score threshold.

14. The computer system of claim 13, wherein based on determining, at an iteration of the iterating, that the alignment processing performed at the iteration results in an alignment of the first sequence to the second sequence with a score within the global alignment score threshold at the iteration, the method further comprises outputting the alignment.IP-2682-PCT15. The computer system of claim 14, wherein the method further comprises continuing the iterating to search for a next-best alignment of the first sequence to the second sequence.

16. The computer system of claim 15, wherein the method further comprises determining whether to halt the iterating at a current iteration, wherein the determining whether to halt is based on (i) the global alignment score threshold at the current iteration and (ii) the global alignment score threshold at a prior iteration, of the iterating, that resulted in an alignment of the first sequence to the second sequence.

17. The computer system of claim 16, wherein the determining whether to halt comprises determining whether an alignment confidence metric, determined as a function of a difference between the global alignment score threshold at the current iteration and the global alignment score threshold at a prior iteration, corresponds to a threshold alignment confidence.

18. The computer system of claim 13, further comprising halting the iterating based on the mismatch tolerance increasing to a threshold mismatch tolerance as a result of changing the global alignment score threshold.

19. The computer system of claim 12, wherein based on determining that the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises outputting at least one alignment of the one or more alignments.

20. The computer system of claim 12, wherein the attempting to align comprises attempting to aggregate a plurality of partial alignments of the first sequence to the second sequence to produce a full alignment of the first sequence to the second sequence within the global alignment score threshold, each partial alignment of the plurality of partial alignments having the first sequence anchored at a respective candidate alignment region of the collection of candidate alignment regions and partially aligning to the second sequence.IP-2682-PCT21. The computer system of claim 20, wherein the attempting to aggregate comprises, for each partial alignment of the plurality of partial alignments: aligning a respective portion of the first sequence to the second sequence in a first direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment, of the partial alignments, in the first direction, or (ii) a first end of the first sequence; and aligning the respective portion of the first sequence to the second sequence in a second direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment in the second direction, or (ii) a second end of the first sequence.

22. The computer system of claim 12, wherein the attempting to align the first sequence to the second sequence for a candidate alignment region of the collection of candidate alignment regions comprises: first aligning a portion of the first sequence to the second sequence in a first direction for one or more nucleobases until reaching a first end of the first sequence; and next aligning the portion of the first sequence to the second sequence in a second direction for one or more nucleobases until reaching a second end of the first sequence.

23. A computer program product comprising: a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit to perform: identifying a collection of candidate alignment regions of the second sequence;IP-2682-PCTperforming alignment processing for the candidate alignment regions using a global alignment score threshold, the alignment processing comprising: attempting, for each candidate alignment region of the collection of candidate alignment regions, to align the first sequence, anchored at the candidate alignment region, to the second sequence with a score within the global alignment score threshold; determining whether the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold; and performing processing based on the determining.

24. The computer program product of claim 23, wherein based on determining that the alignment processing results in no alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises iterating, one or more times, the alignment processing and the determining, wherein at each iteration of the iterating, the method changes the global alignment score threshold to increase a mismatch tolerance of successful alignment of the first sequence to the second sequence, the alignment processing uses the changed global alignment score threshold, and the determining determines whether the alignment processing using the changed global alignment score threshold results in one or more alignments of the first sequence to the second sequence with a score within the changed global alignment score threshold.

25. The computer program product of claim 24, wherein based on determining, at an iteration of the iterating, that the alignment processing performed at the iteration results in an alignment of the first sequence to the second sequence with a score within the global alignment score threshold at the iteration, the method further comprises outputting the alignment.IP-2682-PCT26. The computer program product of claim 25, wherein the method further comprises continuing the iterating to search for a next-best alignment of the first sequence to the second sequence.

27. The computer program product of claim 26, wherein the method further comprises determining whether to halt the iterating at a current iteration, wherein the determining whether to halt is based on (i) the global alignment score threshold at the current iteration and (ii) the global alignment score threshold at a prior iteration, of the iterating, that resulted in an alignment of the first sequence to the second sequence.

28. The computer program product of claim 27, wherein the determining whether to halt comprises determining whether an alignment confidence metric, determined as a function of a difference between the global alignment score threshold at the current iteration and the global alignment score threshold at a prior iteration, corresponds to a threshold alignment confidence.

29. The computer program product of claim 24, further comprising halting the iterating based on the mismatch tolerance increasing to a threshold mismatch tolerance as a result of changing the global alignment score threshold.

30. The computer program product of claim 23, wherein based on determining that the alignment processing results in one or more alignments of the first sequence to the second sequence with a score within the global alignment score threshold, the performing processing comprises outputting at least one alignment of the one or more alignments.

31. The computer program product of claim 23, wherein the attempting to align comprises attempting to aggregate a plurality of partial alignments of the first sequence to the second sequence to produce a full alignment of the first sequence to the second sequence within the global alignment score threshold, each partial alignment of the plurality of partial alignments having the first sequence anchored at a respective candidate alignment region of the collection of candidate alignment regions and partially aligning to the second sequence.IP-2682-PCT32. The computer program product of claim 31 , wherein the attempting to aggregate comprises, for each partial alignment of the plurality of partial alignments: aligning a respective portion of the first sequence to the second sequence in a first direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment, of the partial alignments, in the first direction, or (ii) a first end of the first sequence; and aligning the respective portion of the first sequence to the second sequence in a second direction until reaching (i) a position, in the first sequence, at which to bridge to a next partial alignment in the second direction, or (ii) a second end of the first sequence.

33. The computer program product of claim 23, wherein the attempting to align the first sequence to the second sequence for a candidate alignment region of the collection of candidate alignment regions comprises: first aligning a portion of the first sequence to the second sequence in a first direction for one or more nucleobases until reaching a first end of the first sequence; and next aligning the portion of the first sequence to the second sequence in a second direction for one or more nucleobases until reaching a second end of the first sequence.IP-2682-PCT