Enhanced mapping and alignment of nucleotide reads utilizing an improved haplotype data structure with allele-variant differences
Patent Information
- Application Number
- HK62026125885
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-12-19
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480041058.2 (22) Application Date 2024.12.20 (30) Priority Data 63 / 613574 2023.12.21 US (85) PCT International Application Entering National Phase Date 2025.12.18 (86) PCT International Application Application Data PCT / US2024 / 061580 2024.12.20 (87) PCT International Application Publication Data WO2025 / 137647 EN 2025.06.26 (71) Applicant Inminal Inc. Address California, USA (72) Inventor M. Lule (74) Patent Agency Beijing Panhua Weiye Intellectual Property Agency Co., Ltd. 11280 Patent Attorney Wang Bo (51) Int.Cl. G16B 30 / 10 (2006.01) G16B 20 / 20 (2006.01) (54) Invention Title: Enhanced Narcosome Read Mapping and Alignment Using an Improved Haplotype Data Structure with Allele-Variant Differences (57) Abstract: This disclosure describes methods, non-transitory computer-readable media, and systems for achieving improved mapping and alignment of nucleotide reads with genomic regions of a reference genome. For example, the disclosed system is capable of identifying allele-variant differences between major contiguous sequences and population haplotypes within the corresponding genomic regions for one or more candidate alignments between nucleotide reads from a genomic sample and major contiguous sequences at corresponding genomic regions of a reference genome, thereby generating alignment score adjustments for each population haplotype. To facilitate the disclosed method for improving nucleotide read mapping and alignment, the disclosed system is capable of utilizing a haplotype data structure that includes hierarchically dividing the reference genome into reference bins representing corresponding genomic regions and encoding region-specific allele-variant differences between population haplotypes and major contiguous sequences of the reference genome.Claims 6 pages, Description 43 pages, Drawings 22 pages, CN 121359207 A 2026.01.16 CN 1 21 35 92 07 A 1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: determine a set of candidate alignments between one or more nucleotide reads from a genomic sample and major contiguous sequences at corresponding sets of genomic regions of a reference genome; generate major alignment scores for candidate alignments from the set of candidate alignments; identify one or more allele-variant differences between the major contiguous sequences and one or more population haplotypes corresponding to the candidate alignments; and, based on comparisons of the one or more nucleotide reads with the one or more allele-variant differences, 1. Generate one or more adjusted alignment scores based on the primary alignment score; and select, based on the one or more adjusted alignment scores, one or more nucleotide reads from the set of candidate alignments to align with the primary continuous sequence or with a predicted read from a population haplotype of the one or more population haplotypes. 2. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate substitution alignment scores for the candidate alignments based on the primary alignment score and the one or more adjusted alignment scores; generate additional substitution alignment scores for additional candidate alignments in the set of candidate alignments; and select the predicted read alignment of the one or more nucleotide reads based on comparing the substitution alignment scores with one or more primary alignment scores for one or more candidate alignments with one or more primary continuous sequences, and with the additional substitution alignment scores for the additional candidate alignments of the set of candidate alignments. 3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, for paired-end reads in the one or more nucleotide reads, that a first candidate alignment of a first pair of the paired-end reads with the main continuous sequence is not within a threshold number of nucleotide bases relative to a second candidate alignment of a second pair of the paired-end reads with the main continuous sequence; and, based on the fact that the first candidate alignment is not within the threshold number of nucleotide bases relative to the second candidate alignment, identify a second candidate alignment of the second pair within a predetermined search region relative to the first candidate alignment of the first pair.4. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: identify the one or more allele-variant differences by querying a haplotype data structure, the haplotype data structure comprising a set of bins corresponding to a set of nucleotide reference spans from a reference genome. 5. The system of claim 4, further comprising instructions, which, when executed by the at least one processor, cause the system to: query the haplotype data structure by identifying a reference span in the set of reference spans comprising complete candidate alignments of the one or more nucleotide reads; and identify the one or more allele-variant differences within bins stored in the set of bins corresponding to the identified reference spans. Claims 1 / 6 Page 2 CN 121359207 A 6. The system according to claim 5, further comprising instructions that, when executed by the at least one processor, cause the system to: identify the one or more allele-variant differences stored in the bin by comparing the one or more nucleotide reads with allele-variant differences from one or more local different population haplotype sequences stored in the bin corresponding to the identified reference span. 7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: query a haplotype data structure by identifying a reference span in a set of reference spans that includes a first candidate alignment of the first pair and a second candidate alignment of the second pair, for a first pair and a second pair of paired end reads in the one or more nucleotide reads; for each local different population haplotype encoded by the reference span, generate a first adjusted alignment score for the first pair and a second adjusted alignment score for the second pair based on comparing the first pair and the second pair with the one or more allele-variant differences stored in a set of bins corresponding to the identified reference span; for each local different population haplotype encoded by the reference span, add the first adjusted alignment score of the first pair and the second adjusted alignment score of the second pair; and select, from the set of candidate alignments, a first predicted alignment of the first pair with the major continuous sequence or with the local different population haplotype, and a second predicted alignment of the second pair with the major continuous sequence or with the local different population haplotype.8. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to: generate summation replacement alignment scores for a subset of candidate alignments of the first pair and the second pair based on the primary alignment scores of each local distinct population haplotype encoded by the reference span and the first adjusted alignment scores and the second adjusted alignment scores; generate additional summation replacement alignment scores for an additional subset of candidate alignments of the set of candidate alignments of the first pair and the second pair; and select the first predicted alignment and the second predicted alignment from the set of candidate alignments based on comparing the summation replacement alignment scores with one or more primary alignment scores for one or more candidate alignments of one or more primary consecutive sequences and the additional summation replacement alignment scores for the additional subset of candidate alignments of the set of candidate alignments. 9. The system of claim 1, further comprising instructions, when executed by the at least one processor, to cause the system to: generate the one or more adjusted alignment scores at base positions where there are no allelic-variant differences, without comparing the nucleotide bases of the one or more nucleotide reads with the nucleotide bases of the one or more population haplotypes. 10. The system of claim 1, further comprising instructions, when executed by the at least one processor, to cause the system to: identify the one or more allelic-variant differences by comparing the nucleotide bases within the one or more nucleotide reads with data representing one or more single nucleotide polymorphisms (SNPs) within the one or more population haplotypes corresponding to the corresponding genomic region. 11. The system of claim 1, further comprising instructions, when executed by the at least one processor, to: identify the one or more allelic-variant differences by comparing the one or more nucleotide reads with data representing one or more insertions or deletions (insertions / deletions) within the one or more population haplotypes corresponding to the corresponding genomic region.12. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate at least one adjusted alignment score from the one or more adjusted alignment scores based on the primary alignment score by: determining that the one or more nucleotide reads comprise one or more haplotype variants of a locally different population haplotype, the haplotype variants differing from the primary continuous sequence in the corresponding genomic region; and increasing the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads comprising the one or more haplotype variants. 13. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate at least one adjusted alignment score from the one or more adjusted alignment scores based on the primary alignment score by: determining that the one or more nucleotide reads comprise one or more reference nucleosides of the primary continuous sequence, the reference nucleosides differing from the locally different population haplotypes in the corresponding genomic region; and decreasing the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads comprising one or more reference nucleosides. 14. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a corresponding set of locally distinct population haplotypes corresponding to the corresponding genomic regions of the candidate alignments; select the highest adjusted alignment score from the set of adjusted alignment scores as a replacement alignment score for the candidate alignments; and select the predicted read alignment from the set of candidate alignments based on the replacement alignment score. 15. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a corresponding set of locally distinct population haplotypes corresponding to the corresponding genomic regions of the candidate alignments; convert the set of adjusted alignment scores into a set of alignment likelihood values; adjust the set of alignment likelihood values based on corresponding allele frequencies to generate a set of adjusted alignment likelihood values; convert the sum of the set of adjusted alignment likelihood values into a replacement alignment score for the candidate alignments; and select the predicted read alignment from the set of candidate alignments based on the replacement alignment score.16. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: adjust at least one of the one or more adjusted alignment scores based on the population allele frequency of the population haplotype within the sample population. 17. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate the main alignment score of the candidate alignment based on a given candidate alignment between the one or more nucleotide reads and a modified version of the main continuous sequence, the modified version comprising one or more polybase codes representing one or more single nucleotide polymorphisms (SNPs) or representing one or more insertions or deletions (insertions / deletions). 18. A non-transitory computer-readable medium comprising a haplotype data structure, the haplotype data structure comprising: (a) a base level having a set of base level bins, the base level bins comprising: a set of base level reference spans of a major contiguous sequence of a reference genome, each base level reference span comprising a genomic region of a first length between corresponding genomic coordinates of the reference genome; and variant data of nucleotide variants from corresponding sets of locally distinct population haplotypes, each locally distinct haplotype comprising one or more allele-variant differences of a unique set, the differences being unique relative to other population haplotypes within the genomic region of the corresponding base level reference span; and (b) a contiguous level having a set of higher-level bins, the higher-level bins comprising: A set of higher-level reference spans of the primary continuous sequence, each higher-level reference span including a second-length extended genomic region between corresponding genomic coordinates of the reference genome, the second length being longer than the first length; and a variant data index referencing a combination of variant data from corresponding basic-level bins in the set of basic-level bins. 19. The non-transitory computer-readable medium of claim 18, wherein the variant data of the set of basic-level bins includes data indications of single nucleotide polymorphisms (SNPs) and insertions or deletions (insertions / deletions) at corresponding genomic coordinates of the primary continuous sequence. 20. The non-transitory computer-readable medium of claim 18, wherein the set of basic-level bins includes variant data of nucleotide variants but does not include reference nucleobases of the primary continuous sequence. 21. The non-transitory computer-readable medium of claim 18, wherein population haplotypes having the same nucleotide variant within a given basic-level bin are encoded as a locally distinct population haplotype within the given basic-level bin.22. The non-transitory computer-readable medium of claim 18, wherein each of the set of basic hierarchical bins comprises a matrix, the matrix comprising corresponding variant data, the variant data representing allele-variant differences from local different haplotypes and the variant location of the allele-variant differences. 23. The nontransitory computer-readable medium of claim 18, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: determine a base-level reference span in a set of base-level reference spans including the one or more nucleotide reads from a genome sample and a candidate alignment in a set of candidate alignments to the major continuous sequence; determine one or more alignment score adjustments corresponding to one or more locally distinct haplotypes within a corresponding genomic region of the base-level reference span, based on variant data from base-level bins in the set of base-level bins corresponding to the base-level reference span; and select, from the set of candidate alignments, a predicted alignment of the one or more nucleotide reads to the major continuous sequence or to a population haplotype, based on the one or more alignment score adjustments. 24. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the at least one processor to: adjust the candidate alignment scores based on the one or more alignment scores to generate substitution alignment scores; generate additional substitution alignment scores for additional candidate alignments in the set of candidate alignments; and select the predicted read alignment of the one or more nucleotide reads based on a comparison of the substitution alignment scores with the additional substitution alignment scores. 25. The non-transitory computer-readable medium of claim 18, wherein each corresponding extended genomic region in the set of higher-level reference spans corresponds to a pair of consecutive corresponding genomic regions in the set of base-level reference spans.26. The non-transitory computer-readable medium of claim 18, wherein the continuous hierarchy of the haplotype data structure further comprises a set of offset higher-level bins, the offset higher-level bins comprising: a set of offset higher-level reference spans of the main continuous sequence, each offset higher-level reference span comprising an offset extended genomic region of the second length between corresponding genomic coordinates of the reference genome, wherein the offset extended genomic region corresponds to a pair of consecutive corresponding genomic regions in the set of basic-level reference spans, and wherein the set of offset higher-level reference spans is offset relative to the set of higher-level reference spans by one of the basic-level reference spans in the set of basic-level reference spans. 27. The non-transitory computer-readable medium of claim 26, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: determine, for candidate alignments in a set of candidate alignments between one or more nucleotide reads from a genomic sample and the major contiguous sequence, a higher-level reference span in a set of higher-level reference spans that includes complete candidate alignments of the one or more nucleotide reads; determine, based on variant data indexes of higher-level bins in the set of higher-level bins corresponding to the higher-level reference span, a subset of locally different population haplotypes within a corresponding extended genomic region of the higher-level reference span; and determine, based on variant data of a first base-level bin in the set of base-level bins corresponding to a first corresponding genomic region within the corresponding extended genomic region, a first set of alignment score adjustments for one or more corresponding locally different population haplotypes in the subset of locally different population haplotypes. Based on variant data of the second basic hierarchical binning in the set of basic hierarchical binning corresponding to the second corresponding genomic region within the corresponding extended genomic region, determine a second set of alignment score adjustments for one or more corresponding local different population haplotypes in the subset of local different population haplotypes; and based on a combination of the first set of alignment score adjustments and the second set of alignment score adjustments, select one or more nucleotide reads from the set of candidate alignments to align with the major continuous sequence or with the predicted population haplotype.28. The non-transitory computer-readable medium of claim 26, wherein the haplotype data structure further comprises: at least one additional contiguous level having higher-level reference bins of an additional set, the higher-level reference bins comprising: a set of additional higher-level reference spans of the primary contiguous sequence, each higher-level reference span comprising another extended genomic region of a third length between corresponding genomic coordinates of the reference genome, the third length being longer than the second length; and a variant data index referencing a combination of variant data from a corresponding basic level bin in the set of basic level bins. 29. The non-transitory computer-readable medium of claim 18, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: (Claims 5 / 6, page 6, CN 121359207 A) For candidate alignments in a set of candidate alignments between one or more nucleotide reads from a genomic sample and the major continuous sequence, determine a reference span comprising the complete candidate alignments of the one or more nucleotide reads, the reference span being selected from the lowest level of the haplotype data structure, wherein the one or more nucleotide reads are included in a single reference span in either the set of basic-level reference spans or the set of higher-level reference spans; Based on variant data from one or more bins in the set of basic-level bins corresponding to the reference span, determine one or more alignment score adjustments corresponding to one or more locally distinct haplotypes within the corresponding genomic region of the reference span; and Based on the one or more alignment score adjustments, select from the set of candidate alignments a predicted alignment of the one or more nucleotide reads with the major continuous sequence or with a population haplotype. Claims 6 / 6, page 7, CN 121359207 A Enhanced mapping and alignment of nucleotide reads using improved haplotype data structures with allele-variant differences
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 613,574 (IP-2590-PRV), filed December 21, 2023, entitled “ENHANCED MAPPING AND ALIGNMENT OF NUCLEOTIDE READS UTILIZING AN IMPROVED HAPLOTYPE DATA STRUCTURE WITH ALLELE-VARIANT DIFFERENCES”. The entire contents of the above application are incorporated herein by reference.Background Art
[0003] In recent years, biotechnology companies and research institutions have improved the hardware and software used for sequencing nucleotides and determining variant detection in genomic samples. For example, some existing nucleotide sequencing platforms determine individual nucleotides within a sequence from genomic sample cells using conventional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms can monitor the parallel synthesis of millions to billions of nucleic acid polymers to predict nucleotide detection from a larger dataset of nucleotide detections. For example, cameras in many SBS platforms capture images of irradiated fluorescent tags incorporated into oligonucleotides for the purpose of determining nucleotide detections. After capturing such images, existing sequencing platforms send the nucleotide detection data (or image-based data) to a computing device to apply sequencing data analysis software that determines the nucleotide sequence of a genomic sample or other nucleic acid polymer. For example, such software (i) maps and aligns nucleotide reads determined by the sequencing platform for a sample with (ii) a reference genome containing at least one major contiguous sequence. Based on the differences between the aligned nucleotide reads and the reference genome, existing data analysis software can further utilize variant detectors to identify genotypes and / or variants within the genome sample, such as single nucleotide polymorphisms (SNPs), insertions or deletions (insertions or deletions), or structural variants.
[0004] Despite these recent advances, existing nucleotide sequencing platforms and sequencing data analysis software (collectively and hereinafter referred to as "existing sequencing systems") often utilize reference genomes that do not accurately represent certain populations and result in inaccurate read alignments and erroneous variant detections. For example, some existing sequencing systems use linear reference genomes that supposedly represent common sequences or examples of genes and other nucleotide sequences in an organism. However, for the most common linear human reference genome (GRCh38 from the Genome Reference Consortium), approximately 93% of the primary assembly is based on libraries from only 11 individuals, and 70% of linear human reference genomes are derived from a single individual. Therefore, the linear reference genomes used by many existing systems fail to represent certain populations or common variants.
[0005] To address this lack of genome representation in linear reference genomes, some existing sequencing systems generate or use graph reference genomes. For example, some graph reference genomes consist of a linear reference genome and a graph augmentation, where multinucleotide codes represent SNPs and / or insertions / deletions, while alternative contiguous sequences represent alternative population haplotypes at a given region. In some cases, such graph reference genomes are stacked and the index can span relatively long nucleobase distances (e.g., hundreds to thousands of base pairs in length) of alternative contiguous sequences, thus including redundant reference nucleotides that overlap the same regions.
[0006] While such graph reference genomes better account for the genetics of some populations, the extended representations of existing graph reference genomes are typically large and consume considerable memory and computational resources to implement. In fact, some existing graph reference genomes can contain countless graph augmentations to represent a large number of selectable continuous sequences of SNPs, insertions, deletions, and other variations for various population haplotypes, including some with relatively low allele frequencies (e.g., less than 1% of the population frequency). These seemingly countless alternative pathways can consume unnecessarily much memory and require excessive computational resources to navigate during nucleotide read mapping and alignment of genomic samples. In practice, conventional graph reference genomes often increase the computational time of existing sequencing systems to determine whether to include or exclude matches with graph augmentations when making read alignment inferences. In some cases, an excessive number of candidate alignments can limit the resources available for further alignment procedures on existing sequencing systems, leading to further inaccuracies due to incomplete consideration of potential alignments.
[0007] Furthermore, some existing graph reference genomes contain an extremely large number of allele alternative pathways that are similar to other genomic regions and pathways in the graph reference genome. Therefore, existing sequencing systems may significantly increase the difficulty of accurately predicting degradation from alternative pathways by weakening the uniqueness and usefulness of genomic regions for mapping and alignment, and by increasing confusion between multiple similar genomic regions. For example, some existing sequencing systems utilize ultra-long seed expansion to efficiently locate unique matches for reads within the graph sequence genome. Such excessive seed expansion is less sensitive and may impair alignment accuracy because potential matches are ignored. Furthermore, when dealing with paired-end reads, existing sequencing systems often struggle to locate accurate pairings representing two paired reads within a reasonable distance from each other due to the large amount of overlapping alternative contiguous sequences within the corresponding genomic regions of one or both paired reads.
[0008] In practice, these generic graph reference genomes (with too many alternative pathways representing alternative continuous sequences) frequently lead to incorrect alignments, mismatches, or missed variant detections by existing sequencing systems for large numbers of samples, and increase the chance of alignment mismatches with reads from genomic samples. Because there are multiple similar population haplotypes that perform coordinate transformations on a given genomic region of the main continuous sequence (and because mapping quality (e.g., MAPQ 0) decreases as the number of such population haplotypes in a given genomic region increases), existing sequencing systems are generally unable to expand the candidate population haplotypes in the graph reference genome without slowing down mapping and alignment computations, without reducing mapping quality, and without compromising variant detection accuracy.
[0009] These problems and challenges, as well as others, exist in existing sequencing systems. Summary of the Invention
[0010] This disclosure describes embodiments of methods, non-transitory computer-readable media, and systems that (i) determine a major alignment score for a read aligned to a major continuous sequence, and (ii) adjust the major alignment score based on a comparison between the read and allelic-variant differences representing differences between the major continuous sequence and population haplotypes. Specifically, the disclosed systems can identify candidate alignments between nucleotide reads from a genomic sample and major continuous sequences at corresponding genomic regions of a reference genome. For each candidate alignment, these systems can identify allelic-variant differences between the major continuous sequence and one or more population haplotypes corresponding to a corresponding genomic region of the reference genome. Based on the identified allelic-variant differences, these systems generate adjustments to the corresponding major alignment score. When the alignment score is improved by candidate alignments between nucleotide reads and locally different population haplotypes (such as those represented by one or more allelic-variant differences), these systems can generate alternative alignment scores for such candidate alignments for that locally different population haplotype. Based on the scores of candidate alignments, the disclosed system can identify candidate alignments that exhibit superior primary alignment scores or alternative alignment scores, and determine the predicted read alignments for the corresponding nucleotide reads.
[0011] To facilitate such improved mapping and alignment methods, the disclosed system can utilize a haplotype data structure comprising reference bins that hierarchically divide the regions of a reference genome into corresponding genomic regions (e.g., spans of a set number of nucleotide bases) representing the reference genome. For example, the disclosed haplotype data structure may include a base level having a set of base level bins comprising a corresponding base level reference span of a first length between the corresponding genomic specification 2 / 43 pages 9 CN 121359207 A coordinates of the reference genome, wherein each base level bin comprises variant data of nucleotide variants of locally different population haplotypes within the corresponding genomic region. In addition to such basic-level binning, the disclosed haplotype data structure may also include higher-level binning at consecutive levels, which include corresponding higher-level reference spans that are larger than the basic-level reference span length of the basic-level bins, wherein each higher-level bin includes variant data indexes that reference combinations of variant data from the corresponding basic-level bins in that set of basic-level bins.By utilizing this haplotype data structure, allele-variant differences between major continuous sequences and local population haplotypes can be identified from variant data stored and referenced within one or more bins corresponding to candidate read alignments. These systems can generate alignment scores for nucleotide reads of a genome sample to account for such allele-variant differences and select predicted read alignments based on the corresponding scores of candidate alignments.
[0012] Additional features and advantages of one or more embodiments of this disclosure will be set forth in the following description and will be apparent in part from the description, or may be learned through practice of such exemplary embodiments. Brief Description of the Drawings
[0013] Detailed description refers to the accompanying drawings, which are briefly described below.
[0014] Figure 1 illustrates the environment in which a read alignment adjustment system can operate according to one or more embodiments of this disclosure.
[0015] Figure 2 shows an overview of an existing sequencing system that maps and aligns nucleotide reads from a genome sample using a conventional graph reference genome with graph augmentations representing population haplotypes.
[0016] Figure 3A illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system determines candidate alignments between nucleotide reads and major continuous sequences and generates major alignment scores for the corresponding candidate alignments.
[0017] Figure 3B illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system generates adjusted alignment scores based on allele-variant differences between major continuous sequences and one or more population haplotypes.
[0018] Figure 4 illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system determines alignment score adjustments for candidate alignments of paired-terminal nucleotide reads.
[0019] Figure 5A further illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system determines candidate alignments of nucleotide reads and generates major alignment scores for the candidate alignments.
[0020] Figure 5B illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system generates substitution alignment scores for a given candidate alignment.
[0021] Figure 6 illustrates, according to one or more embodiments of the present disclosure, a read alignment adjustment system determining replacement alignment scores for candidate alignments from a set of adjusted alignment scores.
[0022] Figure 7 illustrates a set of basic hierarchical binning of a haplotype data structure according to one or more embodiments of the present disclosure.
[0023] Figure 8 illustrates the basic hierarchical binning and consecutive higher-level binning of a haplotype data structure according to one or more embodiments of the present disclosure.
[0024] Figures 9A and 9B illustrate experimental results for a set of population haplotype coding variant data using a haplotype data structure according to one or more embodiments of the present disclosure.
[0025] Figure 10 illustrates the alignment score adjustment of a read alignment adjustment system that uses a haplotype data structure to determine candidate alignments of nucleotide reads according to one or more embodiments of the present disclosure.
[0026] Figure 11 illustrates the alignment score adjustment and summation of a read alignment adjustment system that uses a haplotype data structure to determine candidate alignments of paired-terminal nucleotide reads according to one or more embodiments of the present disclosure.
[0027] Figure 12 illustrates an example implementation of the alignment score adjustment of a read alignment adjustment system that uses a haplotype data structure to determine candidate splicing alignments of transcriptome reads according to one or more embodiments of the present disclosure.
[0028] Figures 13A and 13B illustrate comparative experimental results of variant detection from nucleotide reads (i) mapped and aligned to a reference genome using existing sequencing systems, and (ii) mapped and aligned to a reference genome using a read alignment adjustment system and haplotype data structure according to one or more embodiments of the present disclosure.
[0029] Figures 14A and 14B illustrate example implementations of determining alignment scores for candidate nucleotide read alignments, which are (i) mapped and aligned to a reference genome using an existing sequencing system, and (ii) mapped and aligned to a reference genome using a read alignment adjustment system according to one or more embodiments of the present disclosure.
[0030] Figure 15 illustrates a flowchart of a series of actions for selecting predicted read alignments for one or more nucleotide reads from a genome sample according to one or more embodiments of the present disclosure.
[0031] Figure 16 illustrates a flowchart of a series of actions for selecting predicted read alignments for one or more nucleotide reads from a genome sample using a haplotype data structure according to one or more embodiments of the present disclosure.
[0032] Figure 17 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure. Detailed Description
[0033] This disclosure describes an implementation of a read alignment adjustment system that utilizes a haplotype data structure encoding allele-variant differences to determine the alignment of a nucleotide read from a genomic sample with a major contiguous sequence of a reference genome or with a population haplotype represented by allele-variant differences in the data structure. Specifically, the read alignment adjustment system may utilize a haplotype data structure containing graph augmentations that encode population variations in corresponding genomic regions, thereby allowing for scoring of candidate alignments without directly aligning the read with a selected contiguous sequence.For example, for one or more nucleotide reads from a genome sample, the read alignment adjustment system can identify a set of candidate read alignments between these nucleotide reads and major contiguous sequences at a corresponding set of genomic regions of a reference genome, and generate a major alignment score for each candidate alignment. For each candidate read alignment, the read alignment adjustment system can determine an alignment score adjustment to account for allele-variant differences in each local haplotype within the corresponding genomic region. Furthermore, the read alignment adjustment system can adjust the alignment score of the candidate alignments based on the population frequency of the corresponding local haplotypes.
[0034] As mentioned above, implementations of the read alignment adjustment system can utilize haplotype data structures encoding population variations within corresponding genomic regions of a reference genome to facilitate mapping and alignment according to the methods described herein. For example, the read alignment adjustment system can implement a haplotype data structure that includes reference bins that hierarchically divide the reference genome into corresponding genomic regions (e.g., nucleotide spans) representing the reference genome, and encodes allele-variant differences in local population haplotypes within the corresponding genomic regions.
[0035] To facilitate efficient alignment and scoring of both major continuous sequences and locally distinct population haplotypes, the disclosed haplotype data structure may include a base level having a set of base level bins containing a first-length base level reference span between corresponding genomic coordinates of a reference genome. Each base level bin contains variant data of nucleotide variants of locally distinct population haplotypes within a corresponding genomic region. In some cases, each base level bin has a matrix including the corresponding variant data representing allele-variant differences from locally distinct haplotypes and the variant locations of these allele-variant differences.
[0036] In addition to such basic hierarchical binning, the disclosed haplotype data structure may also include sequential hierarchical higher-level binning, which includes a corresponding higher-level reference span that is larger than the basic hierarchical reference span length of the basic binning. Each higher-level binning includes variant data indexes that reference combinations of variant data from the corresponding basic binning in that set of basic binning. As further described below, in some cases, each higher-level binning includes “offset” binning that covers a different nucleotide span than the “non-offset” binning, such that each combination of two sequential binnings from the next level is represented by either a non-offset binning or an offset binning. To query the span of the reference genome, the read alignment adjustment system accesses the lowest-level binning containing the entire candidate alignment of nucleotide reads and the non-offset binning below that lowest-level binning.
[0037] Therefore, in some embodiments, the read alignment adjustment system utilizes this haplotype data structure to identify allelic-variant differences between major continuous sequences and locally distinct population haplotypes. By encoding such locally distinct population haplotypes in variant data stored and referenced within one or more bins corresponding to candidate read alignments, the read alignment adjustment system performs one or more of the disclosed methods for nucleotide read mapping and alignment. In one or more embodiments, for example, the read alignment adjustment system may identify a bin of the haplotype data structure corresponding to a reference span that includes each nucleobase position in candidate alignments of nucleotide reads or multiple linked reads from a genomic sample. Based on the variant data stored or indicated within the selected bin, the read alignment adjustment system may identify allelic-variant differences of locally distinct population haplotypes within the corresponding reference span to determine alignment score adjustments for candidate alignments, thereby facilitating the selection of predicted read alignments for the corresponding nucleotide reads. For example, when improving alignment scores between nucleotide reads and candidate alignments of different local population haplotypes (such as those represented by one or more allele-variant differences), the read alignment adjustment system generates alternative alignment scores for such candidate alignments for those different local population haplotypes.
[0038] As proposed above, the read alignment adjustment system offers several technical advantages, benefits, and / or improvements compared to existing sequencing systems, including systems utilizing conventional graph reference genomes augmented with selectable continuous sequences and other sequencing data analysis software. In some implementations, for example, the read alignment adjustment system can accurately predict read alignments while improving computational speed and memory utilization relative to existing sequencing systems. As noted above, existing sequencing systems use graph reference genomes with general graph augmentations, which include a large number of redundant selectable continuous sequences that consume memory due to repetitive sequences from overlapping portions of the selectable continuous sequences and slow down computer processing speed due to scoring of alignments between reads and such overlapping portions of selectable continuous sequences. Compared to existing systems of this kind, the disclosed read alignment adjustment system accelerates the determination of alignment scores by at least the following means: (i) adjusting the alignment scores of candidate alignments between nucleotide reads and major continuous sequences based on the differences between population haplotypes and major continuous sequences, and (ii) providing a haplotype data structure representing allele-variant differences in genomic regions.
[0039] For example, by determining the alignment score adjustment of different population haplotypes locally based on allele-variant differences between major continuous sequences and each different local haplotype, the disclosed method can accurately determine the predicted read alignments of nucleotide reads, with improved computational speed and less memory footprint compared to graph genomes of existing sequencing systems.Specifically, as mentioned above, existing sequencing systems typically determine predicted read alignments by attempting to align and score nucleotide reads against a robust map genome augmented by alternative contiguous sequences. Instead of determining the alignment score of candidate contiguous sequences that undergo coordinate transformation of the same given primary contiguous sequence (and often re-scoring alignments across spans of the same sequence), the read alignment adjustment system accelerates alignment scoring by first identifying candidate alignments against the primary contiguous sequence and then adjusting the alignment scores of these candidate alignments based on the differences between the primary contiguous sequence and candidate contiguous sequences representing the population haplotype (encoding allele-variant differences). Therefore, the disclosed read alignment adjustment system improves the computational speed of mapping and aligning nucleotide reads of a genome sample with a reference genome representing candidate population haplotypes.
[0040] In addition to increased computational speed and reduced memory footprint, the read alignment adjustment system provides accurate and comprehensive population haplotype information in a scalable manner by utilizing various embodiments of the haplotype data structure described herein. As disclosed herein, for example, the haplotype data structure can be readily extended to include variation and frequency data for virtually any number of population haplotypes because the amount of data required to encode population variations of locally different haplotypes in the corresponding genomic region is minimal, and there is no need to encode nuclei at base positions where there is no allele-variant difference between the corresponding haplotype and the major contiguous sequence. As depicted and described in this disclosure, for example, the read alignment adjustment system can increase the number of population haplotypes represented in the disclosed haplotype data structure from 32 population haplotypes to 128 (or more) population haplotypes without compromising mapping accuracy or variant detection accuracy.
[0041] Furthermore, by initially mapping nucleotide reads to a primary contiguous sequence, rather than utilizing a graph reference genome containing an additional large number of contiguous sequences for selection, the read alignment adjustment system achieves an improved mapping and alignment method. In some implementations, for example, haplotype nucleotide bases are encoded in the primary contiguous sequence (e.g., via polybase encoding) to improve seed mapping sensitivity in hard-to-map regions. Additionally, when performing mapping and alignment of paired reads, rescue scans can be performed as needed by generating candidate alignments using corresponding pairs of primary contiguous sequences for paired reads. Furthermore, for such paired candidate pair alignments, the haplotype data structure can be queried using a reference span covering both paired alignments, and the corresponding alignment scores can be jointly adjusted to further improve the accuracy of predicted read alignments.
[0042] As proposed in the foregoing discussion, this disclosure uses various terms to describe the features and benefits of the read alignment adjustment system and the improved haplotype data structure. Additional details regarding the meaning of these terms as used in this disclosure are provided below.As used in this disclosure, for example, the term "genomic sample" (or simply "sample") refers to a specimen, culture, etc., suspected of containing a target nucleic acid. In some embodiments, the genomic sample contains DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acid as the target. The genomic sample may also contain any biological, clinical, surgical, agricultural, atmospheric, or water-based specimen containing one or more nucleic acids. The genomic sample also includes any nucleic acid sample isolated or extracted from an organism, such as genomic DNA, freshly frozen, or formalin-fixed paraffin-embedded nucleic acid specimens. Thus, in some cases, the genomic sample comprises a complete genome isolated or extracted from an organism (e.g., entirely or partially by a kit) and ready for sequencing or determination in a sequencing device. Genome samples may be derived from: a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, (matched) nucleic acid samples from a single individual (such as tumor samples and normal tissue samples), or samples from a single source containing two different forms of genetic material (such as maternal DNA and fetal DNA obtained from a maternal subject), or samples containing contaminating bacterial DNA in the presence of plant or animal DNA. In some embodiments, the source of nucleic acid material may include nucleic acids obtained from newborns, such as nucleic acids commonly used for newborn screening.
[0043] Genome samples may include high molecular weight substances, such as genomic DNA (gDNA). Genome samples may include low molecular weight substances, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another specific embodiment, low molecular weight substances include enzymatically fragmented or mechanically fragmented DNA. Genome samples may include cell-free circulating DNA. In some specific embodiments, the genome sample may include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microanatomy, surgical resection, and other clinical or laboratory samples. In some implementations, the genome sample may be an epidemiological sample, an agricultural sample, a forensic sample, or a pathogenic sample. In some implementations, the genome sample may include nucleic acid molecules obtained from animals (such as humans or mammalian sources). In another implementation, the sample may include nucleic acid molecules obtained from non-mammal sources (such as plants, bacteria, viruses, or fungi). In some implementations, the source of the nucleic acid molecules may be archived or extinct samples or species.
[0044] Additionally, as used herein, the term "nucleotide read" (or simply "read") refers to a sequence of one or more nucleotide bases (or nucleobase pairs) inferred or predicted from a complete or partial sample genome sequence (e.g., a sample genome sequence, complementary DNA). Such a sample nucleotide sequence may take the form of a sample genome sequence from genomic DNA (gDNA), a transcriptome sequence from complementary DNA (cDNA), a transcriptome sequence from RNA, or other nucleotide sequences. Specifically, a nucleotide read includes a nucleobase detection sequence of a nucleotide fragment (or a group of monoclonal nucleotide fragments) determined or predicted based on a sequencing library corresponding to the genome sample. For example, in some embodiments, the sequencing device determines nucleotide reads by generating nucleobase detection through nanopores in a nucleotide sample slide, via fluorescent tagging, or based on pores in a flow cell. In some cases, a nucleotide read may refer to a specific type of read, such as a nucleotide read synthesized from a sample library fragment shorter than a threshold number of nucleotides (e.g., an SBS read). In these or other cases, another type of nucleotide read may refer to (i) an assembled nucleotide read that has been assembled from shorter nucleotide reads to form a continuous sequence that meets the threshold number of nucleotides (e.g., an assembled nucleotide read), (ii) a cycle shared sequencing (CCS) read that meets the threshold number of nucleotides, or (iii) a nanopore-length read that meets the threshold number of nucleotides.
[0045] Accordingly, as used herein, the term “genomic read” refers to a nucleotide read that represents an inferred sequence of nucleotides (or nucleobase pairs) derived from genomic DNA (gDNA) extracted from a sample. For example, a genomic read includes a read containing gDNA that (i) was extracted or derived from gDNA extracted from a sample, and (ii) is part of a sample library fragment corresponding to the sample.
[0046] Conversely, as used herein, the term “transcriptome read” refers to a nucleotide read that represents a deduced sequence of nucleobases (or nucleobase pairs) complementary to or representing RNA extracted from a sample. For example, a transcriptome read includes a read containing cDNA (i) synthesized from or derived from RNA extracted from a sample by single-stranded messenger RNA (mRNA) or microRNA (miRNA), and (ii) is part of a sample library fragment corresponding to the sample. As another example, a transcriptome read includes a read containing RNA (e.g., mRNA, miRNA, transfer RNA (tRNA)) that (i) is extracted from or derived from RNA extracted from a sample, and (ii) is part of a sample library fragment corresponding to the sample.
[0047] As further used herein, the term "genomic coordinate" (or sometimes simply "coordinate") refers to a specific location or orientation of a nucleobase within a genome (e.g., the genome of an organism or a reference genome). In some cases, genomic coordinates include an identifier of a specific chromosome of the genome and an identifier of the nucleobase location within that chromosome. For example, one or more genomic coordinates may include the number, name, or other identifier of an autosomal chromosome or sex chromosome (e.g., chr1 or chrX), and one or more specific locations, such as the numbered position following the chromosome identifier (e.g., chr1: 1234570 or chr1: 1234570–1234870). In some cases, genomic coordinates refer to genomic coordinates on a sex chromosome (e.g., chrX or chrY). Thus, this read alignment adjustment system can determine the genotype probability of genotype detection (e.g., variant detection) at a given genomic coordinate on a sex chromosome. Additionally, in some embodiments, genomic coordinates refer to the origin of the reference genome (e.g., mt of the mitochondrial DNA reference genome or SARS-CoV-2 of the SARS-CoV-2 virus reference genome) and the orientation of the nucleotide bases of the reference genome (e.g., mt: 16568 or SARS-CoV-2: 29001). In contrast, in some cases, genomic coordinates refer to the location of nucleotide bases in the reference genome without relating to chromosome or origin (e.g., 29727).
[0048] As used herein, a “genomic region” refers to the range of genomic coordinates. Similar to genomic coordinates, in some embodiments, genomic regions can be identified by a chromosome identifier and one or more specific locations (such as numbered locations following the chromosome identifier, e.g., chrl: 1234570–1234870). In various embodiments, genomic coordinates include locations within the reference genome. In some cases, genomic coordinates are specific to a particular reference genome. Relatedly, as used herein, the term “reference span” refers to the span of nucleotide positions within a linear reference genome. In other words, the reference span, as described on page 7 / 43 of the specification (CN 121359207 A), includes a nucleotide span between two corresponding genomic coordinates of a linear reference genome.
[0049] As stated above, genomic coordinates include orientation within the reference genome. Such orientation may be within a specific reference genome. As used herein, the term "reference genome" refers to a digital nucleic acid sequence assembled as a representative example (or multiple representative examples) of genes and other genetic sequences of an organism. Regardless of sequence length, in some cases, a reference genome represents a set of examples of genes or a set of nucleic acid sequences in a digital nucleic acid sequence that scientists have determined to represent an organism of a particular species.For example, a linear human reference genome may be GRCh38 or another version of a reference genome from a genome reference consortium. As mentioned above, in some cases, the reference genome includes polybase codes. Alternatively, the reference genome may include a graph reference genome containing both a linear reference genome and pathways representing nucleic acid sequences from ancestral haplotypes, such as the Illumina DRAGEN graph reference genome hgl9.
[0050] As used herein, the term “major contiguous sequence” (or simply “major contig”) refers to a contiguous sequence representing a reference haplotype of the reference genome. In some embodiments, the major contiguous sequence digitally represents the reference haplotype of the reference genome, but may include additional information from the major assembly of the linear reference genome, such as indications of population variants in certain genomic regions, to help identify candidate alignments of nucleotide reads.
[0051] In contrast, the term “alternative contiguous sequence” (or simply “alternative contig”) refers to a contiguous sequence representing an alternative population haplotype at specific genomic coordinates of the reference genome. For example, in some sequencing systems, a graph reference genome includes alternating contiguous sequences of genomic coordinates mapped to the primary assembly of the linear reference genome. In some cases, the hash table of the graph reference genome includes identifiers that associate alternative continuous sequences representing population haplotypes with the linear reference genome at genomic coordinates. Crucially, as explained and depicted in this disclosure, the disclosed haplotype data structure or corresponding reference genome does not directly include alternative continuous sequences, but rather encodes allelic-variant differences between a predominant continuous sequence and locally distinct haplotypes within a given genomic region.
[0052] Accordingly, as used herein, the term "allelic-variant difference" refers to a difference between corresponding nucleobases of two or more given nucleotide sequences. In some cases, for example, allelic-variant differences are differences between a predominant continuous sequence and at least one population haplotype (e.g., represented by an alternative continuous sequence). In some embodiments, for example, allelic-variant differences within a given genomic region may include single nucleotide variants, multibase differences, and / or insertions and deletions (insertions and deletions) relative to the population haplotype of the predominant continuous sequence. Additionally, allelic-variant differences may refer to differences between a first population haplotype and a second population haplotype.
[0053] As used herein, the term "haplotype data structure" refers to a data structure that encodes variant data of population haplotypes of a sample organism. Specifically, the haplotype data structures disclosed herein include a set of bins that hierarchically divide different genomic regions of a reference genome into bins covering corresponding spans of a linear reference genome (e.g., represented by major continuous sequences).Furthermore, as used herein, the term "basic-level binning" refers to a binning that corresponds to a genomic region of a reference genome and encodes variant data of population haplotypes having allelic-variant differences within that corresponding genomic region. For example, in some cases, basic-level binning includes region-specific data structures, such as matrices, that encode allelic-variant differences of locally distinct population haplotypes from a given genomic region. Relatedly, as used herein, the term "basic-level reference span" refers to the span of nuclei in the genomic region corresponding to a given basic-level binning. As shown below, the basic-level reference span represents or covers multiple nuclei in a given genomic region of the reference genome, but does not need to represent every nuclei in the given genomic region.
[0054] Furthermore, as used herein, the term "higher-level binning" refers to a binning that corresponds to an extended genomic region having a greater length than the corresponding basic-level binning of the haplotype data structure. As shown below, higher-level binning may include a variant data index referencing combinations of variant data from the corresponding basic-level binning. Additionally or alternatively, in some cases, as described on page 8 / 43 of this specification (CN 121359207 A), higher-level bins may include variant data indexes that reference other variant data indexes within the corresponding higher-level bin at the next level below the corresponding higher-level bin, as illustrated below in conjunction with Figure 12. Therefore, the higher-level bin itself does not need to contain variant data, but rather contains indexes that identify variant data encoding allele-variant differences. Relatedly, as used herein, the term "higher-level reference span" refers to the nucleotide span of the genomic region corresponding to a given higher-level bin. Additionally, as used herein, the term "variant data index" refers to coding data within a given higher-level bin that references variant data within the base-level bin corresponding to the given higher-level bin (e.g., as illustrated below in conjunction with Figures 8 and 12).
[0055] Additionally, as used herein, the terms "locally distinct population haplotype" or "locally distinct haplotype" refer to a haplotype that contains at least one set of allele-variant differences, wherein this set is unique relative to other haplotypes within a corresponding genomic region of a reference genome. For example, according to the disclosed embodiments, each bin of the haplotype data structure encodes one or more locally distinct haplotypes that have a unique set of one or more allele-variant differences relative to other population haplotypes within each corresponding genomic region (e.g., as described below in conjunction with Figure 8). Furthermore, in some embodiments, due to the complete overlap of variants within genomic regions, a given set of one or more allele-variant differences within a genomic region corresponding to a candidate read alignment may represent multiple haplotypes.Therefore, in some cases, multiple haplotypes consisting of the same nucleotide bases within a given genomic region can be represented by a single locally distinct haplotype.
[0056] Additionally, as used herein, the term “alignment score” refers to a numerical score, measure, or other quantitative measurement evaluating the accuracy of an alignment between one or more nucleotide reads or fragments of nucleotide reads and another nucleotide sequence from a reference genome. Specifically, alignment scores include measures indicating the degree to which the nucleotide bases of one or more nucleotide reads (or fragments thereof) match or are similar to a reference sequence or alternating contiguous sequence from a reference genome. In some specific implementations, alignment scores take the form of a Smith-Waterman score or a variant or version of a Smith-Waterman score for local alignment, such as the various settings or configurations used by DRAGEN for the Smith-Waterman score by Illumina, Inc.
[0057] Relatedly, as used herein, the term “major alignment score” refers to the alignment score generated for candidate alignments between nucleotide reads and major contiguous sequences. Therefore, in some cases, the primary alignment score does not take into account population haplotypes within the genomic region corresponding to the candidate alignment. Additionally, as used herein, the term "adjusted alignment score" refers to the alignment score for a given candidate alignment of a nucleotide read to a reference genome, which has been adjusted to account for allele-variant differences between the population haplotype and the major contiguous sequence within the genomic region of the given candidate alignment (e.g., as described below in conjunction with Figure 3B).
[0058] As further used herein, the term "replacement alignment score" refers to the alignment score for a given candidate alignment of a nucleotide read to a reference genome, which is generated to replace the primary alignment score of the given candidate alignment, based on one or more adjusted alignment scores determined for the given candidate alignment while taking into account one or more population haplotypes within the genomic region of the given candidate alignment (e.g., as described below in conjunction with Figure 6). For example, when a candidate alignment between a nucleotide read and a local population haplotype (such as one represented by one or more allele-variant differences) improves the primary alignment score, the read alignment adjustment system can generate a replacement alignment score for such a candidate alignment with that local population haplotype, and rely on this replacement alignment score (rather than the primary alignment score) to determine whether the candidate alignment exhibits the highest relative alignment score and is eligible as a predicted read alignment for the nucleotide read. As used herein, the terms “replacement alignment score” and “final adjusted alignment score” are used interchangeably, as in the description below with reference to Figure 14B.Specification 9 / 43 pages 16 CN 121359207 A
[0059] Accordingly, as used herein, the term “mapping quality score” refers to a measure or other measurement that quantifies the quality or certainty of a nucleotide read (or other nucleotide sequence or subsequence) aligned with a reference genome. In some embodiments, for example, the mapping quality score includes the mapping quality (MAPQ) score of nucleotides detected at genomic coordinates, where the MAPQ score represents –10 log10 Pr{mapping position error}, rounded to the nearest integer. As an alternative to mean or median mapping quality, in some implementations, the mapping quality score includes the complete distribution of mapping quality of all nucleotide reads aligned with a reference genome at genomic coordinates.
[0060] As further used herein, the term “genotype detection” refers to identifying or predicting a specific genotype of a genomic sample or sample nucleotide sequence at a genomic locus. Specifically, genotype detection may include predicting a specific genotype of a genomic sample relative to a reference genome or reference sequence at genomic coordinates or genomic regions. For example, in some cases, genotype detection includes determining or predicting at genomic coordinates that a genomic sample includes both nucleobases and complementary nucleobases, which are homozygous or heterozygous with respect to a reference base or variant (e.g., a homozygous reference base is represented as 0|0, or a heterozygous variant on a particular strand is represented as 0|1). Thus, genotype detection may include predicting variants or reference bases of one or more alleles in a genomic sample and indicating conjugation with respect to the variant or reference base. Genotype detection is typically determined at genomic coordinates or genomic regions where SNPs, insertions, deletions, or other variants have been identified for an organismal population.
[0061] As further used herein, the term “nucleobase detection” (or simply “base detection”) refers to a specific nucleobase (or nucleobase pair) that determines or predicts the genomic coordinates of an oligonucleotide (e.g., a nucleotide read) or a sample genome during a sequencing cycle. Specifically, nucleobase detection can indicate the determination or prediction of the nucleobase type within an oligonucleotide incorporated into a nucleotide sample slide (e.g., nucleobase detection based on a read). In some cases, for a nucleotide read, nucleobase detection includes determining or predicting nucleobases based on the intensity values produced by fluorescently tagged nucleotides of oligonucleotides added to the nucleotide sample slide (e.g., in a cluster of flow cells). As mentioned above, a single nucleobase detection can be adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U).
[0062] As used herein, the term "variant" refers to one or more nucleobases that are not aligned with, differ from, or are altered from the corresponding nucleobases (or nucleobases) in a reference sequence or reference genome.For example, variants include SNPs, insertions, deletions, or structural variants, which indicate nucleosides in the sample nucleotide sequence that differ from nucleosides at corresponding genomic coordinates in the reference sequence.
[0063] Along these lines, “variant detection” (or “variant nucleobase detection”) refers to the detection of nucleosides containing mutations or variants relative to a reference at specific genomic coordinates or regions. Specifically, variant detection includes identifying or predicting that a genomic sample contains a specific nucleobase (or nucleobase sequence) at genomic coordinates or regions that differs from a reference nucleobase (or reference nucleobase sequence) at the same genomic coordinates or regions within the reference genome. Conversely, “non-variant detection” (or “non-variant nucleobase detection” or “reference detection”) refers to the detection of nucleosides containing non-variant or reference nucleosides relative to a reference at genomic coordinates or regions. Specifically, non-variant or reference detection includes identifying or predicting that a genomic sample contains a specific nucleobase (or nucleobase sequence) at genomic coordinates or regions that matches a reference nucleobase (or reference nucleobase sequence) at the same genomic coordinates or regions within the reference genome.
[0064] In one or more embodiments, the read alignment adjustment system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term "sequencing data file" refers to a digital file that includes gene sequencing information about genotype detections or nucleotide reads generated by one or more genome sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, etc.
[0065] Furthermore, in one or more embodiments, the read alignment adjustment system identifies or stores one or more sequencing data files, including alignment data files containing information from read processing and mapping processes. As used herein, the term "alignment data file" refers to a digital file indicating the mapping and alignment information of nucleotide reads of a sample nucleotide sequence. For example, the alignment data file may include a binary alignment map (BAM) file, a reference-oriented compressed alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.
[0066] The following paragraphs describe the read alignment adjustment system with reference to the illustrative drawings depicting exemplary embodiments and implementations. For example, FIG1 shows a schematic diagram of a computing system 100, in which a read alignment adjustment system 106 operates according to one or more embodiments. As shown, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server devices 110, and a client device 114. As shown in FIG1, the sequencing device 102, the local device 108, the server device 110, and the client device 114 can communicate with each other via a network 118.Network 118 includes any suitable network through which computing devices can communicate. An example network is discussed in more detail below with respect to Figure 17. While Figure 1 illustrates one embodiment of the read alignment adjustment system 106, alternative embodiments and configurations are described below.
[0067] As indicated by Figure 1, sequencing device 102 includes a computing device or sequencing device system 104 for sequencing a genomic sample or other nucleic acid polymer. In some embodiments, by executing the processor-based sequencing device system 104, sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from a genomic sample to generate nucleotide reads or other data directly or indirectly on sequencing device 102 using computer-implemented methods and systems. More specifically, sequencing device 102 receives a nucleotide sample slide (e.g., a flow cell) comprising nucleotide fragments extracted from a sample, then copies and determines the nucleobase sequence of such extracted nucleotide fragments.
[0068] In one or more embodiments, sequencing device 102 utilizes sequencing-by-synthesis (SBS) technology to sequence nucleotide fragments into nucleotide reads and determine the nucleobase detection of these nucleotide reads. As a supplement or alternative to communication across network 118, in some embodiments, sequencing device 102 bypasses network 118 and communicates directly with local device 108 or client device 114. By executing sequencing device system 104, sequencing device 102 can also store nucleobase detections as part of base detection data formatted as a binary base detection (BCL) file and send the BCL file to local device 108 and / or server device 110.
[0069] As further indicated in FIG1, local device 108 is located in or near the same physical orientation as sequencing device 102. In fact, in some embodiments, local device 108 and sequencing device 102 are integrated into the same computing device. Local device 108 can run read alignment adjustment system 106 to generate, receive, analyze, store, and transmit digital data, such as determining variant detections by receiving base detection data or based on analysis of such base detection data. As shown in Figure 1, sequencing device 102 can send (and local device 108 can receive) base detection data generated during sequencing runs on sequencing device 102. By executing software in the form of read alignment adjustment system 106, local device 108 can align nucleotide reads with a reference genome using haplotype data structure 112 and identify genetic variants based on the aligned nucleotide reads. Local device 108 can also communicate with client device 114. Specifically, local device 108 can send data to client device 114, including binary alignment map (BAM) files, variant detection format (VCF) files, or other information indicating nucleobase detection, sequencing metrics, error data, or other metrics.
[0070] As further indicated in FIG1, server device 110 is located remotely from local device 108 and sequencing device 102. Similar to local device 108, in some embodiments, server device 110 includes (or is otherwise accessible or implementable) a version of read alignment adjustment system 106. Therefore, server device 110 can generate, receive, analyze, store, and transmit digital data, such as determining variant detection by receiving base detection data or by analyzing such base detection data. As described above, sequencing device 102 can send (and server device 110 can receive) base detection data from sequencing device 102. Server device 110 can also communicate with client device 114. Specifically, server device 110 can send data, including BAM files, VCF files, or other sequencing-related information, to client device 114.
[0071] In some embodiments, server device 110 includes a distributed set of servers, wherein server device 110 includes a number of server devices distributed across network 118 and located in the same or different physical locations. Furthermore, server device 110 may include a content server, application server, communication server, web hosting server, or another type of server.
[0072] As noted above, as part of server device 110 or local device 108, read alignment adjustment system 106 may generate, encode, and / or implement haplotype data structure 112 to determine the alignment of nucleotide reads from a genomic sample with a reference genome. For example, read alignment adjustment system 106 may identify candidate alignments of one or more nucleotide reads with a major contiguous sequence, generate a major alignment score for the candidate alignments, and adjust the alignment score based on population variant data indicated in haplotype data structure 112, as described in more detail below in conjunction with the accompanying drawings.
[0073] As further shown and indicated in FIG1, client device 114 may generate, store, receive, and transmit digital data by executing sequencing application 116. Specifically, client device 114 may receive sequencing data from local device 108 or receive detection files (e.g., BCL) and sequencing metrics from sequencing device 102. Furthermore, client device 114 may communicate with local device 108 or server device 110 to receive VCFs, which include genotype or variant detection and / or other metrics (such as base detection quality metrics or filter-through metrics). Client device 114 may accordingly present or display information related to variant detection or other genotype detection to the user associated with client device 114 within the graphical user interface of sequencing application 116.For example, client device 114 may present genotype detection, variant detection, and / or sequencing metrics for a sequenced genome sample within the graphical user interface of sequencing application 116.
[0074] Although Figure 1 depicts client device 114 as a desktop or laptop computer, client device 114 may include various types of client devices. For example, in some embodiments, client device 114 includes non-mobile devices (such as desktop computers or servers) or other types of client devices. In other embodiments, client device 114 includes mobile devices such as laptop computers, tablet computers, mobile phones, or smartphones. Additional details regarding client device 114 are discussed below with respect to Figure 17.
[0075] As further illustrated in Figure 1, client device 114 includes sequencing application 116. Sequencing application 116 may be a web application or a native application (e.g., a mobile application, a desktop application) stored and executed on client device 114. The sequencing application 116 may include instructions (when executed) to cause the client device 114 to receive data from the read alignment adjustment system 106 and to present at the client device 114 to display base detection data or data from the VCF. Furthermore, the sequencing application 116 may instruct the client device 114 to display a summary for multiple sequencing runs.
[0076] As further shown in FIG1, a version of the read alignment adjustment system 106 may (e.g., wholly or partially) be located on and / or implemented on the client device 114 or the sequencing device 102. In other embodiments, the read alignment adjustment system 106 is implemented by one or more other components of the computing system 100, such as the local device 108. Specifically, the read alignment adjustment system 106 can be implemented in a variety of different ways on the sequencing device 102, the local device 108, the server device 110, and the client device 114. For example, read alignment adjustment system 106 can be downloaded from server device 110 to read alignment adjustment system 106 and / or local device 108, wherein all or part of the functionality of read alignment adjustment system 106 is executed at each corresponding device within computing system 100.
[0077] As previously mentioned, in some embodiments, read alignment adjustment system 106 implements and / or utilizes an improved haplotype data structure that encodes allele-variant differences across a linear reference genome between a major continuous sequence and a population haplotype. In contrast, it has also been mentioned that some existing sequencing systems utilize graph reference genomes, which include both a linear reference genome and graph augmentations representing selectable continuous sequences with SNPs and / or insertions and deletions (see specification page 12 / 43, 19 CN 121359207 A).To illustrate this, Figure 2 depicts an example of an existing sequencing system aligning nucleotide reads of a genome sample with a graphical reference genome 212 to determine the nucleobase detection of the genome sample based on the aligned nucleotide reads.
[0078] As shown in Figure 2, for example, the depicted sequencing system identifies or receives nucleotide reads 218 of a genome sample and aligns nucleotide reads 218 with different sequences in the graphical reference genome 212. As may sometimes occur with the graphical reference genome, the graphical reference genome 212 includes a linear reference genome containing reference sequences 216a, 216b, 216c to 216n, which are augmented by various alternative sequential sequences 214a, 214b, 214c to 214n representing various population haplotypes associated with the linear reference genome. As indicated by ellipses (or dots) in Figure 2, the graphical reference genome 212 may contain more reference sequences and / or more alternative sequential sequences than depicted in Figure 2. Although Figure 2 depicts the alternative contiguous sequences 214a-214n as non-overlapping, in some cases, the graph reference genome utilized by existing sequencing systems contains a large number of overlapping alternative contiguous sequences that may undergo coordinate transformation at any given genomic region of the linear reference genome. Therefore, existing sequencing systems (such as those depicted in Figure 2) typically must consider numerous alternative contiguous sequences in addition to the linear reference sequence when mapping and aligning nucleotide reads to the graph reference genome.
[0079] For example, as shown, the depicted sequencing system predicts the alignment of a subset of nucleotide reads 220 from nucleotide read 218 with the alternative contiguous sequence 214b of the graph reference genome 212. As shown in Figure 2, at least a portion of the subset of nucleotide reads 220 overlaps with the alternative contiguous sequence 214b. Although not shown in Figure 2, individual nucleotide reads (or related groups of nucleotide reads) typically overlap with multiple sequences contained in the graph reference genome (such as the graph reference genome 212 depicted in Figure 2). For example, in addition to alignment with the candidate continuous sequence 214b, the subset of nucleotide reads 220 is likely (at least partially) to overlap with one or more other candidate continuous sequences (not shown) of the graph reference genome 212 and the reference sequence 216b of the graph reference genome 212, and / or to overlap with one or more polybase codes not depicted in Figure 2.
[0080] As noted above, in some embodiments, the read alignment adjustment system 106 determines candidate alignments between nucleotide reads from a genome sample and the major continuous sequence, and evaluates these candidate alignments based on the variation between the major continuous sequence and the corresponding population haplotype.For example, Figures 3A and 3B show that the read alignment adjustment system 106 determines candidate alignments 306a, 306b to 306n for nucleotide reads 302, and generates adjusted alignment scores 314a, 314b to 314n from the corresponding major alignment scores 308a, 308b to 308n corresponding to the candidate alignments 306a, 306b to 306n based on the population haplotype 310. In describing Figures 3A and 3B, the following paragraphs summarize how the read alignment adjustment system 106 (i) determines the major alignment scores for read alignments with major continuous sequences, and (ii) adjusts the major alignment scores based on comparisons between the reads and allele-variant differences representing differences between major continuous sequences and population haplotypes. As indicated by the ellipses (or dots) in Figures 3A and 3B, the read alignment adjustment system 106 can identify, determine, generate, or utilize more candidate alignments, major alignment scores, allele-variant differences, and / or adjusted alignment scores than depicted in Figures 3A and 3B. Following the description of Figures 3A to 3B, this disclosure will provide further details and embodiments of the read alignment adjustment system 106 in subsequent paragraphs and figures.
[0081] In one or more embodiments, for example, the read alignment adjustment system 106 identifies or receives nucleotide reads from a genomic sample. In some cases, for example, the read alignment adjustment system 106 receives base detection data (e.g., BCL or FASTQ files) from a sequencing device that has sequenced oligonucleotides extracted from the genomic sample and has determined the individual base detections of the nucleotide reads in the base detection data. Depending on the type of sequencing performed, in some implementations, the read alignment adjustment system 106 identifies or receives single-end or paired-end reads, as well as relatively short nucleotide reads (e.g., <300 base pairs or <10,000 base pairs) or relatively long nucleotide reads (e.g., >300 base pairs or >10,000 base pairs) for mapping and alignment with a reference genome.
[0082] As shown in Figure 3A, the read alignment adjustment system 106 aligns a subset of nucleotide reads 302 from a genome sample with major contiguous sequences 304 at different genomic regions of the reference genome to determine candidate alignments 306a-306n. To show only a few candidate alignment regions, Figure 3A depicts subsets of nucleotide reads 302 at three different genomic regions corresponding to candidate alignments 306a, 306b, and 306n. In one or more embodiments, for example, the primary continuous sequence 304 includes a linear reference sequence that contains a recognized representation of a reference genome (e.g., the human genome) corresponding to the genome sample.In some implementations, the major contiguous sequence 304 is selectively broadened to include data representing population variation in specific genomic regions of the reference genome. For example, the major contiguous sequence 304 may include nucleotide positions encoding multiple bases that represent population variation in regions that are difficult to map (e.g., genomic regions with relatively high-frequency population variation within the reference population).
[0083] As shown, the read alignment adjustment system 106 generates major alignment scores 308a-308n for the corresponding candidate alignments 306a-306n based on a comparison of nucleotide bases within a subset of nucleotide reads 302 with nucleotide bases indicated by the major contiguous sequence 304 at the corresponding genomic regions of the candidate alignments 306a-306n. In some embodiments, the read alignment adjustment system 106 identifies candidate alignments 306a-306n that have a corresponding alignment score exceeding a threshold alignment score relative to the major contiguous sequence 304 to select them as candidate alignments. In some implementations, for example, the read alignment adjustment system 106 uses Smith-Waterman scores, modified versions of Smith-Waterman scores, or similar scoring models or criteria to generate major alignment scores 308a-308n relative to the major contiguous sequence 304.
[0084] Furthermore, as mentioned above, the read alignment adjustment system 106 adjusts the major alignment scores 308a-308n for each corresponding candidate alignment 306a-306n based on population variation at the corresponding genomic region of the reference genome. For example, as shown in FIG3B, the read alignment adjustment system 106 generates one or more adjusted alignment scores 314a-314n for each corresponding candidate alignment 306a-306n based on comparing the nucleotide bases within a subset of nucleotide reads 302 with the variant nucleotide bases of population haplotype 310 at the genomic region of the corresponding candidate alignment 306a-306n.
[0085] Specifically, as shown in FIG3B, the read alignment adjustment system 106 identifies allele-variant differences 312a, 312b to 312n between the major continuous sequence 304 and the population haplotype 310 relative to the corresponding candidate alignments 306a, 306b to 306n. Based on the allele-variant differences 312a-312n, the read alignment adjustment system 106 determines adjustments for the corresponding major alignment scores 308a-308n and generates adjusted alignment scores for each population haplotype (or each locally distinct population haplotype) containing variations at the corresponding genomic region of the reference genome. For example, the allele-variant differences 312a-312n between the population haplotype 310 and the major continuous sequence 304 can include any type of variant, such as, but not limited to, single nucleotide polymorphisms (SNPs), insertions or deletions (insertions or deletions), or other structural variants.
[0086] As further shown in FIG3B, the read alignment adjustment system 106 identifies allele-variant differences 312a-312n between the major continuous sequence 304 and the population haplotype 310 for each corresponding genomic region of candidate alignments 306a-306n. Based on the allele-variant differences 312a-312n, the read alignment adjustment system 106 determines an adjusted alignment score for each population haplotype 310 that contains one or more variations different from the major continuous sequence 304.
[0087] For example, for candidate alignment 306a, the read alignment adjustment system 106 identifies allele-variant differences 312a corresponding to one or more population haplotypes within the corresponding genomic region of the reference genome. Based on the allele-variant differences 312a of each of the one or more population haplotypes containing variations within the corresponding genomic region, the read alignment adjustment system 106 determines one or more adjusted alignment scores corresponding to the one or more population haplotypes, which belong to the adjusted alignment score 314a. Specifically, in some implementations, the read alignment adjustment system 106 improves the primary alignment score 308a for each match between the nucleotide read 302 and a variant nucleotide in a given haplotype (such as that represented by allele-variant difference 312a) within the population haplotype 310. Furthermore, the read alignment adjustment system 106 reduces the primary alignment score 308a for each mismatch between the nucleotide read 302 and a variant nucleotide in a given haplotype (such as that represented by allele-variant difference 312a) within the population haplotype 310. Therefore, as shown in Figure 3B, the read alignment adjustment system 106 generates an adjusted alignment score corresponding to each identified haplotype in the population haplotype 310 within the corresponding genomic region of the candidate alignment 306a, which belongs to the adjusted alignment score 314a. Furthermore, the read alignment adjustment system 106 performs similar steps to determine one or more adjusted alignment scores 314b-314n based on the primary alignment scores 308b-308n of the remaining candidate alignments 306b-306n.
[0088] In some embodiments, in addition to adjusting alignment scores based on read-variant matching and / or mismatch between nucleotide reads 302 and the corresponding population haplotypes 310, the read alignment adjustment system 106 further adjusts the primary alignment scores 308a-308b based on the population frequency (e.g., population allele frequency) of the corresponding population haplotypes 310. For example, the read alignment adjustment system 106 may increase the corresponding adjusted alignment score of a population haplotype with a relatively high frequency in the reference population, or decrease the corresponding adjusted alignment score of a population haplotype with a relatively low frequency in the reference population.
[0089] Therefore, as shown in FIG3B, the read alignment adjustment system 106 can generate multiple adjusted alignment scores 314a-314n for each corresponding candidate alignment 306a-306n. Based on the main alignment scores 308a-308n and the corresponding adjusted alignment scores 314a-314n, the read alignment adjustment system 106 can select nucleotide read 302 to be aligned with the predicted genomic region of the corresponding genomic region of the reference genome represented by the main continuous sequence 304. For example, as described in more detail below (e.g., in conjunction with Figure 6), the read alignment adjustment system 106 can generate substitution alignment scores for one or more of the candidate alignments 306a, 306b, or 306n based on the corresponding primary alignment scores 308a, 308b, or 308n and the corresponding adjusted alignment scores 314a, 314b, or 314n, and select the predicted alignment of the nucleotide read 302 based on these substitution alignment scores (e.g., by selecting the candidate alignment with the highest substitution alignment score).
[0090] As previously mentioned, in one or more embodiments, the read alignment adjustment system 106 determines alignment scores for one or more nucleotide reads, including single-ended nucleotide reads, paired-ended reads, or otherwise grouped nucleotide reads from a genomic sample. For example, Figure 4 shows an overview of a series of actions 400 for determining alignment score adjustments for unpaired reads and / or paired-ended reads. In various embodiments, the read alignment adjustment system 106 performs one or more actions from the series of actions 400 shown in FIG. 4.
[0091] As shown, the series of actions 400 includes an action 402 of generating a seed from one or more nucleotide reads. For example, the read alignment adjustment system 106 identifies one or more nucleotide reads corresponding to genomic regions of a genomic sample. For example, the read alignment adjustment system 106 may identify nucleotide reads corresponding to a sample genomic sequence of a genomic sample. More specifically, the sample genomic sequence comprises a continuous DNA or RNA fragment isolated or extracted from a sample organism, used as a template, and sequenced or generated as nucleotide reads using a single-end or paired-end method. Therefore, the sample genomic sequence is sometimes referred to as a template or template sequence. In the single-end method, a single-end nucleotide read is sequenced from one end (or primer) of the sample genomic sequence. Since the single-end nucleotide read is sequenced from one end of the sample genomic sequence, the single-end nucleotide read represents a complementary sequence of the sample genomic sequence.
[0092] In contrast, in the paired-end method, a first nucleotide read (e.g., R1) is sequenced from one end (or the first primer) of the sample genome sequence toward the middle, and a second nucleotide read (e.g., R2) is sequenced from the other end (or the second primer). This disclosure provides a further example of the first and second nucleotide reads in FIG. 5A, wherein reads R1 and R2 are positioned opposite each other.As discussed herein, two paired nucleotide reads (e.g., R1 and R2) are generally referred to as pairs. In some cases, there is a gap between the two pairs of paired nucleotide reads, while in other cases, there may be overlap between the pairs of paired nucleotide reads. As shown in series of actions 400, the read alignment adjustment system 106 generates a k-mer (i.e., a nucleotide sequence of length k) seed based on the nucleotide bases indicated by one or more nucleotide reads. To illustrate this, the read alignment adjustment system 106 generates a seed S, shown in a cross-shaded pattern in Figure 4. Specification 15 / 43 pages 22 CN 121359207 A
[0093] Figure 4 also illustrates action 404 for identifying candidate alignments with the main continuous sequence of the reference genome. For example, in various embodiments, the read alignment adjustment system 106 uses this seed to identify subsequences of the main continuous sequence that overlap wholly or partially with one or more nucleotide reads used to generate the seed. As shown, the read alignment adjustment system 106 uses the seed to determine candidate positions along the main continuous sequence that match the nucleotide bases of one or more nucleotide reads. In some implementations, the read alignment adjustment system 106 requires a perfect match to the seed. In other implementations, the read alignment adjustment system 106 selects candidate alignments that match the nucleotide bases in a threshold number or proportion.
[0094] As further shown, the series of actions 400 includes determining whether one or more nucleotide reads contain paired-end reads, or in other words, whether the nucleotide reads corresponding to the candidate alignments are paired-end reads. If the nucleotide read is a single-end read (or otherwise unpaired), the read alignment adjustment system 106 performs action 408, which determines the alignment score adjustment for the single-end read according to one or more embodiments described herein (see, for example, Figure 3B and the corresponding text).
[0095] In contrast, in implementations that include paired reads (e.g., as determined or identified in action 406), the series of actions 400 includes action 410, which determines whether a candidate pair of a first pair of paired reads is within a threshold distance of a second pair of paired reads (i.e., separated by a number of nucleotides less than a threshold number of main contiguous sequences). Therefore, in some embodiments, the read alignment adjustment system 106 identifies one or more paired candidate pairs for paired reads.
[0096] For example, as shown in FIG4, the series of actions 400 also includes action 412, which identifies candidate paired reads within a predetermined search region (e.g., a search region defined by a threshold number of nucleotides). Specifically, when the read alignment adjustment system 106 determines at action 410 that a second pair of paired reads is not within the threshold distance of the corresponding first pair, the read alignment adjustment system 106 may search for candidate pairs of the second pair within the search region defined by that threshold distance.In practice, in some implementations, by considering the pairing relationships of the two-end reads, the read alignment adjustment system 106 can thereby identify candidate pairings that might otherwise be ignored (e.g., due to incomplete overlap with the main continuous sequence).
[0097] For paired candidate pairings that are already within a threshold distance, in various embodiments, the read alignment adjustment system 106 proceeds to action 414, which determines the alignment score adjustment for these candidate pairings. Otherwise, after identifying candidate pairings within a predetermined search area (at action 412), the read alignment adjustment system 106 can execute action 414, which determines the alignment score adjustment for the paired candidate pairings. Thus, in one or more embodiments, the read alignment adjustment system 106 scores the paired candidate pairings together to generate an adjusted alignment score corresponding to the two-end reads.
[0098] As previously mentioned, in one or more embodiments, the read alignment adjustment system 106 generates adjusted alignment scores for candidate alignments of nucleotide reads to these corresponding genomic regions based on one or more locally distinct haplotypes at corresponding genomic regions of the reference genome. According to one or more embodiments, Figures 5A to 5B illustrate a series of actions 500a-500b that determine the adjusted alignment scores of candidate alignments based on locally distinct haplotypes and generate replacement alignment scores for each candidate alignment based on the adjusted alignment scores.
[0099] For example, as shown in Figure 5A, the series of actions 500a includes an action 502 that identifies one or more nucleotide reads. As discussed above (e.g., in conjunction with Figure 4), one or more nucleotide reads may include single-terminal nucleotide reads, paired-terminal nucleotide reads, or other subsets of nucleotide reads from a genomic sample. For example, as shown in Figure 5A, the read alignment adjustment system 106 may identify pairs R1 and R2 of paired-terminal nucleotide reads for mapping and alignment with the main contiguous sequence of the reference genome. However, as mentioned, the read alignment adjustment system 106 can perform the disclosed mapping and alignment methods on single-end reads, paired-end reads, or other grouped reads (such as nucleotide read stacks from a genome sample (e.g., as shown in FIG3A)).
[0100] Also as shown in FIG5A, the series of actions 500a includes action 504, which determines candidate alignments 514a-514n between one or more nucleotide reads and major contiguous sequences within the corresponding genomic region of the reference genome (page 16 / 43 of the specification, CN 121359207 A). For example, as shown, the read alignment adjustment system 106 determines candidate alignments 514a-514n of nucleotide read R1 with major contiguous sequences, wherein candidate alignments 514a-514n have different degrees of overlap with the corresponding nucleotide bases of the major contiguous sequences.For example, candidate alignment 514b shown has a shorter read length compared to candidate alignment 514b, which is at least partly due to the shorter span of nucleotide base overlap between candidate alignment 514b and the major continuous sequence. Additionally, candidate alignment 514n shown contains segments within the corresponding read, indicating that this is a partial alignment with non-contiguous overlap with the major continuous sequence. In fact, the read alignment adjustment system 106 can identify candidate alignments with nucleotide reads that have varying degrees and configurational overlap with the major continuous sequence. As indicated by ellipses (or dots) in Figures 5A and 5B, the read alignment adjustment system 106 can identify, determine, generate, or utilize more candidate alignments, major alignment scores, allele-variant differences, locally distinct haplotypes, adjusted alignment scores, substitution alignment scores, and / or predicted read alignments than depicted in Figures 5A and 5B.
[0101] As further shown in FIG5A, the series of actions 500a includes action 506, which generates a major alignment score for candidate alignments of one or more nucleotide reads to major contiguous sequences. For example, the read alignment adjustment system 106 generates an alignment score for each of the candidate alignments 514a-514n based on the amount of overlap between one or more nucleotide reads and major contiguous sequences at corresponding genomic regions of the reference genome. In some embodiments, the major alignment score includes a Smith-Waterman score, an adjusted Smith-Waterman score, or a similar scoring metric. For example, as shown, the read alignment adjustment system 106 determines a major alignment score of 0.92 for candidate alignment 514a and a major alignment score of 0.73 for candidate alignment 514b.
[0102] After generating major alignment scores for candidate alignments 514a-514n, as further shown in FIG5B, the read alignment adjustment system 106 may perform a series of actions 500b to generate replacement alignment scores for one or more candidate alignments. For example, as shown in Figure 5B, the series of actions 500b includes action 508, which identifies allele-variant differences in candidate alignments of one or more nucleotide reads. To illustrate this, in the illustrated implementation, read alignment adjustment system 106 identifies at least two locally distinct haplotypes in the genomic region corresponding to candidate alignment 514a, labeled “Haplotype 1” and “Haplotype 2”, respectively. As shown, read alignment adjustment system 106 identifies allele-variant differences for each corresponding locally distinct haplotype without explicitly identifying reference nucleotides within each corresponding haplotype. In other words, in one or more embodiments, read alignment adjustment system 106 identifies differences between locally distinct population haplotypes and the major continuous sequence, without identifying matching nucleotides between alternative continuous sequences and the major continuous sequence.As noted above, the read alignment adjustment system 106 avoids directly comparing nucleotide reads with selected consecutive sequences and determining alignment scores.
[0103] In various embodiments, if a particular population haplotype includes a unique set of variants (e.g., SNPs or insertions / deletions) relative to other population haplotypes within a given genomic region of the reference genome (e.g., within a genomic region corresponding to a candidate alignment), then that population haplotype is “locally distinct” within that given genomic region of the reference genome. In implementations where two or more population haplotypes contain the same set of variants within a given genomic region, for example, the read alignment adjustment system 106 identifies only one locally distinct haplotype, rather than two or more identical population haplotypes within that given genomic region. Additionally, in implementations where two given haplotypes have one or more identical variants within a given genomic region but also have at least one different variant within that given genomic region, the read alignment adjustment system 106 identifies the two given haplotypes as separate locally distinct haplotypes.
[0104] Also as shown in FIG5B, the series of actions 500b includes action 510, which generates an adjusted alignment score for each corresponding candidate alignment of one or more nucleotide reads, corresponding to each identified population haplotype of the local different haplotype. For example, a major alignment score relative to the major continuous sequence has previously been generated for candidate alignment 514a, and the read alignment adjustment system 106 adjusts the major alignment score based on the allele-variant differences identified for each local different haplotype. By making such adjustments to the major alignment score, the read alignment adjustment system 106 generates an adjusted alignment score corresponding to the corresponding local different haplotype.
[0105] Specifically, also as described above (e.g., in conjunction with FIG. 3B), when the nucleus bases of a given haplotype match the nucleus bases of the corresponding nucleotide read (e.g., as shown with respect to locally different haplotype 1), the read alignment adjustment system 106 increases the major alignment score; when the nucleus bases of a given haplotype do not match the corresponding nucleotide read (e.g., as shown with respect to locally different haplotype 2), the read alignment adjustment system decreases the major alignment score. In various embodiments, the read alignment adjustment system 106 considers additional information, such as, but not limited to, population allele frequencies from each considered haplotype, when adjusting the major alignment score for each locally different haplotype.
[0106] In one or more embodiments, for example, the read alignment adjustment system 106 further adjusts the major alignment score of a given candidate alignment based on the prior probability of the haplotype variant (e.g., to reduce false positives in variant detection from reads aligned to rare haplotypes).Therefore, in some embodiments, the read alignment adjustment system 106 identifies the population frequency (e.g., prior probability) of each allele-variant difference for each locally distinct population haplotype and determines an alignment score adjustment that explains the relative rarity of each allele-variant difference. For example, when the read alignment adjustment system 106 identifies an allele-variant difference with a relatively low prior probability, the read alignment adjustment system 106 may reduce the adjusted alignment score corresponding to that corresponding haplotype relative to the primary alignment score. Furthermore, when the read alignment adjustment system 106 identifies an allele-variant difference with a relatively high prior probability, the read alignment adjustment system 106 may correspondingly increase the adjusted alignment score.
[0107] Alternatively, in some embodiments, the read alignment adjustment system 106 initially determines the adjusted alignment score for locally distinct haplotypes within a genomic region corresponding to a given candidate read, and then further adjusts each adjusted alignment score to take into account the prior probability of each corresponding population haplotype. In one or more embodiments, for example, the read alignment adjustment system 106 converts an initial adjusted alignment score into a likelihood value (e.g., as discussed below in conjunction with FIG. 6), and then increases or decreases the resulting likelihood value based on the prior probability (e.g., population frequency) of the corresponding population haplotype (e.g., increasing the given likelihood value based on a relatively high population frequency, or decreasing the given likelihood value based on a relatively low population frequency).
[0108] Furthermore, in some embodiments, the read alignment adjustment system 106 uses the principal alignment score and the adjusted alignment score of a given candidate alignment to determine a replacement alignment score for that given candidate alignment. For example, a series of actions 500b includes action 511, which generates a replacement alignment score for one or more candidate alignments. To illustrate this, as shown in FIG. 5B, the read alignment adjustment system 106 generates a replacement alignment score for candidate alignment 514a based on the corresponding principal alignment score, the adjusted alignment score of locally different haplotype 1, the adjusted alignment score of locally different haplotype 2, and any additional adjusted alignment scores not depicted in FIG. 5B. Additional details regarding the substitution alignment scores are provided below in conjunction with Figure 6.
[0109] As further shown in Figure 5B, the series of actions 500b includes action 512, which selects a predicted read alignment for one or more nucleotide reads. In some embodiments, the read alignment adjustment system 106 may select a predicted read alignment from candidate read alignments 514a-514n based on the substitution alignment scores generated for each candidate read alignment according to the series of actions 500a and 500b, or alternatively, based on the primary alignment score when the primary alignment score is better than or exceeds the corresponding adjusted alignment score.For example, as shown in the figure, the read alignment adjustment system 106 selects a first candidate read alignment 514a from candidate read alignments 514a-514n because it has the highest corresponding substitution alignment score, and in some embodiments, outputs candidate alignment 514a as a predicted read alignment for one or more nucleotide reads. In some implementations, the read alignment adjustment system 106 may select multiple candidate read alignments for output (e.g., output to a BAM file), such as, but not limited to, in cases where multiple candidate alignments have the same or nearly the same substitution alignment scores.
[0110] As mentioned, in some embodiments, the read alignment adjustment system 106 generates substitution alignment scores for candidate alignments of one or more nucleotide reads based on the corresponding primary alignment scores, as specified on page 18 / 43 of the specification, 25 CN 121359207 A, and one or more adjusted alignment scores generated according to the disclosed method. According to one or more embodiments, FIG6 illustrates that the read alignment adjustment system 106 generates a replacement alignment score 612 for candidate alignments 602 based on the corresponding primary alignment score 604 and the adjusted alignment score 606.
[0111] As shown in FIG6, the read alignment adjustment system 106 determines the primary alignment score 604 of candidate alignments 602 between one or more nucleotide reads from a genomic sample and the primary contiguous sequence at the corresponding genomic region of a reference genome. Based on one or more local dissimilar haplotypes within the corresponding genomic region, the read alignment adjustment system 106 determines the adjusted alignment score 606 (e.g., as described above in conjunction with FIG3B). For example, in the illustrated example shown in FIG6, the read alignment adjustment system 106 determines the adjusted alignment score for each corresponding population haplotype among the local dissimilar haplotypes 1 to N. As indicated by the ellipsis (or dots) in FIG6, the read alignment adjustment system 106 can determine more adjusted alignment scores for local dissimilar haplotypes than depicted in FIG6.
[0112] As further shown in FIG6, the read segment alignment adjustment system 106 determines the replacement alignment score 612 of the candidate alignment 602 based on the primary alignment score 604 and the adjusted alignment score 606. In various embodiments, the read segment alignment adjustment system 106 may utilize a variety of methods to determine the replacement alignment score 612 of the candidate alignment 602. For example, in some implementations, the read segment alignment adjustment system 106 selects the maximum alignment score 608 from the adjusted alignment score 606 and the primary alignment score 604. In such implementations, the maximum alignment score 608 constitutes the replacement alignment score 612.
[0113] In contrast, in some implementations, the read segment alignment adjustment system 106 determines the combined alignment score 610 based on the primary alignment score 604 and the adjusted alignment score 606.In one or more embodiments, the read alignment adjustment system 106 converts each of the primary alignment score 604 and the adjusted alignment score 606 into a likelihood value (e.g., the quantified probability of one or more nucleotide reads corresponding to corresponding major or local distinct population haplotypes). In such embodiments, the combined alignment score 610 constitutes a replacement alignment score 612. For example, in some embodiments, the read alignment adjustment system 106 converts each alignment score into a likelihood value according to the following mathematical relationship:
[0114]
[0115] where C represents a normalization constant, and ∝ represents a base chosen based on the length of one or more nucleotide reads. Thus, as shown in FIG6, the read alignment adjustment system 106 converts the corresponding alignment scores into likelihood values and adjusts and / or combines the resulting likelihood values to determine the overall likelihood value of the candidate alignment 602. Thus, in some cases, the resulting replacement alignment score 612 represents the likelihood value of the corresponding nucleotide read corresponding to the corresponding genomic region of the reference genome. By converting the total / summed likelihood value into an alignment score, the read alignment adjustment system 106 can generate a replacement alignment score 612 for the candidate alignment 602.
[0116] As previously mentioned, the read alignment adjustment system 106 can utilize an enhanced haplotype data structure encoding allele-variant differences to implement the aforementioned mapping and alignment methods. According to one or more embodiments, Figures 7 and 8 illustrate a haplotype data structure that includes a hierarchical division of a reference genome to efficiently and accurately encode population haplotype data of the reference genome. Specifically, Figure 7 illustrates the basic hierarchy of haplotype data structure 700 according to one or more embodiments, while Figure 8 illustrates the basic hierarchy 802 and multiple consecutive hierarchies 806a-806n of haplotype data structure 800 according to one or more embodiments.
[0117] As shown in FIG. 7, the haplotype data structure 700 includes at least one base level containing a set of base level bins 702a, 702b to 702n, which divide the genomic region of the reference genome into a corresponding set of base level reference spans 704a, 704b to 704n. In one or more embodiments, each base level reference span in the set of base level reference spans 704a-704n includes a genomic region of a first specification page 19 / 43 26 CN 121359207 A length between corresponding genomic coordinates of the reference genome, thereby dividing the genomic region of the reference genome into multiple bins, each bin spanning an equal portion / length of the reference genome. In various implementations, the length of the base level reference span may be approximately, for example, the average or maximum length of the nucleotide reads provided to the read alignment adjustment system 106 for mapping and alignment.Alternatively, the basal hierarchical reference span can be selected in other ways to span a predetermined number of nucleotide bases across genomic coordinates or regions from a linear reference sequence, such as, but not limited to, 100 or 1000 base pairs per basal hierarchical bin.
[0118] As further shown in Figure 7, this set of basal hierarchical bins 702a-702n of the haplotype data structure 700 contains coding variant data for nucleotide variants from corresponding groups of locally distinct haplotypes 706a-706n. As previously mentioned, each locally distinct haplotype within a given basal hierarchical bin includes one or more allele-variant differences that are unique relative to other population haplotypes that also have variations within the genomic region of the corresponding basal hierarchical reference span of that given basal hierarchical bin. For example, as shown in Figure 7, each row in this set of locally distinct haplotypes 706a includes a unique set of allele-variant differences (represented by a single letter representing a specific nucleotide) relative to the other rows, such that no two rows are identical—although there may be limited overlap between allele-variant differences, as indicated by the first two rows of the base-level bin 702a. Thus, in one or more embodiments, a population of haplotypes having the same nucleotide variant within a given base-level bin is encoded as a locally distinct haplotype within that given base-level bin.
[0119] In various embodiments, each base-level bin of the haplotype data structure 700 may include a different number of locally distinct haplotypes. For example, as shown in Figure 7, a basic hierarchical bin 702a includes four locally different haplotypes in the set of locally different haplotypes 706a (as indicated by the fourth row of the depicted matrix), a basic hierarchical bin 702b includes five locally different haplotypes in the set of locally different haplotypes 706b, and a basic hierarchical bin 702n includes three locally different haplotypes in the set of locally different haplotypes 706n. In practice, each basic hierarchical bin of the haplotype data structure 700 can include any number of locally different haplotypes, including up to every population haplotype in the dataset, or excluding any population haplotypes (e.g., in the case where there are no haplotypes with allele-variant differences in the genomic region corresponding to a given bin).
[0120] As further shown in FIG7, the set of basic hierarchical bins 702a, 702b to 702n includes allele-variant differences 708a, 708b to 708n for each of the locally different haplotypes 706a, 706b to 706n in the corresponding group. For example, variant data encoded in the basic hierarchical bin 702a includes one or more locally different population haplotypes in the set of locally different haplotypes 706a, and allele-variant differences 708a are included for each corresponding locally different haplotype.In some implementations, for example, each basal hierarchical bin (e.g., basal hierarchical bins 702a-702n belonging to that group) includes a matrix containing corresponding variant data representing allelic-variant differences from locally distinct haplotypes (e.g., locally distinct haplotypes 706a-706n belonging to the corresponding group) and the variant locations of these allelic-variant differences (e.g., as shown in Figures 10-11). In various implementations, the variant data within each basal hierarchical bin includes data indicating single nucleotide polymorphisms (SNPs) and / or insertions or deletions (insertions / deletions) at corresponding genomic coordinates of a reference genome (e.g., for major continuous sequences) (e.g., allelic-variant differences 708a-708n). As indicated by the ellipses (or dots) in Figure 7, the read alignment adjustment system 106 can identify, determine, generate, or utilize more basal hierarchical bins, locally distinct population haplotypes, basal hierarchical reference spans, and / or allelic-variant differences than depicted in Figure 7.
[0121] Furthermore, in some embodiments, the basic hierarchical bins (e.g., the set of basic hierarchical bins 702a-702n) include variant data of nucleotide variants but do not include reference nucleobases of the major continuous sequence. For example, as shown in FIG7, each basic hierarchical bin in the set of basic hierarchical bins 702a-702n contains a matrix whose rows represent multiple sets of locally different haplotypes 706a-706n within the basic hierarchical reference span 704a-704n of the corresponding group, and whose columns represent the allele-variant differences 708a-708n of each locally different haplotype. As shown, the allele-variant differences 708a-708n are labeled with letters representing nucleotides that differ from the major continuous sequence. Alternatively, in various embodiments, allele-variant differences can be represented numerically (e.g., using "0" to represent a nucleotide matching the major continuous sequence, with subsequent values representing variations from the major continuous sequence), or similarly to represent the differences between each population haplotype and the major continuous sequence.
[0122] As mentioned, in some embodiments, the read alignment adjustment system 106 utilizes a haplotype data structure that divides the genomic regions of a reference genome into multiple levels of bins corresponding to nucleotide spans within the reference genome. For example, Figure 8 shows a haplotype data structure 800 having a base level 802 containing a set of base level bins 804 and multiple higher-level bins 806a, 806b, 806c to 806n, which successively span larger nucleotide spans in the reference genome.Specifically, the haplotype data structure 800 includes a set of basic-level bins 804 of basic level 802, which collectively span the major contiguous sequences of the reference genome; and higher-level bins 808a-808c and offset higher-level bins 809a-809c of multiple contiguous levels 806a-806c, which also span the major contiguous sequences of the reference genome. As indicated in Figure 8, contiguous level 806n includes higher-level bins 808n and corresponding offset higher-level bins, but due to space limitations, the corresponding offset higher-level bins are not depicted in Figure 8. As further indicated by the ellipses (or dots) in Figure 8, the read alignment adjustment system 106 can identify, determine, generate, or utilize more basic-level bins, contiguous levels, higher-level bins, and / or offset higher-level bins than depicted in Figure 8.
[0123] As shown in the figure, the basic level 802 of the haplotype data structure 800 includes a set of basic level bins 804, which correspond to a set of basic level reference spans of the main continuous sequence of the reference genome. Each reference span in this set of basic level reference spans corresponds to a genomic region of a first length between corresponding genomic coordinates of the reference genome. In one or more embodiments, for example, each reference span in this set of basic level reference spans includes 1000 base pairs (1 kbp) of the main continuous sequence of the reference genome. Alternatively, the first length of the basic level reference span may be less than or greater than 1 kbp, such as, but not limited to, 250 bp, 500 bp, 1500 bp, 5 kbp, 10 kbp, etc. Thus, in various embodiments, this set of basic level bins 804 collectively spans the entire main continuous sequence or the genomic region of interest, such as, but not limited to, the entire chromosome.
[0124] As further indicated in FIG8, this set of basic hierarchical bins 804 of basic hierarchical 802 includes variant data of nucleotide variants from locally distinct population haplotypes in the corresponding group (e.g., as described above in conjunction with FIG7). As mentioned, each locally distinct population haplotype includes one or more allele-variant differences in a unique group that are unique relative to other population haplotypes within the corresponding basic hierarchical reference span of a given basic hierarchical bin in this set of basic hierarchical bins 804. For example, as shown in FIG8, this set of basic hierarchical bins 804 includes locally distinct population haplotypes in corresponding groups that have different numbers of locally distinct haplotypes, as indicated by the numbers associated with each basic hierarchical reference span in this set of basic hierarchical bins 804.For example, as shown in the figure, the first basic level bin includes three locally different haplotypes (indicated by "3(0..2)"), the second basic level bin includes two locally different haplotypes (indicated by "2(0..1)"), the third basic level bin includes three locally different haplotypes (indicated by "3(0..2)"), and the fourth basic level bin includes four locally different haplotypes (indicated by "4(0..3)"). As previously mentioned, each locally different haplotype within a given basic level bin can represent one or more population haplotypes, because population haplotypes with the same nucleotide variant within a given basic level bin are encoded as a locally different population haplotype within that given basic level bin.
[0125] Also as shown in Figure 8, the haplotype data structure 800 includes higher-level bins 808a-808n of multiple consecutive levels 806a-806n. For example, the first contiguous level 806a includes a first set of higher-level bins 808a, which correspond to the first set of higher-level reference spans of the main contiguous sequences in Group A of the reference genome specification, page 21 / 43, document number 28, CN 121359207. Each reference span in the first set of higher-level reference spans corresponds to a second-length extended genomic region between the corresponding genomic coordinates of the reference genome, wherein these extended genomic regions are extended relative to the genomic region represented by that set of basic-level reference spans, such that the second length (of the corresponding first set of higher-level reference spans) is longer than the first length (of that set of basic-level reference spans). More specifically, as shown in Figure 8, each higher-level bin in the first set of higher-level bins 808a of the first contiguous level 806a corresponds to a pair of consecutive basic-level bins in that set of basic-level bins 804 from the basic level 802 of the haplotype data structure 800.
[0126] Furthermore, as indicated in FIG8, the multiple consecutive levels 806a-806c of the haplotype data structure 800 include corresponding sets of offset higher-level bins 809a-809c, and the consecutive level 806n of the haplotype data structure 800 includes a higher-level bin 808n and a corresponding offset higher-level bin. For example, the first consecutive level 806a includes a set of offset higher-level bins 809a, which correspond to the first set of offset higher-level reference spans of the main consecutive sequence of the reference genome. Each reference span in the first set of offset higher-level reference spans corresponds to an offset extended genomic region of a second length (i.e., the same as the reference span length of the first set of consecutive reference spans) between the corresponding genomic coordinates of the reference genome. Similar to the first set of higher-level bins 808a, the first set of offset higher-level bins 809a corresponds to the corresponding consecutive pairs of basic level bins in this set of basic level bins 804 from the basic level 802 of the haplotype data structure 800.Furthermore, as shown in the figure, the corresponding reference span of the first set of offset higher-level bins 809a is offset relative to the reference span of the first set of higher-level bins 808a, such that each pair of consecutive basic-level bins in this set of basic-level bins 804 is represented by a higher-level bin or an offset higher-level bin from the first consecutive level 806a.
[0127] In addition, each additional consecutive level 806b-806n of the haplotype data structure 800 includes additional higher-level bins 808b-808n, which correspond to the corresponding additional higher-level reference spans corresponding to further extended genomic regions between the genomic coordinates of the main consecutive sequence of the reference genome. Specifically, as shown in Figure 8, each higher-level bin (or offset higher-level bin) of a given consecutive level of the haplotype data structure 800 spans a combined genomic region of a pair of consecutive bins of the previous level of the haplotype data structure 800 (e.g., as indicated by the arrows connecting the individual bins in Figure 8). For example, the first display bin in the set of higher-level bins 808c spans the same genomic region represented by the first two display bins in the set of higher-level bins 808b. Similarly, the first display bin in the set of higher-level bins 808b spans the same genomic region represented by the first two display bins in the set of higher-level bins 808a. In fact, each successive level includes higher-level bins corresponding to a pair of successive bins from the level preceding the haplotype data structure 800.
[0128] Furthermore, in some embodiments, the corresponding higher-level bins of each successive level of the haplotype data structure 800 include variant data indexes that reference combinations of variant data from the base level bins corresponding to the base level 802. Specifically, each higher-level bin and offset higher-level bin in the multiple sets of higher-level bins 808a-808c and offset higher-level bins 809a-809c (as well as each of the higher-level bin 808n and its corresponding offset higher-level bin) includes variant data indexes that reference combinations of variant data from the corresponding base-level bins of that set of base-level bins 804. Furthermore, these variant data indexes include indications of locally distinct haplotypes within each corresponding higher-level bin or offset higher-level bin. For example, as shown in Figure 8, one of the offset higher-level bins 809b of the successive level 806b indicates fifteen locally distinct haplotypes (indicated by "15 haplotypes (0..14)"). As also shown in the figure, two bins in a higher-level bin 808a from the previous consecutive level (e.g., the first consecutive level 806a) indicate three locally different haplotypes (indicated by “3(0..2)”) and five locally different haplotypes (indicated by “5(0..4)”).Additionally, in some embodiments, each bin in the haplotype data structure encodes population frequency data for each corresponding locally distinct haplotype (e.g., the frequency of occurrence of each locally distinct haplotype indicated within a given bin in the sample population, as specified on pages 22 / 43 of the specification, CN 121359207 A).
[0129] In one or more embodiments, each higher-level bin in a successive hierarchy includes variant data indexes that indicate locally distinct haplotypes and link the higher-level bins to variant data within the corresponding base-level bins, without including variant data from the corresponding base-level bins, thereby avoiding redundant encoding of variant data within the haplotype data structure. Referring to successive levels 806b, for example, the bins in the aforementioned offset higher-level bins 809b indicating fifteen locally distinct haplotypes may include variant data indices that reference how locally distinct haplotypes from the corresponding higher-level bins (within higher-level bins 808a) from the previous successive level (e.g., the first successive level 806a) combine to form the aforementioned fifteen locally distinct haplotypes of the bin. Furthermore, each corresponding higher-level bin 808a may include variant data indices that reference the locally distinct haplotypes (and their variant data) indicated within the corresponding base-level bins (belonging to this set of base-level bins 804) from the base level 802. Therefore, by referencing the variant data indices within the prior successive levels of the haplotype data structure 800, the variant data indices of higher-level bins within successive levels 806b-806n may also reference the variant data encoded within this set of base-level bins 804.
[0130] As mentioned above, in some of the described embodiments, the read alignment adjustment system 106 offers improvements over existing systems in terms of efficiency and total data storage. Specifically, in some implementations, the read alignment adjustment system 106 utilizes a haplotype data structure that contains a hierarchical division of population variation relative to the main contiguous sequences of the reference genome (e.g., as described above in conjunction with Figures 7 and 8). To illustrate this, Figures 9A and 9B show experimental results of the read alignment adjustment system 106 encoding reference genome population variation data using a haplotype data structure (according to one or more of the disclosed embodiments).
[0131] For example, Figure 9A shows various metrics of bit utilization efficiency of the haplotype data structure according to one embodiment, as well as an overall spatial comparison between the haplotype data structure and an existing augmented graph reference genome (labeled "Est. Old-Graph Space"). As shown in Table 902, for example, this haplotype data structure allocates 1.79 bits per base, fills 1.30 bits per base, and utilizes 0.37 bits per base in each bin.Furthermore, the illustrated haplotype data structure allocates 0.53 bits per haplotype allele, 0.39 bits per haplotype allele padding, and utilizes 0.11 bits per haplotype allele in each basic bin. Additionally, the illustrated haplotype data structure allocates 5.70 bits per selected allele, 5.15 bits per selected allele padding, and utilizes 1.18 bits per selected allele in each basic bin. Moreover, as shown in Table 904, the illustrated implementation of the haplotype data structure has a total memory allocation of 612 MB, with an additional 1009 MB used for haplotype polymerization, for a total memory allocation of 1.6 GB. In comparison, at least one existing augmented map reference genome has a total memory allocation of 65 GB. Indeed, as shown in Figure 9A, the implementation of the haplotype data structure can achieve improved data storage efficiency when encoding population variation relative to the reference genome.
[0132] Furthermore, Figure 9B illustrates the bit allocation across multiple levels of the haplotype data structure in Figure 9A, including bin size (i.e., reference span length) indications for each corresponding level bin, bit usage at each level and overall for various encoded variant data, and the total number of MBs of data filled at each level and within the haplotype data structure as a whole. For example, as shown in Table 906, each consecutive level of the haplotype data structure presented occupies less memory relative to lower-level bins (e.g., bins spanning fewer nucleotide positions). Indeed, as shown in Figures 9A and 9B, this example implementation of the haplotype data structure achieves improved efficiency and overall data storage in terms of population variation of the reference genome compared to existing systems (such as existing augmented map reference genomes).
[0133] As previously mentioned, in some implementations, the read alignment adjustment system 106 utilizes the haplotype data structure (such as those described above in conjunction with Figures 7 and 8) to determine alignment score adjustments for nucleotide read candidate alignments based on variant data encoded within the haplotype data structure. For example, Figure 10 illustrates an overview of a series of actions 1000 for adjusting one or more alignment scores to determine candidate nucleotide read alignments based on structure, according to one or more embodiments, using haplotype number specification page 23 / 43 30 CN 121359207 A.
[0134] For example, the series of actions 1000 includes action 1002, which generates a primary alignment score for candidate nucleotide reads from a genomic sample. As shown, the read alignment adjustment system 106 identifies candidate alignments between a nucleotide read 1003 from a genomic sample and a primary contiguous sequence of a reference genome. In some embodiments, for example, the read alignment adjustment system 106 determines a set of candidate alignments for the nucleotide read 1003 (or a subset of overlapping nucleotide reads) and generates a corresponding set of primary alignment scores, as described above in conjunction with Figures 3A and 5A.For each candidate alignment in this set of candidate alignments, the read alignment adjustment system 106 can perform a series of actions 1000 to determine alignment score adjustments using the haplotype data structure 1005 (e.g., the haplotype data structure described above in conjunction with Figures 7 and 8).
[0135] Also as shown in Figure 10, the series of actions 1000 includes action 1004, which identifies bins in the haplotype data structure 1005 having corresponding reference spans that include the entire contents of the nucleotide reads 1003 (e.g., bins spanning each genomic coordinate of the candidate alignment with the major contiguous sequence). For example, as similar to that described above in conjunction with Figures 7 and 8, the haplotype data structure 1005 contains base-level bins that include corresponding base-level reference spans that correspond to a first-length genomic region between corresponding genomic coordinates of the reference genome. Furthermore, haplotype data structure 1005 includes one or more consecutive higher-level bins and offset higher-level bins, which include corresponding higher-level reference spans that correspond to extended genomic regions of greater length (relative to the first length) between corresponding genomic coordinates of the reference genome. While Figure 10 illustrates a single consecutive level of haplotype data structure 1005, haplotype data structure 1005 may include additional consecutive levels, such as those shown in Figure 8 (e.g., to provide a sufficient number of bins with sufficiently long reference spans to include all nucleobases of relatively long nucleotide reads).
[0136] As shown, read alignment adjustment system 106 queries haplotype data structure 1005 to identify base-level bins, higher-level bins, or offset higher-level bins whose corresponding reference spans include nucleotide reads 1003. In the illustrated implementation, for example, the read alignment adjustment system 106 identifies offset higher-level bins in the haplotype data structure 1005, which include the entire contents of the nucleotide read 1003. Also as described above in conjunction with Figures 7 and 8, each consecutive higher-level bin and offset higher-level bin in the haplotype data structure 1005 includes variant data indices that indicate combinations of variant data from the corresponding base-level bins. Therefore, the read alignment adjustment system 106 identifies one or more locally different haplotypes within the identified bins and, based on the variant data indices, identifies variant data of the corresponding one or more locally different haplotypes within the corresponding base-level bins.
[0137] Furthermore, as shown in Figure 10, the series of actions 1000 includes action 1006, which determines one or more alignment score adjustments based on the variant data from the identified bins of the haplotype data structure 1005.As mentioned, for example, each given base-level bin of haplotype data structure 1005 includes variant data of locally distinct haplotypes within the corresponding reference span of that given base-level bin, such as allele-variant differences between the corresponding locally distinct haplotypes and the major continuous sequence (e.g., as described above in conjunction with Figure 7). Furthermore, in some embodiments, the bins of haplotype data structure 1005 also include population frequency data of the corresponding locally distinct haplotypes (e.g., population allele frequencies). It has also been mentioned that higher-level bins of haplotype data structure 1005 include variant data indexes that indicate combinations of variant data from corresponding base-level bins. For example, as shown in Figure 10, variant data from identified bins includes a variant data matrix 1007 that represents allele-variant differences from locally distinct haplotypes and the variant locations of these allele-variant differences. As indicated by the ellipsis (or dots) in Figure 10, the read alignment adjustment system 106 can identify, determine, generate, or utilize more local different haplotypes and / or alignment score adjustments than depicted in Figure 10.
[0138] For further illustration, Figure 10 shows the variant data matrix 1007 indicating three allele-variant differences (labeled as "-T--G Specification 24 / 43 pages 31 CN 121359207 AA") between the first local different haplotype (labeled as "Haplotype 1") and the main continuous sequence of the reference genome. Therefore, by comparing the nucleotide bases of nucleotide read 1003 (labeled "AATCGA") with different haplotypes in the first region, read alignment adjustment system 106 determines a first set of alignment score adjustments, including: reducing the major alignment score for the mismatch between adenine in nucleotide read 1003 and thymine in the first haplotype at the second nucleotide position; and increasing the major alignment score for the guanine and adenine matches between nucleotide read 1003 and the first haplotype at the corresponding fifth and sixth nucleotide positions. Furthermore, variant data matrix 1007 indicates two allele-variant differences (labeled "A--C--") between different haplotypes in the second region (labeled "haplotype 2") and the main continuous sequence of the reference genome. Therefore, by comparing the nucleobases of nucleotide read 1003 with different haplotypes in the second locality, read alignment adjustment system 106 determines a second set of alignment score adjustments, including: improving the primary alignment score for adenine and cytosine that match the corresponding first and fourth nucleobase positions of nucleotide read 1003 and the second haplotype.
[0139] Therefore, as shown in FIG10, the read alignment adjustment system 106 determines alignment score adjustments for each locally distinct haplotype indicated by the identified binning based on a comparison of allele-variant differences indicated by the nucleotide bases within the nucleotide read 1003 and the variant data matrix 1007 of the variant data at corresponding nucleotide positions in the major continuous sequence (e.g., as further described above in conjunction with FIG3B and FIG5B).
[0140] As previously mentioned, in some embodiments, the read alignment adjustment system 106 utilizes a haplotype data structure (such as described above in conjunction with FIG7 and FIG8) to determine alignment score adjustments for candidate pairwise nucleotide read alignments based on variant data encoded within that haplotype data structure. For example, FIG11 shows an overview of a series of actions 1100 for determining alignment score adjustments for candidate pairwise nucleotide read alignments using a haplotype data structure according to one or more embodiments.
[0141] For example, the series of actions 1100 includes action 1102, which generates a primary alignment score for candidate alignments of paired nucleotide reads from a genomic sample, the paired reads including a first pair 1103a and a second pair 1103b. As shown, the read alignment adjustment system 106 identifies candidate alignments between paired nucleotide reads 1103a and 1103b from the genomic sample and the primary contiguous sequence of a reference genome. In some embodiments, for example, the read alignment adjustment system 106 determines a set of candidate alignments for the first pair 1103a and the second pair 1103b, wherein the pairs of each candidate alignment are within a threshold distance from each other (e.g., as described above in conjunction with Figure 4). For each candidate alignment of the paired reads 1103a and 1103b, the read alignment adjustment system 106 generates a corresponding set of primary alignment scores, as described above in conjunction with Figures 3A and 5A. For each candidate alignment in this set of candidate alignments, the read alignment adjustment system 106 can perform a series of actions 1100 to determine alignment score adjustments using the haplotype data structure 1105 (e.g., the haplotype data structure described above in conjunction with Figures 7 and 8).
[0142] Also as shown in Figure 11, the series of actions 1100 includes action 1104, which identifies bins in the haplotype data structure 1105 having corresponding reference spans (e.g., bins spanning each genomic coordinate of a candidate alignment between a paired read and a major contiguous sequence) that include two pairs 1103a and 1103b of paired nucleotide reads. For example, as similar to those described above in conjunction with Figures 7 to 8 and Figure 10, the haplotype data structure 1105 contains base-level bins that include corresponding base-level reference spans that correspond to a first-length genomic region between corresponding genomic coordinates of a reference genome.Furthermore, haplotype data structure 1105 includes multiple consecutive levels of higher-level bins and offset higher-level bins, which include corresponding higher-level reference spans that correspond to extended genomic regions of greater length (relative to the first length) between corresponding genomic coordinates of the reference genome. Although Figure 11 shows three consecutive levels of haplotype data structure 1105, haplotype data structure 1005 may include additional consecutive levels (and the number of bins within each corresponding level is significantly greater than illustrated).
[0143] As shown, read alignment adjustment system 106 queries haplotype data structure 1105 to identify base-level bins, higher-level bins, or offset higher-level bins, whose corresponding reference spans include two pairs 1103a and 1103b of the dipole nucleotide read specification, page 25 / 43, 32 CN 121359207 A. In the illustrated implementation, for example, the read alignment adjustment system 106 identifies offset higher-level bins within the third consecutive level of the haplotype data structure 1105, which include two pairs 1103a and 1103b of paired-terminal nucleotide reads. Also as described above in conjunction with Figures 7-8 and 10, each consecutive level of the haplotype data structure 1105 includes higher-level bins and offset higher-level bins containing variant data indices that indicate combinations of variant data from the corresponding base level bins. Therefore, the read alignment adjustment system 106 identifies one or more locally distinct haplotypes within the identified bins and, based on the variant data indices, identifies variant data of the corresponding one or more locally distinct haplotypes within the corresponding base level bins.
[0144] Furthermore, as shown in FIG11, the series of actions 1100 includes action 1106, which determines the alignment score adjustment of the first pair 1103a and the second pair 1103b based on variant data from the identified bins of the haplotype data structure 1105. As mentioned, for example, each given base-level bin of the haplotype data structure 1105 includes variant data of locally different haplotypes within the corresponding reference span of that given base-level bin, such as allele-variant differences between the corresponding locally different haplotypes and the major continuous sequence (e.g., as described above in conjunction with FIG7 and FIG10). Furthermore, in some embodiments, the bins of the haplotype data structure 1105 also include population frequency data of the corresponding locally different haplotypes. It is also mentioned that higher-level bins of the haplotype data structure 1105 include variant data indexes that indicate combinations of variant data corresponding to the base-level bins. For example, as shown in Figure 11, the variant data from the identified bins includes matrix 1107, which represents allele-variant differences from different local haplotypes and the variant locations of these allele-variant differences.
[0145] To further illustrate, FIG11 shows variant data matrix 1107 indicating three allelic-variant differences (labeled "-T--GA") between the first local haplotype (labeled "Haplotype 1") and the main continuous sequence of the reference genome at the nucleotide positions of the candidate alignments of the first pair 1103a corresponding to the paired-end reads. Therefore, by comparing the nucleotides of the first pair 1103a (labeled "AATCGA") with the first local haplotype, read alignment adjustment system 106 determines a first set of alignment score adjustments for the first pair 1103a, including: reducing the main alignment score for the mismatch between adenine in the first pair 1103a and thymine in the first haplotype at the second nucleotide position of the first pair 1103a; and increasing the main alignment score for the guanine and adenine matched by the first pair 1103a and the first haplotype at the corresponding fifth and sixth nucleotide positions of the first pair 1103a.
[0146] Additionally, the variant data matrix 1107 indicates an allelic-variant difference (labeled "---T--") between the first local different haplotype (labeled "haplotype 1") and the main continuous sequence of the reference genome at the nucleotide position of the candidate alignment of the second pair 1103b corresponding to the paired-end read. Therefore, by comparing the nucleotides of the second pair 1103b (labeled "CCGTAC") with the first local different haplotype, the read alignment adjustment system 106 determines a first set of alignment score adjustments for the second pair 1103b, including: improving the main alignment score for the thymine that matches the first haplotype at the fourth nucleotide position of the second pair 1103b.
[0147] Furthermore, the variant data matrix 1107 indicates two allelic-variant differences (labeled "A--C--") between the second local haplotype (labeled "haplotype 2") and the main continuous sequence of the reference genome at the nucleotide positions corresponding to the candidate alignments of the first pair 1103a of the p-terminal nucleotide read. Therefore, by comparing the nucleotides of the first pair 1103a of the p-terminal nucleotide read with the second local haplotype, the read alignment adjustment system 106 determines a second set of alignment score adjustments for the first pair 1103a, including: improving the main alignment score for the adenine and cytosine matched at the corresponding first and fourth nucleotide positions of the first pair 1103a with the second haplotype.
[0148] In addition, variant data matrix 1107 indicates two allelic-variant differences (labeled as “G--T--”) between the second local haplotype (labeled as “haplotype 2”) and the reference genome major continuous sequence at the nucleotide positions of the candidate pairings 1103b corresponding to the paired reads. Therefore, by comparing the nuclei of the second pair of paired nucleotide reads (page 26 / 43, 33 CN 121359207 A) with the second local haplotype, the read alignment adjustment system 106 determines a second set of alignment score adjustments for the second pair 1103b, including: reducing the main alignment score for the mismatch between cytosine in the second pair 1103b and guanine in the second haplotype at the first nucleus position of the second pair 1103b; and increasing the main alignment score for the thymine in the second pair 1103b that matches the second haplotype at the fourth nucleus position of the second pair 1103b.
[0149] Therefore, as shown in FIG11, the read alignment adjustment system 106 determines an alignment score adjustment for each locally distinct haplotype indicated by the identification bin (e.g., as further described above in conjunction with FIG3B, FIG5B and FIG10, for each corresponding locally distinct haplotype, based on a comparison of allele-variant differences indicated at the corresponding nucleotide base positions in the primary continuous sequence within the matrix 1107 of the nucleotide reads in the first pair 1103a and the second pair 1103b).
[0150] Additionally, as shown in FIG11, the series of actions 1100 includes action 1108, which sums the alignment score adjustments corresponding to the first pair 1103a and the second pair 1103b of the paired-end reads for each corresponding locally distinct haplotype. In some embodiments, for example, the read alignment adjustment system 106 adds the alignment score adjustment of the first pair 1103a to the alignment score adjustment of the second pair 1103b for each local distinct population haplotype indicated within the bin identified by action 1104 to determine the adjusted alignment score of the paired read relative to each identified local distinct haplotype. Furthermore, in one or more embodiments, the read alignment adjustment system 106 selects the first and second pairs of the paired read for alignment with the main continuous sequence or with the predicted alignment of the local distinct population haplotype based on the highest sum of the adjusted alignment scores corresponding to each candidate pair within a set of candidate alignments for the paired read.
[0151] Furthermore, in some embodiments, the read alignment adjustment system 106 utilizes a haplotype data structure (such as described above in conjunction with Figures 7 and 8) to determine the alignment score adjustment of candidate alignments for other types of nucleotide reads (such as transcriptome reads representing spliced RNA sequences) based on variant data encoded within that haplotype data structure.For example, Figure 12 illustrates an example implementation of how the read alignment adjustment system 106 uses haplotype data structure 1200 to determine the alignment score adjustment of RNA splicing alignment 1202 for transcriptome reads according to one or more embodiments.
[0152] As shown in Figure 12, RNA splicing alignment 1202 includes: a first candidate read alignment 1204a of approximately 50 nucleotides; a first splice sequence 1206a of approximately 11,250 nucleotides; a second candidate read alignment 1204b of approximately 50 nucleotides; a second splice sequence 1206b of approximately 13,450 nucleotides; and a third candidate read alignment 1204c of approximately 50 nucleotides. As shown in the figure, the read alignment adjustment system 106 identifies the shortest bin (shown as bin number 19 at level 15 in Figure 12) of the complete RNA splicing alignment (e.g., RNA splicing alignment 1202) contained within the haplotype data structure 1200 based on the initial alignment of the RNA splicing alignment 1202 with the major continuous sequence of the reference genome—regardless of whether this shortest bin is a base-level bin, a higher-level bin, or an offset higher-level bin within the haplotype data structure 1200. Therefore, the read alignment adjustment system 106 can determine the alignment score adjustment of the RNA splicing alignment 1202 relative to one or more locally distinct haplotypes identified within the selected bin (shown as bin number 19 at level 15 in Figure 12).
[0153] As further shown in Figure 12, the first candidate read alignment 1204a of the RNA splicing alignment 1202 contains nucleotide positions spanning two consecutive basal level bins (bins 1 and 2 in Figure 12). Therefore, as shown in Figure 12, the read alignment adjustment system 106 first identifies variant data (e.g., allele-variant differences between the major continuous sequence and locally different population haplotypes within the corresponding bin) in the first identified bin (shown as bin 1). Then, the read alignment adjustment system 106 identifies variant data indices in the corresponding bins at the next consecutive level (shown as bin 2), and then in the corresponding bins at the next consecutive level (shown as bin 3) to adjust the alignment score according to locally different haplotypes at each corresponding level. Continuing to the next basic level bin (bin 2) covering the nuclei of the first candidate read alignment 1204a, the read alignment adjustment system 106 further adjusts the alignment scores of the variant data within this bin (bin 2) as well as the locally distinct haplotypes identified by the variant data indexes in the corresponding bins (bins 5 and 6) at each consecutive level. After identifying and adjusting the alignment scores based on the variant data and variant data indexes in bins 1 through 6, the read alignment adjustment system 106 identifies the variant data indexes in the corresponding bin (bin 7) at the next consecutive level.
[0154] Furthermore, following a similar process, the alignment score adjustment for the second candidate read alignment 1204b of the RNA splicing alignment 1202 is determined. The read alignment adjustment system 106 identifies and adjusts the variant data and variant data indexes within bins 8 to 12 shown in FIG. 12. By identifying the variant data indexes within the next consecutive hierarchical bin (bin 13) corresponding to bins 7 and 12, the read alignment adjustment system 106 determines further alignment score adjustments based on locally distinct haplotypes identified within bin 13. Subsequently, the third candidate read alignment 1204c of the RNA splicing alignment 1202 is processed following a similar process. The read alignment adjustment system 106 identifies and adjusts the variant data and variant data indexes within bins 15 to 18. Note that, based on the initial alignment, the third candidate read alignment 1204c falls entirely within a single base hierarchical bin (bin 14). Finally, the read alignment adjustment system 106 identifies variant data indices within a higher-level bin (bin 19) corresponding to a complete RNA splicing alignment (e.g., RNA splicing alignment 1202) and determines one or more final alignment score adjustments relative to one or more locally distinct haplotypes identified within that corresponding bin.
[0155] As mentioned above, in some described embodiments, the read alignment adjustment system 106 achieves efficient and accurate mapping of nucleotide reads from a genome sample to genomic regions of a reference genome. To illustrate this, Figures 13A and 13B show experimental results of the read alignment adjustment system 106 determining predicted alignments of nucleotide reads using haplotype data structures (according to one or more disclosed embodiments). Specifically, Figure 13A shows comparative experimental results of identifying single nucleotide polymorphisms (SNPs) based on read alignments generated according to one or more embodiments, and Figure 13B shows comparative experimental results of identifying insertions or deletions (insertions / deletions) based on read alignments generated according to one or more embodiments.
[0156] As mentioned, Figure 13A provides comparative experimental results for identifying single nucleotide polymorphisms (SNPs) based on read alignments generated according to one or more embodiments and read alignments generated using existing sequencing systems. Specifically, Figure 13A includes an experimental results table showing the SNPs identified by read alignments using existing sequencing systems and read alignment adjustment system 106, reflected in false positives (SNP FP) and false negatives (SNP FN), where each set of three rows corresponds to a standard reference genome sample with identified ground truth variants. Specifically, the ground truth dataset used to generate the provided experimental results includes seven human reference genome samples—Genome in Bottle (GIAB) samples HG001, HG002, HG003, HG004, HG005, and HG007—each with corresponding ground truth variant detections.Furthermore, each row of the illustrated table provides experimental results for identifying SNPs in the corresponding reference sample dataset, reflected in the number of false negatives (FN) and / or false positives (FP). Specifically, the first row in each group of three rows provides experimental results using an augmented map reference genome with existing sequencing systems, while the second and third rows in each group of three rows provide experimental results for two corresponding implementations of the read alignment adjustment system 106 using a haplotype data structure implementation. Additionally, each group of three rows includes an indication of the percentage improvement in accuracy between the implementations represented by the corresponding first and third rows.
[0157] In fact, as shown in Figure 13A, compared to existing sequencing systems, the read alignment adjustment system 106 can efficiently predict read alignments of nucleotide reads from genomic samples and improves accuracy in identifying SNPs, as indicated by the number of false positives (FP) and false negatives (FN) identified in the provided experimental results.
[0158] Furthermore, Figure 13B provides comparative experimental results for identifying insertions or deletions (insertions and deletions) based on read alignments generated according to one or more implementations. Specifically, Figure 13B includes an experimental results table, where the first two rows correspond to existing sequencing systems that use augmented graph reference genomes for mapping and alignment, and the last three rows correspond to an exemplary implementation of read alignment adjustment system 106 that uses the type A data structure implementation scheme described on page 28 / 43 of the haplotype specification (CN 121359207). Additionally, each column of the experimental results table corresponds to a standard reference genome sample—Genome in Bottle (GIAB) samples HG001, HG002, HG005, and HG007—each with a corresponding ground truth variant detection. Specifically, the result row labeled “Graph euro 16” includes experimental results from existing sequencing systems using augmented graph reference genomes containing 16 haplotypes from European population samples; while the result row labeled “Graph global32” includes experimental results from existing sequencing systems using augmented graph reference genomes containing 32 haplotypes from global population samples. In addition, the result rows labeled “HapDB eurl6”, “HapDB global32” and “HapDB global128” respectively include experimental results of the implementation of the read segment alignment adjustment system 106 using a haplotype data structure containing 16 European haplotypes, 32 global haplotypes and 128 global haplotypes.
[0159] In fact, as shown in Figure 13B, compared to existing sequencing systems, the read alignment adjustment system 106 can efficiently predict read alignments of nucleotide reads from genomic samples and achieve comparable accuracy in identifying insertions and deletions, as indicated by the number of false positives (FPs) and false negatives (FNs) identified in the provided experimental results. Furthermore, as shown in the experimental results provided in Figure 13B, the read alignment adjustment system 106 can further improve the accuracy of identifying insertions and deletions within genomic samples as the number of haplotypes realized within the haplotype data structure increases (a capability that is typically unattainable by existing sequencing systems because the augmented reference genome becomes exceptionally large).
[0160] As mentioned above, in some embodiments, the read alignment adjustment system 106 uses an improved haplotype data structure (which encodes allele-variant differences between major continuous sequences and population haplotypes across a linear reference genome) to align and determine the adjusted alignment score of nucleotide reads from genomic samples. In contrast, some existing sequencing systems utilize a graph reference genome (including both a linear reference genome and graph augmentations representing selectable contiguous sequences) to align and determine the alignment scores of nucleotide reads from a genome sample. To further illustrate the different methods and corresponding computational efficiency savings, Figure 14A depicts an example of an existing sequencing system aligning nucleotide reads from a genome sample to a graph reference genome, and Figure 14B depicts an example implementation of a read alignment adjustment system 106, which, according to one or more embodiments, first performs an initial alignment of the same nucleotide read from a genome sample to a primary contiguous sequence (or other reference sequence), and then adjusts the alignment score of the initial alignment relative to the population haplotype encoded within the haplotype data structure.
[0161] As shown in Figure 14A, existing sequencing systems align nucleotide reads (shown as “Read”) from a genome sample to each of the linear reference sequence (shown as “Ref”) of the graph reference genome and the three selectable contiguous sequences (shown as “Alt1”, “Alt2”, and “Alt3”) of the graph reference genome. As shown in Figure 14A, existing sequencing systems must not only store the linear reference sequence and the candidate continuous sequence in memory as part of the graph reference genome, but also determine the individual alignment scores of the nucleotide reads with each of the linear reference sequence and the candidate continuous sequence.
[0162] Specifically, as shown in Figure 14A, existing sequencing systems determine the alignment scores of the nucleotide reads with the linear reference sequence (shown as "Ref"), the first candidate continuous sequence (shown as "Alt1"), the second candidate continuous sequence (shown as "Alt2"), and the third candidate continuous sequence ("Alt3") as 135, 135, 140, and 145, respectively.Existing sequencing systems partially determine such alignment scores by considering mismatches between nucleotide reads and a linear reference sequence or three different candidate contiguous sequences (marked with “X” in Figure 14A)—including mismatches caused by sequencing errors (identified as “error” in Figure 14A). Existing sequencing systems only identify the candidate alignment between the nucleotide read and the third candidate contiguous sequence (shown as “Alt3”) as the highest (maximum) alignment score among various candidate read alignments after scoring the candidate alignments individually against each of the linear reference sequence and the three different candidate contiguous sequences.
[0163] In contrast, as shown in Figure 14B, the read alignment adjustment system 106 aligns nucleotide reads (shown as “Read”) of a genomic sample with a primary contiguous sequence or other reference sequence (shown as “Ref”) and determines the primary alignment score of 135 for the candidate alignment between the nucleotide read and the reference sequence. In fact, the read alignment adjustment system 106 performs a single alignment operation for the nucleotide read in Figure 14B, rather than performing four separate alignment operations for the nucleotide read as in Figure 14A. As further indicated in Figure 14B, instead of storing and scoring alternative continuous sequences separately, the read alignment adjustment system 106 adjusts the primary alignment score (e.g., by "+5" or "-5") based on comparing the nucleotide read with allele-variant differences from different local population haplotypes (shown as "Hap1", "Hap2", and "Hap3"), where the allele-variant differences are encoded against a reference span within the haplotype data structure. Specifically, the read alignment adjustment system 106 (i) adjusts the primary alignment score up and down (displayed as "+5" and "-5") to an adjusted alignment score of 135 to account for allele-variant differences from haplotypes (displayed as "Hap1") from different populations in the first locality; (ii) increases the primary alignment score (displayed as "+5") to an adjusted alignment score of 140 to account for allele-variant differences from haplotypes (displayed as "Hap2") from different populations in the second locality; and (iii) increases the primary alignment score (displayed as "+5" and "+5") to an adjusted alignment score of 145 to account for allele-variant differences from haplotypes (displayed as "Hap3") from different populations in the third locality.
[0164] As further shown in FIG14B, the read alignment adjustment system 106 further (a) converts the adjusted alignment score into an alignment likelihood value, (b) adjusts the alignment likelihood value based on the corresponding allele frequency to generate an adjusted alignment likelihood value, and (c) converts the weighted sum of the adjusted alignment likelihood values into a candidate alignment replacement alignment score corresponding to the main continuous sequence position.Specifically, the read alignment adjustment system 106 converts the adjusted alignment score 135 into a first alignment likelihood value (displayed as "Lik1"), and adjusts the first alignment likelihood value based on the corresponding haplotype frequency of a specific allele combination (displayed as "Freq1") to generate a first adjusted alignment likelihood value (not shown). The read alignment adjustment system 106 also converts the adjusted alignment score 140 into a second alignment likelihood value (displayed as "Lik2"), and adjusts the second alignment likelihood value based on the corresponding haplotype frequency of a specific allele combination (displayed as "Freq2") to generate a second adjusted alignment likelihood value (not shown). Similarly, the read alignment adjustment system 106 converts the adjusted alignment score 145 into a third alignment likelihood value (displayed as "Lik3"), and adjusts the third alignment likelihood value based on the corresponding haplotype frequency of a specific allele combination (displayed as "Freq3") to generate a third adjusted alignment likelihood value (not shown). The read alignment adjustment system 106 further determines a weighted sum (logarithm) of the first, second, and third adjusted alignment likelihood values to generate a substitution alignment score or a final adjusted alignment score (shown as “Adj Score” in Figure 14B) for a specific candidate alignment of the nucleotide read with a major continuous sequence or other reference sequence (shown as “Ref”). As noted above, the terms “substitution alignment score” and “final adjusted alignment score” are used interchangeably.
[0165] As indicated by comparing Figures 14A and 14B, the highest alignment score determined by the existing sequencing system for a candidate alignment between the nucleotide read and a third candidate continuous sequence (shown as “Alt3”) is 145, while the read alignment adjustment system 106 determines a substitution alignment score of approximately 145 for a candidate alignment between the nucleotide read and the major continuous sequence, wherein the adjustment takes into account allele-variant differences in haplotypes of different local populations in the third region. However, the read alignment adjustment system 106 achieves highly similar alignment scores with better computational efficiency by avoiding computationally intensive multiple alignment and complete alignment scoring operations.
[0166] Turning now to Figures 15 and 16, these figures illustrate two example flowcharts of two corresponding series of actions for determining predicted read alignments for one or more nucleotide reads from a genomic sample according to one or more embodiments. While Figures 15 and 16 illustrate actions according to a specific embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the actions shown in Figures 15 and 16. The actions of Figures 15 and / or 16 may be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform the actions depicted in Figures 15 and / or 16.In some other embodiments, the system includes at least one processor and a nontransitory computer-readable medium including instructions that, when executed by one or more processors, cause the system to perform the actions of FIG15 and / or FIG16. Specification 30 / 43 Page 37 CN 121359207 A
[0167] As shown in FIG15, the series of actions 1500 includes: action 1502, which determines a set of candidate alignments between one or more nucleotide reads and a major continuous sequence; action 1504, which generates a major alignment score for the candidate alignments in the set of candidate alignments; action 1506, which identifies allele-variant differences between the major continuous sequence and one or more population haplotypes; action 1508, which generates one or more adjusted alignment scores based on the allele-variant differences; and action 1510, which selects a predicted read alignment from the set of candidate alignments based on one or more adjusted alignment scores.
[0168] As shown in FIG16, the series of actions 1600 includes: action 1602, which determines a reference span for candidate alignments of one or more nucleotide reads with a major continuous sequence; action 1604, which determines one or more alignment score adjustments based on variant data associated with the reference span; and action 1606, which selects a predicted alignment from a set of candidate alignments based on one or more alignment score adjustments.
[0169] For example, series of actions 1500 and / or series of actions 1600 may include actions that perform any of the operations described in the following clauses:
[0170] Clause 1. A computer-implemented method comprising:
[0171] determining a set of candidate alignments between one or more nucleotide reads from a genome sample and a major contiguous sequence at a corresponding set of genomic regions of a reference genome;
[0172] generating a major alignment score for the candidate alignments from the set of candidate alignments;
[0173] identifying one or more allele-variant differences between the major contiguous sequence and one or more population haplotypes corresponding to the candidate alignments;
[0174] generating one or more adjusted alignment scores based on the major alignment scores, based on comparisons of the one or more nucleotide reads with the one or more allele-variant differences; and
[0175] selecting, from the set of candidate alignments, the one or more nucleotide reads to be aligned with the major contiguous sequence or with a predicted read from a population haplotype of the one or more population haplotypes, based on the one or more adjusted alignment scores.
[0176] Clause 2. The computer-implemented method according to Clause 1 further includes:
[0177] generating a substitution alignment score for the candidate alignment based on the primary alignment score and the one or more adjusted alignment scores;
[0178] generating an additional substitution alignment score for additional candidate alignments in the set of candidate alignments; and
[0179] selecting the predicted read alignment of the one or more nucleotide reads based on comparing the substitution alignment score with a primary alignment score for one or more candidate alignments for one or more primary consecutive sequences, and the additional substitution alignment score for the additional candidate alignments of the set of candidate alignments.
[0180] Clause 3. The computer-implemented method according to any one of Clauses 1 to 2 further comprises:
[0181] for a paired-end read in one or more nucleotide reads, determining that a first candidate alignment of a first pair of the paired-end reads with the main continuous sequence is not within a threshold number of nucleotide bases relative to a second candidate alignment of a second pair of the paired-end reads with the main continuous sequence; and
[0182] based on the fact that the first candidate alignment is not within the threshold number of nucleotide bases relative to the second candidate alignment, identifying a second candidate alignment of the second pair relative to the first candidate alignment of the first pair within a predetermined search region.
[0183] Clause 4. The computer-implemented method according to any one of Clauses 1 to 3 further comprises identifying the one or more allele-variant differences by querying a haplotype data structure, the haplotype data structure comprising a set of bins corresponding to a set of nucleotide reference spans from a reference genome.
[0184] Clause 5. The computer-implemented method of Clause 4 further comprises:
[0185] querying the haplotype data structure by identifying a reference span in the set of reference spans comprising complete candidate alignments of the one or more nucleotide reads; and
[0186] identifying the one or more allele-variant differences stored in bins corresponding to the identified reference spans.
[0187] Clause 6. The computer-implemented method of Clause 5 further comprises identifying the one or more allele-variant differences stored in the bins by comparing the one or more nucleotide reads with allele-variant differences from one or more locally distinct population haplotype sequences stored in the bins corresponding to the identified reference spans.
[0188] Clause 7. The computer-implemented method according to any one of Clauses 1 to 6, further comprising:
[0189] querying a haplotype data structure by identifying a reference span in a set of reference spans, including a first candidate alignment of the first pair and a second candidate alignment of the second pair, for a first pair and a second pair of paired end reads in the one or more nucleotide reads;
[0190] generating a first adjusted alignment score for the first pair and a second adjusted alignment score for the second pair for each local different population haplotype encoded by the reference span, based on comparing the first pair and the second pair with the one or more allele-variant differences stored in a set of bins corresponding to the identified reference span;
[0191] summing the first adjusted alignment score of the first pair and the second adjusted alignment score of the second pair for each local different population haplotype encoded by the reference span; and
[0192] Based on the highest value of the sum of the adjusted alignment scores, the first pair is selected from the set of candidate alignments as a first predicted alignment with the main continuous sequence or with a local different population haplotype, and the second pair is selected as a second predicted alignment with the main continuous sequence or with a local different population haplotype.
[0193] Clause 8. The computer-implemented method according to Clause 7 further comprises:
[0194] generating a summation replacement alignment score for a subset of candidate alignments of the first pair and the second pair based on the primary alignment score of each local distinct population haplotype encoded by the reference span and the first adjusted alignment score and the second adjusted alignment score;
[0195] generating an additional summation replacement alignment score for an additional subset of candidate alignments of the set of candidate alignments of the first pair and the second pair; and
[0196] selecting the first predicted alignment and the second predicted alignment from the set of candidate alignments based on comparing the summation replacement alignment score with one or more primary alignment scores for one or more candidate alignments with one or more primary consecutive sequences, and the additional summation replacement alignment score for the additional subset of candidate alignments of the set of candidate alignments.
[0197] Clause 9. The computer-implemented method according to any one of Clauses 1 to 8 further comprises generating the one or more adjusted alignment scores at base positions where there is no allelic-variant difference, without comparing the nuclei of the one or more nucleotide reads with the nuclei of the one or more population haplotypes.
[0198] Clause 10. The computer-implemented method according to any one of Clauses 1 to 9 further comprises identifying the one or more allele-variant differences by comparing nucleotide bases within the one or more nucleotide reads with data representing one or more single nucleotide polymorphisms (SNPs) within the one or more population haplotypes corresponding to the respective genomic region.
[0199] Clause 11. The computer-implemented method according to any one of Clauses 1 to 10 further comprises identifying the one or more allele-variant differences by comparing the one or more nucleotide reads with data representing one or more insertions or deletions (insertions or deletions) within the one or more population haplotypes corresponding to the respective genomic region.
[0200] Clause 12. The computer-implemented method according to any one of Clauses 1 to 11 further comprises generating at least one of the one or more adjusted alignment scores based on the primary alignment score by:
[0201] determining that the one or more nucleotide reads comprise one or more haplotype nucleotide variants of locally different population haplotypes, the haplotypes differing from the primary contiguous sequence in the respective genomic region; and
[0202] improving the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads comprising the one or more haplotypes.
[0203] Clause 13. The computer-implemented method according to any one of Clauses 1 to 12 further comprises generating at least one of the one or more adjusted alignment scores based on the primary alignment score by:
[0204] determining that the one or more nucleotide reads include one or more reference nuclei of the primary continuous sequence, the reference nuclei differing from local haplotypes in the respective genomic region; and
[0205] reducing the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads including one or more reference nuclei.
[0206] Clause 14. The computer-implemented method according to any one of Clauses 1 to 13, further comprising:
[0207] generating the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a corresponding set of locally different population haplotypes corresponding to the corresponding genomic regions of the candidate alignments;
[0208] selecting the highest adjusted alignment score from the set of adjusted alignment scores as a replacement alignment score for the candidate alignments; and
[0209] selecting the predicted read alignment from the set of candidate alignments based on the replacement alignment score.
[0210] Clause 15. The computer-implemented method according to any one of Clauses 1 to 14, further comprising:
[0211] generating the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a corresponding set of locally different population haplotypes corresponding to the corresponding genomic regions of the candidate alignments;
[0212] converting the set of adjusted alignment scores into a set of alignment likelihood values;
[0213] adjusting the set of alignment likelihood values based on corresponding allele frequencies to generate a set of adjusted alignment likelihood values;
[0214] converting the sum of the set of adjusted alignment likelihood values into replacement alignment scores for the candidate alignments; and
[0215] selecting the predicted read alignment from the set of candidate alignments based on the replacement alignment scores.
[0216] Clause 16. The computer-implemented method according to any one of Clauses 1 to 15, further comprising adjusting at least one of the one or more adjusted alignment scores based on population allele frequencies of population haplotypes within the sample population.
[0217] Clause 17. The computer-implemented method according to any one of Clauses 1 to 16 further comprises generating the primary alignment score of the candidate alignment based on a given candidate alignment between the one or more nucleotide reads and a modified version of the primary continuous sequence, the modified version including one or more polybase codes representing one or more single nucleotide polymorphisms (SNPs) or representing one or more insertions or deletions (insertions or deletions).
[0218] Clause 18. A haplotype data structure comprising:
[0219] (a) a base level having a set of base level bins, the base level bins comprising:
[0220] a set of base level reference spans of major contiguous sequences of a reference genome, each base level reference span comprising a genomic region of a first length between corresponding genomic coordinates of the reference genome; and
[0221] variant data of nucleotide variants from locally distinct population haplotypes of corresponding groups, each locally distinct haplotype comprising one or more allele-variant differences of a unique group, the allele-variant differences being unique relative to other population haplotypes within the genomic region of the corresponding base level reference span; and
[0222] (b) a contiguous level having a set of higher-level bins, the higher-level bins comprising:
[0223] A set of higher-level reference spans of the primary continuous sequence, each higher-level reference span including an extended genomic region of a second length between corresponding genomic coordinates of the reference genome, the second length being longer than the first length; and
[0224] a variant data index referencing a combination of variant data from the corresponding basic-level bins in the set of basic-level bins.
[0225] Clause 19. The haplotype data structure of Clause 18, wherein the variant data of the set of basic-level bins includes data indications of single nucleotide polymorphisms (SNPs) and insertions or deletions (insertions or deletions) at the corresponding genomic coordinates of the primary continuous sequence.
[0226] Clause 20. The haplotype data structure of any one of Clauses 18 to 19, wherein the set of basic-level bins includes the variant data of nucleotide variants but does not include reference nucleobases of the primary continuous sequence.
[0227] Clause 21. A haplotype data structure according to any one of Clauses 18 to 20, wherein a population haplotype having the same nucleotide variant within a given basic hierarchical bin is encoded as a locally distinct population haplotype within the given basic hierarchical bin.
[0228] Clause 22. A haplotype data structure according to any one of Clauses 18 to 21, wherein each basic hierarchical bin in the set comprises a matrix comprising corresponding variant data representing allele-variant differences from locally distinct haplotypes and variant positions of the allele-variant differences.
[0229] Clause 23. A haplotype data structure according to any one of Clauses 18 to 22, wherein each corresponding extended genomic region in the set of higher-level reference spans corresponds to a pair of consecutive corresponding genomic regions in the consecutive basic hierarchical reference spans in the set of basic hierarchical reference spans.
[0230] Clause 24. A haplotype data structure according to any one of Clauses 18 to 23, wherein the consecutive levels of the haplotype data structure further include a set of offset higher-level bins, the offset higher-level bins comprising:
[0231] a set of offset higher-level reference spans of the primary consecutive sequence, each offset higher-level reference span comprising an offset extended genomic region of the second length between corresponding genomic coordinates of the reference genome,
[0232] wherein the offset extended genomic region corresponds to a pair of consecutive corresponding genomic regions in the set of basic-level reference spans, and
[0233] wherein the set of offset higher-level reference spans is offset from one of the basic-level reference spans in the set of basic-level reference spans relative to the set of higher-level reference spans.
[0234] Clause 25. The haplotype data structure according to Clause 24 further includes: Specification 34 / 43 Page 41 CN 121359207 A
[0235] at least one additional contiguous level having a higher-level reference bin with an additional set, the higher-level reference bin comprising:
[0236] a set of additional higher-level reference spans of the main contiguous sequence, each higher-level reference span comprising another extended genomic region of a third length between corresponding genomic coordinates of the reference genome, the third length being longer than the second length; and
[0237] a variant data index referencing a combination of variant data from the corresponding basic level bin in the set of basic level bins.
[0238] Clause 26. A computer-implemented method for implementing a haplotype data structure according to any one of Clauses 18 to 25, the computer-implemented method comprising:
[0239] determining a base-level reference span in a set of base-level reference spans including the one or more nucleotide reads from a genome sample and a candidate alignment in a set of candidate alignments to the major continuous sequence;
[0240] determining one or more alignment score adjustments corresponding to one or more locally distinct haplotypes in a corresponding genomic region of the base-level reference span, based on variant data from base-level bins in the set of base-level bins corresponding to the base-level reference span; and
[0241] selecting, from the set of candidate alignments, a predicted alignment of the one or more nucleotide reads to the major continuous sequence or to a population haplotype, based on the one or more alignment score adjustments.
[0242] Clause 27. The computer-implemented method according to Clause 26 further includes:
[0243] adjusting the candidate alignment scores based on the one or more alignment scores to generate replacement alignment scores;
[0244] generating additional replacement alignment scores for additional candidate alignments in the set of candidate alignments; and
[0245] selecting the predicted read alignment of the one or more nucleotide reads based on comparing the replacement alignment scores with the additional replacement alignment scores.
[0246] Clause 28. The computer-implemented method of Clause 27 further comprises:
[0247] determining, for candidate alignments in a set of candidate alignments between one or more nucleotide reads from a genomic sample and the major contiguous sequence, a higher-level reference span comprising complete candidate alignments of the one or more nucleotide reads;
[0248] determining a subset of locally distinct population haplotypes within a corresponding extended genomic region of the higher-level reference span based on variant data indexes of the higher-level bins in the set of higher-level bins corresponding to the higher-level reference span;
[0249] determining a first set of alignment score adjustments for one or more corresponding locally distinct population haplotypes in the subset of locally distinct population haplotypes based on variant data of the first base-level bin in the set of base-level bins corresponding to the first corresponding genomic region within the corresponding extended genomic region;
[0250] Based on variant data of the second basic hierarchical bin in the set of basic hierarchical bins corresponding to the second corresponding genomic region within the corresponding extended genomic region, determine a second set of alignment score adjustments for one or more corresponding local different population haplotypes in the subset of local different population haplotypes; and
[0251] based on the combination of the first set of alignment score adjustments and the second set of alignment score adjustments, select one or more nucleotide reads from the set of candidate alignments to align with the major continuous sequence or with the predicted population haplotype.
[0252] Clause 29. A method for implementing a haplotype data structure according to any one of Clauses 18 to 25, page 35 / 43, 42 CN 121359207 A, the computer-implemented method comprising:
[0253] determining a reference span including the complete candidate alignment of one or more nucleotide reads from a set of candidate alignments between a primary continuous sequence and a candidate alignment of one or more nucleotide reads from a genomic sample, the reference span being selected from the lowest level of the haplotype data structure, wherein the one or more nucleotide reads are included in a single reference span in the set of basic-level reference spans or the set of higher-level reference spans;
[0254] determining one or more alignment score adjustments corresponding to one or more locally different haplotypes in a corresponding genomic region of the reference span, based on variant data from one or more bins in the set of basic-level bins corresponding to the reference span; and
[0255] selecting, from the set of candidate alignments, a predicted alignment of the one or more nucleotide reads with the primary continuous sequence or with a population haplotype based on the one or more alignment score adjustments.
[0256] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which nucleic acids are attached to fixed positions in the array such that their relative positions do not change and in which the array is repeatedly imaged. Embodiments that obtain images in different color channels (e.g., with different markers used to distinguish one nucleobase type from another) are particularly suitable. In some embodiments, the process for determining the nucleotide sequence of the target nucleic acid (i.e., the nucleic acid polymer) can be automated. Preferred embodiments include sequencing-by-synthesis (SBS) technology.
[0257] SBS technology typically involves the enzymatic extension of a nascent nucleic acid chain by repeatedly adding nucleotides to the template strand. In conventional SBS methods, a single nucleotide monomer can be provided to the target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to the target nucleic acid in the presence of a polymerase during delivery.
[0258] SBS can utilize nucleotide monomers with a terminator motif or nucleotide monomers lacking any terminator motif. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is typically variable, and this number depends on the template sequence and the method of nucleotide delivery.For SBS technology utilizing nucleotide monomers with a terminator motif, the terminator may be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or the terminator may be reversible, as in sequencing methods developed by Solexa (now Illumina, Inc.).
[0259] SBS technology may utilize nucleotide monomers with or without a labeled motif. Thus, incorporation events can be detected based on: the characteristics of the label, such as the fluorescence of the label; the characteristics of the nucleotide monomer, such as molecular weight or charge; byproducts of the incorporated nucleotide, such as the release of pyrophosphate; and so on. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or alternatively, the two or more different labels may be indistinguishable under the detection technology used. For example, the different nucleotides present in the sequencing reagent may have different labels, and they may be distinguishable using appropriate optics, as exemplified by sequencing methods developed by Solexa (now Illumina, Inc.).
[0260] Preferred embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001), “Pyrosequencing sheds light on DNA sequencing.”, Genome Res. 11(1), 3-11; 43 CN 121359207 A Ronaghi, M., Uhlen, M. and Nyren, P. (1998), “A sequencing method based on real-time pyrophosphate.”, Science 281(5375). The disclosures of U.S. Patent Nos. 6,210,891, 6,258,568 and 6,274,320 are incorporated herein by reference in their entirety.In pyrosequencing, the released PPi can be detected by being immediately converted to ATP by adenosine triphosphate (ATP) sulfatase, and the level of ATP produced can be detected by photons generated by luciferase. The nucleic acid to be sequenced can be attached to a feature in the array, and the array can be imaged to capture the chemiluminescent signal generated by the incorporation of nucleotides at the feature. Images can be obtained after treating the array with a specific type of nucleotide (e.g., A, T, C, or G). Images obtained after adding each type of nucleotide will differ in which feature in the array is detected. These differences in the images reflect the different sequence contents of the feature on the array. However, the relative position of each feature will remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein. For example, images obtained after treating the array with each different type of nucleotide can be processed in the same manner illustrated herein for images obtained from different detection channels used in reversible terminator-based sequencing methods.
[0261] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, cleavable or photobleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method was commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, the disclosures of each of which are incorporated herein by reference. The availability of fluorescently labeled terminators (where the termination can be reversible and the fluorescent label can be cleaved) facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0262] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection markers can be removable, for example, by cleavage or degradation. Images can be captured after the markers are incorporated into the arrayed nucleic acid signatures. In a specific embodiment, each cycle involves simultaneously delivering four different nucleotide types to the array, and each nucleotide type has a spectrally different marker. Four images are then obtained, each using a detection channel selective for one of the four different markers. Alternatively, different nucleotide types can be added sequentially, and images of the array can be obtained between each addition step. In such embodiments, each image will show the nucleic acid signature with a specific type of nucleotide incorporated. Different signatures may or may not be present in different images due to the different sequence contents of each signature. However, the relative positions of the signatures will remain unchanged in the images.Images obtained by such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. After the image capture step, the markers can be removed, and the reversible terminator portion can be removed for subsequent cycles of nucleotide addition and detection. Removing these markers after they have been detected in a particular cycle and before subsequent cycles provides the advantage of reducing background signal and crosstalk between cycles. Examples of available marker and removal methods are described below.
[0263] In a particular embodiment, some or all of the nucleotide monomers may include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore linked to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other methods have separated the terminator chemistry from the cleavage of the fluorescent marker (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. described the development of reversible terminators that use small 3'-allyl groups to block elongation, but can be easily deblocked by short-time treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker, which can be easily cleaved by exposure to long-wavelength ultraviolet light for 30 seconds. Therefore, disulfide reduction or photocleavage can be used as the cleavable linker. Another method of termination is to use natural termination, which occurs after a bulk dye is placed on a dNTP. The presence of a charged bulk dye on the dNTP can act as an efficient terminator due to steric and / or electrostatic hindrance. The presence of an incorporation event prevents further incorporation unless the dye is removed. The cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entirety.
[0264] Additional exemplary SBS systems and methods that may be used in conjunction with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entirety.
[0265] Some embodiments may use fewer than four different labels to detect four different nucleotides. For example, the methods and systems described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232 can be used to perform SBS. As a first example, a pair of nucleotide types can be detected at the same wavelength, but can be distinguished based on the intensity difference of one member relative to the other, or based on a change in the presence or absence of a signal in one member that results in a significantly different signal compared to the detected signal of the other member of that pair (e.g., by chemical modification, photochemical modification, or physical modification). As a second example, three of four different nucleotide types can be detected under specific conditions, while the fourth nucleotide type lacks a marker that is detectable under those conditions or is minimally detectable under those conditions (e.g., minimal detection due to background fluorescence, etc.). The incorporation of the first three nucleotide types into the nucleic acid can be determined based on the presence of their respective signals, and the incorporation of the fourth nucleotide type into the nucleic acid can be determined based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type may include a marker detected in two different channels, while other nucleotide types are detected in no more than one channel. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary implementation combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type lacking a label or minimally detected in either channel (e.g., unlabeled dGTP).
[0266] Additionally, as described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, the first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0267] Some embodiments may utilize ligation-based sequencing technology. Such technologies utilize DNA ligases to incorporate oligonucleotides and recognize the incorporation of such oligonucleotides. Oligonucleotides typically have different labels associated with the identity of specific nucleotides in the sequence in which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show a nucleic acid feature incorporating a specific type of label. Different features may be present or absent in different images due to the different sequence content of each feature, but the relative positioning of the features will remain unchanged in the images. Images obtained by ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.
[0268] Some embodiments may utilize nanopore sequencing (Deamer, DW and Akeson, M., “Nanopores and nucleic acids: prospects for ultrarapid sequencing.”, Trends Biotechnol. 18, 147–151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis.”, Acc. Chem. Res. 35:817–825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin and J. A. Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope”, Nat. Mater., 2:611–615 (2003), the full text of which is incorporated herein by reference). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore may be a synthetic pore or a biomembrane protein, such as α-hemolysin. When the target nucleic acid passes through the nanopore, each base pair can be identified by measuring the fluctuations in the pore's conductivity. (US Patent Nos. 7,001,792; Soni, GV and Meller, “A. Progress toward ultrafast DNA sequencing using solid-state nanopores.”, Clin. Chem. 53, 1996–2001 (2007); Healy, K., “Nanopore-based single-molecule DNA analysis.”, Nanomed. 2, 459–481 (2007); Cockroft, SL, Chu, J., Amorin, M. and Ghadiri, M. R., “A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.”, J. Am. Chem. Soc. 130, 818–820 (2008), the full text of which is incorporated herein by reference).Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, the data can be processed like images, based on the exemplary processing of optical images and other images described herein.
[0269] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interaction between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414 (each of which is incorporated herein by reference), or by using a zero-mode waveguide, as described, for example, in U.S. Patent No. 7,315,019 (each of which is incorporated herein by reference), and by using fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference). Illumination can be limited to a volume of approximately 1 / 2000 that surrounds the surface-tethered polymerase, allowing for observation of the incorporation of fluorescently labeled nucleotides against a low background (Levene, M. J. et al., “Zero-mode vegetatives for single-molecule analysis at high concentrations.”, Science 299, 682-686 (2003); Lundquist, PM et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures.”, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008)). (Year), the full text of these publications is incorporated herein by reference. Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0270] Some SBS implementations include the detection of protons released during nucleotide incorporation into the elongation product.For example, sequencing based on the detection of released protons can use electrical detectors and related technologies commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies) or sequencing methods and systems described in US 2009 / 0026082 A1, US 2009 / 0127589 A1, specification 39 / 43 pages 46 CN 121359207 A US 2010 / 0137143 A1 or US 2010 / 0282617 A1, each of which is incorporated herein by reference. The method described herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates for proton detection. More specifically, the method described herein can be used to generate a population of amplicon clones for proton detection.
[0271] The above-described SBS method can advantageously be performed in a variety of formats, allowing simultaneous manipulation of multiple different target nucleic acids. In specific embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways. In implementations using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acid can typically bind to the surface in a spatially distinguishable manner. The target nucleic acid can bind via direct covalent attachment, attachment to beads or other particles, or binding to polymerases or other molecules attached to the surface. The array may include a single copy of the target nucleic acid at each site (also referred to as a feature region), or multiple copies having the same sequence may be present at each site or feature region. Multiple copies can be generated by amplification methods such as, for example, bridging amplification or emulsion PCR, which are described in further detail below.
[0272] The methods described herein may use arrays having feature portions at any of a variety of densities, including, for example, at least about 10 feature portions / cm², 100 feature portions / cm², 500 feature portions / cm², 1,000 feature portions / cm², 5,000 feature portions / cm², 10,000 feature portions / cm², 50,000 feature portions / cm², 100,000 feature portions / cm², 1,000,000 feature portions / cm², 5,000,000 feature portions / cm², or higher.
[0273] The advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Therefore, this disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Therefore, the integrated system disclosed herein may include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc.Flow cells in an integrated system can be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference. As illustrated with respect to flow cells, one or more fluid components of the integrated system can be used for amplification and detection methods. Taking a nucleic acid sequencing implementation as an example, one or more fluid components of the integrated system can be used for the amplification methods set forth herein as well as for delivering sequencing reagents in sequencing methods such as those illustrated above. Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of generating amplified nucleic acids and also determining nucleic acid sequences include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and the device described in US Serial No. 13 / 273,666, which is incorporated herein by reference. The sequencing systems described above sequence nucleic acid polymers present in samples received by the sequencing device, as further described above.
[0274] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples containing low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or may include forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel. The nucleic acid sample may be a purified sample or a lysate containing crude DNA, such as derived from oral swabs, paper, fabric, or other substrates that can be impregnated with saliva, blood, or other bodily fluids. Thus, in some embodiments, the nucleic acid sample may include a small amount of DNA (such as genomic DNA) or fragmented portions of DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from the victim's hair, skin, tissue samples, autopsy, or remains. In some embodiments, the nucleic acid containing one or more target sequences may be obtained from a dead animal or human. In some implementations, as described on pages 40 / 43 of CN 121359207 A, the target sequence may include nucleic acids obtained from non-human DNA (such as microbial, plant, or insect DNA). In some embodiments, the target sequence or amplified target sequence is directed towards human identification purposes. In some embodiments, this disclosure relates throughout to methods for identifying characteristics of forensic samples.In some embodiments, this disclosure relates throughout to a human identification method using one or more target-specific primers disclosed herein or one or more target-specific primers designed using the primer design standards outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence may be amplified using any one or more target-specific primers disclosed herein or using the primer standards outlined herein.
[0275] Components of the read alignment adjustment system 106 may include software, hardware, or both. For example, components of the read alignment adjustment system 106 may include one or more instructions stored on a computer-readable storage medium of one or more computing devices (e.g., client device 114) and executable by a processor of one or more computing devices. When executed by one or more processors, the computer-executable instructions of the read alignment adjustment system 106 may cause the computing device to perform the bubble detection method described herein. Alternatively, components of the read alignment adjustment system 106 may include hardware, such as a dedicated processing device for performing a function or group of functions. Additionally or alternatively, components of the read alignment adjustment system 106 may include a combination of computer-executable instructions and hardware.
[0276] Furthermore, components of the read alignment adjustment system 106 that perform the functions described herein with respect to the read alignment adjustment system 106 may be implemented, for example, as part of a standalone application, as a module of an application, as a plugin of an application, as one or more library functions that can be called by other applications, and / or as a cloud computing model. Thus, components of the read alignment adjustment system 106 may be implemented as part of a standalone application on a personal computing device or mobile device. Additionally or alternatively, components of the read alignment adjustment system 106 may be implemented in any application that provides sequencing services, including but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumina", "BaseSpace", "DRAGEN", and "TruSight" are registered trademarks or trademarks of Illumina, Inc. in the U.S. and / or other countries.
[0277] As discussed in more detail below, embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware such as, for example, one or more processors and system memory. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures.Specifically, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more processes described herein.
[0278] A computer-readable medium may be any available medium accessible by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of this disclosure may include at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0279] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSDs) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that is accessible by a general-purpose or special-purpose computer. Specification 41 / 43 pages 48 CN 121359207 A
[0280] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired or wireless), the computer appropriately regards that connection as a transmission medium. The transmission medium may include a network and / or a data link that can be used to carry desired program code means in the form of computer-executable instructions or data structures and that is accessible by a general-purpose or special-purpose computer. The aforementioned combinations should also be included within the scope of computer-readable media.
[0281] Furthermore, upon arrival at various computer system components, program code in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., NIC) and then eventually transferred to computer system RAM and / or to a less lossy computer storage medium (device) at the computer system.Therefore, it should be understood that non-transitory computer-readable storage media (devices) may be included in computer system components that also (or even primarily) utilize transmission media.
[0282] Computer-executable instructions include, for example, instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform certain functions or capabilities. In some embodiments, the computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary numbers, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or methodological actions, it should be understood that the subject matter as defined in the appended claims is not necessarily limited to the described features or actions. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims.
[0283] Those skilled in the art will understand that this disclosure can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, portable computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, etc. This disclosure can also be practiced in distributed system environments, where both local and remote computer systems perform tasks via network links (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may reside in both local and remote memory storage devices.
[0284] Embodiments of this disclosure can also be implemented in cloud computing environments. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. A shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0285] Cloud computing models can consist of various features, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, metered services, etc. Cloud computing models can also exhibit various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). Cloud computing models can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and in the claims, a “cloud computing environment” is an environment in which cloud computing is employed.
[0286] Figure 17 illustrates a block diagram of a computing device 1700 that can be configured to perform one or more of the processes described above.It should be understood that one or more computing devices (such as computing device 1700) may implement read alignment adjustment system 106 and sequencing system 104. As shown in FIG17, computing device 1700 may include processor 1702, memory 1704, storage device 1706, I / O interface 1708, and communication interface 1710, which may be communicatively coupled through communication infrastructure 1712. In some embodiments, computing device 1700 may include fewer or more components than those shown in FIG17. The following paragraphs describe the components of computing device 1700 shown in FIG17 in more detail.
[0287] In one or more embodiments, processor 1702 includes hardware for executing instructions such as those constituting a computer program. By way of example and not limitation, in order to execute instructions for dynamically modifying a workflow, processor 1702 may retrieve (or obtain) instructions from internal registers, internal cache, memory 1704, or storage device 1706, and decode and execute them. Memory 1704 may be volatile or non-volatile memory for storing data, metadata, and programs executed by the processor. Storage device 1706 includes storage means for storing data or instructions for performing the methods described herein, such as a hard disk, flash drive, or other digital storage device.
[0288] I / O interface 1708 allows a user to provide input to computing device 1700, receive output from the computing device, and otherwise transfer data to and receive data from the computing device. I / O interface 1708 may include a mouse, keypad or keyboard, touchscreen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O interfaces. I / O interface 1708 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O interface 1708 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be servicing a particular implementation.
[0289] Communication interface 1710 may include hardware, software, or both. In any case, communication interface 1710 may provide one or more interfaces for communication (such as, for example, packet-based communication) between computing device 1700 and one or more other computing devices or networks.By way of example and not limitation, communication interface 1710 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi.
[0290] Additionally, communication interface 1710 may facilitate communication with various types of wired or wireless networks. Communication interface 1710 may also facilitate communication using various communication protocols. Communication infrastructure 1712 may also include hardware, software, or both that couple components of computing device 1700 to each other. For example, communication interface 1710 may use one or more networks and / or protocols to enable multiple computing devices connected through a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. For illustration, the sequencing process may allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0291] In the foregoing description, this disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of this disclosure have been described with reference to the details discussed herein, and various embodiments are illustrated in the accompanying drawings. The above description and figures are illustrative of this disclosure and should not be construed as limiting it. Numerous specific details have been described to provide a thorough understanding of various embodiments of this disclosure.
[0292] This disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The embodiments described herein should be considered in all respects as exemplary rather than restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other or in parallel with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description. All changes within the equivalent meaning and scope of the claims are to be included within their scope.Instruction manual page 43 / 43, 50 CN 121359207 A, Figure 1; Instruction manual figure 1 / 22 page, 51 CN 121359207 A, Figure 2; Instruction manual figure 2 / 22 page, 52 CN 121359207 A, Figure 3A; Instruction manual figure 3 / 22 page, 53 CN 121359207 A, Figure 3B; Instruction manual figure 4 / 22 page, 54 CN 121359207 A, Figure 4; Instruction manual figure 5 / 22 page, 55 CN 121359207 A, Figure 5A; Instruction manual figure 6 / 22 page, 56 CN 121359207 A, Figure 5B; Instruction manual figure 7 / 22 page, 57 CN 121359207 A, Figure 6; Instruction manual figure 8 / 22 page, 58 CN 121359207 A, Figure 7; Instruction manual figure 9 / 22 page, 59 CN 121359207 A, Figure 8 Instruction manual illustrations, page 10 / 22, figure 9A (CN 121359207 A); page 11 / 22, figure 9B (CN 121359207 A); page 12 / 22, figure 10 (CN 121359207 A); page 13 / 22, figure 11 (CN 121359207 A); page 14 / 22, figure 12 (CN 121359207 A); page 15 / 22, figure 13A (CN 121359207 A); page 16 / 22, figure 13B (CN 121359207 A); page 17 / 22, figure 14A (CN 121359207 A); page 18 / 22, figure 14B (CN 121359207 A); page 19 / 22 Page 69 CN 121359207 A Figure 15 Appendix to the Specification 20 / 22 Page 70 CN 121359207 A Figure 16 Appendix to the Specification 21 / 22 Page 71 CN 121359207 A Figure 17 Appendix to the Specification 22 / 22 Page 72 CN 121359207 A.
Claims
1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: determine a set of candidate alignments between one or more nucleotide reads from a genomic sample and a primary contiguous sequence at a respective set of genomic regions of a reference genome; generate a primary alignment score for a candidate alignment from the set of candidate alignments; identify one or more allele-variant differences between the primary contiguous sequence and one or more population haplotypes corresponding to a respective genomic region of the candidate alignment; generate one or more adjusted alignment scores from the primary alignment score based on comparing the one or more nucleotide reads to the one or more allele-variant differences; and select, based on the one or more adjusted alignment scores, the one or more nucleotide reads to align to the primary contiguous sequence or to a predicted read from a population haplotype from the one or more population haplotypes from the set of candidate alignments.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate a replacement alignment score for the candidate alignment based on the primary alignment score and the one or more adjusted alignment scores; generate an additional replacement alignment score for an additional candidate alignment from the set of candidate alignments; and select the predicted read alignment of the one or more nucleotide reads based on comparing the replacement alignment score to one or more primary alignment scores for one or more candidate alignments to one or more primary contiguous sequences and the additional replacement alignment score for the additional candidate alignment from the set of candidate alignments.
3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: for a paired-end read of the one or more nucleotide reads, determine that a first mate of the paired-end read is not within a threshold number of nucleobases of a first candidate alignment of the primary contiguous sequence relative to a second mate of the paired-end read being not within the threshold number of nucleobases of a second candidate alignment of the primary contiguous sequence; and based on the first candidate alignment being not within the threshold number of nucleobases of the second candidate alignment, identify the second candidate alignment of the second mate within a predetermined search region relative to the first candidate alignment of the first mate.
4. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to identify the one or more allele-variant differences by querying a haplotype data structure comprising a set of bins corresponding to a set of nucleobase reference spans from a reference genome.
5. The system of claim 4, further comprising instructions that, when executed by the at least one processor, cause the system to: querying the haplotype data structure by identifying a reference span in the set of reference spans that includes a complete candidate alignment of the one or more nucleotide reads; and identifying the one or more allele-variant differences stored within the bin in the set of bins corresponding to the identified reference span.
6. The system of claim 5, further comprising instructions that, when executed by the at least one processor, cause the system to identify the one or more allele-variant differences stored within the bin by comparing the one or more nucleotide reads to allele-variant differences from one or more locally distinct population haplotypes stored within the bin corresponding to the identified reference span.
7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: for a first pair and a second pair of paired-end reads in the one or more nucleotide reads, query a haplotype data structure by identifying a reference span in a set of reference spans that includes a first candidate alignment of the first pair and a second candidate alignment of the second pair; for each locally distinct population haplotype encoded by the reference span, generate a first adjusted alignment score for the first pair and a second adjusted alignment score for the second pair based on comparing the first pair and the second pair to the one or more allele-variant differences stored within a bin in a set of bins corresponding to the identified reference span; for each locally distinct population haplotype encoded by the reference span, add the first adjusted alignment score for the first pair and the second adjusted alignment score for the second pair; and select, from the set of candidate alignments, a first predicted alignment of the first pair to the primary contiguous sequence or to a locally distinct population haplotype and a second predicted alignment of the second pair to the primary contiguous sequence or to a locally distinct population haplotype based on a highest value of the adjusted alignment score sums.
8. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to: generate a sum-replacement alignment score for a subset of candidate alignments of the first pair and the second pair based on the primary alignment score for each locally distinct population haplotype encoded by the reference span and the first adjusted alignment score and the second adjusted alignment score; generate an additional sum-replacement alignment score for an additional subset of the set of candidate alignments of the first pair and the second pair; and select the first predicted alignment and the second predicted alignment from the set of candidate alignments based on comparing the sum-replacement alignment score to one or more primary alignment scores for one or more candidate alignments to one or more primary contiguous sequences and the additional sum-replacement alignment score for the additional subset of the set of candidate alignments.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the one or more adjusted alignment scores at base positions without allelic-variant differences without comparing a nucleobase of the one or more nucleotide reads to a nucleobase of the one or more population haplotypes.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to identify the one or more allelic-variant differences by comparing a nucleobase within the one or more nucleotide reads to data representative of one or more single nucleotide polymorphisms (SNPs) within the one or more population haplotypes corresponding to the respective genomic region.
11. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to identify the one or more allelic-variant differences by comparing the one or more nucleotide reads to data representative of one or more insertions or deletions (indels) within the one or more population haplotypes corresponding to the respective genomic region.
12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate at least one of the one or more adjusted alignment scores from the primary alignment score by: determining that the one or more nucleotide reads include one or more haplotype nucleotide variants of a locally different population haplotype, the haplotype nucleotide variants differing from the primary contiguous sequence in the respective genomic region; and increasing the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads including the one or more haplotype nucleotide variants.
13. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate at least one of the one or more adjusted alignment scores from the primary alignment score by: determining that the one or more nucleotide reads include one or more reference nucleobases of the primary contiguous sequence, the reference nucleobases differing from a locally different population haplotype in the respective genomic region; and decreasing the primary alignment score to generate the at least one adjusted alignment score based on the one or more nucleotide reads including the one or more reference nucleobases.
14. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the one or more adjusted alignment scores by generating a respective set of adjusted alignment scores of a respective set of locally different population haplotypes corresponding to the respective genomic region of the candidate alignment; selecting, from the set of adjusted alignment scores, a highest adjusted alignment score as a replacement alignment score for the candidate alignment; and selecting, based on the replacement alignment score, the predicted read alignment from the set of candidate alignments.
15. The system of claim 1, the system further comprising instructions that, when executed by the at least one processor, cause the system to: generate the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a respective set of local distinct population haplotypes corresponding to the respective genomic region of the candidate alignment; convert the set of adjusted alignment scores to a set of alignment likelihoods; adjust the set of alignment likelihoods based on corresponding allele frequencies to generate a set of adjusted alignment likelihoods; convert a sum of the set of adjusted alignment likelihoods to a replacement alignment score for the candidate alignment; and select, based on the replacement alignment score, the predicted read alignment from the set of candidate alignments.
16. The system of claim 1, the system further comprising instructions that, when executed by the at least one processor, cause the system to adjust at least one of the one or more adjusted alignment scores based on population allele frequencies of population haplotypes within a sample population.
17. The system of claim 1, the system further comprising instructions that, when executed by the at least one processor, cause the system to generate the primary alignment score for a given candidate alignment between the one or more nucleotide reads and a modified version of the primary contiguous sequence, the modified version including one or more single nucleotide polymorphisms (SNPs) or one or more multi-base encodings representing one or more insertions or deletions (indels).
18. A non-transitory computer-readable medium comprising a haplotype data structure, the haplotype data structure comprising: (a) a base level having a set of base level bins, the base level bins comprising: a set of base level reference spans of a reference genome primary contiguous sequence, each base level reference span comprising a first length of genomic region between respective genomic coordinates of the reference genome; and variant data for nucleotide variants from a respective set of local distinct population haplotypes, each local distinct haplotype comprising a unique set of one or more allele-variant differences that are unique with respect to other population haplotypes within the genomic region of a respective base level reference span; and (b) a contiguous level having a set of higher level bins, the higher level bins comprising: a set of higher level reference spans of the primary contiguous sequence, each higher level reference span comprising a second length of extended genomic region between respective genomic coordinates of the reference genome, the second length being longer than the first length; and a variant data index referencing a combination of the variant data from a corresponding base level bin of the set of base level bins.
19. A system for predicting a read alignment, the system comprising: a memory; and at least one processor coupled to the memory and configured to execute instructions stored in the memory to perform operations comprising: receiving a set of one or more nucleotide reads and a reference genome primary contiguous sequence; generating a set of candidate alignments between the one or more nucleotide reads and the reference genome primary contiguous sequence; generating, for each candidate alignment, a primary alignment score based on a set of local distinct population haplotypes corresponding to a respective genomic region of the candidate alignment; selecting, from the set of candidate alignments, a highest primary alignment score as a replacement alignment score for the candidate alignment; and selecting, based on the replacement alignment score, a predicted read alignment from the set of candidate alignments.
20. The system of claim 19, the system further comprising instructions that, when executed by the at least one processor, cause the system to: generate the one or more adjusted alignment scores by generating a set of adjusted alignment scores for a respective set of local distinct population haplotypes corresponding to the respective genomic region of the candidate alignment; convert the set of adjusted alignment scores to a set of alignment likelihoods; adjust the set of alignment likelihoods based on corresponding allele frequencies to generate a set of adjusted alignment likelihoods; convert a sum of the set of adjusted alignment likelihoods to a replacement alignment score for the candidate alignment; and select, based on the replacement alignment score, the predicted read alignment from the set of candidate alignments.
21. The system of claim 19, the system further comprising instructions that, when executed by the at least one processor, cause the system to adjust at least one of the one or more adjusted alignment scores based on population allele frequencies of population haplotypes within a sample population.
22. The system of claim 19, the system further comprising instructions that, when executed by the at least one processor, cause the system to generate the primary alignment score for a given candidate alignment between the one or more nucleotide reads and a modified version of the primary contiguous sequence, the modified version including one or more single nucleotide polymorphisms (SNPs) or one or more multi-base encodings representing one or more insertions or deletions (indels).
23. A non-transitory computer-readable medium comprising a haplotype data structure, the haplotype data structure comprising: (a) a base level having a set of base level bins, the base level bins comprising: a set of base level reference spans of a reference genome primary contiguous sequence, each base level reference span comprising a first length of genomic region between respective genomic coordinates of the reference genome; and variant data for nucleotide variants from a respective set of local distinct population haplotypes, each local distinct haplotype comprising a unique set of one or more allele-variant differences that are unique with respect to other population haplotypes within the genomic region of a respective base level reference span; and (b) a contiguous level having a set of higher level bins, the higher level bins comprising: a set of higher level reference spans of the primary contiguous sequence, each higher level reference span comprising a second length of extended genomic region between respective genomic coordinates of the reference genome, the second length being longer than the first length; and a variant data index referencing a combination of the variant data from a corresponding base level bin of the set of base level bins.
19. The non-transitory computer-readable medium of claim 18, wherein the variant data of the set of base-level tiered bins includes data indications of single nucleotide polymorphisms (SNPs) and insertions or deletions (indels) at respective genomic coordinates of the primary contiguous sequence.
20. The non-transitory computer-readable medium of claim 18, wherein the set of base-level tiered bins includes the variant data of nucleotide variants but not reference nucleobases of the primary contiguous sequence.
21. The non-transitory computer-readable medium of claim 18, wherein a population haplotype having identical nucleotide variants within a given base-level tiered bin is encoded as one local distinct population haplotype within the given base-level tiered bin.
22. The non-transitory computer-readable medium of claim 18, wherein each base-level tiered bin of the set of base-level tiered bins includes a matrix including corresponding variant data representing allele-variant differences from local distinct haplotypes and variant positions of the allele-variant differences.
23. The non-transitory computer-readable medium of claim 18, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: for a candidate alignment of a set of candidate alignments between one or more nucleotide reads from a genomic sample and the primary contiguous sequence, determine a base-level reference span of the set of base-level reference spans that includes the one or more nucleotide reads; based on variant data from a base-level tiered bin of the set of base-level tiered bins corresponding to the base-level reference span, determine one or more alignment score adjustments corresponding to one or more local distinct haplotypes within a respective genomic region of the base-level reference span; and based on the one or more alignment score adjustments, select a predicted alignment of the one or more nucleotide reads to the primary contiguous sequence or to a population haplotype from the set of candidate alignments.
24. The non-transitory computer-readable medium of claim 23, further comprising instructions that, when executed by the at least one processor, cause the at least one processor to: generate a replacement alignment score for the candidate alignment based on the one or more alignment score adjustments; generate an additional replacement alignment score for an additional candidate alignment of the set of candidate alignments; and select the predicted read alignment of the one or more nucleotide reads based on comparing the replacement alignment score to the additional replacement alignment score.
25. The non-transitory computer-readable medium of claim 18, wherein each respective extended genomic region of the set of higher-level reference spans corresponds to a pair of consecutive respective genomic regions of consecutive base-level reference spans of the set of base-level reference spans.
26. The non-transitory computer-readable medium of claim 18, wherein the contiguous hierarchy of haplotype data structures further comprises a set of offset higher hierarchy bins, the offset higher hierarchy bins comprising: a set of offset higher hierarchy reference spans of the primary contiguous sequence, each offset higher hierarchy reference span comprising an offset extended genomic region of the second length between respective genomic coordinates of the reference genome, wherein the offset extended genomic region corresponds to a pair of contiguous respective genomic regions in the set of base hierarchy reference spans, and wherein the set of offset higher hierarchy reference spans offsets a base hierarchy reference span in the set of base hierarchy reference spans relative to the set of higher hierarchy reference spans.
27. The non-transitory computer-readable medium of claim 26, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: for a candidate alignment in a set of candidate alignments between one or more nucleotide reads from a genomic sample and the primary contiguous sequence, determine a higher hierarchy reference span in the set of higher hierarchy reference spans comprising a complete candidate alignment of the one or more nucleotide reads; determine a subset of local distinct population haplotypes within a respective extended genomic region of the higher hierarchy reference span according to variant data indices of higher hierarchy bins in the set of higher hierarchy bins corresponding to the higher hierarchy reference span; determine a first set of alignment score adjustments for one or more respective local distinct population haplotypes in the subset of local distinct population haplotypes according to variant data of first base hierarchy bins in the set of base hierarchy bins corresponding to first respective genomic regions within the respective extended genomic region; determine a second set of alignment score adjustments for one or more respective local distinct population haplotypes in the subset of local distinct population haplotypes according to variant data of second base hierarchy bins in the set of base hierarchy bins corresponding to second respective genomic regions within the respective extended genomic region; and select a predicted alignment of the one or more nucleotide reads to the primary contiguous sequence or to a population haplotype from the set of candidate alignments based on a combination of the first set of alignment score adjustments and the second set of alignment score adjustments.
28. The non-transitory computer-readable medium of claim 26, wherein the haplotype data structure further comprises: at least one additional contiguous hierarchy having an additional set of higher hierarchy reference bins, the higher hierarchy reference bins comprising: an additional set of higher hierarchy reference spans of the primary contiguous sequence, each higher hierarchy reference span comprising another extended genomic region of a third length between respective genomic coordinates of the reference genome, the third length being longer than the second length; and a variant data index referencing a combination of the variant data from corresponding base hierarchy bins in the set of base hierarchy bins.
29. The non-transitory computer-readable medium of claim 18, further comprising instructions that, when executed by at least one processor, cause the at least one processor to: for a candidate alignment of a set of candidate alignments between one or more nucleotide reads from a genomic sample and the primary contiguous sequence, determine a reference span comprising a complete candidate alignment of the one or more nucleotide reads, the reference span selected from a lowest level of the haplotype data structure in which the one or more nucleotide reads are included in a single reference span of the set of base level reference spans or the set of higher level reference spans; determine, based on variant data from one or more bins of the set of base level bins corresponding to the reference span, one or more alignment score adjustments corresponding to one or more local distinct haplotypes within a respective genomic region of the reference span; and select, based on the one or more alignment score adjustments, a predicted alignment of the one or more nucleotide reads to the primary contiguous sequence or to a population haplotype from the set of candidate alignments.