A personalized haplotype database for improved mapping and alignment of nucleotide reads and improved genotype calling
Patent Information
- Application Number
- HK62026126123
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-02-25
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202580003270.4 (22) Application Date 2025.02.26 (30) Priority Data 63 / 558754 2024.02.28 US (85) PCT International Application Entering National Phase Date 2025.12.18 (86) PCT International Application Application Data PCT / US2025 / 017424 2025.02.26 (87) PCT International Application Publication Data WO2025 / 184234 EN 2025.09.04 (71) Applicant Inminar Inc. Address California, USA (72) Inventors D.J. Andrews M. Lule T. Moon (74) Patent Agency Beijing Panhua Weiye Intellectual Property Agency Co., Ltd. 11280 Patent Attorney Wang Bo (51) Int.Cl. G16B 30 / 10 (2006.01) (54) Invention Title Personalized Haplotype Database for Improved Mapping and Alignment of Nucleotide Reads and Improved Genotype Detection (57) Abstract This disclosure describes a method, non-transitory computer-readable medium, and system for achieving improved mapping and alignment of nucleotide reads with genomic regions of a reference genome. For example, the disclosed system can identify nucleotide reads and candidate population haplotypes within a set of reference spans of a reference genome, and generate haplotype group scores for groups of candidate population haplotypes within that set of reference spans. Based on the haplotype group scores, the disclosed system can generate a personalized haplotype database including population haplotype subgroups from a population haplotype database, and use the personalized haplotype database to determine one or more personalized alignments of a set of nucleotide reads from a genomic sample.Claims 4 pages, Description 35 pages, Drawings 13 pages, CN 121359209 A 2026.01.16 CN 1 21 35 92 09 A 1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: identify a set of nucleotide reads from a genomic sample and candidate population haplotypes from a population haplotype database within a set of reference spans of a reference genome; and, for the set of reference spans of the reference genome, generate haplotype group scores for groups of the candidate population haplotypes based on comparing the set of nucleotide reads and the candidate population haplotypes; For the genomic sample and based on the haplotype group score, a personalized haplotype database is generated, the personalized haplotype database including population haplotype subgroups from the population haplotype database within the set of reference spans; and using the personalized haplotype database, one or more personalized alignments are determined between the set of nucleotide reads and corresponding genomic regions of the reference genome. 2. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: identify one or more distinct k-mers for each candidate population haplotype within a given reference span in the set of reference spans; and, for the given reference span, generate a corresponding haplotype group score in the haplotype group score within the given reference span based on comparing the k-mers of the set of nucleotide reads with the one or more distinct k-mers of the corresponding candidate population haplotypes in the set of candidate population haplotypes. 3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: determine an initial alignment of the set of nucleotide reads with corresponding genomic regions of the reference genome; identify subgroups of nucleotide reads within the set of nucleotide reads, the subgroups of nucleotide reads being aligned with corresponding reference spans in the set of reference spans according to the initial alignment; and generate haplotype scores for the set of reference spans by comparing the subgroups of nucleotide reads with the candidate population haplotypes within the corresponding reference spans in the set of reference spans.4. The system of claim 3, further comprising instructions that, when executed by the at least one processor, cause the system to: identify allelic variant differences between a population haplotype and a primary continuous sequence at a corresponding genomic region of the reference genome, as indicated in the population haplotype database; and determine one or more initial alignments of the initial alignments of the set of nucleotide reads by re-scoring the candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on the allelic variant differences. 5. The system of claim 3, further comprising instructions, which, when executed by the at least one processor, cause the system to: identify one or more population haplotypes for each nucleotide read in the set of nucleotide reads based on comparisons with haplotype variants within the corresponding genomic regions of the initial alignment; and generate an alignment data file comprising: the initial alignment of the set of nucleotide reads; an alignment score corresponding to the initial alignment; and the one or more population haplotypes identified for each nucleotide read in the set of nucleotide reads. 6. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: determine haplotype group likelihood for the set of reference spans by comparing pairs of the set of nucleotide reads with the candidate population haplotypes within a corresponding reference span in the set of reference spans; and generate the haplotype group score for the set of reference spans based on the haplotype group likelihood. 7. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: identify at least one reference span in the set of reference spans corresponding to a haploid region of the reference genome; and, for the at least one reference span, generate haplotype subsets for each candidate population haplotype based on comparing the set of nucleotide reads and the candidate population haplotypes within the at least one reference span. 8. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system, based on the haplotype subsets, for each reference span in the set of reference spans, select a corresponding subgroup of the candidate population haplotype to be included in the population haplotype subgroup of the personalized haplotype database.9. The system of claim 8, further comprising instructions that, when executed by the at least one processor, cause the system to: generate haplotype group likelihoods for the set of reference spans by comparing the set of nucleotide reads with the candidate population haplotypes; generate a set of haplotype group posterior probabilities for each of the set of reference spans based on the corresponding haplotype group likelihoods of adjacent reference spans in the set of reference spans using a Hidden Markov Model (HMM) algorithm; and generate a corresponding group of haplotype group scores for the set of reference spans based on the set of haplotype group posterior probabilities for each of the set of reference spans. 10. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to: generate a set of forward probabilities for each of the set of reference spans using the haplotype group likelihood in the forward pass of the HMM algorithm across the set of reference spans; update the set of forward probabilities for each of the set of reference spans using the backward pass of the HMM algorithm, generating an updated set of probabilities for each of the set of reference spans; and generate the set of haplotype group posterior probabilities for each of the set of reference spans based on the updated set of probabilities for each of the set of reference spans. 11. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to reduce the variation of the haplotype group posterior probabilities across the set of reference spans using recombination parameters of the HMM algorithm. Claims 2 / 4, page 3, CN 121359209 A 12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: classify each nucleotide read in the set of nucleotide reads within each corresponding reference span of the set of reference spans as inherited from a first parent or a second parent, using a variational Bayesian model; generate individual haplotype scores for population haplotypes selected from the population haplotype database based on each classified nucleotide read in the set of nucleotide reads; and generate the haplotype group scores for pairs of population haplotypes selected from the population haplotype database by combining the individual haplotype scores from nucleotide reads inherited from the first parent with the individual haplotype scores from nucleotide reads inherited from the second parent.13. The system of claim 1, further comprising instructions, when executed by the at least one processor, to generate haplotype subsets for corresponding pairs of locally distinct population haplotypes within a given reference span of the set of reference spans, each locally distinct population haplotype comprising a unique set of one or more allele variant differences relative to other population haplotypes within the given reference span. 14. The system of claim 1, further comprising instructions, when executed by the at least one processor, to: identify allele variant differences between population haplotypes and primary continuous sequences in the population haplotype subset at the corresponding genomic region of the reference genome, as indicated in the personalized haplotype database; and determine one or more personalized alignments of the one or more personalized alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on allele variant differences. 15. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to: generate alignment scores for initial alignments of the set of nucleotide reads to corresponding genomic regions of the reference genome based on comparisons of the set of nucleotide reads to population haplotypes from the population haplotype database; generate an initial alignment data file including the initial alignments and corresponding alignment scores; generate personalized alignment scores for the one or more personalized alignments based on comparisons of the set of nucleotide reads to population haplotype subgroups within the personalized haplotype database; and generate a personalized alignment data file including the one or more personalized alignments and corresponding personalized alignment scores. 16. The system of claim 1, further comprising instructions, which, when executed by the at least one processor, cause the system to determine the genotype detection of the genome sample based on the one or more personalized alignments.17. A computer-implemented method comprising: identifying a set of nucleotide reads from a genome sample and candidate population haplotypes from a population haplotype database within a set of reference spans of a reference genome; generating haplotype group scores for groups of candidate population haplotypes based on comparing the set of nucleotide reads with the candidate population haplotypes for the set of reference spans of the reference genome; generating a personalized haplotype database for the genome sample and based on the haplotype group scores, the personalized haplotype database including population haplotype subgroups from the population haplotype database within the set of reference spans; and determining one or more personalized alignments of the set of nucleotide reads with corresponding genomic regions of the reference genome using the personalized haplotype database. 18. The computer-implemented method of claim 17, further comprising: identifying one or more distinct k-mers for each candidate population haplotype within a given reference span of the set of reference spans; and generating, for the given reference span, a corresponding haplotype group score in the haplotype group score for the corresponding candidate population haplotype within the given reference span, based on comparing the k-mers of the set of nucleotide reads with the one or more distinct k-mers of the corresponding candidate population haplotypes in the set of candidate population haplotypes. 19. The computer-implemented method of claim 17, further comprising: determining an initial alignment of the set of nucleotide reads with a corresponding genomic region of the reference genome; identifying subgroups of nucleotide reads in the set of nucleotide reads, the subgroups of nucleotide reads being aligned with the corresponding reference span of the set of reference spans according to the initial alignment; and generating the haplotype group score for the set of reference spans based on comparing the subgroups of nucleotide reads with the candidate population haplotypes within the corresponding reference span of the set of reference spans. 20. The computer-implemented method of claim 19, further comprising: identifying allelic variant differences between a population haplotype and a primary continuous sequence at a corresponding genomic region of the reference genome, as indicated in the population haplotype database; and determining one or more initial alignments of the initial alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on the allelic variant differences.Claims 4 / 4 Page 5 CN 121359209 A Personalized Haplotype Database for Improving Nucleotide Reads Mapping and Alignment and Improving Genotype Detection
[0001] Cross-Reference to Related Applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 558,754 (IP-2677-PRV), filed February 28, 2024, entitled “A PERSONALIZED HAPLOTYPE DATABASE FOR IMPROVED MAPPING AND ALIGNMENT OF NUCLEOTIDE READS AND IMPROVED GENOTYPE CALLING,” the entire contents of which are incorporated herein by reference. Background Art
[0003] In recent years, biotechnology companies and research institutions have improved the hardware and software used for sequencing nucleotides and identifying variants in genomic samples. For example, some existing nucleotide sequencing platforms determine individual nucleotides within sequences from cellular genomic sample samples using conventional Sanger sequencing or sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms monitor millions to billions of parallel synthesized nucleic acid polymers to predict nucleotide detections based on a larger dataset of nucleotide detections. For example, cameras in many SBS platforms capture images of irradiated fluorescent tags incorporated into oligonucleotides for nucleotide detection. After capturing such images, existing sequencing platforms send nucleotide detection data (or image-based data) to computing devices for application of sequencing data analysis software that determines the nucleotide sequence of a genomic sample or other nucleic acid polymer. For example, such software maps and aligns (i) nucleotide reads determined by the sequencing platform for the genomic sample with (ii) a reference genome comprising at least one primary continuous sequence. Based on the differences between the aligned nucleotide reads and the reference genome, existing data analysis software can further utilize variant detectors to identify genotypes and / or variants within the genomic sample, such as single nucleotide polymorphisms (SNPs), insertions or deletions (indels) or structural variants.
[0004] Despite these recent advances, existing nucleotide sequencing platforms and sequencing data analysis software (hereinafter referred to together as "existing sequencing systems") often utilize reference genomes that distort certain populations and cause inaccurate read mapping and alignment, as well as false variant detection. For example, some existing sequencing systems use a linear reference genome that is claimed to represent a common sequence or example of genes and other nucleotide sequences of an organism.However, for the most common linear human reference genome (GRCh38 from the Genome Reference Consortium), approximately 93% of primary assemblies are based on libraries from only 11 individuals, with 70% of linear human reference genomes originating from a single individual. Therefore, many existing systems use linear reference genomes that do not represent certain populations, common variants, or common population haplotypes.
[0005] To address this lack of genomic representation in linear reference genomes, some existing sequencing systems generate or use graph reference genomes. For example, some graph reference genomes comprise both a linear reference genome and a graph augmentation having multinucleobase codes representing SNPs and / or insertions / deletions, and alternating contiguous sequences representing various alternating population haplotypes at a given genomic region. In some cases, such graph reference genomes stack and index many alternating contiguous sequences, which can each extend relatively long nucleotide distances (e.g., hundreds to thousands of base pairs in length), and thus include redundant reference nucleotides overlapping the same regions. A one-size-fits-all approach to graph reference genomes can correspondingly consume excessive memory to encode alternating contiguous sequences.
[0006] While such graph reference genomes better explain the genetics of some populations, the extended representation of existing graph reference genome specifications (page 1 / 35, CN 121359209 A) typically includes a large number of alternating pathways with alleles similar to those in other genomic regions and pathways in the graph reference genome. Therefore, existing sequencing systems can significantly increase the difficulty of accurately predicting degradation from alternating pathways by compromising the uniqueness and usefulness of the genomic regions used for mapping and aligning nucleotide reads, and by increasing confusion between multiple seemingly similar genomic regions. In fact, a one-size-fits-all approach to graph reference genomes can reduce the quality score of reads used for mapping and undermine the intended utility of a more diverse set of alternating continuous sequences to which reads can be mapped.
[0007] In practice, these universal graph reference genomes have a large number of alternating pathways representing alternating continuous sequences, often leading to incorrect alignments, incorrect matches, or missed variant detections by existing sequencing systems on a large number of samples, and increasing the chance of mismatch alignments with reads from genomic samples. Because multiple seemingly similar population haplotypes exist in a given genomic region of a primary continuous sequence, and because the mapping quality (e.g., MAPQ 0) decreases as the number of population haplotypes in a given genomic region increases, existing sequencing systems are generally unable to scale up candidate population haplotypes in the graph reference genome without adversely affecting the accuracy of mapping and alignment, and thus reducing the accuracy of variant detection.
[0008] Although the human genome is predominantly diploid, many existing sequencing systems utilize human reference genome assemblies that include haploid representations of different reference genome samples.By mapping and aligning diploid genome samples using a haploid reference genome, existing sequencing systems often align nucleotide reads and identify variants negatively affected by reference bias. This reference bias favors alignments of such reads with alleles represented by the reference genome while impairing alignments with alternating alleles, despite allelic variant differences between the sample diploid and the haploid reference genome. This reference bias typically leads to false positives (FP) and false negatives (FN) in variant detection, thereby reducing the accuracy of existing sequencing systems in variant detection of generated diploid genome samples.
[0009] These problems and challenges, along with others, exist in existing sequencing systems. Summary of the Invention
[0010] This disclosure describes embodiments of (i) generating a personalized haploid database for genome samples and (ii) methods, non-transitory computer-readable media, and systems for determining personalized alignments of nucleotide reads from genome samples using a personalized haploid database. Specifically, the disclosed system can generate a haplotype database tailored or personalized for a specific genome sample based on comparisons of nucleotide reads from a genome sample with candidate population haplotypes from a population haplotype database. In some specific implementations, for example, the disclosed system can generate a personalized diploid reference database having a set of customized haplotype pairs of diploid genomic regions of a reference genome. The disclosed system can utilize the personalized haplotype database to determine personalized alignments of nucleotide reads from a genome sample with corresponding genomic regions of the reference genome.
[0011] For example, the disclosed system can identify a set of nucleotide reads from a genome sample and candidate population haplotypes from a population haplotype database within a set of reference spans of a reference genome. For each reference span in this set of reference spans, the disclosed system can generate haplotype group scores for a group of candidate population haplotypes based on comparisons of nucleotide reads and candidate population haplotypes. Based on the haplotype group scores, the disclosed system can generate a personalized haplotype database that includes subgroups of population haplotypes from a population haplotype database. By utilizing a personalized haplotype database generated according to the disclosed method, the disclosed system can determine a personalized alignment of a set of nucleotide reads from a genome sample with corresponding genomic regions of a reference genome.
[0012] Additional features and advantages of one or more embodiments of this disclosure will be set forth in the following description and will be apparent in part from the description, or may be learned by practice of such exemplary embodiments. Specification 2 / 35 pages 7 CN 121359209 A Brief Description of the Drawings
[0013] Detailed description refers to the accompanying drawings, which are briefly described below.
[0014] Figure 1 illustrates an environment in which a personalized sequencing system according to one or more embodiments of this disclosure may operate.
[0015] Figure 2 illustrates a personalized sequencing system for identifying nucleotide reads from a genomic sample within a reference span of a reference genome, according to one or more embodiments of the present disclosure.
[0016] Figure 3A illustrates a personalized sequencing system for determining an initial alignment of a set of nucleotide reads from a genomic sample and generating an alignment data file, according to one or more embodiments of the present disclosure.
[0017] Figure 3B illustrates a personalized sequencing system for generating a personalized haplotype database and personalized alignment data files for a set of nucleotide reads from a genomic sample, according to one or more embodiments of the present disclosure.
[0018] Figure 4 further illustrates a personalized sequencing system for generating a personalized haplotype database for a set of nucleotide reads from a genomic sample, according to one or more embodiments of the present disclosure.
[0019] Figure 5 illustrates a personalized sequencing system for selecting haplotype groups for a set of reference spans of a reference genome and generating a personalized haplotype database, according to one or more embodiments of the present disclosure.
[0020] Figure 6A illustrates a set of basic level bins for a population haplotype database, according to one or more embodiments of the present disclosure.
[0021] Figure 6B illustrates a set of basic level bins for a personalized haplotype database according to one or more embodiments of the present disclosure.
[0022] Figure 7 illustrates the basic level bins and successive higher level bins for a personalized haplotype database according to one or more embodiments of the present disclosure.
[0023] Figure 8 illustrates comparative experimental results of variant detection from nucleotide reads (i) mapped and aligned to a reference genome using an existing sequencing system, and (ii) mapped and aligned to a reference genome using a personalized haplotype database generated by a personalized sequencing system.
[0024] Figure 9 illustrates comparative experimental results of mapping and aligning nucleotide reads from a genomic sample using (i) an existing sequencing system and (ii) a personalized sequencing system according to one or more embodiments of the present disclosure.
[0025] Figure 10 illustrates a flowchart of a series of actions for generating a personalized haplotype database and determining personalized alignments of a set of nucleotide reads from a genomic sample according to one or more embodiments of the present disclosure.
[0026] FIG11 illustrates a block diagram of an exemplary computing device for implementing one or more embodiments of the present disclosure. Detailed Description
[0027] The present disclosure describes embodiments of a personalized sequencing system that generates a personalized haplotype database and utilizes it to determine personalized alignments of nucleotide reads from a genome sample. Specifically, the personalized sequencing system may evaluate and select a set of customized population haplotypes based on comparisons of nucleotide reads from a genome sample with candidate population haplotypes to generate a personalized haplotype database of a genome sample.In some implementations, for example, the personalized sequencing system uses a population haplotype database (e.g., a group of 256 haplotypes) to initially map and align a set of nucleotide reads from a genomic sample to generate an initial alignment data file (e.g., a re-scored binary alignment map (BAM) file). After determining the initial alignment of the set of nucleotide reads, in some implementations, the personalized sequencing system compares the set of nucleotide reads with haplotypes from the population haplotype database to select haplotype subgroups to be included in the personalized haplotype database of the genomic sample (e.g., Personalized Haplotype Group Tag Separation Value (TSV) specification 3 / 35 page 8 CN 121359209 A file).
[0028] Furthermore, in some implementations, the personalized sequencing system utilizes an interpolation model to evaluate candidate population haplotype groups across multiple genomic regions of a reference genome to generate haplotype group scores, and selects a customized population haplotype group based on the generated haplotype group scores. In some specific implementations, the personalized sequencing system generates a personalized haplotype database, which includes customized population haplotype pairs representing predicted diploids of corresponding genomic samples within one or more diploid genomic regions of a corresponding reference genome. After generating the personalized haplotype database, the personalized sequencing system can use the personalized haplotype database of genomic samples to determine personalized alignments of nucleotide reads from the corresponding genomic samples with corresponding genomic regions of the reference genome.
[0029] Prior to determining such personalized alignments, in one or more implementations, the personalized sequencing system identifies a set of nucleotide reads from genomic samples and candidate population haplotypes from a population database within a set of reference spans of the reference genome. Based on comparing this set of nucleotide reads and candidate population haplotypes within each reference span, the personalized sequencing system can generate haplotype set scores for the candidate population haplotype set. Based on the haplotype set scores, the personalized sequencing system can generate a personalized haplotype database, which includes population haplotype subgroups selected from the population haplotype database within the set of reference spans. As mentioned, personalized sequencing systems can utilize personalized haplotype databases to determine one or more personalized alignments of the set of nucleotide reads with the corresponding genomic regions of the reference genome.
[0030] Additionally, as mentioned, personalized sequencing systems can utilize interpolation models to evaluate candidate population haplotypes across multiple reference spans of the reference genome to generate haplotype scores, and select customized population haplotypes based on the generated haplotype scores.In one or more embodiments, for example, the personalized sequencing system generates haplotype set scores, which include (i) haplotype pair scores of population haplotype pairs in a reference span covering diploid regions of a reference genome (e.g., regions corresponding to somatic chromosomes), or (ii) individual haplotype scores of population haplotypes in a reference span covering haploid regions of a reference genome (e.g., regions corresponding to sex chromosomes). Alternatively or additionally, the personalized sequencing system may generate haplotype set scores for groups with more than two population haplotypes (e.g., to customize polyploid genome samples with more than two chromosomes per group, or to account for the uncertainty in selecting customized haplotype sets representing genome samples).
[0031] When using an imputation model to evaluate candidate population haplotype sets across a reference span, the personalized sequencing system may utilize a Hidden Markov Model (HMM) algorithm as the imputation model. For example, in some embodiments, the personalized sequencing system utilizes an HMM algorithm to imput haplotype set posterior probabilities for candidate population haplotype sets across adjacent reference spans in a set of reference spans of a reference genome associated with a genome sample. Based on the posterior probability of haplotype groups, the personalized sequencing system can generate haplotype scores for the corresponding candidate population haplotypes, and based on the haplotype scores, generate a personalized haplotype database of the genome sample.
[0032] Furthermore, in various embodiments, the personalized sequencing system can utilize a variety of methods for scoring candidate population haplotype groups. In some embodiments, for example, the personalized sequencing system uses a variational Bayesian model that implements an iterative algorithm to generate haplotype scores for the corresponding candidate population haplotypes. Additionally, in one or more embodiments, the personalized sequencing system classifies nucleotide reads inherited from the corresponding first and second parents, generates individual haplotype scores for the classified reads, and generates haplotype scores for candidate population haplotype pairs by combining the individual haplotype scores from the nucleotide reads inherited from the first parent with the individual haplotype scores from the nucleotide reads inherited from the second parent.
[0033] To facilitate the determination of haplotype scores for candidate population haplotype groups, in some embodiments, the personalized sequencing system determines an initial alignment of the group of nucleotide reads with the corresponding genomic regions of the reference genome. By identifying nucleotide reads that have been aligned to the corresponding reference span within the set of reference spans based on the initial alignment, the personalized sequencing system can generate haplotype fractions for each corresponding reference span within the set of reference spans (page 4 / 35, CN 121359209 A of the manual).Furthermore, in one or more embodiments, the personalized sequencing system determines the initial alignment by: (1) identifying allelic variant differences between the population haplotype and the primary continuous sequence at the corresponding genomic region of the reference genome, as indicated in the population haplotype database, and (2) re-scoring the candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on the identified allelic variant differences.
[0034] In one or more embodiments, after determining the initial alignment of a set of nucleotide reads via the aforementioned alignment method or by another method, the personalized sequencing system generates an alignment data file that includes information for generating the personalized haplotype database. For example, such an alignment data file (e.g., a personalized binary alignment map (BAM) file) may include the initial alignment of the set of nucleotide reads, the alignment score corresponding to the initial alignment, and one or more population haplotypes identified for each nucleotide read in the set of nucleotide reads. Subsequently, in some embodiments, the personalized sequencing system uses the alignment data file to evaluate and select population haplotypes for the personalized haplotype database. By utilizing a personalized haplotype database, a personalized sequencing system can determine one or more personalized alignments of nucleotide reads from a set of nucleotide reads and output a personalized alignment data file with personalized alignments and corresponding alignment scores.
[0035] As described above, personalized sequencing systems offer several technical advantages, beneficial effects, and / or improvements over existing sequencing systems and methods. For example, personalized sequencing systems improve the accuracy of read alignment and subsequent genomic analysis by utilizing a personalized haplotype database of genomic samples. More specifically, in some embodiments, the personalized sequencing system generates a personalized haplotype database comprising a subset of population haplotypes selected for the genomic sample from a population haplotype database. By utilizing a personalized haplotype database in the personalized alignment of nucleotide reads, personalized sequencing systems can align nucleotide reads to the corresponding reference genome more accurately than existing sequencing systems that utilize a reference genome enhanced by an unfiltered or unnecessarily large population haplotype set (e.g., 15-20 haplotypes per region), particularly in more complex or “difficult” genomic regions (e.g., regions that typically include bases with low confidence detection). Due to the improved alignment with the reference genome, personalized sequencing systems can also determine more accurate genotype and / or variant detection with higher confidence than existing sequencing systems, where such detection matches or differs from reference bases in the reference genome. Examples of such improved genotype and / or variant detection are described and depicted below with reference to Figures 8 through 9.
[0036] Furthermore, in some embodiments, personalized sequencing systems also improve the accuracy of mapping and alignment, as well as subsequent variant detection, in part by avoiding reference bias, which is a bias that facilitates the alignment of nucleotide reads with the alleles of the corresponding reference genome. As mentioned above, for example, existing sequencing systems typically encounter reference bias when mapping nucleotide reads from diploid samples to a haploid reference genome. In contrast, in at least some embodiments, personalized sequencing systems provide a personalized haplotype database that avoids reference bias by including a customized diploid reference genome that includes nucleotide reads in diploid regions for mapping the corresponding reference genome. In fact, by generating a personalized haplotype database for personalized mapping and alignment of nucleotide reads from a variety of polyploid genome samples, personalized sequencing systems can further improve variant detection accuracy relative to existing sequencing systems.
[0037] In addition to improving the accuracy of alignment and related sequencing analyses, personalized sequencing systems also improve computational efficiency compared to existing sequencing systems. For example, by utilizing a personalized haplotype database for a specific genome sample, a personalized sequencing system can accurately determine personalized read alignments of nucleotide reads relative to existing sequencing systems with increased computational speed and less memory. Specifically, as mentioned above, existing sequencing systems typically determine read alignments by attempting to align and score nucleotide reads against a robust graphical genome enhanced with many alternating consecutive sequences representing an unfiltered population haplotype set. In contrast, a personalized sequencing system utilizes a personalized haplotype database that includes discrete subgroups of population haplotypes specifically selected for a given genome sample, thereby improving alignment accuracy, reducing memory consumption, and increasing processing speed.
[0038] Furthermore, in at least some specific embodiments, according to the disclosed method, a personalized sequencing system can accurately predict initial read alignments and / or personalized read alignments while improving computational speed and memory usage relative to existing sequencing systems. As mentioned above, existing sequencing systems use a graph reference genome with universal graph enhancement, which includes many redundant alternating continuous sequences that consume memory from repetitive sequences from overlapping portions of the alternating continuous sequences and slow down computer processing by scoring the alignment between reads and such overlapping portions of the alternating continuous sequences.Compared to existing systems of this kind, personalized sequencing systems can accelerate the mapping and alignment of nucleotide reads at least by: (i) identifying candidate reference alignments of nucleotide reads to primary continuous sequences at corresponding genomic regions of a reference genome, and (ii) determining initial or personalized alignments of nucleotide reads by re-scoring the candidate reference alignments based on allelic variant differences between primary continuous sequences at corresponding genomic regions of a reference genome and population haplotypes.
[0039] In addition to improving computational efficiency, or as an alternative, personalized sequencing systems can generate personalized haplotype databases for determining personalized alignments with greater flexibility compared to existing sequencing systems. For example, in some embodiments, personalized sequencing systems generate personalized haplotype databases of genomic samples based on comparing nucleotide reads from genomic samples to candidate population haplotypes. Thus, personalized sequencing systems can generate nucleotide read alignments with improved accuracy and efficiency without requiring additional information about the genomic sample, such as parental genomic data or demographic data associated with the sample. In some specific implementations, for example, personalized sequencing systems further enhance flexibility by implementing a personalized haploid database that facilitates more accurate mapping and alignment and / or variant detection compared to existing sequencing systems that utilize a haploid reference genome for mapping and alignment of diploid genome samples.
[0040] As proposed in the foregoing discussion, this disclosure uses a variety of terms to describe the features and advantages of personalized sequencing systems. Additional details regarding the meaning of these terms as used in this disclosure are provided below. As used in this disclosure, for example, the term "genome sample" refers to a target genome or a portion of a genome that has undergone determination or sequencing. For example, a genome sample comprises one or more nucleotide sequences (or copies of such isolated or extracted sequences) isolated or extracted from a sample organism. Specifically, a genome sample comprises a whole genome isolated or extracted (in whole or in part) from a sample organism and consisting of nitrogen-containing heterocyclic bases. A genome sample may include fragments or molecules of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other aggregated forms of nucleic acids or chimeric or hybrid forms of nucleic acids as described below. In some cases, the genomic sample is present in a sample prepared or isolated by a kit and received by a sequencing device.
[0041] Additionally, as used herein, the term "nucleotide read" (or simply "read") refers to a sequence of one or more nucleotide bases (or nucleobase pairs) inferred or predicted from all or part of a sample genomic sequence (e.g., a sample genomic sequence, complementary DNA). Specifically, a nucleotide read includes a sequence of nucleobases detected from a nucleotide fragment (or a group of monoclonal nucleotide fragments) determined or predicted based on a sequencing library corresponding to the genomic sample.For example, in some implementations, personalized sequencing systems identify nucleotide reads by generating nucleotide bases through nanopores in a nucleotide sample slide, via fluorescent tagging, or based on the pores in a flow cell. In some cases, a nucleotide read may refer to a specific type of read, such as a nucleotide read synthesized from a sample library fragment shorter than a threshold number of nucleotides (e.g., an SBS read). In these or other cases, another type of nucleotide read may refer to (i) an assembled nucleotide read that has been assembled from shorter nucleotide reads to form a continuous sequence that meets the threshold number of nucleotides (e.g., an assembled nucleotide read), (ii) a circularly shared sequencing (CCS) read that meets the threshold number of nucleotides, or (iii) a nanopore long read that meets the threshold number of nucleotides.
[0042] As further used herein, the term “genomic coordinates” (or sometimes simply “coordinates”) refers to a specific location or orientation of a nucleotide within a genome (e.g., the genome of an organism or a reference genome). In some cases, genomic coordinates include an identifier for a specific chromosome in the genome and an identifier for the location of a specific nucleus base on that chromosome. For example, one or more genomic coordinates may include the number, name, or other identifier of a somatic cell or sex chromosome (e.g., chrl or chrX) and one or more specific locations, such as a numbered location following the chromosome identifier (e.g., chrl: 1234570 or chrl: 1234570–1234870). In some cases, genomic coordinates refer to genomic coordinates on sex chromosomes (e.g., chrX or chrY). Therefore, personalized sequencing systems can determine the genotype probability of genotype detection (e.g., variant detection) for genomic coordinates on sex chromosomes. Additionally, in some specific implementations, genomic coordinates refer to the source of a reference genome (e.g., mt of a mitochondrial DNA reference genome or SARS-CoV-2 of a SARS-CoV-2 virus reference genome) and the location of the source nucleus base of the reference genome (e.g., mt: 16568 or SARS-CoV-2:29001). In contrast, in some cases, genomic coordinates refer to the location of nuclei in a reference genome without relating to chromosomes or origin (e.g., 29727).
[0043] As used herein, a “genomic region” refers to a range of genomic coordinates. Similar to genomic coordinates, in some embodiments, genomic regions can be identified by a chromosome identifier and one or more specific locations (such as numbered locations following the chromosome identifier (e.g., chrl: 1234570–1234870)). In various embodiments, genomic coordinates include locations within a reference genome.In some cases, genomic coordinates are specific to a particular reference genome. Accordingly, as used herein, the term "reference span" refers to the span of nucleotide localization within a linear reference genome. In other words, the reference span includes the nucleotide span between two corresponding genomic coordinates of a linear reference genome.
[0044] As described above, genomic coordinates include localization within a reference genome. This localization may be within a specific reference genome. As used herein, the term "reference genome" refers to a digital nucleic acid sequence assembled as a representative example (or multiple representative examples) of genes and other genetic sequences of an organism. Regardless of sequence length, in some cases, a reference genome represents a set of examples of genes or a set of nucleic acid sequences identified by scientists as representing an organism of a particular species. For example, a linear human reference genome may be GRCh38 or another version of a reference genome from the Genome Reference Consortium. As described above, in some cases, a reference genome includes polybase codes. Alternatively, a reference genome may include a graphical reference genome containing both a linear reference genome and pathways representing nucleic acid sequences from ancestral haplotypes, such as the Illumina DRAGEN graphical reference genome hgl9.
[0045] As used herein, the term “primary contiguous sequence” (or simply “primary contig”) refers to a contiguous sequence representing a reference haplotype of a reference genome. In some embodiments, the primary contiguous sequence digitally represents the reference haplotype of the reference genome, but may include additional information from the primary assembly of the linear reference genome, such as indications of population variants in certain genomic regions, to aid in identifying candidate alignments of nucleotide reads.
[0046] In contrast, the term “alternating contiguous sequence” (or simply “alternating contig”) refers to a contiguous sequence representing a population haplotype at a specific genomic coordinate of a reference genome. For example, in some sequencing systems, a graph reference genome includes alternating contiguous sequences of genomic coordinates mapped to the primary contiguous sequences of a linear reference genome. In some cases, the hash table for a graph reference genome includes identifiers that associate alternating contiguous sequences representing population haplotypes at genomic coordinates relative to the linear reference genome. Crucially, as explained and depicted in this disclosure, a personalized haplotype database includes data corresponding to a limited sampling of population haplotypes selected for a particular genomic sample.
[0047] Accordingly, as used herein, the term “allelic variant difference” refers to the difference between corresponding nucleobases in two or more given nucleotide sequences, as described on page 7 / 35 of the specification, in column 12 CN 121359209 A. In some cases, for example, allelic variant difference is the difference between a primary continuous sequence and at least one population haplotype (e.g., represented by alternating continuous sequences).In some embodiments, for example, allelic variant differences within a given genomic region may include single nucleotide variants, multibase differences, and / or insertions and deletions (insertions and deletions) relative to population haplotypes of a primary continuous sequence. Additionally, allelic variant differences may refer to differences between a first population haplotype and a second population haplotype.
[0048] As used herein, the terms “locally distinct population haplotype” or “locally distinct haplotype” refer to a haplotype comprising a set of at least one allelic variant differences, wherein this set is unique relative to other haplotypes within the corresponding genomic region of a reference genome. For example, each genomic region or reference span of a reference genome may include one or more locally distinct population haplotypes having a unique set of one or more allelic variant differences relative to other population haplotypes within the corresponding genomic region or reference span. Additionally, in some specific embodiments, a given set of one or more allelic variant differences within a genomic region corresponding to a candidate read alignment may represent multiple haplotypes resulting from complete overlap of variants within the genomic region. Therefore, in some cases, multiple haplotypes comprising the same nucleotide bases or consisting of the same nucleotide bases within a given genomic region may be represented by a single locally distinct population of haplotypes.
[0049] Furthermore, as used herein, the term “alignment score” refers to a numerical score, measure, or other quantitative measure that assesses the accuracy of an alignment between one or more nucleotide reads or fragments of nucleotide reads and another nucleotide sequence from a reference genome. Specifically, an alignment score includes a measure indicating the degree to which the nucleotide bases of one or more nucleotide reads (or fragments thereof) match or are similar to a reference sequence or alternating continuous sequence from a reference genome. In some specific implementations, the alignment score takes the form of a Smith-Waterman score or a variant or version of a Smith-Waterman score for local alignment, such as the various settings or configurations used by DRAGEN of Illumina, Inc. for the Smith-Waterman score.
[0050] Relatedly, as used herein, the term “mapped quality score” refers to a measure or other measure that quantifies the quality or certainty of an alignment of a nucleotide read (or other nucleotide sequence or subsequence) to a reference genome. In some implementations, for example, the mapping quality score includes the mapping quality (MAPQ) score of nucleotides detected at genomic coordinates, where the MAPQ score represents –10 log10 Pr{mapping location error}, rounded to the nearest integer. In alternative implementations of mean or median mapping quality, the mapping quality score includes the full distribution of mapping quality across all nucleotide reads aligned to a reference genome at genomic coordinates.In some implementations, MAPQ scores are divided into sequence bins (e.g., Q-score bins denoted by "MAPQ0", "MAPQ10", "MAPQ20", etc.), which represent different MAPQ scores for each bin. Additionally, in some implementations, any MAPQ score above a predetermined threshold score is associated with a maximum value bin (e.g., up to "MAPQ40"). Accordingly, as used herein, the terms "extended mapping quality score" or "extended MAPQ score" refer to MAPQ scores that are divided to include bins with higher maximum values than the conventional implementation (e.g., up to "MAPQ60"). In some implementations, for example, a personalized sequencing system implements an extended MAPQ score when performing initial mapping and alignment of a given set of nucleotide reads, and utilizes the extended MAPQ score to compare nucleotide reads with candidate population haplotypes to generate haplotype set scores (e.g., as described below with respect to Figures 3A-3B and Figure 4).
[0051] As further used herein, the term "genotype detection" refers to identifying or predicting a specific genotype of a genomic sample or sample nucleotide sequence at a genomic locus. Specifically, genotype detection may include predicting a specific genotype of a genome sample relative to a reference genome or reference sequence at genomic coordinates or a genomic region. For example, in some cases, genotype detection includes identifying or predicting that a genome sample includes both nucleobases and complementary nucleobases at genomic coordinates that are homozygous or heterozygous with respect to a reference base or variant (e.g., a homozygous reference base is represented as 0|0, or a heterozygous variant on a particular strand is represented as 0|1). Thus, genotype detection may include predicting variants or reference bases of one or more alleles in a genome sample and indicating conjugation with respect to the variant or reference base. Genotype detection is typically determined against genomic coordinates or a genomic region where SNPs, insertions, deletions, or other variants have been identified for an organism population.
[0052] As further used herein, the term “nucleobase detection” (or simply “base detection”) refers to the determination or prediction of the genomic coordinates of an oligonucleotide (e.g., a nucleotide read) or sample genome during a sequencing cycle. Specifically, nucleobase detection can indicate the determination or prediction of the type of nucleobase within an oligonucleotide incorporated into a nucleotide sample slide (e.g., read-based nucleobase detection). In some cases, for nucleotide reads, nucleobase detection includes determining or predicting nucleobases based on the intensity values produced by fluorescently tagged nucleotides of oligonucleotides added to a nucleotide sample slide (e.g., in a cluster of flow cells).As mentioned above, single nucleobase detection can be adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U).
[0053] As used herein, the term "variant" refers to one or more nucleobases that do not align to, differ from, or are altered from a corresponding nucleobase (or nucleobases) in a reference sequence or reference genome. For example, variants include SNPs, indels, or structural variants that indicate nucleobases in the sample nucleotide sequence that differ from nucleobases at corresponding genomic coordinates in the reference sequence or reference genome.
[0054] Following these lines of thought, "variant detection" (or "variant nucleobase detection") refers to the detection of nucleobases that include mutations or variants at specific genomic coordinates or regions relative to a reference. Specifically, variant detection includes identifying or predicting specific nucleobases (or nucleobase sequences) in a genomic sample that differ from reference nucleobases (or reference nucleobase sequences) at the same genomic coordinates or regions within the reference genome. Conversely, “reference detection” (or “non-variant nucleobase detection” or “non-variant detection”) refers to a nucleobase detection that includes a non-variant or reference nucleobase at genomic coordinates or genomic regions relative to a reference. Specifically, non-variant or reference detection includes identifying or predicting a specific nucleobase (or nucleobase sequence) in a genomic sample that matches a reference nucleobase (or reference nucleobase sequence) at the same genomic coordinates or regions within a reference genome.
[0055] As used herein, the term “alignment data file” refers to a digital file indicating the mapping and alignment information of nucleotide reads of a sample nucleotide sequence. For example, an alignment data file may include a binary alignment map (BAM) file, a reference-oriented compressed alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence. Additionally, as discussed below (e.g., with respect to Figures 3A-3B), an alignment data file may include further information about nucleotide reads, mapping and alignment results, population haplotype data, etc.
[0056] As further used herein, the term "population haplotype" refers to a nucleotide sequence present in an organism (or present in organisms from a population) and inherited from one or more ancestors. Specifically, a population haplotype may include alleles or other nucleotide sequences present in organisms of a population and inherited together by those organisms from a single parent. In one or more embodiments, a population haplotype includes a set of SNPs or other variants that tend to be inherited together on the same chromosome. In some cases, data representing a population haplotype or a set of different population haplotypes are stored on or otherwise accessible on a population haplotype database.As mentioned, in some embodiments, the personalized sequencing system also generates a personalized haplotype database, which includes a customized selection of population haplotypes imported from a specific population haplotype database.
[0057] Accordingly, as used herein, the term "population haplotype database" refers to a database encoding variant data of population haplotypes of a sample organism. Specifically, a population haplotype database refers to an unfiltered compilation of population haplotypes, or in other words, a complete compilation of population haplotypes prior to personalization according to the methods disclosed herein. In one or more embodiments, the population haplotype database includes complete or partially complete nucleotide sequences (e.g., alternating continuous sequences) of population haplotypes of a sample organism. Alternatively, in some embodiments, the population haplotype database encodes variant data of population haplotypes that have allelic variant differences from population haplotypes that are locally different from those in corresponding genomic regions of a corresponding reference genome. For example, in some embodiments, the population haplotype database includes a haplotype data structure that includes bins (e.g., represented by primary continuous sequences) that hierarchically divide different genomic regions of a reference genome into corresponding spans covering the linear reference genome.
[0058] Relatedly, as used herein, the term “personalized haplotype database” is interchangeable with “customized haplotype database” and refers to a haplotype database that includes a personalized (or customized) population haplotype subgroup for a given genome sample. For example, a personalized (or customized) haplotype database may include one or more population haplotypes selected for the sample genome from the population haplotype database as described above. In some embodiments, for example, a personalized (or customized) haplotype database may include two population haplotypes identified as diploid for each genomic region corresponding to one or more genomic regions of a reference genome. Similarly, in some embodiments, for a genomic region corresponding to a sex chromosome of the sample genome, the personalized (or customized) haplotype database may include a single population haplotype. In fact, any number of population haplotypes may be included in the personalized (or customized) haplotype database according to the methods disclosed herein.
[0059] Additionally, as used herein, the term "haplotype fraction" refers to a measure or other measurement that quantifies the likelihood or probability of a nucleotide read from a given genome sample corresponding to a corresponding group of one or more population haplotypes within a specific genomic region. In some embodiments, for example, haplotype fraction quantifies the similarity between haplotype groups that include paired population haplotypes and nucleotide reads within diploid regions corresponding to a reference genome.Similarly, in some embodiments, haplotype components quantify the similarity between nucleotide reads within haploid genomic regions (e.g., sex chromosomes) of a reference genome and individual haplotypes. Furthermore, in some embodiments, haplotype components quantify the similarity between nucleotide reads within polyploid genomic regions of a reference genome and groups of multiple haplotypes (e.g., such that each group of multiple haplotypes includes the same number of haplotypes as the number of chromosome groups in the polyploid genomic regions of the genome sample). Additionally, in some embodiments, haplotype component scores may include haplotype group likelihood or haplotype group posterior probability, as described with respect to Figure 5.
[0060] Additionally, as used herein, the term “personalized alignment” is interchangeable with “customized alignment” and refers to read alignments generated using a personalized (or customized) haplotype database as described herein. For example, a personalized (or customized) haplotype database generated for a specific genome sample according to the methods disclosed herein can be used instead of an unfiltered population haplotype database to generate one or more personalized (or customized) alignments of nucleotide reads from a specific genome sample.
[0061] The following paragraphs describe a personalized sequencing system with reference to exemplary and specific embodiments illustrated in the accompanying drawings. For example, FIG1 illustrates a schematic diagram of a computing system 100 in which a personalized sequencing system 106 according to one or more embodiments operates. As shown, the computing system 100 includes a sequencing device 102 connected to a local device 108 (e.g., a local server device), one or more server devices 110, and a client device 114. As shown in FIG1, the sequencing device 102, the local device 108, the server device 110, and the client device 114 can communicate with each other via a network 118. The network 118 includes any suitable network through which the computing devices can communicate. An example network is discussed in more detail below with respect to FIG11. While FIG1 illustrates an embodiment of the calibration sequencing system 106, this disclosure describes the following alternative embodiments and configurations.
[0062] As indicated in FIG1, the sequencing device 102 includes a computing device or sequencing device system 104 for sequencing genomic samples or other nucleic acid polymers. In some embodiments, by executing a processor-based sequencing device system 104, sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from a genomic sample to generate nucleotide reads or other data directly or indirectly on sequencing device 102 using a computer-implemented method and system described in specification 10 / 35 (page 15, CN 121359209 A). More specifically, sequencing device 102 receives a nucleotide sample slide (e.g., a flow cell) comprising nucleotide fragments extracted from a sample, then copies and determines the nucleobase sequence of such extracted nucleotide fragments.
[0063] In one or more embodiments, sequencing device 102 utilizes sequencing-by-synthesis (SBS) technology to sequence nucleotide fragments into nucleotide reads and determine the nucleobase detections of these nucleotide reads. As a supplement to or alternative to communication across network 118, in some embodiments, sequencing device 102 bypasses network 118 and communicates directly with local device 108 or client device 114. By executing sequencing device system 104, sequencing device 102 may also store nucleobase detections as part of nucleobase detection data formatted as a binary base detection (BCL) file and send the BCL file to local device 108 and / or server device 110.
[0064] As further indicated in FIG1, local device 108 is located at or near the same physical location as sequencing device 102. In fact, in some embodiments, local device 108 and sequencing device 102 are integrated into the same computing device. Local device 108 can run personalized sequencing system 106 to generate, receive, analyze, store, and transmit digital data, such as determining variant detection by receiving base detection data or by analyzing such base detection data. As shown in FIG1, sequencing device 102 can send (and local device 108 can receive) base detection data generated during sequencing runs of sequencing device 102. By executing software in the form of personalized sequencing system 106, local device 108 can use personalized haplotype database 112 to align nucleotide reads with a reference genome and determine genetic variants based on the aligned nucleotide reads. Local device 108 can also communicate with client device 114. Specifically, local device 108 can send data to client device 114, including binary map (BAM) files, variant detection format (VCF) files, or other information indicating nucleobase detection, sequencing metrics, error data, or other metrics.
[0065] As further indicated in FIG1, server device 110 is located remotely from local device 108 and sequencing device 102. Similar to local device 108, in some embodiments, server device 110 includes a version of personalized sequencing system 106 (or is otherwise accessible to or implements that personalized sequencing system). Therefore, server device 110 can generate, receive, analyze, store, and transmit digital data, such as determining variant detections by receiving base detection data or by analyzing such base detection data. As described above, sequencing device 102 can send (and server device 110 can receive) base detection data from sequencing device 102. Server device 110 can also communicate with client device 114. Specifically, server device 110 can send data to client device 114, including BAM files, VCF files, or other sequencing-related information.
[0066] In some embodiments, server device 110 includes a distributed set of servers, wherein server device 110 includes a number of server devices distributed across network 118 and located in the same or different physical locations. Furthermore, server device 110 may include a content server, application server, communication server, web hosting server, or another type of server. Additionally, as shown in FIG1, server device 110 communicates directly or via network 118 with population haplotype database 120, which stores population haplotypes to be evaluated by personalized sequencing system 106 when generating personalized haplotype database 112 for genome samples.
[0067] As described above, as part of server device 110 or local device 108, personalized sequencing system 106 may generate, encode, and / or implement personalized haplotype database 112 to determine personalized alignments of nucleotide reads from genome samples with reference genomes. In some implementations, for example, the personalized sequencing system 106 may generate a personalized haplotype database 112 of a genome sample (e.g., based on a set of nucleotide reads from the genome sample) and utilize the generated personalized haplotype database 112 to determine one or more personalized alignments of nucleotide reads from the genome sample corresponding to the genome sample, as described in more detail below with reference to the following figures.
[0068] As further shown and indicated in FIG1, by executing the sequencing application 116, the client device 114 may generate, store, receive, and transmit digital data. Specifically, the client device 114 may receive sequencing data from the local device 108 or receive detection files (e.g., BCLs) and sequencing metrics from the sequencing device 102. In addition, the client device 114 may communicate with the local device 108 or the server device 110 to receive VCFs, which include genotype or variant detection and / or other metrics, such as base detection quality metrics or passband filter metrics. Client device 114 may accordingly present or display information related to variant detection or other genotype detection to the user associated with client device 114 within the graphical user interface of sequencing application 116. For example, client device 114 may present genotype detection, variant detection, and / or sequencing metrics for a sequenced genome sample within the graphical user interface of sequencing application 116.
[0069] Although Figure 1 depicts client device 114 as a desktop or laptop computer, client device 114 may include various types of client devices. For example, in some embodiments, client device 114 may include non-mobile devices such as desktop computers or servers or other types of client devices.In other embodiments, client device 114 includes mobile devices such as laptops, tablets, mobile phones, or smartphones. Additional details regarding client device 114 are discussed below with respect to FIG11.
[0070] As further illustrated in FIG1, client device 114 includes sequencing application 116. Sequencing application 116 may be a web application or native application (e.g., a mobile application, desktop application) stored and executed on client device 114. Sequencing application 116 may include instructions (when executed) that cause client device 114 to receive data from personalized sequencing system 106 and present base detection data or data from alignment data files or VCFs for display at client device 114. Furthermore, sequencing application 116 may instruct client device 114 to display a summary for multiple sequencing runs.
[0071] As further illustrated in FIG1, a version of personalized sequencing system 106 may be located and / or implemented (e.g., wholly or partially) on client device 114 or sequencing device 102. In yet another embodiment, the personalized sequencing system 106 is implemented by one or more other components of the computing system 100, such as a local device 108. Specifically, the personalized sequencing system 106 may be implemented in a variety of different ways across sequencing device 102, local device 108, server device 110, and client device 114. For example, the personalized sequencing system 106 may be downloaded from server device 110 to the personalized sequencing system 106 and / or local device 108, wherein all or part of the functionality of the personalized sequencing system 106 is performed at each of the respective devices within the computing system 100.
[0072] As previously mentioned, in some embodiments, the personalized sequencing system 106 evaluates a set of reference spans against a reference genome of a genome sample for inclusion in a personalized haplotype database of the genome sample as a candidate population haplotype set. For illustration, Figure 2 depicts an example of a set of nucleotide reads 202 and a set of reference spans 204a, 204b to 204n of a reference genome. Specifically, Figure 2 illustrates a personalized sequencing system 106 that identifies nucleotide reads in the set of nucleotide reads 202 within each reference span in the set of reference spans 204a-204n.
[0073] In one or more embodiments, for example, the personalized sequencing system 106 compares nucleotide reads (such as the set of nucleotide reads 202) with candidate population haplotype sets within one or more genomic regions of a reference genome (such as the set of reference spans 204a-204n) to determine the likelihood that each set of candidate haplotypes represents a genomic sample within each corresponding genomic region.In various implementations, the personalized sequencing system 106 can divide a reference genome into multiple genomic regions, such as, but not limited to, reference spans corresponding to chromosomes of the genome sample, reference spans covering a predetermined number of nucleotides (e.g., 1,000 nucleotides; 10,000 nucleotides; 1 million nucleotides) of primary continuous sequences, or a single reference span covering some or all of the genome sample. Thus, the personalized sequencing system 106 can divide a reference genome (or a portion thereof) into a set of reference spans, such as reference spans 204a–204n, compare population haplotype sets with nucleotide reads aligned to each corresponding reference span, and for each corresponding reference span, select a set of population haplotypes, (see specification page 12 / 35, 17 CN 121359209 A) to be included in a personalized haplotype database (e.g., as described below with respect to Figures 3A–3B and Figures 4–5).
[0074] As shown in FIG2, the personalized sequencing system 106 identifies nucleotide reads within each reference span in the set of reference spans 204a-204n. In some embodiments, for example, the personalized sequencing system 106 determines an initial alignment of the set of nucleotide reads 202 with corresponding genomic regions of a reference genome (e.g., as further described below with respect to FIG3A). Subsequently, the personalized sequencing system 106 may identify one or more nucleotide reads in the initially aligned set of nucleotide reads 202 that are at least partially aligned with corresponding reference spans in the set of reference spans 204a-204n, and compare the identified nucleotide reads (or portions thereof) with candidate population haplotype sets within each corresponding reference span in the set of reference spans 204a-204n.
[0075] Alternatively, in some embodiments, the personalized sequencing system 106 identifies one or more distinct k-mers within the set of nucleotide reads 202 for each candidate population haplotype within a given reference span in the set of reference spans 204a-204n. For example, the personalized sequencing system 106 can identify different nucleotide k-mers (or partially different k-mers) of length k within one or more candidate population haplotypes, and compare the identified different k-mers of the candidate population haplotypes with the set of nucleotide reads 202 within a given reference span in the set of reference spans 204a-204n. In practice, the personalized sequencing system 106 can utilize various methods for identifying nucleotide reads within the corresponding reference spans of the reference genome, and is not limited to the methods described herein.
[0076] As previously mentioned, in some embodiments, the personalized sequencing system 106 generates a personalized haplotype database of the genome sample based on comparing nucleotide reads with population haplotypes from a population haplotype database, and uses the personalized haplotype database to determine personalized alignments of nucleotide reads from the genome sample.For example, Figures 3A and 3B illustrate a personalized sequencing system 106 that performs initial read mapping and alignment 308a of nucleotide read 302 with corresponding genomic regions of a reference genome 304, generates a personalized haplotype database 314 based on the initial alignment of nucleotide read 302, and uses the personalized haplotype database 314 to perform personalized read mapping and alignment 308b of nucleotide read 302.
[0077] As shown in Figure 3A, in some embodiments, the personalized sequencing system 106 performs read mapping and alignment 308a to determine the initial alignment of nucleotide read 302 with corresponding genomic regions of a reference genome 304. For example, the personalized sequencing system 106 uses a population haplotype database 306 associated with the reference genome 304 to generate the initial read alignment of nucleotide read 302. In some embodiments, for example, the population haplotype database 306 includes a graph reference genome comprising primary continuous sequences enhanced by a plurality of alternating continuous sequences representing population haplotypes associated with reference genome 304.
[0078] Alternatively, in some embodiments, the population haplotype database 306 includes a data structure encoding allelic variant differences between population haplotypes and primary continuous sequences of reference genome 304. In such embodiments, to achieve initial read mapping and alignment 308a, the personalized sequencing system 106 determines the initial alignment of nucleotide reads 302 by: (i) determining one or more candidate reference alignments for each nucleotide read against primary continuous sequences of reference genome 304, and (ii) determining the initial alignment of each nucleotide read against a corresponding genomic region of reference genome 304 by re-scoring one or more candidate reference alignments based on comparing each nucleotide read against allelic variant differences indicated within the population haplotype database 306.
[0079] Furthermore, as shown in FIG3A, after determining the initial alignment of nucleotide read 302 through initial read mapping and alignment 308a, the personalized sequencing system 106 generates an alignment data file 310 (e.g., an initial alignment data file, such as a BAM file or other file type) that indicates the initial alignment and, in some embodiments, indicates additional information associated with nucleotide read 302, the initial alignment, and / or the population haplotypes aligned with one or more nucleotide reads in nucleotide read 302 according to the initial alignment.In one or more embodiments, for example, the alignment data file 310 generated via initial read mapping and alignment specification 13 / 35 pages 18 CN 121359209 A 308a includes nucleotide reads 302 (e.g., their nucleobases), initial alignment of nucleotide reads 302 to corresponding genomic regions of the reference genome 304, alignment scores associated with the initial alignment (e.g., MAPQ or extended MAPQ scores), and / or haplotype data associated with the initial alignment of nucleotide reads 302 (e.g., imported from a population haplotype database 306).
[0080] In some embodiments, for example, the haplotype data stored within the alignment data file 310 may include indications of one or more candidate population haplotypes for corresponding reference spans of the reference genome 304. In such embodiments, the personalized sequencing system 106 may identify one or more candidate population haplotypes within the reference span of the reference genome 304 based on the similarity between nucleotide reads 302 and population haplotypes within the corresponding reference span. The personalized sequencing system 106 analyzes these one or more candidate population haplotypes to include them when generating the personalized haplotype database 314. Alternatively, in some embodiments, when selecting haplotypes to include in the personalized haplotype database 314, the personalized sequencing system 106 evaluates all population haplotypes in the population haplotype database 306 as candidate population haplotypes (e.g., as further described below with respect to Figures 4 and 5).
[0081] As shown in Figure 3B, in some embodiments, the personalized sequencing system 106 uses alignment data file 310 (e.g., initial alignment and other information associated therewith) to generate the personalized haplotype database 314 during the haplotype selection process 312. In some implementations, for example, the personalized sequencing system 106 generates haplotype scores for candidate population haplotype groups based on comparisons of nucleotide reads 302 with candidate population haplotypes within one or more reference spans of a reference genome 304. Therefore, as shown in FIG3B, the haplotype selection process 312 includes: (i) comparing nucleotide reads 302 with candidate population haplotypes from a population haplotype database 306 according to an initial read alignment indicated by alignment data file 310, and based on that comparison, (ii) selecting population haplotype subgroups from the population haplotype database 306 to be included in a personalized haplotype database 314. Thus, in one or more implementations, the personalized haplotype database 314 includes population haplotype subgroups selected from the population haplotype database 306 for one or more corresponding reference spans (or other partitions) of the reference genome 304.
[0082] In addition, as shown in FIG3B, after generating a personalized haplotype database 314 of the genome sample with nucleotide reads 302, the personalized sequencing system 106 performs personalized read mapping and alignment 308b.In one or more embodiments, for example, personalized sequencing system 106 determines personalized alignments of nucleotide reads 302 with corresponding genomic regions of reference genome 304 in a manner similar to initial read mapping and alignment 308a, although by utilizing a personalized haplotype database 314 instead of population haplotype database 306. In some embodiments, for example, personalized haplotype database 314 includes a data structure that encodes allelic variant differences between selected population haplotypes and primary continuous sequences of reference genome 304 in each of one or more reference spans of reference genome 304. In such embodiments, to achieve personalized read mapping and alignment 308b, the personalized sequencing system 106 determines the personalized alignment of nucleotide reads 302 by: (i) identifying one or more candidate reference alignments for each nucleotide read against a primary continuous sequence of a reference genome 304, and (ii) determining the personalized alignment of each nucleotide read against a corresponding genomic region of the reference genome 304 by re-scoring the one or more candidate reference alignments based on a comparison of each nucleotide read against indicated allele variant differences within a personalized haplotype database 314. Alternatively, in some embodiments, the personalized haplotype database 314 includes a graph reference genome. Such a graph reference genome may include primary continuous sequences enhanced by a plurality of alternating continuous sequences representing population haplotypes selected from a population haplotype database 306 via a haplotype selection process 312.
[0083] Furthermore, as shown in FIG3B, after determining the personalized alignment of nucleotide read 302 through personalized read mapping and alignment 308b, the personalized sequencing system 106 generates a personalized alignment data file 316 indicating the personalized alignment (e.g., a alignment data file, such as a BAM file or a related file type, as described on pages 14 / 35 of the specification, CN 121359209 A). In addition to the data used for personalized alignment, in some embodiments, the personalized alignment data file 316 also includes additional information associated with nucleotide read 302, personalized alignment, and / or population haplotypes aligned with one or more nucleotide reads 302 according to the personalized alignment. In one or more embodiments, for example, the personalized alignment data file 316 generated via personalized read mapping and alignment 308b includes nucleotide reads 302 (e.g., their nucleobases), personalized alignments of nucleotide reads 302 with corresponding genomic regions of the reference genome 304, alignment scores associated with the personalized alignments (e.g., MAPQ or extended MAPQ scores), and / or haplotype data associated with the personalized alignments of nucleotide reads 302 (e.g., imported from a personalized haplotype database 314).
[0084] As previously mentioned, in some embodiments, the personalized sequencing system 106 generates haplotype scores for candidate population haplotype sets based on comparing nucleotide reads from a genome sample with candidate population haplotype sets within a corresponding genomic region of the genome sample. For example, Figure 4 illustrates a personalized sequencing system 106 that compares nucleotide reads 404 from a genome sample with sets of candidate population haplotypes (e.g., candidate population haplotype 408) from a population haplotype database 406 to generate haplotype scores 410. Based on haplotype scores 410, the personalized sequencing system 106 further determines selected population haplotypes 414 to be included in the personalized haplotype database 412.
[0085] As mentioned, the personalized sequencing system 106 can identify nucleotide reads for comparison with candidate population haplotypes within a given genomic region of the genome sample (e.g., the reference span of a reference genome). As shown in Figure 4, for example, the personalized sequencing system 106 receives an initial alignment 402 that identifies a set of nucleotide reads (e.g., as determined by an initial mapping and alignment process, as described above with respect to Figure 3A), and identifies from this set of nucleotide reads nucleotide reads 404 that, according to the initial alignment 402, at least partially align to a given genomic region. Alternatively, the personalized sequencing system 106 may identify nucleotide reads 404 for comparison with candidate population haplotypes 408 without determining or utilizing the initial alignment 402 (e.g., as described above with respect to Figure 2).
[0086] As further depicted in Figure 4, the personalized sequencing system 106 compares the groups of candidate population haplotypes 408 with nucleotide reads 404 within a given genomic region to determine haplotype group fractions 410. In some embodiments, for example, the candidate population haplotypes 408 include all population haplotypes within a population haplotype database 406 that have variations (e.g., variants relative to a primary continuous sequence) within a given genomic region compared to the corresponding reference genome. Alternatively, in one or more embodiments, the personalized sequencing system 106 identifies a limited number of candidate population haplotypes 408 from a population haplotype database 406 for comparison with nucleotide reads 404 within a given genomic region. In some specific embodiments, for example, the personalized sequencing system 106 limits the candidate population haplotypes 408 to haplotypes that have variations locally different from those in a reference genome within a given genomic region (e.g., population haplotypes comprising a unique group of one or more allele variants that differ from other population haplotypes within a given genomic region or reference span).
[0087] Alternatively, in some embodiments, the personalized sequencing system 106 determines the haplotype likelihood of candidate population haplotypes 408 and / or limits candidate population haplotypes 408 to haplotypes identified during initial mapping and alignment as having a relatively high likelihood of matching the aligned nucleotide reads 404. Such individual haplotype likelihoods may be used as “soft” scores or multi-digit values, which provide the basis or input for the haplotype group score 410, which the personalized sequencing system 106 then determines for haplotype groups (e.g., population haplotype pairs). In some cases, the personalized sequencing system 106 determines individual haplotype likelihoods as Boolean values, multi-digit values, or other values (e.g., normalized Boolean values, normalized multi-digit values), which quantify the number of distinct variants (e.g., SNPs) between the mapped and aligned nucleotide reads and the population haplotypes from the population haplotype database 406.
[0088] To determine such individual haplotype likelihoods, in one or more embodiments, for example, personalized sequencing system 106 determines (or receives) a set of alignment scores corresponding to population haplotypes (or subgroups thereof) within population haplotype database 406, associated with nucleotide reads aligned to a given genomic region according to initial alignment 402. In some such embodiments, personalized sequencing system 106 may identify a single population haplotype exhibiting the highest alignment score in this set of alignment scores and assign (or set) the highest alignment score as a baseline (e.g., by setting the highest alignment score to a baseline value of zero) to determine the individual haplotype likelihoods of the remaining population haplotypes from population haplotype database 406. By using the highest alignment score as a baseline, personalized sequencing system 106 may determine the normalized haplotype likelihoods of the remaining population haplotypes mapped and aligned to a set of nucleotide reads. This individual haplotype likelihood can take the form of a Boolean value, a multi-position value, or other quantitative form such as the number of additional or existing variants (e.g., SNPs) that differ between the nucleotide reads and the remaining population haplotypes relative to the population haplotype exhibiting the highest alignment score. After determining such normalized haplotype likelihood and individual haplotype likelihood, in some cases, the personalized sequencing system 106 selects candidate population haplotypes 408 from the population haplotype database 406 based on their respective normalized haplotype likelihoods.
[0089] To further illustrate, in some specific implementations, the personalized sequencing system 106 selects population haplotypes having a corresponding normalized haplotype likelihood within a threshold (e.g., population haplotypes matched by a threshold number of nucleotides within a given genomic region of nucleotide reads 404) as candidate population haplotypes 408.Additionally or alternatively, in one or more embodiments, the personalized sequencing system 106 utilizes normalized haplotype likelihood (or a subgroup thereof, if less than all population haplotypes are considered candidates) when determining the haplotype group score 410 (e.g., to determine initial likelihood, as described below with respect to Figure 5). Therefore, regardless of whether the personalized sequencing system 106 selects candidate population haplotypes 408 from the population haplotype database 406 based on their respective normalized haplotype likelihoods, the personalized sequencing system 106 can assign or identify individual haplotype likelihoods as candidate population haplotypes 408 in order to determine the haplotype group score 410.
[0090] As just noted, based on haplotype likelihood and / or other comparisons of the nucleotide reads 404 with groups of candidate population haplotypes 408 within a given genomic region, the personalized sequencing system 106 generates a haplotype group score 410 for the corresponding group of candidate population haplotypes 408. In one or more embodiments, for example, haplotype group score 410 represents the likelihood or probability of a genomic sample representing a nucleotide read within a given genomic region, with the corresponding group of candidate population haplotypes 408 representing that region. Therefore, personalized sequencing system 106 determines a selected population haplotype 414 having the highest relative haplotype group score 410 within a given genomic region and generates a personalized haplotype database 412 including the selected population haplotype 414 for determining personalized alignments of nucleotide reads 404 and / or additional nucleotide reads from the genomic sample.
[0091] In various embodiments, personalized sequencing system 106 generates haplotype group scores 410 for candidate population haplotype groups comprising a different number of candidate population haplotypes 408 in each group. In some embodiments, for example, personalized sequencing system 106 identifies a genomic region (e.g., a reference span) corresponding to a diploid region of a reference genome and, in response, generates haplotype group scores 410 for groups including pairs of candidate population haplotypes 408. In contrast, in some embodiments, the personalized sequencing system 106 identifies a genomic region (e.g., a reference span) corresponding to a haploid region of a reference genome (such as the X or Y chromosome). In response to identifying the haploid genomic region, the personalized sequencing system 106 generates haplotype group scores 410 for groups comprising individual haplotypes from candidate population haplotypes 408. Furthermore, the personalized sequencing system 106 may determine haplotype group scores 410 for groups comprising more than two candidate population haplotypes 408.
[0092] As mentioned, in some embodiments, the personalized sequencing system 106 generates haplotype group scores 410 for each combination of candidate population haplotypes 408 (e.g., haplotype group scores for each possible pair of candidate population haplotypes 408).Additionally, in one or more embodiments, the personalized sequencing system 106 considers each nucleotide read in the nucleotide reads 404 identified within a given reference span to adjust each haplotype fraction 410 in the given reference span. Alternatively, in one or more embodiments, the personalized sequencing system 106 considers each adjacent mapped nucleotide read in the nucleotide reads 404 within a given reference span to adjust each haplotype fraction 410 in the given reference span.
[0093] As another alternative, in some embodiments, the personalized sequencing system 106 implements an iterative probabilistic algorithm, such as a variational Bayesian model, to partition the reads within each reference span and generate haplotype fractions 410 based on the partitioned reads within each respective reference span. In some implementations, for example, the personalized sequencing system 106 utilizes a variational Bayesian model to partition corresponding reads based on the classification of each nucleotide read in nucleotide reads 404 inherited from a corresponding first or second parent. In some such implementations, the variational Bayesian model includes an iterative probabilistic algorithm that uses variational inference to predict the posterior distribution of latent variables to determine the corresponding likelihood of the nucleotide read inherited from the corresponding first or second parent. Such corresponding likelihoods of nucleotide reads can be any likelihood (e.g., 0.50 for the first parent and 0.50 for the second parent; 0.25 for the first parent and 0.75 for the second parent).
[0094] After the nucleotide reads 404 are partitioned accordingly, the personalized sequencing system 106 can generate individual haplotype scores for the classified reads relative to each candidate population haplotype in the candidate population haplotype 408, rather than generating haplotype group scores 410 for each group of population haplotypes on a group-by-group basis. Therefore, the personalized sequencing system 106 can generate haplotype group scores 410 for candidate population haplotypes 408 by combining the haplotype scores of each haplotype from the nucleotide reads inherited from the first parent with the haplotype scores of each haplotype from the nucleotide reads inherited from the second parent. Thus, by dividing the nucleotide reads 404 according to parental inheritance, the personalized sequencing system 106 can significantly reduce the number of unique groups of candidate population haplotypes 408 to be scored, thereby avoiding scoring each pair of candidate population haplotypes.
[0095] As previously mentioned, in some embodiments, the personalized sequencing system 106 utilizes an interpolation model to generate haplotype group scores for candidate population haplotype groups across adjacent reference spans of the reference genome.According to one or more implementations, for example, Figure 5 illustrates a personalized sequencing system 106 that uses a Hidden Markov Model (HMM) algorithm 508 to generate a set of haplotype scores for candidate population haplotypes across a set of reference spans 504a, 504b to 504n of a reference genome. Specifically, Figure 5 illustrates a personalized sequencing system 106 that generates haplotype posterior probabilities 510a, 510b to 510n for haplotypes within the set of reference spans 504a, 504b to 504n. Based on the generated haplotype posterior probabilities 510a–510n, the personalized sequencing system 106 selects haplotypes 512a, 512b to 512n for the corresponding reference spans 504a, 504b to 504n to include them in a personalized haplotype database 514.
[0096] As shown in Figure 5, the personalized sequencing system 106 identifies nucleotide reads 502a, 502b to 502n from the genome sample within corresponding reference spans 504a-504n of a reference genome associated with the genome sample. In one or more embodiments, for example, the set of reference spans 504a-504n includes bins that divide the reference genome (or a portion thereof) into bins spanning a selected number of base positions, such as, but not limited to, 1,000 base positions per reference span; 4,000 base positions per reference span; or 16,000 base positions per reference span. However, any number of base positions can be used for the reference span.
[0097] Furthermore, within the reference span 504a-504n, the personalized sequencing system 106 identifies nucleotide reads 502a, 502n to 502n for comparison with various groups of candidate population haplotypes to determine the likelihood of the corresponding haplotype group 506a, 506b to 506n. As indicated above with respect to Figure 4, for example, the personalized sequencing system 106 may compare nucleotide reads 502a-502n with candidate population haplotypes by mapping and aligning nucleotide reads 502a-502n with candidate population haplotypes, determining alignment scores, and determining haplotype likelihood in part based on the corresponding read alignments. In some implementations, for example, the personalized sequencing system 106 further generates haplotype likelihoods (e.g., haplotype likelihood 506a) for a given reference span (e.g., reference span 504a) within the reference spans 504a-504n. In some such implementations, the personalized sequencing system 106 scores each candidate haplotype group based on a comparison of the nucleotide bases of the corresponding nucleotide read (e.g., nucleotide read 502a) within the given reference span with the nucleotide bases of each candidate haplotype group within the given reference span.Therefore, as shown in Figure 5, the personalized sequencing system 106 generates multiple haplotype likelihoods for each corresponding reference span in the reference spans 504a-504n.
[0098] Furthermore, in one or more embodiments, the personalized sequencing system 106 generates haplotype scores for the corresponding candidate population haplotypes in each reference span in the reference spans 504a-504n based on the haplotype likelihoods 506a-506n. Additionally, as shown in Figure 5, for example, the personalized sequencing system 106 utilizes an interpolation model, such as the HMM algorithm 508, to generate haplotype posterior probabilities 510a, 510b to 510n for the candidate population haplotypes within the corresponding reference spans 504a-504n based on the haplotype likelihoods in the haplotype likelihoods 506a-506n of adjacent reference spans in the reference spans 504a-504n.
[0099] As shown in FIG5, in the forward pass of HMM algorithm 508 across the set of reference spans 504a-504n, for example, personalized sequencing system 106 uses haplotype likelihoods 506a-506n to generate a set of positive probabilities for each reference span in the set of reference spans 504a-504n. In some embodiments, for example, the forward pass of HMM algorithm 508 begins with haplotype likelihood 506a of reference span 504a and ends with haplotype likelihood 506n of reference span 504n, thereby generating positive probabilities for each reference span in the set of reference spans 504a-504n using inputs from the corresponding adjacent reference spans (e.g., previous reference spans during the forward pass of HMM algorithm 508).
[0100] In the reverse propagation of the HMM algorithm across the reference spans 504a-504n, the personalized sequencing system 106 uses the forward probability of each candidate population haplotype within the corresponding reference spans 504a-504n to generate updated probabilities for the corresponding candidate population haplotypes. Based on the updated probabilities, the personalized sequencing system 106 generates haplotype group posterior probabilities 510a-510n for the corresponding candidate population haplotypes within the reference spans 504a-504n. Thus, as mentioned with respect to haplotype group likelihood 506a-506n, the personalized sequencing system 106 generates multiple haplotype group posterior probabilities for each corresponding reference span in the reference spans 504a-504n, and in some embodiments, generates haplotype group fractions for the corresponding candidate population haplotypes based on the haplotype group posterior probabilities 510a-510n from each corresponding reference span in the reference spans 504a-504n.
[0101] As further depicted in Figure 5, the personalized sequencing system 106 selects haplogroups 512a-512n for the corresponding reference spans 504a-504n based on the posterior probabilities 510a-510n of the corresponding haplogroups or the haplogroup scores derived therefrom (e.g., selecting one haplogroup for each reference span) to generate a personalized haplotype database 514. Therefore, in some cases, the personalized sequencing system 106 may select different candidate population haplogroups across the reference spans 504a-504n to include in the personalized haplotype database 514, such that haplogroups 512a-512n vary across the reference spans 504a-504n. Furthermore, in some embodiments, the personalized sequencing system 106 utilizes recombination parameters of the HMM algorithm to reduce the variation of haplogroup posterior probabilities 510a-510n across the reference span 504a-504n, thus reducing, in some cases, the variation of selected haplogroups 512a-512n included in the personalized haplogroup database 514 across the reference span 504a-504n.
[0102] As described above, in some embodiments, the personalized sequencing system 106 determines haplogroup scores based on comparisons of nucleotide reads with allele variant differences between candidate population haplogroups and the reference genome, as indicated in the population haplogroup database. Additionally, in some embodiments, the personalized sequencing system 106 generates a personalized haplogroup database for a given genome sample, which includes allele variant differences between population haplogroups and the reference genome selected in a limited number of reference spans covering corresponding genomic regions of the reference genome. In some implementations, for example, personalized sequencing system 106 utilizes a population haplotype database and / or generates a personalized haplotype database encoded as a haplotype data structure, as described in Michael Ruehle, Enhanced Mapping and Alignment of Nucleotide Reads Utilizing an Improved Haplotype Data Structure with Allele-Variant Differences, U.S. Provisional Application No. 63 / 613,574 (filed December 21, 2023) (hereinafter referred to as Specification 18 / 35 pages 23 CN 121359209 A Ruehle), the entire contents of which are incorporated herein by reference.
[0103] For further illustration, Figures 6A and 6B illustrate implementations of haplotype data structures encoding allele variant differences between population haplotypes and a reference genome.Specifically, Figure 6A illustrates a set of basic horizontal bins 602a, 602b to 602n of a population haplotype database 600, which encodes allelic variant differences of population haplotypes within the corresponding reference spans 604a, 604b to 604n of the reference genome, while Figure 6B illustrates a set of basic horizontal bins 612a, 612b to 612n of a personalized haplotype database 610, which encodes allelic variant differences of selected population haplotypes within the corresponding reference spans 604a, 604b to 604n of the reference genome. Although the following paragraphs describe the various bins of either the population haplotype database 600 or the personalized haplotype database 610, the bins of each of these databases may be encoded or otherwise represented as a specific file type, such as a TSV file or a comma-separated value (CSV) file.
[0104] As shown in Figure 6A, the population haplotype database 600 includes at least one base level comprising a set of base level bins 602a-602n that divide the genomic regions of a reference genome into corresponding sets of base level reference spans 604a-604n. In one or more embodiments, each base level reference span in the set of base level reference spans 604a-604n comprises a genomic region of a first length between corresponding genomic coordinates of the reference genome, thus dividing the genomic regions of the reference genome into multiple bins of equal length spanning the reference genome. In various specific embodiments, for example, the length of the base level reference span may be approximately provided to the personalized sequencing system 106 for the average or maximum length of nucleotide reads used for mapping and alignment. Alternatively, the base level reference span may be selected in other ways to span a predetermined number of nucleotides from the genomic coordinates or regions of a linear reference sequence, such as, but not limited to, 100 base pairs or 1,000 base pairs per base level bin.
[0105] As further shown in Figure 6B, the base level boxes 602a-602n of the population haplotype database 600 include encoded variant data of nucleotide variants from locally different population haplotypes 606a-606n within the corresponding base level boxes. Specifically, each locally different population haplotype within a given base level box includes a unique set of one or more allele variant differences relative to other population haplotypes that also have variation in the genomic region of the corresponding base level reference span of the given base level box. As shown in Figure 6A, for example, each row of locally different population haplotype 606a includes a unique set of allele variant differences (represented as a single letter representing a specific nucleotide) relative to other rows, such that no two rows are identical, although there may be limited overlap between allele variant differences, as indicated by the first two rows of base level box 602a.Therefore, in one or more embodiments, population haplotypes having the same nucleotide variant within a given base level box are encoded as a locally distinct population haplotype within that given base level box.
[0106] In various embodiments, each base level box of the population haplotype database 600 may include a different number of locally distinct population haplotypes. As shown in FIG6A, for example, the set of locally distinct population haplotypes 606a included in base level box 602a includes four locally distinct population haplotypes (as indicated by the fourth row of the depicted matrix), the set of locally distinct population haplotypes 606b included in base level box 602b includes five locally distinct population haplotypes, and the set of locally distinct population haplotypes 606n included in base level box 602n includes three locally distinct population haplotypes. In practice, each base level box of the population haplotype database 600 may include any number of locally distinct population haplotypes, including every population haplotype in the dataset or no population haplotypes (e.g., in the case where there are no population haplotypes with allelic variant differences in the genomic region corresponding to a given box).
[0107] As further shown in FIG6A, the basal level boxes 602a, 602b to 602n include corresponding allele variant differences 608a, 608b to 608n for each locally distinct population haplotype group 606a, 606b to 606n. For example, the variant data encoded in the basal level box 602a includes a set of locally distinct population haplotypes 606a, which includes one or more locally distinct population haplotypes, wherein, for each corresponding locally distinct population haplotype, allele variant difference 608a is included. In some embodiments, for example, each basic level box (e.g., within the set of basic level boxes 602a-602n) includes a matrix containing corresponding variant data representing allelic variant differences from (e.g., from locally different population haplotypes 606a-606n in the corresponding set) and the variant locations of these allelic variant differences. In various embodiments, the variant data within each basic level box includes data indicating single nucleotide polymorphisms (SNPs) and / or insertions or deletions (insertions / deletions) at corresponding genomic coordinates of a reference genome (e.g., a primary continuous sequence) (e.g., allelic variant differences 608a-608n).
[0108] As further shown in FIG6A, in some embodiments, each of the basic level boxes 602a, 602b, and 602n also includes population frequencies 609a, 609b, and 609n of the corresponding locally distinct population haplotypes 606a-606n (e.g., relative frequencies of haplotype alleles within a reference population) (or may refer to data including such population frequencies). In some embodiments, for example, when initial read alignments of nucleotide reads are determined using the population haplotype database 600 (e.g., as described above with respect to FIG3A), the personalized sequencing system 106 adjusts the alignment scores based on the population frequencies 609a-609n.
[0109] As shown in FIG6B, in contrast, the personalized haplotype database 610 generated by the personalized sequencing system 106 includes variant data of customized population haplotype groups within each of the basic level boxes 612a-612n. Specifically, Figure 6B depicts each base level bin with two population haplotypes, thus representing the inferred diploid of the genome sample corresponding to the personalized haplotype database 610 (e.g., the genome sample for which the personalized sequencing system 106 generates the personalized haplotype database 610). As previously described, in various embodiments, the personalized sequencing system 106 generates a personalized haplotype database having a different number of selected haplotypes for each reference span (e.g., one, two, or more haplotypes for each selected haplotype group).
[0110] Additionally, as shown in Figure 6B, the personalized haplotype database 610 includes at least one base level comprising base level bins 612a-612n that divide the genomic regions of the reference genome into corresponding base level reference spans 604a-604n. Furthermore, the set of basic bins 612a, 612b to 612n in the personalized haplotype database 610 includes encoded variant data of nucleotide variants from selected population haplotypes 616a, 616b to 616n of the respective group. In some embodiments, for example, each basic bin in the set of basic bins 612a-612n includes a matrix containing corresponding variant data representing allelic variant differences from (e.g., from selected population haplotypes 616a-616n of the respective group) and the variant locations of these allelic variant differences. In various embodiments, the variant data within each basic bin includes data indicating single nucleotide polymorphisms (SNPs) and / or insertions or deletions (insertions / deletions) at corresponding genomic coordinates of a reference genome (e.g., a primary continuous sequence) (e.g., allelic variant differences 618a, 618b to 618n).Additionally, as shown in Figure 6B, the population frequencies 619a, 619b, and 619n of the selected population haplotypes 616a, 616b, and 616n in the corresponding groups are adjusted relative to the corresponding population frequencies 609a, 609b, and 609n (see Figure 6A) to reflect the predicted diploids within each of the basal level bins 612a, 612b, and 612n.
[0111] As described above, in some embodiments, the personalized sequencing system 106 utilizes a haplotype data structure (e.g., to implement a population haplotype database and / or a personalized haplotype database) in which the genomic regions of the reference genome are hierarchically divided into multiple levels of bins corresponding to the nucleotide spans within the reference genome. For example, Figure 7 illustrates a haplotype data structure 700, which has a base level 702 comprising a set of base level boxes 704 and multiple consecutive levels 706a, 706b, 706c to 706n spanning a large continuous span of nucleotides across a reference genome. Specifically, haplotype data structure 700 includes: a base level 702 comprising the set of base level boxes 704 that commonly span the reference genome (see page 20 / 35, 25 CN 121359209 A), and multiple consecutive levels 706a-706n spanning the primary continuous sequences of the reference genome (see page 20 / 35, 25 CN 121359209 A), and multiple consecutive levels 706a-706n spanning the primary continuous sequences (see pages 70 / 35, 25 CN 121359209 A), including higher level boxes 708a, 708b, 708c to 708n and offset higher level boxes 709a, 709b, 709c, etc.). As indicated in Figure 7, consecutive level 706n includes higher level boxes 708n and corresponding offset higher level boxes. However, due to constraints on the attached drawing space, Figure 7 does not depict the corresponding offset higher level bins for continuous level 706n and higher level bins 708n. As further indicated by the ellipses (or dots) in Figure 7, the personalized sequencing system 106 can identify, determine, generate, or utilize more base level bins, continuous levels, higher level bins, and / or offset higher level bins than those depicted in Figure 7. Although the haplotype database 700 in Figure 7 depicts the database structure of an individual haplotype database, a similar data structure can be used for the population haplotype database described by Ruehle. As further indicated above, the haplotype data structure 700 can be encoded or otherwise represented as a specific file type, such as a TSV file or a CSV file.
[0112] As shown in Figure 7, the base level 702 of the haplotype data structure 700 includes a set of base level bins 704 corresponding to the corresponding set of base level reference spans of the primary continuous sequences of the reference genome. Each reference span in the set of base level reference spans corresponds to a genomic region of a first length between the corresponding genomic coordinates of the reference genome.In one or more embodiments, for example, each reference span in the set of base-level reference spans comprises 1,000 base pairs (1 kbp) of a primary continuous sequence of a reference genome. Alternatively, the first length of the base-level reference span may be less than or greater than 1 kbp, such as, but not limited to, 250 bp, 500 bp, 1500 bp, 5 kbp, 10 kbp, etc. Thus, in various embodiments, the set of base-level bins 704 collectively spans the entire primary continuous sequence or a genomic region of interest, such as, but not limited to, the entire chromosome.
[0113] As further indicated in FIG7, the set of base-level bins 704 of base level 702 includes variant data of nucleotide variants from the corresponding group of population haplotypes. In a specific implementation of the population haplotype database, for example, each base-level bin of base level 702 includes variant data of each locally distinct population haplotype. As mentioned, each locally distinct population haplotype comprises a unique set of one or more allelic variant differences relative to other population haplotypes within the corresponding base-level reference span of a given base-level bin in the set of base-level bins 704. However, in a specific implementation of the personalized haplotype database, each basic level bin of basic level 702 includes variant data of population haplotypes selected by the personalized sequencing system 106 (e.g., as described with respect to FIG. 5). As shown in FIG. 7, for example, this set of basic level bins 704 includes corresponding pairs of population haplotypes selected by the personalized sequencing system 106, as indicated by the numbers (“2(0..1)”) associated with each basic level reference span of this set of basic level bins 704.
[0114] Additionally, as shown in FIG. 7, the haplotype data structure 700 includes multiple consecutive levels 706a-706n of higher level bins 708a-708n. For example, the first consecutive level 706a includes a first set of higher level bins 708a, which corresponds to a first set of higher level reference spans of the primary consecutive sequence of the reference genome. Each reference span in the first set of higher-level reference spans corresponds to a second-length extended genomic region between corresponding genomic coordinates of the reference genome, wherein the extended genomic region is extended relative to the genomic region represented by that set of basic-level reference spans, such that the second length (of the corresponding first set of higher-level reference spans) is longer than the first length (of the set of basic-level reference spans). More specifically, as shown in Figure 7, each higher-level bin in the first set of higher-level bins 708a of the first consecutive level 706a corresponds to a pair of consecutive basic-level bins 704 of the set of basic-level bins from the basic level 702 of the haplotype data structure 700.
[0115] Furthermore, as shown in FIG7, the multiple consecutive levels 706a-706c of the haplotype data structure 700 include corresponding sets of offset higher level boxes 709a-709c, and the consecutive level 706n of the haplotype data structure 700 includes a higher level box 708n and a corresponding offset higher level box. For example, the first consecutive level 706a includes a set of offset higher level boxes 709a corresponding to a first set of offset higher level reference spans of the primary consecutive sequence of the reference genome. Each reference span in the first set of offset higher level reference spans corresponds to an offset extended genomic region of a second length (i.e., the same length as the reference span in the first set of consecutive reference spans, as specified on page 21 / 35 of the specification, CN 121359209 A) between the corresponding genomic coordinates of the reference genome. In a similar manner to the first set of higher level boxes 708a, the first set of offset higher level boxes 709a corresponds to the corresponding pair of consecutive basic level boxes 704 of the basic level box 704 from the basic level 702 of the haplotype data structure 700. Furthermore, as shown in the figure, the corresponding reference span of the first set of offset higher level bins 709a is offset relative to the reference span of the first set of higher level bins 708a, such that each pair of consecutive basic level bins from this set of basic level bins 704 is represented by a higher level bin or an offset higher level bin from the first consecutive level 706a.
[0116] In addition, each additional consecutive level 706b-706n of the haplotype data structure 700 includes additional higher level bins 708b-708n corresponding to a corresponding additional higher level reference span, which corresponds to a further extended genomic region between the genomic coordinates of the primary consecutive sequence of the reference genome. Specifically, as shown in Figure 7, each higher level bin (or offset higher level bin) of a given consecutive level of the haplotype data structure 700 spans a combined genomic region of a pair of consecutive bins of the previous level of the haplotype data structure 700 (e.g., as indicated by the arrows linking the various bins in Figure 7). For example, the first illustrated bin in the set of higher level bins 708c spans the same genomic region represented by the first two illustrated bins in the set of higher level bins 708b. Similarly, the first illustrated bin in the set of higher level bins 708b spans the same genomic region represented by the first two illustrated bins in the set of higher level bins 708a. In fact, each successive level includes a higher level bin corresponding to a pair of successive bins from the previous level of the haplotype data structure 700.
[0117] Furthermore, in some embodiments, the corresponding higher level bin of each successive level of the haplotype data structure 700 includes a variant data index that indexes a combination of variant data from the corresponding base level bin of the base level 702.Specifically, each of the higher level boxes 708a-708c and the offset higher level boxes 709a-709c, as well as each of the higher level box 708n and the corresponding offset higher level box of the continuous level 706n, includes a variant data index that indexes a combination of variant data from the corresponding base level box in the base level box 704 of that set. Furthermore, the variant data index includes an indication of locally distinct population haplotypes within each corresponding higher level box or offset higher level box. As shown in Figure 7, for example, one offset higher level box 709b of the continuous level 706b indicates two locally distinct population haplotypes (indicated by "2 haplotypes (0..1)"). Additionally, as shown, two boxes in the higher level box 708a from the previous continuous level (first continuous level 706a) respectively indicate two locally distinct population haplotypes (indicated by "2(0..1)") and one locally distinct population haplotype (indicated by "1(0)").
[0118] In one or more embodiments, each higher-level box of a consecutive level includes a variant data index that indicates locally distinct population haplotypes and links the higher-level box to variant data within the corresponding base-level box, excluding variant data from the respective base-level box, thus avoiding redundant encoding of variant data within the haplotype data structure 700. Referring to consecutive level 706b, for example, the aforementioned box in offset higher-level box 709b indicating four locally distinct population haplotypes may include a variant data index that indexes how the locally distinct population haplotypes of the corresponding higher-level box (in higher-level box 708a) from the previous consecutive level (first consecutive level 706a) combine to form the four locally distinct population haplotypes of the aforementioned box. Furthermore, each higher-level box in the corresponding higher-level box 708a may include a variant data index that indexes the population haplotype (and its variant data) indicated within the corresponding base-level box (in the set of base-level boxes 704) from base level 702.
[0119] Furthermore, while FIG. 7 depicts a haplotype data structure 700 having haplotype pairs selected with a reference span relative to the corresponding base level bins in this set of base level bins 704, the personalized sequencing system 106 may alternatively evaluate and select haplotype pairs across the reference span of bins in corresponding consecutive levels. In some embodiments, for example, the personalized sequencing system 106 evaluates candidate population haplotypes within a reference span corresponding to a higher level bin 708b, and then populates the haplotype data structure 700 with haplotype data from the selected candidate population haplotypes, as described above.Specification 22 / 35 pages 27 CN 121359209 A
[0120] As described above, in some of the described embodiments, the personalized sequencing system 106 achieves mapping and alignment of nucleotide reads from a genomic sample with genomic regions of a reference genome with increased accuracy. For illustration, Figures 8 and 9 show experimental results of the personalized sequencing system 106 generating a personalized haplotype database according to some of the disclosed embodiments and using it to determine personalized alignments of nucleotide reads. Specifically, Figures 8 and 9 illustrate comparison results of identifying single nucleotide polymorphisms (SNPs) based on read alignments generated according to one or more embodiments.
[0121] Specifically, Figure 8 includes a table of experimental results identifying SNPs in nucleotide reads aligned by existing sequencing systems and the personalized sequencing system 106. As shown, the columns of the table correspond to false positives (FP), false negatives (FN), incorrect heterozygous or homozygous genotype detection (Hethom), and combined false positives and false negatives (FP+FN), respectively. In addition, as shown in Figure 8, the first two rows of the experimental results table (excluding the header row) correspond to the results generated by the existing sequencing system using the following databases respectively: (i) a population haplotype database including 128 global population haplotypes (“HapDB Global 128 samples”) and (ii) a population haplotype database including 16 ancestor-specific population haplotypes based on ancestor selection of the tested genome samples (“HapDB Euro 16 samples”).
[0122] In contrast, the last three rows of the experimental results table correspond to the results generated by the personalized sequencing system 106 using the following databases: (i) a first personalized haplotype database, which is generated for the corresponding genome sample using a first probabilistic model (“Personalized (Exhaustive 1)”) that exhaustively scores all possible candidate haplotype pairs for each reference span; (ii) a second personalized haplotype database, which is generated for the corresponding genome sample using a second probabilistic model (“Personalized (Exhaustive 2)”) that exhaustively scores all possible candidate haplotype pairs for each reference span; and (iii) a third personalized haplotype database, which is generated for the corresponding genome sample using a variational Bayesian model (“Personalized (Variational Bayesian)”) that uses parental genetics to divide nucleotide reads to achieve individual scoring of candidate haplotypes.
[0123] In fact, as shown in FIG8, each of the three exemplary embodiments of the personalized sequencing system 106 exhibits improved overall accuracy relative to the existing sequencing systems depicted in terms of identifying SNPs within nucleotide reads (e.g., in terms of FP, FN, Hethom accuracy and combined FP+FN metric).Furthermore, while personalization achieved by matching the ancestry of a genome sample can further improve genotype detection accuracy in some cases (see, for example, results from the “HapDB Euro 16 sample”), the personalized sequencing system 106 can produce similar or improved results without contextual information about a given genome sample, thus providing increased accuracy with increased flexibility and effectiveness compared to many existing systems.
[0124] Additionally, Figure 9 includes an additional illustration of experimental results identifying SNPs in nucleotide reads aligned by the existing sequencing system and the personalized sequencing system 106. Specifically, Figure 9 depicts read stacks 906b and 906c (stacks of nucleotide reads mapped to and aligned to the corresponding genomic regions of the reference genome 902) generated by the existing sequencing system and the personalized sequencing system 106, respectively, compared to a ground truth stack 906a that includes at least one known single nucleotide variant (such as the identified known single nucleotide variant 904a). Figure 9 also depicts reads aligned with relatively low confidence (e.g., with low Mapping Quality (MAPQ) scores) by using unfilled rectangles to represent such relatively low-confidence alignments, and reads aligned with relatively high confidence (e.g., all reads within ground truth stack 906a) by using filled rectangles to represent such relatively high-confidence alignments. As shown, read stack 906b in the existing sequencing system results in false negatives 904b for the identified known single nucleotide variant 904a because multiple reads are aligned with corresponding genomic regions having zero or near-zero Mapping Quality (MAPQ) scores (as indicated by the stack of unfilled rectangles). In contrast, read stack 906c in the personalized sequencing system 106 includes true positives 904c for the identified known single nucleotide variant 904a because multiple reads are aligned with corresponding genomic regions with relatively high Mapping Quality (MAPQ) scores (as indicated by the stack of filled rectangles).
[0125] In fact, as shown in Figures 8 and 9, the personalized sequencing system 106 can generate a personalized haplotype database and, using the specification (pages 23 / 35, 28 CN 121359209 A), effectively determine read alignments of nucleotide reads from a genomic sample, achieving improved accuracy in identifying variants compared to existing sequencing systems, as indicated by the number of false positives, false negatives, and incorrect heterozygous or homozygous genotypes detected in the provided experimental results.
[0126] Turning now to Figure 10, this figure illustrates an exemplary flowchart of a series of actions for generating and utilizing a personalized haplotype database to determine personalized alignments of a set of nucleotide reads from a genomic sample, according to one or more embodiments. While Figure 10 illustrates actions according to a particular embodiment, alternative embodiments may omit, add, reorder, and / or modify any actions shown in Figure 10.The actions of FIG10 can be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform the actions depicted in FIG10. In yet another embodiment, the system includes at least one processor and a non-transitory computer-readable medium that includes instructions that, when executed by one or more processors, cause the system to perform the actions of FIG10.
[0127] As shown in FIG10, the series of actions 1000 includes: action 1002 of identifying nucleotide reads and candidate population haplotypes within a set of reference spans of a reference genome; action 1004 of generating haplotype group scores for candidate population haplotype groups; action 1006 of generating a personalized haplotype database based on the haplotype group scores; and action 1008 of using the personalized haplotype database to determine personalized alignments of the set of nucleotide reads.
[0128] For example, the series of actions 1000 may include actions for performing any of the operations described in the following clauses:
[0129] Clause 1. A computer-implemented method comprising:
[0130] identifying a set of nucleotide reads from a genome sample and candidate population haplotypes from a population haplotype database within a set of reference spans of a reference genome;
[0131] generating haplotype group scores for groups of the candidate population haplotypes for the set of reference spans of the reference genome, based on comparing the set of nucleotide reads and the candidate population haplotypes;
[0132] generating a personalized haplotype database for the genome sample and based on the haplotype group scores, the personalized haplotype database including population haplotype subgroups from the population haplotype database within the set of reference spans; and
[0133] determining one or more personalized alignments of the set of nucleotide reads to corresponding genomic regions of the reference genome using the personalized haplotype database.
[0134] Clause 2. The computer-implemented method according to Clause 1, the computer-implemented method further comprising:
[0135] identifying one or more distinct k-mers for each candidate population haplotype within a given reference span of the set of reference spans; and
[0136] generating a corresponding haplotype group score in the haplotype group score for the corresponding candidate population haplotype within the given reference span, based on comparing the k-mers of the set of nucleotide reads with the one or more distinct k-mers of the corresponding candidate population haplotype in the set of candidate population haplotypes.
[0137] Clause 3. The computer-implemented method according to Clause 1, the computer-implemented method further comprising:
[0138] determining an initial alignment of the set of nucleotide reads with corresponding genomic regions of the reference genome;
[0139] identifying subgroups of nucleotide reads in the set of nucleotide reads, the subgroups of nucleotide reads being aligned with corresponding reference spans in the set of reference spans according to the initial alignment; and
[0140] generating haplotype scores for the set of reference spans by comparing the subgroups of nucleotide reads with the candidate population haplotypes within the corresponding reference spans in the set of reference spans. Specification 24 / 35 pages 29 CN 121359209 A
[0141] Clause 4. The computer-implemented method according to Clause 3, the computer-implemented method further comprising:
[0142] as indicated in the population haplotype database, identifying allelic variant differences between the population haplotype and the primary continuous sequence at a corresponding genomic region of the reference genome; and
[0143] determining one or more initial alignments of the initial alignments of the set of nucleotide reads by re-scoring the candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on the allelic variant differences.
[0144] Clause 5. The computer-implemented method according to any one of Clauses 3 to 4, the computer-implemented method further comprising:
[0145] comparing the set of nucleotide reads with haplotype variants based on the corresponding genomic regions of the initial alignment, identifying one or more population haplotypes for each nucleotide read in the set of nucleotide reads; and
[0146] generating an alignment data file comprising:
[0147] the initial alignment of the set of nucleotide reads;
[0148] an alignment score corresponding to the initial alignment; and
[0149] the one or more population haplotypes identified for each nucleotide read in the set of nucleotide reads.
[0150] Clause 6. The computer-implemented method according to any one of Clauses 1 to 5, the computer-implemented method further comprising:
[0151] determining haplotype group likelihood for the set of reference spans by comparing pairs of the set of nucleotide reads and the candidate population haplotypes within a corresponding reference span in the set of reference spans; and
[0152] generating haplotype group scores for the set of reference spans based on the haplotype group likelihood.
[0153] Clause 7. The computer-implemented method according to any one of Clauses 1 to 6, the computer-implemented method further comprising:
[0154] identifying at least one reference span in the set of reference spans corresponding to a haploid region of the reference genome; and
[0155] generating haplotype group scores for each candidate population haplotype for the at least one reference span based on comparing the set of nucleotide reads and the candidate population haplotypes within the at least one reference span.
[0156] Clause 8. The computer-implemented method according to any one of Clauses 1 to 7, the computer-implemented method further comprising: based on the haplotype group scores, including a corresponding subgroup of the candidate population haplotype within the population haplotype subgroup of the personalized haplotype database for each reference span in the set of reference spans.
[0157] Clause 9. The computer-implemented method according to Clause 8, the computer-implemented method further comprising:
[0158] generating haplotype group likelihoods for the set of reference spans by comparing the set of nucleotide reads with the candidate population haplotypes;
[0159] generating a set of haplotype group posterior probabilities for each of the set of reference spans based on the corresponding haplotype group likelihoods of adjacent reference spans in the set of reference spans using a Hidden Markov Model (HMM) algorithm; and
[0160] generating a corresponding group of haplotype group scores for the set of reference spans based on the set of haplotype group posterior probabilities of each of the reference spans in the set of reference spans.
[0161] Clause 10. The computer-implemented method according to Clause 9, further comprising:
[0162] generating a set of positive probabilities for each reference span in the set of reference spans using the haplotype group likelihood in the forward pass of the HMM algorithm across the set of reference spans;
[0163] updating the set of positive probabilities for each reference span in the set of reference spans using the backward pass of the HMM algorithm, generating a set of updated probabilities for each reference span in the set of reference spans; and
[0164] generating the set of haplotype group posterior probabilities for each reference span in the set of reference spans based on the set of updated probabilities for each reference span in the set of reference spans.
[0165] Clause 11. The computer-implemented method according to any one of Clauses 9 to 10, further comprising: using recombination parameters of the HMM algorithm to reduce the variation of the haplotype group posterior probabilities across the set of reference spans.
[0166] Clause 12. The computer-implemented method according to any one of Clauses 1 to 11, the computer-implemented method further comprising:
[0167] classifying each nucleotide read in the set of nucleotide reads within each respective reference span of the set of reference spans as inherited from a first parent or a second parent using a variational Bayesian model;
[0168] generating individual haplotype scores for population haplotypes selected from the population haplotype database based on each classified nucleotide read in the set of nucleotide reads; and
[0169] generating the haplotype group scores for pairs of population haplotypes selected from the population haplotype database by combining the individual haplotype scores from nucleotide reads inherited from the first parent with the individual haplotype scores from nucleotide reads inherited from the second parent.
[0170] Clause 13. The computer-implemented method according to any one of Clauses 1 to 12, the computer-implemented method further comprising: generating haplotype subsets for corresponding pairs of locally distinct population haplotypes within a given reference span of the set of reference spans, each locally distinct population haplotype comprising a unique set of one or more allele variant differences relative to other population haplotypes within the given reference span.
[0171] Clause 14. The computer-implemented method according to any one of Clauses 1 to 13, the computer-implemented method further comprising:
[0172] identifying allele variant differences between population haplotypes and primary continuous sequences in the population haplotype subset at the corresponding genomic region of the reference genome, as indicated in the personalized haplotype database; and
[0173] determining one or more personalized alignments of the one or more personalized alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary continuous sequence based on allele variant differences.
[0174] Clause 15. The computer-implemented method according to any one of Clauses 1 to 14, the computer-implemented method further comprising:
[0175] generating alignment scores for an initial alignment of the set of nucleotide reads with a corresponding genomic region of the reference genome based on comparing the set of nucleotide reads with a population haplotype from the population haplotype database;
[0176] generating an initial alignment data file including the initial alignment and the corresponding alignment scores;
[0177] generating personalized alignment scores for one or more personalized alignments based on comparing the set of nucleotide reads with a population haplotype subgroup within the personalized haplotype database; and
[0178] generating a personalized alignment data file including the one or more personalized alignments and the corresponding personalized alignment scores.
[0179] Clause 16. The computer-implemented method according to any one of Clauses 1 to 15, the computer-implemented method further comprising: determining the genotype detection of the genome sample based on the one or more personalized alignments.
[0180] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing technologies. Particularly suitable technologies are those in which nucleic acids are attached to fixed positions in the array such that their relative positions do not change and in which the array is repeatedly imaged. Embodiments that obtain images in different color channels (e.g., with different markers used to distinguish one nucleobase type from another) are particularly suitable. In some embodiments, the process for determining the nucleotide sequence of the target nucleic acid (i.e., the nucleic acid polymer) can be automated. Preferred embodiments include sequencing-by-synthesis (SBS) technology.
[0181] SBS technology typically involves the enzymatic extension of a nascent nucleic acid chain by repeatedly adding nucleotides to the template strand. In conventional SBS methods, a single nucleotide monomer can be provided to the target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to the target nucleic acid in the presence of a polymerase in delivery.
[0182] SBS can utilize nucleotide monomers with a terminator motif or nucleotide monomers lacking any terminator motif. Methods utilizing nucleotide monomers lacking terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in further detail below. In methods using nucleotide monomers lacking terminators, the number of nucleotides added in each cycle is typically variable, and this number depends on the template sequence and the method of nucleotide delivery.For SBS technology utilizing nucleotide monomers with a terminator motif, the terminator may be effectively irreversible under the sequencing conditions used, as in conventional Sanger sequencing using dideoxynucleotides, or the terminator may be reversible, as in sequencing methods developed by Solexa (now Illumina, Inc.).
[0183] SBS technology can utilize nucleotide monomers with or without a labeled motif. Therefore, incorporation events can be detected based on: the characteristics of the label, such as the fluorescence of the label; the characteristics of the nucleotide monomer, such as molecular weight or charge; byproducts of the incorporated nucleotide, such as the release of pyrophosphate; and so on. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from each other, or alternatively, the two or more different labels may be indistinguishable under the detection technology used. For example, the different nucleotides present in the sequencing reagent may have different labels, and they may be distinguishable using appropriate optics, as exemplified by sequencing methods developed by Solexa (now Illumina, Inc.).
[0184] Preferred embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into the nascent DNA strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996), “Real-time DNA sequencing using detection of pyrophosphate release.”, Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001), “Pyrosequencing sheds light on DNA sequencing.”, Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998), “A sequencing method based on realtime pyrophosphate.”, Science 281(5375). The publications of U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320, the full text of which are incorporated herein by reference, are cited. In pyrosequencing, the released PPi can be detected by being immediately converted to ATP by adenosine triphosphate (ATP) sulfatedase, and the level of ATP produced can be detected by photons generated by luciferase.The nucleic acid to be sequenced can be attached to a feature region in the array, and the array can be imaged to capture the chemiluminescent signal generated by the incorporation of nucleotides at the feature regions. Images can be obtained after processing the array with a specific nucleotide type (e.g., A, T, C, or G). Images obtained after adding each nucleotide type will differ in which feature regions in the array are detected. These differences in the images reflect the different sequence contents of the feature regions on the array. However, the relative position of each feature region will remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein. For example, images obtained after processing the array with each different nucleotide type can be processed in the same manner as illustrated herein for images obtained from different detection channels used in reversible terminator-based sequencing methods.
[0185] In another exemplary type of SBS, cyclic sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing, for example, cleavable or photobleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This method was commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, the disclosures of each of which are incorporated herein by reference. The availability of fluorescently labeled terminators (where the termination can be reversible and the fluorescent label can be cleaved) facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate these modified nucleotides and extend from them.
[0186] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection markers can be removable, for example, by cleavage or degradation. Images can be captured after the markers are incorporated into the arrayed nucleic acid signatures. In a specific embodiment, each cycle involves the simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally distinct marker. Four images are then obtained, each using a detection channel selective for one of the four different markers. Alternatively, different nucleotide types can be added sequentially, and images of the array can be obtained between each addition step. In such embodiments, each image will show a nucleic acid signature with a specific type of nucleotide incorporated. Different signatures may or may not be present in different images due to the different sequence contents of each signature. However, the relative positioning of the signatures will remain unchanged in the images. Images obtained by such a reversible terminator-SBS method can be stored, processed, and analyzed as described herein.After the image capture step, the markers can be removed, and the reversible terminator portion can be removed for subsequent cycles of nucleotide addition and detection. Removing these markers after they have been detected in a particular cycle and before subsequent cycles provides the advantage of reducing background signal and crosstalk between cycles. Examples of available marker and removal methods are described below.
[0187] In a particular embodiment, some or all of the nucleotide monomers may include a reversible terminator. In such embodiments, the reversible terminator / cleavable fluorophore may include a fluorophore linked to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other methods have separated the terminator chemistry from the cleavage of the fluorescent marker (Ruparel et al., Proc Natl Acad Sci USA 102: 5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. described the development of reversible terminators that use small 3'-allyl groups to block elongation, but can be easily deblocked by short-time treatment with a palladium catalyst. The fluorophore is attached to the base via a photocleavable linker, which can be easily cleaved by exposure to long-wavelength ultraviolet light for 30 seconds. Therefore, disulfide reduction or photocleavage can be used as the cleavable linker. Another method of reversible termination is the use of natural termination, which occurs after the placement of a bulk dye on a dNTP. The presence of a charged bulk dye on the dNTP can act as an efficient terminator due to steric and / or electrostatic hindrance. The presence of an incorporation event prevents further incorporation unless the dye is removed. The cleavage of the dye removes the fluorophore and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entirety.
[0188] Additional exemplary SBS systems and methods that may be used in conjunction with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, specification 28 / 35 pages 33 CN 121359209 A, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entirety.
[0189] Some embodiments may use fewer than four different markers to perform detection of four different nucleotides. For example, the methods and systems described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232 may be used to perform SBS. As a first example, a pair of nucleotide types may be detected at the same wavelength, but may be distinguished based on the intensity difference of one member relative to the other member, or based on a change in the presence or absence of a signal in one member of the pair that results in a signal that is significantly present or absent compared to the signal of the other member of that pair (e.g., by chemical modification, photochemical modification, or physical modification). As a second example, three of the four different nucleotide types may be detectable under specific conditions, while the fourth nucleotide type lacks a marker that is detectable under those conditions or is minimally detectable under those conditions (e.g., minimal detection due to background fluorescence, etc.). The incorporation of the first three nucleotide types into the nucleic acid may be determined based on the presence of their respective signals, and the incorporation of the fourth nucleotide type into the nucleic acid may be determined based on the absence of any signal or minimal detection of any signal. As a third example, one nucleotide type may include a marker detected in two different channels, while other nucleotide types may be detected in no more than one channel. These three exemplary configurations are not considered mutually exclusive and can be used in various combinations.An exemplary implementation combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type detected in a first channel (e.g., dATP with a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP with a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP with at least one label detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type lacking a label or minimally detected in either channel (e.g., unlabeled dGTP).
[0190] Additionally, as described in the material of incorporated U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, the first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0191] Some embodiments may utilize ligation-based sequencing technology. Such technologies utilize DNA ligases to incorporate oligonucleotides and recognize the incorporation of such oligonucleotides. Oligonucleotides typically have different labels associated with the identity of specific nucleotides in the sequence in which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid features with labeled sequencing reagents. Each image will show a nucleic acid feature incorporating a specific type of label. Different features may be present or absent in different images due to the different sequence content of each feature, but the relative positioning of the features will remain unchanged in the images. Images obtained by ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be used with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.
[0192] Some embodiments may utilize nanopore sequencing (Deamer, DW and Akeson, M., “Nanopores and nucleic acids: prospects for ultrarapid sequencing.”, Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, “Characterization of nucleic acids by nanopore analysis.”, Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin and J. A. Golovchenko, “DNA molecules and configurations in a solid-state nanopore microscope”, Nat. Mater., 2:611-615 (2003), the full text of which is incorporated herein by reference). In such embodiments, the target nucleic acid passes through a nanopore. Nanopores can be synthetic pores or biomembrane proteins, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the pore's conductivity.(The full text of these publications is incorporated herein by reference: U.S. Patent Nos. 7,001,792; Soni, GV, and Meller, “A. Progress toward ultrafast DNA sequencing using solid-state nanopores.”, Clin. Chem. 53, 1996–2001 (2007); Healy, K., “Nanopore-based single-molecule DNA analysis.”, Nanomed. 2, 459–481 (2007); Cockroft, SL, Chu, J., Amorin, M., and Ghadiri, M. R., “A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution.”, J. Am. Chem. Soc. 130, 818–820 (2008).) Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. Specifically, data can be processed like images, based on the exemplary processing of optical images and other images described herein.
[0193] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected by fluorescence resonance energy transfer (FRET) interaction between a polymerase carrying a fluorophore and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414 (each of which is incorporated herein by reference), or by using a zero-mode waveguide, as described, for example, in U.S. Patent No. 7,315,019 (each of which is incorporated herein by reference), and by using fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorporated herein by reference).Illumination can be limited to a volume of approximately 1 / 2 oz around the surface-tethered polymerase, allowing observation of the incorporation of fluorescently labeled nucleotides against a low background (Levene, M. J. et al., “Zero-mode waveguides for single-molecule analysis at high concentrations.”, Science 299, 682–686 (2003); Lundquist, PM et al., “Parallel confocal detection of single molecules in real time.”, Opt. Lett. 33, 1026–1028 (2008); Korlach, J. et al., “Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures.”, Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA 105, 1176–1181 (2008), the full text of which is incorporated herein by reference). Images obtained by such methods can be stored, processed, and analyzed as described herein.
[0194] Some SBS implementations include the detection of protons released during nucleotide incorporation into the extension product. For example, sequencing based on the detection of released protons can utilize electrical detectors and related technologies commercially available from Ion Torrent (a subsidiary of Guilford, CT, Life Technologies) or the sequencing methods and systems described in US 2009 / 0026082 A1; US 2009 / 0127589 A1; US 2010 / 0137143 A1; or US 2010 / 0282617 A1, each of which is incorporated herein by reference. The methods described herein for amplifying target nucleic acids using kinetic exclusion can be readily applied to substrates for proton detection. More specifically, the methods described herein can be used to generate a population of amplicon clones for proton detection.
[0195] The SBS method described above can be advantageously performed in a variety of formats, enabling the simultaneous manipulation of multiple different target nucleic acids. In specific embodiments, different target nucleic acids can be processed in a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a variety of ways.In the implementation scheme of the target nucleic acid used in surface-binding target nucleic acid (pages 30 / 35, CN 121359209 A), the target nucleic acid may be in an array format. In an array format, the target nucleic acid can typically bind to the surface in a spatially distinguishable manner. The target nucleic acid can bind by direct covalent attachment, attachment to beads or other particles, or binding to polymerases or other molecules attached to the surface. The array may include a single copy of the target nucleic acid at each site (also referred to as a feature region), or multiple copies having the same sequence may be present at each site or feature region. Multiple copies may be generated by amplification methods such as, for example, bridging amplification or emulsion PCR, which are described in further detail below.
[0196] The methods described herein may use arrays having feature portions at any of a variety of densities, including, for example, at least about 10 feature portions / cm², 100 feature portions / cm², 500 feature portions / cm², 1,000 feature portions / cm², 5,000 feature portions / cm², 10,000 feature portions / cm², 50,000 feature portions / cm², 100,000 feature portions / cm², 1,000,000 feature portions / cm², 5,000,000 feature portions / cm², or higher.
[0197] An advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Therefore, this disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified above. Therefore, the integrated system of this disclosure may include fluid components capable of delivering amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc. A flow cell in the integrated system may be configured for and / or for detecting target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Serial No. 13 / 273,666, each of which is incorporated herein by reference. As illustrated with respect to a flow cell, one or more fluid components of the integrated system may be used for amplification and detection methods. Taking a nucleic acid sequencing implementation as an example, one or more fluid components of the integrated system may be used for the amplification methods set forth herein and for delivering sequencing reagents in sequencing methods (such as those illustrated above). Alternatively, the integrated system may include separate fluid systems for performing amplification and detection methods. Examples of integrated sequencing systems capable of generating amplified nucleic acids and determining nucleic acid sequences include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Serial No. 13 / 273,666, which is incorporated herein by reference.
[0198] The sequencing systems described above sequence nucleic acid polymers present in samples received by the sequencing device.As defined herein, “sample” and its derivatives are used in their broadest sense to include any sample, culture, etc., suspected of containing a target. In some embodiments, a sample includes nucleic acids in the form of DNA, RNA, PNA, LNA, chimeric or hybrid forms. A sample may include any biological, clinical, surgical, agricultural, atmospheric or aquatic plant or animal sample containing one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh frozen or formalin-fixed paraffin-embedded nucleic acid sample. It is also contemplated that a sample may be derived from: a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched to) a single individual (such as tumor samples and normal tissue samples), or a sample from a single source containing two different forms of genetic material (such as maternal DNA and fetal DNA obtained from a maternal subject), or a sample containing contaminating bacterial DNA in the presence of plant or animal DNA. In some embodiments, the source of nucleic acid material may include nucleic acids obtained from newborns, such as nucleic acids commonly used for newborn screening.
[0199] Nucleic acid samples may include high molecular weight substances, such as genomic DNA (gDNA). The sample may include low molecular weight substances, such as nucleic acid molecules obtained from FFPE samples or archived DNA samples. In another embodiment, the low molecular weight substance includes enzymatically fragmented or mechanically fragmented DNA. The sample may include cell-free circulating DNA. In some embodiments, the sample may include nucleic acid molecules obtained from biopsy tissue, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microscopy, surgical resection, and other clinical or laboratory samples. In some embodiments, the sample may be an epidemiological sample, an agricultural sample, a forensic sample, or a pathogenic sample. In some embodiments, the sample may include nucleic acid molecules obtained from animals (such as human or mammalian sources). In another embodiment, the sample may include nucleic acid molecules obtained from non-mammalian sources (such as plants, bacteria, viruses, or fungi). In some embodiments, the source of the nucleic acid molecules may be archived or extinct samples or species.
[0200] Additionally, the methods and compositions disclosed herein can be used to amplify nucleic acid samples containing low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or may include forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel.Nucleic acid samples can be purified samples or lysed products containing crude DNA, such as those derived from oral swabs, paper, fabric, or other substrates that can be impregnated with saliva, blood, or other bodily fluids. Therefore, in some embodiments, the nucleic acid sample may include a small amount of DNA (such as genomic DNA) or fragmented portions of DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including but not limited to blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from a victim's hair, skin, tissue samples, autopsy, or remains. In some embodiments, nucleic acids containing one or more target sequences may be obtained from a deceased animal or human. In some embodiments, the target sequence may include nucleic acids obtained from non-human DNA (such as microbial, plant, or insect DNA). In some embodiments, the target sequence or amplified target sequence is directed for human identification purposes. In some embodiments, this disclosure relates in its entirety to methods for identifying the characteristics of forensic samples. In some embodiments, this disclosure relates in its entirety to methods for human identification using one or more target-specific primers disclosed herein or designed with one or more target-specific primers designed using the primer design standards outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence may be amplified using any one or more target-specific primers disclosed herein or using primer standards outlined herein.
[0201] Components of the personalized sequencing system 106 may include software, hardware, or both. For example, components of the personalized sequencing system 106 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., client device 114, local device 108, or server device 110). When executed by one or more processors, the computer-executable instructions of the personalized sequencing system 106 may cause the computing device to perform the bubble detection method described herein. Alternatively, components of the personalized sequencing system 106 may include hardware, such as a dedicated processing device for performing a function or group of functions. Additionally or alternatively, components of the personalized sequencing system 106 may include a combination of computer-executable instructions and hardware.
[0202] Furthermore, components of the personalized sequencing system 106 that perform the functions described herein with respect to the personalized sequencing system 106 may, for example, be implemented as part of a standalone application, as a module of an application, as a plugin of an application, as one or more library functions that can be called by other applications, and / or as a cloud computing model. Thus, components of the personalized sequencing system 106 may be implemented as part of a standalone application on a personal computing device or mobile device.Additionally or alternatively, components of the personalized sequencing system 106 may be implemented in any application providing sequencing services, including but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. “Illumina,” “BaseSpace,” “DRAGEN,” and “TruSight” are registered trademarks or trademarks of Illumina, Inc., in the U.S. and / or other countries.
[0203] As discussed in more detail below, embodiments of this disclosure may include or utilize a dedicated or general-purpose computer including computer hardware such as, for example, one or more processors and system memory. Embodiments within the scope of this disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Specifically, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein, page 32 / 35, CN 121359209 A Content Access Device). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more processes described herein.
[0204] A computer-readable medium can be any available medium accessible by a general-purpose or special-purpose computer system. A computer-readable medium storing computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium carrying computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of this disclosure may include at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0205] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSDs) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of computer-executable instructions or data structures and that is accessible by a general-purpose or special-purpose computer.
[0206] A “network” is defined as one or more data links that enable the transmission of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred to or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer appropriately regards that connection as a transmission medium.The transmission medium may include a network and / or a data link that can be used to carry desired program code means in the form of computer-executable instructions or data structures, and which can be accessed by a general-purpose or special-purpose computer. The combination described above should also be included within the scope of computer-readable media.
[0207] Furthermore, upon arrival at various computer system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link may be buffered in RAM within a network interface module (e.g., a NIC) and then ultimately transferred to computer system RAM and / or to a less transient computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) may be included in computer system components that also (or even primarily) utilize the transmission medium.
[0208] Computer-executable instructions include, for example, a set of instructions and data that, when executed at a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform certain functions or capabilities. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special-purpose computer that implements the elements of this disclosure. The computer-executable instructions may be, for example, binary numbers, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or actions. Rather, the described features and actions are disclosed as exemplary forms of implementing the claims.
[0209] Those skilled in the art will understand that this disclosure can be practiced in networked computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and so on. This disclosure can also be practiced in distributed system environments where both local and remote computer systems perform tasks via network links (via hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may reside on both local and remote storage devices.
[0210] Embodiments of this disclosure can also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources, as described on pages 33 / 35 of CN 121359209 A.For example, cloud computing can be adopted in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled as a result.
[0211] Cloud computing models can consist of various features such as, for example, on-demand self-service, widespread network access, resource pooling, rapid elasticity, metered services, etc. Cloud computing models can also exhibit various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). Cloud computing models can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and in the claims, "cloud computing environment" is an environment in which cloud computing is employed.
[0212] Figure 11 illustrates a block diagram of computing device 1100 that can be configured to perform one or more of the processes described above. It will be understood that one or more computing devices such as computing device 1100 can implement personalized sequencing system 106 and sequencing device system 104. As shown in FIG11, computing device 1100 may include processor 1102, memory 1104, storage device 1106, I / O interface 1108, and communication interface 1110, which may be communicatively coupled via communication infrastructure 1112. In some embodiments, computing device 1100 may include fewer or more components than those shown in FIG11. The following paragraphs describe the components of computing device 1100 shown in FIG11 in more detail.
[0213] In one or more embodiments, processor 1102 includes hardware for executing instructions such as those constituting a computer program. By way of example and not limitation, in order to execute instructions for dynamically modifying workflows, processor 1102 may retrieve (or obtain) instructions from internal registers, internal cache, memory 1104, or storage device 1106, and decode and execute them. Memory 1104 may be volatile or non-volatile memory for storing data, metadata, and programs executed by the processor. Storage device 1106 includes a storage means for storing data or instructions for performing the methods described herein, such as a hard disk, flash drive, or other digital storage device.
[0214] I / O interface 1108 allows a user to provide input to computing device 1100, receive output from computing device, and otherwise transfer data to and receive data from computing device. I / O interface 1108 may include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or combinations of such I / O interfaces.I / O interface 1108 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O interface 1108 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be servicing a particular implementation.
[0215] Communication interface 1110 may include hardware, software, or both. In any case, communication interface 1110 may provide one or more interfaces for communication (such as, for example, packet-based communication) between computing device 1100 and one or more other computing devices or networks. By way of example and not limitation, communication interface 1110 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as Wi-Fi.
[0216] Additionally, communication interface 1110 may facilitate communication with various types of wired or wireless networks. Communication interface 1110 can also facilitate communication using various communication protocols. Communication infrastructure 1112 may also include hardware, software, or both that couple components of computing device 1100 to each other. For example, communication interface 1110 may use one or more networks and / or protocols to enable multiple computing devices connected through a particular infrastructure to communicate with each other to perform one or more aspects of the process described herein. For illustration, the sequencing process may allow multiple devices (e.g., client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0217] In the foregoing description, this disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of this disclosure have been described with reference to the details discussed herein, and various embodiments are illustrated in the accompanying drawings. The above description and figures are illustrative of this disclosure and should not be construed as limiting this disclosure. Numerous specific details have been described to provide a thorough understanding of various embodiments of this disclosure.
[0218] This disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The embodiments described herein should be considered exemplary in all respects and not limiting. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. Additionally, the steps / actions described herein may be repeated or performed in parallel with each other or in parallel with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description.All changes within the equivalent meaning and scope of the claims shall be included within its scope. Instruction manual page 35 / 35, page 40, CN 121359209 A, Figure 1; Instruction manual figure 1 / 13, page 41, CN 121359209 A, Figure 2; Instruction manual figure 2 / 13, page 42, CN 121359209 A, Figure 3A; Instruction manual figure 3 / 13, page 43, CN 121359209 A, Figure 3B; Instruction manual figure 4 / 13, page 44, CN 121359209 A, Figure 4; Instruction manual figure 5 / 13, page 45, CN 121359209 A, Figure 5; Instruction manual figure 6 / 13, page 46, CN 121359209 A, Figure 6A; Instruction manual figure 7 / 13, page 47, CN 121359209 A, Figure 6B; Instruction manual figure 8 / 13, page 48, CN 121359209 A, Figure 7; Instruction manual figure 9 / 13, page 49, CN 121359209 A, Figure 8. Figure 9 of the instruction manual, page 10 / 13, CN 121359209 A; Figure 10 of the instruction manual, page 11 / 13, CN 121359209 A; Figure 11 of the instruction manual, page 12 / 13, CN 121359209 A; Figure 12 of the instruction manual, page 13 / 13, CN 121359209 A.
Claims
1. A system comprising: at least one processor; and a non-transitory computer-readable medium storing instructions that, when executed by the at least one processor, cause the system to: identify, within a set of reference spans of a reference genome, a set of nucleotide reads from a genomic sample and candidate population haplotypes from a population haplotype database; generate, for the set of reference spans of the reference genome, haplotype group scores for a group of the candidate population haplotypes based on comparing the set of nucleotide reads and the candidate population haplotypes; generate, for the genomic sample and based on the haplotype group scores, a personalized haplotype database comprising a subset of population haplotypes from the population haplotype database within the set of reference spans; and determine, with the personalized haplotype database, one or more personalized alignments of the set of nucleotide reads to respective genomic regions of the reference genome.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: identify, within a given reference span of the set of reference spans, one or more distinct k-mers for each candidate population haplotype; and generate, for the given reference span, a respective haplotype group score of the respective candidate population haplotype of the group of the candidate population haplotypes within the given reference span based on comparing k-mers of the set of nucleotide reads to the one or more distinct k-mers of the respective candidate population haplotype.
3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: determine initial alignments of the set of nucleotide reads to respective genomic regions of the reference genome; identify a subset of nucleotide reads of the set of nucleotide reads that align to respective reference spans of the set of reference spans according to the initial alignments; and generate the haplotype group scores for the set of reference spans based on comparing the subset of nucleotide reads and the candidate population haplotypes within the respective reference spans of the set of reference spans.
4. The system of claim 3, further comprising instructions that, when executed by the at least one processor, cause the system to: identify, as indicated within the population haplotype database, allelic variant differences between population haplotypes and primary contiguous sequences at respective genomic regions of the reference genome; and determine one or more of the initial alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary contiguous sequences according to the allelic variant differences.
5. The system of claim 3, further comprising instructions that, when executed by the at least one processor, cause the system to: and identifying, for each nucleotide read in the set of nucleotide reads, one or more population haplotypes based on comparing the set of nucleotide reads to haplotype variants within the respective genomic region of the initial alignment; generating an alignment data file comprising: the initial alignment of the set of nucleotide reads; an alignment score corresponding to the initial alignment; and the one or more population haplotypes identified for each nucleotide read in the set of nucleotide reads.
6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: determine haplotype block likelihoods for the set of reference spans by comparing pairs of the set of nucleotide reads and the candidate population haplotypes within respective reference spans in the set of reference spans; and generate the haplotype block scores for the set of reference spans based on the haplotype block likelihoods.
7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: identify at least one reference span in the set of reference spans corresponding to a haploid region of the reference genome; and generate haplotype block scores for individual candidate population haplotypes for the at least one reference span based on comparing the set of nucleotide reads and the candidate population haplotypes within the at least one reference span.
8. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to select, based on the haplotype block scores, a respective subset of the candidate population haplotypes for inclusion within the subset of population haplotypes of the personalized haplotype database for each reference span in the set of reference spans.
9. The system of claim 8, further comprising instructions that, when executed by the at least one processor, cause the system to: generate haplotype block likelihoods for the set of reference spans by comparing the set of nucleotide reads and the candidate population haplotypes; generate, using a Hidden Markov Model (HMM) algorithm, a set of haplotype block posterior probabilities for each reference span in the set of reference spans based on respective haplotype block likelihoods of adjacent reference spans in the set of reference spans; and generate respective sets of haplotype block scores for the set of reference spans based on the set of haplotype block posterior probabilities for each reference span in the set of reference spans.
10. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to: generate a set of forward probabilities for each reference span in the set of reference spans using the haplotype block likelihoods in a forward pass of the HMM algorithm across the set of reference spans; update the set of forward probabilities for each reference span in the set of reference spans using a backward pass of the HMM algorithm to generate a set of updated probabilities for each reference span in the set of reference spans; and generate, for each reference span in the set of reference spans, the set of haplotype group posteriors based on the set of update probabilities for each reference span in the set of reference spans.
11. The system of claim 9, further comprising instructions that, when executed by the at least one processor, cause the system to utilize recombination parameters of the HMM algorithm to reduce variation in haplotype group posteriors across the set of reference spans.
12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: classify each nucleotide read in the set of nucleotide reads within each respective reference span in the set of reference spans as inherited from a first parent or a second parent utilizing a variational Bayesian model; generate, for each classified nucleotide read in the set of nucleotide reads, individual haplotype scores for population haplotypes selected from the population haplotype database; and generate the haplotype group score for pairs of population haplotypes selected from the population haplotype database by combining individual haplotype scores from nucleotide reads inherited from the first parent with individual haplotype scores from nucleotide reads inherited from the second parent.
13. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate, for a given reference span in the set of reference spans, the haplotype group score for respective pairs of locally distinct population haplotypes within the given reference span, each locally distinct population haplotype comprising a unique set of one or more allelic variant differences relative to other population haplotypes within the given reference span.
14. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: identify allelic variant differences between population haplotypes in the subset of population haplotypes and a primary contiguous sequence at the respective genomic region of the reference genome as indicated within the personalized haplotype database; and determine one or more of the one or more personalized alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary contiguous sequence according to allelic variant differences.
15. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate alignment scores for initial alignments of the set of nucleotide reads to respective genomic regions of the reference genome based on comparing the set of nucleotide reads to population haplotypes from the population haplotype database; generate an initial alignment data file comprising the initial alignments and corresponding alignment scores; generate personalized alignment scores for the one or more personalized alignments based on comparing the set of nucleotide reads to the subset of population haplotypes within the personalized haplotype database; and generating a personalized alignment data file comprising the one or more personalized alignments and corresponding personalized alignment scores.
16. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine a genotype call for the genomic sample based on the one or more personalized alignments.
17. A computer-implemented method comprising: identifying, within a set of reference spans of a reference genome, a set of nucleotide reads from a genomic sample and candidate population haplotypes from a population haplotype database; generating, for the set of reference spans of the reference genome, haplotype set scores for groups of the candidate population haplotypes based on comparing the set of nucleotide reads and the candidate population haplotypes; generating, for the genomic sample and based on the haplotype set scores, a personalized haplotype database comprising a subset of population haplotypes from the population haplotype database within the set of reference spans; and determining, with the personalized haplotype database, one or more personalized alignments of the set of nucleotide reads to respective genomic regions of the reference genome.
18. The computer-implemented method of claim 17, further comprising: identifying, within a given reference span of the set of reference spans, one or more distinct k-mers for each candidate population haplotype; and generating, for the given reference span and based on comparing k-mers of the set of nucleotide reads to the one or more distinct k-mers of respective candidate population haplotypes of the group of candidate population haplotypes, respective ones of the haplotype set scores for the respective candidate population haplotypes within the given reference span.
19. The computer-implemented method of claim 17, further comprising: determining initial alignments of the set of nucleotide reads to respective genomic regions of the reference genome; identifying a subset of nucleotide reads of the set of nucleotide reads that align to respective ones of the set of reference spans according to the initial alignments; and generating the haplotype set scores for the set of reference spans based on comparing the subset of nucleotide reads and the candidate population haplotypes within the respective ones of the set of reference spans.
20. The computer-implemented method of claim 19, further comprising: identifying, as indicated within the population haplotype database, allelic variant differences between population haplotypes and primary contiguous sequences at respective genomic regions of the reference genome; and determining one or more of the initial alignments of the set of nucleotide reads by re-scoring candidate reference alignments of the set of nucleotide reads to the primary contiguous sequences according to the allelic variant differences.