Accelerator for Genotype Attribution Model
The accelerated genotype attribution system addresses inefficiencies in existing sequencing systems by employing integrated computing and a customized architecture to perform single-pass operations and utilize hot start points, achieving rapid and efficient determination of allele likelihoods, reducing processing time and memory consumption.
Patent Information
- Application Number
- JP2024576789
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-27
- Filing Date
- 2023-06-27
- Publication Date
- 2025-07-23
AI Technical Summary
Existing sequencing systems face inefficiencies in determining nucleotide base calls, particularly in low-read coverage genomic regions, due to high computational load, memory consumption, and processor latency when using Hidden Markov Models (HMMs) for genotype imputation, leading to prolonged processing times and bandwidth bottlenecks.
An accelerated genotype attribution system utilizing integrated computing and a customized architecture performs single-pass simultaneous multiplication operations, determines intermediate allele likelihoods using a subset as a hot start point, and calculates running totals to reduce latency, thereby optimizing memory usage and processing time.
The system significantly reduces processing time from approximately 10 hours to 60 seconds per 40,000 HMM calculation tasks, enhances memory efficiency, and improves throughput by 600 times compared to existing systems.
Smart Images

Figure 2025523560000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit and priority of U.S. Provisional Application No. 63 / 367,105, filed on June 27, 2022, entitled "ACCELERATORS FOR A GENOTYPE IMPUTATION MODEL". The above - mentioned application is hereby incorporated by reference in its entirety.
Background Art
[0002] In recent years, biotechnology companies and research institutions have improved the hardware and software for sequencing nucleotides and determining nucleic acid base calls for genomic samples. For example, some existing sequencing machines and sequencing data analysis software (collectively referred to as "existing sequencing systems") determine individual nucleic acid bases within a sequence by using conventional Sanger sequencing or sequencing - by - synthesis (SBS) methods. When using SBS, existing sequencing systems can monitor thousands of oligonucleotides being synthesized in parallel from a template and predict nucleic acid base calls for increasing nucleotide reads. For example, cameras in many existing sequencing systems capture images of irradiated fluorescent tags incorporated into oligonucleotides. After capturing such images, some existing sequencing systems process the image data from the camera and determine nucleic acid base calls for the nucleotide reads corresponding to the oligonucleotides. Based on the comparison of the nucleic acid base calls for such reads with a reference genome, existing systems utilize variant callers to identify variants in the genomic sample, such as single - nucleotide polymorphisms (SNPs), insertions or deletions (indels), or other variants within the genomic sample.
[0003] Despite these recent advancements, existing sequencing systems may inaccurately call bases or collect an insufficient number of (or seemingly conflicting) nucleotide reads, particularly for nucleotide bases in low-read coverage genomic regions. For a particular genomic region of a genomic sample, existing sequencing systems frequently use genotype imputation models to assign nucleic acid base calls and phase haplotypes based on the detected variants in the genomic sample. For example, existing sequencing systems frequently use various types of hidden Markov models (HMMs) customized for genotype imputation to assign nucleic acid base calls for a particular genomic region, such as by using the Genotype Likelihood Imputation and PhaSing mEthod (GLIMPSE) or IMPUTE. Such HMMs have improved the accuracy of genotype imputation, but existing sequencing systems that use genotype imputation models frequently consume significant computer processing and require significant memory to store the data generated by the genotype imputation models, and execute the genotype imputation models with inefficient wait times for processor downtime.
[0004] As just suggested, existing sequencing systems consume excessive computer processing and time when executing HMMs for genotype imputation. For example, some existing sequencing systems that execute a single thread on a central processing unit (CPU) consume an average of about 17.5 hours to phase and impute haplotype allele likelihoods for a single marker allele corresponding to a genomic region. Approximately 80% of such phasing and imputation calculations are derived from both HMM calculations and the Burrows-Wheeler Transform (BWT), with HMM calculations consuming about 60% of the calculation time and BWT calculations consuming about 20% of the calculation time. BWT calculations can be significantly amortized and reduced as a percentage over large batches of genomic samples, but the HMM calculation time on a single CPU thread can still consume approximately 10 hours or more (e.g., 600 - 640 minutes).
[0005] In addition to consuming a significant amount of time and computer processing, existing genotyping systems can consume a significant amount of memory when running an HMM for genotype attribution. For example, in some cases, existing genotyping systems determine and store values for haplotype allele likelihoods in a haplotype matrix corresponding to 50 million cells for a collection of marker variants and haplotypes from a haplotype reference panel. Given 50 million cells for a single haplotype matrix, an existing genotyping system that determines 40,000 haplotype calls based on 40,000 haplotype matrices over a period of time must determine values corresponding to 2 trillion cells. Some HMMs for genotype attribution, such as GLIMPSE, require that an existing genotyping system determine values once for the alpha path of the HMM and twice for the beta path for each haplotype matrix, so an existing genotyping system can determine and store values over many haplotype matrices of a total of approximately 6 trillion cells in order to calculate HMM-based genotype attribution for multiple genomic regions. The hardware on sequencing devices and servers has increased memory, but chips for field programmable gate arrays (FPGAs) or other configurable processors often contain about 32 or 64 gigabytes of memory on the chip, which is barely sufficient or insufficient memory to store data for a single haplotype matrix.
[0006] Beyond imposing a burden on processing time and memory, some existing genotyping systems inefficiently execute an HMM for genotype attribution due to processor latency. For example, in some cases, some existing genotyping systems determine the sum of all intermediate allele likelihoods for one marker variant from a haplotype reference panel and various haplotypes based on both alpha path values and beta path values, and even before determining the individual intermediate allele likelihoods for subsequent marker variants. Since all intermediate allele likelihoods are summed and the allele likelihoods must be determined for one marker variant before determining the individual intermediate allele likelihoods for another marker variant, the processor used by existing genotyping systems often waits for latency for one or both of the sum of adjacent marker intermediate allele likelihoods and the generation of allele likelihoods. In the case of HMM-based genotype attribution that may require a haplotype matrix of 50 million cells (and 40,000 individual haplotype matrices), such computational latency is inefficient and wastes time that the processor could otherwise use to further calculate allele likelihoods.
[0007] As the above memory sizes and calculation times suggest, existing genotyping systems require a significant throughput of input and output values to efficiently execute HMM-based genotype attribution. Since such attribution may require storing or transferring large amounts of data such as specific input values for haplotype matrices with millions, billions, or trillions of cells, existing HMM-based genotype attribution further burdens the bandwidth of high-speed buses such as Peripheral Component Interconnect Express (PCIe), or other interfaces that connect a processor card to other hardware within a computing device. Bottlenecks to PCIe throughput or other interface throughput can significantly slow down HMM-based genotype attribution.
[0008] These, along with additional problems and challenges, exist in existing genotyping systems.
Summary of the Invention
[0009] The present disclosure describes one or more embodiments of a system, method, and non-transitory computer-readable storage medium that solve one or more of the above problems or provide other advantages over the art. To facilitate computer processing or efficiently redistribute the memory load of a genotype attribution model, the disclosed system uses integrated computing, efficient data transfer, or a customized architecture to determine the allele likelihoods of genomic regions that exhibit specific haplotype alleles. For example, the disclosed system can determine the intermediate allele likelihoods of genomic regions containing haplotype alleles given a marker variant and a reference panel haplotype by performing a single-pass simultaneous multiplication operation on a processor. In some cases, the disclosed system determines and stores a subset of the intermediate allele likelihoods corresponding to a group of marker variants, and uses the subset of intermediate allele likelihoods as a hot start point to generate a full set of intermediate allele likelihoods for a set of marker variants. In a further embodiment, the disclosed system determines a running total of the intermediate allele likelihoods of genomic regions that exhibit haplotype alleles for one or more haplotypes given one marker variant, and uses the running total as a running input to determine the individual intermediate allele likelihoods of genomic regions that exhibit haplotype alleles given another marker variant without latency for summing the intermediate allele likelihoods and / or generating allele likelihoods.
[0010] Additional features and advantages of one or more embodiments of the present disclosure are described in the following description, some of which will be apparent from the description, or may be learned by practice of such exemplary embodiments.
Brief Description of the Drawings
[0011] The "Detailed Description of the Invention" refers to the drawings briefly described below.
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Mode for Carrying Out the Invention
[0012] The present disclosure describes one or more embodiments of an accelerated genotype attribution system that can determine intermediate allele likelihoods of genomic regions that exhibit specific haplotype alleles as part of a genotype attribution model by using efficient data transfer across integrated computing or dedicated hardware. For example, the accelerated genotype attribution system can determine intermediate allele likelihoods of genomic regions that contain haplotype alleles, given a particular marker variant and a haplotype from a haplotype reference panel, by performing a single-pass simultaneous multiplication operation on a processor rather than multiple-pass simultaneous multiplication operations. In some cases, the accelerated genotype attribution system can (i) determine and store a subset of intermediate allele likelihoods corresponding to a group of marker variants, and (ii) generate a set of intermediate allele likelihoods by using the subset of intermediate allele likelihoods as a hot start point for a full pass of intermediate allele likelihoods. The accelerated genotype attribution system can perform (i) and (ii) without storing multiple full sets of intermediate allele likelihoods on the processor chip during real-time processing. In certain further embodiments, the accelerated genotype attribution system determines a running total of intermediate allele likelihoods of genomic regions that exhibit haplotype alleles for one or more haplotypes given one marker variant, and uses the running total as a running input to determine intermediate allele likelihoods of genomic regions that exhibit haplotype alleles for another marker variant. By using such a running total, the accelerated genotype attribution system sums adjacent marker intermediate allele likelihoods and / or avoids processor idle time that generates allele likelihoods that slow down existing sequencing systems.
[0013] As suggested above, an accelerated genotype attribution system applies a genotype attribution model, such as a hidden Markov model (HMM)-based model, to nucleotide reads from a genomic region of a genomic sample to determine posterior genotype likelihoods and haplotype calls for the genomic region. By way of example, in some embodiments, the accelerated genotype attribution system determines prior genotype likelihoods that the genomic region contains a particular genotype (e.g., a reference allele or an alternative allele), where the genomic region corresponds to a variable position or coordinate of a haplotype reference panel. Such prior genotype likelihoods are based on nucleotide reads from the genomic sample and quality scores of the nucleotide reads. The accelerated genotype attribution system further deconvolves a vector of prior genotype likelihoods into two independent vectors of haplotype allele likelihoods (or simply, haplotype likelihoods), where each vector corresponds to one of two complementary haplotypes. Based on the haplotype likelihoods from the independent vectors, the accelerated genotype attribution system uses a haploid version of the HMM to attribute two target haplotypes as haplotype calls. The accelerated genotype attribution system further determines (and updates) the phase of the two attributed haplotypes. In some embodiments, for example, the accelerated genotype attribution system utilizes the Genotype Likelihood Imputation and PhaSing mEthod (GLIMPSE) as the genotype attribution model, which is described by Simone Rubinacci et al., "Efficient Phasing and Imputation of Low-coverage Sequencing Data Using Large Reference Panels," 53 Nature Genetics 120-126 (2021) (hereinafter, Rubinacci), which is hereby incorporated by reference in its entirety.
[0014] The disclosed accelerated genotype attribution system introduces and utilizes an integrated computing or proprietary architecture to efficiently determine the intermediate allele likelihoods of genomic regions that exhibit specific haplotype alleles as part of GLIMPSE or another genotype attribution model. The following paragraphs briefly introduce various embodiments of the accelerated genotype attribution system.
[0015] A. Single Pass - Simultaneous Operation As suggested above, the accelerated genotype attribution system determines the intermediate allele likelihoods of genomic regions containing haplotype alleles by performing a single - pass simultaneous multiplication operation for a given marker variant and haplotype. To perform such an operation, in some embodiments, the accelerated genotype attribution system identifies a haplotype reference panel for a genomic region of a genomic sample as part of the genotype attribution model. The accelerated genotype attribution system further accesses a first transition - recognition allele likelihood factor corresponding to the haplotype alleles from the haplotype reference panel (e.g., Q[m][allele] * P1[m]) and a second transition - recognition allele likelihood factor corresponding to the haplotype alleles (e.g., Q[m][allele] * P0[m]). By combining the first allele likelihood factor with the adjacent - marker intermediate allele likelihood of the genomic region containing the haplotype allele given the adjacent marker variant (e.g., A’[m - 1][k]), the accelerated genotype attribution system performs a single - pass simultaneous multiplication operation and generates an adjacent - marker transition - factor recognition allele likelihood for the marker variant and haplotype (e.g., Q[m][allele] * P1[m] * A’[m - 1]). Based on the adjacent - marker transition - factor recognition allele likelihood and the second transition - recognition allele likelihood factor, the accelerated genotype attribution system further determines the intermediate allele likelihood of the genomic region containing the haplotype allele for a given marker variant and haplotype.
[0016] By performing such single-pass simultaneous multiplication operations, the accelerated genotype attribution system speeds up the computer processing time for determining intermediate allele likelihoods and output allele likelihoods over the slower computer processing times of existing sequencing systems. As noted above, some existing sequencing systems that execute a single thread on a central processing unit (CPU) consume an average of about 17.5 hours to phase and attribute haplotype allele likelihoods for a single marker allele corresponding to a genomic region, and the HMM calculation time on a single CPU thread can consume approximately 10 hours. As further explained below, the several hours of computer processing time is due in part to three multiplication operations for determining the intermediate allele likelihoods for each given pair of marker variants and haplotypes, and 3,000 multiplication operations for each haplotype of a haplotype reference panel (e.g., organized in columns) in a sequencing system.
[0017] In contrast to existing sequencing systems, in some embodiments, the disclosed accelerated genotype attribution system performs a single-pass simultaneous multiplication operation to determine the intermediate allele likelihoods for each given pair of marker variants and haplotypes, and due to such integrated multiplication operations, performs approximately 1,000 multiplication operations for each haplotype of a haplotype reference panel (e.g., organized in columns). Along with other integrated operations or other embodiments, the accelerated genotype attribution system can reduce the computer processing time of a single processor thread for executing approximately 40,000 HMM calculation tasks from approximately 10 hours or more (e.g., 600 - 640 minutes) to approximately 60 seconds, thereby accelerating the processing time by 600 times.
[0018] B. Hot Start Intermediate Allele Likelihood Subset As further described above, in some cases, the accelerated genotype attribution system determines and stores a subset of the intermediate allele likelihoods corresponding to the marker variant groups, and generates a set of intermediate allele likelihoods immediately by using the subset of the intermediate allele likelihoods as a hot start point for determining the full path of the intermediate allele likelihoods. To determine and utilize such hot start likelihoods, in some embodiments, the accelerated genotype attribution system determines the first pass intermediate allele likelihoods of the genomic regions from a genomic sample that includes haplotype alleles corresponding to a given set of haplotypes for a set of marker variants. The accelerated genotype attribution system further stores, on a dynamic random access memory (DRAM) or other memory device, a subset of the first pass intermediate allele likelihoods corresponding to a subset of the marker variants for a group of marker variants. The accelerated genotype attribution system then uses the stored subset of the first pass intermediate allele likelihoods to initialize the allele likelihood determination in the group of marker variants, thereby reproducing the first pass intermediate allele likelihoods. The accelerated genotype attribution system also uses the stored subset of the first pass intermediate allele likelihoods to determine the second pass intermediate allele likelihoods of the genomic regions that include haplotype alleles corresponding to the set of marker variants and the set of haplotypes. Based on the reproduced first pass intermediate allele likelihoods and the second pass intermediate allele likelihoods, the accelerated genotype attribution system generates the allele likelihoods of the genomic regions that include haplotype alleles.
[0019] By determining and using such an intermediate allele likelihood subset as a hot start point, the accelerated genotype attribution system intelligently and efficiently redistributes data among memory devices, reduces data storage, and increases on-chip bandwidth. As noted above, some existing sequencing systems (e.g., GLIMPSE) that use HMMs for genotype attribution determine and store values in a haplotype matrix of approximately 50 million cells. Data for such a haplotype matrix is found to saturate or be too large to store in on-chip memory for a field programmable gate array (FPGA) or other processors of an existing sequencing system. To reduce and redistribute data for such a large haplotype matrix, in some embodiments, the accelerated genotype attribution system determines and stores an intermediate allele likelihood subset corresponding to a marker variant group and uses the intermediate allele likelihood subset as a hot start point for determining the full path of intermediate allele likelihoods.
[0020] By determining and storing an intermediate allele likelihood subset, the accelerated genotype attribution system can exponentially reduce and transfer data depending on the size of the marker variant group or window. In some embodiments, for example, the accelerated genotype attribution system reduces the memory load by 100-fold by determining and storing an intermediate allele likelihood subset corresponding to each marker variant from a 100-count marker variant group, or reduces the memory load by 1000-fold by determining and storing an intermediate allele likelihood subset corresponding to each marker variant from a 1000-count marker variant group. As further described below, in some embodiments, the size of the marker variant group controls the exponent by which the accelerated genotype attribution system reduces memory load and data transfer.
[0021] C. Running Sum of Adjacent Marker Intermediate Allele Likelihoods In addition to the path simultaneous multiplication operation and the hot start intermediate allele likelihood subset, the accelerated genotype attribution system determines a running sum of the intermediate allele likelihoods of genomic regions that indicate haplotype alleles for one or more haplotypes given one marker variant, and uses the running sum as a running input to determine the individual intermediate allele likelihoods of genomic regions that indicate haplotype alleles for a haplotype given another marker variant. To utilize such a running sum, in some embodiments, the accelerated genotype attribution system identifies a haplotype reference panel for genomic regions of a genomic sample as part of a genotype attribution model. The accelerated genotype attribution system further (i) determines a running sum of a first subset of the intermediate allele likelihoods of genomic regions that include haplotype alleles of a first type from one or more haplotypes of the haplotype reference panel for an adjacent marker variant, and (ii) determines a running sum of a second subset of the intermediate adjacent allele likelihoods of genomic regions that include haplotype alleles of a second type from one or more haplotypes for the adjacent marker variant. Based on the running sums, the accelerated genotype attribution system determines the sum of the intermediate allele likelihoods of genomic regions that include haplotype alleles from the haplotypes of the haplotype reference panel for the marker variant.
[0022] By determining and using such a running sum of the intermediate allele likelihoods, the accelerated genotype attribution system eliminates or reduces the latency for one or both of summing adjacent marker intermediate allele likelihoods and generating allele likelihoods. As described above, some existing sequencing systems determine the allele likelihood for one marker variant before summing the intermediate allele likelihoods and determining the individual intermediate allele likelihoods for another marker variant, thereby causing the processor of the existing sequencing system to wait for the latency for one or both of summing adjacent marker intermediate allele likelihoods and generating allele likelihoods for other marker variants.
[0023] In contrast to existing genotyping systems, in some embodiments, an accelerated genotype attribution system determines the running sum of intermediate allele likelihoods of genomic regions indicative of haplotype alleles for one or more haplotypes given one marker variant as a running input, without the conventional latency for determining the individual intermediate allele likelihoods of genomic regions indicative of haplotype alleles given another marker variant. Without such latency of summing adjacent marker intermediate allele likelihoods and generating allele likelihoods, the accelerated genotype attribution system facilitates determining haplotype allele likelihoods for a genotype attribution model faster than existing genotyping systems. Along with other integration operations or other embodiments, the accelerated genotype attribution system can reduce the computer processing time of a single processor thread for performing approximately 40,000 HMM calculation tasks from approximately 10 hours or more (e.g., 600 - 640 minutes) to about 60 seconds, thereby accelerating the processing time by 600 times.
[0024] D. Customized Hardware Architecture To facilitate one or more of integrated computing or data storage, in some embodiments, the accelerated genotype attribution system utilizes a customized architecture. For example, the accelerated genotype attribution system can store (and access from) a subset of intermediate allele likelihoods on dynamic random access memory (DRAM) or another memory device to hot-start the determination of the full pass of intermediate allele likelihoods. As a further example, the accelerated genotype attribution system uses a dataflow engine as part of a configurable processor to (i) queue and manage HMM calculation tasks for corresponding clusters of the accelerated computing engine, and (ii) distribute input values to individual accelerated computing engines from the cluster to determine intermediate allele likelihoods (or other HMM calculation tasks) for columns or matrices. In some cases, for example, the accelerated genotype attribution system transmits each set of input values (e.g., allele likelihood factors, transition coefficients, and haplotype-allele values) from the dataflow engine to each respective accelerated computing engine, and each respective accelerated computing engine is used to determine each set of intermediate allele likelihoods corresponding to each subset of marker variants and each subset of haplotypes based on each set of input values.
[0025] As suggested above, the disclosed accelerated genotype attribution system uses a customized architecture that facilitates a faster throughput of input and output values for allele likelihoods in a genotype attribution model than the conventional and undifferentiated architectures of existing sequencing systems. For example, instead of using off-chip DRAM or other memory devices to store intermediate allele likelihood values and thereby reducing throughput by relying on on-chip memory, the accelerated genotype attribution system stores and can quickly transfer a subset of intermediate allele likelihoods for hot start. As described above, existing HMM-based genotype attribution burdens the bandwidth of high-speed buses such as Peripheral Component Interconnect Express (PCIe), or other interfaces where values for a 50-million-cell haplotype matrix sometimes pass through 40,000 such matrices. By storing on on-chip DRAM (or other on-chip memory) haplotype-allele-index data for the haplotype matrix and accessing from there, in some embodiments, the accelerated genotype attribution system can utilize a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model with more than 4 gigabytes of PCIe bandwidth than existing sequencing systems to generate allele likelihoods.
[0026] In some embodiments, by using a data flow engine as part of a configurable processor and orchestrating the data flow to a cluster of acceleration computing engines, the accelerated genotype attribution system avoids latency and determines allele likelihoods for different haplotypes and marker variants in parallel. In fact, as further described below, the data flow engine of the disclosed accelerated genotype attribution system can efficiently distribute input and output values to different clusters of acceleration computing engines to determine allele likelihoods in about 60 seconds from the equivalent of 6 trillion cells across multiple haplotype matrices.
[0027] As shown by the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of an accelerated genotype attribution system. As used herein, the term "genomic sample" refers to a target genome or a portion of a genome that is to be sequenced. For example, a sample genome can include the sequence of nucleotides isolated or extracted from a sample organism (or a copy of such an isolated or extracted sequence). In particular, a sample genome can include a genome that is (in whole or in part) isolated or extracted from a sample organism and is composed of nitrogenous heterocyclic bases. For example, a nucleic acid polymer can include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acids or segments of chimeric or hybrid forms of nucleic acids described hereinafter. In some cases, a sample genome is one that is prepared or isolated by a kit and found in a sample received by a sequencing device.
[0028] Also as used herein, the term "haplotype" refers to a nucleotide sequence that exists in an organism (or in an organism from a population) and is inherited from one or more ancestors. In particular, a haplotype can include alleles (or other nucleotide sequences) that exist in organisms of a population and are inherited together by such organisms from a single parent. In one or more embodiments, a haplotype includes a set of SNPs on the same chromosome that tend to be inherited together. As described hereinafter, in some cases, a haplotype from a haplotype reference panel can be represented as "k", and a row of different haplotypes from a haplotype reference panel can be represented as "K". Further, an "attributed haplotype" refers to a haplotype that is presumed or statistically inferred to exist in a sample genome. For example, an attributed haplotype can be a haplotype that is statistically inferred for a genomic coordinate or region based on SNPs that surround or are adjacent to the genomic coordinate or region. As described above, an attributed haplotype can surround a target genomic region and can include SNPs or other variant-nucleotide-base calls by which a customized sequencing system attributes the haplotype.
[0029] In this context, the term "haplotype allele" refers to a version of a nucleic acid base or nucleotide sequence at a genomic coordinate or genomic region corresponding to a haplotype, such as a haplotype for a gene or a genomic region encoding a non-coding region. In particular, a haplotype allele includes one of two or more versions of a nucleic acid base or nucleotide sequence at a genomic coordinate or region that tend to be inherited together as part of a haplotype. As part of a haplotype, in some cases, combinations of haplotype alleles can be inherited by an organism as part of a single gene or across multiple genes. In some cases, the present disclosure describes different types of haplotype alleles. For example, in some embodiments, one type of haplotype allele can refer to a sample reference haplotype allele and another type of haplotype allele can refer to a sample alternative haplotype allele. The present disclosure may describe a first type and a second type of haplotype allele corresponding to a particular haplotype, but in some embodiments, a haplotype can include more than two types of haplotype alleles (e.g., a sample reference haplotype allele and multiple sample alternative haplotype alleles).
[0030] In some cases, a haplotype or a haplotype allele that is a component thereof is represented by a haplotype reference panel. As used herein, a "haplotype reference panel" refers to a digital collection or database of haplotypes from genomic samples in which the haplotypes of one or more ancestors or predecessors have been determined. In some cases, a haplotype reference panel includes a digital database of haplotypes from genomic samples that represent (or are common among) a population of organisms, for which multiple ancestral or predecessor haplotypes have been determined. In some cases, an accelerated genotype imputation system uses a haplotype reference panel developed by the Haplotype Reference Consortium (HRM), the 1000 Genomes Project, or Illumina, Inc.
[0031] In this context, the term "genotype imputation model" refers to an algorithm or model for imputing the genotype of a genomic region based on sequencing data from a genomic sample and the haplotypes corresponding to each genomic region. In particular, a genotype imputation model includes a hidden Markov model (HMM)-based algorithm or model for imputing the genotype of a genomic region and for phasing haplotypes based on sequencing data from a genomic sample and the haplotypes corresponding to each genomic region from a haplotype reference panel. As noted above, in some cases, a genotype imputation model includes GLIMPSE. Alternatively, a genotype imputation model includes fastPHASE, BEAGLE, MACH, or IMPUTE.
[0032] As part of genotype attribution, in some cases, an accelerated genotype attribution system determines allele likelihoods. As used herein, the term "allele likelihood" refers to the likelihood that a genomic region indicates or contains a haplotype allele corresponding to a haplotype. For example, in some embodiments, the allele likelihood includes a statistical likelihood that a genomic region of a genomic sample indicates or contains a sample reference haplotype allele or a sample alternative haplotype allele for a particular haplotype from the haplotypes of a haplotype reference panel. As described below, in some cases, the allele likelihood can be represented as (i) R0 for the likelihood that a genomic region of a genomic sample contains a sample reference haplotype allele of a particular haplotype, or (ii) R1 for the likelihood that a genomic region of a genomic sample contains a sample alternative haplotype allele of a particular haplotype. Thus, in some cases, the allele likelihood represents the posterior genotype likelihood generated by a genotype attribution model.
[0033] In connection therewith, the term "intermediate allele likelihood" refers to a value representing a provisional or preliminary likelihood that a genomic region indicates or contains a haplotype allele corresponding to a haplotype. For example, in some embodiments, the intermediate allele likelihood includes a value representing a provisional or preliminary likelihood that a genomic region of a genomic sample indicates or contains a sample reference haplotype allele or a sample alternative haplotype allele for a particular haplotype from the haplotypes of a haplotype reference panel given a target marker variant. As further described below, in some cases, the intermediate allele likelihood can be represented as A[m][k], called an alpha value, or alternatively, can be represented as B[m][k], called a beta value. The present disclosure primarily uses A[m][k] as an exemplary notation for the intermediate allele likelihood in the alpha path, although the notation B[m][k] can be used interchangeably for the intermediate allele likelihood in the beta path.
[0034] In this context, the term "marker variant" refers to a variant at a polymorphic site in a population. In particular, a marker variant includes one of two or more alleles present in a population at a polymorphic genomic coordinate or genomic region at a frequency higher than a threshold frequency, for example, higher than 1% of the population. In some cases, a marker variant includes an SNP present at a polymorphic genomic coordinate among human populations. Further, or alternatively, a marker variant may include an insertion or deletion (indel), a structural variant, or other variant at a polymorphic site in the population. As further described below, in some cases, a marker variant or a target marker variant is represented as m or [m]. In contrast, the term "adjacent marker variant" refers to a marker variant that is ordered before or after a target marker variant according to a particular order. In particular, an adjacent marker variant includes a marker variant represented by an adjacent column that is located one column before or one column after the target column representing the target marker variant in a matrix. As further explained below, in some cases, an adjacent marker variant is represented as m-1 or [m-1] or m+1 or [m+1].
[0035] In this context, as used herein, the term "adjacent marker intermediate allele likelihood" refers to the intermediate allele likelihood for a marker variant adjacent to a target marker variant. In particular, an adjacent marker intermediate allele likelihood includes the intermediate allele likelihood for a marker variant represented by an adjacent column that is located one column before or one column after the target column representing the target marker variant in a matrix. As further explained below, in some cases, an adjacent marker intermediate allele likelihood is represented as A[m-1][k].
[0036] As further used herein, the term "allele likelihood factor" corresponds to a haplotype allele and refers to a factor or parameter applied to a transition coefficient and / or other parameter in a function. In particular, the allele likelihood factor refers to (i) either a sample reference haplotype allele or a sample alternative haplotype allele and a marker variant, and (ii) a factor or parameter applied to a transition linear coefficient, a transition constant coefficient, and / or other parameters in a function for determining allele likelihood. As further described below, in some cases, the allele likelihood factor is generally represented as Q[m][allele], the allele likelihood factor corresponding to the sample reference haplotype allele is represented as Q0, and the allele likelihood factor corresponding to the sample alternative haplotype allele is represented as Q1.
[0037] In connection therewith, the term "transition coefficient" refers to a coefficient or parameter representing the probability of transition or change between marker variants. In particular, the transition coefficient includes a coefficient or parameter representing the probability of transition between rows representing marker variants within a matrix. In some cases, the transition coefficient is classified into several types including a transition linear coefficient and a transition constant coefficient. Hereinafter, in some cases, the transition constant coefficient is denoted as P0 and the transition linear coefficient is denoted as P1.
[0038] In some cases, the accelerated genotype attribution system combines various factors or coefficients (e.g., multiplication, weighted sum). For example, as used herein, the term "transition recognition allele likelihood factor" refers to a value representing a combination of a transition coefficient and an allele likelihood factor. In particular, the transition recognition allele likelihood factor includes a value representing the product of a transition coefficient and an allele likelihood factor. As described below, in some cases, the transition recognition allele likelihood factor is generally Q[m][allele] * represented as P[m], the first transition recognition allele likelihood factor is Q[m][allele] * represented as P1[m], and the second transition recognition likelihood factor is Q[m][allele] * represented as P0[m].
[0039] As used further herein, the term "adjacent marker transition factor recognition allele likelihood" refers to a value representing a combination of an allele likelihood factor, a transition coefficient, and an intermediate allele likelihood for an adjacent marker variant. In particular, the adjacent marker transition factor recognition allele likelihood includes a value representing the product of an allele likelihood factor, a transition linear coefficient, and an intermediate allele likelihood for an adjacent marker variant. As described below, in some cases, the adjacent marker transition factor recognition allele likelihood is generally Q[m][allele] * P1[m] * represented as A’[m - 1].
[0040] As used further herein, the term "total adjacent marker transition recognition allele likelihood factor" refers to a value representing a combination of the sum of an allele likelihood factor, a transition coefficient, and an intermediate allele likelihood for an adjacent marker variant. In particular, the total adjacent marker transition recognition allele likelihood factor includes a value representing the product of an allele likelihood factor, a transition constant coefficient, and the sum of the intermediate allele likelihoods for an adjacent marker variant. As described below, in some cases, the total adjacent marker transition recognition allele likelihood factor is generally Q[m][allele] * P0[m] * represented as Sum’[m - 1].
[0041] As described above, in some embodiments, the accelerated genotype attribution system can immediately generate a set of intermediate allele likelihoods for multiple paths by using an intermediate allele likelihood subset as a hot start point for a full path of intermediate allele likelihoods. As used herein, the term "path" refers to a series of operations for determining intermediate allele likelihoods corresponding to haplotypes from a haplotype reference panel according to a particular direction. In particular, a path includes a series of operations in a direction across a haplotype matrix for determining intermediate allele likelihoods corresponding to different combinations of marker variants and haplotypes from a haplotype reference panel. For example, a path may proceed in a forward or reverse direction across the haplotype matrix. In some cases, a path including a series of operations from left to right of the haplotype matrix constitutes an alpha path, and a path including a series of operations from right to left of the haplotype matrix constitutes a beta path.
[0042] In this regard, the phrase "path intermediate allele likelihoods" refers to a set of intermediate allele likelihoods corresponding to a path. In particular, the first set of path intermediate allele likelihoods includes the set of intermediate allele likelihoods determined by performing a first path of operations in a first direction. In contrast, the second set of path intermediate allele likelihoods includes the set of intermediate allele likelihoods determined by performing a second path of operations in a second direction. For example, the first set of path intermediate allele likelihoods may be determined when the accelerated genotype attribution system performs a first path in a reverse direction across the haplotype matrix, the second set of path intermediate allele likelihoods may be determined when the accelerated genotype attribution system performs a second path in a forward direction across the haplotype matrix, or vice versa.
[0043] As described above, in some embodiments, the accelerated genotype attribution system stores or accesses a subset of first-pass or second-pass intermediate allele likelihoods corresponding to a subset of marker variants from a group of marker variants. As used herein, the term "group of marker variants" refers to a segment or window of marker variants from within a larger set of marker variants. For example, the group of marker variants can include multiple groups of 100, 1000, or 5000 consecutively ordered marker variants out of a set of 50,000 marker variants. Since the haplotype matrix can represent the set of marker variants by columns (where each column represents an individual marker variant), the group of marker variants can similarly correspond to a group of rows. Thus, the subset of first-pass or second-pass intermediate allele likelihoods corresponding to a subset of marker variants can refer to a subset that includes one intermediate allele likelihood for one marker variant from within each group of marker variants, such as one marker variant for every 100, 1000, or 5000 marker variants.
[0044] As further shown above, in some embodiments, an accelerated genotype attribution system determines different running sums of different subsets of intermediate allele likelihoods for genomic regions that include different types of haplotype alleles. As used herein, the term "running sum of a subset of intermediate allele likelihoods" refers to the sum value of one or more intermediate allele likelihoods for a marker variant (e.g., an adjacent marker variant) that can be updated as additional intermediate allele likelihoods are determined. In particular, the running sum of a subset of intermediate allele likelihoods includes the sum value of multiple intermediate allele likelihoods for a genomic region that indicates or includes a particular type of haplotype allele from one or more haplotypes of a haplotype reference panel for an adjacent marker variant, and the sum value can be updated as additional intermediate allele likelihoods corresponding to the adjacent marker variant are determined. Thus, in some embodiments, the accelerated genotype attribution system (i) determines, for an adjacent marker variant, the running sum of a first subset of intermediate allele likelihoods for a genomic region that includes a first type of haplotype allele (e.g., a sample reference haplotype allele) from one or more haplotypes of a haplotype reference panel, and (ii) determines, for the adjacent marker variant, the running sum of a second subset of intermediate adjacent allele likelihoods for a genomic region that includes a second type of haplotype allele (e.g., a sample alternative haplotype allele) from one or more haplotypes.
[0045] Furthermore, as used herein, the term "genomic coordinate" refers to a specific location or position of a nucleotide base within a genome (e.g., the genome of an organism or a reference genome). In some cases, a genomic coordinate includes an identifier for a particular chromosome of the genome and an identifier for the position of a nucleotide base within the particular chromosome. For example, a genomic coordinate (singular or plural) can include a number, name, or other identifier of a chromosome (e.g., chr1 or chrX), and a specific position (singular or plural) such as a numbered position following the identifier of the chromosome (e.g., chr1:1234570 or chr1:1234570~1234870). Further, in certain embodiments, a genomic coordinate refers to a source of the reference genome (e.g., mt for a mitochondrial DNA reference genome, or SARS-CoV-2 for a SARS-CoV-2 virus reference genome), and a position of a nucleotide base within the source for the reference genome (e.g., mt:16568 or SARS-CoV-2:29001). In contrast, in certain cases, a genomic coordinate refers to the position of a nucleotide base within the reference genome without reference to a chromosome or source (e.g., 29727).
[0046] Furthermore, as used herein, "genomic region" refers to a range of genomic coordinates. Similar to genomic coordinates, in certain embodiments, a genomic region can be identified by an identifier for a chromosome and a specific position (singular or plural), e.g., a numbered position following the identifier of the chromosome (e.g., chr1:1234570~1234870).
[0047] As used herein, for example, the term "configurable processor" refers to a circuit or chip that can be configured or customized to execute a particular application. For example, a configurable processor includes an integrated circuit chip designed to be configured or customized on-site by an end-user computing device to execute a particular application. Configurable processors include, but are not limited to, ASICs, ASSPs, coarse-grained reconfigurable arrays (CGRAs), or FPGAs. In contrast, configurable processors do not include CPUs or GPUs. In some embodiments, the accelerated genotype attribution system uses a configurable processor (e.g., an FPGA) or a processor (e.g., a CPU) to implement the various embodiments described herein.
[0048] As further used herein, the term "nucleotide base call" (or simply "base call") refers to the determination or prediction of a particular nucleotide base (or nucleotide base pair) for an oligonucleotide (e.g., a read) during a sequencing cycle or for a genomic locus of a sample genome. In particular, a nucleobase call may refer to (i) the determination or prediction of the type of nucleobase incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., a read-based nucleobase call), or (ii) the determination or prediction of the type of nucleobase present at a genomic locus or region within the genome, including a variant call or non-variant call in a digital output file. In some cases, for a nucleotide fragment read, a nucleotide base call may include the determination or prediction of a nucleotide base based on intensity values obtained from fluorescently tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., within a cluster of a flow cell). Alternatively, a nucleotide base call may include the determination or prediction of a nucleotide base from a chromatogram peak or current change resulting from a nucleotide passing through a nanopore of a nucleotide-sample slide. In contrast, a nucleotide base call may also include the final prediction of a nucleotide base at a genomic locus of a sample genome for a variant call file (VCF) or other base call output file based on a nucleotide fragment read corresponding to the genomic locus. Thus, a nucleotide base call may include a base call corresponding to a genomic locus and a reference genome, e.g., an indication of a variant or non-variant at a particular position corresponding to the reference genome. Indeed, a nucleotide base call may refer to a variant call, including but not limited to single nucleotide variants (SNVs), insertions or deletions (indels), or a base call that is part of a structural variant. As suggested above, a single nucleotide base call may be an adenine (A) call, a cytosine (C) call, a guanine (G) call, or a thymine (T) call.
[0049] As further used herein, the term "nucleotide-sample slide" refers to a plate or slide containing oligonucleotides for sequencing nucleotide sequences derived from genomic samples or other sample nucleic acid polymers. In particular, a nucleotide-sample slide refers to a slide containing fluid channels through which reagents and buffers can move as part of the sequencing. For example, in one or more embodiments, a nucleotide-sample slide includes a flow cell (e.g., a patterned flow cell or an unpatterned flow cell) containing small fluid channels and short oligonucleotides complementary to the binding adapter sequence. As noted above, a nucleotide-sample slide can include wells (e.g., nanowells) containing clusters of oligonucleotides.
[0050] As suggested above, a flow cell or other nucleotide-sample slide may comprise an apparatus having a lid that extends over the reaction structure to form flow channels therebetween that communicate with a plurality of reaction sites of the reaction structure, and may comprise a detection device configured to detect a designated reaction occurring at or proximal to the reaction site. The flow cell or other nucleotide-sample slide may include a solid-state light detection or imaging device, such as a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) (optical) detection device. As one specific example, the flow cell can be electrically coupled to a cartridge (having an integrated pump) that is fluidically configured and configured to be fluidically and / or electrically coupled to a bioassay system. The cartridge and / or bioassay system can deliver a reaction solution to the reaction sites of the flow cell and perform a plurality of imaging events according to a predetermined protocol (e.g., sequencing by synthesis). For example, the cartridge and / or bioassay system can direct one or more reaction solutions through the flow channels of the flow cell and thereby along the reaction sites. At least one of the reaction solutions may include four types of nucleotides having the same or different fluorescent labels. The nucleotides can be bound to the reaction sites of the flow cell, such as corresponding oligonucleotides of the reaction sites. The cartridge and / or bioassay system can then illuminate the reaction sites using an excitation light source (e.g., a solid-state light source such as a light-emitting diode (LED)). The excitation light can provide an emission signal (e.g., light of one or more wavelengths different from the excitation light and potentially different from each other) that can be detected by a light sensor of the flow cell.
[0051] As further used herein, the term "performing sequencing" refers to an iterative process on a sequencing device to determine the primary structure of a nucleotide sequence from a sample (e.g., a genomic sample). In particular, performing sequencing includes the cycles of sequencing chemistry and imaging performed by a sequencing device that incorporates nucleobases into growing oligonucleotides to determine nucleotide fragment reads from a nucleotide sequence (or other sequences within a library fragment) that has been extracted from the sample and seeded across the nucleotide-sample slide. Optionally, performing sequencing includes replicating a nucleotide sequence from one or more genomic samples seeded in clusters across a nucleotide-sample slide (e.g., a flow cell). Once performing sequencing is complete, the sequencing device can generate base call data in a file.
[0052] As just suggested, the term "base call data" refers to data representing nucleobase calls for nucleotide fragment reads and / or corresponding sequencing metrics. For example, base call data includes text data representing the nucleobase calls of nucleotide fragment reads as text (e.g., A, C, G, T), along with corresponding base call quality metrics, depth metrics, and / or other sequencing metrics. Optionally, base call data is formatted in a text file, such as a binary base call (BCL) sequence file, or as a fast-all quality (FASTQ) file.
[0053] As further used herein, the term "nucleotide fragment read" (or simply "read") refers to a deduced sequence of one or more nucleobases (or nucleobase pairs) from all or part of a sample nucleotide sequence (e.g., a sample genomic sequence, cDNA). In particular, a nucleotide fragment read includes the sequence of nucleotide calls determined or predicted for a nucleotide sequence (or a group of monoclonal nucleotide sequences) from a sample library fragment corresponding to a genomic sample. For example, in some cases, a sequencing device determines a nucleotide fragment read by generating nucleotide calls for nucleobases that have passed through nanopores of a nucleotide-sample slide, determined via fluorescent tagging, or determined from clusters within a flow cell.
[0054] The following paragraphs describe an accelerated genotype attribution system with respect to exemplary diagrams depicting exemplary embodiments and implementations. For example, FIG. 1 shows a schematic diagram of a computing system 100 in which an accelerated genotype attribution system 106 operates, according to one or more embodiments. As shown, computing system 100 includes a local device 110 (e.g., a local server device), one or more server devices 120, and a sequencing device 102 connected to a client device 116. As shown in FIG. 1, sequencing device 102, local device 110, server device 120, and client device 116 may communicate with each other via network 122. Network 122 includes any suitable network over which computing devices may communicate. Exemplary networks are considered in more detail below with respect to FIG. 13. Although FIG. 1 shows an embodiment of accelerated genotype attribution system 106, the present disclosure describes the following alternative embodiments and configurations.
[0055] As shown by FIG. 1, the sequencing device 102 includes a computing device and a sequencing device system 104 for sequencing a genomic sample or other nucleic acid polymer. In some embodiments, by executing the sequencing device system 104 on a processor 108 (e.g., a configurable processor), the sequencing device 102 analyzes nucleotide fragments or oligonucleotides extracted from a genomic sample and uses a computer-implemented method and system either directly or indirectly on the sequencing device 102 to generate nucleotide fragment reads or other data. More specifically, the sequencing device 102 receives a nucleotide-sample slide (e.g., a flow cell) containing nucleotide fragments extracted from a sample and additional copies thereof, and determines the nucleic acid base sequence of such extracted nucleotide fragments.
[0056] In one or more embodiments, the sequencing device 102 uses SBS to sequence nucleotide fragments into nucleotide fragment reads and determines nucleic acid base calls for the nucleotide fragment reads. In addition to, or as an alternative to, communicating via the network 122, in some embodiments, the sequencing device 102 bypasses the network 122 and communicates directly with the local device 110 or the client device 116. By executing the sequencing device system 104, the sequencing device 102 can further store nucleic acid base calls as part of base call data formatted as a BCL file and transmit the BCL file to the local device 110 and / or the server device 120.
[0057] As further shown by FIG. 1, the local device 110 is located at or near the same physical location as the sequencing device 102. In fact, in some embodiments, the local device 110 and the sequencing device 102 are integrated into the same computing device. The local device 110 can generate, receive, analyze, store, and transmit digital data by, for example, running a sequencing system 112 to receive base call data or determine variant calls based on analyzing such base call data. As shown in FIG. 1, the sequencing device 102 can transmit (and the local device 110 can receive) base call data generated during the execution of sequencing by the sequencing device 102. By executing software in the form of the sequencing system 112, the local device 110 can align nucleotide fragment reads to a reference genome and determine gene variants based on the aligned nucleotide fragment reads. The local device 110 can also communicate with the client device 116. In particular, the local device 110 can transmit to the client device 116 a variant call file (VCF), or data including nucleic acid base calls, sequencing metrics, error data, or other information indicating other metrics.
[0058] As shown above, as part of the local device 110, the accelerated genotype attribution system 106 can determine the intermediate allele likelihood of a genomic region that exhibits a particular haplotype allele as part of a genotype attribution model by using one or both of data exchanges across integrated computing and dedicated hardware. For example, the accelerated genotype attribution system 106 can determine the intermediate allele likelihood of a genomic region that includes a given haplotype allele from a particular marker variant and a haplotype from a haplotype reference panel by performing a single-pass simultaneous multiplication operation on the processor 114. In some implementations, the processor 114 is a configurable processor. In some cases, the accelerated genotype attribution system 106 (i) determines and stores a subset of intermediate allele likelihoods corresponding to a group of marker variants, and (ii) immediately generates a set of intermediate allele likelihoods for multiple passes by using the intermediate allele likelihood subset as a hot start point for a full pass of intermediate allele likelihoods. In a further embodiment, the accelerated genotype attribution system 106 determines a running total of the intermediate allele likelihoods of genomic regions that exhibit haplotype alleles for one or more haplotypes given on a marker variant as a running input for determining the intermediate allele likelihood of a genomic region that exhibits a haplotype allele for another given marker variant.
[0059] As further shown by FIG. 1, the server device 120 is located remotely from the local device 110 and the sequencing device 102. Similar to the local device 110, in some embodiments, the server device 120 includes a version of the sequencing system 112. Thus, the server device 120 can generate, receive, analyze, store, and transmit digital data, such as by receiving base call data or determining variant calls based on analyzing such base call data. Thus, the sequencing device 102 can transmit base call data from the sequencing device 102 (and the server device 120 can receive the base call data). The server device 120 can also communicate with the client device 116. In particular, the server device 120 can transmit data including VCF or other sequencing-related information to the client device 116.
[0060] In some embodiments, the server device 120 includes a distributed set of servers and the server device 120 includes several server devices that are distributed across the network 122 and located in the same or different physical locations. Further, the server device 120 can include a content server, an application server, a communication server, a web hosting server, or another type of server.
[0061] As further shown in FIG. 1, by executing the alignment determination application 118, the client device 116 can generate, store, receive, and transmit digital data. In particular, the client device 116 can receive alignment determination data from the local device 110 or receive a call file (e.g., BCL) and alignment determination metrics from the alignment determination device 102. Further, the client device 116 can communicate with the local device 110 or the server device 120 to receive a VCF including other metrics such as nucleotide calls and / or base call quality metrics or pass filter metrics. Thus, the client device 116 can present or display information regarding variant calls or other nucleotide calls within the graphical user interface of the alignment determination application 118 to the user associated with the client device 116. For example, the client device 116 can present variant calls and / or alignment determination metrics for the aligned genomic sample within the graphical user interface of the alignment determination application 118.
[0062] FIG. 1 depicts the client device 116 as a desktop or laptop computer, but the client device 116 may include various types of client devices. For example, in some embodiments, the client device 116 includes non-mobile devices such as a desktop computer or server, or other types of client devices. In still other embodiments, the client device 116 includes mobile devices such as a laptop, tablet, mobile phone, or smartphone. Further details regarding the client device 116 are discussed below with respect to FIG. 13.
[0063] As further illustrated in FIG. 1, the client device 116 includes a genotyping application 118. The genotyping application 118 can be a web application or a native application (e.g., a mobile application, a desktop application) stored and executed on the client device 116. The genotyping application 118 can include instructions that, when executed, cause the client device 116 to receive data from the accelerated genotype attribution system 106 and present data from basecalls or VCF for display on the client device 116. Further, the genotyping application 118 can instruct the client device 116 to display an overview of multiple genotyping runs.
[0064] As further shown in FIG. 1, a version of the accelerated genotype attribution system 106 can be located on and (e.g., wholly or partially) implemented on the local device 110. In yet other embodiments, the accelerated genotype attribution system 106 is implemented by one or more other components of the computing system 100, such as the server device 120. In particular, the accelerated genotype attribution system 106 can be implemented in a variety of different ways across the genotyping device 102, the local device 110, the server device 120, and the client device 116. For example, the accelerated genotype attribution system 106 can be downloaded from the server device 120 to the accelerated genotype attribution system 106 and / or the local device 110, and all or part of the functionality of the accelerated genotype attribution system 106 is implemented on each device within the computing system 100.
[0065] As suggested above, in some embodiments, the accelerated genotype attribution system 106 applies a genotype attribution model, such as a hidden Markov model (HMM)-based genotype attribution model, to nucleotide fragment reads corresponding to genomic regions of a genomic sample. By applying the genotype attribution model, the accelerated genotype attribution system 106 can determine posterior genotype likelihoods and haplotype calls for genomic regions. In accordance with one or more embodiments, FIG. 2A shows an accelerated genotype attribution system 106 that applies GLIMPSE as a genotype attribution model to determine posterior genotype likelihoods for genomic regions of a plurality of genomic samples. As part of using an HMM to attribute haplotypes, the accelerated genotype attribution system 106 utilizes a haplotype matrix 220 to determine haplotype allele likelihoods for a genomic region. In accordance with one or more embodiments, FIG. 2B shows a more detailed depiction of an accelerated genotype attribution system 106 that utilizes a haplotype matrix 220 to determine such haplotype allele likelihoods.
[0066] As shown in FIG. 2A, for example, the accelerated genotype attribution system 106 determines prior genotype likelihoods 204 that a genomic region 200 from a plurality of genomic samples exhibits a particular genotype (e.g., a reference allele or an alternative allele). As suggested by FIG. 2A, in some cases, the genomic region 200 corresponds to a substantially same set of genomic coordinates (with respect to a reference genome) for a plurality of genomic samples. As shown by nucleotide fragment reads 202, the genomic region 200 exhibits low coverage (e.g., ≤8X read coverage). In some embodiments, the accelerated genotype attribution system 106 uses a probabilistic call generation model (e.g., a variant caller from DRAGEN) to determine the prior genotype likelihoods 204 based on (i) nucleotide fragment reads 202 from a plurality of genomic samples, and (i) quality scores for base calls of the nucleotide fragment reads 202.
[0067] As further shown by FIG. 2A, the genomic region 200 corresponds to variable positions (or variable genomic coordinates) of the haplotype reference panel 206. The accelerated genotype attribution system 106 further deconvolves a vector of previous genotype likelihoods 204 into two independent vectors of haplotype allele likelihoods (or simply, haplotype likelihoods), each vector corresponding to one of two complementary haplotypes. In some such embodiments, the accelerated genotype attribution system 106 inputs the previous genotype likelihoods 204 in vector form as part of an input matrix.
[0068] Based on the haplotype likelihoods from the independent vectors, in some implementations, the accelerated genotype attribution system 106 uses the haploid version of the HMM in an iterative process to attribute two target haplotypes as haplotype calls. As shown in FIG. 2A, for example, the accelerated genotype attribution system 106 selects a haplotype 210 based on the haplotype reference panel 206 and the target haplotypes 208 estimated for each genomic sample. After selecting a haplotype for a given genomic sample, the accelerated genotype attribution system 106 stores a reference version and a target version of the selected haplotype as a positional Burrows-Wheeler transform (PBWT) 212.
[0069] As further shown in FIG. 2A, in some embodiments, the Accelerated Genotype Attribution System 106 samples haplotypes 214 in PBWT212 format by executing a linear time sampling algorithm based on a haplotype attribution version of the HMM developed by Na Li and Matthew Stephens, “Modeling Linkage Disequilibrium and Identifying Recombination Hotspots Using Single-Nucleotide Polymorphism Data,” 165 Genetics 2213-2233 (2003), which is hereby incorporated by reference in its entirety. By executing the linear time sampling algorithm as part of the sampler iteration, the Accelerated Genotype Attribution System 106 further determines (and updates) the phase of two attributed haplotypes for a genomic region out of the genomic regions 200 for a particular genomic sample.
[0070] Based on the attributed and phased haplotypes, as further shown in FIG. 2A, the Accelerated Genotype Attribution System 106 determines the posterior genotype likelihood 216 that the genomic regions 200 of a plurality of genomic samples exhibit a particular genotype (e.g., a reference allele or an alternative allele). The Accelerated Genotype Attribution System 106 further determines a haplotype call 218 for the genomic regions for each of the plurality of genomic samples. As described above, in some embodiments, the Accelerated Genotype Attribution System 106 uses a modified version of GLIMPSE developed by Rubinacci as the genotype attribution model.
[0071] As part of haplotype selection 210 and haplotype sampling 214, the accelerated genotype attribution system 106 can perform sampler iterations across genomic samples using the haplotype matrix 220. As further described below and further shown in FIG. 2B, the accelerated genotype attribution system 106 can determine the intermediate allele likelihoods of genomic regions containing haplotype alleles in both the forward and reverse directions across the haplotype matrix 220. In the haplotype matrix 220, each column represents a marker variant and each row represents a haplotype from the haplotype reference panel 206. The accelerated genotype attribution system 106 further determines the sum of the intermediate allele likelihoods for each column representing a marker variant. Based on the sum of adjacent marker intermediate allele likelihoods for each column, in some cases, the accelerated genotype attribution system 106 determines the allele likelihoods for the corresponding marker variant and haplotype. Such allele likelihoods represent an example or embodiment of the posterior genotype likelihood 216.
[0072] As shown in FIG. 2B, for example, the accelerated genotype attribution system 106 uses the input haplotype matrix 220a to input various values. As shown in FIG. 2B, the input haplotype matrix 220a and the updated haplotype matrix 220b are organized by "K" rows representing haplotypes from the haplotype reference panel 206 and "M" columns representing marker variants (e.g., SNPs or other variants). Thus, each row represents a haplotype "k" and each column represents a marker variant "m". In some embodiments, both the input haplotype matrix 220a and the updated haplotype matrix 220b include approximately 1000 rows representing approximately 1000 haplotypes from the haplotype reference panel 206 and approximately 50,000 columns representing approximately 50,000 marker variants. Thus, the input haplotype matrix 220a includes approximately 50 million cells. However, other suitable dimensions of larger or smaller columns and rows may be used.
[0073] As further shown by FIG. 2B, in some embodiments, the accelerated genotype attribution system 106 inputs values of transition coefficients (e.g., P0 and P1) and allele likelihood factors (e.g., Q0 and Q1) into each cell of the input haplotype matrix 220a. For example, the accelerated genotype attribution system 106 inputs a specific transition linear coefficient (e.g., P1) and a specific transition constant coefficient (e.g., P0) into each cell, and the transition coefficients generally represent the probability of transition between haplotypes represented by adjacent rows. Further, the accelerated genotype attribution system 106 inputs a specific allele likelihood factor (e.g., Q0) for a first type of haplotype allele for a specific haplotype represented by a row into each cell, and inputs a specific allele likelihood factor (e.g., Q1) for a second type of haplotype allele for the specific haplotype represented by the row. As described above, in some embodiments, one allele likelihood factor (e.g., Q0) corresponds to the sample reference haplotype allele of a specific haplotype represented by a row, and another allele likelihood factor (e.g., Q1) corresponds to the sample alternative haplotype of the specific haplotype.
[0074] In addition to inputting the transition coefficients and allele likelihood factors, as further shown by FIG. 2B, in certain embodiments, the accelerated genotype attribution system 106 inputs values representing haplotype alleles (S bits) into each cell of the input haplotype matrix 220a. In particular, the accelerated genotype attribution system 106 can input a value (or bit) of 0 indicating the sample reference haplotype allele of a specific haplotype represented by a row. Conversely, the accelerated genotype attribution system 106 can input a value (or bit) of 1 indicating the sample alternative haplotype allele of a specific haplotype represented by a row. For the sake of brevity, this disclosure refers to such input values representing haplotype alleles as haplotype-allele-index data for the haplotype matrix, as further described below with respect to FIG. 6.
[0075] After entering the values of the transition coefficient, the allele likelihood factor, and the haplotype allele index, in some embodiments, the accelerated genotype attribution system 106 determines the intermediate allele likelihood in each cell based on the input values. For example, in some embodiments, the accelerated genotype attribution system 106 executes an alpha path and a beta path across the cells of the input haplotype matrix 220a to determine the intermediate allele likelihood represented by darker shading in the updated haplotype matrix 220b. In fact, in a particular embodiment, the alpha value represents the intermediate allele likelihood (e.g., A[m][k]) determined during the alpha path, and the beta value represents the intermediate allele likelihood (e.g., A[m][k]) determined during the beta path. As further described below, in some embodiments, the accelerated genotype attribution system 106 executes two beta paths (including the sacrificial bet path) as part of the HMM calculation task.
[0076] To determine the intermediate allele likelihood (e.g., A[m][k]) for a target cell, in some embodiments, the accelerated genotype attribution system 106 determines a first product of the transition linear coefficient (e.g., P1[m]) for the target marker variant, the normalization value (e.g., Norm[m - 1]) for the column representing the adjacent marker variant, and the adjacent marker intermediate allele likelihood (e.g., A[m - 1][k]) for the adjacent marker variant. The normalization value (e.g., represented by a column) for a given marker variant can be any value that facilitates maintaining the value per cell from overflowing a numerical representation where an intermediate allele likelihood value or the sum of intermediate allele likelihood values exists. The accelerated genotype attribution system 106 further determines a second product of the transition constant coefficient (e.g., P0[m]), the normalization value (e.g., Norm[m - 1]) for the column representing the adjacent marker variant, and the adjacent marker intermediate allele likelihood of the total adjacent marker variant (e.g., Sum[m - 1]). The accelerated genotype attribution system 106 further multiplies the sum of the first product and the second product by the allele likelihood factor (e.g., Q[m][allele]) to determine the intermediate allele likelihood for the target cell.
[0077] As described above, such an allele likelihood factor can constitute an allele likelihood factor (e.g., Q0) corresponding to a sample reference haplotype allele of a specific haplotype represented by a column, or another allele likelihood factor (e.g., Q1) corresponds to a sample alternative haplotype of a specific haplotype. However, as described below, the accelerated genotype attribution system 106 can also perform an improved method of determining such intermediate allele likelihoods.
[0078] As further shown in FIG. 2B, in some embodiments, the accelerated genotype attribution system 106 determines, for each column, the sum of the alpha values for the marker variants and the sum of the beta values for the marker variants. In particular, in some embodiments, the accelerated genotype attribution system 106 determines (i) the sum of the intermediate allele likelihoods for the columns represented by the marker variants in one pass, and (ii) the sum of the intermediate allele likelihoods for the columns represented by the marker variants in another pass.
[0079] Based on the total intermediate allele likelihoods for each marker variant represented by a column, in some embodiments shown in FIG. 2B, the accelerated genotype attribution system 106 further determines a pair of allele likelihoods (e.g., R0 and R1) for each marker variant. For example, in a particular implementation, the accelerated genotype attribution system 106 determines a first allele likelihood (e.g., R0) in which the genomic region includes a sample reference haplotype allele corresponding to various haplotypes represented by various rows. Similarly, the accelerated genotype attribution system 106 determines a second allele likelihood (e.g., R1) in which the genomic region includes a sample alternative haplotype allele corresponding to various haplotypes represented by various rows.
[0080] As described above, in some cases, the accelerated genotype attribution system 106 facilitates intermediate allele likelihood determination by performing a single-pass simultaneous multiplication operation on a given target marker variant and a haplotype from a haplotype reference panel. According to one or more embodiments, FIG. 3A shows an accelerated genotype attribution system 106 that performs a single-pass simultaneous multiplication operation to determine the intermediate allele likelihood of a genomic region containing haplotype alleles when a target cell representing a target marker variant and a target haplotype from a haplotype reference panel are provided. FIG. 3B shows a comparison of the accelerated genotype attribution system 106 that determines such intermediate allele likelihoods for target cells using either (i) a three-pass simultaneous multiplication operation or (ii) a single-pass simultaneous multiplication operation. By pre-determining the transition recognition allele likelihood factor before the processor determines the intermediate allele likelihood for the target marker variant, the accelerated genotype attribution system 106 condenses and facilitates the processing load from a three-pass simultaneous multiplication operation to a single-pass simultaneous multiplication operation for the target cell.
[0081] As shown in FIG. 3A, for example, the accelerated genotype attribution system 106 identifies, from within the memory device 302, a haplotype reference panel 304 corresponding to the genomic regions of one or more genomic samples and a transition recognition allele likelihood factor for executing a genotype attribution model. In particular, in some embodiments, the accelerated genotype attribution system 106 identifies the haplotype reference panel 304 stored in a dynamic random access memory (DRAM), a static random access memory (SRAM), or a cache memory device. Further, the accelerated genotype attribution system 106 identifies a first transition recognition allele likelihood factor 306a and a second transition recognition allele likelihood factor 306b while executing an alpha pass or a beta pass of the haplotype matrix 308. In some cases, when the accelerated genotype attribution system 106 reaches a target cell 300 representing a combination of a target marker variant and a haplotype during a pass of the haplotype matrix 308, the accelerated genotype attribution system 106 identifies the first transition recognition allele likelihood factor 306a and the second transition recognition allele likelihood factor 306b.
[0082] To avoid determining the first and second transition recognition allele likelihood factors 306a and 306b during the pass, in some embodiments, the accelerated genotype attribution system 106 determines the first and second transition recognition allele likelihood factors 306a and 306b in advance before determining the intermediate allele likelihood for the column representing the target marker variant in the haplotype matrix 308. To pre-determine the first transition recognition allele likelihood factor 306a, in some embodiments, the accelerated genotype attribution system 106 combines (e.g., multiplies, weighted sums) the allele likelihood factor for the haplotype allele and the transition constant coefficient for the transition between haplotypes from the haplotype reference panel 304. Similarly, to pre-determine the second transition recognition allele likelihood factor 306b, the accelerated genotype attribution system 106 combines (e.g., multiplies, weighted sums) the allele likelihood factor and the transition linear coefficient for transitioning between haplotypes from the haplotype reference panel 304.
[0083] The accelerated genotype attribution system 106 can generate a predetermined version of the first and second transition recognition allele likelihood factors 306a and 306b. Because the input values are available before the pass over the haplotype matrix 308 or at least before determining the intermediate allele likelihood for the target marker variant. The accelerated genotype attribution system 106 has (and can identify) access to the allele likelihood factor and the transition coefficient for the column representing the target marker variant before determining the intermediate allele likelihood for the target marker variant. Thus, in certain implementations, the accelerated genotype attribution system 106 generates a predetermined version of the first and second transition recognition allele likelihood factors 306a and 306b. Accordingly, in some embodiments, the accelerated genotype attribution system 106 determines the first and second transition recognition allele likelihood factors 306a and 306b in advance before determining one or more intermediate allele likelihoods corresponding to the marker variant as part of the pass of the haplotype matrix 308.
[0084] As part of executing a path for determining an intermediate allele likelihood, in certain cases, the accelerated genotype attribution system 106 determines and accesses values as part of a path across the haplotype matrix 308. To determine the intermediate allele likelihood 316 for the target cell 300, in certain embodiments, the accelerated genotype attribution system 106 identifies from the haplotype matrix 308 the adjacent marker intermediate allele likelihoods 310 for adjacent marker variants relative to the target marker variant. In the haplotype matrix 308, adjacent columns represent adjacent marker variants that are adjacent to the target column representing the target marker variant. As part of a path across the haplotype matrix 308, in some embodiments, the accelerated genotype attribution system 106 determines the adjacent marker intermediate allele likelihoods 310 for combinations of adjacent marker variants and target haplotypes from the haplotype reference panel 304 before determining the intermediate allele likelihood 316.
[0085] After identifying the relevant input values for the multiplication operation, as further shown in FIG. 3A, the accelerated genotype attribution system 106 combines the adjacent marker intermediate allele likelihoods 310 and the first transition recognition allele likelihood factor 306a. In particular, in some embodiments, the accelerated genotype attribution system 106 multiplies the adjacent marker intermediate allele likelihoods 310 and the first transition recognition allele likelihood factor 306a within the path of the haplotype matrix 308. Since the accelerated genotype attribution system 106 determines both the adjacent marker intermediate allele likelihoods 310 and the first transition recognition allele likelihood factor 306a before passing through the cell representing the target marker variant and the target haplotype, the accelerated genotype attribution system 106 can use this single path simultaneous multiplication operation as part of determining the intermediate allele likelihood 316 for the target cell 300. As shown in FIG. 3A, based on combining the adjacent marker intermediate allele likelihoods 310 and the first transition recognition allele likelihood factor 306a, the accelerated genotype attribution system 106 generates the adjacent marker transition factor recognition allele likelihoods 314.
[0086] As further suggested above, in some embodiments, the accelerated genotype attribution system 106 determines an intermediate allele likelihood 316 of a genomic region including haplotype alleles based on an adjacent marker transition factor recognition allele likelihood 314 and a second transition recognition allele likelihood factor 306b. For example, in some embodiments, the accelerated genotype attribution system 106 determines the sum of the adjacent marker transition factor recognition allele likelihood 314 and the second transition recognition allele likelihood factor 306b to determine the intermediate allele likelihood 316. As further described below, in certain implementations, the accelerated genotype attribution system 106 determines the intermediate allele likelihood 316 by determining the sum of (i) the adjacent marker transition factor recognition allele likelihood 314 and (ii) the product of the second transition recognition allele likelihood factor 306b and the total adjacent marker intermediate allele likelihood 312 for the adjacent marker variant.
[0087] As described above, the accelerated genotype attribution system 106 can reduce computer processing from three multiplication operations to one multiplication operation to determine the intermediate allele likelihood for a target cell. According to one or more embodiments, FIG. 3B shows an accelerated genotype attribution system 106 that uses a configurable processor to execute a multiple multiplication model 318 and a single multiplication model 320 for determining an intermediate allele likelihood for a target cell representing a combination of a target marker variant and a haplotype within a haplotype matrix.
[0088] As shown in FIG. 3B, when using the multiplicative model 318, the accelerated genotype attribution system 106 performs multiplication operations 334a, 334b, and 334c as part of determining the intermediate allele likelihood 332a for the target cell. In the following, the multiplication operations 334a, 334b, and 334c are briefly summarized in the order shown in FIG. 3B, although any order can be used. First, the accelerated genotype attribution system 106 performs the multiplication operation 334a by multiplying the transition constant coefficient 322 (e.g., P0) for the column representing the target marker variant by the total adjacent marker intermediate allele likelihood 324 (e.g., Sum[m−1]) for the adjacent marker variant. In some cases, the total adjacent marker intermediate allele likelihood 324 is normalized (e.g., Norm[m−1] * Sum[m−1]). For simplicity, this disclosure uses an apostrophe as an abbreviated representation for indicating a normalized value (e.g., Sum’[m−1]).
[0089] Second, the accelerated genotype attribution system 106 performs the multiplication operation 334b by multiplying the transition linear coefficient 326 (e.g., P1) for the column representing the target marker variant by the adjacent marker intermediate allele likelihood 328a (e.g., A[m−1][k]) for the adjacent marker variant. In some cases, the adjacent marker intermediate allele likelihood 328a is normalized (e.g., Norm[m−1] * A[m−1][k]). As further shown in FIG. 3B, the accelerated genotype attribution system 106 performs the summation operation 340a by summing (i) the product of the transition constant coefficient 322 (P0) and the total adjacent marker intermediate allele likelihood 324 (e.g., Norm[m−1] * Sum[m−1]) and (i) the product of the transition linear coefficient 326 (P0) and the adjacent marker intermediate allele likelihood 328a (e.g., Norm[m−1] * A[m−1][k]).
[0090] Third, the accelerated genotype attribution system 106 performs a multiplication operation 334c by multiplying the allele likelihood factor 330a (e.g., Q0 or Q1) for the column representing the target marker variant by the total product. As suggested above, the allele likelihood factor 330a can constitute an allele likelihood factor (e.g., Q0) corresponding to the sample reference haplotype allele of the target haplotype represented by the column, or another allele likelihood factor (e.g., Q1) corresponding to the sample alternative haplotype of the target haplotype. Based on the multiplication of the allele likelihood factor 330a (e.g., Q0 or Q1) and the total product (P1[m] * Norm[m-1] * A[m-1][k]+P0[m] * Norm[m-1] * Sum[m-1]), the accelerated genotype attribution system 106 determines an intermediate allele likelihood 332a (e.g., A[m][k]) using the multiple multiplication model 318.
[0091] When using the multiple multiplication model 318, in some embodiments, the accelerated genotype attribution system 106 determines the intermediate allele likelihood for the target cell between both the alpha path and the beta path. Thus, the values corresponding to the adjacent marker variant (m-1) are different for the target cell from the alpha path to the beta path. In fact, by using the multiple multiplication model 318, the accelerated genotype attribution system 106 determines one value for the column representing the target marker variant by performing a multiplication operation 334a for the alpha path, and determines another value for the column representing the target marker variant by performing a multiplication operation 334a for the beta path. Further, by using the multiple multiplication model 318, the accelerated genotype attribution system 106 determines one value for each row and column by performing a multiplication operation 334b for the alpha path, and determines another value for each row and column by performing a multiplication operation 334b for the beta path.
[0092] In contrast to the multiple multiplication model 318, when using the single multiplication model 320, the accelerated genotype attribution system 106 performs a multiplication operation 334d as part of determining an intermediate allele likelihood 332b for a target cell. Generally speaking, the accelerated genotype attribution system 106 performs the multiplication operation 334d by multiplying a first transition recognition allele likelihood factor 338 and an adjacent marker intermediate allele likelihood 328b. By further performing a summation operation 340b that sums an adjacent marker transition factor recognition allele likelihood 342 and a total adjacent marker transition recognition allele likelihood factor 336, the accelerated genotype attribution system 106 determines the intermediate allele likelihood 332b for the target cell.
[0093] By using the single multiplication model 320 shown in FIG. 3B, in some embodiments, the accelerated genotype attribution system 106 selects a haplotype allele 330b for a column representing a target variant marker within the haplotype matrix. In some cases, the haplotype allele 330b takes the form of an S-bit that selects a value representing the haplotype allele for passing to downstream logic. For example, in a particular embodiment, the accelerated genotype attribution system 106 selects the haplotype allele 330b by identifying either (i) an allele likelihood factor (e.g., Q0) corresponding to the sample reference haplotype allele of the target haplotype represented by the row, or (ii) another allele likelihood factor (e.g., Q1) corresponding to the sample alternative haplotype of the target haplotype. Based on the identified allele likelihood factor (e.g., Q0 or Q1), the accelerated genotype attribution system 106 passes or transmits a corresponding value representing the downstream haplotype allele for use in the total adjacent marker transition recognition allele likelihood factor 336 and the first transition recognition allele likelihood factor 338. In fact, as further shown in FIG. 3B, the accelerated genotype attribution system 106 uses the selected haplotype allele 330b as part of the total adjacent marker transition recognition allele likelihood factor 336 and uses the first transition recognition allele likelihood factor 338 as part of the single multiplication model 320.
[0094] As suggested above, in some embodiments, the accelerated genotype attribution system 106 determines the first transition recognition allele likelihood factor 338 and the second transition recognition allele likelihood factor (the latter as part of the total adjacent marker transition recognition allele likelihood factor 336) before determining the intermediate allele likelihood for the columns representing the target marker variants within the haplotype matrix. To pre-determine the first transition recognition allele likelihood factor 338, in some embodiments, the accelerated genotype attribution system 106 multiplies the allele likelihood factor corresponding to a particular type of haplotype allele of the haplotype allele 330b (e.g., Q[m][allele]) by the transition constant coefficient (P0) for the transition between haplotypes from the haplotype reference panel. To pre-determine the total adjacent marker transition recognition allele likelihood factor 336, the accelerated genotype attribution system 106 multiplies the allele likelihood factor (e.g., Q[m][allele]), the transition linear coefficient for transitioning between haplotypes from the haplotype reference panel (e.g., P1), and the total adjacent marker intermediate allele likelihood 324 for the adjacent marker variant (e.g., Sum’[m−1]).
[0095] Between the paths of the haplotype matrix, the accelerated genotype attribution system 106 also determines the adjacent marker intermediate allele likelihood 328b for the adjacent variant marker and the adjacent cell representing the target haplotype. In fact, in some embodiments, since the accelerated genotype attribution system 106 executes a path for determining the intermediate allele likelihood for each column of the haplotype matrix, the accelerated genotype attribution system 106 determines the adjacent marker intermediate allele likelihood 328b for the adjacent cell before reaching the target cell.
[0096] When the first transition recognition allele likelihood factor 338 and the adjacent marker intermediate allele likelihood 328b are determined in advance, the accelerated genotype attribution system 106 can perform a single-pass simultaneous multiplication operation on the target cell. In particular, as shown in FIG. 3B, the accelerated genotype attribution system 106 performs a multiplication operation 334d by multiplying the first transition recognition allele likelihood factor 338 (e.g., Q[m][allele] * P1[m]) and the adjacent marker intermediate allele likelihood 328b (e.g., A’[m - 1][k]). As the output of the multiplication operation 334d, the accelerated genotype attribution system 106 generates an adjacent marker transition factor recognition allele likelihood 342 (e.g., Q[m][allele] * P1[m] * A’[m - 1]).
[0097] As further shown in FIG. 3B, the accelerated genotype attribution system 106 further determines the intermediate allele likelihood 332b for the target cell by performing a summation operation 340b. In particular, the accelerated genotype attribution system sums the adjacent marker transition factor recognition allele likelihood 342 (e.g., Q[m][allele] * P1[m] * A’[m - 1]) and the total adjacent marker transition recognition allele likelihood factor 336 (e.g., Q[m][allele] * P0[m] * Sum’[m - 1]) to determine the intermediate allele likelihood 332b (e.g., A[m][k]).
[0098] As suggested above, by using the multiple multiplication model 318 to perform three multiplication operations 334a - 334c for each target cell, the accelerated genotype attribution system 106 performs 3000 multiplication operations for each row representing a haplotype from the haplotype reference panel. In contrast, by using the single multiplication model 320 to perform the multiplication operation 334d for each target cell, the accelerated genotype attribution system 106 reduces the processing to approximately 1000 multiplication operations for each row representing a haplotype from the haplotype reference panel. Since multiplication operations on a configurable processor such as an FPGA consume significant processing, the single multiplication model 320 significantly reduces both the time and computer processing for determining the intermediate allele likelihood and the output allele likelihood.
[0099] In addition to, or alternatively to, performing a single - pass simultaneous multiplication operation on the target cell, in some embodiments, the accelerated genotype attribution system 106 can store and use a subset of the intermediate allele likelihoods to hot - start the determination of specific intermediate allele likelihoods across the haplotype matrix. According to one or more embodiments, FIG. 4A shows the accelerated genotype attribution system 106 that stores and accesses a subset of the intermediate allele likelihoods corresponding to a group of marker variants for hot - starting the intermediate allele likelihood determination between one or more passes across the haplotype matrix. FIG. 4B shows the accelerated genotype attribution system 106 that (i) determines and stores a subset of the intermediate allele likelihoods corresponding to columns of marker variants grouped together, and (ii) generates a set of intermediate allele likelihoods for a pass across the haplotype matrix by using the subset of the intermediate allele likelihoods as a hot - start point.
[0100] As shown in FIG. 4A, in some embodiments, the accelerated genotype attribution system 106 uses a configurable processor 400 to perform a first sacrificial pass 402 that determines intermediate allele likelihoods across the cells of the haplotype matrix 404. Since the accelerated genotype attribution system 106 performs the first sacrificial pass 402 for the purpose of determining a subset of the first pass intermediate allele likelihoods 406 corresponding to a subset of the marker variants, the present disclosure refers to the first sacrificial pass 402 as a "sacrifice." In addition to the hot start points for reproducing the first pass intermediate allele likelihoods, in some embodiments, the accelerated genotype attribution system 106 does not directly use the intermediate allele likelihoods determined during the first sacrificial pass 402.
[0101] When executing the sacrificial first path 402, the accelerated genotype attribution system 106 may execute a forward path or a reverse path (or an alpha path or a beta path). As suggested above, in the forward path, the accelerated genotype attribution system 106 generates a forward intermediate allele likelihood for a genomic region containing haplotype alleles. In contrast, in the reverse path, the accelerated genotype attribution system 106 generates a reverse intermediate allele likelihood for a genomic region containing haplotype alleles. Since the accelerated genotype attribution system 106 executes both the forward path (e.g., the second path) and the reverse path (e.g., the first path) as a basis for generating allele likelihoods regardless of the direction of the sacrificial path, the direction of the sacrificial path should not affect the allele likelihoods (e.g., R0, R1). Regardless of the direction, in some embodiments, the accelerated genotype attribution system 106 determines the intermediate allele likelihoods for each cell and column of the haplotype matrix 404 for each cell representing a combination of a marker variant and a haplotype from the haplotype reference panel, thereby executing the sacrificial first path 402. By executing the sacrificial first path 402, the accelerated genotype attribution system 106 utilizes the configurable processor 400 to determine the first path intermediate allele likelihoods of the genomic regions from a genomic sample containing haplotype alleles corresponding to a given set of haplotypes for a set of marker variants.
[0102] After executing the sacrificial first pass 402, as further shown in FIG. 4A, the accelerated genotype attribution system 106 identifies first-pass intermediate allele likelihoods 406a-406n from among the first-pass intermediate allele likelihoods determined from the sacrificial first pass 402. For example, in some embodiments, the accelerated genotype attribution system 106 identifies a group of marker variants, such as a group of 20, 100, 500, or 1000 marker variants, and (ii) selects the first-pass intermediate allele likelihoods from each group of marker variants for inclusion within a subset of the first-pass intermediate allele likelihoods 406. Thus, in some embodiments, the accelerated genotype attribution system 106 selects the intermediate allele likelihoods for one column of marker variants for every 20, 100, 500, or 1000 columns of marker variants within the haplotype matrix 404.
[0103] As shown in FIG. 4A, the first-pass intermediate allele likelihoods 406a-406n represent the intermediate allele likelihoods from the columns selected for each threshold number of columns representing a group of marker variants. At the same time, up to the first-pass intermediate allele likelihoods 406a, 406b, and 406n constitute a subset of the first-pass intermediate allele likelihoods 406.
[0104] In addition to identifying a subset of the first pass intermediate allele likelihoods 406, as further shown in FIG. 4A, the accelerated genotype attribution system 106 stores a subset of the first pass intermediate allele likelihoods 406 on the memory device 408. As suggested above, the values within the haplotype matrix 404 after the sacrificial first pass 402 are found to saturate or be too numerous to store in the on-chip memory of the configurable processor 400. To reduce and redistribute the large amount of data in the haplotype matrix 404 after the sacrificial first pass 402, the accelerated genotype attribution system 106 stores a subset of the first pass intermediate allele likelihoods 406 on DRAM, SRAM, or other suitable memory for the memory device 408. The memory device 408 may be on-chip with the configurable processor 400 or off-chip from the configurable processor 400. Without saturating the memory of the configurable processor 400, the accelerated genotype attribution system 106 can access a subset of the first pass intermediate allele likelihoods 406 from the memory device 408 as a hot start point for determining intermediate allele likelihoods in the first pass 410.
[0105] As further shown in FIG. 4A, in some embodiments, the accelerated genotype attribution system 106 regenerates the first pass intermediate allele likelihoods from the sacrificial first pass 402 by using a subset of the first pass intermediate allele likelihoods 406 to initialize allele likelihood determination in a group of marker variants. In particular, when executing the first pass 410, the accelerated genotype attribution system 106 (i) uses one of the first pass intermediate allele likelihoods 406a - 406n as the intermediate allele likelihood for one column of marker variants for every 20, 100, 500, or 1000 columns of marker variants, and (ii) uses one of the first pass intermediate allele likelihoods 406a - 406n as a hot start point to determine subsequent intermediate allele likelihoods in subsequent columns during the first pass 410.
[0106] As further shown in FIG. 4A, the accelerated genotype attribution system 106 can further execute a second path 412 that determines the second path intermediate allele likelihood in a direction different from the first path 410. In particular, the accelerated genotype attribution system 106 utilizes the configurable processor 400 to determine the second path intermediate allele likelihood of a genomic region that includes haplotype alleles corresponding to a set of haplotypes given a set of marker variants. Based on the reproduced first path intermediate allele likelihood and the second path intermediate allele likelihood, the accelerated genotype attribution system 106 generates an allele likelihood of a genomic region that includes haplotype alleles.
[0107] FIG. 4B shows a more detailed embodiment of the accelerated genotype attribution system 106 that uses a subset of intermediate allele likelihoods as hot start points. As shown in FIG. 4B, the accelerated genotype attribution system 106 determines and stores a subset of beta path intermediate allele likelihoods 416 corresponding to each column of marker variants grouped into groups of marker variants 1G-6G that include marker variant groups. The accelerated genotype attribution system 106 then accesses the subset of beta path intermediate allele likelihoods 416 and uses the individually stored intermediate allele likelihoods as hot start points to generate intermediate allele likelihoods in both the alpha path and the beta path across the haplotype matrix 404.
[0108] As shown in FIG. 4B, as a sacrificial beta path, the accelerated genotype attribution system 106 executes a continuous beta path 414 that determines a beta path intermediate allele likelihood corresponding to a set of haplotypes and a set of marker variants represented by the haplotype matrix 404. In particular, the accelerated genotype attribution system 106 executes the continuous beta path 414 by determining a beta path intermediate allele likelihood for each cell within the haplotype matrix 404. FIG. 4B uses the continuous beta path 414 as an exemplary sacrificial path, but the accelerated genotype attribution system 106 can similarly use a continuous alpha path as a sacrificial path. However, due to space constraints, FIG. 4B shows the continuous beta path 414 within a horizontal block. However, the continuous beta path 414 generates a beta path intermediate allele likelihood (also known as a beta value) for each cell and each cell column within the haplotype matrix 404. The continuous beta path 414 is generally executed in the reverse direction across the haplotype matrix 404 (typically represented from right to left), but FIG. 4B shows groups of columns 6G to 1G representing groups of marker variants in reverse numerical order along a horizontal processing timeline.
[0109] After executing the continuous beta path 414, the accelerated genotype attribution system 106 identifies the beta path intermediate allele likelihoods 416a - 416e and stores them in the memory device 408 as a subset of the beta path intermediate allele likelihoods 416. As shown by Figure 4B, each of the beta path intermediate allele likelihoods 416a - 416e corresponds to a column representing a marker variant from a group of columns (e.g., one of groups 1G - 5G). For example, the beta path intermediate allele likelihood 416a represents a column of intermediate allele likelihood values selected from group 5G of columns representing a group of marker variants. In contrast, the beta path intermediate allele likelihood 416b represents a column of intermediate allele likelihood values selected from group 4G of columns representing a group of marker variants. Similarly, the beta path intermediate allele likelihoods 416c, 416d, and 416e each represent a column of intermediate allele likelihood values selected from one of groups 3G, 2G, and 1G of columns representing different groups of marker variants, respectively. In some cases, the accelerated genotype attribution system 106 selects the last column of the intermediate allele likelihood (e.g., the beta path intermediate allele likelihood 416e) as the beta path intermediate allele likelihood and stores it for a particular group of columns / marker variants (e.g., 1G).
[0110] After storing the beta path intermediate allele likelihoods 416a - 416e in the memory device 408 as a subset of the beta path intermediate allele likelihoods 416, as further shown in Figure 4B, the accelerated genotype attribution system 106 executes the segmented beta path 417. When executing the segmented beta path 417, the accelerated genotype attribution system 106 reproduces the intermediate allele likelihood values determined in the continuous beta path 414. However, to conserve on-chip memory for the configurable processor or other processors, the accelerated genotype attribution system 106 loads the beta path intermediate allele likelihoods from a subset of the beta path intermediate allele likelihoods 416 in a particular column and initializes (or hot starts) the determination of the beta path intermediate allele likelihoods for adjacent columns without having to re-determine the subset of the beta path intermediate allele likelihoods 416 during the segmented beta path 417.
[0111] As further shown in FIG. 4B, the accelerated genotype attribution system 106 loads a relevant stored subset of the beta-pass intermediate allele likelihoods into associated columns between the segmented beta-paths 417. The segmented beta-paths 417 are generally executed in reverse across the haplotype matrix 404 (typically represented from right to left), but FIG. 4B shows a group of columns representing a group of marker variants proceeding in reverse numerical order along the horizontal processing timeline. As suggested above, if the accelerated genotype attribution system 106 executes a sacrificial alpha-path instead of, or in addition to, the sacrificial beta-path, the accelerated genotype attribution system 106 similarly executes a segmented alpha-path.
[0112] To illustrate the sequence of the segmented beta-paths 417, in some embodiments, the accelerated genotype attribution system 106 determines the beta-pass intermediate allele likelihoods for an initial group 0G of columns and then loads the beta-pass intermediate allele likelihood 416e for the first column of the first group 1G of columns. Based on the beta-pass intermediate allele likelihood 416e, the accelerated genotype attribution system 106 determines the beta-pass intermediate allele likelihoods for the columns adjacent to the first column within the first group 1G of columns. Similarly, the accelerated genotype attribution system 106 determines the beta-pass intermediate allele likelihoods for the entire first group 1G of columns and then loads the beta-pass intermediate allele likelihood 416d for the first column of the second group 2G of columns. Based on the beta-pass intermediate allele likelihood 416d, the accelerated genotype attribution system 106 determines the beta-pass intermediate allele likelihoods for the columns adjacent to the first column within the second group 2G of columns.
[0113] In addition to the segmented beta path 417, as further shown in FIG. 4B, the accelerated genotype attribution system 106 also executes a continuous alpha path 418 that determines alpha path intermediate allele likelihoods corresponding to a set of haplotypes and a set of marker variants represented by the haplotype matrix 404. In particular, the accelerated genotype attribution system 106 executes the continuous alpha path 418 by determining alpha path intermediate allele likelihoods for each cell within the haplotype matrix 404. Since the continuous alpha path 418 is generally executed in a forward direction across the haplotype matrix 404 (typically represented from left to right), FIG. 4B shows groups of columns 0G-6G that represent groups of marker variants in numerical order along a horizontal processing timeline.
[0114] As further shown in FIG. 4B, as both the segmented beta path 417 and the continuous alpha path 418 proceed, the accelerated genotype attribution system 106 determines a segmented allele likelihood 420. To illustrate the sequence of the segmented allele likelihood 420, in some embodiments, the accelerated genotype attribution system 106 determines the allele likelihood for the initial group 0G of columns by multiplying the sum of the corresponding beta path and alpha path intermediate allele likelihoods. When the accelerated genotype attribution system 106 then loads the beta path intermediate allele likelihood 416e for the first column of the first group 1G of columns and determines the alpha path intermediate allele likelihood for the first column as part of the continuous alpha path 418, the accelerated genotype attribution system 106 multiplies the respective sums of the beta path intermediate allele likelihood 416e and the alpha path intermediate allele likelihood for the first column of the first group 1G of columns. Based on such multiplication of the sums, the accelerated genotype attribution system 106 determines the allele likelihoods (R0 and R1) for the first column of the first group 1G of columns. In some embodiments, the accelerated genotype attribution system 106 overwrites the respective sums of the beta path intermediate allele likelihood 416e and the alpha path intermediate allele likelihood for the first column of the first group 1G of columns with the allele likelihood for the first column of the first group 1G of columns.
[0115] As a further example, in some embodiments, the accelerated genotype attribution system 106 determines the allele likelihood for the first group 1G of columns by multiplying the sum of the corresponding beta-path and alpha-path intermediate allele likelihoods. If the accelerated genotype attribution system 106 loads the beta-path intermediate allele likelihood 416d for the first column of the columns of the second group 2G and determines the alpha-path intermediate allele likelihood for the first column as part of the continuous alpha path 418, the accelerated genotype attribution system 106 multiplies each of the sum of the beta-path intermediate allele likelihood 416d and the alpha-path intermediate allele likelihood for the first column of the columns of the second group 2G. Based on such multiplication of the sums, the accelerated genotype attribution system 106 determines the allele likelihoods (R0 and R1) for the first column of the second group 2G of columns and, in some cases, overwrites each of the sums with the allele likelihoods for the first column of the second group 2G of columns.
[0116] In addition to or instead of using an intermediate allele likelihood subset as a hot start point, in some embodiments, the accelerated genotype attribution system 106 determines and uses a running sum of intermediate allele likelihoods to facilitate the execution of a path for determining intermediate allele likelihoods across a haplotype matrix. In accordance with one or more embodiments, FIG. 5A shows an accelerated genotype attribution system 106 that determines a running sum of intermediate allele likelihoods for genomic regions indicating haplotype alleles for one or more haplotypes in column n-1 (representing the first marker variant) as a running input for determining the individual intermediate allele likelihoods of genomic regions indicating haplotype alleles in column n (representing the second marker variant). FIG. 5B shows a comparison of an accelerated genotype attribution system 106 that uses a total sum model and a running sum model to determine the effect of such a model on the column sum of intermediate likelihoods and latency.
[0117] As shown in FIG. 5A, the accelerated genotype attribution system 106 executes a full-column sum model 502 to determine the intermediate allele likelihoods for columns representing different variant markers. For example, when executing the full-column sum model 502, the accelerated genotype attribution system 106 determines the sum of the intermediate allele likelihoods 506 for column n representing the second marker variant before determining the intermediate allele likelihood 508 for column n+1 representing the third marker variant. When executing the full-column sum model 502, the full-column sum model 502 causes the processor to wait for a latency for determining the intermediate allele likelihood 506 for column n- and generating the allele likelihood for column n before starting to determine the intermediate allele likelihood 508 for column n+1. Since the haplotype matrix may need to determine values corresponding to millions, billions, or trillions of cells, and determining the intermediate allele likelihoods for the cells in parallel is more efficient than a sequential approach, such latency has been found to significantly slow down a process that can take on average about 17.5 hours to phase and assign the haplotype allele likelihoods for a single marker allele corresponding to a genomic region.
[0118] In contrast to the full-column sum model 502, in some embodiments, the accelerated genotype attribution system 106 executes a running-column sum model 504 to determine the intermediate allele likelihoods for columns representing different variant markers. As shown in FIG. 5A, for example, the accelerated genotype attribution system 106 determines the running sum of the intermediate allele likelihoods 510 of a genomic region showing the haplotype alleles for one or more haplotypes, given the first marker variant represented by column n-1. By using the running sum of the intermediate allele likelihoods 510 as the running input for column n, the accelerated genotype attribution system 106 determines the sum of the intermediate allele likelihoods 512 of a genomic region showing the haplotype alleles given the second marker variant represented by column n.
[0119] When running the running column sum model 504, the accelerated genotype attribution system 106 facilitates the parallel determination of the intermediate allele likelihoods for the haplotype matrix cells. As further shown in FIG. 5A, given a second marker variant represented by column n, the accelerated genotype attribution system 106 further determines such a running sum of the intermediate allele likelihoods. By using the running sum of the intermediate allele likelihoods for column n as a running input, the accelerated genotype attribution system 106 similarly determines the intermediate allele likelihood 514 of the genomic region showing the given haplotype alleles for the third marker variant represented by column n+1. In fact, the accelerated genotype attribution system 106 can derive (or otherwise determine) the intermediate allele likelihood 514 of column n+1 based on the sum of the intermediate allele likelihoods 512 of column n. Unlike the full column sum model 502, by using the running column sum model 504, the accelerated genotype attribution system 106 does not need to wait to determine the sum of the intermediate allele likelihoods for one column before determining the individual (or total) allele likelihoods for another column in the haplotype matrix.
[0120] FIG. 5B shows in more detail a comparison of the accelerated genotype attribution system 106 that executes the full-column sum model 502 and the running-column sum model 504, along with the relative timing of the input and output values for each column. When executing the full-column sum model 502, the accelerated genotype attribution system 106 can determine the intermediate allele likelihood (e.g., A[m][k]) for a target cell by executing the multiplication operations 334a, 334b, and 334c and the sum operation 340a shown in FIG. 3B and described above. In fact, since one such multiplication operation requires summing the entire column of the intermediate allele likelihood (e.g., the A[m][k] values), the present disclosure refers to the full-column sum model 502 as the "full-column sum." In particular, when executing the multiplication operation 334a shown in FIG. 3B, in some embodiments, the accelerated genotype attribution system 106 multiplies the transition constant coefficient (P0) for the column representing the target marker variant by the normalized sum of the adjacent marker intermediate allele likelihoods (Sum’[m-1]) for the adjacent marker variant represented by the column. Since the sum of the adjacent marker intermediate allele likelihoods (Sum’[m-1]) requires summing the intermediate allele likelihoods for the entire column representing the adjacent marker variant within the haplotype matrix, the full-column sum model 502 imposes the latency shown in FIG. 5B on the processor to determine and sum the intermediate allele likelihoods for the column representing the target marker variant without performing other parallel operations.
[0121] As shown in FIG. 5B, when executing the full-column total model 502, the accelerated genotype attribution system 106 inputs the per-cell column input value 516a into the cells of column n-1 to determine the per-cell column output value 518a for column n-1. As suggested above, in some embodiments of the full-column total model 502, the per-cell column input value 516a includes the allele likelihood factors (Q0 or Q1), transition coefficients (P1[m] and P0[m]), total adjacent marker intermediate allele likelihood (Sum’[m-1]), and normalization value (Norm[m-1]) for each cell within column n-1. Based on the per-cell column input value 516a, in some embodiments, the accelerated genotype attribution system 106 determines the per-cell column output value 518a in the form of an intermediate allele likelihood represented as an alpha value (e.g., A[m][k] value) or a beta value (e.g., B[m][k] value) for each of the alpha path or beta path. FIG. 5B shows the time for determining the per-cell column output value 518a from the cells of column n-1 as the cell update waiting time 524.
[0122] Based on the per-cell column output value 518a, as part of the full-column total model 502, the accelerated genotype attribution system 106 determines the column total output value 520a for column n-1. For example, in some embodiments, the accelerated genotype attribution system 106 determines the sum of the alpha values (
[0123] [Number] ) for column n-1 and the sum of the beta values (
[0124] [Number] ) for column n-1. Based on the column total output value 520a, the accelerated genotype attribution system 106 determines the per-column allele likelihood 522a for column n-1. For example, the accelerated genotype attribution system 106 determines the allele likelihoods (R0 and R1) for column n-1 by using the alpha values (
[0125] [Number] ) sum and beta value (
[0126] [Number] ) sum is multiplied.
[0127] As shown in FIG. 5B, when executing the full-column sum model 502, the accelerated genotype attribution system 106 inputs the column-wise input value 516b into the cells of column n and determines the column sum output value 520a and the allele likelihood 522a for each column before determining the cell-wise output value 518b for column n. Since the processor of the accelerated genotype attribution system 106 determines the column sum output value 520a and the allele likelihood 522a for each column before inputting the column-wise input value 516b, the full-column sum model 502 creates the column sum latency 526a and the allele likelihood latency 528a for each column as shown in FIG. 5B. In other words, the full-column sum model 502 requires the processor to wait for the latency for both summing the intermediate allele likelihoods of adjacent markers and generating the allele likelihoods without performing other parallel operations on adjacent haplotype matrix columns.
[0128] As further shown in FIG. 5B, the full-column sum model 502 similarly creates the column sum latency and the allele likelihood latency for each column between column n and column n + 1. The accelerated genotype attribution system 106 inputs the column-wise input value 516c into the cells of column n + 1 and determines the column sum output value 520b and the allele likelihood 522b for each column before determining the cell-wise output value 518c for column n + 1. Since the processor of the accelerated genotype attribution system 106 determines the column sum output value 520b and the allele likelihood 522b for each column before inputting the column-wise input value 516c, similar to other haplotype matrix columns, the full-column sum model 502 similarly creates the column sum latency 526b and the allele likelihood latency 528b for each column.
[0129] In contrast to the full-column total model 502, the accelerated genotype attribution system 106 eliminates such empty waiting times when executing the running-column total model 504. For example, in some embodiments, the accelerated genotype attribution system 106 calculates the running total of a first subset of intermediate allele likelihoods (e.g.,
[0130]
Number
[0131]
Number
[0132] Based on the running total of the first subset of intermediate allele likelihoods and the running total of the second subset of intermediate allele likelihoods, the accelerated genotype attribution system 106 determines, for column n representing the target marker variant, the sum of the intermediate allele likelihoods (e.g., Sum[m]) of the genomic region containing the haplotype alleles from the haplotypes of the haplotype reference panel. For example, in some embodiments, the accelerated genotype attribution system 106 determines the sum of the intermediate allele likelihoods from the alpha path and the sum of the intermediate allele likelihoods from the beta path. Based on the sum of the intermediate allele likelihoods, the accelerated genotype attribution system 106 generates the allele likelihoods (R0 and R1) of the genomic region containing the haplotype alleles for column n representing the target marker variant.
[0133] As described above, the accelerated genotype attribution system 106 can pre-determine certain variables before the path of the haplotype matrix in order to facilitate the path. In some cases, for example, the accelerated genotype attribution system 106 pre-determines and describes the column input values for each of the various cells as part of the running column total model 504. For example, in some embodiments, the accelerated genotype attribution system 106 has a first transition recognition allele likelihood factor (e.g., Q0[m] * P0[m] * (K-S1)) corresponding to the row of the first type of haplotype allele and a second transition recognition allele likelihood factor (e.g., Q1[m] * P0[m] * S1) corresponding to the row of the second type of haplotype allele. Thus, in addition to the running total, the accelerated genotype attribution system 106 can determine the sum of the intermediate allele likelihoods (e.g., Sum[m]) based further on the first transition recognition allele likelihood factor corresponding to the row of the first type of haplotype allele and the second transition recognition allele likelihood factor corresponding to the row of the second type of haplotype allele.
[0134] As a further example, as shown above, the accelerated genotype attribution system 106 can estimate the sum of adjacent marker heterozygous likelihoods (e.g., Sum[m - 1]) for determining the sum of adjacent marker heterozygous likelihoods (e.g., Sum[m - 1]) among all adjacent markers, instead of summing all adjacent marker heterozygous likelihoods (e.g., A[m][1][k] values). Thus, in some embodiments, the accelerated genotype attribution system 106 determines, for column n - 1 representing an adjacent marker variant, the sum of adjacent marker heterozygous likelihoods (e.g., Sum[m - 1]) of a genomic region containing haplotype alleles, based on the running sum of the first subset of heterozygous likelihoods (e.g.,
[0135] [Number] ) and the running sum of the second subset of heterozygous likelihoods (e.g.,
[0136] [Number] ).
[0137] Thus, in some embodiments, the accelerated genotype attribution system 106 determines the sum of heterozygous likelihoods (Sum[m]) based on a combination of (i) the sum of adjacent marker heterozygous likelihoods, (ii) the first transition recognition heterozygous likelihood factor corresponding to the row of the first type of haplotype alleles, (iii) the running sum of the first subset of heterozygous likelihoods, (iv) the running sum of the second subset of heterozygous likelihoods, and (v) the second transition recognition heterozygous likelihood factor corresponding to the row of the second type of haplotype alleles, for column n representing a marker variant. In some such cases, for example, the accelerated genotype attribution system 106 determines the sum of adjacent marker heterozygous likelihoods (Sum[m - 1]) and the first transition recognition heterozygous likelihood factor (Q0[m] * P0[m] *(K-S1)) and determine the product, and add the product to the second transition recognition allele likelihood factor (Q1[m] * P0[m] * S1).
[0138] In addition to determining the running total, in some cases, the accelerated genotype attribution system 106 multiplies the running total of a subset of the intermediate allele likelihoods by the transition recognition allele likelihood factor as part of determining the sum (Sum[m]) of the intermediate allele likelihoods for column n representing the target marker variant. For example, in some embodiments, the accelerated genotype attribution system 106 (i) multiplies the running total of the first subset of the intermediate allele likelihoods by the first transition recognition allele likelihood factor (e.g.,
[0139]
Number
[0140]
Number
[0141] Thus, in some embodiments, the accelerated genotype attribution system 106 sums, for column n representing marker variants, (a) the multiplied running total of a first subset of intermediate allele likelihoods, (b) the multiplied running total of a second subset of intermediate allele likelihoods, and (c) the product of (i) the normalized value of an adjacent marker variant, (ii) the product of the adjacent marker total of the intermediate allele likelihoods, and (iii) the sum of a first transition recognition allele likelihood factor corresponding to the row of the first type of haplotype allele and a second transition recognition allele likelihood factor corresponding to the row of the second type of haplotype allele, to determine the sum of the intermediate allele likelihoods (Sum[m]).
[0142] Figure 5B shows the effect of the running column total model 504 on various waiting times. As shown in Figure 5B, when running the running column total model 504, the accelerated genotype attribution system 106 inputs the per-cell column input value 530b into the cells of column n to determine the per-cell column output value 532b from the cells of column n. As suggested above, in some embodiments of the running column total model 504, the per-cell column input value 530b for each cell within column n is (i) the normalized value of an adjacent marker variant (e.g., Normm-1), (ii) the estimated adjacent marker total of the intermediate allele likelihoods (e.g., Sum[m-1), (iii) a first transition recognition allele likelihood factor corresponding to the row of the first type of haplotype allele (e.g., Q0[m] * P0[m] * (K-S1), (iv) a second transition recognition allele likelihood factor corresponding to the row of the second type of haplotype allele (e.g., Q1[m] * P0[m] * S1), (v) the multiplied running total of a first subset of intermediate allele likelihoods (e.g.,
[0143] [Number] ), and (v) the multiplied running total of a second subset of intermediate allele likelihoods (e.g.,
[0144]
Number
[0145] For the running column total model 504, as shown in FIG. 5B, before finishing the determination of all intermediate allele likelihoods in the output value 532b for each cell, the accelerated genotype attribution system 106 determines the column total output value 534b for column n in the form of the sum of intermediate allele likelihoods (e.g., Sum[m]). In fact, as further shown by FIG. 5B, while determining the column total output value 534b, the accelerated genotype attribution system 106 also determines the column total output value 534a for column n - 1 and the allele likelihood for each column 536a for column n - 1. Therefore, the accelerated genotype attribution system 106 inputs the input value 530b for each cell between both (i) the column total waiting time 540 for column n - 1 and (ii) the allele likelihood waiting time 542 for each column for column n - 1, and determines the column total output value 534b. Therefore, the running column total model 504 ensures that the processor of the accelerated genotype attribution system 106 determines the intermediate allele likelihood for column n between the column total waiting time 540 for column n - 1 and the allele likelihood waiting time 542 for each column for column n - 1 (instead of waiting).
[0146] As further shown by FIG. 5B, in some embodiments, the accelerated genotype attribution system 106 applies the running column total model 504 to columns n-1 and n+1. For example, the accelerated genotype attribution system 106 inputs the cell-by-cell column input value 530c for column n+1 and determines the column total output value 532c for column n+1, while also determining the column total output value 534b for column n and the allele likelihood 536b for each cell in column n, thereby ensuring that the processor of the accelerated genotype attribution system 106 does not wait for the column total latency and the allele likelihood latency for each cell for column n without performing parallel operations for other columns.
[0147] Although not shown in FIG. 5B, in some embodiments, the accelerated genotype attribution system 106 can use the running total of the intermediate allele likelihoods from column n-2 to determine the total of the intermediate allele likelihoods for column n-1. In particular, the accelerated genotype attribution system 106 inputs the cell-by-cell column input value 530a for column n-1 and determines the column total output value 534a for column n-1, while also determining the column total output value of column n-2 and the allele likelihood for each cell in column n-2. Accordingly, FIG. 5B shows the cell update latency 538 representing the time and processing for determining the cell-by-cell column output value 532a from the cells of column n-1, but in some cases, the accelerated genotype attribution system 106 uses the processor of the accelerated genotype attribution system 106 to determine other values of the haplotype matrix during the cell update latency 538.
[0148] As described above, in some embodiments, the accelerated genotype attribution system 106 intelligently transfers data to increase the throughput on a configurable processor or other processor. In accordance with one or more embodiments, FIG. 6 shows the accelerated genotype attribution system 106 that stores the haplotype-allele-index data for the haplotype matrix on a memory device, accesses the stored haplotype-allele index data, and determines values as part of a path across the haplotype matrix.
[0149] As described above, HMM-based genotype attribution may require determining and storing a vast amount of data (e.g., values for millions, billions, or trillions of cells in a haplotype matrix). For example, in some embodiments, the accelerated genotype attribution system 106 inputs values representing haplotype alleles into each cell of the haplotype matrix, such as (i) one "S" bit indicating a sample reference haplotype allele for a particular haplotype, and (ii) another "S" bit indicating a sample alternative haplotype allele for a particular haplotype. As noted above, the present disclosure refers to such input values representing haplotype alleles as haplotype-allele-index data for the haplotype matrix. Since haplotype-allele-index data for a haplotype matrix having millions, billions, or trillions of cells can consume many gigabytes of memory, the haplotype allele index data taxes the bandwidth of a high-speed bus for a configurable processor, such as a peripheral component interconnect express (PCIe), or other interfaces that connect a processor card to other hardware within a computing device.
[0150] To conserve bandwidth on PCIe or other interfaces, as shown in FIG. 6, the accelerated genotype attribution system 106 stores haplotype-allele-index data 602a for the haplotype matrix on a memory device 600. In some cases, the accelerated genotype attribution system 106 stores the haplotype-allele-index data 602a on on-chip DRAM, SRAM, or other suitable memory. Since the haplotype-allele-index data 602a is readily accessible, the accelerated genotype attribution system 106 can access the haplotype-allele-index data 602a and transfer it from the memory device 600 to the configurable processor 604 to execute a path for determining intermediate allele likelihoods across the haplotype matrix. For example, in some embodiments, the accelerated genotype attribution system 106 uses the configurable processor 604 to access the haplotype-allele-index data 602a for the haplotype matrix from the memory device 600 to generate allele likelihoods for the genotype attribution model.
[0151] The haplotype-allele-index data 602a (or "S" bit data) is in the same format as the haploid or diploid genotype attribution, so the accelerated genotype attribution system 106 can store and access the haplotype-allele-index data 602a for the hidden Markov haploid or diploid genotype attribution model. Thus, the accelerated genotype attribution system 106 can use the configurable processor 604 to access the haplotype-allele-index data 602a for the haplotype matrix from the memory device 600 to generate allele likelihoods using either a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model. When input into the haplotype matrix for the path, FIG. 6 shows the input data as haplotype-allele-index data 602b on the matrix being analyzed by the configurable processor 604.
[0152] To execute approximately 40,000 HMM computational tasks for a single processor thread in about 60 seconds, in some embodiments, the accelerated genotype attribution system 106 requires a PCIe throughput of about 10 gigabytes per second, along with a margin of 6 gigabytes per second available during the pass. By storing the haplotype-genotype-index data 602a of the haplotype matrix on on-chip DRAM (or other on-chip memory) and accessing it therefrom, in some embodiments, the accelerated genotype attribution system 106 saves more than 4 gigabytes per second of PCIe bandwidth.
[0153] As described above, in some embodiments, the accelerated genotype attribution system 106 includes and uses an architecture customized to execute a genotype attribution model such as GLIMPSE. According to one or more embodiments, FIG. 7 shows an accelerated computing engine 700 that includes various customized engines and memory devices for determining allele likelihoods 722 using a genotype attribution model. The following paragraphs describe the various memory devices and operations used to determine allele likelihoods 722. The accelerated computing engine 700 shown in FIG. 7 represents memory devices and engines for haploid HMM computations, although a similar accelerated computing engine may be used for diploid HMM computations.
[0154] As shown in FIG. 7, for example, the accelerated computing engine 700 includes an alpha column memory 704a and a beta column memory 704b. In some embodiments, the alpha column memory 704a and the beta column memory 704b store the pre-normalized intermediate allele likelihoods for the alpha and beta paths, respectively. In particular, in some implementations, the alpha column memory 704a and the beta column memory 704b store one column of pre-normalized alpha values (e.g., A[m][k] values) and one column of pre-normalized beta values (e.g., B[m][k] values), respectively. With respect to the format, the alpha column memory 704a and the beta column memory 704b are each K×Z ABwideValues encoded by bits, i.e., K rows representing haplotypes with a Z-bit width for the stored pre-normalized alpha or beta values, can be stored.
[0155] As further shown in FIG. 7, the acceleration computing engine 700 includes a haplotype - allele - index memory 708. The haplotype - allele - index memory 708 stores haplotype - allele - index data (or "S" - bit data) including input values representing haplotype alleles for each cell of the haplotype matrix. The present disclosure describes the above - mentioned haplotype - allele - index data with respect to FIGS. 2B and 6. In terms of format, the haplotype - allele - index memory 708 can store values or bits of haplotype - allele - index data organized as M×K bits, i.e., M columns representing marker variants and K rows representing haplotypes from a haplotype reference panel. As described above, in some embodiments, the accelerated genotype attribution system 106 transfers haplotype - allele - index data from on - chip DRAM or another memory device to the haplotype - allele - index memory 708 to execute the path of the haplotype matrix.
[0156] In addition to the haplotype - allele - index memory 708, the acceleration computing engine 700 includes a transition coefficient memory 710. The transition coefficient memory 710 stores transition coefficients (e.g., P0 and P1 values) corresponding to the columns or cells of the haplotype matrix. The transition coefficient memory 710 is 2×M×Z p bits, i.e., two sections or blocks of values for M columns representing marker variants with a Z - bit width of the input P0 and P1 values (e.g., one section for P0 values and one section for P1 values), and can store the values of the transition coefficients organized as such. p
[0157] In addition to the transition coefficient memory 710, the acceleration calculation engine 700 includes an allele likelihood factor memory 712. The allele likelihood factor memory 712 stores allele likelihood factors (e.g., Q0 and Q1 values) corresponding to the columns or cells of the haplotype matrix. The allele likelihood factor memory 712 is 2×M×Z Q bits, that is, two sections or blocks of values for M columns representing the marker variants of the Z Q bit width of the input Q0 and Q1 values (e.g., one section of Q0 values and one section of Q1 values) can store the values of the allele likelihood factors organized as.
[0158] As further shown in FIG. 7, the acceleration calculation engine 700 also includes an intermediate allele likelihood memory 716. The intermediate allele likelihood memory 716 stores the intermediate allele likelihoods for the haplotype matrix. For example, in some cases, the intermediate allele likelihood memory 716 stores the alpha and beta values determined over the entire haplotype matrix. Regarding the organization, the intermediate allele likelihood memory 716 is W×K×Z AB bits, that is, it can store the intermediate allele likelihoods organized as W columns within the marker variant group, K rows representing the haplotypes, and Z bit widths for the stored normalized alpha or beta values. Thus, in some embodiments, the intermediate allele likelihood memory 716 organizes the alpha or beta values by the group of marker variants to be compatible with a subset of the path intermediate allele likelihoods that initialize the determination of the intermediate allele likelihoods at the hot start point.
[0159] By using the customized architecture of the acceleration computing engine 700, in some embodiments, the accelerated genotype attribution system 106 determines allele likelihoods 722 for one or more of cells, columns, or haplotype matrices. As shown in FIG. 7, for example, the acceleration computing engine 700 uses SNIFF702a to generate alpha normalization values and uses SNIFF702a to generate beta normalization values. The acceleration computing engine 700 further applies the normalization values from SNIFF702a to normalize the adjacent marker intermediate likelihood values from the column of alpha values stored in the alpha column memory 704a (706a). Similarly, the acceleration computing engine 700 applies the normalization values from SNIFF702b to normalize the adjacent marker intermediate likelihood values from the column of beta values stored in the beta column memory 704b (706b).
[0160] As further shown in FIG. 7, the acceleration computing engine 700 uses the joint engine 714 to determine the intermediate likelihood values for the target cells having the haplotype matrix. Specifically, the acceleration computing engine 700 (i) receives the normalized adjacent marker intermediate allele likelihoods indirectly from the alpha column memory 704a and the beta column memory 704b, (ii) combines the haplotype allele indicators from the haplotype-allele-indicator memory 708, the transition coefficients from the transition coefficient memory 710, and the allele likelihood factors from the allele likelihood factor memory 712 with the normalized adjacent marker intermediate allele likelihoods, and (iii) determines the intermediate allele likelihoods for the target cells stored in the intermediate allele likelihood memory 716. The acceleration computing engine 700 further uses the allele likelihood engine 718 to determine the allele likelihoods 722 for the target cells based on the intermediate allele likelihoods stored in the intermediate allele likelihood memory 716.
[0161] As further shown in FIG. 7, in some embodiments, the acceleration computing engine 700 receives from the memory device an intermediate allele likelihood subset 720a corresponding to the marker variant group. Consistent with the above disclosure, in some cases, the acceleration computing engine 700 uses the intermediate allele likelihood subset to initialize allele likelihood determination in the group of marker variants, thereby reproducing the first-pass intermediate allele likelihood. As further shown, the acceleration computing engine 700 also performs a sacrificial first pass to determine an intermediate allele likelihood subset 720b corresponding to the marker variant group, which is stored in the memory device and can be accessed later to initialize allele likelihood determination in the corresponding group of marker variants.
[0162] In fact, in some cases, the accelerated genotype attribution system 106 can use the acceleration computing engine 700 to determine and access an intermediate allele likelihood subset as described above with respect to FIGS. 4A-4B. Further, in some embodiments, the accelerated genotype attribution system 106 uses the acceleration computing engine 700 to determine a single-pass simultaneous multiplication operation, determine and use a running total of a subset of intermediate allele likelihoods, or perform other embodiments as described above with respect to FIGS. 3A-3B and FIGS. 5A-5B.
[0163] In addition to the acceleration computing engine as part of the customized architecture, in some embodiments, the accelerated genotype attribution system 106 queues and distributes HMM computing tasks to the acceleration computing engines within the cluster of acceleration computing engines, and includes a data flow engine that can manage data communication with a central processing unit (CPU), memory, and the acceleration computing engines. According to one or more embodiments, FIG. 8 shows a configurable processor board 800 including a data flow engine 802 for executing a genotype attribution model, a cluster of acceleration computing engines 804, and an on-board memory device 822. As shown in FIG. 8, the data flow engine 802 interacts and interfaces with the cluster of acceleration computing engines 804, the on-board memory device 822, and the CPU to queue, distribute, or otherwise manage data for HMM computing tasks. The following paragraphs describe the interaction and data exchange between the data flow engine 802 and the acceleration computing engine 804a from the cluster of acceleration computing engines 804, but the same interaction and data exchange can be executed by the data flow engine 802 with each of the acceleration computing engines 804b - 804n.
[0164] As shown by FIG. 8, for example, the configurable processor board 800 is part of a local server device (e.g., the local device 110 shown in FIG. 1) or part of a sequencing device (e.g., the sequencing device 102 shown in FIG. 1). As part of such a computing device, in some embodiments, the data flow engine 802 on the configurable processor board 800 includes a PCIe interface for an FPGA and a double data rate (DDR) interface for interfacing with an on-board memory device 822 such as DRAM.
[0165] In addition to, or as part of, functioning as an interface, in some embodiments, data flow engine 802 transmits and receives data between the CPU, on-board memory device 822, and other hardware on the accelerated genotype attribution system 106 to determine intermediate allele likelihoods, allele likelihoods, or other HMM calculations. As part of CPU communication 818, in some embodiments, data flow engine 802 receives data pointers from the CPU and performs genotype attribution for one or more genomic regions of a genomic sample based on previous genotype likelihoods derived from nucleotide fragment reads. As part of memory communication 820, in some cases, data flow engine 802 transmits and receives input or output requests to and from on-board memory device 822 to store or access data for genotype attribution or phasing. Such requests can include, for example, transmitting and receiving a column of intermediate allele likelihoods (e.g., one column of alpha or beta values) or a subset of intermediate allele likelihoods as a hot start point.
[0166] As previously shown, in some embodiments, the accelerated genotype attribution system 106 can exchange a subset of intermediate allele likelihoods as a hot start point between the data flow engine 802 and the on-board memory device 822. For example, in some cases, the accelerated genotype attribution system 106 (i) transmits a subset of first pass intermediate allele likelihoods from the on-board memory device 822 to the data flow engine 802, and (ii) transmits a subset of first pass intermediate allele likelihoods from the data flow engine 802 to the accelerated calculation engine 804a of the cluster of accelerated calculation engines 804 to reproduce the first pass intermediate allele likelihoods based on the subset of first pass intermediate allele likelihoods.
[0167] In addition to CPU communication 818 and memory communication 820, in some embodiments, data flow engine 802 distributes HMM calculation tasks from a cluster of acceleration calculation engines 804 to individual acceleration calculation engines. By way of example, in some cases, the data flow engine assigns a single HMM calculation task for a haplotype matrix of about 50 million cells that results in about 40,000 haplotype calls from a single acceleration calculation engine from a cluster of acceleration calculation engines 804. Other HMM calculation tasks may be larger or smaller than the foregoing example, but in some embodiments, each of the individual HMM calculation tasks includes input and output values for such a haplotype matrix.
[0168] As shown in FIG. 8, for example, data flow engine 802 can send input values 806 for a target sequence or haplotype matrix for genotype attribution to acceleration calculation engine 804a, or receive output values 808 for a target sequence or haplotype matrix such as allele likelihoods or a subset of intermediate allele likelihoods from acceleration calculation engine 804a. Data flow engine 802 can similarly (i) receive a subset 810b of intermediate allele likelihoods as a hot start point from a sacrificial first pass of acceleration calculation engine 804a, or (ii) send a subset 810a of intermediate allele likelihoods as a hot start point to acceleration calculation engine 804a to reproduce a sequence of intermediate allele likelihoods first determined in a sacrificial first pass.
[0169] As examples of input values 806 and output values 808, in some embodiments, accelerated genotype attribution system 106 sends each set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from data flow engine 802 to each acceleration calculation engine of a cluster of acceleration calculation engines 804. Based on each set of input values, each acceleration calculation engine determines a respective set of intermediate allele likelihoods corresponding to a respective subset of marker variants and a subset of haplotypes.
[0170] More specifically, in certain implementations, the accelerated genotype attribution system 106 transmits (i) a first set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from the data flow engine 802 to the acceleration computing engine 804a, and (ii) a second set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from the data flow engine 802 to the acceleration computing engine 804b. Based on the first set of input values, the acceleration computing engine 804a determines a first set of intermediate allele likelihoods corresponding to a first subset of marker variants and a first subset of haplotypes. Similarly, based on the second set of input values, the acceleration computing engine 804b determines a second set of intermediate allele likelihoods corresponding to a second subset of marker variants and a second subset of haplotypes.
[0171] As an example of the intermediate allele likelihood subsets 810a and 810b, in some embodiments, the accelerated genotype attribution system 106 transmits from the data flow engine 802 to the acceleration computing engine 804a a subset of the first path intermediate allele likelihoods for the acceleration computing engine 804a to reproduce the first path intermediate allele likelihoods from the sacrificial path. Similarly, the accelerated genotype attribution system 106 transmits from the data flow engine 802 to the acceleration computing engine 804b an additional subset of the first path intermediate allele likelihoods for the acceleration computing engine 804b to reproduce additional first path intermediate allele likelihoods from an additional sacrificial path.
[0172] In addition to distributing specific data for the HMM calculation task, as further shown in FIG. 8, the data flow engine 802 queues the HMM calculation tasks for individual acceleration calculation engines from the cluster of acceleration calculation engines 804 and performs further data exchange with the on-board memory device 822. As shown in FIG. 8, for example, the data flow engine 802 transmits configuration and control signals 814, such as data indicators regarding the timing and order of the HMM calculation tasks queued for the acceleration calculation engine 804a, to the acceleration calculation engine 804a. Similarly, in some embodiments, the data flow engine 802 receives a status signal 816 regarding the status or completion of a specific HMM calculation task from the acceleration calculation engine 804a. Based on the status signal 816 from the acceleration calculation engine 804a, the data flow engine 802 queues additional HMM calculation tasks for the acceleration calculation engine 804a or reorganizes or reorders other HMM calculation tasks for the acceleration calculation engines 804b - 804n. As part of such HMM calculation tasks, in some embodiments, the data flow engine 802 also receives and responds to DDR input or output requests from the on-board memory device 822.
[0173] As described above, in some embodiments, the accelerated genotype attribution system 106 can execute approximately 40,000 HMM calculation tasks in about 60 seconds, thereby accelerating the processing time by 600 times. The configurable processor board 800 shown in FIG. 8 can be implemented to facilitate such speed. When the accelerated genotype attribution system 106 determines the 1×alpha values and 2×beta values across a haplotype matrix of 2 trillion cells, the accelerated genotype attribution system 106 must determine the equivalents of the values for 6 trillion cells. Given 16 acceleration calculation engines, the customized architecture within the configurable processor board 800 can execute approximately 40,000 HMM calculation tasks in about 60 seconds.
[0174] For illustration purposes, let "L" represent the level of parallelism of a given acceleration calculation engine for calculating "L" alpha and beta values per clock cycle. If a given acceleration calculation engine has a core clock speed of 400 MHz, a single acceleration calculation engine can calculate L cells / cycle × 400 M cycles per second in 60 seconds, which is equivalent to L × 24 billion alpha or beta cells. To calculate the value of 600 million cells with 24 billion cells per single acceleration calculation engine, L (or the level of parallelism) needs to be equal to 16. Therefore, a set of 16 acceleration calculation engines using the architecture within the configurable processor board 800 of FIG. 8 can execute approximately 40,000 HMM calculation tasks in about 60 seconds.
[0175] In some embodiments, the acceleration calculation engine can be part of a larger hardware structure. According to one or more embodiments, FIG. 9 shows a schematic diagram 900 of an acceleration calculation engine core 914 having a surrounding interface and other hardware.
[0176] As shown in FIG. 9, the acceleration calculation engine core 914 includes an input first-in-first-out (FIFO) for receiving data from a card DRAM advanced extensible interface (AXI) interface 902 and from an address read meta FIFO 912. The acceleration calculation engine core 914 also includes an output FIFO for outputting HMM calculated values to the write channel of the card DRAM AXI interface 902. As further shown within the acceleration calculation engine core 914 of FIG. 9, each of the input FIFO and the output FIFO includes a corresponding converter for down-sizing and up-sizing data, respectively.
[0177] On each side of the acceleration computing engine core 914, the schematic diagram 900 includes a buffer 910 and a buffer 916. As part of the buffer 910, the read parameter buffer and the read status buffer transmit or receive data from the block read state machine 920. As further shown in FIG. 9, the read parameter buffer receives data from the input job FIFO 908. As part of the buffer 916, the write parameter buffer and the write status buffer transmit or receive data from the block write state machine 922. Further, the address write meta FIFO 918 transmits and receives data with the block write state machine 922 and, optionally, with the address write channel of the card DRAM AXI interface 902.
[0178] As further shown in FIG. 9, the card DRAM AXI interface 902 includes a plurality of different channels. In particular, the card DRAM AXI interface 902 includes an address read (AR) channel that receives data from the block read state machine 920, a write (W) channel that receives output values from the acceleration computing engine core 914, and an address write (AW) channel that receives data from the block write state machine 922. Further, the card DRAM AXI interface 902 includes a read (R) channel that receives data from the common engine wrapper (CEW) 904 and a write response (B) channel on which response information is signaled for write transactions.
[0179] Finally, as further shown in FIG. 9, the CEW 904 provides access to a job control infrastructure (e.g., configuration and control signals from the data flow engine 802), a card DRAM AXI interface 902, and a host memory (e.g., the on-board memory device 822). Thus, by using the CEW 904, the accelerated genotype attribution system 106 can exchange data with the card DRAM AXI interface 902 and the streaming CEW interface 906. For example, the CEW 904 transmits configuration and control signals to and from the acceleration computing engine core 914.
[0180] Referring now to FIG. 10, this figure shows a flowchart of a series of operations 1000 for determining intermediate allele likelihoods of genomic regions containing haplotype alleles by performing integrated operations on a processor, according to one or more embodiments of the present disclosure. FIG. 10 shows operations according to one embodiment, although alternative embodiments can omit, add, reorder, and / or modify any of the operations shown in FIG. 10. The operations of FIG. 10 can be implemented as part of a method. Alternatively, a non-transitory computer-readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device or system to perform the operations shown in FIG. 10. In still further embodiments, the system comprises at least one processor and a non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the system to perform the operations of FIG. 10.
[0181] As shown in FIG. 10, the operation 1000 includes an operation 1002 of identifying a haplotype reference panel for a genomic region of a genomic sample. In particular, in some embodiments, the operation 1002 includes using a genotype attribution model to identify a haplotype reference panel for a genomic region of a genomic sample. In some cases, the genotype attribution model includes a hidden Markov genotype attribution model.
[0182] As further shown in FIG. 10, operation 1000 includes an operation 1004 of accessing a first allele likelihood factor corresponding to a haplotype allele and a second allele likelihood factor corresponding to the haplotype allele. In particular, in some embodiments, operation 1004 includes accessing, from a memory device, for a marker variant, a first allele likelihood factor corresponding to a haplotype allele from a haplotype reference panel and a second allele likelihood factor corresponding to the haplotype allele. In this regard, in some embodiments, operation 1004 includes accessing, from a memory device, for a marker variant, a first transition recognition allele likelihood factor corresponding to a haplotype allele from a haplotype reference panel and a second transition recognition allele likelihood factor corresponding to the haplotype allele. Further, in some cases, the memory device includes a dynamic random access memory (DRAM), a static random access memory (SRAM), or a cache memory device.
[0183] For example, in some embodiments, accessing, from a memory device, a first allele likelihood factor and a second allele likelihood factor for a marker variant includes accessing, from a memory device, for a marker variant, a first transition recognition allele likelihood factor corresponding to a haplotype allele from a haplotype reference panel and a second transition recognition allele likelihood factor corresponding to the haplotype allele. In some cases, determining the first transition recognition allele likelihood factor includes combining an allele likelihood factor and a transition linear coefficient. For example, in a particular implementation, the first allele likelihood factor includes an allele likelihood factor for a sample reference haplotype allele or a sample alternative haplotype allele, and the second allele likelihood factor includes an allele likelihood factor for the sample reference haplotype allele or the sample alternative haplotype allele.
[0184] In connection with this, in some embodiments, operation 1000 further includes pre-determining a first transition recognition allele likelihood factor and a second transition recognition allele likelihood factor before determining one or more intermediate allele likelihoods corresponding to marker variants as part of a path across the haplotype matrix. Similarly, in some cases, operation 1000 includes pre-determining a first transition recognition allele likelihood factor and a second transition recognition allele likelihood factor before determining one or more intermediate allele likelihoods corresponding to marker variants. For example, in some embodiments, operation 1004 pre-determines the first transition recognition allele likelihood factor by combining an allele likelihood factor for a haplotype allele with a transition constant coefficient for transitioning between haplotypes from a haplotype reference panel, and pre-determines the second transition recognition allele likelihood factor by combining the allele likelihood factor with a transition linear coefficient for transitioning between haplotypes from the haplotype reference panel.
[0185] As further shown in FIG. 10, operation 1000 includes operation 1006 of combining a first allele likelihood factor with an adjacent marker intermediate allele likelihood to generate an adjacent marker factor recognition allele likelihood. In particular, in a specific implementation, operation 1006 includes combining the first allele likelihood factor with an adjacent marker intermediate allele likelihood of a genomic region including a haplotype allele given an adjacent marker variant to generate an adjacent marker factor recognition allele likelihood for the marker variant and the haplotypes from the haplotype reference panel.
[0186] Furthermore, in some cases, operation 1006 includes the configurable processor combining a first transition recognition allele likelihood factor with an adjacent marker intermediate allele likelihood of a genomic region including a haplotype allele with an adjacent marker variant, to generate an adjacent marker transition factor recognition allele likelihood for the marker variant and the haplotype from a marker variant and a haplotype reference panel. For example, in some embodiments, the configurable processor includes an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a coarse grained reconfigurable array (CGRA), or a field programmable gate array (FPGA).
[0187] As a further illustration, in some embodiments, combining the first allele likelihood factor and the adjacent marker intermediate allele likelihood includes multiplying the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood without a further multiplication operation to determine the intermediate allele likelihood. In connection therewith, in certain implementations, combining the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood includes multiplying the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood without a further multiplication operation to determine the intermediate allele likelihood.
[0188] As further shown in FIG. 10, operation 1000 includes operation 1008 of determining an intermediate allele likelihood based on an adjacent marker factor recognition allele likelihood and a second allele likelihood factor. In particular, in certain implementations, operation 1008 includes determining an intermediate allele likelihood of a genomic region including a haplotype allele for a marker variant and a haplotype based on the adjacent marker factor recognition allele likelihood and the second allele likelihood factor. Furthermore, in some cases, operation 1008 includes the configurable processor determining an intermediate allele likelihood of a genomic region including a haplotype allele for a marker variant and a haplotype based on an adjacent marker transition factor recognition allele likelihood and a second transition recognition allele likelihood factor.
[0189] Furthermore, in some cases, determining the intermediate allele likelihood involves determining the intermediate allele likelihood of a genomic region that includes a sample reference haplotype allele or a sample alternative haplotype allele. In this regard, in certain cases, determining the intermediate allele likelihood based on an adjacent marker factor recognition allele likelihood and a second allele likelihood factor includes summing an adjacent marker transition factor recognition allele likelihood and a total adjacent marker transition recognition allele likelihood factor.
[0190] As further shown in FIG. 10, operation 1000 includes operation 1010 of generating an allele likelihood based on an intermediate allele likelihood. In particular, in some implementations, operation 1008 includes generating, for a set of marker variants corresponding to a genomic region, an allele likelihood of a genomic region that includes a haplotype allele from a haplotype reference panel based on the intermediate allele likelihood. Furthermore, in some cases, operation 1010 includes generating, by a configurable processor, for a set of marker variants corresponding to a genomic region, an allele likelihood of a genomic region that includes a haplotype allele from a haplotype reference panel based on the intermediate allele likelihood.
[0191] In addition to, or instead of, operations 1002 - 1010, in certain implementations, operation 1000 further includes transmitting, from a dataflow engine to each respective acceleration computing engine of a cluster of acceleration computing engines, each respective set of input values that includes an allele likelihood factor, a transition coefficient, and a haplotype - allele value; and determining, by each respective acceleration computing engine, for each respective subset of marker variants and each respective subset of haplotypes, each respective set of intermediate allele likelihoods based on each respective set of input values. In some embodiments, the dataflow engine corresponds to a cluster of acceleration computing engines.
[0192] To further illustrate, in some cases, operation 1000 includes transmitting each set of input values from the data flow engine to each acceleration calculation engine, from the data flow engine to the first acceleration calculation engine of the cluster of acceleration calculation engines, a first set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values; transmitting from the data flow engine to the second acceleration calculation engine of the cluster of acceleration calculation engines a second set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values; determining, by the first acceleration calculation engine, a first set of intermediate allele likelihoods corresponding to a first subset of marker variants and a first subset of haplotypes based on the first set of input values; and determining, by the second acceleration calculation engine, a second set of intermediate allele likelihoods corresponding to a second subset of marker variants and a second subset of haplotypes based on the second set of input values, thereby further including determining each set of intermediate allele likelihoods.
[0193] As suggested above, in some cases, operation 1000 further includes accessing a second transition recognition allele likelihood factor as part of the adjacent marker transition recognition allele likelihood factor; and determining an intermediate allele likelihood based on the adjacent marker transition factor recognition allele likelihood and the total adjacent marker transition recognition allele likelihood factor. In this regard, in some implementations, operation 1000 includes pre-determining the total adjacent marker transition recognition allele likelihood factor by combining the allele likelihood factor for the haplotype allele, the transition constant coefficient for the transition between haplotypes from the haplotype reference panel, and the total adjacent marker intermediate allele likelihood for the adjacent marker variant. As further suggested above, in some cases, the allele likelihood factor for the haplotype allele includes a reference allele likelihood factor for the sample reference haplotype allele or an alternative allele likelihood factor for the sample alternative haplotype allele.
[0194] Additionally, in certain implementations, operation 1000 further includes determining one or more nucleotide calls for a genomic region from a genomic sample based on the allelic likelihoods of the genomic region and one or more variant nucleotide calls surrounding the genomic region.
[0195] Referring now to FIG. 11, this figure shows a flowchart of a series of operations 1100 for determining and storing an intermediate allelic likelihood subset as a hot start point corresponding to a marker variant group, and immediately generating a set of intermediate allelic likelihoods for a set of marker variants by using the intermediate allelic likelihood subset, in accordance with one or more embodiments of the present disclosure. FIG. 11 shows operations according to one embodiment, although alternative embodiments can omit, add, reorder, and / or modify any of the operations shown in FIG. 11. The operations of FIG. 11 can be implemented as part of a method. Alternatively, a non-transitory computer-readable storage medium can include instructions that, when executed by one or more processors, cause a computing device or system to perform the operations shown in FIG. 11. In yet further embodiments, a system can include at least one processor and a non-transitory computer-readable medium that includes instructions that, when executed by the one or more processors, cause the system to perform the operations of FIG. 11.
[0196] As shown in FIG. 11, operation 1100 includes an operation 1102 of determining a first-pass intermediate allele likelihood. In particular, in some embodiments, operation 1102 includes determining a first-pass intermediate allele likelihood of a genomic region from a genomic sample including haplotype alleles corresponding to a set of haplotypes given a set of marker variants by performing a first pass. Further, in some cases, operation 1102 includes determining a first-pass intermediate allele likelihood of a genomic region from a genomic sample including haplotype alleles corresponding to a set of haplotypes given a set of marker variants using a configurable processor that performs the first pass. In some cases, the configurable processor includes an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a coarse grained reconfigurable array (CGRA), or a field programmable gate array (FPGA).
[0197] As further shown in FIG. 11, operation 1100 includes an operation 1104 of storing a subset of the first-pass intermediate allele likelihoods. In particular, in some embodiments, operation 1104 includes storing, on a memory device, a subset of the first-pass intermediate allele likelihoods corresponding to a subset of marker variants for a group of marker variants. Further, in some cases, operation 1104 includes storing a subset of the first-pass intermediate allele likelihoods corresponding to a subset of marker variants for a group of marker variants. In some cases, the memory device includes a dynamic random access memory (DRAM), a static random access memory (SRAM), or a cache memory device.
[0198] As further shown in FIG. 11, operation 1100 includes operation 1106 of reproducing the first-pass intermediate allele likelihoods based on a stored subset of the first-pass intermediate allele likelihoods. In particular, in certain implementations, operation 1106 includes reproducing the first-pass intermediate allele likelihoods by using the stored subset of the first-pass intermediate allele likelihoods to initialize allele likelihood determination in a group of marker variants. Further, in some embodiments, operation 1106 includes using a configurable processor to reproduce the first-pass intermediate allele likelihoods by using the stored subset of the first-pass intermediate allele likelihoods to initialize allele likelihood determination in a group of marker variants.
[0199] In connection therewith, in some cases, initializing allele likelihood determination in a group of marker variants by using a stored subset of the first-pass intermediate allele likelihoods includes determining a first subset of the first-pass intermediate allele likelihoods for a first group of marker variants based on a first stored column of the first-pass intermediate allele likelihoods for an initial marker variant from the first group of marker variants, and determining a second subset of the first-pass intermediate allele likelihoods for a second group of marker variants based on a second stored column of the first-pass intermediate allele likelihoods for an initial marker variant from the second group of marker variants.
[0200] In connection therewith, in some cases, operation 1100 includes storing a subset of the first-pass intermediate allele likelihoods by storing the subset of the first-pass intermediate allele likelihoods in a dynamic random access memory (DRAM), and initializing allele likelihood determination in a group of marker variants by using the stored subset of the first-pass intermediate allele likelihoods includes accessing the stored subset of the first-pass intermediate allele likelihoods from the DRAM.
[0201] As further shown in FIG. 11, operation 1100 includes an operation 1108 of determining second-pass intermediate allele likelihoods. In particular, in certain implementations, operation 1108 includes determining the second-pass intermediate allele likelihoods of a genomic region including haplotype alleles corresponding to a set of haplotypes given a set of marker variants by executing a second pass. Further, in some cases, operation 1108 includes determining the second-pass intermediate allele likelihoods of a genomic region including haplotype alleles corresponding to a set of haplotypes given a set of marker variants using a configurable processor that executes the second pass.
[0202] As suggested above, in some cases, operation 1100 includes determining first-pass intermediate allele likelihoods, which includes determining reverse intermediate allele likelihoods of a genomic region including haplotype alleles using a reverse pass, and determining second-pass intermediate allele likelihoods, which includes determining forward intermediate allele likelihoods of a genomic region including haplotype alleles using a forward pass.
[0203] As further shown in FIG. 11, operation 1100 includes an operation 1110 of generating allele likelihoods based on the reproduced first-pass intermediate allele likelihoods and second-pass intermediate allele likelihoods. In particular, in certain implementations, operation 1110 includes generating allele likelihoods of a genomic region including haplotype alleles based on the reproduced first-pass intermediate allele likelihoods and second-pass intermediate allele likelihoods. Further, in some embodiments, operation 1110 includes generating allele likelihoods of a genomic region including haplotype alleles based on the reproduced first-pass intermediate allele likelihoods and second-pass intermediate allele likelihoods using an output engine.
[0204] By way of example, in some embodiments, generating allele likelihoods based on the reproduced first-pass intermediate allele likelihoods and the second-pass intermediate allele likelihoods includes determining a total first-pass intermediate allele likelihood for a set of marker variants based on the reproduced first-pass intermediate allele likelihoods, determining a total second-pass intermediate allele likelihood for the set of marker variants based on the second-pass intermediate allele likelihoods, and determining allele likelihoods based on the total first-pass intermediate allele likelihood and the total second-pass intermediate allele likelihood.
[0205] In addition to, or instead of, operations 1102-1110, in certain implementations, operation 1000 further includes storing haplotype-allele-metric data in a haplotype-allele-metric memory, storing transition coefficients in a transition coefficient memory, and storing allele likelihood factors in an allele likelihood factor memory. Further, in some embodiments, operation 1000 includes determining intermediate allele likelihood values using a joint engine.
[0206] As suggested above, in some cases, operation 1100 further includes transmitting a respective set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from a dataflow engine to each of the acceleration computing engines of a cluster of acceleration computing engines, and determining, by each of the acceleration computing engines, a respective set of intermediate allele likelihoods corresponding to a respective subset of marker variants and a respective subset of haplotypes based on the respective set of input values. In some embodiments, the dataflow engine corresponds to a cluster of acceleration computing engines.
[0207] To further illustrate, in some cases, operation 1100 includes transmitting each set of input values from the data flow engine to each acceleration calculation engine, from the data flow engine to the first acceleration calculation engine of the cluster of acceleration calculation engines, a first set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values; transmitting from the data flow engine to the second acceleration calculation engine of the cluster of acceleration calculation engines, a second set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values; determining, by the first acceleration calculation engine, a first set of intermediate allele likelihoods corresponding to a first subset of marker variants and a first subset of haplotypes based on the first set of input values; and determining, by the second acceleration calculation engine, a second set of intermediate allele likelihoods corresponding to a second subset of marker variants and a second subset of haplotypes based on the second set of input values, thereby determining each set of intermediate allele likelihoods.
[0208] As further suggested above, in some cases, operation 1100 includes transmitting from the data flow engine to the first acceleration calculation engine of the cluster of acceleration calculation engines, a subset of the first path intermediate allele likelihoods for the first acceleration calculation engine to reproduce the first path intermediate allele likelihoods; and transmitting from the data flow engine to the second acceleration calculation engine of the cluster of acceleration calculation engines, an additional subset of the first path intermediate allele likelihoods for the second acceleration calculation engine to reproduce additional first path intermediate allele likelihoods.
[0209] Furthermore, in certain implementations, operation 1100 includes sending a subset of the first-pass intermediate allele likelihoods from the memory device to the data flow engine and sending the subset of the first-pass intermediate allele likelihoods from the data flow engine to the acceleration calculation engine to reproduce the first-pass intermediate allele likelihoods based on the subset of the first-pass intermediate allele likelihoods. Further, in some cases, operation 1100 includes storing haplotype-allele-index data for the haplotype matrix on the memory device and accessing the haplotype-allele-index data for the haplotype matrix from the memory device to generate allele likelihoods using a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model.
[0210] As suggested above, in some cases, operation 1100 includes determining one or more nucleotide calls for a genomic region from a genomic sample based on the allele likelihoods for the genomic region and one or more variant nucleotide calls surrounding the genomic region.
[0211] As further suggested above, in certain implementations, operation 1100 includes storing haplotype-allele-index data for the haplotype matrix on the dynamic random access memory (DRAM) and accessing the haplotype-allele-index data for the haplotype matrix by a configurable processor from the DRAM to generate allele likelihoods using a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model.
[0212] In addition to, or instead of, in certain embodiments, operation 1100 includes, for an adjacent marker variant, determining a running sum of a first subset of intermediate allele likelihoods of a genomic region that includes haplotype alleles from one or more haplotypes of a haplotype reference panel; for the adjacent marker variant, determining a running sum of a second subset of intermediate allele likelihoods of a genomic region that includes haplotype alleles of a second type from one or more haplotypes; and for the marker variant, determining a sum of intermediate allele likelihoods of a genomic region that includes haplotype alleles from a haplotype of the haplotype reference panel, based on the running sum of the first subset of intermediate allele likelihoods and the running sum of the second subset of intermediate allele likelihoods.
[0213] Referring now to FIG. 12, this figure shows a flowchart of a series of operations 1200 for determining a running sum of intermediate allele likelihoods of a genomic region showing haplotype alleles for one or more haplotypes given one marker variant, and using the running sum as a running input to determine individual intermediate allele likelihoods of a genomic region showing haplotype alleles for a haplotype given another marker variant, in accordance with one or more embodiments of the present disclosure. FIG. 12 shows operations according to one embodiment, although alternative embodiments can omit, add, reorder, and / or modify any of the operations shown in FIG. 12. The operations of FIG. 12 can be implemented as part of a method. Alternatively, a non-transitory computer-readable storage medium can include instructions that, when executed by one or more processors, cause a computing device or system to perform the operations shown in FIG. 12. In yet further embodiments, a system can include at least one processor and a non-transitory computer-readable medium that includes instructions that, when executed by the one or more processors, cause the system to perform the operations of FIG. 12.
[0214] As shown in FIG. 12, operation 1200 includes an operation 1202 of identifying a haplotype reference panel for a genomic region of a genomic sample. In particular, in some embodiments, operation 1202 includes identifying a haplotype reference panel for a genomic region of a genomic sample using a genotype attribution model.
[0215] As further shown in FIG. 12, operation 1200 includes an operation 1204 of determining a running total of a first subset of intermediate allele likelihoods for adjacent marker variants. In particular, in some embodiments, operation 1204 includes determining a running total of a first subset of intermediate allele likelihoods for a genomic region that includes first-type haplotype alleles from one or more haplotypes of a haplotype reference panel for adjacent marker variants.
[0216] As further shown in FIG. 12, operation 1200 includes an operation 1206 of determining a running total of a second subset of intermediate allele likelihoods for adjacent marker variants. In particular, in certain implementations, operation 1206 includes determining a running total of a second subset of intermediate allele likelihoods for a genomic region that includes second-type haplotype alleles from one or more haplotypes for adjacent marker variants.
[0217] As described above, in some embodiments, the first-type haplotype alleles include sample reference haplotype alleles, and the second-type haplotype alleles include sample alternative haplotype alleles.
[0218] As further shown in FIG. 12, operation 1200 includes an operation 1208 of determining the total intermediate allele likelihood based on the running total of the first subset of intermediate allele likelihoods and the running total of the second subset of intermediate allele likelihoods for a marker variant. In particular, in certain implementations, operation 1208 includes determining the total intermediate allele likelihood of a genomic region including haplotype alleles from a haplotype of a haplotype reference panel based on the running total of the first subset of intermediate allele likelihoods and the running total of the second subset of intermediate allele likelihoods for a marker variant.
[0219] As described above, in some cases, determining the total intermediate allele likelihood includes, by a configurable processor, determining an initial intermediate allele likelihood from the intermediate allele likelihoods for a marker variant based on the intermediate allele likelihoods from the first subset or the second subset of intermediate allele likelihoods, and then summing the adjacent marker intermediate allele likelihoods of a genomic region including haplotype alleles for adjacent marker variants.
[0220] Additionally or alternatively, in certain implementations, determining the total intermediate allele likelihood includes, by a configurable processor, determining an initial intermediate allele likelihood from the intermediate allele likelihoods for a marker variant based on the intermediate allele likelihoods from the first subset or the second subset of intermediate allele likelihoods, and then generating an allele likelihood of a genomic region including haplotype alleles for adjacent marker variants for adjacent marker variants.
[0221] As further shown in FIG. 12, operation 1200 includes an operation 1210 of generating an allele likelihood based on the total intermediate allele likelihood. In particular, in certain implementations, operation 1210 includes generating an allele likelihood of a genomic region including haplotype alleles based on the total intermediate allele likelihood.
[0222] In addition to, or instead of, operations 1202 - 1210, in certain implementations, operation 1000 further includes determining a first transition recognition allele likelihood factor corresponding to a row of first - type haplotype alleles and a second transition recognition allele likelihood factor corresponding to a row of second - type haplotype alleles in advance, and determining a sum of intermediate allele likelihoods further based on the first transition recognition allele likelihood factor corresponding to the row of first - type haplotype alleles and the second transition recognition allele likelihood factor corresponding to the row of second - type haplotype alleles.
[0223] In connection therewith, in some cases, for adjacent marker variants, determining a sum of adjacent marker intermediate allele likelihoods of a genomic region containing haplotype alleles, and for marker variants, determining a sum of intermediate allele likelihoods further based on a combination of the sum of adjacent marker intermediate allele likelihoods, a first transition recognition allele likelihood factor corresponding to a row for first - type haplotype alleles, and a second transition recognition allele likelihood factor corresponding to a row for second - type haplotype alleles.
[0224] As suggested above, in some cases, operation 1200 further includes multiplying a running sum of a first subset of intermediate allele likelihoods by a first transition recognition allele likelihood factor, multiplying a running sum of a second subset of intermediate allele likelihoods by a second transition recognition allele likelihood factor, and for marker variants, determining a sum of intermediate allele likelihoods based on the multiplied running sum of the first subset of intermediate allele likelihoods and the multiplied running sum of the second subset of intermediate allele likelihoods.
[0225] As further suggested above, in some embodiments, operation 1200 may include pre-determining a first transition recognition allele likelihood factor by combining a first allele likelihood factor for a first type of haplotype allele with a transition linear coefficient for transitioning between haplotypes from a haplotype reference panel, and pre-determining a second transition recognition allele likelihood factor by combining a second allele likelihood factor for a second type of haplotype allele with the transition linear coefficient.
[0226] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those in which nucleic acids are attached to fixed positions within an array such that their relative positions do not change and the array is repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleobase type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of a target nucleic acid can be an automated process. Preferred embodiments include sequencing by synthesis (SBS) techniques.
[0227] SBS techniques generally involve the enzymatic extension of a nascent nucleic acid strand by the iterative addition of nucleotides to a template strand. In conventional methods of SBS, single nucleotide monomers can be provided to the target nucleotides in the presence of polymerase at each delivery. However, in the methods described herein, two or more types of nucleotide monomers can be provided to the target nucleic acid in the presence of polymerase during delivery.
[0228] SBS can utilize nucleotide monomers with a terminator part or nucleotide monomers lacking any terminator part. As methods of using nucleotide monomers lacking a terminator, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in more detail below, can be mentioned. In methods of using nucleotide monomers without a terminator, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS technologies that utilize nucleotide monomers with a terminator part, the terminator can be effectively irreversible under the sequencing conditions used as in the case of conventional Sanger sequencing that utilizes dideoxynucleotides, or the terminator can be reversible as in the case of the sequencing method developed by Solexa (now Illumina, Inc.).
[0229] SBS technology can use nucleotide monomers with a label part or nucleotide monomers lacking a label part. Thus, incorporation events can be detected based on properties of the label such as fluorescence of the label, properties of the nucleotide monomer such as molecular weight or charge, by-products of nucleotide incorporation such as the release of pyrophosphate, etc. In embodiments where two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, two or more different labels can be distinguishable under the detection technology used. For example, different nucleotides present in the sequencing reagent can have different labels, and they can be distinguished using an appropriate optical system exemplified by the sequencing method developed by Solexa (now Illumina, Inc.).
[0230] Preferred embodiments include pyrosequencing technology. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9, Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11, Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375), 363, U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568 and U.S. Patent No. 6,274,320, the entire disclosures of which are incorporated herein by reference). In pyrosequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of generated ATP is detected via photons generated by luciferase. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture the chemiluminescent signal generated by incorporating nucleotides into the features of the array. An image can be obtained after treating the array with a specific nucleotide type (e.g., A, T, C, or G). The images obtained after the addition of each nucleotide type are different with respect to which features in the array are detected. These differences in the images reflect the different sequence contents of the features on the array. However, the relative positions of each feature remain unchanged within the image. The images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after processing an array with each different nucleotide type can be processed in the same manner as those exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.
[0231] In another exemplary type of SBS, cycle sequencing is achieved, for example, by stepwise addition of cleavable or photobleachable dye-labeled reversible terminator nucleotides as described in International Publication No. WO 04 / 018497 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference. This approach has been commercialized by Solexa (now Illumina, Inc.) and is also described in International Publication No. WO 91 / 06678 and International Publication No. WO 07 / 123,744, each of which is incorporated herein by reference. The availability of fluorescently labeled terminators whose fluorescence labels can be cleaved allows for efficient cyclic reversible termination (CRT) sequencing. Polymerases can also co-operate to efficiently incorporate and extend from these modified nucleotides.
[0232] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be taken after incorporation of the label into the arrayed nucleic acid features. In certain embodiments, each cycle involves the simultaneous delivery of four different nucleotide types to the array, and each nucleotide type has spectrally distinct labels. Next, four images can be obtained, each using a detection channel selective for one of the four different labels. Alternatively, the different nucleotide types can be added sequentially, and an image of the array can be obtained between each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature is different, different features will be present in, or absent from, different images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the imaging step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removing the label after detection in a particular cycle and prior to subsequent cycles has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are described below.
[0233] In certain embodiments, some or all of the nucleotide monomers can include reversible terminators. In such embodiments, the reversible terminator / cleavable fluor can include a fluor attached to the ribose moiety via a 3'-ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches separate the chemistry of the terminator from the cleavage of the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. describe the development of reversible terminators that use a small amount of 3'-allyl groups to block extension but can be easily de-blocked by treatment with a palladium catalyst for a short time. The fluor is attached to the group via a photocleavable linker that can be easily cleaved by exposure to long wavelength UV light for 30 seconds. Thus, either disulfide reduction or photocleavage can be used as the cleavable linker. Another approach to reversible termination is the use of natural termination following placement of a bulky dye on the dNTP. The presence of the charged bulky dye on the dNTP can act as an effective terminator via steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Patent No. 7,427,673 and U.S. Patent No. 7,057,026, the disclosures of which are incorporated herein by reference in their entirety.
[0234] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Patent No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, International Publication No. WO05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, International Publication No. WO06 / 064199, International Publication No. WO07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are hereby incorporated by reference in their entireties.
[0235] Some embodiments can utilize the detection of four different nucleotides using less than four different labels. For example, SBS can be implemented using the methods and systems described in incorporated reference US Patent Application Publication No. 2013 / 0079232. As a first example, nucleotide type pairs can be detected at the same wavelength, but based on the difference in intensity for one member of the pair, or based on a change (e.g., via chemical modification, photochemical modification, or physical modification) to one member of the pair that causes a distinct signal to appear or disappear as compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types can be detected under certain conditions, while the fourth nucleotide type has no detectable label or is minimally detected under those conditions (e.g., minimal detection by background fluorescence). Incorporating the first three nucleotide types into a nucleic acid can be determined based on the presence of their respective signals, and incorporating the fourth nucleotide type into a nucleic acid can be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while other nucleotide types are detected in one or fewer channels. The three exemplary configurations described above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a first nucleotide type detected in a first channel (e.g., dATP having a label detected in the first channel when excited by a first excitation wavelength), a second nucleotide type detected in a second channel (e.g., dCTP having a label detected in the second channel when excited by a second excitation wavelength), a third nucleotide type detected in both the first and second channels (e.g., dTTP having at least one label detected in both channels when excited by the first and / or second excitation wavelength), and a fourth nucleotide type that is not detected in any channel or lacks a label that is minimally detected (e.g., unlabeled dGTP), in a fluorescence-based SBS method.
[0236] Furthermore, as described in U.S. Patent Application Publication No. 2013 / 0079232, which is incorporated herein by reference, sequencing data can be obtained using a single channel. In such a so-called one-dye sequencing method, the first nucleotide type is labeled, but the label is removed after the first image is generated, and the second nucleotide type is labeled only after the first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.
[0237] Some embodiments can utilize ligation-based sequencing techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence to which the oligonucleotide hybridizes. As with other SBS methods, after processing an array of nucleic acid features with labeled sequencing reagents, an image can be obtained. Each image shows nucleic acid features incorporating a particular type of label. Because the sequence content of each feature is different, different features may or may not be present in different images, but the relative positions of the features remain unchanged within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent No. 6,969,488, U.S. Patent No. 6,172,218, and U.S. Patent No. 6,306,597, the disclosures of which are incorporated herein by reference in their entirety.
[0238] Some embodiments can utilize nanopore sequencing (Deamer, D. W. & Akeson, M., "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147 - 151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35: 817 - 825 (2002); Li, J., M. Gershow, D. Stein, E. Brardin, and J. A. Golovchenko, "DNA molecules and configurations in a solid - state nanopore microscope" Nat. Mater. 2: 611 - 615 (2003), the disclosures of which are incorporated herein by reference in their entirety). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore such as α - hemolysin or a biomembrane protein. When the target nucleic acid passes through the nanopore, each base pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Patent No. 7,001,792; Soni, G. V. & Meller, "A. Progress toward ultrafast DNA sequencing using solid - state nanopores." Clin. Chem. 53, 1996 - 2001 (2007); Healy, K., "Nanopore - based single - molecule DNA analysis." Nanomed. 2, 459 - 481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R., "A single - molecule nanopore device detects DNA polymerase activity with single - nucleotide resolution." J. Am Chem. Soc. 130, 818 - 820 (2008), the disclosures of which are incorporated herein by reference in their entirety).Data obtained from nanopore array determination can be stored, processed, and analyzed as described herein. Specifically, the data can be processed as an image according to the exemplary processing of the optical and other images described herein.
[0239] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected, for example, via fluorescence resonance energy transfer (FRET) interactions between a fluorophore-containing polymerase and a γ-phosphate labeled nucleotide as described in, for example, U.S. Patent No. 7,329,492 and U.S. Patent No. 7,211,414, each of which is incorporated herein by reference, or nucleotide incorporation can be detected, for example, using zero-mode waveguides as described in U.S. Patent No. 7,315,019, which is incorporated herein by reference, and fluorescent nucleotide analogs and engineered polymerases as described in, for example, U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference. Illumination can be restricted to zeptoliter-scale volumes surrounding surface-tethered polymerases such that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, M.J. et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science, 299, 682 - 686 (2003), Lundquist, P.M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026 - 1028 (2008), Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176 - 1181 (2008), the disclosures of which are incorporated herein by reference in their entireties).Images obtained from such methods can be stored, processed, and analyzed as described herein.
[0240] Some SBS embodiments include the detection of protons released upon incorporation of nucleotides into an extension product. For example, sequencing based on the detection of released protons may use an electrical detector and related technology commercially available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082 (A1), 2009 / 0127589 (A1), 2010 / 0137143 (A1), or 2010 / 0282617 (A1), each of which is incorporated herein by reference. The methods described herein for amplifying a target nucleic acid using kinetic exclusion can be readily applied to the substrate used for detecting protons. More specifically, the methods described herein can be used to generate a clonal population of amplicons used for detecting protons.
[0241] The above-described SBS method can be advantageously implemented in a multiplex format such that multiple different target nucleic acids are manipulated simultaneously. In certain embodiments, the different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This enables convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a multiplex manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in array format. In array format, the target nucleic acids can typically be bound to the surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent bonding, binding to beads or other particles, or binding to a polymerase or other molecule bound to the surface. The array can contain a single copy of the target nucleic acid at each site (also referred to as a feature), or multiple copies having the same sequence can be present at each site or feature. The multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, which are described in more detail below.
[0242] The methods described herein can use arrays having features at various densities, for example, including at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or more.
[0243] The advantage of the methods described herein is that they provide for the rapid and efficient detection of multiple target nucleic acids in parallel. Accordingly, the present disclosure provides an integrated system that can prepare and detect nucleic acids using techniques known in the art such as those exemplified above. Thus, the integrated system of the present disclosure can include fluid components that can deliver amplification reagents and / or sequencing reagents to one or more immobilized DNA fragments, and the system can include components such as pumps, valves, reservoirs, fluid lines, and the like. A flow cell can be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent Application Publication No. 2010 / 0111768 (A1) and U.S. Patent Application No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluid components of the integrated system can be used in amplification methods and detection methods. Taking an embodiment of nucleic acid sequencing as an example, one or more of the fluid components of the integrated system can be used for delivering sequencing reagents in the amplification methods described herein and the sequencing methods exemplified above. Alternatively, the integrated system can include separate fluid systems for performing the amplification method and for performing the detection method. Examples of integrated sequencing systems that can create amplified nucleic acids and also determine the sequence of the nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA), and the apparatus described in U.S. Patent Application No. 13 / 273,666, which is incorporated herein by reference.
[0244] The above-described array determination system determines the sequence of nucleic acid polymers present in a sample received by the array determination apparatus. As defined herein, "sample" and its derivatives are used in the broadest sense and include any sample, culture, etc. that is suspected of containing the target. In some embodiments, the sample contains nucleic acids in the form of DNA, RNA, PNA, LNA, chimeras, or hybrids. The sample can include any biological sample, clinical sample, surgical sample, agricultural sample, atmospheric sample, or water sample that contains one or more nucleic acids. The term also includes any isolated nucleic acid sample, e.g., genomic DNA, freshly frozen or formalin-fixed paraffin-embedded nucleic acid samples. The sample can include nucleic acid samples from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from non-genetically related members, nucleic acid samples (matched) from a single individual such as tumor samples and normal tissue samples, or a sample from a single source that includes two different forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or can be derived from the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of the nucleic acid material can include nucleic acids obtained from a neonate, such as is typically used in neonatal screening.
[0245] The nucleic acid sample can contain high molecular weight substances such as genomic DNA (gDNA). The sample can contain low molecular weight substances such as nucleic acid molecules obtained from FFPE or stored DNA samples. In another embodiment, the low molecular weight substance contains enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can contain nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical excisions, and other clinical or laboratory-obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, the sample can contain nucleic acid molecules obtained from animals such as human or mammalian sources. In another embodiment, the sample can contain nucleic acid molecules obtained from non-mammalian sources such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecule can be a preserved or extinct sample or species.
[0246] Furthermore, the methods and compositions disclosed herein can be useful for amplifying nucleic acid samples having low-quality nucleic acid molecules such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory associated with a forensic investigation, or forensic samples obtained by law enforcement agencies, one or more military services, or any such personnel. The nucleic acid sample can be, for example, a purified sample or crude DNA containing a lysate derived from a buccal swab, paper, cloth, or other substrate that can be impregnated with saliva, blood, or other body fluids. Thus, in some embodiments, the nucleic acid sample can contain a small amount of DNA or a fragmented portion of DNA, such as genomic DNA. In some embodiments, the target sequence can be present in one or more body fluids including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence can be obtained from the hair, skin, tissue samples, autopsy, or remains of a victim. In some embodiments, the nucleic acid containing one or more target sequences can be obtained from a deceased animal or human. In some embodiments, the target sequence can contain nucleic acids obtained from non-human DNA such as microbial, plant, or entomological DNA. In some embodiments, the target sequence or amplified target sequence is for the purpose of human identification. In some embodiments, the present disclosure generally relates to methods for identifying the characteristics of forensic samples. In some embodiments, the present disclosure generally relates to methods of human identification using one or more of the target-specific primers disclosed herein, or one or more target-specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identification sample containing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein, or using the primer criteria outlined herein.
[0247] The components of the sequencing system 112 or the accelerated genotype attribution system 106 can include software, hardware, or both. For example, the components of the sequencing system 112 or the accelerated genotype attribution system 106 can be stored in a computer-readable storage medium and include one or more instructions executable by a processor of one or more computing devices (e.g., the client device 116). When executed by one or more processors, the computer-executable instructions of the sequencing system 112 or the accelerated genotype attribution system 106 can cause the computing device to perform the bubble detection methods described herein. Alternatively, the components of the sequencing system 112 or the accelerated genotype attribution system 106 can include hardware such as a dedicated processing device for performing a particular function or group of functions. Further, or alternatively, the components of the sequencing system 112 or the accelerated genotype attribution system 106 can include a combination of computer-executable instructions and hardware.
[0248] Furthermore, the components of the accelerated genotype attribution system 106 that perform the functions described herein with respect to the accelerated genotype attribution system 106 may be implemented, for example, as part of a stand-alone application, as a module of an application, as a plug-in of an application, as a library function that can be called by another application, and / or as a cloud computing model. Thus, the components of the accelerated genotype attribution system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the accelerated genotype attribution system 106 may be implemented in any application that provides a sequencing service, including but not limited to Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumina", "BaseSpace", "DRAGEN", and "TruSight" are registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0249] Embodiments of the present disclosure may include, or utilize, a special purpose or general purpose computer including, for example, one or more processors and system memory, as will be discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be embodied in a non-transitory computer-readable medium and at least partially implemented as executable instructions by one or more computing devices (e.g., any of the media content access devices described herein). Generally, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., memory, etc.), executes those instructions, thereby performing one or more processes including one or more of the processes described herein.
[0250] A computer-readable medium can be any available medium that can be accessed by a general-purpose computer system or a dedicated computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that conveys computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure can include at least two distinctly different types of computer-readable media, namely non-transitory computer-readable storage media (devices) and transmission media.
[0251] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs, e.g., based on RAM), flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or dedicated computer.
[0252] "Network" is defined as one or more data links that enable the transfer of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer via a network or another communication connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly recognizes the connection as a transmission medium. Transmission media can be used to convey desired program code means in the form of computer-executable instructions or data structures and can include networks and / or data links that can be accessed by a general-purpose or dedicated computer. Combinations of the above should also be included within the scope of computer-readable media.
[0253] Furthermore, when reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be automatically transferred from a transmission medium to a non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received via a network or data link are buffered in RAM within a network interface module (e.g., NIC), and then can ultimately be transferred to the computer system RAM and / or a less volatile computer storage medium (device) in the computer system. Thus, it should be understood that a non-transitory computer-readable storage medium (device) can be included in computer system components that also (or further primarily) utilize a transmission medium.
[0254] Computer-executable instructions, for example, when executed by a processor, include instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or group of functions. In some embodiments, the computer-executable instructions are executed on a general-purpose computer, converting the general-purpose computer into a special-purpose computer implementing the elements of the present disclosure. The computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in a language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts above. Rather, the described features and acts are disclosed as exemplary forms for implementing the claims.
[0255] Those skilled in the art will understand that the present disclosure can be implemented in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, cellular telephones, PDAs, tablets, pagers, routers, switches, and the like. The present disclosure can also be implemented in a distributed system environment where local and remote computer systems, linked via a network (either by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links), both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0256] Embodiments of the present disclosure can also be implemented in a cloud computing environment. As used herein, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the marketplace to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization, exposed with low management effort or service provider interaction, and then scaled accordingly.
[0257] The cloud computing model can be composed of various characteristics such as, for example, on-demand self-service, wide area network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can also expose various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). The cloud computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and the claims, a "cloud computing environment" is an environment in which cloud computing is adopted.
[0258] FIG. 13 shows a block diagram of a computing device 1300 that can be configured to perform one or more of the above processes. It will be understood that one or more computing devices, such as computing device 1300, can implement the accelerated genotype attribution system 106. As shown by FIG. 13, the computing device 1300 can include a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310, which can be communicatively coupled by a communication infrastructure 1312. In certain embodiments, the computing device 1300 can include fewer or more components than those shown in FIG. 13. The following paragraphs describe the components of the computing device 1300 shown in FIG. 13 in more detail.
[0259] In one or more embodiments, processor 1302 includes hardware for executing instructions, such as instructions that make up a computer program. By way of example and not limitation, to execute instructions for dynamically modifying a workflow, processor 1302 can retrieve (or fetch) instructions from internal registers, internal caches, memory 1304, or storage device 1306, decode them, and execute them. Memory 1304 may be volatile or non-volatile memory used to store data, metadata, and programs for execution by the processor. Storage device 1306 includes storage such as a hard disk, flash disk drive, or other digital storage device for storing data or instructions for implementing the methods described herein.
[0260] I / O interface 1308 enables a user to provide input to, receive output from, transfer data to, or receive data from computing device 1300. I / O interface 1308 can include a mouse, keypad or keyboard, touch screen, camera, optical scanner, network interface, modem, other known I / O devices, or a combination of such I / O interfaces. I / O interface 1308 can include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, I / O interface 1308 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be useful in a particular implementation.
[0261] The communication interface 1310 can include hardware, software, or both. In any case, the communication interface 1310 can provide one or more interfaces for communication (such as packet-based communication) between the computing device 1300 and one or more other computing devices or a network. By way of non-limiting example, the communication interface 1310 can include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as WI-FI.
[0262] Additionally, the communication interface 1310 can facilitate communication with various types of wired or wireless networks. The communication interface 1310 can also facilitate communication using various communication protocols. The communication infrastructure 1312 can also include hardware, software, or both that couples the components of the computing device 1300 to each other. For example, the communication interface 1310 can enable multiple computing devices connected by a particular infrastructure to communicate with each other to implement one or more aspects of the processes described herein using one or more networks and / or protocols. By way of illustration, the sequencing process can enable multiple devices (such as client devices, sequencing devices, and server devices) to exchange information such as sequencing data and error notifications.
[0263] In the foregoing specification, the present disclosure has been described with reference to its specific exemplary embodiments. Various embodiments and aspects of the present disclosure are described with reference to the details discussed herein, and the accompanying drawings illustrate the various embodiments. The above description and drawings are examples of the present disclosure and should not be construed as limiting the present disclosure. Numerous specific details are described to provide a complete understanding of the various embodiments of the present disclosure.
[0264] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be implemented with fewer or more steps / operations, or the steps / operations may be performed in a different order. Additionally, the steps / operations described herein may be repeated or performed in parallel with each other, or in parallel with different occurrences of the same or similar steps / operations. Accordingly, the scope of the present application is shown not by the foregoing description but by the appended claims. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
Claim 1 A method comprising: identifying a haplotype reference panel for a genomic region of a genomic sample using a genotype attribution model; accessing, from a memory device, for a marker variant, a first transition recognition allele likelihood factor corresponding to a haplotype allele from the haplotype reference panel and a second transition recognition allele likelihood factor corresponding to the haplotype allele; combining, by a configurable processor, the first transition recognition allele likelihood factor with an adjacent marker intermediate allele likelihood of the genomic region including the haplotype allele provided with an adjacent marker variant to generate an adjacent marker transition factor recognition allele likelihood for the marker variant and the haplotype from the haplotype reference panel; determining, by the configurable processor, an intermediate allele likelihood of the genomic region including the haplotype allele based on the adjacent marker transition factor recognition allele likelihood and the second transition recognition allele likelihood factor for the marker variant and the haplotype; generating, by the configurable processor, an allele likelihood of the genomic region including the haplotype allele from the haplotype reference panel based on the intermediate allele likelihood for a set of marker variants corresponding to the genomic region. Claim 2 The method of claim 1, further comprising pre-determining the first transition recognition allele likelihood factor and the second transition recognition allele likelihood factor before determining one or more intermediate allele likelihoods corresponding to the marker variant. Claim 3 Pre-determining the first transition recognition allele likelihood factor includes combining an allele likelihood factor for the haplotype allele with a transition constant coefficient for a transition between haplotypes from the haplotype reference panel; Pre-determining the second transition recognition allele likelihood factor includes combining the allele likelihood factor with a transition linear coefficient for a transition between haplotypes from the haplotype reference panel. Claim 4 Combining the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood includes multiplying the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood without a further multiplication operation for determining the intermediate allele likelihood, the method according to claim 1.
5. Accessing the second transition recognition allele likelihood factor as part of the total adjacent marker transition recognition allele likelihood factor, and further comprising determining the intermediate allele likelihood based on the adjacent marker transition factor recognition allele likelihood and the total adjacent marker transition recognition allele likelihood factor, the method according to claim 1.
6. Further comprising pre-determining the total adjacent marker transition recognition allele likelihood factor by combining the allele likelihood factor for the haplotype allele, the transition constant coefficient for the transition between haplotypes from the haplotype reference panel, and the total adjacent marker intermediate allele likelihood for the adjacent marker variant, the method according to claim 5.
7. The method according to claim 6, wherein the allele likelihood factor for the haplotype allele includes a reference allele likelihood factor for a sample reference haplotype allele or an alternative allele likelihood factor for a sample alternative haplotype allele.
8. The method according to claim 1, wherein the genotype attribution model includes a hidden Markov genotype attribution model.
9. The method according to claim 1, wherein the configurable processor includes an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a coarse-grained reconfigurable array (CGRA), or a field programmable gate array (FPGA).
10. The method according to claim 1, wherein the memory device includes a dynamic random access memory (DRAM), a static random access memory (SRAM), or a cache memory device.
11. A system, comprising at least one processor, a memory device, and a non-transitory computer-readable medium, wherein when the non-transitory computer-readable medium is executed by the at least one processor, the system is caused to utilize a genotype attribution model to identify a haplotype reference panel for a genomic region of a genomic sample, Access, from the memory device, for a marker variant, to a first allele likelihood factor corresponding to a haplotype allele from the haplotype reference panel and a second allele likelihood factor corresponding to the haplotype allele. Combine the first allele likelihood factor with an adjacent marker intermediate allele likelihood of the genomic region including the haplotype allele given an adjacent marker variant, to generate an adjacent marker factor recognition allele likelihood for the marker variant and the haplotype from the haplotype reference panel. For the marker variant and the haplotype, determine an intermediate allele likelihood of the genomic region including the haplotype allele, based on the adjacent marker factor recognition allele likelihood and the second allele likelihood factor. A system including instructions to generate an allele likelihood of the genomic region including a haplotype allele from the haplotype reference panel, based on the intermediate allele likelihood, for a set of marker variants corresponding to the genomic region.
12. The system according to claim 11, further including instructions that, when executed by the at least one processor, cause the system to access, from the memory device, for the marker variant, a first transition recognition allele likelihood factor corresponding to the haplotype allele from the haplotype reference panel and a second transition recognition allele likelihood factor corresponding to the haplotype allele, thereby accessing, from the memory device, for the marker variant, the first allele likelihood factor and the second allele likelihood factor.
13. The system according to claim 12, further including instructions that, when executed by the at least one processor, cause the system to pre-determine the first transition recognition allele likelihood factor and the second transition recognition allele likelihood factor, before determining one or more intermediate allele likelihoods corresponding to the marker variant as part of a path across a haplotype matrix.
14. When executed by the at least one processor, the system By combining the allele likelihood factor for the haplotype allele and the transition constant coefficient for the transition between haplotypes from the haplotype reference panel, the first transition recognition allele likelihood factor is determined in advance. The system according to claim 13, further comprising an instruction to determine the second transition recognition allele likelihood factor in advance by combining the allele likelihood factor and the transition linear coefficient for the transition between haplotypes from the haplotype reference panel. **Claim 15** The system according to claim 12, further comprising an instruction to cause the system to determine the first transition recognition allele likelihood factor by combining an allele likelihood coefficient and a transition linear coefficient when executed by the at least one processor. **Claim 16** The first allele likelihood factor includes an allele likelihood factor for a sample reference haplotype allele or a sample alternative haplotype allele. The system according to claim 15, wherein the second allele likelihood factor includes the allele likelihood factor for the sample reference haplotype allele or the sample alternative haplotype allele. **Claim 17** The system according to claim 11, further comprising an instruction to cause the system to combine the first allele likelihood factor and the adjacent marker intermediate allele likelihood by multiplying the first transition recognition allele likelihood factor and the adjacent marker intermediate allele likelihood without a further multiplication operation for determining the intermediate allele likelihood when executed by the at least one processor. **Claim 18** A data flow engine, and when executed by the at least one processor, the system sends a respective set of input values including an allele likelihood factor, a transition coefficient, and a haplotype-allele value from the data flow engine to each acceleration calculation engine of a cluster of acceleration calculation engines. The system according to claim 11, further comprising an instruction to cause each of the respective acceleration calculation engines to determine a respective set of intermediate allele likelihoods corresponding to each subset of marker variants and each subset of haplotypes based on each set of the input values. **Claim 19** When executed by the at least one processor, the system each respective set of input values from the data flow engine to each respective acceleration calculation engine, transmitting, from the data flow engine, a first set of input values including an allele likelihood factor, a transition coefficient, and a haplotype-allele value to a first acceleration calculation engine of a cluster of acceleration calculation engines; causing transmission by transmitting, from the data flow engine, a second set of input values including an allele likelihood factor, a transition coefficient, and a haplotype-allele value to a second acceleration calculation engine of a cluster of acceleration calculation engines; each respective set of intermediate allele likelihoods, determining, by the first acceleration calculation engine, a first set of intermediate allele likelihoods corresponding to a first subset of marker variants and a first subset of haplotypes based on the first set of input values; The system of claim 18, further comprising instructions for causing determination by causing the second acceleration calculation engine to determine a second set of intermediate allele likelihoods corresponding to a second subset of marker variants and a second subset of haplotypes based on the second set of input values. Claim 20 The system of claim 11, further comprising instructions for causing the system to determine the intermediate allele likelihood by determining the intermediate allele likelihood of the genomic region including the sample reference haplotype allele or the sample alternative haplotype allele when executed by the at least one processor. Claim 21 The system of claim 11, further comprising instructions for causing the system to determine the intermediate allele likelihood based on the adjacent marker factor recognition allele likelihood and the second allele likelihood factor by summing the adjacent marker transition factor recognition allele likelihood and the total adjacent marker transition recognition allele likelihood factor when executed by the at least one processor. Claim 22 The system of claim 11, further comprising instructions for causing the system to determine one or more nucleotide calls for the genomic region from the genomic sample based on the allele likelihood of the genomic region and one or more variant nucleotide calls surrounding the genomic region when executed by the at least one processor. Claim 23 A method, Determining a first-pass intermediate allele likelihood of a genomic region from a genomic sample containing haplotype alleles corresponding to a set of haplotypes given a set of marker variants using a configurable processor that executes a first pass; Storing a subset of the first-pass intermediate allele likelihoods corresponding to a subset of the marker variants for a group of marker variants; Using the configurable processor to reproduce the first-pass intermediate allele likelihoods by initializing allele likelihood determination in the group of marker variants using the stored subset of the first-pass intermediate allele likelihoods; Determining a second-pass intermediate allele likelihood of the genomic region containing the haplotype alleles corresponding to the set of haplotypes given the set of marker variants using the configurable processor that executes a second pass; Generating an allele likelihood of the genomic region containing the haplotype alleles based on the reproduced first-pass intermediate allele likelihoods and the second-pass intermediate allele likelihoods. A method comprising.
24. The method according to claim 23, wherein the configurable processor includes an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a coarse-grained reconfigurable array (CGRA), or a field-programmable gate array (FPGA).
25. Determining the first-pass intermediate allele likelihood includes determining a reverse intermediate allele likelihood of the genomic region containing the haplotype alleles using a reverse pass, The method according to claim 23, wherein determining the second-pass intermediate allele likelihood includes determining a forward intermediate allele likelihood of the genomic region containing the haplotype alleles using a forward pass.
26. Storing the subset of the first-pass intermediate allele likelihoods includes storing the subset of the first-pass intermediate allele likelihoods in a dynamic random access memory (DRAM), Initializing the allele likelihood determination in the set of marker variants by using the stored subset of the first-pass intermediate allele likelihoods includes accessing the stored subset of the first-pass intermediate allele likelihoods from the DRAM, the method according to claim 23.
27. Generating the allele likelihood based on the reproduced first-pass intermediate allele likelihood and the second-pass intermediate allele likelihood includes Determining a total first-pass intermediate allele likelihood for the set of marker variants based on the reproduced first-pass intermediate allele likelihood; Determining a total second-pass intermediate allele likelihood for the set of marker variants based on the second-pass intermediate allele likelihood; Determining the allele likelihood based on the total first-pass intermediate allele likelihood and the total second-pass intermediate allele likelihood, the method according to claim 23.
28. For adjacent marker variants, determining a running total of a first subset of intermediate allele likelihoods of the genomic region including first-type haplotype alleles from one or more haplotypes of a haplotype reference panel; For adjacent marker variants, determining a running total of a second subset of intermediate allele likelihoods of the genomic region including second-type haplotype alleles from the one or more haplotypes; For marker variants, further including determining a total of intermediate allele likelihoods of the genomic region including haplotype alleles from haplotypes of the haplotype reference panel based on the running total of the first subset of intermediate allele likelihoods and the running total of the second subset of intermediate allele likelihoods, the method according to claim 23.
29. Storing haplotype-allele-index data for a haplotype matrix in a dynamic random access memory (DRAM); To generate the allele likelihoods using a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model, the configurable processor from the DRAM further accesses the haplotype-allele-index data of the haplotype matrix. The method according to claim 23.
30. Initializing allele likelihood determination in the group of marker variants using the stored subset of first-pass intermediate allele likelihoods, Determining a first subset of the first-pass intermediate allele likelihoods for the first group of marker variants based on a first stored column of the first-pass intermediate allele likelihoods for the initial marker variants from the first group of marker variants, Determining a second subset of the first-pass intermediate allele likelihoods for the second group of marker variants based on a second stored column of the first-pass intermediate allele likelihoods for the initial marker variants from the second group of marker variants. The method according to claim 23.
31. A system, At least one processor; A memory device; A non-transitory computer-readable medium. When the non-transitory computer-readable medium is executed by the at least one processor, the system is caused to: Execute a first pass to determine first-pass intermediate allele likelihoods for a genomic region from a genomic sample including haplotype alleles corresponding to a given set of haplotypes for a set of marker variants; Store on the memory device a subset of the first-pass intermediate allele likelihoods corresponding to a subset of marker variants for the group of marker variants; Reproduce the first-pass intermediate allele likelihoods by initializing allele likelihood determination in the group of marker variants using the stored subset of the first-pass intermediate allele likelihoods; Execute a second pass to determine second-pass intermediate allele likelihoods for the genomic region including the haplotype alleles corresponding to the given set of haplotypes for the set of marker variants; A system including instructions to generate allele likelihoods of the genomic region including the haplotype alleles based on the reproduced first path intermediate allele likelihoods and the second path intermediate allele likelihoods using an output engine.
32. A haplotype-allele-index memory for storing haplotype-allele-index data, A transition coefficient memory for storing transition coefficients, and An allele likelihood factor memory for storing allele likelihood factors, the system according to claim 31.
33. The system according to claim 31, further including a joint engine for determining intermediate allele likelihood values.
34. A data flow engine, and when executed by the at least one processor, causing the system to send each set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from the data flow engine to each acceleration calculation engine of a cluster of acceleration calculation engines, and instructions for causing each of the acceleration calculation engines to determine each set of intermediate allele likelihoods corresponding to each subset of marker variants and each subset of haplotypes based on each set of the input values, the system according to claim 31.
35. When executed by the at least one processor, causing the system to send each set of the input values from the data flow engine to each of the acceleration calculation engines, send a first set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from the data flow engine to a first acceleration calculation engine of the cluster of acceleration calculation engines, send a second set of input values including allele likelihood factors, transition coefficients, and haplotype-allele values from the data flow engine to a second acceleration calculation engine of the cluster of acceleration calculation engines, each set of the intermediate allele likelihoods, determine, by the first acceleration calculation engine, a first set of intermediate allele likelihoods corresponding to a first subset of marker variants and a first subset of haplotypes based on the first set of the input values, The system according to claim 34, further comprising an instruction to cause determination by the second acceleration calculation engine to determine a second set of intermediate allele likelihoods corresponding to a second subset of marker variants and a second subset of haplotypes based on the second set of input values.
36. A dataflow engine corresponding to a cluster of acceleration calculation engines, and when executed by the at least one processor, cause the system to send a subset of the first pass intermediate allele likelihoods from the dataflow engine to a first acceleration calculation engine among the cluster of acceleration calculation engines for the first acceleration calculation engine to reproduce the first pass intermediate allele likelihoods, The system according to claim 31, further comprising an instruction to send an additional subset of the first pass intermediate allele likelihoods from the dataflow engine to a second acceleration calculation engine from the cluster of acceleration calculation engines for the second acceleration calculation engine to reproduce additional first pass intermediate allele likelihoods.
37. A dataflow engine, and when executed by the at least 31 processors, cause the system to send the subset of the first pass intermediate allele likelihoods from the memory device to the dataflow engine, The system according to claim 31, further comprising an instruction to send the subset of the first pass intermediate allele likelihoods from the dataflow engine to an acceleration calculation engine to reproduce the first pass intermediate allele likelihoods based on the subset of the first pass intermediate allele likelihoods.
38. A dataflow engine, and when executed by the at least 31 processors, cause the system to store haplotype-allele-index data for a haplotype matrix on the memory device, The system according to claim 31, further comprising an instruction to access the haplotype-allele-index data for the haplotype matrix from the memory device to generate the allele likelihoods using a hidden Markov haploid genotype attribution model or a hidden Markov diploid genotype attribution model.
39. The system of claim 31, wherein the memory device includes a dynamic random access memory (DRAM), a static random access memory (SRAM), or a cache memory device.
40. The system of claim 31, further comprising instructions that, when executed by the at least one processor, cause the system to determine one or more nucleotide calls for the genomic region from the genomic sample based on the allelic likelihood of the genomic region and one or more variant nucleotide calls surrounding the genomic region.
41. A method comprising: identifying a haplotype reference panel for a genomic region of a genomic sample using a genotype imputation model; determining, for adjacent marker variants, a running total of a first subset of intermediate allelic likelihoods of the genomic region that includes a first type of haplotype alleles from one or more haplotypes of the haplotype reference panel; determining, for the adjacent marker variants, a running total of a second subset of intermediate allelic likelihoods of the genomic region that includes a second type of haplotype alleles from the one or more haplotypes; determining, for a marker variant, a total of intermediate allelic likelihoods of the genomic region that includes haplotype alleles from a haplotype of the haplotype reference panel, based on the running total of the first subset of intermediate allelic likelihoods and the running total of the second subset of intermediate allelic likelihoods; generating an allelic likelihood for the genomic region that includes the haplotype alleles, based on the total of the intermediate allelic likelihoods.
42. The method of claim 41, wherein the first type of haplotype alleles includes sample reference haplotype alleles and the second type of haplotype alleles includes sample alternative haplotype alleles.
43. Determining the sum of the intermediate allele likelihoods, by a configurable processor, for the marker variant, based on the intermediate allele likelihoods from the first subset of intermediate allele likelihoods or the second subset of intermediate allele likelihoods, determining an initial intermediate allele likelihood from the intermediate allele likelihoods, and then, for the adjacent marker variant, summing the adjacent marker intermediate allele likelihoods of the genomic region including the haplotype allele, the method according to claim 41.
44. Determining the sum of the intermediate allele likelihoods, by a configurable processor, for the marker variant, based on the intermediate allele likelihoods from the first subset of intermediate allele likelihoods or the second subset of intermediate allele likelihoods, determining an initial intermediate allele likelihood from the intermediate allele likelihoods, and then, for the adjacent marker variant, generating an allele likelihood of the genomic region including the haplotype allele, the method according to claim 41.
45. Pre-determining a first transition recognition allele likelihood factor corresponding to a row for the first type of haplotype allele and a second transition recognition allele likelihood factor corresponding to a row for the second type of haplotype allele; Further comprising determining the sum of the intermediate allele likelihoods based further on the first transition recognition allele likelihood factor corresponding to the row for the first type of haplotype allele and the second transition recognition allele likelihood factor corresponding to the row for the second type of haplotype allele, the method according to claim 41.
46. Determining, for the adjacent marker variant, the adjacent marker sum of the intermediate allele likelihoods of the genomic region including the haplotype allele; Further comprising determining the sum of the intermediate allele likelihoods based further on a combination of the adjacent marker sum of the intermediate allele likelihoods for the marker variant, the first transition recognition allele likelihood factor corresponding to the row for the first type of haplotype allele, and the second transition recognition allele likelihood factor corresponding to the row for the second type of haplotype allele, the method according to claim 45.
47. multiplying the running total of the first subset of the intermediate allele likelihoods by a first transition recognition allele likelihood factor; multiplying the running total of the second subset of the intermediate allele likelihoods by a second transition recognition allele likelihood factor; the method of claim 41, further comprising determining the total of the intermediate allele likelihoods for the marker variant based on the multiplied running total of the first subset of the intermediate allele likelihoods and the multiplied running total of the second subset of the intermediate allele likelihoods. **Claim 48** pre-determining the first transition recognition allele likelihood factor includes combining a first allele likelihood factor for the first type of haplotype allele and a transition linear coefficient for transitions between haplotypes from the haplotype reference panel; the method of claim 47, further comprising pre-determining the second transition recognition allele likelihood factor includes combining a second allele likelihood factor for the second type of haplotype allele and the transition linear coefficient. **Claim 49** A non-transitory computer-readable medium that, when executed by at least one processor, causes a computing device to utilize a genotype attribution model to identify a haplotype reference panel for a genomic region of a genomic sample; for an adjacent marker variant, determine a running total of a first subset of intermediate allele likelihoods of the genomic region including haplotype alleles of the first type from one or more haplotypes of the haplotype reference panel; for the adjacent marker variant, determine a running total of a second subset of intermediate allele likelihoods of the genomic region including haplotype alleles of the second type from the one or more haplotypes; for a marker variant, determine the total of the intermediate allele likelihoods of the genomic region including haplotype alleles from the haplotypes of the haplotype reference panel based on the running total of the first subset of the intermediate allele likelihoods and the running total of the second subset of the intermediate allele likelihoods; A non - transitory computer - readable medium storing instructions to generate allele likelihoods of the genomic region including the haplotype alleles based on the sum of the intermediate allele likelihoods.
50. The non - transitory computer - readable medium according to claim 49, wherein the first type of haplotype alleles includes sample reference haplotype alleles and the second type of haplotype alleles includes sample alternative haplotype alleles.
51. When executed by the at least one processor, the computing device, by a configurable processor, for the marker variant, based on the intermediate adjacent allele likelihoods from the first subset of intermediate allele likelihoods or the second subset of intermediate allele likelihoods, determines an initial intermediate allele likelihood from the intermediate allele likelihoods, and for the adjacent marker variant, before summing the intermediate allele likelihoods of the genomic region including the haplotype alleles, further includes instructions to cause the sum of the intermediate allele likelihoods to be determined. The non - transitory computer - readable medium according to claim 49.
52. When executed by the at least one processor, the computing device, by a configurable processor, for the marker variant, based on the intermediate adjacent allele likelihoods from the first subset of intermediate allele likelihoods or the second subset of intermediate allele likelihoods, determines an initial intermediate allele likelihood from the intermediate allele likelihoods, and for the adjacent marker variant, before generating the allele likelihoods of the genomic region including the haplotype alleles, further includes instructions to cause the sum of the intermediate allele likelihoods to be determined. The non - transitory computer - readable medium according to claim 49.
53. When executed by the at least one processor, the computing device, pre - determines a first transition - recognition allele likelihood factor corresponding to a row for the first type of haplotype alleles and a second transition - recognition allele likelihood factor corresponding to a row for the second type of haplotype alleles. The non - transitory computer - readable medium according to claim 49, further comprising an instruction for causing determination of the sum of the intermediate allele likelihoods, further based on the first transition recognition allele likelihood factor corresponding to the row for the haplotype allele of the first type and the second transition recognition allele likelihood factor corresponding to the row for the haplotype allele of the second type.
54. When executed by the at least one processor, cause the computing device to for the adjacent marker variant, determine the sum of the adjacent marker intermediate allele likelihoods of the genomic region including the haplotype allele, The non - transitory computer - readable medium according to claim 53, further comprising an instruction for causing determination of the sum of the intermediate allele likelihoods, further based on a combination of the sum of the adjacent marker intermediate allele likelihoods for the marker variant, the first transition recognition allele likelihood factor corresponding to the row for the haplotype allele of the first type, and the second transition recognition allele likelihood factor corresponding to the row for the haplotype allele of the second type.
55. When executed by the at least one processor, cause the computing device to multiply the running sum of the first subset of the intermediate allele likelihoods by a first transition recognition allele likelihood factor, multiply the running sum of the second subset of the intermediate allele likelihoods by a second transition recognition allele likelihood factor, The non - transitory computer - readable medium according to claim 49, further comprising an instruction for causing determination of the sum of the intermediate allele likelihoods, based on the multiplied running sums of the first subset and the multiplied running sums of the second subset of the intermediate allele likelihoods for the marker variant.
56. When executed by the at least one processor, cause the computing device to pre - determine the first transition recognition allele likelihood factor, including combining a first allele likelihood factor for the haplotype allele of the first type and a transition linear coefficient for the transition between haplotypes from the haplotype reference panel. The non-transitory computer-readable medium according to claim 55, further comprising an instruction for pre-determining the second transition recognition allele likelihood factor, including combining a second allele likelihood factor for the second type of haplotype allele and the transition linear coefficient.
57. The non-transitory computer-readable medium according to claim 56, wherein the first allele likelihood factor includes an allele likelihood factor for a sample reference haplotype allele, and the second allele likelihood factor includes an allele likelihood factor for a sample alternative haplotype allele.