Machine learning models for recalibrating nucleotide base calls corresponding to targeted variants
Patent Information
- Application Number
- JP2023579835
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-28
- Filing Date
- 2022-12-23
- Publication Date
- 2026-01-07
AI Technical Summary
Existing nucleotide base sequence determination systems struggle with inaccuracies in variant calling, particularly for complex genomic regions like multiple-to-paired gene coordinates, due to insufficient training data, reliance on limited data sources, and inefficient computational resources, leading to false negative and false positive variant calls, and lack of interpretability in deep learning models.
A machine learning model is employed to re-correct nucleotide base calls using a call re-comparative system that trains on specific genomic coordinates, incorporating both internal and external array determination metrics to improve accuracy, efficiency, and interpretability, generating a call re-comparable model that updates initial nucleotide base calls.
The system enhances the accuracy and speed of nucleotide base calling, reduces false positive and false negative variant calls, and provides interpretable results, improving computational efficiency by using a lightweight architecture that can process nucleotide sequences in under 30 minutes compared to traditional deep learning methods that take hours or days.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. patent application Ser. No. 17 / 563,934, filed Dec. 28, 2021, entitled "MACHINE-LEARNING MODEL FOR RECALIBRATING NUCLEOTIDE BASE CALLS CORRESPONDING TO TARGET VARIANTS," the contents of which are incorporated herein by reference in their entirety. [Background technology]
[0002] In recent years, biotechnology companies and research institutions have improved hardware and software for sequencing nucleotides and determining nucleotide base calls (e.g., variant calls) for genomic samples. For example, some existing nucleotide base sequencing platforms determine individual nucleotide bases in a sequence by using traditional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms can monitor thousands of nucleic acid polymers that are synthesized in parallel to predict nucleotide base calls from a larger base call dataset. For example, a camera in many SBS platforms captures images of illuminated fluorescent tags incorporated into oligonucleotides to determine nucleotide base calls. After capturing such images, existing SBS platforms transmit the base call data (or image data) to a computing device to apply sequencing data analysis software that determines the nucleotide base sequence of the nucleic acid polymer. In certain cases, some conventional systems further utilize a variant caller to identify variants, such as single nucleotide polymorphisms (SNPs), insertions or deletions (indels), or other variants within the nucleic acid sequence of a sample.
[0003] Despite these recent advances in sequencing and variant calling, existing nucleotide-based sequencing platforms and sequencing data analysis software (collectively, hereinafter, existing sequencing systems) often include variant callers that inaccurately determine nucleotide base calls (and / or corresponding variant calls). For example, existing sequencing systems inaccurately or are unable to determine nucleotide base calls for multi-allelic genomic coordinates. Indeed, for regions of nucleotide sequences, such as multi-allelic regions that are more challenging than bi-allelic regions, some existing systems struggle (or are unable) to accurately determine genotypes when alleles cover or correspond to a given genomic coordinate. For example, some machine learning-based sequencing systems struggle to determine genotypes for multi-allelic coordinates because the training data is primarily bi-allelic data. Thus, in the case of pile-ups or large insertions, existing sequencing systems often fail to accurately determine nucleotide base calls and / or genotypes from multiple possible alleles at a given genomic coordinate.
[0004] In addition, existing sequencing systems inaccurately determine nucleotide base calls (e.g., variant calls) for haploid genomic coordinates in genomic samples or other nucleotide sequences. For example, many existing sequencing systems inaccurately determine nucleotide base calls in sex chromosomes, often due to the sparseness or complete lack of good haploid training data. In particular, existing sequencing systems often learn parameters for determining nucleotide base calls exclusively from unmodified diploid data (e.g., as described in PrecisionFDA Truth Data from the PrecisionFDA Truth Challenge, https: / / precision.fda.gov / challenges / truth) and lack models or training for identifying nucleotide bases or genotypes of coordinates other than diploid coordinates. As a result, many of these existing sequencing systems cannot accurately determine nucleotide base calls or variant calls for haploid genomic coordinates.
[0005] Furthermore, in some situations, existing sequencing systems apply variant callers that inaccurately identify an excessive number of false-negative variant calls. For example, existing sequencing systems may determine that a genomic coordinate indicates a homozygous reference genotype (and therefore does not contain a variant) when in fact the coordinate contains a variant. Indeed, existing variant callers have achieved a certain level of accuracy, but their limitations leave room for improvement in recovering false-negative variant calls. To illustrate the impact of such inaccuracies, a variant call that identifies a particular single nucleotide polymorphism (SNP) in the hemoglobin beta (HBB) gene can have significant implications. For example, if a variant caller identifies a SNP at rs344 on chromosome 11, the variant caller can either correctly identify the genetic cause of sickle cell anemia or miss the cause of the disease. As a further example, variant calling that accurately or inaccurately identifies a deletion in one or more copies of the hemoglobin subunit alpha 1 (HbA1) or hemoglobin subunit alpha 2 (HbA2) genes can result in either accurately identifying the genetic cause of an inherited blood disorder or missing the gene deletion entirely.
[0006] As a contributing factor to the aforementioned inaccuracies, many existing sequencing systems only utilize a limited set of data when determining nucleotide base calls. For example, existing sequencing systems often rely exclusively on information extracted directly from the nucleotide reads of the sample sequence, such as read depth, number of mismatches, sequence alignment score, and mapping quality, to determine nucleotide base calls. Although sequence information from nucleotide reads can provide valuable insights for determining nucleotide base calls, existing sequencing systems that rely only on these data may perform poorly when determining nucleotide base calls. In fact, some existing sequencing systems that rely on raw sequence data inaccurately determine SNPs, indels, or other variants in genomic sample sequences compared to more complex models. In fact, existing sequencing systems often identify false negative or false positive variants in the Truth Challenges of the US Food and Drug Administration (FDA), and reliable haploid data is often difficult to obtain for testing or training variant callers.
[0007] In addition to making inaccurate variant calls, some existing sequencing systems also waste computational resources inefficiently by using overly complex models. Specifically, the variant callers of some existing sequencing systems are computationally expensive and slow. In fact, some existing sequencing systems utilize variant callers with deep learning architectures or some other neural network architectures that require large computational resources (e.g., computation time, processing power, and memory) to train and apply. For example, some existing sequencing systems utilize deep learning architectures that require a lot of time across multiple computing devices to generate nucleotide base calls for a single sample sequence, even after training.
[0008] As a further drawback of existing sequencing systems with complex networks, many such systems utilize model architectures that make sequence data uninterpretable. More specifically, some existing deep neural networks transform and manipulate sequence data multiple times, changing it from one vector to another across various layers and neurons as a basis for generating variant calls. In many cases, the internal data of these deep neural networks is uninterpretable and cannot be utilized in any way outside the neural network architecture itself. Summary of the Invention [Means for solving the problem]
[0009] The present disclosure describes embodiments of methods, non-transitory computer readable media, and systems that can utilize machine learning models to recalibrate nucleotide base calls (e.g., variant calls) of a call generation model. For example, the disclosed system can train and utilize a call recalibration machine learning model to generate a set of classification predictions (e.g., variant call classifications) to improve nucleotide base calls in certain scenarios, such as nucleotide base calls for coordinates that were incorrectly identified by an existing sequencing system as indicative of a multi-allelic coordinate, a haploid coordinate, and / or a homozygous reference genotype. As disclosed, the disclosed system can (i) determine sequencing metrics for a particular genomic coordinate (e.g., a multi-allelic coordinate, a haploid coordinate, or a homozygous reference coordinate that was incorrectly identified), and (ii) utilize a call recalibration machine learning model to generate classification predictions to update or recalibrate the initial nucleotide base calls for the genomic coordinate. After recalibration, the disclosed system can output the updated or recalibrated nucleotide base calls as final nucleotide base calls (e.g., final variant calls) in a variant call file or other base call output file.
[0010] By utilizing the call recalibration machine learning model to update the sequencing metrics for generating nucleotide base calls, the disclosed system can improve accuracy, efficiency, and speed over existing sequencing systems. As described further below, for example, the disclosed call recalibration machine learning model determines variant calls with better accuracy than traditional hidden Markov model (HMM)-based or probability-based variant callers and more complex neural networks (e.g., deep neural network-based variant callers) for variant calling at multi-allelic coordinates, haploid coordinates, or inaccurately identified homozygous reference coordinates. The disclosed call recalibration machine learning model also determines variant calls at such genomic coordinates with faster computation time than complex neural networks. In addition, the disclosed system can improve the interpretability of factors that impact accurate variant calling at such genomic coordinates compared to complex neural networks by utilizing the call recalibration machine learning model that processes data in an accessible, interpretable format. Indeed, for improved interpretability of the disclosed systems, in some embodiments, the disclosed systems can generate and provide visualizations of various contribution measures associated with individual sequencing metrics to visually depict the respective measures of impact that sequencing metrics have on the resulting nucleotide base calls. [Brief description of the drawings]
[0011] The detailed description refers to the drawings, which are briefly described below. [Figure 1] FIG. 1 illustrates a block diagram of a sequencing system including a call recalibration system in accordance with one or more embodiments. [Diagram 2] 1 shows an overview of a call recalibration system that utilizes a call recalibration system to generate nucleotide base calls, in accordance with one or more embodiments. [Figure 3A]1 illustrates a call recalibration system for generating nucleotide base calls for multi-allelic genomic coordinates, according to one or more embodiments. [Figure 3B] 1 illustrates a call recalibration system for generating nucleotide base calls for multi-allelic genomic coordinates, according to one or more embodiments. [Figure 4A] 1 illustrates a call recalibration system for generating nucleotide base calls for haploid genome coordinates, according to one or more embodiments. [Figure 4B] 1 illustrates a call recalibration system for generating nucleotide base calls for haploid genome coordinates, according to one or more embodiments. [Diagram 5] 1 illustrates a call recalibration system for generating variant calls for homozygous reference genome coordinates, according to one or more embodiments. [Figure 6A] 1 illustrates a call recalibration system for generating or determining sequencing metrics in accordance with one or more embodiments. [Figure 6B] 1 illustrates a call recalibration system for generating or determining sequencing metrics in accordance with one or more embodiments. [Figure 6C] 1 illustrates a call recalibration system for generating or determining sequencing metrics in accordance with one or more embodiments. [Figure 7] 1 illustrates a call recalibration system that utilizes a call recalibration machine learning model to generate variant call classifications and recalibrate nucleotide base calls, according to one or more embodiments. [Figure 8] 1 illustrates an example process of a call recalibration system for training a call recalibration machine learning model in accordance with one or more embodiments. [Figure 9] 1 illustrates an exemplary contribution measure interface displayed on a client device in accordance with one or more embodiments. [Figure 10A] 13A-13C depict graphs and tables illustrating the improvement in accuracy associated with a call recalibration system for diploid coordinates, in accordance with one or more embodiments. [Figure 10B]13A-13C depict graphs and tables illustrating the improvement in accuracy associated with a call recalibration system for diploid coordinates, in accordance with one or more embodiments. [Figure 11A] 13A-13C depict graphs and tables illustrating the improvement in accuracy associated with a call recalibration system for haploid coordinates, in accordance with one or more embodiments. [Figure 11B] 13A-13C depict graphs and tables illustrating the improvement in accuracy associated with a call recalibration system for haploid coordinates, in accordance with one or more embodiments. [Figure 12] 1 depicts a flowchart of a series of operations for generating nucleotide base calls associated with multi-allelic genomic coordinates in accordance with one or more embodiments. [Figure 13] 1 depicts a flowchart of a series of operations for generating nucleotide base calls associated with haploid genome coordinates in accordance with one or more embodiments. [Figure 14] 1 depicts a flowchart of a series of operations for generating variant calls associated with homozygous reference genomic coordinates, according to one or more embodiments. [Figure 15] 1 illustrates a block diagram of an exemplary computing device in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The present disclosure describes an embodiment of a call recalibration system that utilizes a call recalibration machine learning model to generate and recalibrate nucleotide base calls for a sample nucleotide sequence. In particular, the call recalibration system can utilize the call recalibration machine learning model to update, recalibrate, or correct the initial nucleotide base calls generated by the call generation model. For example, the call recalibration system can utilize the call recalibration machine learning model to recalibrate and improve the accuracy of the initial nucleotide base calls by updating various call metrics, such as call quality, genotypes associated with the calls, genotype quality associated with the genotypes, Phred-scaled Likelihood (PL), and / or other metrics with corresponding fields. By utilizing the call recalibration machine learning model to update the metrics, the call recalibration system can improve the accuracy of nucleotide base calls at certain genomic coordinates, such as multi-allelic coordinates, haploid coordinates, and coordinates that were erroneously determined (in the initial call or by an existing sequencing system) to indicate a homozygous reference genotype.
[0013] As just mentioned, in certain implementations, the call recalibration system improves the nucleotide base calls and corresponding variant calls for multi-allelic coordinates of the sample nucleotide sequence. To facilitate the generation of multi-allelic nucleotide base calls, in some embodiments, the call recalibration system utilizes a call recalibration machine learning model specialized and adaptable to generate nucleotide base calls for both bi-allelic and multi-allelic coordinates. For example, the call recalibration system can generate a set of variant call classifications from the sequencing metrics associated with the multi-allelic genomic coordinates, including a probability of a homozygous reference genotype at the multi-allelic genomic coordinate (i.e., reference probability), a probability of a genotype error at the multi-allelic genomic coordinate (i.e., different genotype probability), and a probability of a correct variant call genotype at the multi-allelic genomic coordinate (i.e., correct variant probability). The call recalibration system can further determine a final nucleotide base call for the multi-allelic genomic coordinate from the set of variant call classifications. Further details regarding the generation of calls for multi-allelic coordinates are provided below with reference to the figures.
[0014] As mentioned, in one or more embodiments, the call recalibration system improves the nucleotide base calls and corresponding variant calls for haploid genomic coordinates of the sample nucleotide sequence. In particular, the call recalibration system can utilize a call recalibration machine learning model adapted to determine haploid genotypes based on diploid data. For example, the call recalibration system can train the call recalibration machine learning model by modifying the diploid data (e.g., diploid sequencing metrics) to simulate haploid data (e.g., haploid sequencing metrics). In addition, the call recalibration system can utilize the trained call recalibration machine learning model to generate three outputs for a given genomic coordinate: (i) a first confidence score for a homozygous reference genotype (0 / 0), (ii) a second confidence score for a heterozygous genotype (0 / 1), and (iii) a third confidence score for a homozygous alternative genotype (1 / 1).
[0015] The call recalibration system can further reduce or remove the second confidence score (e.g., the 0 / 1 confidence score) and utilize a softmax model or tier to normalize across the other two confidence scores and convert the confidence score to a haploid probability. Thus, utilizing a softmax model or tier, the call recalibration system can (i) determine a haploid reference probability (0) from the homozygous reference confidence score (0 / 0) and (ii) determine a haploid substitution probability (1) from the homozygous substitution confidence score (1 / 1). Further details regarding the generation of calls for haploid coordinates are provided below with reference to the figures.
[0016] As further mentioned above, the call recalibration system improves nucleotide base calls and corresponding variant calls for genomic coordinates of a sample nucleotide sequence determined to be indicative of a homozygous reference genotype. More specifically, the call recalibration system can recover false negative variant calls for genomic coordinates that were initially determined to be indicative of a homozygous reference genotype (e.g., as determined by a call generation model) when, in fact, the genotype of those coordinates is not homozygous with respect to the reference sequence. In contrast to existing sequencing systems that filter out data associated with homozygous reference coordinates, the call recalibration system can determine sequencing metrics for such homozygous reference coordinates and utilize a call recalibration machine learning model to generate variant call classifications from the sequencing metrics. Furthermore, the call recalibration system can generate final nucleotide base calls for the homozygous reference coordinates based on the variant call classifications and can change variant calls that were indicative of a homozygous reference genotype to indicate a different genotype (thereby recovering false negative variant calls). Further details regarding correcting or updating variant calls for genomic coordinates that would have been incorrectly identified as indicative of a homozygous reference genotype are provided below with reference to the figures.
[0017] As mentioned above, in some embodiments, the call recalibration system may more generally utilize a machine learning model to generate variant call classifications based on sequencing metrics of nucleotide base calls corresponding to genomic coordinates. To generate such classifications, the call recalibration system extracts or determines sequencing metrics from sample nucleotide sequences. For example, the call recalibration system determines sequencing metrics from nucleotide base calls of nucleotide reads from the sample nucleotide sequences. Indeed, in some cases, the call recalibration system generates or determines a set of initial nucleotide base calls from nucleotide reads captured or determined via fluorescent imaging of the sample nucleotide sequences (e.g., at specific genomic coordinates). From the read-based nucleotide base calls, in some embodiments, the call recalibration system determines or extracts various sequencing metrics (e.g., various types of sequencing metrics obtained from the reads and / or from different components of the call generation model).
[0018] More specifically, in certain implementations, the call recalibration system determines different types of sequencing metrics associated with different sources. For example, the call recalibration system determines read-based sequencing metrics, including metrics derived from nucleotide reads of a sample nucleotide sequence. In addition, the call recalibration system determines sequencing metrics of external sources identified from one or more external databases representing genomic sequences associated with various nucleotide attributes, mapping challenges, and sequencing biases. Furthermore, the call recalibration system determines call model-generated sequencing metrics generated via a variant caller or other call generation model, such as variables internal to the call recalibration system that are not accessible to other systems or parties (e.g., proprietary quality scores, base context, read filtering, proprietary hypothesis scores, and other metrics). Indeed, in some cases, the call recalibration system determines call model-generated sequencing metrics in the form of variant calling sequencing metrics and mapping alignment sequencing metrics, each type being extracted by a different component of the call generation model.
[0019] As further noted, in certain implementations, the call recalibration system generates a set of predicted classifications from the sequencing metrics to correct or improve the nucleotide base call or variant call data or fields associated with the nucleotide base calls. More specifically, the call recalibration system utilizes a call recalibration machine learning model to generate a set of three variant call classifications from the sequencing metrics that impact or reflect the accuracy of identifying variants at a particular genomic coordinate (e.g., a genomic coordinate corresponding to a nucleotide base call of a nucleotide read from a sample nucleotide sequence). Optionally, the call recalibration system can utilize a call recalibration machine learning model to generate variant call classifications for multi-allelic coordinates that are different from a reference coordinate that may be, for example, a haploid coordinate or a pseudohomozygous coordinate.
[0020] For example, when generating variant call classifications for a multi-allelic genomic coordinate, the call recalibration system can utilize a call recalibration machine learning model to generate a set that includes (i) a reference probability of a homozygous reference genotype at the multi-allelic genomic coordinate, (ii) different genotype probabilities of a genotype error at the multi-allelic genomic coordinate, and (iii) a correct variant probability of a correct variant call genotype at the multi-allelic genomic coordinate. As another example, for a haploid coordinate, the call recalibration system can utilize a call recalibration machine learning model to generate a set of variant call classifications that include (i) a first genotype probability of a first genotype at the genomic coordinate, and (ii) a second genotype probability of a second genotype at the genomic coordinate. Further, for reference coordinates that would be homozygous, the call recalibration system can utilize the call recalibration machine learning model to generate a set of variant call classifications that include: (i) a false positive classification (e.g., the probability that the nucleotide base call is a false positive variant), (ii) a genotype error classification (e.g., a heterozygous genotype classification indicating the probability of identifying the correct alternative allele but with a genotype error - e.g., 0 / 1 instead of 1 / 1, or 1 / 1 instead of 0 / 1 -, or the probability of inaccurately identifying the genotype of the nucleotide base call), and (iii) a true positive classification (e.g., a homozygous alternative classification indicating the probability that the nucleotide base call or genotype call is a true positive variant). Thus, in some cases, the variant call classification represents an intermediate scoring metric associated with the variant caller.
[0021] From the variant call classification, the call recalibration system can further correct or update one or more final nucleotide base calls for the genomic coordinates (e.g., a final nucleotide base call indicative of a variant call or a non-variant call). For example, the call recalibration system uses the variant call classification to update a data field in a digital call file (e.g., a variant call format file or other base call output file) that indicates or represents the final nucleotide base call and / or variant call. Indeed, as mentioned above, in some embodiments, the call recalibration system uses a call generation model to generate or determine a final nucleotide base call from sequencing metrics for the genomic coordinates.
[0022] In addition, the call recalibration system can utilize the variant call classification to update the nucleotide base calls and / or variant calls to improve accuracy. In certain implementations, the call recalibration system updates nucleotide base calls for certain genomic coordinates, such as multi-allelic genomic coordinates, haploid genomic coordinates, and / or homozygous reference coordinates that would be incorrectly identified (i.e., genomic coordinates that were previously or would have been incorrectly identified by the variant caller as indicating a homozygous reference genotype). Indeed, in some embodiments, the call recalibration system (i) utilizes a call generation model to generate initial nucleotide base calls, and (ii) utilizes a call recalibration machine learning model to correct data fields corresponding to the variant call file of nucleotide base calls. In some cases, the call recalibration system further corrects the nucleotide base calls based on one or more of the data fields and generates a variant call file having the corrected nucleotide base calls. In certain embodiments, the call recalibration system can utilize a call recalibration machine learning model to generate variant call classifications, while also utilizing a call generation model to generate nucleotide base calls based on the variant call classifications.
[0023] In contrast, in some cases, the call recalibration system determines a final nucleotide base call or variant call for a genomic coordinate based on both the sequencing metrics for the call generation model and the variant call classification from the call recalibration machine learning model, without an initial nucleotide base call (e.g., an initial nucleotide variant call) from the call generation model. For example, the call generation model may not output an initial nucleotide base call, but instead may evaluate the genomic coordinate and then generate sequencing metrics that the call recalibration machine learning model can use in combination with the call generation model to generate a variant call. In some embodiments, the call generation model may output a final variant call that takes into account the variant call classification from the call recalibration machine learning model (without generating an initial variant call that is updated). In contrast, in certain cases, the call generation model may initially determine that the confidence or quality corresponding to a potential variant call does not meet a threshold for inclusion in the variant call file, but may determine to include the variant call in the variant call file (after taking into account the variant call classification that updates the base call quality metrics). As a result of implementing the call recalibration machine learning model and the call generation model in this manner, the call recalibration system recovers false negative calls, corrects variant genotype errors, and / or removes false positive calls originally made by the call generation model.
[0024] In one or more embodiments, the call recalibration system further determines a contribution measure associated with one or more of the sequencing metrics. In particular, the call recalibration system determines a measure of the impact or influence that each sequencing metric or a subset of sequencing metrics has on the final nucleotide base call. For example, some metrics may be weighted more heavily than other metrics in determining a call at one genomic coordinate versus another. Indeed, due to the accessibility and interpretability of the call generation model and the call recalibration machine learning model, the call recalibration system can access the internal sequencing metrics used to generate the nucleotide base calls and determine the respective contribution measures in ultimately determining which metrics are causing or driving the recalibration of the final nucleotide base call (or variant call). In some cases, the call recalibration system also generates and provides a visualization of the contribution measures for display on the client device.
[0025] As alluded to above, the call recalibration system provides several advantages, benefits, and / or improvements over existing sequencing systems, including variant callers and other sequencing data analysis software. For example, the call recalibration system generates more accurate nucleotide base calls and / or variant calls than existing sequencing systems. While some existing sequencing systems are either unable to generate or generate inaccurate nucleotide base calls for multi-allelic coordinates, in some embodiments, the call recalibration system generates more accurate calls for multi-allelic genomic coordinates. Specifically, the call recalibration system can utilize or adapt a call recalibration machine learning model having parameters trained or tuned to generate a set of variant call classifications specific to the multi-allelic genomic coordinates. From the set of variant call classifications, the call recalibration system can further generate one or more final nucleotide base calls for the multi-allelic genomic coordinates to indicate the genotype of the multi-allelic coordinate, indicate whether the genotype is variant with respect to the reference sequence, and / or indicate whether the genotype is accurate (e.g., a genotype quality metric in a GQ field indicating the likelihood or probability that the genotype is accurate). Similarly, from a set of variant call classifications, the call recalibration system can also improve the accuracy of other fields, such as the quality field and the PL.
[0026] In some embodiments, the call recalibration system generates more accurate nucleotide base calls and / or variant calls for haploid coordinates of a sample nucleotide sequence compared to existing sequencing systems. Unlike some existing sequencing systems that cannot recalibrate nucleotide base calls for haploids, the call recalibration system can utilize a call recalibration machine learning model that can be adapted to haploid regions of a sample nucleotide sequence. In certain cases, the call recalibration system learns parameters for the call recalibration machine learning model by fitting diploid data to simulate haploid data. Furthermore, the call recalibration system can generate nucleotide base calls for haploid coordinates by reducing certain machine learning outputs (e.g., confidence scores) of the call recalibration machine learning model that are not related to haploid calls for a particular genomic coordinate and normalizing across the remaining two outputs (e.g., confidence scores). By reducing and normalizing the outputs that are adapted to diploid data to the outputs that are adapted to haploid data, the call recalibration system can determine the probability of indicating a haploid reference genotype and a haploid alternative genotype at the coordinate.
[0027] In one or more embodiments, the call recalibration system generates more accurate nucleotide base calls and / or variant calls for (otherwise incorrectly identified) homozygous reference coordinates of a sample nucleotide sequence compared to existing sequencing systems. For example, some existing sequencing systems generate an inordinate number of false negative variant calls by incorrectly identifying certain genomic coordinates as indicative of homozygous reference genotypes when in fact their genotypes are not homozygous references. In contrast, the call recalibration system identifies fewer false negative variant calls (or recovers more false negative variant calls) by determining sequencing metrics for genomic coordinates shown to be indicative of homozygous reference genotypes and utilizing a call recalibration machine learning model to generate variant call classifications for these coordinates. The call recalibration system can further generate one or more final nucleotide base calls from the variant call classifications of the homozygous reference coordinates.
[0028] The call recalibration system improves the accuracy of an existing sequencing system (e.g., in each of the scenarios described above) by utilizing a call recalibration machine learning model to remove a large number of false positive variant calls and / or recover a large number of false negative variant calls. By editing the initial nucleotide base calls or generating final nucleotide base calls based on variant call classifications from the call recalibration machine learning model, the call recalibration system can use unique machine learning output to recalibrate base calls with better accuracy than an existing variant caller or an existing machine learning model. For example, the call recalibration system utilizes a call recalibration machine learning model to generate variant call classifications from both internal (e.g., unique and model-specific) and external sequencing metrics, which results in recovery of previously filtered out variant nucleotide base calls and / or removal of previously unfiltered non-variant nucleotide base calls.
[0029] To achieve the aforementioned accuracy improvements, as shown, the call recalibration system utilizes an improved specific machine learning model - a call recalibration machine learning model - that is trained to perform the new application. Unlike existing variant callers that generate nucleotide base calls from general sequencing data (without any particular emphasis on one genomic coordinate or another), the call recalibration system utilizes a specific call recalibration machine learning model that generates specific variant call classifications for specific scenarios (e.g., multi-allelic genomic coordinates, haploid genomic coordinates, and pseudohomozygous reference coordinates). In some cases, the call recalibration system utilizes the call recalibration machine learning model to update the nucleotide base calls generated by the call generation model from the same metrics (or a subset of the same metrics) used by the call recalibration machine learning model to generate the variant call classifications.
[0030] Attributing at least in part to the improved accuracy, the call recalibration system exhibits improved flexibility over existing sequencing systems. For example, while many existing sequencing systems are limited to application at certain genomic coordinates and / or are incompatible with other genomic coordinates, in some embodiments the call recalibration system flexibly adapts to many of these previously incompatible coordinates. Specifically, unlike some existing sequencing systems, the call recalibration system can generate nucleotide base calls and / or variant calls for multi-allelic genomic coordinates, haploid genomic coordinates, and pseudohomozygous reference genomic coordinates.
[0031] As another example of improved flexibility, as mentioned above, existing sequencing systems may utilize variant callers that rely exclusively on internal sequencing metrics for a particular base call to generate nucleotide base calls, without re-engineering or modifying such internal sequencing metrics or analyzing external source sequencing metrics associated with the genomic coordinates of the corresponding nucleotide base call. In contrast, in some embodiments, the call recalibration system generates and manipulates both external and internal sequencing metrics. Indeed, in some cases, the call recalibration system determines the sequencing metrics of the call model generation from the variant caller component and the mapping and alignment component of the call generation model by efficiently combining a Bayesian probability model with machine learning techniques. In addition, the call recalibration system utilizes a call recalibration machine learning model to generate updated nucleotide base calls (e.g., from variant call classifications) from one or more sequencing metrics.
[0032] In addition to improving accuracy and flexibility, in certain embodiments, the call recalibration system improves efficiency and speed. As described above, some existing sequencing systems utilize computationally expensive and slow neural network architectures (e.g., deep learning architectures such as convolutional neural networks) that require many hours (e.g., 5-8 hours on multiple processors running on a server) and large amounts of computational resources even to implement and generate a file with variant calls from the sequencing run. Such deep learning architectures may further require days (or weeks) to train. Conversely, the call recalibration system utilizes relatively lightweight and fast architectures for both the call generation model and the call recalibration machine learning model. Indeed, in contrast to the many hours across multiple processors required by existing sequencing systems, the call recalibration system often requires less than 30 minutes of runtime (for both the call generation model and the call recalibration machine learning model together) on a single field programmable gate array or single processor to generate nucleotide base calls for a sample nucleotide sequence. Thus, the call recalibration system is much faster and less computationally expensive than many deep learning approaches to variant calling. Not only are the models of the Call Recalibration System faster and less computationally expensive to implement than many existing deep learning-based systems, the models of the Call Recalibration System are also much faster and less computationally expensive to train.
[0033] As part of the improved speed and efficiency, in some embodiments, the call recalibration system recalibrates the nucleotide base calls for each call as it is processed by the call generation model. In fact, the call recalibration system can generate variant call classifications for recalibrating the nucleotide base calls (e.g., utilizing a call recalibration machine learning model), while also generating nucleotide base calls from the variant call classifications along with one or more sequencing metrics. In some embodiments, the call recalibration system utilizes a call generation model in parallel with the call recalibration machine learning model to simultaneously generate initial nucleotide base calls and variant call classifications for correcting or recalibrating the initial nucleotide base calls.
[0034] As an additional advantage over existing sequencing systems, in certain implementations, the call recalibration system can identify or facilitate modifications to individual metrics that affect the accuracy of a nucleotide base call. While the neural network architecture of many existing sequencing systems does not allow any interpretation of the internal model data with latent features, the call recalibration system utilizes a model architecture that facilitates the interpretation of the effect of individual sequencing metrics. More specifically, in some cases, the call recalibration system utilizes a call generation model and a call recalibration machine learning model that allow the extraction and analysis of individual sequencing metrics used throughout the process of generating a nucleotide base call. In effect, the call recalibration system can determine the respective contribution measures for the sequencing metrics involved in determining a nucleotide base call at a particular genomic coordinate.
[0035] As suggested by the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of the call recalibration system. Further details regarding the meaning of these terms as used in the present disclosure are provided below. For example, the term "sample nucleotide sequence" or "sample sequence" as used in the present disclosure refers to a sequence of nucleotides isolated or extracted from a sample organism (or a copy of such an isolated or extracted sequence). In particular, the sample nucleotide sequence is isolated or extracted from a sample organism and includes a segment of a nucleic acid polymer composed of nitrogenous heterocyclic bases. For example, the sample nucleotide sequence can include a segment of deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acid or chimeric or hybrid forms of nucleic acid as described below. More specifically, in some cases, the sample nucleotide sequence is one that is found in a sample prepared or isolated by the kit and received by the sequencing device.
[0036] As further used herein, the term "nucleotide base call" (or sometimes simply "call") refers to the determination or prediction of a particular nucleotide base (or nucleotide base pair) for a genomic coordinate or oligonucleotide of a sample genome during a sequencing cycle. In particular, a nucleotide base call can refer to (i) a determination or prediction of the type of nucleotide base incorporated within an oligonucleotide on a nucleotide sample slide (e.g., a read-based nucleotide base call), or (ii) a determination or prediction of the type of nucleotide base present at a genomic coordinate or region in a sample genome, including a variant or non-variant call in a digital output file. In some cases, for a nucleotide read, a nucleotide base call includes a nucleotide base determination or prediction based on an intensity value resulting from a fluorescently tagged nucleotide added to an oligonucleotide on a nucleotide sample slide (e.g., in a well of a flow cell). Alternatively, a nucleic acid base call includes a nucleotide base determination or prediction to a chromatogram peak or current change resulting from a nucleotide passing through a nanopore of a nucleotide sample slide. In contrast, a nucleotide base call can also include an initial or final prediction of a nucleotide base at a genomic coordinate of a sample genome for a variant call file or other base call output file based on a nucleotide read corresponding to the genomic coordinate. Thus, a nucleotide base call can include a base call corresponding to a genomic coordinate and a reference genome, such as an indication of a variant or non-variant at a particular position corresponding to a reference genome. In practice, a nucleotide base call can refer to a variant call, including but not limited to a base call that is part of a single nucleotide polymorphism (SNP), an insertion or deletion (indel), or a structural variant. By using a nucleotide base call, a sequencing system determines the sequence of a nucleic acid polymer.For example, single nucleotide base calls can include adenine, cytosine, guanine, or thymine calls (abbreviated as A, C, G, T) for DNA, or uracil (instead of thymine) (abbreviated as U) for RNA.
[0037] Relatedly, as used herein, the term "nucleotide read" refers to a predicted sequence of one or more nucleotide bases (or nucleotide base pairs) from all or a portion of a sample nucleotide sequence. In particular, a nucleotide read includes a sequence of determined or predicted nucleotide base calls for a nucleotide fragment (or a group of monoclonal nucleotide fragments) from a sequencing library corresponding to a genomic sample. For example, a call recalibration system determines a nucleotide read by generating a nucleotide base call for a nucleotide base that has passed through a nanopore of a nucleotide sample slide, that has been determined via fluorescent tagging, or that has been determined from a well in a flow cell.
[0038] As described above, in some embodiments, the call recalibration system determines a sequencing metric for the nucleotide base call of the nucleotide read. As used herein, the term "sequencing metric" refers to a quantitative measurement or score that indicates the degree to which an individual nucleotide base call (or a sequence of nucleotide base calls) aligns, compares, or quantifies with respect to a genomic coordinate or genomic region of a reference genome, with respect to a nucleotide base call from a nucleotide read, or with respect to external genomic sequencing or genomic structure. For example, a sequencing metric includes a quantitative measurement or score that indicates the degree to which (i) an individual nucleotide base call aligns, maps, or covers a genomic coordinate or reference base of a reference genome, (ii) the degree to which a nucleotide base call compares with a reference or alternative nucleotide read in terms of mapping, mismatches, base call quality, or other raw sequencing metrics, or (iii) the degree to which a genomic coordinate or region corresponding to a nucleotide base call demonstrates mappability, repeat base call content, DNA structure, or other generalized metrics.
[0039] Relatedly, the term "diploid sequencing metrics" refers to sequencing metrics determined for a nucleotide base call at a diploid genomic coordinate. For example, diploid sequencing metrics include sequencing metrics for a particular genomic coordinate of a nucleotide sequence that is derived (or shown to be derived) from a diploid chromosome or a diploid nucleotide sequence (e.g., has two alleles at the genomic region corresponding to the genomic coordinate). Furthermore, the term "haploid sequencing metrics" refers to sequencing metrics determined for a nucleotide base call at a haploid genomic coordinate. For example, haploid sequencing metrics include sequencing metrics for a particular genomic coordinate of a nucleotide sequence that is derived (or shown to be derived) from a haploid chromosome or a haploid nucleotide sequence (e.g., has a single allele at the genomic region corresponding to the genomic coordinate).
[0040] As further used herein, the term "genomic coordinate (or sometimes simply "coordinate")" refers to a specific location or position of a nucleotide base within a genome (e.g., the genome of an organism or a reference genome). In some cases, the genomic coordinate includes an identifier for a particular chromosome of the genome and an identifier for a nucleotide base position within the particular chromosome. For example, the genomic coordinate(s) may include a chromosome number, name, or other identifier (e.g., chr1 or chrX) and a specific position(s) such as a numbered position (e.g., chr1:1234570 or chr1:1234570-1234870) following the chromosome identifier. Furthermore, in certain implementations, the genomic coordinate refers to the source of the reference genome (e.g., mt for a mitochondrial DNA reference genome, or SARS-CoV-2 for a reference genome of the SARS-CoV-2 virus), and the nucleotide base position within the source for the reference genome (e.g., mt:16568 or SARS-CoV-2:29001). In contrast, in certain cases, genomic coordinates refer to nucleotide base positions within a reference genome, without reference to a chromosome or source (e.g., 29727).
[0041] Relatedly, as used herein, the term "multi-allelic genomic coordinate" refers to a genomic coordinate associated with three or more alleles. For example, a multi-allelic genomic coordinate includes a genomic coordinate of a nucleotide sequence, where the nucleotide read indicates three or more possible alleles (e.g., a reference allele, a first alternative allele, a second alternative allele, etc.) corresponding to the coordinate. In some cases, a multi-allelic genomic coordinate corresponds to a genomic coordinate where a read pile-up occurs or an insertion occurs. For example, a multi-allelic genomic coordinate can indicate a multi-allelic genotype (e.g., a 1 / 2 genotype), where the first allele at the coordinate corresponds to an allele from a first alternative nucleotide sequence and the second allele corresponds to an allele from a second alternative nucleotide sequence.
[0042] As mentioned above, in some embodiments, the call recalibration system generates nucleotide base calls for haploid genomic coordinates or genomic coordinates within a haploid nucleotide sequence. As used herein, the term "haploid nucleotide sequence" refers to a sequence of one or more nucleotide bases from a haploid chromosome (e.g., a sex chromosome in a male) or a single chromosome without a corresponding chromosome. For example, a haploid nucleotide sequence can include a haploid region of a sample nucleotide sequence in which each of the genomic coordinates covers nucleotide bases from a haploid chromosome or a single chromosome without a corresponding chromosome. Thus, a haploid coordinate within a haploid nucleotide sequence has a haploid genotype, such as a haploid reference genotype (0) or a haploid alternative genotype (1).
[0043] Other coordinates in the nucleotide sequence may indicate different genotypes. For example, a "homozygous reference genotype" refers to a genotype (represented as 0 / 0) in which both nucleotide bases at a given coordinate of a sample nucleotide sequence match the reference nucleotide base of a reference sequence or reference genome. As another example, a "homozygous alternative genotype" refers to a genotype (represented as 1 / 1) in which both nucleotide bases at a given coordinate are different from the reference nucleotide base of a reference sequence or reference genome. As a further example, a "heterozygous genotype" refers to a genotype in which the nucleotide bases at a given coordinate are not the same. In some cases, a heterozygous genotype includes a genotype (represented as 0 / 1 or 1 / 0) in which one nucleotide base matches the reference nucleotide base and the other nucleotide base is different from the reference nucleotide base of a reference genome. For a multi-allelic genomic coordinate, the genotype may indicate a nucleotide base derived from two or more alternative nucleotide bases that are different from the reference nucleotide base of a reference genome. For example, a multi-allelic heterozygous genotype can be represented as 1 / 2, where one nucleotide base call corresponds to a first alternative nucleotide base that differs from the reference nucleotide base, and the other nucleotide base call corresponds to a second alternative nucleotide base that differs from the reference nucleotide base.
[0044] As described above, a genome coordinate includes a location within a reference genome. Such a location may be within a particular reference genome. As used herein, the term "reference genome" refers to a digital nucleic acid sequence assembled as a representative (or representatives) of genes and other gene sequences of an organism. Regardless of sequence length, in some cases, a reference genome represents an exemplary set of genes or a set of nucleic acid sequences in a digital nucleic acid sequence that have been determined by scientists as representative of a particular species of organism. For example, a linear human reference genome may be GRCh38 or other versions of the reference genome from the Genome Reference Consortium. As a further example, a reference genome may include a reference graph genome, such as Illumina DRAGEN Graph Reference Genome hg19, that includes both a linear reference genome and paths that represent nucleic acid sequences from ancestral haplotypes.
[0045] In some embodiments, the call recalibration system determines various types of sequencing metrics from different sources, such as read-based sequencing metrics, external source sequencing metrics, and call model generation sequencing metrics. As used herein, the term "read-based sequencing metrics" refers to sequencing metrics derived from nucleotide reads of a sample nucleotide sequence. For example, read-based sequencing metrics include sequencing metrics determined by applying statistical tests to detect differences between a reference sequence and a nucleotide read. For example, read-based sequencing metrics can include comparative mapping quality distribution metrics indicating a comparison between mapping qualities, or comparative mismatch count metrics indicating a comparison between mismatch counts.
[0046] In contrast, "external source sequencing metrics" refer to sequencing metrics identified or obtained from one or more external databases. For example, external source sequencing metrics include metrics related to nucleotide mappability, replication timing, or DNA structure that are available outside the call recalibration system.
[0047] Further, "sequencing metrics of the call model generation" refers to internal model-specific sequencing metrics generated or extracted by the call generation model. For example, the sequencing metrics of the call model generation include variant calling sequencing metrics extracted or determined via the variant caller component of the call generation model, and mapping and alignment sequencing metrics extracted or determined via the mapping and alignment component of the call generation model. As indicated above, the sequencing metrics of the call model generation can include alignment metrics, such as deletion size metrics or mapping quality metrics, that quantify the degree to which the sample nucleic acid sequence aligns with the genomic coordinates of the exemplary nucleic acid sequence. Furthermore, the sequencing metrics of the call model generation can include depth metrics, such as forward-reverse depth metrics or normalized depth metrics, that quantify the depth of the nucleotide base calls for the sample nucleic acid sequence at the genomic coordinates of the exemplary nucleic acid sequence. The sequencing metrics of the call model generation can also include call quality metrics, such as nucleotide base call quality metrics, callability metrics, or somatic quality metrics, that quantify the quality or accuracy of the nucleotide base calls.
[0048] As used herein, the term "base call quality metric" refers to a particular score or other measure that indicates the accuracy of a nucleotide base call. In particular, a base call quality metric includes a value that indicates the likelihood that one or more predicted nucleotide base calls for a genomic coordinate will contain an error. For example, in certain implementations, a base call quality metric can include a Q-score (e.g., a Phred quality score) that predicts the probability of error for any given nucleotide base call. To illustrate, a quality score (or Q-score) can indicate that the probability of an incorrect nucleotide base call at a genomic coordinate is equal to 1 in 100 for a Q20 score, 1 in 1,000 for a Q30 score, 1 in 10,000 for a Q40 score, etc.
[0049] Relatedly, as used herein, the term "re-engineered sequencing metrics" refers to sequencing metrics that have been updated, modified, augmented, improved, or re-engineered to measure or compare nucleotide base calls (e.g., nucleotide base calls or variant calls for a read) with respect to other nucleotide base calls (standards or references) or targeted to a particular purpose or task. For example, re-engineered sequencing metrics can include modifications to raw sequencing metrics or combinations of raw sequencing metrics. In some embodiments, for example, the call recalibration system generates one or more of the read-based sequencing metrics, external source sequencing metrics, and / or call model generation sequencing metrics as re-engineered sequencing metrics. In some cases, re-engineered sequencing metrics refer to sequencing metrics that are generated by the call recalibration system and thus are unique or internal to the call recalibration system and are not available to third party systems. Exemplary re-engineered sequencing metrics include a comparative mapping quality distribution metric indicating a comparison between mapping quality distributions associated with the reference sequence and the alternative supporting nucleotide read, or a comparative base quality metric indicating a comparison between the base qualities of the reference sequence and the alternative supporting nucleotide read.
[0050] As alluded to above, the call recalibration system can utilize machine learning models to correct sequencing metrics and update nucleotide base calls. As used herein, the term "machine learning model" refers to a computer algorithm or collection of computer algorithms that automatically improves for a particular task through experience based on the use of data. For example, the machine learning model can utilize one or more learning techniques to improve accuracy and / or effectiveness. Exemplary machine learning models include various types of decision trees, support vector machines, Bayesian networks, or neural networks. In some cases, the call recalibration machine learning model is a series of gradient boosting decision trees (e.g., XGBoost algorithm), and in other cases, the call recalibration machine learning model is a random forest model, a multi-layer perceptron, linear regression, support vector machine, a deep table learning architecture, a deep learning transformer (e.g., a self-attention based table transformer), or a logistic regression.
[0051] In some cases, the call recalibration system utilizes a call recalibration machine learning model to correct or update nucleotide base calls based on sequencing metrics. As used herein, the term "call recalibration machine learning model" refers to a machine learning model that generates variant call classifications. For example, in some cases, the call recalibration machine learning model is trained to generate variant call classifications that indicate various probabilities or predictions of variant calls based on sequencing metrics. Thus, in some cases, the call recalibration machine learning model is a variant call recalibration machine learning model. In certain embodiments, the call recalibration machine learning model includes multiple sub-models or works in conjunction with another call recalibration machine learning model. For example, a first call recalibration machine learning model (e.g., an ensemble of gradient boosted trees) generates a first set of variant call classifications, and a second call recalibration machine learning model (e.g., a random forest) generates a second set of variant call classifications.
[0052] Relatedly, the term "variant call classification" refers to a predicted classification from a call recalibration machine learning model that indicates a probability, score, or other quantitative measure associated with some aspect of a nucleotide base call based on one or more sequencing metrics. The variant call classification can include a special prediction depending on the application of the call recalibration machine learning model. In an embodiment for generating a nucleotide base call (or variant call) for a multi-allelic genomic coordinate, the variant call classification can include (i) a reference probability of a homozygous reference genotype at the multi-allelic genomic coordinate, (ii) a different genotype probability of a genotype error at the multi-allelic genomic coordinate, and (iii) an accurate variant probability of an accurate variant call genotype at the multi-allelic genomic coordinate.
[0053] In embodiments for generating nucleotide base calls (or variant calls) for a haploid genomic coordinate, the variant call classification may include (i) a first genotype probability of a first genotype at the genomic coordinate and (ii) a second genotype probability of a second genotype at the genomic coordinate. As alluded to above, the first genotype probability may be the probability that the genotype at the genotypic coordinate is a haploid reference genotype and the second genotype probability may be the probability that the genotype at the genotypic coordinate is a haploid alternative genotype. In these or other embodiments, such as those for generating nucleotide base calls (or variant calls) for genomic coordinates that have been shown to indicate a homozygous reference genotype, the variant call classification may include (i) a false positive classification or a homozygous reference classification indicating the probability that the nucleotide base call is a false positive or homozygous reference genotype, respectively, (ii) a genotype error classification or a heterozygous genotype classification indicating the probability that the genotype (e.g., an indication of a heterozygous or homozygous genotype for a variant call at a particular position) is an incorrect genotype or a heterozygous genotype, respectively, and / or (iii) a true positive classification or a homozygous alternative classification indicating the probability that the nucleotide base call is a true positive or a homozygous alternative genotype, respectively. Thus, in some cases, the variant call classification represents an intermediate scoring metric and / or a predicted probability that the genotype for a nucleotide base call is correct.
[0054] As mentioned, in some embodiments, the call recalibration machine learning model can be a neural network. The term "neural network" refers to a machine learning model that can be trained and / or adjusted based on inputs to determine classification or approximate unknown functions. For example, a neural network includes a model of interconnected artificial neurons (e.g., organized into layers) that learn to communicate, approximate complex functions, and generate outputs (e.g., generated digital images) based on multiple inputs provided to the neural network. In some cases, a neural network refers to an algorithm (or set of algorithms) that implements deep learning techniques to model high-level abstractions in data. For example, a neural network can include a convolutional neural network, a recurrent neural network (e.g., LSTM), a graph neural network, a self-attention transform neural network, or a generative adversarial neural network.
[0055] As described above, the call recalibration system can generate variant call classifications that indicate or reflect the likelihood of identifying a variant at a genomic coordinate. As used herein, the term "variant" refers to a nucleotide base or multiple nucleotide bases that do not align with, differ from, or change from the corresponding nucleotide base (or multiple nucleotide bases) in a reference sequence or reference genome. For example, variants include SNPs, indels, or structural variants that indicate a nucleotide base in a sample nucleotide sequence that is different from the nucleotide base at the corresponding genomic coordinate of the reference sequence. Along these lines, a "variant nucleotide base call" (or simply "variant call") refers to a nucleotide base call that includes a variant at a particular genomic coordinate. Conversely, a "non-variant nucleotide base call" (or simply "non-variant call") refers to a nucleotide base call that includes a non-variant at a genomic coordinate.
[0056] As mentioned, in some embodiments, the call recalibration system modifies data fields corresponding to the variant call file. As used herein, the term "variant call file" refers to a digital file that indicates or represents one or more nucleotide base calls (e.g., variant calls) compared to a reference genome along with other information about the nucleotide base calls (e.g., variant calls). For example, a variant call format (VCF) file refers to a text file format that has information about variants at specific genomic coordinates, including a meta information row, a header row, and a data row, where each data row has information about a single nucleotide base call (e.g., a single variant). As described further below, the call recalibration system can generate different versions of the variant call file, including a pre-filter variant call file that includes variant nucleotide base calls that pass or do not pass a quality filter for a base call quality metric, or a post-filter variant call file that includes variant nucleotide base calls that pass a quality filter but exclude variant nucleotide base calls that do not pass the quality filter.
[0057] In some embodiments, the call recalibration system modifies data fields corresponding to metrics of nucleotide base calls associated with the variant call file, such as fields for call quality, genotype, and genotype quality. As used herein, the term "call quality," when used with respect to a data field in a variant call file, refers to a measure or index of the likelihood or probability that a variant is present at a given location. Thus, a call quality field (or QUAL field) corresponding to a VCF file may include a base call quality metric, such as a Phred scale quality or Q score, that represents the probability that a genomic coordinate of a sample genome contains a variant. Similarly, "genotype quality," when used with respect to a field, refers to the likelihood or probability that a particular predicted genotype for a nucleotide base call is accurate.
[0058] As described, in some embodiments, the call recalibration system utilizes a call generation model to generate nucleotide base calls for genomic coordinates. As used herein, the term "call generation model" refers to a probabilistic model that generates sequencing data from nucleotide reads of a sample nucleotide sequence, including nucleotide base calls and associated metrics. Thus, in some cases, the call generation model may be a variant call generation model. For example, in some cases, the call generation model refers to a Bayesian probability model that generates variant calls based on nucleotide reads of a sample nucleotide sequence. Such models can process or analyze sequencing metrics corresponding to read pileups (e.g., multiple nucleotide reads corresponding to a single genomic coordinate), including mapping quality, base quality, and various hypotheses including extraneous reads, missing reads, joint detection, and the like. The call generation model may also include multiple components, including, but not limited to, different software applications or components for mapping and alignment, sorting, duplicate marking, calculation of read pileup depth, and variant calling. In some cases, the call generation model refers to ILLUMINA DRAGEN models for variant calling functions and mapping and alignment functions.
[0059] As mentioned above, in certain described embodiments, the call recalibration system generates or determines a contribution measure associated with each sequencing metric. As used herein, the term "contribution measure" refers to a measure of the effect, influence, or impact that a sequencing metric has on a base call output file (e.g., a variant call file), a given nucleotide base call in a base call output file, or (particularly) a given recalibration of a field for a given variant call. For example, the contribution measure indicates how much of a role one sequencing metric plays over a different nucleotide base call (and compared to other sequencing metrics) in determining a nucleotide base call.
[0060] The following paragraphs describe the call recalibration system with respect to example diagrams depicting example embodiments and implementations. For example, Figure 1 shows a schematic diagram of a system environment (or "environment") 100 in which a call recalibration system 106 operates, according to one or more embodiments. As shown, the environment 100 includes one or more server devices 102 connected to user client devices 108 and a sequencing device 114 via a network 112. While Figure 1 illustrates one embodiment of the call recalibration system 106, this disclosure below describes alternative embodiments and configurations.
[0061] 1, server device 102, client device 108, and sequencing device 114 may communicate with each other via network 112. Network 112 may include any suitable network with which computing devices may communicate. An exemplary network is discussed in further detail below with respect to FIG.
[0062] As illustrated by FIG. 1, the sequencing device 114 includes a device for sequencing nucleic acid polymers. In some embodiments, the sequencing device 114 analyzes nucleic acid segments or oligonucleotides extracted from a genomic sample to generate nucleotide reads or other data using computer-implemented methods and systems (described herein) either directly or indirectly on the sequencing device 114. More specifically, the sequencing device 114 receives and analyzes nucleic acid sequences extracted from a sample in a nucleotide sample slide (e.g., a flow cell). In one or more embodiments, the sequencing device 114 utilizes SBS to sequence the nucleic acid polymer into nucleotide reads. In some embodiments, the sequencing device 114, in addition to or as an alternative to communicating via the network 112, bypasses the network 112 and communicates directly with the client device 108.
[0063] As further illustrated by Figure 1, the server device 102 can generate, receive, analyze, store, and transmit digital data, such as data for determining nucleotide base calls or for sequencing a nucleic acid polymer. As illustrated in Figure 1, the sequencing device 114 can transmit call data (and the server device 102 can receive call data) from the sequencing device 114. The server device 102 can also communicate with the client device 108. In particular, the server device 102 can transmit data to the client device 108, including variant call files or other information indicative of nucleotide base calls, sequencing metrics, error data, or other metrics associated with the nucleotide base calls.
[0064] In some embodiments, server device 102 comprises a distributed collection of servers, where server device 102 includes several server devices distributed across network 112 and located at the same or different physical locations. Also, server device 102 may include a content server, an application server, a communication server, a web hosting server, or another type of server. In some cases, server device 102 may be located at the same physical location as sequencer 114.
[0065] As further shown in FIG. 1 , the server device 102 can include a sequencing system 104. Generally, the sequencing system 104 analyzes call data, such as sequencing data received from the sequencing device 114, to determine a nucleotide-based sequence for the nucleic acid polymer. For example, the sequencing system 104 can receive raw data from the sequencing device 114 and determine a nucleotide-based sequence for the nucleic acid segment. In some embodiments, the sequencing system 104 determines a nucleotide-based sequence in a DNA and / or RNA segment or oligonucleotide. In addition to processing and determining a sequence for the nucleic acid polymer, the sequencing system 104 also generates a variant call file indicating one or more nucleotide base calls and / or variant calls for one or more genomic coordinates.
[0066] As just described and illustrated in FIG. 1 , the call recalibration system 106 analyzes call data, such as sequencing metrics from the sequencer 114, to determine nucleotide base calls for sample nucleic acid sequences. The call recalibration system 106 includes a call generation model and a call recalibration machine learning model. In some embodiments, the call recalibration system 106 determines sequencing metrics for the sample nucleotide sequences. Based on data derived or prepared from the sequencing metrics, the call recalibration system 106 trains and applies the call generation model to determine nucleotide base calls for the sample sequences corresponding to the genomic coordinates. The call recalibration system 106 further utilizes the call recalibration machine learning model to generate a set of variant call classifications to update or correct the nucleotide base calls (and / or variant calls). Based on such data, for example, the call recalibration system 106 can update data fields corresponding to a variant call file to update the nucleotide base calls and / or variant calls to improve accuracy.
[0067] As further illustrated and shown in FIG. 1, the client device 108 can generate, store, receive, and transmit digital data. In particular, the client device 108 can receive sequencing metrics from the sequencing device 114. Additionally, the client device 108 can communicate with the server device 102 to receive variant call files including nucleotide base calls and / or other metrics such as call quality, genotype index, and genotype quality. Thus, the client device 108 can present or display information regarding the nucleotide base calls in a graphical user interface to a user associated with the client device 108. For example, the client device 108 can present a contribution measure interface that includes a visualization or depiction of various contribution measures associated with or resulting from individual sequencing metrics for a particular nucleotide base call.
[0068] The client devices 108 illustrated in Figure 1 can include various types of client devices. For example, in some embodiments, the client devices 108 include non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client devices 108 include mobile devices, such as laptops, tablets, cell phones, or smartphones. Further details regarding the client devices 108 are discussed below with respect to Figure 15.
[0069] 1, the client device 108 includes a sequencing application 110. The sequencing application 110 can be a web application or a native application (e.g., a mobile application, a desktop application) that is stored and executed on the client device 108. The sequencing application 110 can include instructions that (when executed) cause the client device 108 to receive data from the call recalibration system 106 and present data from a variant call file for display at the client device 108. Additionally, the sequencing application 110 can instruct the client device 108 to display a visualization of the contribution measures for the sequencing metrics of the nucleotide base calls.
[0070] 1, the call recalibration system 106 may be located on the client device 108 or on the sequencing device 114 as part of the sequencing application 110. Thus, in some embodiments, the call recalibration system 106 is implemented (e.g., fully or partially located) on the client device 108. In yet other embodiments, the call recalibration system 106 is implemented by one or more other components of the environment 100, such as the sequencing device 114. In particular, the call recalibration system 106 may be implemented in a variety of different ways across the server device 102, the network 112, the client device 108, and the sequencing device 114. For example, the call recalibration system 106 may be downloaded from the server device 102 to the client device 108 and / or the sequencing device 114, with all or a portion of the functionality of the call recalibration system 106 being implemented on the respective devices in the environment 100.
[0071] As further illustrated in FIG. 1 , the environment 100 includes a database 116. The database 116 can store information such as variant call files, sample nucleotide sequences, nucleotide reads, nucleotide base calls, variant calls, and sequencing metrics. In some embodiments, the server device 102, the client device 108, and / or the sequencing device 114 communicate (e.g., via the network 112) with the database 116 to store and / or access information such as variant call files, sample nucleotide sequences, nucleotide reads, nucleotide base calls, variant calls, and sequencing metrics. In some cases, the database 116 also stores one or more models, such as a call recalibration machine learning model and / or a call generation model.
[0072] 1 illustrates the components of environment 100 communicating over network 112, in certain implementations, the components of environment 100 may also communicate directly with one another, bypassing the network. For example, as discussed above, in some implementations, client device 108 may communicate directly with sequencer 114. Additionally, in some embodiments, client device 108 communicates directly with call recalibration system 106. Furthermore, call recalibration system 106 may access one or more databases housed in or accessed by server device 102 or elsewhere in environment 100.
[0073] As indicated above, the call recalibration system 106 can determine nucleotide base calls based on one or more variant call classifications. In particular, the call recalibration system 106 can utilize a call recalibration machine learning model to determine variant call classifications from sequencing metrics, and can determine or update various metrics associated with nucleotide base calls from the generated variant call classifications. Figure 2 shows an exemplary overview of determining nucleotide base calls based on variant call classifications, according to one or more embodiments.
[0074] As illustrated in FIG. 2, the call recalibration system 106 performs operation 202 to determine sequencing metrics. In particular, the call recalibration system 106 determines sequencing metrics, such as read-based sequencing metrics, external source sequencing metrics, and call model generated sequencing metrics. For example, the call recalibration system 106 determines sequencing metrics indicative of various attributes or data regarding various nucleotide base calls of nucleotide reads from a sample nucleotide sequence. Further details regarding determining various types of sequencing metrics are provided below with reference to FIGS. 6A-6C.
[0075] As further illustrated in FIG. 2, the call recalibration system 106 performs operation 204 to generate variant call classifications. More specifically, the call recalibration system 106 utilizes a call recalibration machine learning model to generate (or update or refine) variant call classifications from sequencing metrics. In particular, the call recalibration system 106 utilizes a call recalibration machine learning model to process or analyze one or more sequencing metrics and generate a set of classifications (e.g., predicted probabilities associated with genotypes). For example, the call recalibration system 106 utilizes a call recalibration machine learning model to generate a set of variant call classifications (represented in FIG. 2 as “Class 1,” “Class 2,” and “Class 3”) that indicate certain probabilities associated with genotypes of corresponding nucleotide base calls based on the sequencing metrics.
[0076] In some embodiments, the call recalibration system 106 generates different variant call classifications for different applications and / or different genomic coordinates. For example, the call recalibration system 106 generates a first set of variant call classifications for multi-allelic genomic coordinates, a second set of variant call classifications for haploid genomic coordinates, and a third set of variant call classifications for genomic coordinates for which homozygous reference genotypes are indicated. In certain embodiments, the call recalibration system 106 generates the same variant call classifications for different applications and / or different genomic coordinates, but utilizes them separately or utilizes different information associated with the variant call classifications. Further details regarding generating variant call classifications are provided below with reference to subsequent figures.
[0077] 2, the call recalibration system 106 also performs an operation 206 to determine a final nucleotide base call (or variant call) based on the variant call classification. More specifically, the call recalibration system 106 determines or updates a nucleotide base call for a sample nucleotide sequence at a genomic coordinate in a reference genome. To determine or generate a final nucleotide base call, in some embodiments, the call recalibration system 106 utilizes a call generation model to determine an initial nucleotide base call and edits or updates certain initial nucleotide base calls based on the variant call classification generated by the call recalibration machine learning model.
[0078] In particular, the call recalibration system 106 utilizes a call generation model to process or analyze sequencing metrics (e.g., one or more of the same sequencing metrics used to generate the variant call classification in operation 204) and determine nucleotide base calls (e.g., initial nucleotide base calls) from the sequencing metrics. For example, the call recalibration system 106 applies a number of Bayesian probability models or algorithms to derive various probabilities for different nucleotide bases, quality metrics, mapping metrics, joint metrics, and other data occurring within the sample nucleotide sequence for inclusion within the variant call file. From the probability models, the call recalibration system 106 determines nucleotide base calls indicative of predicted nucleotide bases of the sample genome at corresponding genomic coordinates (e.g., calls indicative of differences or identities relative to a reference base from a reference genome).
[0079] 2, in certain implementations, the call recalibration system 106 utilizes the initial variant call classifications (e.g., as determined via operation 204) to generate, recalibrate, determine, modify, or augment nucleotide base calls. In particular, the call recalibration system 106 utilizes the probabilities associated with the variant call classifications to determine or update certain metrics associated with the nucleotide base calls. For example, the call recalibration system 106 modifies data fields corresponding to the variant call file for metrics such as call quality, genotypes, and genotype quality (or others described below).
[0080] In some cases, the call recalibration system 106 extrapolates from the variant call classification to determine metrics corresponding to the variant call file, such as call quality, genotype, and genotype quality associated with the nucleotide base call. For example, by utilizing the genotype error classification, the call recalibration system 106 can repair certain errors in or associated with the initial nucleotide base call. Indeed, if the call recalibration system 106 determines a high false positive probability for a nucleotide base call, the call recalibration system 106 applies a call recalibration machine learning model to act as a variant filter to modify (e.g., reduce) the call quality associated with the nucleotide base call. As another example, the call recalibration system 106 utilizes the genotype error probability to modify the genotype and / or genotype quality of the nucleotide base call if the system previously filtered out or doubly penalized heterozygous / homozygous (het / hom) errors (e.g., if the system generated an incorrect nucleotide base call that would have missed a more accurate nucleotide base call).
[0081] In certain embodiments, the call recalibration system 106 considers a single variant call classification to modify the data fields for a nucleotide base call (e.g., call quality, genotype, or genotype quality). In other embodiments, the call recalibration system 106 considers multiple variant call classifications at once (e.g., in a weighted combination) to modify or update one or more data fields for call quality, genotype, and / or genotype quality. Further details regarding the generation and modification of nucleotide base calls are provided below with reference to subsequent figures.
[0082] In one or more implementations, the call recalibration system 106 generates variant call classifications (e.g., via operation 204) during or during the process of determining nucleotide base calls. For example, the call recalibration system 106 simultaneously implements the call recalibration machine learning model and the call generation model to generate nucleotide base calls and variant call classifications for correcting the nucleotide base calls. The call recalibration system 106 further corrects data fields corresponding to the variant call files of the nucleotide base calls to generate final nucleotide base calls (e.g., in the pre-filter or post-filter variant call files). In effect, the call recalibration system 106 generates final (e.g., recalibrated) nucleotide base calls from the variant call classifications and sequencing metrics processed by the call generation model (e.g., one or more of the same sequencing metrics used to generate the variant call classifications). As explained above, this simultaneous or parallel operation provides improved computational efficiency and increased speed to the call recalibration system 106 by recalibrating nucleotide base calls as they are initially generated (rather than performing one operation before the other).
[0083] In one or more implementations, the call recalibration system 106 determines the nucleotide base call as a part of a SNP, deletion, insertion, or structural variation. For example, the call recalibration system 106 determines that the nucleotide base call represents a SNP at a genomic coordinate (e.g., chr1:151863125) by identifying a G in the sample nucleotide sequence where an A is present in the reference sequence. As another example, the call recalibration system 106 determines that the nucleotide base call around one or more genomic coordinates (e.g., chr1:49263256) represents a deletion by identifying a single G in the sample nucleotide sequence where GTAAC is present in the reference sequence.
[0084] As a further example, the call recalibration system 106 determines that a sequence of nucleotide base calls represents an insertion at a genomic coordinate (e.g., chr1:7602080) by identifying a sequence of TTTCC in the sample nucleotide sequence where a T is present in the reference sequence. Indeed, in some cases, the insertion comprises a sequence of nucleotide base calls that replaces a single reference base in the genomic coordinate of the reference sequence.
[0085] In some embodiments, the call recalibration system 106 sets quality thresholds (e.g., customized quality thresholds) for base call quality metrics at genomic coordinates of the genomic sample (e.g., including one or more of diploid coordinates, haploid coordinates, multi-allelic coordinates, and genomic coordinates incorrectly identified as indicative of a homozygous reference genotype). The base call quality metrics may vary significantly between the call generation model and the call recalibration machine learning model. To accommodate for the potentially wide range and significant changes in base call quality metrics, the call recalibration system 106 can determine or set a stringent filter QUAL threshold for the variant call file output that results in (or corresponds to) a preferred F1 position as a measure of performance (e.g., a preferred tradeoff between false positives and false negatives).
[0086] Such a preferred F1 position may include a score or position with a preferred (e.g., best) tradeoff between precision and recall of calling variants. In some cases, for example, the F1 position (or F1 score) is proportional to the combination (e.g., sum) of false positive and false negative variants (meaning that a preferred F1 score corresponds to a low FP+FN metric). As described below, for example, FIG. 10B shows an example of an FP+FN metric based on which the call recalibration system 106 sets a quality threshold for the base call quality metric at haploid genome coordinates. Thus, in some embodiments, the call recalibration system 106 utilizes a quality filter for the base call quality metric that results in a preferred F1 position for the call recalibration system 1 or 2 shown in FIG. 10B.
[0087] However, as indicated above, the call recalibration system 106 may set such quality thresholds for base call quality metrics at any or all genomic coordinates that result in favorable F1 positions when using a call recalibration machine learning model. Indeed, in some embodiments, the call recalibration system 106 generates F1 scores and applies filtering logic associated with QUAL scores (as described above) to various genomic coordinates, including haploid coordinates, diploid coordinates, multi-allelic coordinates, genomic coordinates incorrectly identified as indicating a homozygous reference genotype, or other genomic coordinates.
[0088] Thus, in some cases, rather than discarding certain variant nucleotide base calls for which the call generation model does not pass a previous quality filter, the call recalibration system 106 performs the following series of operations: (i) utilizing a call generation model to generate variant nucleotide base calls across various regions or coordinates; (ii) utilizing a call recalibration machine learning model to recalibrate the variant nucleotide base calls and corresponding metrics (e.g., one or more of a base call quality metric, a genotype quality metric, or a genotype metric in the corresponding VCF field); (iii) generating a pre-filtered VCF that includes variant nucleotide base calls that are above a quality threshold, either because the call generation model called a variant nucleotide base call for the corresponding genomic coordinate or because the call recalibration machine learning model called a variant nucleotide base call for a genomic coordinate for which the call generation model determined that the variant nucleotide base call, etc., did not pass a previous quality filter; and (iv) utilizing a stringent quality threshold filter to select quality variant nucleotide base calls from the pre-filtered VCF. Such a stringent quality threshold is configured such that the filtered output of the call recalibration system 106 is close to the preferred F1 position (resulting in a post-filter VCF that includes only variant nucleotide base calls that meet the stringent quality threshold). The call recalibration system 106 can vary the QUAL threshold depending on whether a call recalibration machine learning model is active or whether the call generation model (e.g., DRAGEN) is running without a call recalibration machine learning model.
[0089] As mentioned above, in certain described embodiments, the call recalibration system 106 generates variant call classifications for multi-allelic genomic coordinates. In addition, the call recalibration system 106 generates or updates a variant call file for the multi-allelic coordinates based on the variant call classifications. Figures 3A-3B show an exemplary flow of the call recalibration system 106 generating a variant call file from variant call classifications of multi-allelic genomic coordinates, according to one or more embodiments. For example, Figure 3A shows the call recalibration system 106 generating variant call classifications for multi-allelic genomic coordinates, according to one or more embodiments. Then, Figure 3B shows the call recalibration system 106 generating a variant call file from the variant call classifications, according to one or more embodiments.
[0090] As illustrated in FIG. 3A, the call recalibration system 106 identifies the multi-allelic genomic coordinate 302. For example, the call recalibration system 106 identifies the multi-allelic genomic coordinate 302 from nucleotide base calls corresponding to a sample nucleotide sequence or based on haplotype data corresponding to the multi-allelic genomic coordinate 302. In some cases, the call recalibration system 106 identifies the multi-allelic genomic coordinate 302 by determining that (i) the nucleotide base calls from the nucleotide reads covering the genomic coordinate include three or more possible nucleotide base calls from the corresponding alleles, and (ii) the nucleotide base calls meet one or more threshold sequencing metrics (e.g., a base call quality metric of Q30). Additionally or alternatively, in certain embodiments, the call recalibration system 106 identifies the genomic coordinate from a database that includes a haplotype reference panel that correlates with the particular genomic coordinate. Based on the different haplotype probabilities in the database, the call recalibration system 106 identifies the genomic coordinate as a possible candidate for the multi-allelic genomic coordinate. Regardless of the identification method, in some cases, the call recalibration system 106 uses a call generation model (eg, a variant caller within the call generation model) to identify multi-allelic genomic coordinates 302.
[0091] As shown in FIG. 3A, for example, the call recalibration system 106 analyzes nucleotide base sequences corresponding to coordinates 1-5 to identify genomic coordinate 4 as a multi-allelic genomic coordinate 302. For ease of explanation, each of the nucleotide base sequences maps to a different allele and constitutes a representative of a different nucleotide read corresponding to coordinates 1-5. Although coordinates 1-5 each have three possible alleles corresponding to them (one from the reference genome, one from a first possible allele "alternative allele 1", and one from a second possible allele "alternative allele 2"), only coordinate 4 exhibits three or more different nucleotide base calls from different alleles that can potentially be assigned. Specifically, coordinate 4 may exhibit a G from the reference genome, a C from the first alternative allele, or a T from the second alternative allele.
[0092] As further illustrated in Figure 3A, the call recalibration system 106 determines sequencing metrics 304 for the multi-allelic genomic coordinates 302. In particular, the call recalibration system 106 determines sequencing metrics associated with nucleotide reads generated by the call generation model or retrieved from an external source. Further details regarding determining the sequencing metrics 304 are provided below with particular reference to Figures 6A-6C.
[0093] In addition, as shown in FIG. 3A, the call recalibration system 106 utilizes a call recalibration machine learning model 306 to generate a variant call classification 308. Specifically, the call recalibration system 106 utilizes the call recalibration machine learning model 306 to generate a reference probability 310 indicating the probability of a homozygous reference genotype at the multi-allelic genomic coordinate 302. In particular, the call recalibration system 106 generates a variant call classification indicating the probability that the multi-allelic genomic coordinate 302 indicates a homozygous genotype with respect to the reference genome based on its sequencing metrics. As shown, the reference probability 310 indicates the probability of a homozygous reference genotype (0 / 0) at coordinate 4 from the sample nucleotide sequence (represented as P(0 / 0)@4).
[0094] In one or more implementations, the call recalibration system 106 also utilizes the call recalibration machine learning model 306 to generate distinct genotype probabilities 312 indicating the probability of a genotype error at the multi-allelic genomic coordinate 302. For example, the call recalibration system 106 determines the probability that the predicted genotype for the multi-allelic genomic coordinate 302 is an incorrect genotype (e.g., a genotype incorrectly identified by the call generation model) or includes an incorrect allele in the predicted genotype. Specifically, in some cases, the call recalibration system 106 determines the probability that any het / hom errors are present in the multi-allelic genomic coordinate 302 (e.g., where the alternative bases are correct but the genotype is incorrect), or that the nucleotide base calls represent either all incorrect genotypes or incorrect alleles in the predicted genotype. For example, when determining the probability that a het / hom error exists, the call recalibration system 106 determines the probability that an alternative base call, represented as "1," is correct but the genotype is incorrect, such as the probability of incorrectly determining a 0 / 1 genotype call (e.g., A / T) instead of a correct 1 / 1 genotype call (e.g., T / T) (or vice versa if the correct genotype call is 0 / 1).
[0095] By determining the different genotype probabilities 312, the call recalibration system 106 can correct the inaccuracies of existing sequencing systems where the inaccurate calls are often indels. In particular, the call recalibration system 106 can more accurately generate nucleotide base calls for genomic coordinates that correspond to indels when the existing sequencing system determines that the nucleotide base calls represent inaccurate genotypes that represent inaccurate alleles resulting from long insertion or deletion sequences. As shown, the different genotype probabilities 312 indicate the probability of a different genotype belonging to coordinate 4 (represented as P(diff genotype)@4).
[0096] 3A , the call recalibration system 106 utilizes the call recalibration machine learning model 306 to generate a correct variant probability 314 indicating the probability of a correct variant call genotype at the multi-allelic genomic coordinate 302. For example, the call recalibration system 106 generates a probability that a predicted genotype for the multi-allelic genomic coordinate 302 is correct, as determined by the call generation model. As shown, the call recalibration system 106 determines a correct variant probability 314 indicating the probability that a predicted variant call from the call generation model is correct for coordinate 4 (represented as P(correct)@4).
[0097] Continuing with FIG. 3B , in some embodiments, the call recalibration system 106 utilizes the variant call classification 308 to update one or more data fields or variant call file fields ("VCF" fields) associated with a variant call file (e.g., variant call file 324). For example, the call recalibration system 106 generates an updated VCF field 316 that indicates updated sequencing metrics for the final nucleotide base call. In some cases, the call recalibration system 106 modifies or updates only certain VCF fields and does not update other fields based on the variant call classification 308. In other cases, the call recalibration system 106 does not update a VCF field based on the variant call classification 308. For example, when generating nucleotide base calls for a multi-allelic genomic coordinate 302, the call recalibration system 106 does not update certain fields, such as a genotype (GT) field, based on the variant call classification 308. Thus, in contrast to biallelic genomic coordinates, in some cases the call recalibration system 106 does not correct or update the GT field because in some cases there is not enough information to determine a new or updated genotype at the multiallelic genomic coordinate.
[0098] To illustrate one embodiment, FIG. 3B shows the call recalibration system 106 generating an updated VCF field 316 for a 1 / 2 genotype (GT), where cytosine represents a reference base (shown as "Ref:C") at the bi-allelic genomic coordinate for an allele corresponding to the reference genome, adenine represents a first alternative base ("Alt 1:A") at the bi-allelic genomic coordinate for a different allele, and thymine represents a second alternative base ("Alt 2:T") at the bi-allelic genomic coordinate for yet another different allele. However, FIG. 3B only shows examples of possible reference bases and possible alternative bases at bi-allelic genomic coordinates. The call recalibration system 106 can generate variant call classifications and correct corresponding metrics in the VCF field for various other reference bases and alternative bases at other bi-allelic genomic coordinates.
[0099] As further illustrated in FIG. 3B , the call recalibration system 106 generates an updated base call quality (QUAL) field 318. More specifically, the call recalibration system 106 modifies or updates a base call quality metric based on the variant call classification 308 to indicate the accuracy of the nucleotide base call at the multi-allelic genomic coordinate 302. As shown, the updated base call quality field 318 indicates a QUAL score of 48 for the variant at the corresponding genomic coordinate. In this example, the updated base call quality metric (e.g., a QUAL score of 48) represents a score for any type of variant at the corresponding multi-allelic genomic coordinate. Additionally, the call recalibration system 106 generates a modified or updated genotype quality (GQ) field 320. For example, based on the variant call classification 308, the call recalibration system 106 generates a modified or updated genotype quality metric that indicates the likelihood or probability that a predicted genotype at the multi-allelic genomic coordinate 302 is accurate. As shown, for example, the updated genotype quality field 320 indicates a genotype quality metric for genotype calls with heterozygous genotypes (e.g., a GQ score of 4 for a 1 / 2 genotype at a multi-allelic genomic coordinate).
[0100] In one or more embodiments, the call recalibration system 106 further generates or updates genotype likelihoods 322 and (in some cases) uses the genotype likelihoods 322 to rank the alleles. In particular, the call recalibration system 106 generates the updated genotype likelihoods as genotype likelihoods 322 by ordering the candidate nucleotide base calls at the multi-allelic genomic coordinate 302 according to their respective probabilities of attribution at the multi-allelic genomic coordinate 302. For example, the call recalibration system 106 determines the probability that each diploid genotype is associated with a plurality of genotypes composed of a pair of alleles. As another example, the call recalibration system 106 determines the relative probabilities associated with a plurality of alleles (e.g., from a reference genome, a first alternative allele, and a second alternative allele) belonging to the multi-allelic genomic coordinate 302 of the sample nucleotide sequence. In some embodiments, the call recalibration system 106 generates a metric for a PHRED scaled likelihood (PL) field as part of the updated VCF field 316. For example, the call recalibration system 106 generates metrics for the PL field that can indicate, for example, genotypes of a homozygous reference genotype, a heterozygous genotype, and a homozygous alternative genotype (e.g., each having a PL field nomenclature of 9 / 0 / 3).
[0101] In effect, the call recalibration system 106 generates allele-specific probabilities or likelihoods based on the relative probability of the nucleotide base calls corresponding to the alleles from the call generation model relative to any other (non-reference) genotypes identified by the call recalibration machine learning model 306. For example, in some embodiments, the call recalibration system 106 displays a relative probability score for each allele corresponding to each nucleotide base call in a PL field, which indicates the normalized PHRED-scaled likelihood for the genotype, and / or a genotype likelihood (GL) field, which indicates the log-scaled likelihood (e.g., log10 scale) of the data (e.g., sequencing metrics) given the called genotype.
[0102] As an example of generating updated genotype likelihoods and correcting a particular VCF field, in some cases, the call recalibration system 106 utilizes the call recalibration machine learning model 306 to generate a variant call classification 308 as a set of three variant call classifications whose probabilities sum to 1. In particular, the call recalibration machine learning model 306 can generate a reference probability 310 as 0.1, a different genotype probability 312 as 0.2, and a correct variant probability 314 as 0.7. Based on the reference probability 310, the different genotype probability 312, and the correct variant probability 314 in such an example, the call recalibration system 106 generates an updated genotype likelihood 322 by updating GT=0 / 0 using the reference probability 310, updating GT=1 / 2 using the correct variant probability 314, and updating other genotype positions in the PL field using a combination of information from the call recalibration machine learning model 306 and the call generation model. To use such combinations, in some embodiments, the call recalibration system 106 combines (e.g., sums) all of the probabilities of the alternative genotypes (as determined by the call generation model) and scales the combination to match the different genotype probabilities 312.
[0103] As illustrated in FIG. 3B, the call recalibration system 106 generates genotype likelihoods by determining normalized PL scores for different genotypes (GT). According to the normalized scale of PL scores, a relatively lower score for a genotype (e.g., PL0) represents a relatively higher likelihood that the genotype is present at the genomic coordinate, and a relatively higher score for a genotype (e.g., PL101) represents a relatively lower likelihood that the genotype is present at the genomic coordinate. For example, the call recalibration system 106 determines a PL score of 111 for the 0 / 0 genotype, a PL score of 52 for the 0 / 1 genotype, a PL score of 49 for the 0 / 2 genotype, a PL score of 42 for the 1 / 1 genotype, a PL score of 0 for the 1 / 2 genotype, and a PL score of 30 for the 2 / 2 genotype. Thus, in Figure 3B, a PL score of 0 indicates the genotype with the highest likelihood or selected genotype (e.g., 1 / 2 genotype) and a PL score of 111 represents the lowest likelihood (e.g., 0 / 0 genotype). Thus, in this example, the order of genotypes by likelihood (from most likely to least likely) is 1 / 2, 2 / 2, 1 / 1, 0 / 2, 0 / 1, and 0 / 0.
[0104] In some cases, the call recalibration system 106 generates the updated genotype likelihoods 322 as rankings of multiple alleles identified via the call generation model (without utilizing the call recalibration machine learning model 306). In other cases, the call recalibration system 106 utilizes a specialized version of the call recalibration machine learning model 306 that was trained to generate the updated genotype likelihoods 322 based on the variant call classifications 308.
[0105] As further illustrated in FIG. 3B, the call recalibration system 106 generates or updates a variant call file 324. The call recalibration system 106 can generate the variant call file 324 to include an updated VCF field 316, including base call quality metrics, genotype quality metrics, and / or updated genotype likelihoods. As mentioned, in some cases, the call recalibration system 106 updates only certain fields based on the multi-allelic analysis of FIGS. 3A-3B, while other fields, such as the genotype (GT) field, are left unchanged. For example, the call recalibration system 106 updates the genotype quality field and the base call quality field.
[0106] For other data fields, such as the normalized PHRED scaled likelihood (PL) for the genotype and posterior genotype probability (GP), the call recalibration system 106 either (i) keeps the field as is, (ii) removes the field, or (iii) updates only the field to reflect the GQ for the called genotype and class 0 output 0 / 0. In some cases, the call recalibration system 106 maintains the relative probability of other genotypes with respect to the called genotype to ensure consistent updates and the called genotype is the highest. By updating only the 0 / 0 and 1 / 2 values, the call recalibration system 106 maintains the distance of other genotypes from the called genotype.
[0107] Within the variant call file 324, the call recalibration system 106 can include or update one or more final nucleotide base calls (e.g., variant nucleotide base calls) associated with the multi-allelic genomic coordinate 302, determined based on the updated VCF field 316. Indeed, to generate a final nucleotide base call for the multi-allelic genomic coordinate 302, the call recalibration system 106 can predict two nucleotide bases (e.g., according to their respective probabilities) from three or more candidate alleles at the multi-allelic genomic coordinate.
[0108] As mentioned, in certain described embodiments, the call recalibration system 106 generates final nucleotide base calls (e.g., variant calls) for genomic coordinates within a haploid nucleotide sequence from a genomic sample. In particular, the call recalibration system 106 determines a haploid genotype for the haploid coordinate of the sample nucleotide sequence and further determines whether the haploid genotype is variant. FIGS. 4A-4B illustrate the generation of final nucleotide base calls for haploid genomic coordinates, according to one or more embodiments. For example, FIG. 4A illustrates the call recalibration system 106 utilizing a call recalibration machine learning model to generate final nucleotide base calls, according to one or more embodiments. FIG. 4B then illustrates a process for the call recalibration system 106 to train, tune, test, and / or apply the call recalibration machine learning model to generate nucleotide base calls for the haploid coordinates, according to one or more embodiments.
[0109] As illustrated in FIG. 4A, the call recalibration system 106 identifies a haploid nucleotide sequence 402. In particular, the call recalibration system 106 identifies a haploid nucleotide sequence 402 as a region of a sample nucleotide sequence that contains only haploids (as opposed to diploids). For example, the call recalibration system 106 determines, via a call generation model, that a region of the sample nucleotide sequence is located on a haploid sex chromosome (e.g., chr:Y). In some cases, the call recalibration system 106 determines or identifies the haploid nucleotide sequence 402 by determining nucleotide base calls for nucleotide reads corresponding to the haploid nucleotide sequence 402 and aligning the nucleotide reads to a reference genome. Although the haploid nucleotide sequence 402 is determined from the nucleotide base calls for the nucleotide reads, for ease of explanation, FIG. 4A shows the nucleobases for the underlying haploid nucleotide sequence 402 at a given genomic coordinate. This disclosure describes a process for determining nucleotide base calls for the following nucleotide reads and corresponding sequencing metrics with respect to Figures 6A-6B: As shown in Figure 4A, the call recalibration system 106 identifies a haploid nucleotide sequence 402 to include genomic coordinates 1-4 based on nucleotide reads, each having a single nucleotide base: 1.A 2.A 3.T 4.G.
[0110] As further illustrated in Figure 4A, the call recalibration system 106 determines sequencing metrics 404 for the haploid nucleotide sequence 402. In particular, the call recalibration system 106 determines read-based sequencing metrics, call model generation sequencing metrics, and / or external source sequencing metrics associated with particular genomic coordinates within the haploid nucleotide sequence 402. Further details regarding determining sequencing metrics are provided below with reference to Figures 6A-6C.
[0111] Based on the sequencing metrics 404, the call recalibration system 106 utilizes a call recalibration machine learning model 406 (e.g., call recalibration machine learning model 306) to generate a first genotype probability 408 and a second genotype probability 410 for a genomic coordinate within the haploid nucleotide sequence 402 based on the sequencing metrics 404. For example, the call recalibration system 106 generates the first genotype probability 408 indicating the probability that the genomic coordinate indicates a first genotype (e.g., a haploid reference genotype) and generates the second genotype probability 410 indicating the probability that the genomic coordinate indicates a second genotype (e.g., a haploid alternative genotype). As used herein, in some cases, the first genotype probability 408 and the second genotype probability 410 are examples of types of variant call classifications.
[0112] In some cases, the call recalibration system 106 generates the first genotype probability 408 and the second genotype probability 410 by transforming the input and / or output of the call recalibration machine learning model 406 to fit the model to a haploid scenario. For example, in some cases, the call recalibration system 106 transforms certain sequencing metrics or features as inputs to the call recalibration machine learning model 406 from haploid inputs to diploid inputs. More specifically, the call recalibration system 106 transforms haploid reference genotype calls generated by the call generation model to diploid homozygous reference genotype calls as inputs for the call recalibration machine learning model 406 (e.g., transforming haploid 0VC GT as input to diploid 0 / 0 GT). In addition, the call recalibration system 106 converts haploid alternative genotype calls generated by the call generation model to diploid homozygous alternative genotype calls (e.g., converting haploid 1VC GT to diploid 1 / 1 GT as input) as input for the call recalibration machine learning model 406. Further, in some cases, the call recalibration system 106 filters out, removes, or ignores heterozygous genotype calls generated by the call generation model as input for the call recalibration machine learning model 406.
[0113] In one or more embodiments, the call recalibration system 106 also (or alternatively) converts the output of the call recalibration machine learning model 406 from a diploid output to a haploid output. For example, in some cases, the call recalibration system 106 utilizes a softmax model or layer (e.g., as a layer within the call recalibration machine learning model 406) to convert from a diploid output to a haploid output. In some cases, the call recalibration system 106 utilizes a softmax layer to modify the confidence scores of diploid genotypes to simulate (or convert to) the probability of a haploid genotype for a genomic coordinate. For example, the call recalibration system 106 utilizes a softmax layer to modify the homozygous reference confidence scores of homozygous reference genotypes at a genomic coordinate to generate a haploid reference probability of the reference genotype at the genomic coordinate. Additionally, the call recalibration system 106 utilizes a softmax layer to modify the homozygous alternative confidence scores of homozygous alternative genotypes at a genomic coordinate to generate a haploid alternative probability of the alternative genotype at the genomic coordinate.
[0114] In one or more embodiments, the call recalibration system 106 reduces or removes one of the three model outputs. For example, when determining a nucleotide base call for a haploid genomic coordinate, the call recalibration system 106 removes confidence scores where the genotype of the genomic coordinate is heterozygous (or there is a het / hom error at the coordinate) and does not input such confidence scores into the softmax layer. Based on the first confidence score that the genomic coordinate indicates a haploid reference genotype and the second confidence score (or the third confidence score) that the genomic coordinate indicates a haploid alternative genotype, the call recalibration system 106 normalizes these two remaining confidence scores (so that their sum is 1) using a softmax layer to generate a first genotype probability 408 and a second genotype probability 410 for haploids based on the corresponding diploid probabilities.
[0115] As shown in Figure 4A, the first genotype probability 408 indicates that there is an 80% probability that the genotype at haploid coordinate 3 is 0 (or constitutes the haploid reference) (represented as 0@3→80%). Similarly, the second genotype probability 410 indicates that there is a 20% probability that the genotype at haploid coordinate 3 is 1 (or constitutes the haploid surrogate) (represented as 1@3→20%). Further details regarding the conversion of outputs are provided below with reference to Figure 4B.
[0116] 4A, the call recalibration system 106 generates or updates a variant call file 412 based on the first genotype probability 408 (e.g., haploid reference genotype probability) and the second genotype probability 410 (e.g., haploid alternative genotype probability). For example, the call recalibration system 106 updates the variant call file 412 to reflect or indicate a final nucleotide base call 414 associated with the haploid genomic coordinate based on the first genotype probability 408 and the second genotype probability 410.
[0117] In certain embodiments, the call recalibration system 106 determines a final nucleotide base call 414 indicating a haploid genotype for the genomic coordinate based on comparing the first genotype probability 408 and the second genotype probability 410 and selecting the highest genotype from among the first genotype probability 408 and the second genotype probability 410. In some cases, the call recalibration system 106 updates additional fields associated with the variant call file 412, such as a base call quality field, a genotype quality field, and / or a genotype field, based on comparing the first genotype probability 408 and the second genotype probability 410.
[0118] For example, based on determining that the second genotype probability 410 is highest (i.e., exceeds the first genotype probability 408) or that the nucleotide base call (or variant call) is most likely to be a true positive, the call recalibration system 106 determines a haploid alternative genotype for the genomic coordinate. For example, if the second genotype probability 410 (e.g., haploid alternative genotype probability) exceeds the first genotype probability 408 (e.g., haploid reference type probability), the call recalibration system 106 further determines a revised base call quality metric, a revised genotype metric, and / or a revised genotype quality metric (for inclusion in the variant call file 412). In some cases, the above will modify the genotype quality metric to reflect the likelihood that the nucleotide base call or variant call is inaccurate (in PHRED format) with the existing genotype.
[0119] Based on a determination that the second genotype probability 410 is not the highest (i.e., the first genotype probability 408 exceeds the second genotype probability 410), the call recalibration system 106 determines a haploid reference genotype for the genomic coordinate. For example, if the first genotype probability 408 (e.g., haploid reference type probability) exceeds the second genotype probability 410 (e.g., haploid alternate genotype probability), the call recalibration system 106 further determines a revised genotype quality metric and / or a revised base call quality metric. For example, if the call recalibration system 106 predicts a reference genotype call, the call recalibration system 106 retains the called genotype and sets the score to the value output by the call recalibration machine learning model. However, if the call recalibration system 106 uses a call recalibration machine learning model to determine a corrected base call quality metric for a genotype call at a haploid genomic coordinate, the call recalibration system 106 modifies the quality field for the genotype call to include the corrected base call quality metric. Alternatively, in some cases, if the base call quality metric falls below a quality threshold, the call recalibration system 106 can remove the nucleotide base call, or at least not include the nucleotide base call for the genomic coordinate in the variant call file.
[0120] In some embodiments, the call recalibration system 106 generates a final nucleotide base call 414 based on comparing the first genotype probability 408 and the second genotype probability 410. As shown, the call recalibration system 106 determines that the first genotype probability 408 is higher than the second genotype probability 410, and therefore generates a final nucleotide base call 414 indicating that the genotype at the particular haploid coordinate (coordinate 3) is most likely the haploid reference genotype (represented as 3→0).
[0121] As illustrated in Figure 4B, the call recalibration system 106 modifies the inputs and outputs of the call recalibration machine learning model 406 to facilitate the generation of a final nucleotide base call (e.g., a variant call) for the genomic coordinates of the haploid nucleotide sequence. In some embodiments, the process illustrated in Figure 4B represents the training and / or tuning of the call recalibration machine learning model 406 to learn parameters for generating a nucleotide base call (e.g., a variant call). In other embodiments, some or all of the process illustrated in Figure 4B represents the application of or inference using the call recalibration machine learning model 406.
[0122] As shown, the call recalibration system 106 performs downsampling 418 of (a subset of) diploid nucleotide reads 420 to simulate haploid nucleotide reads. More specifically, the call recalibration system 106 downsamples (or otherwise modifies) the diploid data to mimic or simulate haploid data in order to train or tune the call recalibration machine learning model 406. Indeed, because ground truth haploid data is sparse, the call recalibration system 106 cannot rely solely on diploid data to learn robust parameters for generating nucleotide base calls for haploid coordinates. Thus, unlike some existing sequencing systems that cannot generate calls for haploid coordinates (due to lack of training data), in some embodiments, the call recalibration system 106 fits the haploid scenario by simulating haploid data from diploid data.
[0123] For example, the call recalibration system 106 determines (or receives) diploid nucleotide reads 420 via the call generation model 416. In addition, the call recalibration system 106 (randomly) selects a subset of the diploid nucleotide reads 420 (e.g., random selection of 50% of the reads) to use as training or test data. As shown, the diploid nucleotide reads 420 include reads for four genomic coordinates 1-4, as follows: 1) AA, 2) AA, 3) CC, 4) TT. In addition, the call recalibration system 106 determines a diploid sequencing metric 422 from (a subset of) the diploid nucleotide reads 420. In some embodiments, the call recalibration system 106 determines or identifies one or more genomic coordinates of diploid nucleotide reads 420 indicative of a homozygous genotype, such as a homozygous reference genotype or a homozygous alternative genotype, based on truth data (e.g., PrecisionFDA truth data, Platinum Genomes, or some other high-confidence truth set, such as a truth set from Genome in a Bottle (GIAB), Global Alliance for Genomic Health (GA4GH), or the Telomere to Telomere Consortium) and / or diploid sequencing metrics 422.
[0124] As further illustrated in FIG. 4B , the call recalibration system 106 generates (or simulates) a haploid sequencing metric 424 from the diploid sequencing metric 422 via downsampling 418. For example, the call recalibration system 106 corrects the homozygous genotype of the diploid nucleotide read 420 to simulate a haploid genotype. Specifically, the call recalibration system 106 converts the homozygous reference genotype to a haploid reference genotype (represented as 0 / 0→0) and converts the homozygous alternate genotype to a haploid alternate genotype (represented as 1 / 1→1). The call recalibration system 106 further selects as the haploid sequencing metric 424 the sequencing metric of the diploid nucleotide read 420 used to simulate the haploid nucleotide read. Based on these haploid sequencing metrics 424, the call recalibration system 106 can train and / or test a call recalibration machine learning model 406 to accurately generate final nucleotide base calls (e.g., variant calls) for the haploid coordinates.
[0125] Indeed, in training, testing, and / or inference, the call recalibration system 106 utilizes a call recalibration machine learning model 406 to generate final nucleotide base calls based on sequencing metrics, such as haploid sequencing metrics 424. As mentioned above, as part of generating final nucleotide base calls 432 via the call recalibration machine learning model 406 (for either training, testing, or inference), the call recalibration system 106 modifies the output of the call recalibration machine learning model 406. For example, the call recalibration system 106 modifies the confidence scores generated by one or more classifier layers 426 of the call recalibration machine learning model 406.
[0126] In some embodiments, the call recalibration system 106 does not simulate haploid data from diploid data during the inference process (as opposed to the training or testing process). Indeed, when applying the call recalibration machine learning model 406 to generate predictions, the call recalibration system 106 need only modify the inputs and outputs of the call recalibration machine learning model 406 once the model has been trained for the haploid scenario using simulated haploid data. When using the call recalibration machine learning model 406, for example, the call recalibration system 106 inputs a sequencing metric that indicates the data is haploid during the inference process.
[0127] Specifically, as shown in FIG. 4B, the call recalibration system 106 generates three separate confidence scores, each representing a confidence level of a different genotype belonging to a given genomic coordinate: (i) a first confidence score (z0) for a haploid reference genotype, (ii) a second confidence score (z1) for a heterozygous genotype, and (iii) a third confidence score (z2) for a haploid alternative genotype. Because heterozygous genotypes cannot be simulated from diploid data to haploid data and (during implementation) haploid coordinates do not indicate heterozygous genotypes, the call recalibration system 106 further reduces, ignores, or removes the second confidence score (z1). In some embodiments, the call recalibration system 106 utilizes a softmax model 428 as another layer in the call recalibration machine learning model 406 to generate the final probability from the confidence scores.
[0128] In particular, as shown in FIG. 4B, the call recalibration system 106 utilizes a softmax model 428 to generate two genotype probabilities from the corrected confidence scores (e.g., a first genotype probability 408 of a haploid reference genotype and a second genotype probability 410 of a haploid alternative genotype). In particular, after ignoring or discarding the second confidence score (z1), the call recalibration system 106 utilizes a softmax model 428 to normalize across the first confidence score to the third confidence score (so that they sum to 1). The call recalibration system 106 further generates the first genotype probability 408 represented by σ0=p(0) and the second genotype probability 410 represented by σ1=p(1). As described, each probability score represents the probability of a respective haploid genotype at a genomic coordinate.
[0129] Based on the probability scores, the call recalibration system 106 further generates a variant call file 430 (e.g., variant call file 412) that includes a final nucleotide base call 432 (e.g., final nucleotide base call 414). For example, the call recalibration system 106 determines the final nucleotide base call 432 from the two genotype probabilities. As shown, for example, the final nucleotide base call 432 is haploid A for the given genomic coordinate. However, the final nucleotide base call 432 can be a different predicted nucleotide base in other embodiments. Further details regarding generating variant call files are provided throughout this disclosure.
[0130] As noted above, in certain described embodiments, the call recalibration system 106 generates final nucleotide base calls (e.g., variant calls) for homozygous reference genomic coordinates (as originally predicted by the call generation model). In particular, the call recalibration system 106 generates final nucleotide base calls for genomic coordinates of the sample nucleotide sequence that were (or would be) determined by the call generation model to indicate a homozygous reference genotype. Figure 5 illustrates the generation of variant calls for genomic coordinates that were or could have been incorrectly identified as a homozygous reference genotype, according to one or more embodiments.
[0131] As illustrated in FIG. 5 , the call recalibration system 106 utilizes a call generation model 502 to generate initial nucleotide base calls for a sample nucleotide sequence 504. In particular, the call recalibration system 106 generates nucleotide base calls indicative of alleles or genotypes associated with particular genomic coordinates. As shown, the call generation model 502 determines genotypes for the sample nucleotide sequence 504 at coordinates 1-4 as follows: 1) 0 / 1, 2) 1 / 1, 3) 0 / 0, 4) 0 / 1. Additionally, the call recalibration system 106 identifies or determines genomic coordinates within the sample nucleotide sequence 504 that are indicative of a homozygous reference genotype as determined by the call generation model 502. In the illustrated example, the call recalibration system 106 identifies coordinate 3 as a homozygous reference coordinate. In contrast, in some embodiments, the call recalibration system 106 does not generate initial nucleotide base calls indicative of a homozygous reference genotype, but rather determines sequencing metrics 506 for nucleotide reads that cover genomic coordinates that match the homozygous reference genotype.
[0132] In many cases, existing sequencing systems have ignored homozygous reference coordinates such as coordinate 3, treating them as true negative variant calls not necessary for further processing. However, such processing depends on the accuracy of the call generation model 502 to initially make the proper nucleotide base calls, which is not always the case. In fact, the call generation model 502 may, in some scenarios, generate a large number of false negative variant calls. Thus, the call recalibration system 106 recovers some of these false negative variant calls by not ignoring genomic coordinates that were initially identified (or would have been identified) as homozygous reference genotypes, but by forcing further analysis at these loci (e.g., to subsequently update or modify their determined genotypes).
[0133] Specifically, as illustrated in Figure 5, the call recalibration system 106 extracts or determines sequencing metrics 506 for homozygous reference coordinates 3. For example, the call recalibration system 106 determines read-based sequencing metrics 508, external source sequencing metrics 510, and call model generation sequencing metrics 512. Additional details regarding determining or extracting sequencing metrics are provided below with reference to Figures 6A-6C.
[0134] 5, the call recalibration system 106 utilizes a call recalibration machine learning model 514 (e.g., call recalibration machine learning model 306 or 406) to generate one or more variant call classifications 516 from the sequencing metrics 506. In particular, the call recalibration system 106 generates the variant call classification 516 that is indicative of (or defines a level of) accuracy for identifying a variant at a genomic coordinate (e.g., coordinate 3).
[0135] The following paragraphs describe examples of variant call classifications 516. As an exemplary variant call classification, the call recalibration system 106 utilizes the call recalibration machine learning model 514 to generate a false positive classification. For example, the call recalibration system 106 generates a false positive classification indicating the probability that a nucleotide base call (e.g., a genotype call) is a false positive variant, or the probability that a nucleotide base call indicates a variant where no variant is actually present in the sample nucleotide sequence 504. The call recalibration system 106 generates a false positive classification from one or more of the sequencing metrics 506 considered together by the call recalibration machine learning model 514.
[0136] In certain implementations, the call recalibration system 106 also (or alternatively) generates a genotype error classification (or heterozygous genotype classification) as part of the variant call classification 516. More specifically, the call recalibration system 106 utilizes the call recalibration machine learning model 514 to determine the probability that a genotype associated with a nucleotide base call is incorrect or that a heterozygous genotype exists (e.g., for coordinate 3). For example, the call recalibration system 106 determines the probability that a het / hom error exists at coordinate 3, where the nucleotide base call may indicate a heterozygous genotype (e.g., 0 / 1) in the sample nucleotide sequence 504 and the genotype is actually a homozygous alternative (e.g., 1 / 1) with respect to the reference genome. Conversely, the call recalibration system 106 determines the probability of determining that the genotype at coordinate 3 is a homozygous alternative (e.g., 1 / 1) when in fact the nucleotide base is heterozygous with respect to the reference genome (e.g., 0 / 1).
[0137] In one or more embodiments, the call recalibration system 106 also (or alternatively) generates a true positive classification (or a homozygous alternative classification) for coordinate 3 as part of the variant call classification 516. In particular, the call recalibration system 106 utilizes a call recalibration machine learning model 514 to determine the probability that a nucleotide base call for coordinate 3 is a true positive variant call, or that the nucleotide base call indicates a true variant if a variant is actually present with respect to the reference genome, or the probability that a homozygous alternative genotype is present at the genomic coordinate.
[0138] As further illustrated in FIG. 5 , the call recalibration system 106 generates or updates the variant call file 518 to indicate a variant call 520. More specifically, the call recalibration system 106 generates the variant call 520 based on the variant call classification 516 to indicate whether there is a variant at coordinate 3. In some cases, the call recalibration system 106 updates one or more of the call quality field, the genotype field, or the genotype quality field corresponding to the variant call file 518 based on the one or more variant call classifications 516. The call quality field, the genotype field, and / or the genotype quality field can indicate the updated variant call as variant call 520. As shown, the variant call 520 indicates a variant at coordinate 3 and changes the initial nucleotide base call for coordinate 3 from one indicative of a homozygous reference genotype (0 / 0) to one indicative of a heterozygous genotype (0 / 1). In other examples, the call recalibration system 106 does not change the initial nucleotide base call for coordinate 3, or changes the initial nucleotide base call to a different genotype, such as a homozygous alternate genotype (1 / 1).
[0139] In one or more embodiments, the call recalibration system 106 determines a genotype for the indicated genomic coordinate (e.g., coordinate 3) based on a comparison of the probabilities of the variant call classifications. For example, the call recalibration system 106 determines a homozygous alternative genotype based on determining that a true positive classification (or a homozygous alternative classification) has the highest probability among one or more variant call classifications. Specifically, the call recalibration system 106 updates the genotype quality field while also updating the genotype field (e.g., to 1 / 1) and the PL field.
[0140] Alternatively, the call recalibration system 106 determines the heterozygous genotype based on determining that the genotype error classification (e.g., the heterozygous genotype classification) has the highest probability among one or more variant call classifications. Specifically, the call recalibration system 106 updates the genotype quality field while also updating the genotype field (e.g., to 0 / 1) and the PL field.
[0141] Further alternatively, the call recalibration system 106 determines a homozygous reference genotype based on determining that neither a true positive classification (e.g., a homozygous alternative classification) nor a genotype error classification (e.g., a heterozygous genotype) has the highest probability among one or more variant call classifications. In some cases, the call recalibration system 106 removes or discards a record that compares the probability of a variant classification if both the call generation model 502 and the call recalibration machine learning model 514 determine that the genomic coordinate has a homozygous reference genotype.
[0142] In one or more embodiments, updating the variant calls for homozygous reference coordinates provides or improves forced genotype functionality (e.g., for queries of genotypes and genotype probabilities at a particular genomic coordinate). In particular, the call recalibration system 106 can initially genotype genomic coordinates that do not meet a variant quality threshold (e.g., as indicated by the call generation model 502). Indeed, the call recalibration system 106 can output genotypes to the variant call file 518 even if the variant quality of the genomic coordinate falls below a threshold typically required to identify structural or other hard-to-determine variants.
[0143] As mentioned above, in certain described embodiments, the call recalibration system 106 determines or extracts sequencing metrics for nucleotide base calls at particular genomic coordinates. In particular, the call recalibration system 106 determines sequencing metrics, such as read-based sequencing metrics, external source sequencing metrics, and call model generation sequencing metrics, from calls corresponding to nucleotide reads from a sample nucleotide sequence. Figures 6A-6C illustrate determining sequencing metrics, according to one or more embodiments. Specifically, Figure 6A illustrates determining read-based sequencing metrics, Figure 6B illustrates determining call model generation sequencing metrics, and Figure 6C illustrates determining external source sequencing metrics.
[0144] As illustrated in FIG. 6A, the call recalibration system 106 accesses, retrieves, obtains, determines, or generates nucleotide reads 602. In particular, the call recalibration system 106 utilizes a sequencer 114 to determine nucleotide reads 602, including nucleotide base calls for regions from a sample nucleotide sequence (e.g., a sample genome). For example, the call recalibration system 106 utilizes sequencing-by-synthesis (SBS) and / or Sanger sequencing techniques to generate a plurality of nucleotide reads 602 to determine nucleotide base calls for oligonucleotide clusters from wells in a flow cell and / or via fluorescent tagging. More specifically, the call recalibration system 106 utilizes cluster generation and SBS chemistry to sequence millions or billions of clusters in a flow cell. During SBS chemistry, for each cluster, the call recalibration system 106 stores the nucleotide base calls from the nucleotide reads 602 for each cycle of sequencing via real-time analysis (RTA) software.
[0145] As further illustrated in FIG. 6A, in some embodiments, the call recalibration system 106 performs read processing and mapping 604. For example, the call recalibration system 106 utilizes RTA software to store base call data in the form of individual base call data files (or BCLs). In some cases, the call recalibration system 106 further converts the BCL files into sequence data 608 (e.g., via a BCL to FASTQ conversion), as illustrated in FIG. 6B. As illustrated in FIG. 6A, the call recalibration system 106 generates a multiple read coverage (e.g., a read pileup) that includes multiple nucleotide reads 602 or nucleotide base calls that correspond to a single genomic coordinate.
[0146] In particular, in certain embodiments, the call recalibration system 106 aligns nucleotide reads to a reference genome or receives information regarding read alignment. Specifically, the call recalibration system 106 determines (or receives information indicating) which nucleotide bases of a given read align with which genomic coordinates of a reference sequence. Different reads have different lengths and contain different nucleotide bases. Thus, in some cases, the call recalibration system 106 analyzes each nucleotide of each read to determine (or receives information indicating) where the read "fits" with respect to the reference sequence, e.g., where a base in the read aligns with a base in the reference. In some cases, the call recalibration system 106 aligns many reads at a single genomic coordinate, thus resulting in read pile-up.
[0147] In certain embodiments, the call recalibration system 106 performs additional statistical tests to determine or detect differences between the metrics associated with the reference nucleotide sequence and the metrics associated with the alternative supporting nucleotide reads. Through these statistical tests, the call recalibration system 106 re-operates the raw sequencing metrics to determine the read-based sequencing metrics 606. In some cases, the call recalibration system 106 determines or extracts raw sequencing metrics including one or more of: (i) an alignment metric to quantify the alignment of the sample nucleotide sequence with the genomic coordinates of the exemplary nucleotide sequence (e.g., a nucleotide sequence from a reference genome or an ancestral haplotype); (ii) a depth metric to quantify the depth of a nucleotide base call for the sample nucleotide sequence at the genomic coordinates of the exemplary nucleotide sequence; or (iii) a call quality metric to quantify the quality of a nucleotide base call for the sample nucleotide sequence at the genomic coordinates of the exemplary nucleotide sequence. For example, the call recalibration system 106 determines a mapping quality metric (e.g., the MAPQ metric shown in FIG. 6A), a soft clipping metric, or other alignment metric that measures the alignment of the sample sequence with the reference genome. As another example, the call recalibration system 106 determines a forward-reverse depth metric (or other such depth metric) or a callability metric (or other such call quality metric) for the variant nucleotide base calls.
[0148] As just mentioned, in some embodiments, the call recalibration system 106 re-operates raw sequencing metrics to generate read-based sequencing metrics 606 that are more useful for comparing metrics associated with a reference nucleotide sequence with metrics associated with various supporting alternative nucleotide reads. For example, the call recalibration system 106 determines various metrics for a sample sequence with respect to a reference sequence, and further determines various metrics for a sample sequence with respect to an alternative supporting sequence. In addition, the call recalibration system 106 performs a comparative analysis between the metrics associated with the reference sequence and the metrics associated with the alternative supporting reads.
[0149] For example, the call recalibration system 106 compares how nucleotide bases of a sample nucleotide sequence (e.g., a sample genome) map to a reference sequence with how the nucleotide bases map to various alternative supporting reads. In some cases, the call recalibration system 106 determines a mapping quality associated with the reference sequence to compare with the mapping quality associated with the alternative supporting reads. For example, the call recalibration system 106 determines a mapping quality statistic that reflects differences in the distribution of reads supporting the reference sequence versus reads supporting alternative alleles.
[0150] In these or other cases, the call recalibration system 106 determines mismatch counts between the sample sequence and the reference sequence and between the reference sequence and the alternate support read. The call recalibration system 106 further compares the mismatch counts to determine a comparative mismatch count metric. Furthermore, the call recalibration system 106 determines a soft clipping metric for the sample sequence relative to the reference sequence and further determines a soft clipping metric for the alternate support read. The call recalibration system 106 also compares the soft clipping metric between the reference sequence and the alternate support read to generate a comparative soft clipping metric. Furthermore, the call recalibration system 106 compares base call quality metrics for the reference sequence and the alternate support read, and / or compares the query position of the sample sequence relative to the reference sequence with the query position for the alternate support read.
[0151] As further illustrated in FIG. 6A, the call recalibration system 106 utilizes comparisons and / or other statistical tests to generate read-based sequencing metrics 606 including: i) a comparative mapping quality distribution metric indicative of a mapping quality distribution comparing the mapping quality for the reference sequence to the mapping quality for the alternate supporting read; ii) a comparative secondary mapping alignment metric indicative of a comparison of the secondary mapping of bases in the reference sequence to the secondary mapping of bases in the alternate supporting read; iii) a comparative mismatch count metric indicative of a comparison of mismatched nucleotide bases for the reference sequence to the mismatched bases for the alternate supporting read; iv) a comparative soft clipping metric indicative of a comparison of the soft clipping metric for the reference sequence to the soft clipping metric for the alternate supporting read; v) a read depth of the nucleotide read 602 and one or more one or more read depth comparison metrics indicative of a comparison between average read depths (e.g., a local average read depth at a particular genomic coordinate and a global average read depth across multiple genomic coordinates in a region); vi) one or more comparative base quality metrics indicative of a comparison between base quality for the reference sequence and base quality for the alternative supporting reads (e.g., overall base quality, initial base quality, and later base quality in the nucleotide reads 602); vii) one or more comparative query position metrics indicative of a comparison between a query position for the reference sequence and a query position for the alternative supporting reads; viii) one or more context information metrics indicative of homopolymer and periodicity of the nucleotide base calls; ix) a strand bias metric indicative of a strand bias associated with one or more of the nucleotide reads 602; and x) a read direction bias metric indicative of a read direction bias associated with the nucleotide reads 602. In some cases, the call recalibration system 106 generates or re-engineers additional or alternative read-based sequencing metrics as part of the read-based sequencing metrics 606.
[0152] In addition to the read-based sequencing metrics 606, as illustrated in FIG. 6B, the call recalibration system 106 generates call model-generated sequencing metrics 612. In particular, the call recalibration system 106 utilizes the call generation model 610 to generate the call model-generated sequencing metrics from the sequence data 608. For example, the call recalibration system 106 extracts or determines the sequence data 608 based on the read processing and mapping 604 described in connection with FIG. 6A. In some cases, the call recalibration system 106 generates the sequence data 608 as part of one or more digital files, such as BCL and FASTQ files.
[0153] To generate such a file, in some embodiments, the sequencing device 114 (or call recalibration system 106) sequences millions or billions of clusters in a flow cell using cluster generation and SBS chemistry. During the SBS chemistry, for each cluster, the sequencing device 114 (or call recalibration system 106) stores nucleotide base calls from the nucleotide reads 602 for each cycle of sequencing via real-time analysis (RTA) software. The sequencing device 114 (or call recalibration system 106) further stores the base call data in the form of individual base call data files (or BCLs) using the RTA software. In some cases, the sequencing device 114 (or call recalibration system 106) further converts the BCL files into sequence data 608 (e.g., via a BCL to FASTQ conversion). For example, the sequencing device 114 (or call recalibration system 106) generates a FASTQ file from the nucleotide reads 602, where the FASTQ file includes the sequence data 608.
[0154] In some cases, the call recalibration system 106 generates sequence data 608 for each cluster that passes the initial quality filter from the sample sequence. For example, the call recalibration system 106 generates an entry for each cluster, with each entry including four rows (or four items of sequence data): i) a sequence identifier with information about the sequencing run and the cluster, ii) the nucleotide base calls that make up the sequence (e.g., a sequence of A, C, T, G, and / or N calls), iii) a separator (e.g., a "+" symbol), and iv) a base call quality metric indicating the probability of accuracy for the nucleotide base call (Phred+33 coded).
[0155] As further illustrated in FIG. 6B , the call recalibration system 106 implements, utilizes, or applies a call generation model 610 to process or analyze the sequence data 608. Indeed, in some embodiments, the call recalibration system 106 utilizes the call generation model 610 to re-manipulate raw sequencing metrics (e.g., raw sequencing metrics in the sequence data 608) to generate call model-generated sequencing metrics 612. In particular, the call generation model 610 includes a mapping and alignment component for mapping and aligning nucleotide base calls from the sequence data 608. In addition, the call generation model 610 includes a variant calling component for generating nucleotide base calls (e.g., reference base calls, such as variant calls or non-variant calls) from the sequence data 608. In some cases, the call recalibration system 106 extracts the call model-generated sequencing metrics 612 that have been generated utilizing the mapping and variant calling components of the call generation model 610.
[0156] To illustrate examples of call model-generated sequencing metrics 612, in some cases, the call recalibration system 106 may generate a call model-generated sequencing metric 612 that includes: i) a base call quality metric (e.g., a DRAGEN QUAL score) indicating a quality score for the nucleotide base calls generated via the call generation model 610; ii) a call model-generated foreign read detection metric (e.g., a foreign read detection (FRD) score) indicating the probability that one or more of the nucleotide reads 602 in the pileup may be a foreign read (e.g., their true location is elsewhere in the reference sequence); iii) a call model-generated base quality dropoff metric (e.g., a base quality dropoff (BQD) score) indicating the probability of a base quality dropoff based on one or more of strand bias, error location in a thread, or low average base quality across a subset of the nucleotide reads 602; iv) average read depth; v) indel statistics (e.g., a polymerase chain reaction (PCR) score; The call recalibration system 106 generates variant calling metrics 612, including one or more of: reaction, "PCR" (reaction), "PCR" (reaction curves), and / or vi) hidden Markov model (HMM) statistics, vii) secondary alignment metrics indicating the probability that a secondary nucleotide base call is correct, viii) base context metrics indicating context information for nucleotides surrounding a nucleotide base call, iv) neighborhood call metrics indicating the proximity of a nucleotide base call (e.g., adjacent or within a threshold degree of separation from it), x) joint detection metrics indicating the probability of detecting a joint corresponding to two or more overlapping nucleotide base calls, xii) read filtering metrics indicating a threshold quality metric or other metric for filtering out nucleotide base calls having low mapping quality, base quality, or other quality metrics, etc. The call recalibration system 106 generates sequencing metrics 612 for call model generation from internal (e.g., unique and model-specific) variables reflecting interacting processing pathways, corner cases, and difficult predictions / decisions.
[0157] As indicated above, in some cases, the call recalibration system 106 determines the FRD score according to the method described in U.S. Patent Application No. 16 / 280,022 to Eric Jon Ojard, entitled System and Method for Correlated Error Event Mitigation for Variant Calling, which is incorporated herein by reference in its entirety. In certain implementations, the call recalibration system 106 also (or alternatively) determines the BQD score, FRD score, HMM statistics, and / or other variant calling metrics according to the methods described in U.S. Patent Application Nos. 17 / 165,828, 15 / 643,381, and 14 / 811,836, which are incorporated herein by reference in their entireties.
[0158] 6B, the sequencing metrics 612 for call model generation include, but are not limited to, variant calling metrics extracted via the variant calling component of the call generation model 610. In addition to or as an alternative to the examples of sequencing metrics 612 for call model generation described above, in some cases, the call recalibration system 106 may also use a call recalibration function to generate a call model file that includes a call model ... x) the number of de novo indels (e.g., SNPs with a de novo quality metric that meets a threshold level); x) the number of de novo MNPs (e.g., SNPs with a de novo quality metric that meets a threshold level); xi) the number of SNPs in the first chromosome divided by the number of SNPs in the second chromosome; xii) the number of SNP transitions; xiii) the number of SNP transversions; xiv) the number of heterozygous variants; xv) the number of homozygous variants; xvi) the ratio between the number of heterozygous variants and the number of homozygous variants; xvii) the number of variants detected in the dbSNP reference file; and / or xviii) the total number of variants minus the number detected in the dbSNP file.
[0159] Additionally, the sequencing metrics for call model generation 612 can include mapping and alignment sequencing metrics extracted via the mapping and alignment component of the call generation model 610. For example, the base pair caller recalibration system 106 may generate a sequence metric for each of the following: i) the number of total input reads, ii) the number of duplicate mark reads, iii) the number of mate reads with duplicate marks removed, iv) the number of unique reads, v) the number of reads with mate sequences, vi) the number of reads without mate sequences, vii) an index of reads that fail a quality check, viii) an index of mapped reads, ix) the number of unique and mapped reads, x) the number of unmapped reads, xi) the number of singleton reads, xi) the number of unique reads, ii) the number of unique reads, iii) the number of unmapped reads, xi ... xii) number of paired reads; xiii) number of properly paired reads (e.g., when both reads of a pair are mapped and fall within an acceptable range of each other based on the estimated insert length distribution); xiv) number of mismatched reads (e.g., number of reads that are not properly paired); xv) number of paired reads that are mapped to different chromosomes; xvi) number of paired reads that are mapped to different chromosomes and have a mapping quality metric of 10 or higher; xvii) indels R1 and R2 Generate or extract (e.g., via metric re-operation) mapping and alignment metrics including one or more of: the percentage of reads in R1 and R2; xviii) the percentage of soft-clipped bases in R1 and R2; xix) the number of mismatched bases in indels R1 and R2; xx) the number of bases (e.g., total and / or R1 or R2) with a base quality of at least 30; xxi) the number of alignments (e.g., total alignments, secondary alignments, and / or supplemental alignments); xxii) estimated read length; and xxiii) estimated sample contamination.
[0160] 6C, the call recalibration system 106 generates, extracts, or determines externally sourced sequencing metrics 616. In particular, the call recalibration system 106 determines the externally sourced sequencing metrics 616 from one or more databases external to the call recalibration system 106, such as a sequencing information database 614 (e.g., database 116). For example, the call recalibration system 106 accesses sequencing metrics that are general or generally applicable to nucleotide sequencing. In addition, the call recalibration system 106 accesses or determines sequencing information for a particular reference sequence (e.g., stored in the sequencing information database 614). In some cases, the call recalibration system 106 determines external source sequencing metrics 616 including: i) a mappability metric indicating the ease or difficulty of mapping a particular nucleotide sequence (or a particular nucleotide read or nucleotide base call); ii) a guanine-cytosine content metric indicating the count (or dropout or average) of guanine-cytosine content in a reference nucleotide sequence (e.g., a reference genome); iii) a replication timing metric indicating the time required to replicate a particular number of nucleotides from a reference sequence; iv) one or more DNA structure metrics indicating the DNA structure of the reference sequence (e.g., a reference genome); v) a conservation metric indicating a measure of sequence conservation across multiple species (e.g., a measure of change relative to the average); and / or others.
[0161] As mentioned, in certain described embodiments, the call recalibration system 106 utilizes a call recalibration machine learning model in conjunction with a call generation model to generate nucleotide base calls. In particular, the call recalibration system 106 utilizes the call recalibration machine learning model to modify data fields corresponding to the variant call file that represents the nucleotide base calls. Figure 7 illustrates generating nucleotide base calls by modifying the variant call file using a call recalibration machine learning model and a call generation model according to one or more embodiments.
[0162] As illustrated in FIG. 7, the call recalibration system 106 accesses a sequencing information database 702 (e.g., sequencing information database 614), a reference sequence 704, and sequence data 706 (e.g., sequence data 608) extrapolated from one or more nucleotide reads. In practice, the call recalibration system 106 performs sequencing metric extraction 712 to extract or re-engineer sequencing metrics as described above in connection with FIGS. 6A-6C. For example, the call recalibration system 106 generates read-based sequencing metrics, external source sequencing metrics, and call model generated sequencing metrics. In some cases, the call recalibration system 106 utilizes a mapping and alignment component 708 of a call generation model 722 (e.g., call generation model 610) to determine mapping and alignment sequencing metrics as described above. In addition, the call recalibration system 106 utilizes a variant caller component 710 of the call generation model 722 to generate variant calling metrics as described above. Additionally, the call recalibration system 106 determines read-based sequencing metrics and external source sequencing metrics (eg, from the sequencing information database 702 and / or the reference sequences 704).
[0163] As further illustrated in FIG. 7, the call recalibration system 106 generates variant call classifications 716. More specifically, the call recalibration system 106 utilizes a call recalibration machine learning model 714 to generate the variant call classifications 716 from the sequencing metrics. For example, the call recalibration machine learning model 714 generates the variant call classifications 716 including a false positive classification, a genotype error classification, and a true positive classification. Specifically, the false positive classification indicates the probability that the nucleotide base call (e.g., variant call) is a false positive. Conversely, the true positive classification indicates the probability that the nucleotide base call (e.g., variant call) is a true positive. Furthermore, the genotype error classification indicates the probability of an error associated with a genotype for the nucleotide base call (e.g., variant call).
[0164] In some cases, the call recalibration machine learning model 714 is an ensemble of gradient boosted trees that process sequencing metrics to generate variant call classifications 716. For example, the call recalibration machine learning model 714 includes a series of weak learners, such as non-linear decision trees, that are trained in logistic regression to generate variant call classifications 716. In some cases, the call recalibration machine learning model 714 includes metrics within the various trees that define how the call recalibration machine learning model 714 processes sequencing metrics to generate variant call classifications 716. Further details regarding training of the call recalibration machine learning model 714 are provided below with reference to FIG. 8.
[0165] In certain embodiments, the call recalibration machine learning model 714 is a different type of machine learning model, such as a neural network, a support vector machine, or a random forest. For example, if the call recalibration machine learning model 714 is a neural network, the call recalibration machine learning model 714 includes one or more layers, each having neurons that constitute the layer for processing the sequencing metrics. In some cases, the call recalibration machine learning model 714 generates the variant call classifications 716 by extracting latent vectors from the sequencing metrics, passing the latent vectors from layer to layer (or neuron to neuron), and manipulating the vectors until it generates the variant call classifications 716 (e.g., as a set of three distinct classifications) utilizing an output layer (e.g., one or more fully connected layers).
[0166] As alluded to above, in some embodiments, the call recalibration system 106 can utilize multiple call recalibration machine learning models together. For example, the call recalibration system 106 utilizes the call recalibration machine learning model 714 to generate a first set of variant call classifications and further utilizes a second call recalibration machine learning model (e.g., having the same or a different architecture) to generate a second set of variant call classifications. For example, the call recalibration system 106 utilizes two (or more) different call recalibration machine learning models in parallel, each trained with a different random seed (e.g., for different biases to process the data differently) to result in different variant call classifications from the same sequencing metrics.
[0167] In some embodiments, the call recalibration system 106 further generates a combined set of variant call classifications from the different variant call classifications generated via the different call recalibration machine learning models. In some cases, the base call recalibration system 106 generates a variant call classification (e.g., variant call classification 716) from the first set of variant call classifications and the second set of variant call classifications generated from the first call recalibration machine learning model and the second call recalibration machine learning model, respectively. For example, the call recalibration system 106 determines an average or weighted combination of the first set of variant call classifications and the second set of variant call classifications to generate a combined variant call classification for recalibrating the nucleotide base calls. In some embodiments, the call recalibration system 106 determines an average of each variant call classification across each call recalibration machine learning model and renormalizes the average variant call classification. In other embodiments, the call recalibration system 106 learns linear weights and adapts the weights to minimize an overall error or loss for the variant call classifications. In yet another embodiment, the call recalibration system 106 weights the variant call classification for each call recalibration machine learning model based on the inverse of the average error across the models.
[0168] In one or more implementations, the call recalibration system 106 further utilizes a meta-model following the call recalibration machine learning models. For example, the call recalibration system 106 utilizes a classification combiner machine learning model to combine the variant call classifications generated from each call recalibration machine learning model, such as by selecting weights to apply to the variant call classifications generated by each call recalibration machine learning model. Indeed, in some cases, the call recalibration system 106 trains the classification combiner machine learning model to determine, select, or predict the respective weights for the call recalibration machine learning models that result in the highest accuracy or the lowest loss.
[0169] When generating the variant call classification 716, in some embodiments, the call recalibration system 106 generates the variant call classification by utilizing statistics to summarize the mapping quality distributions (e.g., comparative mapping quality distribution metrics) of the reference and alternative supporting reads. For example, the call recalibration system 106 can determine and utilize the average of the MAPQ for the reads supporting the alternative allele as the variant call classification. In these or other embodiments, the call recalibration machine learning model 714 learns from the data that if the MAPQ of the alternative allele is low and the depth metric is high relative to the other MAPQs and depth metrics in the distribution, the resulting nucleotide base call is more likely to be a false positive variant. In fact, as the probability of a false positive variant increases, the MAPQ metric may decrease.
[0170] As a further example of utilizing the call recalibration machine learning model 714 to generate variant call classifications 716, in some cases, the base call recalibration system 106 compares a mapping quality (e.g., MAPQ) associated with a nucleotide read (e.g., from a sequencing metric) to a mapping quality threshold. For example, the call recalibration system 106 utilizes a mapping quality threshold, such as a threshold difference between the best alignment score and the next best alignment score. Upon determining that the mapping quality does not meet the threshold, the call recalibration system 106 adjusts one or more of the variant call classifications 716 accordingly. For example, the call recalibration system 106 increases the probability of a genotype error and / or a false positive error based on whether the mapping quality meets a corresponding threshold.
[0171] In addition to (or as an alternative to) the method of generating variant call classifications 716 just described, the call recalibration system 106 can (i) utilize an accumulation of statistical analyses across complex functions (depending on the architecture of the call recalibration machine learning model 714) to determine how to best fit the data (e.g., based on relationships between various metrics) or (ii) compare other metrics, such as read depth, base quality, or others associated with nucleotide base calls (e.g., from sequencing metrics), to corresponding thresholds. The call recalibration system 106 further generates variant call classifications 716 accordingly. For example, in some embodiments, the call recalibration system 106 trains the call recalibration machine learning model 714 to minimize losses generated from several (different types of) sequencing metrics to determine weights and biases that best fit (e.g., result in reduced or minimized losses) the data to generate variant call classifications 716. As another example, upon determining that the read depth does not meet a read depth threshold (e.g., a maximum read depth corresponding to a particular genomic coordinate or across all genomic coordinates generally), the call recalibration system 106 increases the genotype error probability and / or increases or decreases the false positive probability and true positive probability for the corresponding nucleotide base call.
[0172] In addition to generating variant call classifications 716, as further illustrated in FIG. 7, the call recalibration system 106 performs data field generation 718. More specifically, the call recalibration system 106 utilizes the variant caller component 710 of the call generation model 722 to generate data fields for the nucleotide base calls corresponding to the variant call file and modifies or maintains values of such data fields based on the variant call classifications 716. For example, the call recalibration system 106 modifies various metrics, such as quality metrics, mapping metrics, or other metrics associated with the nucleotide base calls. In certain embodiments, the nucleotide base calls are represented or defined by a variant call file 720 that includes metrics corresponding to data fields, such as call quality metrics corresponding to the call quality fields, genotype metrics corresponding to the genotype fields, and genotype quality metrics corresponding to the genotype quality fields.
[0173] In certain embodiments, the call recalibration system 106 utilizes the variant caller component 710 in conjunction with the variant call classification 716 to generate (data fields for) the nucleotide base calls. For example, the call recalibration system 106 utilizes the variant caller component 710 to generate data fields for various metrics of the nucleotide base calls, such as the nucleotides included in the call, the call quality (QUAL), the genotype (GT), and the genotype quality (GQ).
[0174] In addition to generating nucleotide base calls via the call generation model 722, the call recalibration system 106 also recalibrates or corrects the nucleotide base calls via the variant call classifications 716 from the call recalibration machine learning model 714. In one or more implementations, the call recalibration system 106 corrects the nucleotide base calls by correcting or recalibrating data fields for one or more of the metrics associated with the nucleotide base calls (e.g., as included in the variant call file 720). For example, the call recalibration system 106 determines updated values for metrics such as call quality, genotypes, and genotype quality from the variant call classifications 716. In effect, the call recalibration system 106 combines or compares the variant call classifications 716 to recalibrate the corresponding metrics of the nucleotide base calls included in the variant call file 720.
[0175] To update or recalibrate the call quality metrics associated with the nucleotide base calls, the call recalibration system 106 determines how each of the variant call classifications 716 impacts or influences the base call quality metrics and adjusts the base call quality metrics accordingly. For example, the call recalibration system 106 determines that a high probability for a genotype error results in a lower overall genotype quality and possibly a different overall call quality. As another example, the call recalibration system 106 determines that a high probability for a false positive variant results in a lower overall call quality. As yet another example, the call recalibration system 106 determines that a high probability for a true positive variant results in a higher overall (variant) call quality. As a further example, if the call recalibration system 106 determines a high probability for a genotype error (e.g., higher than the other two variant call classifications 716), the call recalibration system 106 determines that the nucleotide base call is most likely a true variant with an incorrect genotype. Thus, the call recalibration system 106 updates the genotypes along with the genotype quality and call quality associated with the nucleotide base calls.
[0176] In one or more implementations, the call recalibration system 106 generates a combination (e.g., a weighted combination or average) of variant call classifications 716 to recalibrate the call quality metrics. In particular, the call recalibration system 106 weights the false positive classification, the genotype error classification, and the true positive classification according to their respective impact on the (variant) call quality. In some cases, the call recalibration system 106 weights each variant call classification equally, while in other cases, the call recalibration system 106 determines different weights for each variant call classification. In any case, the call recalibration system 106 determines a weighted combination or weighted average of the variant call classifications 716 to recalibrate (increase or decrease) the call quality metrics for the nucleotide base call (e.g., the initial variant call).
[0177] To update or recalibrate the genotypic metrics associated with the nucleotide base calls (e.g., in the GT field of the variant call file 720), the call recalibration system 106 utilizes one or more of the variant call classifications 716. For example, the call recalibration system 106 compares three variant call classifications 716 (e.g., a false positive classification, a genotypic error classification, and a true positive classification) to determine which of the variant call classifications 716 has the highest probability. In some cases, the call recalibration system 106 utilizes the variant call classification with the highest probability to recalibrate the genotypic metrics (e.g., from 0 corresponding to the reference base to 1 corresponding to the first alternative supporting read). For example, if the call recalibration system 106 determines the highest probability for a false positive classification, the call recalibration system 106 recalibrates the genotypic metrics accordingly. As another example, if the call recalibration system 106 determines the highest probability for a true positive classification, the call recalibration system 106 recalibrates (or refrains from recalibrating) the genotype metrics.
[0178] In other embodiments, the call recalibration system 106 utilizes only the genotype error probability to modify the genotype metric. For example, if the call recalibration system 106 determines a high genotype error probability, the call recalibration system 106 recalibrates the genotype metric to indicate a different genotype for the nucleotide base call.
[0179] To update or recalibrate a genotype quality metric associated with a nucleotide base call (e.g., in the GQ field of the variant call file 720), the call recalibration system 106 utilizes one or more of the variant call classifications 716. More specifically, the call recalibration system 106 determines how each of the variant call classifications 716 affects the genotype quality metric and recalibrates the genotype quality metric accordingly (e.g., by increasing or decreasing the quality score between 0-10 or 0-100, or some other scale). For example, the call recalibration system 106 determines that a higher genotype error probability indicates (generally) a lower genotype quality metric, and the call recalibration system 106 reduces the metric accordingly.
[0180] In some cases, the call recalibration system 106 determines a combination (e.g., a weighted combination or weighted average) of the variant call classifications 716 to modify the genotype quality metric. For example, the call recalibration system 106 determines the combined effect of the variant call classifications 716 on the genotype quality metric. As another example, the call recalibration system 106 determines the individual impact each variant call classification has on the genotype quality metric and weights each variant call classification accordingly. The call recalibration system 106 further recalibrates the genotype quality metric by increasing or decreasing its value based on the indicated probability associated with each of the variant call classifications 716.
[0181] As described, the call recalibration system 106 generates variant call classifications 716 and nucleotide base calls from the same set of sequencing metrics (or a subset of sequencing metrics shared between the call recalibration machine learning model 714 and the call generation model 722). In effect, the call recalibration system 106 utilizes the call recalibration machine learning model 714 to generate variant call classifications 716 from sequencing metrics while also generating nucleotide base calls for the sample sequences. In effect, the call recalibration system 106 can operate the call recalibration machine learning model 714 in parallel with the call generation model 722 to generate metrics for nucleotide base calls and variant call classifications 716 to recalibrate the generated metrics.
[0182] As further illustrated in FIG. 7, the call recalibration system 106 generates a variant call file 720. In particular, the call recalibration system 106 generates the variant call file 720 that represents or defines nucleotide base calls from sequencing metrics corresponding to genomic coordinates. As shown, the variant call file 720 includes various call metrics, such as a call quality metric (QUAL), a genotype metric (GT), and a genotype quality metric (GQ). To generate the variant call file 720, the call recalibration system 106 utilizes a call generation model 722 to generate metrics for the nucleotide base calls and utilizes variant call classifications 716 from a call recalibration machine learning model 714 to recalibrate the nucleotide base calls, as described.
[0183] In one or more implementations, the call recalibration system 106 updates or otherwise modifies data fields for the variant call file 720 according to a particular algorithm. After modifying such data fields, the call recalibration system 106 can generate the variant call file 720 (e.g., a post-filter variant call file) to include metrics that reflect the updated data fields for QUAL, GT, and GQ. For example, in some cases, the call recalibration system 106 updates the QUAL field for each variant based on the probability of a false positive variant (e.g., a false positive classification). As indicated above, in some cases, QUAL indicates the probability of a certain variant (or other nucleotide base call) being present at a given position, as measured on the PHRED scale.
[0184] In addition, if the call recalibration system 106 determines that the highest probability among the three variant call classifications as the variant call classification 716 is a genotype error classification (e.g., the probability of a het / hom error), the call recalibration system 106 updates the GQ field while preserving or maintaining the GT field. Specifically, in some embodiments, the call recalibration system 106 updates the GQ field based on the true positive classification (e.g., the probability of a true genotype).
[0185] Additionally, if the call recalibration system 106 determines that the highest probability among the variant call classifications 716 is a true positive classification, then in some cases the call recalibration system 106 updates both the GQ and GT fields. Specifically, the call recalibration system 106 updates the GQ field based on the genotype error classification and also updates the GT field to switch the genotype depending on whether the existing GT is 0 / X or X / X (X is a non-zero value).
[0186] If the call recalibration system 106 determines that neither the true positive classification nor the genotype error classification has the highest probability among the variant call classifications 716, in some embodiments, the call recalibration system 106 updates the GQ field. In other words, if the call recalibration system 106 determines that the false positive classification has the highest probability, the call recalibration system 106 updates the GQ field. In particular, the call recalibration system 106 updates the GQ field based on the probability indicated by the true positive classification.
[0187] As alluded to above, in some embodiments, the call recalibration system 106 increases or decreases a base call quality metric (e.g., Q-score) for a nucleotide base call. Based on the variant call classification 716, for example, the call recalibration system 106 increases a base call quality metric for a nucleotide base call that would not have previously passed the quality filter and determines that the increased base call quality metric now passes the quality filter. In some such cases, the call recalibration system 106 includes the nucleotide base call having such an increased base call quality metric (that passes the quality filter) in the post-filter variant call file. In contrast, in other cases, the call recalibration system 106 decreases a base call quality metric for a nucleotide base call that would have previously passed the quality filter and determines that the decreased base call quality metric now does not pass the quality filter. In some such cases, the call recalibration system 106 excludes nucleotide base calls with reduced base call quality metrics (that do not pass the quality filter) from the post-filter variant call file, but includes such nucleotide base calls with reduced base call quality metrics in the pre-filter variant call file.
[0188] For example, the base call recalibration system 106 can remove false positive variant calls and recover false negative variant calls by modifying the corresponding base call quality metric. To remove false positives, in some cases, the call recalibration system 106 reduces the base call quality metric of a nucleotide base call that initially passed the quality filter based on the variant call classification 716 from the call recalibration machine learning model 714. Based on determining that the reduced base call quality metric falls below a threshold metric (e.g., a Q-score of 3.0 or 10.0), the call recalibration system 106 determines that the nucleotide base call no longer passes the quality filter. Thus, the call recalibration system 106 filters out or removes a false positive nucleotide base call that initially passed the filter by modifying its base call quality metric.
[0189] In addition to removing false positive variant calls based on changes to the base call quality metrics, the call recalibration system 106 can remove false positive variant calls based on changes to the genotype. To remove false positives, in some cases, the call recalibration system 106 changes the genotype of an initial nucleotide base call that indicates a different nucleotide base from the reference base (e.g., GT=1 or 2) to the genotype of an updated nucleotide base call that indicates the same nucleotide base as the reference base (e.g., GT=0) based on the variant call classification 716 from the call recalibration machine learning model 714. Based on the genotype being the same as the reference base, the call recalibration system 106 does not identify the nucleotide base call as a variant, and in some cases excludes the data for the nucleotide base call from the variant call file.
[0190] To recover the false negatives, the call recalibration system 106 increases the base call quality metric of the nucleotide base call that did not initially pass the quality filter based on the variant call classification 716 from the call recalibration machine learning model 714. Based on determining that the increased base call quality metric exceeds a threshold metric, the call recalibration system 106 determines that the nucleotide base call passes the quality filter. Thus, the call recalibration system 106 recovers the false negative nucleotide base call that was initially filtered out by modifying its base call quality metric.
[0191] In addition to recovering false negatives based on changes to the base call quality metrics, the call recalibration system 106 can recover false negative variant calls based on changes to the genotype. To recover false negatives, in some cases, the call recalibration system 106 changes the genotype of an initial nucleotide base call that indicates the same nucleotide base as the reference base (e.g., GT=0) to a different genotype of an updated nucleotide base call that indicates a different nucleotide base from the reference base (e.g., GT=1 or 2) based on the variant call classification 716 from the call recalibration machine learning model 714. Based on the different genotype of the updated nucleotide base call and the passed base call quality metrics, the call recalibration system 106 identifies the nucleotide base call as a variant and includes the nucleotide base call in the variant call file.
[0192] Indeed, in some implementations, the call recalibration system 106 utilizes the call generation model 722 and the call recalibration machine learning model 714 to operate in a particular order. For example, the call recalibration system 106 generates a FASTQ file by converting a BCL file to a FASTQ. In addition, the call recalibration system 106 (then) utilizes the mapping and alignment component 708 of the call generation model 722 to map and align nucleotide bases from the sample nucleotide sequence. In some cases, the call recalibration system 106 maps and aligns nucleotide bases of the sample sequence relative to a reference sequence (e.g., a reference genome) and / or various alternative supporting reads.
[0193] As described herein, after mapping and alignment, the call recalibration system 106 then utilizes the variant caller component 710 of the call generation model 722 to generate initial nucleotide base calls for the sample sequence corresponding to the particular genomic coordinates based on various sequencing metrics. Thereafter, or simultaneously, the call recalibration system 106 also applies a call recalibration machine learning model 714 to generate variant call classifications 716 from the mapping and alignment, the sequencing metrics extracted through variant calling, and / or from other sources as described above. Based on the variant call classifications 716, the call recalibration system 106 recalibrates the nucleotide base calls (e.g., by modifying various data fields corresponding to particular metrics of the nucleotide base calls, such as QUAL, GT, and GQ).
[0194] In some cases, the call recalibration system 106 further applies a quality filter to the nucleotide base calls to determine whether the nucleotide base calls pass the quality filter (e.g., a Q20 or other Q score hard pass filter). The call recalibration system 106 then identifies a subset of the nucleotide base calls that represent variants from the reference base and that pass the quality filter. The call recalibration system 106 further generates a modified or updated variant call file (e.g., variant call file 720) that includes the subset of the nucleotide base calls and recalibrated metrics for the subset of the nucleotide base calls, such as an updated QUAL metric, an updated GT metric, and / or an updated GQ metric.
[0195] As mentioned above, in certain embodiments, the call recalibration system 106 trains or tunes a call recalibration machine learning model (e.g., the call recalibration machine learning model 714). In particular, the call recalibration system 106 utilizes an iterative training process to adapt the call recalibration machine learning model by adjusting or adding decision trees or learning parameters that result in an accurate variant call classification (e.g., the variant call classification 716). Figure 8 illustrates training a call recalibration machine learning model in accordance with one or more embodiments.
[0196] As illustrated in FIG. 8, the call recalibration system 106 accesses sample sequencing metrics 804 from a database 802 (e.g., database 116). For example, the call recalibration system 106 accesses sample sequencing metrics including sample read-based metrics, sample external source sequencing metrics, and sample call model generation sequencing metrics. In some cases, the sample sequencing metrics 804 have corresponding ground truth variant call files 816 associated with them, which indicate the actual nucleotide base calls and various metrics thereof resulting from the sample sequencing metrics 804. For example, the call recalibration system 106 utilizes the sample sequencing metrics 804 and a ground truth variant call file from a training dataset from the Food and Drug Administration, referred to as the PrecisionFDA dataset. In some cases, the sample sequencing metrics 804 include a subset of the sample sequencing metrics for each nucleotide base call in the ground truth variant call file. The ground truth variant call file can have ground truth variant calls (e.g., genotype metrics in the genotype field) and / or ground truth base calls corresponding to each subset of sample sequencing metrics.
[0197] 8, the call recalibration system 106 generates predicted variant call classifications 808 based on the sample sequencing metrics 804. Specifically, the call recalibration system 106 utilizes a call recalibration machine learning model 806 (e.g., call recalibration machine learning model 714) to generate the predicted variant call classifications 808. Indeed, in some embodiments, the call recalibration machine learning model 806 generates a set of three predicted variant call classifications 808 including a predicted false positive classification, a predicted genotype error classification, and a predicted true positive classification. Thus, the predicted variant call classifications 808 can take the form of any of the variant call classifications described above.
[0198] Based on the predicted variant call classification 808, the call recalibration system 106 determines nucleotide base calls and generates a modified variant call file 810 including the nucleotide base calls and corresponding fields. As indicated above, the call recalibration system 106 can (i) utilize a call generation model to generate initial nucleotide base calls and (ii) utilize a call recalibration machine learning model 806 to modify data fields corresponding to the variant call file for the nucleotide base calls. Such modified or recalibrated values are output in the modified variant call file 810, for example, by the call generation model. For example, the call recalibration system 106 determines recalibration values for certain metrics in the modified variant call file 810, including a call quality metric (QUAL), a genotype metric (GT), and a genotype quality metric (GQ).
[0199] As further illustrated in FIG. 8 , the call recalibration system 106 performs a comparison 812. Specifically, the call recalibration system 106 performs a comparison 812 between (i) the variant nucleotide base calls and / or data fields in the corrected variant call file 810 and (ii) the variant nucleotide base calls and / or data fields in the ground truth variant call file 816. In some embodiments, the call recalibration system 106 utilizes a loss function 814 to compare (e.g., determine a measure of error or loss between) the variant nucleotide base calls and / or data fields from the two variant call files. For example, if the call recalibration machine learning model 806 is an ensemble of gradient boosted trees, the call recalibration system 106 utilizes a mean squared error loss function (e.g., for regression) and / or a logarithmic loss function (e.g., for classification) as the loss function 814.
[0200] In contrast, in embodiments in which the call recalibration machine learning model 806 is a neural network, the call recalibration system 106 may utilize a cross-entropy loss function, an L1 loss function, or a mean squared error loss function as the loss function 814. For example, the call recalibration system 106 utilizes the loss function 814 to determine the differences between variant nucleotide base calls and / or data fields from the corrected variant call file 810 and the ground truth variant call file 816.
[0201] 8, the call recalibration system 106 performs model fitting 818. In particular, the call recalibration system 106 adapts the call recalibration machine learning model 806 based on the comparison 812. For example, the call recalibration system 106 makes modifications or adjustments to the call recalibration machine learning model 806 to reduce a measure of loss from a loss function 814 for subsequent training iterations.
[0202] In the case of gradient boosted trees, for example, the call recalibration system 106 trains the call recalibration machine learning model 806 on the gradient of the error determined by the loss function 814. For example, the call recalibration system 106 solves a (e.g., infinite-dimensional) convex optimization problem while regularizing the objective function to avoid over-fitting. In one particular implementation, the call recalibration system 106 scales the gradient to emphasize correction for under-represented classes (e.g., when there are significantly more true positive variant calls than false positive variant calls).
[0203] In some embodiments, the call recalibration system 106 adds a new weak learner (e.g., a new boosted tree) to the call recalibration machine learning model 806 at each successive training iteration as part of solving the optimization problem. For example, the call recalibration system 106 finds a feature (e.g., a sequencing metric) that minimizes the loss from the loss function 814 and adds that feature to the current iteration's tree or starts building a new tree with that feature.
[0204] In addition to or as an alternative to gradient boosting decision trees, the call recalibration system 106 trains a logistic regression to learn parameters for generating one or more variant call classifications, such as true positive classifications. To avoid over-fitting, the call recalibration system 106 further regularizes based on hyperparameters such as learning rate, stochastic gradient boosting, number of trees, tree depth, complexity penalization, and L1 / L2 regularization.
[0205] In embodiments where the Cole recalibrated machine learning model 806 is a neural network, the Cole recalibration system 106 performs model fitting 818 by modifying the internal parameters (e.g., weights) of the Cole recalibrated machine learning model 806 to reduce a loss measure for the loss function 814. In effect, the Cole recalibration system 106 modifies how the Cole recalibrated machine learning model 806 analyzes and passes data between layers and neurons by modifying the internal network parameters. Thus, over multiple iterations, the Cole recalibration system 106 improves the accuracy of the Cole recalibrated machine learning model 806.
[0206] Indeed, in some cases, the call recalibration system 106 repeats the training process illustrated in FIG. 8 multiple times. For example, the call recalibration system 106 repeats the iterative training by selecting a new set of sequencing metrics for each nucleotide base call along with the corresponding ground truth nucleotide base call in the corresponding ground truth variant call file. The call recalibration system 106 further generates a new set of predicted variant call classifications for each iteration along with the new modified variant call file. As described above, the call recalibration system 106 also compares the variant nucleotide base calls and / or data fields from the modified variant call file in each iteration with the corresponding variant-nucleotide base calls and / or data fields from the corresponding ground truth variant call file, and further performs model fitting 818. The call recalibration system 106 repeats this process until the call recalibration machine learning model 806 generates predicted variant call classifications that result in variant calls that meet a threshold measure of loss. In some embodiments, the call recalibration system 106 performs the training process of FIG. 8 on homozygous reference coordinates to update or correct variant calls at those coordinates, thereby recovering false negative variant calls (based on simulating haploid data from diploid data and correcting the inputs and outputs of the call recalibration machine learning model 806, as described).
[0207] As mentioned above, in certain described embodiments, the call recalibration system 106 generates and provides contribution measures associated with the sequencing metrics. In particular, the call recalibration system 106 determines each contribution measure that indicates how impactful an individual sequencing metric is in determining a particular nucleotide base call. Figure 9 shows an example visualization of the contribution measures for sequencing metrics associated with nucleotide base calls, according to one or more embodiments.
[0208] As illustrated in Figure 9, the client device 108 displays a contribution measure interface 902 that includes individual depictions of the contribution measures associated with corresponding sequencing metrics. In effect, the call recalibration system 106 determines the contribution measures of a sequencing metric based on how much impact or influence the sequencing metric has on the final nucleotide base call. Unlike many existing sequencing systems that utilize deep learning architectures, the structure of the call generation model used by the call recalibration system 106 facilitates the determination of such a contribution measure for each metric.
[0209] For example, the call recalibration system 106 determines the contribution measure by determining a Shapley Additive Description (SHAP) value for each of the sequencing metrics for a nucleotide base call. Specifically, the call recalibration system 106 determines the SHAP value by determining the impact of the sequencing metric compared to the results of a baseline value (e.g., the baseline value of the sequencing metric). As illustrated in FIG. 9, the call recalibration system 106 determines the contribution measure for several enumerated sequencing metrics, where the thicker (e.g., bulbier) portion of the graph for each sequencing metric (roughly) indicates its contribution measure.
[0210] 9, the call recalibration system 106 may similarly rank the sequencing metrics according to their contribution measures. For example, the call recalibration system 106 may determine that the contribution for the mapq_p metric is the highest among those displayed in the contribution-measure interface 902, followed by the qual metric, the gt0 metric, etc. down the list.
[0211] As mentioned above, in certain described embodiments, the call recalibration system 106 improves accuracy over existing sequencing systems. In particular, the call recalibration system 106 reduces false positive variant nucleotide base calls and false negative variant nucleotide base calls compared to existing sequencing systems. In fact, by utilizing a call recalibration machine learning model to recalibrate nucleotide base calls, the call recalibration system 106 further improves over previous versions of the call generation model that did not utilize a call recalibration machine learning model (but still outperformed other systems). Figures 10A-10B show graphs and tables of experiments that demonstrate the improvement in accuracy of the call recalibration system 106 compared to several existing systems.
[0212] For reference, as shown in Figures 10A-10B and 11A-11B, the name "non-recalibrated system 1" refers to an existing sequencing system that uses a linear reference genome for variant calling. In contrast, the name "non-recalibrated system 2" refers to an existing sequencing system that uses a graph reference genome for variant calling. Additionally, the name "call recalibration system 1" refers to an embodiment of a call recalibration system 106 that is not configured for nucleotide base calling at multi-allelic genome coordinates, haploid genome coordinates, and reference genome coordinates that may be homozygous. In contrast, the name "call recalibration system 2" refers to an embodiment of a call recalibration system 106 that is not configured for nucleotide base calling at multi-allelic genome coordinates, haploid genome coordinates, and reference genome coordinates that may be homozygous.
[0213] As illustrated in FIG. 10A, graph 1002 shows several receiver operating characteristic (ROC) curves comparing SNP false positives for two variations of the call recalibration system 106 with SNP false positives for two non-recalibrated systems. Graph 1002 depicts the portion of the ROC curve that represents the sensitivity to detected false positive variants, where sensitivity represents the number of correctly determined true positive variant calls divided by the sum of true positive variant calls and false positive variant calls. In particular, graph 1002 shows ROC curves for different embodiments of the call recalibration system 106 that utilize a call recalibration machine learning model, namely "Call Recalibration System 1" and "Call Recalibration System 2" as described above. The experiments were performed using PrecisionFDA truth sets (e.g., PrecisionFDA HG002 high confidence truth set). In general, curves moving toward the upper left in graph 1002 are more accurate. As shown, embodiments of call recalibration system 106 exhibit improved accuracy over each of the three non-recalibrated systems, with higher sensitivity and fewer false positive variant calls in comparison. The increased sensitivity, as shown by the improvement between the ROC curves of call recalibration system 1 and call recalibration system 2, is due in part to the recovery of false negative variant calls at genomic coordinates that would have been identified as homozygous reference genotypes by another sequencing system.
[0214] In addition, graph 1004 shows several ROC curves comparing non-SNP (e.g., indel) false positive variant calls for different embodiments of call recalibration system 106 with those of paired non-recalibrated systems (non-recalibrated system 1 and non-recalibrated system 2). Graph 1004 shows ROC curves that represent sensitivity to detected false positive variants. In particular, graph 1004 shows ROC curves for one embodiment of call recalibration system 106 configured for nucleotide base calling at multi-allelic genomic coordinates, haploid genomic coordinates, and reference genomic coordinates that would be homozygous, which removes or reduces the bumps or jogs that are prevalent in the non-recalibrated system (instead continuing smoothly upward on a nearly vertical trajectory) with a sensitivity of about 0.4. Indeed, due at least in part to improvements in multi-allelic genomic coordinates, one embodiment of the call recalibration system 106 (here, call recalibration system 2) exhibits fewer false positive variant calls with similar sensitivity compared to one or more non-recalibrated systems that do not recalibrate multi-allelic variants (e.g., non-recalibrated system 2). Experiments were performed using a PrecisionFDA truth set (e.g., PrecisionFDA HG002 high confidence truth set).
[0215] 10B, table 1006 corresponds to graph 1002 and table 1008 corresponds to graph 1004. The numbers in table 1006 and table 1008 are taken at the best F-value point for the curves in each of graphs 1002 and 1004, respectively. As shown in table 1006, both embodiments of call recalibration system 106 make fewer false negative variant calls (FN), fewer false positive variant calls (FP), and more true positives (TP) than either of the non-recalibrated systems. For example, at the best F-value point, an embodiment of call recalibration system 106 (shown as call recalibration system 2 and configured for nucleotide base calling at multi-allelic genomic coordinates, haploid genomic coordinates, and reference genomic coordinates that would be homozygous) produces 7309 false negative variant calls and 2801 false positive variant calls, as shown in table 1006. Another embodiment of the call recalibration system 106 (shown as call recalibration system 1, but not configured for nucleotide base calling at multi-allelic and homozygous reference genome coordinates) produces 7717 false negative variant calls and 3216 false positive variant calls. Similarly, the call recalibration system 106 also has fewer het / hom errors, better recall, and higher precision.
[0216] As illustrated in table 1008, embodiments of the call recalibration system 106 also outperform non-recalibrated systems for non-SNP scenarios. For example, at the best F-value point in table 1008, an embodiment of the call recalibration system 106 (shown as call recalibration system 2 and configured for nucleotide base calling at multi-allelic genomic coordinates, haploid genomic coordinates, and reference genomic coordinates that would be homozygous) generates 513 false positive variant calls, while another embodiment of the call recalibration system 106 generates 618 false positive variant calls. Both non-recalibrated systems generate far more false positive variant calls. An embodiment of the call recalibration system 106 also has a higher accuracy than either of the non-recalibrated systems.
[0217] In addition to the improved diploid accuracy shown in Figures 10A-10B, Figures 11A-11B show an improved haploid accuracy. Specifically, graphs 1102 and 1104 each show two ROC curves, one for the call recalibrated system 106 and one for the non-recalibrated system. For example, graph 1102 shows the ROC curve for SNPs, while graph 1104 shows the ROC curve for non-SNPs (e.g., indels). In both cases, as a result of the improved accuracy in the haploid coordinates, the call recalibrated system 106 has a higher sensitivity and fewer false positive variant calls compared to the non-recalibrated system. Indeed, in each of graphs 1102 and 1104, the ROC curve of the call recalibrated system 106 is improved, and the best F-value point located at the cross section shows fewer false positive variant calls at (approximately) the same sensitivity compared to the non-recalibrated system. The experiments in graphs 1102 and 1104 were performed using the PrecisionFDA truth set.
[0218] 11B, table 1106 corresponds to graph 1102, and table 1108 corresponds to graph 1104. Indeed, table 1106 shows SNP results for the call recalibrated system 106 compared to a non-recalibrated system at the best F-point. As shown, the call recalibrated system 106 has more true positives, fewer false negative variant calls, fewer false positive variant calls, higher recall, and higher precision for SNPs. Looking at table 1108, the call recalibrated system 106 also produces more true positives, fewer false negative variant calls, fewer false positive variant calls, higher recall, and higher precision for non-SNPs than the non-recalibrated system (at the best F-point).
[0219] 12-14, which show exemplary flow charts of each of a series of operations for generating final nucleotide base calls based on variant call classifications from a call recalibration machine learning model, according to one or more embodiments. Although FIGS. 12-14 illustrate operations according to one embodiment, alternative embodiments may omit, add, rearrange, and / or modify any of the operations shown in FIGS. 12-14. The operations of FIGS. 12-14 may be performed as part of a method. Alternatively, a non-transitory computer-readable storage medium may include instructions that, when executed by one or more processors, cause a computing device to perform the operations shown in FIGS. 12-14. In a further embodiment, a system includes at least one processor and a non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the system to perform the operations of FIGS. 12-14.
[0220] 12, the series of operations 1200 includes an operation 1202 of determining sequencing metrics for a multi-allelic genomic coordinate. In particular, operation 1202 may include determining sequencing metrics for nucleotide base calls of nucleotide reads that correspond to the genomic coordinate of the sample nucleotide sequence.
[0221] Additionally, the set of operations 1200 includes an operation 1204 of generating a set of variant call classifications for the multi-allelic genomic coordinate. In particular, operation 1204 may include utilizing a call recalibration machine learning model to generate a set of variant call classifications based on sequencing metrics, the set including a reference probability of a homozygous reference genotype at the multi-allelic genomic coordinate, a different genotype probability of a genotype error at the multi-allelic genomic coordinate, and an accurate variant probability of an accurate variant call genotype at the multi-allelic genomic coordinate.
[0222] For example, generating a reference probability may include determining a probability that a genotype at a multi-allelic genomic coordinate is a homozygous genotype with respect to a reference genome. Generating different genotype probabilities may include determining a probability that a predicted genotype for a multi-allelic genomic coordinate is an incorrect genotype or an incorrect allele in the predicted genotype. Generating accurate variant probabilities may include determining a probability that a predicted genotype for a multi-allelic genomic coordinate is correct as originally determined by the call generation model.
[0223] 12, the series of operations 1200 includes an operation 1206 of determining a final nucleotide base call for the multi-allelic genomic coordinate. In particular, operation 1206 may include determining a final nucleotide base call for the multi-allelic genomic coordinate based on the set of one or more variant call classifications. For example, operation 1206 may include predicting two nucleotide bases from three or more candidate alleles at the multi-allelic genomic coordinate.
[0224] The series of operations 1200 may also include an operation of modifying a base call quality metric or a genotype quality metric based on the set of variant call classifications. Further, the series of operations 1200 may include an operation of generating a variant call file including the modified base call quality metric or the modified genotype quality metric. In addition, the series of operations 1200 may include an operation of generating updated genotype likelihoods for candidate nucleotide base calls of alleles at the multi-allelic genomic coordinates. In some embodiments, the series of operations 1200 includes an operation of generating a variant call file including the updated genotype likelihoods.
[0225] 13, the series of operations 1300 includes an operation 1302 of determining sequencing metrics for nucleotide base calls corresponding to genomic coordinates of a haploid nucleotide sequence. In particular, operation 1302 may include determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to genomic coordinates of a haploid nucleotide sequence from a sample.
[0226] The set of operations 1300 may also include an operation 1304 of generating a first genotype probability and a second genotype probability. In particular, operation 1304 may include generating a first genotype probability of the first genotype at the genomic coordinate and a second genotype probability of the second genotype at the genomic coordinate utilizing a call recalibration machine learning model and based on the sequencing metric. In some cases, operation 1304 includes the operation of generating the first genotype probability including generating a probability that the first genotype at the genomic coordinate is a haploid reference genotype and the operation of generating the second genotype probability including generating a probability that the second genotype at the genomic coordinate is a haploid substitute genotype.
[0227] Generating the first genotype probability may include utilizing a layer of a call recalibration machine learning model to modify a homozygous reference probability of a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability of the reference genotype at the genomic coordinate. Generating the second genotype probability may include utilizing a layer of a call recalibration machine learning model to modify a homozygous substitution probability of a homozygous alternative genotype at the genomic coordinate to generate a haploid substitution probability of an alternative genotype at the genomic coordinate.
[0228] In some cases, operation 1304 includes utilizing one or more layers of a call recalibration machine learning model to generate, for the genomic coordinate, a first confidence score corresponding to the first genotype, a second confidence score corresponding to the second genotype, and a third confidence score corresponding to the third genotype. Operation 1304 may also include excluding the second confidence score corresponding to the second genotype and normalizing the first confidence score and the third confidence score utilizing a softmax model to generate the first genotype probability and the second genotype probability.
[0229] As further shown, the set of operations 1300 may include operation 1306 of determining a final nucleotide base call indicative of a haploid genotype. In particular, operation 1306 may include determining a final nucleotide base call indicative of a haploid genotype for the genomic coordinate based on the first genotype probability and the second genotype probability. For example, operation 1306 may include determining one of a haploid substitute genotype, a modified base call quality metric, a modified genotype metric, and a modified genotype quality metric for the genomic coordinate based on determining that the second genotype probability exceeds the first genotype probability, or determining one of a haploid reference genotype, a modified base call quality metric, and a modified genotype quality metric for the genomic coordinate based on determining that the first genotype probability exceeds the second genotype probability.
[0230] In some embodiments, the series of operations 1300 includes converting the haploid reference genotype calls generated by the call generation model to diploid homozygous reference genotype calls as inputs for a call recalibration machine learning model. The series of operations 1300 may include converting the haploid alternative genotype calls generated by the call generation model to diploid homozygous alternative genotype calls as inputs for a call recalibration machine learning model. Additionally, the series of operations 1300 may include utilizing the call recalibration machine learning model to generate first genotype probabilities and second genotype probabilities further based on the diploid homozygous reference genotype calls or the diploid homozygous alternative genotype calls.
[0231] In certain embodiments, the series of operations 1300 includes downsampling diploid sequencing metrics to simulate haploid sequencing metrics corresponding to haploid nucleotide sequences. Downsampling diploid sequencing metrics to simulate haploid sequencing metrics may include selecting a subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads, and selecting a subset of genomic coordinates that indicate homozygous reference genotypes or homozygous alternative genotypes based on nucleotide base calls of the subset of diploid nucleotide reads, as indicated by a call generation model or as indicated by a ground truth base call dataset (e.g., a curated truth set such as PrecisionFDA v4.2.1).
[0232] 14, the series of operations 1400 includes an operation 1402 of determining one or more nucleotide base calls indicative of a homozygous reference genotype. In particular, operation 1402 may include determining, for one or more nucleotide reads, one or more nucleotide base calls indicative of a homozygous reference genotype at genomic coordinates of a sample nucleotide sequence.
[0233] The series of operations 1400 may include an operation 1404 of determining sequencing metrics for one or more nucleotide base calls. In particular, operation 1404 may include determining sequencing metrics for one or more nucleotide base calls corresponding to a genomic coordinate. For example, operation 1404 may include determining one or more of read-based sequencing metrics, external source sequencing metrics, or call model generation sequencing metrics for a genomic coordinate that is indicated as having a homozygous reference genotype.
[0234] As shown, the series of operations 1400 may include an operation 1406 of generating one or more variant call classifications. In particular, operation 1406 may include generating one or more variant call classifications that indicate an accuracy of identifying a variant at a genomic coordinate utilizing a call recalibration machine learning model and based on sequencing metrics from the one or more nucleotide base calls.
[0235] 14, the series of operations 1400 may include an operation 1408 of determining a variant call from the one or more variant call classifications. In some cases, operation 1408 may include determining a variant call for the genomic coordinate based on the one or more variant call classifications. For example, operation 1408 may include receiving an indication of a homozygous reference genotype at the genomic coordinate from a call generation model and determining a variant call for the genomic coordinate by correcting the homozygous reference genotype to a different genotype based on the one or more variant call classifications.
[0236] In some embodiments, the series of operations 1400 includes identifying a previous homozygous reference genotype call from the call generation model for the sample at the genomic coordinate. Additionally, the series of operations 1400 includes identifying a ground truth base call for the sample at the genomic coordinate and modifying the call recalibration machine learning model based on a comparison of the variant call for the genomic coordinate to the ground truth base call for the genomic coordinate. The series of operations 1400 may include updating one or more of a call quality field, a genotype field, or a genotype quality field corresponding to the variant call file based on the one or more variant call classifications.
[0237] In one particular implementation, the series of operations 1400 includes operations for determining, for the genomic coordinate, one of a homozygous alternative genotype based on a determination that a true positive classification (e.g., a homozygous alternative classification) has the highest probability among one or more variant call classifications, a heterozygous genotype based on a determination that a genotype error classification (e.g., a heterozygous genotype classification) has the highest probability among one or more variant call classifications, or a homozygous reference genotype based on a determination that neither the true positive classification nor the genotype error classification has the highest probability among one or more variant call classifications.
[0238] The methods described herein can be used in conjunction with various nucleic acid sequencing techniques. Particularly applicable techniques are those in which the nucleic acids are attached to fixed positions within an array such that their relative positions do not change, and the array is repeatedly imaged. For example, embodiments in which images are obtained in different color channels that correspond to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of the target nucleic acid can be an automated process. A preferred embodiment includes sequencing-by-synthesis ("SBS") techniques.
[0239] SBS technology generally involves the enzymatic extension of nascent nucleic acid strand by repeated addition of nucleotide to template strand.In the conventional method of SBS, a single nucleotide monomer can be provided to target nucleic acid in the presence of polymerase in each delivery.However, in the method described herein, multiple types of nucleotide monomers can be provided to target nucleic acid in the presence of polymerase during delivery.
[0240] SBS can utilize nucleotide monomers with terminator moieties or nucleotide monomers that lack any terminator moiety. Methods that utilize nucleotide monomers that lack terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in more detail below. In methods that use nucleotide monomers that do not contain terminators, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques that utilize nucleotide monomers with terminator moieties, the terminator can be effectively irreversible under the sequencing conditions used, as in the case of conventional Sanger sequencing that utilizes dideoxynucleotides, or the terminator can be reversible, as in the case of the sequencing method developed by Solexa (now Illumina, Inc.).
[0241] SBS techniques can use nucleotide monomers with or without a label moiety. Thus, incorporation events can be detected based on the properties of the label, such as the fluorescence of the label, the properties of the nucleotide monomer, such as the molecular weight or charge, the by-products of nucleotide incorporation, such as the release of pyrophosphate, and the like. In embodiments in which two or more different nucleotides are present in the sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be distinguishable under the detection technique used. For example, the different nucleotides present in the sequencing reagent can have different labels, which can be distinguished using appropriate optical systems, as exemplified by the sequencing method developed by Solexa (now Illumina, Inc.).
[0242] A preferred embodiment includes the technique of pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281(5375),363, U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320, the disclosures of which are incorporated herein by reference in their entirety). In pyrosequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of ATP generated is detected via luciferase-generated photons. The nucleic acid to be sequenced can be bound to features in an array, and the array can be imaged to capture chemiluminescent signals generated by incorporation of nucleotides into the features of the array. Images can be obtained after treatment of the array with a particular nucleotide type (e.g., A, T, C, or G). Images obtained after addition of each nucleotide type differ with respect to which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative position of each feature remains unchanged in the image. The images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after treating the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.
[0243] In another exemplary type of SBS, cycle sequencing is accomplished by stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, for example as described in WO 04 / 018497 and U.S. Pat. No. 7,057,026, the disclosures of which are incorporated by reference. This approach has been commercialized by Solexa (now Illumina Inc.) and is also described in WO 91 / 06678 and WO 07 / 123,744, each of which is incorporated by reference herein. The availability of fluorescently labeled terminators, both of which can be reversed and from which the fluorescent labels are cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.
[0244] Preferably, in reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be taken after incorporation of the label into the arrayed nucleic acid features. In certain embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, each nucleotide type having a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, and images of the array can be obtained during each addition step. In such embodiments, each image shows nucleic acid features that incorporate a particular type of nucleotide. Different features are present or absent in different images, since the sequence content of each feature is different. However, the relative positions of the features remain unchanged within the images. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the image taking step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removal of the label after detection in a particular cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described below.
[0245] In certain embodiments, some or all of the nucleotide monomers can include reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore can include a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005), which is incorporated herein by reference). Other approaches separate the terminator chemistry from the cleavage of the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al. describe the development of reversible terminators that use a small amount of 3' allyl group to block extension, but can be easily deblocked by brief treatment with a palladium catalyst. The fluorophore was attached to the group via a photocleavable linker that can be easily cleaved by 30 seconds of exposure to long wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as a cleavable linker. Another approach to reversible termination is the use of a natural terminus followed by placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluor, effectively reversing the terminus. Examples of modified nucleotides are also described in U.S. Pat. Nos. 7,427,673 and 7,057,026, the disclosures of which are incorporated herein by reference in their entireties.
[0246] Additional exemplary SBS systems and methods that may be utilized with the methods and systems described herein are described in U.S. Patent Application Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, WO 06 / 064199, WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305, and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorporated herein by reference in their entireties.
[0247] Some embodiments may utilize detection of four different nucleotides using fewer than four different labels. For example, SBS may be performed using the methods and systems described in incorporated document US Patent Application Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types may be detected at the same wavelength but may be distinguished based on differences in intensity for one member of the pair, or based on a change to one member of the pair (e.g., via making a chemical modification, photochemical modification, or physical modification) that results in the appearance or disappearance of a distinct signal compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types may be detected under certain conditions, while the fourth nucleotide type may have no detectable label under those conditions or may be minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid may be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid may be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while the other nucleotide type is detected in no more than one of the channels. The three exemplary configurations above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and second channels (e.g., dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelengths), and a fourth nucleotide type that is not detected in any channel or that is minimally devoid of a label (e.g., unlabeled dGTP).
[0248] Furthermore, as described in incorporated document U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called one-dye sequencing methods, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.
[0249] Some embodiments may utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides. The oligonucleotides typically have different labels that correlate with the identity of a particular nucleotide in the sequence to which the oligonucleotide hybridizes. As with other SBS methods, images can be obtained after treating an array of nucleic acid sequences with labeled sequencing reagents. Each image shows nucleic acid features that incorporate a particular type of label. Because the sequence content of each feature is different, different features are present or absent in different images, but the relative positions of the features remain unchanged within the images. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that may be utilized with the methods and systems described herein are described in U.S. Pat. No. 6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.
[0250] Some embodiments may utilize nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis." Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and JA Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope." Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through the nanopore. The nanopore may be a synthetic pore or a biological membrane protein, such as α-hemolysin. As the target nucleic acid passes through the nanopore, each base pair can be identified by measuring the fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties).Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein. In particular, the data can be processed as images according to the exemplary processing of optical and other images described herein.
[0251] Some embodiments may utilize methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation may be detected via fluorescence resonance energy transfer (FRET) interactions between fluorophore-containing polymerases and γ-phosphate-labeled nucleotides, for example, as described in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 211,414, each of which is incorporated herein by reference, or nucleotide incorporation may be detected using zero-mode waveguides, for example, as described in U.S. Pat. No. 7,315,019, each of which is incorporated herein by reference, and fluorescent nucleotide analogs and engineered polymerases, for example, as described in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference. Illumination can be restricted to a zeptoliter-scale volume around the surface-tethered polymerase so that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, MJ et al. "Zero-mode waveguides for single-molecule analysis at high concentrations." Science, 299, 682-686 (2003); Lundquist, PM et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nano structures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties).Images resulting from such methods can be stored, processed, and analyzed as described herein.
[0252] Some SBS embodiments include detection of protons released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons may use electrical detectors and related technology available from Ion Torrent (Guilford, CT, a subsidiary of Life Technologies), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082(A1), 2009 / 0127589(A1), 2010 / 0137143(A1), or 2010 / 0282617(A1), each of which is incorporated herein by reference. The methods described herein for amplifying target nucleic acids using kinetic exclusion can be easily adapted to substrates used to detect protons. More specifically, the methods described herein can be used to generate clonal populations of amplicons used to detect protons.
[0253] The SBS method described above can be advantageously performed in a multiplex format, such that multiple different target nucleic acids are manipulated simultaneously. In certain embodiments, the different target nucleic acids can be processed in a common reaction vessel or on the surface of a particular substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and detection of incorporation events in a multiplexed manner. In embodiments using surface-bound target nucleic acids, the target nucleic acids can be in an array format. In an array format, the target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent binding, binding to beads or other particles, or binding to a polymerase or other molecule bound to the surface. The array can include a single copy of the target nucleic acid at each site (also referred to as a feature), or multiple copies with the same sequence can be present at each site or feature. The multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, which are described in more detail below.
[0254] The methods described herein can use arrays having any of a variety of densities of features, including, for example, at least about 10 features / cm2, 100 features / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 features / cm2, 100,000 features / cm2, 1,000,000 features / cm2, 5,000,000 features / cm2, or more.
[0255] An advantage of the methods described herein is that they provide rapid and efficient detection of multiple target nucleic acids in parallel. Thus, the present disclosure provides an integrated system that can prepare and detect nucleic acids using techniques known in the art, such as those exemplified above. Thus, the integrated system of the present disclosure can include fluidic components that can deliver amplification and / or sequencing reagents to one or more immobilized DNA fragments, the system including components such as pumps, valves, reservoirs, fluid lines, etc. A flow cell can be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Patent No. 2010 / 0111768(A1) and U.S. Patent Application No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for the flow cell, one or more of the fluidic components of the integrated system can be used for amplification and detection methods. Taking the nucleic acid sequencing embodiment as an example, one or more of the fluidic components of the integrated system can be used for delivery of sequencing reagents in the amplification methods described herein and in the sequencing methods as exemplified above. Alternatively, an integrated system may include separate fluidic systems for performing the amplification method and for performing the detection method. Examples of integrated sequencing systems capable of producing amplified nucleic acids and sequencing the nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina Inc., San Diego, Calif.) and the devices described in U.S. Patent Application No. 13 / 273,666, which is incorporated herein by reference.
[0256] The sequencing system described above sequences the nucleic acid polymers present in the sample received by the sequencing device. As defined herein, "sample" and its derivatives are used in the broadest sense and include any sample, culture, etc. suspected of containing a target. In some embodiments, the sample includes DNA, RNA, PNA, LNA, chimeric or hybrid forms of nucleic acid. A sample can include any biological, clinical, surgical, agricultural, air or water sample that contains one or more nucleic acids. The term also includes any isolated nucleic acid sample, such as genomic DNA, fresh frozen or formalin-fixed paraffin-embedded nucleic acid samples. It is also envisioned that the sample can be derived from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples from a single individual such as a tumor sample and a normal tissue sample (matched), or a sample from a single source containing two different forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample containing plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acid obtained from a newborn, for example, as typically used for newborn screening.
[0257] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, the low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, and other clinical or laboratory obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some embodiments, the sample can include nucleic acid molecules obtained from animals, such as human or mammalian sources. In another embodiment, the sample can include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecule can be an archived or extinct sample or species.
[0258] Additionally, the methods and compositions disclosed herein may be useful for amplifying nucleic acid samples having low quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, the forensic sample may include nucleic acid obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or may include forensic samples obtained by a law enforcement agency, one or more military services, or members thereof. The nucleic acid sample may be crude DNA, including purified samples or lysates, for example, from buccal swabs, paper, cloth, or other substrates that may be impregnated with saliva, blood, or other bodily fluids. Thus, in some embodiments, the nucleic acid sample may include small amounts of DNA, such as genomic DNA, or fragmented portions of DNA. In some embodiments, the target sequence may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from hair, skin, tissue samples, autopsies, or remains of victims. In some embodiments, the nucleic acid including one or more target sequences may be obtained from deceased animals or humans. In some embodiments, the target sequence may comprise nucleic acid obtained from non-human DNA, such as microbial, plant or entomological DNA. In some embodiments, the target sequence or the amplified target sequence is for human identification. In some embodiments, the present disclosure generally relates to a method for identifying characteristics of a forensic sample. In some embodiments, the present disclosure generally relates to a human identification method using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodiment, a forensic sample or human identification sample comprising at least one target sequence can be amplified using any one or more of the target specific primers disclosed herein or using the primer criteria outlined herein.
[0259] The components of the call recalibration system 106 may include software, hardware, or both. For example, the components of the call recalibration system 106 may include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., client device 108). When executed by one or more processors, the computer-executable instructions of the call recalibration system 106 may cause the computing device to perform the call recalibration methods described herein. Alternatively, the components of the call recalibration system 106 may include hardware, such as a dedicated processing device for performing a particular function or group of functions. Additionally or alternatively, the components of the call recalibration system 106 may include a combination of computer-executable instructions and hardware.
[0260] Further, the components of the call recalibration system 106 that perform the functions described herein with respect to the call recalibration system 106 may be implemented, for example, as part of a stand-alone application, as a module of an application, as a plug-in of an application, as a library function that can be called by other applications, and / or as a cloud computing model. Thus, the components of the call recalibration system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally or alternatively, the components of the call recalibration system 106 may be implemented in any application that provides sequencing services, including, but not limited to, Illumina BaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumina", "BaseSpace", "DRAGEN", and "TruSight" are registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0261] Embodiments of the present disclosure may include or utilize special purpose or general purpose computers including, for example, computer hardware such as one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer readable media for carrying or storing computer executable instructions and / or data structures. In particular, one or more of the processes described herein may be embodied in a non-transitory computer readable medium and implemented at least in part as instructions executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer readable medium (e.g., a memory, etc.) and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0262] A computer-readable medium may be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present disclosure may include at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage medium (device) and transmission media.
[0263] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid state drives (SSD) (e.g., based on RAM), flash memory, phase change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0264] A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly recognizes the connection as a transmission medium. A transmission medium may include a network and / or data links that may be used to carry desired program code means in the form of computer-executable instructions or data structures and that may be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
[0265] Furthermore, upon reaching various computer system components, program code means in the form of computer executable instructions or data structures may be automatically transferred from transmission media to non-transitory computer readable storage media (devices) (or vice versa). For example, computer executable instructions or data structures received over a network or data link may be buffered in a RAM in a network interface module (e.g., a NIC) and then eventually transferred to the computer system RAM and / or to less volatile computer storage media (devices) in the computer system. It should therefore be understood that non-transitory computer readable storage media (devices) may be included in computer system components that also (or even primarily) utilize transmission media.
[0266] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to transform the general-purpose computer into a special-purpose computer that implements elements of the present disclosure. Computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological operations, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or operations described above. Rather, the described features and operations are disclosed as example forms of implementing the claims.
[0267] Those skilled in the art will appreciate that the present disclosure may be implemented in a networked computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, cell phones, PDAs, tablets, pagers, routers, switches, etc. The present disclosure may also be implemented in a distributed system environment where both local and remote computer systems perform tasks that are linked through a network (either by hardwired data links, wireless data links, or a combination of hardwired and wireless data links). In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0268] Embodiments of the present disclosure may also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing may be used in the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources may be quickly configured through virtualization, exposed with low management effort or service provider interaction, and then scaled accordingly.
[0269] The cloud computing model may consist of various characteristics such as, for example, on-demand self-service, wide area network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model may also expose various service models such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). The cloud computing model may also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and claims, a "cloud computing environment" is an environment in which cloud computing is employed.
[0270] FIG. 15 illustrates a block diagram of a computing device 1500 that may be configured to perform one or more of the processes described above. It will be understood that one or more computing devices, such as the computing device 1500, may implement the call recalibration system 106 and the sequencing system 104. As illustrated by FIG. 15, the computing device 1500 may include a processor 1502, a memory 1504, a storage device 1506, an I / O interface 1508, and a communication interface 1510, which may be communicatively coupled by a communication infrastructure 1512. In certain embodiments, the computing device 1500 may include fewer or more components than those illustrated in FIG. 15. The following paragraphs describe in more detail the components of the computing device 1500 illustrated in FIG. 15.
[0271] In one or more embodiments, the processor 1502 includes hardware for executing instructions, such as those that make up a computer program. By way of example and not limitation, to execute instructions for dynamically modifying a workflow, the processor 1502 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 1504, or storage device 1506, decode them, and execute them. The memory 1504 may be a volatile or non-volatile memory used to store data, metadata, and programs for execution by the processor. The storage device 1506 includes a storage device, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0272] The I / O interface 1508 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from the computing device 1500. The I / O interface 1508 may include a mouse, a keypad or keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of such I / O interfaces. The I / O interface 1508 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In a particular embodiment, the I / O interface 1508 is configured to provide graphical data to a display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be useful in a particular implementation.
[0273] Communications interface 1510 may include hardware, software, or both. In any case, communications interface 1510 may provide one or more interfaces for communications (e.g., packet-based communications, etc.) between computing device 1500 and one or more other computing devices or networks. By way of example and not limitation, communications interface 1510 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as WI-FI.
[0274] Further, the communication interface 1510 can facilitate communication with various types of wired or wireless networks. The communication interface 1510 can also facilitate communication using various communication protocols. The communication infrastructure 1512 can also include hardware, software, or both that couple the components of the computing device 1500 to one another. For example, the communication interface 1510 can enable multiple computing devices connected by a particular infrastructure to communicate with one another to perform one or more aspects of the processes described herein using one or more networks and / or protocols. To illustrate, a sequencing process can enable multiple devices (e.g., a client device, a sequencing device, and a server device) to exchange information such as sequencing data and error notifications.
[0275] In the foregoing specification, the present disclosure has been described with reference to certain exemplary embodiments thereof. Various embodiments and aspects of the present disclosure are described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The above description and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0276] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects as illustrative only and not restrictive. For example, the methods described herein may be implemented with fewer or more steps / actions, or the steps / actions may be performed in a different order. Further, the steps / actions described herein may be repeated or performed in parallel with each other, or with different occurrences of the same or similar operations. The scope of the present application is therefore indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope. [Explanation of symbols]
[0277] 100 System Environment (or "Environment") 102 Server device 104 Sequencing System 106 Cole Recalibration System 108 User client device 110 Sequencing Applications 112 Network 114 Sequencing Instrument 116 Databases 302 Multi-allelic Genomic Coordinates 304 Sequencing Metrics 306 Call Recalibration Machine Learning Model 308 Variant Call Classification 310 Reference Probability 312 Genotype Probability 314 Variant Probability 316 VCF Field 318 Base call quality (QUAL) field 320 Genotype quality (GQ) field 322 Genotype Likelihood 324 variant call file 402 Haploid nucleotide sequence 404 Sequencing Metrics 406 Call Recalibration Machine Learning Model 408 First Genotype Probability 410 Second Genotype Probability 412 variant call file 414 final nucleotide base call 416 Call Generation Model 418 Downsampling 420 diploid nucleotide reads 422 Diploid Sequencing Metrics 424 Haploid Sequencing Metrics 426 classifier layer 428 Softmax Model 430 variant call files 432 final nucleotide base call 502 Call Generation Model 504 Sample Nucleotide Sequence 506 Sequencing Metrics 508 Sequencing Metrics 510 Sequencing Metrics 512 Sequencing Metrics 514 Call Recalibration Machine Learning Model 516 Variant Call Classification 518 variant call file 520 Variant Calling 602 nucleotide reads 604 Mapping 606 Sequencing Metrics 608 Sequence Data 608 Sequence Data 610 Call Generation Model 612 Sequencing Metrics 614 Sequencing Information Database 616 Sequencing Metrics 702 Sequencing Information Database 704 Reference Sequence 706 Sequence Data 708 Alignment Components 710 Variant Caller Component 712 Sequencing Metric Extraction 714 Call Recalibration Machine Learning Model 716 Variant Call Classification 718 Data Field Generation 720 variant call files 722 Call Generation Model 802 Database 804 Sample Sequencing Metrics 806 Call Recalibration Machine Learning Model 806 (ii) Call Recalibration Machine Learning Model 808 Predictive variant calling classification 810 variant call file 812 comparison 814 Loss Function 816 ground truth variant call files 816 (ii) Ground Truth Variant Call File 818 Model Fitting 902 Contribution Scale Interface 1500 Computing Devices 1502 Processor 1504 Memory 1506 Storage device 1508 I / O Interface 1510 Communication Interface 1512 Communications Infrastructure
Claims
1. 1. A system comprising: at least one processor; A non-transitory computer-readable medium that, when executed by the at least one processor, provides the system with: determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to multi-allelic genomic coordinates of the sample nucleotide sequence; utilizing a call recalibration machine learning model to process the sequencing metrics; a reference probability of a homozygous reference genotype at said multi-allelic genomic coordinate with respect to a reference genome; different genotypic probabilities of genotypic errors at the multi-allelic genomic coordinates; and generating a set of variant call classifications, the variant call classifications comprising variant call probabilities for variant called genotypes that are accurate at the multi-allelic genomic coordinates with respect to the reference genome; determining a final nucleotide base call for the multi-allelic genomic coordinate based on the set of variant call classifications.
2. 2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the final nucleotide base calls by modifying one or more initial nucleotide base calls generated by a call generation model for a variant call file based on the set of variant call classifications.
3. When executed by the at least one processor, the system: generating updated genotype likelihoods for candidate nucleotide base calls of alleles at the multi-allelic genomic coordinates; and generating a variant call file that includes the updated genotype likelihoods.
4. The sequencing metric of the nucleotide base call is read-based sequencing metrics comprising sequencing metrics derived from nucleotide reads of the sample nucleotide sequence; externally sourced sequencing metrics, including sequencing metrics stored in one or more databases external to the system; call model-generated sequencing metrics, including sequencing metrics generated by the call generation model; The system of claim 1 , comprising:
5. 2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the reference probability by determining the probability that a genotype at the multi-allelic genomic coordinate is a homozygous genotype with respect to the reference genome.
6. 2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the different genotype probabilities by determining the probability that a predicted variant call genotype initially determined with respect to the reference genome by a call generation model for the multi-allelic genomic coordinate is an incorrect genotype or represents an incorrect allele.
7. 10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the correct variant call probability by determining a probability that a predicted genotype for the reference genome, as initially determined by a call generation model for the multi-allelic genomic coordinate, is a correct genotype.
8. 1. A computer-implemented method comprising: determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to genomic coordinates of a haploid nucleotide sequence from the sample; generating a set of variant call classifications utilizing a call recalibration machine learning model for processing the sequencing metrics, the set of variant call classifications comprising: a first genotype probability of a haploid reference genotype at the genomic coordinate with respect to a reference genome; and a second genotype probability for a haploid alternative genotype at the genome coordinate with respect to the reference genome; generating a set of variant call classifications; determining a final nucleotide base call indicative of a haploid genotype for the genomic coordinate based on the first genotype probability and the second genotype probability.
9. generating the first genotype probability includes utilizing a layer of the call-recalibration machine learning model to modify a homozygous reference probability for a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability for the haploid reference genotype at the genomic coordinate; 9. The computer-implemented method of claim 8, wherein generating the second genotype probability comprises utilizing the layer of the call-recalibration machine learning model to modify homozygous substitution probabilities for homozygous alternative genotypes at the genomic coordinates to generate haploid substitution probabilities for the haploid alternative genotypes at the genomic coordinates.
10. generating the first genotype probability and the second genotype probability; utilizing one or more layers of the call-recalibration machine learning model to generate, for the genomic coordinate, a first confidence score corresponding to a first genotype, a second confidence score corresponding to a second genotype, and a third confidence score corresponding to a third genotype; removing the second confidence score corresponding to the second genotype; and normalizing the first and third confidence scores utilizing a softmax model to generate the first and second genotype probabilities.
11. determining the final nucleotide base call indicative of the haploid genotype for the genomic coordinates; the haploid alternate genotype, a corrected base call quality metric, a corrected genotype metric, and a corrected genotype quality metric for the genomic coordinate based on determining that the second genotype probability exceeds the first genotype probability; or 9. The computer-implemented method of claim 8, comprising determining one of the haploid reference genotype, a corrected base call quality metric, and a corrected genotype quality metric for the genomic coordinate based on determining that the first genotype probability exceeds the second genotype probability.
12. converting the haploid reference genotype calls generated by the call generation model to diploid homozygous reference genotype calls as input for the call recalibration machine learning model; or converting haploid alternative genotype calls generated by the call generation model to diploid homozygous alternative genotype calls as input for the call recalibration machine learning model; 9. The computer-implemented method of claim 8, further comprising utilizing the call recalibration machine learning model to generate the first genotype probability and the second genotype probability further based on the diploid homozygous reference genotype call or the diploid homozygous alternative genotype call.
13. Downsampling diploid sequencing metrics selecting a subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads; 10. The computer-implemented method of claim 8, further comprising: simulating haploid sequencing metrics corresponding to the haploid nucleotide sequence by: selecting, based on the nucleotide base calls of the subset of diploid nucleotide reads, a subset of genomic coordinates that represent a homozygous reference genotype or a homozygous alternative genotype, as indicated by a call generation model or as indicated by a ground truth base call dataset.
14. The computer-implemented method of claim 13, further comprising selecting the diploid sequencing metrics corresponding to the subset of genomic coordinates representing a homozygous reference genotype or a homozygous alternative genotype to simulate the haploid sequencing metrics corresponding to the haploid nucleotide sequence.
15. A non-transitory computer-readable medium that, when executed by at least one processor, causes a computing device to: determining one or more initial nucleotide base calls indicative of a homozygous reference genotype at a genomic coordinate of the sample nucleotide sequence using a call generation model for the one or more nucleotide reads; determining sequencing metrics for the one or more initial nucleotide base calls corresponding to the genomic coordinates; and utilizing a call recalibration machine learning model to process sequencing metrics from the one or more initial nucleotide base calls to generate a set of variant call classifications indicative of an accuracy for identifying variants at the genomic coordinates, wherein the set of variant call classifications comprises: a homozygous reference classification indicating the probability of a homozygous reference genotype with respect to a reference genome at said genomic coordinates; a heterozygous genotype classification indicating the probability of a heterozygous genotype at said genomic coordinates with respect to a reference genome; and a homozygous alternative classification indicating the probability of a homozygous alternative genotype at the genomic coordinate with respect to the reference genome. and and determining a variant call for the genomic coordinate based on the set of variant call classifications by modifying at least one initial nucleotide base call generated by the call generation model.
16. When executed by the at least one processor, the computing device: receiving an indication of the homozygous reference genotype at the genomic coordinate from the call generation model; and determining the variant call for the genomic coordinate by correcting the homozygous reference genotype to a different genotype based on the set of variant call classifications.
17. 16. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the sequencing metrics for the genomic coordinates indicated as having a homozygous reference genotype relative to the reference genome by determining one or more of read-based sequencing metrics, external source sequencing metrics, or call model generation sequencing metrics.
18. When executed by the at least one processor, the computing device: identifying a previous homozygous reference genotype call with respect to the reference genome from the call generation model for the sample nucleotide sequence at the genome coordinates; identifying a ground truth base call for the sample nucleotide sequence at the genomic coordinates; and modifying the call recalibration machine learning model based on a comparison of the variant call for the genomic coordinate and the ground truth base call for the genomic coordinate.
19. When executed by the at least one processor, the method causes the computing device to perform the following steps for the genomic coordinates: the homozygous alternative genotype based on determining that the homozygous alternative classification has the highest probability among the set of variant call classifications; the heterozygous genotype based on determining that the heterozygous genotype classification has the highest probability among the set of variant call classifications; or 16. The non-transitory computer-readable medium of claim 15, further comprising instructions to determine one of the homozygous reference genotypes based on determining that neither the homozygous alternative classification nor the heterozygous genotype classification has the highest probability among the set of variant call classifications.
20. 16. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to update one or more of a call quality field, a genotype field, or a genotype quality field corresponding to a variant call file based on the set of variant call classifications.