Machine learning model for recalibrating nucleotide base calls corresponding to target variants

HK40135051APending Publication Date: 2026-07-17ILLUMINA INC

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2026-05-29
Publication Date
2026-07-17

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This disclosure describes methods, non-transitory computer readable media, and systems that can utilize a machine learning model to recalibrate nucleotide base calls (e.g., variant calls) of a call generation model. For instance, the disclosed systems can train and utilize a call recalibration machine learning model to generate a set of predicted variant call classifications based on sequencing metrics associated with a sample nucleotide sequence. Leveraging the set of variant call classifications, the disclosed systems can further update or modify nucleotide base calls (e.g., variant calls) corresponding to genomic coordinates, such as multiallelic genomic coordinates, haploid genomic coordinates, and genomic coordinates indicated (by the call generation model) to exhibit homozygous reference genotypes.
Need to check novelty before this filing date? Find Prior Art

Description

(19) *EP004693301A2* (11) EP 4 693 301 A2 (12) EUROPEAN PATENT APPLICATION (43) Date of publication: 11.02.2026 Bulletin 2026 / 07 (21) Application number: 25224089.0 (22) Date of filing: 23.12.2022 (51) International Patent Classification (IPC): G16B 40 / 20 (2019.01) (52) Cooperative Patent Classification (CPC): G16B 20 / 20; G16B 30 / 20; G16B 40 / 10; G16B 40 / 20 (84) Designated Contracting States: AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR (30) Priority: 28.12.2021 US 202117563934 (62) Document number(s) of the earlier application(s) in accordance with Art. 76 EPC: 22857067.7 / 4 457 822 (71) Applicant: ILLUMINA, INC. San Diego, CA 92122 (US) (72) Inventor: PARNABY, Gavin San Diego, 92122 (US) (74) Representative: Robinson, David Edward Ashdown Marks & Clerk LLP Wytham Court, 11 West Way Oxford OX2 0JB (GB) Remarks: This application was filed on 16‑12‑2025 as a divisional application to the application mentioned under INID code 62. (54) MACHINE LEARNING MODEL FOR RECALIBRATING NUCLEOTIDE BASE CALLS CORRESPONDING TO TARGET VARIANTS (57) This disclosure describes methods, non-transi- tory computer readable media, and systems that can utilize amachine learningmodel to recalibrate nucleotide base calls (e.g., variant calls) of a call generation model. For instance, thedisclosedsystemscan train andutilizea call recalibration machine learning model to generate a set of predicted variant call classifications based on sequencingmetrics associated with a sample nucleotide sequence. Leveraging the set of variant call classifica- tions, the disclosed systems can further update ormodify nucleotide base calls (e.g., variant calls) corresponding to genomic coordinates, such as multiallelic genomic coordinates, haploid genomic coordinates, and genomic coordinates indicated (by the call generation model) to exhibit homozygous reference genotypes. EP 4 69 3 30 1 A 2 Processed by Luminess, 75001 PARIS (FR) 2 1 EP 4 693 301 A2 2 Description CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of, and prior- ity to, U.S. Application No. 17 / 563,934, entitled "MA- CHINE-LEARNING MODEL FOR RECALIBRATING NUCLEOTIDE BASE CALLS CORRESPONDING TO TARGET VARIANTS," filed December 28, 2021, the contents of which are hereby incorporated by reference in their entirety. BACKGROUND

[0002] In recent years, biotechnology firms and re- search institutions have improved hardware and soft- ware for sequencing nucleotides and determining nu- cleotide base calls (e.g., variant calls) for genomic sam- ples. For instance, some existing nucleotide base se- quencing platforms determine individual nucleotide bases within sequences by using conventional Sanger sequencing or by using sequencing-by-synthesis (SBS) methods. When using SBS, existing platforms canmoni- tor many thousands of nucleic acid polymers being synthesized in parallel to predict nucleotide base calls from a larger base call dataset. For instance, a camera in many SBS platforms captures images of irradiated fluor- escent tags incorporated into oligonucleotides for deter- mining the nucleotide base calls. After capturing such images, existing SBS platforms send base call data (or image data) to a computing device to apply sequencing data analysis software that determines anucleotide base sequence for a nucleic acid polymer. In certain cases, some prior systems further utilize a variant caller to identify variants, such as single nucleotide polymorph- isms (SNPs), insertions or deletions (indels), or other variants within a sample’s nucleic acid sequence.

[0003] Despite these recent advances in sequencing and variant calling, existing nucleotide base sequencing platforms and sequencing data analysis software (to- gether and hereinafter, existing sequencing systems) often include variant callers that inaccurately determine nucleotide base calls (and / or corresponding variant calls). For example, existing sequencing systems either inaccurately determine-or are incapable of determining- nucleotide base calls for multiallelic genomic coordi- nates. Indeed, for regionsof a nucleotide sequence, such as multiallelic regions, that are more challenging than biallelic regions, some existing systems struggle to (or cannot) accurately determine genotypes when alleles cover or correspond to a given genomic coordinate. For instance, some machine learning based sequencing systems struggle to determine genotypes for multiallelic coordinatesbecause trainingdata is largelybiallelic data. Thus, in the case of a pileup or a large insertion, existing sequencing systems often fail to correctly determine nucleotide base calls and / or a genotype from multiple possible alleles at the given genomic coordinate.

[0004] In addition, existing sequencing systems inac- curately determine nucleotide base calls (e.g., variant calls) for haploid genomic coordinates within a genomic sample or other nucleotide sequence. For instance, many existing sequencing systems inaccurately deter- mine nucleotide base calls within sex chromosomes, often due to the sparsity or complete lack of good haploid training data. Specifically, existing sequencing systems often learn parameters for determining nucleotide base calls exclusively from unmodified diploid data (e.g., Pre- cisionFDA truth data from the PrecisionFDATruth Chal- lenge, described at https: / / precision.fda.gov / challenges / truth) and lack models or training to identify nucleotide bases or genotypes for coordinates other than diploid coordinates. Consequently, many of these existing se- quencing systems cannot accurately determine nucleo- tide base calls or variant calls for haploid genomic co- ordinates.

[0005] Further, in some circumstances, existing se- quencing systems apply a variant caller that inaccurately identifies excessive numbers of false negative variant calls. For instance, existing sequencing systems some- times determine a genomic coordinate exhibits a homo- zygous reference genotype (and therefore not include a variant) when, in fact, the coordinate includes a variant. Indeed, existing variant callers achieve a certain level of accuracy but, due to their limitations, still leave room for improvement in recovering false negative variant calls. To illustrate the impact of such inaccuracy, a variant call identifying a particular single nucleotide polymorphism (SNP) in the hemoglobin beta (HBB) gene can have significant implications. When a variant caller identifies an SNP at rs344 on chromosome 11, for instance, the variant caller can either correctly identify the genetic cause of sickle cell anemia or miss the cause of the disease. As a further example, a variant call that correctly or incorrectly identifies the deletion of one ormore copies of hemoglobin subunit alpha 1 (HbA1) or hemoglobin subunit alpha 2 (HbA2) genes can result in either cor- rectly identifying a genetic cause of an inherited blood disorder or miss the gene deletion entirely.

[0006] As a contributing factor to the aforementioned inaccuracies, many existing sequencing systems lever- age only limited sets of data in determining nucleotide base calls. For instance, existing sequencing systems frequently rely exclusively on information extracted di- rectly from nucleotide reads of a sample sequence, such as read depth, mismatch counts, sequence alignment scores, and mapping quality, to determine nucleotide base calls. While sequence information from nucleotide reads can provide valuable insight for determining nu- cleotide base calls, existing sequencing systems that solely rely on these data can underperform when deter- mining nucleotide base calls. Indeed, some existing se- quencing systems that rely on raw sequence data incor- rectly determine SNPs, indels, or other variants in a genomic sample sequence in comparison to more com- plex models. Indeed, existing sequencing systems fre- 5 10 15 20 25 30 35 40 45 50 55 3 3 EP 4 693 301 A2 4 quently identify false negative variants or false positive variants in the Truth Challenges of the U.S. Food and Drug Administration (FDA), and reliable haploid data is often difficult to acquire for testing or training a variant caller.

[0007] In addition to inaccurately determining variant calls, some existing sequencing systems also ineffi- ciently expend computing resources with overly complex models. Specifically, the variant callers of some existing sequencing systems are computationally expensive and slow. Indeed, some existing sequencing systems utilize variant callers with a deep learning architecture or some other neural network architecture that require extensive computational resources (e.g., computing time, proces- singpower, andmemory) to train andapply. For example, some existing sequencing systems utilize deep learning architectures that, even after training, take many hours across multiple computing devices to generate nucleo- tide base calls for a single sample sequence.

[0008] As an added drawback of existing sequencing systems with complex networks, many such systems utilize model architectures that render sequence data uninterpretable. More specifically, some existing deep neural networks transform andmanipulate the sequence data many times over, changing from one vector to an- other across the various layers and neurons, as the basis for generating a variant call. In many cases, the internal data of these deep neural networks is uninterpretable and impossible to utilize in any way outside of the neural network architecture itself. SUMMARY

[0009] This disclosure describes embodiments of methods, non-transitory computer readable media, and systems that can utilize a machine learning model to recalibrate nucleotide base calls (e.g., variant calls) of a call generation model. For example, the disclosed systems can train and utilize a call recalibration machine learning model to generate a set of classification predic- tions (e.g., variant call classifications) to improve nucleo- tide base calls in specific scenarios, such as generating nucleotide base calls for multiallelic coordinates, haploid coordinates, and / or coordinates incorrectly identified by existing sequencing systems as exhibiting homozygous reference genotypes. As disclosed, the disclosed sys- temscan (i) determinesequencingmetrics for aparticular genomic coordinate, such as a multiallelic coordinate, a haploid coordinate, or an incorrectly identified homozy- gous reference coordinate and (ii) utilize a call recalibra- tion machine learning model to generate classification predictions for updating or recalibrating an initial nucleo- tide base call for the genomic coordinate. After recali- brating, the disclosed systems can output the updated or recalibrated nucleotide base call as a final nucleotide base call (e.g., a final variant call) in a variant call file or other base call output file.

[0010] By utilizing a call recalibrationmachine learning model to update sequencing metrics for generating nu- cleotide base calls, the disclosed systems can improve accuracy, efficiency, and speed over existing sequencing systems. As described further below, for instance, the disclosed call recalibration machine learning model de- termines variant calls with better accuracy than conven- tional hidden Markov model (HMM)‑based or probabil- istic-based variant callers and more complex neural net- works (e.g., deepneural network-base variant callers) for variant calling at a multiallelic coordinate, a haploid co- ordinate, or an incorrectly identified homozygous refer- ence coordinate. The disclosed call recalibration ma- chine learning model also determines variant calls at such genomic coordinates with faster computing times than complex neural networks. Additionally, the dis- closed systems can improve interpretability of factors impacting accurate variant calls at such genomic coordi- nates in comparison to complex neural networks by utilizing a call recalibration machine learning model that processes data in an accessible, interpretable format. Indeed, because of the improved interpretability of the disclosed systems, in some embodiments, the disclosed systems can generate and provide a visualization of various contributionmeasures associatedwith individual sequencing metrics to visually depict respective mea- sures of impact that the sequencing metrics have on a resultant nucleotide base call. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The detailed description refers to the drawings briefly described below. FIG. 1 illustrates a block diagram of a sequencing system including a call recalibration system in ac- cordance with one or more embodiments. FIG. 2 illustrates an overview of the call recalibration systemgenerating anucleotide base call utilizing the call recalibration system in accordance with one or more embodiments. FIGS. 3A‑3B illustrate the call recalibration system generating nucleotide base calls for multiallelic genomiccoordinates inaccordancewithoneormore embodiments. FIGS. 4A‑4B illustrate the call recalibration system generatingnucleotidebasecalls for haploid genomic coordinates in accordance with one or more embo- diments. FIG. 5 illustrates the call recalibration system gen- erating variant calls for homozygous reference genomiccoordinates inaccordancewithoneormore embodiments. FIGS. 6A‑6C illustrate the call recalibration system generating or determining sequencing metrics in accordance with one or more embodiments. FIG. 7 illustrates the call recalibration system gen- erating variant call classifications and recalibrating a nucleotide base call utilizing a call recalibration ma- 5 10 15 20 25 30 35 40 45 50 55 4 5 EP 4 693 301 A2 6 chine learningmodel in accordancewith one ormore embodiments. FIG. 8 illustrates an example process for the call recalibration system training a call recalibration ma- chine learningmodel in accordancewith one ormore embodiments. FIG. 9 illustrates an example contribution-measure interface displayed on a client device in accordance with one or more embodiments. FIGS.10A‑10B illustrategraphsand tablesdepicting accuracy improvements associated with the call re- calibration system for diploid coordinates in accor- dance with one or more embodiments. FIGS. 11A‑11B illustrate a graphs and tables depict- ing accuracy improvements associated with the call recalibration system for haploid coordinates in ac- cordance with one or more embodiments. FIG. 12 illustrates a flowchart of a series of acts for generating nucleotide base calls associated with multiallelic genomic coordinates in accordance with one or more embodiments. FIG. 13 illustrates a flowchart of a series of acts for generating nucleotide base calls associated with haploid genomic coordinates in accordance with one or more embodiments. FIG. 14 illustrates a flowchart of a series of acts for generating variant calls associated with homozy- gous reference genomic coordinates in accordance with one or more embodiments. FIG. 15 illustrates a block diagram of an example computing device for implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0012] Thisdisclosuredescribesembodimentsofacall recalibration system that generates and recalibrates nu- cleotide base calls (e.g., variant calls) for a sample nu- cleotide sequence utilizing a call recalibration machine learningmodel. In particular, the call recalibration system can utilize a call recalibration machine learning model to update, recalibrate, or modify an initial nucleotide base call generated by a call generation model. For example, the call recalibration system can recalibrate the initial nucleotide base call to improve its accuracy by utilizing a call recalibration machine learning model to update various call metrics, such as a call quality, a genotype associated with the call, a genotype quality associated with the genotype, Phred-scaled Likelihood (PL), and / or other metrics with corresponding fields. By utilizing the call recalibration machine learning model to update me- trics, the call recalibration system can improve the accu- racy of nucleotide base calls at particular genomic co- ordinates, such as multiallelic coordinates, haploid co- ordinates, and coordinates falsely determined (in an initial call or by an existing sequencing system) to exhibit homozygous reference genotypes.

[0013] As just mentioned, in certain implementations, the call recalibration system improves nucleotide base calls and corresponding variant calls for multiallelic co- ordinates of a sample nucleotide sequence. To facilitate generating multiallelic nucleotide base calls, in some embodiments, the call recalibration system utilizes a call recalibration machine learning model that is specialized and adaptable to generate nucleotide base calls for both biallelic andmultiallelic coordinates.For instance, thecall recalibration system can generate, from sequencingme- trics associated with a multiallelic genomic coordinate, a set of variant call classifications that includes a probabil- ity of ahomozygous referencegenotypeat themultiallelic genomic coordinate (i.e., a reference probability), a prob- ability of a genotype error at the multiallelic genomic coordinate (i.e., a differing genotype probability), and a probability of a correct variant call genotype at the multi- allelic genomic coordinate (i.e., a correct variant prob- ability). The call recalibration system can further deter- mine a final nucleotide base call for the multiallelic geno- mic coordinate from the set of variant call classifications. Additional detail regardinggenerating calls formultiallelic coordinates is provided below with reference to the fig- ures.

[0014] Asmentioned, inoneormoreembodiments, the call recalibration system improves nucleotide base calls and corresponding variant calls for haploid genomic co- ordinates of a sample nucleotide sequence. In particular, the call recalibration systemcanutilizea call recalibration machine learning model adapted to determine haploid genotypes based on diploid data. For instance, the call recalibration system can train a call recalibration ma- chine learning model by modifying diploid data (e.g., diploid sequencing metrics) to simulate haploid data (e.g., haploid sequencing metrics). In addition, the call recalibration system can utilize the trained call recalibra- tionmachine learningmodel togenerate threeoutputs for a given genomic coordinate: (i) a first confidence score for a homozygous reference genotype (0 / 0), (ii) a second confidence score for a heterozygous genotype (0 / 1), and (iii) a third confidence score for a homozygous alternate genotype (1 / 1).

[0015] The call recalibration system can further prune or remove the second confidence score (e.g., the 0 / 1 confidence score) and can utilize a softmax model or layer to normalize across the other two confidence scores and convert the confidence scores to haploid probabilities. Utilizing the softmaxmodel or layer, the call recalibration system can thus determine: (i) from the homozygous reference confidence score (0 / 0), a haploid reference probability (0) and (ii) from the homozygous alternate confidence score (1 / 1), a haploid alternate probability (1). Additional detail regarding generating calls for haploid coordinates is provided below with re- ference to the figures.

[0016] As further mentioned above, the call recalibra- tion system improves nucleotide base calls and corre- sponding variant calls for genomic coordinates of a sam- ple nucleotide sequence that are determined to exhibit 5 10 15 20 25 30 35 40 45 50 55 5 7 EP 4 693 301 A2 8 homozygous reference genotypes. More specifically, the call recalibration system can recover false negative var- iant calls for genomic coordinates that are initially deter- mined as exhibiting homozygous reference genotypes (e.g., as determined by a call generationmodel) when, in fact, the genotypes of these coordinates are not homo- zygous with respect to the reference sequence. As op- posed to existing sequencing systems that filter out data associated with homozygous reference coordinates, the call recalibration system can determine sequencing me- trics for such homozygous reference coordinates and can utilize a call recalibration machine learning model to generate variant call classifications from the sequen- cing metrics. Further, the call recalibration system can generate final nucleotide base calls for the homozygous reference coordinates based on the variant call classifi- cations, changing a variant call that would have indicated a homozygous reference genotype to indicating a differ- ent genotype (and thereby recovering false negative variant calls). Additional detail regarding correcting or updating variant calls for genomic coordinates thatwould have been incorrectly identified as exhibiting homozy- gous reference genotypes is provided below with refer- ence to the figures.

[0017] As mentioned above, in some embodiments, the call recalibration system can more generally utilize a machine learning model to generate variant call classi- fications based on sequencing metrics for nucleotide base calls corresponding to genomic coordinates. To generate such classifications, the call recalibration sys- tem extracts or determines sequencing metrics from a sample nucleotide sequence. For example, the call re- calibration system determines sequencing metrics from nucleotide base calls of nucleotide reads from a sample nucleotide sequence. Indeed, in some cases, the call recalibration system generates or determines a set of initial nucleotide base calls from nucleotide reads cap- tured or determined via fluorescent imaging of a sample nucleotide sequence (e.g., at a particular genomic co- ordinate). From the read-based nucleotide base calls, in some embodiments, the call recalibration system deter- mines or extracts various sequencing metrics (e.g., se- quencing metrics of various types obtained from reads and / or from different components of a call generation model).

[0018] To elaborate, in certain implementations, the call recalibration system determines different types of sequencing metrics associated with different sources. For example, the call recalibration system determines read-based sequencing metrics including metrics de- rived from nucleotide reads of the sample nucleotide sequence. In addition, the call recalibration system de- termines externally sourced sequencing metrics identi- fied from one or more external databases that indicate various nucleotide attributes, mapping challenges, and genomic sequences associated with sequencing biases. Further, the call recalibration system determines call model generated sequencing metrics generated via a variant caller or other call generation model, such as variables internal to the call recalibration system that are not accessible to other systems or parties (e.g., proprietary quality scores, base contexts, read filtering, proprietary hypothesis scores, and other metrics). In- deed, in some cases, the call recalibration system de- termines call model generated sequencingmetrics in the form of variant calling sequencing metrics and mapping- and-alignment sequencing metrics, where each type is extracted by different components of the call generation model.

[0019] As further mentioned, in certain implementa- tions, the call recalibration system generates a set of predicted classifications from the sequencing metrics for modifying or improving a nucleotide base call or variant call data or fields associated with a nucleotide base call. More specifically, the call recalibration system utilizes a call recalibration machine learning model to generate, from the sequencing metrics, a set of three variant call classifications that impact or reflect the accuracy of identifying a variant at a particular genomic coordinate (e.g., a genomic coordinate corresponding to nucleotide base calls of nucleotide reads from a sample nucleotide sequence). Depending on the circumstances, the call recalibration system can utilize the call recalibration ma- chine learning model to, for example, generate different variant call classifications formultiallelic coordinates than for haploid coordinates or would-be-false homozygous reference coordinates.

[0020] For instance, when generating variant call clas- sifications for a multiallelic genomic coordinate, the call recalibration system can utilize the call recalibration ma- chine learning model to generate a set including: (i) a reference probability of a homozygous reference geno- type at the multiallelic genomic coordinate, (ii) a differing genotype probability of a genotype error at themultiallelic genomic coordinate, and (iii) a correct variant probability of a correct variant call genotype at the multiallelic geno- mic coordinate. As another example, for haploid coordi- nates, the call recalibration system can utilize the call recalibrationmachine learningmodel to generate a set of variant call classifications including: (i) a first genotype probability of a first genotype at the genomic coordinate and (ii) a second genotype probability of a second geno- type at the genomic coordinate. Further, for would-be homozygous referencecoordinates, thecall recalibration system can utilize the call recalibrationmachine learning model to generate a set of variant call classifications including: (i) a false positive classification (e.g., a prob- ability that a nucleotide base call is a false positive variant), (ii) a genotype error classification (e.g., a het- erozygous genotype classification indicating a probabil- ity of identifying a correct alt allele but with a genotype error-e.g., 0 / 1 instead of 1 / 1 or 1 / 1 instead of 0 / 1‑ or a probability of incorrectly identifying a genotype of a nu- cleotide base call), and a (iii) true-positive classification (e.g., homozygous alternate classification indicating a probability that a nucleotide base call or a genotype call 5 10 15 20 25 30 35 40 45 50 55 6 9 EP 4 693 301 A2 10 is a true positive variant). In some cases, the variant call classifications accordingly represent intermediate scor- ing metrics associated with a variant caller.

[0021] From the variant call classifications, the call recalibration systemcan furthermodify or updatemetrics for one or more final nucleotide base calls for a genomic coordinate (e.g., final nucleotide base calls that indicates a variant call or a non-variant call). For example, the call recalibration system utilizes the variant call classifica- tions to update data fields within a digital call file (e.g., a variant call format file or other base call output file) that indicates or represents a final nucleotide base call and / or a variant call. Indeed, as mentioned above, in some embodiments, the call recalibration system utilizes a call generation model to generate or determine a final nu- cleotide base call from the sequencing metrics for the genomic coordinate.

[0022] Additionally, the call recalibration system can utilize the variant call classifications to update a nucleo- tide base call and / or a variant call for improved accuracy. In certain implementations, the call recalibration system updates nucleotide base calls for specific genomic co- ordinates, such asmultiallelic genomic coordinates, hap- loid genomic coordinates, and / or would-be falsely iden- tified homozygous reference coordinates (i.e., genomic coordinates that previously or would have been falsely identified by a variant caller to exhibit homozygous re- ference genotypes). Indeed, in some embodiments, the call recalibration system utilizes (i) the call generation model to generate an initial nucleotide base call and (ii) the call recalibration machine learning model to modify data fields corresponding to a variant call file for the nucleotide base call. In some cases, the call recalibration system furthermodifies thenucleotidebasecall basedon one ormore of the data fields and generates a variant call file with the modified nucleotide base call. In certain embodiments, the call recalibration system can generate the variant call classifications utilizing the call recalibra- tion machine learning model while also utilizing the call generation model to generate the nucleotide base call based on the variant call classifications.

[0023] By contrast, in some cases, the call recalibra- tion system determines a final nucleotide base call or a variant call for a genomic coordinate based on both sequencing metrics for a call generation model and var- iant call classifications from thecall recalibrationmachine learning model-without an initial nucleotide base call (e.g., an initial variant call) from thecall generationmodel. For example, thecall generationmodelmaynot output an initial nucleotide base call, but may instead evaluate the genomic coordinate and generate sequencing metrics that the call recalibration machine learning model can then use to generate a variant call in combinationwith the call generation model. In some embodiments, the call generation model may output a final variant call that accounts for the variant call classifications from the call recalibration machine learning model (without generat- ing an initial variant call that is updated). By contrast, in certain cases, the call generation model may initially determine a confidence or quality corresponding to a potential variant call fails to satisfy a threshold for includ- ing in a variant call file but (after accounting for variant call classifications that updates a base call quality metric) determine to include a variant call in the variant call file. As a result of implementing the call recalibrationmachine learning model and the call generation model in this way, the call recalibration system recovers false negative calls, fixes variant genotype errors, and / or removes false positive calls initially made by the call generation model.

[0024] In one or more embodiments, the call recalibra- tion system further determines contribution measures associated with one or more of the sequencing metrics. In particular, the call recalibration system determines measures of impact or influence that each sequencing metric or a subset of sequencing metrics has on a final nucleotide base call. For example, somemetrics may be moreheavilyweighted thanothers in determininga call at one genomic coordinate versus another. Indeed, due to the accessibility and interpretability of the call generation model and the call recalibrationmachine learningmodel, the call recalibration system can access internal sequen- cing metrics used to generate a nucleotide base call and can determine their respective contribution measures in ultimately determining which metrics are causing or driv- ing the recalibration of the final nucleotide base calls (or variant calls). Insomecases, thecall recalibrationsystem further generates and provides a visualization of the contribution measures for display on a client device.

[0025] As suggested above, the call recalibration sys- tem provide several advantages, benefits, and / or im- provements over existing sequencing systems, including variant callers and other sequencing data analysis soft- ware. For instance, the call recalibration system gener- ates more accurate nucleotide base calls and / or variant calls than existing sequencing systems. While some existing sequencing systems are either incapable of generating, or inaccurately generating, nucleotide base calls for multiallelic coordinates, in some embodiments, the call recalibration system generates more accurate calls for multiallelic genomic coordinates. Specifically, the call recalibration system can utilize or adapt a call recalibration machine learning model with parameters trained or tuned to generate a set of variant call classi- fications specific to multiallelic genomic coordinates. From the set of variant call classifications, the call recali- bration system can further generate one or more final nucleotide base calls for a multiallelic genomic coordi- nate to indicate a genotype of the multiallelic coordinate, indicate whether the genotype is a variant with respect to a reference sequence, and / or indicate whether the gen- otype is correct (e.g., a genotypequalitymetric inGQfield indicating a likelihood or probability that a genotype is correct). Similarly, from the set of variant call classifica- tions, the call recalibration system can also improve accuracy of quality fields and other fields, such as PL.

[0026] In some embodiments, the call recalibration 5 10 15 20 25 30 35 40 45 50 55 7 11 EP 4 693 301 A2 12 system generates more accurate nucleotide base calls and / or variant calls for haploid coordinates of a sample nucleotide sequence, as compared to an existing se- quencing system. Unlike some existing sequencing sys- tems that cannot recalibrate nucleotide base calls for haploids, the call recalibration system can utilize a call recalibration machine learning model adaptable to hap- loid regions of a sample nucleotide sequence. In certain cases, the call recalibration system learnsparameters for the call recalibrationmachine learningmodel by adapting diploid data to simulate haploid data. Additionally, the call recalibration system can generate nucleotide base calls for haploid coordinates by pruning, for a particular geno- mic coordinate, a particular machine learning output (e.g., confidence score) of the call recalibration machine learning model not pertinent to haploid calls and by normalizing across the remaining two outputs (e.g., con- fidence scores). By pruning and normalizing outputs compatible with diploid data to outputs compatible with haploid data, the call recalibration system can determine probabilities indicating a haploid reference genotype and a haploid alternate genotype at the coordinate.

[0027] In one or more embodiments, the call recalibra- tion system generates more accurate nucleotide base calls and / or variant calls for (would-be-falsely identified) homozygous reference coordinates of a sample nucleo- tide sequence, as compared to an existing sequencing system.For instance, someexistingsequencingsystems generate an inordinate number of false negative variant calls by incorrectly identifying certain genomic coordi- nates as exhibiting homozygous reference genotypes when, in actuality, their genotypes are not homozygous reference. By contrast, the call recalibration system iden- tifies fewer false negative variant calls (or recovers more false negative variant calls) by determining sequencing metrics for genomic coordinates indicated to exhibit homozygous reference genotypes and utilizing a call recalibrationmachine learningmodel to generate variant call classifications for these coordinates. The call recali- bration system can further generate one or more final nucleotide base calls from the variant call classifications of the homozygous reference coordinates.

[0028] Thecall recalibrationsystem improvesupon the accuracies of existing sequencing systems (e.g., in each of the scenarios described above) by removing large numbers of false positive variant calls and / or recovering large numbers of false negative variant calls utilizing the call recalibration machine learning model. By editing an initial nucleotide base call or generating a final nucleotide basecall basedonvariant call classifications from thecall recalibration machine learning model, the call recalibra- tion system can use unique machine learning outputs to recalibrate base calls with better accuracy than existing variant callers or existing machine learning models. For instance, the call recalibration system utilizes the call recalibrationmachine learningmodel to generate variant call classifications from both internal (e.g., proprietary and model-specific) and external sequencing metrics, which results in recovering variant nucleotide base calls that were previously filtered out and / or removing non- variant nucleotide base calls that were previously not filtered out.

[0029] To accomplish the aforementioned improved accuracies, as indicated, the call recalibration system utilizesan improvedanduniquemachine learningmodel- the call recalibration machine learning model-that is trained to perform new applications. Unlike existing var- iant callers that generate nucleotide base calls from general sequencing data (without any particular empha- sis on one genomic coordinate or another), the call recalibration system utilizes a unique call recalibration machine learning model that generates specific variant call classifications for specific scenarios, such as multi- allelic genomic coordinates, haploid genomic coordi- nates, and false homozygous reference coordinates. In some cases, the call recalibration system utilizes the call recalibrationmachine learningmodel to update a nucleo- tide base call generated by a call generation model from the same (or a subset of the same) metrics used by the call recalibrationmachine learningmodel to generate the variant call classifications.

[0030] Contributing at least in part to the improved accuracy, the call recalibration system exhibits improved flexibility over existing sequencing systems. For exam- ple, while many existing sequencing systems are limited to application at certain genomic coordinates and / or are incompatible with other genomic coordinates, in some embodiments, the call recalibration system flexibly adapt to many of these previously incompatible coordinates. Specifically, unlike some existing sequencing systems, the call recalibration system can generate nucleotide base calls and / or variant calls for multiallelic genomic coordinates, haploid genomic coordinates, and false homozygous reference genomic coordinates.

[0031] As another example of improved flexibility, as mentioned above, existing sequencing systems some- times utilize variant callers that rely exclusively on inter- nal sequencing metrics for particular base calls to gen- erate a nucleotide base call-without re-engineering or modifying such internal sequencing metrics or analyzing externally sourced sequencing metrics relevant to the genomic coordinates of corresponding nucleotide base calls. By contrast, in some embodiments, the call recali- bration system generates andmanipulates both external and internal sequencing metrics. Indeed, in some cases, the call recalibration system determines call model gen- erated sequencing metrics from variant caller compo- nents and mapping-and-alignment components of a call generation model by combining Bayesian probabilistic models with machine learning techniques in an efficient manner. In addition, the call recalibration system utilizes a call recalibration machine learning model to generate an updated nucleotide base call (e.g., from variant call classifications) from one or more sequencing metrics.

[0032] In addition to improved accuracy and flexibility, in certain embodiments, the call recalibration system 5 10 15 20 25 30 35 40 45 50 55 8 13 EP 4 693 301 A2 14 improves efficiency and speed. As noted above, some existing sequencing systems utilize computationally ex- pensive, slow neural network architectures (e.g., deep learning architectures such as convolutional neural net- works) that require many hours (e.g., 5‑8 hours with multiple processors executing on a server) and large amounts of computational resources to even implement and generate a file with variant calls from a sequencing run. Such deep learning architectures can further require several days (or weeks) to train. Conversely, the call recalibration system utilizes comparatively lightweight, fast architectures for both the call generation model and the call recalibration machine learning model. Indeed, contrasting with the many hours across multiple proces- sors required by existing sequencing systems, the call recalibration system, in many cases, requires under 30 minutes (for both the call generation model and the call recalibration machine learning model together) of run- time on a single field-programmable-gate array or a single processor to generate nucleotide base calls for a sample nucleotide sequence. Thus, the call recalibra- tion system is far faster and less computationally expen- sive than many deep learning approaches to variant calling. Not only are the models of the call recalibration system faster and less computationally expensive to implement, but themodelsof thecall recalibrationsystem are alsomuch faster and less computationally expensive to train than many existing deep-learning-based sys- tems.

[0033] As part of the improved speed and efficiency, in some embodiments, the call recalibration system recali- brates nucleotide base calls on a call-by-call basis as each call is processed by the call generation model. Indeed, thecall recalibrationsystemcangenerate variant call classifications for recalibrating a nucleotide base call (e.g., utilize the call recalibration machine learning mod- el) while also generating thenucleotide base call from the variant call classifications along with one or more se- quencing metrics. In some embodiments, the call recali- bration system utilizes the call generation model in par- allel with the call recalibration machine learning model to contemporaneously generate an initial nucleotide base call and variant call classifications for modifying or reca- librating the initial nucleotide base call.

[0034] As a further advantage over existing sequen- cing systems, in certain implementations, the call recali- bration system can identify or facilitate changes to in- dividual metrics that affect the accuracy of nucleotide base calls. While the neural network architectures of many existing sequencing systems render any interpre- tation of internal model data impossible with latent fea- tures, the call recalibration system utilizes model archi- tectures that facilitate interpretation of the effect of in- dividual sequencing metrics. More specifically, in some cases, the call recalibration system utilizes a call gen- eration model and a call recalibration machine learning model that enable extraction and analysis of individual sequencing metrics used throughout the process of gen- erating a nucleotide base call. Indeed, the call recalibra- tion system can determine respective contribution mea- sures for sequencing metrics involved in determining a nucleotide base call at a particular genomic coordinate.

[0035] As suggested by the foregoing discussion, this disclosure utilizes a variety of terms to describe features and benefits of the call recalibration system. Additional detail is hereafter provided regarding the meaning of these terms as used in this disclosure. As used in this disclosure, for instance, the term "sample nucleotide sequence" or "sample sequence" refers to a sequence of nucleotides isolated or extracted froma sample organ- ism (or a copyof suchan isolatedor extractedsequence). In particular, a sample nucleotide sequence includes a segment of a nucleic acid polymer that is isolated or extracted from a sample organism and composed of nitrogenous heterocyclic bases. For example, a sample nucleotide sequence can include a segment of deoxyr- ibonucleic acid (DNA), ribonucleic acid (RNA), or other polymeric forms of nucleic acids or chimeric or hybrid forms of nucleic acids noted below. More specifically, in some cases, the sample nucleotide sequence is found in a sample prepared or isolated by a kit and received by a sequencing device.

[0036] As further used herein, the term "nucleotide base call" (or sometimes simply "call") refers to a deter- mination or prediction of a particular nucleotide base (or nucleotide base pair) for a genomic coordinate of a sample genome or for an oligonucleotide during a se- quencing cycle. In particular, a nucleotide base call can indicate (i) a determination or prediction of the type of nucleotide base that has been incorporated within an oligonucleotide on a nucleotide-sample slide (e.g., read-based nucleotide base calls) or (ii) a determination or prediction of the type of nucleotide base that is present at a genomic coordinate or region within a sample gen- ome, including a variant call or a non-variant call in a digital output file. In some cases, for a nucleotide read, a nucleotide base call includes a determination or a pre- diction of a nucleotide base based on intensity values resulting from fluorescent-tagged nucleotides added to an oligonucleotide of a nucleotide-sample slide (e.g., in a well of a flow cell). Alternatively, a nucleotide base call includes a determination or a prediction of a nucleotide base to chromatogram peaks or electrical current changes resulting from nucleotides passing through a nanopore of a nucleotide-sample slide. By contrast, a nucleotide base call can also include an initial or final predictionofanucleotidebaseatagenomiccoordinateof a sample genome for a variant call file or other base call output file-based on nucleotide reads corresponding to the genomic coordinate. Accordingly, a nucleotide base call can include a base call corresponding to a genomic coordinate and a reference genome, such as an indica- tion of a variant or a non-variant at a particular location corresponding to the reference genome. Indeed, a nu- cleotide base call can refer to a variant call, including but not limited to, a single nucleotide polymorphism (SNP), 5 10 15 20 25 30 35 40 45 50 55 9 15 EP 4 693 301 A2 16 an insertion or a deletion (indel), or base call that is part of a structural variant. By using nucleotide base call, a sequencing system determines a sequence of a nucleic acid polymer. For example, a single nucleotide base call can comprise an adenine call, a cytosine call, a guanine call, or a thymine call for DNA (abbreviated as A, C, G, T) or a uracil call (instead of a thymine call) for RNA (ab- breviated as U).

[0037] Relatedly, as used herein, the term "nucleotide read" refers to an inferred sequence of one or more nucleotide bases (or nucleotide base pairs) from all or part of a sample nucleotide sequence. In particular, a nucleotide read includes a determined or predicted se- quence of nucleotide base calls for a nucleotide fragment (or group of monoclonal nucleotide fragments) from a sequencing library corresponding to a genome sample. For example, the call recalibration system determines a nucleotide read by generating nucleotide base calls for nucleotide bases passed through a nanopore of a nu- cleotide-sample slide, determined via fluorescent tag- ging, or determined from a well in a flow cell.

[0038] As noted above, in some embodiments, the call recalibration system determines sequencing metrics for nucleotide base calls of nucleotide reads. As used here- in, the term "sequencing metric" refers to a quantitative measurement or score indicating a degree to which an individual nucleotide base call (or a sequence of nucleo- tide base calls) aligns, compares, or quantifies with re- spect to a genomic coordinate or genomic region of a reference genome, with respect to nucleotide base calls from nucleotide reads, or with respect to external geno- mic sequencing or genomic structure. For instance, a sequencing metric includes a quantitative measurement or score indicating a degree towhich (i) individual nucleo- tide base calls align, map, or cover a genomic coordinate or reference base of a reference genome; (ii) nucleotide base calls compare to reference or alternative nucleotide reads in termsofmapping,mismatch, base call quality, or other raw sequencing metrics; or (iii) genomic coordi- nates or regions corresponding to nucleotide base calls demonstrate mappability, repetitive base call content, DNA structure, or other generalized metrics.

[0039] Relatedly, the term "diploid sequencing metric" refers to a sequencingmetric determined for a nucleotide base call at a diploid genomic coordinate. For example, a diploid sequencing metric includes a sequencing metric for a particular genomic coordinate of a nucleotide se- quence from (or is indicated to be from) a diploid chromo- some or a diploid nucleotide sequence (e.g., with two alleles at genomic regions corresponding to the genomic coordinate). Additionally, the term "haploid sequencing metric" refers to a sequencing metric determined for a nucleotide base call at a haploid genomic coordinate. For example, a haploid sequencing metric includes a se- quencing metric for a particular genomic coordinate of a nucleotide sequence from (or is indicated to be from) a haploid chromosome or a haploid nucleotide sequence (e.g., with a single allele at a genomic region correspond- ing to the genomic coordinate).

[0040] As further used herein, the term "genomic co- ordinate" (or sometimes simply "coordinate") refers to a particular locationorpositionofanucleotidebasewithina genome (e.g., an organism’s genome or a reference genome). In some cases, a genomic coordinate includes an identifier for a particular chromosome of a genome and an identifier for a position of a nucleotide base within the particular chromosome. For instance, a genomic coordinate or coordinates may include a number, name, or other identifier for a chromosome (e.g., chrl or chrX) and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chr1:1234570 or chr1:1234570‑1234870). Further, in certain implementations, a genomic coordinate refers to a source of a reference genome (e.g., mt for a mito- chondrial DNA reference genome or SARS-CoV‑2 for a reference genome for the SARS-CoV‑2 virus) and a position of a nucleotide base within the source for the reference genome (e.g., mt:16568 or SARS- CoV‑2:29001). By contrast, in certain cases, a genomic coordinate refers to a position of a nucleotide basewithin a reference genome without reference to a chromosome or source (e.g., 29727).

[0041] Relatedly, as used herein, the term "multiallelic genomic coordinate" refers to a genomic coordinate associated with three or more alleles. For example, a multiallelic genomic coordinate includes a genomic co- ordinate of a nucleotide sequence where nucleotide reads indicate three ormore possible alleles correspond- ing to the coordinate, such as a reference allele, a first alternate allele, a second alternate allele, and so forth. In some cases, a multiallelic genomic coordinate corre- sponds to a genomic coordinate where a read pileup occurs or where an insertion occurs. For instance, a multiallelic genomic coordinate can exhibit a multiallelic genotype, suchasa1 / 2 genotype,where the first allele at the coordinate corresponds to an allele from a first alter- nate nucleotide sequence and the second allele corre- sponds to an allele from a second alternate nucleotide sequence.

[0042] As mentioned above, in some embodiments, the call recalibration system generates nucleotide base calls for haploid genomic coordinates, or genomic co- ordinates within a haploid nucleotide sequence. As used herein, the term "haploidnucleotidesequence" refers toa sequenceof oneormorenucleotidebases fromahaploid chromosome (e.g., sexchromosome inmales) or asingle chromosome without a counterpart chromosome. For instance, a haploid nucleotide sequence can include a haploid region of a sample nucleotide sequence in which each of the genomic coordinate cover a nucleotide base from a haploid chromosome or a single chromosome without a counterpart chromosome. Thus, a haploid co- ordinate within a haploid nucleotide sequence has a haploid genotype, such as a haploid reference genotype (0) or a haploid alternate genotype (1).

[0043] Other coordinateswithin anucleotide sequence 5 10 15 20 25 30 35 40 45 50 55 10 17 EP 4 693 301 A2 18 can exhibit different genotypes. For example, a "homo- zygous reference genotype" refers to a genotype where both nucleotide bases at a given coordinate of a sample nucleotide sequence match a reference nucleotide base of a reference sequence or a reference genome (repre- sented as 0 / 0). As another example, a "homozygous alternate genotype" refers to a genotype at a given co- ordinate where both nucleotide bases differ from a re- ference nucleotide base of a reference sequence or a reference genome (represented as 1 / 1). As a further example, a "heterozygous genotype" refers to a geno- typewhere thenucleotidebasesatagivencoordinateare not the same. In some cases, a heterozygous genotype includes a genotype in which one nucleotide base matches a reference nucleotide base and the other nu- cleotide base differs from the reference nucleotide base (represented as 0 / 1 or 1 / 0). For multiallelic genomic coordinates, genotypes can exhibit nucleotide bases from more than one alternate nucleotide base differing froma reference nucleotide base of a reference genome. For instance, a multiallelic heterozygous genotype can be represented as 1 / 2, where one nucleotide base call matches a first alternate nucleotide base differing from a reference nucleotide base and the other nucleotide base callmatches a second alternate nucleotide base differing from the reference nucleotide base.

[0044] As noted above, a genomic coordinate includes a position within a reference genome. Such a position may be within a particular reference genome. As used herein, the term "reference genome" refers to a digital nucleic acid sequence assembled as a representative example (or representative examples) of genes and other genetic sequences of an organism. Regardless of the sequence length, in some cases, a reference genome represents an example set of genes or a set of nucleic acid sequences in a digital nucleic acid se- quenceddeterminedbyscientists as representativeof an organism of a particular species. For example, a linear human reference genome may be GRCh38 or other versions of reference genomes from the Genome Re- ference Consortium. As a further example, a reference genome may include a reference graph genome that includes both a linear reference genome and paths re- presenting nucleic acid sequences from ancestral hap- lotypes, such as Illumina DRAGEN Graph Reference Genome hg19.

[0045] In some embodiments, the call recalibration system determines various types of sequencing metrics from different sources, such as read-based sequencing metrics, externally sourced sequencing metrics, and call model generated sequencing metrics. As used herein, the term "read-based sequencing metrics" refers to se- quencing metrics derived from nucleotide reads of a sample nucleotide sequence. For example, read-based sequencing metrics include sequencing metrics deter- mined by applying statistical tests to detect differences betweena referencesequenceandnucleotide reads.For example, read-based sequencing metrics can include a comparative-mapping-quality-distribution metric that in- dicates a comparison between mapping qualities or a comparative-mismatch-count metric that indicates a comparison between mismatch counts.

[0046] By contrast, "externally sourced sequencing metrics" refer to sequencing metrics identified or ob- tained from one or more external databases. For exam- ple, externally sourced sequencing metrics include me- trics relating to mappability of nucleotides, replication timing, or DNA structure that are available outside of the call recalibration system.

[0047] Further, "call model generated sequencing me- trics" refer to internal, model-specific sequencingmetrics generated or extracted by a call generation model. For example, call model generated sequencing metrics in- clude variant calling sequencing metrics extracted or determined via variant caller components of a call gen- eration model and mapping-and-alignment sequencing metrics extracted or determined via mapping-and-align- ment components of a call generation model. As indi- cated above, call model generated sequencing metrics can include alignment metrics that quantify a degree to which sample nucleic acid sequences alignwith genomic coordinates of an example nucleic acid sequence, such as deletion-size metrics or mapping-quality metrics. Further, call model generated sequencing metrics can include depth metrics that quantify the depth of nucleo- tide base calls for sample nucleic acid sequences at genomic coordinates of an example nucleic acid se- quence, such as forward-reverse-depth metrics or nor- malized-depth metrics. Call model generated sequen- cing metrics can also include call-quality metrics that quantify a quality or accuracy of nucleotide base calls, such as nucleotide base call quality metrics, callability metrics, or somatic-quality metrics.

[0048] As used herein, the term "base call quality me- tric" refers to a specific score or other measurement indicating an accuracy of a nucleotide base call. In parti- cular, a base call quality metric comprises a value indi- cating a likelihood that one or more predicted nucleotide base calls for a genomic coordinate contain errors. For example, in certain implementations, a base call quality metric can comprise a Q score (e.g., a Phred quality score) predicting the error probability of any given nu- cleotide base call. To illustrate, a quality score (or Q score) may indicate that a probability of an incorrect nucleotide base call at a genomic coordinate is equal to 1 in 100 for aQ20 score, 1 in 1,000 for aQ30 score, 1 in 10,000 for a Q40 score, etc.

[0049] Relatedly, as used herein, the term "re-engi- neeredsequencingmetrics" refers tosequencingmetrics that have been updated, modified, augmented, refined, or re-engineered tomeasureor comparenucleotidebase calls (e.g., nucleotide base calls for reads or variant calls) with respect to other nucleotide base calls, a standard or reference, or for targeted for a particular objectiveor task. For example, re-engineered sequencing metrics can in- clude modifications to, or combinations of, raw sequen- 5 10 15 20 25 30 35 40 45 50 55 11 19 EP 4 693 301 A2 20 cingmetrics. In some embodiments, for instance, the call recalibration system generates one or more of the read- based sequencing metrics, the externally sourced se- quencing metrics, and / or the call model generated se- quencing metrics as re-engineered sequencing metrics. In some cases, re-engineered sequencing metrics refer to sequencing metrics that are generated by the call recalibration system and are therefore proprietary or internal to the call recalibration system and not available to third-party systems. Example re-engineered sequen- cing metrics include a comparative-mapping-quality-dis- tribution metric indicating a comparison between map- ping quality distributions associated with a reference sequence and alternatives supporting nucleotide reads or a comparative-base-quality metric indicating compar- isons between base qualities of a reference sequence and alternative supporting nucleotide reads.

[0050] As suggested above, the call recalibration sys- tem can utilize a machine learning model to modify sequencing metrics and update a nucleotide base call. As usedherein, the term "machine learningmodel" refers to a computer algorithm or a collection of computer algorithms that automatically improve for a particular task through experience based on use of data. For example, a machine learning model can utilize one or more learning techniques to improve in accuracy and / or effectiveness. Examplemachine learningmodels include various types of decision trees, support vector machines, Bayesian networks, or neural networks. In some cases, the call recalibration machine learning model is a series of gra- dient boosted decision trees (e.g., XGBoost algorithm), while in other cases the call recalibration machine learn- ing model is a random forest model, a multilayer percep- tron, a linear regression, a support vector machine, a deep tabular learning architecture, a deep learning trans- former (e.g., self-attention-based-tabular transformer), or a logistic regression.

[0051] In some cases, the call recalibration system utilizes a call recalibration machine learning model to modifyorupdateanucleotidebasecall basedonsequen- cing metrics. As used herein, the term "call recalibration machine learning model" refers to a machine learning model that generates variant call classifications. For example, in some cases, the call recalibration machine learning model is trained to generate variant call classi- fications indicating various probabilities or predictions for variant calls based on the sequencing metrics. Accord- ingly, in somecases, a call recalibrationmachine learning model a variant call recalibration machine learning mod- el. In certain embodiments, a call recalibration machine learningmodel includesmultiple sub-models or operates in tandem with another call recalibration machine learn- ing model. For instance, a first call recalibration machine learning model (e.g., an ensemble of gradient boosted trees) generates a first set of variant call classifications and a second call recalibration machine learning model (e.g., a random forest) generates a second set of variant call classifications.

[0052] Relatedly, the term "variant call classification" refers toapredictedclassification fromacall recalibration machine learning model that indicates a probability, score, or other quantitative measurement associated with some aspect of a nucleotide base call based on one ormore sequencingmetrics. A variant call classifica- tion can include a specialized prediction depending on the application of a call recalibration machine learning model. In embodiments for generating nucleotide base calls (or variant calls) for multiallelic genomic coordi- nates, variant call classifications can include: (i) a refer- ence probability of a homozygous reference genotype at amultiallelic genomic coordinate, (ii) a differing genotype probability of a genotype error at a multiallelic genomic coordinate, and (iii) a correct variant probability of a correct variant call genotype at a multiallelic genomic coordinate.

[0053] In embodiments for generating nucleotide base calls (or variant calls) for a haploid genomic coordinate, variant call classifications can include: (i) a first genotype probability of a first genotype at the genomic coordinate and (ii) a second genotype probability of a second geno- type at the genomic coordinate. As suggested above, the first genotype probability can be a probability that a genotype at a genotype coordinate is a haploid reference genotype, and the second genotype probability can be a probability that a genotype at the genotype coordinate is a haploid alternate genotype. In these or other embodi- ments, such as embodiments for generating nucleotide base calls (or variant calls) for genomic coordinates indicated to exhibit homozygous reference genotypes, variant call classifications can include: (i) a false positive classification or a homozygous reference classification indicating a probability that a nucleotide base call is a false positive or a homozygous reference genotype, respectively; (ii) a genotype error classification or a het- erozygous genotype classification indicating a probabil- ity that a genotype (e.g., an indication of a heterozygous or homozygous genotype for a variant call at a particular location) is incorrect or a heterozygous genotype, re- spectively; and / or (iii) a true-positive classification or a homozygous alternate classification indicating a prob- ability that a nucleotide base call is a true positive or a homozygous alternate genotype, respectively. In some cases, the variant call classifications accordingly repre- sent intermediate scoring metrics and / or a predicted probability that a genotype for a nucleotide base call is accurate.

[0054] As mentioned, in some embodiments, the call recalibration machine learning model can be a neural network. The term the term "neural network" refers to a machine learningmodel that can be trained and / or tuned based on inputs to determine classifications or approx- imate unknown functions. For example, a neural network includes a model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate complex functions and generate outputs (e.g., generated digital images) based on a plurality of 5 10 15 20 25 30 35 40 45 50 55 12 21 EP 4 693 301 A2 22 inputs provided to the neural network. In some cases, a neural network refers to an algorithm (or set of algo- rithms) that implements deep learning techniques to model high-level abstractions in data. For example, a neural network can include a convolutional neural net- work, a recurrent neural network (e.g., anLSTM), a graph neural network, a self-attention transformer neural net- work, or a generative adversarial neural network.

[0055] As noted above, the call recalibration system can generate variant call classifications that indicate or reflect a likelihood of identifying a variant at a genomic coordinate. As used herein, the term "variant" refers to a nucleotide base or multiple nucleotide bases that do not align with, differs from, or varies from a corresponding nucleotide base (or nucleotide bases) in a reference sequence or a reference genome. For example, a variant includes a SNP, an indel, or a structural variant that indicates nucleotide bases in a sample nucleotide se- quence that differ from nucleotide bases in correspond- ing genomic coordinates of a reference sequence. Along these lines, a "variant nucleotide base call" (or simply "variant call") refers toanucleotidebasecall comprisinga variant at a particular genomic coordinate. Conversely, a "non-variant nucleotide base call" (or simply "non-variant call") refers to a nucleotide base call comprising a non- variant at a genomic coordinate.

[0056] As mentioned, in some embodiments, the call recalibration system modifies data fields corresponding to a variant call file. As used herein, the term "variant call file" refers to a digital file that indicates or represents one or more nucleotide base calls (e.g., variant calls) com- pared toa referencegenomealongwithother information pertaining to thenucleotidebasecalls (e.g., variant calls). For example, a variant call format (VCF) file refers to a text file format that contains information about variants at specific genomic coordinates, including meta-informa- tion lines, a header line, and data lines where each data line contains information about a single nucleotide base call (e.g., a single variant). As described further below, the call recalibration system can generate different ver- sions of variant call files, including a pre-filter variant call file comprising variant nucleotide base calls that either pass or fail a quality filter for base call quality metrics or a post-filter variant call file comprising variant nucleotide base calls that pass the quality filter but excludes variant nucleotide base calls that fail the quality filter.

[0057] In some embodiments, the call recalibration system modifies data fields corresponding to metrics of a nucleotide base call associated with a variant call file, such as fields for call quality, genotype, and genotype quality. As used herein, the term "call quality" when used with respect to a data field in a variant call file refers to a measure or an indication of a likelihood or a probability that a variant exists at a given location. Accordingly, a call quality field (or QUAL field) corresponding to a VCF file may include a base call quality metric, such as a Phred- scaledquality orQscore, representingaprobability that a genomic coordinate of a sample genome includes a variant. Similarly, a "genotype quality" when used with respect to a field refers to a likelihood or a probability that a particular predicted genotype for a nucleotide base call is correct.

[0058] As noted, in some embodiments, the call reca- libration system utilizes a call generation model to gen- erate a nucleotide base call for a genomic coordinate. As used herein, the term "call generation model" refers to a probabilistic model that generates sequencing data from nucleotide reads of a sample nucleotide sequence, in- cluding nucleotide base calls and associated metrics. Accordingly, in some cases, a call generation model may be a variant call generation model. For example, in some cases, a call generation model refers to a Baye- sian probability model that generates variant calls based on nucleotide reads of a sample nucleotide sequence. Such a model can process or analyze sequencing me- trics corresponding to read pileups (e.g.,multiple nucleo- tide reads corresponding to a single genomic coordi- nate), including mapping quality, base quality, and var- ious hypotheses including foreign reads, missing reads, joint detection, and more. A call generation model may likewise include multiple components, including, but not limited to, different software applications or components for mapping and aligning, sorting, duplicate marking, computing read pileup depths, and variant calling. In some cases, the call generation model refers to the ILLUMINA DRAGEN model for variant calling functions and mapping and alignment functions.

[0059] As mentioned above, in certain described em- bodiments, the call recalibration system generates or determines contribution measures associated with indi- vidual sequencing metrics. As used herein, the term "contribution measure" refers to a measure of effect, influence, or impact that a sequencing metric has on a given recalibration of fields for a base call output file (e.g., a variant call file), a given nucleotide base call in a base call output file, or (in particular) a given variant call. For example, a contribution measure indicates how much of a role one sequencing metric plays in determining a nucleotide base call over a different nucleotide base call (and compared to other sequencing metrics).

[0060] The following paragraphs describe the call re- calibration system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a system environment (or "environment") 100 in which a call recalibrationsystem106operates inaccordancewith one or more embodiments. As illustrated, the environ- ment 100 includes one or more server device(s) 102 connected toaclient device108andasequencingdevice 114 via a network 112. While FIG. 1 shows an embodi- ment of the call recalibration system 106, this disclosure describes alternative embodiments and configurations below.

[0061] As shown in FIG. 1, the server device(s) 102, the client device 108, and the sequencing device 114 can communicate with each other via the network 112. The 5 10 15 20 25 30 35 40 45 50 55 13 23 EP 4 693 301 A2 24 network 112 comprises any suitable network over which computing devices can communicate. Example net- works are discussed in additional detail below with re- spect to FIG. 15.

[0062] As indicated by FIG. 1, the sequencing device 114 comprises a device for sequencing a nucleic acid polymer. In some embodiments, the sequencing device 114 analyzes nucleic acid segments or oligonucleotides extracted from genomic samples to generate nucleotide readsorotherdatautilizingcomputer implementedmeth- ods and systems (described herein) either directly or indirectly on the sequencing device 114. More particu- larly, the sequencing device 114 receives and analyzes, within nucleotide-sample slides (e.g., flow cells), nucleic acid sequences extracted from samples. In one or more embodiments, the sequencing device 114 utilizes SBS to sequence nucleic acid polymers into nucleotide reads. In addition or in the alternative to communicating across the network 112, in some embodiments, the sequencing device 114bypasses thenetwork 112and communicates directly with the client device 108.

[0063] As further indicated by FIG. 1, the server de- vice(s) 102 may generate, receive, analyze, store, and transmit digital data, suchasdata for determiningnucleo- tide base calls or sequencing nucleic acid polymers. As shown in FIG. 1, the sequencing device 114 may send (and the server device(s) 102may receive) call data from the sequencing device 114. The server device(s) 102 may also communicate with the client device 108. In particular, the server device(s) 102 can send data to the client device 108, including a variant call file or other information indicating nucleotide base calls, sequencing metrics, error data, or other metrics associated with a nucleotide base call.

[0064] In someembodiments, the server device(s) 102 comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 112 and located in the same or different physical locations. Further, the server device(s) 102 can comprise a content server, an applica- tion server, a communication server, a web-hosting ser- ver, or another type of server. In some cases, the server device(s) 102 are located at a same physical location as the sequencing device 114.

[0065] As further shown in FIG. 1, the server device(s) 102 can include a sequencing system104.Generally, the sequencing system 104 analyzes call data, such as sequencingmetrics received from thesequencingdevice 114, to determine nucleotide base sequences for nucleic acid polymers. For example, the sequencing system 104 can receive rawdata from thesequencingdevice114and can determine a nucleotide base sequence for a nucleic acid segment. In some embodiments, the sequencing system 104 determines the sequences of nucleotide bases inDNAand / or RNA segments or oligonucleotides. In addition to processing and determining sequences for nucleic acid polymers, the sequencing system 104 also generates a variant call file indicating one or more nu- cleotide base calls and / or variant calls for one or more genomic coordinates.

[0066] As just mentioned, and as illustrated in FIG. 1, the call recalibration system 106 analyzes call data, such as sequencing metrics from the sequencing device 114, to determine nucleotide base calls for sample nucleic acid sequences. The call recalibration system 106 in- cludes a call generation model and a call recalibration machine learning model. In some embodiments, the call recalibration system106 determines sequencingmetrics for sample nucleotide sequences. Based on data derived or prepared from the sequencing metrics, the call recali- bration system 106 trains and applies a call generation model to determine nucleotide base calls for the sample sequence corresponding to genomic coordinates. The call recalibration system 106 further utilizes a call recali- bration machine learning model to generate sets of var- iant call classifications to update ormodify the nucleotide base calls (and / or variant calls). Based on such data, for example, the call recalibration system 106 can update data fields corresponding to a variant call file to update a nucleotide base call and / or a variant call for improved accuracy.

[0067] As further illustrated and indicated in FIG. 1, the client device 108 can generate, store, receive, and send digital data. Inparticular, theclient device108can receive sequencing metrics from the sequencing device 114. Furthermore, the client device 108 may communicate with the server device(s) 102 to receive a variant call file comprising nucleotide base calls and / or other metrics, such as a call-quality, a genotype indication, and a gen- otype quality. The client device 108 can accordingly present or display information pertaining to the nucleo- tide base call within a graphical user interface to a user associated with the client device 108. For example, the client device 108 can present a contribution-measure interface that includes a visualization or a depiction of various contribution measures associated with, or attrib- uted to, individual sequencing metrics with respect to a particular nucleotide base call.

[0068] The client device 108 illustrated in FIG. 1 may comprise various types of client devices. For example, in some embodiments, the client device 108 includes non- mobiledevices, suchasdesktopcomputersorservers, or other types of client devices. In yet other embodiments, the client device 108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 108 are discussed below with respect to FIG. 15.

[0069] As further illustrated in FIG. 1, the client device 108 includes a sequencing application 110. The sequen- cing application 110may be aweb application or a native application stored and executed on the client device 108 (e.g., a mobile application, desktop application). The sequencing application 110 can include instructions that (when executed) cause the client device 108 to receive data from the call recalibration system 106 and present, for display at the client device 108, data fromavariant call 5 10 15 20 25 30 35 40 45 50 55 14 25 EP 4 693 301 A2 26 file. Furthermore, the sequencing application 110 can instruct the client device 108 to display a visualization of contribution measures for sequencing metrics of a nucleotide base call.

[0070] As further illustrated in FIG. 1, the call recalibra- tion system 106 may be located on the client device 108 as part of the sequencing application 110 or on the sequencing device 114. Accordingly, in some embodi- ments, the call recalibration system 106 is implemented by (e.g., located entirely or in part) on the client device 108. In yet other embodiments, the call recalibration system 106 is implemented by one or more other com- ponents of the environment 100, such as the sequencing device 114. In particular, the call recalibration system106 can be implemented in a variety of different ways across the server device(s) 102, the network 112, the client device 108, and the sequencing device 114. For exam- ple, the call recalibration system 106 can be downloaded from the server device(s) 102 to the client device 108 and / or to the sequencing device 114 where all or part of the functionality of the call recalibration system 106 is performed at each respective device within the environ- ment 100.

[0071] As further illustrated in FIG. 1, the environment 100 includes a database 116. The database 116 can store information, such as variant call files, sample nu- cleotide sequences, nucleotide reads, nucleotide base calls, variant calls, and sequencing metrics. In some embodiments, the server device(s) 102, the client device 108, and / or the sequencing device 114 communicate with the database 116 (e.g., via the network 112) to store and / or access information, such as variant call files, sample nucleotide sequences, nucleotide reads, nucleo- tide base calls, variant calls, and sequencing metrics. In some cases, the database 116 also stores one or more models, such as a call recalibration machine learning model and / or a call generation model.

[0072] Though FIG. 1 illustrates the components of environment 100 communicating via the network 112, in certain implementations, the components of environ- ment 100 can also communicate directly with each other, bypassing the network 112. For instance, and as pre- viously mentioned, in some implementations, the client device 108 communicates directly with the sequencing device 114.Additionally, in someembodiments, the client device 108 communicates directly with the call recalibra- tion system 106. Moreover, the call recalibration system 106 can access one or more databases housed on or accessed by the server device(s) 102 or elsewhere in the environment 100.

[0073] As indicated above, the call recalibration sys- tem 106 can determine a nucleotide base call based on one or more variant call classifications. In particular, the call recalibration system 106 can determine variant call classifications from sequencing metrics utilizing a call recalibrationmachine learningmodel and can determine or update various metrics associated with a nucleotide base call from the generated variant call classifications. FIG. 2 illustrates an example overview of determining a nucleotide base call based on variant call classifications in accordance with one or more embodiments.

[0074] As illustrated in FIG. 2, the call recalibration system 106 performs an act 202 to determine sequen- cing metrics. In particular, the call recalibration system 106 determines sequencing metrics such as read-based sequencing metrics, externally sourced sequencing me- trics, and call model generated sequencing metrics. For example, the call recalibration system 106 determines sequencing metrics that indicate various attributes or data in relation to various nucleotide base calls of nucleo- tide reads from a sample nucleotide sequence. Addi- tional detail regarding determining the various types of sequencing metrics is provided below with reference to FIGS. 6A‑6C.

[0075] As further illustrated in FIG. 2, the call recalibra- tion system 106 performs an act 204 to generate variant call classifications. More specifically, the call recalibra- tion system 106 generates (or updates or refines) variant call classifications fromsequencingmetricsutilizingacall recalibration machine learning model. To elaborate, the call recalibration system106 utilizes the call recalibration machine learning model to process or analyze one or more sequencing metrics and to generate a set of clas- sifications (e.g., predicted probabilities associated with genotypes). For instance, the call recalibration system 106 generates, utilizing the call recalibration machine learning model, a set of variant call classifications (re- presented inFIG. 2 as "Class 1," "Class 2," and "Class 3") that indicate certain probabilities associated with a gen- otype of a corresponding nucleotide base call based on the sequencing metrics.

[0076] In some embodiments, the call recalibration system106generatesdifferent variant call classifications for different applications and / or for different genomic coordinates. For example, the call recalibration system 106 generates a first set of variant call classifications for multiallelic genomic coordinates, generates a second set of variant call classifications for haploid genomic coordi- nates, and generates a third set of variant call classifica- tion for genomic coordinates indicated to exhibit homo- zygous reference genotypes. In certain embodiments, the call recalibration system 106 generates the same variant call classifications fordifferent applicationsand / or for different genomic coordinates but utilizes them differ- ently or utilizes different information associated with the variant call classifications. Additional detail regarding generating variant call classifications is provided below with reference to subsequent figures.

[0077] As further illustrated in FIG. 2, the call recalibra- tion system 106 also performs an act 206 to determine a final nucleotide base call (or a variant call) based on the variant call classifications. More particularly, the call re- calibration system 106 determines or updates a nucleo- tide base call for a sample nucleotide sequence at a genomic coordinate within a reference genome. To de- termineorgenerate thefinal nucleotidebasecall, in some 5 10 15 20 25 30 35 40 45 50 55 15 27 EP 4 693 301 A2 28 embodiments the call recalibration system 106 deter- mines initial nucleotide base calls utilizing a call genera- tion model and edits or updates certain initial nucleotide base calls based on the variant call classifications gen- erated by the call recalibration machine learning model.

[0078] To elaborate, the call recalibration system 106 utilizes a call generation model to process or analyze sequencing metrics (e.g., one or more of the same se- quencing metrics used to generate the variant call clas- sifications in act 204) to determine a nucleotide base call (e.g., an initial nucleotide base call) from the sequencing metrics. For example, the call recalibration system 106 applies a number of Bayesian probabilistic models or algorithms to derive various probabilities for different nucleotide bases, quality metrics, mapping metrics, joint metrics, and other data occurring within the sample nu- cleotide sequence to include within a variant call file. From the probabilistic models, the call recalibration sys- tem 106 determines a nucleotide base call (e.g., a call indicating a difference or sameness to a reference base from a reference genome) that indicates a predicted nucleotide base for the sample genome at a correspond- ing genomic coordinate.

[0079] As further illustrated in FIG. 2, in certain imple- mentations, the call recalibration system 106 utilizes the initial variant call classifications (e.g., as determined via the act 204) to generate, recalibrate, determine, modify, or augment thenucleotidebasecall. Toelaborate, thecall recalibration system 106 utilizes probabilities associated with the variant call classifications to determine or update certain metrics associated with a nucleotide base call. For example, the call recalibration system 106 modifies data fields corresponding to a variant call file for metrics, such as call quality, genotype, and genotype quality (or others as described below).

[0080] In somecases, the call recalibration system106 extrapolates from the variant call classifications to de- termine metrics corresponding to a variant call file, such as call quality, genotype, and genotype quality asso- ciated with the nucleotide base call. For instance, by utilizing a genotype error classification, the call recalibra- tion system 106 can remedy certain errors in or asso- ciated with an initial nucleotide base call. Indeed, if the call recalibration system 106 determines a high false positive probability for a nucleotide base call, then the call recalibration system106applies the call recalibration machine learning model to function as a variant filter to modify (e.g., reduce) a call quality associated with the nucleotide base call. As another example, the call recali- bration system106utilizes a genotype error probability to modify a genotype and / or a genotype quality of a nucleo- tide base call in cases where systems would previously filter out or doubly penalize heterozygous / homozygous (het / hom) errors (e.g., where the system generates a nucleotide base call that is incorrect which further results in missing a nucleotide base call that is correct).

[0081] In certain embodiments, the call recalibration system106considersasinglevariant call classification to modify a data field for a nucleotide base call (e.g., a call quality, a genotype, or a genotype quality). In other embodiments, the call recalibration system 106 consid- ers multiple variant call classifications at once (e.g., in a weighted combination) to modify or update one or more data fields for call quality, genotype, and / or genotype quality. Additional detail regarding generating and mod- ifying nucleotide base calls is provided below with refer- ence to subsequent figures.

[0082] In one or more implementations, the call recali- bration system 106 generates the variant call classifica- tions (e.g., via the act 204)while, or during the process of, determining a nucleotide base call. For example, the call recalibration system 106 simultaneously implements the call recalibration machine learning model and the call generation model to generate a nucleotide base call and variant call classifications for modifying the nucleotide base call. The call recalibration system 106 further modi- fies data fields corresponding to a variant call file of the nucleotide base call to generate a finalized nucleotide base call (e.g., within a pre-filter or post-filter variant call file). Indeed, the call recalibration system 106 generates the finalized (e.g., recalibrated) nucleotide base call from the variant call classifications as well as sequencing metrics processed by the call generation model (e.g., one or more of the same sequencing metrics used to generate the variant call classifications). As described above, this simultaneous or parallel operation affords the call recalibration system 106 improved computational efficiency and increased speed by recalibrating nucleo- tide base calls as they are initially generated (rather than performing one operation before the other).

[0083] In one or more implementations, the call recali- bration system 106 determines a nucleotide base call as part of a SNP, a deletion, an insertion, or a structural variation. For example, the call recalibration system 106 determines a nucleotide base call represent an SNP at a genomic coordinate (e.g., chr1:151863125) by identify- ing a G in the sample nucleotide sequence where an A exists in the reference sequence. As another example, the call recalibration system 106 determines nucleotide base calls surrounding one ormore genomic coordinates (e.g., chr1:49263256) indicate a deletion by identifying a single G in the sample nucleotide sequence where GTAAC exists in the reference sequence.

[0084] As a further example, the call recalibration sys- tem 106 determines a sequence of nucleotide base calls represent an insertion at a genomic coordinate (e.g., chr1:7602080) by identifying a sequence of TTTCC in the sample nucleotide sequence where a Texists in the reference sequence. Indeed, in some cases, an insertion includes a sequence of nucleotide base calls that replace a single reference base at a genomic coordinate of a reference sequence.

[0085] In some embodiments, the call recalibration system 106 sets a quality threshold (e.g., a customized quality threshold) for base call quality metrics at genomic coordinates for a genomic sample (e.g., including one or 5 10 15 20 25 30 35 40 45 50 55 16 29 EP 4 693 301 A2 30 more of diploid coordinates, haploid coordinates, multi- allelic coordinates, and genomic coordinates incorrectly identified as exhibiting homozygous reference geno- types). The base call quality metrics can change signifi- cantly between a call generation model and a call recali- bration machine learning model. To adjust for the poten- tially broad range and significant changes of base call quality metrics, the call recalibration system 106 can determine or set a hard filter QUAL threshold for a variant call file output that results in (or corresponds to) a favor- able F1 position as a measure of performance (e.g., a favorable trade-off between false positives and false negatives).

[0086] Such a favorable F1 position can include a score or a position with a favorable (e.g., best) tradeoff between precision and recall of calling variants. In some cases, for instance, an F1 position (or an F1 score) is proportional to a combination (e.g., sum) of false positive variants and false negative variants (which means that a favorable F1 score corresponds to a lowFP+FNmetric). As described below, for instance, FIG. 10B depicts ex- amples of FP+FN metrics based upon which the call recalibration system106 sets a quality threshold for base call quality metrics at haploid genomic coordinates. In some embodiments, the call recalibration system 106 thusutilizesaquality filter for base call qualitymetrics that result in the favorable F1 position for the Call Recalibra- tion System 1 or 2 depicted in FIG. 10B.

[0087] As indicated above, however, the call recalibra- tion system 106 can set such a quality threshold for base call quality metrics at any or all genomic coordinates resulting in a favorable F1 position when using a call recalibration machine learning model. Indeed, in some embodiments, the call recalibration system 106 gener- atesF1scoresandapplies relatedfiltering logic forQUAL scores (as described above) for various genomic coor- dinates, including haploid coordinates, diploid coordi- nates, multiallelic coordinates, genomic coordinates in- correctly identified as exhibiting homozygous reference genotypes, or other genomic coordinates.

[0088] Thus, in some cases, rather than a call genera- tion model discarding certain variant nucleotide base calls that do not pass a previous quality filter, the call recalibration system 106 executes a pipeline of the fol- lowing acts: (i) utilizing a call generation model to gen- erate variant nucleotide base calls across various re- gions or coordinates; (ii) utilizing a call recalibration ma- chine learning model to recalibrate variant nucleotide base calls and corresponding metrics, such as one or more of base call quality metrics, genotype quality me- trics, or genotype metrics in corresponding VCF fields; (iii) generating a prefiltered VCF that includes variant nucleotide base calls above a quality threshold either because the call generation model called a variant nu- cleotide base call at a correspondinggenomic coordinate or because the call recalibrationmachine learningmodel called a variant nucleotide base call at a genomic co- ordinate at which the call generation model had deter- mined such as variant nucleotide base call did not pass a previous quality filter; and (iv) utilizing a hard quality threshold filter to select quality variant nucleotide base calls from the prefiltered VCF. Such a hard quality thresh- old is configured such that the filtered output of the call recalibration system 106 is close to a favorable F1 posi- tion (thereby resulting in a post-filter VCF that contains only variant nucleotide base calls satisfying the hard quality threshold). The call recalibration system 106 can change the QUAL threshold depending on whether the call recalibration machine learning model is active or the call generation model (e.g., DRAGEN) is executing without the call recalibration machine learning model.

[0089] As mentioned above, in certain described em- bodiments, the call recalibration system 106 generates variant call classifications for multiallelic genomic coor- dinates. In addition, the call recalibration system 106 generates or updates a variant call file for the multiallelic coordinate based on the variant call classifications. FIGS. 3A‑3B illustrate an example flow of the call recali- bration system 106 generating a variant call file from variant call classifications of a multiallelic genomic co- ordinate in accordance with one or more embodiments. For example, FIG. 3A illustrates the call recalibration system 106 generating variant call classifications for multiallelic genomic coordinates in accordance with oneormoreembodiments. Thereafter, FIG. 3B illustrates the call recalibration system 106 generating a variant call file from variant call classifications in accordance with one or more embodiments.

[0090] As illustrated in FIG. 3A, the call recalibration system 106 identifies a multiallelic genomic coordinate 302. For instance, the call recalibration system 106 identifies the multiallelic genomic coordinate 302 from nucleotide base calls corresponding to a sample nucleo- tide sequenceor basedonhaplotypedata corresponding to the multiallelic genomic coordinate 302. In some cases, the call recalibration system 106 identifies the multiallelic genomic coordinate 302 by determining (i) that nucleotide base calls fromnucleotide reads covering the genomic coordinate include more than two possible nucleotide base calls from corresponding alleles and (ii) that the nucleotide base calls satisfy one or more thresh- old sequencingmetrics (e.g., a base-call-qualitymetric of Q30). Additionally or alternatively, in certain embodi- ments, the call recalibration system 106 identifies the genomic coordinate by from a database comprising a haplotype reference panel correlated with specific geno- mic coordinates. Based on different haplotype probabil- ities in the database, the call recalibration system 106 identifies genomic coordinates as probable candidates for multiallelic genomic coordinates. Regardless of the identification method, in some cases, the call recalibra- tion system 106 uses a call generation model (e.g., a variant caller within a call generation model) to identify the multiallelic genomic coordinate 302.

[0091] As depicted in FIG. 3A, for instance, the call recalibration system 106 analyzes the nucleotide base 5 10 15 20 25 30 35 40 45 50 55 17 31 EP 4 693 301 A2 32 sequences corresponding to coordinates 1 through 5 to identify genomic coordinate 4 as themultiallelic genomic coordinate 302. For simplicity of illustration, each of the nucleotide base sequences constitute representatives of different nucleotide reads thatmap todifferent allelesand correspond to coordinates 1 through 5. While each of coordinates 1 through 5 have three possible alleles cor- responding to them-one from the reference genome, another from a first possible allele "Alternate allele 1," and another from a second possible allele "Alternate allele 2"-only coordinate 4 exhibits three ormore different nucleotide base calls from different alleles that could possibly be assigned. Specifically, coordinate 4 could exhibit a G from the reference genome, a C from the first alternate allele, or a T from the second alternate allele.

[0092] As further illustrated in FIG. 3A, the call recali- bration system 106 determines sequencing metrics 304 for the multiallelic genomic coordinate 302. In particular, the call recalibration system 106 determines sequencing metrics associatedwith nucleotide reads, generated by a call generation model, or retrieved from an external source. Additional detail regarding determining the se- quencing metrics 304 is provided below with specific reference to FIGS. 6A‑6C.

[0093] Additionally, as shown in FIG. 3A, the call re- calibration system 106 utilizes a call recalibration ma- chine learning model 306 to generate variant call classi- fications 308. Specifically, the call recalibration system 106 utilizes the call recalibrationmachine learningmodel 306 to generate a reference probability 310 indicating a probability of a homozygous reference genotype at the multiallelic genomiccoordinate302.Toelaborate, thecall recalibration system 106 generates a variant call classi- fication that indicates a probability that, based on its sequencing metrics, the multiallelic genomic coordinate 302 exhibits a homozygous genotype with respect to a reference genome. As shown, the reference probability 310 indicates a probability of a homozygous reference genotype (0 / 0) at coordinate 4 from the sample nucleo- tide sequence (represented as P(0 / 0)@4).

[0094] In one or more implementations, the call recali- bration system 106 also utilizes the call recalibration machine learning model 306 to generate a differing gen- otype probability 312 indicating a probability of a geno- type error at the multiallelic genomic coordinate 302. For instance, the call recalibration system 106 determines a probability that a predicted genotype for the multiallelic genomic coordinate 302 is an incorrect genotype (e.g., a genotype incorrectly identified by a call generation mod- el) or includes an incorrect allele in the predicted geno- type. To elaborate, in some cases, the call recalibration system 106 determines a probability that any het / hom error exists at the multiallelic genomic coordinate 302- e.g., where the alternate base is correct but the genotype is wrong-or a probability that the nucleotide base calls represent either the wrong genotype altogether or the wrong allele(s) in the predicted genotype. For example, when determining a probability that a het / hom error ex- ists, the call recalibration system 106 determines a prob- ability that an alternate base call represented as "1" is correct, but the genotype is incorrect, such as a prob- ability of incorrectly determining a 0 / 1 genotype call (e.g., A / T) instead of a correct 1 / 1 genotype call (e.g., T / T) (or vice versa when the correct genotype call is 0 / 1).

[0095] By determining the differing genotype probabil- ity 312, the call recalibration system 106 can fix inac- curacies of existing sequencing systemswhere incorrect calls are often indels. In particular, the call recalibration system 106 can more accurately generate nucleotide base calls for genomic coordinates corresponding to indels where existing sequencing systems would deter- mine a nucleotide base call represent an incorrect geno- type that represents an incorrect allele resulting from a long inserted or deleted sequence. As shown, the differ- ing genotype probability 312 indicates a probability of a different genotype belonging at coordinate 4 (repre- sented as P(diff genotype)@4).

[0096] As further illustrated in FIG. 3A, the call recali- bration system 106 utilizes the call recalibrationmachine learning model 306 to generate a correct variant prob- ability 314 indicating a probability of a correct variant call genotype at the multiallelic genomic coordinate 302. For example, the call recalibration system 106 generates a probability that a predicted genotype for the multiallelic genomic coordinate302 is correct asdeterminedbyacall generation model. As shown, the call recalibration sys- tem 106 determines the correct variant probability 314 indicating a probability that the predicted variant call from the call generation model is correct for coordinate 4 (represented as P(correct)@4).

[0097] Continuing to FIG. 3B, in some embodiments, the call recalibration system 106 utilizes the variant call classifications 308 to update one or more data fields or variant call file fields ("VCF" fields) associated with a variant call file (e.g., a variant call file 324). For example, the call recalibration system106generates updatedVCF fields 316 that indicate updated sequencing metrics for a final nucleotide base call. In some cases, the call recali- bration system 106modifies or updates only certain VCF fields and does not update others based on the variant call classifications 308. In other cases, the call recalibra- tion system106 does not update VCF fields based on the variant call classifications 308.When generating nucleo- tide base calls for the multiallelic genomic coordinate 302, for instance, the call recalibration system 106 does not update certain fields, such as a genotype (GT) field, based on the variant call classifications 308. Thus, in contrast to biallelic genomic coordinates, in some cases, the call recalibration system 106 does not modify or update a GT field because, in some cases, there is not enough information to determine a new or updated gen- otype at a multiallelic genomic coordinate.

[0098] To illustrate one embodiment, FIG. 3B depicts the call recalibration system 106 generating updated VCFfields 316 for a genotype (GT) of 1 / 2, where cytosine represents a reference base (shown as "Ref: C") at a 5 10 15 20 25 30 35 40 45 50 55 18 33 EP 4 693 301 A2 34 multiallelic genomic coordinate for an allele correspond- ing to the reference genome, adenine represents a first alternate base ("Alt 1: A") at the multiallelic genomic coordinate for a different allele, and thymine represents a second alternate base ("Alt 2: T") at the multiallelic genomic coordinate for yet a different allele. But FIG. 3B merely depicts examples of a possible reference base and possible alternate bases at a multiallelic genomic coordinate. The call recalibration system 106 can gen- erate variant call classifications and modify correspond- ing metrics in VCF fields for various other reference bases and alternate bases at other multiallelic genomic coordinates.

[0099] As further illustrated in FIG. 3B, the call recali- bration system 106 generates an updated base call quality (QUAL) field 318. More specifically, the call re- calibration system 106 modifies or updates a base call qualitymetric basedon the variant call classifications308 to indicate an accuracy of a nucleotide base call at the multiallelic genomic coordinate 302. As shown, the up- dated base call quality field 318 indicates a QUAL score of 48 for a variant at the corresponding genomic coordi- nate. In this example, the updated base call qualitymetric (e.g., QUAL score of 48) represents a score for any type of variant at the corresponding multiallelic genomic co- ordinate. In addition, the call recalibration system 106 generates a modified or updated genotype quality (GQ) field 320. For instance, based on the variant call classi- fications 308, the call recalibration system106 generates a modified or updated genotype quality metric indicating a likelihood or a probability that a predicted genotype at the multiallelic genomic coordinate 302 is correct. As shown, for instance, the updated genotype quality field 320 indicates a genotype quality metric for a genotype call with a heterozygous genotype (e.g., a GQ score of 4 foragenotypeof1 / 2atamultiallelic genomiccoordinate).

[0100] In one or more embodiments, the call recalibra- tion system 106 further generates or updates genotype likelihoods 322 and (in some cases) uses the genotype likelihoods 322 to rank alleles. To elaborate, the call recalibration system 106 generates updated genotype likelihoods as the genotype likelihoods 322 by ordering candidate nucleotide base calls at the multiallelic geno- mic coordinate302according to the candidatenucleotide base calls’ respective probabilities of belonging at the multiallelic genomic coordinate 302. For example, the call recalibration system 106 determines probabilities associated with a plurality of genotypes where each diploid genotype is composed of a pair of alleles. As another example, the call recalibration system 106 de- termines relative probabilities associated with a plurality of alleles (e.g., from a reference genome, a first alternate allele, and a second alternate allele) of belonging at the multiallelic genomic coordinate 302 of the sample nu- cleotide sequence. In some embodiments, the call reca- libration system 106 generates metrics for a PHRED- scale Likelihood (PL) field as part of the updated VCF fields 316. For example, the call recalibration system106 generates metrics for a PL field that can indicate geno- types, such as homozygous reference, heterozygous, and homozygous alternate genotypes (e.g., with PL field nomenclature 9 / 0 / 3, respectively).

[0101] Indeed, the call recalibration system 106 gen- erates the allele-specific probabilities or likelihoods based on a relative probability of a nucleotide base call corresponding to an allele from a call generation model versus any other (non-reference) genotype identified by the call recalibration machine learning model 306. For instance, in some embodiments, the call recalibration system 106 indicates relative probability scores for each allele corresponding to respective nucleotide base calls in PL fields indicating normalized PHRED-scale likeli- hoods for genotypes and / or Genotype Likelihood (GL) fields indicating log-scaled likelihoods (e.g., log10- scaled) of data (e.g., sequencing metrics) given a called genotype.

[0102] As an example of generating updated genotype likelihoods and modifying certain VCF fields, in some cases, the call recalibration system 106 utilizes the call recalibrationmachine learningmodel 306 togenerate the variant call classifications308asaset of threevariant call classifications (whose probabilities sum to 1). In particu- lar, the call recalibration machine learning model 306 may generate the reference probability 310 as 0.1, the differing genotype probability 312 as 0.2, and the correct variant probability 314 as 0.7. Based on the reference probability 310, the differing genotype probability 312, and the correct variant probability 314 in such an exam- ple, the call recalibration system 106 generates the up- dated genotype likelihoods 322 by updating GT=0 / 0 using the reference probability 310, updating GT=1 / 2 using the correct variant probability 314, and updating othergenotypepositions in aPLfieldusingacombination of information from the call recalibration machine learn- ingmodel 306 and a call generationmodel. To use such a combination, in someembodiments, thecall recalibration system 106 combines (e.g., sums) the probabilities of all of the alternative genotypes (as determined by the call generation model) and scales the combination to match the differing genotype probability 312.

[0103] As illustrated in FIG. 3B, the call recalibration system106 generates genotype likelihoods by determin- ing a normalized PL score for different genotypes (GT). According to the normalized scale of a PL score, a relatively lower score (e.g., PL 0) for a genotype repre- sents a relatively higher likelihood of the genotype being present at a genomic coordinate; and a relatively higher score (e.g., PL 101) for the genotype represents a rela- tively lower likelihood of the genotype being present at the genomic coordinate. For example, the call recalibra- tion system 106 determines a PL score of 111 for the 0 / 0 genotype, a PL score of 52 for the 0 / 1 genotype, a PL score of 49 for the 0 / 2 genotype, a PL score of 42 for the 1 / 1 genotype, a PL score of 0 for the 1 / 2 genotype, and a PL score of 30 for the 2 / 2 genotype. Accordingly, in FIG. 3B, the PL score of 0 indicates a genotype with the 5 10 15 20 25 30 35 40 45 50 55 19 35 EP 4 693 301 A2 36 highest likelihood or the selected genotype (e.g., a 1 / 2 genotype) and the PL score of 111 represents the lowest likelihood (e.g., a 0 / 0 genotype). Thus, in the example, the order of genotypes according to likelihood (frommost likely to least likely) is as follows: 1 / 2, 2 / 2, 1 / 1, 0 / 2, 0 / 1, and 0 / 0.

[0104] In somecases, the call recalibration system106 generates the updated genotype likelihoods 322 as a ranking of a plurality of alleles identified via the call generation model (without utilizing the call recalibration machine learning model 306). In other cases, the call recalibration system 106 utilizes a specialized version of the call recalibration machine learning model 306 that is trained togenerate theupdatedgenotype likelihoods322 based on the variant call classifications 308.

[0105] As further illustrated in FIG. 3B, the call recali- brationsystem106generatesorupdatesavariant call file 324. The call recalibration system 106 can generate the variant call file 324 to include the updatedVCFfields 316, including a base call quality metric, a genotype quality metric, and / or updated genotype likelihoods. As men- tioned, in some cases, the call recalibration system 106 updates only certain fields while other fields, such as a genotype (GT) field remain unchanged based on the multiallelic analysis of FIGS. 3A‑3B. For instance, the call recalibration system 106 updates the genotype qual- ity field and the based call quality field.

[0106] For other data fields such as normalized PHRED-scale likelihoods (PL) for genotypes and poster- ior genotype probability (GP), the call recalibration sys- tem 106 either: (i) maintains the field as-is, (ii) removes the field, or (iii) only updates fields to reflect GQ for the called genotype and Class 0 output 0 / 0. In some cases, the call recalibration system 106 maintains the relative probabilities of other genotypeswith respect to the called genotype to ensure consistent updates and that the called genotype is highest. By updating only the values for 0 / 0 and 1 / 2, the call recalibration system 106 main- tains distances of other genotypes from the called geno- type.

[0107] Within the variant call file 324, the call recalibra- tion system 106 can include or update one or more final nucleotide base calls (e.g., variant nucleotide base calls) associated with the multiallelic genomic coordinate 302, as determined based on the updated VCF fields 316. Indeed, to generate a final nucleotide base call for the multiallelic genomic coordinate 302, the call recalibration system 106 can predict two nucleotide bases from three or more candidate alleles at the multiallelic genomic coordinate (e.g., according to their respective probabil- ities).

[0108] As mentioned, in certain described embodi- ments, the call recalibration system 106 generates final nucleotide base calls (e.g., variant calls) for genomic coordinates within a haploid nucleotide sequence from a genomic sample. In particular, the call recalibration system 106 determines a haploid genotype for a haploid coordinate of a sample nucleotide sequence and further determines whether the haploid genotype is a variant. FIGS. 4A‑4B illustrate generating a final nucleotide base call for a haploid genomic coordinate in accordance with one or more embodiments. For example, FIG. 4A illus- trates the call recalibration system 106 generating a final nucleotide base call utilizing a call recalibration machine learning model in accordance with one or more embodi- ments. Thereafter, FIG. 4B illustrates a process for the call recalibration system 106 training, tuning, testing, and / or applying a call recalibration machine learning model to generate nucleotide base calls for haploid co- ordinates in accordance with one or more embodiments.

[0109] As illustrated in FIG. 4A, the call recalibration system106 identifiesahaploidnucleotidesequence402. In particular, the call recalibration system 106 identifies the haploid nucleotide sequence 402 as a region of a sample nucleotide sequence that only includes haploids (as opposed to diploids). For example, the call recalibra- tion system 106 determines, via a call generation model, that a region of a sample nucleotide sequence is located on a haploid sex chromosome (e.g., chr:Y). In some cases, the call recalibration system 106 determines or identifies the haploid nucleotide sequence 402 by deter- mining nucleotide base calls for nucleotide reads corre- sponding to the haploid nucleotide sequence 402 and aligning the nucleotide reads with a reference genome. While thehaploidnucleotidesequence402 isdetermined from nucleotide base calls for nucleotide reads, for sim- plicity of illustration, FIG. 4A depicts the nucleobases for the underlying haploid nucleotide sequence 402 at given genomic coordinates. This disclosure describes the pro- cess of determining nucleotide base calls for nucleotide reads and corresponding sequencingmetrics below with respect to FIGS. 6A‑6B. As shown in FIG. 4A, the call recalibration system106 identifies the haploid nucleotide sequence 402 to include genomic coordinates 1 through 4 based on nucleotide reads, each with a single nucleo- tide base: 1. A 2. A 3. T 4. G.

[0110] As further illustrated in FIG. 4A, the call recali- bration system 106 determines sequencing metrics 404 for thehaploid nucleotide sequence402. Inparticular, the call recalibration system 106 determines read-based sequencing metrics, call model generated sequencing metrics, and / or externally sourced sequencing metrics associated with a particular genomic coordinate within the haploid nucleotide sequence 402. Additional detail regarding determining sequencing metrics is provided below with reference to FIGS. 6A‑6C.

[0111] Based on the sequencing metrics 404, the call recalibration system 106 utilizes a call recalibration ma- chine learning model 406 (e.g., the call recalibration machine learning model 306) to generate, for a genomic coordinatewithin the haploid nucleotide sequence 402, a first genotype probability 408 and a second genotype probability 410basedon thesequencingmetrics404.For instance, the call recalibration system 106 generates the first genotype probability 408 indicating a probability that the genomic coordinate exhibits a first genotype (e.g., a 5 10 15 20 25 30 35 40 45 50 55 20 37 EP 4 693 301 A2 38 haploid reference genotype) and generates the second genotype probability 410 indicating a probability that the genomic coordinate exhibits a second genotype (e.g., a haploid alternate genotype). As used herein, in some cases, the first genotype probability 408 and the second genotypeprobability 410are examples of types of variant call classifications.

[0112] In somecases, the call recalibration system106 generates the first genotype probability 408 and the second genotype probability 410 by converting inputs and / or outputs of the call recalibration machine learning model 406 to adapt the model to haploid scenarios. For example, in some cases, the call recalibration system 106 converts certain sequencing metrics or features as inputs of the call recalibration machine learning model 406 from haploid inputs to diploid inputs. More specifi- cally, the call recalibration system 106 converts a haploid reference genotype call generated by a call generation model to a diploid homozygous reference genotype call as an input for the call recalibration machine learning model 406 (e.g., converts a haploid 0 VC GT to a diploid 0 / 0 GT as an input). In addition, the call recalibration system 106 converts a haploid alternate genotype call generatedby the call generationmodel to adiploid homo- zygous alternate genotype call as an input for the call recalibrationmachine learningmodel 406 (e.g., converts ahaploid 1VCGT toadiploid 1 / 1GTasan input). Further, in some cases, the call recalibration system 106 ex- cludes, removes, or ignores a heterozygous genotype call generatedby thecall generationmodel asan input for the call recalibration machine learning model 406.

[0113] In one or more embodiments, the call recalibra- tion system106 also (or alternatively) converts outputs of the call recalibration machine learning model 406 from diploid outputs to haploid outputs. For instance, in some cases, the call recalibration system 106 converts from diploid outputs to haploid outputs utilizing a softmax model or layer (e.g., as a layerwithin the call recalibration machine learning model 406). In some cases, the call recalibration system 106 utilizes the softmax layer to modify confidence scores of diploid genotypes to simu- late (or transform into) probabilities of haploid genotypes for the genomic coordinate. For instance, the call recali- bration system 106 utilizes a softmax layer to modify a homozygous reference confidence score of a homozy- gous reference genotype at the genomic coordinate to generate a haploid reference probability of a reference genotype at the genomic coordinate. Further the call recalibration system 106 utilizes a softmax layer to mod- ify a homozygous alternate confidence score of a homo- zygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of an alternate genotype at the genomic coordinate.

[0114] In one or more embodiments, the call recalibra- tion system 106 prunes or removes one of the three model outputs. For instance,when determining a nucleo- tide base call for a haploid genomic coordinate, the call recalibration system 106 removes a confidence score that a genotype of the genomic coordinate is heterozy- gous (or that a het / homerror exists at the coordinate) and does not input such a confidence score into the softmax layer. Based on a first confidence score that the genomic coordinate exhibits a haploid reference genotype and a second confidence score (or a third confidence score) that the genomic coordinate exhibits a haploid alternate genotype, the call recalibration system 106 uses a soft- max layer to normalize these remaining two confidence scores (so that they sum to 1) to generate the first genotype probability 408 and the second genotype prob- ability 410. Thus, the call recalibration system 106 gen- erates the first genotype probability 408 and the second genotype probability 410 for haploids based on corre- sponding diploid probabilities.

[0115] As shown in FIG. 4A, the first genotype prob- ability 408 indicates an 80%probability that the genotype of haploid coordinate 3 is 0 or, in other words, constitutes a haploid reference (represented as 0@3→80%). Simi- larly, the second genotype probability 410 indicates a 20%probability that the genotype of haploid coordinate 3 is 1 or, in other words, constitutes a haploid alternate (represented as 1@3→20%). Additional detail regarding converting the outputs is provided below with reference to FIG. 4B.

[0116] As further illustrated in FIG. 4A, the call recali- brationsystem106generatesorupdatesavariant call file 412 based on the first genotype probability 408 (e.g., a haploid-reference-genotype probability) and the second genotype probability 410 (e.g., a haploid-alternate-gen- otype probability). For example, the call recalibration system 106 updates the variant call file 412 to reflect or indicate a final nucleotide base call 414 associated with the haploid genomic coordinate based on the first genotype probability 408 and the second genotype prob- ability 410.

[0117] In certain embodiments, the call recalibration system 106 determines the final nucleotide base call 414 to indicateahaploid genotype for thegenomic coordinate based on comparing the first genotype probability 408 and the second genotype probability 410 and selecting a highest genotype from among the first genotype prob- ability 408 and the second genotype probability 410. In some cases, the call recalibration system 106 updates additional fields associated with the variant call file 412, such as a base call quality field, a genotype quality field, and / or a genotype field based on comparing the first genotype probability 408 and the second genotype prob- ability 410.

[0118] Basedondetermining that thesecondgenotype probability 410 is highest (i.e., exceeds the first genotype probability 408) or that the nucleotide base call (or the variant call) is most likely a true positive, for instance, the call recalibration system 106 determines a haploid alter- nate genotype for the genomic coordinate. When the second genotype probability 410 (e.g., a haploid-alter- nate-genotype probability) exceeds the first genotype probability 408 (e.g., a haploid-reference-type probabil- 5 10 15 20 25 30 35 40 45 50 55 21 39 EP 4 693 301 A2 40 ity), for example, the call recalibration system 106 further determinesamodifiedbase call qualitymetric, amodified genotype metric, and / or a modified genotype quality metric (to include within the variant call file 412). In some cases, the above modifies the genotype quality metric to reflect a likelihood that the nucleotide base call or the variant call is incorrect (in PHRED format) with the ex- isting genotype.

[0119] Basedondetermining that thesecondgenotype probability 410 is not highest (i.e., that the first genotype probability 408 exceeds the second genotype probability 410), the call recalibration system 106 determines a haploid reference genotype for the genomic coordinate. When the first genotype probability 408 (e.g., a haploid- reference-type probability) exceeds the second geno- type probability 410 (e.g., a haploid-alternate-genotype probability), for example, the call recalibration system 106 further determines a modified genotype quality me- tric and / or a modified base call quality metric. For in- stance, if the call recalibration system 106 predicts a reference genotype call, the call recalibration system 106 keeps the called genotype and sets the score to the value output by the call recalibration machine learn- ing model. If, however, the call recalibration system 106 uses the call recalibration machine learning model to determine a modified base call quality metric for the genotype call at a haploid genomic coordinate, the call recalibration system 106 changes a quality field for the genotype call to include the modified base call quality metric. Alternatively, in some cases, when a base call quality metric falls below a quality threshold, the call recalibration system 106 can drop the nucleotide base call or at least not include the nucleotide base call for the genomic coordinate in a variant call file.

[0120] In some embodiments, the call recalibration system 106 generates a final nucleotide base call 414 based on the comparison of the first genotype probability 408 and the second genotype probability 410. As shown, the call recalibration system 106 determines that the first genotype probability 408 is higher than the second gen- otype probability 410 and therefore generates the final nucleotide base call 414 to indicate that the genotype for the specific haploid coordinate (coordinate 3) is most likely a haploid reference genotype (represented as 3→0).

[0121] As illustrated in FIG. 4B, the call recalibration system 106 modifies input and outputs of the call recali- bration machine learningmodel 406 to facilitate generat- ing final nucleotide base calls (e.g., variant calls) for a genomic coordinate of a haploid nucleotide sequence. In some embodiments, the process illustrated in FIG. 4B represents a training and / or tuning of the call recalibra- tion machine learning model 406 to learn parameters for generating nucleotide base calls (e.g., variant calls). In other embodiments, some or all of the process illustrated in FIG. 4B represents application, or inference using, of the call recalibration machine learning model 406.

[0122] As shown, the call recalibration system 106 performs a downsampling 418 of (a subset of) diploid nucleotide reads 420 to simulate haploid nucleotide reads. More specifically, the call recalibration system 106 downsamples (or otherwise modifies) diploid data tomimic or simulate haploid data for training or tuning the call recalibration machine learning model 406. Indeed, because ground truth haploid data is sparse, the call recalibration system 106 cannot rely on diploid data alone to learn robust parameters for generating nucleo- tidebase calls for haploid coordinates. Thus, unlike some existing sequencing systems that cannot generate calls for haploid coordinates (due to the lack of training data), in some embodiments, the call recalibration system 106 adapts to haploid scenarios by simulating haploid data from diploid data.

[0123] For example, the call recalibration system 106 determines (or receives) diploid nucleotide reads 420 via a call generation model 416. Additionally, the call recali- bration system 106 (randomly) selects a subset of the diploid nucleotide reads 420 to use as training or testing data (e.g., a random selection of 50% of the reads). As depicted, the diploid nucleotide reads 420 include reads for four genomic coordinates 1 through 4, as follows: 1) AA 2) AA 3) CC 4) TT. In addition, the call recalibration system 106 determines diploid sequencing metrics 422 from the (subset of the) diploid nucleotide reads 420. In some embodiments, the call recalibration system 106 determines or identifies, based on truth data (e.g., Pre- cisionFDA truth data, Platinum Genomes, or some other high confidence truth set, such as truth sets from the Genome in a Bottle (GIAB), Global Alliance for Genomic Health (GA4GH), or Telomere to Telomere Consortium) and / or the diploid sequencing metrics 422, one or more genomic coordinates of the diploid nucleotide reads 420 that exhibit homozygous genotypes, such as a homo- zygous reference genotype or a homozygous alternate genotype.

[0124] As further illustrated in FIG. 4B, the call recali- bration system 106 generates (or simulates) haploid sequencing metrics 424 from the diploid sequencing metrics 422 via the downsampling 418. For instance, the call recalibration system 106 modifies the homozy- gous genotypes of the diploid nucleotide reads 420 to simulate haploid genotypes. Specifically, the call recali- bration system 106 transforms a homozygous reference genotype to a haploid reference genotype (represented as 0 / 0→0) and transforms a homozygous alternate gen- otype to a haploid alternate genotype (represented as 1 / 1→1). Thecall recalibration system106 further selects, as the haploid sequencing metrics 424, the sequencing metrics of the diploid nucleotide reads 420 used to simu- late haploid nucleotide reads. Based on these haploid sequencing metrics 424, the call recalibration system 106 can train and / or test the call recalibration machine learning model 406 to accurately generate final nucleo- tidebasecalls (e.g., variant calls) for haploid coordinates.

[0125] Indeed, in training, testing, and / or inference, the call recalibration system106 utilizes the call recalibration 5 10 15 20 25 30 35 40 45 50 55 22 41 EP 4 693 301 A2 42 machine learning model 406 to generate final nucleotide base calls based on sequencing metrics, such as the haploid sequencingmetrics424.Asmentionedabove,as part of generating a final nucleotide base call 432 via the call recalibration machine learning model 406 (either for training, testing, or inference), the call recalibration sys- tem 106 modifies outputs of the call recalibration ma- chine learning model 406. For example, the call recali- bration system 106 modifies confidence scores gener- ated by one or more classifier layers(s) 426 of the call recalibration machine learning model 406.

[0126] In some embodiments, the call recalibration system 106 does not simulate haploid data from diploid data during an inference process (as opposed to a train- ing or testing process). Indeed, when applying the call recalibration machine learning model 406 to generate predictions, the call recalibration system 106 may only modify the inputs and outputs of the call recalibration machine learningmodel 406once themodel is trained for haploid scenarios with simulated haploid data. When using the call recalibration machine learning model 406, for instance, the call recalibration system 106 inputs a sequencing metric indicating the data is haploid during the inference process.

[0127] Specifically, and as depicted in FIG. 4B, the call recalibration system 106 generates three separate con- fidence scores, each representing confidence levels of a different genotype belonging to a given genomic coordi- nate: (i) a first confidence score for a haploid reference genotype (z0), (ii) a second confidence score for a het- erozygous genotype (z1), and (iii) a third confidence score for a haploid alternate genotype (z2). The call recalibration system 106 further prunes, ignores, or re- moves the second confidence score (z1) because het- erozygous genotypes cannot be simulated from diploid data to haploid data and (during implementation) haploid coordinates do not exhibit heterozygous genotypes. In some embodiments, the call recalibration system 106 utilizes a softmax model 428 as another layer in the call recalibration machine learning model 406 to generate final probabilities from the confidence scores.

[0128] In particular, and as shown in FIG. 4B, the call recalibration system 106 utilizes the softmax model 428 to generate two genotype probabilities (e.g., the first genotypeprobability 408of a haploid referencegenotype and the second genotype probability 410 of a haploid alternategenotype) from themodifiedconfidencescores. To elaborate, after ignoring or discarding the second confidence score (z1), the call recalibration system 106 utilizes the softmax model 428 to normalize across the first confidence score and the third confidence score (so they sum to 1). The call recalibration system 106 further generates the first genotype probability 408 represented by σ0 = p(0) and the second genotype probability 410 represented by σ1 = p(1). As described, each probability score represents a probability of a respective haploid genotype at a genomic coordinate.

[0129] Based on the probability scores, the call recali- bration system 106 further generates the variant call file 430 (e.g., the variant call file 412) including the final nucleotide base call 432 (e.g., the final nucleotide base call 414). For example, the call recalibration system 106 determines the final nucleotidebase call 432 from the two genotype probabilities. As shown, for example, the final nucleotide base call 432 is a haploid A for the given genomic coordinate. But the final nucleotide base call 432 could bedifferent predictednucleotidebases in other embodiments. Additional detail regarding generating a variant call file is provided throughout this disclosure.

[0130] As mentioned above, in certain described em- bodiments, the call recalibration system 106 generates final nucleotide base calls (e.g., variant calls) for homo- zygous reference genomic coordinates (as initially pre- dicted by a call generation model). In particular, the call recalibration system 106 generates final nucleotide base calls for genomic coordinates of a sample nucleotide sequence determined (or would be determined) by a call generation model to exhibit homozygous reference gen- otypes. FIG. 5 illustrates generating a variant call for a genomic coordinate that would or could have been in- correctly identified as a homozygous reference genotype in accordance with one or more embodiments.

[0131] As illustrated in FIG. 5, the call recalibration system 106 utilizes a call generation model 502 to gen- erate initial nucleotide base calls for a sample nucleotide sequence 504. In particular, the call recalibration system 106 generates nucleotide base calls that indicate alleles or genotypes associated with particular genomic coordi- nates. As shown, the call generation model 502 deter- mines genotypes for the sample nucleotide sequence 504 of coordinates 1 through 4 as follows: 1) 0 / 1 2) 1 / 1 3) 0 / 0 4) 0 / 1. Additionally, the call recalibration system 106 identifies or determines genomic coordinates within the sample nucleotide sequence 504 that exhibit homozy- gous reference genotypes as determined by the call generationmodel 502. In the illustrated example, the call recalibration system 106 identifies coordinate 3 as a homozygous reference coordinate. By contrast, in some embodiments, the call recalibration system 106 does not generatean initial nucleotidebase call indicatingahomo- zygous reference genotype, but rather determines se- quencing metrics 506 for nucleotide reads covering the genomic coordinate consistent with a homozygous re- ference genotype.

[0132] In many cases, existing sequencing systems ignored homozygous reference coordinates, such as coordinate 3, and treated them as true negative variant calls that were not necessary for further processing. However, such treatment relies on the accuracy of the call generation model 502 making the proper nucleotide base call initially, which is not always the case. Indeed, the call generation model 502 can generate large num- bers of false negative variant calls in some scenarios. Thus, the call recalibration system 106 recovers some of these false negative variant calls by not ignoring genomic coordinates that were initially (or would have been) iden- 5 10 15 20 25 30 35 40 45 50 55 23 43 EP 4 693 301 A2 44 tified as homozygous reference genotypes and forcing further analysis at these loci (e.g., to consequently up- date or modify their determined genotypes).

[0133] Specifically, as illustrated in FIG. 5, the call recalibration system 106 extracts or determines sequen- cing metrics 506 for the homozygous reference coordi- nate 3. For example, the call recalibration system 106 determines read-based sequencing metric 508, exter- nally sourced sequencing metrics 510, and call model generated sequencing metrics 512 for coordinate 3. Ad- ditional detail regarding determining or extracting se- quencing metrics is provided below with reference to FIGS. 6A‑6C.

[0134] As further illustrated in FIG. 5, the call recalibra- tion system 106 utilizes a call recalibration machine learning model 514 (e.g., the call recalibration machine learning model 306 or 406) to generate one or more variant call classifications 516 from the sequencing me- trics 506. To elaborate, the call recalibration system 106 generates variant call classifications 516 indicating (or defining a level of) an accuracy of identifying a variant at the genomic coordinate (e.g., coordinate 3).

[0135] The following paragraphs describe examples of the variant call classifications 516. As an example variant call classification, the call recalibration system 106 gen- erates a false positive classification utilizing the call recalibration machine learning model 514. For example, the call recalibration system 106 generates a false posi- tive classification that indicates a probability that a nu- cleotide base call (e.g., genotype call) is a false positive variant, or that the nucleotide base call indicates a variant where no variant actually exists within the sample nu- cleotide sequence 504. The call recalibration system106 generates the false positive classification from one or more of the sequencingmetrics 506 considered together by the call recalibration machine learning model 514.

[0136] In certain implementations, the call recalibra- tion system 106 also (or alternatively) generates a gen- otype error classification (or a heterozygous genotype classification) as part of the variant call classifications 516. More specifically, the call recalibration system 106 determines, utilizing the call recalibrationmachine learn- ing model 514, a probability that a genotype associated with a nucleotide base call is incorrect or that a hetero- zygous genotype exists (e.g., for coordinate 3). For in- stance, the call recalibration system 106 determines a probability that a het / hom error exists at coordinate 3, where the nucleotide base call may indicate a hetero- zygous genotype (e.g., 0 / 1) within the sample nucleotide sequence 504 and the genotype is actually homozygous alternate (e.g., 1 / 1) with respect to the reference gen- ome. Conversely, the call recalibration system 106 de- termines a probability of determining that a genotype for coordinate 3 is homozygous alternate (e.g., 1 / 1) when, in fact, the nucleotide base(s) are heterozygous with re- spect to the reference genome (e.g., 0 / 1).

[0137] In one or more embodiments, the call recalibra- tion system 106 also (or alternatively) generates, as part of the variant call classifications 516, a true positive classification (or a homozygous alternate classification) for coordinate 3. In particular, the call recalibration sys- tem 106 determines, utilizing the call recalibration ma- chine learning model 514, a probability that a nucleotide base call for coordinate 3 is a true positive variant call, or that the nucleotide base call indicates a true variant where a variant does indeed exist in relation to a refer- ence genome, or that a homozygous alternate genotype exists at the genomic coordinate.

[0138] As further illustrated in FIG. 5, the call recalibra- tion system 106 generates or updates a variant call file 518 to indicate a variant call 520. More specifically, the call recalibration system 106 generates the variant call 520 based on the variant call classifications 516 to in- dicate whether there is a variant at coordinate 3. In some cases, the call recalibration system 106 updates one or more of a call quality field, a genotype field, or a genotype quality field corresponding to the variant call file 518 based on the one or more variant call classifications 516. The call quality field, the genotype field, and / or the genotype quality field can indicate an updated variant call as the variant call 520. As shown, the variant call 520 indicates a variant at coordinate 3, changing the initial nucleotide base call for coordinate 3 from indicating a homozygous reference genotype (0 / 0) to indicate a het- erozygous genotype (0 / 1). In other examples, the call recalibration system 106 does not change the initial nucleotide base call for coordinate 3 or changes the initial nucleotide base call to a different genotype, such as a homozygous alternate genotype (1 / 1).

[0139] In one or more embodiments, the call recalibra- tion system 106 determines a genotype for the indicated genomic coordinate (e.g., coordinate 3) based on com- paring the probabilities of the variant call classifications. For example, the call recalibration system 106 deter- mines a homozygous alternate genotype based on de- termining that a true positive classification (or a homo- zygous alternate classification) has a highest probability from among the one or more variant call classifications. Specifically, the call recalibration system106updates the genotype quality field while also updating the genotype field (e.g. to 1 / 1) and the PL field.

[0140] Alternatively, the call recalibration system 106 determines a heterozygous genotype based on deter- mining that a genotype error classification (e.g., a hetero- zygous genotype classification) has the highest prob- ability from among the one ormore variant call classifica- tions. Specifically, the call recalibration system 106 up- dates the genotype quality field while also updating the genotype field (e.g., to 0 / 1) and the PL field.

[0141] Alternatively still, the call recalibration system 106 determines a homozygous reference genotype based on determining that neither the true positive clas- sification (e.g., the homozygous alternate classification) nor the genotype error classification (e.g., the heterozy- gous genotype) has the highest probability from among the one or more variant call classifications. In some 5 10 15 20 25 30 35 40 45 50 55 24 45 EP 4 693 301 A2 46 cases, the call recalibration system 106 removes or dis- cards a record of comparing probabilities for variant classifications when both the call generation model 502 and the call recalibration machine learning model 514 determine that the genomic coordinate has a homo- zygous reference genotype.

[0142] In one or more embodiments, updating variant calls for homozygous reference coordinates provides or improves forced genotype functionality (e.g., for query of a genotype andgenotype probabilities at a specific geno- mic coordinate). To elaborate, the call recalibration sys- tem 106 can determine genotypes of genomic coordi- nates that initially (e.g., as indicatedby thecall generation model 502) fail to satisfy a variant quality threshold. Indeed, the call recalibration system 106 can output genotypes to the variant call file 518 even if the variant quality of the genomic coordinate falls below a threshold typically required to identify a structural variant or other difficult-to-determine variants.

[0143] As mentioned above, in certain described em- bodiments, the call recalibration system 106 determines or extracts sequencing metrics for nucleotide base calls at particular genomic coordinates. In particular, the call recalibration system106 determines sequencingmetrics such as read-based sequencing metrics, externally sourced sequencing metrics, and call model generated sequencing metrics from calls corresponding to nucleo- tide reads from a sample nucleotide sequence. FIGS. 6A‑6C illustrate determining sequencing metrics in ac- cordance with one or more embodiments. Specifically, FIG. 6A illustrates determining read-based sequencing metrics while FIG. 6B illustrates determining call model generated sequencing metrics, and FIG. 6C illustrates determining externally sourced sequencing metrics.

[0144] As illustrated in FIG. 6A, the call recalibration system 106 accesses, retrieves, obtains, determines, or generates nucleotide reads 602. In particular, the call recalibration system 106 determines the nucleotide reads602utilizing the sequencingdevice 114comprising nucleotide base calls for regions from a sample nucleo- tide sequence (e.g., sample genome). For example, the call recalibration system 106 generates a plurality of nucleotide reads 602 utilizing sequencing-by-synthesis (SBS) techniques and / or Sanger sequencing techniques to determine nucleotide base calls for oligonucleotide clusters from wells in a flow cell and / or via fluorescent tagging. More specifically, the call recalibration system 106 utilizes cluster generation and SBS chemistry to sequence millions or billions of clusters in a flow cell. During SBS chemistry, for each cluster, the call recali- bration system 106 stores nucleotide base calls from the nucleotide reads 602 for every cycle of sequencing via real-time analysis (RTA) software.

[0145] As further illustrated in FIG. 6A, in some embo- diments, the call recalibration system 106 performs read processing and mapping 604. For example, the call recalibration system 106 utilizes RTA software to store base call data in the form of individual base call data files (or BCLs). In some cases, the call recalibration system 106 further converts theBCLfiles into sequencedata608 (e.g., viaBCL toFASTQconversion), as illustrated inFIG. 6B.Asshown inFIG.6A, thecall recalibrationsystem106 generates multiple-read coverage (e.g., read pileups) that include multiple nucleotide reads 602 or nucleotide base calls corresponding to a single genomic coordinate.

[0146] In particular, in certain embodiments, the call recalibration system 106 aligns nucleotide reads with a reference genome or receives information pertaining to the read alignment. Specifically, the call recalibration system 106 determines which nucleotide base(s) of a given read align with which genomic coordinate of a reference sequence (or receives information indicating alignment). Different reads have different lengths and include different nucleotide bases. Accordingly, in some cases, the call recalibration system 106 analyzes each nucleotide of each read to determine (or receives infor- mation indicating) where the read "fits" in relation to a reference sequence-e.g., where the bases within the read align with bases in the reference. In some cases, the call recalibration system 106 aligns many reads at a single genomic coordinate, thus resulting a read pileup.

[0147] In certain embodiments, the call recalibration system 106 performs additional statistical tests to deter- mine or detect differences between metrics associated with a reference nucleotide sequence and metrics asso- ciated with alternative supporting nucleotide reads. Through these statistical tests, the call recalibration sys- tem 106 re-engineers raw sequencing metrics to deter- mine read-based sequencing metrics 606. In some cases, the call recalibration system 106 determines or extracts raw sequencingmetrics that includeoneormore of (i) alignment metrics for quantifying alignment of sam- plenucleotidesequenceswithgenomiccoordinatesofan example nucleotide sequence (e.g., a reference genome or a nucleotide sequence from an ancestral haplotype), (ii) depthmetrics for quantifying depth of nucleotide base calls for sample nucleotide sequences at genomic co- ordinates of the example nucleotide sequence, or (iii) call-quality metrics for quantifying quality of nucleotide base calls for sample nucleotide sequences at genomic coordinates of the example nucleotide sequence. For instance, the call recalibration system 106 determines mapping-quality metrics (e.g., the MAPQ metrics indi- cated in FIG. 6A), soft-clipping metrics, or other align- ment metrics that measure an alignment of sample se- quences with a reference genome. As another example, the call recalibration system 106 determines forward- reverse-depth metrics (or other such depth metrics) or callability metrics for variant nucleotide base calls (or other such call-quality metrics).

[0148] As just mentioned, in some embodiments, the call recalibration system 106 re-engineers the raw se- quencing metrics to generate the read-based sequen- cingmetrics 606 that aremore informative for comparing metrics associatedwith a referencenucleotide sequence with metrics associated with various supporting alterna- 5 10 15 20 25 30 35 40 45 50 55 25 47 EP 4 693 301 A2 48 tive nucleotide reads. For example, the call recalibration system 106 determines various metrics for a sample sequence in relation to a reference sequence and further determines various metrics for the sample sequence in relation to alternative supporting sequences. In addition, the call recalibration system 106 performs comparative analyses between metrics associated with the reference sequence and themetrics associatedwith the alternative supporting reads.

[0149] For instance, the call recalibration system 106 compares how nucleotide bases of a sample nucleotide sequence (e.g., sample genome) map to a reference sequence with how the nucleotide bases map to various alternative supporting reads. In some cases, the call recalibration system 106 determines mapping qualities associated with the reference sequence to compare with mapping qualities associated with alternative supporting reads. For example, the call recalibration system 106 determines mapping quality statistics reflecting differ- ences in the distribution of reads supporting a reference sequence versus reads supporting alternative alleles.

[0150] In these or other cases, the call recalibration system 106 determines mismatch counts between the sample sequence and the reference sequence and be- tween the reference sequence and alternative support- ing reads. The call recalibration system 106 further com- pares the mismatch counts to determine a comparative- mismatch-count metric. Further, the call recalibration system 106 determines soft-clippingmetrics for the sam- ple sequence in relation to the reference sequence and further determines soft-clipping metrics in relation to alternative supporting reads. The call recalibration sys- tem106 also compares the soft clippingmetrics between the reference sequence and the alternative supporting reads to generate a comparative-soft-clipping metric. Further still, the call recalibration system 106 compares base call quality metrics in relation to the reference sequence and alternative supporting reads and / or com- pares query positions of the sample sequence in relation to the reference sequence with those in relation to alter- native supporting reads.

[0151] As further illustrated in FIG. 6A, the call recali- bration system106 utilizes the comparisons and / or other statistical tests to generate the read-based sequencing metrics606, including: i) a comparative-mapping-quality- distribution metric indicating a mapping quality distribu- tion comparing mapping qualities in relation to the refer- ence sequence andmapping qualities in relation to alter- native supporting reads, ii) a comparative-secondary- mapping-alignment metric indicating a comparison be- tween secondary mapping in relation to bases in the reference sequence and bases in alternative supporting reads, iii) a comparative-mismatch-count metric indicat- ing a comparisonbetweenmismatchednucleotide bases in relation to the reference sequence and mismatched bases in relation to alternative supporting reads, iv) a comparative-soft-clippingmetric indicatinga comparison between soft-clipping metrics in relation to the reference sequence and soft-clipping metrics in relation to alter- native supporting reads, v) one or more comparative- read-depth metrics indicating comparisons between read depths of the nucleotide reads 602 and one or more average read depths (e.g., local average read depths at a particular genomic coordinate and global average read depths across a number genomic coordinates in a re- gion), vi) one or more comparative-base-quality metric indicating comparisons between base qualities in rela- tion to the reference sequence and base qualities in relation to alternative supporting reads (e.g., for overall base quality, early base quality, and late base quality in the nucleotide reads 602), vii) a comparative-query-posi- tion metric indicating a comparison between query posi- tions in relation to the reference sequence and query positions in relation to alternative supporting reads, viii) one or more contextual-information metrics indicating homopolymers and periodicity of nucleotide base calls, ix) a strand-bias metric indicating a strand bias asso- ciated with one or more of the nucleotide reads 602, and x) a read-direction-bias metric indicating a read direction bias associated with the nucleotide reads 602. In some cases, the call recalibration system 106 generates or re- engineers additional or alternative read-based sequen- cing metrics as part of the read-based sequencing me- trics 606.

[0152] In addition to the read-based sequencing me- trics 606, as illustrated in FIG. 6B, the call recalibration system 106 generates call model generated sequencing metrics 612. In particular, the call recalibration system 106 generates the call model generated sequencing metrics from sequence data 608 utilizing a call genera- tion model 610. For example, the call recalibration sys- tem 106 extracts or determines sequence data 608 based on the read processing and mapping 604 de- scribed in relation to FIG. 6A. In some cases, the call recalibration system 106 generates the sequence data 608 as part of one or more digital files, such as BCL and FASTQ files.

[0153] To generate such files, in some embodiments, the sequencing device 114 (or the call recalibration sys- tem106) utilizes cluster generationandSBSchemistry to sequence millions or billions of clusters in a flow cell. During SBS chemistry, for each cluster, the sequencing device 114 (or the call recalibration system 106) stores nucleotide base calls from the nucleotide reads 602 for every cycle of sequencing via real-time analysis (RTA) software. The sequencing device 114 (or the call recali- bration system 106) utilizes RTA software to further store base call data in the form of individual base call data files (or BCLs). In some cases, the sequencing device 114 (or the call recalibration system 106) further converts the BCLfiles into sequencedata608 (e.g., viaBCL toFASTQ conversion). For instance, the sequencing device 114 (or the call recalibration system106) generates aFASTQfile from the nucleotide reads 602, where the FASTQ file includes sequence data 608.

[0154] In somecases, the call recalibration system106 5 10 15 20 25 30 35 40 45 50 55 26 49 EP 4 693 301 A2 50 generates the sequence data 608 for each cluster that passes an initial quality filter from a sample sequence. For example, the call recalibration system106 generates entries for each cluster, where each entry includes four lines (or four items of sequence data): i) a sequence identifier with information about the sequencing run and the cluster, ii) nucleotide base calls that make up the sequence (e.g., a sequence of A, C, T, G, and / or N calls), iii) a separator (e.g., a "+" sign), and iv) base call quality metrics indicating probabilities of correctness for the nucleotide base calls (Phred +33 encoded).

[0155] As further illustrated in FIG. 6B, the call recali- bration system 106 implements, utilizes, or applies the call generation model 610 to process or analyze the sequence data 608. Indeed, in some embodiments, thecall recalibration system106generates thecallmodel generated sequencing metrics 612 by utilizing the call generation model 610 to re-engineer raw sequencing metrics (e.g., raw sequencing metrics within the se- quence data 608). In particular, the call generationmodel 610 includes mapping-and-alignment components to map and align nucleotide base calls from the sequence data 608. In addition, the call generation model 610 includes variant calling components to generate nucleo- tide base calls (e.g., reference-base calls such as variant calls or non-variant calls) from the sequence data 608. In some cases, the call recalibration system 106 extracts the call model generated sequencing metrics 612 that have been generated utilizing the mapping-and-align- ment components and the variant calling components of the call generation model 610.

[0156] To illustrate examples of the call model gener- ated sequencing metrics 612, in some cases, the call recalibration system 106 generates (variant calling me- trics including one or more of: i) a base call quality metric (e.g., DRAGEN QUAL score) indicating a quality score for nucleotide base calls generated via the call genera- tion model 610, ii) a call model generated-foreign-read- detection metric (e.g., foreign read detection (FRD) score) indicating a probability that one or more of the nucleotide reads 602 in a pileup might be foreign reads (e.g., their true location is elsewhere in the reference sequence), iii) a callmodel generated-base-quality-drop- off metric (e.g., base quality dropoff (BQD) score) indi- cating a probability of base quality dropoff based on one or more of strand bias, error position in a thread, or low mean base quality over a subset of the nucleotide reads 602, iv) average read depths, v) indel statistics (e.g., a polymerase chain reaction or "PCR" curve) and / or vi) hidden Markov model (HMM) statistics, vii) a secondary- alignmentmetric indicating a probability that a secondary nucleotide base call is correct, viii) a base-context metric indicating contextual information for nucleotide around a nucleotide base call, iv) a nearby-call metric indicating nearby (e.g., adjacent or within a threshold degree of separation from) a nucleotide base call, x) a joint-detec- tion metric indicating a probability of detecting a joint corresponding to two or more overlapping nucleotide base calls, xii) read-filtering metrics indicating threshold quality metrics or other metrics for filtering out nucleotide base calls with lowmapping quality, base quality, or other quality metrics, or others. The call recalibration system 106 generates the call model generated sequencing metrics 612 from internal (e.g., proprietary, and model- specific) variables that reflect interacting processing paths, corner cases, and difficult predictions / decisions.

[0157] As indicated above, in some cases, the call recalibrationsystem106determinesFRDscoresaccord- ing to the methods described in U.S. Patent Application No. 16 / 280,022 to Eric Jon Ojard, entitled System and Method for Correlated Error Event Mitigation for Variant Calling, which is incorporated by reference herein in its entirety. In certain implementations, the call recalibration system 106 also (or alternatively) determines BQD scores, FRD scores, HMMstatistics, and / or other variant calling metrics according to the methods described in U.S. Patent Application Nos. 17 / 165,828, 15 / 643,381, and 14 / 811,836, which are incorporated by reference herein in their entireties.

[0158] As illustrated in FIG. 6B, the call model gener- ated sequencing metrics 612 include, but are not limited to, variant calling metrics extracted via the variant calling components of the call generationmodel 610. In addition or in the alternative to the examples of the call model generated sequencing metrics 612 described above, in some cases, the call recalibration system 106 generates (e.g., via metric re-engineering) variant calling metrics including one or more of: i) a number of samples in a population, ii) a number of reads processed for generat- ing nucleotide base calls, a number of variants (e.g., SNPs, indels, and MNPs), iii) a number of biallelic sites (e.g., genomic coordinates that contain two observed alleles), iv) a number of multiallelic sites (e.g., a number of sites in a variant call file that contain three or more observed alleles), v) a number of SNPs, vi) numbers of different types of indels (e.g., homozygous insertions, heterozygous insertions, and heterozygous deletions), vii) a total number of heterozygous indels (e.g., insertion + deletion, insertion + SNP, or deletion + SNP), viii) a number of denovoSNPs (e.g., SNPswith denovoquality metrics that satisfy a threshold level), ix) a number of de novo indels (e.g., indels with de novo quality metrics that satisfy a threshold level), x) a number of de novo MNPs (e.g., MNPs with de novo quality metrics that satisfy a threshold level, xi) a number of SNPs in a first chromo- some divided by a number of SNPs in a second chromo- some, xii) a number of SNP transitions, xiii) a number of SNP transversions, xiv) a number of heterozygous var- iants, xv) a number of homozygous variants, xvi) a ratio between the number of heterozygous variants and the number of homozygous variants, xvii) a number of var- iants detected within a dbSNP reference file, and / or xviii) a total number of variants minus the number detected within the dbSNP file.

[0159] Additionally, the call model generated sequen- cing metrics 612 can include mapping-and-alignment 5 10 15 20 25 30 35 40 45 50 55 27 51 EP 4 693 301 A2 52 sequencing metrics extracted via the mapping-and- alignment components of the call generation model 610. For instance, the call recalibration system 106 gen- erates or extracts (e.g., via metric re-engineering) map- ping-and-alignment metrics including one or more of: i) a number of total input reads, ii) a number of duplicate marked reads, iii) a number of duplicate marked and mate reads removed, iv) a number of unique reads, v) a number of reads with mate sequenced, vi) a number of reads without mate sequenced, vii) indications of reads that fail quality checks, viii) indications of mapped reads, ix) a number of unique andmapped reads, x) a number of unmapped reads, xi) a number of singleton reads (e.g., where the read is mapped but the paired mate could not be read), xii) a number of paired reads, xiii) a number of properly paired reads (e.g., where both reads in a pair are mapped and fall within an acceptable range from each other based on an estimated insert length distribution), xiv) a number of discordant reads (e.g., not properly paired reads), xv) a number of paired reads mapped to different chromosomes, xvi) a number of paired reads mapped to different chromosomes that also have amap- ping-quality metric of 10 or greater, xvii) percentages of reads within indels R1 and R2, xviii) percentages of bases in R1 and R2 that are soft clipped, xix) a number of mismatched bases in R1 and R2, xx) a number of bases with a base quality of at least 30 (e.g., total and / or in R1 or R2), xxi) a number of alignments (e.g., total alignments, secondary alignments, and / or supplemen- tary alignments), xxii) an estimated read length, and xxiii) an estimated sample contamination.

[0160] Turning now to FIG. 6C, the call recalibration system106generates, extracts, or determines externally sourced sequencing metrics 616. In particular, the call recalibration system 106 determines externally sourced sequencing metrics 616 from one or more databases external to the call recalibration system 106, such as a sequencing information database 614 (e.g., the data- base 116). For example, the call recalibration system 106 accesses sequencing metrics that are generic or applicable to sequencing nucleotides generally. In addi- tion, the call recalibration system 106 accesses or de- termines sequencing information about a particular re- ference sequence (e.g., stored within the sequencing information database 614). In some cases, the call re- calibration system 106 determines externally sourced sequencingmetrics 616 including: i) amappability metric indicating an ease or difficult of mapping a particular nucleotide sequence (or a particular nucleotide read or nucleotide base call), ii) a guanine-cytosine-content me- tric indicatingacount (or adropout or amean) of guanine- cytosine content in a reference nucleotide sequence (e.g., reference genome), iii) a replication-timing metric indicating a time required to replicate a particular number of nucleotides from a reference sequence, iv) one or more DNA-structure-metrics indicating DNA structures of a reference sequence (e.g., reference genome), v) a conservation metric indicating a measure of sequence conservation acrossmultiple species (e.g., a measure of change relative to an average), and / or others.

[0161] As mentioned, in certain described embodi- ments, the call recalibration system 106 utilizes a call recalibrationmachine learning model together with a call generation model to generate a nucleotide base call. In particular, the call recalibration system 106 utilizes the call recalibration machine learning model to modify data fields corresponding to a variant call file representing a nucleotide base call. FIG. 7 illustrates generating a nu- cleotide base call bymodifying a variant call file utilizing a call recalibration machine learning model and call gen- eration model in accordance with one or more embodi- ments.

[0162] As illustrated in FIG. 7, the call recalibration system 106 accesses a sequencing information data- base 702 (e.g., the sequencing information database 614), a reference sequence 704, and sequence data 706 (e.g., the sequence data 608) extrapolated from one or more nucleotide reads. Indeed, the call recalibra- tion system 106 performs sequencing-metric extraction 712 to extract or re-engineer sequencing metrics as describedabove in relation toFIGS.6A‑6C.For example, the call recalibration system 106 generates read-based sequencing metrics, externally sourced sequencing me- trics, and call model generated sequencing metrics. In some cases, the call recalibration system 106 utilizes mapping-and-alignment components 708 of a call gen- erationmodel 722 (e.g., the call generationmodel 610) to determine mapping-and-alignment sequencing metrics as described above. In addition, the call recalibration system 106 utilizes variant caller components 710 of the call generation model 722 to generate variant calling metricsasdescribedabove.Further, thecall recalibration system 106 determines read-based sequencing metrics and externally source sequencing metrics as well (e.g., from sequencing information database 702 and / or the reference sequence 704).

[0163] As further illustrated in FIG. 7, the call recalibra- tion system 106 generates variant call classifications 716. More specifically, the call recalibration system 106 utilizes a call recalibrationmachine learningmodel 714 to generate the variant call classifications 716 from the sequencing metrics. For example, the call recalibration machine learning model 714 generates variant call clas- sifications 716 including a false positive classification, a genotype error classification, and a true-positive classi- fication. Specifically, the false positive classification in- dicates a probability that a nucleotide base call (e.g., a variant call) is a false positive. Conversely, a true-positive classification indicates a probability that a nucleotide base call (e.g., a variant call) is a true positive. Addition- ally, a genotype error classification indicates a probability of error associated with a genotype for a nucleotide base call (e.g., a variant call).

[0164] In some cases, the call recalibration machine learning model 714 is an ensemble of gradient boosted trees that processes the sequencing metrics to generate 5 10 15 20 25 30 35 40 45 50 55 28 53 EP 4 693 301 A2 54 the variant call classifications 716. For instance, the call recalibration machine learning model 714 includes a series of weak learners such as non-linear decision trees that are trained in a logistic regression to generate the variant call classifications 716. In some cases, the call recalibration machine learning model 714 includes me- trics within various trees that define how the call recali- bration machine learning model 714 processes the se- quencing metrics to generate the variant call classifica- tions 716. Additional detail regarding the training of the call recalibrationmachine learningmodel 714 is provided below with reference to FIG. 8.

[0165] In certain embodiments, the call recalibration machine learningmodel 714 isadifferent typeofmachine learning model such as a neural network, a support vector machine, or a random forest. For example, in caseswhere the call recalibrationmachine learningmod- el 714 is a neural network, the call recalibration machine learningmodel 714 includesoneormore layerseachwith neurons that make up the layer for processing the se- quencing metrics. In some cases, the call recalibration machine learning model 714 generates the variant call classifications 716 by extracting latent vectors from the sequencingmetrics, passing the latent vectors from layer to layer (or neuron to neuron) to manipulate the vectors until utilizing an output layer (e.g., one or more fully connected layers) to generate the variant call classifica- tions 716 (e.g., as a set of three separate classifications).

[0166] As suggested above, in some embodiments, the call recalibration system 106 can utilize multiple call recalibration machine learning models together. For ex- ample, the call recalibration system 106 utilizes the call recalibration machine learning model 714 to generate a first set of variant call classifications and further utilizes a second call recalibration machine learning model (e.g., with the same or a different architecture) to generate a secondsetof variant call classifications.Forexample, the call recalibration system 106 utilizes two (or more) dif- ferent call recalibration machine learning models in par- allel, each trained with different random seeds (e.g., for different biases to process data differently), resulting in different variant call classifications from the same se- quencing metrics.

[0167] In some embodiments, the call recalibration system 106 further generates a combined set of variant call classifications from the different variant call classifi- cations generated via the different call recalibration ma- chine learning models. In some cases, the call recalibra- tion system 106 generates variant call classifications (e.g., the variant call classifications 716) from a first set and a second set of variant call classifications generated from a first call recalibrationmachine learningmodel and a second call recalibration machine learning model, re- spectively. For instance, the call recalibration system106 determines an average or a weighted combination of the first and second set of variant call classifications to gen- erate the combined variant call classifications for recali- brating a nucleotide base call. In someembodiments, the call recalibration system106determines amean for each variant call classification across each call recalibration machine learning model and renormalizes the mean variant call classification. In other embodiments, the call recalibration system 106 learns linear weights and adapts the weights to minimize overall error or loss for the variant call classifications. In still other embodiments, the call recalibration system 106 weights the variant call classifications for each call recalibration machine learn- ing model based on the inverse of average error across the models.

[0168] In one or more implementations, the call recali- bration system 106 further utilizes a metamodel subse- quent to the call recalibration machine learning models. For example, the call recalibration system 106 utilizes a classification-combiner-machine learning model to com- bine variant call classifications generated from each call recalibration machine learning model-such as by select- ing weights to apply to the variant call classifications generated by each call recalibration machine learning model. Indeed, in some cases, the call recalibration system 106 trains the classification-combiner-machine learningmodel to determine, select, or predict respective weights for call recalibration machine learning models to result in a highest accuracy or a minimized loss.

[0169] When generating the variant call classifications 716, in some embodiments, the call recalibration system 106 generates variant call classifications by utilizing sta- tistics to summarizeamappingquality distribution (e.g., a comparative-mapping-quality-distribution metric) of re- ference supporting reads and alternative supporting reads. For example, the call recalibration system 106 can determine and utilize the mean of the MAPQ for reads supporting an alternative allele as a variant call classification. In these or other embodiments, the call recalibrationmachine learningmodel 714 learns from the data that, when the MAPQ of an alternative allele is low and a depth metric is high relative to other MAPQ and depthmetrics in distributions, a resultant nucleotide base call is more likely to be a false positive variant. Indeed, as the probability of a false positive variant increases, the MAPQ metrics would likely decrease.

[0170] As a further example of generating the variant call classifications 716 utilizing the call recalibration ma- chine learning model 714, in some cases, the call recali- bration system 106 compares a mapping quality (e.g., MAPQ) associated with a nucleotide read (e.g., from the sequencing metrics) with a mapping-quality threshold. For instance, the call recalibration system 106 utilizes a mapping-quality threshold such as a threshold difference between best and secondbest alignment scores. Upon determining that the mapping quality does not satisfy the threshold, the call recalibration system 106 adjusts one ormore of the variant call classifications 716 accordingly. For instance, the call recalibration system 106 increases a probability of genotype error and / or false positive error based on whether the mapping quality satisfies the cor- responding threshold. 5 10 15 20 25 30 35 40 45 50 55 29 55 EP 4 693 301 A2 56

[0171] In addition (or in thealternative) to themethodof generating the variant call classifications 716 just de- scribed, the call recalibration system 106 can (i) utilize an accumulation of statistical analyses over complex functions (depending on the architecture of the call re- calibration machine learning model 714) to determine how to best fit the data (e.g., based on relationship between the various metrics) or (ii) compare other me- trics, such as read depth, base quality, or others asso- ciated with a nucleotide base call (e.g., from the sequen- cing metrics) with corresponding thresholds. The call recalibration system 106 further generates variant call classifications 716 accordingly. For example, in some embodiments, the call recalibration system 106 trains the call recalibration machine learning model 714 to minimize a loss generated from a number of (different types of) sequencing metrics to determine weights and biases that best fit the data (e.g., that result in a reduced or minimized loss) for generating the variant call classi- fications716.Asanother example, upondetermining that a readdepth fails to satisfy a read-depth threshold (e.g., a maximum readdepth corresponding to a particular geno- mic coordinate or generally across all genomic coordi- nates), the call recalibration system 106 increases a genotype error probability and / or increases or decreases a false positive probability and a true-positive probability for a corresponding nucleotide base call.

[0172] In addition to generating the variant call classi- fications 716, as further illustrated in FIG. 7, the call recalibration system 106 performs data field generation 718. More specifically, the call recalibration system 106 generates data fields for a nucleotide base call corre- sponding to a variant call file utilizing the variant caller components 710 of the call generation model 722 and modifies or maintains values for such data fields based the variant call classifications 716. For instance, the call recalibration system 106 modifies various metrics such as quality metrics, mapping metrics, or other metrics associated with the nucleotide base call. In certain em- bodiments, the nucleotide base call is represented or defined by the variant call file 720 which includesmetrics corresponding to the data fields, such as a call-quality metric corresponding to a call-quality field, a genotype metric corresponding to a genotype field, and a geno- type-quality metric corresponding to a genotype-quality field.

[0173] In certain embodiments, the call recalibration system 106 generates (data fields for) a nucleotide base call utilizing the variant caller components 710 together with the variant call classifications 716. For instance, the call recalibration system 106 generates, utilizing the variant caller components 710, data fields for various metrics of a nucleotide base call such as nucleotide(s) included in the call, a call quality (QUAL), a genotype (GT), and a genotype quality (GQ).

[0174] In addition to generating a nucleotide base call via the call generation model 722, the call recalibration system 106 also recalibrates or modifies the nucleotide base call via the variant call classifications 716 from the call recalibration machine learning model 714. In one or more implementations, the call recalibration system 106 modifies the nucleotide base call by modifying or recali- brating data fields for one or more of the metrics asso- ciated with the nucleotide base call (e.g., as included within the variant call file 720). For example, the call recalibration system 106 determines updated values for metrics such as the call quality, the genotype, and the genotype quality from the variant call classifications 716. Indeed, the call recalibration system 106 combines or compares the variant call classifications 716 to recali- brate the corresponding metrics of the nucleotide base call included in the variant call file 720.

[0175] To update or recalibrate the call-quality metric associated with a nucleotide base call, the call recalibra- tion system 106 determines how each of the variant call classifications 716 impact or affect the base call quality metric and adjusts the base call quality metric accord- ingly. For example, the call recalibration system 106 determines that a high probability for a genotype error results in a lower overall genotype quality and possibly a different overall call quality. As another example, the call recalibration system 106 determines that a high prob- ability for a false positive variant results in a lower overall call quality. As yet another example, the call recalibration system 106 determines that a high probability for a true positive variant results in a higher overall (variant) call quality. As a further example, if the call recalibration system 106 determines a high probability for a genotype error (e.g., higher than for the other two variant call classifications 716), then the call recalibration system 106 determines that nucleotide base call is most likely a true variant with the wrong genotype. The call recalibra- tion system 106 accordingly updates the genotype along with the genotype quality and the call quality associated with the nucleotide base call.

[0176] In one or more implementations, the call recali- bration system 106 generates a combination (e.g., a weighted combination or an average) of the variant call classifications 716 to recalibrate the call-qualitymetric. In particular, the call recalibration system 106 weights the false positive classification, the genotype error classifi- cation, and the true-positive classification according to their respective impact on (variant) call quality. In some cases, the call recalibration system 106 weights each variant call classification evenly, while in other cases the call recalibration system 106 determines different weights for each variant call classification. In any event, the call recalibration system 106 determines a weighted combination or a weighted average of the variant call classifications716 to recalibrate (increaseor decrease) a call-qualitymetric for anucleotidebasecall (e.g., an initial variant call).

[0177] To update or recalibrate the genotype metric (e.g., within the GT field of the variant call file 720) associated with a nucleotide base call, the call recalibra- tion system 106 utilizes one or more of the variant call 5 10 15 20 25 30 35 40 45 50 55 30 57 EP 4 693 301 A2 58 classifications 716. For example, the call recalibration system 106 compares the three variant call classifica- tions as the variant call classifications 716 (e.g., the false positive classification, the genotype error classification, and the true-positive classification) to determinewhich of the variant call classifications 716 has a highest prob- ability. In some cases, the call recalibration system 106 utilizes the variant call classification with the highest probability to recalibrate the genotype metric (e.g., from 0 as corresponding to the reference base to 1 as corre- sponding to a first alternative supporting read). For in- stance, if the call recalibration system 106 determines a highest probability for the false positive classification, then the call recalibration system 106 recalibrates the genotype metric accordingly. As another example, if the call recalibration system 106 determines a highest prob- ability for the true-positive classification, then the call recalibration system 106 recalibrates (or refrains from recalibrating) the genotype metric.

[0178] In other embodiments, the call recalibration system 106 utilizes only the genotype error probability to modify the genotype metric. For example, if the call recalibration system 106 determines a high genotype error probability, then the call recalibration system 106 recalibrates the genotype metric to indicate a different genotype of a nucleotide base call.

[0179] To update or recalibrate the genotype-quality metric (e.g., within theGQ field of the variant call file 720) associated with a nucleotide base call, the call recalibra- tion system 106 utilizes one or more of the variant call classifications 716. More specifically, the call recalibra- tion system 106 determines how each of the variant call classifications 716 affect the genotype-qualitymetric and recalibrates the genotype-quality metric accordingly (e.g., by increasing or decreasing the quality score be- tween 0 to 10 or 0 to 100 or on some other scale). For example, the call recalibration system 106 determines that a higher genotype error probability (generally) indi- cates a lower genotype-quality metric, and the call re- calibration system 106 reduces the metric accordingly.

[0180] In somecases, the call recalibration system106 determines a combination (e.g., a weighted combination or a weighted average) of the variant call classifications 716 to modify the genotype-quality metric. For example, the call recalibration system 106 determines a combined effect that the variant call classifications 716 have on the genotype-quality metric. As another example, the call recalibration system 106 determines individual impacts that each variant call classification has on the genotype- quality metric and weights each variant call classification accordingly. The call recalibration system 106 further recalibrates the genotype-quality metric by increasing or decreasing its value based on the indicated probabil- ities associated with each of the variant call classifica- tions 716.

[0181] As described, the call recalibration system 106 generates variant call classifications 716 and a nucleo- tide base call from the sameset of sequencingmetrics (or a subset of the sequencing metrics that are shared between the call recalibration machine learning model 714 and the call generation model 722). Indeed, the call recalibration system 106 utilizes the call recalibration machine learning model 714 to generate the variant call classifications 716 from sequencing metrics while also generating a nucleotide base call for a sample sequence. Indeed, the call recalibration system 106 can operate the call recalibration machine learning model 714 in parallel with the call generationmodel 722 to generatemetrics for a nucleotide base call and variant call classifications 716 for recalibrating the generated metrics.

[0182] As further illustrated in FIG. 7, the call recalibra- tion system 106 generates a variant call file 720. In particular, the call recalibration system 106 generates a variant call file 720 that represents or defines a nucleo- tidebasecall from thesequencingmetrics corresponding to a genomic coordinate. As shown, the variant call file 720 includes various call metrics such as a call-quality metric (QUAL), a genotype metric (GT), and a genotype- quality metric (GQ). To generate the variant call file 720, as described, the call recalibration system106generates metrics for a nucleotide base call utilizing the call gen- eration model 722 and recalibrates the nucleotide base call utilizing the variant call classifications 716 from the call recalibration machine learning model 714.

[0183] In one or more implementations, the call recali- bration system 106 updates or otherwise modifies the data fields for the variant call file 720 according to parti- cular algorithms. Aftermodifying such data fields, the call recalibration system106 can generate the variant call file 720 (e.g., a post-filter variant call file) to include metrics reflecting the updated data fields for QUAL, GT, and GQ. For instance, in somecases, the call recalibration system 106 updates the QUAL field for every variant based on the probability of a false positive variant (e.g., the false positive classification). As indicated above, in some cases, QUAL indicates the probability that there is some kind of variant (or other nucleotide base call) at a given location, measured in PHRED scale.

[0184] In addition, if the call recalibration system 106 determines that the highest probability from among the three variant call classifications (as the variant call clas- sifications 716) is the genotype error classification (e.g., the probability of a het / hom error), then the call recalibra- tion system 106 updates the GQ field while preserving or maintaining the GT field. Specifically, in some embodi- ments, the call recalibration system 106 updates the GQ field based on the true-positive classification (e.g., the probability of a true genotype).

[0185] Further, if the call recalibration system 106 de- termines that the highest probability from among the variant call classifications 716 is the true-positive classi- fication, in some cases, the call recalibration system 106 updates both the GQ field and the GT field. Specifically, the call recalibration system 106 updates the GQ field based on the genotype error classification and further updates theGTfield to switch thegenotypedependingon 5 10 15 20 25 30 35 40 45 50 55 31 59 EP 4 693 301 A2 60 whether the existing GT is 0 / X or X / X (where X is a non- zero value).

[0186] If the call recalibration system 106 determines that neither the true-positive classification nor the geno- type error classification has the highest probability among the variant call classifications 716, in some em- bodiments, the call recalibration system 106 updates the GQ field. In other words, if the call recalibration system 106 determines that the false positive classification has the highest probability, the call recalibration system 106 updates the GQ field. In particular, the call recalibration system106updates theGQfield basedon theprobability indicated by the true-positive classification.

[0187] As suggested above, in some embodiments, the call recalibration system 106 increases or decreases a base call quality metric (e.g., Q score) for a nucleotide basecall. Basedon thevariant call classifications716, for example, the call recalibration system 106 increases base call quality metrics for nucleotide base calls that would not have previously passed a quality filter and determines that the increased base call quality metrics now passes the quality filter. In some such cases, the call recalibration system 106 includes nucleotide base calls with such increasedbase call qualitymetrics (passing the quality filter) in a post-filter variant call file. By contrast, in other cases, the call recalibration system 106 decreases base call quality metrics for nucleotide base calls that previously would have passed a quality filter and deter- mines that the decreased base call quality metrics now fail the quality filter. In some such cases, the call recali- bration system 106 excludes nucleotide base calls with decreased base call quality metrics (failing the quality filter) from a post-filter variant call file, but includes the nucleotide base calls with such decreased base call quality metrics in a pre-filter variant call file.

[0188] For example, the call recalibration system 106 can remove false positive variant calls and recover false negative variant calls by changing corresponding base call quality metrics. To remove a false positive, in some cases, the call recalibration system 106 decreases the base call quality metric of a nucleotide base call that initially passed a quality filter-based on the variant call classifications 716 from the call recalibration machine learning model 714. Based on determining the de- creased base call quality metric falls below a threshold metric (e.g., aQscoreof 3.0or10.0), thecall recalibration system 106 determines that the nucleotide base call no longer passes the quality filter. The call recalibration system 106 thus filters out, or removes, the false posi- tive-nucleotide base call that initially passed the filter by changing its base call quality metric.

[0189] In addition to removing false positive variant calls based on changes to base call quality metrics, the call recalibration system 106 can remove false posi- tive variant calls based on changes to genotype. To remove a false positive, in some cases, the call recali- bration system 106 changes a genotype of an initial nucleotidebase call indicatingadifferent nucleotidebase than a reference base (e.g., GT = 1 or 2) to a genotype of an updated nucleotide base call indicating a same nu- cleotide base as the reference base (e.g., GT = 0)‑based on the variant call classifications 716 from the call recali- bration machine learning model 714. Based on the gen- otype being the same as the reference base, the call recalibration system 106 does not identify the nucleotide base call as a variant and, in some cases, excludes data for the nucleotide base call from a variant call file.

[0190] To recover a false negative, the call recalibra- tion system106 increases thebase call qualitymetric of a nucleotide base call that initially failed a quality filter- based on the variant call classifications 716 from the call recalibration machine learning model 714. Based on determining the increased base call quality metric ex- ceeds a threshold metric, the call recalibration system 106 determines that the nucleotide base call passes the quality filter. The call recalibration system 106 thus re- covers a false-negative-nucleotide base call that was initially filteredout bychanging its basecall qualitymetric.

[0191] In addition to recovering false negative variant calls based on changes to base call quality metrics, the call recalibration system 106 can recover false negative variant calls based on changes to genotype. To recover a false negative, in some cases, the call recalibration sys- tem 106 changes a genotype of an initial nucleotide base call indicating the same nucleotide base as a reference base (e.g., GT = 0) to a different genotype of an updated nucleotidebase call indicatingadifferent nucleotidebase than the reference base (e.g., GT = 1 or 2)‑based on the variant call classifications 716 from the call recalibration machine learning model 714. Based on the differing genotype of the updated nucleotide base call and a passing base call quality metric, the call recalibration system106 identifies thenucleotide base call as a variant and includes the nucleotide base call within a variant call file.

[0192] Indeed, in some implementations, the call re- calibration system 106 operates in a specific sequential order utilizing the call generation model 722 and the call recalibration machine learning model 714. For example, the call recalibration system 106 generates a FASTQ file by converting a BCL file to FASTQ. In addition, the call recalibrationsystem106 (subsequently) utilizes themap- ping-and-alignment components 708 of the call genera- tion model 722 to map and align nucleotide bases from a sample nucleotide sequence. In some cases, the call recalibration system 106maps and aligns the nucleotide bases of the sample sequence in relation to a reference sequence (e.g., reference genome) and / or various alter- native supporting reads.

[0193] After mapping and aligning, as described here- in, the call recalibration system 106 then utilizes the variant caller components 710 of the call generation model 722 to generate an initial nucleotide base call for the sample sequence corresponding to a particular genomic coordinatebased on various sequencing me- trics. After or at the same time, the call recalibration 5 10 15 20 25 30 35 40 45 50 55 32 61 EP 4 693 301 A2 62 system 106 also applies the call recalibration machine learningmodel 714 to generate the variant call classifica- tions 716 from sequencing metrics extracted via the mapping and aligning, the variant calling, and / or from other sources as described above. Based on the variant call classifications 716, the call recalibration system 106 recalibrates the nucleotide base call (e.g., by modifying various data fields corresponding to specific metrics of the nucleotide base call such as QUAL, GT, and GQ).

[0194] In somecases, the call recalibration system106 further applies a quality filter to the nucleotide base call to determine whether the nucleotide base call passes the quality filter (e.g., a hard pass filter of Q20 or other Q score). The call recalibration system 106 subsequently identifies a subset of nucleotide base calls that represent variants from reference bases and pass the quality filter. The call recalibration system 106 further generates a modified or updated variant call file (e.g., the variant call file 720) that includes the subset of nucleotide base calls and recalibratedmetrics for the subset of nucleotidebase calls, such as updated QUAL metrics, updated GT me- trics, and / or updated GQ metrics.

[0195] As mentioned above, in certain embodiments, the call recalibration system 106 trains or tunes a call recalibration machine learning model (e.g., the call re- calibration machine learning model 714). In particular, the call recalibration system 106 utilizes an iterative trainingprocess to fit acall recalibrationmachine learning model by adjusting or adding decision trees or learning parameters that result in accurate variant call classifica- tions (e.g., variant call classifications 716). FIG. 8 illus- trates trainingacall recalibrationmachine learningmodel in accordance with one or more embodiments.

[0196] As illustrated in FIG. 8, the call recalibration system 106 accesses sample sequencing metrics 804 from a database 802 (e.g., the database 116). For ex- ample, the call recalibration system 106 accesses sam- ple sequencing metrics including sample read-based metrics, sample externally sourced sequencing metrics, and sample call model generated sequencingmetrics. In some cases, the sample sequencing metrics 804 have a corresponding ground truth variant call file 816 asso- ciated with them, where the ground truth variant call file 816 indicates an actual nucleotide base call and its various metrics that result from the sample sequencing metrics 804. For instance, the call recalibration system 106 utilizes sample sequencing metrics 804 and ground truth variant call files froma training dataset from the food and drug administration, called the PrecisionFDA data- set. In some cases, the sample sequencing metrics 804 include a subset of sample sequencing metrics for each nucleotide base call in a ground truth variant call file. The ground truth variant call file can have a ground truth variant call (e.g., genotype metric in a genotype field) and / or a ground truth base call corresponding to each subset of sample sequencing metrics.

[0197] As further illustrated in FIG. 8, the call recalibra- tion system 106 generates predicted variant call classi- fications 808 based on the sample sequencing metrics 804.Specifically, thecall recalibrationsystem106utilizes a call recalibrationmachine learningmodel 806 (e.g., the call recalibration machine learning model 714) to gener- ate the predicted variant call classifications 808. Indeed, in some embodiments, the call recalibration machine learning model 806 generates a set of three predicted variant call classifications 808 including a predicted false positive classification, a predicted genotype error classi- fication, and a predicted true-positive classification. The predicted variant call classifications 808 can accordingly take the form of any of the variant call classifications described above.

[0198] Based on the predicted variant call classifica- tions 808, the call recalibration system 106 determines nucleotide base calls and generates a modified variant call file 810 comprising the nucleotide base calls and corresponding fields. As indicated above, the call recali- bration system 106 can utilize (i) a call generation model to generate an initial nucleotide base call and (ii) the call recalibration machine learning model 806 to modify data fields corresponding to a variant call file for the nucleotide basecall.Suchmodifiedor recalibratedvaluesareoutput in themodified variant call file 810 by, for example the call generation model. For example, the call recalibration system 106 determines recalibrated values for particular metrics within the modified variant call file 810, including a call-qualitymetric (QUAL), a genotypemetric (GT), and a genotype-quality metric (GQ).

[0199] As further illustrated in FIG. 8, the call recalibra- tion system106performsa comparison 812.Specifically, the call recalibration system 106 performs the compar- ison 812 between (i) variant nucleotide base calls and / or data fields in the modified variant call file 810 and (ii) variant nucleotide base calls and / or data fields in the ground truth variant call file 816. In some embodiments, the call recalibration system 106 utilizes a loss function 814 to compare variant nucleotide base calls and / or data fields from the two variant call files (e.g., to determine an error or ameasureof lossbetween them).For instance, in caseswhere the call recalibrationmachine learningmod- el 806 is an ensemble of gradient boosted trees, the call recalibration system 106 utilizes a mean squared error loss function (e.g., for regression) and / or a logarithmic loss function (e.g., for classification) as the loss function 814.

[0200] By contrast, in embodiments where the call recalibration machine learning model 806 is a neural network, the call recalibration system 106 can utilize a cross entropy loss function, an L1 loss function, or a mean squared error loss function as the loss function 814. For example, the call recalibration system 106 uti- lizes the loss function 814 to determine a difference between variant nucleotide base calls and / or data fields from themodified variant call file 810and theground truth variant call file 816.

[0201] As further illustrated in FIG. 8, the call recalibra- tion system 106 performs model fitting 818. In particular, 5 10 15 20 25 30 35 40 45 50 55 33 63 EP 4 693 301 A2 64 the call recalibration system 106 fits the call recalibration machine learning model 806 based on the comparison 812. For instance, the call recalibration system 106 per- forms modifications or adjustments to the call recalibra- tion machine learning model 806 to reduce the measure of loss from the loss function 814 for a subsequent train- ing iteration.

[0202] For gradient boosted trees, for example, the call recalibration system 106 trains the call recalibration ma- chine learning model 806 on the gradients of the errors determinedby the loss function814.For instance, thecall recalibration system 106 solves a convex optimization problem (e.g., of infinite dimensions) while regularizing the objective to avoid overfitting. In certain implementa- tions, the call recalibration system 106 scales the gradi- ents to emphasize corrections to under-represented classes (e.g., where there are significantly more true positives than false positive variant calls).

[0203] In some embodiments, the call recalibration system 106 adds a new weak learner (e.g., a new boosted tree) to the call recalibration machine learning model 806 foreachsuccessive training iterationaspart of solving the optimization problem. For example, the call recalibration system 106 finds a feature (e.g., a sequen- cing metric) that minimizes a loss from the loss function 814 and either adds the feature to the current iteration’s tree or starts to build a new tree with the feature.

[0204] In addition or in the alternative to gradient boosted decision trees, the call recalibration system 106 trains a logistic regression to learn parameters for generating one or more variant call classifications such as a true-positive classification. To avoid overfitting, the call recalibrationsystem106 further regularizesbasedon hyperparameters such as the learning rate, stochastic gradient boosting, the number of trees, the tree-depth(s), complexity penalization, and L1 / L2 regularization.

[0205] In embodiments where the call recalibration machine learning model 806 is a neural network, the call recalibration system 106 performs the model fitting 818 by modifying internal parameters (e.g., weights) of the call recalibration machine learning model 806 to reduce the measure of loss for the loss function 814. Indeed, the call recalibration system 106 modifies how the call re- calibration machine learning model 806 analyzes and passes data between layers and neurons by modifying the internal network parameters. Thus, over multiple iterations, the call recalibration system 106 improves the accuracy of the call recalibration machine learning model 806.

[0206] Indeed, in some cases, the call recalibration system 106 repeats the training process illustrated in FIG. 8 for multiple iterations. For example, the call recali- bration system 106 repeats the iterative training by se- lecting a new set of sequencing metrics for each nucleo- tide base call along with a corresponding ground truth nucleotide base call in a corresponding ground truth variant call file. The call recalibration system 106 further generates a new set of predicted variant call classifica- tions for each iteration along with a new modified variant call file. As described above, the call recalibration system 106 also compares a variant nucleotide base calls and / or data fields from the modified variant call file at each iteration with the corresponding variant nucleotide base calls and / or data fields from the corresponding ground truth variant call file and further performs model fitting 818. The call recalibration system 106 repeats this pro- cess until the call recalibration machine learning model 806 generates predicted variant call classifications that result in variant calls that satisfies a thresholdmeasure of loss. In some embodiments, the call recalibration system 106 performs the training process of FIG. 8 for homo- zygous reference coordinates to update ormodify variant calls of these coordinates and to thereby recover false negative variant calls (based on simulating haploid data from diploid data andmodifying inputs and outputs of the call recalibration machine learning model 806 as de- scribed).

[0207] As mentioned above, in certain described em- bodiments, the call recalibration system 106 generates and provides contribution measures associated with se- quencing metrics. In particular, the call recalibration sys- tem 106 determines respective contribution measures indicating how impactful individual sequencing metrics are indeterminingaparticular nucleotidebasecall. FIG.9 illustrates an example visualization of contribution mea- sures for sequencing metrics associated with a nucleo- tide base call in accordance with one or more embodi- ments.

[0208] As illustrated in FIG. 9, the client device 108 displays a contribution-measure interface 902 that in- cludes individual depictions of contribution measures associated with corresponding sequencing metrics. In- deed, the call recalibration system 106 determines a contribution measure for a sequencing metric based on how impactful or influential the sequencing metric is on a final nucleotide base call. Unlike many existing sequencing systems that utilize deep learning architec- tures, the structure of the call generation model used by the call recalibration system 106 facilitates the determi- nation of such contribution measures on a metric-by- metric basis.

[0209] For example, the call recalibration system 106 determines contributionmeasures by determining Shap- ley Additive Explanation (SHAP) values for each of the sequencing metrics for a nucleotide base call. Specifi- cally, the call recalibration system 106 determines a SHAP value by determining an impact of a sequencing metric as compared to the results of a baseline value (e.g., a baseline value for the sequencing metric). As shown in FIG. 9, the call recalibration system 106 deter- mines contribution measures for a number of listed se- quencingmetrics, where the thicker (e.g., more bulbous) portions of the graphs for each sequencing metric (roughly) indicate its contribution measure.

[0210] As further shown in FIG. 9, the call recalibration system 106 can rank the sequencing metrics according 5 10 15 20 25 30 35 40 45 50 55 34 65 EP 4 693 301 A2 66 to contribution measures as well. For instance, the call recalibration system106determines that the contribution for the mapq_p metric is highest among those displayed within the contribution-measure interface 902, followed by the qual metric, the gt0 metric, and so forth down the list.

[0211] As mentioned above, in certain described em- bodiments, the call recalibration system 106 improves in accuracy over existing sequencing systems. In particu- lar, the call recalibration system 106 reduces false posi- tive variant nucleotide base calls and false negative variant nucleotide base calls compared to existing se- quencing systems. Indeed, by utilizing a call recalibration machine learning model to recalibrate nucleotide base calls, the call recalibration system 106 even improves over previous versions of the call generation model that did not utilize a call recalibrationmachine learningmodel (but which still outperform other systems). FIGS. 10A‑10B illustrate graphs and tables of experiments demonstrating the accuracy improvements of the call recalibration system 106 as compared to some existing systems.

[0212] For reference and as depicted in FIGS. 10A‑10B and 11A‑11B, the name "Non-Recalibrated System 1" refers to an existing sequencing system that uses a linear reference genome for variant calling. By contrast, the name "Non-Recalibrated System 2" refers to an existing sequencing system that uses a graph reference genome for variant calling. Further, the name "Call Recalibration System1" refers to an embodiment of the call recalibration system 106 that is not configured for nucleotide base calls atmultiallelic genomic coordinates, haploid genomic coordinates, and would-be homozy- gous reference genomic coordinates. By contrast, the name "Call Recalibration System 2" refers to an embodi- ment of the call recalibration system 106 that is not configured for nucleotide base calls at multiallelic geno- mic coordinates, haploid genomic coordinates, and would-be homozygous reference genomic coordinates.

[0213] As illustrated in FIG. 10A, a graph 1002 depicts a number of receiver operating characteristic (ROC) curves that compare SNP false positives for two varia- tions of the call recalibration system106with those of two non-recalibrated systems. The graph 1002 depicts por- tions of ROC curves representing sensitivity over false positive variants detected, where sensitivity represents a number of correctly determined true positive variant calls divided by the sum of true positive variant calls and false positive variant calls. In particular, thegraph1002depicts ROC curves for different embodiments of the call recali- bration system106utilizing the call recalibrationmachine learningmodel-that is, "Call Recalibration System 1" and "Call Recalibration System 2," as described above. The experiment was performed using the PrecisionFDA truth set (e.g., the Precision FDA HG002 high confidence truthset). Generally, curves that trend upward and to the left in the graph 1002 are more accurate. As shown, embodimentsof thecall recalibrationsystem106exhibits improved accuracy over each of the three non-recali- brated systems, with higher sensitivity and fewer false positive variant calls comparatively. As shown by the improvements between the ROC curves for the Call Recalibration System 1 to the Call Recalibration System 2, the gain in sensitivity is due in part to recovering false negative variant calls at genomic coordinates that would have been identified as homozygous reference geno- types by another sequencing system.

[0214] Additionally, the graph 1004 depicts a number of ROC curves that compare non-SNP (e.g., indel) false positive variant calls for different embodiments of the call recalibration system 106 with those of a couple non- recalibrated systems, Non-Recalibrated System 1 and Non-Recalibrated System 2. The graph 1004 depicts ROC curves representing sensitivity over false positive variantsdetected. Inparticular, thegraph1004depictsan ROC curve for an embodiment of the call recalibration system 106-configured for nucleotide base calls at multi- allelic genomic coordinates, haploid genomic coordi- nates, and would-be homozygous reference genomic coordinates-that removes or reduces the bump or jog prevalent in the non-recalibrated systems at a sensitivity of ~0.4 (instead continuing smoothly upward on a nearly vertical trajectory). Indeed, due at least in part to the improvements at multiallelic genomic coordinates, an embodiment of the call recalibration system 106 (here, Call Recalibration System 2) exhibits fewer false positive variant calls at similar sensitivities, as compared tooneor more non-recalibrated systems that do not recalibrate multiallelic variants (e.g., the Non-Recalibrated System 2). The experiment was performed using the PrecisionF- DA truth set (e.g., the Precision FDA HG002 high con- fidence truth set).

[0215] As illustrated in FIG. 10B, table 1006 corre- sponds to the graph 1002, while the table 1008 corre- sponds to the graph1004. The numbers of the table 1006 and the table1008are takenatabestF-measurepoint for the curves in each of the graphs 1002 and 1004, respec- tively. As shown in table 1006, both embodiments of the call recalibration system 106 have fewer false negative variant calls (FN), fewer false positive variant calls (FP), and more true positives (TP) than any of the non-recali- brated systems. For example, at the best F-measure point, an embodiment of the call recalibration system 106-shown as Call Recalibration System 2 and is con- figured for nucleotide base calls at multiallelic genomic coordinates, haploid genomic coordinates, andwould-be homozygous reference genomic coordinates-produces 7309 false negative variant calls and 2801 false positive variant calls, as shown in the table 1006. The other embodiment of the call recalibration system 106-shown as Call Recalibration System 1 but is not configured for nucleotide base calls at multiallelic genomic coordinates and would-be homozygous reference genomic coordi- nates-generates 7717 false negative variant calls and 3216 false positive variant calls. Similarly, the call recali- bration system 106 has fewer het / hom errors, better 5 10 15 20 25 30 35 40 45 50 55 35 67 EP 4 693 301 A2 68 recall, and higher precision as well.

[0216] As illustrated in table 1008, the embodiments of the call recalibration system 106 outperforms the non- recalibrated systems for non-SNP scenarios as well. For example, at the best F-measure point of the table 1008, an embodiment of the call recalibration system 106- shown as Call Recalibration System 2 and is configured for nucleotide base calls at multiallelic genomic coordi- nates, haploid genomic coordinates, and would-be homozygous reference genomic coordinates-produces 513 false positive variant calls while the other embodi- ment of the call recalibration system 106 produces 618 false positive variant calls. Both non-recalibrated sys- tems produce far more false positive variant calls. The embodiments of the call recalibration system 106 also have higher precision than any of the non-recalibrated systems.

[0217] In addition to the diploid accuracy improve- ments shown inFIGS. 10A‑10B, FIGS. 11A‑11B illustrate haploid accuracy improvements. Specifically, the graphs 1102 and 1104 each depict two ROC curves, one for the call recalibration system 106 and one for a non-recali- bratedsystem.For instance, thegraph1102depictsROC curves for SNPs, while the graph 1104 depicts ROC curves for non-SNPs (e.g., indels). In each case, as a result of the accuracy improvements at haploid coordi- nates, the call recalibration system 106 has higher sen- sitivity and fewer false positive variant calls as compared to the non-recalibrated system. Indeed, in each of the graphs 1102 and 1104, the ROC curve for the call recali- bration system 106 is improved, with a best F-measure point that is located at a cross-section indicating fewer false positive variant calls at (approximately) the same sensitivity as compared to the non-recalibrated system. The experiments for the graphs 1102 and 1104 were performed using the PrecisionFDA truth set.

[0218] As illustrated in FIG. 11B, the table 1106 corre- sponds to the graph 1102, and the table 1108 corre- sponds to the graph 1104. Indeed, the table 1106 indi- cates SNP results for the call recalibration system 106 compared to the non-recalibrated system at a best F- measure point. As shown, the call recalibration system 106 hasmore true positives, fewer false negative variant calls, fewer false positive variant calls, higher recall, and higher precision for SNPs. Looking to the table 1108, the call recalibration system 106 produces (at the best F- measure point) more true positives, fewer false negative variant calls, fewer false positive variant calls, higher recall, and higher precision than the non-recalibrated system for non-SNPs as well.

[0219] Turning now to FIGS. 12‑14, these figures illus- trate example flowcharts, each of a series of acts of generating a final nucleotide base call or variant call based on variant call classifications from a call recalibra- tion machine learning model in accordance with one or more embodiments. While FIGS. 12‑14 illustrate acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIGS. 12‑14. The acts of FIGS. 12‑14 can be performed as part of a method. Alternatively, a non- transitory computer readable storage medium can com- prise instructions that, when executed by one or more processors, cause a computing device to perform the actsdepicted inFIGS.12‑14. Instill furtherembodiments, a system comprising at least one processor and a non- transitory computer readable medium comprising in- structions that, when executed by one or more proces- sors, cause the system to perform the acts of FIGS. 12‑14.

[0220] As shown in FIG. 12, the series of acts 1200 includes an act 1202 of determining sequencing metrics for amultiallelic genomic coordinate. In particular, the act 1202 can include determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to a multiallelic genomic coordinate of a sample nucleo- tide sequence.

[0221] In addition, the series of acts 1200 includes an act 1204 of generating a set of variant call classifications for the multiallelic genomic coordinate. In particular, the act 1204 can involve generating, utilizing a call recalibra- tion machine learning model and based on the sequen- cing metrics, a set of variant call classifications compris- ing a reference probability of a homozygous reference genotype at the multiallelic genomic coordinate, a differ- ing genotype probability of a genotype error at the multi- allelic genomic coordinate, and a correct variant prob- ability of a correct variant call genotype at the multiallelic genomic coordinate.

[0222] For example, generating the reference prob- ability can include determining a probability that a geno- type at the multiallelic genomic coordinate is a homo- zygous genotype with respect to a reference genome. Generate the differing genotype probability can include determining a probability that a predicted genotype for the multiallelic genomic coordinate is an incorrect geno- type or an incorrect allele in the predicted genotype. Generating the correct variant probability can include determining a probability that a predicted genotype for the multiallelic genomic coordinate is correct as initially determined by a call generation model.

[0223] As further illustrated in FIG. 12, the series of acts 1200 includes an act 1206 of determining final nucleotide base calls for the multiallelic genomic coordi- nate. In particular, the act 1206 can involve determining final nucleotide base calls for the multiallelic genomic coordinate based on the set of variant call classifications. For example, the act 1206 can involve predicting two nucleotide bases from three or more candidate alleles at the multiallelic genomic coordinate.

[0224] The series of acts 1200 can also include an act of modifying a base call quality metric or a genotype quality metric based on the set of variant call classifica- tions.Further, theseriesofacts1200can includeanact of generating a variant call file that includes the modified base call quality metric or the modified genotype quality metric. In addition, the series of acts 1200 can include an 5 10 15 20 25 30 35 40 45 50 55 36 69 EP 4 693 301 A2 70 act of generating updated genotype likelihoods for can- didate nucleotide base calls of alleles at the multiallelic genomic coordinate. In someembodiments, the series of acts 1200 includes an act of generating a variant call file that includes the updated genotype likelihoods.

[0225] As shown in FIG. 13, the series of acts 1300 includes an act 1302 of determining sequencing metrics for nucleotide base calls corresponding to a genomic coordinate of a haploid nucleotide sequence. In particu- lar, the act 1302 can involve determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to a genomic coordinate of a haploid nucleotide sequence from a sample.

[0226] The series of acts 1300 can also include an act 1304 of generating a first genotype probability and a second genotype probability. In particular, the act 1304 can involve generating, utilizing a call recalibration ma- chine learning model and based on the sequencing me- trics, a first genotype probability of a first genotype at the genomiccoordinateandasecondgenotypeprobability of a second genotype at the genomic coordinate. In some cases, the act 1304 includes acts of generating the first genotype probability comprises generating a probability that the first genotype at the genomic coordinate is a haploid reference genotype and generating the second genotype probability comprises generating a probability that the second genotype at the genomic coordinate is a haploid alternate genotype.

[0227] Generating the first genotype probability can include utilizing a layer of the call recalibration machine learning model to modify a homozygous reference prob- ability of a homozygous reference genotype at the geno- mic coordinate to generate a haploid reference probabil- ity of a reference genotype at the genomic coordinate. Generating the second genotype probability can include utilizing the layer of the call recalibration machine learn- ing model to modify a homozygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of an alternate genotype at the genomic coordinate.

[0228] In some cases, the act 1304 involves generat- ing, for the genomic coordinate utilizing one or more layers of the call recalibration machine learning model, a first confidencescore corresponding to a first genotype, a second confidence score corresponding to a second genotype, and a third confidence score corresponding to a third genotype. Theact 1304 canalso involveexcluding the second confidence score corresponding to the sec- ond genotype and normalizing the first confidence score and the third confidence score utilizing a softmax model to generate the first genotype probability and the second genotype probability.

[0229] As further shown, the series of acts 1300 can includeanact 1306of determininga final nucleotide base call indicating a haploid genotype. In particular, the act 1306 can involve determining a final nucleotide base call indicating a haploid genotype for the genomic coordinate based on the first genotype probability and the second genotype probability. For example, the act 1306 can involve determining one of: a haploid alternate genotype for the genomic coordinate, a modified base call quality metric, a modified genotype metric, and a modified gen- otype quality metric based on determining that the sec- ond genotype probability exceeds the first genotype probability or a haploid reference genotype for the geno- mic coordinate, a modified base call quality metric, and a modified genotype quality metric based on determining that the first genotype probability exceeds the second genotype probability.

[0230] In some embodiments, the series of acts 1300 includes an act of converting a haploid reference geno- type call generated by a call generationmodel to a diploid homozygous reference genotype call as an input for the call recalibration machine learning model. The series of acts 1300 can include an act of converting a haploid alternate genotype call generated by the call generation model toadiploidhomozygousalternategenotypecall as an input for the call recalibrationmachine learningmodel. Additionally, the series of acts 1300 can include an act of generating, utilizing the call recalibration machine learn- ing model, the first genotype probability and the second genotype probability based further on the diploid homo- zygous reference genotype call or the diploid homozy- gous alternate genotype call.

[0231] In certain embodiments, the series of acts 1300 includes an act of downsampling diploid sequencing metrics to simulate haploid sequencing metrics corre- sponding to thehaploid nucleotide sequence.Downsam- pling diploid sequencing metrics to simulate haploid se- quencingmetrics can includeacts of selectinga subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads and selecting, basedonnucleo- tide base calls of the subset of diploid nucleotide reads, a subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate geno- types as indicated by a call generation model or as indicated by a ground-truth base-call dataset (e.g., a well-curated truth set such as PrecisionFDA v4.2.1).

[0232] As shown in FIG. 14, the series of acts 1400 includes an act 1402 of determining one or more nucleo- tide base calls indicating a homozygous reference gen- otype. In particular, the act 1402 can involve determining, for one ormore nucleotide reads, one ormore nucleotide base calls indicating a homozygous reference genotype at a genomic coordinate of a sample nucleotide se- quence.

[0233] The series of acts 1400 can include an act 1404 of determining sequencing metrics for the one or more nucleotide base calls. In particular, the act 1404 can involve determining sequencing metrics for the one or morenucleotidebasecalls corresponding to thegenomic coordinate. For example, the act 1404 can involve de- termining one or more of read-based sequencing me- trics, externally sourced sequencingmetrics, or call mod- el generated sequencing metrics for the genomic coor- dinate indicated as having a homozygous reference 5 10 15 20 25 30 35 40 45 50 55 37 71 EP 4 693 301 A2 72 genotype.

[0234] As shown, the series of acts 1400 can include an act 1406 of generating one or more variant call clas- sifications. In particular, the act 1406 can involve gen- erating, utilizing a call recalibration machine learning model and based on the sequencing metrics from the one or more nucleotide base calls, one or more variant call classifications indicating an accuracy of identifying a variant at the genomic coordinate.

[0235] As further illustrated in FIG. 14, the series of acts 1400 can include an act 1408 of determining a variant call from the one or more variant call classifica- tions. In particular, theact 1408can involvedetermininga variant call for the genomic coordinate based on the one or more variant call classifications. For example, the act 1408 can involve receiving, from a call generationmodel, an indication of the homozygous reference genotype at the genomic coordinate and determining the variant call for thegenomiccoordinatebymodifying thehomozygous reference genotype to a different genotype based on the one or more variant call classifications.

[0236] In some embodiments, the series of acts 1400 includes an act of identifying a previous homozygous reference genotype call from a call generation model for the sample at the genomic coordinate. Further, the series of acts 1400 includes an act of identifying a ground truth base call for the sample at the genomic coordinate and an act of modifying the call recalibration machine learning model based on a comparison of the variant call for the genomic coordinate and the ground truth base call for the genomic coordinate. The series of acts 1400 can include an act of updating one or more of a call quality field, a genotype field, or a genotype quality field corre- sponding to a variant call file based on the one or more variant call classifications.

[0237] In certain implementations, the series of acts 1400 includes an act of determining, for the genomic coordinate, one of: a homozygous alternate genotype based on determining that a true positive classification (e.g., a homozygous alternate classification) has a high- est probability from among the one or more variant call classifications, a heterozygous genotype based on de- termining that a genotype error classification (e.g., a heterozygous genotype classification) has the highest probability from among the one or more variant call classifications, or a homozygous reference genotype based on determining that neither the true positive clas- sification nor the genotype error classification has the highest probability from among the one or more variant call classifications.

[0238] The methods described herein can be used in conjunction with a variety of nucleic acid sequencing techniques. Particularly applicable techniques are those wherein nucleic acids areattachedat fixed locations in an array such that their relative positions do not change and wherein the array is repeatedly imaged. Embodiments in which images are obtained in different color channels, for example, coinciding with different labels used to distin- guish one nucleotide base type from another are parti- cularly applicable. In some embodiments, the process to determine the nucleotide sequence of a target nucleic acid (i.e., a nucleic acid polymer) can be an automated process. Preferred embodiments include sequencing- by-synthesis (SBS) techniques.

[0239] SBS techniques generally involve the enzy- matic extension of a nascent nucleic acid strand through the iterative addition of nucleotides against a template strand. In traditional methods of SBS, a single nucleotide monomer may be provided to a target nucleotide in the presence of a polymerase in each delivery. However, in the methods described herein, more than one type of nucleotide monomer can be provided to a target nucleic acid in the presence of a polymerase in a delivery.

[0240] SBS can utilize nucleotidemonomers that have a terminator moiety or those that lack any terminator moieties.Methods utilizing nucleotidemonomers lacking terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as set forth in further detail below. In methods using nucleo- tide monomers lacking terminators, the number of nu- cleotides added in each cycle is generally variable and dependent upon the template sequence and themode of nucleotide delivery. For SBS techniques that utilize nu- cleotide monomers having a terminator moiety, the ter- minator can be effectively irreversible under the sequen- cing conditions used as is the case for traditional Sanger sequencing which utilizes dideoxynucleotides, or the terminator canbe reversibleas is thecase for sequencing methods developed by Solexa (now Illumina, Inc.).

[0241] SBS techniques can utilize nucleotide mono- mers that have a label moiety or those that lack a label moiety. Accordingly, incorporation events can be de- tected based on a characteristic of the label, such as fluorescence of the label; a characteristic of the nucleo- tide monomer such as molecular weight or charge; a byproduct of incorporation of the nucleotide, such as release of pyrophosphate; or the like. In embodiments, where two or more different nucleotides are present in a sequencing reagent, the different nucleotides can be distinguishable from each other, or alternatively, the two or more different labels can be the indistinguishable under the detection techniquesbeing used. For example, the different nucleotides present in a sequencing reagent can have different labels and they can be distinguished using appropriate optics as exemplified by the sequen- cing methods developed by Solexa (now Illumina, Inc.).

[0242] Preferred embodiments include pyrosequen- cing techniques. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleo- tides are incorporated into the nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M. and Nyren, P. (1996) "Real-time DNA sequencing using de- tection of pyrophosphate release." Analytical Biochem- istry 242(1), 84‑9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3‑11; Ronaghi, M., Uhlen, M. and Nyren, P. (1998) "A 5 10 15 20 25 30 35 40 45 50 55 38 73 EP 4 693 301 A2 74 sequencingmethod based on real-time pyrophosphate." Science 281(5375), 363; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320, the disclosures of which are incorporated herein by refer- ence in their entireties). In pyrosequencing, released PPi can be detected by being immediately converted to ade- nosine triphosphate (ATP) by ATP sulfurylase, and the level of ATP generated is detected via luciferase-pro- duced photons. The nucleic acids to be sequenced can be attached to features in an array and the array can be imaged to capture the chemiluminescent signals that are produced due to incorporation of a nucleotides at the features of the array. An image can be obtained after the array is treated with a particular nucleotide type (e.g., A, T, C orG). Images obtained after addition of each nucleo- tide type will differ with regard to which features in the array are detected. Thesedifferences in the image reflect the different sequence content of the features on the array. However, the relative locations of each feature will remain unchanged in the images. The images can be stored, processed and analyzed using the methods set forth herein. For example, images obtained after treat- ment of the array with each different nucleotide type can be handled in the same way as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0243] In another exemplary type of SBS, cycle se- quencing is accomplished by stepwise addition of rever- sible terminator nucleotides containing, for example, a cleavable or photobleachable dye label as described, for example, inWO04 / 018497 andU.S. Pat. No. 7,057,026, the disclosures of which are incorporated herein by re- ference. This approach is being commercialized by So- lexa (now Illumina Inc.), and is also described in WO 91 / 06678 andWO 07 / 123,744, each of which is incorpo- rated herein by reference. The availability of fluorescen- tlylabeled terminators in which both the termination can be reversed and the fluorescent label cleaved facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be coengineered to efficiently incorporate and extend from thesemodified nucleotides.

[0244] Preferably in reversible terminator-based se- quencing embodiments, the labels do not substantially inhibit extension under SBS reaction conditions. How- ever, the detection labels canbe removable, for example, by cleavage or degradation. Images can be captured following incorporation of labels into arrayed nucleic acid features. In particular embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array and each nucleotide type has a spectrally distinct label. Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially and an image of the array can be obtained between each addition step. In such embodiments, each image will show nucleic acid features that have incorporated nucleotides of a particu- lar type. Different features are present or absent in the different images due the different sequence content of each feature. However, the relative position of the fea- tures will remain unchanged in the images. Images ob- tained fromsuch reversible terminator-SBSmethods can be stored, processed and analyzed as set forth herein. Following the image capture step, labels can be removed and reversible terminator moieties can be removed for subsequent cycles of nucleotide addition and detection. Removal of the labels after they have been detected in a particular cycle and prior to a subsequent cycle can provide the advantage of reducing background signal and crosstalk between cycles. Examples of useful labels and removal methods are set forth below.

[0245] In particular embodiments some or all of the nucleotidemonomerscan include reversible terminators. In such embodiments, reversible terminators / cleavable fluors can include fluor linked to the ribosemoiety via a 3’ ester linkage (Metzker, Genome Res. 15:1767‑1776 (2005), which is incorporated herein by reference). Other approaches have separated the terminator chemistry from the cleavage of the fluorescence label (Ruparel et al., Proc Natl Acad Sci USA 102: 5932‑7 (2005), which is incorporated herein by reference in its entirety). Ruparel et al described the development of reversible terminators that used a small 3’ allyl group to block extension, but could easily be deblocked by a short treatment with a palladium catalyst. The fluorophore was attached to the base via a photocleavable linker that could easily be cleaved by a 30 second exposure to long wavelength UV light. Thus, either disulfide reduction or photoclea- vage can be used as a cleavable linker. Another ap- proach to reversible termination is the use of natural termination that ensues after placement of a bulky dye on a dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one in- corporation event prevents further incorporations unless the dye is removed. Cleavage of the dye removes the fluor and effectively reverses the termination. Examples of modified nucleotides are also described in U.S. Pat. No. 7,427,673, and U.S. Pat. No. 7,057,026, the disclo- sures of which are incorporated herein by reference in their entireties.

[0246] Additional exemplary SBS systems and meth- ods which can be utilized with the methods and systems describedhereinaredescribed inU.S.PatentApplication Publication No. 2007 / 0166705, U.S. Patent Application Publication No. 2006 / 0188901, U.S. Pat. No. 7,057,026, U.S. Patent Application Publication No. 2006 / 0240439, U.S. Patent Application Publication No. 2006 / 0281109, PCT Publication No. WO 05 / 065814, U.S. Patent Appli- cation Publication No. 2005 / 0100900, PCT Publication No. WO 06 / 064199, PCT Publication No. WO 07 / 010,251, U.S. Patent Application Publication No. 2012 / 0270305 and U.S. Patent Application Publication No. 2013 / 0260372, the disclosures of which are incorpo- rated herein by reference in their entireties.

[0247] Someembodiments can utilize detection of four 5 10 15 20 25 30 35 40 45 50 55 39 75 EP 4 693 301 A2 76 different nucleotides using fewer than four different la- bels. For example, SBS can be performed utilizingmeth- ods and systems described in the incorporatedmaterials of U.S. Patent Application Publication No. 2013 / 0079232. As a first example, a pair of nucleotide types can be detected at the same wavelength, but distinguished based on a difference in intensity for one member of the pair compared to the other, or based on a change to one member of the pair (e.g. via chemical modification, photochemical modification or physical modification) that causes apparent signal to appear or disappear compared to the signal detected for the other member of the pair. As a second example, three of four different nucleotide types can be detected under parti- cular conditions while a fourth nucleotide type lacks a label that is detectable under those conditions, or is minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incor- poration of the first three nucleotide types into a nucleic acid can be determined based on presence of their respective signals and incorporation of the fourth nucleo- tide type into the nucleic acid can be determined based on absence or minimal detection of any signal. As a third example, one nucleotide type can include label(s) that are detected in two different channels, whereas other nucleotide types are detected in no more than one of the channels. The aforementioned three exemplary config- urations are not considered mutually exclusive and can be used in various combinations. An exemplary embodi- ment that combines all three examples, is a fluorescent- basedSBSmethod that usesa first nucleotide type that is detected in a first channel (e.g. dATPhavinga label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g. dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and the second channel (e.g. dTTP having at least one label that is detected in both channelswhenexcitedby the first and / or second excitation wavelength) and a fourth nucleotide type that lacks a label that is not, orminimally, detected in either channel (e.g. dGTP having no label).

[0248] Further, as described in the incorporated mate- rials of U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called one-dye sequencing approaches, the first nucleotide type is labeled but the label is removed after the first image is generated, and the second nucleotide type is labeled only after a first image is generated. The third nucleotide type retains its label in both the first and second images, and the fourth nucleotide type remains unlabeled in both images.

[0249] Some embodiments can utilize sequencing by ligation techniques. Such techniques utilize DNA ligase to incorporate oligonucleotides and identify the incor- poration of such oligonucleotides. The oligonucleotides typically have different labels that are correlated with the identity of a particular nucleotide in a sequence to which the oligonucleotides hybridize. As with other SBS meth- ods, images can be obtained following treatment of an arrayof nucleic acid featureswith the labeled sequencing reagents. Each imagewill shownucleic acid features that have incorporated labels of a particular type. Different features are present or absent in the different images due the different sequence content of each feature, but the relative position of the features will remain unchanged in the images. Images obtained from ligation-based se- quencing methods can be stored, processed and ana- lyzed as set forth herein. Exemplary SBS systems and methods which can be utilized with the methods and systems described herein are described in U.S. Pat. No. 6,969,488, U.S. Pat. No. 6,172,218, and U.S. Pat. No. 6,306,597, the disclosures of which are incorporated herein by reference in their entireties.

[0250] Some embodiments can utilize nanopore se- quencing (Deamer, D. W. & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147‑151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis". Acc. Chem. Res. 35:817‑825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A.Golovchenko, "DNAmolecules and configurations in a solid-state nanopore microscope" Nat. Mater. 2:611‑615 (2003), the disclosures of which are incorporated herein by reference in their entireties). In such embodiments, the target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or biological mem- brane protein, such asα-hemolysin. As the target nucleic acid passes through the nanopore, each base-pair can be identified by measuring fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni, G.V.&Meller, "A.Progress towardultrafastDNAsequen- cing using solid-state nanopores." Clin. Chem. 53, 1996‑2001 (2007); Healy, K. "Nanopore-based single- molecule DNA analysis." Nanomed. 2, 459‑481 (2007); Cockroft, S. L., Chu, J., Amorin, M. & Ghadiri, M. R. "A single-molecule nanopore device detects DNA polymer- ase activity with single-nucleotide resolution." J. Am. Chem. Soc. 130, 818‑820 (2008), the disclosures of whichare incorporatedherein by reference in their entire- ties). Data obtained from nanopore sequencing can be stored, processed and analyzed as set forth herein. In particular, the data can be treated as an image in accor- dance with the exemplary treatment of optical images and other images that is set forth herein.

[0251] Some embodiments can utilize methods invol- ving the real-timemonitoring of DNApolymerase activity. Nucleotide incorporations can be detected through fluor- escence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and γ-phos- phate-labeled nucleotides as described, for example, in U.S. Pat. No. 7,329,492 and U.S. Pat. No. 7,211,414 (each of which is incorporated herein by reference) or nucleotide incorporations can be detected with zero- mode waveguides as described, for example, in U.S. 5 10 15 20 25 30 35 40 45 50 55 40 77 EP 4 693 301 A2 78 Pat. No. 7,315,019 (which is incorporated herein by reference) and using fluorescent nucleotide analogs and engineered polymerases as described, for example, in U.S. Pat. No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082 (each of which is incorpo- rated herein by reference). The illumination can be re- stricted to a zeptoliter-scale volume around a surface- tethered polymerase such that incorporation of fluores- cently labeled nucleotides can be observed with low background (Levene,M.J. etal. "Zero-modewaveguides for single-molecule analysis at high concentrations." Science 299, 682‑686 (2003); Lundquist, P. M. et al. "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026‑1028 (2008); Korlach, J. et al. "Selective aluminum passivation for targeted immobili- zation of single DNA polymerase molecules in zero- mode waveguide nano structures." Proc. Natl. Acad. Sci.USA105,1176‑1181 (2008), thedisclosuresofwhich are incorporated herein by reference in their entireties). Images obtained from such methods can be stored, processed and analyzed as set forth herein.

[0252] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commer- cially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 A1; US 2009 / 0127589 A1; US 2010 / 0137143 A1; or US 2010 / 0282617 A1, each of which is incorporated herein by reference. Methods set forth herein for amplifying target nucleic acids using kinetic exclusion canbe readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons that are used to detect protons.

[0253] The above SBS methods can be advanta- geously carriedout inmultiplex formats such thatmultiple different target nucleic acids are manipulated simulta- neously. In particular embodiments, different target nu- cleic acids can be treated in a common reaction vessel or on a surface of a particular substrate. This allows con- venient delivery of sequencing reagents, removal of un- reacted reagents anddetection of incorporation events in a multiplex manner. In embodiments using surface- bound target nucleic acids, the target nucleic acids can be inanarray format. Inanarray format, the target nucleic acids can be typically bound to a surface in a spatially distinguishable manner. The target nucleic acids can be bound by direct covalent attachment, attachment to a bead or other particle or binding to a polymerase or other molecule that is attached to the surface. The array can include a single copy of a target nucleic acid at each site (also referred to as a feature) or multiple copies having the same sequence can be present at each site or fea- ture. Multiple copies can be produced by amplification methods such as, bridge amplification or emulsion PCR as described in further detail below.

[0254] The methods set forth herein can use arrays having features at any of a variety of densities including, for example, at least about 10 features / cm2, 100 fea- tures / cm2, 500 features / cm2, 1,000 features / cm2, 5,000 features / cm2, 10,000 features / cm2, 50,000 fea- tures / cm2, 100,000 features / cm2, 1,000,000 fea- tures / cm2, 5,000,000 features / cm2, or higher.

[0255] An advantage of the methods set forth herein is that they provide for rapid and efficient detection of a plurality of target nucleic acid in parallel. Accordingly the present disclosure provides integrated systems capable of preparing and detecting nucleic acids using techni- ques known in the art such as those exemplified above. Thus, an integrated system of the present disclosure can include fluidic components capable of delivering ampli- fication reagents and / or sequencing reagents to one or more immobilized DNA fragments, the system compris- ing components such as pumps, valves, reservoirs, flui- dic lines and the like. A flow cell can be configured and / or used in an integrated system for detection of target nucleic acids. Exemplary flow cells are described, for example, in US 2010 / 0111768 A1 and US Ser. No. 13 / 273,666, each of which is incorporated herein by reference. As exemplified for flow cells, one or more of the fluidic components of an integrated system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated system can be used for an amplification methodset forth hereinand for thedelivery of sequencing reagents in a sequencing method such as those exem- plified above. Alternatively, an integrated system can include separate fluidic systems to carry out amplification methods and to carry out detection methods. Examples of integrated sequencing systems that are capable of creating amplified nucleic acids and also determining the sequence of the nucleic acids include, without limitation, theMiSeqTMplatform (Illumina, Inc.,SanDiego,CA)and devices described in US Ser. No. 13 / 273,666, which is incorporated herein by reference.

[0256] The sequencing system described above se- quences nucleic acid polymers present in samples re- ceived by a sequencing device. As defined herein, "sam- ple" and its derivatives, is used in its broadest sense and includes any specimen, culture and the like that is sus- pected of including a target. In some embodiments, the sample comprises DNA, RNA, PNA, LNA, chimeric or hybrid formsofnucleicacids.Thesamplecan includeany biological, clinical, surgical, agricultural, atmospheric or aquatic-based specimen containing one or more nucleic acids. The term also includes any isolated nucleic acid sample such a genomic DNA, fresh-frozen or formalin- fixed paraffin-embedded nucleic acid specimen. It is also envisioned that the sample can be from a single indivi- dual, a collection of nucleic acid samples fromgenetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) 5 10 15 20 25 30 35 40 45 50 55 41 79 EP 4 693 301 A2 80 from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.

[0257] The nucleic acid sample can include high mo- lecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, low molecular weight material includes enzymatically or me- chanically fragmentedDNA.Thesamplecan includecell- free circulating DNA. In some embodiments, the sample can include nucleic acid molecules obtained from biop- sies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory ob- tained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic or patho- genic sample. In some embodiments, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another embodiment, the sample can include nucleic acid mole- cules obtained from a non-mammalian source such as a plant, bacteria, virus or fungus. In some embodiments, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0258] Further, the methods and compositions dis- closed herein may be useful to amplify a nucleic acid sample having low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from a forensic sample. In one embodiment, forensic samples can include nucleic acids obtained from a crime scene, nucleic acids obtained from a missing persons DNA database, nucleic acids obtained from a laboratory as- sociated with a forensic investigation or include forensic samples obtained by law enforcement agencies, one or moremilitary servicesoranysuchpersonnel. Thenucleic acid sample may be a purified sample or a crude DNA containing lysate, for example derived from a buccal swab, paper, fabric or other substrate that may be im- pregnated with saliva, blood, or other bodily fluids. As such, in some embodiments, the nucleic acid sample may comprise low amounts of, or fragmented portions of DNA, such as genomic DNA. In some embodiments, target sequences can be present in one or more bodily fluids including but not limited to, blood, sputum, plasma, semen, urine and serum. In some embodiments, target sequences can be obtained from hair, skin, tissue sam- ples, autopsy or remains of a victim. In some embodi- ments, nucleic acids including one or more target se- quences can be obtained from a deceased animal or human. In some embodiments, target sequences can include nucleic acids obtained from non-human DNA such a microbial, plant or entomological DNA. In some embodiments, target sequences or amplified target se- quences are directed to purposes of human identifica- tion. In some embodiments, the disclosure relates gen- erally to methods for identifying characteristics of a for- ensic sample. In some embodiments, the disclosure relates generally to human identification methods using one or more target specific primers disclosed herein or one or more target specific primers designed using the primer design criteria outlined herein. In one embodi- ment, a forensic or human identification sample contain- ing at least one target sequence can be amplified using any one or more of the target-specific primers disclosed herein or using the primer criteria outlined herein.

[0259] Thecomponents of thecall recalibrationsystem 106 can include software, hardware, or both. For exam- ple, the components of the call recalibration system 106 can include one or more instructions stored on a com- puter-readable storage medium and executable by pro- cessorsofoneormorecomputingdevices (e.g., theclient device 108). When executed by the one or more proces- sors, the computer-executable instructions of the call recalibration system 106 can cause the computing de- vices to perform the bubble detectionmethods described herein. Alternatively, the components of the call recali- bration system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alterna- tively, the components of the call recalibration system 106 can include a combination of computer-executable instructions and hardware.

[0260] Furthermore, the components of the call recali- bration system 106 performing the functions described herein with respect to the call recalibration system 106 may, for example, be implemented as part of a stand- alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the call recalibration system 106 may be implemented as part of a stand-alone application on a personal computing de- vice or a mobile device. Additionally, or alternatively, the components of the call recalibration system 106 may be implemented in anyapplication that provides sequencing services including, but not limited to IlluminaBaseSpace, Illumina DRAGEN, or Illumina TruSight software. "Illumi- na," "BaseSpace," "DRAGEN," and "TruSight," are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.

[0261] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, asdiscussed ingreater detail below.Embodimentswithin the scope of the present disclosure also include physical and other computer-readable media for carrying or stor- ing computer-executable instructions and / or data struc- tures. In particular, one or more of the processes de- 5 10 15 20 25 30 35 40 45 50 55 42 81 EP 4 693 301 A2 82 scribed herein may be implemented at least in part as instructions embodied in a non-transitory computer- readable medium and executable by one or more com- puting devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., amicroprocessor) receives instructions, fromanon-tran- sitory computer-readablemedium, (e.g., amemory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the pro- cesses described herein.

[0262] Computer-readablemedia canbeanyavailable media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are non-transitory computer-readable storage media (de- vices). Computer-readable media that carry computer- executable instructionsare transmissionmedia. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory com- puter-readable storage media (devices) and transmis- sion media.

[0263] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD- ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired pro- gram code means in the form of computer-executable instructions or data structures and which can be ac- cessed by a general purpose or special purpose compu- ter.

[0264] A "network" is defined as one ormore data links that enable the transport of electronic data between computer systems and / or modules and / or other electro- nic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hard- wired or wireless) to a computer, the computer properly views the connection as a transmission medium. Trans- missions media can include a network and / or data links which can be used to carry desired program codemeans in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0265] Further, upon reaching various computer sys- tem components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (de- vices) (or viceversa).Forexample, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network inter- face module (e.g., a NIC), and then eventually trans- ferred to computer system RAM and / or to less volatile computer storagemedia (devices) at a computer system. Thus, it should be understood that non-transitory com- puter-readable storage media (devices) can be included in computer system components that also (or even pri- marily) utilize transmission media.

[0266] Computer-executable instructions comprise, for example, instructions and datawhich, when executed at a processor, cause a general purpose computer, spe- cial purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instruc- tions are executed on a general-purpose computer to turn the general-purpose computer into a special pur- pose computer implementing elements of the disclosure. The computer executable instructions may be, for exam- ple, binaries, intermediate format instructions such as assembly language, or even source code. Although the subjectmatter hasbeendescribed in languagespecific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the ap- pended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0267] Those skilled in the art will appreciate that the disclosure may be practiced in network computing en- vironments with many types of computer system config- urations, including, personal computers, desktop com- puters, laptop computers, message processors, hand- held devices, multiprocessor systems, microprocessor- based or programmable consumer electronics, network PCs,minicomputers, mainframe computers, mobile tele- phones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distrib- uted system environments where local and remote com- puter systems,which are linked (either by hardwired data links, wireless data links, or by a combination of hard- wired and wireless data links) through a network, both perform tasks. In a distributed system environment, pro- gram modules may be located in both local and remote memory storage devices.

[0268] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the mar- ketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0269] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service 5 10 15 20 25 30 35 40 45 50 55 43 83 EP 4 693 301 A2 84 models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a "cloud-computing environment" is an environment in which cloud computing is employed.

[0270] FIG. 15 illustrates a block diagram of a comput- ing device 1500 thatmay be configured to perform one or more of the processes described above. One will ap- preciate that one or more computing devices such as the computing device 1500 may implement the call recali- bration system 106 and the sequencing system 104. As shown by FIG. 15, the computing device 1500 can com- priseaprocessor1502, amemory1504, a storagedevice 1506, an I / O interface 1508, and a communication inter- face 1510, which may be communicatively coupled by way of a communication infrastructure 1512. In certain embodiments, the computing device 1500 can include fewer or more components than those shown in FIG. 15. The following paragraphs describe components of the computing device 1500 shown in FIG. 15 in additional detail.

[0271] In one or more embodiments, the processor 1502 includes hardware for executing instructions, such as thosemaking up a computer program.As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1502 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1504, or the storage device 1506 and decode and execute them. The memory 1504 may be a volatile or nonvolatile mem- ory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1506 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instruc- tions for performing the methods described herein.

[0272] The I / O interface 1508 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1500. The I / O interface 1508 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1508 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1508 is configured to provide graphical data to a display for presentation to a user. The graphical datamay be representative of one or moregraphical user interfacesand / oranyothergraphical content as may serve a particular implementation.

[0273] The communication interface 1510 can include hardware, software, or both. In any event, the commu- nication interface 1510 can provide one or more inter- faces for communication (such as, for example, packet- based communication) between the computing device 1500 and one or more other computing devices or net- works. As an example, and not by way of limitation, the communication interface 1510 may include a network interface controller (NIC) or network adapter for commu- nicatingwithanEthernet orotherwire-basednetworkora wireless NIC (WNIC) or wireless adapter for communi- cating with a wireless network, such as a WI-FI.

[0274] Additionally, the communication interface 1510 may facilitate communicationswith various typesofwired or wireless networks. The communication interface 1510 may also facilitate communications using various com- munication protocols. The communication infrastructure 1512 may also include hardware, software, or both that couples components of the computing device 1500 to each other. For example, the communication interface 1510 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes de- scribed herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequen- cing device, and server device(s)) to exchange informa- tion such as sequencing data and error notifications.

[0275] In the foregoing specification, the present dis- closure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the ac- companying drawings illustrate the various embodi- ments. The description above and drawings are illustra- tive of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described toprovidea thoroughunderstandingof various embodiments of the present disclosure.

[0276] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scopeof the present application is, there- fore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0277] Aspects can be understood from the following numbered paragraphs: 1. A system comprising: at least one processor; and a non-transitory computer readable medium 5 10 15 20 25 30 35 40 45 50 55 44 85 EP 4 693 301 A2 86 comprising instructions that, when executed by the at least one processor, cause the system to: determine sequencing metrics for nucleo- tide base calls of nucleotide reads corre- sponding to a multiallelic genomic coordi- nate of a sample nucleotide sequence; generate, utilizing a call recalibration ma- chine learning model and based on the sequencing metrics, a set of variant call classifications comprisinga referenceprob- ability of a homozygous referencegenotype at the multiallelic genomic coordinate, a differing genotype probability of a genotype error at the multiallelic genomic coordinate, and a correct variant probability of a correct variant call genotype at the multiallelic genomic coordinate; and determine final nucleotide base calls for the multiallelic genomic coordinate based on the set of variant call classifications. 2. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to: modify a base call quality metric or a genotype quality metric based on the set of variant call classifications; and generate a variant call file that includes the modified base call quality metric or the modified genotype quality metric. 3. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate updated genotype likelihoods for can- didate nucleotide base calls of alleles at the multiallelic genomic coordinate; and generate a variant call file that includes the updated genotype likelihoods. 4. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the final nucleotide base calls for the multiallelic genomic coordinate by predicting two nucleotide bases from three or more candidate alleles at the multiallelic genomic coordinate. 5. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the refer- ence probability by determining a probability that a genotype at the multiallelic genomic coordinate is a homozygous genotype with respect to a reference genome. 6. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the differ- ing genotype probability by determining a probability that a predictedgenotype for themultiallelic genomic coordinate is an incorrect genotype or an incorrect allele in the predicted genotype. 7. The system of paragraph 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the correct variant probability by determining a probability that a predicted genotype for the multiallelic genomic co- ordinate is correct as initially determined by a call generation model. 8. A computer-implemented method comprising: determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to a genomic coordinate of a haploid nucleotide sequence from a sample; generating, utilizing a call recalibration machine learning model and based on the sequencing metrics, a first genotype probability of a first genotype at the genomic coordinate and a sec- ond genotype probability of a second genotype at the genomic coordinate; and determining a final nucleotide base call indicat- ing a haploid genotype for the genomic coordi- nate based on the first genotype probability and the second genotype probability. 9. The computer-implementedmethod of paragraph 8, wherein: generating the first genotype probability com- prises utilizing a layer of the call recalibration machine learning model to modify a homozy- gous reference probability of a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability of a reference genotype at the genomic coordinate; and generating the second genotype probability comprises utilizing the layer of the call recalibra- tion machine learning model to modify a homo- zygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of an alternate genotype at the genomic coordinate. 10. The computer-implemented method of para- graph 8, wherein generating the first genotype prob- ability and the second genotype probability com- prises: generating, for the genomic coordinate utilizing one or more layers of the call recalibration ma- chine learning model, a first confidence score corresponding to a first genotype, a second 5 10 15 20 25 30 35 40 45 50 55 45 87 EP 4 693 301 A2 88 confidence score corresponding to a second genotype, and a third confidence score corre- sponding to a third genotype; excluding the second confidence score corre- sponding to the second genotype; and normalizing the first confidence score and the third confidence score utilizing a softmax model to generate the first genotypeprobability and the second genotype probability. 11. The computer-implemented method of para- graph 8, wherein determining the final nucleotide base call indicating the haploid genotype for the genomic coordinate comprises determining one of: a haploid alternate genotype for the genomic coordinate, amodified base call qualitymetric, a modified genotype metric, and a modified gen- otype quality metric based on determining that the second genotype probability exceeds the first genotype probability; or a haploid reference genotype for the genomic coordinate, a modified base call quality metric, andamodifiedgenotypequalitymetric basedon determining that the first genotype probability exceeds the second genotype probability. 12. The computer-implemented method of para- graph 8, further comprising: converting a haploid reference genotype call generatedby a call generationmodel to a diploid homozygous reference genotype call as an in- put for the call recalibration machine learning model; or converting a haploid alternate genotype call generated by the call generation model to a diploid homozygous alternate genotype call as an input for the call recalibration machine learn- ing model; and generating, utilizing the call recalibration ma- chine learning model, the first genotype prob- ability and the second genotype probability based further on the diploid homozygous refer- ence genotype call or the diploid homozygous alternate genotype call. 13. The computer-implemented method of para- graph 8, further comprising downsampling diploid sequencing metrics to simulate haploid sequencing metrics corresponding to the haploid nucleotide se- quence by: selecting a subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads; and selecting, based on nucleotide base calls of the subset of diploid nucleotide reads, a subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate genotypes as indicated by a call generation model or as indicated by a ground-truth base- call dataset. 14. The computer-implemented method of para- graph 8, wherein: generating the first genotype probability com- prises generating a probability that the first gen- otype at the genomic coordinate is a haploid reference genotype; and generating the second genotype probability comprises generating a probability that the sec- ond genotype at the genomic coordinate is a haploid alternate genotype. 15. A non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause a computing device to: determine, for one or more nucleotide reads, one or more nucleotide base calls indicating a homozygous reference genotype at a genomic coordinate of a sample nucleotide sequence; determine sequencing metrics for the one or more nucleotide base calls corresponding to the genomic coordinate; generate, utilizing a call recalibration machine learning model and based on the sequencing metrics from the one or more nucleotide base calls, one or more variant call classifications indicating an accuracy of identifying a variant at the genomic coordinate; and determine a variant call for the genomic coordi- nate based on the one or more variant call classifications. 16.Thenon-transitory computer readablemediumof paragraph 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to: receive, from a call generationmodel, an indica- tion of the homozygous reference genotype at the genomic coordinate; and determine the variant call for the genomic co- ordinate by modifying the homozygous refer- ence genotype to a different genotype based on the one or more variant call classifications. 17.Thenon-transitory computer readablemediumof paragraph 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine the sequencing metrics by determining one or more of read-based sequencing metrics, externally sourced sequencing 5 10 15 20 25 30 35 40 45 50 55 46 89 EP 4 693 301 A2 90 metrics, or call model generated sequencingmetrics for the genomic coordinate indicated as having a homozygous reference genotype. 18.Thenon-transitory computer readablemediumof paragraph 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to: identify a previous homozygous reference gen- otype call from a call generation model for the sample nucleotide sequence at the genomic coordinate; identify a ground truth base call for the sample nucleotide sequenceat the genomic coordinate; and modify the call recalibration machine learning model based on a comparison of the variant call for the genomic coordinate and the ground truth base call for the genomic coordinate. 19.Thenon-transitory computer readablemediumof paragraph 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to determine, for the genomic coordinate, one of: a homozygous alternate genotype based on determining that a homozygous alternate clas- sification has a highest probability from among the one or more variant call classifications; a heterozygous genotype based on determining that a heterozygous genotype classification has the highest probability from among the one or more variant call classifications; or a homozygous reference genotype based on determining that neither the homozygous alter- nate classification nor the heterozygous geno- type classification has the highest probability from among the one or more variant call classi- fications. 20.Thenon-transitory computer readablemediumof paragraph 15, further comprising instructions that, when executed by the at least one processor, cause the computing device to update one or more of a call quality field, a genotype field, or a genotype quality field corresponding to a variant call file based on the one or more variant call classifications. Claims 1. A computer-implemented method comprising: determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to a genomic coordinate of a haploid nucleotide sequence from a sample; generating, utilizing a call recalibration machine learning model to process the sequencing me- trics, a set of variant call classifications compris- ing: a first genotype probability of a haploid re- ference genotype, with respect to a refer- ence genome, at the genomic coordinate; and a second genotype probability of a haploid alternate genotype, with respect to the re- ference genome, at the genomic coordi- nate; and determining a final nucleotide base call indicat- ing a haploid genotype for the genomic coordi- nate based on the first genotype probability and the second genotype probability. 2. The computer-implemented method of claim 1, wherein: generating the first genotype probability com- prises utilizing a layer of the call recalibration machine learning model to modify a homozy- gous reference probability of a homozygous reference genotype at the genomic coordinate togenerateahaploid referenceprobability of the haploid reference genotype at the genomic co- ordinate; and generating the second genotype probability comprises utilizing the layer of the call recalibra- tion machine learning model to modify a homo- zygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of the haploid alternate genotype at the genomic co- ordinate. 3. The computer-implemented method of claim 1 or 2, wherein generating the first genotypeprobability and the second genotype probability comprises: generating, for the genomic coordinate utilizing one or more layers of the call recalibration ma- chine learning model, a first confidence score corresponding to a first genotype, a second confidence score corresponding to a second genotype, and a third confidence score corre- sponding to a third genotype; excluding the second confidence score corre- sponding to the second genotype; and normalizing the first confidence score and the third confidencescoreutilizingasoftmax layer to generate the first genotype probability and the second genotype probability. 4. The computer-implemented method of any one of 5 10 15 20 25 30 35 40 45 50 55 47 91 EP 4 693 301 A2 92 claims 1‑3, wherein determining the final nucleotide base call indicating the haploid genotype for the genomic coordinate comprises determining one of: the haploid alternate genotype for the genomic coordinate, amodified base call qualitymetric, a modified genotype metric, and a modified gen- otype quality metric based on determining that the second genotype probability exceeds the first genotype probability; or the haploid reference genotype for the genomic coordinate, a modified base call quality metric, andamodifiedgenotypequalitymetric basedon determining that the first genotype probability exceeds the second genotype probability. 5. The computer-implemented method of any one of claims 1‑4, further comprising: converting a haploid reference genotype call generatedby a call generationmodel to a diploid homozygous reference genotype call as an in- put for the call recalibration machine learning model; or converting a haploid alternate genotype call generated by the call generation model to a diploid homozygous alternate genotype call as an input for the call recalibration machine learn- ing model; and generating, utilizing the call recalibration ma- chine learning model, the first genotype prob- ability and the second genotype probability based further on the diploid homozygous refer- ence genotype call or the diploid homozygous alternate genotype call. 6. The computer-implemented method of any one of claims 1‑5, further comprising downsampling diploid sequencing metrics to simulate haploid sequencing metrics corresponding to the haploid nucleotide se- quence by: selecting a subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads; and selecting, based on nucleotide base calls of the subset of diploid nucleotide reads, a subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate genotypes as indicated by a call generation model or as indicated by a ground-truth base- call dataset. 7. The computer-implemented method of claim 6, further comprising selecting, to simulate the haploid sequencing metrics corresponding to the haploid nucleotidesequence, thediploid sequencingmetrics corresponding to the subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate genotypes. 8. The computer-implemented method of any one of claims1‑7,wherein thehaploid nucleotide sequence comprises a sequence of one or more nucleotide bases from a haploid chromosome or a single chro- mosome without a counterpart chromosome. 9. The computer-implemented method of any one of claims1‑8,wherein thehaploid nucleotide sequence comprises a sequence of one or more nucleotide bases from a sex chromosome. 10. The computer-implemented method of any one of claims 1‑9, wherein: determining the sequencing metrics comprises determining one ormore of read-based sequen- cing metrics, externally sourced sequencing metrics, or call model generated sequencing metrics for the nucleotide base calls of nucleo- tide reads corresponding to the genomic coor- dinate of the haploid nucleotide sequence; and generating the set of variant call classifications comprises utilizing the call recalibration ma- chine learningmodel to process the one ormore of the read-based sequencing metrics, the ex- ternally sourced sequencing metrics, or the call model generated sequencing metrics to gener- ate the set of variant call classifications. 11. The computer-implemented method of claim 10, wherein: the read-based sequencing metrics comprise sequencing metrics derived from nucleotide reads of the sample; the externally sourced sequencingmetrics com- prise sequencing metrics stored in one or more databases external to a system executing the computer-implemented method; and the call model generated sequencing metrics comprise sequencing metrics generated by a call generation model. 12. The computer-implemented method of any one of claims 1‑11, wherein to determine the final nucleo- tide base call by modifying an initial nucleotide base call generated by a call generation model for a var- iant call file. 13. The computer-implemented method of any one of claims 1‑12, wherein: generating the first genotype probability com- prises utilizing a layer of the call recalibration machine learning model to modify a homozy- 5 10 15 20 25 30 35 40 45 50 55 48 93 EP 4 693 301 A2 94 gous reference probability of a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability of a reference genotype at the genomic coordinate; and generating the second genotype probability comprises utilizing the layer of the call recalibra- tion machine learning model to modify a homo- zygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of an alternate genotype at the genomic coordinate. 14. A system comprising at least one processor and a non-transitory computer readable medium compris- ing instructions that, when executed by the at least one processor, cause the system to perform ameth- od according to any one of claims 1‑13. 15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a system to perform a method according to any one of claims 1‑13. 5 10 15 20 25 30 35 40 45 50 55 49 EP 4 693 301 A2 50 EP 4 693 301 A2 51 EP 4 693 301 A2 52 EP 4 693 301 A2 53 EP 4 693 301 A2 54 EP 4 693 301 A2 55 EP 4 693 301 A2 56 EP 4 693 301 A2 57 EP 4 693 301 A2 58 EP 4 693 301 A2 59 EP 4 693 301 A2 60 EP 4 693 301 A2 61 EP 4 693 301 A2 62 EP 4 693 301 A2 63 EP 4 693 301 A2 64 EP 4 693 301 A2 65 EP 4 693 301 A2 66 EP 4 693 301 A2 67 EP 4 693 301 A2 68 EP 4 693 301 A2 69 EP 4 693 301 A2 REFERENCES CITED IN THE DESCRIPTION This list of references cited by the applicant is for the reader’s convenience only. It does not form part of the European patent document. Even though great care has been taken in compiling the references, errors or omissions cannot be excluded and the EPO disclaims all liability in this regard. Patent documents cited in the description • US 56393421

[0001] • US 280022

[0157] • US 165828

[0157] • US 15643381 B

[0157] • US 14811836 B

[0157] • US 6210891 B

[0242] • US 6258568 B

[0242] • US 6274320 B

[0242] • WO 04018497 A

[0243] • US 7057026 B

[0243]

[0245]

[0246] • WO 9106678 A

[0243] • WO 07123744 A

[0243] • US 7427673 B

[0245] • US 20070166705

[0246] • US 20060188901

[0246] • US 20060240439

[0246] • US 20060281109

[0246] • WO 05065814 A

[0246] • US 20050100900

[0246] • WO 06064199 A

[0246] • WO 07010251 A

[0246] • US 20120270305

[0246] • US 20130260372

[0246] • US 20130079232

[0247]

[0248] • US 6969488 B

[0249] • US 6172218 B

[0249] • US 6306597 B

[0249] • US 7001792 B

[0250] • US 7329492 B

[0251] • US 7211414 B

[0251] • US 7315019 B

[0251] • US 7405281 B

[0251] • US 20080108082

[0251] • US 20090026082 A1

[0252] • US 20090127589 A1

[0252] • US 20100137143 A1

[0252] • US 20100282617 A1

[0252] • US 20100111768 A1

[0255] • US 273666

[0255] Non-patent literature cited in the description • METZKER.Genome Res., 2005, vol. 15, 1767-1776

[0245] • RUPAREL et al. Proc Natl Acad Sci USA, 2005, vol. 102, 5932-7

[0245] • DEAMER, D. W. ; AKESON, M. Nanopores and nucleic acids: prospects for ultrarapid sequencing.. Trends Biotechnol., 2000, vol. 18, 147-151

[0250] • DEAMER, D. ; D. BRANTON. Characterization of nucleicacidsbynanoporeanalysis.Acc.Chem.Res., 2002, vol. 35, 817-825

[0250] • LI, J. ;M.GERSHOW ;D.STEIN ;E.BRANDIN ;J.A. GOLOVCHENKO. DNA molecules and configura- tions in a solid-state nanopore microscope. Nat. Mater., 2003, vol. 2, 611-615

[0250] • SONI, G. V. ; MELLER. A. Progress toward ultrafast DNA sequencing using solid-state nanopores.. Clin. Chem., 2007, vol. 53, 1996-2001

[0250] • HEALY, K. Nanopore-based single-molecule DNA analysis.. Nanomed., 2007, vol. 2, 459-481

[0250] • COCKROFT, S. L. ; CHU, J. ; AMORIN, M. ; GHADIRI, M. R. A single-molecule nanopore device detects DNA polymerase activity with single-nucleo- tide resolution.. J. Am. Chem. Soc., 2008, vol. 130, 818-820

[0250] • LEVENE, M. J. et al. Zero-mode waveguides for single-molecule analysis at high concentrations.. Science, 2003, vol. 299, 682-686

[0251] • LUNDQUIST, P. M. et al. Parallel confocal detection of single molecules in real time..Opt. Lett., 2008, vol. 33, 1026-1028

[0251] • KORLACH, J. et al. Selective aluminum passivation for targeted immobilization of singleDNApolymerase molecules in zero-modewaveguide nano structures.. Proc.Natl. Acad.Sci.USA, 2008, vol. 105, 1176-1181

[0251] (19) *EP004693301A3* (11) EP 4 693 301 A3 (12) EUROPEAN PATENT APPLICATION (88) Date of publication A3: 29.04.2026 Bulletin 2026 / 18 (43) Date of publication A2: 11.02.2026 Bulletin 2026 / 07 (21) Application number: 25224089.0 (22) Date of filing: 23.12.2022 (51) International Patent Classification (IPC): G16B 20 / 20 (2019.01) G16B 30 / 20 (2019.01) G16B 40 / 10 (2019.01) G16B 40 / 20 (2019.01) (52) Cooperative Patent Classification (CPC): G16B 20 / 20; G16B 30 / 20; G16B 40 / 10; G16B 40 / 20 (84) Designated Contracting States: AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR (30) Priority: 28.12.2021 US 202117563934 (62) Document number(s) of the earlier application(s) in accordance with Art. 76 EPC: 22857067.7 / 4 457 822 (71) Applicant: ILLUMINA, INC. San Diego, CA 92122 (US) (72) Inventor: PARNABY, Gavin San Diego, 92122 (US) (74) Representative: Robinson, David Edward Ashdown Marks & Clerk LLP Wytham Court, 11 West Way Oxford OX2 0JB (GB) (54) MACHINE LEARNING MODEL FOR RECALIBRATING NUCLEOTIDE BASE CALLS CORRESPONDING TO TARGET VARIANTS (57) This disclosure describes methods, non-transi- tory computer readable media, and systems that can utilize amachine learningmodel to recalibrate nucleotide base calls (e.g., variant calls) of a call generation model. For instance, thedisclosedsystemscan train andutilizea call recalibration machine learning model to generate a set of predicted variant call classifications based on sequencingmetrics associated with a sample nucleotide sequence. Leveraging the set of variant call classifica- tions, the disclosed systems can further update ormodify nucleotide base calls (e.g., variant calls) corresponding to genomic coordinates, such as multiallelic genomic coordinates, haploid genomic coordinates, and genomic coordinates indicated (by the call generation model) to exhibit homozygous reference genotypes. EP 4 69 3 30 1 A 3 Processed by Luminess, 75001 PARIS (FR) 2 EP 4 693 301 A3 5 10 15 20 25 30 35 40 45 50 55 3 EP 4 693 301 A3 5 10 15 20 25 30 35 40 45 50 55 用於重新校準與目標變異對應的核苷酸鹼基調用的機器學習模型 摘要 本揭露描述了能夠利用機器學習模型重新校準調用生成模型的核苷酸鹼基調用(例 如,變異調用)的方法、非瞬態計算機可讀介質和系統。例如,所揭露的系統可以訓 練和利用調用重新校準機器學習模型,以基於與樣本核苷酸序列相關的定序指標來產 生一組預測的變異調用分類。利用這組變異調用分類,所公開的系統可以進一步更新 或修改與基因組坐標(例如,多等位基因基因組坐標、單倍體基因組坐標以及(由調 用生成模型指示的)用於表示純合參考基因型的基因組坐標)對應的核苷酸鹼基調用 (例如,變異調用)。 摘 要

Claims

1. A computer-implemented method comprising: determining sequencing metrics for nucleotide base calls of nucleotide reads corresponding to a genomic coordinate of a haploid nucleotide sequence from a sample; generating, utilizing a call recalibration machine learning model to process the sequencing metrics, a set of variant call classifications comprising: a first genotype probability of a haploid reference genotype, with respect to a reference genome, at the genomic coordinate; and a second genotype probability of a haploid alternate genotype, with respect to the reference genome, at the genomic coordinate; and determining a final nucleotide base call indicating a haploid genotype for the genomic coordinate based on the first genotype probability and the second genotype probability.

2. The computer-implemented method of claim 1, wherein: generating the first genotype probability comprises utilizing a layer of the call recalibration machine learning model to modify a homozygous reference probability of a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability of the haploid reference genotype at the genomic coordinate; and generating the second genotype probability comprises utilizing the layer of the call recalibration machine learning model to modify a homozygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of the haploid alternate genotype at the genomic coordinate.

3. The computer-implemented method of claim 1 or 2, wherein generating the first genotype probability and the second genotype probability comprises: generating, for the genomic coordinate utilizing one or more layers of the call recalibration machine learning model, a first confidence score corresponding to a first genotype, a second confidence score corresponding to a second genotype, and a third confidence score corresponding to a third genotype; excluding the second confidence score corresponding to the second genotype; and normalizing the first confidence score and the third confidence score utilizing a softmax layer to generate the first genotype probability and the second genotype probability.

4. The computer-implemented method of any one of claims 1-3, wherein determining the final nucleotide base call indicating the haploid genotype for the genomic coordinate comprises determining one of: the haploid alternate genotype for the genomic coordinate, a modified base call quality metric, a modified genotype metric, and a modified genotype quality metric based on determining that the second genotype probability exceeds the first genotype probability; or the haploid reference genotype for the genomic coordinate, a modified base call quality metric, and a modified genotype quality metric based on determining that the first genotype probability exceeds the second genotype probability.

5. The computer-implemented method of any one of claims 1-4, further comprising: converting a haploid reference genotype call generated by a call generation model to a diploid homozygous reference genotype call as an input for the call recalibration machine learning model; or converting a haploid alternate genotype call generated by the call generation model to a diploid homozygous alternate genotype call as an input for the call recalibration machine learning model; and generating, utilizing the call recalibration machine learning model, the first genotype probability and the second genotype probability based further on the diploid homozygous reference genotype call or the diploid homozygous alternate genotype call.

6. The computer-implemented method of any one of claims 1-5, further comprising downsampling diploid sequencing metrics to simulate haploid sequencing metrics corresponding to the haploid nucleotide sequence by: selecting a subset of diploid nucleotide reads from the sample to simulate haploid nucleotide reads; and selecting, based on nucleotide base calls of the subset of diploid nucleotide reads, a subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate genotypes as indicated by a call generation model or as indicated by a ground-truth base-call dataset.

7. The computer-implemented method of claim 6, further comprising selecting, to simulate the haploid sequencing metrics corresponding to the haploid nucleotide sequence, the diploid sequencing metrics corresponding to the subset of genomic coordinates exhibiting homozygous reference genotypes or homozygous alternate genotypes.

8. The computer-implemented method of any one of claims 1-7, wherein the haploid nucleotide sequence comprises a sequence of one or more nucleotide bases from a haploid chromosome or a single chromosome without a counterpart chromosome.

9. The computer-implemented method of any one of claims 1-8, wherein the haploid nucleotide sequence comprises a sequence of one or more nucleotide bases from a sex chromosome.

10. The computer-implemented method of any one of claims 1-9, wherein: determining the sequencing metrics comprises determining one or more of read-based sequencing metrics, externally sourced sequencing metrics, or call model generated sequencing metrics for the nucleotide base calls of nucleotide reads corresponding to the genomic coordinate of the haploid nucleotide sequence; and generating the set of variant call classifications comprises utilizing the call recalibration machine learning model to process the one or more of the read-based sequencing metrics, the externally sourced sequencing metrics, or the call model generated sequencing metrics to generate the set of variant call classifications.

11. The computer-implemented method of claim 10, wherein: the read-based sequencing metrics comprise sequencing metrics derived from nucleotide reads of the sample; the externally sourced sequencing metrics comprise sequencing metrics stored in one or more databases external to a system executing the computer-implemented method; and the call model generated sequencing metrics comprise sequencing metrics generated by a call generation model.

12. The computer-implemented method of any one of claims 1-11, wherein to determine the final nucleotide base call by modifying an initial nucleotide base call generated by a call generation model for a variant call file.

13. The computer-implemented method of any one of claims 1-12, wherein: generating the first genotype probability comprises utilizing a layer of the call recalibration machine learning model to modify a homozygous reference probability of a homozygous reference genotype at the genomic coordinate to generate a haploid reference probability of a reference genotype at the genomic coordinate; and generating the second genotype probability comprises utilizing the layer of the call recalibration machine learning model to modify a homozygous alternate probability of a homozygous alternate genotype at the genomic coordinate to generate a haploid alternate probability of an alternate genotype at the genomic coordinate.

14. A system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to perform a method according to any one of claims 1-13.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a system to perform a method according to any one of claims 1-13.