Variant calling without a targeted reference genome

JP2024546515A5Pending Publication Date: 2026-01-14ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024539855
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2022-12-28
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

The challenge of performing accurate variant calling in the absence of a target reference genome, which often leads to numerous false positive calls due to the lack of available reference genomes for many species and the limitations of existing computational methods.

Method used

The use of non-target and pseudo-target reference genomes for variant calling, combined with ensemble comparison and machine learning classifiers, to identify true positive variants and filter out false positives, utilizing techniques such as random forest models, logistic regression, and neural networks to improve variant quality.

Benefits of technology

This approach significantly reduces false positive variants, enhancing the accuracy of variant calling by leveraging homology thresholds and evolutionary relationships, even when a target reference genome is unavailable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosed technology relates to determining the feasibility of using a reference genome of a non-target species for variant calling of a sample of a target species. In particular, the disclosed technology relates to mapping the sequenced reads of the sample of the target species to a reference genome of a non-target species to detect a first set of variants in the sequenced reads of the sample of the target species, and mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo-target species to detect a second set of variants in the sequenced reads of the sample of the target species.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Priority Application This application claims the benefit and priority of the following:

[0002] U.S. patent application Ser. No. 17 / 952,192, entitled “VARIANT CALLING WITHOUT A TARGET REFERENCE GENOME,” filed on Sept. 23, 2022 (Attorney Docket No. ILLM 1064-2 / IP-2297-US1);

[0003] U.S. Patent Application No. 17 / 952,194, entitled “QUALITY DETECTION OF VARIANT CALLING USING A MACHINE LEARNING CLASSIFIER,” filed on September 23, 2022 (Attorney Docket No. ILLM 1064-3 / IP-2297-US2);

[0004] U.S. Patent Application No. 17 / 952,198, entitled “UNIQUE MAPPER TOOL FOR EXCLUDING REGIONS WITHOUT ONE-TO-ONE MAPPING BETWEEN A SET OF TWO REFERENCE GENOMES,” filed on September 23, 2022 (Attorney Docket No. ILLM 1064-4 / IP-2297-US3);

[0005] U.S. Provisional Patent Application No. 63 / 294,813, entitled “PERIODIC MASK PATTERN FOR REVELATION LANGUAGE MODELS,” filed on December 29, 2021 (Attorney Docket No. ILLM 1063-1 / IP-2296-PRV);

[0006] U.S. Provisional Patent Application No. 63 / 294,816, entitled “CLASSIFYING MILLIONS OF VARIANTS OF UNCERTAIN SIGNIFICANCE USING PRIMATE SEQUENCING AND DEEP LEARNING,” filed on December 29, 2021 (Attorney Docket No. ILLM 1064-1 / IP-2297-PRV);

[0007] U.S. Provisional Patent Application No. 63 / 294,820, entitled “IDENTIFYING GENES WITH DIFFERENTIAL SELECTIVE CONSTRAINT BETWEEN HUMANS AND NONHUMAN PRIMATES,” filed on December 29, 2021 (Attorney Docket No. ILLM 1065-1 / IP-2298-PRV);

[0008] U.S. Provisional Patent Application No. 63 / 294,827, entitled “DEEP LEARNING NETWORK FOR EVOLUTIONARY CONSERVATION,” filed on December 29, 2021 (Attorney Docket No. ILLM 1066-1 / IP-2299-PRV);

[0009] U.S. Provisional Patent Application No. 63 / 294,828, entitled “INTER-MODEL PREDICTION SCORE RECALIBRATION,” filed on December 29, 2021 (Attorney Docket No. ILLM 1067-1 / IP-2301-PRV), and

[0010] U.S. Provisional Patent Application No. 63 / 294,830, entitled “SPECIES-DIFFERENTIABLE EVOLUTIONARY PROFILES,” filed on December 29, 2021 (Attorney Docket No. ILLM 1068-1 / IP-2302-PRV).

[0011] These priority applications are incorporated by reference into this disclosure as if fully set forth.

[0012] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep convolutional neural networks to analyze ordinal data.

[0013] Built-in The following are incorporated by reference for all purposes as if fully set forth herein and should be considered part of this provisional patent application:

[0014] Sundaram,L.et al.Predicting the clinical impact of human mutation with deep neural networks.Nat.Genet.50,1161-1170(2018),

[0015] Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019),

[0016] U.S. patent application Ser. No. 62 / 573,144, entitled “TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV);

[0017] U.S. Patent Application No. 62 / 573,149, entitled “PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV);

[0018] U.S. patent application Ser. No. 62 / 573,153, entitled “DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV);

[0019] U.S. Patent Application No. 62 / 582,898, entitled “PATHOGENICITY CLASSIFICATION OF GENOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV);

[0020] U.S. Patent Application No. 16 / 160,903, entitled “DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-5 / IP-1611-US);

[0021] U.S. Patent Application Serial No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed on October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US);

[0022] U.S. Patent Application Serial No. 16 / 160,968, entitled “SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US);

[0023] U.S. patent application Ser. No. 16 / 160,978, entitled “DEEP LEARNING-BASED SPLICE SITE CLASSIFICATION,” filed on October 15, 2018 (Attorney Docket No. ILLM 1001-4 / IP-1680-US);

[0024] U.S. Patent Application No. 16 / 407,149, entitled “DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed May 8, 2019 (Attorney Docket No. ILLM 1010-1 / IP-1734-US);

[0025] U.S. Patent Application No. 17 / 232,056, entitled “DEEP CONVOLUTIONAL NEURAL NETWORKS TO PREDICT VARIANT PATHOGENICITY USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURES,” filed on April 15, 2021 (Attorney Docket No. ILLM 1037-2 / IP-2051-US);

[0026] U.S. Patent Application No. 63 / 175,495, entitled “MULTI-CHANNEL PROTEIN VOXELIZATION TO PREDICT VARIANT PATHOGENICITY USING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV);

[0027] U.S. Patent Application No. 63 / 175,767, entitled “EFFICIENT VOXELIZATION FOR DEEP LEARNING,” filed on April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV);

[0028] U.S. Patent Application No. 17 / 468,411, entitled “ARTIFICIAL INTELLIGENCE-BASED ANALYSIS OF PROTEIN THREE-DIMENSIONAL (3D) STRUCTURES,” filed on September 7, 2021 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US);

[0029] U.S. Provisional Patent Application No. 63 / 253,122, entitled “PROTEIN STRUCTURE-BASED PROTEIN LANGUAGE MODELS,” filed on October 6, 2021 (Attorney Docket No. ILLM 1050-1 / IP-2164-PRV);

[0030] U.S. Provisional Patent Application No. 63 / 281,579, entitled “PREDICTING VARIANT PATHOGENICITY FROM EVOLUTIONARY CONSERVATION USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURE VOXELS,” filed on November 19, 2021 (Attorney Docket No. ILLM 1060-1 / IP-2270-PRV), and

[0031] U.S. Provisional Patent Application No. 63 / 281,592, entitled “COMBINED AND TRANSFER LEARNING OF A VARIANT PATHOGENICITY PREDICTOR USING GAPED AND NON-GAPED PROTEIN SAMPLES,” filed on November 19, 2021 (Attorney Docket No. ILLM 1061-1 / IP-2271-PRV). [Background technology]

[0032] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to embodiments of the claimed technology.

[0033] The proliferation of available biological sequence data has led to multiple computational approaches to infer the three-dimensional structure, biological function, fitness, and evolutionary history of proteins from sequence data. So-called protein language models, such as models based on the Transformer architecture, have been trained on large ensembles of protein sequences by using a masked language modeling strategy that fills in masked amino acids in the sequence by considering the surrounding amino acids.

[0034] Protein language models capture long-range dependencies, learn rich representations of protein sequences, and can be employed for multiple tasks. For example, protein language models can predict structural contacts from a single sequence in an unsupervised manner.

[0035] Protein sequences can be grouped into families of homologous proteins that are derived from ancestral proteins and share similar structures and functions. Analysis of multiple sequence alignments (MSA) of homologous proteins provides important information about functional and structural constraints. Statistics of MSA strings, which represent amino acid sites, identify functional residues that are conserved during evolution. Correlation of amino acid usage between MSA strings contains important information about functional sectors and structural contacts.

[0036] Language models were initially developed for natural language processing and operate on a simple but powerful principle. Language models acquire language understanding by learning to fill in missing words in sentences, similar to sentence completion tasks in standardized tests. By applying this principle across large text corpora, language models develop powerful inference capabilities. The Bidirectional Encoder Representations from Transformers (BERT) mode exemplified this principle using Transformers, a class of neural networks in which attention is a key component of the learning system. In Transformers, each token in the input sentence can "join" every other token by exchanging the corresponding activation patterns with the intermediate outputs of neurons in the neural network.

[0037] Protein language models such as the MSA Transformer have been trained to make inferences from the MSAs of evolutionarily related sequences. The MSA Transformer alternates sequence-wise ("row") attention with site-wise ("column") attention to incorporate co-evolution. The combination of row-attention heads with the MSA Transformer has led to state-of-the-art unsupervised structural contact prediction.

[0038] An end-to-end deep learning approach for variant effect prediction is applied to predict the pathogenicity of missense variants from protein sequence and sequence conservation data (see Sundaram, L. et al. Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), referred to herein as "PrimateAI"). PrimateAI uses deep neural networks trained on variants of known pathogenicity with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences and uses a trained deep neural network to compare the differences and determine the pathogenicity of the variant. Such an approach utilizing protein sequences for pathogenicity prediction is promising because it can avoid the circular problem and overfitting to prior knowledge. However, the number of clinical data available in ClinVar is relatively small compared to the sufficient number of data to effectively train a deep neural network. To overcome this data scarcity, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.

[0039] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in reproducing prior knowledge in ClinVar. These results suggest that PrimateAI is an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.

[0040] Therefore, an opportunity arises to use protein language models and MSA for variant pathogenicity prediction. More accurate variant pathogenicity predictions may be obtained. [Brief description of the drawings]

[0041] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.

[0042] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Figure 1] FIG. 1 is a flow diagram illustrating the process of the system for variant calling for a specific target species 120 without the use of a target reference genome. [Diagram 2] FIG. 1 is a sequence flow diagram illustrating the process of identifying false positive variants by ensemble comparison to identify true positive variants from sequenced reads. [Diagram 3] Illustrated is an example of a genetic variant from an exemplary reference gene sequence A with two exemplary variant sequences each having a single nucleotide variant at a single base position, but otherwise having the same composition as the reference sequence. [Figure 4] FIG. 1 illustrates the process of base-resolved variant calling sensitivity detection in a series of exemplary flow diagrams for mapping sequenced reads to a reference genome. [Diagram 5] FIG. 1 is a schematic illustrating the process of mapping sequenced reads to a reference genome. [Figure 6]FIG. 1 depicts a graphical flow diagram of a process for surrogate detection of a second set of variants in one embodiment of the disclosed technology where a target species reference genome is available. [Figure 7] FIG. 1 is a schematic illustrating the evolutionary relationships between target and pseudo-target species using a simplified phylogenetic tree diagram. [Figure 8] FIG. 1 is a schematic illustrating the evolutionary relationships between target and pseudo-target species using a simplified phylogenetic tree diagram. [Figure 9] FIG. 1 is a schematic diagram depicting genome-resolved variant calling sensitivity detection of sequenced reads. [Figure 10] FIG. 1 is a schematic flow diagram demonstrating how output data generated by a set comparison of sequenced reads mapped to non-targeted and pseudo-targeted reference genomes can be used as a training data set for a quality classifier to predict variant quality when a reference genome is not available, in contrast to methods where quality is determined by mapping to a non-targeted or pseudo-targeted genome. [Figure 11] Illustrative examples of variant features in a plurality of variant features describing guanine-cytosine content. [Figure 12] Illustrative examples of variant features in a plurality of variant features describing local compositional complexity. [Figure 13] Illustrative examples of variant features in a multivariate feature describing allele counts. [Figure 14] Illustrative examples of variant features in a plurality of variant features illustrating the process of mapping sequenced reads to a reference genome, with an additional step added to calculate a quality metric of the mapping. [Figure 15] Illustrative examples of variant features in a plurality of variant features that, when mapped to a reference genome, account for strand bias in sequenced reads. [Figure 16]Illustrative examples of variant features in a plurality of variant features that account for the depth and coverage of sequenced reads mapped to a reference genome. [Figure 17] FIG. 1 is an illustrated flow diagram representing a variant quality classifier configured as a random forest model for classifying target variants as belonging to either a high quality class or a low quality class. [Figure 18] FIG. 1 is an illustrated flow diagram showing a variant quality classifier configured as a logistic regression model for classifying target variants as belonging to either a high quality class or a low quality class. [Figure 19] FIG. 1 is an illustrated flow diagram showing a variant quality classifier configured as a neural network model for classifying a target variant as belonging to either a high quality class or a low quality class. [Figure 20] FIG. 13 is a flow diagram of the unique mapper overview process to further improve the quality of the set of called variants following variant calling or machine learning variant classification. [Figure 21] 1 illustrates a schematic of a gene annotation filter in a series of cascaded filters for variant quality filtering. [Figure 22] 1 illustrates generally the process of codon transcription and translation and filtering of codon matches. [Figure 23] 1 illustrates the process of filtering genes based on the distribution of machine learning scores. [Figure 24] Illustrates deviations from Hardy-Weinberg equilibrium in an exemplary population. [Diagram 25] FIG. 1 is an illustration of nonsense variants. [Figure 26] Included is a graph of collected results demonstrating the effect of the cascade filter on the number of nonsense variants per sample. [Figure 27]Included is a graph of collected results demonstrating the effect of the cascade filter on the missense:synonymous ratio of called variants per sample. [Figure 28] Included is a graph of collected results demonstrating the cascade filter effect on the number of insertion-deletion variants (indels) per sample. [Figure 29] FIG. 1 illustrates an exemplary computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0043] The following discussion is presented to enable those skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0044] The detailed description of the various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures show diagrams of functional blocks of the various embodiments, the functional blocks do not necessarily show a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., a module, a processor, or a memory) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or a block of random access memory, a hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine in an operating system, may be a function in an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentalities shown in the figures.

[0045] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided in exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers or servers, or even spread among many different processors, computers or servers. In addition, it will be understood that some of the modules can be operated in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures can also be considered as flow chart steps in a method. Also, a module does not necessarily have to have all code located contiguously in memory. Some portions of code can be separated from other portions of code, with code from other modules or other functions located between them.

[0046] The disclosed technology can be used to improve the quality of pathogenic variant calling. The disclosed technology can be used to improve the quality of variant calling in scenarios where the desired reference genome is not available. There are 8.7 million species in the world, but very few have reference genome constructs. In many scenarios, we need to perform variant calling in the absence of reference genome constructs. In some instances, we could select closely related species as reference genomes for variant calling. However, this often leads to many false positive calls. Therefore, we developed various methods to reduce false positives, including random forest classifiers, linear regression models, and neural network models. We also devised a unique mapper score to identify regions that are not one-to-one mapping between species, which further reduces variant calling errors.

[0047] In some embodiments, the disclosed technology relates to determining the feasibility of using a non-target species reference genome for variant calling of a target species sample. In particular, the disclosed technology relates to mapping the sequenced reads of the target species sample to the non-target species reference genome to detect a first set of variants in the sequenced reads of the target species sample, and mapping the sequenced reads of the target species sample to the pseudo target species reference genome to detect a second set of variants in the sequenced reads of the target species sample.

[0048] Further, in certain embodiments, the disclosed technology relates to variant calling of sequenced reads of a sample of a target species against a reference genome of a pseudo target species. Low quality variants are identified as false positive variants that are present in the second set of variants but not present in the first set of variants.

[0049] Additionally, in some cases, the disclosed technology relates to a variant quality classifier configured to process multiple features of a target variant and generate a quality indicator for a quality variant. The variant quality classifier is trained on a set of high quality variants and a set of low quality variants. The high quality variants are identified as true positive variants that are common between a first set of variants detected by variant calling of sequenced reads of a sample of a target species against a reference genome of a non-target species and a second set of variants detected by variant calling of sequenced reads of a sample of a target species against a reference genome of a pseudo target species. The low quality variants are identified as false positive variants that are present in the second set of variants but not present in the first set of variants.

[0050] Variant calling using an untargeted reference genome FIG. 1 is a flow diagram illustrating a process 100 of a system for variant calling of a particular target species 120 without the use of a target reference genome. A first set of variants in sequenced reads of a target species 104 is detected by mapping sequenced reads from the target species 120 to a non-target reference genome 102. The non-target reference genome 102 is derived from a non-target species other than the target species 120. In some embodiments of the disclosed technology, the non-target reference genome 102 is non-homologous to the genome of the target species 120 as determined by a homology threshold (e.g., percentage homology less than 30%, 40%, or 50%, or a double bound range of acceptable homology percentages, such as 30-40% or 40-50%). In certain embodiments, the non-target species and the target species 120 belong to the same taxonomic genus, family, order, or class.

[0051] A second set of variants in the sequenced reads of the target species 144 is detected by mapping the sequenced reads from the target species 120 to the pseudo target reference genome 142. The pseudo target reference genome 142 is derived from a pseudo target species other than the target species 120. In some embodiments of the disclosed technology, the pseudo target reference genome 142 is homologous to the genome of the target species 120 as determined by a homology threshold (e.g., percentage homology greater than 80%, 90%, or 95%, or a double bound range of acceptable homology percentages, such as 85-90% or 80-89%). The homology threshold set for determining the degree of homology between the pseudo target species and the target species may be the same as the homology threshold set for determining the degree of homology between the non-target species and the target species, or the respective homology thresholds may be different. In some embodiments, the homology threshold set for determining the degree of homology between the non-target species and the target species may be informed by the degree of homology between the pseudo target species and the target species, or vice versa. Comparison 126 of the first set of variants and the second set of variants identifies a subset of false positive variants 128 (i.e., overlapping variants identified by mapping to the pseudo target reference genome 142 cannot be considered reliable positive variants based on homology if the variant is also identified by mapping to the non-target reference genome 102).

[0052] FIG. 2 is a sequence flow diagram 200 depicting a process of identifying false positive variants by set comparison to identify true positive variants from sequenced reads. A first set of variants 223 is detected by mapping sequenced reads from a target species 202 to a non-target species reference genome 204. A second set of variants 227 is detected by mapping sequenced reads from a target species 206 to a pseudo target species reference genome 208. The Venn diagram represents the union variant set 1 ∪ variant set 2, where the difference set variant set 1 - variant set 2 is represented by the lower left shaded region 244 and the difference set variant set 2 - variant set 1 is represented by the upper right shaded region 246. The intersection of variant set 1 ∩ variant set 2 is represented by the central diamond cross-hatched region for the set of true positive variants 266. Intersection variant set 1 ∩ variant set 2 (i.e., called variants identified in both the first set 223 of variants detected by mapping to the non-target reference genome 204 and the second set 227 of variants detected by mapping to the pseudo target reference genome 208) translates to a set of true positive variants 266. Difference set variant set 2 minus variant set 1 (i.e., called variants identified in the second set 227 of variants detected by mapping to the pseudo target reference genome 208 but not in the first set 223 of variants detected by mapping to the non-target reference genome 204) translates to a set of false positive variants 268.

[0053] 3 illustrates an example of a genetic variant 300 from an exemplary reference genetic sequence A302 having two exemplary genetic variant sequences, 1A322 and 1B342, where the variant sequences have a respective single nucleotide variant at a single base position, but otherwise have the same composition as the reference sequence. For example, single nucleotide substitutions are shown as adenine 326 in variant 1A322 and thymine 336 in variant 1B342, compared to cytosine 306 in reference genetic sequence A302.

[0054] FIG. 4 diagrams the process 400 of base-resolved variant calling sensitivity detection in a series of exemplary flow diagrams of mapping sequenced reads to a reference genome. Sequenced read X 410 is mapped to the pseudo target reference genome 402 and called as a variant due to the A→C single nucleotide variant at position 5. Sequenced read X 410 is also mapped to the non-target reference genome 404 and is not called as a variant because no single nucleotide variant is identified. As a result, sequenced read X 410 belongs to the difference set between the called variant set from mapping to the pseudo target reference genome 402 and the variant set from mapping to the non-target reference genome 404, and thus sequenced read X 410 is a false positive. Sequenced read Y 412 is mapped to the pseudo target reference genome 422 and called as a variant due to the A→C single nucleotide variant at position 5. Sequenced read Y412 is also mapped to the non-targeted reference genome 424 and called as variant due to the A→C single nucleotide variant at position 5. As a result, sequenced read Y412 belongs to the intersection set between the variant set called from mapping to the pseudo-targeted reference genome 422 and the variant set from mapping to the non-targeted reference genome 424, and thus sequenced read Y412 is a true positive.

[0055] Sequenced read Z is mapped to the pseudo target reference genome 442 and is not called as a variant even though cytosine and guanine are not equivalent at position 5. Due to base pairing, the complementary strand of the pseudo target reference genome 442 has a cytosine at position 5 and the complementary strand of sequenced read Z 414 has a guanine at position 5. As a result, this sequenced read Z 414 is not a variant when mapped to the pseudo target reference genome 442. Sequenced read Z 414 is also mapped to the non-target reference genome 444 and is not called as a variant due to the complementary base present at position 5. As a result, sequenced read Z 414 belongs to the complement of both the called variant set from mapping to the pseudo target reference genome 442 and the called variant set from the non-target reference genome 444, and therefore sequenced read Z 414 is a true negative variant.

[0056] 5 is a schematic diagram 500 illustrating the process of mapping sequenced reads to a reference genome. Sequenced reads 502 are mapped to a reference genome 505, resulting in a mapping 555. In the mapping 555, each sequenced read from the set of sequenced reads 502 is aligned with a given genomic region in the reference genome 505. As seen in the mapping 555, the mapped sequenced reads may be aligned with genomic regions that are mutually exclusive (i.e., non-overlapping, such as the left-most sequenced read and the middle sequenced read in the mapping 555) or that are not mutually exclusive (i.e., overlapping, such as the middle sequenced read and the right-most sequenced read in the mapping 555).

[0057] 6 depicts a graphical flow diagram 600 of a process for alternative detection of variant set 2 in one embodiment of the disclosed technology when a target species reference genome 622 is available. Sequenced reads from a target species 602 are mapped to the target species reference genome 622. The mapped sequence reads of the target species are lifted over to a non-target species reference genome 642. Variants detected in the non-target species reference genome 642 but not in the target species reference genome 622 comprise variant set 2 662, and multiple variants in variant set 2 662 are false positive variants. Alternative detection of variant set 2 662 is useful in generating training data that includes known ground truth data for use in training machine learning classifiers for detection of true positive and false positive variants.

[0058] FIG. 7 is a schematic diagram 700 illustrating the evolutionary relationship between the target species and the pseudo target species using a simplified phylogenetic tree diagram. In phylogenetic tree A702, the pseudo target species and the target species are different species that are not orthologous. In one embodiment of the disclosed technology, the evolutionary relationship between the target species and the pseudo target species reflects that shown in phylogenetic tree A702. In phylogenetic tree B704, the pseudo target species and the target species are different species that are orthologous. In one embodiment of the disclosed technology, the evolutionary relationship between the target species and the pseudo target species reflects that shown in phylogenetic tree B704. In phylogenetic tree C706, the pseudo target species and the target species are the same species. In one embodiment of the disclosed technology, the evolutionary relationship between the target species and the pseudo target species reflects that shown in phylogenetic tree C706.

[0059] FIG. 8 is a schematic diagram 800 illustrating the evolutionary relationship between the target species and the pseudo target species using a simplified phylogenetic tree diagram. In the phylogenetic tree D802, the non-target species and the target species are different species, where the non-target species is human and the target species is a non-human primate. In one embodiment of the disclosed technology, the evolutionary relationship between the target species and the non-target species reflects the evolutionary relationship of the phylogenetic tree D802. In one embodiment of the disclosed technology, the primate species sample and the reference genome are utilized to infer the pathogenicity of orthologous human variants by variant calling against closely related primate species genomes, variant calling against non-target non-homologous primate species genomes, and contrasting the results as set forth in FIGs. 1-5. In some embodiments of the disclosed technology, a machine learning classifier is trained to detect false positive variants to further refine the identification process of true positive variants.

[0060] 9 is a schematic diagram 900 showing genome-resolved variant calling sensitivity detection of sequenced reads. A sequenced read A 902 from a target species is equivalent to a sequenced read A 942 from a target species and is located in region 1 922 in a non-targeted reference genome 924 and region 2 962 in a pseudo-targeted reference genome 964. Region 1 922 and region 2 962 are not equivalent (i.e., not orthologous), and thus sequenced read A is not located in the same genomic region in the non-targeted reference genome 924 and pseudo-targeted reference genome 964. Despite mapping to both the non-targeted reference genome 924 and pseudo-targeted reference genome 964, sequenced read A 902 does not result in a called variant in a genomic region in the non-targeted reference genome 964 that is orthologous to the genomic region where sequenced read A 942 is located in the pseudo-targeted reference genome 964. As a result, this variant belongs to the set of variants called from mapping to the pseudo-targeted reference genome 964, but does not belong to the set of variants called from mapping to the non-targeted reference genome 924, resulting in a false positive.

[0061] Sequenced read B982 from the target species, sequenced read B984 from the target species, and sequenced read B986 from the target species are equivalent. Region 3 972, region 4 978, and region 5 980 belong to the non-targeted reference genome and are not equivalent. Sequenced read B982 from the target species is located in multiple regions in the non-targeted reference genome. Similar to sequenced read A902, sequenced read B982 is located in a genomic region in the non-targeted species reference genome that is different from the orthologous genomic region where sequenced read B982 is located in the pseudo-targeted reference genome due to the diversity of variant calling in the non-targeted reference genome. Then, a sequenced read located in three or more genomic regions in the non-targeted reference genome will result in a false positive.

[0062] Machine Learning Classifier FIG. 10 is a schematic flow diagram 1000 demonstrating how output data generated by a set comparison of sequenced reads mapped to non-target and pseudo-target reference genomes can be used as a training data set for a quality classifier to predict variant quality when a reference genome is not available, as opposed to a method in which quality is determined by mapping to a non-target or pseudo-target genome. Sequenced reads from a target species 1002 are mapped to a reference genome as previously described in FIG. 1, FIG. 2, and FIG. 3. The intersection of variant set 1 1004 and variant set 2 1006 corresponds to a set of true positive variants 1022. The difference set between variant set 2 1006 and variant set 1 1004 (i.e., variants present in variant set 2 1006 but absent from variant set 1 1004) corresponds to a set of false positive variants 1008. The set of true positive variants 1022 is further coded as a set of high quality variants 1024. The set of false positive variants 1008 is further coded as a set of low quality variants 1010. The combined set of high quality variants 1024 and low quality variants 1010 comprises a set of ground truth data 1020.

[0063] The quality classifier 1064 undergoes a model training process 1040 on the ground truth data 1020. The quality classifier 1064 is configured to receive a vector {x1:x2} that includes a set of variant features in the plurality of variant features. n It takes an input target variant 1062 represented as x, where each value of x is a variant feature in a set of variant features in a plurality of variant features that describe the target variant 1062. In some embodiments of the disclosed technology, additional variant features can be extracted from a variant call format (.vcf) file. The quality classifier 1064 is a binary classification model with output classes for high quality 1066 and low quality 1068.

[0064] FIG. 11 is an illustrative example 1100 of a variant feature in a plurality of variant features describing guanine-cytosine content. A short gene sequence B1102 contains a percentage of adenine, thymine, guanine, and cytosine nucleic acids. The guanine-cytosine content (GC) of a gene sequence corresponds to the percentage of guanine and cytosine nucleic acids in the sequence. GC content is a physiochemical descriptor of a nucleic acid sequence that can be used as a proxy for the thermal stability of the sequence due to the different behavior of chemical bonds compared to the behavior of adenine-thymine bonds. GC content affects read coverage in next generation sequencing applications. Formula 1122 is used for a sample calculation of the GC content of gene sequence B1102, where GC content is equal to the ratio of guanine and cytosine counts to the total count of all nucleic acids. Formula 1124 is used for a sample calculation of gene skew for gene sequence B 1102, where GC skew is determined as the ratio of the difference between guanine and cytosine counts to the sum of guanine and cytosine counts for a given window size. Window size example 1164 illustrates a window size of 5. When a window size of 5 is applied to gene sequence B 1102, GC skew is calculated as shown in table 1184.

[0065] FIG. 12 is an illustrative example 1200 of a variant feature in a plurality of variant features that describes local compositional complexity. Local compositional complexity is a measure of entropy within a gene sequence. Gene sequence X 1202 does not contain variability in nucleic acid composition and therefore has low entropy. Gene sequence Z 1242 has high variability in nucleic acid composition and therefore has high entropy. Gene sequence Y 1222 contains more variability than gene sequence X 1202, but less than gene sequence Z 1242 and therefore can be described as having a moderate (i.e., moderate) level of entropy. Equation 1224 calculates the entropy of a gene sequence in the form of a local compositional complexity. A sample calculation for gene sequence B 1204 results in an entropy value of 1.92, where entropy is equivalent to the sum of the log probabilities scaled by the same respective probability for each nucleic acid.

[0066] 13 is an illustrative example 1300 of a variant feature in a plurality of variant features describing allele counts. Variant 1 1302 is shown shaded gray and variant 2 1304 is shown in white. Population 1322 contains a number of gene sequence samples that belong to either variant 1 1302 or variant 2 1304. Within population 1322, there are a total of six samples that belong to variant 1 1302, and therefore the total allele count of variant 1 1302 is six. Within population 1322, there are a total of nine samples that belong to variant 2 1304, and therefore the total allele count of variant 1 1304 is nine. The error rate of detecting heterozygous called variants is higher than the corresponding error rate of homozygous called variants (i.e., heterozygous false positives occur at a higher rate than homozygous false positives).

[0067] FIG. 14 is an illustrative example 1400 of variant features in a plurality of variant features that describes the process of mapping sequenced reads to a reference genome, with an additional step added to calculate a quality metric of the mapping. A sequenced read 1402 is mapped to a reference genome 1404 to generate a mapping 1444. The mapping quality score quantifies the likelihood of a sequenced read being misaligned to the reference genome. The mapping quality is judged by the sum of possible alignments for a given sequenced read and the count of mismatched base pairs within the alignment. The mapping quality score is reported in Phred scale, a commonly used logarithmic data scaling technique for error rates in sequence analysis.

[0068] FIG. 15 is an illustrative example 1500 of a variant feature in a plurality of variant features that, when mapped to a reference genome, describes strand bias in a sequenced read. A sequenced read 1502 includes reads with different strand orientations (i.e., a strand oriented in a 5'→3' direction and a strand oriented in a 3'→5' direction). When the sequenced read 1502 is mapped to a reference genome 1504, the generated mapping 1544 displays a sequencing bias based on strand orientation in which one DNA strand is favored over the other. Strand bias can result in a higher error rate for allele counts.

[0069] FIG. 16 is an illustrative example 1600 of variant features in a plurality of variant features describing the depth and coverage of sequenced reads mapped to a reference genome. The depth and coverage of a particular mapping are measures of mapping quality, and both the sequencing coverage and the depth of sequencing coverage are metrics proportional to the quality of a particular mapping. Sequenced reads 1602 are mapped to a reference genome 1604 at various genomic regions along the X-axis. The total percentage of target bases in the reference genome to which the sequenced reads are mapped is quantified as the coverage of the genome. The average depth of sequencing coverage is the ratio of the number of reads scaled by the read length to the total reference genome length. This concept is illustrated by visualizing the X-axis as the length of the reference genome 1604 with coverage corresponding to the total breadth of the aligned sequenced reads 1602, while the Y-axis visualizes the depth to which the reference genome 1604 is covered.

[0070] 17 is an example flow diagram 1700 illustrating a variant quality classifier configured as a random forest model for classifying a target variant 1762 as belonging to either a high quality class 1766 or a low quality class 1768. As a quality classifier, the random forest model 1744 is configured to generate a vector {x1:x2} that includes a set of variant features in a plurality of variant features. n It takes an input target variant 1762 represented as {x,y,y,z}, where each value of x is a variant feature in a set of variant features in a plurality of variant features that explain the target variant 1762, and generates a classification from a random forest model 1744. In the random forest model 1744, each of a plurality of decision trees generates an output result for each of the target variant classes, and a final result is generated via a majority average.

[0071] 18 is an example flow diagram 1800 illustrating a variant quality classifier configured as a logistic regression model for classifying a target variant 1862 as belonging to either a high quality class 1866 or a low quality class 1868. The quality classifier 1844 receives a vector {x1:x2} that includes a set of variant features in the plurality of variant features. n The logistic regression model 1844 takes an input target variant 1862 represented as {x,y}, where each value of x is a variant feature in a set of variant features in a plurality of variant features that describe the target variant 1862, and generates a classification from a logistic regression model 1844. In the logistic regression model 1844, the model generates output values ​​in the range {0,1}, and a decision threshold boundary determines whether the input value (i.e., the target variant 1862) is classified as an output of 0 or 1 (e.g., a decision threshold boundary of 0.5 results in values ​​in the range {0,0.4} generating an output of 0 and values ​​in the range {0.5,1} generating an output of 1). The determination of the optimal decision threshold boundary may be determined based on optimizing a particular performance metric, such as accuracy, precision, recall, or a particular error function, when training the logistic regression model. The binary output values ​​0 and 1 are assigned to two output classes, a high quality class 1866 or a low quality class 1868.

[0072] 19 is an example flow diagram 1900 illustrating a variant quality classifier configured as a neural network for classifying a target variant 1962 as belonging to either a high quality class 1966 or a low quality class 1968. The quality classifier 1944 receives a vector {x1:x2} that includes a set of variant features in the plurality of variant features. nThe neural network model takes an input target variant 1962 represented as {x,y,y}, where each value of x is a variant feature in a set of variant features in a plurality of variant features that describe the target variant 1962, and generates a classification from the neural network 1944. The neural network model processes the input target variant 1962 through a series of connected layers of nodes, each performing a respective weighted data transformation. Backpropagation through the network iteratively updates the weights of each node during the training process, and the final trained model generates an output that for the input target variant 1962 belongs to a high quality class 1966 or a low quality class 1968. At this stage, variants identified as high quality or low quality may be subjected to further filtering steps in certain embodiments, as described further below.

[0073] Unique Mapper FIG. 20 is a flow diagram 2000 of an overview process of a unique mapper to further improve the quality of the set of called variants following variant calling or machine learning variant classification. The sequenced reads 2002 of the target species sample undergo a filtering process via cascade filters 2004 to remove low quality sequenced reads 2024. The sequence data can be obtained from a binary alignment map (.bam) file. The series of cascade filters 2004 includes filters to remove variants with incorrect codon matches between primates and humans, filters to remove variants with annotation errors, gene-specific filters (e.g., biased distribution of variant machine learning classifier quality scores compared to the whole exosome score or deviation from Hardy-Weinberg equilibrium), and removal of variants that do not meet certain machine learning classifier performance metric thresholds. The resulting intermediate set of sequenced reads 2006 is mapped to a pseudo-targeted reference genome 2008 and a non-targeted reference genome 2026. The pseudo target reference genome 2008 is divided into a number of bins (i.e., contiguous non-overlapping genomic regions of a specified equal length). The non-target reference genome 2026 is also divided into an equal number of bins of equal size compared to the pseudo target reference genome 2008 bins. The bins are compared on a one-to-one basis to determine the degree of mapping homology between corresponding bins. The best mapped bin is identified as the bin whose degree of match (i.e., alignment between the mapped genomes for the bin) is used to generate a unique mapper score 2040. In one embodiment of the disclosed technology, the unique mapper score 2040 is unique to each particular sample, and the unique mapper scores across all samples for a particular reference target species are averaged to obtain a single average unique mapper score that is applied to all variants of the reference target species that fall into each respective bin.

[0074] FIG. 21 illustrates a schematic of a gene annotation filter 2100 in a series of cascaded filters for variant quality filtering. Gene annotation includes labeling of genomes for features such as gene location, coding and non-coding regions, and various descriptors of gene function. Incorrect gene annotation may result in errors in the variant calling process. Gene A 2102 is located in a genomic region and contains feature X at a specific location within its structure. Gene A 2104 is the correct gene annotation for gene A 2102, and the genomic structure is correctly annotated with properly placed feature X. Gene B 2106 is located in a different location and contains a different structure (i.e., contains feature Y but not feature X). In the case of incorrect gene annotation due to gene prediction error, gene A 2102 may be incorrectly annotated as gene B 2106. As a result, any resulting mapping to gene B 2106 (i.e., the genomic sequence belongs to gene A despite being labeled as gene B) is incorrect. Called variants that map to genes with annotation errors are excluded.

[0075] FIG. 22 illustrates generally the codon transcription and translation and filtering for codon matches of process 2200. Gene sequence A2202 is comprised of nucleic acids. The nucleic acid sequence within gene sequence A2202 undergoes transcription to produce mRNA transcript A2242. After transcription, mRNA transcript A2242 is translated to produce amino acid sequence A2262. Each amino acid is translated from three nucleic acid sequences, called codons, as highlighted by grey shaded boxes for a total of five codons across gene sequence A2202, mRNA transcript A2242, and amino acid sequence A2262. Codon B2282 and codon C2284 comprise the same nucleic acid and are therefore transcribed and translated into the same amino acid. If the non-target reference genome contains codon B2282 and the pseudo target reference genome contains codon C2284 at the same aligned position, then these codons match and the codon mismatch filter will not remove called variants that align to the genomic regions corresponding to codon B2282 and codon C2284. Codons D2286 and E2288 differ at a third nucleic acid position and are not transcribed and translated into the same amino acid. If the non-target reference genome contains codon D2286 and the pseudo target reference genome contains codon E2288 at the same aligned position, then these aligned codons do not match and called variants that align to the genomic regions corresponding to codon D2286 and codon E2288 are excluded.

[0076] FIG. 23 illustrates a process 2300 for filtering genes based on the distribution of machine learning scores. Scores from the variant quality classifier are plotted on a graph measuring frequency for both specific genes, represented by the specific gene distribution 2304, and the entire exosome, represented by the entire exosome distribution 2302. The Wilcoxon rank sum test determines whether the specific gene distribution 2304 is biased compared to the entire exosome distribution 2302 via a significance test that the probability that a randomly selected machine learning score from the specific gene distribution 2304 is greater than a randomly selected machine learning score from the entire exosome distribution 2302 is equal to the probability that a randomly selected machine learning score from the entire exosome distribution 2302 is greater than a randomly selected machine learning score from the specific gene distribution 2304. Called variants that map to genes determined to be biased relative to the entire exosome distribution 2302 are filtered out. Genes are identified as outliers with potential errors by determining bias when comparing the overall exosome distribution 2302 with the distribution of a particular gene 2304.

[0077] FIG. 24 illustrates deviations from Hardy-Weinberg equilibrium in an exemplary population 2400. Filters in the cascade filter 2004 remove variants that deviate from Hardy-Weinberg equilibrium. Dominant alleles are represented by the letter "p" and recessive alleles are represented by the letter "q". Homozygous dominant genotypes (i.e., "pp" 2402) are represented by circles with an upward slash. Heterozygous genotypes ("pq" 2404) are represented by diamond-shaped cross-hatched circles. Homozygous recessive genotypes ("qq" 2406) are represented by circles with a downward slash. The population shown includes 25 samples with each genotype. In the first generation 2442, each genotype has a respective frequency counted as a proportion of each genotype relative to the total population count. In the second generation 2444, each genotype has an updated respective frequency counted as a proportion of each genotype relative to the total population count. A population in which genotype frequencies do not change in successive generations is considered to be in Hardy-Weinberg equilibrium. The genotype frequencies for the example population shown in Figure 24 are different in the second generation 2444 and the first generation 2442, and therefore the population deviates from Hardy-Weinberg equilibrium. Deviance from Hardy-Weinberg equilibrium can result in an overcalling of heterozygous genotypes, resulting in the elimination of called variants that map to genes that are not in Hardy-Weinberg equilibrium as determined by the large population database.

[0078] FIG. 25 is a diagram 2500 of nonsense variants. Filters in the cascade filter 2004 remove nonsense variants. Nonsense variants (also called "stop-gain variants") result from single nucleotide polymorphisms that change the codon sequence such that a codon that previously translated into an amino acid translates into a stop codon as a result of a new type of mutated amino acid sequence. The premature stop codon prevents the remainder of the mRNA transcript from being translated, resulting in the amino acid sequence being prematurely terminated. Gene sequence B2502 is transcribed into mRNA transcript B2522, which is translated into amino acid sequence B2542 for a total of five codons. Single nucleotide polymorphism 2540 at position 12 results in a change from a guanine nucleic acid to a thymine nucleic acid. As a result, the fourth codon changes from ACG to ACT, which is then transcribed into a stop codon rather than being transcribed and translated into a cysteine ​​amino acid residue. The premature stop codon terminates translation and the fifth codon is never translated. As further shown by Figure 25, as a result of single nucleotide polymorphism 2540, gene sequence B2562 is transcribed into mRNA transcript B2582, which is translated into amino acid sequence B2510 for a total of four codons, where the fourth codon is a stop codon.

[0079] Figure 26 includes a graph 2600 of collected results demonstrating the effect of cascade filters on the number of nonsense variants per sample. In one embodiment of the disclosed technology, the number of nonsense variants per sample is compared between samples from nonhuman primate species and humans. Not filtering the variants called from samples from nonhuman primate species results in a significantly higher number of nonsense variants per sample compared to the corresponding human level of nonsense variants per sample. The variants called from samples from nonhuman primate species are subjected to cascade filters including a codon match filter, a gene annotation error filter, a machine learning distribution skew filter, a Hardy-Weinberg imbalance filter, a unique mapper filter (called variants with a unique mapper score of less than 0.6 are removed), and a random forest score filter (called variants with a random forest score of more than 0.17 are removed). Box plots showing the average number of stop-gain variants per sample for each primate reference species, gradually decreasing to near human levels after a series of variant filtering steps, including requiring codon matches, removing SNPs from poorly annotated genes or genes with biased random forest (RF) score distributions or that deviate from Hardy Weinberg equilibrium, and removing SNPs with unique mapper scores <0.6 or RF scores >0.17. Each dot represents the average number of stop-gain variants for each primate reference species. The horizontal line indicates the average number of stop-gain variants for human samples from the Platinum Genome Project.

[0080] Figure 27 includes a graph 2700 of collected results demonstrating the effect of cascade filters on the missense:synonymous ratio of called variants per sample. In one embodiment of the disclosed technology, the missense:synonymous ratio (MSR), a ratio used to estimate the balance of benign and pathogenic variants present in a particular cohort, is compared between samples from non-human primate species and samples from humans. The called variants from samples from non-human primate species are subjected to cascade filters including a codon match filter, a gene annotation error filter, a machine learning distribution skew filter, a Hardy-Weinberg imbalance filter, a unique mapper filter (called variants with a unique mapper score less than 0.6 are removed), and a random forest score filter (called variants with a random forest score greater than 0.17 are removed). The box plot showing the missense:synonymous ratio was reduced after the variant filtering step. Each dot represents the MSR of each primate reference species. The black line represents the MSR of human samples.

[0081] Figure 28 includes a graph 2800 of collected results demonstrating the effect of cascade filters on the number of insertion-deletion variants (indels) per sample. Called variants from samples derived from non-human primates are subjected to cascade filters including a gene annotation error filter, a machine learning distribution skew filter, a Hardy-Weinberg imbalance filter, and a unique mapper filter (called variants with a unique mapper score less than 0.6 are removed). The average number of indels per sample for each primate reference species was reduced after the filtering steps.

[0082] Computer Systems 29 illustrates an exemplary computer system 2900 that can be used to implement the disclosed techniques. The computer system 2900 includes at least one central processing unit (CPU) 2924 that communicates with several peripheral devices via a bus subsystem 2922. These peripheral devices can include, for example, a storage subsystem 2910 including memory devices and a file storage subsystem 2918, user interface input devices 2920, user interface output devices 2928, and a network interface subsystem 2926. The input and output devices enable user interaction with the computer system 2900. The network interface subsystem 2926 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0083] In one embodiment, the random forest model 1744 is communicatively linked to a storage subsystem 2910 and a user interface input device 2920.

[0084] User interface input devices 2920 can include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into computer system 2900.

[0085] The user interface output devices 2928 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 2900 to a user or to another machine or computer system.

[0086] The storage subsystem 2910 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 2930.

[0087] The processor 2930 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 2930 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 2930 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX29 Rackmount Series™, NVIDIA DGX-1™, Microsoft' Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.

[0088] The memory subsystem 2912 used in the storage subsystem 2910 may include several memories including a main random access memory (RAM) 2914 for storing instructions and data during program execution, and a read only memory (ROM) 2916 in which fixed instructions are stored. The file storage subsystem 2918 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 2918 in the storage subsystem 2910 or in another machine accessible by the processor.

[0089] The bus subsystem 2922 provides a mechanism for allowing the various components and subsystems of the computer system 2900 to communicate with each other as intended. Although the bus subsystem 2922 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0090] The computer system 2900 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed collection of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 2900 depicted in Figure 29 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 2900 can have more or fewer components than the computer system depicted in Figure 29.

[0091] Terms The disclosed technology, particularly the provisions disclosed in this section, can be implemented as a system, method, or product. One or more features of the embodiments can be combined with the base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission of repetitive descriptions of these options from some embodiments should not be interpreted as limiting the combinations taught in the previous section. These descriptions are incorporated herein by reference into each of the following embodiments.

[0092] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing the specific technology described herein, and the software module is stored in a computer-readable storage medium (or multiple such media).

[0093] The clauses described in this section can be combined as features. For the sake of brevity, combinations of features are not listed separately and are not repeated for each base set of features. The reader will understand how features specified in the clauses described in this section can be easily combined with sets of basic features specified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or restrictive, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.

[0094] Other implementations of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.

[0095] The inventors disclose the following provisions: Clause 1. A computer-implemented method for determining the feasibility of using a non-target species reference genome for variant calling of a target species sample, comprising: mapping sequenced reads of the target species sample to a non-target species reference genome to detect a first set of variants in the sequenced reads of the target species sample; and mapping sequenced reads of the target species sample to a pseudo-target species reference genome to detect a second set of variants in the sequenced reads of the target species sample; comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not present in the first set of variants; and determining the feasibility of using the reference genome of the non-target species for variant calling of the target species based on the count of the subset of false positive variants. Clause 2. The computer-implemented method of clause 1, wherein the pseudo target species is the target species. Clause 3. The computer-implemented method of clause 1, wherein the pseudo target species is different from the target species. Clause 4. The computer-implemented method of clause 3, wherein the pseudo target species is homologous to the target species. Clause 5. The computer-implemented method of clause 1, wherein the non-target species is human. Clause 6. The computer-implemented method of clause 1, wherein the target species is a non-human primate. Clause 7. The computer-implemented method of clause 1, further comprising detecting a second set of variants by mapping the sequenced reads of the target species sample to a reference genome of the target species and then lifting over the mapped sequenced reads of the target species sample to a reference genome of a non-target species. Clause 8. The computer-implemented method of clause 1, further comprising applying a first filter to filter out low quality variants from the first set of variants and the second set of variants. Clause 9. The computer-implemented method of clause 1, further comprising applying a second filter to filter out from the first set of variants and the second set of variants fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo target species. Clause 10. The computer-implemented method of clause 1, wherein a false positive variant in the subset of false positive variants occurs because a particular region in a sequenced read of a sample of a target species maps to a first region in a reference genome of a non-target species and to a second region in a reference genome of a pseudo-target species, and the first region and the second region are different. Clause 11. The computer-implemented method of clause 10, wherein a false positive variant arises because a particular region in the sequenced reads of a sample of a target species maps to multiple regions in a reference genome of a non-target species. Clause 12. A system comprising one or more processors coupled to a memory, the memory being loaded with computer instructions for determining the feasibility of using a reference genome of a non-target species for variant calling of a sample of the target species, the instructions, when executed on the processor, mapping the sequenced reads of the target species sample to a non-target species reference genome to detect a first set of variants in the sequenced reads of the target species sample; and mapping the sequenced reads of the target species sample to a pseudo-target species reference genome to detect a second set of variants in the sequenced reads of the target species sample. comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not present in the first set of variants; and determining the feasibility of using a reference genome of a non-target species for variant calling of the target species based on the count of a subset of false positive variants. Clause 13. The system of clause 12, wherein the pseudo target species is the target species. Clause 14. The system of clause 12, wherein the pseudo target species is different from the target species. Clause 15. The system of clause 12, wherein the pseudo target species is homologous to the target species. Clause 16. The system of clause 12, wherein the non-target species is human. Clause 17. The system of clause 12, wherein the target species is a non-human primate. Clause 18. The system of clause 12, further comprising detecting a second set of variants by mapping sequenced reads of the target species sample to a reference genome of the target species and then lifting over the mapped sequenced reads of the target species sample to a reference genome of a non-target species. Clause 19. The system of clause 12, further comprising applying a first filter to filter out low quality variants from the first set of variants and the second set of variants. Clause 20. The system of clause 12, further comprising applying a second filter to filter out fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo target species from the first set of variants and the second set of variants. Clause 21. The system of clause 12, wherein a false positive variant in a subset of false positive variants occurs because a particular region in a sequenced read of a sample of a target species is located in a first region in a reference genome of a non-target species and in a second region in a reference genome of a pseudo-target species, and the first region and the second region are different. Clause 22. The system of clause 12, wherein a false positive variant occurs because a particular region in the sequenced read of a sample of a target species maps to multiple regions in a reference genome of a non-target species. Clause 23. A non-transitory computer-readable storage medium having computer program instructions embodied therein for determining the feasibility of using a reference genome of a non-target species for variant calling of a sample of a target species, the instructions, when executed on a processor, comprising: mapping the sequenced reads of the target species sample to a non-target species reference genome to detect a first set of variants in the sequenced reads of the target species sample; and mapping the sequenced reads of the target species sample to a pseudo-target species reference genome to detect a second set of variants in the sequenced reads of the target species sample. comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not present in the first set of variants; and determining the feasibility of using a reference genome of a non-target species for variant calling of the target species based on the count of a subset of false positive variants. Clause 24. The non-transitory computer-readable storage medium of clause 23, wherein the pseudo target species is a target species. Clause 25. The non-transitory computer-readable storage medium of clause 23, wherein the pseudo target species is different from the target species. Clause 26. The non-transitory computer-readable storage medium of clause 23, wherein the pseudo target species is homologous to the target species. Clause 27. The non-transitory computer-readable storage medium of clause 23, wherein the non-target species is human. Clause 28. The non-transitory computer-readable storage medium of clause 23, wherein the target species is a non-human primate. Clause 29. The non-transitory computer-readable storage medium of clause 23, further comprising detecting a second set of variants by mapping the sequenced reads of the target species sample to a reference genome of the target species and then lifting over the mapped sequenced reads of the target species sample to a reference genome of a non-target species. Clause 30. The non-transitory computer-readable storage medium of clause 23, further comprising applying a first filter to filter out low quality variants from the first set of variants and the second set of variants. Clause 31. The non-transitory computer-readable storage medium of clause 23, further comprising applying a second filter to filter out fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo target species from the first set of variants and the second set of variants. Clause 32. The non-transitory computer-readable storage medium of clause 23, wherein a false positive variant in a subset of false positive variants occurs because a particular region in a sequenced read of a sample of a target species is located in a first region in a reference genome of a non-target species and in a second region in a reference genome of a pseudo target species, and the first region and the second region are different. Clause 33. The non-transitory computer-readable storage medium of clause 32, wherein a false positive variant occurs because a particular region in the sequenced read of a sample of a target species maps to multiple regions in a reference genome of a non-target species. Clause 34. A system comprising: a variant quality classifier configured to process a plurality of features of the target variant and generate a quality indicator for the target variant; A variant quality classifier is trained on a set of high quality variants and a set of low quality variants; high quality variants in the set of high quality variants are identified as true positive variants that are common between the first set of variants and the second set of variants; low-quality variants in the set of low-quality variants are identified as false positive variants that are present in the second set of variants but not in the first set of variants; a first set of variants is detected by variant calling of sequenced reads of a sample of a target species against a reference genome of a non-target species; The system, wherein a second set of variants is detected by variant calling of sequenced reads of the target species sample against a reference genome of a pseudo target species. Clause 35. The system of clause 34, wherein the variant quality classifier is a random forest model. Clause 36. The system of clause 34, wherein the variant quality classifier is a logistic regression model. Clause 37. The system of clause 34, wherein the variant quality classifier is a neural network model. Clause 38. The system of clause 34, wherein one of the multiple features of the target variant is a guanine-cytosine (GC) content within the sequenced reads of the target variant. Clause 39. The system of clause 34, wherein one feature of the plurality of features of the target variant is guanine-cytosine (GC) skew within a sequenced read of the target variant, the GC skew representing a normalized excess of cytosine relative to guanine in a given sequenced read of the target variant. Clause 40. The system of clause 34, wherein one of the multiple features of the target variant is local compositional complexity within 100 base pairs upstream or downstream of the target variant. Clause 41. The system of clause 34, wherein one feature of the plurality of features of the target variant is an allele count of sequenced reads of the target variant. Clause 42. The system of clause 34, wherein one of the plurality of features of the target variant is a mapping quality of the sequenced reads of the target variant. Clause 43. The system of clause 34, wherein one feature of the plurality of features of the target variant is a p-value of a Fisher's exact test for detecting strand bias in sequenced reads of the target variant. Clause 44. The system of clause 34, wherein one feature of the plurality of features of the target variant is a symmetric odds ratio for detecting strand bias in sequenced reads of the target variant. Clause 45. The system of clause 34, wherein one of the multiple features of the target variant is a variant quality according to the depth of the sequencing read of the target variant. Clause 46. The system of clause 34, wherein one of the plurality of features of the target variant is a genotype quality of the sequenced reads of the target variant. Clause 47. The system of clause 34, wherein one feature of the plurality of features of the target variant is a read depth of the target variant normalized by the average coverage of the sequenced reads of the target variant. Clause 48. The system of clause 34, wherein one feature of the plurality of features of the target variant is a read depth of an alternative allele fragment from the target variant coverage of sequenced reads of the target variant. Clause 49. The system of clause 34, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 5 base pairs upstream or downstream of the sequenced read of the target variant. Clause 50. The system of clause 34, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 10 base pairs upstream or downstream of the sequenced read of the target variant. Clause 51. The system of clause 34, wherein one of the multiple features of the target variant is the average coverage of a flanking region 100 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 52. The system of clause 34, wherein one of the multiple features of the target variant is the average coverage of a flanking region 500 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 53. The system of clause 34, wherein one of the multiple features of the target variant is the number of heterozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 54. The system of clause 34, wherein one of the multiple features of the target variant is the number of heterozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 55. The system of clause 34, wherein one of the multiple features of the target variant is the number of homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 56. The system of clause 34, wherein one of the multiple features of the target variant is the number of homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 57. The system of clause 34, wherein one feature of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 58. The system of clause 34, wherein one feature of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 59. A computer-implemented method for processing a plurality of features of a target variant and generating a quality index for the target variant, comprising: training a variant quality classifier on a set of high quality variants and a set of low quality variants; identifying high quality variants in the set of high quality variants as true positive variants that are common between the first set of variants and the second set of variants; identifying low quality variants in the low quality set as false positive variants that are present in the second set of variants but not in the first set of variants; detecting a first set of variants by variant calling sequenced reads of a sample of a target species against a reference genome of a non-target species; and detecting a second set of variants by variant calling the sequenced reads of the sample of the target species against a reference genome of the pseudo target species. Clause 60. The computer-implemented method of clause 59, wherein the variant quality classifier is a random forest model. Clause 61. The computer-implemented method of clause 59, wherein the variant quality classifier is a logistic regression model. Clause 62. The computer-implemented method of clause 59, wherein the variant quality classifier is a neural network model. Clause 63. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a guanine-cytosine (GC) content within the sequenced reads of the target variant. Clause 64. One of the features of the target variant is guanine-cytosine (GC) skew within the sequenced reads of the target variant; 60. The computer-implemented method of clause 59, wherein GC skew represents a normalized excess of cytosine over guanine in a given sequenced read of a target variant. Clause 65. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a local compositional complexity within 100 base pairs upstream or downstream of the target variant. Clause 66. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is an allele count of sequenced reads of the target variant. Clause 67. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a mapping quality of the sequenced reads of the target variant. Clause 68. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a p-value of a Fisher's exact test for detecting strand bias in sequenced reads of the target variant. Clause 69. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a symmetric odds ratio for detecting strand bias in sequenced reads of the target variant. Clause 70. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a variant quality according to the depth of the sequencing read of the target variant. Clause 71. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a genotype quality of a sequenced read of the target variant. Clause 72. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a read depth of the target variant normalized by the average coverage of the sequenced reads of the target variant. Clause 73. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is a read depth of an alternative allelic fragment from the target variant coverage of sequenced reads of the target variant. Clause 74. The computer-implemented method of clause 59, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 5 base pairs upstream or downstream of the sequenced read of the target variant. Clause 75. The computer-implemented method of clause 59, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 10 base pairs upstream or downstream of the sequenced read of the target variant. Clause 76. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the average coverage of a flanking region 100 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 77. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the average coverage of a flanking region 500 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 78. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of heterozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 79. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of heterozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 80. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 81. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 82. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 83. The computer-implemented method of clause 59, wherein one feature of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 84. A non-transitory computer-readable storage medium having computer program instructions for processing a plurality of features of a target variant and generating a quality indicator for the target variant, the instructions, when executed on a processor, Implementing a method including a variant quality classifier trained on a set of high quality variants and a set of low quality variants; high quality variants in the set of high quality variants are identified as true positive variants that are common between the first set of variants and the second set of variants; low-quality variants in the set of low-quality variants are identified as false positive variants that are present in the second set of variants but not in the first set of variants; a first set of variants is detected by variant calling of sequenced reads of a sample of a target species against a reference genome of a non-target species; A non-transitory computer-readable storage medium, wherein the second set of variants is detected by variant calling of sequenced reads of the target species sample against a reference genome of a pseudo target species. Clause 85. The non-transitory computer-readable storage medium of clause 84, wherein the variant quality classifier is a random forest model. Clause 86. The non-transitory computer-readable storage medium of clause 84, wherein the variant quality classifier is a logistic regression model. Clause 87. The non-transitory computer-readable storage medium of clause 84, wherein the variant quality classifier is a neural network model. Clause 88. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is a guanine-cytosine (GC) content within the sequenced reads of the target variant. Clause 89. One of the features of the target variant is guanine-cytosine (GC) skew within the sequenced reads of the target variant; 85. The non-transitory computer-readable storage medium of claim 84, wherein GC skew represents a normalized excess of cytosine over guanine in a given sequenced read of a target variant. Clause 90. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is a local compositional complexity within 100 base pairs upstream or downstream of the target variant. Clause 91. The non-transitory computer-readable storage medium of clause 84, wherein one feature of the plurality of features of the target variant is an allele count of sequenced reads of the target variant. Clause 92. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is a mapping quality of the sequenced reads of the target variant. Clause 93. The non-transitory computer-readable storage medium of clause 84, wherein one feature of the plurality of features of the target variant is a p-value of a Fisher's exact test for detecting strand bias in sequenced reads of the target variant. Clause 94. The non-transitory computer-readable storage medium of clause 84, wherein one feature of the plurality of features of the target variant is a symmetric odds ratio for detecting strand bias in sequenced reads of the target variant. Clause 95. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is a variant quality according to the depth of the sequenced read of the target variant. Clause 96. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is a genotypic quality of a sequenced read of the target variant. Clause 97. The non-transitory computer-readable storage medium of clause 84, wherein one feature of the plurality of features of the target variant is a read depth of the target variant normalized by the average coverage of the sequenced reads of the target variant. Clause 98. The non-transitory computer-readable storage medium of clause 84, wherein one feature of the plurality of features of the target variant is a read depth of an alternative allele fragment from the target variant coverage of sequenced reads of the target variant. Clause 99. The non-transitory computer-readable storage medium of clause 84, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 5 base pairs upstream or downstream of the sequenced read of the target variant. Clause 100. The non-transitory computer-readable storage medium of clause 84, wherein one of the multiple features of the target variant is the presence of an insertion and / or deletion (indel) mutation within 10 base pairs upstream or downstream of the sequenced read of the target variant. Clause 101. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is an average coverage of a flanking region 100 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 102. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is an average coverage of a flanking region 500 base pairs upstream or downstream of the sequenced reads of the target variant, normalized by the average coverage of the sequenced reads of the target variant. Clause 103. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of heterozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 104. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of heterozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 105. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 106. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 107. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 100 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 108. The non-transitory computer-readable storage medium of clause 84, wherein one of the plurality of features of the target variant is the number of alternative homozygous single nucleotide polymorphisms within 500 base pairs upstream or downstream of the sequenced read of the target variant, normalized by the median count of variants within the same length region of the sequenced read of the target variant. Clause 109. A computer-implemented method for identifying and eliminating regions that do not have a one-to-one mapping between a first reference genome and a second reference genome, comprising: Accessing sequenced reads for a sample of a target species; Identifying and removing low quality sequenced reads from the sequenced reads based on applying a mapping quality filter to the sequenced reads, thereby removing high quality sequenced reads from the sequenced reads; Segmenting a non-targeted reference genome of a non-targeted species into a plurality of bins and then mapping high quality sequenced reads for each bin to the plurality of bins in the non-targeted reference genome; Segmenting a pseudo-targeted reference genome of a pseudo-targeted species into a number of bins, and then mapping high quality sequenced reads for each bin to a number of bins in the pseudo-targeted reference genome; Identifying a best mapped bin in the pseudo target reference genome based on a maximum degree of concordance between the best mapped bin in the pseudo target reference genome and a corresponding bin in the non-target reference genome, wherein the degree of concordance between corresponding bins in the pseudo target reference genome and the non-target reference genome is determined by the number of reads mapped between the corresponding bins; generating a unique mapper score for the pseudo-targeted reference genome based on the number of reads mapped between the best mapped bin in the pseudo-targeted reference genome and a corresponding bin in the non-targeted reference genome; and using the unique mapper score to identify and eliminate low quality sequenced reads. Clause 110. The computer-implemented method of clause 109, wherein the low quality sequenced reads include stop gain variants. Clause 111. The computer-implemented method of clause 109, wherein a plurality of cascade filters are applied to the sequenced reads of a sample of a target species to filter low-quality sequenced reads. Clause 112. The computer-implemented method of clause 111, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genetic regions having inaccurate genetic annotation in the reference genome. Clause 113. The computer-implemented method of clause 111, wherein one filter of the plurality of cascade filters is configured to detect and eliminate codons that do not match between the pseudo target species reference genome and the non-target species reference genome. Clause 114. The computer-implemented method of clause 111, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having a skewed distribution of variant classifier scores compared to the distribution of variant classifier scores for the complete reference genome. Clause 115. The computer-implemented method of clause 111, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having deviations from Hardy-Weinberg equilibrium. Clause 116. The computer-implemented method of clause 111, wherein one filter of the plurality of cascade filters is configured to detect and remove single nucleotide polymorphisms having a Random Forest score greater than 0.17. Clause 117. The computer-implemented method of clause 109, describing a fraction of the number of reads in one bin in the non-targeted reference genome map to a single corresponding region in the pseudo-targeted reference genome by one-to-one mapping. Clause 118. The computer-implemented method of clause 117, wherein consecutive identical bins are considered collectively as a single bin, without precluding the possibility of one-to-one mapping. Clause 119. The computer-implemented method of clause 117, wherein three or more non-contiguous identical bins are considered as overlapping regions, eliminating the possibility of one-to-one mapping. Clause 120. The computer-implemented method of clause 109, wherein the bins describe 1 kilobase (kb) regions within the reference genome. Clause 121. The computer-implemented method of clause 109, wherein fragments of the mapped reads are detected for the best mapped regions as determined by the mapper score. Clause 122. The computer-implemented method of clause 109, wherein the unique mapper score is determined by the average of the top fragments across samples for each reference genome. Clause 123. The computer-implemented method of clause 109, wherein the unique mapper score is configured as a filter to eliminate sequenced reads having a mapper score less than 20. Clause 124. The computer-implemented method of clause 109, wherein the pseudo target species is human. Clause 125. The computer-implemented method of clause 109, wherein the simulated target species is a non-human primate. Clause 126. The computer-implemented method of clause 109, wherein the non-target species is human. Clause 127. The computer-implemented method of clause 109, wherein the non-target species is a non-human primate. Clause 128. The computer-implemented method of clause 109, wherein the target species is human. Clause 129. The computer-implemented method of clause 109, wherein the target species is a non-human primate. Clause 130. The computer-implemented method of clause 109, wherein the target species and the non-target species are homologous. Clause 131. The computer-implemented method of clause 109, wherein the target species and the pseudo-target species are homologous. Clause 132. The computer-implemented method of clause 109, wherein the quality of the variants identified from the sequenced reads of the target species is a proxy for evolutionary constraints on the variant genes. Clause 133. A system including one or more processors coupled to a memory, the memory being loaded with computer instructions for identifying and eliminating regions that do not have a one-to-one mapping between a first reference genome and a second reference genome, the instructions, when executed on the processor, performing: Accessing sequenced reads for a sample of a target species; Identifying and removing low quality sequenced reads from the sequenced reads based on applying a mapping quality filter to the sequenced reads, thereby removing high quality sequenced reads from the sequenced reads; Segmenting a non-targeted reference genome of a non-targeted species into a plurality of bins, and then mapping high quality sequenced reads for each bin to the plurality of bins in the non-targeted reference genome; Segmenting a pseudo-targeted reference genome of a pseudo-targeted species into a number of bins, and then mapping high quality sequenced reads for each bin to the number of bins in the pseudo-targeted reference genome; Identifying a best mapped bin in the pseudo target reference genome based on a maximum degree of concordance between the best mapped bin in the pseudo target reference genome and a corresponding bin in the non-target reference genome, wherein the degree of concordance between corresponding bins in the pseudo target reference genome and the non-target reference genome is determined by the number of reads mapped between the corresponding bins; generating a unique mapper score for the pseudo-targeted reference genome based on the number of reads mapped between the best mapped bin in the pseudo-targeted reference genome and a corresponding bin in the non-targeted reference genome; and using the unique mapper score to identify and eliminate low quality sequenced reads. Clause 134. The system of clause 133, wherein the low quality sequenced reads include stop gain variants. Clause 135. The system of clause 133, wherein multiple cascade filters are applied to the sequenced reads of a sample of a target species to filter out low quality sequenced reads. Clause 136. The system of clause 135, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genetic regions having inaccurate genetic annotation in the reference genome. Clause 137. The system of clause 135, wherein one filter of the plurality of cascade filters is configured to detect and eliminate codons that do not match between the pseudo target species reference genome and the non-target species reference genome. Clause 138. The system of clause 135, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having a biased distribution of variant classifier scores compared to the distribution of variant classifier scores for the complete reference genome. Clause 139. The system of clause 135, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having deviations from Hardy-Weinberg equilibrium. Clause 140. The system of clause 135, wherein one filter of the plurality of cascade filters is configured to detect and remove single nucleotide polymorphisms having a Random Forest score greater than 0.17. Clause 141. The system described in clause 133, describing a fraction of the number of reads in one bin in the non-targeted reference genome map to a single corresponding region in the pseudo-targeted reference genome by one-to-one mapping. Clause 142. The system of clause 141, wherein consecutive identical bins are considered collectively as a single bin, and does not preclude the possibility of one-to-one mapping. Clause 143. The system of clause 141, wherein three or more non-contiguous identical bins are considered to be overlapping regions, eliminating the possibility of one-to-one mapping. Clause 144. The system of clause 133, wherein the bins describe 1 kilobase (kb) regions within the reference genome. Clause 145. The system of clause 133, wherein fragments of mapped reads are detected for best mapped regions as determined by mapper scores. Clause 146. The system of clause 133, wherein the unique mapper score is determined by the average of the top fragments across samples for each reference genome. Clause 147. The system of clause 133, wherein the unique mapper score is configured as a filter to eliminate sequenced reads having a mapper score less than 20. Clause 148. The system according to clause 133, wherein the pseudo target species is human. Clause 149. The system of clause 133, wherein the pseudo target species is a non-human primate. Clause 150. The system according to clause 133, wherein the non-target species is human. Clause 151. The system according to clause 133, wherein the non-target species is a non-human primate. Clause 152. The system according to clause 133, wherein the target species is human. Clause 153. The system according to clause 133, wherein the target species is a non-human primate. Clause 154. The system of clause 133, wherein the target species and the non-target species are homologous. Clause 155. The system according to clause 133, wherein the target species and the pseudo-target species are homologous. Clause 156. The system of clause 133, wherein the quality of variants identified from sequenced reads of the target species is a proxy for evolutionary constraints on the variant genes. Clause 157. A non-transitory computer readable storage medium having computer program instructions for identifying and eliminating regions that do not have a one-to-one mapping between a first reference genome and a second reference genome, the instructions, when executed on a processor, Accessing sequenced reads for a sample of a target species; Identifying and removing low quality sequenced reads from the sequenced reads based on applying a mapping quality filter to the sequenced reads, thereby removing high quality sequenced reads from the sequenced reads; Segmenting a non-targeted reference genome of a non-targeted species into a plurality of bins, and then mapping high quality sequenced reads for each bin to the plurality of bins in the non-targeted reference genome; Segmenting a pseudo-targeted reference genome of a pseudo-targeted species into a number of bins, and then mapping high quality sequenced reads for each bin to the number of bins in the pseudo-targeted reference genome; Identifying a best mapped bin in the pseudo target reference genome based on a maximum degree of concordance between the best mapped bin in the pseudo target reference genome and a corresponding bin in the non-target reference genome, wherein the degree of concordance between corresponding bins in the pseudo target reference genome and the non-target reference genome is determined by the number of reads mapped between the corresponding bins; generating a unique mapper score for the pseudo-targeted reference genome based on the number of reads mapped between the best mapped bin in the pseudo-targeted reference genome and a corresponding bin in the non-targeted reference genome; A system implementing a method comprising: using unique mapper scores to identify and eliminate low quality sequenced reads. Clause 158. The non-transitory computer-readable storage medium of clause 157, wherein the low quality sequenced reads include stop gain variants. Clause 159. The non-transitory computer-readable storage medium of clause 157, wherein a plurality of cascade filters are applied to the sequenced reads of a sample of a target species to filter out low-quality sequenced reads. Clause 160. The non-transitory computer-readable storage medium of clause 159, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genetic regions having inaccurate genetic annotations in the reference genome. Clause 161. The non-transitory computer-readable storage medium of clause 159, wherein one filter of the plurality of cascade filters is configured to detect and eliminate codons that do not match between the pseudo target species reference genome and the non-target species reference genome. Clause 162. The non-transitory computer-readable storage medium of clause 159, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having a skewed distribution of variant classifier scores compared to the distribution of variant classifier scores for the complete reference genome. Clause 163. The non-transitory computer-readable storage medium of clause 159, wherein one filter of the plurality of cascade filters is configured to detect and eliminate genes in the reference genome having deviations from Hardy-Weinberg equilibrium. Clause 164. The non-transitory computer-readable storage medium of clause 159, wherein one filter of the plurality of cascade filters is configured to detect and remove single nucleotide polymorphisms having a Random Forest score greater than 0.17. Clause 165. A non-transitory computer-readable storage medium as described in clause 157, describing a fraction of the number of reads in one bin in a non-targeted reference genome map to a single corresponding region in a pseudo-targeted reference genome by one-to-one mapping. Clause 166. The non-transitory computer-readable storage medium of clause 165, wherein consecutive identical bins are collectively considered as a single bin, not precluding the possibility of one-to-one mapping. Clause 167. The non-transitory computer-readable storage medium of clause 165, wherein three or more non-contiguous identical bins are considered an overlapping region, eliminating the possibility of one-to-one mapping. Clause 168. The non-transitory computer-readable storage medium of clause 157, wherein the bins describe 1 kilobase (kb) regions within the reference genome. Clause 169. The non-transitory computer-readable storage medium of clause 157, wherein fragments of the mapped reads are detected for best mapped regions as determined by the mapper score. Clause 170. The non-transitory computer-readable storage medium of clause 157, wherein the unique mapper score is determined by an average of the top fragments across samples for each reference genome. Clause 171. The non-transitory computer-readable storage medium of clause 157, wherein the unique mapper score is configured as a filter to eliminate sequenced reads having a mapper score less than 20. Clause 172. The non-transitory computer-readable storage medium of clause 157, wherein the pseudo target species is human. Clause 173. The non-transitory computer-readable storage medium of clause 157, wherein the simulated target species is a non-human primate. Clause 174. The non-transitory computer-readable storage medium of clause 157, wherein the non-target species is human. Clause 175. The non-transitory computer-readable storage medium of clause 157, wherein the non-target species is a non-human primate. Clause 176. The non-transitory computer-readable storage medium of clause 157, wherein the target species is human. Clause 177. The non-transitory computer-readable storage medium of clause 157, wherein the target species is a non-human primate. Clause 178. The non-transitory computer-readable storage medium of clause 157, wherein the target species and the non-target species are homologous. Clause 179. The non-transitory computer-readable storage medium of clause 157, wherein the target species and the pseudo-target species are homologous. Clause 180. The non-transitory computer-readable storage medium of clause 157, wherein the quality of variants identified from the sequenced reads of the target species is a proxy for evolutionary constraints on the variant genes.

Claims

1. 1. A computer-implemented method using a reference genome of a non-target species for variant calling of a sample of a target species, comprising: mapping the sequenced reads of a sample of a target species to a reference genome of a non-target species belonging to the same taxonomic class as the target species to detect a first set of variants in the sequenced reads of the sample of the target species; mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo target species that is homologous to the target species to detect a second set of variants in the sequenced reads of the sample of the target species; comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not in the first set of variants; and using the reference genome of the non-target species for variant calling of the target species based on the count of the subset of false positive variants.

2. The computer-implemented method of claim 1 , wherein the pseudo-target species is orthologous to the target species.

3. The computer-implemented method of claim 1 , wherein the pseudo target species is different from the target species.

4. The computer-implemented method of claim 3 , wherein the pseudo-target species is homologous to the target species by a homology threshold.

5. The computer-implemented method of claim 1 , wherein the non-target species is a human and the target species is a non-human primate.

6. 2. The computer-implemented method of claim 1, further comprising detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.

7. 7. The computer-implemented method of claim 1, wherein a false positive variant in the subset of false positive variants occurs because a particular region in the sequenced reads of the sample of the target species maps to a first region in the reference genome of the non-target species and to a second region in the reference genome of the pseudo-target species, and the first region and the second region are different.

8. 8. The computer-implemented method of claim 7, wherein the false positive variants arise because the particular region in the sequenced reads of the sample of the target species maps to multiple regions in the reference genome of the non-target species.

9. 1. A system comprising one or more processors coupled to a memory, the memory loaded with computer instructions for using a reference genome of a non-target species for variant calling of a sample of a target species, the computer instructions, when executed on the one or more processors, performing: mapping the sequenced reads of a sample of a target species to a reference genome of a non-target species belonging to the same taxonomic class as the target species to detect a first set of variants in the sequenced reads of the sample of the target species; mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo target species that is homologous to the target species to detect a second set of variants in the sequenced reads of the sample of the target species; comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not in the first set of variants; using the reference genome of the non-target species for variant calling of the target species based on a count of a subset of the false positive variants.

10. The system of claim 9 , wherein the pseudo target species is different from the target species.

11. 10. The system of claim 9, wherein the non-target species is a human and the target species is a non-human primate.

12. 10. The system of claim 9, further comprising detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.

13. The system of claim 9 , further comprising applying a first filter to filter out low quality variants from the first set of variants and the second set of variants.

14. 10. The system of claim 9, further comprising applying a second filter to filter out fixed substitutions shared between the reference genome of the non-target species and the reference genome of the pseudo-target species from the first set of variants and the second set of variants.

15. 15. The system of any one of claims 9 to 14, wherein a false positive variant in the subset of false positive variants occurs because a particular region in the sequenced reads of the sample of the target species maps to a first region in the reference genome of the non-target species and a second region in the reference genome of the pseudo-target species, and the first region and the second region are different.

16. 16. The system of claim 15, wherein the false positive variant occurs because the particular region in the sequenced read of the sample of the target species maps to multiple regions in the reference genome of the non-target species.

17. 1. A non-transitory computer-readable storage medium having computer program instructions thereon for using a reference genome of a non-target species for variant calling of a sample of a target species, the computer program instructions, when executed on a processor, comprising: mapping the sequenced reads of a sample of a target species to a reference genome of a non-target species belonging to the same taxonomic class as the target species to detect a first set of variants in the sequenced reads of the sample of the target species; mapping the sequenced reads of the sample of the target species to a reference genome of a pseudo target species that is homologous to the target species to detect a second set of variants in the sequenced reads of the sample of the target species; comparing the first set of variants to the second set of variants and identifying a subset of true positive variants that are common between the first set of variants and the second set of variants; comparing the first set of variants to the second set of variants and identifying a subset of false positive variants that are present in the second set of variants but not in the first set of variants; and using the reference genome of the non-target species for variant calling of the target species based on a count of a subset of the false positive variants.

18. 20. The non-transitory computer-readable storage medium of claim 17, wherein the pseudo-target species is orthologous to the target species.

19. 20. The non-transitory computer-readable storage medium of claim 17, wherein the non-target species is a human and the target species is a non-human primate.

20. 20. The non-transitory computer-readable storage medium of any one of claims 17 to 19, further comprising detecting the second set of variants by mapping the sequenced reads of the sample of the target species to the reference genome of the target species, and then lifting over the mapped sequenced reads of the sample of the target species to the reference genome of the non-target species.