Computer-implemented method for identifying rare variants that cause extreme levels of gene expression

An AI-driven chromatin model addresses the challenge of identifying rare genetic variants causing extreme gene expression by integrating epigenetic signals and controlling confounding factors, enhancing the accuracy of variant classification and disease understanding.

JP2025527970APending Publication Date: 2025-08-26ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024557742
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-05
Filing Date
2023-07-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing genomics analysis methods struggle to accurately identify rare genetic variants causing extreme levels of gene expression due to confounding factors and incomplete data representation, which affects the classification accuracy and reliability of pathogenicity predictions.

Method used

An artificial intelligence-based approach that utilizes a chromatin model to analyze chromatin sequences, incorporating epigenetic signals and controlling for multiple confounding factors to determine the causal relationship between rare variants and extreme gene expression levels, using a processing pipeline and causal models to generate causality scores.

Benefits of technology

Enhances the identification of rare variants associated with extreme gene expression by improving signal-to-noise ratio and accuracy, allowing for better understanding of genetic diseases and potential therapeutic interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025527970000001_ABST
    Figure 2025527970000001_ABST
Patent Text Reader

Abstract

The disclosed technology relates to reliably identifying variants that cause extreme levels of gene expression. Extreme levels of gene expression include under-expression and over-expression. These variants can then be used to train artificial intelligence-based models for various prediction tasks. One example of a prediction task is generating base-by-base resolution for chromatin sequences. Another example of a chromatin task is generating gene expression changes caused by reliably identified variants.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to artificial intelligence-based epigenetics at base resolution.

[0002] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is related to a concurrently filed U.S. patent application entitled "ARTIFICIAL INTELLIGENCE-BASED DETECTION OF GENE CONSERVATION AND EXPRESSION PRESERVATION AT BASE RESOLUTION" (Attorney Docket No. ILLM 1036-1 / IP-2045-PRV), which is incorporated by reference for all purposes as if fully set forth herein.

[0003] (built-in) The following are incorporated by reference for all purposes as if fully set forth herein: U.S. Patent Application No. 62 / 903,700, entitled "ARTIFICIAL INTELLIGENCE-BASED EPIGENETICS," filed September 20, 2019 (Attorney Docket No. ILLM 1025-1 / IP-1898-PRV); Sundaram,L.et al.Predicting the clinical impact of human mutation with deep neural networks.Nat.Genet.50,1161-1170(2018); Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019); U.S. Patent Application No. 62 / 573,144, entitled "TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV); U.S. Patent Application No. 62 / 573,149, entitled "PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV); U.S. Patent Application No. 62 / 573,153, entitled "DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV); U.S. Patent Application No. 62 / 582,898, entitled "PATHOGENICITY CLASSIFICATION OF GEOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV); U.S. Patent Application No. 16 / 160,903, entitled "DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM1000-5 / IP-1611-US); U.S. Patent Application No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US); U.S. Patent Application No. 16 / 160,968, entitled "SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US); U.S. Patent Application No. 16 / 407,149, entitled "DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS," filed May 8, 2019 (Attorney Docket No. ILLM 1010-1 / IP-1734-US); U.S. Patent Application No. 17 / 232,056, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS TO PREDICT VARIANT PATHOGENICITY USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURES," filed April 15, 2021 (Attorney Docket No. ILLM 1037-2 / 1P-2051-US); U.S. Patent Application No. 63 / 175,495, entitled "MULTI-CHANNEL PROTEIN VOXELIZATION TO PREDICT VARIANT PATHOGENICITY USING DEEP CONVOLUTIONAL NEURAL NETWORKS," filed April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV); U.S. Patent Application No. 63 / 175,767, entitled "EFFICIENT VOXELIZATION FOR DEEP LEARNING," filed April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV); and U.S. Patent Application No. 17 / 468,411, filed September 7, 2021 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US), entitled "ARTIFICIAL INTELLIGENCE-BASED ANALYSIS OF PROTEIN THREE-DIMENSIONAL (3D) STRUCTURES." [Background technology]

[0004] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section or problems associated with the subject matter provided as background have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which, as such, may also correspond to embodiments of the claimed technology.

[0005] Genomics, broadly defined, also known as functional genomics, aims to characterize the function of all genomic elements in an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science and operates not by testing preconceived models and hypotheses, but by discovering novel properties from the examination of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene function, and mapping biochemically active genomic regions and residues, such as transcriptional enhancers and single nucleotide polymorphisms (SNPs).

[0006] Genomics data are too large and complex to be mined solely by visual inspection of pairwise correlations. For example, protein sequences can be grouped into families of homologous proteins that are derived from ancestral proteins and share similar structures and functions. Analysis of multiple sequence alignments (MSA) of homologous proteins provides important information about functional and structural constraints. Statistics of MSA columns, representing amino acid positions, identify functional residues that are conserved during evolution. Correlations of amino acid usage between MSA columns contain important information about functional sectors and structural contacts.

[0007] Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some algorithms in which assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically discover patterns in data. Therefore, machine learning algorithms are well-suited for data-driven science, particularly genomics. However, the performance of machine learning algorithms can strongly depend on how the data is represented, i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from fluorescence microscopy images, a preprocessing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.

[0008] A machine learning model can take estimated cell counts, which are an example of hand-designed features, as input features for classifying tumors. A central problem is that classification performance depends heavily on the quality and relevance of these features. For example, relevant visual features such as cell morphology, distance between cells, or localization within an organ are not captured in cell counts, and this incomplete representation of the data can reduce classification accuracy.

[0009] Deep learning, a subdiscipline of machine learning, addresses this problem by embedding feature computation within the machine learning model itself, generating an end-to-end model. This result has been achieved through the development of deep neural networks, machine learning models that involve successive primitive operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of high complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in algorithms, and substantial increases in computing power, particularly through the use of graphical processing units (GPUs).

[0010] The goal of supervised learning is to obtain a model that takes features as input and returns a prediction of a so-called target variable. An example of a supervised learning problem is predicting whether an intron will be spliced ​​(target) or not, given RNA features such as the presence or absence of canonical splice site sequences, the location of splicing branch points, or the length of the intron. Training a machine learning model refers to learning its parameters, which generally involves minimizing a loss function on the training data with the goal of making accurate predictions on unknown data.

[0011] For many supervised learning problems in computational biology, input data can be represented as a table with multiple columns, or features, each of which contains numerical or categorical data potentially useful for making predictions. While some input data are naturally represented as tabular features (e.g., temperature or time), other input data must first be transformed using a process called feature extraction (e.g., converting deoxyribonucleic acid (DNA) sequences to k-mer counts) to fit a tabular representation. For intron-splicing prediction problems, the presence or absence of canonical splice site sequences, splicing branch point locations, and intron lengths can be preprocessed features collected in tabular format. Tabular data is the standard for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible nonlinear models such as neural networks and many others.

[0012] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of a positive class by calculating a weighted sum of input features mapped to the [0, 1] interval using a sigmoid activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when the weighted sum of input features cannot sufficiently distinguish between classes, such as whether an intron is spliced ​​out or not. To improve prediction performance, new input features can be manually added by transforming or combining existing features in new ways, such as by taking exponentiation or pairwise products.

[0013] Neural networks automatically learn these nonlinear feature transformations using hidden layers. Each hidden layer can be thought of as multiple linear models with their output transformed by a nonlinear activation function, such as a sigmoid function or the more general rectified-linear unit (ReLU). Together, these layers organize the input features into related complex patterns, easing the task of distinguishing between two classes.

[0014] Deep neural networks use many hidden layers, and when each neuron receives input from all neurons in the previous layer, the layer is said to be fully connected. Neural networks are typically trained using stochastic gradient descent, an algorithm suitable for training models on very large datasets. Implementation of neural networks using modern deep learning frameworks allows for rapid prototyping with different architectures and datasets. Fully connected neural networks can be used in several genomics applications, including predicting the proportion of exons spliced ​​into a given sequence from sequence features such as the presence of splice factor binding motifs or sequence conservation, prioritizing potentially disease-causing genetic variants, and predicting cis-regulatory elements in a given genomic region using features such as chromatin marks, gene expression, and evolutionary conservation.

[0015] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling DNA sequences or image pixels severely disrupts information patterns. These local dependencies set spatial or longitudinal data apart from tabular data, where feature ordering is arbitrary. Consider the problem of classifying genomic regions as bound or unbound by a specific transcription factor, where binding regions are defined as high-confidence binding events in chromatin immunoprecipitation following sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. Fully connected layers based on sequence-derived features, such as the number of k-mer instances in a sequence or position weight matrix (PWM) matches, can be used for this task. Because k-mer or PWM instance frequencies are robust to shifting motifs within a sequence, such models can generalize well to sequences with the same motif located at different positions. However, they cannot recognize patterns where transcription factor binding depends on the combination of multiple motifs with distinct intervals. Furthermore, the number of possible k-mers increases exponentially with k-mer length, which poses both conservation and overfitting challenges.

[0016] A convolutional layer is a special form of a fully connected layer in which the same fully connected layer is applied locally to every sequence position, for example, within a 6-bp window. This approach can also be viewed as scanning a sequence using multiple PWMs, for example, for the transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced, allowing the network to detect motifs at positions not seen during training. Each convolutional layer scans the sequence with several filters by generating a scalar value at every position that quantizes the match between the filter and the sequence. As in a fully connected neural network, a nonlinear activation function (typically ReLU) is applied in each layer. Next, a pooling operation is applied, which aggregates activations within successive bins across the position axis, typically taking the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers can then construct the output of the previous layer and detect whether the GATA1 and TAL1 motifs were present within a certain distance range. Finally, the output of the convolutional layer can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.

[0017] Convolutional neural networks (CNNs) can predict various molecular phenotypes based solely on DNA sequence. Applications include classification of transcription factor binding sites and prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequence, CNNs can be applied to more technical tasks traditionally addressed by hand-designed bioinformatics pipelines. For example, CNNs can predict guide RNA specificity, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequence, and call genetic variants. CNNs have also been used to model long-range dependencies in genomes. Although interacting regulatory elements may be located far apart on unfolded linear DNA sequences, these elements are often proximal in actual 3D chromatin structures. Thus, modeling molecular phenotypes from linear DNA sequences can be improved by allowing long-range dependencies, albeit with a coarse approximation of chromatin, allowing the model to implicitly learn aspects of 3D organization such as promoter-enhancer loops. This is achieved by using dilated convolutions with a receptive field of up to 32 kb. Dilated convolutions also allow splice sites to be predicted from the sequence using a receptive field of 10 kb, thereby enabling the integration of gene sequences over distances as long as a typical human intron (see Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).

[0018] Different types of neural networks can be characterized by their parameter-sharing schemes. For example, fully connected layers have no parameter sharing, while convolutional layers impose translational invariance by applying the same filter at every position in their input. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data, such as DNA sequences or time series, that implement a different parameter-sharing scheme. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. It updates the memory and optionally emits an output, which is either passed to subsequent layers or used directly as a model prediction. By applying the same model to each sequence element, recurrent neural networks are invariant to positional indices in the processed sequence. For example, recurrent neural networks can detect open reading frames in DNA sequences, regardless of their position in the sequence. This task requires the recognition of a specific sequence of inputs, such as a start codon followed by an in-frame stop codon.

[0019] The main advantage of recurrent neural networks over convolutional neural networks is that, theoretically, they can inherit information through infinitely long sequences via memory. Furthermore, recurrent neural networks can naturally process sequences of widely varying lengths, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can achieve performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the outputs of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, because recurrent neural networks apply sequential operations, they cannot be easily parallelized and are therefore much slower to compute than convolutional neural networks.

[0020] Although most of the human genetic code is common to all humans, each human has a unique genetic code. In some cases, the human genetic code may contain outliers, called genetic variants, that may be common among a relatively small group of individuals in the human population. For example, a particular human protein may contain a specific sequence of amino acids, but variants of that protein may differ by one amino acid in an otherwise identical specific sequence.

[0021] Genetic variants can be pathogenic and can result in disease. Although most such genetic variants have been eliminated from the genome by natural selection, the ability to identify which genetic variants are likely pathogenic can help researchers focus on these genetic variants to gain an understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human genetic variants remains unclear. Some of the most frequent pathogenic variants are single-nucleotide missense mutations that change the amino acid of a protein. However, not all missense mutations are pathogenic.

[0022] Models that can directly predict molecular phenotypes from biological sequences can be used as in silico perturbation tools to investigate the association between genetic and phenotypic variation and have emerged as new methods for quantitative trait locus identification and variant prioritization. These approaches are crucial given that the majority of variants identified by genome-wide association studies of complex phenotypes are non-coding, making it difficult to estimate their effect and contribution to the phenotype. Furthermore, linkage disequilibrium results in blocks of co-inherited variants, which makes it difficult to accurately identify individual causal variants. Therefore, sequence-based deep learning models that can be used as matching tools to assess the impact of such variants offer a promising approach for discovering potential drivers of complex phenotypes. One example is predicting the effects of non-coding single-nucleotide variants and short insertions or deletions (indels) indirectly from the differences between two variants on transcription factor binding, chromatin accessibility, or gene expression prediction. Another example is predicting novel splice site generation from the quantitative effects of genetic variants on sequence or splicing.

[0023] To predict the pathogenicity of missense variants from protein sequence and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (see Sundaram, L. et al. Predicting the clinical impact of human mutations with deep neural networks. Nat. Genet. 50, 1161-1170 (2018) (referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on known pathogenic variants with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences to compare differences and determine the pathogenicity of the variant using a trained deep neural network. Such an approach using protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to prior knowledge. However, the amount of clinical data available in ClinVar is relatively small compared to the amount of data sufficient to effectively train a deep neural network. To overcome this data shortage, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.

[0024] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in replicating prior knowledge in ClinVar. These results suggest that PrimateAI represents an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.

[0025] Central to protein biology is understanding how structural elements give rise to observed functions. The plethora of protein structural data enables the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods critically depends on the choice of protein structural representation.

[0026] Protein sites are microenvironments within a protein structure that are distinguished by their structural or functional role. Sites can be defined by their three-dimensional (3D) location and the local neighborhood around this location where the structure or function resides. Central to rational protein engineering is understanding how the structural arrangement of amino acids creates functional features within a protein site. Determining the structural and functional roles of individual amino acids in a protein provides information to aid in the manipulation and modification of protein function. Identifying functionally or structurally important amino acids enables focused engineering efforts, such as site-directed mutagenesis, to alter the functional properties of a target protein. Alternatively, this knowledge can help avoid engineering designs that abolish desired functions.

[0027] Because it is well established that structure is much more conserved than sequence, the increase in protein structural data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using data-driven approaches. A fundamental aspect of any computational protein analysis is how protein structural information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, whereas a poor representation produces a noisy distribution lacking the underlying pattern.

[0028] The plethora of protein structures and the recent success of deep learning algorithms provide an opportunity to develop tools for automatically extracting task-specific representations of protein structures.

[0029] Computational analysis of genomics studies is challenged by confounding variations unrelated to the genetic factors of interest. Identifying variants that cause extreme levels of gene expression, either high or low, is crucial for diagnosing the pathogenicity of genetic diseases. However, numerous confounding factors exist that can interfere with the identification of pathogenic variants. Isolating variants by examining rare variants that may be associated with specific disease states can simplify the problem. Furthermore, removing the noise introduced by confounding factors can increase the signal-to-noise ratio.

[0030] Therefore, an opportunity arises to apply artificial intelligence to epigenetics to obtain genetic information between variable loci and the expression levels of individual genes. [Brief explanation of the drawings]

[0031] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.

[0032] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Figure 1] FIG. 1 shows one embodiment of a processing pipeline that identifies rare variants that cause extreme levels of gene expression. [Figure 2] 1 illustrates one embodiment of a processing system for generating causality scores for outlier variants that cause extreme levels of gene expression. [Figure 3]FIG. 1 shows one embodiment of a fitted causal model that determines the causal relationship between rare variants and extreme levels of gene expression while controlling for multiple confounding factors. [Figure 4] 10 shows an example of a causality score generated by the disclosed technology for a sample of rare variants. [Figure 5] We show that performance outcomes, measured as z-scores across individuals and genes, progressively improve by sequentially correcting for multiple confounders using fitted causal models. [Figure 6] For a given p-value from the fitted causal model, the counts of true rare variants and shuffled (or randomly selected) variants identified as causing under- and over-gene expression are compared. [Figure 7] For a given p-value from the fitted causal model, odds ratios comparing the causal impact of true rare variants and shuffled variants on under- and over-gene expression are shown. [Figure 8] 1 shows a first exemplary architecture of the disclosed chromatin model. [Figure 9] 1 shows a second exemplary architecture of the disclosed chromatin model. [Figure 10] 1 shows a third exemplary architecture of the disclosed chromatin model. [Figure 11] 1 shows a fourth exemplary architecture of the disclosed chromatin model. [Figure 12] 1 shows a fifth exemplary architecture of the disclosed chromatin model. [Figure 13] The input generation logic that accesses a sequence database and generates an input base sequence is shown. [Figure 14] 1 shows one embodiment of base-resolution evolutionary conservation prediction using a chromatin model. [Figure 15] An example of an output sequence corresponding to a target base sequence is shown below. [Figure 16]1 shows one embodiment of the disclosed gene expression model. [Figure 17] Examples of reference and alternative sequences are shown. [Figure 18] 1 illustrates one embodiment of variant classification logic. [Figure 19] 1 shows one embodiment of the disclosed pathogenicity prediction logic. [Figure 20] 1 is an exemplary computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE INVENTION

[0033] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0034] The detailed description of various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks are not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., modules, processors, or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or block of random access memory, hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine within an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentality shown in the figures.

[0035] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers, or servers, or may be spread across multiple different processors, computers, or servers. In addition, it will be understood that some of the modules may operate in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures may also be considered flowchart steps in a method. Also, a module need not necessarily have all its code located contiguously in memory. Some portions of code may be separated from other portions of code, with code from other modules or other functions located between them.

[0036] Exemplary Applications of the Disclosed Technology Gene expression is the process by which instructions in DNA are converted into functional products such as RNA molecules or proteins. The disclosed technology identifies rare variants that cause extreme levels of gene expression, including both underexpression and overexpression. Rare variants are identified by their association with nearby genes that have extreme levels of gene expression. In one embodiment, the disclosed technology identifies individuals who have a specific variant in the promoter region of a gene and also have significantly different gene expression for that gene compared to individuals who do not have the specific variant. Based on the identification of such individuals, the disclosed technology classifies the specific variant as a gene expression modifier variant. The disclosed technology further uses artificial intelligence to train multiple models using the identified rare variants as training data and their underexpression and overexpression phenotypes as ground truth labels.

[0037] In this application, the terms chromatin sequence, chromatin input sequence, and output sequence are used throughout this application. Chromatin is DNA with bound proteins and / or RNA. By using the term chromatin sequence, we refer to the DNA sequence of chromatin. DNA sequences in sections of chromatin can be protected by proteins and RNA and then sequenced, as in DNA footprinting. Chromatin sequences can also be chemically modified. For example, DNA sequences often have methyl groups attached to nucleotides in the sequence and are therefore methylated.

[0038] As highlighted by a comparison of various embodiments of the chromatin model 124, many embodiments share overlaps in architectural components. Each element of the chromatin model 124 has multiple embodiments that can be combined in numerous configurations. The multiple permutations that can be implemented for the disclosed technology provide both a broader range of utility, performance efficiency, and performance accuracy. The data transformations applied to input sequences in many embodiments of the disclosed technology to generate multiple additional sequence formats from the perspective of base sequence and chromatin structure are innovative strategies that result in a wealth of output signals with broad applicability to a wide range of genomics, protein analysis, and pathogenicity research questions. Previous versions of PrimateAI have used multiple tools to classify variant pathogenicity with high performance. This chromatin model 124 introduces another tool in this methodology and an additional dimension for studying epigenetic signals that influence biological replication and transcription processes.

[0039] While there is clear utility in the added advantage of chromatin structure for studying gene expression to the various tools provided by PrimateAI, the true impact provided by the disclosed technology lies in the addition of epigenetic signals to the overall gene expression prediction logic. Both the DNA sequence and histone protein components of chromatin can undergo a plethora of chemical modifications. Enzymes that directly bind to and catalyze chemical modifications of chromatin components can alter chromatin structure, and changes in chromatin structure can also alter the ability of chromatin-interacting enzymes to access and function with their target ligands. Chromatin structure and enzymes that alter its structure directly affect the accessibility of genes for transcription and expression. DNA variants can cause changes in chromatin structure, which can subsequently alter epigenetic effects such as transcription factor binding and enzymatic reactions necessary for the proper regulation of gene expression and gene repression.

[0040] Conversely, epigenetic effects on chromatin, such as methylation and protein binding events, can affect mutation rates, potentially introducing silent or pathogenic variants. The study of evolutionary constraints on genes and the pathogenicity of their variants is significantly more comprehensive and accurate when augmented by epigenetic signatures, as demonstrated in many embodiments of the disclosed technology. Overall, the disclosed technology has several permutations that follow a series of training and learning strategies to generate several outputs that can be applied to predicting gene expression and gene pathogenicity for target gene sequences. The disclosed chromatin-focused strategies are useful in the study of genetic and environmental exposure-related diseases, drug development, and the influence of epigenetics on the transcription and translation of nucleic acid sequences into proteins.

[0041] Identification of rare variants Processing Pipeline 1 shows one embodiment of a processing pipeline 100 for identifying rare variants that cause extreme levels of gene expression. In operation 110, gene expression levels are accessed for a group of individuals. In one embodiment, the gene expression levels are accessed from Genotype-Tissue Expression (GTEx).

[0042] In operation 120, the gene expression levels are normalized, for example, by calculating the mean and multiple standard deviations from the mean.

[0043] In operation 130, outlier individuals having extreme levels of gene expression are identified from the group of individuals. The extreme levels of gene expression are determined from the tail quantiles 124 of the normalized gene expression levels 122. Examples of tail quantiles 124 include one or more standard deviations from the mean, both positive and negative. For example, outlier individuals have gene expression levels that have a z-score of at least 1.2 from the mean.

[0044] In operation 140, rare variants are selected from the gene sequences of the outlier individuals. The rare variants are selected based on an allele frequency cutoff. For example, the rare variants have a minor allele frequency (MAF) of less than 0.1%.

[0045] In operation 150, a causal model is fitted to determine the causal relationship between rare variants and extreme levels of gene expression in outlier individuals while controlling for multiple confounding factors. In one embodiment, the fitted causal model determines the causal relationship by predicting a specific gene expression level for a specific gene in a specific chromosome depending on the variant-driven gene expression level caused by a specific rare variant. In one embodiment, the fitted causal model measures the contribution of variant-driven gene expression levels as a variant effect covariate.

[0046] Examples of causal models include logistic regression models, linear regression models, analysis of covariance (ANCOVA) models, and / or multivariate analysis of covariance (MANCOVA) models. Examples of multiple confounders include distal trans-expression quantitative trait loci (eQTL) effects, local cis-eQTL effects, genotype-based principal components (gPCs), expression residual (PEER) effects, environmental effects, population structure and ancestry effects, as well as sex effects, batch effects, genotyping platform effects, and library construction protocol effects. PEER stands for "probabilistic estimation of expression residual." It is a collection of Bayesian approaches for inferring hidden determinants and their effects from gene expression profiles using factor analysis methods.

[0047] In operation 160, a causality score is generated for the rare variant based on the determined causality. A particular causality score for a particular rare variant indicates the likelihood that the particular rare variant causes extreme levels of gene expression in outlier individuals whose gene sequences contain the particular rare variant. In one embodiment, the causality score is a probability value (p-value). In one embodiment, the p-value is determined by the Pearson correlation coefficient.

[0048] Processing System 2 illustrates one embodiment of a processing system 200 for generating causality scores for outlier variants that cause extreme levels of gene expression. A data store 202 contains gene expression data, e.g., from GTEx, RNA-Seq, or whole genome sequencing (WGS). A normalizer 204 normalizes the gene expression data and stores the normalized gene expression data in a data store 206. The normalized gene expression data can be measured by a mean and one or more standard deviations from the mean. In one embodiment, extreme levels of gene expression are determined from the tail quantiles of the normalized gene expression data.

[0049] The causality model 212 is fitted to determine a causal relationship between the variant 216 and extreme levels of gene expression while controlling for multiple confounders 226. In one embodiment, the fitted causality model 212 determines causality by predicting a specific gene expression level of a specific gene in a specific chromosome depending on the variant-driven gene expression level caused by the specific variant, referred to herein as gene expression caused by the variant 214. In one embodiment, the fitted causality model 212 measures the contribution of the variant-driven gene expression level 214 (caused by the variant 216) as a variant effect covariate (shown later in FIG. 3 as 308). In one embodiment, the fitted causality model 212 measures the contribution of the confounder-driven gene expression level 224 (caused by the confounder 226) as multiple confounder effect covariates (shown later in FIG. 3 as 304, 306, and 310).

[0050] Examples of causal models 212 include logistic regression models, linear regression models, analysis of covariance (ANCOVA) models, and / or multivariate analysis of covariance (MANCOVA) models. Examples of confounders 226 include distal trans-expression quantitative trait loci (eQTL) effects, local cis-cQTL effects, genotype-based principal components (gPCs), expression residual (PEER) effects, environmental effects, population structure and ancestry effects, as well as sex effects, batch effects, genotyping platform effects, and library construction protocol effects.

[0051] The causal model 212 produces as output a confounder-corrected normalized gene expression 232 caused by the variant 216 on which the effects of the confounder 226 are regressed.

[0052] The rare variant identifier 234 identifies rare variants 236 from among the variants 216. The rare variants 216 can be selected based on an allele frequency cutoff. For example, the rare variants 236 can have a minor allele frequency (MAF) of less than 0.1%.

[0053] The causality score generator 242 generates a causality score 246 for the rare variant 236 based on the confounder-corrected normalized gene expression 232. A particular causality score for a particular rare variant indicates the likelihood that the particular rare variant causes extreme levels of gene expression in outlier individuals whose gene sequences contain the particular rare variant. In one embodiment, the causality score 246 is a probability value (p-value). In one embodiment, the p-value is determined by the Pearson correlation coefficient.

[0054] The rare variant ranker 256 ranks the rare variants 236 based on the causality scores 246 and stores the ranked rare variants in a data store 252 .

[0055] Causal Model 3 shows one embodiment of a fitted causal model 300 that determines the causal relationship between rare variants and extreme levels of gene expression while controlling for multiple confounding factors. The fitted causal model 300 has a dependent variable 302 titled "G" (gene expression) and multiple independent variables 304, 306, 308, and 310, respectively titled as follows: "G1" (gene expression driven by distal trans-eQTL effects) "G2" (gene expression driven by local cis-eQTL effects) "G3" (gene expression caused by rare variants), and "G4" (gene expression driven by genotype-based principal components (gPCs) 320, expression residual (PEER) effects 330, environmental effects 340, population structure and ancestry effects 350, and other effects such as sex effects, batch effects, genotyping platform effects, and library construction protocol effects 360).

[0056] The disclosed technique removes G1, G2, and G4 from G, i.e., G-G1-G2-G4=G3, eliminating (or correcting for) confounding effects and accurately determining gene expression caused by rare variants.

[0057] The fitted causal model 300 controls for distal trans eQTL effects by predicting a particular gene expression level (G) 302 depending on transgene expression levels (G1) 304 caused by other genes in other chromosomes. In one embodiment, the fitted causal model 300 measures the contribution of transgene expression level as a trans-effect covariate.

[0058] The fitted causal model 300 controls for local cis eQTL effects by predicting specific gene expression levels (G) 302 depending on cis gene expression levels (G) 306 caused by the presence of multiple common variants in the vicinity of a specific gene. In some embodiments, the vicinity is defined by an offset from the transcription start site (TSS) in the specific gene. In one embodiment, the fitted causal model 300 measures the contribution of cis gene expression levels as a cis-effect covariate.

[0059] The fitted causal model 300 controls for population structure and ancestry effects by predicting specific gene expression levels (G) 302 depending on the gPC-driven gPC gene expression levels (G4) 310. In one embodiment, the fitted causal model 300 measures the contribution of gPC gene expression levels (G4) 310 as a population structure and ancestry effect covariate.

[0060] The fitted causal model 300 controls for the PEER effect by predicting specific gene expression levels (G) 302 as a function of the PEER-induced PEER gene expression levels (G4) 310. In one embodiment, the fitted causal model 300 measures the contribution of the PEER gene expression levels (G4) 310 as a PEER effect covariate.

[0061] The fitted causal model 300 controls for environmental effects by predicting specific gene expression levels (G) 302 depending on the environmental gene expression levels (G4) 310 caused by the environmental effects. In one embodiment, the fitted causal model 300 measures the contribution of the environmental gene expression levels (G4) 310 as an environmental effect covariate.

[0062] In some embodiments, extreme levels of gene expression include over-gene expression and under-gene expression.

[0063] Overexpression In one embodiment, the fitted causality model 300 determines the causal relationship between rare variants and excessive gene expression while controlling for multiple confounding factors. In some embodiments, the fitted causality model 300 generates a so-called "excess causality score" for a rare variant. A particular excess causality score for a particular rare variant indicates the likelihood that the particular rare variant causes excessive gene expression in outlier individuals whose gene sequence contains the particular rare variant.

[0064] In one embodiment, excess causality score is excess probability value (excess p-value). In one embodiment, excess p-value is determined by Pearson correlation coefficient. In some embodiments, excess p-value specifies the statistically unconfounded likelihood that rare variants increase gene expression in genes that otherwise have lower gene expression compared to other genes in the gene set.

[0065] Underexpression In one embodiment, the fitted causality model 300 determines the causal relationship between rare variants and under-gene expression while controlling for multiple confounding factors. In some embodiments, the fitted causality model 300 generates a so-called "under-causality score" for a rare variant. A particular under-causality score for a particular rare variant indicates the likelihood that the particular rare variant causes under-gene expression in outlier individuals whose gene sequence contains the particular rare variant.

[0066] In one embodiment, the under-causation score is an under-probability value (under-p-value).In one embodiment, the under-probability p-value is determined by Pearson correlation coefficient.In some embodiments, the under-probability p-value identifies the statistically unaffected likelihood that rare variants reduce gene expression in genes that otherwise have higher gene expression compared to other genes in gene set.

[0067] In some embodiments, the rare variants are non-coding variants, which can include 5 prime untranslated region (UTR) variants, 3 prime UTR variants, enhancer variants, and promoter variants.

[0068] In some embodiments, the gene expression levels are further stratified into tissue-specific gene expression levels for multiple tissues. In one embodiment, the gene expression levels of each gene in each tissue are normalized using quantile normalization. In some embodiments, a causal model is fitted for each tissue separately. In some embodiments, a causal model is fitted using stratification.

[0069] Ranking In some embodiments, the ranking of rare variants is generated based on a causality score. In one embodiment, the ranking of rare variants is generated based on an over-causality score. In one embodiment, the ranking of rare variants is generated based on an under-causality score.

[0070] In some embodiments, the rare variant is a singleton variant. In one embodiment, the singleton variant occurs in only one of the outlier individuals.

[0071] Causality score

[0013] Figure 4 shows an example of a causality score 400 generated by the disclosed technology for a sample of rare variants. Figure 4 shows a "chrom" column 402, which identifies the chromosome on which the rare variant is located. Figure 4 also shows a "position" column 404, which identifies the position of the rare variant. Figure 4 also shows a "REF" column 406, which identifies the reference nucleotide (or base) that corresponds to the rare variant. Figure 4 also shows an "ALT" column 408, which identifies an alternative nucleotide (or base) that represents the rare variant.

[0072] 4 also shows a "p_under" column 410 that identifies the under-causation score for the rare variant. The p_under value is inversely proportional to the under-causation score, i.e., the higher the p_under value, the less likely the rare variant is to cause under-gene expression, and the lower the p_under value, the more likely the rare variant is to cause under-gene expression.

[0073] 4 also shows a "p_over" column 412 that identifies the excess causality score for the rare variant. The p_over value is inversely proportional to the excess causality score, i.e., the higher the p_over value, the less likely the rare variant is to cause excess gene expression, and the lower the p_over value, the more likely the rare variant is to cause excess gene expression.

[0074] Performance results as objective indicators of non-obviousness and inventive step Comparison and detection of differences between a sample distribution and a reference distribution, or sample outliers from a reference distribution, can include the use of parametric and non-parametric statistical tests, such as (one-sided or two-sided) t-tests, Mann-Whitney Rank Sum tests, and others, including the use of z-scores, such as z-scores based on Median Absolute Deviation (e.g., those used by Stumm et al. 2014, Prenat Diagn 34:185). When comparing a distribution to a reference distribution (or outliers from the reference distribution), in certain embodiments, the comparison is distinguished (and / or identified as significantly different) if the mean, median, or separation of an individual sample is greater than about 1.5, 1.6, 1.7, 1.8, 1.9, 1.95, 1.97, 2.0, or about 2.0 standard deviation ("SD") of the reference distribution, and / or if an individual sample separates from the reference distribution by a z-score greater than about 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.5, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.75, 4.0, 4.5, 5.0, or greater than about 5.0.

[0075] In certain embodiments, a parameter (such as a mean, median, standard deviation, median absolute deviation, or z-score) is calculated for a set of samples. In some such embodiments, such calculated parameters are used to identify outliers from those test samples that have been detected / analyzed. In certain embodiments, such parameters are calculated from all test samples without knowledge of the identity of any outliers (e.g., a "masked" analysis). In other certain embodiments, such parameters are calculated from a set of reference samples known to be normal (not outlying) or test samples that are presumed to be such normal (or not).

[0076] In certain embodiments, in the context of a dataset, a z-score (or an equivalent statistic based on the distribution pattern of replicates of a parameter) can be calculated to identify outlying data points (e.g., representing extreme levels of gene expression (under- or over-)), data representing such data points are removed from the dataset, and subsequent z-score analyses are performed on the dataset to attempt to identify further outliers. Such iterative z-score analyses can be particularly useful when two or more samples may skew a single z-score analysis and / or when follow-up tests are available to confirm false positives, and thus avoiding false negatives is potentially more important than the (initial) identification of false positives.

[0077] 5 shows that performance results 500, measured as z-scores across individuals and genes, progressively improve by successively correcting for multiple confounding factors using the fitted causal model 300. In particular, under-representation z-scores are determined to measure the correlation between rare variants and under-gene expression, and over-representation z-scores are determined to measure the correlation between rare variants and over-gene expression.

[0078] At 502, the underestimated z-score is 1.7 and the overestimated z-score is 1.2. None of the confounders are corrected for at this stage.

[0079] At 512, the under-reported z-score is increased to 2.25 and the over-reported z-score is increased to 2. At this stage, the PEER coefficient is corrected.

[0080] At 522, the under-reported z-score increases to 2.6 and the over-reported z-score increases to 2.2. At this stage, 30 gPCs are corrected.

[0081] At 532, the under-represented z-scores are increased to 3 and the over-represented z-scores are decreased to 2. At this stage, local cis eQTL effects are corrected.

[0082] At 542, the under-z score increases to 3.5 and the over-z score increases to 3. At this stage, distal trans eQTL effects are corrected.

[0083] Figure 6 compares, for a given p-value 604 from the fitted causal model 300, the counts 611, 613 of actual rare variants 622 and shuffled (or randomly selected) variants 632 identified as causing under- and over-expression 612 and 616 gene expression. As shown in Figure 6, the count-based correlation for the true rare variants 622 is consistently higher than for the shuffled variants 632. In Figure 6, the x-axis is the distance 662, 666 of the true rare variants 622 and shuffled variants 632 from the transcription start site (TSS). In Figure 6, the y-axis is the counts 611, 613 of the true rare variants 622 and shuffled variants 632.

[0084] Figure 7 shows odds ratios 711, 713 comparing the causal role of the true rare variant 622 and the shuffled variant 632 with respect to under- and over-expression 612 and over-expression 616 for a given p-value 604 from the fitted causal model 300. As shown in Figure 7, the consistently high odds ratios 711, 713 for different TSS distances 662, 666 demonstrate the strong causal role of the true rare variant 622.

[0085] Base-resolution evolutionary conservation prediction chromatin model Figure 8 shows a first exemplary architecture 800 of the disclosed chromatin model 802. Figure 9 shows a second exemplary architecture 900 of the disclosed chromatin model 802. Figure 10 shows a third exemplary architecture 1000 of the disclosed chromatin model 802. Figure 11 shows a fourth exemplary architecture 1100 of the disclosed chromatin model 802. Figure 12 shows a fifth exemplary architecture 1200 of the disclosed chromatin model 802.

[0086] As shown in FIG. 8, the chromatin model 802 includes groups of residual blocks arranged in order from lowest to highest. Each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block. In some embodiments, the dilated convolution rate progresses non-exponentially from lower residual block groups to higher residual block groups. In other embodiments, the dilated convolution rate progresses exponentially. The size of the convolution window varies between groups of residual blocks, and each residual block includes at least one batch normalization layer, at least one rectified linear unit (ReLU) layer, at least one dilated convolution layer, and at least one residual connection.

[0087] As shown in Figure 9, the dimensionality of the input is (Cu + L + Cd) x 4, where Cu is the number of upstream flanking context bases, Cd is the number of downstream flanking context bases, and L is the number of bases in the input promoter sequence. The dimensionality of the output is 4 x L. In some embodiments, each group of residual blocks produces an intermediate output by processing the previous input, and the dimensionality of the intermediate output is (I - [{(W - 1) * D} *A] × N, where I is the dimensionality of the preceding input, W is the convolution window size of the residual block, D is the dilation convolution rate of the residual block, A is the number of dilated convolution layers in the group, and N is the number of convolution filters in the residual block.

[0088] The exemplary architecture 1000 is used when the input has 200 upstream adjacent context bases (Cu) to the left of the input sequence and 200 downstream adjacent context bases (Cd) to the right of the input sequence. The length (L) of the input sequence can be any length, such as 3001. In the exemplary architecture 1000, each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and one dilated convolution rate, and each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and four dilated convolution rates. In other architectures, each residual block has 32 convolution filters, 11 convolution window sizes, and one dilated convolution rate.

[0089] The exemplary architecture 1100 is used when the input has 1000 upstream adjacent context bases (Cu) on the left side of the input sequence and 1000 downstream adjacent context bases (Cd) on the right side of the input sequence. The length (L) of the input sequence can be any length, such as 3001. In the exemplary architecture 1100, there are at least three groups of four residual blocks and at least three skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate. Each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates. Each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates.

[0090] The exemplary architecture 1200 is used when the input has 5000 upstream adjacent context bases (Cu) to the left of the input sequence and 5000 downstream adjacent context bases (Cd) to the right of the input sequence. The length (L) of the input sequence can be any length, such as 3001. In the exemplary architecture 1200, there are at least four groups of four residual blocks and at least four skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate; each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates; each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates; and each residual block in the fourth group has 32 convolution filters, 41 convolution window sizes, and 25 dilated convolution rates.

[0091] Generally speaking, the chromatin model 802 can be a rule-based model, a tree-based model, or a machine learning model, examples of which include a multilayer perceptron (MLP), a feed-forward neural network, a fully connected neural network, a fully convolutional neural network, a sequence-to-sequence (Seq2Seq) model such as ResNet or WaveNet, a semantic segmentation neural network, and a generative adversarial network (GAN) (e.g., CycleGAN, StyleGAN, pixelRNN, text-2-image, DiscoGAN, IsGAN).

[0092] In some embodiments, the chromatin model 802 is a Transformer, Vision Transformer (ViT), Bidirectional Transformer (BERT), Detection Transformer (DETR), Deformable DETR, UP-DETR, DeiT, Swin, GPT, iGPT, GPT-2, GPT-3, BERT, SpanBERT, RoBERTa, XLNet, ELECTRA, UniLM, BART, T5, ERNIE (THU), KnowBERT, DeiT-Ti, DeiT-S, DeiT-B, T2T-ViT-14, T2T-ViT-19, T2T-ViT-24, PVT-Small, PVT-Medium, PV T-Large, TNT-S, TNT-B, CPVT-S, CPVT-S-GAP, CPVT-B, Swin-T, Swin-S, Swin-B, Twins-SVT-S, Twins-SVT-B, Twins- SVT-L, Shuffle-T, Shuffle-S, Shuffle-B, XCiT-S12 / 16, CMT-S, CMT-B, VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, MoCo v3, ACT, TSP, Max-DeepLab, VisTR, SETR, Hand-Transformer, HOT-Net, METRO, Image Transformer, Taming transformer, TransGAN, IPT, TTSR, STTN, Masked Transformer, CLIP, DALL-E, Cogview, UniT, ASH, TinyBert, FullyQT, ConvBert, FCOS, Faster R-CNN+FPN, DETR-DC5, TSP-FCOS, TSP-RCNN, ACT+MKDD(L=32), ACT+MKDD(L=16), SMCA, Efficient These may include self-attention mechanisms such as DETR, UP-DETR, UP-DETR, ViTB / 16-FRCNN, ViT-B / 16-FRCNN, PVT-Small+RetinaNet, Swin-T+RetinaNet, Swin-T+ATSS, PVT-Small+DETR, TNT-S+DETR, YOLOS-Ti, YOLOS-S, and YOLOS-B.

[0093] In some embodiments, examples of chromatin model 802 include a convolutional neural network (CNN) with multiple convolutional layers, a recurrent neural network (RNN) such as a long short-term memory network (LSTM), a bi-directional LSTM (Bi-LSTM), or gated recurrent units, as well as a combination of both CNNs and RNNs.

[0094] In some embodiments, the chromatin model 802 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, point-wise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The chromatin model 802 can use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 oss, L2 loss, smoothed L1 loss, and Huber loss. The chromatin model 802 can use any parallelism, efficiency, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharding, parallel calls for map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The chromatin model 802 can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual joins, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions (e.g., rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid, and hyperbolic tangent (tanh)), etc.), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.

[0095] In some embodiments, chromatin model 802 can be a linear regression model, a logistic regression model, an elastic net model, a support vector machine (SVM), a random forest (RF), a decision tree, or a boosted decision tree (e.g., XGBoost), or some other tree-based logic (e.g., metric tree, kd tree, R-tree, universal B-tree, X-tree, ball tree, locality-sensitive hashing, and inverted index). Chromatin model 802 can, in some embodiments, be an ensemble of multiple models.

[0096] In some embodiments, the chromatin model 802 can be trained using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that can be used to train the model include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the model include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0097] Input sequence 13 shows input generation logic 1304 that accesses a sequence database 1302 (e.g., GTEx, RNA-Seq, WGS) and generates input sequences 1314. The input sequences 1314 include a target sequence 1324. The target sequence 1324 is flanked by a right sequence 1322 with downstream context bases and a left sequence 1326 with upstream context bases.

[0098] Base resolution output sequence 14 shows one embodiment of base-resolution evolutionary conservation prediction 1400 by chromatin model 802. Chromatin model 802 processes an input base sequence 1314 and generates an alternative representation 1406 of the input base sequence 1314. In one embodiment, alternative representation 1406 is a convolutional representation of input base sequence 1314 as it is processed through a cascade of convolutional layers in chromatin model 802.

[0099] The chromatin output generation logic 1408 processes the alternate representation 1406 of the input base sequence 1314 and generates an output sequence 1410 of chromatin output for each base for each target base in the target base sequence 1324 .

[0100] 15 shows an example of an output sequence 1410 corresponding to a target base sequence 1324. The chromatin output for each given base in the output sequence 1410 for a given target base at a given position in the target base sequence 1324 specifies a measure of evolutionary conservation of the given target base across multiple species. The multiple species can include homologous species.

[0101] Predicting the functional consequences of variants relies, at least in part, on the assumption that amino acids important to a protein family have been conserved through evolution by negative selection (i.e., amino acid changes at these sites have been deleterious in the past), and that mutations at these sites are likely to be pathogenic (disease-causing) in humans. Homologous sequences of a target protein are collected and aligned in a multiple sequence alignment (MSA). A conservation metric is calculated based on the weighted scores and / or frequencies of different amino acids observed at target positions in the MSA.

[0102] MSA is generally an alignment of three or more biological sequences, proteins, or nucleic acids of similar length. From the alignment, the degree of homology can be inferred and evolutionary relationships between sequences can be studied. MSA is also a tool used to identify evolutionary relationships and common patterns between genes. Alignments are generated and analyzed using computational algorithms. Most MSA algorithms use dynamic and heuristic methods. One of the goals of alignment is to detect structural or functional identities and similarities between residues in a protein sequence relative to other protein sequences.

[0103] The homology information for aligned sequences in MSA can be represented by two matrices (evolutionary conservation metrics): a position-specific scoring matrix (PSSM) and a position-specific frequency matrix (PSFM). PSSM and PSFM reflect the conservation of residues at specific positions in protein chains based on evolutionary information.

[0104] In some embodiments, the chromatin output for each given base further specifies a measure of transcription initiation for a given target base at a given position.

[0105] In one embodiment, the measure of evolutionary conservation is a phylogenetic P-value (phyloP) score that identifies deviations from the null model of neural substitution to detect a decrease in the rate of substitution of a given target base at a given position as conservation and an increase in the rate of substitution of a given target base at a given position as acceleration.

[0106] In one embodiment, the measure of evolutionary conservation is the phastCons score, which specifies the posterior probability of a given target base at a given position having a conserved or non-conserved state.

[0107] In one embodiment, the measure of evolutionary conservation is a genomic evolutionary rate profiling (GERP) score, which identifies the reduction in the number of substitutions of a given target base at a given position across multiple species.

[0108] In one embodiment, the measure of transcription initiation is a cap analysis of gene expression (CAGE) score, which specifies the frequency of transcription initiation of a given target base at a given position.

[0109] In some embodiments, the chromatin output for each given base further specifies a confounder signal level for a given target base at a given position. In one embodiment, the confounder signal level identifies a DNase I-hypersensitive site (DHS). In one embodiment, the confounder signal level identifies an assay for transposase-accessible chromatin using sequencing (ATAC-Seq). In another embodiment, the confounder signal level identifies transcription factor (TF) binding. In yet another embodiment, the confounder signal level identifies a histone modification (HM) mark. In yet another embodiment, the confounder signal level identifies a DNA methylation mark.

[0110] Base-resolution gene expression level prediction Gene Expression Models 16 illustrates one embodiment of the disclosed gene expression model 1600. Generally speaking, the gene expression model 1602 can be a rule-based model, a tree-based model, or a machine learning model. Examples include multilayer perceptrons (MLPs), feed-forward neural networks, fully connected neural networks, fully convolutional neural networks, sequence-to-sequence (Seq2Seq) models such as ResNet and WaveNet, semantic segmentation neural networks, and generative adversarial networks (GANs) (e.g., CycleGAN, StyleGAN, pixelRNN, text-2-image, DiscoGAN, IsGAN).

[0111] In some embodiments, the gene expression model 1602 is a Transformer, Vision Transformer (ViT), Bidirectional Transformer (BERT), Detection Transformer (DETR), Deformable DETR, UP-DETR, DeiT, Swin, GPT, iGPT, GPT-2, GPT-3, BERT, SpanBERT, RoBERTa, XLNet, ELECTRA, UniLM, BART, T5, ERNIE (THU), KnowBERT, DeiT-Ti, DeiT-S, DeiT-B, T2T-ViT-14, T2T-ViT-19, T2T-ViT-24, PVT-Small, PVT-Medium, PV T-Large, TNT-S, TNT-B, CPVT-S, CPVT-S-GAP, CPVT-B, Swin-T, Swin-S, Swin-B, Twins-SVT-S, Twins-SVT-B, Twins- SVT-L, Shuffle-T, Shuffle-S, Shuffle-B, XCiT-S12 / 16, CMT-S, CMT-B, VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, MoCo v3, ACT, TSP, Max-DeepLab, VisTR, SETR, Hand-Transformer, HOT-Net, METRO, Image Transformer, Taming transformer, TransGAN, IPT, TTSR, STTN, Masked Transformer, CLIP, DALL-E, Cogview, UniT, ASH, TinyBert, FullyQT, ConvBert, FCOS, Faster R-CNN+FPN, DETR-DC5, TSP-FCOS, TSP-RCNN, ACT+MKDD(L=32), ACT+MKDD(L=16), SMCA, Efficient These may include self-attention mechanisms such as DETR, UP-DETR, UP-DETR, ViTB / 16-FRCNN, ViT-B / 16-FRCNN, PVT-Small+RetinaNet, Swin-T+RetinaNet, Swin-T+ATSS, PVT-Small+DETR, TNT-S+DETR, YOLOS-Ti, YOLOS-S, and YOLOS-B.

[0112] In some embodiments, examples of gene expression models 1602 include convolutional neural networks (CNNs) with multiple convolutional layers, recurrent neural networks (RNNs) such as long short-term memory networks (LSTMs), bidirectional LSTMs (Bi-LSTMs), or gated recurrent units, as well as combinations of both CNNs and RNNs.

[0113] In some embodiments, gene expression model 1602 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, point-wise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. Gene expression model 1602 can use one or more loss functions such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. The gene expression model 1602 can use any parallelism, efficiency, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharding, parallel calls for map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The gene expression model 1602 can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions (rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid, and hyperbolic tangent (tanh)), etc.), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.

[0114] In some embodiments, gene expression model 1602 can be a linear regression model, a logistic regression model, an elastic net model, a support vector machine (SVM), a random forest (RF), a decision tree, or a boosted decision tree (e.g., XGBoost), or some other tree-based logic (e.g., metric tree, kd tree, R-tree, universal B-tree, X-tree, ball tree, locality sensitive hashing, and inverted index). Gene expression model 1602 can, in some embodiments, be an ensemble of multiple models.

[0115] In some embodiments, the gene expression model 1602 can be trained using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that can be used to train the model include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the model include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMS Grad.

[0116] Gene expression model 1602 processes output sequence 1410 and generates an alternate representation 1612 of output sequence 1410. In one embodiment, alternate representation 1612 is a convolutional representation of output sequence 1410 when output sequence 1410 is processed by a cascade of convolutional layers of gene expression model 1602.

[0117] The gene expression model output generation logic 1622 processes the alternate representation 1612 of the output sequence 1410 to generate a gene expression output sequence 1632 of gene expression outputs for each base for each target base in the target base sequence 1324 .

[0118] The gene expression output per given base in gene expression output sequence 1632 for a given target base at a given position specifies a measure of the gene expression level of the given target base at the given position. In one embodiment, the gene expression level is measured in a per-base metric, such as CAGE transcription start site (CTSS). In another embodiment, the gene expression level is measured in a per-gene metric, such as transcripts per million (TPM) or reads per kilobase of transcript (RPKM). In yet another embodiment, the gene expression level is measured in a per-gene metric, such as fragments per kilobase million (FPKM).

[0119] Variant pathogenicity classification 17 shows an example of a reference sequence 1702 and an alternative sequence 1722 (or substitute sequence). Alternative sequence 1712 differs from reference sequence 1702 by variant nucleotide 1722.

[0120] 18 illustrates one embodiment of the disclosed variant classification logic 1800. The variant classification logic 1800 is further configured to include reference input generation logic 1802 that accesses the sequence database 1302 and generates a reference base sequence 1702. The reference base sequence 1702 includes a reference target base sequence. The reference target base sequence includes a reference base at a position to be analyzed. The reference base is flanked by a right base sequence with a downstream context base and a left base sequence with an upstream context base.

[0121] The variant classification logic 1800 is further configured to include alternative input generation logic 1812 that accesses the sequence database 1302 and generates an alternative base sequence 1712. The alternative base sequence 1712 includes an alternative target base sequence. The alternative target base sequence includes an alternative base 1722 at the position under analysis. The alternative base 1722 is flanked by a right base sequence with a downstream context base and a left base sequence with an upstream context base.

[0122] The variant classification logic 1800 is further configured to include reference processing logic 1822 that causes the chromatin model 802 to process the reference base sequence 1702 and generate an alternate representation 1832 of the reference base sequence 1702, and further causes the chromatin output generation logic 1408 to process the alternate representation 1832 of the reference base sequence 1702 and generate a reference output sequence 1842 of reference chromatin outputs for each base for each reference target base in the reference target base sequence.

[0123] The reference chromatin output for each given base in the reference output sequence 1842 for a given reference target base at a given position in the reference target base sequence specifies a measure of evolutionary conservation of the given reference target base across multiple species.

[0124] The variant classification logic 1800 is further configured to include alternative processing logic 1852 that causes the chromatin model 802 to process the alternative base sequence 1712 and generate alternative representations 1862 of the alternative base sequence 1712, and further causes the chromatin output generation logic 1408 to process the alternative representations 1862 of the alternative base sequence 1712 and generate alternative output sequences 1872 of alternative chromatin outputs for each base for each alternative target base in the alternative target base sequence.

[0125] The alternative chromatin output for each given base in the alternative output sequence 1872 for a given alternative target base at a given position in the alternative target base sequence specifies a measure of evolutionary conservation of the given alternative target base across multiple species.

[0126] 19 illustrates one embodiment of the disclosed pathogenicity prediction logic 1900. The variant classification logic 1800 is further configured with pathogenicity prediction logic 1900 that compares the reference output sequence 1842 and the alternate output sequence 1872 position by position and generates a delta sequence 1912 having position-by-position sequence diffs for positions in the reference output sequence 1842 and the alternate output sequence 1872.

[0127] The pathogenicity prediction logic 1900 is further configured to generate a pathogenicity prediction 1922 for the alternative base 1722 depending on the delta sequence 1912. In one embodiment, the pathogenicity prediction logic 1900 is further configured to accumulate the sequence diffs for the positions into a cumulative sequence value and generate a pathogenicity prediction 1922 for the alternative base 1722 according to the cumulative sequence value. In some embodiments, the cumulative sequence value is the average or maximum of the sequence diffs for the positions. In other embodiments, the cumulative sequence value is the sum of the sequence diffs for the positions.

[0128] In some embodiments, the pathogenicity prediction logic 1900 is further configured to compare each portion of the reference output sequence 1842 and the alternative output sequence 1872 position by position and generate a delta subsequence having a position-by-position subsequence diff for positions within each portion.

[0129] In one embodiment, each portion spans adjacent positions to the left and right around the position under analysis. In some embodiments, the pathogenicity prediction logic 1900 is further configured to generate a pathogenicity prediction for the alternative base 1722 depending on the delta subsequence. In one embodiment, the pathogenicity prediction can be a score between 0 and 1, with 0 representing absolute benign and 1 representing absolute pathogenic. In other embodiments, a cutoff can be used, for example, a pathogenicity score above 5 can be considered pathogenic and below 5 can be considered benign.

[0130] In some embodiments, the pathogenicity prediction logic 1900 is further configured to accumulate the subsequence diffs for the positions into a cumulative subsequence value and generate a pathogenicity prediction for the alternative base 1722 according to the cumulative subsequence value. In one embodiment, the cumulative subsequence value is the average of the subsequence diffs for the positions. In another embodiment, the cumulative subsequence value is the sum or max of the subsequence diffs for the positions.

[0131] Computer Systems 20 is an exemplary computer system 2000 that can be used to implement various aspects of the disclosed technology. Computer system 2000 includes at least one central processing unit (CPU) 2024 that communicates with a number of peripheral devices via a bus subsystem 2022. These peripheral devices may include, for example, a storage subsystem 2010, including memory devices and a file storage subsystem 2018, a user interface input device 2020, a user interface output device 2028, and a network interface subsystem 2026. The input and output devices enable user interaction with computer system 2000. Network interface subsystem 2026 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0132] In one embodiment, the chromatin model 802 is communicatively linked to a storage subsystem 2010 and a user interface input device 2020 .

[0133] The user interface input devices 2020 can include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, audio input devices such as a voice recognition system and a microphone, and other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computer system 2000.

[0134] The user interface output devices 2028 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and ways for outputting information from the computer system 2000 to a user or to another machine or computer system.

[0135] The storage subsystem 2010 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 2030.

[0136] The processor 2030 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 2030 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 2078 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX20 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.

[0137] The memory subsystem 2012 used in the storage subsystem 2010 may include several memories, including a main random access memory (RAM) 2014 for storing instructions and data during program execution, and a read only memory (ROM) 2016 in which fixed instructions are stored. The file storage subsystem 2018 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 2018 in the storage subsystem 2010 or in another machine accessible by the processor.

[0138] Bus subsystem 2022 provides a mechanism for allowing the various components and subsystems of computer system 2000 to communicate with each other as intended. Although bus subsystem 2022 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0139] The computer system 2000 itself can be of a variety of types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 2000 shown in Figure 20 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 2000 can have more or fewer components than the computer system shown in Figure 20.

[0140] Terms The disclosed technology can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following embodiments.

[0141] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be embodied in the form of a computer product including a non-transitory computer-readable storage medium with computer-usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be embodied in the form of an apparatus including a memory and at least one processor, coupled to the memory, operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be embodied in the form of a means for performing one or more of the method steps described herein, which means can include (i) hardware modules, (ii) software modules executing on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which (i)-(iii) implements particular technology described herein, and the software modules are stored on a computer-readable storage medium (or multiple such media).

[0142] The clauses described in this section can be combined as features. For the sake of brevity, combinations of features are not individually listed and are not repeated for each base set of features. The reader will understand how features identified in clauses described in this section can be readily combined with sets of basic features identified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or limiting, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.

[0143] Another implementation of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.

[0144] The inventors disclose the following provisions:

[0145] Clause Set 1 1. A computer-implemented method for identifying rare variants that cause extreme levels of gene expression, comprising: accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, the extreme levels of gene expression being determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; Fitting a causal model to determine the causal relationship between rare variants and extreme levels of gene expression in outlier individuals while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein the particular causality score for a particular rare variant indicates the likelihood that the particular rare variant causes extreme levels of gene expression in outlier individuals whose gene sequences include the particular rare variant. 2. The computer-implemented method of clause 1, wherein the causality score is a probability value (p-value). 3. The computer-implemented method of clause 1, wherein the p-value is determined by Pearson correlation coefficient. 4. The computer-implemented method of clause 1, wherein the causal model is a logistic regression model, a linear regression model, an analysis of covariance (ANCOVA) model, and / or a multivariate analysis of covariance (MANCOVA) model. 5. The computer-implemented method of clause 4, wherein the fitted causal model determines causality by predicting a specific gene expression level of a specific gene in a specific chromosome depending on a variant-driven gene expression level caused by a specific rare variant. 6. The computer-implemented method of clause 5, wherein the fitted causal model measures the contribution of variant-driven gene expression levels as a variant effect covariate. 7. The computer-implemented method of clause 1, wherein the plurality of confounding factors comprises distal trans-expression quantitative trait loci (eQTL) effects. 8. The computer-implemented method of clause 7, wherein the fitted causal model controls for distal trans eQTL effects by predicting specific gene expression levels dependent on transgene expression levels caused by other genes in other chromosomes. 9. The computer-implemented method of clause 8, wherein the fitted causal model measures the contribution of transgene expression level as a trans-effect covariate. 10. The computer-implemented method of clause 1, wherein the plurality of confounding factors comprises local cis-eQTL effects. 11. The computer-implemented method of clause 10, wherein the fitted causal model controls for local cis eQTL effects by predicting specific gene expression levels dependent on cis gene expression levels caused by the presence of multiple common variants in the vicinity of a specific gene. 12. The computer-implemented method of clause 11, wherein the neighborhood is defined by an offset from a transcription start site (TSS) in a particular gene. 13. The computer-implemented method of clause 11, wherein the fitted causal model measures the contribution of cis gene expression levels as cis-effect covariates. 14. The computer-implemented method of clause 1, wherein the plurality of confounding factors includes population structure and ancestry effects. 15. The computer-implemented method of clause 14, wherein the population structure and ancestry effects are represented by one or more genotype-based principal components (gPCs). 16. The computer-implemented method of clause 15, wherein the fitted causal model controls for population structure and ancestry effects by predicting specific gene expression levels depending on gPC gene expression levels caused by gPC. 17. The computer-implemented method of clause 16, wherein the fitted causal model measures the contribution of gPC gene expression levels as a population structure and ancestry effect covariate. 18. The computer-implemented method of clause 1, wherein the plurality of confounding factors comprises a probabilistic estimate of expression residual (PEER) effect. 19. The computer-implemented method of clause 18, wherein the fitted causal model controls for the PEER effect by predicting specific gene expression levels depending on the PEER-induced gene expression levels. 20. The computer-implemented method of clause 19, wherein the fitted causal model measures the contribution of PEER gene expression levels as a PEER effect covariate. 21. The computer-implemented method of clause 1, wherein the plurality of confounding factors includes environmental effects. 22. The computer-implemented method of clause 21, wherein the fitted causal model controls for environmental effects by predicting specific gene expression levels depending on the environmental gene expression levels caused by the environmental effects. 23. The computer-implemented method of clause 22, wherein the fitted causal model measures the contribution of environmental gene expression levels as an environmental effect covariate. 23. The computer-implemented method of clause 1, wherein the plurality of confounding factors comprises a gender effect, a batch effect, a genotyping platform effect, and a library construction protocol effect. 24. The computer-implemented method of clause 1, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 25. The computer-implemented method of clause 24, further comprising determining a causal relationship between rare variants and over-gene expression while controlling for multiple confounding factors. 26. The computer-implemented method of clause 25, further comprising generating an excess causality score for the rare variant, wherein the particular excess causality score for the particular rare variant indicates the likelihood that the particular rare variant causes excess gene expression in an outlier individual whose gene sequence includes the particular rare variant. 27. The computer-implemented method of clause 26, wherein the excess causality score is an excess probability value (excess p-value). 28. The computer-implemented method of clause 27, wherein the excess p-value is determined by the Pearson correlation coefficient. 29. The computer-implemented method of clause 27, wherein the excess p-value identifies the statistically confounding-free likelihood that a rare variant increases gene expression in a gene that otherwise has lower gene expression compared to other genes in the gene set. 30. The computer-implemented method of clause 24, further comprising determining a causal relationship between rare variants and under-expression of genes while controlling for multiple confounding factors. 31. The computer-implemented method of clause 30, further comprising generating an under-causation score for the rare variant, wherein the particular under-causation score for the particular rare variant indicates the likelihood that the particular rare variant causes under-gene expression in an outlier individual whose gene sequence includes the particular rare variant. 32. The computer-implemented method of clause 31, wherein the under-causation score is an under-causation probability value (under-causation p-value). 33. The computer-implemented method of clause 32, wherein the undervalued p-value is determined by Pearson correlation coefficient. 34. The computer-implemented method of clause 32, wherein the underrepresentation p-value identifies the statistically unconfounded likelihood that a rare variant reduces gene expression in genes that otherwise have higher gene expression compared to other genes in the gene set. 35. The computer-implemented method of clause 1, wherein the rare variant is a non-code variant. 36. The computer-implemented method of clause 17, wherein the non-coding variants include 5 prime untranslated region (UTR) variants, 3 prime UTR variants, enhancer variants, and promoter variants. 37. The computer-implemented method of clause 1, wherein the gene expression levels are further stratified into tissue-specific gene expression levels for a plurality of tissues. 38. The computer-implemented method of clause 37, wherein the gene expression level of each gene in each tissue is normalized using quantile normalization. 39. The computer-implemented method of clause 38, wherein the causal model is fitted separately for each tissue. 40. The computer-implemented method of clause 1, wherein the causal model is fitted using stratification. 41. The computer-implemented method of clause 1, further comprising generating a ranking of rare variants based on the causality scores. 42. The computer-implemented method of clause 26, further comprising generating a ranking of rare variants based on excess causality scores. 43. The computer-implemented method of clause 31, further comprising generating a ranking of rare variants based on the under-causation scores. 44. The computer-implemented method of clause 1, wherein the rare variant is a singleton variant. 45. The computer-implemented method of clause 44, wherein the singleton variant occurs in only one of the outlier individuals. 46. ​​A system comprising one or more processors coupled to a memory, the memory being loaded with computer instructions for identifying rare variants that cause extreme levels of gene expression, the instructions, when executed on the processor, accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, the extreme levels of gene expression being determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; Fitting a causal model to determine the causal relationship between rare variants and extreme levels of gene expression in outlier individuals while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein the particular causality score for the particular rare variant indicates the likelihood that the particular rare variant causes extreme levels of gene expression in outlier individuals whose gene sequences include the particular rare variant. 47. The system of clause 46, further performing actions including performing clauses 2 to 45. 48. A non-transitory computer-readable storage medium having stored thereon computer program instructions for identifying rare variants that cause extreme levels of gene expression, the instructions, when executed on a processor, performing: accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, the extreme levels of gene expression being determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; Fitting a causal model to determine the causal relationship between rare variants and extreme levels of gene expression in outlier individuals while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein the particular causality score for the particular rare variant indicates the likelihood that the particular rare variant causes extreme levels of gene expression in outlier individuals whose gene sequences include the particular rare variant. 49. The non-transitory computer-readable storage medium of clause 46, wherein implementing the method further comprises performing clauses 2 to 45.

[0146] Clause Set 2 1. An artificial intelligence-based system for detecting changes in gene expression with base-by-base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a chromatin model that processes an input sequence and generates alternative representations of the input sequence; chromatin output generation logic that processes the alternative representations of the input base sequence and generates an output sequence of chromatin outputs for each base for each target base in the target base sequence; An artificial intelligence-based system in which the chromatin output for each given base in the output sequence for a given target base at a given position in the target sequence specifies a measure of evolutionary conservation of the given target base across multiple species. 2. The artificial intelligence-based system of clause 1, wherein the chromatin output for each given base further specifies a measure of transcription initiation of the given target base at the given position. 3. The artificial intelligence-based system described in clause 1, wherein the measure of evolutionary conservation is a phylogenetic P-value (phyloP) score that identifies deviations from a null model of neural substitution to detect a decrease in the rate of substitution of a given target base at a given position as conservation and an increase in the rate of substitution of a given target base at a given position as acceleration. 4. The artificial intelligence-based system of clause 1, wherein the measure of evolutionary conservation is a phastCons score that identifies the posterior probability of a given target base at a given position having a conserved or non-conserved state. 5. The artificial intelligence-based system described in clause 1, wherein the measure of evolutionary conservation is a Genomic Evolutionary Rate Profiling (GERP) score that identifies a decrease in the number of substitutions of a given target base at a given position across multiple species. 6. The artificial intelligence-based system described in clause 2, wherein the measure of transcription initiation is a Cap Analysis of Gene Expression (CAGE) score, which identifies the frequency of transcription initiation of a given target base at a given position. 7. The artificial intelligence-based system of clause 1, wherein the chromatin output for each given base further identifies a confounder signal level for a given target base at a given position. 8. The artificial intelligence-based system described in clause 7, wherein confounder signal levels identify DNase I hypersensitive sites (DHS). 9. The artificial intelligence-based system of clause 7, wherein the confounder signal level identifies transcription factor (TF) binding. 10. The artificial intelligence-based system described in clause 7, wherein confounder signal levels identify histone modification (HM) marks. 11. The artificial intelligence-based system described in clause 7, wherein confounder signal levels identify DNA methylation marks. 12. a gene expression model that processes the output sequence and generates alternative representations of the output sequence; gene expression model output generation logic that processes the alternative representations of the output sequences and generates a gene expression output sequence of gene expression outputs for each base for each target base in the target base sequence; 10. The artificial intelligence-based system of clause 1, wherein the gene expression output for each given base in the gene expression output sequence for a given target base at a given position specifies a measure of the gene expression level of the given target base at the given position. 13. The artificial intelligence-based system of clause 12, wherein gene expression levels are measured in a per-base metric such as CAGE transcription start site (CTSS). 14. The artificial intelligence-based system of clause 12, wherein gene expression levels are measured in a pcr-gene metric such as transcripts per million (TPM) or reads per kilobase of transcript (RPKM). 15. The artificial intelligence-based system of clause 12, wherein gene expression levels are measured in a per-gene metric, such as fragments per million kilobases (FPKM). 16. The artificial intelligence-based system of clause 1, further configured to include variant classification logic. 17. The artificial intelligence-based system of clause 16, wherein the variant classification logic is further configured to include reference input generation logic that accesses a sequence database and generates a reference base sequence, the reference base sequence comprising a reference target base sequence, the reference target base sequence comprising a reference base at a position under analysis, the reference base being flanked by a right base sequence having a downstream context base and a left base sequence having an upstream context base. 18. The artificial intelligence-based system of clause 17, wherein the variant classification logic is further configured to include alternative input generation logic that accesses a sequence database to generate an alternative base sequence, the alternative base sequence comprising an alternative target base sequence, the alternative target base sequence comprising an alternative base at the position being analyzed, the alternative base flanked by a right base sequence with a downstream context base and a left base sequence with an upstream context base. 19. The variant classification logic is further configured to include reference processing logic that causes the chromatin model to process the reference base sequence and generate alternative representations of the reference base sequence, and causes the chromatin output generation logic to process the alternative representations of the reference base sequence and further generate a reference output sequence of reference chromatin outputs for each base for each reference target base in the reference target base sequence; 18. The artificial intelligence-based system of clause 17, wherein the reference chromatin output for each given base in the reference output sequence for a given reference target base at a given position in the reference target base sequence specifies a measure of evolutionary conservation of the given reference target base across multiple species. 20. The variant classification logic is further configured to include alternative processing logic that causes the chromatin model to process the alternative base sequences and generate alternative representations of the alternative base sequences, and further causes the chromatin output generation logic to process the alternative representations of the alternative base sequences and generate alternative output sequences of alternative chromatin outputs for each base for each alternative target base in the alternative target base sequences; 19. The artificial intelligence-based system of clause 18, wherein the alternative chromatin output for each given base in the alternative output sequence for a given alternative target base at a given position in the alternative target base sequence specifies a measure of evolutionary conservation of the given alternative target base across multiple species. 21. The artificial intelligence-based system of clause 20, wherein the variant classification logic is further configured to include pathogenicity prediction logic that compares the reference output sequence and the alternative output sequence position by position and generates a delta sequence having a position-by-position sequence diff for positions within the reference output sequence and the alternative output sequence. 22. The artificial intelligence-based system of clause 21, wherein the pathogenicity prediction logic is further configured to generate pathogenicity predictions for alternative bases dependent on the delta sequence. 23. The artificial intelligence-based system of clause 21, wherein the pathogenicity prediction logic is further configured to accumulate the sequence diff for the position into a cumulative sequence value and generate a pathogenicity prediction for the alternative base according to the cumulative sequence value. 24. The artificial intelligence-based system of clause 23, wherein the cumulative alignment value is the average or maximum of the alignment diff for the position. 25. The artificial intelligence-based system of clause 23, wherein the cumulative array value is the sum of the array diffs with respect to the position. 26. The artificial intelligence-based system of clause 21, wherein the pathogenicity prediction logic is further configured to positionally compare portions of the reference output sequence and the alternative output sequence and generate a delta subsequence having a positional subsequence diff for positions in each portion. 27. The artificial intelligence-based system of clause 26, wherein each portion spans adjacent positions to the left and right around the position being analyzed. 28. The artificial intelligence-based system of clause 26, wherein the pathogenicity prediction logic is further configured to generate pathogenicity predictions for alternative bases dependent on the delta subsequence. 29. The artificial intelligence-based system of clause 26, wherein the pathogenicity prediction logic is further configured to accumulate position-wise subsequence diffs into a cumulative subsequence value and generate a substitution-based pathogenicity prediction dependent on the cumulative subsequence value. 30. The artificial intelligence-based system of clause 29, wherein the cumulative subsequence value is an average of the subsequence diffs for the position. 31. The artificial intelligence-based system of clause 29, wherein the cumulative subarray value is the sum or max of the subarray diffs for the position. 32. The artificial intelligence-based system according to clause 1, wherein the target base sequence is a coding region of a gene. 33. The artificial intelligence-based system according to clause 1, wherein the target base sequence is a non-coding region of a gene. 34. The artificial intelligence-based system of clause 33, wherein the non-coding regions span a transcription start site, a 5 prime untranslated region (UTR), a 3 prime UTR, an enhancer, and a promoter. 35. The alternative base is a rare variant of an outlier in a cohort of outliers, 19. The artificial intelligence-based system of clause 18, wherein an outlier individual in the cohort of outlier individuals exhibits extreme levels of gene expression. 36. The artificial intelligence-based system of clause 35, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels. 37. The artificial intelligence-based system of clause 36, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 38. The artificial intelligence-based system according to clause 35, wherein the rare variant is a code variant. 39. The artificial intelligence-based system according to clause 35, wherein the rare variant is a non-code variant. 40. The artificial intelligence-based system of clause 39, wherein the non-coding variant is a promoter variant. 41. The artificial intelligence-based system of clause 39, wherein the non-coding variant is an enhancer variant. 42. The artificial intelligence-based system of clause 1, wherein the chromatin model has a first set of weights and the chromatin output generation logic has a second set of weights. 43. During training, a first set of chromatin model weights is trained from scratch to process input sequences and generate alternative representations of the input sequences; 43. The artificial intelligence based system of clause 42, wherein a second set of chromatin output generation logic weights is trained end-to-end from scratch with the first set of chromatin model weights to process alternative representations of the input base sequence to generate an output sequence. 44. During inference, the chromatin model uses the first set of trained weights, 64. The artificial intelligence-based system of clause 63, wherein during inference, the chromatin output generation logic uses the second set of trained weights. 45. The gene expression model has a third set of weights: 10. The artificial intelligence-based system of claim 1, wherein the gene expression output generation logic has a fourth set of weights. 46. ​​A third set of weights for the gene expression model is trained from scratch to process the output sequence and generate alternative representations of the output sequence. 46. ​​The artificial intelligence based system of clause 45, wherein a fourth set of weights for the gene expression output generation logic is trained end-to-end from scratch with the third set of weights for the gene expression model to process alternative representations of the output sequences and generate the gene expression output sequences. 47. During inference, the gene expression model uses a third set of trained weights, 47. The artificial intelligence-based system of clause 46, wherein during inference, the gene expression output generation logic uses the fourth set of trained weights. 48. During training, a first set of chromatin model weights is first trained to process input base sequences and generate alternative representations of the input base sequences, and then retrained on behalf of a third set of gene expression model weights to process output sequences and generate alternative representations of the output sequences; 46. ​​The artificial intelligence based system of clause 45, wherein the fourth set of weights of the gene expression output generation logic is trained end-to-end from scratch using the trained first set of weights substituted in the gene expression model to process alternative representations of the output sequences generated by the trained first set of weights substituted in the gene expression model and generate gene expression output sequences. 49. During inference, the chromatin model uses the first set of retrained weights, During inference, the chromatin output generation logic uses the second set of trained weights, During inference, the gene expression model uses the first set of retrained weights, 49. The artificial intelligence-based system of clause 48, wherein during inference, the gene expression output generation logic uses the fourth set of trained weights. 50. During training, a first set of chromatin model weights is initially trained from scratch to process input base sequences and generate alternative representations of the input base sequences; During training, a second set of weights for the chromatin output generation logic is trained end-to-end from the beginning with the first set of weights for the chromatin model to process alternative representations of the input base sequence and generate an output sequence; During training, the first set of trained weights of the chromatin model are then retrained to process the reference base sequence to generate alternative representations of the reference base sequence and to process the alternative base sequences to generate alternative representations of the alternative base sequences; 20. The artificial intelligence-based system of clause 19, wherein during training, the second set of trained weights of the chromatin output generation logic are then retrained end-to-end with the first set of trained weights of the chromatin model to process alternative representations of the reference base sequence to generate reference output sequences and to process alternative representations of the alternative base sequences to generate alternative output sequences. 51. During inference, the chromatin model uses the first set of retrained weights, 51. The artificial intelligence-based system of clause 50, wherein during inference, the chromatin output generation logic uses the retrained second set of weights. 52. The artificial intelligence-based system described in clause 21, wherein the pathogenicity prediction logic has a fifth set of weights. 53. During training, a first set of chromatin model weights is initially trained from scratch to process input base sequences and generate alternative representations of the input base sequences; During training, a second set of weights for the chromatin output generation logic is trained end-to-end from the beginning with the first set of weights for the chromatin model to process alternative representations of the input base sequence and generate an output sequence; During training, the first set of trained weights of the chromatin model and the second set of trained weights of the chromatin output generation logic are then retrained end-to-end to generate pathogenicity predictions for the alternative bases, the artificial intelligence based system of clause 52. 54. During inference, the chromatin model uses the first set of retrained weights, During inference, the chromatin output generation logic uses the second set of retrained weights, 54. The artificial intelligence-based system of clause 53, wherein during inference, the pathogenicity prediction logic uses the fifth set of trained weights. 55. The artificial intelligence-based system of clause 19, wherein the reference chromatin output for each given base in the reference output sequence for a given reference target base at a given position in the reference target base sequence specifies a confounder signal level for the given reference target base at the given position. 56. The artificial intelligence-based system described in clause 20, wherein the alternative chromatin output for each given base in the alternative output sequence for a given alternative target base at a given position in the alternative target base sequence specifies a confounder signal level for the given alternative target base at the given position. 57. The artificial intelligence-based system described in clause 1, wherein during training, the chromatin model and chromatin output generation logic are first trained end-to-end from scratch to translate analysis of input base sequences into base-specific evolutionary conserved chromatin sequences, and then retrained end-to-end to translate analysis of input base sequences into base-specific transcription start frequency chromatin sequences. 58. The artificial intelligence-based system described in clause 1, wherein during training, the chromatin model and chromatin output generation logic are first trained end-to-end from scratch to translate analysis of input base sequences into base-specific confounder signal-level chromatin sequences, and then retrained end-to-end to translate analysis of input base sequences into base-specific evolutionarily conserved chromatin sequences. 59. The artificial intelligence-based system described in clause 1, wherein during training, the chromatin model and chromatin output generation logic are first trained end-to-end from scratch to convert analysis of input base sequences into base-based confounder signal level chromatin sequences, and then retrained end-to-end to convert analysis of input base sequences into base-based transcription start frequency chromatin sequences. 60. The artificial intelligence-based system described in clause 1, wherein during training, the chromatin model and chromatin output generation logic are first trained end-to-end from scratch to convert the analysis of input base sequences into base-based confounder signal-level chromatin sequences, and then retrained end-to-end to convert the analysis of input base sequences into base-based evolutionary conserved chromatin sequences and base-based transcription start frequency chromatin sequences. 61. The artificial intelligence-based system of clause 1, further configured to include a first training set of training input base sequences that include variants confounded by multiple confounder effects. 62. The artificial intelligence-based system described in clause 61, wherein the confounding factor effects in the multiple confounding factor effects include interchromosomal effects, intragenic effects, population structure and ancestry effects, probabilistic estimation of expression residuals (PEER) effects, environmental effects, sex effects, batch effects, genotyping platform effects, and / or library construction protocol effects. 63. The artificial intelligence-based system of clause 61, further configured to include a second training set of training input sequences that includes variants that are not confounded by multiple confounder effects. 64. The artificial intelligence-based system of clause 63, wherein variants in the second training set are reliably determined to alter gene expression and cause extreme levels of gene expression. 65. The artificial intelligence-based system of clause 64, wherein the variants in the second training set include variants that cause overexpression, which increases gene expression levels. 66. The artificial intelligence-based system of clause 64, wherein the variants in the second training set include variants that cause underexpression, resulting in reduced gene expression levels. 67. The artificial intelligence-based system of clause 65, wherein the second training set identifies overexpression probabilities for variants that identify the likelihood of overexpression of the causal gene. 68. The artificial intelligence-based system of clause 66, wherein the second training set identifies an under-expression probability for a variant that identifies the likelihood that the variant causes the gene to be under-expressed. 69. Each variant in the second training set is a rare variant that occurs in an outlier individual of a cohort of outlier individuals; 65. The artificial intelligence-based system of clause 64, wherein an outlier individual in the cohort of outlier individuals exhibits extreme levels of gene expression. 70. The artificial intelligence-based system of clause 69, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels. 71. The artificial intelligence-based system of clause 70, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 72. The artificial intelligence-based system described in clause 69, wherein the rare variant is a code variant. 73. The artificial intelligence-based system described in clause 69, wherein the rare variant is a non-code variant. 74. The artificial intelligence-based system of clause 73, wherein the non-coding variant is a 5 prime untranslated region (UTR) variant, a 3 prime UTR variant, an enhancer variant, or a promoter variant. 75. The artificial intelligence-based system of clause 64, wherein the variants in the second training set span multiple tissue types. 76. The artificial intelligence-based system of clause 64, wherein the variants in the second training set span multiple cell types. 77. The artificial intelligence-based system of clause 1, wherein the input base sequence and the plurality of output sequences span a plurality of tissue types. 78. The artificial intelligence-based system of clause 1, wherein the input base sequence and the plurality of output sequences span a plurality of cell types. 79. The artificial intelligence-based system of clause 12, wherein the gene expression output sequences span multiple tissue types. 80. The artificial intelligence-based system of clause 12, wherein the gene expression output sequences span multiple cell types. 81. The artificial intelligence-based system of clause 1, wherein the chromatin model and chromatin output generation logic are first trained end-to-end on a first training set and then retrained on a second training set. 82. The artificial intelligence-based system of clause 1, wherein the variants in the second training set are used as a pathogenic set labeled with a first ground truth label indicating a gene expression change, and the common variants are used as a benign set labeled with a second ground truth label indicating no gene expression change. 83. The artificial intelligence-based system of clause 102, wherein the benign set is balanced for trinucleotide context, homopolymer, k-mer, neighborhood GC frequency, and sequencing depth. 84. The artificial intelligence-based system of clause 102, wherein, based on cutoff probabilities applied to the overexpression and underexpression probabilities, the variants in the second training set are divided into an overexpressed variant training set having a first ground truth label indicating increased gene expression, an overexpressed variant training set having a second ground truth label indicating decreased gene expression, and a neurally expressed variant training set indicating maintained gene expression. 85. The artificial intelligence-based system of clause 12, wherein the gene expression model and gene expression output generation logic are first trained end-to-end on a first training set and then retrained on a second training set. 86. The artificial intelligence-based system described in clause 1, wherein the chromatin model and chromatin output generation logic are first trained end-to-end on a first training set and then retrained on variants in a second training set that occur on odd-numbered chromosomes. 87. The artificial intelligence-based system of clause 12, wherein the gene expression model and gene expression output generation logic are first trained end-to-end on a first training set and then retrained on variants in a second training set that occur on odd chromosomes. 88. The artificial intelligence-based system described in clause 64, wherein the variants in the second training set are not used for training, but instead are used as a validation set to evaluate the performance of the trained chromatin model, the trained chromatin output generation logic, the trained gene expression model, and the trained gene expression output generation logic. 89. The artificial intelligence-based system of clause 88, wherein variants in the second training set present on even-numbered chromosomes are used as a validation set. 90. The artificial intelligence-based system described in clause 1, wherein the size of the target base sequence is varied during training to account for variations in the offset position of the transcription start site (TSS). 91. An artificial intelligence-based system for detecting changes in gene expression with base-by-base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a chromatin model that processes an input sequence and generates alternative representations of the input sequence; output generation logic that processes the alternative representations of the input base sequence and generates an output sequence of base-by-base chromatin outputs for each target base in the target base sequence; An artificial intelligence based system in which the chromatin output for each given base in the output sequence for a given target base at a given position in the target base sequence specifies a measure of transcription initiation for the given target base at the given position. 92. An artificial intelligence-based system for detecting changes in gene expression with base-by-base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a gene expression model that processes an input sequence and generates alternative representations of the input sequence; gene expression model output generation logic that processes the alternative representations of the input base sequence and generates a gene expression output sequence of gene expression outputs for each base in the target base sequence; An artificial intelligence based system in which the gene expression output for a given target base at a given position, per given base in a sequence, specifies a measure of the gene expression level of the given target base at the given position. 93. An artificial intelligence-based system for detecting changes in gene expression with base-by-base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a chromatin model that processes an input sequence and generates alternative representations of the input sequence; and chromatin output generation logic that processes alternative representations of the input base sequence and generates an output sequence of chromatin outputs for each base for each target base in the target base sequence. 94. The artificial intelligence-based system described in clause 93, wherein the chromatin output for each given base in the output sequence for a given target base at a given position in the target base sequence specifies a measure of evolutionary conservation of the given target base across multiple species. 95. The artificial intelligence-based system of clause 93, wherein the chromatin output for each given base further specifies a measure of transcription initiation of the given target base at the given position.

[0147] While the present invention has been disclosed with reference to the above-described preferred embodiments and examples, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims.

Claims

1. 1. A computer-implemented method for identifying rare variants that cause extreme levels of gene expression, comprising: accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, wherein the extreme levels of gene expression are determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; fitting a causal model to determine a causal relationship between the rare variant and the extreme level of gene expression in the outlier individual while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein a particular causality score for a particular rare variant indicates a likelihood that the particular rare variant causes extreme levels of gene expression in an outlier individual whose gene sequence includes the particular rare variant.

2. The computer-implemented method of claim 1 , wherein the causality score is a probability value (p-value).

3. The computer-implemented method of claim 1 , wherein the p-value is determined by a Pearson correlation coefficient.

4. The computer-implemented method of claim 1 , wherein the causal model is a logistic regression model, a linear regression model, an analysis of covariance (ANCOVA) model, and / or a multivariate analysis of covariance (MANCOVA) model.

5. 5. The computer-implemented method of claim 4, wherein the fitted causal model determines the causal relationship by predicting a specific gene expression level of a specific gene in a specific chromosome depending on a variant-driven gene expression level caused by a specific rare variant.

6. 6. The computer-implemented method of claim 5, wherein the fitted causal model measures the contribution of the variant-driven gene expression level as a variant effect covariate.

7. The computer-implemented method of claim 1 , wherein the plurality of confounding factors comprises distal trans-expressed quantitative trait loci (eQTL) effects.

8. 8. The computer-implemented method of claim 7, wherein the fitted causal model controls for the distal trans eQTL effects by predicting the particular gene expression level depending on transgene expression levels caused by other genes in other chromosomes.

9. 9. The computer-implemented method of claim 8, wherein the fitted causal model measures the contribution of the transgene expression level as a trans-effect covariate.

10. The computer-implemented method of claim 1 , wherein the plurality of confounding factors comprises local cis eQTL effects.

11. 11. The computer-implemented method of claim 10, wherein the fitted causal model controls for the local cis eQTL effects by predicting the expression level of a particular gene depending on the cis gene expression level caused by the presence of multiple common variants in the vicinity of the particular gene.

12. 12. The computer-implemented method of claim 11, wherein the vicinity is defined by an offset from a transcription start site (TSS) in the particular gene.

13. 12. The computer-implemented method of claim 11, wherein the fitted causal model measures the contribution of the cis gene expression levels as cis-effect covariates.

14. The computer-implemented method of claim 1 , wherein the plurality of confounding factors comprises population structure and ancestry effects.

15. 15. The computer-implemented method of claim 14, wherein the population structure and ancestry effects are represented by one or more genotype-based principal components (gPCs).

16. 16. The computer-implemented method of claim 15, wherein the fitted causal model controls for the population structure and ancestry effects by predicting the specific gene expression levels depending on the gPC gene expression levels caused by the gPC.

17. 17. The computer-implemented method of claim 16, wherein the fitted causal model measures the contribution of the gPC gene expression levels as population structure and ancestry effect covariates.

18. The computer-implemented method of claim 1 , wherein the plurality of confounding factors comprises probabilistic estimates of expression residuals (PEER) effects.

19. 20. The computer-implemented method of claim 18, wherein the fitted causal model controls the PEER effect by predicting the specific gene expression levels depending on the PEER gene expression levels caused by the PEER.

20. 20. The computer-implemented method of claim 19, wherein the fitted causal model measures the contribution of the PEER gene expression level as a PEER effect covariate.

21. The computer-implemented method of claim 1 , wherein the plurality of confounding factors includes environmental effects.

22. 22. The computer-implemented method of claim 21, wherein the fitted causal model controls the environmental effect by predicting the particular gene expression level depending on the environmental gene expression level caused by the environmental effect.

23. 23. The computer-implemented method of claim 22, wherein the fitted causal model measures the contribution of the environmental gene expression levels as an environmental effect covariate.

24. The computer-implemented method of claim 1 , wherein the plurality of confounding factors comprises a gender effect, a batch effect, a genotyping platform effect, and a library construction protocol effect.

25. The computer-implemented method of claim 1 , wherein the extreme levels of gene expression include over- and under-gene expression.

26. 25. The computer-implemented method of claim 24, further comprising determining the causal relationship between the rare variant and the overexpression of a gene while controlling for the plurality of confounding factors.

27. 26. The computer-implemented method of Claim 25, further comprising generating an excess causality score for the rare variant, wherein the particular excess causality score for the particular rare variant indicates the likelihood that the particular rare variant causes excess gene expression in an outlier individual whose gene sequence includes the particular rare variant.

28. 27. The computer-implemented method of claim 26, wherein the excess causality score is an excess probability value (excess p-value).

29. 28. The computer-implemented method of claim 27, wherein the excess p-value is determined by a Pearson correlation coefficient.

30. 28. The computer-implemented method of Claim 27, wherein the excess p-value identifies a statistically confounding-free likelihood that the rare variant increases gene expression in genes that otherwise have lower gene expression compared to other genes in a gene set.

31. 25. The computer-implemented method of claim 24, further comprising determining the causal relationship between the rare variant and the under-expressed gene while controlling for the plurality of confounding factors.

32. 31. The computer-implemented method of Claim 30, further comprising generating an under-causation score for the rare variant, wherein the particular under-causation score for the particular rare variant indicates the likelihood that the particular rare variant causes under-gene expression in an outlier individual whose gene sequence includes the particular rare variant.

33. 32. The computer-implemented method of claim 31, wherein the under-causation score is an under-causation probability value (under-causation p-value).

34. 33. The computer-implemented method of claim 32, wherein the undervalued p-value is determined by a Pearson correlation coefficient.

35. 33. The computer-implemented method of Claim 32, wherein the underrepresentation p-value identifies a statistically confounding-free likelihood that the rare variant reduces gene expression in genes that otherwise have higher gene expression relative to other genes in a gene set.

36. The computer-implemented method of claim 1 , wherein the rare variant is a non-coding variant.

37. 18. The computer-implemented method of claim 17, wherein the non-coding variants comprise 5 prime untranslated region (UTR) variants, 3 prime UTR variants, enhancer variants, and promoter variants.

38. The computer-implemented method of claim 1 , wherein the gene expression levels are further stratified into tissue-specific gene expression levels for a plurality of tissues.

39. 38. The computer-implemented method of claim 37, wherein the gene expression levels of each gene in each tissue are normalized using quantile normalization.

40. 39. The computer-implemented method of claim 38, wherein the causality model is fitted separately for each tissue.

41. The computer-implemented method of claim 1 , wherein the causal model is fitted using stratification.

42. The computer-implemented method of claim 1 , further comprising generating a ranking of the rare variants based on the causality scores.

43. 27. The computer-implemented method of claim 26, further comprising generating a ranking of the rare variants based on the excess causality scores.

44. 32. The computer-implemented method of claim 31, further comprising generating a ranking of the rare variants based on the under-causal scores.

45. The computer-implemented method of claim 1 , wherein the rare variant is a singleton variant.

46. 45. The computer-implemented method of claim 44, wherein a singleton variant occurs in only one of the outlier individuals.

47. 1. A system comprising one or more processors coupled to a memory, the memory loaded with computer instructions for identifying rare variants that cause extreme levels of gene expression, the instructions, when executed on the processor, accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, wherein the extreme levels of gene expression are determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; fitting a causal model to determine a causal relationship between the rare variant and the extreme level of gene expression in the outlier individual while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein the particular causality score for a particular rare variant indicates a likelihood that the particular rare variant causes extreme levels of gene expression in an outlier individual whose gene sequence includes the particular rare variant.

48. The system of claim 46, further performing actions including carrying out claims 2-45.

49. 1. A non-transitory computer-readable storage medium having stored thereon computer program instructions for identifying rare variants that cause extreme levels of gene expression, the instructions, when executed on a processor, comprising: accessing gene expression levels for a group of individuals; normalizing the gene expression levels and identifying outlier individuals from the group of individuals having extreme levels of gene expression, wherein the extreme levels of gene expression are determined from tail quantiles of the normalized gene expression levels; selecting rare variants from the gene sequences of the outlier individuals, wherein the rare variants are selected based on an allele frequency cutoff; fitting a causal model to determine a causal relationship between the rare variant and the extreme level of gene expression in the outlier individual while controlling for multiple confounding factors; generating a causality score for the rare variant based on the determined causality, wherein a particular causality score for a particular rare variant indicates a likelihood that the particular rare variant causes an extreme level of gene expression in an outlier individual whose gene sequence includes the particular rare variant.

50. 47. The non-transitory computer-readable storage medium of claim 46, wherein implementing the method further comprises carrying out claims 2-45.