Artificial intelligence-based detection of gene and expression conservation at base resolution
An AI-based biological mass model using deep learning techniques addresses the challenge of accurately analyzing genomic data by predicting gene expression and pathogenicity at base resolution, enhancing the understanding of genetic variants' impact.
Patent Information
- Application Number
- JP2024557743
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-05
- Filing Date
- 2023-08-04
- Publication Date
- 2025-08-26
AI Technical Summary
Existing genomics analysis methods struggle to accurately represent and analyze large, complex genomic data, leading to reduced classification accuracy and difficulty in identifying pathogenic genetic variants and understanding their impact on gene expression.
An artificial intelligence-based biological mass model that integrates deep learning techniques, including convolutional neural networks and residual blocks, to predict gene expression and pathogenicity by analyzing genomic sequences at base resolution, incorporating both evolutionary and epigenetic features.
Enhances the accuracy of predicting gene expression and pathogenicity by effectively capturing complex patterns in genomic data, improving the identification of pathogenic variants and their effects on gene function.
Smart Images

Figure 2025527971000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The disclosed technology relates to artificial intelligence-based computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to artificial intelligence-based detection of gene conservation and expression conservation at base resolution.
[0002] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is related to a concurrently filed U.S. patent application entitled "ARTIFICIAL INTELLIGENCE-BASED EPIGENETICS AT BASE RESOLUTION" (Attorney Docket No. ILLM 1040-1 / IP-2045-PRV), which is incorporated by reference for all purposes as if fully set forth herein.
[0003] (built-in) The following are incorporated by reference for all purposes as if fully set forth herein:
[0004] U.S. Patent Application No. 62 / 903,700, entitled "ARTIFICIAL INTELLIGENCE-BASED EPIGENETICS," filed September 20, 2019 (Attorney Docket No. ILLM 1025-1 / IP-1898-PRV).
[0005] Sundaram, L. et al. Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018).
[0006] Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019).
[0007] U.S. Patent Application No. 62 / 573,144, entitled "TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV).
[0008] U.S. Patent Application No. 62 / 573,149, entitled "PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV).
[0009] U.S. Patent Application No. 62 / 573,153, entitled "DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV).
[0010] U.S. Patent Application No. 62 / 582,898, entitled "PATHOGENICITY CLASSIFICATION OF GEOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV).
[0011] U.S. Patent Application No. 16 / 160,903, entitled "DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM1000-5 / IP-1611-US).
[0012] U.S. Patent Application No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US).
[0013] U.S. Patent Application No. 16 / 160,968, entitled "SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US).
[0014] U.S. Patent Application No. 16 / 407,149, filed May 8, 2019, entitled "DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS" (Attorney Docket No. ILLM 1010-1 / IP-1734-US).
[0015] U.S. Patent Application No. 17 / 232,056, filed April 15, 2021 (Attorney Docket No. ILLM 1037-2 / IP-2051-US), entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS TO PREDICT VARIANT PATHOGENICITY USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURES."
[0016] U.S. Patent Application No. 63 / 175,495, filed April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV), entitled "MULTI-CHANNEL PROTEIN VOXELIZATION TO PREDICT VARIANT PATHOGENICITY USING DEEP CONVOLUTIONAL NEURAL NETWORKS."
[0017] U.S. Patent Application No. 63 / 175,767, entitled "EFFICIENT VOXELIZATION FOR DEEP LEARNING," filed April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV).
[0018] U.S. Patent Application No. 17 / 468,411, filed September 7, 2021 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US), entitled "ARTIFICIAL INTELLIGENCE-BASED ANALYSIS OF PROTEIN THREE-DIMENSIONAL (3D) STRUCTURES." [Background technology]
[0019] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section or problems associated with the subject matter provided as background have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which, as such, may also correspond to embodiments of the claimed technology.
[0020] Genomics, broadly defined, also known as functional genomics, aims to characterize the function of all genomic elements in an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science and operates not by testing preconceived models and hypotheses, but by discovering novel properties from the examination of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene function, and mapping biochemically active genomic regions and residues, such as transcriptional enhancers and single nucleotide polymorphisms (SNPs).
[0021] Genomics data are too large and complex to be mined solely by visual inspection of pairwise correlations. For example, protein sequences can be grouped into families of homologous proteins that are derived from ancestral proteins and share similar structures and functions. Analysis of multiple sequence alignments (MSA) of homologous proteins provides important information about functional and structural constraints. Statistics of MSA columns, representing amino acid positions, identify functional residues that are conserved during evolution. Correlations of amino acid usage between MSA columns contain important information about functional sectors and structural contacts.
[0022] Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some algorithms in which assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically discover patterns in data. Therefore, machine learning algorithms are well-suited for data-driven science, particularly genomics. However, the performance of machine learning algorithms can strongly depend on how the data is represented, i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from fluorescence microscopy images, a preprocessing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.
[0023] A machine learning model can take estimated cell counts, which are an example of hand-designed features, as input features for classifying tumors. A central problem is that classification performance depends heavily on the quality and relevance of these features. For example, relevant visual features such as cell morphology, distance between cells, or localization within an organ are not captured in cell counts, and this incomplete representation of the data can reduce classification accuracy.
[0024] Deep learning, a subdiscipline of machine learning, addresses this problem by embedding feature computation within the machine learning model itself, generating an end-to-end model. This result has been achieved through the development of deep neural networks, machine learning models that involve successive primitive operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of high complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in algorithms, and substantial increases in computing power, particularly through the use of graphical processing units (GPUs).
[0025] The goal of supervised learning is to obtain a model that takes features as input and returns a prediction of a so-called target variable. An example of a supervised learning problem is predicting whether an intron will be spliced (target) or not, given RNA features such as the presence or absence of canonical splice site sequences, the location of splicing branch points, or the length of the intron. Training a machine learning model refers to learning its parameters, which generally involves minimizing a loss function on the training data with the goal of making accurate predictions on unknown data.
[0026] For many supervised learning problems in computational biology, input data can be represented as a table with multiple columns, or features, each of which contains numerical or categorical data potentially useful for making predictions. While some input data are naturally represented as tabular features (e.g., temperature or time), other input data must first be transformed using a process called feature extraction to fit the tabular representation (e.g., converting deoxyribonucleic acid (DNA) sequences to k-mer counts). For intron-splicing prediction problems, the presence or absence of canonical splice site sequences, the location of splicing branch points, and intron length can be preprocessed features collected in tabular format. Tabular data is the standard for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible nonlinear models such as neural networks and many others.
[0027] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of a positive class by calculating a weighted sum of input features mapped to the [0, 1] interval using a sigmoid activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when the weighted sum of input features cannot sufficiently distinguish between classes, such as whether an intron is spliced out or not. To improve prediction performance, new input features can be manually added by transforming or combining existing features in new ways, such as by taking exponentiation or pairwise products.
[0028] Neural networks automatically learn these nonlinear feature transformations using hidden layers. Each hidden layer can be thought of as multiple linear models with their output transformed by a nonlinear activation function, such as a sigmoid function or the more general rectified-linear unit (ReLU). Together, these layers organize the input features into related complex patterns, easing the task of distinguishing between two classes.
[0029] Deep neural networks use many hidden layers, and when each neuron receives input from all neurons in the previous layer, the layer is said to be fully connected. Neural networks are typically trained using stochastic gradient descent, an algorithm suitable for training models on very large datasets. Implementation of neural networks using modern deep learning frameworks allows for rapid prototyping with different architectures and datasets. Fully connected neural networks can be used in several genomics applications, including predicting the proportion of exons spliced into a given sequence from sequence features such as the presence of splice factor binding motifs or sequence conservation, prioritizing potentially disease-causing genetic variants, and predicting cis-regulatory elements in a given genomic region using features such as chromatin marks, gene expression, and evolutionary conservation.
[0030] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling DNA sequences or image pixels severely disrupts information patterns. These local dependencies set spatial or longitudinal data apart from tabular data, where feature ordering is arbitrary. Consider the problem of classifying genomic regions as bound versus unbound by a specific transcription factor, where binding regions are defined as high-confidence binding events in chromatin immunoprecipitation followed by sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. Fully connected layers based on sequence-derived features, such as the number of k-mer instances in a sequence or position weight matrix (PWM) matches, can be used for this task. Because k-mer or PWM instance frequencies are robust to shifting motifs within a sequence, such models can generalize well to sequences with the same motif located at different positions. However, they cannot recognize patterns where transcription factor binding depends on the combination of multiple motifs with distinct intervals. Furthermore, the number of possible k-mers increases exponentially with k-mer length, which poses both conservation and overfitting challenges.
[0031] A convolutional layer is a special form of a fully connected layer in which the same fully connected layer is applied locally to every sequence position, for example, within a 6-bp window. This approach can also be viewed as scanning a sequence using multiple PWMs, for example, for the transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced, and the network can detect motifs at positions not seen during training. Each convolutional layer scans the sequence with several filters by generating a scalar value at every position that quantizes the match between the filter and the sequence. As in a fully connected neural network, a nonlinear activation function (typically ReLU) is applied in each layer. Next, a pooling operation is applied, which aggregates activations within successive bins across the position axis, typically taking the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers can then construct the output of the previous layer and detect whether the GATA1 and TAL1 motifs were present within a certain distance range. Finally, the output of the convolutional layer can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.
[0032] Convolutional neural networks (CNNs) can predict various molecular phenotypes based solely on DNA sequence. Applications include classification of transcription factor binding sites and prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequence, CNNs can be applied to more technical tasks traditionally addressed by hand-designed bioinformatics pipelines. For example, CNNs can predict guide RNA specificity, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequence, and call genetic variants. CNNs have also been used to model long-range dependencies in genomes. Although interacting regulatory elements may be located far apart on unfolded linear DNA sequences, these elements are often proximal in actual 3D chromatin structures. Thus, modeling molecular phenotypes from linear DNA sequences can be improved by allowing long-range dependencies, albeit with a coarse approximation of chromatin, allowing the model to implicitly learn aspects of 3D organization such as promoter-enhancer loops. This is achieved by using dilated convolutions with receptive fields up to 32 kb. Dilated convolutions also allow splice sites to be predicted from sequences using receptive fields as small as 10 kb, thereby enabling integration of gene sequences over distances as long as a typical human intron (see Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).
[0033] Different types of neural networks can be characterized by their parameter-sharing schemes. For example, fully connected layers have no parameter sharing, while convolutional layers impose translational invariance by applying the same filter at every position in their input. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data, such as DNA sequences or time series, that implement a different parameter-sharing scheme. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. It updates the memory and optionally emits an output, which is either passed to subsequent layers or used directly as a model prediction. By applying the same model to each sequence element, recurrent neural networks are invariant to positional indices in the processed sequence. For example, recurrent neural networks can detect open reading frames in DNA sequences, regardless of their position in the sequence. This task requires the recognition of a specific sequence of inputs, such as a start codon followed by an in-frame stop codon.
[0034] The main advantage of recurrent neural networks over convolutional neural networks is that, theoretically, they can inherit information through infinitely long sequences via memory. Furthermore, recurrent neural networks can naturally process sequences of widely varying lengths, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can achieve performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the outputs of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, because recurrent neural networks apply sequential operations, they cannot be easily parallelized and are therefore much slower to compute than convolutional neural networks.
[0035] Although most of the human genetic code is common to all humans, each human has a unique genetic code. In some cases, the human genetic code may contain outliers, called genetic variants, that may be common among a relatively small group of individuals in the human population. For example, a particular human protein may contain a specific sequence of amino acids, but variants of that protein may differ by one amino acid in an otherwise identical specific sequence.
[0036] Genetic variants can be pathogenic and can result in disease. Although most such genetic variants have been eliminated from the genome by natural selection, the ability to identify which genetic variants are likely pathogenic can help researchers focus on these genetic variants to gain an understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human genetic variants remains unclear. Some of the most frequent pathogenic variants are single-nucleotide missense mutations that change the amino acid of a protein. However, not all missense mutations are pathogenic.
[0037] Models that can directly predict molecular phenotypes from biological sequences can be used as in silico perturbation tools to investigate the association between genetic and phenotypic variation and have emerged as new methods for quantitative trait locus identification and variant prioritization. These approaches are crucial given that the majority of variants identified by genome-wide association studies of complex phenotypes are non-coding, making it difficult to estimate their effect and contribution to the phenotype. Furthermore, linkage disequilibrium results in blocks of co-inherited variants, which makes it difficult to accurately identify individual causal variants. Therefore, sequence-based deep learning models that can be used as matching tools to assess the impact of such variants offer a promising approach for discovering potential drivers of complex phenotypes. One example is predicting the effects of non-coding single-nucleotide variants and short insertions or deletions (indels) indirectly from the differences between two variants on transcription factor binding, chromatin accessibility, or gene expression prediction. Another example is predicting novel splice site generation from the quantitative effects of genetic variants on sequence or splicing.
[0038] To predict the pathogenicity of missense variants from protein sequence and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (Sundaram, L. et al. Predicting the clinical impact of human mutations with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), herein referred to as "PrimateAI"). PrimateAI uses a deep neural network trained on known pathogenic variants with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences to compare differences and determine the pathogenicity of the variant using a trained deep neural network. Such an approach using protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to prior knowledge. However, the amount of clinical data available in ClinVar is relatively small compared to the amount of data sufficient to effectively train a deep neural network. To overcome this data shortage, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.
[0039] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in replicating prior knowledge in ClinVar. These results suggest that PrimateAI represents an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.
[0040] Central to protein biology is understanding how structural elements give rise to observed functions. The plethora of protein structural data enables the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods critically depends on the choice of protein structural representation.
[0041] Protein sites are microenvironments within a protein structure that are distinguished by their structural or functional role. Sites can be defined by their three-dimensional (3D) location and the local neighborhood around this location where the structure or function resides. Central to rational protein engineering is understanding how the structural arrangement of amino acids creates functional features within a protein site. Determining the structural and functional roles of individual amino acids in a protein provides information to aid in the manipulation and alteration of protein function. Identifying functionally or structurally important amino acids enables focused engineering efforts, such as site-directed mutagenesis, to alter the functional properties of a target protein. Alternatively, this knowledge can help avoid engineering designs that abolish desired functions.
[0042] Because it is well established that structure is much more conserved than sequence, the increase in protein structural data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using data-driven approaches. A fundamental aspect of any computational protein analysis is how protein structural information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, whereas a poor representation produces a noisy distribution lacking the underlying pattern.
[0043] The plethora of protein structures and the recent success of deep learning algorithms provide an opportunity to develop tools for automatically extracting task-specific representations of protein structures.
[0044] Computational analysis of genomics studies is challenged by confounding variations unrelated to the genetic factors of interest. Identifying variants that cause extreme levels of gene expression, either high or low, is crucial for diagnosing the pathogenicity of genetic diseases. However, numerous confounding factors exist that can interfere with the identification of pathogenic variants. Isolating variants by examining rare variants that may be associated with specific disease states can simplify the problem. Furthermore, removing the noise introduced by confounding factors can increase the signal-to-noise ratio.
[0045] Therefore, an opportunity arises to apply artificial intelligence to epigenetics to obtain genetic information between variable loci and the expression levels of individual genes.
[0046] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. Color drawings may also be available in PAIR via the Supplemental Content tab.
[0047] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Brief explanation of the drawings]
[0048] [Figure 1] FIG. 1 is a flow diagram showing the process of a system for determining evolutionary and epigenetic features of gene sequences. [Figure 2] 1 shows a schematic representation of an exemplary input sequence comprising nucleotide bases extracted from a sequence database, where the target sequence is flanked by a left sequence containing upstream context bases and a right sequence containing downstream context bases. [Figure 3] 1 shows an example of an alternative sequence from an exemplary reference gene sequence with two exemplary alternative sequences, each with a single nucleotide variant at a single base position, but otherwise with the same composition as the reference sequence. [Figure 4] 1 shows the genetic composition of sequences belonging to a training dataset developed for one embodiment of the disclosed technology. [Figure 5] 4A and 4B illustrate schematically one embodiment of a training procedure applied to the system of FIG. 1, in which a model is trained using a first training set as described in FIG. 4, and then retrained using a second training set as described in FIG. 4. [Figure 6] 1 is a schematic diagram of another embodiment of a training procedure applied to the system of FIG. 1 using a first training set described in FIG. 4, then undergoing retraining using a subset of the second training set described in FIG. 4, followed by model validation using the remaining subset of samples from the second training set. [Figure 7] FIG. 2 is a schematic diagram of an embodiment of the system of FIG. 1 for variant classification, where the system is used to compare a reference sequence and alternative sequences at base resolution by comparing the respective model outputs for each sequence. [Figure 8] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which comparison of a reference sequence to an alternative sequence is quantified by an average delta value. [Figure 9] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which the comparison of a reference sequence to an alternative sequence is quantified by a final total delta value. [Figure 10] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which a biomass model generates two biomass output sequences from an input base sequence via a first set of weights and a second set of weights trained end-to-end. [Figure 11] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which a biomass model generates three biomass output sequences from an input base sequence via a first set of weights and a second set of weights trained end-to-end. [Figure 12] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model includes a first set of weights that generate alternative representations of an input base sequence, and a second set of weights that are trained from scratch to generate multiple biomass output sequences from the alternative representations of the input base sequence. [Figure 13] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model generates one biomass output sequence from an input base sequence via a first set of weights and a second set of weights that are trained end-to-end and retrained to generate subsequent biomass output sequences on a single basis. [Figure 14] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model includes a first set of weights and a second set of weights that are trained end-to-end to generate a plurality of biomass output sequences from an input base sequence, and a third and fourth set of weights that are trained end-to-end to generate a gene expression output sequence from the plurality of input biomass output sequences. [Figure 15] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model includes a first set of weights that generate alternative representations of an input base sequence, a second set of weights that is trained from scratch to generate multiple biomass output sequences from the alternative representations of the input base sequence, and third and fourth sets of weights that are trained end-to-end to generate gene expression output sequences from the multiple input biomass output sequences. [Figure 16] 1 is a flow diagram of one embodiment of the disclosed technology, in which the biological mass model includes a first set of weights that generate alternative representations of an input base sequence, a second set of weights that are trained from scratch to generate a plurality of biological mass output sequences, a third set of weights that are trained from scratch to generate alternative biological mass representations from the plurality of biological mass output sequences, and a fourth set of weights that are trained from scratch to generate gene expression output sequences from the alternative biological mass representations. [Figure 17]1 is a flow diagram of one embodiment of the disclosed technology, wherein the biological mass model includes a first set of weights and a second set of weights that are trained end-to-end to generate a plurality of biological mass output sequences from an input base sequence, a third set of weights that is trained from scratch to generate alternative biological mass representations from the plurality of biological mass output sequences, and a fourth set of weights that is trained from scratch to generate gene expression output sequences from the alternative biological mass representations. [Figure 18] 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model is trained to generate a plurality of biomass output sequences from an input base sequence, and then a first set of weights is retrained as a substitute for a third set of weights to generate alternative biomass representations from the plurality of biomass output sequences, and a fourth set of weights is trained end-to-end using the substituted first set of weights to generate gene expression output sequences from the alternative biomass representations. [Figure 19] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which a biomass model includes a first set of weights that are trained to generate alternative sequence representations from input base sequences and then retrained as substitutes for a third set of weights to generate alternative biomass representations from a plurality of biomass output sequences, and a fourth set of weights that are trained from scratch to generate gene expression output sequences from the alternative biomass representations. [Figure 20] 1 is a flow diagram of one embodiment of the disclosed technology, wherein the biomass model includes a first set of weights trained to generate alternative sequence representations from an input base sequence; a second set of weights trained to generate a plurality of biomass output sequences from the alternative representations of the input bases, wherein the retrained first set of weights is used as a substitute for the third set of weights to generate the alternative biomass representations from the plurality of biomass output sequences; and a fourth set of weights trained from scratch to generate gene expression output sequences from the alternative biomass representations. [Figure 21]1 is a flow diagram of one embodiment of the disclosed technology, wherein the biomass model includes a first set of weights that are trained to generate alternative sequence representations from an input base sequence; a second set of weights that are trained to generate a plurality of biomass output sequences from the alternative representations of the input bases, wherein the retrained first set of weights are used as substitutes for the third set of weights to generate the alternative biomass representations from the plurality of biomass output sequences; and a fourth set of weights that are trained end-to-end using the replaced first set of weights to generate gene expression output sequences from the alternative biomass representations. [Figure 22] FIG. 10 is a flow diagram of one embodiment of the disclosed technology, wherein the model is further configured to include pathogenicity prediction logic that compares the reference biological quantity output sequence and the alternative biological quantity output sequence at base resolution, a fifth set of weights that generates alternative sequence pathogenicity predictions from the plurality of biological quantity output sequences, and the first and second sets of weights are each trained from scratch. [Figure 23] FIG. 10 is a flow diagram of one embodiment of the disclosed technology, wherein the model is further configured to include pathogenicity prediction logic that compares the reference biological quantity output sequence and the alternative biological quantity output sequence at base resolution, a fifth set of weights generates alternative sequence pathogenicity predictions from the plurality of biological quantity output sequences, and the first and second sets of weights are trained end-to-end. [Figure 24] FIG. 1 is a schematic diagram of a measure of evolutionary conservation that can be generated from a biomass model from an input base sequence as a value of a first biomass output sequence. [Figure 25] FIG. 10 is a schematic diagram of a measure of transcription initiation that can be generated from a biomass model from an input base sequence as a value of a second biomass output sequence. [Figure 26] FIG. 10 is a schematic diagram of an epigenetic signal that can be generated from a biomass model from an input base sequence as a third biomass output sequence value. [Figure 27] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which an expression change classifier is configured to predict the effect of a variant on gene expression. [Figure 28] FIG. 10 is a flow diagram of one embodiment of the disclosed technology, wherein the change expression classifier is further configured as a down-expression classifier to predict whether a variant reduces gene expression or does not reduce gene expression. [Figure 29] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, wherein the change expression classifier is further configured as a down-expression classifier to predict whether a variant increases gene expression or does not increase gene expression. [Figure 30] FIG. 1 is a flow diagram of one embodiment of the disclosed technology, in which the expression change classifier is further configured into a multi-class expression classifier that predicts whether a variant preserves gene expression, decreases gene expression, or increases gene expression. [Figure 31] FIG. 1 is a flow diagram of one embodiment of the disclosed technology in which gene expression classifier training is employed for comparison of ground truth causality scores with inferred causality scores. [Figure 32] 1 illustrates an exemplary computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE INVENTION
[0049] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0050] The detailed description of various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks are not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., modules, processors, or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or block of random access memory, hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine within an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentality shown in the figures.
[0051] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers, or servers, or may be spread across multiple different processors, computers, or servers. In addition, it will be understood that some of the modules may operate in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures may also be considered flowchart steps in a method. Also, a module need not necessarily have all its code located contiguously in memory. Some portions of code may be separated from other portions of code, with code from other modules or other functions located between them.
[0052] Exemplary Applications of the Disclosed Technology As highlighted by numerous comparisons of various embodiments of the biological mass model 124, many embodiments share overlap in architectural components. With respect to the phrase "biological mass model," the biological mass model predicts multiple classes of biological mass from genome sequences. Examples of biological mass classes include protein (transcription factor) binding, methylation, histone modification, DNA accessibility, and conservation. Of these, methylation and histone modification are counted as epigenetics. In contrast, chromatin refers to histone modification and DNA accessibility.
[0053] Thus, in some embodiments, the biological mass model may be referred to as an "epigenetics model." In other embodiments, the biological mass model may be referred to as a "chromatin model." In still other embodiments, the biological mass model may be referred to as a "chromatin and epigenetics model" or an "epigenetics and chromatin model."
[0054] Each element of the biological mass model 124 has multiple implementations that can be combined in numerous configurations. The multiple permutations that can be implemented for the disclosed technology provide both a broader range of utility, performance efficiency, and performance accuracy. The data transformations applied to input sequences in many embodiments of the disclosed technology to generate multiple additional sequence formats, both from the perspective of base sequence and chromatin structure, are innovative strategies that result in a wealth of output signals with broad applicability to a wide range of genomics, protein analysis, and pathogenicity research questions. Previous versions of PrimateAI have used multiple tools to classify variant pathogenicity with high performance. This biological mass model 124 introduces another tool in this methodology and an additional dimension for studying epigenetic signals that affect biological replication and transcription processes.
[0055] While there is clear utility in the added advantage of chromatin structure for studying gene expression to the various tools provided by PrimateAI, the true impact provided by the disclosed technology lies in the addition of epigenetic signals to the overall gene expression prediction logic. Both the DNA sequence and histone protein components of chromatin can undergo a plethora of chemical modifications. Enzymes that directly bind to and catalyze chemical modifications of chromatin components can alter chromatin structure, and changes in chromatin structure can also alter the ability of chromatin-interacting enzymes to access and function with their target ligands. Chromatin structure and enzymes that alter its structure directly affect the accessibility of genes for transcription and expression. DNA variants can cause changes in chromatin structure, which can subsequently alter epigenetic effects such as transcription factor binding and enzymatic reactions necessary for the proper regulation of gene expression and gene repression.
[0056] Conversely, epigenetic effects on chromatin, such as methylation and protein binding events, can affect mutation rates, potentially introducing silent or pathogenic variants. The study of evolutionary constraints on genes and the pathogenicity of their variants is significantly more comprehensive and accurate when augmented by epigenetic signatures, as demonstrated in many embodiments of the disclosed technology. Overall, the disclosed technology has several permutations that follow a series of training and learning strategies to generate several outputs that can be applied to predicting gene expression and gene pathogenicity for target gene sequences. The disclosed chromatin-focused strategies are useful in the study of genetic and environmental exposure-related diseases, drug development, and the influence of epigenetics on the transcription and translation of nucleic acid sequences into proteins.
[0057] Biological Mass Model Overview 1 is a flow diagram illustrating a system process 100 for determining evolutionary and epigenetic features of genetic sequences. An input sequence 122 is extracted from a sequence database 110 and processed by a biological mass model 124, which generates alternative representations of the input sequence 126. The alternative representations of the input sequence 126 are converted into alternative biological mass representations in the form of a plurality of biological mass output sequences 136.
[0058] Biological mass model input data 2 shows a schematic representation of an exemplary input base sequence 200 containing nucleotide bases extracted from a sequence database 202, where a target base sequence 226 is flanked by a left sequence 224 containing upstream context bases and a right sequence 228 containing downstream context bases. The upstream context bases 224 are a sequence of nucleotide bases {x1, x2, x3, ..., x n}, which may be equal to adenine, thymine, cytosine, or guanine. The target base sequence 226 then follows the upstream context bases 224. The target base sequence 226 is a sequence of nucleotide bases {y1, y2, y3, ..., y n}, which can be equal to adenine, thymine, cytosine, or guanine. Downstream context bases 228 then follow the target base sequence 226. The downstream context bases are a sequence of nucleotide bases {z1, z2, z3, ..., z n}, which may be equal to adenine, thymine, cytosine, or guanine. The input base sequence 200 may also include unknown or missing gap positions in the upstream context bases 224, the target base sequence 226, or the downstream context bases 228.
[0059] 3 shows an example of an alternative sequence 300 from an exemplary reference gene sequence 302 with two exemplary alternative sequences 322 and 342, each with a single nucleotide variant at a single base position, but otherwise with the same composition as the reference sequence. For example, the single nucleotide substitutions are shown as variants 326 and 336 compared to nucleotide 306. In addition to having identical target base sequences, upstream sequences 304, 324, and 344 are identical to each other, and downstream sequences 308, 328, and 348 are identical to each other.
[0060] 4 shows the genetic composition of sequences belonging to training datasets 400 developed for one embodiment of the disclosed technology. One training dataset 422 contains alternative sequences that are confounded by epigenetic effects (e.g., alternative sequence A 432 with single nucleotide variant 433), and a second training dataset 452 contains alternative sequences that are not confounded by epigenetic effects (e.g., sequence 462 with single nucleotide variant 463). Single nucleotide variants 433 and 463 differ in composition from reference base position 403. However, all other base positions within the reference sequence, alternative sequence A 432, and alternative sequence B 462, do not differ.
[0061] Biological mass model structure FIG. 5 illustrates a schematic representation of one embodiment 500 of a training procedure applied to the system 100 of FIG. 1, in which the biological mass model 124 is trained using a first training set as described in FIG. 4 and then retrained using a second training set as described in FIG. 4.
[0062] Sequences obtained from the sequence database 502 are first obtained from a first training dataset in a first training iteration set 566 to generate a plurality of biological mass output sequences 548 from input base sequences 524, and the biological mass model 124 is configured to detect changes in gene expression at base resolution and includes a biological mass model 124 that processes the input base sequences 524 and generates alternative representations (e.g., convoluted representations) of the input base sequences 524, and a biological mass output sequence generator 528. The biological mass model 124 is then subjected to a second training iteration set 586 on sequences obtained from a second training dataset without changing the model configuration of the biological mass model 124.
[0063] In one embodiment, the biological quantity model 124 includes groups of residual blocks arranged in sequence from lowest to highest. Each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block. In some embodiments, the dilated convolution rate progresses non-exponentially from lower to higher residual block groups. In other embodiments, it progresses exponentially. The size of the convolution window varies between groups of residual blocks, and each residual block includes at least one batch normalization layer, at least one rectified linear unit (ReLU) layer, at least one dilated convolution layer, and at least one residual connection.
[0064] In one embodiment, the dimensionality of the input is (Cu+L+Cd)×4, where Cu is the number of upstream adjacent context bases, Cd is the number of downstream adjacent context bases, and L is the number of bases in the input promoter sequence. The dimensionality of the output is 4×L. In some embodiments, each group of residual blocks produces an intermediate output by processing the previous input, and the dimensionality of the intermediate output is (I−[{(W−1) * D} *A] × N, where I is the dimensionality of the preceding input, W is the convolution window size of the residual block, D is the dilation convolution rate of the residual block, A is the number of dilated convolution layers in the group, and N is the number of convolution filters in the residual block.
[0065] In one embodiment, the input has 200 upstream adjacent context bases (Cu) to the left of the input sequence and 200 downstream adjacent context bases (Cd) to the right of the input sequence. The length (L) of the input sequence can be any length, such as 3001. In one embodiment, each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and one dilated convolution rate, and each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and four dilated convolution rates. In other architectures, each residual block has 32 convolution filters, 11 convolution window sizes, and one dilated convolution rate.
[0066] In one embodiment, the input has 1000 upstream adjacent context bases (Cu) to the left of the input sequence and 1000 downstream adjacent context bases (Cd) to the right of the input sequence. The length (L) of the input sequence can be any length, such as 3001. In one embodiment, there are at least three groups of four residual blocks and at least three skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate; each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates; and each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates.
[0067] In one embodiment, the input has 5,000 upstream adjacent context bases (Cu) to the left of the input sequence and 5,000 downstream adjacent context bases (Cd) to the right of the input sequence. The length (L) of the input sequence can be any length, such as 3,001. In one embodiment, there are at least four groups of four residual blocks and at least four skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate; each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates; each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates; and each residual block in the fourth group has 32 convolution filters, 41 convolution window sizes, and 25 dilated convolution rates.
[0068] Generally speaking, the biological quantity model 124 can be a rule-based model, a tree-based model, or a machine learning model, examples of which include multi-layer perceptrons (MLPs), feed-forward neural networks, fully connected neural networks, fully convolutional neural networks, sequence-to-sequence (Seq2Seq) models such as ResNet and WaveNet, semantic segmentation neural networks, and generative adversarial networks (GANs) (e.g., CycleGAN, StyleGAN, pixelRNN, text-2-image, DiscoGAN, IsGAN).
[0069] In some embodiments, the biological mass model 124 is a Transformer, Vision Transformer (ViT), Bidirectional Transformer (BERT), Detection Transformer (DETR), Deformable DETR, UP-DETR, DeiT, Swin, GPT, iGPT, GPT-2, GPT-3, BERT, SpanBERT, RoBERTa, XLNet, ELECTRA, UniLM, BART, T5, ERNIE (THU), KnowBERT, DeiT-Ti, DeiT-S, DeiT-B, T2T-ViT-14, T2T-ViT-19, T2T-ViT-24, PVT-Small, PVT-Medium, PV T-Large, TNT-S, TNT-B, CPVT-S, CPVT-S-GAP, CPVT-B, Swin-T, Swin-S, Swin-B, Twins-SVT-S, Twins-SVT-B, Twins- SVT-L, Shuffle-T, Shuffle-S, Shuffle-B, XCiT-S12 / 16, CMT-S, CMT-B, VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, MoCo v3, ACT, TSP, Max-DeepLab, VisTR, SETR, Hand-Transformer, HOT-Net, METRO, Image Transformer, Taming transformer, TransGAN, IPT, TTSR, STTN, Masked Transformer, CLIP, DALL-E, Cogview, UniT, ASH, TinyBert, FullyQT, ConvBert, FCOS, Faster R-CNN+FPN, DETR-DC5, TSP-FCOS, TSP-RCNN, ACT+MKDD(L=32), ACT+MKDD(L=16), SMCA, Efficient These may include self-attention mechanisms such as DETR, UP-DETR, UP-DETR, ViTB / 16-FRCNN, ViT-B / 16-FRCNN, PVT-Small+RetinaNet, Swin-T+RetinaNet, Swin-T+ATSS, PVT-Small+DETR, TNT-S+DETR, YOLOS-Ti, YOLOS-S, and YOLOS-B.
[0070] In some embodiments, examples of biological quantity model 124 include convolutional neural networks (CNNs) with multiple convolutional layers, recurrent neural networks (RNNs) such as long short-term memory networks (LSTMs), bidirectional LSTMs (Bi-LSTMs), or gated recurrent units, as well as combinations of both CNNs and RNNs.
[0071] In some implementations, the biological mass model 124 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, point-wise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The biological mass model 124 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. The biological mass model 124 can use any parallelism, efficiency, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharding, parallel calls for map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The biological mass model 124 can include nonlinear transformation functions, such as upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions (e.g., rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid, and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or mean pooling), global mean pooling layers, and attention mechanisms.
[0072] In some embodiments, the biological mass model 124 may be a linear regression model, a logistic regression model, an elastic net model, a support vector machine (SVM), a random forest (RF), a decision tree, or a boosted decision tree (e.g., XGBoost), or some other tree-based logic (e.g., metric tree, kd-tree, R-tree, universal B-tree, X-tree, ball tree, locality-sensitive hashing, and inverted index). The biological mass model 124 may, in some embodiments, be an ensemble of multiple models.
[0073] In some implementations, the biological mass model 124 can be trained using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that can be used to train the model include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the model include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.
[0074] 6 schematically illustrates another embodiment 600 of a training procedure applied to the system 100 of FIG. 1 using the first training set described in FIG. 4 , then undergoing retraining using a subset of the second training set described in FIG. 4 , followed by model validation of the biomass model 124 using the remaining subset of samples from the second training set. Sequences obtained from the sequence database 602 are first obtained from a first training dataset in a first training iteration set 666, generating a plurality of biomass output sequences 648 from input base sequences 624, the biomass model 124 being configured to detect changes in gene expression at base resolution, processing the input base sequences 624 and generating alternative representations (e.g., convoluted representations) of the input base sequences 624, and a biomass output sequence generator 628. The biomass model 124 then undergoes a second training iteration set 667 on a subset of sequences obtained from the second training dataset without changing the model configuration of the biomass model 124. Following training iteration 666 and training iteration 667, the biological mass model 124 undergoes a validation process 686 using a second subset of the second training dataset, where the first and second subsets of training dataset 2 do not overlap 668, which allows the biological mass model 124 to be trained with more diverse and unknown training examples.
[0075] FIG. 7 is a schematic diagram of an implementation 700 of the system 100 from FIG. 1 for variant classification, where the system is used to compare a reference sequence 702 and alternative sequences 704 at base resolution by comparing the respective model outputs of the biological mass models 124 for each sequence, represented by 783. The reference sequence 702 is processed separately by the biological mass models 124. An alternative representation generator 722 processes the reference input base sequence 702 to generate alternative representations (e.g., convoluted representation sequences), and a biological mass output sequence generator 742 processes the alternative representations to generate multiple biological mass output sequences 762. The alternative sequence 704 is processed separately by the biological mass models 124. The alternative representation generator 724 processes the reference input base sequence 704 to generate alternative representations (e.g., convoluted representation sequences), and a biological mass output sequence generator 744 processes the alternative representations to generate multiple biological mass output sequences 764. As demonstrated in the process depicted in completing flow diagram 783, the plurality of biological quantity output arrays 762 and the plurality of biological quantity output arrays 764 are compared at base resolution.
[0076] 8 is a flow diagram of one embodiment 800 of the disclosed technology in which the comparison between the reference sequence and the alternative sequence is quantified by a final average delta value 833. To determine the final average delta value 833, the base resolution pathogenicity classification logic 826 is configured to process the differences between multiple biological mass output sequences 810 predicted for the reference sequence and the alternative sequence. The final average delta value 833 is taken as the average from a first cumulative average delta value 822 comparing the reference sequence and the alternative sequence and a second average delta value 824 comparing the reference sequence and the alternative sequence. A first delta sequence 812 is generated as the base-by-base difference between the first reference biological mass output sequence 802 predicted from the reference sequence and the first alternative biological mass output sequence 804 predicted from the alternative sequence. A second delta sequence 814 is generated as the base-by-base difference between the second reference biological mass output sequence 806 and the second alternative biological mass output sequence 808. A first cumulative average delta value 822 is taken as the average of the delta values obtained for each base position in the first delta sequence 812. A second cumulative average delta value 824 is taken as the average of the delta values obtained for each base position in the second delta sequence 814. Based on the final average delta value 833, the variants represented by the alternative sequences can be classified into a conserved state 846, and the variants can be classified as belonging to a conserved state 842 or a non-conserved state 844.
[0077] 9 is a flow diagram of one embodiment 900 of the disclosed technology in which the comparison between the reference sequence and the alternative sequence is quantified by a final sum delta value 933. To determine the final sum delta value 933, the base-resolution pathogenicity classification logic 926 is configured to process the differences between multiple biological mass output sequences 910 predicted for the reference sequence and the alternative sequence. The final sum delta value 933 is obtained as the sum of a first accumulated sum delta value 922 comparing the reference sequence and the alternative sequence and a second sum delta value 924 comparing the reference sequence and the alternative sequence. A first delta sequence 912 is generated as the base-by-base difference between the first reference biological mass output sequence 902 predicted from the reference sequence and the first alternative biological mass output sequence 904 predicted from the alternative sequence. A second delta sequence 914 is generated as the base-by-base difference between the second reference biological mass output sequence 906 and the second alternative biological mass output sequence 909. A first cumulative total delta value 922 is taken as the sum of the delta values obtained for each base position in the first delta sequence 912. A second cumulative total delta value 924 is taken as the sum of the delta values obtained for each base position in the second delta sequence 914. Based on the final total delta value 933, the variants represented by the alternative sequences can be classified into a conserved state 946, and the variants can be classified as belonging to a conserved state 942 or a non-conserved state 944.
[0078] 10 is a flow diagram of one embodiment 1000 of the disclosed technology, in which a biological mass model 124 generates two biological mass output sequences 1084 from an input base sequence 1022 via a first set of weights 1042 and a second set of weights 1062 that are trained end-to-end. The first set of weights 1042 includes an alternative representation generator 1044 that processes the input base sequence 1022 and generates alternative representations of the input base sequence 1022. The second set of weights 1062 includes a biological mass output sequence generator 1064 that processes the alternative representations of the input base sequence 1022 and generates multiple biological mass output sequences 1084. In the embodiment 1000 of the disclosed technology, two output sequences 1081 and 1083 are generated from the biological mass output sequence generator 1064.
[0079] 11 is a flow diagram of one embodiment 1100 of the disclosed technology, in which a biological mass model 124 generates three biological mass output sequences 1168 from an input base sequence 1102 via an end-to-end trained first set of weights 1122 and a second set of weights 1142. The first set of weights 1122 includes an alternate representation generator 1124 that processes the input base sequence 1102 and generates alternate representations of the input base sequence 1102. The second set of weights 1142 includes a biological mass output sequence generator 1144 that processes the alternate representations of the input base sequence 1102 and generates multiple biological mass output sequences 1168.
[0080] The embodiment 1100 of Figure 11 differs from the embodiment 1000 of Figure 10 in that three output sequences 1162, 1164, and 1166 are generated from the biological mass output sequence generator 1144, compared to two output sequences 1081 and 1083 generated from the biological mass output sequence generator 1064. In Figure 10, the first output sequence 1081 can be a base-by-base measure of evolutionary conservation. The second output sequence 1083 can be a base-by-base measure of transcription start. In Figure 11, the first output sequence 1162 can be a base-by-base measure of evolutionary conservation. The second output measure 1164 can be a base-by-base measure of transcription start.
[0081] Those skilled in the art will appreciate that multiple output sequences can be predicted simultaneously and can represent different permutations and combinations of information such as measures of evolutionary conservation, measures of transcription initiation represented, and base-by-base epigenetic signals.
[0082] 12 is a flow diagram of an embodiment 1200 of the disclosed technology in which the biomass model 124 includes a first set of weights 1222 that generates alternative representations of the input base sequence 1206 and a second set of weights 1226 that are trained from scratch to generate multiple biomass output sequences 1246 from the alternative representations of the input base sequence 1242. Compared to the embodiment 1000 of FIG. 10 and the embodiment 1100 of FIG. 11, the first set of weights and the second set of weights are not trained end-to-end in the embodiment 1200. As is similarly done in the embodiment 1000 of FIG. 10 and the embodiment 1100 of FIG. 11, the first set of weights 1222 for the embodiment 1200 includes an alternative representation generator 1224 that processes the input base sequence 1202 to generate alternative representations of the input base sequence 1242. However, in contrast to embodiments in which weights are trained end-to-end, the alternative sequence representation output 1242 from the alternative representation generator 1224 is mapped to input 1206 for a subsequent model configured as a biological quantity output sequence generator 1228 that is used to train a second set of weights 1226 to generate multiple biological quantity output sequences 1246.
[0083] 13 is a flow diagram of one embodiment 1300 of the disclosed technology, in which a biological mass model 124 generates one biological mass output sequence 1362 from an input base sequence 1302 via a first set of weights 1322 and a second set of weights 1342 that are trained end-to-end and retrained to generate subsequent biological mass output sequences on a single basis. In the given example shown in FIG. 13 , the biological mass model 124 is trained with the first weight set 1322 and the second weight set 1342 to generate a first output sequence 1362 in a plurality of biological mass output sequences 1366. Following generation of the first output sequence, the biological mass model 124 is retrained end-to-end on the first weight set 1324 and the second weight set 1344 to process the input base sequence 1304 and generate a second output sequence 1364. In the retraining process, the input base sequence 1304 is the same as the input base sequence 1302, the first weight set 1324 is the same as the first weight set 1322, and the second weight set 1344 is the same as the second weight set 1362. However, the second output sequence 1364 is a different biological quantity output sequence from the first biological quantity output sequence. For all the multiple biological quantity output sequences 1366, the first and second weights are retrained end-to-end to generate each subsequent biological quantity output sequence.
[0084] 14 is a flow diagram of one embodiment 1400 of the disclosed technology, in which a biological quantity model 124 includes a first set of weights 1422 and a second set of weights 1442 that are trained end-to-end to generate a plurality of biological quantity output sequences 1406 from an input base sequence, and a third set of weights 1426 and a fourth set of weights 1446 that are trained end-to-end to generate a gene expression output sequence 1466 from a plurality of input biological quantity output sequences 1462. The generation of the biological quantity output sequence 1426 from the input base sequence 1402 is similar to the embodiment 1000 of FIG. 10 and the embodiment 1100 of FIG. 11, in which the input base sequence 1402 is processed by the first set of weights 1422 configured as an alternative representation generator 1404 and the second set of weights 1442 configured as a biological quantity output sequence generator 1444 to generate a plurality of biological quantity output sequences 1462. The plurality of output sequences 1462, as an output of the second set of weights 1442, are then mapped to inputs 1406 for a subsequent model configured to generate gene expression output sequences 1466. The input biological quantity output sequences 1406 are processed by a third set of weights 1426 and a fourth set of weights 1446, which are trained end-to-end. The third set of weights 1426 is configured as a biological quantity alternative representation generator 1428, and the fourth set of weights is configured as a gene expression output generator 1448. The resulting gene expression output sequences 1466 are a measure of base-resolution gene expression 1468.
[0085] 15 is a flow diagram of one embodiment 1500 of the disclosed technology in which the biological mass model 124 includes a first set of weights 1522 that generate alternative representations of the input base sequence 1562, a second set of weights 1582 that are trained from scratch to generate a plurality of biological mass output sequences 1512 from the alternative representations of the input base sequence 1542, and third and fourth sets of weights 1124 and 1546 that are trained end-to-end to generate a gene expression output sequence 1566 from a plurality of input biological mass output sequences 1506. The generation of the biological mass output sequence 1512 from the input base sequence 1562 is similar to the embodiment 1200 of FIG. 12 in which the biological mass model 124 includes a first set of weights 1522 that generate alternative representations of the input base sequence 1502, and a second set of weights 1582 that are trained from scratch to generate a plurality of biological mass output sequences 1512 from the alternative representations of the input base sequence 1542.
[0086] As in embodiments 1000 from Figure 10, 1100 from Figure 11, 1200 from Figure 12, 1300 from Figure 13, and 1400 from Figure 14, a first set of weights 1522 is configured as an alternative representation generator 1524, and a second set of weights 1582 is configured as a biological quantity output sequence generator 1584. A plurality of output sequences 1512 as an output of the second set of weights 1582 are then mapped to inputs 1506 for a subsequent model configured to generate gene expression output sequences 1566. The input biological quantity output sequences 1506 are processed by a third set of weights 1124 and a fourth set of weights 1546, which are trained end-to-end. The third set of weights 1124 is configured as a biological quantity alternative representation generator 1528, and the fourth set of weights is configured as a gene expression output generator 1548. The resulting gene expression output sequence 1566 is a measure of base-resolution gene expression 1568.
[0087] 16 is a flow diagram of an embodiment 1600 of the disclosed technology in which the biological quantity model 124 includes a first set of weights 1622 that generate alternative representations of the input base sequence 1642, a second set of weights 1682 that are trained from scratch to generate a plurality of biological quantity output sequences 1606, a third set of weights 1124 that are trained from scratch to generate alternative biological quantity representations 1646 from a plurality of biological quantity output sequences 1612, and a fourth set of weights 1686 that are trained from scratch to generate a gene expression output sequence 1616 from the alternative biological quantity representations 1666. The generation of the biological quantity output sequence 1612 from the input base sequence 1662 is similar to the embodiment 1600 of FIG. 16 , in which the biological quantity model 124 includes a first set of weights 1622 that generate alternative representations of the input base sequence 1602, and a second set of weights 1682 that are trained from scratch to generate a plurality of biological quantity output sequences 1612 from the alternative representations of the input base sequence 1642.
[0088] Similar to the embodiment 1000 of Figure 10 , the embodiment 1100 of Figure 11 , the embodiment 1200 of Figure 12 , the embodiment 1300 of Figure 13 , the embodiment 1400 of Figure 14 , and the embodiment 1500 of Figure 15 , the first set of weights 1622 is configured as an alternative representation generator 1624, and the second set of weights 1682 is configured as a biological quantity output array generator 1684. The plurality of biological quantity output arrays 1612 as the output of the second set of weights 1682 are then mapped to inputs 1606 for a subsequent model configured to generate alternative biological quantity representations 1646. The subsequent model is configured as a biological quantity alternative representation generator 1628 that includes a third set of weights 1124 that is trained from scratch to generate the alternative biological quantity representations 1646. The alternative biological quantity representation output 1646 as the output of the third set of weights 1124 is then mapped to the input 1666 for a second subsequent model to generate the gene expression output sequence 1616. The second subsequent model is configured as a gene expression output generator 1688 that includes a fourth set of weights 1686 that is trained from scratch to generate the gene expression output sequence 1616. The resulting gene expression output sequence 1616 is a measure of base-resolution gene expression 1618.
[0089] Figure 17 is a flow diagram of one embodiment 1700 of the disclosed technology, in which a biological mass model 124 includes a first set of weights 1722 and a second set of weights 1742 that can be trained end-to-end to generate multiple biological mass output sequences 1762 from an input base sequence 1702. A third set of weights 1726 can be trained from scratch to generate alternative biological mass representations 1746 from the multiple biological mass output sequences 1706. A fourth set of weights 1786 can be trained from scratch to generate gene expression output sequences 1716 from the alternative biological mass representations 1766. The gene expression output per given base in the gene expression output sequence 1716 for a given target base at a given position specifies a measure of the gene expression level of the given target base at the given position. In one embodiment, the gene expression level is measured by a per-base metric, such as CAGE transcription start site (CTSS). In another embodiment, gene expression levels are measured in a per-gene metric such as transcripts per million (TPM) or reads per kilobase of transcript (RPKM). In yet another embodiment, gene expression levels are measured in a per-gene metric such as fragments per kilobase million (FPKM).
[0090] The generation of biological quantity output sequences 1762 from the input base sequence 1702 is similar to the embodiment 1000 of FIG. 10 and the embodiment 1100 of FIG. 11, in which the input base sequence 1702 can be processed by a first set of weights 1722 applied by an alternative representation generator 1724 and a second set of weights 1742 applied by a biological quantity output sequence generator 1744 to generate multiple biological quantity output sequences 1762.
[0091] The plurality of biological quantity output arrays 1762 may then be mapped as an input 1706 for a subsequent model configured to generate an alternative biological quantity representation 1746 as an output of a second set of weights 1742. The subsequent model may be configured as a biological quantity alternative representation generator 1728 including a third set of weights 1726 trained from scratch to generate the alternative biological quantity representation 1746. The alternative biological quantity representation output 1746 as an output of the third set of weights 1726 may then be mapped to an input 1766 for a second subsequent model to generate a gene expression output array 1716. The second subsequent model may be configured as a gene expression output generator 1788 including a fourth set of weights 1786 trained from scratch to generate the gene expression output array 1716. The resulting gene expression output array 1716 is a measure of base-resolution gene expression 1718.
[0092] Applying transfer learning to biological quantity models 18 is a flow diagram of one embodiment 1800 of the disclosed technology, in which the biological mass model 124 includes a first set of weights 1812 that is trained to generate alternative sequence representations 1822 from an input base sequence 1802 and then retrained as alternatives for a third set of weights 1842 to generate alternative biological mass representations from a plurality of biological mass output sequences 1832, and a fourth set of weights 1852 that is trained end-to-end using the replaced first set of weights to generate gene expression output sequences 1862 from the biological mass output sequences 1832. The optimized weight scalar values of weight set 1 1812 learned from the alternative representation generator 1813 can be transferred as the alternative scalar values for each weight in the third set of weights 1842. Similar to embodiment 1400 of Figure 14 and embodiment 1500 of Figure 15, biological quantity surrogate representation generator 1843 includes a fourth set of weights 1852 that includes a gene expression output generator 1863 and a third set of weights 1842 that are trained end-to-end. The resulting gene expression output sequence 1862 is a measure of base-resolution gene expression 1863.
[0093] 19 is a flow diagram of one embodiment 1900 of the disclosed technology, in which the biological mass model 124 includes a first set of weights 1912 that is trained to generate an alternative sequence representation 1922 from an input base sequence 1902 and then retrained as a replacement for a third set of weights 1932 to generate an alternative biological mass representation 1942 from the alternative sequence representation 1922, and a fourth set of weights 1962 that is trained from scratch to generate a gene expression output sequence 1972 from the alternative biological mass representation 1952. The optimized weight scalar values for weight set 1 1912 learned from the alternative representation generator 1913 can be transferred as the alternative scalar values for each weight in the third set of weights 1932.
[0094] The alternative biological quantity representation output 1942 as the output of the third set of weights 1932 is then mapped to the input 1952 for the subsequent model to generate the gene expression output sequence 1716. Similar to the embodiment 1600 of Figure 16 and the embodiment 1700 of Figure 17, the subsequent model is configured as a gene expression output generator 1963 including a fourth set of weights 1962 that is trained from scratch to generate the gene expression output sequence 1972. The resulting gene expression output sequence 1972 is a measure of base-resolution gene expression 1973.
[0095] 20 is a flow diagram of one embodiment 2000 of the disclosed technology, in which the biological mass model 124 includes a first set of weights 2052 trained to generate an alternative sequence representation 2062 from an input base sequence 2042, a second set of weights 2004 trained to generate a plurality of biological mass output sequences 2014, a retrained first set of weights 2052 used as a substitute for the third set of weights 2034 to generate an alternative biological mass representation 2044 from the plurality of biological mass output sequences, and a fourth set of weights 2064 trained from scratch to generate a gene expression output sequence 2074 from the alternative biological mass representations 2054. Similar to the embodiment 1900 of FIG. 19 , the optimized weight scalar values of weight set 1 2052 learned from the alternative representation generator 2053 can be transferred as the alternative scalar values for each weight in the third set of weights 2034.
[0096] The second set of weights 2004 is configured to generate a plurality of biological quantity output arrays 2014. Similar to the embodiment 1400 of Figure 14 , the embodiment 1500 of Figure 15 , the embodiment 1600 of Figure 16 , and the embodiment 1700 of Figure 17 , the biological quantity output arrays 2014 are mapped to inputs of a subsequent model configured as a biological quantity surrogate representation generator 2035 including a third set of weights 2034. In embodiment 2000, the scalar values of the third set of weights 2034 are substituted from optimized scalar values from the trained first set of weights 2052. Similar to the embodiment 1600 of Figure 16 and the embodiment 1700 of Figure 17 , the surrogate biological quantity representation output 2044 as an output of the third set of weights 2034 is then mapped to inputs 2044 for the subsequent model to generate gene expression output arrays 2074. The subsequent model is configured as a gene expression output generator 2065 that includes a fourth set of weights 2064 that is trained from scratch to generate a gene expression output sequence 2074. The resulting gene expression output sequence 2074 is a measure of base-resolution gene expression 2075.
[0097] 21 is a flow diagram of one embodiment 2100 of the disclosed technology, in which the biological mass model 124 includes a first set of weights 2142 trained to generate alternative sequence representations 2152 from an input base sequence 2132, a second set of weights 2104 trained to generate a plurality of biological mass output sequences 2124 from the alternative representations of the input bases, a first set of retrained weights 2142 used as a substitute for the third set of weights 2134 to generate alternative biological mass representations from a plurality of biological mass output sequences 2114, and a fourth set of weights 2144 trained end-to-end using the first set of weights 2142 substituted to generate gene expression output sequences 2154 from the alternative biological mass representations. Similar to the embodiment 1900 of FIG. 19 and the embodiment 2000 of FIG. 20, the optimized weight scalar values of weight set 1 2142 learned from the alternative representation generator 2152 can be transferred as the alternative scalar values for each weight in the third set of weights 2134.
[0098] Similar to the embodiment 2000 of Figure 20, the second set of weights 2104 is configured to generate a plurality of biological quantity output arrays 2114. Similar to the embodiment 1400 of Figure 14, the embodiment 1500 of Figure 15, the embodiment 1600 of Figure 16, and the embodiment 1700 of Figure 17, the biological quantity output arrays 2114 are mapped to the input of a subsequent model configured as a biological quantity surrogate representation generator 2135 including a third set of weights 2134. Similar to the embodiment 2000 of Figure 20, the scalar values of the third set of weights 2134 are substituted from the optimized scalar values of the first set of trained weights 2142. Similar to the embodiment 1400 of Figure 14, the embodiment 1500 of Figure 15, and the embodiment 1800 of Figure 18, the biological quantity surrogate representation generator 2135 includes a fourth set of weights 2144 including a gene expression output generator 2145 and the third set of weights 2134 that are trained end-to-end. The resulting gene expression output sequence 2154 is a measure of base-resolution gene expression 2155 .
[0099] Pathogenicity prediction at base resolution 22 is a flow diagram of an embodiment 2200 of the disclosed technology, wherein the biological mass model 124 is further configured to include pathogenicity prediction logic that compares a reference biological mass output sequence 2253 and an alternative biological mass output sequence 2256 at base resolution, a fifth set of weights 2223 that generates an alternative sequence pathogenicity prediction 2233 from the plurality of biological mass output sequences 2253 and 2256, and the first set of weights 2212 and 2214 and the second set of weights 2242 and 2244 are each trained from scratch. Similar to embodiment 1200 of FIG. 12 , embodiment 1500 of FIG. 15 , and embodiment 1600 of FIG. 16 , the first set of weights 2212 of embodiment 2200 includes an alternative representation generator 2213 that processes a reference input base sequence 2202 to generate an alternative representation 2222 of the reference input base sequence 2202. The alternative sequence representation output 2222 from the alternative representation generator 2213 is mapped to input 2232 for a subsequent model configured as a biological quantity output sequence generator 2243 that is used to train a second set of weights 2242 to generate a plurality of biological quantity output sequences 2253.
[0100] In parallel, the first set of weights 2214 for embodiment 2200 includes an alternate representation generator 2216 that processes the alternate input base sequence 2204 to generate an alternate representation 2224 of the alternate input base sequence 2204. The alternate sequence representation output 2224 from the alternate representation generator 2216 is mapped to an input 2234 for a subsequent model configured as a biological mass output sequence generator 2246 that is used to train a second set of weights 2244 to generate a plurality of biological mass output sequences 2256. The optimized scalar values of the weights for each respective biological mass model 124 for the reference sequence 2202 and the alternate sequence 2204 can be transferred to a fifth set of weights 2223 that generates base resolution alternate sequence pathogenicity predictions 2233.
[0101] In one embodiment, the pathogenicity prediction can be a score between 0 and 1, where 0 represents absolute benign and 1 represents absolute pathogenic. In other embodiments, a cutoff can be used, for example, a pathogenicity score above 5 can be considered pathogenic and below 5 can be considered benign.
[0102] FIG. 23 is a flow diagram of one embodiment 2300 of the disclosed technology, wherein the biological quantity model 124 is further configured to include pathogenicity prediction logic that compares a reference biological quantity output sequence 2362 and an alternative biological quantity output sequence 2364 at base resolution, a fifth set of weights 2353 generates an alternative sequence pathogenicity prediction 2363 from the multiple biological quantity output sequences 2362 and 2364, and the first set of weights 2322 and 2324 and the second set of weights 2342 and 2344 are trained end-to-end. As in embodiments 1100 of Figure 11 , 1200 of Figure 12 , 1300 of Figure 13 , 1400 of Figure 14 , and 1700 of Figure 17 , a reference input base sequence 2302 is processed by a first set of weights 2322 configured as an alternative representation generator 2323 and a second set of weights 2342 configured as a biological quantity output sequence generator 2343 to generate a plurality of biological quantity output sequences 2362. The first set of weights 2322 and the second set of weights 2342 are trained end-to-end.
[0103] In parallel, a first set of weights 2324 processes the alternate input base sequence 2304 with an alternate representation generator 2326, and a second set of weights 2344 includes a biological mass output sequence generator 2346, to generate multiple biological mass output sequences 2364 in the same manner that multiple biological mass output sequences 2362 are generated from the reference input base sequence 2302. The optimized scalar values of the weights for each respective biological mass model 124 for the reference sequence 2302 and the alternate sequence 2304 can be transferred to a fifth set of weights 2353, which generates base resolution alternate sequence pathogenicity predictions 2363.
[0104] Biological abundance output sequence data 24 is a schematic diagram of a measure of evolutionary conservation 2400 that can be generated from the biomass model 124 from an input base sequence as a value of a first biomass output sequence. After the input base sequence is processed by an alternative representation generator 2401, a biomass output sequence generator 2402 can generate a first biomass output sequence in the form of a phyloP score 2403. After the input base sequence is processed by an alternative representation generator 2404, a biomass output sequence generator 2405 can generate a first biomass output sequence in the form of a phastCons value 2406. After the input base sequence is processed by an alternative representation generator 2407, a biomass output sequence generator 2408 can generate a first biomass output sequence in the form of a phastCons value 2409.
[0105] 25 is a schematic diagram of a measure 2500 of transcription initiation that can be generated from the biological mass model 124 from an input sequence as a value of a second biological mass output sequence. After the input sequence is processed by the alternative representation generator 2502, the biological mass output sequence generator 2504 can generate the second biological mass output sequence in the form of a cap analysis of gene expression (CAGE) value 2506.
[0106] 26 is a schematic diagram of an epigenetic signal 2600 that can be generated from the biological mass model 124 from an input base sequence as a value of a third biological mass output sequence. After the input base sequence is processed by an alternative representation generator 2601, a biological mass output sequence generator 2602 can generate a first biological mass output sequence in the form of a DNase I hypersensitive site predicted output 2603. After the input base sequence is processed by an alternative representation generator 2604, a biological mass output sequence generator 2605 can generate a first biological mass output sequence in the form of a transcription factor binding site predicted output 2606. After the input base sequence is processed by an alternative representation generator 2607, a biological mass output sequence generator 2608 can generate a first biological mass output sequence in the form of a histone modification mark predicted output 2609.
[0107] Gene Expression Classification Model 27 is a flow diagram of one embodiment of the disclosed technology in which an expression change classifier 2700 is configured to predict the effect of variants on gene expression. A variant expression conserving causality score validation dataset 2702 is used to generate a ground truth branch of a set of variants of a binary classification model 2704 with a specified decision threshold (e.g., a p-value cutoff less than 0.01, 0.0001, or 1e-14) that can be used to classify variants into a gene expression changing class 2724, i.e., variants that alter gene expression, or a gene expression conserving class 2744, i.e., variants that do not alter gene expression. The classification of variants is learned from assigned causality scores that specify a statistically unconfounded likelihood of altering gene expression.
[0108] 28 is a flow diagram of one embodiment of the disclosed technology, in which the expression change classifier is further configured as a down-expression classifier 2800 for predicting whether a variant reduces gene expression or does not reduce gene expression. The variant under-expression causality score validation set is used to generate an under-expression ground truth branch of the set of variants for a binary classification model 2804 with a specified decision threshold (e.g., a p-value cutoff less than 0.01, 0.0001, or 1e-14) that can be used to classify variants into a reduced gene expression class 2844, i.e., variants that reduce gene expression, or a non-reduced gene expression class 2824, i.e., variants that do not reduce gene expression. The classification of variants is learned from assigned under-expression causality scores that specify a statistically unconfounded likelihood of reducing gene expression.
[0109] 29 is a flow diagram of one embodiment of the disclosed technology in which the expression change classifier is further configured as an increased expression classifier 2900 to predict whether a variant increases gene expression or does not increase gene expression. The variant overexpression causality score validation dataset 2902 is used to generate an overexpression ground truth branch of the set of variants in a binary classification model 2904 with a specified decision threshold (e.g., a p-value cutoff less than 0.01, 0.0001, or 1e-14) that can be used to classify variants into a decreased gene expression class 2924 or a non-decreased gene expression class 2944. The classification of variants is learned from their assigned overexpression causality scores, which specify a statistically unconfounded likelihood of increasing gene expression.
[0110] FIG. 30 is a flow diagram of one embodiment of the disclosed technology, in which the expression change classifier is further configured into a multi-class expression classifier 3000 that predicts whether a variant preserves gene expression, decreases gene expression, or increases gene expression. The variant dataset is processed by system 3000, and a first performance measure is generated for the comparison of the inferred branch to the ground truth branch via a binary classifier 3026 having a decision threshold that generates an output corresponding to a causality score of changing gene expression 3024 into a class of changing gene expression 3028 or a class of maintaining gene expression 3038; a second performance measure is generated for the comparison of the inferred branch to the ground truth branch via a binary classifier 3046 having a decision threshold that generates an output corresponding to a causality score of decreasing gene expression 3044 into a class of decreasing gene expression 3048 or a class of not decreasing gene expression 3058; and a third performance measure is generated for the comparison of the inferred branch to the ground truth branch via a binary classifier 3066 having a decision threshold that generates an output corresponding to a causality score of increasing gene expression 3064 into a class of increasing gene expression 3068 or a class of not increasing gene expression 3078.
[0111] The system 3000 is configured to require the first, second, and third inferred branches to classify the same number of variants into genetic variation classes, thereby making the first, second, and third performance measures comparable to each other. The system 3000 is further configured to compare the performance of each of the first binary classifier on the validation data 3002 with a decision threshold 3026, a second binary classifier 3046, and a third binary classifier 3066 based on a comparison of the first, second, and third performance measures. The decision thresholds of the first binary classifier 3026, the second binary classifier 3046, and the third binary classifier 3066 with decision thresholds may be the same. The decision thresholds of the first binary classifier 3026, the second binary classifier 3046, and the third binary classifier 3066 with decision thresholds may be different.
[0112] The system 3000 is further configured to generate a ground truth trichotomy and an inferred trichotomy of the set of variants 3002 into a gene expression conserved class 3082, a gene expression decreased class 3084, and a gene expression increased class 3086 from a multi-class classifier 3080 that processes one-hot coded vectors including the ground truth and inferred bichotomies from the first binary classifier 3026, the second binary classifier 3046, and the third binary classifier 3066 using a decision threshold.
[0113] 31 is a flow diagram of one implementation of the disclosed technology in which gene expression classifier training 3100 is employed to compare ground truth causality scores with inferred causality scores 3144. A set of ground truth causality scores 3122 is generated for a variant training dataset 3102. A binary classifier with a decision threshold 3142 processes the variant training dataset 3102 to generate inferred causality scores classified into an inferred first class 3161 and an inferred second class 3163. The gene expression classifier training protocol 3100 performs backpropagation on the weights of the binary classifier 3142 for several iterations to optimize a loss function.
[0114] Computer Systems 32 illustrates an exemplary computer system 3200 that can be used to implement the disclosed techniques. The computer system 3200 includes at least one central processing unit (CPU) 3272 that communicates with a number of peripheral devices via a bus subsystem 3255. These peripheral devices may include, for example, a storage subsystem 3210, including memory devices and a file storage subsystem 3232, a user interface input device 3238, a user interface output device 3276, and a network interface subsystem 3274. The input and output devices enable user interaction with the computer system 3200. The network interface subsystem 3274 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0115] In one embodiment, the biological mass model 124 is communicatively linked to a storage subsystem 3210 and a user interface input device 3238 .
[0116] The user interface input devices 3238 can include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and ways of inputting information into the computer system 3200.
[0117] The user interface output devices 3276 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and ways for outputting information from the computer system 3200 to a user or to another machine or computer system.
[0118] The storage subsystem 3210 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 3278.
[0119] The processor 3278 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 3278 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 3278 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX32 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0120] The memory subsystem 3222 used in the storage subsystem 3210 may include several memories, including a main random access memory (RAM) 3232 for storing instructions and data during program execution, and a read only memory (ROM) 3234 in which fixed instructions are stored. The file storage subsystem 3232 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 3232 in the storage subsystem 3210 or in another machine accessible by the processor.
[0121] Bus subsystem 3255 provides a mechanism for allowing the various components and subsystems of computer system 3200 to communicate with each other as intended. Although bus subsystem 3255 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0122] The computer system 3200 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 3200 depicted in Figure 32 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 3200 can have more or fewer components than the computer system depicted in Figure 32.
[0123] Terms The disclosed technology can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeating list of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following embodiments.
[0124] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer-usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor, coupled to the memory, operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be realized in the form of a means for performing one or more of the method steps described herein, which means can include (i) hardware modules, (ii) software modules executing on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which (i)-(iii) implements particular technology described herein, and the software modules are stored on a computer-readable storage medium (or multiple such media).
[0125] The clauses described in this section can be combined as features. For purposes of brevity, combinations of features are not individually listed and are not repeated for each base set of features. The reader will understand how features identified in clauses described in this section can be readily combined with sets of basic features identified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or limiting, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.
[0126] Another implementation of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.
[0127] The inventors disclose the following provisions:
[0128] Clause Set 1 1. An artificial intelligence-based system for detecting changes in gene expression at base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a biological quantity model that processes an input sequence and generates an alternative representation of the input sequence; and biological quantity output generation logic for processing the alternative representations of the input sequence and generating a plurality of biological quantity output sequences; a first biological mass output sequence in the plurality of biological mass output sequences includes a first respective per-base biological mass output for each target base in the target base sequence; the first respective per-base biological abundance output identifying a respective measure of evolutionary conservation of each target base across a plurality of species; a second biological mass output sequence in the plurality of biological mass output sequences comprising a second respective per-base biological mass output for each target base in the target base sequence; The second respective per-base biological abundance output identifies respective measurements of transcription initiation of each target base at each position within the target base sequence, an artificial intelligence-based system. 2. The artificial intelligence-based system of clause 1, wherein each measure of evolutionary conservation is a phylogenetic P-value (phyloP) score that identifies deviations from a null model of neural substitution to detect as conservation a decrease in the rate of substitution of a given target base at a given position within the target base sequence and to detect as acceleration an increase in the rate of substitution of a given target base at a given position. 3. The artificial intelligence-based system of clause 2, wherein each measure of evolutionary conservation is a phastCons score that identifies the posterior probability of a given target base at a given position having a conserved or non-conserved state. 4. The artificial intelligence-based system of clause 2, wherein each measure of evolutionary conservation is a genomic evolutionary rate profiling (GERP) score that identifies a decrease in the number of substitutions of a given target base at a given position across multiple species. 5. The artificial intelligence-based system of clause 1, wherein each measure of transcription initiation is a Cap Analysis of Gene Expression (CAGE) score that identifies the frequency of transcription initiation of a given target base at a given position. 6. A third biological quantity output sequence in the plurality of biological quantity output sequences includes a third respective per-base biological quantity output for each target base in the target base sequence; 10. The artificial intelligence-based system of claim 1, wherein the third per-base biological abundance output identifies a respective measure of an epigenetic signal level for each target base at each position within the target base sequence. 7. The artificial intelligence-based system described in clause 6, wherein the epigenetic signal levels identify DNase I-hypersensitive sites (DHS) or assay for transposase-accessible chromatin by sequencing (ATAC-Seq). 8. The artificial intelligence-based system of clause 6, wherein epigenetic signal levels identify transcription factor (TF) binding. 9. The artificial intelligence-based system described in clause 6, wherein the epigenetic signal level identifies histone modification (HM) marks. 10. a gene expression model that processes a plurality of biological quantity output sequences and generates alternative representations of the plurality of biological quantity output sequences; gene expression output generation logic that processes the alternative representations of the plurality of biological abundance output sequences to generate a gene expression output sequence of a respective per-base gene expression output for each target base in the target base sequence; 10. The artificial intelligence based system of clause 1, wherein the gene expression output for each given base in the gene expression output sequence for a given target base at a given position identifies a measure of the gene expression level of the given target base at the given position. 11. The artificial intelligence-based system of clause 10, wherein gene expression levels are measured in a per-base metric such as CAGE transcription start site (CTSS). 12. The artificial intelligence-based system of clause 10, wherein gene expression levels are measured in a per-gene metric such as transcripts per million (TPM) or reads per kilobase of transcript (RPKM). 13. The artificial intelligence-based system of clause 10, wherein gene expression levels are measured in a per-gene metric, such as fragments per million kilobases (FPKM). 14. The artificial intelligence-based system of clause 1, further configured to include variant classification logic. 15. The artificial intelligence-based system of clause 14, wherein the variant classification logic is further configured to include reference input generation logic that accesses a sequence database and generates a reference base sequence, the reference base sequence comprising a reference target base sequence, the reference target base sequence comprising a reference base at the position being analyzed, the reference base being flanked by a right base sequence having a downstream context base and a left base sequence having an upstream context base. 16. The artificial intelligence-based system of clause 15, wherein the variant classification logic is further configured to include alternative input generation logic that accesses a sequence database and generates an alternative base sequence, the alternative base sequence comprising an alternative target base sequence, the alternative target base sequence comprising an alternative base at the position being analyzed, the alternative base flanked by a right base sequence having a downstream context base and a left base sequence having an upstream context base. 17. The variant classification logic is further configured to include reference processing logic that causes the biomass model to process the reference base sequence and generate alternative representations of the reference base sequence, and further causes the biomass output generation logic to process the alternative representations of the reference base sequence and generate a plurality of reference biomass output sequences; 16. The artificial intelligence based system of clause 15, wherein each reference biological mass output sequence in the plurality of reference biological mass output sequences comprises a respective base-by-base reference biological mass output for each reference target base in the reference target base sequence. 18. A first reference biological quantity output sequence in the plurality of reference biological quantity output sequences includes a first respective base-by-base reference biological quantity output for each reference target base in the reference target base sequence; 18. The artificial intelligence-based system of clause 17, wherein the first respective base-by-base reference biological abundance output identifies a respective measure of evolutionary conservation of each reference target base across a plurality of species. 19. A second reference biological quantity output sequence in the plurality of reference biological quantity output sequences includes a second respective base-by-base reference biological quantity output for each reference target base in the reference target base sequence; 18. The artificial intelligence-based system of clause 17, wherein the second per-each-base reference biological abundance output identifies a respective measure of transcription initiation of each reference target base at each position within the reference target base sequence. 20. The variant classification logic is further configured to include alternative processing logic that causes the biomass model to process the alternative base sequences and generate alternative representations of the alternative base sequences, and further causes the biomass output generation logic to process the alternative representations of the alternative base sequences and generate a plurality of alternative biomass output sequences; 17. The artificial intelligence based system of clause 16, wherein each alternative biological mass output sequence in the plurality of alternative biological mass output sequences comprises an alternative biological mass output for each base for each alternative target base in the alternative target base sequence. 21. A first alternative biological mass output sequence in the plurality of alternative biological mass output sequences includes a first respective base-by-base alternative biological mass output for each alternative target base in the alternative target base sequence; 21. The artificial intelligence-based system of clause 20, wherein the first per-each-base alternative biological abundance output identifies a respective measure of evolutionary conservation of each alternative target base across a plurality of species. 22. A second alternative biological mass output sequence in the plurality of alternative biological mass output sequences includes a second respective base-by-base alternative biological mass output for each alternative target base in the alternative target base sequence; 21. The artificial intelligence-based system of clause 20, wherein the second per-respective base alternative biological abundance output identifies a respective measure of transcription initiation of each alternative target base at a respective position within the alternative target base sequence. 23. The artificial intelligence based system of clause 20, wherein the variant classification logic is further configured to include pathogenicity prediction logic that compares the first reference biological mass output sequence and the first alternative biological mass output sequence position-wise and generates a first delta sequence having a first position-wise sequence difference for positions in the first reference biological mass output sequence and the first alternative biological mass output sequence. 24. The artificial intelligence-based system of clause 23, wherein the pathogenicity prediction logic is further configured to compare the second reference biological quantity output sequence and the second alternative biological quantity output sequence position-by-position and generate a second delta sequence having second position-by-position sequence differences for positions in the second reference biological quantity output sequence and the second alternative biological quantity output sequence. 25. The artificial intelligence-based system of clause 24, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the first delta sequence and the second delta sequence. 26. The artificial intelligence-based system of clause 24, wherein the pathogenicity prediction logic is further configured to accumulate the sequence differences per first position into a first cumulative sequence value and to accumulate the sequence differences per second position into a second cumulative sequence value. 27. The artificial intelligence-based system of clause 26, wherein the first cumulative sequence value is an average of sequence differences per first position, and the second cumulative sequence value is an average of sequence differences per second position. 28. The artificial intelligence-based system of clause 26, wherein the first cumulative sequence value is a sum of sequence differences per first position, and the second cumulative sequence value is a sum of sequence differences per second position. 29. The artificial intelligence-based system of clause 26, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the first cumulative sequence value and the second cumulative sequence value. 30. The artificial intelligence-based system of clause 29, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on an average of the first cumulative sequence value and the second cumulative sequence value. 31. The artificial intelligence-based system of clause 29, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the sum of the first cumulative sequence value and the second cumulative sequence value. 32. The artificial intelligence-based system described in clause 23, wherein the pathogenicity prediction logic is further configured to classify positions within the first delta sequence as belonging to a conserved state or a non-conserved state based on the sequence differences for each of the first positions. 33. The artificial intelligence-based system described in clause 32, wherein the pathogenicity prediction logic is further configured to classify positions in the second delta sequence that match positions in the first delta sequence classified as belonging to a conserved state as belonging to a signal state, and to classify positions in the second delta sequence that match positions in the first delta sequence classified as belonging to a non-conserved state as belonging to a noise state. 34. The pathogenicity prediction logic is further configured to accumulate a subset of the second positional sequence differences into a modulated accumulated sequence value; 34. The artificial intelligence-based system of clause 33, wherein the second positional sequence differences in the subset of second positional sequence differences are located at positions in the second delta sequence that are classified as belonging to the signal state. 35. The artificial intelligence-based system of clause 34, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the modulation cumulative sequence value. 36. The artificial intelligence-based system of clause 34, wherein the modulated cumulative sequence value is an average of the sequence differences per second position within the subset of sequence differences per second position. 37. The artificial intelligence-based system of clause 34, wherein the modulated cumulative sequence value is a sum of sequence differences per second position within a subset of sequence differences per second position. 38. The artificial intelligence-based system described in clause 24, wherein the pathogenicity prediction logic is further configured to compare the respective portions of the first reference biological quantity output sequence and the first alternative biological quantity output sequence position by position, and generate a first delta subsequence having first position-by-position subsequence differences for the positions in the respective portions. 39. The artificial intelligence-based system described in clause 38, wherein the pathogenicity prediction logic is further configured to compare the respective portions of the second reference biological quantity output sequence and the second alternative biological quantity output sequence position by position, and generate a second delta subsequence having second position-by-position subsequence differences for the positions in the respective portions. 40. The artificial intelligence-based system of clause 39, wherein each portion spans adjacent positions to the right and left around the position being analyzed. 41. The artificial intelligence-based system of clause 40, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the first delta subsequence and the second delta subsequence. 42. The artificial intelligence-based system of clause 40, wherein the pathogenicity prediction logic is further configured to accumulate the subsequence differences per first position into a first cumulative subsequence value and to accumulate the subsequence differences per second position into a second cumulative subsequence value. 43. The artificial intelligence-based system of clause 42, wherein the first cumulative subsequence value is an average of the subsequence differences per first position, and the second cumulative subsequence value is an average of the subsequence differences per second position. 44. The artificial intelligence-based system of clause 42, wherein the first cumulative subsequence value is a sum of subsequence differences per first position, and the second cumulative subsequence value is a sum of subsequence differences per second position. 45. The artificial intelligence-based system of clause 42, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the first cumulative subsequence value and the second cumulative subsequence value. 46. The artificial intelligence-based system of clause 45, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on an average of the first cumulative subsequence value and the second cumulative subsequence value. 47. The artificial intelligence-based system of clause 45, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the sum of the first cumulative subsequence value and the second cumulative subsequence value. 48. The artificial intelligence-based system described in clause 38, wherein the pathogenicity prediction logic is further configured to classify positions within the first delta subsequence as belonging to a conserved state or a non-conserved state based on the subsequence differences for each of the first positions. 49. The artificial intelligence-based system described in clause 48, wherein the pathogenicity prediction logic is further configured to classify positions in the second delta subsequence that match positions in the first delta subsequence classified as belonging to the conserved state as belonging to the signal state, and to classify positions in the second delta subsequence that match positions in the first delta subsequence classified as belonging to the non-conserved state as belonging to the noise state. 50. The pathogenicity prediction logic is further configured to accumulate a subset of the second position-wise subsequence differences into a modulated accumulated subsequence value; The artificial intelligence-based system of clause 49, wherein a second positional subsequence difference within the subset of second positional subsequence differences is located at a position within the second delta subsequence that is classified as belonging to the signal state. 51. The artificial intelligence-based system of clause 50, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base depending on the modulated cumulative subsequence value. 52. The artificial intelligence-based system of clause 50, wherein the modulated cumulative subsequence value is an average of the subsequence differences per second position within the subset of the subsequence differences per second position. 53. The artificial intelligence-based system of clause 50, wherein the modulated cumulative subsequence value is a sum of the subsequence differences per second position within the subset of the subsequence differences per second position. 54. The artificial intelligence-based system according to clause 1, wherein the target base sequence is a coding region of a gene. 55. The artificial intelligence-based system according to clause 1, wherein the target base sequence is a non-coding region of a gene. 56. The artificial intelligence-based system of clause 55, wherein the non-coding regions span the transcription start site, the 5 prime untranslated region (UTR), the 3 prime UTR, the enhancer, and the promoter. 56. An alternative base is a singleton variant that occurs in only one outlier in a cohort of outliers, 17. The artificial intelligence-based system of clause 16, wherein an outlier individual in the cohort of outlier individuals exhibits extreme levels of gene expression. 57. The artificial intelligence-based system of clause 56, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels. 58. The artificial intelligence-based system of clause 57, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 59. The artificial intelligence-based system according to clause 56, wherein the singleton variant is a code variant. 60. The artificial intelligence-based system according to clause 56, wherein the singleton variant is a non-code variant. 61. The artificial intelligence-based system according to clause 60, wherein the non-coding variant is a promoter variant. 62. The artificial intelligence-based system of clause 60, wherein the non-coding variant is an enhancer variant. 62. The biological mass model has a first set of weights, 10. The artificial intelligence-based system of claim 1, wherein the biological quantity output generation logic has a second set of weights. 63. During training, a first set of weights of the biological quantity model is trained from scratch to process an input base sequence and generate alternative representations of the input base sequence; 63. The artificial intelligence based system of clause 62, wherein a second set of weights for the bioquantity output generation logic is trained end-to-end from scratch with the first set of weights for the bioquantity model to process alternative representations of the input base sequence and generate a plurality of bioquantity output sequences. 64. During inference, the biological mass model uses the first set of trained weights, 64. The artificial intelligence based system of clause 63, wherein during inference, the biological quantity output generation logic uses the second set of trained weights. 65. The gene expression model has a third set of weights: 10. The artificial intelligence-based system of claim 1, wherein the gene expression output generation logic has a fourth set of weights. 66. A third set of weights for the gene expression model is trained from scratch to process the plurality of biological quantity output sequences and generate alternative representations of the plurality of biological quantity output sequences; 66. The artificial intelligence based system of clause 65, wherein a fourth set of weights for the gene expression output generation logic is trained end-to-end from scratch with the third set of weights for the gene expression model to process alternative representations of the plurality of biological quantity output sequences and generate the gene expression output sequences. 67. During inference, the gene expression model uses a third set of trained weights, 67. The artificial intelligence-based system of clause 66, wherein during inference, the gene expression output generation logic uses the fourth set of trained weights. 68. During training, a first set of weights of the biological abundance model is first trained from scratch to process an input base sequence and generate alternative representations of the input base sequence, and then retrained as a substitute for a third set of weights of the gene expression model to process a plurality of biological abundance output sequences and generate alternative representations of the plurality of biological abundance output sequences; 66. The artificial intelligence based system of clause 65, wherein the fourth set of weights of the gene expression output generation logic is trained end-to-end from scratch using the first set of trained weights substituted in the gene expression model to process alternative representations of the plurality of biological quantity output sequences generated by the first set of trained weights substituted in the gene expression model and generate gene expression output sequences. 69. During inference, the biological mass model uses the first set of retrained weights, During inference, the biological quantity output generation logic uses the second set of trained weights, During inference, the gene expression model uses the first set of retrained weights, 69. The artificial intelligence-based system of clause 68, wherein during inference, the gene expression output generation logic uses the fourth set of trained weights. 70. During training, a first set of weights for the biological quantity model is first trained from scratch to process an input base sequence and generate alternative representations of the input base sequence; During training, a second set of weights for the bioquantity output generation logic is first trained end-to-end from scratch using the first set of weights for the bioquantity model to process alternative representations of the input base sequence and generate a plurality of bioquantity output sequences; During training, the first set of trained weights of the biological mass model are then retrained to process a reference base sequence, generate alternative representations of the reference base sequence, process alternative base sequences, and generate alternative representations of the alternative base sequences; and wherein during training, the second set of trained weights of the biomass output generation logic is then retrained end-to-end with the first set of trained weights of the biomass model to process alternative representations of the reference base sequence to generate a plurality of reference biomass output sequences, and to process alternative representations of the alternative base sequence to generate a plurality of alternative biomass output sequences. 71. During inference, the biological mass model uses the first set of retrained weights, 71. The artificial intelligence based system of clause 70, wherein during inference, the biological quantity output generation logic uses the retrained second set of weights. 72. The artificial intelligence-based system described in clause 23, wherein the pathogenicity prediction logic has a fifth set of weights. 73. During training, a first set of weights for the biological quantity model is first trained from scratch to process an input base sequence and generate alternative representations of the input base sequence; During training, a second set of weights for the bioquantity output generation logic is first trained end-to-end from scratch using the first set of weights for the bioquantity model to process alternative representations of the input base sequence and generate a plurality of bioquantity output sequences; During training, the first set of trained weights of the biomass model and the second set of trained weights of the biomass output generation logic are then retrained end-to-end to generate pathogenicity predictions for the alternative bases. 74. During inference, the biological mass model uses the first set of retrained weights, During inference, the biological quantity output generation logic uses the retrained second set of weights, 74. The artificial intelligence-based system of clause 73, wherein during inference, the pathogenicity prediction logic uses the fifth set of trained weights. 75. The artificial intelligence-based system of clause 18, wherein the first reference biological abundance output for each respective base identifies a respective measurement of a first reference epigenetic signal level for each reference target base at a respective position within the reference target base sequence. 76. The artificial intelligence-based system described in clause 29, wherein the reference biological abundance output for each second base identifies a respective measurement of a second reference epigenetic signal level for each reference target base at a respective position within the reference target base sequence. 77. The artificial intelligence-based system described in clause 21, wherein the first alternative biological quantity output for each respective base identifies a respective measurement of a first alternative epigenetic signal level for each alternative target base at a respective position within the alternative target base sequence. 78. The artificial intelligence-based system described in clause 22, wherein the second alternative biological quantity output for each respective base identifies a respective measurement of a second alternative epigenetic signal level for each alternative target base at a respective position within the alternative target base sequence. 79. The artificial intelligence-based system described in clause 1, wherein during training, the biological mass model and biological mass output generation logic are first trained end-to-end from scratch to translate the analysis of input base sequences into evolutionarily conserved chromatin sequences per base, and then retrained end-to-end to translate the analysis of input base sequences into transcription start frequency chromatin sequences per base. 80. The artificial intelligence-based system described in clause 1, wherein during training, the biological mass model and biological mass output generation logic are first trained end-to-end from scratch to translate analysis of input base sequences into base-by-base epigenetic signal-level chromatin sequences, and then retrained end-to-end to translate analysis of input base sequences into base-by-base evolutionarily conserved chromatin sequences. 81. The artificial intelligence-based system described in clause 1, wherein during training, the biological mass model and biological mass output generation logic are first trained end-to-end from scratch to translate analysis of input base sequences into base-by-base epigenetic signal level chromatin sequences, and then retrained end-to-end to translate analysis of input base sequences into base-by-base transcription start frequency chromatin sequences. 82. The artificial intelligence-based system described in clause 1, wherein during training, the biological mass model and biological mass output generation logic are first trained end-to-end from scratch to translate analysis of input base sequences into base-by-base epigenetic signal level chromatin sequences, and then retrained end-to-end to translate analysis of input base sequences into base-by-base evolutionarily conserved chromatin sequences and base-by-base transcription start frequency chromatin sequences. 83. The artificial intelligence-based system of clause 1, further configured to include a first training set of training input base sequences that include variants confounded by multiple epigenetic effects. 84. The artificial intelligence-based system described in Clause 83, wherein the epigenetic effects in the plurality of epigenetic effects include interchromosomal effects, intragenic effects, population structure and ancestry effects, probabilistic estimation of expression residual (PEER) effects, environmental effects, sex effects, batch effects, genotyping platform effects, and / or library construction protocol effects. 85. The artificial intelligence-based system of clause 83, further configured to include a second training set of training input base sequences that includes variants that are not confounded by multiple epigenetic effects. 86. The artificial intelligence-based system of clause 85, wherein variants in the second training set are reliably determined to alter gene expression and cause extreme levels of gene expression. 87. The artificial intelligence-based system of clause 86, wherein the variants in the second training set include variants that cause overexpression, increasing gene expression levels. 88. The artificial intelligence-based system of clause 86, wherein the variants in the second training set include variants that cause underexpression, resulting in reduced gene expression levels. 89. The artificial intelligence-based system of clause 87, wherein the second training set identifies overexpression probabilities for variants that identify the likelihood of overexpression of the causal gene. 90. The artificial intelligence-based system of clause 88, wherein the second training set identifies expression probabilities for variants that identify the likelihood that the variant causes underexpression of the gene. 91. Each variant in the second training set is a singleton variant that occurs in only one outlier individual in the cohort of outlier individuals; 87. The artificial intelligence-based system of clause 86, wherein an outlier individual in the cohort of outlier individuals exhibits extreme levels of gene expression. 92. The artificial intelligence-based system according to clause 91, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels. 93. The artificial intelligence-based system of clause 91, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 94. The artificial intelligence-based system according to clause 91, wherein the singleton variant is a code variant. 95. The artificial intelligence-based system described in clause 91, wherein the singleton variant is a non-code variant. 96. The artificial intelligence-based system according to clause 95, wherein the non-coding variant is a 5 prime untranslated region (UTR) variant, a 3 prime UTR variant, an enhancer variant, or a promoter variant. 97. The artificial intelligence-based system of clause 85, wherein the variants in the second training set span multiple tissue types. 98. The artificial intelligence-based system of clause 85, wherein the variants in the second training set span multiple cell types. 99. The artificial intelligence-based system of clause 1, wherein the input base sequence and the plurality of biological abundance output sequences span a plurality of tissue types. 100. The artificial intelligence-based system of clause 1, wherein the input base sequence and the plurality of biological abundance output sequences span a plurality of cell types. 101. The artificial intelligence-based system of clause 10, wherein the gene expression output sequences span multiple tissue types. 102. The artificial intelligence-based system of clause 10, wherein the gene expression output sequences span multiple cell types. 103. The artificial intelligence-based system of clause 1, wherein the biological quantity model and biological quantity output generation logic are first trained end-to-end on a first training set and then retrained on a second training set. 104. The artificial intelligence-based system of clause 1, wherein the variants in the second training set are used as a pathogenic set labeled with a first ground truth label indicating a gene expression change, and the common variants are used as a benign set labeled with a second ground truth label indicating no gene expression change. 105. The artificial intelligence-based system of clause 104, wherein the benign set is balanced for trinucleotide context, homopolymer, k-mer, neighborhood GC frequency, and sequencing depth. 106. The artificial intelligence-based system of clause 104, wherein, based on cutoff probabilities applied to the overexpression and underexpression probabilities, the variants in the second training set are divided into an overexpressed variant training set having a first ground truth label indicating increased gene expression, an overexpressed variant training set having a second ground truth label indicating decreased gene expression, and a neurally expressed variant training set indicating maintained gene expression. 107. The artificial intelligence-based system of clause 10, wherein the gene expression model and gene expression output generation logic are first trained end-to-end on a first training set and then retrained on a second training set. 108. The artificial intelligence-based system described in clause 1, wherein the biological quantity model and biological quantity output generation logic are first trained end-to-end on a first training set and then retrained on variants in a second training set that occur on odd-numbered chromosomes. 109. The artificial intelligence-based system of clause 10, wherein the gene expression model and gene expression output generation logic are first trained end-to-end on a first training set and then retrained on variants in a second training set that occur on odd-numbered chromosomes. 110. The artificial intelligence-based system described in clause 85, wherein the variants in the second training set are not used for training, but instead are used as a validation set to evaluate the performance of the trained biological quantity model 124, the trained biological quantity output generation logic, the trained gene expression model, and the trained gene expression output generation logic. 111. The artificial intelligence-based system of clause 110, wherein variants in the second training set that occur on even-numbered chromosomes are used as a validation set. 112. The artificial intelligence-based system described in clause 1, wherein the size of the target base sequence is varied during training to account for variations in the offset position of the transcription start site (TSS). 113. An artificial intelligence-based system for detecting changes in gene expression at base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a biological quantity model that processes an input sequence and generates an alternative representation of the input sequence; and biological quantity output generation logic for processing the alternative representations of the input sequence and generating a plurality of biological quantity output sequences; An artificial intelligence based system, wherein each biological mass output sequence in the plurality of biological mass output sequences includes a respective per-base biological mass output for each target base in the target base sequence. 114. A first biological quantity output array in the plurality of biological quantity output arrays includes a first respective base-by-base biological quantity output for each target base in the target base sequence; 114. The artificial intelligence-based system of clause 113, wherein the first per-base biological abundance output identifies a respective measure of evolutionary conservation of each target base across a plurality of species. 115. A second biological quantity output sequence in the plurality of biological quantity output sequences includes a second respective base-by-base biological quantity output for each target base in the target base sequence; 114. The artificial intelligence-based system of clause 113, wherein the second per-each-base biological abundance output identifies a respective measure of transcription initiation of each target base at each position within the target base sequence.
[0129] Clause Set 2 1. A system comprising: validation data having a set of variants with a set of causality scores, the causality scores specifying a statistically unconfounded likelihood of altering gene expression; validation set discretization logic configured to classify each causal relationship score in the set of causal relationship scores into a gene expression change class or a gene expression conservation class based on application of a cutoff to the set of causal relationship scores, thereby generating a ground truth split of the set of variants into gene expression change classes and gene expression conservation classes; inference logic configured to cause a model to generate a set of prediction scores for the set of variants, wherein a prediction score in the set of prediction scores specifies an inferred likelihood of altering gene expression, and the model is trained to determine the gene expression altering potential of the variants; model score discretization logic configured to classify each prediction score in the set of prediction scores into a gene expression change class or a gene expression conservation class based on application of a threshold to the set of prediction scores, thereby generating an inferred bifurcation of the set of variants into gene expression change classes and gene expression conservation classes; validation logic configured to determine a performance measure of the model based on a comparison of the inferred branches and the ground truth branches. 2. The system of clause 1, wherein the ground truth branch assigns a first label (e.g., 0) to variants that fall into a gene expression alteration class and a second label (e.g., 1) to variants that fall into a gene expression conservation class. 3. The system of clause 2, wherein the ground truth branch assigns a first label (e.g., 0) to variants that fall into a gene expression alteration class and a second label (e.g., 1) to variants that fall into a gene expression conservation class. In other embodiments, the ground truth branching splits variants into three categories: -1, 0, and 1, corresponding to decreased gene expression, no change in gene expression, and increased gene expression, i.e., in such embodiments, a single classifier can perform three-way classification. 4. The system of clause 3, further configured to encode the ground truth branch into a first vector and the inferred branch into a second vector. 5. The system of clause 4, further configured to determine a performance measure of the model based on an element-by-element comparison of the first vector and the second vector. 6. The system of clause 1, further configured to determine a performance measure of the model based on an odds ratio of the number of variants classified into the gene expression altered class by the inferred branch, the number of variants classified into the gene expression conserved class by the inferred branch, the number of variants classified into the gene expression altered class by the ground truth branch, and the number of variants classified into the gene expression conserved class by the ground truth branch. 7. The system of clause 1, wherein the set of variants has a set of under-expression causality scores, the under-expression causality scores identifying a statistically unconfounded likelihood of decreasing gene expression. 8. The system of clause 7, wherein the validation set discretization logic is further configured to classify each underexpression causality score in the set of underexpression causality scores into a decreased gene expression class or a non-decreased gene expression class based on application of an underexpression cutoff to the set of underexpression causality scores, thereby generating an underexpression ground truth branching of the set of variants into decreased gene expression classes and non-decreased gene expression classes. 9. The inference logic is further configured to cause the model to generate a set of underexpression prediction scores for the set of variants, wherein an underexpression prediction score in the set of underexpression prediction scores specifies an inferred likelihood of reducing gene expression, and the model is trained to determine the likelihood of a variant to reduce gene expression; the model score discretization logic is further configured to classify each underexpression prediction score in the set of underexpression prediction scores into a reduced gene expression class or a non-reduced gene expression class based on application of an underexpression threshold to the set of underexpression prediction scores, thereby generating underexpression inferential branches of the set of variants into reduced gene expression and non-reduced gene expression classes; 9. The system of clause 8, wherein the validation logic is further configured to determine an underrepresentation performance measure of the model based on a comparison of the underrepresented inferred branch and the underrepresented ground truth branch. 10. The system of clause 9, wherein the model score discretization logic is further configured to sort the set of underexpression prediction scores in descending order, classify a subset of N lowest underexpression prediction scores in the sorted set of underexpression prediction scores into a decreased gene expression class, and classify a subset of remaining underexpression prediction scores in the sorted set of underexpression prediction scores into a non-decreased gene expression class. 11. The system of clause 8, wherein the underexpression ground truth branch assigns a first label (e.g., 0) to variants that fall into a reduced gene expression class and a second label (e.g., 1) to variants that fall into a non-reduced gene expression class. 12. The system of clause 11, wherein the underexpression inference branch assigns a first label (e.g., 0) to variants that fall into a reduced gene expression class and a second label (e.g., 1) to variants that fall into a non-reduced gene expression class. 13. The system of clause 12, further configured to encode the under-represented ground truth branches into a first vector and encode the inferred branches into a second vector. 14. The system of clause 13, further configured to determine an underrepresentation performance measure of the model based on an element-by-element comparison of the first vector and the second vector. 15. The system of clause 9, further configured to determine an underexpression performance measure of the model based on an odds ratio of the number of variants classified into the reduced gene expression class by the underexpression inference branch, the number of variants classified into the non-reduced gene expression class by the underexpression inference branch, the number of variants classified into the reduced gene expression class by the underexpression ground truth branch, and the number of variants classified into the non-reduced gene expression class by the underexpression ground truth branch. 16. The system of clause 1, wherein the set of variants has a set of overexpression causality scores, the overexpression causality scores identifying a statistically unconfounded likelihood of increasing gene expression. 17. The system of clause 16, wherein the validation set discretization logic is further configured to classify each overexpression causality score in the set of overexpression causality scores into an increased gene expression class or a non-increased gene expression class based on application of an overexpression cutoff to the set of overexpression causality scores, thereby generating an overexpression ground truth branching of the set of variants into increased gene expression classes and non-increased gene expression classes. 18. The inference logic is further configured to cause the model to generate a set of overexpression prediction scores for the set of variants, wherein an overexpression prediction score in the set of overexpression prediction scores identifies an inferred likelihood of increasing gene expression, and the model is trained to determine the likelihood of the variants increasing gene expression; the model score discretization logic is further configured to classify each overexpression prediction score in the set of overexpression prediction scores into an increased gene expression class or a non-increased gene expression class based on application of an overexpression threshold to the set of overexpression prediction scores, thereby generating an overexpression inferential branch of the set of variants into an increased gene expression class and a non-increased gene expression class; 18. The system of clause 17, wherein the validation logic is further configured to determine an overexpression performance measure for the model based on a comparison of the overexpressed inferred branch and the overexpressed ground truth branch. 19. The system of clause 18, wherein the model score discretization logic is further configured to sort the set of overexpression prediction scores in descending order, classify a subset of the N highest overexpression prediction scores in the sorted set of overexpression prediction scores into an increased gene expression class, and classify a subset of the remaining overexpression prediction scores in the sorted set of overexpression prediction scores into a non-increased gene expression class. 20. The system of clause 17, wherein the overexpression ground truth branch assigns a first label (e.g., 0) to variants that fall into an increased gene expression class and a second label (e.g., 1) to variants that fall into a non-increased gene expression class. 21. The system of clause 20, wherein the overexpression inference branch assigns a first label (e.g., 0) to variants that fall into an increased gene expression class and a second label (e.g., 1) to variants that fall into a non-increased gene expression class. 22. The system of clause 21, further configured to encode the over-represented ground truth branch into a first vector and the inferred branch into a second vector. 23. The system of clause 22, further configured to determine an overrepresentation performance measure of the model based on an element-by-element comparison of the first vector and the second vector. 24. The system of clause 18, further configured to determine an overexpression performance measure of the model based on an odds ratio of the number of variants classified into the increased gene expression class by the overexpression inference branch, the number of variants classified into the non-increased gene expression class by the overexpression inference branch, the number of variants classified into the increased gene expression class by the overexpression ground truth branch, and the number of variants classified into the non-increased gene expression class by the overexpression ground truth branch. 25. The inference logic is further configured to cause the first model to generate a first set of prediction scores for the set of variants, wherein the prediction scores in the first set of prediction scores specify an inferred likelihood of altering gene expression, and the first model is trained to determine the gene expression altering likelihood of the variants; the model score discretization logic is further configured to classify each prediction score in the first set of prediction scores into a gene expression change class or a gene expression conservation class based on application of a first threshold to the first set of prediction scores, thereby generating a first inferred branching of the set of variants into gene expression change classes and gene expression conservation classes; 10. The system of claim 1, wherein the validation logic is further configured to determine a first performance measure of the first model based on a comparison of the first inferred branch to a ground truth branch. 26. The inference logic is further configured to cause a second model to generate a second set of prediction scores for the set of variants, wherein the prediction scores in the second set of prediction scores identify an inferred likelihood of altering gene expression, and the second model is trained to determine the gene expression altering potential of the variants; 26. The system of clause 25, wherein the model score discretization logic is further configured to classify each prediction score in the second set of prediction scores into a gene expression change class or a gene expression conservation class based on application of a second threshold to the second set of prediction scores, thereby generating a second inferred branch of the set of variants into gene expression change classes and gene expression conservation classes, and the validation logic is further configured to determine a second performance measure of the second model based on a comparison of the second inferred branch to a ground truth branch. 27. The inference logic is further configured to cause a third model to generate a third set of prediction scores for the set of variants, the prediction scores in the third set of prediction scores specifying an inferred likelihood of altering gene expression, and the third model is trained to determine the gene expression altering potential of the variants; the model score discretization logic is further configured to classify each prediction score in the third prediction score set into a gene expression change class or a gene expression conservation class based on application of a third threshold to the third prediction score set, thereby generating a third inferred bifurcation of the variant set into gene expression change classes and gene expression conservation classes; 27. The system of clause 26, wherein the validation logic is further configured to determine a third performance measure of the third model based on a comparison of the third inferred branch and the ground truth branch. 28. The system of clause 27, further configured to require the first, second, and third inferred branches to classify the same number of variants into gene expression change classes, thereby making the first, second, and third performance measures comparable to each other. 29. The system of clause 28, further configured to compare performance of each of the first, second, and third models on the validation data based on a comparison of the first, second, and third performance measures. 30. The system of clause 27, wherein the first, second, and third thresholds are different. 31. The system of clause 27, wherein at least some of the first, second, and third thresholds are the same. 32. The system of clause 1, further configured to generate a ground truth branching of each of the sets of variants into gene expression change classes and gene expression conservation classes based on the application of each of different cutoffs to the set of causality scores. 33. The system of clause 1, further configured to generate a ground truth trifurcation of the set of variants into a decreased gene expression class, an increased gene expression class, and a preserved gene expression class. 34. The system of clause 33, further configured to generate an inferred trifurcation of the set of variants into a decreased gene expression class, an increased gene expression class, and a conserved gene expression class. 35. Each variant in the set of variants is a singleton variant that occurs in only one outlier individual in a cohort of outlier individuals; 10. The system of claim 1, wherein an outlier individual in the cohort of outlier individuals exhibits extreme levels of gene expression. 36. The system of clause 35, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels. 37. The system of clause 35, wherein extreme levels of gene expression include over-gene expression and under-gene expression. 38. The system of clause 35, wherein the singleton variant is a code variant. 39. The system of clause 35, wherein the singleton variant is a non-coding variant. 40. The system of clause 39, wherein the non-coding variant is a 5 prime untranslated region (UTR) variant, a 3 prime UTR variant, an enhancer variant, or a promoter variant. 41. The system of clause 1, wherein the set of variants spans multiple tissue types. 42. The system of clause 1, wherein the set of variants spans multiple cell types. 43. A system comprising: validation data having a set of observations having a set of ground truth scores; validation set discretization logic configured to classify each ground truth score in the set of ground truth scores into a first class or a second class based on application of a cutoff to the set of ground truth scores, thereby generating a ground truth split of the set of observations into the first class and the second class; inference logic configured to cause the trained model to generate a set of predicted scores for a set of observations; model score discretization logic configured to classify each prediction score in the set of prediction scores into a first class or a second class based on application of a threshold to the set of prediction scores, thereby generating an inferred bifurcation of the set of observations into the first class and the second class; validation logic configured to determine a performance measure of the trained model based on a comparison of the inferred branch and the ground truth branch. 44. The system of clause 43, further configured to generate respective inferred branches for each model such that each of the respective inferred branches is required to classify the same number of observations into the first class, thereby making the respective performance measures of the respective models comparable to each other. 45. The system of clause 44, further configured to compare the performance of each of the models on the validation data based on a comparison of the respective performance measures.
[0130] While the present invention has been disclosed with reference to the above-described preferred embodiments and examples, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims. [Explanation of symbols]
[0131] 100 systems 110 Sequence Database 122 input sequences 124 Biological Mass Model 126 input sequences 136 Biological Quantity Output Array 200 input sequences 202 Sequence Database 224 upstream context bases 226 target base sequence 228 downstream context bases 300 Alternative expressions 302 reference gene sequences 304 upstream sequence 306 nucleotides 308 downstream sequence 322 Alternate Arrays 324 upstream sequence 326 variants 328 downstream sequence 336 variants 342 Alternate Arrays 400 training datasets 403 reference base positions 422 training dataset 433 single nucleotide variants 452 Second training dataset 462 sequences 463 single nucleotide variants 500 One embodiment 502 sequence database 524 input sequences 528 Biological Quantity Output Sequence Generator 548 Biological Quantity Output Array 566 First training iteration set 586 Second training iteration set 600 Implementation 602 sequence database 624 input sequences 628 Biological Quantity Output Sequence Generator 648 Biological Quantity Output Array 666 First training iteration set 667 Second training iteration set 686 Verification Process 700 Implementation 702 reference sequences 704 Alternate Array 722 Alternative representation generator 724 Alternative Expression Generator 742 Biological Quantity Output Sequence Generator 744 Biological Quantity Output Sequence Generator 762 Biological Quantity Output Array 764 Biological Quantity Output Array 783 Flow Diagram 800 One embodiment 802 First Reference Biological Quantity Output Sequence 804 First Alternate Biological Quantity Output Array 806 Second Reference Biological Quantity Output Sequence 808 Second Alternate Biological Quantity Output Array 810 Biological Quantity Output Array 812 First Delta Array 814 Second Delta Array 822 First cumulative average delta value 824 Second Average Delta Value 826 base resolution pathogenicity classification logic 833 Final average delta value 842 Preservation status 844 Unpreserved 846 Preservation status 900 One embodiment 902 First Reference Biological Quantity Output Sequence 904 First Alternate Biological Quantity Output Sequence 906 Second Reference Biological Quantity Output Sequence 909 Second Alternate Biological Quantity Output Sequence 910 Biological Quantity Output Array 912 First Delta Array 914 Second Delta Array 922 First cumulative total delta value 924 Second Total Delta Value 924 Second cumulative total delta value 926 Base resolution pathogenicity classification logic 933 Final Total Delta Value 942 State of preservation 944 Unpreserved 946 State of preservation 1000 Implementation 1022 input sequences 1042 First Set 1044 Alternative representation generator 1062 Second Set 1064 Biological Quantity Output Sequence Generator 1081 First Output Array 1083 Second Output Array 1084 Biological Quantity Output Array 1100 Implementation 1102 Input sequence 1122 First set 1124 Alternative representation generator 1142 Second Set 1144 Biological Quantity Output Sequence Generator 1162 First Output Array 1164 Second Output Scale 1168 Biological Quantity Output Array 1200 Implementation 1202 input sequences 1206 input sequences 1222 First set 1224 Alternative representation generator 1226 Second Set 1228 Biological Quantity Output Sequence Generator 1242 input sequences 1246 Biological Quantity Output Array 1300 Implementation 1302 input sequences 1304 Input base sequence 1322 First Set 1324 sets 1342 Second Set 1344 sets 1362 Biological Quantity Output Array 1364 Second Output Array 1366 Biological Quantity Output Array 1400 Implementation 1402 input sequences 1404 Alternative representation generator 1406 Biological Quantity Output Array 1422 First Set 1426 Third Set 1428 Biological quantity alternative expression generator 1442 Second Set 1444 Biological Quantity Output Sequence Generator 1446 Fourth Set 1448 Gene Expression Output Generator 1462 Biological Quantity Output Array 1462 Output Array 1466 Gene Expression Output Sequences 1468 base resolution gene expression 1500 Implementation 1502 input sequences 1506 Biological Quantity Output Array 1512 Biological Quantity Output Array 1522 First set 1524 Alternative representation generator 1528 Biological quantity alternative expression generator 1542 input sequences 1546 Fourth Set 1548 Gene Expression Output Generator 1562 input sequences 1566 Gene Expression Output Sequences 1568 base resolution gene expression 1582 Second Set 1584 Biological Quantity Output Sequence Generator 1600 Implementation 1602 input sequences 1606 Biological Quantity Output Array 1612 Biological Quantity Output Array 1616 Gene Expression Output Sequence 1618 base resolution gene expression 1622 First set 1624 Alternative representation generator 1628 Biological quantity alternative expression generator 1642 input sequences 1646 Alternative biological quantity expression 1662 input sequences 1666 Alternative biological quantity expression 1682 Second set 1684 Biological Quantity Output Sequence Generator 1686 4th set 1688 Gene Expression Output Generator 1700 Implementation 1702 input sequences 1706 Biological Quantity Output Array 1716 Gene Expression Output Sequence 1718 base resolution gene expression 1722 First set 1724 Alternative representation generator 1726 Third Set 1728 Biological quantity alternative expression generator 1742 Second Set 1744 Biological Quantity Output Sequence Generator 1746 Alternative biological quantity expression 1762 Biological Quantity Output Array 1766 Alternative biological quantity expression 1786 Fourth Set 1788 Gene Expression Output Generator 1800 Implementation 1802 base sequences 1812 1st set 1813 Alternative representation generator 1822 Alternative Array Representations 1832 Biological Quantity Output Array 1842 Third set 1843 Biological quantity alternative expression generator 1852 Fourth set 1862 Gene Expression Output Sequence 1863 Gene Expression Output Generator 1900 Implementation 1902 input sequences 1912 First set 1913 Alternative representation generator 1922 Alternative Array Representations 1932 Third set 1942 Alternative biological quantity expression 1952 Alternative biological quantity expression 1962 Fourth set 1963 Gene Expression Output Generator 1972 Gene Expression Output Sequence 1973 Base Resolution Gene Expression 2000 Implementation 2004 2nd set 2014 Biological Quantity Output Array 2034 Third Set 2034 sets 2035 Biological quantity alternative expression generator 2042 input sequences 2044 Alternative biological quantity expression 2052 First Set 2053 Alternative representation generator 2054 Alternative biological quantity expression 2062 Alternative Array Representations 2064 Fourth Set 2065 Gene Expression Output Generator 2074 gene expression output sequences 2075 base resolution gene expression 2100 One embodiment 2104 Second Set 2114 Biological Quantity Output Array 2124 Biological Quantity Output Array 2132 input sequences 2134 Third Set 2134 sets 2135 Biological quantity alternative expression generator 2142 First Set 2144 Fourth Set 2145 Gene Expression Output Generator 2152 Alternative Array Representations 2154 Gene Expression Output Sequence 2155 base resolution gene expression 2200 Implementation 2202 Reference input sequence 2204 Alternative Input Sequences 2212 First Set 2213 Alternative representation generator 2214 First Set 2216 Alternative representation generator 2222 Alternative expression 2223 5th set 2224 Alternative expression 2232 input 2233 Alternative Sequence Pathogenicity Prediction 2234 input 2242 Second Set 2243 Biological Quantity Output Sequence Generator 2244 Second Set 2246 Biological Quantity Output Sequence Generator 2253 Reference Biological Quantity Output Sequence 2256 Alternative Biological Quantity Output Sequence 2300 One embodiment 2302 Reference input sequence 2304 Alternative Input Sequences 2322 First set 2323 Alternative representation generator 2324 First Set 2326 Alternative representation generator 2342 Second Set 2343 Biological Quantity Output Sequence Generator 2344 Second Set 2346 Biological Quantity Output Sequence Generator 2353 5th set 2362 Reference Biological Quantity Output Sequence 2363 Alternative Sequence Pathogenicity Prediction 2364 Alternative Biological Quantity Output Sequence 2400 scale 2401 Alternative representation generator 2402 Biological Quantity Output Sequence Generator 2403 phyloP score 2404 Alternative representation generator 2405 Biological Quantity Output Sequence Generator 2406 phastCons value 2407 Alternative representation generator 2408 Biological Quantity Output Sequence Generator 2409 phastCons value 2500 scale 2502 Alternative representation generator 2504 Biological Quantity Output Sequence Generator 2600 Epigenetic Signaling 2601 Alternative representation generator 2602 Biological Quantity Output Sequence Generator 2603 I Hypersensitive site prediction output 2604 Alternative representation generator 2605 Biological Quantity Output Sequence Generator 2606 Transcription factor binding site prediction output 2607 Alternative representation generator 2608 Biological Quantity Output Sequence Generator 2609 Histone modification mark prediction output 2700 Expression Change Classifier 2700 Gene Expression Change Probability Classifier 2702 Variant Expression Conserved Causality Score Validation Dataset 2704 Binary Classification Model 2724 Gene Expression Change Class 2744 Gene Expression Conservation Class 2800 Down-Regulation Classifier 2804 Binary Classification Model 2824 gene expression non-decreasing class 2844 Decreased Gene Expression Class 2900 Up-Regulation Classifier 2902 variant overexpression causality score validation dataset 2904 Binary Classification Model 2924 gene expression reduction class 2944 gene expression non-decreasing class 3000 Multi-class Expression Classifier 3002 Validation Data 3024 Causality Score 3026 Binary Classifier 3028 class 3044 Causality Score 3046 Binary Classifier 3048 class 3058 class 3064 Causality Score 3066 Binary Classifier 3068 class 3078 class 3080 Multi-class Classifier 3082 Gene Expression Conservation Class 3084 Decreased Gene Expression Class 3086 Gene Expression Increase Class 3100 Gene Expression Classifier Training 3102 variant training dataset 3122 Ground Truth Causality Score 3142 Binary Classifier 3144 Causality Score 3161 First Class 3163 Second Class 3200 Computer System 3210 Storage Subsystem 3222 Memory Subsystem 3232 File Storage Subsystem 3234 dedicated memory (read only memory, ROM) 3238 User Interface Input Devices 3255 Bus Subsystem 3274 Network Interface Subsystem 3276 User Interface Output Device 3278 processor
Claims
1. 1. An artificial intelligence based system for detecting changes in gene expression at base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a biological mass model that processes the input sequence and generates alternative representations of the input sequence; and biological mass output generation logic for processing the alternative representations of the input base sequence to generate a plurality of biological mass output sequences; a first biological mass output sequence in the plurality of biological mass output sequences comprising a first respective base-by-base biological mass output for each of the target bases in the target base sequence; the first per-base biological abundance output identifying a respective measure of evolutionary conservation of the respective target base across a plurality of species; a second biological mass output sequence in the plurality of biological mass output sequences comprising a second respective base-by-base biological mass output for each target base in the target base sequence; The second respective base-by-base biological abundance output identifies a respective measure of transcription initiation of the respective target base at each position within the target base sequence.
2. 2. The artificial intelligence-based system of claim 1, wherein each measure of evolutionary conservation is a phylogenetic P-value (phyloP) score that identifies deviations from a null model of neural substitution to detect a decrease in the substitution rate of a given target base at a given position within the target base sequence as conservation and an increase in the substitution rate of the given target base at the given position as acceleration.
3. 3. The artificial intelligence-based system of claim 2, wherein each measure of evolutionary conservation is a phastCons score that specifies a posterior probability of the given target base at the given position having a conserved or non-conserved state.
4. 3. The artificial intelligence-based system of claim 2, wherein each measure of evolutionary conservation is a Genome Evolutionary Rate Profiling (GERP) score that identifies a decrease in the number of substitutions of the given target base at the given position across the plurality of species.
5. 2. The artificial intelligence-based system of claim 1, wherein each measure of transcription initiation is a Cap Analysis of Gene Expression (CAGE) score that specifies the transcription initiation frequency of the given target base at the given position.
6. a third biological mass output sequence in the plurality of biological mass output sequences comprising a third respective base-by-base biological mass output for each target base in the target base sequence; 2. The artificial intelligence-based system of claim 1, wherein the third per-base biological abundance output identifies a respective measure of an epigenetic signal level for the respective target base at a respective position within the target base sequence.
7. 7. The artificial intelligence-based system of claim 6, wherein the epigenetic signal levels identify DNase I hypersensitive sites (DHS) or assay for transposase accessible chromatin by sequencing (ATAC-Seq).
8. The artificial intelligence-based system of claim 6 , wherein the epigenetic signal level identifies transcription factor (TF) binding.
9. 7. The artificial intelligence-based system of claim 6, wherein the epigenetic signal level identifies histone modification (HM) marks.
10. a gene expression model that processes the plurality of biological abundance output sequences and generates alternative representations of the plurality of biological abundance output sequences; gene expression output generation logic that processes the alternative representations of the plurality of biological abundance output sequences to generate a gene expression output sequence of a base-by-base gene expression output for each target base in the target base sequence; 2. The artificial intelligence based system of claim 1, wherein the gene expression output for each given base in the gene expression output sequence for the given target base at the given position specifies a measure of the gene expression level of the given target base at the given position.
11. 11. The artificial intelligence-based system of claim 10, wherein the gene expression levels are measured in a base-by-base metric such as CAGE transcription start site (CTSS).
12. 11. The artificial intelligence-based system of claim 10, wherein the gene expression levels are measured in a per-gene metric such as transcripts per million (TPM) or reads per kilobase of transcript (RPKM).
13. 11. The artificial intelligence-based system of claim 10, wherein the gene expression levels are measured in a per-gene metric such as fragments per million kilobases (FPKM).
14. The artificial intelligence-based system of claim 1 , further configured to include variant classification logic.
15. 15. The artificial intelligence-based system of claim 14, wherein the variant classification logic is further configured to include reference input generation logic that accesses the sequence database and generates a reference base sequence, the reference base sequence comprising a reference target base sequence, the reference target base sequence comprising a reference base at a position to be analyzed, the reference base being flanked by a right base sequence having a downstream context base and a left base sequence having an upstream context base.
16. 16. The artificial intelligence-based system of claim 15, wherein the variant classification logic is further configured to include alternative input generation logic that accesses the sequence database and generates an alternative base sequence, the alternative base sequence comprising an alternative target base sequence, the alternative target base sequence comprising an alternative base at the analyzed position, the alternative base flanking the right base sequence with the downstream context base and the left base sequence with the upstream context base.
17. the variant classification logic is further configured to include reference processing logic that causes the biomass model to process the reference base sequence and generate alternative representations of the reference base sequence, and further causes the biomass output generation logic to process the alternative representations of the reference base sequence and generate a plurality of reference biomass output sequences; 16. The artificial intelligence-based system of claim 15, wherein each reference biological quantity output sequence in the plurality of reference biological quantity output sequences includes a reference biological quantity output for each base for each reference target base in the reference target base sequence.
18. a first reference biological mass output sequence in the plurality of reference biological mass output sequences comprising a first respective base-by-base reference biological mass output for each of the reference target bases in the reference target base sequence; 18. The artificial intelligence-based system of claim 17, wherein the first respective base-by-base reference biological abundance output identifies a respective measure of evolutionary conservation of the respective reference target base across the plurality of species.
19. a second reference biological mass output sequence in the plurality of reference biological mass output sequences comprising a second respective base-by-base reference biological mass output for each of the reference target bases in the reference target base sequence; 18. The artificial intelligence-based system of claim 17, wherein the second respective base-by-base reference biological abundance output identifies a respective measurement of transcription initiation of the respective reference target base at a respective position within the reference target base sequence.
20. the variant classification logic is further configured to include alternative processing logic that causes the biomass model to process the alternative base sequences and generate alternative representations of the alternative base sequences, and further causes the biomass output generation logic to process the alternative representations of the alternative base sequences and generate a plurality of alternative biomass output sequences; 17. The artificial intelligence-based system of claim 16, wherein each alternative biological mass output sequence in the plurality of alternative biological mass output sequences includes an alternative biological mass output for each base for each alternative target base in the alternative target base sequence.
21. a first alternative biological mass output sequence in the plurality of alternative biological mass output sequences comprising a first respective base-by-base alternative biological mass output for the respective alternative target base in the alternative target base sequence; 21. The artificial intelligence based system of claim 20, wherein the first respective base-by-base alternative biological abundance output identifies a respective measure of evolutionary conservation of the respective alternative target base across the plurality of species.
22. a second alternative biological mass output sequence in the plurality of alternative biological mass output sequences comprising a second respective base-by-base alternative biological mass output for each of the alternative target bases in the alternative target base sequence; 21. The artificial intelligence-based system of claim 20, wherein the second per-each-base alternative biological abundance output identifies a respective measure of transcription initiation of the respective alternative target base at a respective position within the alternative target base sequence.
23. 21. The artificial intelligence based system of claim 20, wherein the variant classification logic is further configured to include pathogenicity prediction logic that compares the first reference biological mass output sequence and the first alternative biological mass output sequence position by position and generates a first delta sequence having first position-by-position sequence differences for positions in the first reference biological mass output sequence and the first alternative biological mass output sequence.
24. 24. The artificial intelligence based system of claim 23, wherein the pathogenicity prediction logic is further configured to compare the second reference biological mass output sequence and the second alternative biological mass output sequence position by position and generate a second delta sequence having second position-by-position sequence differences for positions in the second reference biological mass output sequence and the second alternative biological mass output sequence.
25. 25. The artificial intelligence-based system of claim 24, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base dependent on the first delta sequence and the second delta sequence.
26. 25. The artificial intelligence-based system of claim 24, wherein the pathogenicity prediction logic is further configured to accumulate the first position-by-position sequence differences into a first cumulative sequence value and accumulate the second position-by-position sequence differences into a second cumulative sequence value.
27. 27. The artificial intelligence-based system of claim 26, wherein the first cumulative sequence value is an average of sequence differences for each of the first positions, and the second cumulative sequence value is an average of sequence differences for each of the second positions.
28. 27. The artificial intelligence-based system of claim 26, wherein the first cumulative sequence value is a sum of sequence differences for each of the first positions, and the second cumulative sequence value is a sum of sequence differences for each of the second positions.
29. 27. The artificial intelligence-based system of claim 26, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on the first cumulative sequence value and the second cumulative sequence value.
30. 30. The artificial intelligence-based system of claim 29, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on an average of the first cumulative sequence value and the second cumulative sequence value.
31. 30. The artificial intelligence-based system of claim 29, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on a sum of the first cumulative sequence value and the second cumulative sequence value.
32. 24. The artificial intelligence-based system of claim 23, wherein the pathogenicity prediction logic is further configured to classify positions within the first delta sequence as belonging to a conserved state or a non-conserved state based on the sequence differences for each of the first positions.
33. 33. The artificial intelligence-based system of claim 32, wherein the pathogenicity prediction logic is further configured to classify positions in the second delta sequence that match positions in the first delta sequence classified as belonging to the conserved state as belonging to a signal state, and to classify positions in the second delta sequence that match positions in the first delta sequence classified as belonging to the non-conserved state as belonging to a noise state.
34. the pathogenicity prediction logic is further configured to accumulate the subset of second positional sequence differences into a modulated accumulated sequence value; 34. The artificial intelligence-based system of claim 33, wherein a second positional sequence difference in the subset of second positional sequence differences is located at a position in the second delta sequence that is classified as belonging to the signal state.
35. 35. The artificial intelligence-based system of claim 34, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on the modulated cumulative sequence value.
36. 35. The artificial intelligence-based system of claim 34, wherein the modulated cumulative sequence value is an average of the second position-by-position sequence differences within the subset of the second position-by-position sequence differences.
37. 35. The artificial intelligence-based system of claim 34, wherein the modulated cumulative array value is a sum of the array differences per second position within the subset of array differences per second position.
38. 25. The artificial intelligence-based system of claim 24, wherein the pathogenicity prediction logic is further configured to compare the first reference biological mass output sequence and the first alternative biological mass output sequence, position by position, in respective portions thereof, and generate a first delta subsequence having first position-by-position subsequence differences for positions in the respective portions.
39. 39. The artificial intelligence-based system of claim 38, wherein the pathogenicity prediction logic is further configured to compare the second reference biological mass output sequence and the second alternative biological mass output sequence, position by position, in respective portions thereof, and generate a second delta subsequence having second position-by-position subsequence differences for positions in the respective portions.
40. 40. The artificial intelligence-based system of claim 39, wherein the respective portions span adjacent locations to the right and left around the location of interest.
41. 41. The artificial intelligence-based system of claim 40, wherein the pathogenicity prediction logic is further configured to generate a pathogenicity prediction for the alternative base dependent on the first delta subsequence and the second delta subsequence.
42. 41. The artificial intelligence-based system of claim 40, wherein the pathogenicity prediction logic is further configured to accumulate the first position-wise subsequence differences into a first cumulative subsequence value and to accumulate the second position-wise subsequence differences into a second cumulative subsequence value.
43. 43. The artificial intelligence-based system of claim 42, wherein the first cumulative subsequence value is an average of the subsequence differences for each of the first positions, and the second cumulative subsequence value is an average of the subsequence differences for each of the second positions.
44. 43. The artificial intelligence-based system of claim 42, wherein the first cumulative subsequence value is a sum of subsequence differences for each of the first positions, and the second cumulative subsequence value is a sum of subsequence differences for each of the second positions.
45. 43. The artificial intelligence-based system of claim 42, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on the first cumulative subsequence value and the second cumulative subsequence value.
46. 46. The artificial intelligence-based system of claim 45, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on an average of the first cumulative subsequence value and the second cumulative subsequence value.
47. 46. The artificial intelligence-based system of claim 45, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on a sum of the first cumulative subsequence value and the second cumulative subsequence value.
48. 39. The artificial intelligence-based system of claim 38, wherein the pathogenicity prediction logic is further configured to classify positions within the first delta subsequence as belonging to a conserved state or a non-conserved state based on the first position-by-position subsequence differences.
49. 49. The artificial intelligence-based system of claim 48, wherein the pathogenicity prediction logic is further configured to classify positions in the second delta subsequence that match positions in the first delta subsequence classified as belonging to the conserved state as belonging to a signal state, and to classify positions in the second delta subsequence that match positions in the first delta subsequence classified as belonging to the non-conserved state as belonging to a noise state.
50. the pathogenicity prediction logic is further configured to accumulate a subset of the second positional subsequence differences into a modulated accumulated subsequence value; 50. The artificial intelligence-based system of claim 49, wherein a second positional subsequence difference within the subset of second positional subsequence differences is located at a position within the second delta subsequence that is classified as belonging to the signal state.
51. 51. The artificial intelligence-based system of claim 50, wherein the pathogenicity prediction logic is further configured to generate the pathogenicity prediction for the alternative base dependent on the modulated cumulative subsequence value.
52. 51. The artificial intelligence-based system of claim 50, wherein the modulated cumulative subsequence value is an average of the second position-wise subsequence differences within the subset of second position-wise subsequence differences.
53. 51. The artificial intelligence-based system of claim 50, wherein the modulated cumulative subsequence value is a sum of the subsequence differences per second position within the subset of subsequence differences per second position.
54. The artificial intelligence-based system of claim 1 , wherein the target base sequence is a coding region of a gene.
55. The artificial intelligence-based system of claim 1 , wherein the target base sequence is a non-coding region of a gene.
56. 56. The artificial intelligence-based system of claim 55, wherein the non-coding regions span a transcription start site, a 5 prime untranslated region (UTR), a 3 prime UTR, an enhancer, and a promoter.
57. the alternative base is a singleton variant that occurs in only one outlier individual in a cohort of outlier individuals; 17. The artificial intelligence-based system of claim 16, wherein outlier individuals in the cohort of outlier individuals exhibit extreme levels of gene expression.
58. 57. The artificial intelligence-based system of claim 56, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels.
59. 58. The artificial intelligence-based system of claim 57, wherein the extreme levels of gene expression include over-gene expression and under-gene expression.
60. 57. The artificial intelligence-based system of claim 56, wherein the singleton variant is a chord variant.
61. 57. The artificial intelligence-based system of claim 56, wherein the singleton variant is a non-coding variant.
62. 61. The artificial intelligence-based system of claim 60, wherein the non-coding variant is a promoter variant.
63. 61. The artificial intelligence-based system of claim 60, wherein the non-coding variant is an enhancer variant.
64. the biological mass model has a first set of weights; The artificial intelligence-based system of claim 1 , wherein the biological quantity output generation logic has a second set of weights.
65. During training, the first set of weights of the biomass model are trained from scratch to process the input sequence and generate the alternative representations of the input sequence; 63. The artificial intelligence based system of claim 62, wherein the second set of weights of the biological mass output generation logic is trained end-to-end from scratch using the first set of weights of the biological mass model to process the alternative representations of the input base sequence and generate the plurality of biological mass output sequences.
66. During inference, the biological mass model uses the first set of trained weights; 64. The artificial intelligence based system of claim 63, wherein during said inference, said biological quantity output generation logic uses said trained second set of weights.
67. the gene expression model has a third set of weights; The artificial intelligence-based system of claim 1 , wherein the gene expression output generation logic has a fourth set of weights.
68. the third set of weights of the gene expression model is trained from scratch to process the plurality of biological quantity output sequences and generate the alternative representations of the plurality of biological quantity output sequences; 66. The artificial intelligence based system of claim 65, wherein the fourth set of weights of the gene expression output generation logic is trained end-to-end from scratch using the third set of weights of the gene expression model to process the alternative representations of the plurality of biological abundance output sequences and generate the gene expression output sequences.
69. during inference, the gene expression model uses the third set of trained weights; 67. The artificial intelligence based system of claim 66, wherein during said inference, said gene expression output generation logic uses said trained fourth set of weights.
70. During training, the first set of weights of the biomass model are first trained from scratch to process the input base sequence and generate the alternative representation of the input base sequence, and then retrained as a substitute for the third set of weights of the gene expression model to process the plurality of biomass output sequences and generate the alternative representation of the plurality of biomass output sequences; 66. The artificial intelligence based system of claim 65, wherein the fourth set of weights of the gene expression output generation logic is trained end-to-end from scratch using the first set of trained weights substituted in the gene expression model to process the alternative representations of the plurality of biological quantity output sequences generated by the first set of trained weights substituted in the gene expression model and generate the gene expression output sequences.
71. During inference, the biological mass model uses the retrained first set of weights; during the inference, the biological quantity output generation logic uses the second set of trained weights; during inference, the gene expression model uses the retrained first set of weights; 69. The artificial intelligence based system of claim 68, wherein during said inference, said gene expression output generation logic uses said fourth set of trained weights.
72. During training, the first set of weights of the biomass model are first trained from scratch to process the input sequence and generate the alternative representations of the input sequence; during the training, the second set of weights of the biomass output generation logic is first trained end-to-end from scratch using the first set of weights of the biomass model to process the alternative representations of the input base sequence and generate the plurality of biomass output sequences; During the training, the first set of trained weights of the biomass model are then retrained to process the reference base sequence and generate the alternative representations of the reference base sequence, and to process the alternative base sequences and generate the alternative representations of the alternative base sequences; 18. The artificial intelligence based system of claim 17, wherein during the training, the second set of trained weights of the biomass output generation logic are then retrained end-to-end with the first set of trained weights of the biomass model to process the alternative representations of the reference base sequence and generate the plurality of reference biomass output sequences, and to process the alternative representations of the alternative base sequences and generate the plurality of alternative biomass output sequences.
73. During inference, the biological mass model uses the retrained first set of weights; 71. The artificial intelligence based system of claim 70, wherein during said inference, said biological quantity output generation logic uses said retrained second set of weights.
74. 24. The artificial intelligence-based system of claim 23, wherein the pathogenicity prediction logic has a fifth set of weights.
75. During training, the first set of weights of the biomass model are first trained from scratch to process the input sequence and generate the alternative representations of the input sequence; during the training, the second set of weights of the biomass output generation logic is first trained end-to-end from scratch using the first set of weights of the biomass model to process the alternative representations of the input base sequence and generate the plurality of biomass output sequences; 73. The artificial intelligence based system of claim 72, wherein during the training, the first set of trained weights of the biomass model and the second set of trained weights of the biomass output generation logic are then retrained end-to-end to generate the pathogenicity predictions for the alternative bases.
76. During inference, the biological mass model uses the retrained first set of weights; during the inference, the biological quantity output generation logic uses the retrained second set of weights; 74. The artificial intelligence-based system of claim 73, wherein during said inference, said pathogenicity prediction logic uses said fifth set of trained weights.
77. 20. The artificial intelligence-based system of claim 18, wherein the first reference biological abundance output for each respective base identifies a respective measure of a first reference epigenetic signal level for the respective reference target base at the respective position within the reference target base sequence.
78. 30. The artificial intelligence-based system of claim 29, wherein the second reference biological abundance output for each respective base identifies a respective measure of a second reference epigenetic signal level for the respective reference target base at the respective position within the reference target base sequence.
79. 22. The artificial intelligence based system of claim 21, wherein the first per-each base alternative biological abundance output identifies a respective measure of a first alternative epigenetic signal level for the respective alternative target base at the respective position within the alternative target base sequence.
80. 23. The artificial intelligence based system of claim 22, wherein the second per-each-base alternative biological abundance output identifies a respective measure of a second alternative epigenetic signal level for the respective alternative target base at the respective position within the alternative target base sequence.
81. 2. The artificial intelligence-based system of claim 1, wherein during training, the biomass model and the biomass output generation logic are first trained end-to-end from scratch to translate an analysis of input base sequences into base-by-base evolutionarily conserved chromatin sequences, and then retrained end-to-end to translate an analysis of input base sequences into base-by-base transcription start frequency chromatin sequences.
82. 2. The artificial intelligence-based system of claim 1, wherein during training, the biomass model and the biomass output generation logic are first trained end-to-end from scratch to translate an analysis of input base sequences into base-by-base epigenetic signal-level chromatin sequences, and then retrained end-to-end to translate an analysis of input base sequences into base-by-base evolutionarily conserved chromatin sequences.
83. 2. The artificial intelligence-based system of claim 1, wherein during training, the biological mass model and the biological mass output generation logic are first trained end-to-end from scratch to translate an analysis of input base sequences into base-by-base epigenetic signal level chromatin sequences, and then retrained end-to-end to translate an analysis of input base sequences into base-by-base transcription start frequency chromatin sequences.
84. 2. The artificial intelligence-based system of claim 1, wherein during training, the biomass model and the biomass output generation logic are first trained end-to-end from scratch to translate an analysis of input base sequences into base-by-base epigenetic signal-level chromatin sequences, and then retrained end-to-end to translate an analysis of input base sequences into base-by-base evolutionarily conserved chromatin sequences and base-by-base transcription start frequency chromatin sequences.
85. 10. The artificial intelligence-based system of claim 1, further configured to include a first training set of training input sequences that include variants confounded by multiple epigenetic effects.
86. 84. The artificial intelligence-based system of claim 83, wherein the epigenetic effects in the plurality of epigenetic effects comprise inter-chromosomal effects, intra-genic effects, population structure and ancestry effects, probabilistic estimation of expression residuals (PEER) effects, environmental effects, gender effects, batch effects, genotyping platform effects, and / or library construction protocol effects. PEER stands for "Probabilistic Estimation of Expression Residuals," which is a collection of Bayesian approaches for inferring hidden determinants and their effects from gene expression profiles using factor analysis methods.
87. 84. The artificial intelligence-based system of Claim 83, further configured to include a second training set of training input sequences that includes variants that are not confounded by the plurality of epigenetic effects.
88. 86. The artificial intelligence based system of Claim 85, wherein the variants in the second training set are determined with certainty to alter gene expression and cause extreme levels of gene expression.
89. 87. The artificial intelligence-based system of Claim 86, wherein the variants in the second training set include variants that cause overexpression, which increases gene expression levels.
90. 87. The artificial intelligence-based system of Claim 86, wherein the variants in the second training set include variants that cause underexpression, which reduces gene expression levels.
91. 88. The artificial intelligence based system of Claim 87, wherein said second training set identifies an overexpression probability for said variant that identifies a likelihood of causal gene overexpression.
92. 89. The artificial intelligence based system of Claim 88, wherein the second training set identifies an underexpression probability for the variant that identifies the likelihood that the variant causes underexpression of the gene.
93. each variant in the second training set is a singleton variant that occurs in only one outlier individual among a cohort of outlier individuals; 87. The artificial intelligence-based system of claim 86, wherein outlier individuals in said cohort of outlier individuals exhibit extreme levels of gene expression.
94. 92. The artificial intelligence-based system of claim 91, wherein the extreme levels of gene expression are determined from tail quantiles of normalized gene expression levels.
95. 92. The artificial intelligence-based system of claim 91, wherein the extreme levels of gene expression include over-gene expression and under-gene expression.
96. 92. The artificial intelligence-based system of claim 91, wherein the singleton variant is a code variant.
97. 92. The artificial intelligence-based system of claim 91, wherein the singleton variant is a non-coding variant.
98. 96. The artificial intelligence-based system of claim 95, wherein the non-coding variant is a 5 prime untranslated region (UTR) variant, a 3 prime UTR variant, an enhancer variant, or a promoter variant.
99. 86. The artificial intelligence-based system of claim 85, wherein the variants in the second training set span multiple tissue types.
100. 86. The artificial intelligence-based system of Claim 85, wherein the variants in the second training set span multiple cell types.
101. The artificial intelligence-based system of claim 1 , wherein the input base sequence and the plurality of biological abundance output sequences span the plurality of tissue types.
102. The artificial intelligence-based system of claim 1 , wherein the input base sequence and the plurality of biological abundance output sequences span the plurality of cell types.
103. The artificial intelligence-based system of claim 10 , wherein the gene expression output sequences span the multiple tissue types.
104. The artificial intelligence-based system of claim 10 , wherein the gene expression output sequences span the multiple cell types.
105. 2. The artificial intelligence-based system of claim 1, wherein the biological mass model and the biological mass output generation logic are first trained end-to-end on the first training set and then retrained on the second training set.
106. 10. The artificial intelligence-based system of claim 1, wherein the variants in the second training set are used as a pathogenic set labeled with a first ground truth label indicating gene expression change, and common variants are used as a benign set labeled with a second ground truth label indicating no gene expression change.
107. 105. The artificial intelligence-based system of claim 104, wherein the good set is balanced for trinucleotide context, homopolymer, k-mer, neighborhood GC frequency, and sequencing depth.
108. 105. The artificial intelligence based system of claim 104, wherein based on cutoff probabilities applied to the overexpression and underexpression probabilities, the variants in the second training set are divided into an overexpressed variant training set with a first ground truth label indicating increased gene expression, an overexpressed variant training set with a second ground truth label indicating decreased gene expression, and a neurally expressed variant training set indicating maintained gene expression.
109. 11. The artificial intelligence-based system of claim 10, wherein the gene expression model and the gene expression output generation logic are first trained end-to-end on the first training set and then retrained on the second training set.
110. 2. The artificial intelligence-based system of claim 1, wherein the biological mass model and the biological mass output generation logic are first trained end-to-end on the first training set and then retrained on variants in the second training set that occur on odd-numbered chromosomes.
111. 11. The artificial intelligence-based system of claim 10, wherein the gene expression model and the gene expression output generation logic are first trained end-to-end on the first training set and then retrained on variants in the second training set that occur on odd-numbered chromosomes.
112. 86. The artificial intelligence-based system of claim 85, wherein the variants in the second training set are not used for training, but instead are used as a validation set to evaluate the performance of the trained biological quantity model 124, the trained biological quantity output generation logic, the trained gene expression model, and the trained gene expression output generation logic.
113. 111. The artificial intelligence-based system of claim 110, wherein variants in the second training set that occur on even-numbered chromosomes are used as the validation set.
114. 2. The artificial intelligence-based system of claim 1, wherein the size of the target base sequence is varied during training to account for variations in the offset position of the transcription start site (TSS).
115. 1. An artificial intelligence based system for detecting changes in gene expression at base resolution, comprising: input generation logic that accesses a sequence database and generates an input sequence, the input sequence including a target sequence, the target sequence flanked by a right sequence having downstream context bases and a left sequence having upstream context bases; a biological mass model that processes the input sequence and generates alternative representations of the input sequence; and biological mass output generation logic for processing the alternative representations of the input base sequence to generate a plurality of biological mass output sequences; An artificial intelligence based system, wherein each biological mass output sequence in the plurality of biological mass output sequences includes a respective base-by-base biological mass output for each target base in the target base sequence.
116. a first biological mass output sequence in the plurality of biological mass output sequences comprising a first respective base-by-base biological mass output for each target base in the target base sequence; 114. The artificial intelligence-based system of claim 113, wherein the first respective base biological abundance output identifies a respective measure of evolutionary conservation of the respective target base across multiple species.
117. a second biological mass output sequence in the plurality of biological mass output sequences comprising a second respective base-by-base biological mass output for each target base in the target base sequence; 114. The artificial intelligence-based system of claim 113, wherein the second per-base biological abundance output identifies a respective measure of transcription initiation of the respective target base at a respective position within the target base sequence.