Artificial intelligence-based epigenetics
Through deep convolutional neural network technology, combined with void convolution, residual connection and batch normalization, a pathogenicity classifier was constructed, which solved the problem of explaining the functional consequences of non-coding variants and achieved efficient pathogenicity prediction and model performance improvement with limited labeled data.
Patent Information
- Application Number
- CN202080065003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2020-09-18
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Existing technologies struggle to effectively interpret the functional consequences of non-coding variants, especially when biases exist in labeled datasets, resulting in poor performance of pathogenicity classification models in clinical applications.
A deep convolutional neural network is used, combined with techniques such as void convolution, residual connection, batch normalization and global average pooling, to construct a pathogenicity classifier, which is trained on large-scale datasets to identify the pathogenicity of non-coding variants.
It achieves accurate prediction of the pathogenicity of non-coding variants in the presence of limited labeled data, reduces human determination bias, and improves the classification accuracy and generalization ability of the model in clinical practice.
Smart Images

Figure CN114402393B_ABST
Abstract
Description
[0001] Priority data
[0002] This patent application claims priority to, or the benefit of, U.S. Provisional Patent Application No. 62 / 903,700, filed September 20, 2019, entitled “ARTIFICIAL INTELLIGENCE-BASED EPIGENETICS” (Attorney Docket No. ILLM1025-1 / IP-1898-PRV).
[0003] Literature Incorporation
[0004] The following documents are incorporated by reference as if fully set out herein, for all purposes:
[0005] PCT Patent Application No. PCT / US18 / 55915 (Attorney Docket No. ILLM 1001-7 / IP-1610-PCT) entitled “Deep Learning-Based Splice Site Classification” to Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, filed October 15, 2018, subsequently published as PCT Publication No. WO 2019 / 079198;
[0006] PCT Patent Application No. PCT / US18 / 55919 (Attorney Docket No. ILLM 1001-8 / IP-1614-PCT) entitled “Deep Learning-Based Aberrant Splicing Detection” to Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, filed October 15, 2018, subsequently published as PCT Publication No. WO 2019 / 079200;
[0007] PCT Patent Application No. PCT / US18 / 55923, entitled “Aberrant Splicing Detection Using Convolutional Neural Networks (CNNs)” by Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, filed on October 15, 2018 (Attorney Docket No. ILLM1001-9 / IP-1615-PCT), subsequently published as PCT Publication No. WO 2019 / 079202;
[0008] U.S. patent application Ser. No. 16 / 160,978, filed October 15, 2018, by Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, titled “Deep Learning-Based SpliceSite Classification” (Attorney Docket No. ILLM 1001-4 / IP-1610-US);
[0009] U.S. patent application Ser. No. 16 / 160,980, filed October 15, 2018, by Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, titled “Deep Learning-Based Aberrant Splicing Detection” (Attorney Docket No. ILLM 1001-5 / IP-1614-US);
[0010] U.S. patent application Ser. No. 16 / 160,984, filed on October 15, 2018, by Kishore Jaganathan, Kai-How Farh, Sofia Kyriazopoulou Panagiotopoulou, and Jeremy Francis McRae, titled “Aberrant Splicing Detection Using Convolutional Neural Networks (CNNs)” (Attorney Docket No. ILLM 1001-6 / IP-1615-US);
[0011] Document 1—S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WAVENET: AGENERATIVE MODEL FOR RAW AUDIO,” arXiv:1609.03499, 2016;
[0012] Document 2—S. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “DEEP VOICE: REAL-TIMENEURAL TEXT-TO-SPEECH,” arXiv:1702.07825, 2017;
[0013] Document 3 — F. Yu and V. Koltun, “Multi-scale context aggregation by diluted conversations,” arXiv:1511.07122, 2016;
[0014] Document 4 — K. He, X. Zhang, S. Ren, and J. Sun, “DEEP RESIDUAL LEARNING FOR IMAGE RECOGNITION,” arXiv:1512.03385, 2015.
[0015] Document 5 — R.K. Srivastava, K. Greff, and J. Schmidhuber, “HIGHWAY NETWORKS,” arXiv:1505.00387, 2015;
[0016] Document 6 — G. Huang, Z. Liu, L. van der Maaten, and K.Q. Weinberger, “DENSELY CONNECTED CONVOLUTIONAL NETWORKS,” arXiv:1608.06993, 2017;
[0017] Document 7—C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” arXiv:1409.4842, 2014.
[0018] Document 8 — S. Ioffe and C. Szegedy, “BATCH NORMALIZATION: ACCELERATING DEEPNETWORK TRAINING BY REDUCING INTERNAL COVARIATE SHIFT,” arXiv:1502.03167, 2015.
[0019] Document 9—JM Volterink, T. Leiner, MA Viergever, and I. "DILATEDCONVOLUTIONAL NEURAL NETWORKS FOR CARDIOVASCULAR MR SEGMENTATION INCONGENITAL HEART DISEASE", arXiv:1704.03669, 2017;
[0020] Document 10—LC Piqueras, “AUTOREGRESSIVE MODEL BASED ON ADEEPCONVOLUTIONAL NEURAL NETWORK FOR AUDIO GENERATION,” Tampere University of Technology, 2016;
[0021] Document 11—J. Wu, “Introduction to Convolutional Neural Networks,” Nanjing University, 2017.
[0022] Document 12—I.J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio, “Convolutional Networks,” Deep Learning, MIT Press, 2016; and
[0023] Document 13 — J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, and G. Wang, “Recent Advances in Convolutional Neural Networks,” arXiv:1512.07108, 2017.
[0024] Document 1 describes a deep convolutional neural network architecture that uses residual block groups with convolution filters having the same convolution window size, batch normalization layers, rectified linear unit (ReLU) layers, dimension change layers, dilated convolution layers with exponentially increasing dilated convolution rates, skip connections, and softmax classification layers to accept an input sequence and produce an output sequence that scores the entries in the input sequence. The disclosed technology uses the neural network components and parameters described in Document 1. In one specific implementation, the disclosed technology modifies the parameters of the neural network components described in Document 1. For example, unlike Document 1, the dilated convolution rate in the disclosed technology progresses non-exponentially from a lower residual block group to a higher residual block group. In another example, unlike Document 1, the convolution window size in the disclosed technology differs between residual block groups.
[0025] Document 2 describes the details of the deep convolutional neural network architecture described in Document 1.
[0026] Document 3 describes the dilated convolution used by the disclosed technology. As used herein, dilated convolution is also referred to as "expanded convolution". Dilated / expanded convolution allows for a large receptive field with few trainable parameters. Dilated / expanded convolution is a convolution in which a kernel is applied over an area larger than the kernel length by skipping input values with a specific step size (also referred to as the dilated convolution rate or dilation factor). Dilated / expanded convolution adds spacing between elements of the convolution filter / kernel so that larger spacings of adjacent input entries (e.g., nucleotides, amino acids) are considered when performing the convolution operation. This enables the incorporation of long-range contextual dependencies in the input. Dilated convolution saves part of the convolution computation for reuse when processing adjacent nucleotides.
[0027] Document 4 describes the residual blocks and residual connections used by the disclosed technology.
[0028] Document 5 describes the hop connections used by the disclosed technology. As used herein, hop connections are also referred to as "highway networks."
[0029] Document 6 describes the densely connected convolutional network architecture used by the disclosed technology.
[0030] Document 7 describes the dimension-changing convolution layer and the module-based processing pipeline used by the disclosed technology. An example of a dimension-changing convolution is a 1×1 convolution.
[0031] Document 8 describes the batch normalization layer used by the disclosed technology.
[0032] Document 9 also describes the dilated / atrous convolutions used by the disclosed technique.
[0033] Document 10 describes various architectures of deep neural networks that can be used by the disclosed technology, including convolutional neural networks, deep convolutional neural networks, and deep convolutional neural networks with dilated / atrous convolutions.
[0034] Document 11 describes details of convolutional neural networks that can be used by the disclosed technology, including algorithms for training convolutional neural networks with subsampling layers (e.g., pooling) and fully connected layers.
[0035] Document 12 describes the details of various convolution operations that can be used by the disclosed technology.
[0036] Document 13 describes various architectures of convolutional neural networks that can be used by the disclosed technology. Technical Field
[0037] The technology disclosed herein relates to artificial intelligence-type computers and digital data processing systems, as well as corresponding data processing methods and products for simulating intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems). The technology also includes systems for inferring uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. Specifically, the technology disclosed herein relates to using deep learning-based techniques to train deep convolutional neural networks. Background Art
[0038] The subject matter discussed in this section should not be considered prior art simply because it is mentioned in this section. Similarly, problems mentioned in this section or associated with the subject matter provided as background technology should not be considered to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to specific implementations of the claimed technology.
[0039] Machine Learning
[0040] In machine learning, input variables are used to predict output variables. Input variables are usually called features and are represented by X = (X1, X2, ..., X k ) means that each X i(i∈1,...,k) are features. The output variable is often called the response or dependent variable and is represented by the variable Y i The relationship between Y and its corresponding X can be written in the general form:
[0041] Y=f(X)+∈
[0042] In the above equation, f is the feature (X1,X2,...,X k ), and ∈ is a random error term. The error term is independent of X and has a mean of zero.
[0043] In practice, feature X is available without having Y or knowing the exact relationship between X and Y. Since the error term has zero mean, the goal is to estimate f.
[0044]
[0045] In the above equation, is an estimate of ∈, which is usually considered a black box, meaning that only The relationship between the input and output of is known, but the question of why it works remains unanswered.
[0046] Use learning to find functions Supervised learning and unsupervised learning are two approaches used in machine learning for this task. In supervised learning, labeled data is used for training. By showing the input and the corresponding output (= label), the function It is optimized so that it approximates the output. In unsupervised learning, the goal is to find hidden structure from unlabeled data. The algorithm has no measure of accuracy on the input data, which distinguishes it from supervised learning.
[0047] Neural Networks
[0048] Single-layer perception (SLP) is the simplest model of a neural network. It consists of an input layer and an activation function, such as Figure 1 The input passes through a weighted graph. The function f uses the sum of the inputs as an argument and compares this to a threshold θ.
[0049] Figure 2 One specific implementation of a fully connected neural network with multiple layers is shown. A neural network is a system of interconnected artificial neurons (e.g., a1, a2, a3) that exchange messages between each other. The neural network is shown with three inputs, two neurons in the hidden layer, and two neurons in the output layer. The hidden layer has an activation function f(·) and the output layer has an activation function g(·). The connections have digital weights (e.g., w) that are tuned during the training process. 11 、w21 、w 12 、w 31 、w 22 、w 32 、v 11 、v 22 ) so that when fed an image for recognition, a properly trained network responds correctly. The input layer processes the raw input, and the hidden layer processes the output from the input layer based on the weights of the connections between the input layer and the hidden layer. The output layer obtains the output from the hidden layer and processes it based on the weights of the connections between the hidden layer and the output layer. The network includes multiple layers of feature detection neurons. Each layer has many neurons that respond to different combinations of inputs from the previous layer. The layers are constructed so that the first layer detects a set of primitive patterns in the input image data, the second layer detects patterns of patterns, and the third layer detects patterns of those patterns.
[0050] Genetic variation can help explain many diseases. Each person has a unique genetic code, and many genetic variants exist within a population. Natural selection has eliminated most harmful genetic variants from the genome. Identifying which genetic variants are likely pathogenic or harmful is crucial. This will help researchers focus on potentially pathogenic genetic variants and accelerate the diagnosis and treatment of many diseases.
[0051] Modeling the properties and functional effects (e.g., pathogenicity) of variants is an important but challenging task in the field of genomics. Despite the rapid progress of functional genomic sequencing technology, the interpretation of the functional consequences of non-coding variants remains a huge challenge due to the complexity of cell type-specific transcriptional regulatory systems. In addition, a limited number of non-coding variants have been functionally verified experimentally.
[0052] Previous efforts to explain genomic variants have primarily focused on variants in coding regions. However, noncoding variants also play an important role in complex diseases. Identifying pathogenic functional noncoding variants from a large population of neutral noncoding variants may be important in genotype-phenotype relationship studies and precision medicine.
[0053] Furthermore, a large proportion of known pathogenic noncoding variants reside in promoter regions or conserved sites, leading to a positive diagnosis bias in the training set, as simple or obvious cases with known pathogenicity trends may be enriched in the labeled dataset relative to the entire population of pathogenic noncoding variants. If not addressed, this bias in labeled pathogenicity data will lead to unrealistic model performance, as the model can achieve relatively high test-set performance simply by predicting all core variants as pathogenic and all other variants as benign. However, in the clinic, such models will incorrectly classify pathogenic, non-core variants as benign at an unacceptably high rate.
[0054] Advances in biochemical techniques over the past few decades have produced next generation sequencing (NGS) platforms that rapidly produce genomic data at a much lower cost than ever before. Sequencing DNA at such overwhelming volumes remains difficult to annotate. Supervised machine learning algorithms often perform well when large amounts of labeled data are available. In bioinformatics and many other data-rich disciplines, the process of labeling examples is expensive; however, unlabeled examples are inexpensive and readily available. For cases in which the amount of labeled data is relatively small and the amount of unlabeled data is substantially larger, semi-supervised learning represents a cost-effective alternative to manual labeling.
[0055] An opportunity arises to construct a deep learning-based pathogenicity classifier that accurately predicts the pathogenicity of non-coding variants. A database of pathogenic non-coding variants free of human determination bias can be produced. BRIEF DESCRIPTION OF DRAWINGS
[0056] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. The color drawing(s) can also be obtained from the Office's website.
[0057] In the drawings, like reference numerals in all different views generally refer to like parts. Further, the drawings are not necessarily to scale, emphasis instead being placed on illustrating the principles of the disclosed technology. In the following description, various specific embodiments of the disclosed technology are described, with reference to the following drawings, in which:
[0058] Figure 1 A single layer perception (SLP) is shown.
[0059] Figure 2 One embodiment of a feedforward neural network with multiple layers is shown.
[0060] Figure 3 One embodiment of the workings of a convolutional neural network is depicted.
[0061] Figure 4 A block diagram of training a convolutional neural network is depicted according to one embodiment of the disclosed technology.
[0062] Figure 5 One embodiment of a ReLU nonlinearity layer is shown according to one embodiment of the disclosed technology.
[0063] Figure 6 Dilated convolutions are shown.
[0064] Figure 7 is one embodiment of a subsampling layer (average / max pooling) according to one embodiment of the disclosed technology.
[0065] Figure 8 Depicting a concrete implementation of two convolution layers in a convolutional layer.
[0066] Figure 9 Depicts a residual connection that reinjects prior information downstream via feature map addition.
[0067] Figure 10 Depicts a specific implementation of residual blocks and skip connections.
[0068] Figure 11 Shown is a specific implementation of stacked dilated convolutions.
[0069] Figure 12 Showing the batch normalization forward pass.
[0070] Figure 13 Showing the batch normalization transformation at test time.
[0071] Figure 14 Showing the batch normalization backward pass.
[0072] Figure 15 Describes the use of batch normalization layers with convolutional or densely connected layers.
[0073] Figure 16 A specific implementation of 1D convolution is shown.
[0074] Figure 17 Shows how global average pooling (GAP) works.
[0075] Figure 18 One specific implementation is shown of a computing environment with training servers and production servers that can be used to implement the disclosed technology.
[0076] Figure 19 Depicts one specific implementation of the architecture of a model that can be used by a sequence-to-sequence epigenetic model and / or pathogenicity determiner.
[0077] Figure 20 One specific implementation of a residual block that can be used by a sequence-to-sequence epigenetic model and / or pathogenicity determiner is shown.
[0078] Figure 21 Another specific implementation of the architecture of a sequence-to-sequence epigenetic model and / or pathogenicity determiner is depicted, referred to herein as "Model80."
[0079] Figure 22 Another specific implementation of the architecture of a sequence-to-sequence epigenetic model and / or pathogenicity determiner is depicted, referred to herein as "Model400."
[0080] Figure 23 Another specific implementation of an architecture depicting a sequence-to-sequence epigenetic model and / or pathogenicity determiner is referred to herein as "Model2000."
[0081] Figure 24 Another specific implementation of the architecture of a sequence-to-sequence epigenetic model and / or pathogenicity determiner is depicted, referred to herein as "Model 10000."
[0082] Figure 25 A one-hot encoder is shown.
[0083] Figure 25B Example promoter sequences for genes are shown.
[0084] Figure 26 An input preparation module is shown that accesses a sequence database and generates an input base sequence.
[0085] Figure 27 An example of a sequence-to-sequence epigenetic model is shown.
[0086] Figure 28 Shows the per-track processor for the sequence-to-sequence model.
[0087] Figure 29 Reference and alternative sequences are shown.
[0088] Figure 30 One specific implementation of using the sequence-to-sequence model 2700 to generate position-by-position comparison results is shown.
[0089] Figure 31 An example pathogenicity classifier is shown.
[0090] Figure 32 Depicted is one implementation of training a pathogenicity classifier using training data comprising a set of pathogenic non-coding variants annotated with a pathogenic label (eg, "1") and a set of benign non-coding variants annotated with a benign label (eg, "0").
[0091] Figure 33 is a simplified block diagram of a computer system that can be used to implement the disclosed technology.
[0092] Figure 34 One specific implementation of how to generate a panel of pathogenic noncoding variants is shown.
[0093] Figure 35 Describes how a training dataset is generated for the disclosed techniques. DETAILED DESCRIPTION
[0094] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed implementations will be apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Therefore, the disclosed technology is not intended to be limited to the specific implementations shown but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0095] Residual Connection
[0096] Figure 9 Depicting residual connections that reinject prior information downstream via feature map addition. Residual connections involve reinjecting prior representations into the downstream data flow by adding past output tensors to later output tensors, which helps prevent information loss along the data processing stream. Residual connections address two common problems that plague any large-scale deep learning model: vanishing gradients and representation bottlenecks. In general, adding residual connections to any model with more than 10 layers can be beneficial. As discussed above, residual connections involve using the output of an earlier layer as input to a later layer, effectively creating shortcuts in a sequential network. Instead of being concatenated to the later activation, the earlier output is summed with the later activation, assuming that the two activations are of the same size. If they are of different sizes, a linear transformation can be used to reshape the earlier activation to the target shape.
[0097] Residual learning and skip connections
[0098] Figure 10 Depicts a specific implementation of residual blocks and skip connections. The key concept of residual learning is that residual mappings are easier to learn than the original mappings. Residual networks stack multiple residual units to mitigate the degradation of training accuracy. Residual blocks utilize special additive skip connections to mitigate vanishing gradients in deep neural networks. At the beginning of a residual block, the data flow is split into two streams: the first stream carries the unchanged input of the block, while the second stream applies the weights and nonlinearity. At the end of the block, the two streams are merged using an element-wise summation. The main advantage of this construction is that it allows gradients to flow more easily through the network.
[0099] Thanks to residual networks, deep convolutional neural networks (CNNs) can be easily trained and have achieved improved accuracy in image classification and object detection. th The output of the layer is connected as input to (l+1) th layer, which produces the following layer transformation: x l =H l (x l-1). The residual block adds a skip connection that bypasses the nonlinear transformation using the identity function: x l =H l (x l-1 )+x l-1 The advantage of the residual block is that the gradient can flow directly from later layers to earlier layers through the identity function. However, the identity function and H l The outputs of are combined by summing, which may block the flow of information in the network.
[0100] Wave Network
[0101] WaveNet is a deep neural network used to generate raw audio waveforms. WaveNet distinguishes itself from other convolutional networks because it can take a relatively large "field of view" at a low cost. Furthermore, it can condition the signal both locally and globally, allowing WaveNet to be used as a text-to-speech (TTS) engine for multiple voices, with both local conditioning applied to the TTS and global conditioning applied to the specific voice.
[0102] The main building block of WaveNet is the causal dilated convolution. As an extension of the causal dilated convolution, WaveNet also allows stacking of these convolutions, such as Figure 11 As shown in the figure. To obtain the same receptive field with dilated convolutions, another dilated layer is needed. Stacking is a repetition of dilated convolutions, connecting the outputs of the dilated convolutional layers to a single output. This allows wave networks to obtain a large "visual" field of view for one output node at a relatively low computational cost. For comparison, to obtain a field of view of 512 inputs, a fully convolutional network (FCN) would require 511 layers. In the case of a dilated convolutional network, we would need eight layers. Stacked dilated convolutions only need to have seven layers of two stacks or six layers of four stacks. To get an idea of the difference in computational power required to cover the same field of view, the number of weights required in the network is shown below, assuming one filter per layer and a filter width of two. Furthermore, it is assumed that the network uses 8-bit binary encoding.
[0103]
[0104] WaveNet adds skip connections before making residual connections, which bypass all following residual blocks. Each of these skip connections is summed before passing through a series of activation functions and convolutions. Intuitively, this is the sum of the information extracted in each layer.
[0105] Batch Normalization
[0106] Batch Normalization is a method that accelerates the training of deep networks by normalizing the data over the integral part of the network architecture. Batch Normalization adaptively normalizes the data as the mean and variance change over time during training. It works by internally maintaining an exponential moving average of the batch-wise mean and variance of the data seen during training. The primary benefit of Batch Normalization is that it aids gradient propagation—much like residual connections—and thus enables deep networks. Some very deep networks can be trained only if they include multiple Batch Normalization layers.
[0107] Batch Normalization can be viewed as another layer that can be inserted into the model architecture, just like a fully connected or convolutional layer. Batch Normalization layers are typically used after convolutional or densely connected layers. They can also be used before convolutional or densely connected layers. Two specific implementations can be used with the disclosed technology and in Figure 15 As shown. Batch Normalization layers take an axis argument that specifies the feature axis along which normalization should be performed. This argument defaults to -1, the last axis in the input tensor. This is the correct value when using Dense, Conv1D, RNN, and Conv2D layers with data_format set to "channels_last". However, in the niche use of Conv2D layers with data_format set to "channels_first", the feature axis is axis 1; the axis argument in Batch Normalization can be set to 1.
[0108] Batch normalization provides a definition of the gradients with respect to the feedforward input and calculates them with respect to the parameters and its own input via the backward pass. In practice, a batch normalization layer is inserted after the convolutional or fully connected layer, but before the output is fed into the activation function. For convolutional layers, different elements (i.e., activations) of the same feature map at different positions are normalized in the same way to follow the convolutional properties. Therefore, all activations in the mini-batch are normalized at all positions, rather than per activation.
[0109] Internal covariance shift is the main reason why deep architectures are notoriously slow to train. This stems from the fact that deep networks must not only learn new representations at each layer, but also take into account changes in their distribution.
[0110] Covariance shift is a known problem in deep learning and frequently occurs in real-world problems. A common covariance shift problem is the difference in the distribution of the training and test sets, which can lead to suboptimal generative performance. This problem is often addressed with a normalization or whitening preprocessing step. However, whitening, in particular, is computationally expensive and therefore impractical in online settings, especially when covariance shift occurs across different layers.
[0111] Internal covariance shift is a phenomenon in which the distribution of network activations changes across layers due to changes in network parameters during training. Ideally, each layer should be transformed into a space where they have the same distribution but functional relationships remain unchanged. To avoid the expensive computation of the covariance matrix to decorrelate and whiten the data at each layer and step, we normalize the distribution of each input feature in each layer across each mini-batch to have zero mean and standard deviation 1.
[0112] Forward pass
[0113] During the forward pass, the mini-batch mean and variance are calculated. Using these mini-batch statistics, the data is normalized by subtracting the mean and dividing by the standard deviation. Finally, the data is scaled and offset using the learned scale and offset parameters. Figure 12 Depict the batch normalization forward pass f BN .
[0114] exist Figure 12 , respectively, μ β is the batch average and is the batch variance. The learned scale and bias parameters are denoted by γ and β, respectively. For clarity, this paper describes the batch normalization procedure for each activation and omits the corresponding index.
[0115] Since normalization is a differentiable transformation, the error is propagated into these learned parameters and it is therefore possible to restore the representational power of the network by learning the identity transformation. Conversely, by learning the same scaling and offset parameters as the corresponding batch statistics, the batch normalization transformation will have no effect on the network if it is the best operation to perform. At test time, the batch mean and variance are replaced by the corresponding population statistics since the input does not depend on the other samples from the mini-batch. Another approach is to maintain a running average of the batch statistics during training and use the running average to compute the network output at test time. At test time, the batch normalization transformation can be as follows Figure 13 The expression shown. Figure 13 ,μ D as well as These represent the population mean and variance, respectively, rather than batch statistics.
[0116] Backward pass
[0117] Since normalization is a differentiable operation, the backward pass can be computed as Figure 14 Depicted.
[0118] 1D convolution
[0119] 1D convolution extracts local 1D patches or subsequences from a sequence, such as Figure 16 As shown in Figure 2. 1D convolution obtains each output time step from a temporal patch in the input sequence. 1D convolution layers recognize local patterns in the sequence. Because the same input transformation is performed on each patch, patterns learned at a specific position in the input sequence can be recognized later at different positions, making 1D convolution layer translations invariant to temporal translations. For example, a 1D convolution layer that processes a base sequence using a convolution window of size 5 should be able to learn bases or base sequences of length 5 or less, and it should be able to recognize base motifs in any context in the input sequence. Therefore, 1D convolution at the base level is able to learn base morphology.
[0120] Global average pooling
[0121] Figure 17 Shows how global average pooling (GAP) works. Global average pooling can be used to replace fully connected (FC) layers for classification by taking the spatial average of the features in the last layer for scoring. Global average pooling reduces training load and avoids overfitting problems. Global average pooling applies a structure before the model and is equivalent to a linear transformation with predefined weights. Global average pooling reduces the number of parameters and eliminates fully connected layers. Fully connected layers are typically the most parameter- and connection-intensive layers, and global average pooling provides a much lower-cost way to achieve similar results. The main concept of global average pooling is to generate an average from each last layer feature map as a confidence factor for scoring that is fed directly into the softmax layer.
[0122] Global average pooling has three benefits: (1) there are no extra parameters in the global average pooling layer, so overfitting is avoided at the global average pooling layer; (2) since the output of global average pooling is the average of the entire feature map, global average pooling will be more robust to spatial translation; and (3) since a large number of parameters in the fully connected layers usually account for more than 50% of all the parameters in the entire network, replacing them by global average pooling layers can significantly reduce the size of the model, and this makes global average pooling very useful in model compression.
[0123] Global average pooling makes sense because stronger features in the last layer are expected to have higher averages. In some implementations, global average pooling can be used as a proxy for classification scores. Feature maps under global average pooling can be interpreted as confidence maps and facilitate correspondence between feature maps and categories. Global average pooling can be particularly effective if the last layer features are abstract enough to be directly classified; however, if multi-level features should be combined into groups, such as component models (this is best performed by adding a simple fully connected layer or other classifier after global average pooling), then global average pooling alone is not enough.
[0124] Deep Learning in Genomics
[0125] Genetic variation can help explain many diseases. Each person has a unique genetic code, and many genetic variants exist within a population. Natural selection has eliminated most harmful genetic variants from the genome. Identifying which genetic variants are likely pathogenic or harmful is crucial. This will help researchers focus on potentially pathogenic genetic variants and accelerate the diagnosis and treatment of many diseases.
[0126] Modeling the properties and functional effects (e.g., pathogenicity) of variants is an important but challenging task in the field of genomics. Despite the rapid progress of functional genomic sequencing technologies, the interpretation of the functional consequences of variants remains a huge challenge due to the complexity of cell type-specific transcriptional regulatory systems.
[0127] Regarding pathogenicity classifiers, deep neural networks are a type of artificial neural network that continuously models high-level features using multiple layers of nonlinear and complex transformations. Deep neural networks provide feedback via backpropagation, which carries the difference between observed and predicted outputs to adjust parameters. Deep neural networks have evolved with the availability of large training datasets, the power of parallel and distributed computing, and sophisticated training algorithms. They have driven significant advances in many fields, such as computer vision, speech recognition, and natural language processing.
[0128] Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are components of deep neural networks. Convolutional neural networks are particularly successful in image recognition with an architecture comprising convolutional layers, nonlinear layers, and pooling layers. Recurrent neural networks are designed to exploit the sequential information of input data and have cyclic connections between building blocks such as perceptrons, long short-term memory units, and gated recurrent units. In addition, many other emerging deep neural networks for limited contexts have been proposed, such as deep spatiotemporal neural networks, multidimensional recurrent neural networks, and convolutional autoencoders.
[0129] The goal of training a deep neural network is to optimize the weight parameters in each layer, gradually combining simpler features into complex ones, so that the most appropriate layered representation can be learned from the data. A single cycle of the optimization process proceeds as follows. First, given a training dataset, a forward pass sequentially computes the output of each layer and propagates the function signal forward through the network. In the final output layer, the target loss function measures the error between the inferred output and the given label. To minimize the training error, a backward pass uses the chain rule to backpropagate the error signal and calculate the gradient with respect to all weights in the entire neural network. Finally, an optimization algorithm based on stochastic gradient descent is used to update the weight parameters. While batch gradient descent performs parameter updates for each complete dataset, stochastic gradient descent provides a stochastic approximation by performing updates for each small data example. Several optimization algorithms are derived from stochastic gradient descent. For example, the Adagrad and Adam training algorithms perform stochastic gradient descent while adaptively modifying the learning rate based on the update frequency and momentum of the gradient of each parameter, respectively.
[0130] Another core element in deep neural network training is regularization, which refers to strategies designed to avoid overfitting and thus achieve good generalization performance. For example, weight decay adds a penalty factor to the objective loss function so that the weight parameters converge to a small absolute value. Dropout randomly removes hidden units from a neural network during training and can be thought of as an ensemble of possible subnetworks. To enhance the capabilities of dropout, new activation functions, maximum output, and a dropout variant of recurrent neural networks (called rnnDrop) have been proposed. In addition, batch normalization provides a new regularization method by normalizing the scalar features of each activation within a mini-batch and learning each mean and variance as parameters.
[0131] Given that sequence data is multidimensional and high-dimensional, deep neural networks have great prospects in bioinformatics research due to their wide applicability and enhanced predictive power. Convolutional neural networks have been used to solve sequence-based problems in genomics, such as motif discovery, pathogenic variant identification, and gene expression inference. Convolutional neural networks use a weight sharing strategy, which is particularly useful for studying DNA because it can capture sequence motifs, which are short and recurring local patterns in DNA that are assumed to have significant biological functions. The hallmark of convolutional neural networks is the use of convolutional filters. Unlike traditional classification methods based on precisely designed features and hand-crafted features, convolutional filters perform adaptive learning of features, similar to the process of mapping raw input data to information representations of knowledge. In this sense, convolutional filters act as a series of motif scanners because a set of such filters can identify relevant patterns in the input and update themselves during the training process. Recurrent neural networks can capture long-range dependencies in sequence data of different lengths (such as protein or DNA sequences).
[0132] Therefore, powerful computational models for predicting the pathogenicity of variants can be of great benefit to both basic science and translational research.
[0133] Specific implementation
[0134] We have described systems, methods, and articles for artificial intelligence-based epigenetics. One or more features of a specific implementation can be combined with the basic specific implementation. Specific implementations that are not mutually exclusive are taught to be combinable. One or more features of a specific implementation can be combined with other specific implementations. The present disclosure periodically reminds users of these options. Omitting the repetition of these options from some specific implementations should not be considered to limit the combinations taught in the preceding sections, and these statements will be incorporated by reference into each of the following specific implementations.
[0135] In one embodiment, the sequence-to-sequence epigenetic model and / or pathogenicity determiner is a convolutional neural network. In another embodiment, the sequence-to-sequence epigenetic model and / or pathogenicity determiner is a recurrent neural network. In another embodiment, the sequence-to-sequence epigenetic model and / or pathogenicity determiner is a residual neural network with residual blocks and residual connections. In other embodiments, the sequence-to-sequence epigenetic model and / or pathogenicity determiner is a combination of a convolutional neural network and a recurrent neural network.
[0136] Those skilled in the art will appreciate that sequence to sequence epigenetic models and / or pathogenicity determiners can use various filling and stride configurations. It can use different output functions (e.g., classification or regression) and may or may not include one or more fully connected layers. It can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, expansion or hole convolution, transposed convolution, depth-separable convolution, point-by-point convolution, 1×1 convolution, grouped convolution, flat convolution, space and cross-channel convolution, shuffled grouped convolution, space-separable convolution and deconvolution. It can use one or more loss functions, such as logistic regression / logarithmic loss function, multi-class cross entropy / softmax loss function, binary cross entropy loss function, mean square error loss function, L1 loss function, L2 loss function, smooth L1 loss function and Huber loss function. It can use any parallelism, efficiency and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel detection of map transformations, batching, prefetching, model parallelism, data parallelism and synchronous / asynchronous SGD. It can include upsampling layers, downsampling layers, recurrent connections, gate and gate memory units (such as LSTM or GRU), residual blocks, residual connections, high-speed connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential lining units (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout layers, pooling layers (e.g., max or average pooling), global average pooling layers and attention mechanisms.
[0137] Promoter sequence
[0138] Figure 25B An example promoter sequence 2501 for a gene is shown. The input to the pathogenicity classifier is a promoter sequence, which is a regulatory region upstream (towards the 5' region) of a gene adjacent to the transcription start site (TSS). They do not encode for protein, but rather provide a starting and control point for regulating gene transcription.
[0139] In a specific implementation, the length of the promoter sequence is 3001 bases. In other specific implementations, the length can be, for example, reduced or increased from 200 to 20,000 bases, or it can be suitable for a specific promoter region (for example, centered around TSS). The promoter sequence flanks the right end context and the left end context extending outside the promoter region, including the gene sequence (for example, 5'UTR region 2502, start codon and stop codon 2503, 3'UTR region 2504, transcription terminator 2505) after the promoter region. The flank context can be 100 to 5000 bases. Typically, upstream and downstream flank contexts are equal, but this is not essential.
[0140] The promoter sequence contains reference bases from one or more reference genome databases. The reference bases are one-hot encoded to save position-specific information for each individual base in the promoter sequence. In one-hot encoding, each reference base is encoded with a four-bit binary vector, where one bit is hot (i.e., 1) and the other bits are off (i.e., 0). For example, as shown, T = (1, 0, 0, 0), G = (0, 1, 0, 0), C = (0, 0, 1, 0), and A = (0, 0, 0, 1). In some implementations, an undetermined base is encoded as N = (0, 0, 0, 0). Figure 25 Figure 25B An example promoter sequence (in yellow) is shown with reference bases represented using one-hot encoding. When the pathogenicity classifier receives the one-hot encoded reference bases as a convolutional neural network, they are able to maintain the spatial positional relationships within the promoter sequence.
[0141] Sequence-to-sequence epigenetic model
[0142] Figure 26 An input preparation module 2610 is shown accessing a sequence database 2608 and generating an input base sequence 2602. The input base sequence 2602 includes (i) a target base sequence having a target base. The target base sequence is flanked by (ii) a right end base sequence having a downstream context base, and (iii) a left end base sequence having an upstream context base. In some implementations, the target base sequence is a promoter sequence 2501.
[0143] Figure 27 An example of a sequence-to-sequence genetic model 2700 is shown.
[0144] The sequence-to-sequence model 2700 processes the input base sequence 2602 and generates a substitute representation 2702 of the input base sequence 2602. An output module processes the substitute representation 2702 of the input base sequence 2602 and produces an output 2704 having at least one per-base output for each target base in the target base sequence, where the per-base output specifies signal levels for a plurality of epigenetic tracks for the corresponding target base. Details of the sequence-to-sequence model 2700 are described in Figure 19 、 Figure 20 、 Figure 21 、 Figure 22 、 Figure 23 and Figure 24 .
[0145] Epigenetic trajectories and their signaling levels
[0146] The plurality of epigenetic tracks includes deoxyribonucleic acid (DNA) methylation alterations (e.g., CpG), histone modifications, non-coding ribonucleic acid (ncRNA) expression, and chromatin structure alterations (e.g., nucleosome positioning). The plurality of epigenetic tracks includes deoxyribonuclease (DNase) tracks. The plurality of epigenetic tracks includes histone 3 lysine 27 acetylation (H3K27ac) tracks. The combination of these epigenetic tracks across different tissues, cell types, and cell lines yields over one thousand different epigenetic tracks, and our sequence-to-sequence model 2700 can produce an output that specifies, for each base in an input base sequence, a signal level for each of the one thousand epigenetic tracks.
[0147] In one implementation, our sequence-to-sequence model 2700 produces an output that specifies, for each base in an input base sequence, a signal level for each of 151 epigenetic tracks. These 151 epigenetic tracks are produced as a cell type and cell line combination of the following epigenetic signals: GM12878 Roadmap tracks (DNase, H2A.Z, H3K27ac, H3K27me3, H3K36me3, H3K4mel, H3K4me2, H3K4me3, H3K79me2, H3K9ac, H3K9me3, H4K20ml).
[0148] For training purposes, the baseline true signal level of the epigenetic track is obtained from sources such as the Roadmap Epigenetic Project (https: / / egg2.wustl.edu / roadmap / web_portal / index.html) and / or ENCODE (https: / / www.encodeproject.org / ). In one embodiment, the epigenetic track is a genome-wide signal coverage track, which is found here at https: / / egg2.wustl.edu / roadmap / web_portal / processed_data.html#ChipSeq_DNaseSeq (which is incorporated herein by reference). In some embodiments, the epigenetic track is a -log10 (p-value) signal track, which is found here at https: / / egg2.wustl.edu / roadmap / data / byFileType / signal / consolidated / macs2signal / pval / (which is incorporated herein by reference). In other specific implementations, the epigenetic track is a fold enrichment signal track found here at https: / / egg2.wustl.edu / roadmap / data / byFileType / signal / consolidated / macs2signal / foldChange / (which is incorporated herein by reference).
[0149] In a specific implementation, we use the signal processing engine of MACSv2.0.10 peak caller to generate genome wide signal coverage tracks (https: / / github.com / taoliu / MACS / , incorporated herein by reference). Whole cell extracts are used as the control for the signal normalization of histone ChIP-seq coverage. Each DNase-seq data set is normalized using a simulated background data set that is generated by evenly distributing an equal number of readings across a mappable genome.
[0150] In a specific implementation, we generate two types of trajectories that use different statistical values based on the Poisson background model to represent the signal score per base. In short, the reads extend the estimated fragment length in the 5' to 3' direction. At each base, the counts of the ChIP-seq / DNaseI-seq extended reads observed to overlap with the base are compared with the corresponding dynamic expected background counts (local) estimated from the control data set. Local is defined as max(BG, 1K, 5K, 10K), where BG is the expected count per base, assuming that the control reads are evenly distributed across all mappable bases in the genome and 1K, 5K, 10K are the expected counts estimated from the 1kb, 5kb, and 10kb windows centered on the base. Local is adjusted for the ratio of the sequencing depth of the ChIP-seq / DNase-seq data set relative to the control data set. The two types of signal score statistics calculated per base are as follows.
[0151] (1) Fold enrichment ratios of ChIP-seq or DNase counts relative to expected background counts. These scores provide a direct measure of the effect size of enrichment at any base in the genome.
[0152] (2) Poisson p-values for ChIP-seq or DNase counts relative to the negative log10 of the expected background counts local. These signal confidence scores provide a measure of the statistical significance of the observed enrichment.
[0153] Additional information on how to use ChIP-seq or DNase and peak calling to measure signal levels such as p-values, fold enrichment values, etc. can be found in Appendix B, which is incorporated in its entirety into priority provisional application No. 62 / 903,700.
[0154] Processor per track
[0155] Figure 28The per-track processors of the sequence-to-sequence model 2700 are shown. In one embodiment, the sequence-to-sequence model 2700 has a plurality of per-track processors 2802, 2812, and 2822 corresponding to respective epigenetic tracks in a plurality of epigenetic tracks. Each per-track processor also includes at least one processing module 2804a, 2814a, and 2824a (e.g., a plurality of residual blocks in each per-track processor) and an output module 2808, 2818, and 2828 (e.g., a linear module, a rectified linear unit (ReLU) module). The sequence-to-sequence model processes an input base sequence and generates an alternative representation 2702 of the input base sequence. The processing module of each per-track processor processes the alternative representation and generates additional alternative representations 2806, 2816, and 2826 that are specific to the particular per-track processor. The output module of each per-track processor processes the additional alternative representation generated by its corresponding processing module and produces as outputs 2810 , 2820 and 2830 the signal level for each base in the corresponding epigenetic track and input base sequence.
[0156] Position-by-position comparison
[0157] Figure 29 Reference sequence 2902 and alternative / variant sequence 2912 are depicted. Figure 30 One specific implementation of using the sequence-to-sequence model 2700 to generate position-by-position comparison results is shown.
[0158] In one embodiment, the input preparation module 2610 accesses a sequence database and generates (i) a reference sequence 2902 containing a base at a target position 2932, wherein the base is flanked by a downstream context base 2942 and an upstream context base 2946; and (ii) an alternative sequence 2912 containing a variant 2922 of the base at the target position, wherein the variant is flanked by a downstream context base and an upstream context base. The sequence-to-sequence model 2700 processes the reference sequence and generates a reference output 3014, wherein the reference output specifies a signal level of multiple epigenetic tracks for each base 2902 in the reference sequence, and processes the alternative sequence 2912 and generates an alternative output 3024, wherein the alternative output 3024 specifies a signal level of multiple epigenetic tracks for each base in the alternative sequence 2912.
[0159] The comparator 3002 applies a position-by-position deterministic comparison function to the reference output 3014 and the alternative output 3024 generated for the bases in the reference sequence 2902 and the alternative sequence 2912, and generates a position-by-position comparison result 3004 based on the difference between the signal level of the reference output 3014 and the signal level of the alternative output 3024 caused by the variant in the alternative sequence 2912.
[0160] The per-position determinism comparison function computes a per-element difference between the reference output 3014 and the alternative output 3024. The per-position determinism comparison function computes a per-element sum of the reference output 3014 and the alternative output 3024. The per-position determinism comparison function computes a per-element ratio between the reference output 3014 and the alternative output 3024.
[0161] Pathogenicity determiner
[0162] The pathogenicity determiner 3100 processes the per-position comparison results 3004 and produces an output that scores the variant in the alternative sequence 2912 as pathogenic or benign. The pathogenicity determiner 3100 processes the per-position comparison results 3004 and generates an alternative representation of the per-position comparison results 3004. The per-position comparison results 3004 are based on differences caused by the variant in the alternative sequence 2912 between signal levels of a plurality of epigenetic tracks determined for a base in the reference sequence and signal levels of the plurality of epigenetic tracks determined for a base in the alternative sequence. An output module (e.g., sigmoid) processes the alternative representation of the per-position comparison results and produces an output that scores the variant in the alternative sequence as pathogenic or benign. In a particular implementation, if the output is above a threshold (above 0.5), the variant is classified as pathogenic.
[0163] Gene expression-based pathogenicity markers
[0164] The pathogenicity determiner 3100 is trained using training data that includes a set of pathogenic non-coding variants 3202 annotated with a pathogenicity label 3312 (e.g., “1”) and a set of benign non-coding variants 3204 annotated with a benignity label 3314 (e.g., “0”);
[0165] Figure 34 One particular implementation is shown for how a set of pathogenic non-coding variants is generated. The set of pathogenic non-coding variants includes a singleton that only appears in a single individual among a group of individuals and the single individual exhibits low expression across multiple tissue / organ tissues 3402 for a gene adjacent to the non-coding variant in the pathogenic set. The set of pathogenic non-coding variants is considered to be an “expression outlier” in Figure 34 .
[0166] The set of benign non-coding variants is a singleton that only appears in a single individual among a group of individuals and the single individual does not exhibit low expression across multiple tissues for a gene adjacent to the non-coding variant in the benign set.
[0167] Underexpression is determined by analyzing the distribution of expression exhibited by a cohort of individuals across each of the plurality of tissues for each of the genes, and calculating a median z-score for a single individual based on the distribution.
[0168] Within each tissue, we calculate the z-score for each individual. That is, if x_i is the value for individual i, m is the mean of the x_i values across all individuals, and s is the standard deviation, then the z-score for individual i is (x_i – m) / s. We then test whether z_i is below a certain threshold (e.g., -1.5 as mentioned in the claims). We also need this to occur for the same individual in multiple tissues (e.g., at least two tissues with z_i < -1.5).
[0169] Additional details regarding gene expression-based pathogenicity markers can be found in Appendix B, which is incorporated in its entirety into priority provisional application No. 62 / 903,700.
[0170] The disclosed system implementations and other systems optionally include one or more of the following features. The systems may also include the features described in conjunction with the disclosed methods. For the sake of brevity, alternative combinations of system features are not individually enumerated. Features applicable to the systems, methods, and articles of manufacture are not repeated for each set of essential features of the statutory classification. The reader will understand how the features identified in this section can be easily combined with essential features in other statutory classifications.
[0171] Other implementations may include a non-transitory computer-readable storage medium storing instructions that are executable by a processor to perform the actions of the above-described system. Another implementation may include a method of performing the actions of the above-described system.
[0172] Additional details regarding the disclosed technology can be found in Appendix A, which is incorporated in its entirety into priority provisional application No. 62 / 903,700.
[0173] Model Architecture
[0174] We trained several ultra-deep convolutional neural network-based models for the sequence-to-sequence epigenetic model 2700 and / or pathogenicity determiner 3100. We designed four architectures, Model-80nt, Model-400nt, Model-2k, and Model-10k, which used 40, 200, 1,000, and 5,000 nucleotides on each side of the position of interest as input, respectively. The input to the model is a sequence of one-hot encoded nucleotides, where A, C, G, and T (or equivalently, U) are encoded as [1,0,0,0], [0,1,0,0], [0,0,1,0], and [0,0,0,1].
[0175] The model architecture may be used in the sequence-to-sequence epigenetic model 2700 and / or the pathogenicity determiner 3100 .
[0176] The model architecture has a modified wavenet-type architecture that iterates over a specific position in the input promoter sequence and over three base changes from a reference base found at that specific position. The modified wavenet-type architecture can calculate up to 9,000 outputs for 3,000 positions in the input because each position has up to three single base changes. The modified wavenet-type architecture is relatively well calibrated because intermediate calculations are reused. The pathogenicity classifier determines the pathogenicity likelihood score of at least one of the three base changes at multiple specific positions in the input promoter sequence in a single call of the modified wavenet-like architecture and stores the pathogenicity likelihood score determined in the single call. Determining at least one of the three base changes also includes determining all three variations. The multiple specific positions are at least 500, or 1,000, or 1,500, or 2,000, or 90% of the input promoter sequence.
[0177] The basic unit of the model architecture is the residual block (He et al., 2016b), which consists of a batch normalization layer (Ioffe and Szegedy, 2015), a rectified linear unit (ReLU), and convolutional units organized in a specific way ( Figure 21 、 Figure 22 、 Figure 23 and Figure 24 Residual blocks are commonly used when designing deep neural networks. Before the development of residual blocks, deep neural networks consisting of multiple convolutional units stacked one after another were very difficult to train due to the problem of exploding / vanishing gradients (Glorot and Bengio, 2010), and increasing the depth of such neural networks generally led to higher training errors (He et al., 2016a). Through a comprehensive set of computational experiments, an architecture consisting of many residual blocks stacked one after another was shown to overcome these problems (He et al., 2016a).
[0178] exist Figure 21 、 Figure 22 、 Figure 23 and Figure 24 The complete model architecture is provided. The architecture consists of K stacked residual blocks connecting the input layer to the penultimate layer, and convolutional units with softmax activation connecting the penultimate layer to the output layer. The residual blocks are stacked so that i th The output of the residual block is connected to i+1 thThe output of the residual block is added to the input of the penultimate layer. Such “skip connections” are often used in deep neural networks to increase the convergence rate during training (Oord et al., 2016).
[0179] Each residual block has three hyperparameters N, W, and D, where N represents the number of convolution kernels, W represents the window size, and D represents the dilation rate of each convolution kernel (Yu and Koltun, 2016). Since a convolution kernel with window size W and dilation rate D extracts features across (W-1)D neighboring locations, a residual block with hyperparameters N, W, and D extracts features across 2(W-1)D neighboring locations. Therefore, the total neighbor span of the model architecture is given by: where N i 、W i and D i is i th Hyperparameters of residual blocks. For the Model-80nt, Model-400nt, Model-2k, and Model-10k architectures, the number of residual blocks and the hyperparameters of each residual block are chosen so that S is equal to 80, 400, 2,000, and 10,000, respectively.
[0180] In addition to the convolutional units, the model architecture also has only normalization and nonlinear activation units. Therefore, the model can be used for sequence-to-sequence patterns with variable sequence lengths (Oord et al., 2016). For example, the input of the Model-10k model (S=10,000) is a single-hot encoded nucleotide sequence of length S / 2+l+S / 2, and the output is a l×3 matrix of three scores corresponding to the l-center positions in the input (i.e., the positions remaining after excluding the first and last S / 2 nucleotides). This feature can be used to obtain a huge amount of computational savings during training and testing. This is due to the fact that most of the computations for positions close to each other are common, and shared computations need to be performed only once by the model when they are used for the sequence-to-sequence epigenetic model 2700 and / or the pathogenicity determiner 3100.
[0181] Our model uses a residual block architecture, which has become widely adopted due to its success in image classification. A residual block consists of repeated convolutional units interspersed with skip connections that allow information from earlier layers to skip the residual block. In each residual block, the input layer is first batch normalized, followed by an activation layer using rectified linear units (ReLUs). The activations are then passed through a 1D convolutional layer. This intermediate output from the 1D convolutional layer is again batch normalized and ReLU activated, followed by another 1D convolutional layer. At the end of the second 1D convolution, we sum the output of the second 1D convolution with the original input into the residual block, which acts as a skip connection by allowing the original input information to bypass the residual block. In this type of architecture (called Deep Residual Learning Networks by its authors), the input is preserved in its original state and the residual connections remain free of nonlinear activations from the model, allowing for efficient training of deeper networks.
[0182] After the residual block, a softmax layer calculates the probability of the three states for each amino acid, where the maximum softmax probability determines the state of the amino acid. The model is trained using the ADAM optimizer with the cumulative class mutual entropy loss function of the entire protein sequence.
[0183] Dilated convolutions allow for large receptive fields with few trainable parameters. Dilated convolutions are convolutions in which the kernel is applied over an area larger than the kernel length by skipping input values at a specific step size (also known as the dilated convolution rate or dilation factor). Dilated convolutions add spacing between the elements of the convolution filter / kernel so that when performing the convolution operation, larger intervals of adjacent input entries (e.g., nucleotides, amino acids) are considered. This enables the incorporation of long-range contextual dependencies in the input. Dilated convolutions save part of the convolution calculations for reuse when processing adjacent nucleotides.
[0184] The example shown uses 1D convolution. In other implementations, the model can use different types of convolutions, such as 2D convolution, 3D convolution, dilated or dilated convolution, transposed convolution, separable convolution, and depthwise separable convolution. Some layers also use the ReLU activation function, which greatly accelerates the convergence of stochastic gradient descent compared to saturating nonlinearities (such as sigmoid or hyperbolic tangent). Other examples of activation functions that can be used by the disclosed technology include parametric ReLU, leaky ReLU, and exponential linear unit (ELU).
[0185] Some layers also use batch normalization (Ioffe and Szegedy 2015). Regarding batch normalization, the distribution of each layer in a convolutional neural network (CNN) changes during training and varies from layer to layer. This slows down the convergence of the optimization algorithm. Batch normalization is a technique to overcome this problem. Let x denote the input to a batch normalization layer and z denote its output. Batch normalization applies the following transformation to x:
[0186]
[0187] Batch Normalization applies mean-variance normalization to the input x using μ and σ, and linearly scales and offsets it using γ and β. The normalization parameters μ and σ for the current layer are calculated over the training set using a method called exponential moving averages. In other words, they are not trainable parameters. In contrast, γ and β are trainable parameters. During inference, the values of μ and σ calculated during training are used in the forward pass.
[0188] like Figure 19 As shown, the model can include groups of residual blocks arranged in a sequence from lowest to highest. Each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block.
[0189] like Figure 20 As shown, each residual block may include at least one batch normalization layer, at least one rectified linear unit (ReLU) layer, at least one dilated convolution layer, and at least one residual connection. In such a specific implementation, each residual block includes two batch normalization layers, two ReLU nonlinearity layers, two dilated convolution layers, and one residual connection.
[0190] like Figure 21 、 Figure 22 、 Figure 23 and Figure 24 As shown, in the model, the dilated convolution rate proceeds non-exponentially from the lower residual block group to the higher residual block group.
[0191] like Figure 21 、 Figure 22 、 Figure 23 and Figure 24 As shown, in the model, the convolution window size varies between groups of residual blocks.
[0192] The model can be configured to evaluate an input comprising a target nucleotide sequence further flanked by 40 upstream context nucleotides and 40 downstream context nucleotides. In such an implementation, the model comprises a set of four residual blocks and at least one skip connection. Each residual block has 32 convolutional filters, 11 convolution window sizes, and 1 dilated convolution rate. This implementation of the model is referred to herein as "SpliceNet80" and is Figure 21 Shown.
[0193] The model can be configured to evaluate an input comprising a target nucleotide sequence further flanked by 200 upstream context nucleotides and 200 downstream context nucleotides. In such an implementation, the model comprises at least two groups of four residual blocks and at least two skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate. Each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates. This implementation of the model is referred to herein as "SpliceNet400" and is described in detail in the accompanying drawings. Figure 22 Shown.
[0194] The model can be configured to evaluate an input comprising a target nucleotide sequence further flanked by 1000 upstream context nucleotides and 1000 downstream context nucleotides. In such an implementation, the model comprises at least three groups of four residual blocks and at least three skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate. Each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates. Each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates. This implementation of the model is referred to herein as "SpliceNet2000" and is described in detail in
[15] . Figure 23 Shown.
[0195] The model can be configured to evaluate an input comprising a target nucleotide sequence further flanked by 5000 upstream context nucleotides and 5000 downstream context nucleotides. In such an implementation, the model comprises at least four groups of four residual blocks and at least four skip connections. Each residual block in the first group has 32 convolution filters, 11 convolution window sizes, and 1 dilated convolution rate. Each residual block in the second group has 32 convolution filters, 11 convolution window sizes, and 4 dilated convolution rates. Each residual block in the third group has 32 convolution filters, 21 convolution window sizes, and 19 dilated convolution rates. Each residual block in the fourth group has 32 convolution filters, 41 convolution window sizes, and 25 dilated convolution rates. This implementation of the model is referred to herein as "SpliceNet10000" and is described in detail in
[15] . Figure 24 Shown.
[0196] The trained model can be deployed on one or more production servers, which receive input sequences from requesting clients, such as Figure 18 In such implementations, the production server processes the input sequence through the input and output stages of the model to produce output that is transmitted to the client, as shown in Figure 18 shown.
[0197] Figure 35 Describes how to generate a training data set for the disclosed technology. First, according to one embodiment, promoter sequences in 19,812 genes are identified. In some embodiments, each of the 19,812 promoter sequences has 3001 base positions (excluding flanking context outside the promoter region), which results in a total of 59,455,812 base positions 3501 (gray).
[0198] In one implementation, from a total of 59,455,812 base positions 3501, 8,048,977 observed pSNV positions 3502 are suitable as benign positions. 8,048,977 benign positions 3502 yield 8,701,827 observed pSNVs according to one implementation, which form the final benign group. In some implementations, benign pSNVs are observed in humans and non-human primate species such as chimpanzees, bonobos, gorillas, orangutans, rhesus monkeys, and marmosets.
[0199] In some implementations, the criterion for inclusion in the benign group is that the minimum allele frequency of the observed pSNV should be greater than 0.1%. According to one implementation, such a criterion resulted in 600,000 observed pSNVs. In other implementations, the inclusion criterion did not consider the minimum allele frequency of the observed pSNV. That is, as long as a pSNV was observed in humans and non-human primate species, it was included in the benign group and therefore labeled as benign. According to one implementation, the second inclusion strategy resulted in a much larger benign group containing 8,701,827 observed pSNVs.
[0200] Additionally, from a total of 59,455,812 base positions 3501, 15,406,835 unobserved pSNV positions 3503 belonging to homopolymer regions, low complexity regions, and overlapping coding positions (e.g., start or stop codons) were removed, as they were considered unreliable due to sequence-specific errors or irrelevant to the analysis of non-coding variants.
[0201] Thus, in some implementations, the result is 36,000,000 unobserved pSNV positions 3504, from which, by mutating each of the 36,000,000 loci to three alternative single-base alleles, a total of 108,000,000 unobserved pSNVs 3505 were obtained. According to one implementation, these 108,000,000 unobserved pSNVs form the final pool 3505 of unobserved pSNVs generated by substitution.
[0202] Computer system
[0203] Figure 33 This is a simplified block diagram of a computer system that can be used to implement the disclosed technology. The computer system typically includes at least one processor that communicates with a plurality of peripheral devices via a bus subsystem. These peripheral devices may include a storage subsystem, including, for example, a memory device and a file storage subsystem, a user interface input device, a user interface output device, and a network interface subsystem. The input and output devices allow a user to interact with the computer system. The network interface subsystem provides an interface to an external network, including interfaces to corresponding interface devices in other computer systems.
[0204] In one implementation, a neural network, such as an ACNN and a CNN, is communicatively linked to a storage subsystem and a user interface input device.
[0205] User interface input devices may include keyboards; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; scanners; touch screens incorporated into displays; audio input devices such as voice recognition systems and microphones; and other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways of entering information into a computer system.
[0206] User interface output devices may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as an audio output device. Generally speaking, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from a computer system to a user or to another machine or computer system.
[0207] The storage subsystem stores programming and data structures that provide the functionality and methods of some or all modules described herein. These software modules are typically executed by a processor alone or in combination with other processors.
[0208] The memory used in the storage subsystem may include multiple memories, including a primary random access memory (RAM) for storing instructions and data during program execution and a read-only memory (ROM) in which fixed instructions are stored. The file storage subsystem may provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media tape. Modules that implement the functionality of certain specific implementations may be stored by the file storage subsystem in the storage subsystem or in other machines accessible to the processor.
[0209] The bus subsystem provides a mechanism for the various components and subsystems of the computer system to communicate with each other as intended. Although the bus subsystem is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0210] The computer system itself can be of different types, including a personal computer, a laptop, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely distributed group of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the Figure 33 The description of the computer system depicted in the is intended only as a specific example for illustrating the disclosed technology. Many other configurations of the computer system are possible, with more Figure 33 More or fewer components may be used in the computer system depicted.
[0211] The deep learning processor can be a GPU or FPGA and can be hosted by deep learning cloud platforms such as Google Cloud Platform, Xilinx, and Cirrascale. Examples of deep learning processors include Google's Tensor Processing Unit (TPU), rack-mounted solutions (such as the GX4 rack-mounted series and the GX8 rack-mounted series), NVIDIA DGX-1, Microsoft's Stratix V FPGA, Graphcore's Intelligent Processor Unit (IPU), Qualcomm's Zeroth Platform with Snapdragon processors, NVIDIA's Volta, NVIDIA's DRIVE PX, NVIDIA's JETSON TX1 / TX2 MODULE, Intel's Nirvana, Movidius VPU, Fujitsu DPI, ARM's DynamicIQ, IBM TrueNorth, etc.
[0212] The foregoing description is presented to enable making and using the disclosed technology. Various modifications to the disclosed implementations will be apparent, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the disclosed technology. Therefore, the disclosed technology is not intended to be limited to the specific implementations shown but is to be accorded the widest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims.
[0213] Terms
[0214] This document contains the following terms.
[0215] Artificial intelligence-based pathogenicity classification of noncoding variants
[0216] 1. An artificial intelligence-based system, comprising:
[0217] An input preparation module, the input preparation module accesses a sequence database and generates an input base sequence, wherein the input base sequence includes
[0218] (i) a targeting base sequence having a targeting base, wherein the targeting base sequence is flanked by
[0219] (ii) a right-end base sequence having a downstream context base,
[0220] and
[0221] (iii) a left-end base sequence having upstream context bases;
[0222] a sequence-to-sequence model that processes the input base sequence and generates an alternative representation of the input base sequence; and
[0223] an output module that processes the alternative representation of the input base sequence and generates at least one per-base output for each of the targeted bases in the targeted base sequence;
[0224] wherein the per-base output specifies the signal levels of the plurality of epigenetic loci for the corresponding targeted base.
[0225] 2. The artificial intelligence-based system of clause 1 , wherein the per-base output uses a continuous value to specify the signal level of the plurality of epigenetic loci.
[0226] 3. The artificial intelligence-based system of claim 1 , wherein the output module further comprises a plurality of per-track processors, wherein each per-track processor corresponds to a respective epigenetic track among the plurality of epigenetic tracks and generates a signal level for the corresponding epigenetic track.
[0227] 4. An artificial intelligence-based system according to claim 3, wherein each of the plurality of per-track processors processes the alternative representation of the input base sequence, generates an additional alternative representation specific to the particular per-track processor, and generates the signal level of the corresponding epigenetic track based on the additional alternative representation.
[0228] 5. The artificial intelligence-based system of clause 4, wherein each per-track processor further comprises at least one residual block and a linear module and / or a rectified linear unit (ReLU) module, and
[0229] wherein the residual block generates the further alternative representation and the linear module and / or the ReLU module processes the further alternative representation and produces the signal level of the corresponding epigenetic trajectory.
[0230] 6. An artificial intelligence-based system according to claim 3, wherein each of the plurality of per-track processors processes the alternative representation of the input base sequence and generates the signal level of the corresponding epigenetic track based on the alternative representation.
[0231] 7. The artificial intelligence-based system of clause 6, wherein each per-track processor further comprises a linear module and / or a ReLU module.
[0232] 8. An artificial intelligence-based system according to claim 1, wherein the multiple epigenetic tracks include changes in deoxyribonucleic acid (DNA) methylation (e.g., CpG), histone modifications, non-coding RNA (ncRNA) expression, and chromatin structure changes (e.g., nucleosome positioning).
[0233] 9. An artificial intelligence-based system according to claim 1, wherein the plurality of epigenetic tracks include deoxyribonuclease (DNase) tracks.
[0234] 10. The artificial intelligence-based system of clause 1 , wherein the plurality of epigenetic tracks comprises a histone 3 lysine 27 acetylation (H3K27ac) track.
[0235] 11. An artificial intelligence-based system according to claim 1, wherein the sequence-to-sequence model is a convolutional neural network.
[0236] 12. An artificial intelligence-based system according to claim 11, wherein the convolutional neural network further comprises a residual block group.
[0237] 13. An artificial intelligence-based system according to clause 12, wherein each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block.
[0238] 14. An artificial intelligence-based system according to clause 12, wherein the convolutional neural network is parameterized by the number of residual blocks, the number of skip connections, and the number of residual connections.
[0239] 15. The artificial intelligence-based system of clause 12, wherein each set of residual blocks generates an intermediate output by processing the aforementioned input, wherein the dimension of the intermediate output is (I-[{(W-1)*D}*A])×N, where
[0240] I is the dimension of the aforementioned input,
[0241] W is the convolution window size of the residual block,
[0242] D is the dilated convolution rate of the residual block,
[0243] A is the number of dilated convolutional modules in the group, and
[0244] N is the number of convolutional filters in the residual block.
[0245] 16. The artificial intelligence-based system of clause 15, wherein the dilated convolution rate progresses non-exponentially from a lower group of residual blocks to a higher group of residual blocks.
[0246] 17. The artificial intelligence-based system of clause 16, wherein the dilated convolution saves some of the convolution computations for reuse in processing adjacent bases.
[0247] 18. An artificial intelligence-based system according to clause 13, wherein the convolution window size differs between groups of residual blocks.
[0248] 19. An artificial intelligence-based system according to clause 1, wherein the dimension of the input base sequence is (Cu+L+Cd)×4, where
[0249] Cu is the number of upstream context bases in the left end base sequence;
[0250] Cd is the number of downstream context bases in the right end base sequence, and
[0251] L is the number of targeted bases in the targeted base sequence.
[0252] 20. The artificial intelligence-based system of clause 11, wherein the convolutional neural network further comprises a dimension-changing convolutional module that reshapes spatial dimensions and feature dimensions of the preceding input.
[0253] 21. The artificial intelligence-based system of clause 12, wherein each residual block further comprises at least one batch normalization module, at least a ReLU module, at least one dilated convolutional module, and at least one residual connection.
[0254] 22. The artificial intelligence-based system of clause 21, wherein each residual block further comprises two batch normalization modules, two ReLU non-linear modules, two dilated convolutional modules, and one residual connection.
[0255] 23. The artificial intelligence-based system of clause 1, wherein the sequence-to-sequence model is trained with training data comprising both coding bases and non-coding bases.
[0256] 24. An artificial intelligence-based system, the artificial intelligence-based system comprising:
[0257] a sequence-to-sequence model;
[0258] a plurality of per-track processors corresponding to respective epigenetic tracks of a plurality of epigenetic tracks,
[0259] wherein each per-track processor further comprises at least one processing module (e.g., a residual block) and an output module (e.g., a linear module, a rectified linear unit (ReLU) module);
[0260] the sequence-to-sequence model processes an input base sequence and generates a surrogate representation of the input base sequence;
[0261] the processing module of each per-track processor processes the surrogate representation and generates a further surrogate representation specific to the particular per-track processor; and
[0262] the output module of each per-track processor processes the further surrogate representation generated by the corresponding processing module of the per-track processor and produces, as output, a signal level for each base in the input base sequence and the corresponding epigenetic track.
[0263] Other implementations can include a non-transitory computer readable storage medium storing instructions executable by a processor to perform actions of the systems described above. Yet another implementation can include a method of performing actions of the systems described above.
[0264] Artificial intelligence-based classification of pathogenicity of non-coding variants
[0265] 1. An artificial intelligence-based system, the artificial intelligence-based system comprising:
[0266] an input preparation module that accesses a sequence database and generates
[0267] (i) a reference sequence that contains a base at a targeted position, wherein the base is flanked by a downstream context base and an upstream context base, and
[0268] (ii) an alternative base sequence that contains a variant of the base at the targeted position, wherein the variant is flanked by the downstream context base and the upstream context base;
[0269] a sequence-to-sequence model,
[0270] the sequence-to-sequence model processes the reference sequence and generates a reference output, wherein the reference output specifies, for each base in the reference sequence, a signal level of a plurality of epigenetic tracks, and
[0271] processes the alternative sequence and generates an alternative output, wherein the alternative output specifies, for each base in the alternative sequence, the signal level of the plurality of epigenetic tracks;
[0272] a comparator that
[0273] applies a position-wise determinantal comparison function to the reference output and the alternative output generated for bases in the reference sequence and the alternative sequence, and
[0274] generates a position-wise comparison result based on differences between the signal levels of the reference output and the alternative output caused by the variant in the alternative sequence; and
[0275] a pathogenicity determiner that processes the position-wise comparison result and produces an output that scores the variant in the alternative sequence as pathogenic or benign.
[0276] 2. The artificial intelligence-based system of clause 1, wherein the position-wise determinantal comparison function computes an element-wise difference between the reference output and the alternative output.
[0277] 3. The artificial intelligence-based system of clause 1, wherein the position-wise determinantal comparison function computes an element-wise sum of the reference output and the alternative output.
[0278] 4. The artificial intelligence-based system of clause 1, wherein the position-wise deterministic comparison function computes a position-wise element-wise ratio between the reference output and the alternative output.
[0279] 5. The artificial intelligence-based system of clause 1, further comprising a post-processing module, the post-processing module
[0280] processing the reference output and producing a further reference output; and
[0281] processing the alternative output and producing a further alternative output.
[0282] 6. The artificial intelligence-based system of clause 5, wherein the comparator
[0283] applying the position-wise deterministic comparison function to the further reference output and the further alternative output for elements in the reference output and the alternative output, and
[0284] generating the position-wise comparison result.
[0285] 7. The artificial intelligence-based system of clause 5, wherein the post-processing module is a convolutional neural network having one or more convolutional layers.
[0286] 8. The artificial intelligence-based system of clause 1, wherein the sequence-to-sequence model further comprises a plurality of intermediate layers, and one of the intermediate layers
[0287] processing the reference sequence and generating an intermediate reference output;
[0288] processing the alternative sequence and generating an intermediate alternative output; and
[0289] the comparator
[0290] applying the position-wise deterministic comparison function to the intermediate reference output and the intermediate alternative output for elements in the intermediate reference sequence and the intermediate alternative sequence, and
[0291] generating the position-wise comparison result.
[0292] 9. The artificial intelligence-based system of clause 1, wherein the reference sequence is further flanked by a right-end flanking base sequence and a left-end flanking base sequence, and the sequence-to-sequence model processes the reference sequence together with the right-end flanking base sequence and the left-end flanking base sequence.
[0293] 10. An artificial intelligence-based system according to claim 1, wherein the alternative sequence is also flanked by a right-side flanking base sequence and a left-side flanking base sequence, and the sequence-to-sequence model processes the alternative sequence together with the right-side flanking base sequence and the left-side flanking base sequence.
[0294] 11. An artificial intelligence-based system according to claim 1, wherein the reference sequence and the alternative sequence are non-coding base sequences.
[0295] 12. An artificial intelligence-based system according to clause 11, wherein the reference sequence and the alternative sequence are promoter sequences, and the variant is a promoter sequence.
[0296] 13. The artificial intelligence-based system of clause 1 , wherein the input feed module includes the reference sequence and the alternative sequence in addition to the position-by-position comparison results in the input to the pathogenicity determiner.
[0297] 14. The artificial intelligence-based system of clause 13, wherein the reference sequence and the alternative sequence are one-hot encoded.
[0298] 15. An artificial intelligence-based system according to claim 1, wherein the sequence-to-sequence model is a convolutional neural network.
[0299] 16. An artificial intelligence-based system according to clause 15, wherein the convolutional neural network further comprises a residual block group.
[0300] 17. An artificial intelligence-based system according to clause 16, wherein each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block.
[0301] 18. An artificial intelligence-based system according to clause 16, wherein the convolutional neural network is parameterized by the number of residual blocks, the number of skip connections, and the number of residual connections.
[0302] 19. The artificial intelligence-based system of clause 1 , wherein the pathogenicity determiner further comprises an output module that generates the output as a pathogenicity score.
[0303] 20. The artificial intelligence-based system of clause 19, wherein the output module is a sigmoid processor and the pathogenicity score is between zero and one.
[0304] 21. The artificial intelligence-based system of clause 20, further comprising identifying whether the variant in the alternative sequence is pathogenic or benign based on whether the pathogenicity score is above or below a preset threshold.
[0305] 22. The artificial intelligence-based system of clause 1 , further comprising jointly training the sequence-to-sequence model and the pathogenicity determiner using continuous back-propagation.
[0306] 23. The artificial intelligence-based system of clause 1 , further comprising using transfer learning to initialize weights of the pathogenicity determiner during its training based on weights of the sequence-to-sequence model learned during its training.
[0307] 24. An artificial intelligence-based system, comprising:
[0308] a pathogenicity determiner that processes the position-by-position comparison results and generates an alternative representation of the position-by-position comparison results,
[0309] The position-by-position comparisons are based on differences in signal levels caused by variants in the alternative sequences:
[0310] the signal levels of multiple epigenetic loci determined for bases in the reference sequence, and
[0311] signal levels of the plurality of epigenetic loci determined for bases in the replacement sequence; and
[0312] An output module processes the alternative representation of the position-by-position comparison results and generates an output that scores the variant in the alternative sequence as pathogenic or benign.
[0313] 25. The artificial intelligence-based system of clause 24, further comprising providing the reference sequence as input to the pathogenicity determiner.
[0314] 26. The artificial intelligence-based system of clause 24, further comprising providing the alternative sequence as input to the pathogenicity determiner.
[0315] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform the actions of the above-described system. Another implementation may include a method of performing the actions of the above-described system.
[0316] Gene expression-based labeling for non-coding variants trained with artificial intelligence
[0317] 1. An artificial intelligence-based method for training a pathogenicity determiner, the method comprising:
[0318] training the pathogenicity determiner using training data comprising a set of pathogenic non-coding variants annotated with a pathogenicity label (e.g., “1”) and a set of benign non-coding variants annotated with a benign label (e.g., “0”);
[0319] wherein the set of pathogenic non-coding variants is a singleton pattern that occurs only in a single individual among a group of individuals, and the single individual exhibits low expression across multiple tissues of genes adjacent to the non-coding variants in the pathogenic set;
[0320] wherein the set of benign non-coding variants is a singleton pattern that occurs only in a single individual among the group of individuals, and the single individual does not exhibit the low expression across the plurality of tissues for genes adjacent to the non-coding variants in the benign set; and
[0321] For a specific non-coding variant in the training data,
[0322] generating (i) an alternative sequence containing the specific noncoding variant at the targeted position, the specific noncoding variant flanked by downstream context bases and upstream context bases, and (ii) a reference sequence containing a reference base at the targeted position, the reference base flanked by the downstream context bases and the upstream context bases;
[0323] processing the alternative sequence and the reference sequence through an epigenetic model and determining a signal level for a plurality of epigenetic loci for each base in the alternative sequence and the reference sequence;
[0324] generating a position-by-position comparison based on differences between the signal levels determined for bases in the reference sequence and the alternative sequence caused by the specific non-coding variant;
[0325] processing the position-by-position comparison results through the pathogenicity determiner and generating a pathogenicity prediction for the particular non-coding variant; and
[0326] Backpropagation is used to modify the weights of the pathogenicity determiner based on the calculated error between the pathogenicity prediction and the pathogenicity signature when the particular non-coding variant is from the pathogenic group, and based on the calculated error between the pathogenicity prediction and the benign signature when the particular non-coding variant is from the benign group.
[0327] 2. An artificial intelligence-based method according to claim 1, wherein the low expression is determined by analyzing the distribution of expression exhibited by the group of individuals for each of the genes across each of the multiple tissues, and calculating the median z-score of the single individual based on the distribution.
[0328] 3. An artificial intelligence-based method according to claim 1, wherein the proximity of the gene adjacent to the non-coding variant in the pathogenicity group is measured based on the number of bases between the transcription start site (TSS) on the gene and the non-coding variant.
[0329] 4. The artificial intelligence-based system of clause 3, wherein the number of bases is 1500 bases.
[0330] 5. The artificial intelligence-based system of clause 1 , further comprising inferring low expression when the median z-score is below a threshold.
[0331] 6. The artificial intelligence-based system of clause 5, wherein the threshold is -1.5.
[0332] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform the actions of the above-described system. Another implementation may include a method of performing the actions of the above-described system.
Claims
1. An artificial intelligence-based system, comprising: An input preparation module, the input preparation module accesses a sequence database and generates an input base sequence, wherein the input base sequence includes (i) a targeting base sequence having a targeting base, wherein the targeting base sequence is flanked by (ii) a right-end base sequence having downstream context bases, and (iii) a left-end base sequence having upstream context bases; a sequence-to-sequence model based on a convolutional neural network, wherein the sequence-to-sequence model processes the input base sequence and generates an alternative representation of the input base sequence; wherein the convolutional neural network has been trained on the input base sequences of the target base sequence, the right end base sequence, and the left end base sequence and on the reference truth signal level in sequences of multiple epigenetic tracks; and an output module that processes the alternative representation of the input base sequence and generates at least one per-base output for each of the targeted bases in the targeted base sequence; wherein the per-base output specifies the signal levels of the plurality of epigenetic loci for the corresponding targeted base.
2. The artificial intelligence-based system of claim 1 , wherein the per-base output uses a continuous value to specify the signal level of the plurality of epigenetic loci.
3. The artificial intelligence-based system of claim 1 , wherein the output module further comprises a plurality of per-track processors, wherein each per-track processor corresponds to a respective epigenetic track in the plurality of epigenetic tracks and generates a signal level of the corresponding epigenetic track.
4. The artificial intelligence-based system of claim 3 , wherein each of the plurality of per-track processors processes the alternative representation of the input base sequence, generates an additional alternative representation specific to the particular per-track processor, and generates the signal level of the corresponding epigenetic track based on the additional alternative representation.
5. The artificial intelligence-based system according to claim 4, wherein each per-track processor further comprises at least one residual block and a linear module and / or a rectified linear unit (ReLU) module, and wherein the residual block generates the further alternative representation and the linear module and / or the ReLU module processes the further alternative representation and produces the signal level of the corresponding epigenetic trajectory.
6. The artificial intelligence-based system of claim 3, wherein each of the plurality of per-track processors processes the alternative representation of the input base sequence and generates the signal level of the corresponding epigenetic track based on the alternative representation.
7. The artificial intelligence-based system according to claim 6, wherein each per-track processor further comprises a linear module and / or a ReLU module.
8. The artificial intelligence-based system of claim 1, wherein the plurality of epigenetic tracks comprises changes in DNA methylation, histone modifications, non-coding RNA (ncRNA) expression, and chromatin structure changes.
9. The artificial intelligence-based system of claim 8, wherein the DNA methylation changes are CpG and the chromatin structure changes are nucleosome positioning.
10. The artificial intelligence-based system of claim 1, wherein the plurality of epigenetic tracks comprises a deoxyribonuclease (DNase) track.
11. The artificial intelligence-based system of claim 1 , wherein the plurality of epigenetic tracks comprises a histone 3 lysine 27 acetylation H3K27ac track.
12. The artificial intelligence-based system of claim 1, wherein the convolutional neural network further comprises a residual block group.
13. The artificial intelligence-based system of claim 12, wherein each group of residual blocks is parameterized by the number of convolution filters in the residual block, the convolution window size of the residual block, and the dilated convolution rate of the residual block.
14. The artificial intelligence-based system of claim 12, wherein the convolutional neural network is parameterized by the number of residual blocks, the number of skip connections, and the number of residual connections.
15. The artificial intelligence-based system of claim 12, wherein each group of residual blocks generates an intermediate output by processing the aforementioned input, wherein the dimension of the intermediate output is (I-[{(W-1) * D} * A]) × N, where I is the dimension of the aforementioned input, W is the convolution window size of the residual block, D is the dilated convolution rate of the residual block, A is the number of dilated convolutional modules in the group, and N is the number of convolutional filters in the residual block.
16. The artificial intelligence based system of claim 15, wherein the dilated convolution rate progresses non-exponentially from a lower residual block group to a higher residual block group.
17. The artificial intelligence-based system according to claim 16, wherein the dilated convolution saves part of the convolution calculation for reuse in processing adjacent bases.
18. The artificial intelligence based system of claim 13, wherein the convolution window size differs between groups of residual blocks.
19. The artificial intelligence-based system according to claim 1, wherein the dimension of the input base sequence is (Cu + L + Cd) × 4, where Cu is the number of upstream context bases in the left end base sequence, Cd is the number of downstream context bases in the right end base sequence, and L is the number of targeted bases in the targeted base sequence.
20. The artificial intelligence-based system of claim 1, wherein the convolutional neural network further comprises a dimension-changing convolution module, wherein the dimension-changing convolution module reshapes the spatial dimension and feature dimension of the aforementioned input.
21. The artificial intelligence-based system of claim 12, wherein each residual block further comprises at least one batch normalization module, at least one ReLU module, at least one dilated convolution module, and at least one residual connection.
22. The artificial intelligence-based system according to claim 21, wherein each residual block further comprises two batch normalization modules, two ReLU nonlinear modules, two dilated convolution modules and one residual connection.
23. The artificial intelligence-based system of claim 1, wherein the sequence-to-sequence model is trained with training data comprising both coding bases and non-coding bases.
24. An artificial intelligence-based system, comprising: Sequence-to-sequence model based on convolutional neural network; a plurality of per-track processors corresponding to respective epigenetic tracks of the plurality of epigenetic tracks, Each track processor further comprises at least one processing module and an output module; and Each per-track processor has been trained on an input base sequence and at a ground truth signal level in sequences of multiple epigenetic tracks; The sequence-to-sequence model processes an input base sequence and generates an alternative representation of the input base sequence; The processing module of each per-track processor processes the alternative representation and generates a further alternative representation specific to the particular per-track processor; and The output module of each per-track processor processes the additional alternative representation generated by the corresponding processing module of the per-track processor and produces as output the corresponding epigenetic track and the signal level of each base in the input base sequence.
25. The artificial intelligence-based system of claim 24, wherein the at least one processing module is at least one residual block, and the output module is a linear module and / or a rectified linear unit (ReLU) module.
Citation Information
Patent Citations
Aberrant splicing detection using convolutional neural networks (CNNs)
US11397889B2
Deep Learning-Based Aberrant Splicing Detection
US20190114391A1
Deep Learning-Based Splice Site Classification
US20190114547A1
Deep learning-based splice site classification
WO2019079198A1
Deep learning-based aberrant splicing detection
WO2019079200A1