Bioinformatics-based method and system for classifying the pathogenicity of single nucleotide variations
By adopting channel attention mechanism and improved variational autoregression model in the classification of single nucleotide variant pathogenicity, the problem of relying on complex biological experiments and lack of accurate pathogenicity labels in the prior art is solved, and efficient and accurate pathogenicity assessment is achieved.
Patent Information
- Application Number
- CN202311082624.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-08-24
AI Technical Summary
When determining the pathogenicity of single nucleotide variants, the prior art relies on complex biological experiments, which leads to high cost, time-consuming and lacks accurate pathogenicity labels, resulting in insufficient model learning effect.
The preprocessing method based on the channel attention mechanism was adopted, combined with the improved variational autoregression model, and the evolution index of each single nucleotide variant was calculated. The evolution index was clustered using the Gaussian mixed model to finally obtain a pathogenic refinement classification of each single nucleotide variant.
The need to use sparse, biased and noisy tagged data reduces experimental operations, improves the model's ability to capture important patterns and distributions of DNA sequences, and enhances the accuracy of pathogenicity assessment of rare variants.
Smart Images

Figure CN117292753B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics computing, and particularly to a method and system for classifying the pathogenicity of single nucleotide variations based on bioinformatics. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the development of gene sequencing technology, a large amount of gene sequence data has been generated, such as Multiple Sequence Alignment (MSA). Among them, single nucleotide variants (SNVs) can cause more than 6,000 diseases such as cystic fibrosis, Marfan syndrome, Alzheimer's disease, and cancer. Single nucleotide variants have an important impact on an individual's genetic characteristics and disease risk. When determining the impact of single nucleotide variants on an individual through traditional biological experimental methods, complex experimental processes such as cell culture and transfection, expression analysis, and function identification are required, resulting in high costs and long time consumption.
[0004] When using machine learning, taking the supervised learning method as an example, a large amount of labeled data is required as the training data set. The quality and number of labels have an important impact on the effect of machine learning. The pathogenicity labels of single nucleotide variants need to be verified by biological experiments or manually annotated, which consumes a large amount of time, resources, and manpower.
[0005] Secondly, there are approximately 3 million single nucleotide variants in the human genome, but the number of currently determined pathogenicity labels is only a few thousand. For most single nucleotide variants, especially rare variants, accurate pathogenicity labels are lacking, which makes the labels sparse and unbalanced, resulting in insufficient learning effect of the trained model for minority classes (such as pathogenicity) and prone to errors.
[0006] Thirdly, the selection and feasibility of biological experimental methods may also affect the accuracy of labels. Some experimental methods may not be able to fully simulate the role of single nucleotide variants in the real biological environment, resulting in label deviation. At the same time, due to differences in genetic background, environmental factors, and disease characteristics among different patients, there may be inconsistent clinical manifestations and pathogenicity assessments for the same single nucleotide variant. This heterogeneity may lead to noise and deviation in labels, affecting the learning effect of the model. Summary of the Invention
[0007] To solve the technical problems existing in the above-mentioned background art, the present invention provides a method and system for classifying the pathogenicity of single nucleotide variations based on biological information. The method uses a channel attention mechanism to preprocess the data, and then uses an improved Variational Auto Regressive (VAR) model to learn the sequence distribution of DNA and learn the evolutionary index of each single nucleotide variation of the DNA. A Gaussian mixture model is used to cluster the evolutionary indices, and finally a refined classification of the pathogenicity of each single nucleotide variation is obtained.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] The first aspect of the present invention provides a method for classifying the pathogenicity of single nucleotide variations based on biological information, including the following steps:
[0010] Obtain the multiple sequence alignment of each DNA, and through preprocessing, obtain the position of each nucleotide and the importance of each nucleotide sequence. The obtained importance is used as a weight and assigned to the corresponding nucleotide sequence to form an input matrix;
[0011] Sample the DNA sequence according to the obtained input matrix, and based on a variational auto-regressive model with a self-attention mechanism, learn the probability distribution of the DNA sequence;
[0012] Calculate the evolutionary index of each single nucleotide variation sequence according to the obtained probability distribution. The evolutionary index is the logarithmic likelihood difference between the single nucleotide variation sequence and the wild-type sequence;
[0013] Fit the obtained evolutionary indices into multiple clusters, which respectively correspond to the pathogenicity probabilities of the single nucleotide variation sequences. The pathogenicity probabilities of the single nucleotide variation sequences are divided into five categories: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance.
[0014] Obtain the multiple sequence alignment of each DNA, and through preprocessing, obtain the position of each nucleotide and the importance of each nucleotide sequence; specifically:
[0015] Obtain the multiple sequence alignment file of each DNA, perform one-hot encoding on the sequence, where the encoding includes at least the sequence length and the number of nucleotide types, and perform a transpose process;
[0016] Based on parallel processing with a channel attention mechanism, obtain the position of each nucleotide and the importance of each nucleotide sequence.
[0017] Based on parallel processing with a channel attention mechanism, obtain the position of each nucleotide and the importance of each nucleotide sequence. The obtained importance is used as a weight and assigned to the corresponding nucleotide sequence to form an input matrix; specifically:
[0018] Based on global average pooling, compress the dimension of each channel to one dimension;
[0019] Based on sparse Softmax coding, obtain the relative importance of the one-dimensional matrix;
[0020] Take the obtained relative importance as weights and assign them to the corresponding sequences and nucleotide columns to obtain the preprocessed data as the input matrix.
[0021] The variational auto-regressive modeler with self-attention mechanism includes an encoder and a decoder corresponding to the parameters;
[0022] For each sequence S in the given multiple sequence alignment S, the encoder i obeys a posterior distribution. Based on the self-attention mechanism and the linear layer, fit the mean and standard deviation of the posterior distribution to achieve encoding;
[0023] The decoder uses an auto-regressive model to reconstruct the input latent variable into an approximate sequence s i ', and output the approximate probability distribution of the DNA sequence.
[0024] The specific process of the encoder is as follows:
[0025] Map the sequence distributed feature representation learned by the linear layer to the query matrix, key matrix, and value matrix through cross-channel information integration and one-dimensional convolution;
[0026] Perform matrix multiplication on the query matrix and the key matrix to calculate their correlation;
[0027] Perform normalization processing to obtain the attention probability distribution;
[0028] Take the attention probability distribution as the weight coefficient of the value matrix and perform weighted summation with the original features to obtain the output of the attention module;
[0029] Obtain the mean μ and standard deviation σ through the linear layer, and use the reparameterization trick to sample to get the latent variable z = μ + rnv·σ, where rnv is a random variable sampled from the standard normal distribution N(0,1).
[0030] The specific process of the decoder is as follows:
[0031] Concatenate the latent variable z with the input sequence, embed the concatenated enhanced input, convert it into a vector representation of a fixed dimension, and perform positional encoding on the embedded vector to retain the relative positions of the latent variable and the sequence;
[0032] The encoder models the relationships between different positions in the sequence and further processes the representation of the sequence;
[0033] Through the iteration of the masked multi-attention mechanism, self-attention mechanism and feed-forward neural network in the decoder, an approximate sequence of DNA is gradually generated;
[0034] Probability modeling of nucleotides at each position is performed through the output layer.
[0035] The obtained evolutionary indices are fitted into multiple clusters, corresponding to the pathogenic probabilities of single nucleotide variant sequences respectively. The pathogenic probabilities of single nucleotide variant sequences are divided into five categories: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance; specifically:
[0036] A Gaussian mixture model, i.e., the global Gaussian, is trained on the evolutionary indices of all DNAs. The parameters of this global model are used to initialize the Gaussian mixture model of a specific DNA, i.e., the local Gaussian;
[0037] Based on the global and local Gaussians, the pathogenicity score of the single nucleotide variant sequence of this specific DNA is calculated;
[0038] The obtained pathogenicity scores divide the single nucleotide variant sequences into four categories: benign, likely benign, likely pathogenic, and pathogenic;
[0039] Based on the prediction entropy, the uncertainty of the classification result of each single nucleotide variant sequence category is calculated, and the category with a set percentage in the highest uncertainty is taken as of uncertain significance.
[0040] The second aspect of the present invention provides a system for implementing the above method, including:
[0041] A preprocessing module, configured to: obtain the multiple sequence alignment of each DNA, and through preprocessing, obtain the position of each nucleotide and the importance of each nucleotide sequence. The obtained importance is used as a weight and assigned to the corresponding nucleotide sequence to form an input matrix;
[0042] A DNA distribution learning module, configured to: sample the DNA sequence according to the obtained input matrix, and based on a variational auto-regressive modeler with a self-attention mechanism, learn the probability distribution of the DNA sequence;
[0043] An evolutionary index module, configured to: calculate the evolutionary index of each single nucleotide variant sequence according to the obtained probability distribution. The evolutionary index is the logarithmic likelihood difference between the single nucleotide variant sequence and the wild-type sequence;
[0044] A classification module, configured to: fit the obtained evolutionary indices into multiple clusters, corresponding to pathogenic probabilities in different intervals respectively.
[0045] The third aspect of the present invention provides a computer-readable storage medium.
[0046] A computer-readable storage medium stores a computer program thereon. When the program is executed by a processor, it implements the steps in the above-mentioned method for classifying the pathogenicity of single nucleotide variations based on biological information.
[0047] The fourth aspect of the present invention provides a computer device.
[0048] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned method for classifying the pathogenicity of single nucleotide variations based on biological information.
[0049] Compared with the prior art, the above one or more technical solutions have the following beneficial effects:
[0050] 1. By learning the sequence distribution of DNA from the multiple sequence alignment data of DNA (derived from an MSA file) and learning the evolutionary index of each single nucleotide variation of the DNA, and classifying it into different categories according to the pathogenicity probability of each single nucleotide variation after clustering. The overall method is an unsupervised method. Compared with the supervised learning method, it does not need to use sparse, biased, and noisy labeled data, thus improving the learning effect of the model.
[0051] 2. Due to the adoption of machine learning, the experimental operations of researchers are reduced. There is no pre-induction preparation and long-term tracking experiment. By using the self-attention mechanism and variational auto-regression model, it is possible to better capture the important patterns and distributions of DNA sequences, and it is easier to migrate to a large amount of DNA sequence data compared with high-throughput biological experimental techniques.
[0052] 3. The parallel channel attention mechanism is used, which can simultaneously weight the importance of sequences and columns, improving the attention to key sequences and columns, and thus better capturing the important features of sequences.
[0053] 4. Autonomously find the focused columns and focused sequences. During data preprocessing, the gap threshold is abandoned, and the parallel channel attention mechanism is used instead to find the focused columns and focused sequences, and re-weight each sequence. In particular, a sparse Softmax is used instead of the original activation function Softmax in the attention mechanism. This operation only retains the top k calculated probabilities, where k is a hyperparameter selected manually, sparsifying the result of Softmax, which helps to avoid overfitting and improve the accuracy of subsequent classification.
[0054] 5. Reduce information attenuation during long sequence modeling. High-quality biological sequences are as long as possible and have the largest possible proportion of gap-free parts. Therefore, a self-attention mechanism is added to the encoder part, which can aggregate similar network nodes, calculate potential features, and integrate context information at the same time, effectively avoiding the problem of information attenuation during long sequence modeling, thereby enhancing the model's ability to capture long-term dependencies and improving the quality of the finally generated sequence.
[0055] 6. Combine the channel attention mechanism, one-dimensional long short-term memory network layer and encoder. By leveraging the advantages of these models and methods, it is possible to provide channel-level importance weights, process the temporal information of the sequence, and the variational auto-regressive model can learn the distribution characteristics of the sequence. The strategy of comprehensively using different models and methods can improve the expressive ability of the model to a certain extent.
[0056] 7. Compared with traditional unsupervised models, more single amino acid variation data, including synonymous mutation sequences, are used, expanding the dataset, and the classification categories are refined based on the public dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0058] Figure 1 is a schematic diagram of the pathogenicity classification process provided by one or more embodiments of the present invention;
[0059] Figure 2 is a schematic diagram of parallelly obtaining weights using the channel attention mechanism with fused sparse Softmax provided by one or more embodiments of the present invention;
[0060] Figure 3 is a schematic diagram of using VAR with an added attention mechanism to learn the DNA sequence distribution provided by one or more embodiments of the present invention;
[0061] Figure 4 is a schematic diagram of the modeling strategy provided by one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The present invention will be further described below in conjunction with the drawings and embodiments.
[0063] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0064] As described in the background art, single nucleotide variants (SNVs) can cause some diseases. By giving a multiple sequence alignment (MSA) of a DNA, the pathogenicity category of all SNVs of the DNA is judged. If pathogenic SNVs can be discovered in advance from the sequence information of the gene, it is beneficial for targeted prevention clinically.
[0065] Therefore, the following embodiments provide a method and system for classifying the pathogenicity of single nucleotide variants based on biological information. The channel attention mechanism is used to preprocess the data, and then an improved variational autoregressive (VAR) model is used to learn the sequence distribution of the DNA and learn the evolutionary index of each single nucleotide variant of the DNA. The Gaussian mixture model is used to cluster the evolutionary indices, and finally a refined classification of the pathogenicity of each single nucleotide variant is obtained. The DNA sequence after single nucleotide variation is divided into five categories according to the pathogenicity probability: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance.
[0066] Term explanation:
[0067] SNV, single nucleotide variants (Single Nucleotide Variants, SNV).
[0068] MSA, multiple sequence alignment (Multiple Sequence Alignment, MSA).
[0069] MVMA, method for classifying pathogenicity based on variational multi-attention autoregressive model (Method based on Variational Multi-Attention autoregressive, MVMA).
[0070] Embodiment 1:
[0071] (1) Traditional biological experiments include the following steps:
[0072] Determine the research objective: clarify the SNV to be studied, such as a specific gene or gene region, and the purpose of the study, such as functional analysis, mutation impact, etc.;
[0073] Data collection and selection: collect relevant SNV data, which can come from public databases, literature reports, or sequencing data of the research group itself. Select appropriate SNVs for further experiments according to the research objective;
[0074] Cloning construction: Use cloning technology to construct a recombinant vector containing the target SNV. This may involve steps such as PCR amplification of the target DNA fragment, ligation to an appropriate expression vector or reporter gene, etc.;
[0075] Cell culture and transfection: Transfect the recombinant vector into a suitable cell line or cell model for further study of the function of SNV. This may include common cell lines such as HEK293, CHO, etc., or specific cell lines such as cancer cell lines, etc.;
[0076] Expression analysis and functional identification: Evaluate the impact of SNV on gene function by methods such as protein expression analysis (such as immunoblotting, immunofluorescence), functional enzyme activity analysis, cell proliferation or apoptosis analysis, etc., so as to determine the function and possible pathogenicity of SNV.
[0077] It can be found that the following problems exist in traditional biological experimental methods:
[0078] (A) Time-consuming and expensive: Traditional biological experimental methods usually require a large amount of time and resources. For example, processes such as cloning, expression, purification, and functional identification may take weeks or even months. In addition, purchasing and maintaining experimental equipment and performing complex experimental operations usually incur extremely high costs.
[0079] (B) Dependence on cell and animal models: Many biological experimental methods require the use of cell or animal models to study the function and impact of SNV. The establishment and maintenance of these models require specialized technologies and facilities, and there may be ethical and animal protection issues.
[0080] (C) May have technical limitations: The functions of certain SNVs may be difficult to directly measure or detect by traditional biological experimental methods. For example, some SNVs may have subtle effects on protein stability, interaction partners, or cell signaling pathways, which may require advanced experimental techniques or special analysis methods to detect.
[0081] (D) Limited coverage: Traditional biological experimental methods are usually designed and implemented for specific SNVs, so their coverage is limited. This means that in large-scale SNV research, traditional experimental methods may not be able to evaluate and analyze a large number of SNVs in a high-throughput manner.
[0082] (F) May have technical difficulties and complexities: Some SNVs may involve complex mechanisms such as gene regulation, epigenetic modification, or non-coding RNA. For these complex SNVs, traditional biological experimental methods may not be able to provide sufficient solutions or technical support.
[0083] (2) Machine learning methods include the following steps:
[0084] (A) Supervised learning method:
[0085] (a) DEOGEN2 uses a Support Vector Machine (SVM) model to combine and process the predicted numbers of early folding residues of proteins. For each mutant sequence, various heterogeneous information is further integrated to obtain its early folding, and the difference from the wild-type prediction is calculated to explain the relationship between the obtained toxicity score and other mutations.
[0086] (b) REVEL uses a random forest with 1000 binary classification trees to integrate multiple predictors using different training data and features, and uses recently discovered pathogenic and neutral missense mutations that do not overlap with the predictors as the training set for prediction.
[0087] (c) CADD uses comprehensive annotations of more than 60 genomic features and a machine learning model to perform binary classification learning on simulated de novo mutations that have not been subject to natural selection and fixed mutations that have occurred in human and chimpanzee populations. Finally, a score for ranking single nucleotide variations in the reference assembly of the human genome can be obtained.
[0088] Supervised learning methods rely on labeled data, and the quality and quantity of labels have an important impact on the performance of the methods:
[0089] Difficulty in data annotation: The pathogenicity annotation of SNVs requires biological experiment verification or manual annotation by experts. This usually takes a large amount of time, resources, and manpower. An experiment can take several weeks or even months, and in addition, work such as verifying the experimental results needs to be considered to ensure the consistency and reliability of the results. When it is necessary to annotate a large-scale SNV dataset, the cost will increase significantly;
[0090] Sparse and unbalanced labels: There are approximately 3 million SNVs in a person's genome, but the number of pathogenicity labels in the database is only a few thousand. For most SNVs, especially rare mutations, we lack accurate pathogenicity labels. And the pathogenicity of SNVs is usually unbalanced, that is, the number of pathogenic SNVs is relatively small, while the number of benign SNVs is relatively large. This will lead to insufficient learning of the model for the minority class (such as pathogenicity) during the training of the supervised learning model, and it is easy to deviate;
[0091] Label bias and noise: The selection and feasibility of biological experimental methods may also affect the accuracy of labels. Certain experimental methods may not fully simulate the role of SNVs in the real biological environment, resulting in label bias. On the other hand, due to differences in genetic backgrounds, environmental factors, and disease characteristics among different patients, there may be inconsistent clinical manifestations and pathogenicity assessments for the same SNV. This heterogeneity may lead to label noise and uncertainty. Moreover, the functions and pathogenicity of SNVs are usually complex and may be affected by multiple factors. Functional annotation and pathogenicity assessment of SNVs involve complex data interpretation and integration processes, which may be subjective and error-prone.
[0092] (B) Unsupervised methods include the following steps:
[0093] Eigen uses SNVs and functional genomic annotations as datasets and models the relationship between them as a spectral problem. Spectral clustering is used to capture the relationship between different functional genomic annotations and SNVs and is used to predict the functional impact of SNVs.
[0094] EVE uses the MSA sequences of 3,218 proteins as a dataset. Then, the sequence distribution of each protein is modeled through Variational Auto Encoders (VAEs), and finally, the pathogenicity score of each single amino acid variant is obtained.
[0095] Unsupervised methods have the following problems:
[0096] (a) Eigen: Data dependence: The performance of Eigen depends on the quality and reliability of the functional genomic annotations used. If the annotations are incorrect or inaccurate, it will affect the model performance. Moreover, some annotations may lack data in specific genomic regions or specific variant types, and the sparsity of the data will affect the prediction results of the model to a certain extent.
[0097] (b) EVE:
[0098] i. Preprocessing will ignore valuable variant sites: EVE uses a gap threshold to find focused columns and focused sequences. This data preprocessing method is somewhat subjective and will ignore some variant sites with important pathogenic research value, thereby losing the corresponding data. The subsequent learned nucleotide distribution will be biased, so the evolutionary indices sampled from the distribution will not be accurate enough, which will affect the effect of Gaussian clustering to a certain extent and thus weaken the final prediction performance.
[0099] ii. Insufficiently clear classification of pathogenicity categories: Single nucleotide variations in the public dataset are divided into 5 categories: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance. However, EVE classifies both pathogenic and likely pathogenic as pathogenic, and both benign and likely benign as benign. According to the rating rules of the American College of Medical Genetics and Genomics (ACMG) and the guidelines for using diagnostic evidence of variant sites, the pathogenic probabilities of each category of variants are significantly different, and the corresponding evidence strengths also vary.
[0100] iii. Information diffusion: All input information generates a latent space representation through the mean and variance of the encoder. This may lead to the diffusion of information during the encoding process and the loss of local structure and important features in the input.
[0101] iv. Fixed weight assignment: The processing method for different input sequences is fixed and cannot be flexibly adjusted according to the characteristics and importance of the input sequences. This may result in limited encoding ability of the model for different sequences and the inability to fully capture the key features of the input.
[0102] vi. Incomplete data used: EVE uses single amino acid variant sequences of proteins, but there are a large number of synonymous mutations (which do not change the amino acid sequence) in the public dataset, and synonymous mutations are also related to the occurrence of various diseases.
[0103] Therefore, the bioinformatics-based method for classifying the pathogenicity of single nucleotide variations proposed in this implementation, as Figure 1 shown, includes four steps: data preprocessing, learning the nucleotide distribution of each DNA, calculating the evolutionary index, and classification.
[0104] Specifically:
[0105] 1. Data preprocessing:
[0106] (1) Obtain the multiple sequence alignment (MSA) file D = {D 1 , D 2 ,..., D N} for each DNA, where N is the number of sequences in the MSA.
[0107] (2) As Figure 2 shown, perform one-hot encoding on the sequences to obtain where S is the sequence length and H is the number of nucleotide types, and then transpose it to obtain
[0108] (3) Use the Squeeze-and-Excitation Network (SENet) of the channel attention mechanism to obtain the importance of each nucleotide position and each nucleotide sequence in parallel:
[0109] (A) Convert each channel into a scalar value using global average pooling:
[0110]
[0111]
[0112] That is, compress the dimension to 1×1×N.
[0113] (B) Use sparse Softmax to encode the relative importance of the one-dimensional matrix:
[0114]
[0115] where m = 1, 2 is the subscript of parameter T, corresponding to different channels; p is a vector;
[0116] where, p i is each component in vector p;
[0117]
[0118]
[0119] where [K] = {1, 2,...K}, and k is an element in set [K].
[0120] (C) Assign the relative importance as weights to the corresponding sequences and nucleotide columns on.
[0121] (4) Obtain the preprocessed input matrix D SE .
[0122] 2 Learn the nucleotide distribution of each DNA:
[0123] (1) Sample the sequences according to the reallocated weights, and the sampled batch data will pass through a one-dimensional LSTM layer.
[0124] (2) As Figure 3 shown, use a variational auto-regressive modeler with self-attention mechanism to learn the probability distribution of DNA sequences. This network has two components: an encoder with parameter φ and a decoder with parameter θ. Specifically, for a given MSA sequence S = {s 1 , s 2 ,..., s n} of DNA, assume that each sequence s i of S follows a posterior distribution Q φ (z|s i), where \(z = \{z 1 , z 2 ,..., z m \}\) is a latent variable, and it is assumed that this distribution is an independent multivariate normal distribution.
[0125] (A) Encoder: Given the MSA sequence \(S\) and assuming using the self-attention mechanism and a linear layer to fit the mean and standard deviation of \(Q φ (z|s i ), the encoding is achieved as follows:
[0126] (a) Map the sequence distributional feature representation learned by the linear layer to the query matrix, key matrix, and value matrix through cross-channel information integration and one-dimensional convolution. Specifically:
[0127] Query matrix \(Q = reshape(F CNN (D SE ; \(\theta 1 ));
[0128] Key matrix \(K = reshape(F CNN (D SE ; \(\theta 2 ));
[0129] Value matrix \(V = reshape(D SE );
[0130] (b) Perform matrix multiplication on \(Q\) and \(K\) to calculate their correlation;
[0131] (c) Use the Softmax function for normalization to obtain the attention probability distribution:
[0132]
[0133] (d) Use \(A\) as the weight coefficient of \(V\) and perform weighted summation with the original features to obtain the output of the attention module where \(\beta\) is initialized to 0 and gradually assigned a larger weight during the learning process;
[0134] (e) Obtain the mean \(\mu\) and standard deviation \(\sigma\) through a linear layer, and use the reparameterization trick to sample the latent variable \(z=\mu + rnv\cdot\sigma\), where \(rnv\) is a random variable sampled from the standard normal distribution \(N(0, 1)\).
[0135] (B) The decoder uses the autoregressive model Transformer to model the conditional distribution \(P θ (s i |z)\) to generate the approximate sequence \(s\) from the \(z\) sampled from the distribution \(Q φ (z|s i ) i i ', as follows:
[0136] (a) Concatenate the latent variable z with the input sequence D SE Perform embedding on the concatenated augmented input, convert it into a vector representation of a fixed dimension, perform positional encoding on the embedded vector, and retain the relative positions of the latent variable and the sequence;
[0137] (c) Pass through the encoder of the Transformer, which includes self-attention mechanism and feed-forward neural network, model the relationships between different positions in the sequence, and further process the representation of the sequence;
[0138] (d) Pass through the decoder of the Transformer, which includes three components: masked multi-attention mechanism, self-attention mechanism, and feed-forward neural network. Through the iteration of these three components, gradually generate an approximate sequence of DNA;
[0139] (e) Perform probability modeling on the nucleotides at each position through the output layer and the Softmax function;
[0140] (C) The loss function uses the variational likelihood Evidence Lower Bound (ELBO):
[0141]
[0142] 3 Calculate the evolutionary index
[0143] Sample from the approximate posterior distribution Q φ (S|z) learned in step 2, and calculate the evolutionary index of each SNV sequence s i where the evolutionary index is the log-likelihood difference between s i and the wild-type sequence w i . In fact, the exact log-likelihood calculation is difficult to handle, so ELBO is used for approximation:
[0144]
[0145] 4 Classification
[0146] Use the Gaussian Mixture Model (GMM) to fit the evolutionary indices into four clusters.
[0147] (1) Train a global four-component GMM on the evolutionary indices of all DNA.
[0148] (2) Initialize the GMM for specific DNA using the parameters of the global GMM. Among the four resulting clusters, the cluster with a higher mean contains SNV sequences with a higher evolutionary index, that is, the pathogenic probability of this sequence is higher.
[0149] (3) Calculate the MVMA score of s according to the global GMM and the GMM of specific DNA, that is: MVMA s = α * p(X s =(1,0),(1,1)|E s , θ p )+(1 - α) * p(X s =(0,1),(0,0)|E s , θ g );
[0150] Among them, X s is a two - bit binary random variable, α is the relative weight of the GMM of specific DNA in the global GMM, θ p and θ g are the parameters of the GMM of specific DNA and the global GMM respectively. This MAE score quantifies the pathogenic tendency of s.
[0151] (4) Specifically, the classification in step (3) here is a forced classification, and this embodiment allows classification results with uncertain meanings. Measure the total uncertainty of the clustering assignment of DNA by using the prediction entropy PE:
[0152]
[0153] Set the 25% of SNV categories with the highest uncertainty as uncertain, which will help improve the accuracy of classification.
[0154] Figure 4 In:
[0155] For each protein, for each protein.
[0156] Bayesian variational autoencoder, Bayesian variational autoencoder.
[0157] Inferring constraints at each protein by learning the distribution of sequences in evolutionary data, Inferring constraints at each protein by learning the distribution of sequences in evolutionary data.
[0158] One-hot encoding of MSA sequences, the one-hot encoding of MSA sequences.
[0159] VAE reconstruction, the reconstruction of the variational autoencoder.
[0160] We sample from the approx posterior, We sample from the approximate posterior distribution.
[0161] Evolutionary index, Evolutionary index.
[0162] Approximating the negative log-likelihood ratio of mutant versus wildtype, Approximating the negative log-likelihood ratio of mutant compared to wildtype.
[0163] Gaussian mixture model, Gaussian mixture model.
[0164] Computing EVE pathogenicity scores and filtering out most uncertain predictions, Calculating EVE pathogenicity scores and filtering out the most uncertain predictions.
[0165] Effect comparison:
[0166] 1. MVMA uses a channel attention mechanism that integrates sparse Softmax during preprocessing
[0167] As Figure 4 shown, EVE (evolutionary model of variant effect) does not present preprocessing as a separate step in the modeling strategy diagram. It calculates the number of gaps and sets a gap threshold α gap_sequence 、α gap_column to determine the focused sequences and focused columns of the input sequence. For all MSA sequences of 3218 protein families, it calculates the number of gaps in each sequence and each column, deletes the sequences with a gap ratio exceeding α gap_sequence , and takes the columns with a gap ratio exceeding α gap_column as non-focused columns, lowercase the amino acids on the non-focused columns. Perform one-hot encoding on the sequences and use a formula to re-weight the sequences to correct data bias.
[0168] 2. MVMA uses VAR when learning sequence distribution, and EVE uses Bayesian VAE to learn the sequence distribution of proteins, where the encoder of Bayesian VAE is a fully connected neural network and the decoder part is a Bayesian neural network.
[0169] 3. MVMA refines SNV into five categories, while EVE classifies single amino acid variant sequences into three categories. EV uses Gaussian clustering to separate pathogenic and benign variants: a global Gaussian with two components is fitted on the evolutionary index of all variant sequences, and a local Gaussian with two components is fitted on the evolutionary index of a specific protein. The EVE score is calculated to quantify the pathogenicity score of a given variant sequence. The predicted entropy is then used to predict the uncertainty of the cluster assignment, and the category of the variant sequence with the highest 25% uncertainty is set to uncertain.
[0170] The preprocessing of this embodiment uses a parallel channel attention mechanism, which can weight the importance of sequences and columns at the same time, improve the attention to key sequences and columns, and thus better capture the important features of the sequence.
[0171] This embodiment integrates sparse Softmax with SENet. Through sparsity constraints, the key parts of the sequence can be further highlighted, thereby enhancing the ability to learn important information.
[0172] This embodiment uses a one-dimensional LSTM to process the input data before passing it to the encoder, which can help the model learn the context information and sequence dependency in the sequence data, thereby better modeling the distribution of DNA sequences.
[0173] This embodiment adds a self-attention mechanism to the encoder, which can aggregate similar network nodes, calculate potential features, and integrate context information, which can effectively avoid the information attenuation problem in the long sequence modeling process, thereby enhancing the model's ability to capture long-term dependencies and improving the quality of the final generated sequence.
[0174] This embodiment uses Transformer as the decoder of VAR, which can calculate information at different positions in the sequence in parallel and allow each position to directly pay attention to other positions in the sequence to perform global information interaction, which is helpful for learning the distribution of the sequence.
[0175] This embodiment performs refined classification of SNVs, and provides more fine-grained classification results according to the ACMG guidelines and the standards of public data sets, so as to facilitate the analysis of SNVs in the fields of medicine and biological research.
[0176] Therefore, the method provided in this embodiment is applicable to the learning of DNA sequences and is easy to popularize. It does not require the operation of professional biological researchers, without prior induction preparation and long-term tracking experiments. By using the self-attention mechanism and variational auto-regressive model, it can better capture the important patterns and distributions of DNA sequences and is easier to migrate to a large amount of DNA sequence data compared to high-throughput biological experimental techniques.
[0177] No labeled data is required for classification: The method provided in this embodiment uses an unsupervised deep generative model VAE. Compared with supervised models, it does not need to use sparse, biased, and noisy labeled data.
[0178] The parallel channel attention mechanism is used: It can simultaneously perform importance weighting on sequences and columns, improving the attention to key sequences and columns, and thus better capturing the important features of the sequences.
[0179] Autonomously find the focused columns and focused sequences: During data preprocessing, the gap threshold is abandoned, and instead, the parallel channel attention mechanism is used to find the focused columns and focused sequences, and re-weight each sequence. In particular, sparse Softmax is used instead of the original activation function Softmax in the attention mechanism. This operation only retains the top k calculated probabilities, where k is a hyperparameter manually selected, sparsifying the result of Softmax. This helps to avoid overfitting and improve the accuracy of subsequent classification.
[0180] Reduce the information decay during the long sequence modeling process: High-quality biological sequences are as long as possible and have a large proportion of non-gap parts. Therefore, a self-attention mechanism is added to the encoder part, combined with the LSTM and Transformer architectures, which can aggregate similar network nodes, calculate potential features, and integrate context information at the same time, effectively avoiding the information decay problem during the long sequence modeling process, thereby enhancing the model's ability to capture long-term dependencies and improving the quality of the finally generated sequences.
[0181] Combine multiple models and methods to improve the model's expressive ability: Combine the channel attention mechanism SENet, one-dimensional LSTM layer, and improved VAR, and utilize the advantages of these models and methods: SENet can provide channel-level importance weights, the LSTM layer can process the temporal information of sequences, and the variational auto-regressive model can learn the distribution characteristics of sequences. This strategy of comprehensively using different models and methods can improve the model's expressive ability to a certain extent.
[0182] Increased data volume and refined class division: Compared with EVE, more SNV mutation data including synonymous mutation sequences is used, expanding the dataset, and the classification categories are refined according to the ACMG guidelines and public datasets.
[0183] Embodiment 2:
[0184] A system for implementing the above method, comprising:
[0185] A preprocessing module, configured to: obtain a multiple sequence alignment of each DNA, obtain the importance of each nucleotide position and each nucleotide sequence through preprocessing, and assign the obtained importance as weights to the corresponding nucleotide sequences to form an input matrix;
[0186] A DNA distribution learning module, configured to: sample DNA sequences according to the obtained input matrix, and learn the probability distribution of DNA sequences based on a variational auto-regressive model with a self-attention mechanism;
[0187] An evolutionary index module, configured to: calculate the evolutionary index of each single nucleotide variant sequence according to the obtained probability distribution, and the evolutionary index is the log-likelihood difference between the single nucleotide variant sequence and the wild-type sequence;
[0188] A classification module, configured to: fit the obtained evolutionary indices into multiple clusters, respectively corresponding to pathogenic probabilities in different intervals.
[0189] Example 3:
[0190] This example provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the method for classifying the pathogenicity of single nucleotide variants based on biological information as described in Example 1 above.
[0191] Example 4:
[0192] This example provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for classifying the pathogenicity of single nucleotide variants based on biological information as described in Example 1 above.
[0193] The steps or networks involved in Examples 2 to 4 above correspond to those in Example 1, and the specific implementation details can be referred to the relevant description part of Example 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0194] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for classifying the pathogenicity of single nucleotide variations based on biological information, characterized in that, it includes the following steps: Obtain the multiple sequence alignment of each DNA, and through preprocessing, obtain the position of each nucleotide and the importance of each nucleotide sequence. Assign the obtained importance as weights to the corresponding nucleotide sequences to form an input matrix; Sample the DNA sequences according to the obtained input matrix, and based on a variational auto-regressive model with a self-attention mechanism, learn the probability distribution of the DNA sequences; Calculate the evolutionary index of each single nucleotide variant sequence according to the obtained probability distribution. The evolutionary index is the logarithmic likelihood difference between the single nucleotide variant sequence and the wild-type sequence; Fit the obtained evolutionary indices into multiple clusters, which respectively correspond to the pathogenicity probabilities of the single nucleotide variant sequences. Divide the pathogenicity probabilities of the single nucleotide variant sequences into five categories: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance. Specifically: Train a Gaussian mixture model on the evolutionary indices of all DNAs, that is, a global Gaussian. The parameters of this global model are used to initialize the Gaussian mixture model of a specific DNA, that is, a local Gaussian; Calculate the pathogenicity score of the single nucleotide variant sequences of this specific DNA according to the global and local Gaussians; The obtained pathogenicity scores divide the single nucleotide variant sequences into four categories: benign, likely benign, likely pathogenic, and pathogenic; Calculate the uncertainty of the class division of each single nucleotide variant sequence based on the prediction entropy, and use the class with the highest uncertainty and a set percentage as of uncertain significance.
2. The method for classifying the pathogenicity of single nucleotide variations based on biological information according to claim 1, characterized in that, Obtain the multiple sequence alignment of each DNA, and through preprocessing, obtain the position of each nucleotide and the importance of each nucleotide sequence. Specifically: Obtain the multiple sequence alignment file of each DNA, perform one-hot encoding on the sequences, where the encoding includes at least the sequence length and the number of nucleotide types, and perform a transpose process; Based on the channel attention mechanism for parallel processing, obtain the position of each nucleotide and the importance of each nucleotide sequence.
3. The method for classifying the pathogenicity of single nucleotide variations based on biological information according to claim 2, characterized in that, Based on the channel attention mechanism for parallel processing, obtain the position of each nucleotide and the importance of each nucleotide sequence. Assign the obtained importance as weights to the corresponding nucleotide sequences to form an input matrix; Specifically: Based on global average pooling, compress the dimension of each channel to one dimension; Based on sparse Softmax encoding, obtain the relative importance of the one-dimensional matrix; Use the obtained relative importance as weights and assign them to the corresponding sequences and nucleotide columns to obtain the preprocessed data as the input matrix.
4. The method for classifying the pathogenicity of single nucleotide variations based on biological information according to claim 1, characterized in that, The variational auto-regressive model with a self-attention mechanism includes an encoder and a decoder corresponding to the parameters; The encoder is based on a given multiple sequence alignment S In which, each sequence Si obeys a posterior distribution, and the mean and standard deviation of the posterior distribution are fitted based on the self-attention mechanism and the linear layer to achieve encoding; The decoder uses an autoregressive model to reconstruct the input latent variable into an approximate sequence and outputs an approximate probability distribution of the DNA sequence.
5. The method for classifying the pathogenicity of single nucleotide variations based on biological information according to claim 4, characterized in that, The specific process of the encoder for encoding is: Map the sequence distributed feature representation learned by the linear layer to the query matrix, key matrix, and value matrix through cross-channel information integration and one-dimensional convolution; Perform matrix multiplication on the query matrix and the key matrix to calculate their correlation; Normalize to obtain the attention probability distribution; Use the attention probability distribution as the weight coefficient of the value matrix and perform weighted summation with the original features to obtain the output of the attention module; Obtain the mean through a linear layer and the standard deviation , and sample the latent variable using the reparameterization trick , where rnv is a random variable sampled from the standard normal distribution .
6. The method for classifying the pathogenicity of single nucleotide variants based on biological information according to claim 4, characterized in that, The specific process of the decoder outputting the probability distribution of the DNA sequence is as follows: Concatenate the latent variable with the input sequence, embed the concatenated augmented input, convert it into a vector representation of a fixed dimension, perform positional encoding on the embedded vector, and retain the relative positions of the latent variable and the sequence; The encoder models the relationships between different positions in the sequence and further processes the representation of the sequence; Through the iteration of the masked multi-attention mechanism, self-attention mechanism, and feed-forward neural network in the decoder, an approximate sequence of DNA is gradually generated; The probability of each nucleotide at each position is modeled through the output layer.
7. A system for classifying the pathogenicity of single nucleotide variants based on biological information, characterized in that, comprising: A preprocessing module configured to: obtain the multiple sequence alignment of each DNA, obtain the importance of each nucleotide position and each nucleotide sequence through preprocessing, and assign the obtained importance as weights to the corresponding nucleotide sequences to form an input matrix; A DNA distribution learning module configured to: sample the DNA sequence according to the obtained input matrix, and learn the probability distribution of the DNA sequence based on a variational auto-regressive modeler with a self-attention mechanism; An evolutionary index module configured to: calculate the evolutionary index of each single nucleotide variant sequence according to the obtained probability distribution, and the evolutionary index is the logarithmic likelihood difference between the single nucleotide variant sequence and the wild-type sequence; A classification module configured to: fit the obtained evolutionary indices into multiple clusters, corresponding to pathogenicity probabilities in different intervals respectively, and divide the pathogenicity probabilities of single nucleotide variant sequences into five categories: benign, likely benign, likely pathogenic, pathogenic, and of uncertain significance; specifically: Train a Gaussian mixture model on the evolutionary indices of all DNAs, that is, a global Gaussian, and the parameters of this global model are used to initialize the Gaussian mixture model of a specific DNA, that is, a local Gaussian; Calculate the pathogenicity score of the single nucleotide variant sequence of this specific DNA according to the global and local Gaussians; The obtained pathogenicity score divides the single nucleotide variant sequence into four categories: benign, likely benign, likely pathogenic, and pathogenic; Calculate the uncertainty of the class division of each single nucleotide variant sequence based on the prediction entropy, and use the class with the highest uncertainty and a set percentage as of uncertain significance.
8. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the method for classifying the pathogenicity of single nucleotide variants based on biological information according to any one of claims 1-6 above.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the method for classifying the pathogenicity of single nucleotide variants based on biological information according to any one of claims 1-6.
Citation Information
Patent Citations
Automatic distinguishing system for autism spectrum disorder, storage medium and equipment
CN113724863A
Knowledge enhancement-based text generation model and training method thereof
CN115345169A