Method and System for Predicting the Pathogenicity of Splicing Variants Based on Contrastive Learning
Through a method based on contrast learning, the embedding representation is directly extracted from the DNA sequence information, and the embedding distance is adjusted for splicing variant pathogenicity prediction, solving the problems of time-consuming and easy noise introduction in the prior art, and achieving efficient and accurate pathogenic prediction.
Patent Information
- Application Number
- CN202411990089.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The prior art requires reference to a large number of complex biological characteristics in the prediction of pathogenicity of splicing variants, which leads to a long time consumption and easy introduction of noise, making it difficult to meet the needs of efficient large-scale gene screening.
Using a method based on contrast learning, the embedded representation is directly extracted from the DNA sequence information through the DNA pre-trained model and the contrast learning module, and the embedding distance is adjusted through the contrast learning module, simplifying feature input, and using the natural relative characteristics between variants for pathogenic prediction.
It realizes efficient and accurate pathogenic prediction, reduces model bias caused by feature selection, improves the generalization and robustness of the model, and is suitable for processing large-scale gene data.
Smart Images

Figure CN119400239B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and particularly to a method and system for predicting the pathogenicity of splicing variants based on contrastive learning. Background Art
[0002] RNA splicing is a process of removing introns from pre-mRNA and ligating exons to obtain mature mRNA, which is controlled by a variety of cis-acting and trans-acting regulatory elements. When these regulatory elements are affected by mutations, the spliceosome may misidentify splice sites, generating abnormal splicing variants. Splicing variants may be associated with rare pathogenic variants, cancers, or various genetic diseases. The prediction of the pathogenicity of splicing variants is of great significance for genetic counseling and clinical treatment.
[0003] Traditional biological methods are difficult to simulate the splicing environment when identifying the pathogenicity of splicing variants and cannot process a large amount of data information, resulting in difficulty in accurate assessment.
[0004] Some existing prediction models based on machine learning need to consider multiple complex biological features such as sequence conservation, splice site strength, and secondary structure simultaneously. These biological features are obtained through biological experiments or other bioinformatics tools, which consume a large amount of time and are extremely prone to introducing noise. Summary of the Invention
[0005] The method and system for predicting the pathogenicity of splicing variants based on contrastive learning provided by the embodiments of the present invention at least solve the problems that predicting the pathogenicity of splicing variants requires referring to a large number of complex biological features, consuming a large amount of time, and being extremely prone to introducing noise.
[0006] In a first aspect, the embodiments of the present invention provide a method for predicting the pathogenicity of splicing variants based on contrastive learning, including:
[0007] Inputting the original sequence information and the variant sequence information, and respectively obtaining a first original embedding vector and a first variant embedding vector; wherein, the first original embedding vector is obtained according to the original sequence information, and the first variant embedding vector is obtained according to the variant sequence information;
[0008] Inputting the first original embedding vector and the first variant embedding vector into a contrastive learning module, adjusting the embedding distances of the first original embedding vector and the first variant embedding vector in the contrastive learning module, and outputting a second original embedding vector and a second variant embedding vector; wherein, the second original embedding vector is obtained according to the first original embedding vector, and the second variant embedding vector is obtained according to the first variant embedding vector;
[0009] Perform feature fusion on the second original embedding vector and the second mutated embedding vector to obtain a fusion tensor;
[0010] Input the fusion tensor into a classifier for classification prediction, and output a pathogenicity prediction result.
[0011] Before inputting the original sequence information and the mutated sequence information, the method for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiments of the present invention further includes:
[0012] Create a classification prediction model according to a DNA pre-training model, the contrast learning module, and the classifier; wherein, the DNA pre-training model is used to extract corresponding embedding representations according to sequence information;
[0013] Train the DNA pre-training model and the contrast learning module according to training data in the case of not connecting the classifier; wherein, the training data includes sample pairs composed of multiple groups of the sequence information;
[0014] Freeze the DNA pre-training model and the contrast learning module, connect the classifier, and train the classifier according to the training data.
[0015] The method for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiments of the present invention, training the DNA pre-training model and the contrast learning module according to training data includes:
[0016] Input the training data into the DNA pre-training model and perform max pooling to obtain a group of embedding representations; wherein, the DNA pre-training model includes initial pre-training parameters;
[0017] Input the group of embedding representations into the contrast learning module to obtain a group of embedding vectors;
[0018] Calculate the contrast learning loss of the group of embedding vectors using a contrast learning loss function in the contrast learning module; wherein, the formula of the contrast learning loss function is:
[0019] ;
[0020] ;
[0021] ;
[0022] is the total contrast learning loss; is the contrast learning loss of positive sample pairs, is the contrast learning loss of negative sample pairs; is the label of the sample pair; is the Euclidean distance between the pre- and post-mutation embedded vectors in the embedded vector group, is the maximum allowable Euclidean distance in the negative sample pair, is the number of sample pairs in the training data;
[0023] Train the DNA pre-training model and the contrastive learning module until the total contrastive learning loss converges; obtain the trained DNA pre-training model and the contrastive learning module.
[0024] The method for predicting the pathogenicity of splicing variants based on contrastive learning provided by the embodiments of the present invention, training the classifier according to the training data includes:
[0025] Input the training data into the trained DNA pre-training model and take the maximum pooling;
[0026] Input the result of the maximum pooling into the trained contrastive learning module;
[0027] Input the output result of the trained contrastive learning module into the classifier; wherein, the classifier uses the cross-entropy loss function;
[0028] Train the classifier until the loss of the classification prediction model converges, and obtain the trained classifier.
[0029] The method for predicting the pathogenicity of splicing variants based on contrastive learning provided by the embodiments of the present invention, inputting the original sequence information and the mutated sequence information, and respectively obtaining the first original embedded vector and the first mutated embedded vector includes:
[0030] Input the original sequence information and the mutated sequence information into the trained DNA pre-training model, and respectively obtain the first original embedded representation and the first mutated embedded representation;
[0031] Take the maximum pooling of the first original embedded representation and the first mutated embedded representation, and respectively obtain the first original embedded vector and the first mutated embedded vector; wherein, both the first original embedded vector and the first mutated embedded vector are 768-dimensional vectors.
[0032] The method for predicting the pathogenicity of splicing variants based on contrastive learning provided by the embodiments of the present invention, adjusting the embedding distance between the first original embedded vector and the first mutated embedded vector in the contrastive learning module, and outputting the second original embedded vector and the second mutated embedded vector includes:
[0033] Adjust the embedding distance between the first original embedding vector and the first variant embedding vector according to different tags; wherein, the embedding distance is set as the Euclidean distance;
[0034] If the splicing variant is a pathogenic variant, make the Euclidean distance between the second original embedding vector and the second variant embedding vector greater than the Euclidean distance between the first original embedding vector and the first variant embedding vector;
[0035] If the splicing variant is a benign variant, make the Euclidean distance between the second original embedding vector and the second variant embedding vector less than the Euclidean distance between the first original embedding vector and the first variant embedding vector.
[0036] The method for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiments of the present invention creates a feature fusion of the second original embedding vector and the second variant embedding vector, and the obtained fusion tensor includes:
[0037] Perform a stacking operation on the second original embedding vector and the second variant embedding vector to complete feature fusion, and obtain the fusion tensor; wherein, both the second original embedding vector and the second variant embedding vector are 128-dimensional vectors, and the fusion tensor is a two-dimensional tensor of (2, 128).
[0038] The method for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiments of the present invention creates inputs the fusion tensor into a classifier for classification prediction, and the output pathogenicity prediction results include:
[0039] Input the fusion tensor into the trained classifier for binary classification prediction, and perform non-linear division through the activation function of the classifier to output the classification result; wherein, the binary classification prediction result of the classifier is 0 or 1;
[0040] If the binary classification prediction result is 1, the pathogenicity prediction result is that the splicing variant has a pathogenic variant;
[0041] If the binary classification prediction result is 0, the pathogenicity prediction result is that the splicing variant has a benign variant.
[0042] In a second aspect, the embodiments of the present invention also provide a system for predicting the pathogenicity of splicing variants based on contrast learning, including: a user input module, a sequence information matching module, a splicing variant pathogenicity prediction module, and a result feedback and notification module; wherein, the splicing variant pathogenicity prediction module includes the method for predicting the pathogenicity of splicing variants based on contrast learning described in any of the above embodiments;
[0043] The user input module is used to input genomic sequence information;
[0044] The sequence information matching module is used to identify the variant sites in the genomic sequence information, and extract the DNA sequence information centered on the variant site before and after the variation, with a window size of 1024 bases, as the original sequence information and the variant sequence information respectively, and then input the original sequence information and the variant sequence information into the splicing variant pathogenicity prediction module;
[0045] The splicing variant pathogenicity prediction module outputs a pathogenicity prediction result according to the original sequence information and the variant sequence information, and transmits the pathogenicity prediction result to the result feedback and notification module;
[0046] The result feedback and notification module is used to collect and record the pathogenicity prediction result, and display the pathogenicity prediction result.
[0047] In a third aspect, an embodiment of the present invention also provides an electronic device, including: a processor, and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method for predicting the pathogenicity of splicing variants based on contrast learning according to any one of the above embodiments.
[0048] The method and system for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiment of the present invention use sequence information as input to achieve the prediction of the pathogenicity of splicing variants completely based on sequence information. Compared with the existing methods, it can avoid introducing other biological features such as structural conservation and splicing site strength. On the one hand, directly learning from the sequence can reduce the model bias caused by different feature selections, and thus better adapt to new variant types and data sets, with better model generalization and robustness. On the other hand, it greatly simplifies the feature input of the model, does not rely on complex feature engineering, avoids the cumbersome process of screening and optimizing among various biological features, makes the model more automated and efficient. It reduces the computational time and the complexity of feature engineering, and is suitable for processing large-scale gene data.
[0049] The method and system for predicting the pathogenicity of splicing variants based on contrast learning provided by the embodiment of the present invention apply contrast learning to the variant prediction task, and utilize the natural relative characteristics between variants to achieve more accurate pathogenicity prediction. Through contrast learning, the model enhances its learning ability of key features in the case of similar variant characteristics, thus exceeding the traditional classification prediction model, effectively identifying the subtle differences between different pathogenic variants, and improving the prediction effect of the model. Description of the Drawings
[0050] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can also be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the principle of the classification prediction model in the splicing variant pathogenicity prediction method based on contrast learning in the first embodiment of the present invention.
[0052] Figure 2 It is a flowchart of the splicing variant pathogenicity prediction method based on contrast learning in the first embodiment of the present invention.
[0053] Figure 3 It is a schematic diagram of the splicing variant pathogenicity prediction system based on contrast learning in the second embodiment of the present invention. Detailed implementation manners
[0054] The following will describe the embodiments of the present invention in more detail with reference to the drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention. Embodiment 1
[0055] Genes in eukaryotes are chimeric, with coding regions separated by non-coding regions. Among them, the coding regions are called exons, and the non-coding regions are called introns. Removing introns and joining exons together in mature mRNA is called RNA splicing. Splicing variants mainly include exon skipping, 5′ splice site alteration, 3′ splice site alteration, intron retention, and exon mutual exclusion according to different forms.
[0056] The prior art has confirmed that splicing variants are related to the occurrence of various diseases, including various genetic diseases, neurological diseases, and cancers. In cancers, splicing variants may cause the activation of oncogenes or the inactivation of tumor suppressor genes, promoting the growth and spread of tumors. In some neurodegenerative diseases, abnormal splicing processes affect the normal functions of neurons, thereby aggravating the progression of the diseases.
[0057] Some pathogenic variants will affect the function of cis-regulatory elements, thus disrupting the normal splicing process. Approximately 15-60% of rare pathogenic variants are located in these cis-regulatory regions, leading to abnormal gene expression by generating new splicing sites or interfering with the use of normal splicing sites.
[0058] In addition, splicing variants also have other uncertain clinical significances in diagnosis, making it difficult to accurately determine their pathogenicity. These uncertainties pose challenges to genetic diagnosis and gene screening and may affect the formulation of genetic counseling and treatment plans.
[0059] Therefore, splicing variants pose a great harm to human health, and predicting the pathogenicity of splicing variants is crucial for disease prevention and treatment.
[0060] Traditional prediction of the pathogenicity of splicing variants is carried out through biological methods. Traditional biological methods include Reverse Transcription-Polymerase Chain Reaction (RT-PCR), RNA sequencing (RNA-seq), and minigene assay, etc. Although these biological methods can provide direct experimental evidence for the pathogenicity of splicing variants, they usually take a long time and are costly, making it difficult to be popularized and used in large-scale gene screening. Traditional methods are difficult to perform high-throughput variant screening and analysis and cannot quickly process a large number of samples. With the development of genome sequencing technology, the generation of massive data requires more efficient analysis methods, and traditional methods obviously cannot meet this demand.
[0061] In addition, the splicing process is tissue- or cell-type specific, and the splicing patterns in different tissues may be different. It is difficult to fully simulate the splicing environment of specific tissues or cell types in in vitro experiments using traditional biological methods, resulting in experimental results that may not represent the actual situation in vivo. Especially some key disease-related splicing events may only occur in specific tissues. Many splicing variants are located in deep intronic regions or non-classical splicing sites, and it is difficult for traditional experimental methods to accurately evaluate their effects on splicing. These variants are usually considered to be benign or variants of uncertain significance (VUS), and traditional biological methods are difficult to interpret these variants and cannot provide a clear pathogenicity judgment.
[0062] In view of the limitations of traditional biological methods in processing power and experimental results, machine learning has been introduced into the pathogenicity prediction of splicing variants in the prior art. The prediction model CADD uses up to 60 biological features, covering conservation and structural features, etc. Machine learning algorithms such as Support Vector Machine (SVM) are used to train the model to integrate these different types of biological features, and finally generate prediction scores.
[0063] dbscSNV utilizes ensemble learning algorithms, including the iterative algorithm AdaBoost and RandomForest, and combines multiple features such as splice site strength, sequence conservation, and secondary structure information. However, this algorithm is only applicable to identifying the impact of Single Nucleotide Variations (SNV) on splice sites.
[0064] These machine learning-based algorithms require the pre-computation of a variety of complex biological features, such as sequence conservation, splice site strength, secondary structure, and the ratio of guanine and cytosine in the region (GC content). The calculation and selection of these features not only require a large amount of time, but also have high professional requirements for feature engineering. Usually, it requires the support of genomic annotation databases and complex computing resources. Some of these biological features do not directly come from experimental results, but rely on other bioinformatics tools. A large number of biological features related to splicing need to be calculated in advance, introducing additional noise and affecting the prediction efficiency and accuracy of the model.
[0065] With the continuous development of deep learning, architectures such as convolutional neural networks, recurrent neural networks, and ResNet have begun to be applied to the pathogenicity prediction of splicing variants. SpliceAI developed by the DeepVariant team is a deep learning-based splicing variant prediction tool that uses a convolutional neural network to process the nucleotide sequence information of the variant context. The maximum range of the nucleotide sequence information of the variant context is 10 kb. Although this can extract local features of the sequence to a certain extent, it is difficult to predict and judge long-range dependencies, destroying the relevance and integrity of the information.
[0066] Accordingly, Embodiment 1 of the present invention provides a method for predicting the pathogenicity of splicing variants based on contrastive learning, which uses only the DNA sequences before and after the splicing variant mutation as input. Based on the DNA pre-trained model, embeddings are obtained respectively. Utilizing the prior knowledge of the DNA pre-trained model, high-dimensional and multi-level nucleotide local features and global information are directly extracted from the original sequences, comprehensively considering the influence of the splicing process on the nucleotide sequence before and after mutation and the existence state of biological elements in the DNA sequence. The embeddings obtained through the DNA pre-trained model are further processed through a contrastive learning module respectively. For the variants that show pathogenicity after mutation, the contrastive learning module processes them into more distinct embeddings; for the variants that are benign after mutation, the difference between the two is reduced, making it easier for the classifier to capture the difference, thereby achieving efficient pathogenicity prediction.
[0067] Compared with traditional biological detection methods, the present invention can realize the analysis and prediction of massive data and meet the requirements of large-scale gene screening. Compared with machine learning prediction models, it can avoid calculating a large number of biological features related to splicing and mutation, with short time consumption and no additional noise introduced, and has higher prediction efficiency and accuracy.
[0068] The method for predicting the pathogenicity of splicing variants based on contrastive learning provided in Embodiment 1 of the present invention includes the following steps.
[0069] Step S100, create a classification prediction model and train the classification prediction model.
[0070] As an implementable manner, step S100 specifically includes the following steps:
[0071] Step S110, create a classification prediction model according to the DNA pre-trained model, the contrastive learning module and the classifier.
[0072] Among them, the DNA pre-trained model is used to extract the corresponding embedding representation according to the sequence information. In Embodiment 1, the DNA pre-trained model is preferably set to DNABert-2.
[0073] To address the problem that traditional languages have difficulty capturing information between DNA semantics, Professor Ramana V. Davuluri of Northwestern University and others published an article "DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome" in the Bioinfomatics journal. The authors of the article proposed a pre-trained bidirectional encoding representation DNABERT to perform global or transfer analysis on DNA sequences through context information. DNABERT can directly rank the importance of nucleotide molecules and analyze the relationships between input sequence contexts, thereby obtaining better visualization information and accurate motif extraction. DNABert-2 is a publicly available follow-up study of DNABERT, which improves the word segmentation efficiency, computational efficiency, and reduces the problems of sampling redundancy and inefficiency.
[0074] In other embodiments of the present invention, the DNA pre-trained model can also use other models that meet the requirements of sequence information extraction, not limited to DNABert-2.
[0075] Step S120, without connecting to a classifier, train the DNA pre-trained model and the contrast learning module according to the training data.
[0076] For the classification prediction model, in step S120, only connect the DNA pre-trained model and the contrast learning module, and do not connect to a classifier.
[0077] The training data includes sample pairs composed of multiple sets of sequence information. In the sample pairs, the sample pairs with pathogenicity after splicing variation are set as positive sample pairs, the corresponding expected output result for the positive sample pairs is pathogenicity, and the label of the positive sample pairs . In the sample pairs, the sample pairs that are benign after splicing variation are set as negative sample pairs, the corresponding expected output result for the negative sample pairs is benign, and the label of the negative sample pairs .
[0078] Input the training data into the DNA pre-trained model and take the maximum pooling to obtain a group of embedded representations of the training data; among them, the DNA pre-trained model includes initial pre-trained parameters.
[0079] Input the group of embedded representations of the training data into the contrast learning module to obtain a group of embedded vectors of the training data.
[0080] In the contrast learning module, use the contrast learning loss function to calculate the contrast learning loss of the group of embedded vectors of the training data. The contrast learning loss function includes the loss of positive sample pairs and the loss of negative sample pairs.
[0081] The loss calculation formula for positive sample pairs is as follows:
[0082] ;
[0083] ;
[0084] Among them, is the contrastive learning loss for positive sample pairs, is the label of the sample pair. For positive sample pairs, .
[0085] is the Euclidean distance between the pre- and post-mutation embedding vectors in the embedding vector group of the training data, is the embedding vector of the original sequence before mutation in the training data, is the embedding vector of the mutated sequence after mutation in the training data.
[0086] The loss calculation formula for negative sample pairs is as follows:
[0087] ;
[0088] Among them, is the contrastive learning loss for negative sample pairs, is the label of the sample pair. For negative sample pairs, . is the maximum allowable Euclidean distance between the pre- and post-mutation embedding vectors in the negative sample pair.
[0089] The loss calculation formula for the total contrastive learning loss is as follows:
[0090] .
[0091] Among them, is the total contrastive learning loss. is the number of sample pairs in the training data.
[0092] Train the DNA pre-training model and the contrastive learning module until the total contrastive learning loss converges, then the training of the DNA pre-training model and the contrastive learning module is completed. The initial pre-training parameters in the DNA pre-training model change, and the trained DNA pre-training model and contrastive learning module are obtained as the training results.
[0093] Step S130: Freeze the DNA pre-training model and the contrastive learning module according to the training results of Step S120. After freezing, connect a classifier and train the classifier. Step S130 specifically includes:
[0094] Input the training data in step S120 into the DNA pre-training model trained in step S120 and perform max pooling; input the result of max pooling into the contrastive learning module trained in step S120.
[0095] Input the output result of the contrastive learning module trained in step S120 into the classifier. Among them, the classifier uses the cross-entropy loss function.
[0096] Preferably, the cross-entropy loss function used by the classifier is the Binary CrossEntropy Loss (BCELoss) function. The binary cross-entropy loss function is one of the most commonly used loss functions for dealing with binary classification problems. The binary cross-entropy loss function is used to calculate the difference between the probability distribution output by the model and the probability distribution of the true label. The formula of the binary cross-entropy loss function is:
[0097] ;
[0098] Among them, is the true label, is the predicted probability. When the predicted probability is close to the true label, the value of the binary cross-entropy will decrease; otherwise, it will increase. This enables the classification prediction model to improve the prediction accuracy by minimizing the loss function.
[0099] Train the classifier until the loss of the classification prediction model converges to obtain the trained classifier.
[0100] Step S140, connect the DNA pre-training model and the contrastive learning module trained in step S120 to the classifier trained in step S130 to complete the training of the classification prediction model.
[0101] In this embodiment, a step-by-step freezing training framework is designed in step S100. First, train the DNA pre-training model and the contrastive learning module in stages. After the loss function of the contrastive learning converges, freeze the relevant parameters, and then connect the classifier to perform the classification task. Only consider the loss of the classifier for backpropagation and train until the model reaches the optimal performance, so as to improve the training effect and convergence speed of the entire classification prediction model.
[0102] The step-by-step freezing training strategy can effectively ensure that the pre-training model and the contrastive learning module stably generate different embedding representations under the guidance of labels, thus realizing the functional decoupling and complementarity between the contrastive learning module and the classifier inside the model. By fixing the model parameters of the pre-training model and the contrastive learning part, this strategy reduces the overfitting risk of the pre-training model to downstream tasks and enhances the sensitivity of the contrastive learning module to label information. It not only improves the overall prediction accuracy of the model but also achieves better performance balance between positive and negative samples, and is applicable to pathogenicity prediction tasks of various complex mutation types.
[0103] Step S200, input the original sequence information and the mutant sequence information, and obtain the first original embedding vector and the first mutant embedding vector respectively; wherein, the first original embedding vector is obtained according to the original sequence information, and the first mutant embedding vector is obtained according to the mutant sequence information.
[0104] As an implementable manner, step S200 specifically includes the following steps:
[0105] Step S210, input the original sequence information before mutation and the mutant sequence information after mutation into the DNA pre-trained model, and obtain the first original embedding representation and the first mutant embedding representation respectively.
[0106] Among them, the DNA pre-trained model is in the classification prediction model obtained in step S100, specifically the result obtained after being trained in step S120. The first original embedding representation is output by the DNA pre-trained model from the original sequence information before mutation, and the first mutant embedding representation is output by the DNA pre-trained model from the mutant sequence information after mutation.
[0107] Step S220, perform max pooling on the first original embedding representation and the first mutant embedding representation, and obtain the first original embedding vector and the first mutant embedding vector respectively.
[0108] Among them, the first original embedding vector is a 768-dimensional vector obtained by performing max pooling on the first original embedding representation. The first mutant embedding vector is a 768-dimensional vector obtained by performing max pooling on the first mutant embedding representation.
[0109] Max pooling compresses the embedding representation of the entire sequence into an embedding vector with a fixed dimension. In the first embodiment, the window width of the sequence information before and after mutation is 1024 bases, and the max pooling operation compresses (1024, 768) to 768 dimensions.
[0110] Through step S200, the nucleotide sequence information can be converted into a vector form composed of real numbers that can be processed by deep learning by capturing potential relationships and structures.
[0111] Step S300, input the first original embedding vector and the first mutant embedding vector into the contrast learning module, adjust the embedding distance between the first original embedding vector and the first mutant embedding vector in the contrast learning module, and output the second original embedding vector and the second mutant embedding vector. Among them, the second original embedding vector is obtained according to the first original embedding vector, and the second mutant embedding vector is obtained according to the first mutant embedding vector.
[0112] As an implementable manner, step S300 specifically includes:
[0113] Input the first original embedding vector and the first mutated embedding vector into the contrastive learning module; the contrastive learning module is in the classification prediction model obtained in step S100, specifically the result obtained after being trained in step S120. The contrastive learning module includes an Encoder layer of a Transformer and a linear layer.
[0114] According to the different labels of the sample pairs, adjust the embedding distance between the first original embedding vector and the first mutated embedding vector in the contrastive learning module. Among them, the embedding distance is set as the Euclidean distance.
[0115] The adjustment result obtained in the contrastive learning module is:
[0116] Make the Euclidean distance between the embedding vectors before and after the mutation of the pathogenic variant as far as possible, and make the Euclidean distance between the embedding vectors before and after the mutation of the benign variant as close as possible.
[0117] That is to say, if the splicing variant is a pathogenic variant, then expand the Euclidean distance between the first original embedding vector and the first mutated embedding vector, so that the Euclidean distance between the output second original embedding vector and the second mutated embedding vector is at least greater than the Euclidean distance between the first original embedding vector and the first mutated embedding vector, and the greater the better.
[0118] If the splicing variant is a benign variant, then reduce the Euclidean distance between the first original embedding vector and the first mutated embedding vector, so that the Euclidean distance between the output second original embedding vector and the second mutated embedding vector is at least less than the Euclidean distance between the first original embedding vector and the first mutated embedding vector, and the smaller the better.
[0119] Adjust the embedding distance in the contrastive learning module, corresponding to pathogenic variants and benign variants to make the embedding distance differentiation more significant, which helps to more efficiently and simply distinguish the two variants in the subsequent classification task, make a judgment on whether it is pathogenic, and improve the classification accuracy and classification efficiency.
[0120] It should be noted that both the second original embedding vector and the second mutated embedding vector are 128-dimensional vectors. During the processing of the contrastive learning module, not only the distance relationship between the embedding vectors is adjusted, but also the dimension of the embedding vectors is reduced from 768 to 128 through a dimensionality reduction operation. The dimensionality reduction process can effectively extract the information most relevant to the prediction task in the sequence, while removing noise and irrelevant variables, thereby simplifying the learning task of the classification prediction model and achieving a significant noise reduction effect. In addition, dimensionality reduction also reduces the computational complexity of the classification layer, significantly improves the training speed of the model, and at the same time reduces the overfitting risk brought by high-dimensional features of the model, further improving the generalization performance of the model.
[0121] Step S400: Feature fusion is performed on the second original embedding vector and the second variant embedding vector output after adjustment in step S300 to obtain a fusion tensor. The fusion tensor is a two-dimensional tensor of (2, 128).
[0122] Preferably, in this embodiment, the operation of feature fusion is to stack the two vectors. In other embodiments, other fusion operations that meet the prediction requirements can also be used for feature fusion, not limited to this.
[0123] Step S500: The fusion tensor output in step S400 is input into a classifier for classification prediction, and a pathogenicity prediction result is output. The classifier is in the classification prediction model obtained in step S100, specifically the result obtained after training in step S130. The classifier consists of a convolutional layer and two linear layers.
[0124] As an implementable manner, step S500 specifically includes:[[]]
[0125] Input the fusion tensor into the trained classifier for binary classification prediction, and perform non-linear division through the activation function of the classifier to output a classification result. The activation function is the PReLU activation function. The binary classification prediction result of the classifier is 0 or 1.
[0126] If the binary classification prediction result is 1, the pathogenicity prediction result is that the splicing variant has a pathogenic mutation.
[0127] If the binary classification prediction result is 0, the pathogenicity prediction result is that the splicing variant has a benign mutation.
[0128] Taking the SPiCE model, dbscSNV_RF model, SpliceAI model, and kipoiSplice model in the prior art as four comparative examples, a performance comparison is made with the classification prediction model applied in the splicing variant pathogenicity prediction method based on contrast learning provided in the first embodiment of the present invention, and the same test data is used for testing.
[0129] Specifically, the test data is set as the standard test set of SPCards, and this test set has excluded the training data that other splicing prediction tools may use, and has high reliability. The obtained performance test results are shown in Table 1. Among them, the classification prediction model applied in the splicing variant pathogenicity prediction method based on contrast learning provided in the first embodiment is represented by the ACE model.
[0130] Table 1
[0131]
[0132] The results of the performance test include six metrics: Positive Predictive Value (PPV), Negative Predictive Value (NPV), Specificity (denoted as Spec in Table 1), Sensitivity (denoted as Sens in Table 1), Accuracy, Matthews Correlation Coefficient (MCC), and Coverage. These are all commonly used metrics in binary classification tasks.
[0133] Among them, the positive predictive value is also called the positive prediction rate. Specifically, it refers to the proportion of samples with an actual value of "positive" among those with a predicted result of "positive" by the model. The formula for calculating the positive predictive value is:
[0134] ;
[0135] In the above formula, TP (True Positive) represents the number of samples that are actually "positive" and are predicted to be "positive"; FP (False Positive) represents the number of samples that are actually "negative" but are predicted to be "positive".
[0136] The negative predictive value is also called the negative prediction rate. Specifically, it refers to the proportion of samples with an actual value of "negative" among those with a predicted result of "negative" by the model. The formula for calculating the negative predictive value is:
[0137] ;
[0138] In the above formula, TN (True Negative) represents the number of samples that are actually negative examples and are predicted to be negative examples; FN (False Negative) represents the number of samples that are actually "positive" but are predicted to be "negative".
[0139] Specificity refers to the proportion of samples that are correctly predicted as "negative" by the model among all samples that are actually "negative". The formula for calculating specificity is:
[0140] .
[0141] Sensitivity refers to the proportion of samples that are correctly predicted as "positive" by the model among all samples that are actually "positive". The formula for calculating sensitivity is:
[0142] .
[0143] Accuracy refers to the proportion of correct predictions in the total number of predictions. The formula for calculating accuracy is:
[0144] .
[0145] The Matthews correlation coefficient is an index for comprehensively evaluating classification performance. The higher it is, the stronger the robustness of the model to imbalanced data. The calculation formula of the Matthews correlation coefficient is:
[0146] .
[0147] As shown in Table 1, the classification prediction model in Example 1 is superior to other existing prediction models in six indicators: positive predictive value (PPV), negative predictive value (NPV), specificity (Spec), accuracy, Matthews correlation coefficient (MCC), and coverage rate, and also shows a sub-optimal performance in sensitivity (Sens).
[0148] Furthermore, the positive predictive value (PPV) of the classification prediction model has increased by 3.1% compared to the sub-optimal value, the negative predictive value (NPV) has increased by 7.1%, and the Matthews correlation coefficient (MCC) has reached a performance improvement of 9.8%. Due to the imbalance between positive and negative samples in the standard test set of SPCards, the higher MCC value demonstrates the overall stability and comprehensiveness of the model's accurate prediction for different categories, thus indicating that the model has good robustness in dealing with data imbalance problems.
[0149] The prediction coverage rate of the classification prediction model provided in Example 1 of the present invention reaches 100%. Restricted by the computational defects of biometric features, other models can only predict the pathogenicity of 26.1% - 87.9% of partial variants, with poor generality. The pathogenicity prediction method provided in Example 1 of the present invention is not restricted by biometric calculations and can give highly accurate predictions for all variants, with high generality and high accuracy. Example 2
[0150] Example 2 of the present invention provides a splicing variant pathogenicity prediction system based on contrastive learning. The system includes: a user input module, a sequence information matching module, a splicing variant pathogenicity prediction module, and a result feedback and notification module.
[0151] Among them, the splicing variant pathogenicity prediction module includes the splicing variant pathogenicity prediction method based on contrastive learning provided in Example 1, and specifically applies the classification prediction model provided in Example 1.
[0152] Specifically, the user input module is used to input genomic sequence information. Users can upload the sequence information before and after variation according to their own needs.
[0153] The uploaded sequence information is determined according to the predicted variant data volume. The specific upload format is the standard format for variants formulated by the Human Genome Variation Society (HGVS), which can be recognized by computers. Users paste the sequence information in the standard format corresponding to the variant into the text box for uploading, or directly upload a file in csv format.
[0154] Specifically, the sequence information matching module is used to process and extract sequence information.
[0155] The sequence information matching module matches the sequence information of the variant in the reference genome sequence stored in the cloud according to the user's uploaded data, identifies the variant positions in the genome sequence information, and extracts at least part of the bases centered on the variant positions before and after the variant, and temporarily stores them as the original sequence information and the variant sequence information respectively. The temporarily stored information will be forwarded to the splicing variant pathogenicity prediction module.
[0156] It should be noted that the window width of the base sequence extracted for temporary storage is preferably set to 1024 bases. In other embodiments, the number of bases in the window can also be 128, 256, 512, or 2048, and is not limited thereto. Through research by the inventors of this case, it is found that the setting of the window width affects the time required for the DNA pre-training model to output the embedding representation from obtaining the sequence information. When the window width is set to 1024 bases, the training time and prediction time can be controlled within the acceptable range of the model.
[0157] Specifically, the splicing variant pathogenicity prediction module outputs a pathogenicity prediction result according to the original sequence information and the variant sequence information output by the sequence information matching module.
[0158] When the splicing variant pathogenicity prediction module is working, if there is a graphics processing unit GPU currently, the GPU is preferentially used to accelerate the prediction task of the classification prediction model; if there is no GPU or the GPU is already occupied, the central processing unit CPU is used to load the classification prediction model for pathogenicity prediction.
[0159] The process of judging and outputting the pathogenicity prediction result applies the splicing variant pathogenicity prediction method. Since the splicing variant pathogenicity prediction method has been described in Embodiment 1, the system in this embodiment has all the technical effects and content of the method in Embodiment 1, and will not be elaborated here.
[0160] After obtaining the pathogenicity prediction result, the splicing variant pathogenicity prediction module transmits it to the result feedback and notification module.
[0161] Specifically, the result feedback and notification module is used to collect and record the pathogenicity prediction results. On the one hand, according to the user's selection, the prediction results are presented to the user in the form of an email attachment or online display. On the other hand, the calculation records and prediction result files are saved for the user to the cloud. Embodiment III
[0162] Embodiment III of the present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the splicing variant pathogenicity prediction method provided in Embodiment I of the present invention. Embodiment IV
[0163] Embodiment IV of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the splicing variant pathogenicity prediction method provided in Embodiment I of the present invention. Embodiment V
[0164] Embodiment V of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and the computer program, when executed by the at least one processor, is used to cause the electronic device to execute the splicing variant pathogenicity prediction method provided in Embodiment I.
[0165] The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0166] The computer program for implementing the method of the embodiment of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing devices, such that the computer programs, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0167] In the context of embodiments of the present inventive concept, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0168] It should be noted that the term "including" and its variations used in the embodiments of the present inventive concept are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "an embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present inventive concept are illustrative rather than restrictive, and those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".
[0169] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present inventive concept are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.
[0170] The various steps described in the method embodiments provided by the embodiments of the present inventive concept may be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present inventive concept is not limited in this regard.
[0171] As used in this specification, the term "embodiment" means that the specific features, structures or characteristics described in connection with an embodiment may be included in at least one embodiment of the present invention. The phrase appears at various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. The embodiments in this specification are all described in a related manner, and the same or similar parts among the embodiments are referred to each other. In particular, for embodiments of devices, equipment, and systems, since they are basically similar to embodiments of methods, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiments.
[0172] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.
Claims
1. A method for predicting the pathogenicity of splicing variants based on contrastive learning, characterized in that, Including: Inputting the original sequence information and the mutant sequence information, and respectively obtaining a first original embedding vector and a first mutant embedding vector; wherein, the first original embedding vector is obtained according to the original sequence information, and the first mutant embedding vector is obtained according to the mutant sequence information; Inputting the first original embedding vector and the first mutant embedding vector into a contrastive learning module, adjusting the embedding distance between the first original embedding vector and the first mutant embedding vector in the contrastive learning module, and outputting a second original embedding vector and a second mutant embedding vector; wherein, the second original embedding vector is obtained according to the first original embedding vector, and the second mutant embedding vector is obtained according to the first mutant embedding vector; Performing feature fusion on the second original embedding vector and the second mutant embedding vector to obtain a fusion tensor; Inputting the fusion tensor into a classifier for classification prediction, and outputting a pathogenicity prediction result; Before inputting the original sequence information and the mutant sequence information, the splicing variant pathogenicity prediction method based on contrastive learning further includes: Creating a classification prediction model according to a DNA pre-training model, the contrastive learning module, and the classifier; wherein, the DNA pre-training model is used to extract corresponding embedding representations according to sequence information; Training the DNA pre-training model and the contrastive learning module according to training data in the case of not connecting the classifier; wherein, the training data includes sample pairs composed of multiple groups of the sequence information; Freezing the DNA pre-training model and the contrastive learning module, connecting the classifier, and training the classifier according to the training data.
2. The method for predicting the pathogenicity of splicing variants based on contrastive learning according to claim 1, wherein Training the DNA pre-training model and the contrastive learning module according to the training data includes: Inputting the training data into the DNA pre-training model and performing max pooling to obtain a group of embedding representations; wherein, the DNA pre-training model includes initial pre-training parameters; Inputting the group of embedding representations into the contrastive learning module to obtain a group of embedding vectors; Calculating the contrastive learning loss of the group of embedding vectors using a contrastive learning loss function in the contrastive learning module; wherein, the formula of the contrastive learning loss function is: ; ; ; is the total contrastive learning loss; is the contrastive learning loss of the positive sample pair, is the contrastive learning loss of the negative sample pair; is the label of the sample pair; is the Euclidean distance between the pre - and post - mutation embedding vectors in the embedding vector group, is the maximum allowable Euclidean distance in the negative sample pair, is the number of sample pairs in the training data; Training the DNA pre-training model and the contrastive learning module until the total contrastive learning loss converges; obtaining the trained DNA pre-training model and contrastive learning module.
3. The method for predicting the pathogenicity of splicing variants based on contrastive learning according to claim 2, wherein Training the classifier according to the training data includes: Inputting the training data into the trained DNA pre-training model and performing max pooling; Inputting the result of max pooling into the trained contrastive learning module; Inputting the output result of the trained contrastive learning module into the classifier; wherein, the classifier uses a cross-entropy loss function; Training the classifier until the loss of the classification prediction model converges, and obtaining the trained classifier.
4. The splicing variant pathogenicity prediction method based on contrastive learning according to claim 1, wherein Inputting the original sequence information and the mutant sequence information, and respectively obtaining a first original embedding vector and a first mutant embedding vector includes: Input the original sequence information and the variant sequence information into the trained DNA pre-trained model to obtain a first original embedding representation and a first variant embedding representation respectively; Perform max pooling on the first original embedding representation and the first variant embedding representation to obtain the first original embedding vector and the first variant embedding vector respectively; wherein, both the first original embedding vector and the first variant embedding vector are 768-dimensional vectors.
5. The splicing variant pathogenicity prediction method based on contrastive learning according to claim 1, wherein Adjust the embedding distances of the first original embedding vector and the first variant embedding vector in the contrast learning module, and output the second original embedding vector and the second variant embedding vector, including: Adjust the embedding distances of the first original embedding vector and the first variant embedding vector according to different labels; wherein, the embedding distance is set as the Euclidean distance; If the splicing variant is a pathogenic variant, make the Euclidean distance between the second original embedding vector and the second variant embedding vector greater than the Euclidean distance between the first original embedding vector and the first variant embedding vector; If the splicing variant is a benign variant, make the Euclidean distance between the second original embedding vector and the second variant embedding vector less than the Euclidean distance between the first original embedding vector and the first variant embedding vector.
6. The method for predicting the pathogenicity of splicing variants based on contrastive learning according to claim 1, wherein Perform feature fusion on the second original embedding vector and the second variant embedding vector to obtain a fusion tensor, including: Perform a stacking operation on the second original embedding vector and the second variant embedding vector to complete feature fusion and obtain the fusion tensor; wherein, both the second original embedding vector and the second variant embedding vector are 128-dimensional vectors, and the fusion tensor is a two-dimensional tensor of (2, 128).
7. The splicing variant pathogenicity prediction method based on contrastive learning according to claim 1, wherein Input the fusion tensor into a classifier for classification prediction, and output a pathogenicity prediction result, including: Input the fusion tensor into the trained classifier for binary classification prediction, and perform non-linear partitioning through the activation function of the classifier to output a classification result; wherein, the binary classification prediction result of the classifier is 0 or 1; If the binary classification prediction result is 1, the pathogenicity prediction result is that the splicing variant has a pathogenic variant; If the binary classification prediction result is 0, the pathogenicity prediction result is that the splicing variant has a benign variant.
8. A splicing variant pathogenicity prediction system based on contrastive learning, characterized in that, Including: A user input module, a sequence information matching module, a splicing variant pathogenicity prediction module, and a result feedback and notification module; wherein, the splicing variant pathogenicity prediction module includes the method for predicting the pathogenicity of splicing variants based on contrast learning according to any one of claims 1 to 7; The user input module is used to input genomic sequence information; The sequence information matching module is used to identify the variant positions in the genomic sequence information, and extract the DNA sequence information centered on the variant positions before and after the variant, with a window size of 1024 bases, as the original sequence information and the variant sequence information respectively, and then input the original sequence information and the variant sequence information into the splicing variant pathogenicity prediction module; The splicing variant pathogenicity prediction module outputs a pathogenicity prediction result according to the original sequence information and the variant sequence information, and transmits the pathogenicity prediction result to the result feedback and notification module; The result feedback and notification module is used to collect and record the pathogenicity prediction result and display the pathogenicity prediction result.
9. An electronic device, comprising: A processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the splicing variant pathogenicity prediction method based on contrastive learning according to any one of claims 1 to 7.
Citation Information
Patent Citations
Prediction model establishment and prediction method for disease-related variable splicing isomers
CN116825178A