Method, device and equipment for predicting function influence of single nucleotide variation in non-coding region
By extracting and fusing the global semantic features and local features of single nucleotide variation in the non-coding region in the prediction model, the problem of low prediction accuracy in the prior art is solved, and the accuracy of the prediction results is significantly improved.
Patent Information
- Application Number
- CN202510224933.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has low prediction accuracy when predicting the impact of single nucleotide variation in non-coding regions on gene function.
A prediction model including feature extraction module, multi-level feature fusion module and prediction output module is adopted to extract global semantic features and local features through a single-hot coding sequence, and perform multi-level feature fusion to improve the accuracy of prediction results.
By combining global semantic features and local features, the prediction accuracy of the impact of single nucleotide variants in non-coding regions is improved.
Smart Images

Figure CN120072054A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene technology, and particularly to a method, device, and equipment for predicting the functional impact of single nucleotide variations in non-coding regions. Background Art
[0002] Currently, with the in-depth study of the relationship between variations and diseases, it has been found that gene mutations in non-coding regions are closely related to many complex diseases. Although the non-coding region of the genome occupies more than 98% of the human genome, although it does not directly encode proteins, the regulatory sequences it contains are crucial for the regulation of gene expression. Specifically, the non-coding region contains key functional elements, such as enhancers and promoters, which play a crucial role in gene expression and regulation. When there are single nucleotide variations in the non-coding region, it may disrupt key regulatory elements, such as enhancers, promoters, or RNA binding sites, resulting in abnormal gene expression. Further, such abnormalities may lead to the occurrence of various complex diseases. Therefore, accurately predicting the functional impact of single nucleotide variations in non-coding regions has important scientific value and clinical significance.
[0003] In the prior art, one method is to be able to predict the functional impact of single nucleotide variations in non-coding regions based on deep learning using different prediction models. Another method is to perform k-mer processing on the one-hot encoded gene sequences or directly use the one-hot encoding method for subsequent processing to obtain the prediction results.
[0004] However, using the prior art, there is a problem of low prediction accuracy for the prediction results of the functional impact of single nucleotide variations in non-coding regions. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, and equipment for predicting the functional impact of single nucleotide variations in non-coding regions for the above technical problems.
[0006] In a first aspect, an embodiment of the present invention provides a method for predicting the functional impact of single nucleotide variations in non-coding regions, the method comprising:
[0007] Obtaining a to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide;
[0008] Inputting the to-be-predicted one-hot encoded sequence into a trained prediction model to obtain a target prediction result for the functional impact of the to-be-predicted single nucleotide; wherein, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module, the feature extraction module is used to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence, and the multi-level feature fusion module is used to fuse the global semantic features and local features.
[0009] In one embodiment, the step of inputting the to-be-predicted one-hot encoded sequence into the trained prediction model to obtain the target prediction result for the functional impact of the to-be-predicted single nucleotide includes:
[0010] Input the to-be-predicted one-hot encoded sequence into the feature extraction module to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence;
[0011] Input the global semantic features and the local features into the multi-level feature fusion module for fusion processing to obtain the target prediction features;
[0012] Input the target prediction features into the prediction output module to obtain the target prediction result.
[0013] In one embodiment, the feature extraction module includes: a global semantic feature extraction sub-module and a local feature extraction sub-module. The step of inputting the to-be-predicted one-hot encoded sequence into the feature extraction module to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence includes:
[0014] Input the to-be-predicted one-hot encoded sequence into the global semantic feature extraction sub-module to extract the global semantic features corresponding to the to-be-predicted one-hot encoded sequence;
[0015] Input the to-be-predicted one-hot encoded sequence into the local feature extraction sub-module to extract the local features corresponding to the to-be-predicted one-hot encoded sequence.
[0016] In one embodiment, the global semantic feature extraction sub-module includes: a convolutional network, a long short-term memory network, and an attention mechanism layer. The step of inputting the to-be-predicted one-hot encoded sequence into the global semantic feature extraction sub-module to extract the global semantic features corresponding to the to-be-predicted one-hot encoded sequence includes:
[0017] Input the to-be-predicted one-hot encoded sequence into the convolutional network to extract convolutional features;
[0018] Input the convolutional features into the long short-term memory network to extract global context semantic features;
[0019] Input the global context semantic features into the attention mechanism layer to extract the global semantic features corresponding to the to-be-predicted one-hot encoded sequence.
[0020] In one embodiment, the step of inputting the global context semantic features into the attention mechanism layer to extract the global semantic features corresponding to the to-be-predicted one-hot encoded sequence includes:
[0021] Input the global context semantic features into the attention mechanism layer, and obtain the query vector, key vector, and value vector corresponding to the global context semantic features through the attention mechanism layer;
[0022] Determine the global semantic features corresponding to the to-be-predicted one-hot encoding sequence according to the query vector, the key vector, and the value vector.
[0023] In one embodiment, the local feature extraction sub-module includes: a k-mer feature extraction layer and a convolutional network. Inputting the to-be-predicted one-hot encoding sequence into the local feature extraction sub-module to extract the local features corresponding to the to-be-predicted one-hot encoding sequence includes:
[0024] Input the to-be-predicted one-hot encoding sequence into the k-mer feature extraction layer to extract position embedding features;
[0025] Input the position embedding features into the convolutional network to extract the local features corresponding to the to-be-predicted one-hot encoding sequence.
[0026] In one embodiment, the multi-level feature fusion module includes: a feature splicing sub-module and an attention mechanism sub-module. Inputting the global semantic features and the local features into the multi-level feature fusion module for fusion processing to obtain target prediction features includes:
[0027] Input the global semantic features and the local features into the feature splicing sub-module for splicing processing to obtain initial prediction features;
[0028] Input the initial prediction features into the attention mechanism sub-module to obtain target prediction features.
[0029] In one embodiment, the attention mechanism sub-module includes: a channel attention mechanism layer and a spatial attention mechanism layer. Inputting the initial prediction features into the attention mechanism sub-module to obtain target prediction features includes:
[0030] Input the initial prediction features into the channel attention mechanism layer for attention processing of channel information to obtain channel attention features;
[0031] Input the channel attention features into the spatial attention mechanism layer for attention processing of spatial information to obtain target prediction features.
[0032] In a second aspect, an apparatus for predicting the functional impact of non-coding region single nucleotide variations according to an embodiment of the present invention includes:
[0033] A one-hot encoded sequence acquisition module to be predicted, which is used to acquire a one-hot encoded sequence to be predicted corresponding to a single nucleotide to be predicted;
[0034] A target prediction result acquisition module, which is used to input the one-hot encoded sequence to be predicted into a trained prediction model to obtain a target prediction result for the functional impact of the single nucleotide to be predicted; wherein, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module, and the feature extraction module is used to extract the global semantic features and local features corresponding to the one-hot encoded sequence to be predicted, and the multi-level feature fusion module is used to fuse the global semantic features and local features.
[0035] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the method for predicting the functional impact of single nucleotide variation in the non-coding region described in the first aspect are implemented.
[0036] The technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0037] A method for predicting the functional impact of single nucleotide variation in the non-coding region provided by the embodiment of the present invention. In this way, by acquiring a one-hot encoded sequence to be predicted corresponding to a single nucleotide to be predicted, and inputting the one-hot encoded sequence to be predicted into a trained prediction model, a target prediction result for the functional impact of the single nucleotide to be predicted is obtained. Among them, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module. For the feature extraction module, it is used to extract the global semantic features and local features corresponding to the one-hot encoded sequence to be predicted. Further, the multi-level feature fusion module is used to fuse the global semantic features and local features, so that when predicting the functional impact of the single nucleotide to be predicted, the global semantic features and local features can be combined at the same time, thereby improving the accuracy of the target prediction result for the functional impact of the single nucleotide to be predicted. Description of the Drawings
[0038] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1Schematic flowchart of a method for predicting the functional impact of single nucleotide variations in non-coding regions provided by an embodiment of the present invention;
[0041] Figure 2 Schematic structural diagram of a prediction model provided by an embodiment of the present invention;
[0042] Figure 3 Schematic structural diagram of another prediction model provided by an embodiment of the present invention;
[0043] Figure 4 Schematic structural diagram of a device for predicting the functional impact of single nucleotide variations in non-coding regions provided by an embodiment of the present invention. Detailed implementation manners
[0044] In order to more clearly understand the above objects, features and advantages of the present invention, the solutions of the present invention will be further described below. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0045] In the following description, many specific details are set forth in order to fully understand the present invention, but the present invention may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0046] Currently, with the in-depth study of variations and diseases, it is found that gene mutations in non-coding regions are closely related to many complex diseases. The non-coding region of the genome accounts for more than 98% of the human genome. Although the non-coding region does not directly encode proteins, the regulatory sequences it contains are crucial for the regulation of gene expression. Specifically, the non-coding region contains key functional elements, such as enhancers and promoters, which play a crucial role in gene expression and regulation. When a single nucleotide in the non-coding region varies, it may disrupt key regulatory elements, such as enhancers, promoters or RNA binding sites, thereby leading to abnormal gene expression. Further, this abnormality may lead to the occurrence of various complex diseases. Therefore, accurately predicting the functional impact of single nucleotide variations in non-coding regions on genes has important scientific value and clinical significance.
[0047] In the prior art, based on deep learning, different prediction models are used to predict the functional impact of single nucleotide variations in non-coding regions. Or the one-hot encoded gene sequences are subjected to k-mer processing or directly processed in the one-hot encoding manner to obtain the prediction results.
[0048] However, when using the existing technology, it is necessary to process large-scale data, the prediction model structure is complex, and only local feature information of a single gene sequence or global context feature information is considered. There is a problem of low prediction accuracy for the prediction results of the functional impact of single nucleotide variations in non-coding regions.
[0049] Therefore, the present invention provides a method for predicting the functional impact of single nucleotide variations in non-coding regions. By obtaining the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide, and inputting the to-be-predicted one-hot encoded sequence into the trained prediction model, the target prediction result of the functional impact on the to-be-predicted single nucleotide is obtained. Among them, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module. For the feature extraction module, it is used to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence. Further, the multi-level feature fusion module is used to fuse the global semantic features and local features, so that when predicting the functional impact of the to-be-predicted single nucleotide, the global semantic features and local features can be combined simultaneously, thereby improving the accuracy of the target prediction result of the functional impact on the to-be-predicted single nucleotide.
[0050] In one embodiment, as Figure 1 shown, Figure 1 is a schematic flowchart of a method for predicting the functional impact of single nucleotide variations in non-coding regions provided by an embodiment of the present invention, specifically including the following steps:
[0051] S10: Obtain the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide.
[0052] Among them, the to-be-predicted single nucleotide refers to the single nucleotide in the non-coding region that has undergone mutation. The to-be-predicted one-hot encoding, that is, OneHot encoding, also known as one-hot encoding, generally uses an N-bit status register to encode N states, and at any time, only one of them is valid. Exemplarily, for color features, there are 3 types: red, green, and yellow, then the converted one-hot encodings are respectively represented as: 001, 010, 100. The to-be-predicted one-hot encoded sequence refers to the sequence obtained by encoding the gene sequence of the to-be-predicted single nucleotide. Exemplarily, for example, the gene sequence of the to-be-predicted single nucleotide is 000110000100...0001, but it is not limited to this. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.
[0053] Specifically, encode the to-be-predicted single nucleotide that has undergone mutation in the non-coding region to obtain the corresponding to-be-predicted one-hot encoded sequence.
[0054] S11: Input the to-be-predicted one-hot encoded sequence into the trained prediction model to obtain the target prediction result of the functional impact on the to-be-predicted single nucleotide.
[0055] Among them, the target prediction result is used to predict whether the single nucleotide pair regulatory element in the mutated non-coding region is affected. Figure 2 FIG. Figure 2 is a schematic structural diagram of a prediction model provided by an embodiment of the present invention. The prediction model includes: a feature extraction module 21, a multi-level feature fusion module 22, and a prediction output module 23.
[0056] Among them, the feature extraction module is used to extract the global semantic features and local features corresponding to the one-hot encoded sequence to be predicted. The multi-level feature fusion module is used to fuse the global semantic features and local features, so that the global semantic features and local features can be combined simultaneously when predicting the functional impact of the single nucleotide to be predicted, thereby improving the accuracy of the prediction result. The prediction output module is used to output the target prediction result for the functional impact of the single nucleotide to be predicted.
[0057] Specifically, the one-hot encoded sequence to be predicted corresponding to the single nucleotide to be predicted with a mutation in the non-coding region obtained is input into the trained prediction model, and the target prediction result for the functional impact of the single nucleotide to be predicted is output through the trained prediction model.
[0058] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, before performing S11, it further includes:
[0059] Obtain a training sample set. Use the training sample set to train the initial prediction model until the model converges to obtain the trained prediction model.
[0060] Among them, the training sample set includes: a non-overlapping training subset, a validation subset, and a test subset. The training sample set can be, for example, the publicly available dataset D ori . For the training of the non-overlapping training subset, the validation subset, and the test subset, for example, 8,000 samples on chromosome 7 can be used as the validation subset, and 45,524 samples on chromosomes 8 and 9 can be used as the test subset. 4.4 million samples on the remaining chromosomes are selected as the non-overlapping training subset, but it is not limited to this. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.
[0061] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, one implementation manner of S11 can be:
[0062] S111: Input the one-hot encoded sequence to be predicted into the feature extraction module to extract the global semantic features and local features corresponding to the one-hot encoded sequence to be predicted.
[0063] Specifically, the one-hot encoded sequence to be predicted corresponding to the single nucleotide to be predicted with a mutation in the non-coding region obtained is input into the feature extraction module, and the global semantic features and local features corresponding to the one-hot encoded sequence to be predicted are extracted through the feature extraction module.
[0064] S112: Input the global semantic feature and the local feature into the multi-level feature fusion module for fusion processing to obtain the target prediction feature.
[0065] Specifically, input the global semantic feature and the local feature corresponding to the single nucleotide to be predicted with a mutation in the non-coding region into the multi-level feature fusion module, and through the multi-level feature fusion module, perform fusion processing on the global semantic feature and the local feature to obtain the target prediction feature corresponding to the one-hot encoded sequence to be predicted.
[0066] S113: Input the target prediction feature into the prediction output module to obtain the target prediction result.
[0067] Specifically, input the target prediction feature corresponding to the single nucleotide to be predicted with a mutation in the non-coding region into the prediction output module, and through the prediction output module, obtain the target prediction result regarding the functional impact on the single nucleotide to be predicted.
[0068] In this way, for the method for predicting the functional impact of single nucleotide variations in the non-coding region provided in this embodiment, by obtaining the one-hot encoded sequence to be predicted corresponding to the single nucleotide to be predicted, inputting the one-hot encoded sequence to be predicted into the trained prediction model, the target prediction result regarding the functional impact on the single nucleotide to be predicted is obtained. Among them, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module. For the feature extraction module, it is used to extract the global semantic feature and the local feature corresponding to the one-hot encoded sequence to be predicted. Further, the multi-level feature fusion module is used to fuse the global semantic feature and the local feature, so that when predicting the functional impact of the single nucleotide to be predicted, both the global semantic feature and the local feature can be combined, thereby improving the accuracy of the target prediction result regarding the functional impact on the single nucleotide to be predicted.
[0069] Optionally, based on the above embodiment, in some embodiments of the present invention Figure 3 is a schematic structural diagram of another prediction model provided by an embodiment of the present invention. The feature extraction module 21 includes: a global semantic feature extraction sub-module 31 and a local feature extraction sub-module 32. Based on this, one implementation manner of S111 may be:
[0070] S21: Input the one-hot encoded sequence to be predicted into the global semantic feature extraction sub-module to extract the global semantic feature corresponding to the one-hot encoded sequence to be predicted.
[0071] Specifically, input the one-hot encoded sequence to be predicted corresponding to the single nucleotide to be predicted with a mutation in the non-coding region into the global semantic feature extraction sub-module, and through the global semantic feature extraction sub-module, extract the global semantic feature corresponding to the one-hot encoded sequence to be predicted.
[0072] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 , the global semantic feature extraction sub-module 31 includes a convolutional network 311, a long short-term memory network 312, and an attention mechanism layer 313. Based on this, one implementation of S21 can be:
[0073] S211: Input the to-be-predicted one-hot encoded sequence into the convolutional network to extract convolutional features.
[0074] Among them, the convolutional network includes a convolutional operation layer, an activation function layer, a dropout layer, a batch normalization layer, and a local pooling layer. By using the convolutional operation layer through the local perception and weight sharing mechanisms, it can improve the learning efficiency of the network, reduce the number of network parameters, reduce the risk of overfitting, and help reduce the dimension of high-dimensional input data and extract the core features of the input data. Using the local pooling layer for pooling operations, by using the feature invariance principle for input data with translation, rotation, and scale invariance, a large amount of redundant information is removed to reduce the dimension, speed up the training speed, prevent overfitting, and improve the overall generalization ability of the prediction model. Based on this, using the convolutional network to perform convolutional processing on the to-be-predicted one-hot encoded sequence can improve the model training efficiency.
[0075] Specifically, input the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide with a mutation in the non-coding region obtained into the convolutional network, and extract the convolutional features corresponding to the to-be-predicted one-hot encoded sequence through the convolutional network.
[0076] S212: Input the convolutional features into the long short-term memory network to extract global context semantic features.
[0077] Among them, the long short-term memory network (LSTM) can perform context modeling on the convolutional features, capture global context dependencies, and at the same time understand the context information before and after the text, so as to provide a more comprehensive understanding of the text, thereby obtaining global context semantic features.
[0078] Specifically, input the obtained convolutional features into the long short-term memory network, and extract the global context semantic features corresponding to the to-be-predicted one-hot encoded sequence through the long short-term memory network.
[0079] S213: Input the global context semantic features into the attention mechanism layer to extract the global semantic features corresponding to the to-be-predicted one-hot encoded sequence.
[0080] Among them, the attention mechanism layer is used to further process and enhance the focusing ability for the global context semantic features to obtain the key features in the global context semantic features.
[0081] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation manner of S213 may be:
[0082] Input the global context semantic features into the attention mechanism layer, and obtain the query vector, key vector, and value vector corresponding to the global context semantic features through the attention mechanism layer.
[0083] Determine the global semantic features corresponding to the to-be-predicted one-hot encoded sequence according to the query vector, key vector, and value vector.
[0084] Specifically, input the obtained global context semantic features into the attention mechanism layer to obtain the corresponding query vector and key vector. Further, perform dot product operations on the query vector and key vector to calculate the attention scores between word vectors, and then normalize and weight the attention scores through the Softmax function to obtain the value vector. Finally, determine the global semantic features corresponding to the to-be-predicted one-hot encoded sequence according to the query vector, key vector, and value vector.
[0085] In this way, the method for predicting the functional impact of non-coding region single nucleotide mutations provided in this embodiment extracts convolutional features by inputting the to-be-predicted one-hot encoded sequence into the convolutional network, extracts global context semantic features by inputting the convolutional features into the long short-term memory network, extracts the global semantic features corresponding to the to-be-predicted one-hot encoded sequence by inputting the global context semantic features into the attention mechanism layer, and then uses the global semantic features to obtain the target prediction result subsequently, thereby improving the model training efficiency and the accuracy of the target prediction result.
[0086] S22: Input the to-be-predicted one-hot encoded sequence into the local feature extraction sub-module to extract the local features corresponding to the to-be-predicted one-hot encoded sequence.
[0087] Specifically, input the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide with a mutation in the non-coding region obtained into the local feature extraction sub-module, and extract the local features corresponding to the to-be-predicted one-hot encoded sequence through the local feature extraction sub-module.
[0088] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 , the local feature extraction sub-module 32 includes: a k-mer feature extraction layer 321 and a convolutional network 322. Based on this, one implementation manner of S22 may be:
[0089] S221: Input the to-be-predicted one-hot encoded sequence into the k-mer feature extraction layer to extract positional embedding features.
[0090] Among them, the k-mer feature extraction layer refers to performing k-mer processing on the input one-hot encoded sequence to be predicted, and after the k-mer processing, using the dna2vec technology for word embedding processing and adding positional encoding information. The dna2vec technology can convert variable-length k-mers into vector representations. Through this method, the original information is retained, and the similarity between the vectors after k-mer processing is measured, thereby generating vectors with semantic information.
[0091] Specifically, the one-hot encoded sequence to be predicted corresponding to the single nucleotide to be predicted with a mutation in the non-coding region obtained is input into the k-mer feature extraction layer, and the positional embedding features are obtained through the k-mer feature extraction layer.
[0092] S222: Input the positional embedding features into the convolutional network to extract the local features corresponding to the one-hot encoded sequence to be predicted.
[0093] Specifically, the obtained positional embedding features are input into the convolutional network, and the local features are extracted through the convolutional network extraction layer.
[0094] In this way, for the method for predicting the functional impact of single nucleotide variations in the non-coding region provided in this embodiment, by inputting the one-hot encoded sequence to be predicted into the k-mer feature extraction layer, the positional embedding features are extracted. The positional embedding features are input into the convolutional network to extract the local features corresponding to the one-hot encoded sequence to be predicted. Then, the local features are used to obtain the target prediction result subsequently, thereby improving the model training efficiency and the accuracy of the target prediction result.
[0095] Optionally, on the basis of the above embodiment, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the multi-level feature fusion module 22 includes: a feature splicing sub-module 33 and an attention mechanism sub-module 34. Based on this, one implementation manner of S112 can be:
[0096] S31: Input the global semantic features and the local features into the feature splicing sub-module for splicing processing to obtain initial prediction features.
[0097] Specifically, the obtained global semantic features and local features corresponding to the single nucleotide to be predicted with a mutation in the non-coding region are input into the feature splicing sub-module, and the global semantic features and local features are spliced through the feature splicing sub-module to obtain initial prediction features.
[0098] S32: Input the initial prediction features into the attention mechanism sub-module to obtain the target prediction features.
[0099] Specifically, the initial prediction features are input into the attention mechanism sub-module, and the target prediction features are obtained through the attention mechanism sub-module.
[0100] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the attention mechanism sub-module 34 includes: a channel attention mechanism layer 341 and a spatial attention mechanism layer 342. Based on this, one implementation manner of S32 can be:
[0101] S321: Input the initial prediction feature into the channel attention mechanism layer for attention processing of channel information to obtain channel attention features.
[0102] Among them, the channel attention mechanism layer is used to calculate the weight of each channel, give an adaptive weighting coefficient, so as to obtain the key feature information in the initial prediction feature, thereby facilitating the improvement of the accuracy of the target prediction result during subsequent prediction. The channel attention mechanism layer can be, for example, a squeeze-and-excitation (SE) module, but is not limited thereto. The present invention does not specifically limit it, and those skilled in the art can set it according to the actual situation.
[0103] Specifically, input the obtained initial prediction feature into the channel attention mechanism layer for attention processing of channel information, and obtain channel attention features through the channel attention mechanism layer.
[0104] S322: Input the channel attention features into the spatial attention mechanism layer for attention processing of spatial information to obtain target prediction features.
[0105] Among them, the spatial attention mechanism layer is used to obtain spatial feature information, thereby facilitating the improvement of the accuracy of the target prediction result during subsequent prediction.
[0106] Specifically, input the obtained channel attention features into the spatial attention mechanism layer for attention processing of spatial information, and obtain target prediction features through the spatial attention mechanism layer.
[0107] In this way, the method for predicting the functional impact of non-coding region single nucleotide variations provided by the present invention splices the global semantic feature and the local feature through the feature splicing sub-module to obtain the initial prediction feature, and further uses the attention mechanism layer and the spatial attention mechanism layer to obtain the key target prediction features, thereby improving the accuracy of the target prediction result for the functional impact of the single nucleotide to be predicted.
[0108] It should be understood that although Figures 1-3The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 1-3 at least a part of the steps in
[0109] can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps. Figure 4 In one embodiment, as
[0110] shown, a device for predicting the functional impact of non-coding region single nucleotide variations is provided, including: a to-be-predicted one-hot encoded sequence acquisition module 10 and a target prediction result acquisition module 11.
[0111] Among them, the to-be-predicted one-hot encoded sequence acquisition module 10 is used to acquire the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide;
[0111] The target prediction result acquisition module 11 is used to input the to-be-predicted one-hot encoded sequence into the trained prediction model to obtain the target prediction result of the functional impact on the to-be-predicted single nucleotide; among them, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module. The feature extraction module is used to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence, and the multi-level feature fusion module is used to fuse the global semantic features and local features.
[0112] In the above embodiment, the to-be-predicted one-hot encoded sequence acquisition module is used to acquire the to-be-predicted one-hot encoded sequence corresponding to the to-be-predicted single nucleotide, and the target prediction result acquisition module is used to input the to-be-predicted one-hot encoded sequence into the trained prediction model to obtain the target prediction result of the functional impact on the to-be-predicted single nucleotide. Among them, the prediction model includes a feature extraction module, a multi-level feature fusion module, and a prediction output module. For the feature extraction module, it is used to extract the global semantic features and local features corresponding to the to-be-predicted one-hot encoded sequence. Further, the multi-level feature fusion module is used to fuse the global semantic features and local features, so that when predicting the functional impact of the to-be-predicted single nucleotide, the global semantic features and local features can be combined at the same time, thereby improving the accuracy of the target prediction result of the functional impact on the to-be-predicted single nucleotide.
[0113] For the specific limitations of the device for predicting the functional impact of non-coding region single nucleotide variations, reference may be made to the limitations of the method for predicting the functional impact of non-coding region single nucleotide variations in the foregoing text, which will not be elaborated herein. Each module in the above server can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0114] An embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for predicting the functional impact of non-coding region single nucleotide variations provided by the embodiment of the present invention can be implemented. For example, when the processor executes the computer program, the technical solution of any of the method embodiments shown can be implemented. The implementation principle and technical effect are similar, and will not be elaborated herein. Figures 1-3 Any person of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, a database, or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. The non-volatile memory can include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0115]
[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0117] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A method for predicting the functional impact of a single nucleotide variation in a non-coding region, characterized in that: include: Obtain the one-hot encoding sequence to be predicted corresponding to the single nucleotide to be predicted; The one-hot coding sequence to be predicted is input into the trained prediction model to obtain the target prediction result for the functional impact of the single nucleotide to be predicted; wherein the prediction model includes a feature extraction module, a multi-level feature fusion module and a prediction output module, the feature extraction module is used to extract the global semantic features and local features corresponding to the one-hot coding sequence to be predicted, and the multi-level feature fusion module is used to fuse the global semantic features and local features.
2. The method according to claim 1, characterized in that The step of inputting the one-hot encoding sequence to be predicted into a trained prediction model to obtain a target prediction result for the functional impact of the single nucleotide to be predicted comprises: Inputting the one-hot coded sequence to be predicted into the feature extraction module, extracting global semantic features and local features corresponding to the one-hot coded sequence to be predicted; Inputting the global semantic features and the local features into the multi-level feature fusion module for fusion processing to obtain target prediction features; The target prediction feature is input into the prediction output module to obtain the target prediction result.
3. The method according to claim 1, characterized in that: The feature extraction module includes: a global semantic feature extraction submodule and a local feature extraction submodule, wherein the one-hot encoding sequence to be predicted is input into the feature extraction module, and the global semantic features and local features corresponding to the one-hot encoding sequence to be predicted are extracted, including: Inputting the one-hot coded sequence to be predicted into the global semantic feature extraction submodule to extract the global semantic feature corresponding to the one-hot coded sequence to be predicted; The one-hot encoding sequence to be predicted is input into the local feature extraction submodule to extract the local features corresponding to the one-hot encoding sequence to be predicted.
4. The method according to claim 3, characterized in that The global semantic feature extraction submodule includes: a convolutional network, a long short-term memory network and an attention mechanism layer. The inputting the one-hot encoding sequence to be predicted into the global semantic feature extraction submodule to extract the global semantic features corresponding to the one-hot encoding sequence to be predicted includes: Inputting the one-hot encoding sequence to be predicted into the convolutional network to extract convolutional features; Inputting the convolutional features into the long short-term memory network to extract global context semantic features; The global context semantic features are input into the attention mechanism layer to extract the global semantic features corresponding to the one-hot encoding sequence to be predicted.
5. The method according to claim 4, characterized in that The step of inputting the global context semantic features into the attention mechanism layer and extracting the global semantic features corresponding to the one-hot encoding sequence to be predicted includes: Inputting the global context semantic feature into the attention mechanism layer, and obtaining the query vector, key vector and value vector corresponding to the global context semantic feature through the attention mechanism layer; A global semantic feature corresponding to the one-hot encoded sequence to be predicted is determined according to the query vector, the key vector, and the value vector.
6. The method according to claim 3, characterized in that The local feature extraction submodule includes: a k-mer feature extraction layer and a convolutional network. The inputting the one-hot encoding sequence to be predicted into the local feature extraction submodule to extract the local features corresponding to the one-hot encoding sequence to be predicted includes: Input the one-hot encoding sequence to be predicted into the k-mer feature extraction layer to extract position embedding features; The position embedding feature is input into the convolutional network to extract the local features corresponding to the one-hot encoding sequence to be predicted.
7. The method according to claim 2, characterized in that The multi-level feature fusion module includes: a feature splicing submodule and an attention mechanism submodule. The global semantic features and the local features are input into the multi-level feature fusion module for fusion processing to obtain target prediction features, including: Inputting the global semantic features and the local features into the feature splicing submodule for splicing processing to obtain initial prediction features; The initial prediction features are input into the attention mechanism submodule to obtain the target prediction features.
8. The method according to claim 7, characterized in that The attention mechanism submodule includes: a channel attention mechanism layer and a spatial attention mechanism layer. The initial prediction feature is input into the attention mechanism submodule to obtain the target prediction feature, including: Inputting the initial prediction features into the channel attention mechanism layer to perform attention processing on the channel information to obtain channel attention features; The channel attention features are input into the spatial attention mechanism layer to perform attention processing on spatial information to obtain target prediction features.
9. A device for predicting the functional impact of single nucleotide variations in non-coding regions, characterized in that: include: A module for obtaining a one-hot coding sequence to be predicted, used to obtain a one-hot coding sequence to be predicted corresponding to a single nucleotide to be predicted; The target prediction result acquisition module is used to input the one-hot coding sequence to be predicted into the trained prediction model to obtain the target prediction result for the functional impact of the single nucleotide to be predicted; wherein the prediction model includes a feature extraction module, a multi-level feature fusion module and a prediction output module, the feature extraction module is used to extract the global semantic features and local features corresponding to the one-hot coding sequence to be predicted, and the multi-level feature fusion module is used to fuse the global semantic features and local features.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for predicting the functional impact of a single nucleotide variation in a non-coding region according to any one of claims 1 to 8 are implemented.