A Method for Identifying Transcription Factors with Preferential Binding to Methylated DNA Based on ProtBERT
By processing transcription factor sequences based on ProtBERT, the problem of difficulty in identifying transcription factors that prefer binding to methylated DNA in the prior art is solved, and efficient and accurate transcription factor recognition and prediction are achieved.
Patent Information
- Application Number
- CN202410783592.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-06-18
AI Technical Summary
It is difficult for the prior art to effectively identify transcription factors that prefer to bind methylated DNA, especially as the number of protein sequences is rapidly increasing, and experimental methods with high cost and long-term effectiveness are difficult to meet the needs.
Using a multi-layer Transformer architecture based on ProtBERT, the transcription factor sequence is processed through AutoTokenizer, and the pre-trained ProtBERT model is loaded for sequence classification to identify transcription factors that prefer to bind methylated DNA.
It improves the accuracy and efficiency of identification of transcription factors that prefer to bind methylated DNA, enhances the model's generalization ability on new data, and improves the sensitivity and specificity of processing efficiency and prediction.
Smart Images

Figure CN118430661B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnological data research, and particularly to a method for identifying transcription factors that preferentially bind to methylated DNA based on ProtBERT. Background Art
[0002] Transcription factors are DNA-binding proteins that regulate gene expression. Transcription factors mediate long-range interactions between DNA sequences during the formation of three-dimensional genomes, guiding the formation of TADs and loop structures, the conversion of A-B compartments, and nuclear repositioning, thereby affecting transcriptional regulation. The transcriptional regulation mechanism plays an important role in cellular processes and affects aspects such as cancer development and plant yield.
[0003] The traditional view is that transcription factors tend to bind to unmethylated DNA, while high levels of methylation of CpG dinucleotides hinder their binding. Recent studies have shown that many transcription factors, such as KLF4, TET, CEBPA, and ZFP57, preferentially bind to methylated DNA and significantly regulate gene expression, promoting transcriptional initiation and RNA splicing.
[0004] The specific interaction and function between methylated DNA and transcription factors are not yet clear. Identifying transcription factors that preferentially bind to methylated DNA and revealing the interaction mechanism are of great significance for understanding methylation-mediated biological processes and related diseases.
[0005] Currently, high-throughput experimental methods such as tandem mass spectrometry, functional protein microarrays, DNA microarrays, ChIP-BS-seq, and HT-SELEX can be used to identify transcription factors that preferentially bind to methylated DNA. However, given the rapid growth of the number of protein sequences in the post-genomic era, experimental methods are difficult to meet the requirements due to their high cost and long time-consuming nature. Therefore, it is particularly important to develop new computational methods to identify these transcription factors. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for identifying transcription factors that preferentially bind to methylated DNA based on ProtBERT.
[0007] The purpose of the present invention is achieved by the following technical solutions:
[0008] 1. A method for identifying transcription factors that preferentially bind to methylated DNA based on ProtBERT, comprising the following steps:
[0009] S1. Obtain a transcription factor sequence dataset and divide it into a training set and a test set;
[0010] S2. All sequences are padded or truncated to the same length and tokenized by AutoTokenizer;
[0011] S3. Use BertForSequenceClassification to load the pre-trained ProtBERT model for sequence classification.
[0012] Furthermore, in the step S1, during the processing of the sequence dataset, the processing of transcription factors preferring methylated DNA and transcription factors preferring non-methylated DNA is included, and the processing steps are as follows:
[0013] S201. Exclude sequences containing non-standard amino acid residues;
[0014] S202. Use CD-HIT20 to remove redundant samples;
[0015] S203. Transcription factors preferring methylated DNA sequences are extracted as positive samples, and transcription factors preferring non-methylated DNA are extracted as negative samples.
[0016] Furthermore, in the step S2, according to the distribution of transcription factors, a unified target length is set to cover most sequence lengths. For sequences longer than the target length, they are truncated from the end; for sequences shorter than the target length, specific padding tokens [PAD] need to be added to the specified positions of the sequence. Then each transcription factor sequence is decomposed into basic unit amino acids and mapped to numerical identifiers recognizable by the model according to a predefined vocabulary, and special tokens [CLS] and [SEP] are added at the beginning and end of the sequence respectively to indicate the start and end of the whole sequence.
[0017] Furthermore, in the step S3, when using BertForSequenceClassification to load the pre-trained ProtBERT model for sequence classification, the initial training set is:
[0018] ,
[0019] The initial test set is:
[0020] ,
[0021] and represent the protein sequences in the training set and the test set respectively, and are respectively and The corresponding tags, where N and M represent the number of samples in the training set and the test set respectively; the protein sequences need to be tokenized by AutoTokenizer and marked as , and each sequence is converted into a vector through the Tokenizer , where the integer IDs contained correspond to each amino acid in the sequence. Since the sequence lengths need to be unified, the tokenized vector is transformed into , where L is the predefined target sequence length. If the length is greater than L, the sequence is truncated; if it is less than L, it is padded with a specific token [PAD] until the length reaches L.
[0022] Furthermore, in the step S3, the model calculates the attention of each element in the input sequence to other elements in the sequence, and the model captures the complex dependencies within the sequence in this way. First, it is necessary to obtain the Query (Q), Key (K), and Value (V) matrices through a linear transformation of the input sequence. The mathematical expression is:
[0023]
[0024]
[0025]
[0026] where X is the input sequence, , and the weight matrices are parameters learned during the model training process.
[0027] The mathematical representation for calculating self-attention is:
[0028] ,
[0029] where Q, K, and V are the query, key, and value matrices respectively, is the dimension of the key, and the softmax function is used to normalize the weights;
[0030] Each Transformer encoder layer includes a feed-forward neural network that performs two-layer linear transformation, and the ReLU activation function is used during the transformation process:
[0031] ,
[0032] where , are the updated weight matrices, , are the updated biases, which are used to introduce non-linearity.
[0033] Furthermore, the sequence is converted into an embedding vector through an embedding layer. The embedding vector includes a word embedding and a position embedding. The word embedding and the position embedding are added together to obtain the final input formula ; The word embedding is obtained by looking up in a randomly initialized embedding matrix according to the integer encoding and combining the corresponding word embedding vectors, while the position embedding is calculated through a sine-cosine function. Given a position p and a dimension i, it is expressed as:
[0034]
[0035]
[0036] Accordingly, the corresponding embedding value is calculated to generate a d-dimensional position embedding vector, and then they are combined;
[0037] Then is input into the Transformer layer and updated layer by layer. The update formula is:
[0038] ,
[0039] where is the output result of the th layer, is the output result after updating the th layer;
[0040] Among them, in the first layer, the hidden state , and the final hidden state h [ CLS ] L of the first token CLS is used as the summary vector of the entire sequence. After being processed by the Droupout layer and the fully connected layer, the original scores of each category are obtained. The processing formula of the Droupout layer is: h drop = Dropout ( h [ CLS ] L ) , randomly discarding the output of neurons with a specified probability. The processing formula of the fully connected layer is: ; Then, through the softmax function the original scores are converted into a probability distribution of the same dimension;
[0041] where is the processed vector, W and b are the weight matrix and bias vector of the linear layer respectively, which are parameters learned during the model training process, z is the original score of each category, and p is the probability vector.
[0042] The beneficial effects of the present invention are:
[0043] (1) The ProtBERT model adopts a multi-layer Transformer architecture, which can capture deep features and complex dependencies in the sequence. The application of the Dropout layer enhances the generalization ability of the model on new data. The linear classification layer converts the output of the model into class probabilities, and has high accuracy and reliability in classifying transcription factors that preferentially bind to methylated DNA.
[0044] (2) Compared with traditional sequence-based prediction techniques, by combining large model techniques, the processing efficiency is improved, and the inherent features of the sequence are adaptively learned, improving indicators such as prediction accuracy, sensitivity, specificity, Matthews correlation coefficient, and area under the ROC curve. Brief Description of the Drawings
[0045] Figure 1 It is a flowchart of a method for identifying transcription factors that preferentially bind to methylated DNA based on ProtBERT;
[0046] Figure 2 It is a diagram of the ProtBERT model;
[0047] Figure 3 It is the forward propagation diagram of this embodiment. Detailed Embodiment
[0048] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.
[0049] Refer to Figure 1 , a method for identifying transcription factors that preferentially bind to methylated DNA based on ProtBERT.
[0050] Obtain a transcription factor sequence dataset and divide it into a training set and a test set. Obtain 601 transcription factors in humans that preferentially bind to methylated DNA and 286 transcription factors that preferentially bind to non-methylated DNA, and then further process these sequences in three steps:
[0051] First, exclude sequences containing non-standard amino acid residues, such as B, X, or Z;
[0052] Secondly, use CD-HIT20 to remove redundant samples with a threshold of 25%; finally, remove sequences less than 50 amino acids;
[0053] Finally, 339 transcription factors with a preference for methylated DNA sequences were extracted as positive samples, and 183 transcription factors with a preference for non-methylated DNA were extracted as negative samples.
[0054] All sequences were padded or truncated to the same length and tokenized by AutoTokenizer. A unified target length was set according to the distribution of transcription factors. For sequences longer than the target length, they were truncated starting from the end or other specific positions; for sequences shorter than the target length, specific padding tokens PAD were added to the specified positions of the sequences. Then each transcription factor sequence was decomposed into basic unit amino acids and mapped to numerical identifiers recognizable by the model according to a predefined vocabulary, and then special sequence tokens, namely [CLS] and [SEP], were added to indicate the start and end of the whole sequence.
[0055] [CLS] and [SEP] are special tokens, the same as [PAD]. The [CLS] token is added at the beginning of the sequence and is usually used to indicate the start of the whole sequence. [SEP] is at the end of the sequence, indicating the end position of the sequence. [PAD] is the padding token. When the sequence length is not enough for the unified length, [PAD] is placed in the empty positions. Suppose there is a protein sequence 'ACGT' that needs to be processed:
[0056] Integer encoding: First, 'A' is mapped to 1, 'C' is mapped to 2, and so on, to obtain the integer-encoded sequence {1, 2, 3, 4}.
[0057] Padding and truncation: Suppose the target length is 6:
[0058] If the sequence is 'ACGT' (length 4), two [PAD] tokens will be added, and the result is {1, 2, 3, 4, [PAD], [PAD]}.
[0059] If the sequence is 'ACGTACGT' (length 8), the last two characters will be truncated, and the result is {1, 2, 3, 4, 1, 2}.
[0060] Adding special tokens:
[0061] Adding the [CLS] token at the start of the sequence and the [SEP] token at the end of the sequence, the result is {[CLS], 1, 2, 3, 4, [SEP], [PAD], [PAD]}
[0062] In using BertForSequenceClassification to load the pre-trained ProtBERT model for sequence classification, the initial training set is:
[0063] ,
[0064] The initial test set is:
[0065] ,
[0066] and represent the protein sequences in the training set and the test set respectively, and are respectively and corresponding labels, where N and M represent the number of samples in the training set and the test set respectively; the protein sequences need to be tokenized by AutoTokenizer and marked as , and each sequence is converted into a vector by the Tokenizer, where the integer IDs contained correspond to each amino acid in the sequence. Since the sequence lengths need to be unified, the tokenized vectors are transformed into , where L is the predefined target sequence length. If is longer than L, the sequence is truncated; if it is shorter than L, it is padded with a specific marker PAD until the length reaches L.
[0067] The processed sequences are organized into batches for input into the model. The processed sequences are: , where B is the batch size. The sequences first enter the ProtBERT model, which is a model based on the Transformer architecture and mainly consists of multiple layers of Transformer encoders.
[0068] Each layer contains two main parts: the multi-head self-attention mechanism and the position-wise fully connected feed-forward network, both of which include layer normalization and residual connections. Layer normalization adjusts and scales the activation values of each layer to keep the data distributions between different layers consistent, reducing the training time and improving the generalization ability of the model. Residual connections, on the other hand, alleviate the vanishing gradient problem in deep networks by adding skip connections that directly add to the input in each layer, while retaining the original input information, enhancing the expressive power and training efficiency of the model.
[0069] Refer to Figure 2 , the ProtBERT model calculates the attention of each element in the input sequence to other elements in the sequence, and the model captures the complex dependencies within the sequence in this way. Mathematically, it is expressed as:
[0070] ,
[0071] where Q, K, and V are the query, key, and value matrices respectively, is the dimension of the key.
[0072] Each Transformer encoder layer includes a feed-forward neural network that performs two-layer linear transformation, and the ReLU activation function is used during the transformation process:
[0073] ,
[0074] where , are the updated weight matrices, , are the updated biases, which are used to introduce non-linearity.
[0075] Refer to Figure 3 , during the forward propagation process, the sequence is converted into an embedding vector, including word embedding and position embedding. The word embedding is determined by looking up the embedding of each token, while the position embedding ensures that even in the self-attention mechanism lacking sequential relationships, the model can understand the order of the tokens. The embedding vectors are then added together to form the final input representation ; then is input into the Transformer layer, and each layer applies the above self-attention mechanism and is updated layer by layer. The update formula is:
[0076] ,
[0077] where, is the output result of the th layer, is the updated output result of the th layer; in the first layer, , the final hidden state h [ CLS ] L of the first token is used h drop = Dropout ( h [ CLS ] L ) as the summary of the entire sequence. After being processed by the Droupout layer and the fully connected (linear) layer, the raw scores for each category are obtained. The processing formula of the Droupout layer is: ; the processing formula of the fully connected (linear) layer is: ; then through the softmax function
[0078] the raw scores are converted into a probability distribution of the same dimension; where, is the processed vector, W and b are the weight matrix and bias vector of the classification layer respectively, z is the raw score for each category, where p is the probability vector, where is the probability that the input sequence belongs to the
[0079] This method can capture deep features and complex dependencies in the sequence. The application of the Dropout layer enhances the generalization ability of the model on new data. The linear classification layer converts the output of the model into class probabilities, and has high accuracy and reliability in classifying transcription factors that preferentially bind methylated DNA. Compared with traditional sequence-based prediction techniques, by combining large model techniques, the processing efficiency is improved, and the intrinsic features of the sequence are adaptively learned, improving indicators such as prediction accuracy, sensitivity, specificity, Matthews correlation coefficient, and area under the ROC curve.
[0080] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. And the changes and alterations made by those skilled in the art that do not depart from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for identifying methylated DNA-preferentially binding transcription factors based on ProtBERT, characterized in that: The following steps are involved: S1. Obtain a transcription factor sequence dataset and divide it into a training set and a test set; S2. Pad or trim all sequences to the same length and tokenize them through AutoTokenizer; S3. Use BertForSequenceClassification to load the pre-trained ProtBERT model for sequence classification; In step S1, the processing of the sequence data set includes processing of transcription factors that prefer methylated DNA and transcription factors that prefer unmethylated DNA, and the processing steps include: S201. Exclude sequences containing non-standard amino acid residues; S202. Use CD-HIT20 to remove redundant samples; S203. Transcription factors that prefer methylated DNA sequences are extracted as positive samples, and transcription factors that prefer unmethylated DNA are extracted as negative samples; will sequence The embedding layer converts it into an embedding vector, which includes word embedding and position embedding. The word embedding and position embedding are added together to get the final input formula. ; The word embedding is obtained by looking up the integer code in the randomly initialized embedding matrix and merging the corresponding word embedding vectors, while the position embedding is calculated by the sine and cosine functions. Given a position p and dimension i, it is expressed as: ; ; Based on this, the corresponding embedding value is calculated to generate a d-dimensional position embedding vector, which is then merged; Then Enter the Transformer layer and update it layer by layer. The update formula is: , in, It is The output of the layer, It is Output result after layer update; Among them, in the first layer, the hidden state , using the final hidden state of the first labeled CLS , as the summary vector of the entire sequence, is processed by the Droupout layer and the fully connected layer to obtain the original score of each category. The Droupout layer processing formula is: , randomly discard the output of neurons according to the specified probability, and the processing formula of the fully connected layer is: ; Then through the softmax function Convert the original scores into probability distributions of the same dimension; in, is the processed vector, W and b are the weight matrix and bias vector of the linear layer respectively, are the parameters learned during model training, z is the original score of each category, and p is the probability vector.
2. A method for identifying methylated DNA-preferentially binding transcription factors based on ProtBERT according to claim 1, characterized in that: In step S2, a uniform target length is set according to the distribution of transcription factors to cover most sequence lengths. For sequences longer than the target length, they are trimmed from the ends. For sequences shorter than the target length, each transcription factor sequence is decomposed into basic unit amino acids by adding a specific padding marker [PAD] to a specified position of the sequence, mapped to a numerical identifier recognizable by the model according to a defined vocabulary, and special markers [CLS] and [SEP] are added at the beginning and end of the sequence, respectively, to indicate the beginning and end of the entire sequence.
3. A method for identifying methylated DNA-preferentially binding transcription factors based on ProtBERT according to claim 1, characterized in that: In step S3, BertForSequenceClassification is used to load the pre-trained ProtBERT model for sequence classification, and the initial training set is: , The initial test set is: , and represent the protein sequences in the training set and the test set respectively, and They are and The corresponding labels, N and M represent the number of samples in the training set and the test set respectively; the protein sequence needs to be tokenized by AutoTokenizer and marked as , each sequence Converted into a vector through Tokenizer , which contains integer IDs corresponding to each amino acid in the sequence, and the tokenized vector is transformed into , L is the predefined target sequence length, if If the length of is greater than L, the sequence is truncated; if it is less than L, a specific marker [PAD] is padded after it until the length reaches L.
4. A method for identifying methylated DNA-preferentially binding transcription factors based on ProtBERT according to claim 1, characterized in that: In step S3, the processed sequences are organized into batches to prepare for input into the model. The processed sequences are: , B is the batch size, and the sequence first enters the ProtBERT model.
5. A method for identifying methylated DNA-preferentially binding transcription factors based on ProtBERT according to claim 1, characterized in that: In step S3, the model calculates the attention of each element in the input sequence to other elements in the sequence. The model uses this to capture the complex dependencies within the sequence. First, it is necessary to obtain the Query (Q), Key (K), and Value (V) matrices by linearly transforming the input sequence. The mathematical expression is: ; ; ; Where X is the input sequence, , and The weight matrix is the parameters learned during model training; The mathematical representation of calculating self-attention is: , Where Q, K, and V are query, key, and value matrices, respectively. is the dimension of the key, and the softmax function is used to normalize the weights; Each Transformer encoder layer consists of a feed-forward neural network that performs two layers of linear transformations using the ReLU activation function: , in , , , is the parameter learned during model training, and max(0,) is the ReLU activation function, which is used to introduce nonlinearity.
Citation Information
Patent Citations
TF-DNA combination recognition method based on multi-feature fusion
CN115810398A
Cancer detection, classification, prognostication, therapy prediction and therapy monitoring using methylome analysis
WO2019084659A1