A dataset construction method for b cell antigen epitope prediction
By constructing an integrated learning framework based on deep Q_learning and using the IEDB and Uniport databases to screen high-quality non-epitope samples, the problem of unscientific acquisition of non-epitope samples in the existing technology is solved, and the accuracy of B cell antigen epitope prediction is improved.
Patent Information
- Application Number
- CN202411871926.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-12-18
AI Technical Summary
In the existing technology, there is a lack of scientific and effective methods for obtaining non-epitope samples for B cell antigen epitope prediction, resulting in low data set quality and affecting the accuracy of the antigen epitope classification model.
An integrated learning framework based on deep Q_learning was used to construct epitope and non-epitope sample collections through the IEDB database, epitope prediction literature and Uniport database. The CD-hit redundancy removal method and deep Q_learning algorithm were used to train the non-epitope automatic filter to screen out high-quality non-epitope samples and construct a dataset for B cell antigen epitope prediction.
It improves the quality of non-epitope samples, enhances the accuracy of antigen epitope classification, breaks the shortcomings of traditional random selection of non-epitope samples, and realizes the acquisition of high-quality data sets.
Smart Images

Figure CN119864082B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data set construction method for predicting B cell antigen epitopes, and belongs to the technical field of data processing. Background Art
[0002] The present invention is an invention made with the use of the Natural Science Foundation of Hebei Province. The project name of the fund is “Research on B cell antigen epitope prediction based on self-learning mechanism, project number: F2022302001”.
[0003] B cell antigen epitope prediction based on machine learning has important applications in vaccine development and disease test kit development. It can quickly identify antigen epitopes, accelerate R&D progress, and save R&D investment.
[0004] Machine learning-based prediction of B cell antigen epitopes often uses supervised learning, that is, classification based on a dataset. Classification model training must be based on a high-quality sample dataset. Chinese invention patent publication number CN114242169A discloses a method for predicting B cell antigen epitopes. The method first constructs a pre-training set (PT). In each episode of the Q-learning algorithm, the Q agent takes any eight consecutive amino acid residues in the protein primary sequence as a state and selects k residues from the 12 consecutive residues following each state to incorporate into the state as the first action. The Q agent selects one of n complementary classifiers as the second action option and searches the PT using a continuous action search method. The searched amino acid sequences are given an immediate reward using a biased reward rule. The Q value is calculated and updated until the change in the value function is less than 1%. The trained strategy is then used to search for amino acid sequences in the protein primary sequence, and the selected classifier performs classification. This invention significantly enhances the prediction capability of B cell antigen epitopes and improves the accuracy of epitope classification through automatic iteration. Training B cell antigen epitope models typically requires two types of samples: epitope and non-epitope samples. Epitope samples used for B cell antigen epitope prediction are mostly derived from the IEDB (http: / / www.iedb.org / ) database, which is highly reliable. However, there is no reliable source for non-epitope samples. Literature generally uses random selection from the remaining sequence after removing experimentally verified epitopes from the protein primary sequence. This random selection method is unscientific and imprecise. First, the literature only presents the epitope sequences deemed most likely by the experimenter, omitting many possible epitope results. Random selection may result in these missed epitopes being selected as non-epitope samples. Second, our current level of scientific understanding of protein sequences makes it difficult to randomly select high-quality non-epitope samples. In fact, many literatures have pointed out the unreliability of this non-epitope sample selection method. To date, there is a lack of scientifically effective methods for obtaining non-epitope samples. The low quality of non-epitope samples results in low-quality datasets for classification model training, which seriously affects the accuracy of antigen epitope classification models. Summary of the Invention
[0005] The purpose of the present invention is to address the shortcomings of the existing technology and provide a data set construction method for B cell antigen epitope prediction to obtain a sample data set and improve the accuracy of antigen epitope classification.
[0006] The problem described in the present invention is solved by the following technical solutions:
[0007] A method for constructing a dataset for predicting B cell antigen epitopes, the method comprising the following steps:
[0008] a. Establish epitope sample set EPT using IEDB database, epitope prediction literature and uniport database L , non-epitope sample set NEPT L and candidate sequence set CPT L ;
[0009] b. Using CD-hit redundancy removal method to analyze the epitope sample set EPT L and non-epitope sample set NEPT L De-redundancy operations are performed separately, and the epitope sample set EPT after de-redundancy is L and non-epitope sample set NEPT L Randomly extract the same number of samples from the set to form a high-quality positive sample seed set. P And negative sample seed set Seed NP ;
[0010] c. Set the candidate sequence CPT L As a training data set, it is trained based on the deep Q_learning algorithm and uses the epitope sample set EPT. L and non-epitope sample set NEPT L The Q network is optimized and after training is completed, a non-epitope automated filter PC is obtained;
[0011] d. Extract the primary sequence of the antigen protein from the Uniport database and perform automatic non-epitope screening on the sequence according to the set sequence length L using the non-epitope automatic screener PC. The screened non-epitope sequences are combined with the sequences of sequence length L in the IEDB database as a sample data set. The sample data set is de-redundanted using the cd-hit de-redundancy method, and the same number of epitope and non-epitope sequences are selected to form a data set for B cell antigen epitope prediction.
[0012] The above-mentioned method for constructing a dataset for predicting B cell antigen epitopes, wherein the epitope sample set EPT is established L , non-epitope sample set NEPT L and candidate sequence set CPT L The specific methods are as follows:
[0013] Download B cell epitope sequences from the IEDB database, retrieve B cell epitope sequence data with more than 5 entries, and form epitope sample sets EPT according to the sequence length L. L ; Collect non-epitope data from epitope prediction literature, extract non-epitope data that appear more than 5 times from the collected data as negative samples, and form non-epitope sample sets NEPT according to sequence length L L; Collect several protein primary sequences from the uniport database and divide them into sequence segments according to the length L to form the candidate sequence set CPT L .
[0014] The above-mentioned dataset construction method for B cell antigen epitope prediction is used to establish the epitope sample set EPT L , non-epitope sample set NEPT L and candidate sequence set CPT L When the sequence length L is: 7 <L<25。
[0015] The above-mentioned dataset construction method for B cell antigen epitope prediction uses the following specific method for training the deep Q-learning algorithm:
[0016] Each time the training is carried out according to a fixed length L, during the training, the candidate sequence set CPT L As the state space, a sequence is randomly selected as the initial state. The action space contains two actions, "yes" and "no". The batch value is k. The training process is as follows:
[0017] Use s t represents the state at time step t, a t represents the action at time step t; θ i represents the learning network parameters at the i-th iteration, Q(s t ,a t θ i ) represents the state s at the i-th iteration t Next, perform action a t At the beginning of training, initialize a learning rate r between 0 and 1, a discount factor γ between 0 and 1, a state list, a set number of rounds, and learning network parameters θ i ;
[0018] At time step t = 1, the initial action a0, initial state s0, and batch value k are input into a convolutional neural network consisting of 3 convolutional and 2 linear layers to obtain an initial state and action corresponding to Q(s0, a0; θ0), where θ0 represents the initial value of the learning network parameters;
[0019] At each time step t, t≥1, the agent observes the current state s t , and select an action a from the action space t , then the environment calculates the state s at the i-th iteration t Next, perform action a t The value reward Q(s t ,a t θ i ), and then the state st Corresponding action a t The reward value is updated to the state list. When the set number of time steps is reached, it is called a round. When the entire training reaches the set number of rounds, a training is completed.
[0020] The above-mentioned method for constructing a dataset for predicting B cell antigen epitopes, wherein the environment calculates the state s at the i-th iteration t Next, perform action a t The value reward Q(s t ,a t θ i ) are as follows:
[0021] The agent selects m based on each state and the reward value of the action at that time. a axQ t (s t ,a t θ i ) action, the sequence set P1 that performs the "yes" action is merged into Seed P Form a new set Seed P ', the sequence set P2 that performs the "No" action is merged into Seed NP Form a new set Seed NP ', extract the collection Seed P ' and Seed NP 'Based on the physicochemical characteristics of amino acids in the sequence, the support vector machine SVM is used to train the classifier C', and then the classification scores of the sets P1 and P2 are normalized by C', and the classifier scores are divided into four levels from small to large, and the corresponding weights P are assigned to each level. t : 1, 0.8, 0.8, 1, and these weights are combined into a vector in sequence order and multiplied with the calculated reward value, and then iterative optimization learning is performed in the Q network. In the learning network, the state s at the i-th iteration t Next, perform action a t The value reward Q(s t ,a t θ i ) is calculated as follows:
[0022] Q(s t ,a t θ i )=P t R t +γQ(s t+1 ,a t+1 θ i )
[0023] where R t Is to perform action a tThe immediate reward after Q(s t+1 ,a t+1 θ i ) is the next state s t+1 Next, perform action a t+1 The maximum value reward;
[0024] The state s at the i-1th iteration t The value reward V(θ i-1 ) is calculated as:
[0025]
[0026] in, is the expectation, V(s t+1 ,a t+1 θ i-1 ) is the next state s in the learning network t+1 The value reward below;
[0027] The formula for minimizing the loss of gradient descent of network parameters at the i-th iteration is:
[0028]
[0029] The above-mentioned dataset construction method for B cell antigen epitope prediction is used to construct a dataset for any state sequence s t The maximum consecutive overlapping matching reward is calculated according to the formula:
[0030]
[0031]
[0032] in and For state s t Epitope sample set EPT L and non-epitope sample set NEPT L Sequence matching reward, and State s t In EPT L and NEPT L The maximum number of consecutive coincidences in , and State s t In EPT L and NEPT L The maximum continuous overlap ratio in state s, L is the maximum continuous overlap ratio in state s t The number of amino acids in the corresponding sequence;
[0033] At time step t, perform action a tThe immediate reward R t The calculation formula is:
[0034]
[0035] in
[0036] Beneficial effects
[0037] The present invention adopts an integrated learning framework based on deep Q-learning to learn non-epitope data in noisy protein sequence fragment data. Compared with traditional methods, it can effectively improve the quality of non-epitope samples, thereby obtaining high-quality sample data sets and improving the accuracy of antigen epitope classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention will be further described below in conjunction with the accompanying drawings.
[0039] Figure 1 This is a schematic diagram of the composition principle of this method.
[0040] Figure 2 This is a diagram of the dataset construction framework.
[0041] The symbols in the text represent the following: IEDB database: Immune Epitope Database (IEDB) funded by the National Institute of Allergy and Infectious Diseases (NIAID);
[0042] Uniport database: It is the protein database with the richest information and the broadest resources, which is composed of the data of Swiss-Prot, TrEMBL and PIR-PSD databases;
[0043] EPT L : epitope sample collection;
[0044] NEPT L : non-epitope sample collection;
[0045] CPT L : candidate sequence set;
[0046] cd-hit: a redundancy removal method;
[0047] Seed P : positive sample seed set;
[0048] Seed NP : negative sample seed set;
[0049] Q network: refers to the network trained with the deep Q_learning algorithm;
[0050] PC: non-epitope automated screener;
[0051] Sequence length L: refers to the number of amino acids in the sequence (from the primary sequence of the protein) Batch value k: refers to the number of samples in each batch of batch training;
[0052] Time step t: refers to the state value of a time series;
[0053] s t : represents the state at time step t;
[0054] a t : represents the action at time step t;
[0055] θ i : represents the learning network parameters at the i-th iteration;
[0056] Q(s t ,a t θ i ): indicates the state s at the i-th iteration t Next, perform action a t Value rewards;
[0057] Learning rate r: is a number between 0 and 1;
[0058] Discount factor γ: a number between 0 and 1;
[0059] a0: is the initial action;
[0060] s0: is the initial state;
[0061] Q(s0,a0;θ0): a set of network parameters obtained by training the convolutional neural network;
[0062] θ0: represents the initial value of the learning network parameters;
[0063] Q t (s t ,a t θ i ) The corresponding action with the largest action value;
[0064] C': is a classifier obtained through training;
[0065] Pt: refers to weight;
[0066] Refers to expectations;
[0067] For state s t Epitope sample set EPT L Sequence matching reward in ;
[0068] For state s t Non-epitope sample set NEPT L Sequence matching reward in ;
[0069] For state s t In EPT L The maximum number of consecutive coincidences in ;
[0070] For state s t In NEPT L The maximum number of consecutive coincidences in ;
[0071] For state s t In EPT L The proportion of the maximum continuous overlap in ;
[0072] For state s t In NEPT L The proportion of the maximum continuous overlap in ;
[0073] L: state s t The number of amino acids in the corresponding sequence. DETAILED DESCRIPTION
[0074] The present invention provides a dataset construction method for predicting B cell antigen epitopes. This method uses an integrated learning framework based on deep Q-learning to learn non-epitope data from noisy protein sequence fragment data. The method includes a candidate set of protein sequence fragments, a deep Q-learning-based strategy learner, and a non-epitope sample filter. This method can automatically label non-epitope samples. The method includes the following steps:
[0075] a. Download B cell epitope sequences from the IEDB database, and then retrieve B cell epitope sequence data with more than 5 entries, and form epitope sample sets EPT according to the sequence length L. L , usually 7 <L<25;从表位预测文献中采集非表位数据,从采集到的数据中抽取出现超过5次的非表位数据作为负样本,按照序列长度L分别组成非表位样本集合NEPT L , usually 7 <L<25;从uniport数据库中采集若干条蛋白质一级序列,按照长度L统一分割成序列片段分别组成候选序列集合CPT L ;
[0076] b. Using the CD-hit redundancy removal method to remove the epitope sample set EPT obtained in step a L and non-epitope sample set NEPT L Perform redundancy removal operations separately to find samples with high similarity. Then delete samples with high similarity and update the epitope sample set EPT. L and non-epitope sample set NEPT L From the redundancy-free epitope sample set EPT L and non-epitope sample set NEPT L Randomly extract the same number of samples from the set to form the positive sample seed set Seed P And negative sample seed set Seed NP .
[0077] c. Set the candidate sequence CPT L As a training data set, the deep Q_learning algorithm is trained, and each training is carried out according to a fixed length L. <L<25。在训练中,将候选序列集合CPT L As the state space, a sequence is randomly selected as the initial state. The action space contains two actions, "yes" and "no". The batch value is k, that is, k candidate sequences are input each time. <k<100。基于深度Q_learning算法的训练流程如下:
[0078] Use s t represents the state at time step t, a t represents the action at time step t; θ i represents the learning network parameters at the i-th iteration, Q(s t ,a t θ i ) represents the state s at the i-th iteration t Next, perform action a t At the beginning of training, initialize a learning rate r between 0 and 1, a discount factor γ between 0 and 1, a state list, a set number of rounds, and learning network parameters θ i .
[0079] The deep Q_learning algorithm is a new algorithm that combines Q_learning and convolutional neural networks. At time step t=1, the initial action a0, initial state s0, and batch value k are input into a convolutional neural network containing 3 convolutional layers and 2 linear layers to obtain Q(s0, a0; θ0) corresponding to the initial state and action, where θ0 represents the initial value of the learning network parameters.
[0080] At each time step t, t≥1, the agent observes the current state s t, and select an action a from the action space t , then the environment calculates the state s at the i-th iteration according to the “synchronous optimization strategy” and “matching reward rule” t Next, perform action a t The value reward Q(s t ,a t θ i ), and then the state s t Corresponding action a t The reward value is updated to the state list, and the agent updates its state by executing actions. In a given training dataset, this method is used to complete a set number of time steps, which is called an episode. When the entire training process reaches the set number of episodes, the training is complete, and the trained deep Q-learning algorithm becomes an excellent non-epitope automated filter PC.
[0081] d. Extract several primary antigen protein sequences from the Uniport database and perform automated non-epitope screening using the non-epitope automatic screener PC according to a set sequence length, L. The screened non-epitope sequences are combined with sequences of length L in the IEDB database as a sample data set. The sample data set is de-redundanted using the cd-hit de-redundancy method, identifying and deleting samples with high similarity. By selecting the same number of epitope and non-epitope sequences, an excellent data set can be formed.
[0082] The "synchronization optimization strategy" in step c is to use the "action selection strategy" to select the action, and the sequence set P1 of the agent executing the "yes" action is incorporated into the Seed P Form a new collection Seed P ', the sequence set P2 that performs the "No" action is merged into Seed NP Form a new set Seed NP '. Extract collection Seed P ' and Seed NP 'Based on the physicochemical characteristics of amino acids in the sequence, the support vector machine SVM is used to train the classifier C', and then the classification scores of the sets P1 and P2 are normalized by C', and the classifier scores are divided into four levels from small to large, and the corresponding weights P are assigned to each level. t : 1, 0.8, 0.8, 1, and organize these weights into a vector in sequence order and perform point multiplication with the reward value calculated in the "matching reward rule", and then perform iterative optimization learning in the Q network.
[0083] In the learning network, the state s at the i-th iteration t Next, perform action a t The value reward Q(s t ,at θ i ) is calculated as follows:
[0084] Q(s t ,a t θ i )=P t R t +γQ(s t+1 ,a t+1 θ i )(1)
[0085] where R t Is to perform action a t The immediate reward after Q(s t+1 ,a t+1 θ i ) is the next state s t+1 Next, perform action a t+1 The maximum value reward.
[0086] In the learning network, the state s at the i-1th iteration is t The value reward V(θ i-1 ) is calculated as:
[0087]
[0088] in, is the expectation, V(s t+1 ,a t+1 θ i-1 ) is the next state s in the learning network t+1 The value reward below.
[0089] In the learning network, the formula for minimizing the loss of the gradient descent of the network parameters at the i-th iteration is:
[0090]
[0091] The "matching reward rule" in step c follows the state sequence s t EPT with epitope sample collection L and non-epitope sample set NEPT L The proportion and number of the maximum continuous overlap are calculated.
[0092] For any state sequence s t The maximum consecutive overlapping matching reward is calculated according to the formula:
[0093]
[0094]
[0095] in and is the state s t respectively in epitope sample set EPT L and non-epitope sample set NEPT L the matching reward of the sequence, and is the state s t the maximum number of consecutive coincident occurrences in EPT L and NEPT L and is the state s t the proportion of the maximum number of consecutive coincident occurrences in EPT L and NEPT L L is the number of amino acids corresponding to the state s t .
[0096] The calculation formula of the immediate reward R t after performing the action a t at the time step t is:
[0097]
[0098] wherein
[0099] The "action selection strategy" in the above steps is that the agent selects the action with the maximum reward value according to each state and the action reward value at the time, and selects the action as the result.
[0100] Advantages of the present application:
[0101] The epitope prediction method based on continuous action search adopted by the present application breaks the epitope prediction scheme based on the "window method" on one hand, and realizes autonomous selection of epitope sequences on the other hand. Through the introduction of a complementary classifier, the classification accuracy is improved.
[0102] The method is suitable for the construction of different length data sets required for window method prediction of epitopes, and formulates the seeds required for policy learning of deep Q_learning.
[0103] Considering that the number of sequences as states is huge, the present application adopts a deep learning strategy, which not only realizes easy training but also guarantees excellent learning effect.
[0104] The method adopts a "synchronous optimization strategy" and a "matching reward rule", which can guarantee iterative optimization of the learning strategy, and also uses the respective features of non-epitope and epitope samples in learning.
[0105] The automatic screening of non-epitope samples in the method has the characteristics of rapidity, effectiveness and accuracy.
[0106] Introduction to Q-Learning algorithm:
[0107] (1) In the Q-Learning algorithm, each state-action has a corresponding Q value. Therefore, the learning process of the Q-Learning algorithm is to repeatedly calculate the Q value of the learning state-action pair. Finally, the optimal action strategy obtained by the learner is to choose the action corresponding to the maximum Q value in state s. The Q value Q(s, a) of action a in state s is defined as the cumulative reward value obtained after the learner performs action a in state s and continues to execute according to a certain action strategy. The basic equation for its Q value update is:
[0108] Q(s t ,a t )=Q(s t ,a t )+α[Rs t +γmaxQ(s t+1 ,a)-Q(s t ,at)](6-1)
[0109] In formula (6-1): a is the optional action under the state; Rs t is the immediate reward given by the environment in state s at time t; α is the learning rate; Q(s t ,a t ), the evaluation value of the state-action (s, a) at time t.
[0110] (2) The steps of the Q-Learning algorithm are shown in Table 1:
[0111] Table 1. Pseudocode of Q_learning algorithm
[0112]
[0113] (3) Deep Q-Learning Algorithm
[0114] In deep Q-Learning, due to the powerful expressive power of neural networks, traditional Q tables are replaced by a deep neural network, which approximates the Q function. The network's input is the state of the environment, and its output is the expected reward for each action. The problem of labels and loss functions arises. The goal of reinforcement learning is to develop a relatively good policy, which requires a relatively appropriate value estimation function.
[0115] One implementation step of the deep Q-Learning algorithm:
[0116] 1) Network Initialization: Initialize the Q network Q(s, a; θ) and the target network Q(s, a; θ-). These two networks have the same structure but independent parameters. The target network's parameters θ- are periodically copied and updated from the Q network.
[0117] 2) Observation initialization: Obtain the initial state information s_1 by randomly selecting actions in the environment.
[0118] 3) For each step t, perform the following operations:
[0119] Strategy execution: select action a according to the ε-greedy strategy, randomly select an action with probability ε, and select argmax with probability 1-ε a Q(st,a;θ).
[0120] Take action: Perform action a in the environment and observe the reward r and the new state s_{t+1}.
[0121] Calculate the target Q value: For each sample, calculate the corresponding target Q value: Q(s,a) = r + γmaxa′Q(s′,a′). Where r is the reward, s′ is the next state, γ is the discount factor, and a′ is the optimal action for the next state. Using the target network to calculate the target Q value can improve the stability of the algorithm.
[0122] Gradient descent: Perform gradient descent on the network parameters to minimize the loss function:
[0123] L(θ)=(1 / N)*Σj(yj-Q(sj,aj;θ)) 2, Where N is the number of samples drawn.
[0124] 4) Target network update: After a certain number of steps, the parameters θ of the Q network are used to update the parameters θ- of the target network.
[0125] 5) Loop: Repeat step 4 until the agent shows stable learning performance or reaches the predetermined number of iterations.
[0126] 6) Termination: When the agent's performance meets certain criteria or reaches the iteration limit, the training process ends.
[0127] (4) CD-hit redundancy removal method
[0128] CD-HIT is a dataset de-redundancy method that identifies highly similar sequences within a dataset, facilitating their removal to ensure data specificity. CD-HIT works by clustering all sequences according to predefined parameters and outputting the longest sequence within each cluster as the representative sequence. It also provides the names of each sequence within each cluster for similarity analysis. Below is a brief introduction to its usage.
[0129] cd-hit needs to be run under Linux. After unzipping the package, go to the software directory and directly enter the command "make" to compile it.
[0130] The input file of cd-hit is only a fasta format file. Generally speaking, cd-hit clusters the gene or protein sequences of several samples, so the sequences of these samples need to be aggregated together as an input file. This can be achieved through the cat command in Linux system:
[0131] cat a.fasta b.fasta c.fasta>all.fasta
[0132] Among them, a.fasta, b.fasta, and c.fasta are three sample protein sequences in fasta format, and all.fasta is the summarized sequence, which is used as the input sequence of cd-hit in the analysis.
[0133] It's important to note that no sequence with the same name can exist in any of the three sample sequences, as this will result in an error. Therefore, the sample name is typically prefixed to each sample sequence name during analysis to avoid duplication. The sequence name is the content of the line beginning with a ">" sign and preceding the space in the FASTA file.
[0134] There are many parameters that can be adjusted when running cd-hit. The command to run it is as follows (the parameters are only examples):
[0135] / path-to-cdhit-4.6.8 / cd-hit-i all.fasta-o new.fa-c 0.8-aS 0.8-d 0
[0136] Important parameters of cd-hit:
[0137] -i: input file, fasta format
[0138] -o: Output file prefix. There are two output files: fasta format sequence file and clustering information file ending with .clstr
[0139] -c: If the ratio of the bp of the shorter sequence aligned to the longer sequence to the bp of the shorter sequence itself exceeds this value, the shorter sequence will be clustered into one group. The default value is 0.9
[0140] -d: The length of the sequence name in each cluster group in the cluster information file. If it is set to 0, the full sequence name will be used.
[0141] -aL: controls the stringency of the sequence alignment. The default value is 0. If it is set to 0.8, it means that the alignment interval should account for 80% of the representative (long) sequence.
[0142] -AL: Parameter that controls the stringency of sequence alignment. The default value is 99999999. If it is set to 40, it means that the non-aligned interval of the sequence is shorter than 40bp.
[0143] -aS: Parameter that controls the strictness of short sequence alignment. The default value is 0. If it is set to 0.8, it means that the alignment interval will account for 80% of the short sequence.
[0144] -AS: Parameter that controls the strictness of short sequence alignment. The default value is 99999999. If it is set to 40, it means that the non-alignment interval of the short sequence must be shorter than 40bp.
[0145] Explanation of professional terms
[0146] (1) Protein primary sequence, a sequence consisting of 20 amino acids, such as ADFCEGHIKLST
[0147] (2) B cell epitopes are part of the primary sequence of a protein and can be composed of partial sequences
[0148] (3) An “episode” is the process from the beginning to the end of an agent executing a strategy in an environment.
[0149] Agent: A software or hardware mechanism that interacts with its environment to take appropriate actions. Also known as an intelligent agent in Chinese, and often referred to as a Q-agent in Q-learning algorithms, it is the subject that explores and learns within the environment.
[0150] Actions: The various possible actions that the agent can take. While actions are somewhat self-explanatory, we still need to allow the agent to choose from a set of discrete, possible actions.
[0151] Environment: The external environment interacts with and responds to the agent. It takes the agent's current state and action as input and outputs the agent's reward and next state. The environment is everything outside the agent.
[0152] State: A state is a specific, immediate situation discovered by an agent, including a specific place, moment, and the instantaneous configuration that relates the agent to other important things.
[0153] Reward: A reward is a type of feedback that we can use to measure the success or failure of various actions by an agent in a given state.
[0154] Discount factor: A discount factor is a multiplier. Future rewards discovered by an agent are multiplied by this factor to weaken the cumulative impact of such rewards on the agent's current action selection. This is the core of reinforcement learning, i.e., by gradually reducing the value of future rewards, more weight is given to recent actions. This is crucial for paradigms based on the "delayed action" principle.
[0155] Policy: A policy is a function that takes an input state observation and outputs an action. It is the strategy that an agent uses to determine the next action based on the current state. It is able to map different states to various actions in order to promise the highest reward.
[0156] Value: It is defined as the long-term expected reward (not the short-term reward) of a current state with a discount under a particular policy. The short-term reward is the instantaneous reward obtained by the agent when it is in a certain state and takes a particular action. The value is the total amount of reward that the agent expects to obtain from a certain state until the future.
[0157] Q-value or action-value: The difference between "value" and "Q-value" is that Q-value requires an additional parameter, which is the current action. It refers to the long-term reward generated by a particular action from a current state under a particular policy.
[0158] Bellman equation: It is a set of equations that decompose the value function into the immediate reward plus the discounted future value.
[0159] Value iteration: This is an algorithm that iteratively improves estimates of the value to calculate a function with the best state values. The algorithm initializes the value function to any random value, then repeatedly updates the values of Q and the value function until they converge.
[0160] Policy iteration: Since the agent only focuses on finding the optimal policy, and the optimal policy sometimes converges before the value function, policy iteration should not repeatedly improve estimates of the value function. Instead, it needs to redefine the policy at each step and calculate the value based on the new policy until the policy converges.
[0161] episode: Reinforcement learning describes the interaction between an agent and an environment, where at a certain time the agent takes an action a t Give action a t, the environment gives feedback r according to the action taken by the agent t , and the environment also updates s t = 1, the agent makes a new action according to the feedback. This cycle repeats, from start to end, s1, a1, r1, s2, a2, r2, … is called an episode.
[0162] Timestep: In reinforcement learning, each interaction between the agent and the environment constitutes a timestep.
[0163] Primary sequence of the protein in uniport:
[0164] For example A0A0A0MQN9:
[0165] MEQPPPLAPEPASARSRRRREPESPPAPIPLFGARTVVQRSPDEPALSKAEFVEKVRQSNQACHDGDF
[0166] HTAIVLYNEALAVDPQNCILYSNRSAAYMKTQQYHKALDDAIKARLLNPKWPKAYFRQGVALQYLGRH
[0167] ADALAAFASGLAQDPKSLQLLVGMVEAAMKSPMRDTLEPTYQQLQKMKLDKSPFVVVSVVGQELLTAG
[0168] HHGASVVVLEAALKIGTCSLKLRGSVFSALSSAHWSLGNTEKSTGYMQQDLDVAKTLGDQTGECRAHG
[0169] NLGSAFFSKGNYREALTNHRHQLVLAMKLKDREETIVCVSRGRYTATSSQLHTGWGETSQGLPSAASS
[0170] ALSSLGHVYTAIGDYPNALASHKQCVLLAKQSKDDLSEARELGNMGAVYIAMGDFENAVQCHEQHLRI
[0171] AKDLGSKREEARAYSNLGSAYHYRRNFDKAMSYHNCVLELAQELMEKPIEMRAYAGLGHAARCMQDLE
[0172] RAKQYHEQQLGIAEDLKDRAAEGRASSNLGIIHQMKGDYDTALKLHKTHLCIAQELSDYAAQGRAYGN
[0173] MGNAYNALGMYDQAVKYRHQELQISMEVNDRASQASTHGNLAVAYQALGAHDRALQHYQNHNLIAREL
[0174] RDIQSEARALSNLGNFHCSRGEYVQAAPYYEQYLRLAPDLQDMEGGKVCHNLGYAHYCLGNYQEAVK
[0175] YYEQDLALAKDLHDKLSQAKAYCNLGLAFKALLNFAKAEECQKYLLSLAQSLDNSQAKFRALGNLGDI
[0176] FICKKDINGAIKFYEQQLGLSHHVKDRRLEASAYAALGTAYRMVQKYDKALGYHTQELEVYQELSDLP
[0177] GECRAHGHLAAVYMALGKYTMAFKCYQEQLELGRKLKEPSLEAQVYGNMGITKMNMNVMEDAIGYFEQ
[0178] QLAMLQQLSGNESVLDRGRAYGNLGDCYEALGDYEEAIKYYEQYLSVAQSLNRMQDQAKAYRGLGNGH
[0179] RATGSLQQALVCFEKRLVVAHELGEASNKAQAYGELGSLHSQLGNYEQAISCLERQLNIARDMKDRAL
[0180] ESDAACGLGGVYQQMGEYDTALQYHQLDLQIAEETDNPTCQGRAYGNLGLTYESLGTFERAVVYQEQH
[0181] LSIAAQMNDLVAKTVSYSSLGRTHHALQNYSQAVMYLQEGLRLAEQLGRREDEAKIRHGLGLSLWASG
[0182] NLEEAQHQLYRASALFETIRHEAQLSTDYKLSLFDLQTSSYQALQRVLVSLGHHDEALAVAERGRTRA
[0183] FADLLVERQTGQQDSDPYSPITIDQILEMVNAQRGLVLYYSLAAGYLYSWLLAPGAGILKFHEHYLGD
[0184] NSVESSSDFQAGSSAALPVATNSTLEQHIASVREALGVESYYSRACASSETESEAGDIMEQQLEEMNK
[0185] QLNSVTDPTGFLRMVRHNNLLHRSCQSMTSLFSGTVSPSKDGTSSLPRRQNSLAKPPLRALYDLLIAP
[0186] MEGGLMHSSGPVGRHRQLVLVLEGELYFVPFALLKGSASNEYLYERFTLIAVPAVRSLGPHSKCHLRK
[0187] TPPTYSSSTTMAAVIGNPKLPSAVMDRWLWGPMPSAEEEAFMVSELLGCQPLVGSMATKERVMSALTQ
[0188] AECVHFATHVSWKLSALVLTPNTEGNPAGSKSSFGHPYTIPESLRVQDDASDVESISDCPPLRELLLT
[0189] AADLLDLRLSVKLVVLSSSQEANGRVTADGLVALTRAFLAAGAQCVLVALWPVPVAASKMFVHAFYSS
[0190] LLNGLKASASLGEAMKVVQSSKAFSHPSNWAGFTLIGSDVKLNSPSSLIGQALTEILQHPERARDALR
[0191] VLLHLVEKSLQRIQNGQRNAMYTSQQSVENKVGGIPGWQALLTAVGFRLDPAASGLPAAVFFPTSDPG
[0192] DRLQQCSSTLQALLGLPNPALQALCKLITASETGEQLISRAVKNMVGMLHQVLVQLQACEKEQDFASA
[0193] PIPVSLSVQLWRLPGCHEFLAALGFDLCEVGQEEVILKTGKQASRRTTHFALQSLLSLFDSTELPKRL
[0194] SLDSSSSLESLASAQSVSNALPLGYQHPPFSPTGADSIASDAISVYSLSSIASSMSFVSKPEGGLEGG
[0195] GPRGRQDYDRSKSTHPQRATLPRRQTSPQARRGASKEEEEYEGFSIISMEPLATYQGEGKTRFSPDPK
[0196] QPCVKAPGGVRLSVSSKGSVSTPNSPVKMTLIPSPNSPFQKVGKLASSDTGESDQSSTETDSTVKSQE
[0197] ESTPKLDPQELAQRILEETKSHLLAVERLQRSGGPAGPDREDSVVAPSSTTVFRASETSAFSKPILSH
[0198] QRSQLSPLTVKPQPPARSSSLPKVSSPATSEVSGKDGLSPPGSSHPSPGRDTPVSPADPPLFRLKYPS
[0199] SPYSAHISKSPRNTSPACSAPSPALSYSSAGSARSSPADAPDEKVQAVHSLKMLWQSTPQPPRGPRKT
[0200] CRGAPGTLTSKRDVLSLLNLSPRHGKEEGGADRLELKELSVQRHDEVPPKVPTNGHWCTDTATLTTAG
[0201] GRSTTAAPRPLRLPLANGYKFLSPGRLFPSSKC
[0202] 表位序列样本:
[0203] 例如:LASSDTGESDQSSTE
[0204] 非表位序列样本:
[0205] For example: GTLTSKRDVLSLLNLSPR.
Claims
1. A method for constructing a dataset for predicting B cell antigen epitopes, characterized in that: The method comprises the following steps: a. Establish epitope sample set EPT using IEDB database, epitope prediction literature and uniport database L , non-epitope sample set NEPT L and candidate sequence set CPT L ; b. Using CD-hit redundancy removal method to analyze the epitope sample set EPT L and non-epitope sample set NEPT L De-redundancy operations are performed separately, and the epitope sample set EPT after de-redundancy is L and non-epitope sample set NEPT L Randomly extract the same number of samples from the set to form the positive sample seed set Seed P And negative sample seed set Seed NP ; c. Set the candidate sequence CPT L As a training data set, the deep Q_learning algorithm is trained and the epitope sample set EPT is used. L and non-epitope sample set NEPT L Optimize the Q network to make it a non-epitope automated screener PC based on the deep Q_learning algorithm; d. Extract the primary sequence of the antigen protein from the Uniport database and perform automatic non-epitope screening on the sequence according to the set sequence length L using the non-epitope automatic screener PC. The screened non-epitope sequences are combined with the sequences of sequence length L in the IEDB database as a sample data set. The sample data set is de-redundanted using the cd-hit de-redundancy method, and the same number of epitope and non-epitope sequences are selected to form a data set for B cell antigen epitope prediction.
2. The method for constructing a dataset for predicting B cell antigen epitopes according to claim 1, wherein: Establishing epitope sample set EPT L , non-epitope sample set NEPT L and candidate sequence set CPT L The specific methods are as follows: Download B cell epitope sequences from the IEDB database, retrieve B cell epitope sequence data with more than 5 entries, and form epitope sample sets EPT according to the sequence length L. L ; Collect non-epitope data from epitope prediction literature, extract non-epitope data that appear more than 5 times from the collected data as negative samples, and form non-epitope sample sets NEPT according to sequence length L L ; Several protein primary sequences were collected from the uniport database and divided into sequence segments according to the length L to form the candidate sequence set CPT. L .
3. The method for constructing a dataset for predicting B cell antigen epitopes according to claim 2, wherein: Epitope Sample Collection EPT L , non-epitope sample set NEPT L and candidate sequence set CPT L When the sequence length L is: 7 <L<25。 4. The method for constructing a dataset for predicting B cell antigen epitopes according to claim 1, wherein: The specific method for training the deep Q_learning algorithm is: Each time the training is carried out according to a fixed length L, during the training, the candidate sequence set CPT L As the state space, a sequence is randomly selected as the initial state. The action space contains two actions, "yes" and "no". The batch value is k. The training process is as follows: use Represents the time step The state of Represents the time step Actions when Indicates the The learning network parameters at the iteration, Indicates the The state at the iteration Next action At the beginning of training, initialize a learning rate r of any number between 0 and 1 and a discount factor of any number between 0 and 1. , a state list, a set number of rounds, and learning network parameters ; At time step In the initial action , initial state , batch value k is input into a convolutional neural network containing 3 convolutional layers and 2 linear layers to obtain an initial state and action corresponding to ,in Represents the learning network parameters at iteration 0; At each time step , the agent observes the current state s t , and select an action from the action space , then the environment calculates The state at the iteration Next action Value Reward , and then the state s t Corresponding actions The reward value is updated to the state list. When the set number of time steps is reached, it is called a round. When the entire training reaches the set number of rounds, a training is completed.
5. The method for constructing a dataset for predicting B cell antigen epitopes according to claim 4, wherein: The environment calculates The state at the iteration Next action Value Reward The specific methods are as follows: The agent selects the reward value of each state and action at that time. Action, the sequence set P1 of executing the "yes" action is merged into Seed P Form a new set Seed P ', the sequence set P2 that performs the "no" action is merged into Seed NP Form a new set Seed NP ', extract the collection Seed P ' and Seed NP 'Based on the physicochemical characteristics of amino acids in the sequence, use support vector machine SVM to train classifier C', then use C' to normalize the classification scores of set P1 and set P2, and divide the classifier scores into four levels from small to large, and assign corresponding weights to each level. P t : 1, 0.8, 0.8, 1, and these weights are combined into a vector in sequence order and multiplied with the calculated reward value, and then iterative optimization learning is performed in the Q network. In the learning network, The state at the iteration Next action Value Reward The calculation formula is: ; in Is to perform an action Instant rewards after Is the next state Next action The maximum value reward; No. The state at the iteration Value Reward The calculation formula is: ; in, It's expectation. is the next state in the learning network The value reward below; No. The formula for minimizing the loss of gradient descent of network parameters at the iteration is: 。 6. The method for constructing a dataset for predicting B cell antigen epitopes according to claim 5, wherein: For any state sequence The maximum consecutive coincidence matching reward is calculated according to the formula: ; in and Status Epitope sample set EPT L and non-epitope sample set NEPT L Matching reward for the sequence, Status In EPT L and NEPT L The maximum number of consecutive coincidences in , Status In EPT L and NEPT L The maximum proportion of consecutive overlaps in Status The number of amino acids in the corresponding sequence; At time step , perform the action Instant rewards after The calculation formula is: ; in .
Citation Information
Patent Citations
Prediction method of protein antigenic epitope
CN107341363A
Antigenic epitope prediction method for B cells
CN114242169A