Method for designing mhc class ii binding peptides based on evolutionary information and transformer neural network algorithm

By combining evolutionary information with the Transformer neural network algorithm, a highly efficient and reliable MHCII binding peptide was designed, which solved the problems of long computation time and low efficiency in traditional methods. The generated peptide sequence has high similarity to the natural sequence, strong binding affinity, and high structural reliability.

CN120727089BActive Publication Date: 2025-11-21WENZHOU INST UNIV OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511211222.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-21
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively design MHCII-binding peptides. Traditional methods are not applicable to the structural design of short peptides, and the Mond Carlo simulated annealing algorithm is computationally time-consuming and inefficient, failing to simultaneously optimize for first-order and second-order conservation.

Method used

A method based on evolutionary information and Transformer neural network algorithm is adopted. By fusing convolutional neural network and Transformer module, amino acid frequency and joint frequency features are extracted, global dependencies are extracted using self-attention mechanism, the probability distribution of MHCII binding peptide is generated, and the target sequence is generated by random sampling.

Benefits of technology

It improves the reliability and efficiency of short peptide design, can simultaneously satisfy first-order and second-order conservation, generates peptide sequences with high similarity to natural sequences, strong binding affinity, and high structural reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120727089B_ABST
    Figure CN120727089B_ABST
Patent Text Reader

Abstract

The application discloses a design method of MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm, relates to the field of protein design, and comprises the following steps: S1, extracting evolutionary information features of alleles of each MHCII molecule and binding core sequences of corresponding binding peptides; S2, establishing a neural network model based on convolution modules and Transformer modules, taking two frequency feature tensors extracted in S11 and S12 as double inputs, and finally obtaining probability distributions of 20 amino acids at each position of each sequence; and S3, according to an output result of the neural network model, performing random sampling according to probability to generate binding core sequences of MHCII-peptides satisfying a target distribution. The method introduces evolutionary information such as sequence position amino acid frequency (first-order conservation analysis) and joint frequency of amino acid pairs (second-order conservation analysis) to design new short peptide sequences, solves the problem that short peptides cannot be designed based on structures, and provides reliability of designing short peptide sequences based on evolutionary information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of protein design, and particularly relates to a method for designing MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm. BACKGROUND

[0002] MHCII molecules play a key role in immune responses, are widely expressed on the surface of antigen-presenting cells, can capture exogenous antigen peptides and present them to helper T cells, thereby activating humoral and cellular immune responses. Therefore, precisely designing peptide segments that can bind to MHCII molecules not only helps to understand the mechanism of immune responses, but also can be used to identify T cell epitopes and promote the development of new vaccines and immunotherapy.

[0003] In recent years, with the development of bioinformatics and protein engineering, studies have shown that even without structural information, functional proteins can be designed through evolutionary characteristics at the sequence level. By aligning and analyzing the amino acid sequences of homologous proteins, the evolutionary information can be effectively tracked. For a target peptide sequence family, statistical modeling methods can be used to quantitatively analyze MSA data to capture key sequence features related to function: at the single-residue level, conservation analysis can identify sites that are critical to the function of the peptide sequence or the structure of the MHCII complex; at the residue pair level, a two-site co-evolution model can capture the cooperative variation pattern between homologous peptide sequences, thereby revealing the functional or structural coupling relationship between residues.

[0004] The peptide binding groove of MHCII is composed of its alpha chain and beta chain, and is in an open state, which can accommodate peptides of different lengths. There are nine characteristic pockets in the binding interface, corresponding to the nine core residues in the binding peptide. These nine sites together form the binding core of MHCII-peptide, which is the decisive region of the peptide-MHCII interaction.

[0005] Traditional design methods are based on protein structure design methods, such as AlphaFold structure prediction combined with Rosetta or ProteinMPNN tools, which are widely used in the rational design of stable protein structures. However, this method cannot be directly applied to the design of binding peptides. The main reason for this dilemma is the special structural characteristics of the MHCII-peptide complex: the length of the peptide sequence is relatively short (the length of the binding core sequence of MHCII molecules is even shorter, typically 9-11 amino acids), which limits the formation of stable secondary structures by the peptide segment. This inherent instability in structure limits the applicability of traditional structure design methods.

[0006] Meanwhile, since short peptides are usually disordered and there are a large number of good peptide sequence data, the design can make full use of the evolutionary information extracted by multiple sequence alignment (MSA). However, the traditional Monte Carlo simulated annealing algorithm based on evolutionary information has a long calculation time and can only design a specific MHCII-peptide binding core each time, and the calculation time of each algorithm is also long, which increases the calculation cost and is low in efficiency. In addition, the Monte Carlo simulated annealing algorithm cannot achieve the expected performance effect if only the second-order conservation is used to optimize the objective function without the constraint of amino acid frequency (first-order conservation). The algorithm needs to meet the first-order conservation and second-order conservation of natural peptide sequences to effectively design new peptide sequences. SUMMARY

[0007] In view of the defects in the prior art, the purpose of the present application is to provide a design method of MHCII binding peptide based on evolutionary information and Transformer neural network algorithm.

[0008] To achieve the above purpose, the present application provides the following technical scheme:

[0009] A design method of MHCII binding peptide based on evolutionary information and Transformer neural network algorithm, comprising the following steps:

[0010] S1, obtain the data set of the binding core of the major histocompatibility complex class II molecule MHCII and the peptide, and extract the evolutionary information features of the alleles of each MHCII molecule and the binding core sequence of the corresponding peptide, which comprises:

[0011] S11, calculate the unit point amino acid frequency in the sequence f i a , generate a two-dimensional position amino acid frequency tensor F, F∈R 9×20 , wherein f i a represents the amino acid type at the i-th position of the binding core, i=1,…,9, a corresponding to the frequency of occurrence of 20 standard amino acids, i=1,…,9, a =1,…,20,;

[0012] S12, calculate the joint frequency of amino acid pairs in the sequence f ij ab , generate a four-dimensional amino acid pair joint frequency tensor T, T∈R 9 ×20×9×20 , wherein fij ab represents the amino acid type at the i-th position of the binding core a co-occurrence frequency of the j-th position being amino acid type b;

[0013] S2, a neural network model based on the fusion of convolution modules and Transformer modules is established, and the two frequency feature tensors extracted in S11 and S12 are taken as double inputs;

[0014] S21, two-dimensional tensor F and four-dimensional tensor T are respectively subjected to convolution operation, convolution features are extracted, and the obtained convolution features are connected along the feature dimension for feature fusion;

[0015] S22, the data fused in S21 is input into the Transformer module of the deep learning model to extract data features, and global dependency relationships are extracted through the self-attention mechanism;

[0016] S23, after the feature extraction via the Transformer module, one-dimensional convolution is performed to map the data to the dimension of the output;

[0017] S24, a Softmax function is added to ensure that the last dimension of the output result is a probability, and finally the probability P of each position of each sequence for the 20 kinds of amino acids is obtained i a , a probability distribution tensor P ∈ R N×9×20 is generated, where N is the number of generated sequences;

[0018] S3, according to the output result of the neural network model, random sampling is performed according to the probability to generate a binding core sequence of MHCII-peptide that meets the target distribution.

[0019] Further, each MHCII-peptide in the data set in S1 contains M aligned binding core sequence matrices X= { x µi | i = 1,..., 9, µ = 1,..., M}, wherein x_µi represents the amino acid type at the i-th position of the µ-th sequence;

[0020] In S11, the frequency values of all positions of the sequence matrix and the amino acid types are arranged in position order to form a two-dimensional position amino acid frequency tensor F with a size of 9x20, reflecting the conservation distribution of amino acids at each site in the multiple sequence alignment MSA;

[0021] In S12, the joint frequency values of all position pairs and amino acid type pairs in the sequence matrix are arranged in index order to form a four-dimensional amino acid pair joint frequency tensor T with a size of 9x20x9x20, reflecting the co-evolution relationship of amino acid pairs in the multiple sequence alignment MSA.

[0022] Further, in S21,

[0023] The two-dimensional position amino acid frequency tensor F ∈ R 9×20 After 1D CNN processing by a convolution kernel with a size of 1x1, a high-dimensional frequency feature matrix F' is obtained.

[0024] The four-dimensional amino acid pair joint frequency tensor T ∈ R 9×20×9×20 After 2D CNN processing by a convolution kernel with a size of 16x16, a high-dimensional joint frequency feature matrix T' is obtained.

[0025] The obtained convolution features are connected along the feature dimension by Z = concat(F', T') to perform feature fusion.

[0026] Further, in S22, the fusion data of the processed two-dimensional amino acid frequency tensor F and the four-dimensional amino acid pair joint frequency tensor T is input into a stacked structure composed of 12 Transformer Blocks.

[0027] Further, the core of each Transformer Block is a multi-head self-attention mechanism, and the steps of the multi-head self-attention mechanism are as follows: multiply each input matrix by three learnable matrices to create three vectors: Query matrix, Key matrix and Value matrix.

[0028] Further, in S24

[0029] When only the amino acid pair joint frequency is used as the loss of the network, the loss function is as follows:

[0030]

[0031] Wherein respectively represent the predicted joint frequency and the natural joint frequency of the amino acid type a appearing in position i and the residue type b appearing in position j in the binding core, B is the batch size, and respectively represent the first n MHCII allele corresponding to the binding core, the first i , j amino acid pair in the position a , bpredicted joint frequency and natural joint frequency of amino acid pairs;

[0032] When the frequency of amino acids and the joint frequency of amino acid pairs are used as loss, the loss function is as follows:

[0033]

[0034] wherein is a coefficient used to adjust the balance of the loss, set to 65, and respectively represent the predicted frequency and natural frequency of amino acid types i at positions a in the binding core. and respectively represent the predicted frequency and natural frequency of amino acid types a at position n in the binding core corresponding to the allele of the MHCII of the first kind. i

[0035] Loss1 is to minimize the difference between the joint frequency of amino acid pairs of the designed sequence and the natural sequence, and Loss2 is to minimize the difference between the amino acid frequency and the joint frequency of amino acid pairs of the designed sequence and the natural sequence.

[0036] A system for designing MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm, comprising:

[0037] A sequence data acquisition module for acquiring a dataset of binding cores of different types of major histocompatibility complex class II molecules MHCII and peptides peptide;

[0038] A sequence feature extraction module for extracting evolutionary information features of each allele of the MHCII molecule and its corresponding binding peptide binding core sequence.

[0039] A neural network prediction module based on two frequency feature tensors extracted by the sequence feature extraction module as double input, and outputs the probability distribution of 20 amino acids at each position of each sequence.

[0040] A prediction evaluation module for evaluating whether the designed binding core meets the design parameters.

[0041] A computer readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the above-mentioned method for designing MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm. ​

[0042] An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the processor implements the above-mentioned method for designing MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm when executing the computer program.

[0043] The beneficial effects of the present application are:

[0044] Reliability: The method of introducing evolutionary information such as sequence position amino acid frequency (first-order conservation analysis) and joint frequency of amino acid pairs (second-order conservation analysis) to design new short peptide sequences solves the problem of short peptides that cannot be designed based on structure, providing the reliability of designing short peptide sequences based on evolutionary information.

[0045] Innovation: The use of a CNN and Transformer hybrid neural network architecture can better understand the long-range dependencies and local features of sequence frequencies, improving recognition ability and prediction accuracy. In particular, by considering the interactions and dependencies between amino acids in the sequence based on the calculation of amino acid frequencies at each position, more stable and efficient peptide sequences can be designed, providing a new perspective for designing peptide sequences.

[0046] Efficiency: By automatically learning sequence evolutionary information features through a deep learning model, compared to traditional Monte Carlo simulated annealing algorithms, the calculation time and cost are reduced, and the neural network method can effectively design peptide sequences without relying on amino acid site frequency constraints.

[0047] Generalization: On the basis of ensuring design reliability, it has higher computational efficiency and generalization ability, and can be applied to the design of various types of MHCII-peptide binding cores. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 The flowchart of the design of MHCII-peptide binding cores of the present application.

[0049] Figure 2 The first-order conservation analysis of the binding core for the HLA-DRB5*01:01 allele of the present application. (A) Natural sequence. (B) FS sequence. (C) L1 sequence. (D) L2 sequence. The overall height of the letters represents the conservation of the amino acid at that position, and the height of each amino acid represents the relative frequency at that position.

[0050] Figure 3Figure 6. Secondary conservation analysis of the natural binding core of HLA-DRB5*01:01 allele and the binding cores designed with different strategies. (A) represents the secondary conservation matrix of the natural binding core. (B) represents the secondary conservation matrix of the FS sequence. (C) shows the PCC value of the secondary conservation values between the FS sequence and the natural sequence is 0.82. (D) represents the secondary conservation matrix of the L1 sequence. (E) shows the PCC value of the secondary conservation values between the L1 sequence and the natural sequence is 0.98. (F) represents the secondary conservation matrix of the L2 sequence. (G) shows the PCC value of the secondary conservation values between the L2 sequence and the natural sequence is 0.99.

[0051] Figure 4 Figure 7. The distribution of binding affinities of the HLA-DRB5*01:01 allele and the designed peptides, natural peptides, random peptides.

[0052] Figure 5 Figure 8. The distribution of pLDDT scores of the five types of binding cores of the present application, each sampling 200 sequences. (A) For MHCII-peptide, the average pLDDT of the binding cores designed with different strategies is >90, much higher than the pLDDT score of the randomly generated peptide sequences. (B) For MHCII-Peptide-SPE-C, the average pLDDT of the binding cores designed with different strategies is >80, much higher than the pLDDT score of the randomly generated peptide sequences. DETAILED DESCRIPTION

[0053] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0054] The present application provides a design method of MHCII binding peptides based on evolutionary information and Transformer neural network algorithm, which comprises the following steps:

[0055] S1, obtaining a data set of binding cores of different types of major histocompatibility complex class II molecules MHCII and peptides, extracting evolutionary information features of alleles of each MHCII molecule and binding peptide sequences corresponding to the binding peptide, which comprises:

[0056] S11, calculating the frequency of amino acids at each site in the sequence fi a , generate a two-dimensional position amino acid frequency tensor F, F ∈ R 9×20 wherein f i a represents the amino acid type at the i-th position of the binding core a corresponding to the occurrence frequency of the 20 standard amino acids, i = 1, …, 9, a = 1, …, 20;

[0057] S12, calculate the joint frequency of amino acid pairs in the sequence f ij ab , generate a four-dimensional amino acid pair joint frequency tensor T, T ∈ R 9 ×20×9×20 wherein f ij ab represents the co-occurrence frequency of the amino acid type a at the i-th position of the binding core and the amino acid type b at the j-th position a

[0058] Each MHCII-peptide in the data set in S1 contains a matrix of M aligned binding core sequences X = {x_µi | i = 1,..., 9; µ = 1,..., M}, wherein x_µi represents the amino acid type at the i-th position of the µ-th sequence;

[0059] In S11, the frequency values of all positions of the sequence matrix and the amino acid types are arranged in position order to form a two-dimensional position amino acid frequency tensor F of size 9 × 20, reflecting the conservative distribution of amino acids at each site in the MSA.

[0060] In S12, the joint frequency values of all position pairs of the sequence matrix and the amino acid type pairs are arranged in index order to form a four-dimensional amino acid pair joint frequency tensor T of size 9 × 20 × 9 × 20, reflecting the coevolution relationship of amino acid pairs in the MSA.

[0061] S2, establish a neural network model based on the fusion of convolution modules and Transformer modules, and take the two frequency feature tensors extracted in S11 and S12 as double inputs;

[0062] S21, respectively, perform convolution operations on the two-dimensional tensor F and the four-dimensional tensor T, extract convolution features, and connect the obtained convolution features along the feature dimension for feature fusion, wherein in S21,

[0063] The two-dimensional position amino acid frequency tensor F ∈ R 9×20 ​After 1D CNN processing through a convolution kernel with a size of 1x1, a high-dimensional frequency feature matrix F' is obtained.

[0064] The four-dimensional amino acid pair joint frequency tensor T e R 9×20×9×20 After 2D CNN processing through a convolution kernel with a size of 16x16, a high-dimensional joint frequency feature matrix T' is obtained.

[0065] The obtained convolution features are connected along the feature dimension by Z = concat(F', T'), and feature fusion is performed.

[0066] S22, the data fused in S21 is input into the Transformer module of the deep learning model to extract data features, and the global dependency relationship is extracted through the self-attention mechanism.

[0067] The processed amino acid frequency and amino acid pair joint frequency are input into a stacked structure composed of 12 Transformer blocks to deeply mine high-order frequency dependency relationships and capture complex patterns in the sequence.

[0068] The core of each Transformer block is the multi-head self-attention mechanism, and the steps of the multi-head self-attention mechanism are as follows: multiply each input matrix by three learnable matrices to create three vectors: Query matrix, Key matrix, and Value matrix.

[0069] The attention score is calculated by calculating the dot product of the Q matrix and the transpose of the K matrix, and then dividing by the square root of the dimension of the K matrix to ensure the stability of the gradient during training. Then, the softmax (normalized exponential function) normalization processing is performed to make them all positive and the sum is 1, and the attention score is obtained. The calculation formula is as follows:

[0070]

[0071] Then multiply the attention score by the V matrix, so that the position with a high attention score has more attention. After 12 transformer modules, the data features are further extracted.

[0072] S23, after feature extraction by the Transformer module, one-dimensional convolution is performed to map the data to the output dimension;

[0073] S24, a Softmax function is added to ensure that the last dimension of the output result is a probability, and the final result is the probability P of each position of each sequence for the 20 amino acids i a , generating a probability distribution tensor P e R N×9×20where N is the number of generated sequences;

[0074] To better guide the network to get accurate probability, the present application calculates the amino acid frequency and the joint frequency of amino acid pairs based on the output probability, and respectively fuses two kinds of loss to better train the network.

[0075] When only the joint frequency of amino acid pairs is used as the loss of the network, the loss function is as follows:

[0076]

[0077] where respectively represent the predicted joint frequency and the natural joint frequency of amino acid pair in the binding core corresponding to the allele of the MHCII molecule. and respectively represent the predicted frequency and the natural frequency of amino acid a in the binding core corresponding to the allele of the MHCII molecule. n i , j a , b

[0078] When the frequency of amino acid and the joint frequency of amino acid pairs are used as the loss, the loss function is as follows:

[0079]

[0080] where is a coefficient used to adjust the balance of the loss, which is set to 65. and respectively represent the predicted frequency and the natural frequency of amino acid a in the binding core corresponding to the allele of the MHCII molecule. i a and respectively represent the predicted frequency and the natural frequency of amino acid a in the binding core corresponding to the allele of the MHCII molecule. n i

[0081] ​​​​​​​Loss1 is to minimize the difference between the designed sequence and the natural sequence of the joint frequency of amino acid pairs, and Loss2 is to minimize the difference between the designed sequence and the natural sequence of the joint frequency of amino acid pairs. Specifically, the network model trained using the loss function Loss1 is named L1 sequence, and the network model trained using the loss function Loss2 is named L2 sequence. The L1 sequence generated by the network model can accurately reproduce the interaction mode of the natural peptide amino acid pair, while the L2 sequence can maintain both the single-point amino acid conservation mode and the interaction mode of the amino acid pair.

[0082] In the optimization process, we use the Adam optimizer (learning rate 1e-4) to train for 1000 epochs, optimize the network parameters through back propagation, and observe the trend of loss reduction. When it tends to be stable and no longer decreases, the network training is considered complete.

[0083] S3, according to the output results of the neural network model, randomly sample according to the probability to generate MHCII-peptide binding core sequences that meet the target distribution.

[0084] As shown in Figures 2-5 , the present study takes the HLA-DRB5*01:01 allele and the natural binding core of the peptide as an example, because it can form a resolved PDB structure (such as PDB ID: 1HQR) with the binding peptide and the interacting antigen, which is convenient for subsequent experimental comparison. The technical effect details are as follows:

[0085] Effect one: The primary conservation distribution of the binding core of HLA-DRB5*01:01 designed based on different strategies is highly consistent with the natural sequence, among which positions 1, 4, 6 and 9 show high conservation and are determined as key functional sites, which is consistent with the general view that P1, P4, P6 and P9 are the four main anchor points.

[0086] Effect two: The natural sequence reveals highly conserved sites and secondary conservation, and the Pearson correlation coefficient (PCC) analysis shows that the FS sequence maintains primary conservation but lacks significant secondary conservation (PCC=0.82); L1 and L2 sequences retain both primary and secondary conservation (PCC=0.98, 0.99, respectively). The coupling matrix of L1, L2 sequences and the natural sequence is highly similar. It is verified that the neural network can learn the evolutionary characteristics of the natural sequence.

[0087] Effect three: The present application uses a deep learning model DeepMHCII to predict the MHCII-peptide binding affinity. The results show that the binding affinity of the peptide segment designed based on evolutionary information and the MHCII molecule is similar to that of the natural MHCII-peptide complex, and the average affinity is about 0.5. The predicted affinity of the randomly generated peptide segment and MHCII is about 0.15. These findings not only verify the theoretical hypothesis that the analysis of amino acid frequency and amino acid pair joint frequency can be synergistically optimized, but also prove the feasibility of using amino acid pair joint frequency (secondary conservation) to develop MHCII binding peptides.

[0088] Effect four: The structural reliability of MHCII peptides is evaluated using the confidence index pLDDT (Predicted Local Structure Confidence Score) of AlphaFold3 (pLDDT>70 indicates high confidence, and pLDDT<50 indicates low confidence). The pLDDT score quantifies the structural confidence through atomic and pairwise error estimates, where the higher the value, the higher the model reliability. 200 binding cores are randomly selected from each peptide category (natural / FS / L1 / L2 / random) and subjected to structure prediction with HLA-DRB5*01:01. For each binding core, its pLDDT score is calculated as the average of all predicted atomic confidence values in the sequence. The results show that the pLDDT score of the designed peptide is comparable to that of the natural peptide, and the average of the 200 pLDDT values of different categories is greater than 90, while the pLDDT score of the random peptide is significantly lower, with an average of about 50, and the L2 sequence obtains the highest pLDDT score and the smallest fluctuation range. This large-scale validation not only confirms the structural reliability of the designed peptide, but also highlights its potential for more extensive applications.

[0089] Effect five: To evaluate the structural performance of our designed peptides, we replaced the natural peptide in the MHCII-peptide-SPE-C (PDB ID: 1HQR) complex structure with the designed peptide. 200 binding cores that bind to HLA-DRB5*01:01 are randomly selected from each peptide category (natural / FS / L1 / L2 / random), and their ternary complex structures with MHCII and SPE-C superantigen are predicted using AlphaFold3. The average pLDDT score of the peptide in the predicted complex is calculated to evaluate the structural confidence. The results show that the average of the 200 pLDDT values of the designed peptides of different categories is about 85, while the pLDDT score of the random peptide is significantly lower than that of the designed peptide, with an average of about 45.

[0090] The present application specifically includes:

[0091] (1) Single residue evolutionary conservation design sequence Extract the position-specific amino acid frequency matrix from the good binding core by comparison, and randomly sample according to the frequency distribution at the single residue level to generate FS sequences with natural evolutionary conservation characteristics. At the same time, according to the background frequency of 20 kinds of amino acids in nature, a random sequence (random) is generated for comparison and reference of the sequence designed by the neural network.

[0092] (2) Neural network design based on two-residue co-evolution

[0093] The amino acid site frequency and amino acid pair joint frequency information characteristics of the natural binding core are taken as the input of the network, and the CNN-Transformer hybrid architecture is trained, and the probability distribution of 20 kinds of amino acids at each position of each sequence is output.

[0094] (3) Evolutionary loss design of neural network

[0095] The neural network model is trained by using different loss functions, and the model parameters with different performance can be obtained. In the present application, the mean square error (MSE) between the joint frequency of the amino acid pair of the binding core predicted by the model and the natural amino acid pair joint frequency is taken as the loss function, thereby generating L1 sequence. Another is to take the MSE between the unit point amino acid frequency and the two-site amino acid pair joint frequency predicted by the model and the unit point amino acid frequency and the two-site amino acid pair joint frequency in the natural sequence as the loss function, thereby generating L2 sequence.

[0096] (4) Verification and evaluation of the designed peptide sequence

[0097] A core perception deep interaction model DeepMHCII is trained to predict the binding affinity of MHCII-binding core, and an advanced AlphaFold3 tool is used to accurately predict the structural accuracy of the designed binding core in the binding groove of MHCII two chains.

[0098] Embodiment

[0099] In the design of 27 MHCII-peptide binding cores, we selected the design of the natural binding core combined with the HLA-DRB5*01:01 allele as an example, and the design details are as follows:

[0100] Step one: Use the predictor NetMHCIIpan-4.1 to predict the binding core of HLA-DRB5*01:01-peptide.

[0101] Step two: Calculate the frequency of amino acids and the joint frequency of amino acid pairs in the binding core in step one, and take the two frequencies as the input of the network. For example: for three binding core sequences (sequences 1, 2, and 3), if A (alanine) appears twice and E (glutamic acid) appears once in the first position, then the frequency = 2 / 3 ≈ 0.67, = 1 / 3 ≈ 0.33, and the frequency of the remaining amino acids is 0. If the amino acid pair (A, E) appears twice in positions 1 and 2 (sequences 1 and 3), and (S, I) appears once in position 2 (sequence 2), then the joint frequency = 2 / 3 ≈ 0.67, = 1 / 3 ≈ 0.33, and the frequency of the remaining combinations is 0.

[0102] Step three: Perform CNN operation on the two frequencies respectively, and then perform feature fusion on the high-dimensional frequency matrices obtained after convolution in the feature dimension.

[0103] Step four: The high-dimensional matrix after feature fusion is input into a 12-layer Transformer module for training.

[0104] Step five: After feature extraction by the Transformer module, perform one-dimensional convolution, and finally use the softmax activation function to ensure that the output is a probability and to maintain gradient continuity.

[0105] Step six: The output is the probability of each of the 20 amino acids at each position of each sequence, and the amino acid is selected according to the probability to generate the sequence. For example, for a binding core of length 9, the output of the first position is: A (alanine): 0.10, C (cysteine): 0.05, D (aspartic acid): 0.40, E (glutamic acid): 0.10,... (the remaining amino acids). This distribution indicates that the first position has a probability of 0.40 for D, a probability of 0.10 for A or E, and a lower probability for the remaining amino acids. According to this distribution, a sampling operation is performed for each position to obtain the complete amino acid sequence.

[0106] As Figure 1 shown, the application also discloses a design system for MHCII binding peptides based on evolutionary information and a Transformer neural network algorithm, which comprises a sequence data acquisition module, a sequence feature extraction module, a neural network prediction module, and a prediction evaluation module for designed sequences. The specific implementation of each module is as follows:

[0107] Sequence data acquisition module in the embodiment, sequence data acquisition module is mainly responsible for obtaining training data from public papers. The implementation process of the module specifically includes the following aspects:

[0108] (1) Data source selection

[0109] Natural MHCII molecule binding peptide data sources include two independent data sets: (1) BD2020 (literature data before 2020): extracted from IEDB by Venkatesh et al, used to train MHCAttnNet. (2) Dset pos (Newly added data after 2020): downloaded from IEDB by Yang, Yaqing et al, used to predict the performance of different predictor models. All the data finally come from the public database IEDB.

[0110] (2) Data preprocessing steps

[0111] From BD2020 and Dset pos Each MHCII subtype in the data set is filtered to have no less than 200 corresponding polypeptide sequence samples, and after integrating the two data sets, a total of 46145 polypeptide sequences that can bind to 27 HLA alloantigens (9 DR, 1 DQ and 17 DP) will be obtained. The collected all MHCII-peptides complexes are predicted by NetMHCIIpan-4.1 predictor to obtain the binding core data as the natural MSA sequence data set of this experiment, named Dset_bc.

[0112] Sequence feature extraction module The binding core of each type of MHCII-peptide in Dset_bc data set is constructed into a M-aligned binding core sequence matrix a = {aµi |i = 1,…, 9,µ= 1,…,M}, and the amino acid frequency f i a and amino acid pair joint frequency f ij ab are calculated respectively. The amino acid frequency of each position i is obtained by formula f i a =N ( i , a ) / M , where N ( i , a ) is the amino acid i a ​The frequency of occurrence constitutes a two-dimensional positional amino acid frequency tensor F (F∈R) of size 9×20. 9×20 ); through formula f ij ab =N ( i, j , a, b ) / M calculates the amino acid pair combination frequency at each position (i, j), where N( i, j , a, b The co-occurrence counts of amino acid pairs at positions i = a and j = b constitute a four-dimensional amino acid pair joint frequency tensor T (T∈R) of size 9×20×9×20. 9×20×9×20 The two extracted frequency tensors are used as training datasets and fed into the neural network for training.

[0113] The neural network prediction module (1) extracts tensors F∈R from MSA. 9×20 First flatten it into R 1×180 Then, a one-dimensional convolution is performed using a 1×1 kernel to obtain the high-dimensional feature matrix F′. The tensor T∈R is then... 9×20×9×20 First refactor into R 1 ×180×180 Then, a two-dimensional convolution is performed using a 16×16 kernel to obtain a high-dimensional feature matrix T′. The two feature matrices are concatenated along the feature dimension, and the calculation formula is: Z=concat(F′, T′).

[0114] (2) The fused feature Z is input into the Transformer module for further processing. A linear transformation maps Z to a higher dimension, forming multiple independent subspaces (i.e., multi-head). The feature Z is mapped to a higher dimension, possessing more subspaces. The multi-head self-attention mechanism performs self-attention calculations in different subspaces, focusing on different feature patterns, and finally merges them. Subsequently, residual connections and normalization are used to stabilize training, alleviating gradient vanishing and overfitting. The feedforward neural network consists of two consecutive fully connected layers, and also undergoes residual connections and normalization.

[0115] (3) The specific steps of the multi-head self-attention mechanism are as follows: Each input matrix is ​​multiplied by three learnable matrices to create three vectors: a Query (Q) matrix, a Key (K) matrix, and a Value (V) matrix. The attention score is calculated by taking the dot product of the Q and K matrices, and then divided by the square root of the K matrix dimension to ensure gradient stability during training. SoftMax (normalized exponential function) normalization is then performed to obtain the attention scores, ensuring they are all positive and sum to 1. The calculation formula is as follows:

[0116]

[0117] Then multiply the attention score with the V matrix, so that the position with high attention score has more attention.

[0118] (4) After 12 transformer modules, then a 1x1 convolution kernel is used to reduce the dimension of the data.

[0119] (5) The output after convolution is processed by the Softmax function to obtain the probability distribution tensor P ∈ R N×9×20 of each amino acid at each position of each sequence, and the sequence is generated by randomly sampling amino acids according to the probability distribution, where N is the number of output sequences.

[0120] The prediction evaluation module of the designed sequence is relatively short, and the affinity with MHCII is usually lower than that of the full-length peptide. In order to ensure that there is a significant difference in affinity between the randomly generated peptide and the natural peptide, the high-affinity binding core (predicted by NetMHCIIpan-4.1, parameter IC50<50 nM) is strictly screened from Dset pos and BD2020 data sets. Then, these natural high-affinity binding cores are designed according to the above steps of designing sequences to design new binding cores. The DeepMHCII model trained by the five-fold cross-validation method is used to calculate the binding affinity of the MHCII-binding core, and AlphaFold3 is used to predict the structural confidence of the binding core in the MHCII-binding core complex, so as to evaluate whether the designed binding core has similar properties to the natural binding core (reference Figure 2 , Figure 3 Experimental results).

[0121] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the above-mentioned method for designing MHCII-binding peptides based on evolutionary information and a Transformer neural network algorithm.

[0122] The readable storage medium can include a readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory, and can further include a read only memory (ROM). Program code is stored thereon, and the program code is executed by the processor to cause the processor to perform the steps of the embodiments as described in the specification.

[0123] The memory may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0124] A bus can represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus structures.

[0125] An electronic device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described design method for MHC11 binding peptides based on evolutionary information and the Transformer neural network algorithm.

[0126] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0127] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0128] The embodiments should not be regarded as limitations on the present invention, but any improvements made based on the spirit of the present invention should be within the protection scope of the present invention.

Claims

1. A method for designing MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm, characterized in that: It includes the following steps: S1. Obtain datasets of binding cores for different types of major histocompatibility complex II molecules (MHCII) and peptides. Extract evolutionary information features from the alleles of each MHCII molecule and their corresponding binding core sequences, including: S11. Calculate the frequency of amino acids per unit point in the sequence. f i a Generate a two-dimensional positional amino acid frequency tensor F, F∈R 9×20 ,in f i a This indicates the amino acid type at the i-th position of the binding core. a The frequencies of occurrence of the 20 standard amino acids, i=1,…,9, a =1,…,20; S12. Calculate the amino acid pair combination frequency in the sequence. f ij ab Generate a four-dimensional amino acid pair joint frequency tensor T, T∈R 9 ×20×9×20 ,in f ij ab This indicates that the amino acid type is at the i-th position of the binding core. a The co-occurrence frequency with amino acid type b at position j; S2. Establish a neural network model based on the fusion of convolutional and Transformer modules, using the two frequency feature tensors extracted in S11 and S12 as dual inputs. S21. Perform convolution operations on the two-dimensional tensor F and the four-dimensional tensor T respectively, extract convolution features, and connect the obtained convolution features along the feature dimensions to perform feature fusion. S22. Input the data fused from S21 into the Transformer module of the deep learning model to extract data features, and extract global dependencies through the self-attention mechanism; S23. After feature extraction via the Transformer module, a one-dimensional convolution is performed to map the data to the output dimension. S24. Add the Softmax function to ensure that the last dimension of the output is probability, ultimately yielding the probability P of each of the 20 amino acids at each position in each sequence. i a This generates the probability distribution tensor P∈R N×9×20 , where N is the number of generated sequences; S3. Based on the output of the neural network model, random sampling is performed according to probability to generate a binding core sequence of MHCII-peptide that satisfies the target distribution.

2. The design method for MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm according to claim 1, characterized in that: Each MHCII-peptide in the dataset S1 contains M aligned binding core sequence matrices. X= { x µi |i = 1,…, 9, µ= 1,…,M },in x_µi This indicates the amino acid type at the i-th position of the µ-th sequence; In S11, the frequency values ​​of all positions and amino acid types in the sequence matrix are arranged in positional order to form a two-dimensional position amino acid frequency tensor F with a size of 9×20, which reflects the conserved distribution of amino acids at each site in multiple sequence alignment (MSA). In S12, the joint frequency values ​​of all position pairs and amino acid type pairs in the sequence matrix are arranged in index order to form a four-dimensional amino acid pair joint frequency tensor T with a size of 9×20×9×20, which reflects the co-evolutionary relationship of amino acid pairs in multiple sequence alignment (MSA).

3. The design method for MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm according to claim 1, characterized in that: In S21, Two-dimensional positional amino acid frequency tensor F∈R 9×20 After performing 1D CNN processing with a 1×1 kernel, a high-dimensional frequency feature matrix F' is obtained; Four-dimensional amino acid pairs, joint frequency tensor T∈R 9×20×9×20 After performing 2D CNN processing with a convolution kernel of size 16×16, a high-dimensional joint frequency feature matrix T' is obtained; The convolutional features obtained are concatenated along the feature dimension using Z=concat(F', T') to perform feature fusion.

4. The method for designing MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm according to claim 1, characterized in that: In S22, the fused data of the processed amino acid frequency tensor F and the amino acid pair joint frequency tensor T are input into a stacked structure consisting of 12 Transformer Blocks.

5. The method for designing MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm according to claim 4, characterized in that: At the core of each Transformer Block is the multi-head self-attention mechanism. The specific steps of the multi-head self-attention mechanism are as follows: multiply each input matrix by three learnable matrices to create three vectors: Query matrix, Key matrix, and Value matrix.

6. The design method for MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm according to claim 1, characterized in that: S24 When only the amino acid pair joint frequency is used as the network loss, the loss function is as follows: in Let B represent the predicted and natural binding frequencies of amino acid type a appearing at position i and amino acid type b appearing at position j in the binding cores, respectively, where B is the batch size. and They represent the first n In the binding cores corresponding to the alleles of MHCII, the first i , j Amino acid pairs at each position a , b The predicted joint frequency and the natural joint frequency; When the frequency of an amino acid and the joint frequency of an amino acid pair are used as the loss, the loss function is as follows: in This is a coefficient used to adjust the balance of losses; it's set to 65. and These represent the positions in the binding cores. i amino acid type a The predicted frequency and the natural frequency, and These represent the amino acid types. a In the n The binding cores corresponding to the alleles of MHCII i The predicted frequency and natural frequency of the bit; Loss1 minimizes the difference in the combination frequency of amino acid pairs between the designed sequence and the natural sequence, while Loss2 minimizes the difference in both the amino acid frequency and the combination frequency of amino acid pairs between the designed sequence and the natural sequence.

7. A design system based on the design method of MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm as described in any one of claims 1 to 6, characterized in that: It includes: The sequence data acquisition module is used to acquire datasets of binding cores of different types of major histocompatibility complex II molecules (MHCII) and peptides. The sequence feature extraction module extracts evolutionary information features from the binding core sequences of the alleles and their corresponding binding peptides of each MHCII molecule. The neural network prediction module uses two frequency feature tensors extracted by the sequence feature extraction module as dual inputs and outputs the probability distribution of 20 amino acids at each position of each sequence. The design sequence prediction and evaluation module is used to evaluate whether the binding core of the design meets the design parameters.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the design method of MHC11 binding peptide based on evolutionary information and Transformer neural network algorithm as described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the processor executes the computer program, it implements the design method for MHC11 binding peptides based on evolutionary information and Transformer neural network algorithm as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Transform-based cross-modal fusion multi-modal emotion recognition method

    CN120508972A

  • Compilation method and apparatus for neural network model, inference method and apparatus for neural network model, and device and medium

    WO2024239971A1