Prediction method and prediction system for combination of CD4 + T cell receptor and polypeptide based on deep learning

Through deep learning combined with multimodal feature fusion, the accuracy of CD4+ T cell receptors and peptide binding force prediction is solved, the accuracy and reliability of the prediction model are improved, and more comprehensive support is provided for immunotherapy and vaccine development.

CN120412705APending Publication Date: 2025-08-01BEIJING YUEKANGKECHUANG PHARM TECH CO LTD

Patent Information

Application Number
CN202510896911.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the binding power of CD4+ T cell receptors and polypeptides under limited data, and ignores the important role of CD4+ T cells in immunotherapy.

Method used

Through a deep learning-based method, combining the characteristics of CD4+TCR sequences and polypeptide sequences, a multimodal feature fusion strategy is adopted, including sequence features, descriptor features and spatial features, and a Transformer model and convolutional neural network are used to extract and fusion features to build a multi-dimensional prediction model.

Benefits of technology

It improves the accuracy and reliability of CD4+ T cell receptor binding prediction and provides a more comprehensive prediction tool to support the research of immune recognition mechanisms and vaccine development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412705A_ABST
    Figure CN120412705A_ABST
Patent Text Reader

Abstract

The invention provides a CD4 + T cell receptor and polypeptide combination prediction method and prediction system based on deep learning. The prediction method comprises the following steps: respectively obtaining TCR sequence features and polypeptide sequence features based on a CD4 + TCR sequence and a polypeptide sequence, and fusing the TCR sequence features and the polypeptide sequence features to obtain fused sequence features; obtaining TCR descriptor features and polypeptide descriptor features based on the CD4 + TCR sequence and the polypeptide sequence, and fusing the TCR descriptor features and the polypeptide descriptor features to obtain fused descriptor features; obtaining TCR spatial characteristics and polypeptide spatial characteristics based on the CD4 + TCR sequence and the polypeptide sequence, and fusing the TCR spatial characteristics and the polypeptide spatial characteristics to obtain fused spatial characteristics; and predicting the binding force based on the fusion sequence features, the fusion descriptor features and the fusion spatial features. Therefore, the accuracy of combination prediction of the CD4 + T cell receptor and the polypeptide can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of neoantigen immunotherapy, and particularly to a prediction method and a prediction system for the binding of CD4 + T cell receptor to polypeptide based on deep learning. Background Art

[0002] In neoantigen immunotherapy, the interaction between T cell receptor and polypeptide presented by major histocompatibility complex MHC is the basis of immune response. However, current prediction of the binding of T cell receptor to polypeptide mainly focuses on CD8 + T cells, while the immune function of CD4 + T cells is often overlooked. In actual immunotherapy, CD4 + T cells and CD8 + T cells usually act synergistically.

[0003] Deep learning networks have been used for predicting the binding specificity of neoantigens to CD8 + T cell receptors. However, since the immune function of CD4 + T cells is often overlooked, and currently, the sample data set of the binding of CD4 + T cell receptor to polypeptide is small, it is difficult to accurately predict the binding affinity of CD4 + T cell receptor to polypeptide based on limited specific data. How to accurately predict the binding specificity of CD4 + T cell receptor to polypeptide using deep learning under limited data has become a difficult point and an urgent problem to be solved. Summary of the Invention

[0004] In view of the above technical problems existing in the prior art, this application is proposed. This application aims to provide a prediction method and a prediction system for the binding of CD4 + T cell receptor to polypeptide based on deep learning, which can improve the accuracy and reliability of predicting the binding affinity of CD4 + T cell receptor to polypeptide when the specific sample data of the binding of CD4 + T cell receptor to polypeptide is small.

[0005] According to the first aspect of this application, a prediction method for the binding of CD4 + T cell receptor to polypeptide based on deep learning is provided. The prediction method includes obtaining TCR sequence features and polypeptide sequence features respectively based on CD4 + TCR sequence and polypeptide sequence, and fusing the TCR sequence features and the polypeptide sequence features to obtain fused sequence features; based on the CD4 +The TCR sequence and the polypeptide sequence are respectively used to obtain TCR descriptor features and polypeptide descriptor features, and the TCR descriptor features and the polypeptide descriptor features are fused to obtain fused descriptor features; based on the CD4 + The TCR sequence and the polypeptide sequence are respectively used to obtain TCR spatial features and polypeptide spatial features, and the TCR spatial features and the polypeptide spatial features are fused to obtain fused spatial features; based on the fused sequence features, the fused descriptor features, and the fused spatial features, the CD4 + binding force between the TCR and the polypeptide is predicted.

[0006] According to the second aspect of the present application, there is provided a prediction system for the binding of a CD4 + T cell receptor to a polypeptide based on deep learning. The prediction system includes: a sequence processing unit configured to: respectively obtain TCR sequence features and polypeptide sequence features based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR sequence features and the polypeptide sequence features to obtain fused sequence features; a descriptor processing unit configured to: respectively obtain TCR descriptor features and polypeptide descriptor features based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR descriptor features and the polypeptide descriptor features to obtain fused descriptor features; a spatial processing unit configured to: respectively obtain TCR spatial features and polypeptide spatial features based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR spatial features and the polypeptide spatial features to obtain fused spatial features; a fusion unit configured to predict the CD4 + binding force between the TCR and the polypeptide based on the fused sequence features, the fused descriptor features, and the fused spatial features.

[0007] According to the third aspect of the present application, there is provided a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the prediction method for the binding of a CD4 + T cell receptor to a polypeptide based on deep learning according to various embodiments of the present application are implemented.

[0008] According to the fourth aspect of the present application, there is provided an electronic device. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, the steps in the prediction method for the binding of a CD4 + T cell receptor to a polypeptide based on deep learning according to various embodiments of the present application are implemented.

[0009] Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows: Prediction method for the binding of CD4 + T cell receptor to polypeptide based on deep learning. Based on CD4 + TCR sequences and polypeptide sequences, TCR sequence features and polypeptide sequence features are obtained respectively, and fused to obtain fused sequence features; Based on CD4 + TCR sequences and polypeptide sequences, TCR descriptor features and polypeptide descriptor features are obtained respectively, and fused to obtain fused descriptor features; Based on CD4 + TCR sequences and polypeptide sequences, TCR spatial features and polypeptide spatial features are obtained respectively, and fused to obtain fused spatial features. The sequence features, descriptor features and spatial features are concatenated and fused into multi-dimensional features for predicting the binding of CD4 + T cell receptor to polypeptide, improving the prediction accuracy and reliability. This prediction method not only considers the CD4 + sequence features of TCR sequences and polypeptide sequences, but also considers the descriptor features of the sequences and the spatial features of the sequences. Through a multi-modal deep learning model, the correlation features, physicochemical properties and sequence structure features between sequences are combined to capture more comprehensive prediction features, further improving the accuracy of predicting the binding of polypeptides to TCR sequences.

[0010] In addition, through the cross-analysis of multi-dimensional features, this prediction method systematically models the interaction between CD4 + TCR and polypeptide from three levels of "sequence - physicochemical properties - structure", providing a more accurate and reliable quantitative tool for the study of immune recognition mechanisms.

[0011] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above description and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Brief Description of the Drawings

[0012] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. Similar reference numerals with letter suffixes or different letter suffixes may represent different examples of similar components. The drawings generally illustrate various embodiments by way of example and not limitation, and are used in conjunction with the specification and the claims to explain the disclosed embodiments. Such embodiments are illustrative and exemplary and are not intended to be an exhaustive or exclusive embodiment of the method, apparatus, system or non-transitory computer-readable medium having instructions for implementing the method.

[0013] Figure 1 Shows the CD4 based on deep learning according to an embodiment of the present application +Flowchart of a method for predicting the binding of a T cell receptor to a polypeptide.

[0014] Figure 2 Shows CD4 according to an embodiment of the present application + Schematic diagram of a network model for predicting the binding of a T cell receptor to a polypeptide.

[0015] Figure 3 Shows a schematic diagram of the Transformer model according to an embodiment of the present application.

[0016] Figure 4 Shows a schematic diagram of a convolutional neural network model according to an embodiment of the present application.

[0017] Figure 5 Shows a schematic diagram of an amino acid token dictionary according to an embodiment of the present application.

[0018] Figure 6 Shows a schematic diagram of a descriptor vector of a polypeptide sequence according to an embodiment of the present application.

[0019] Figure 7 Shows a schematic diagram of the three-dimensional structure information of a sequence according to an embodiment of the present application.

[0020] Figure 8 Shows a schematic diagram of an amino acid distance matrix according to an embodiment of the present application.

[0021] Figure 9 Shows CD4 based on deep learning according to an embodiment of the present application + Schematic diagram of the structure of a prediction system for the binding of a T cell receptor to a polypeptide.

[0022] Figure 10 Shows a schematic diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0023] To enable those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific implementation manners. The embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific examples, but this is not a limitation to the present application.

[0024] The "first", "second" and similar terms used in this application do not denote any order, quantity or importance, but are only used for distinction and convenience in expression, without strictly limiting that the "first" and "second" must be different. For example, the "first preset length" and the "second preset length" can be the same or different. Among them, the "first", "second" and similar terms can be replaced with each other. The terms such as "including" or "comprising" used in this application mean that the elements before this word are covered by the elements listed after this word, and it does not exclude the possibility of also covering other elements. In this application, the arrows shown in the figures for each step are only examples of the execution order and not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. Each step in the execution order can be executed jointly, can be decomposed, and can be reordered as long as the logical relationship of the execution content is not affected.

[0025] All terms used in this application (including technical terms or scientific terms) have the same meaning as understood by those of ordinary skill in the art to which this application belongs, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such here. Technologies and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies and devices should be regarded as part of the specification.

[0026] An embodiment of this application provides a prediction method for the binding of CD4 + T cell receptor and polypeptide, specifically, as shown in Figure 1 Steps S101 - S104, where the arrows shown in the figures for each step are only examples of the execution order and not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. Each step in the execution order can be executed jointly, can be decomposed, and can be reordered as long as the logical relationship of the execution content is not affected.

[0027] In step S101, based on the CD4 + TCR sequence and polypeptide sequence, TCR sequence features and polypeptide sequence features are respectively obtained, and the TCR sequence features and polypeptide sequence features are fused to obtain fused sequence features.

[0028] CD4 + T cell receptor (TCR) is an important molecular structure on the surface of CD4 + T lymphocytes, and is a key component for such cells to recognize antigens, playing a core role in the human immune system in recognizing foreign antigens and initiating immune responses.

[0029] Specifically, CD4 can be collected based on VDJdb (2019), IEDB-AR public database + T and peptide sample datasets. Sample data collected from public databases have not been filtered, optimized, or subjected to other sequence modifications. VDJdb (2019) and IEDB-AR are both public databases that play an important role in the field of immunity.

[0030] It should be noted that in each embodiment of this application, for CD4 + The prediction of T cell receptor-peptide binding affinity is independent of the specific sequence.

[0031] The sample data set is divided into a training set and a validation set. The peptides in the validation set do not appear in the training set. The prediction model with the best training effect is selected based on the training set training model, and the selected prediction model is verified based on the validation set.

[0032] In some embodiments, the amino acid word dictionary is used to identify CD4 + The amino acid sequences of the TCR sequence and the polypeptide sequence are digitized to obtain their respective digitized sequences.

[0033] Specifically, amino acid word-merge dictionaries are mapping tools used to convert amino acid sequences into computer-processable numerical or symbolic representations. For example, Python libraries include a built-in mapping table for amino acid single-letter / three-letter abbreviations, as well as command-line tools that support sequence format conversion and word-merge encoding. Furthermore, amino acid word-merge dictionaries can be obtained based on physicochemical property databases, such as the AAindex database, which contains physicochemical property parameters for over 500 amino acids and can be used to construct multidimensional feature vectors.

[0034] Figure 5 A custom amino acid word dictionary is shown. Each character identifier can be used to represent a single-letter amino acid abbreviation or special symbol. For example, the character identifier "X" may represent an unknown or random amino acid. The numbers represent the numeric IDs to which the character identifiers are mapped, which are used to convert the character identifiers into computer-processable values. For example, "R" is mapped to the number "1" and "N" is mapped to the number "2."

[0035] in, Figure 5 As just one implementation method, the amino acid word-unit dictionary may be configured, and no specific limitation is given to this.

[0036] Specifically, CD4 + The original amino acid sequence of TCR is digitally encoded, the original amino acid sequence of the peptide is digitally encoded, and the sequence is digitally mapped according to the amino acid word dictionary to obtain CD4+ The digital sequences of the TCR and the digital sequences of the polypeptides.

[0037] In some embodiments, a preset CD4 + The digital length of the amino acid sequence of the TCR sequence is a first preset length. If the CD4 + The actual digital length of the amino acid sequence of the TCR sequence is less than the first preset length, then it is padded with a preset number; the actual digital length of the amino acid sequence of the preset polypeptide sequence is a second preset length. If the digital length of the amino acid sequence of the polypeptide sequence is less than the second preset length, then it is padded with a preset number.

[0038] Among them, no specific limitation is made on the first preset length and the second preset length. For example, the first preset length may be equal to the second preset length.

[0039] As a preferred embodiment, the first preset length is greater than the second preset length. For example, the first preset length is 30 and the second preset length is 22.

[0040] Exemplarily, CD4 can be set + The digital length of the amino acid sequence of the TCR sequence is 30, and the digital length of the amino acid sequence of the polypeptide sequence is set to 22. During the digital processing based on the amino acid token dictionary, if the CD4 + The actual digital length after digital processing of the amino acid sequence of the TCR is less than 30, then it can be padded with the number "22". If the actual digital length after digital processing of the amino acid sequence of the polypeptide is less than 22, then it can be padded with the number "22".

[0041] In this way, it helps to align and compare the features of different sequences in the same position dimension. For example, for the polypeptide sequence, a fixed length of 22 is set, and for the CD4 + The TCR sequence is set to a fixed length of 30, which is beneficial for the subsequent extracted sequence features (such as the information represented by the amino acid digital encodings at different positions) to be in a unified position framework, facilitating the model to learn position-related patterns across sequences, such as the influence of certain amino acid combinations at specific positions on the binding force and other rules, and improving the model's ability to capture features and make predictions.

[0042] Among them, the numbers used for padding should ensure that they do not repeat with the digital encodings of the existing amino acids, so that the prediction model can distinguish which are the amino acid encodings of the original sequence and which are the padding values, preventing the padded numbers from interfering with the learning and recognition of the true sequence features.

[0043] In some embodiments, based on CD4 +The digitized sequence obtained from the TCR sequence uses a Transformer model to obtain the TCR sequence features; the digitized sequence obtained from the polypeptide sequence uses a Transformer model to obtain the polypeptide sequence features; based on the TCR sequence features and the polypeptide sequence features, a cross-attention model is used to obtain the fused sequence features.

[0044] Specifically, as Figure 2 shown, the polypeptide sequence is input into the Transformer encoder to obtain polypeptide sequence features, the TCR sequence is input into the Transformer encoder to obtain TCR sequence features, and then the obtained TCR sequence features and polypeptide sequence features are both input into the cross-attention model to obtain the fused sequence features.

[0045] Exemplarily, the Transformer model is as Figure 3 shown. Each layer of the encoder consists of two parts. The first sub-layer includes multi-head attention, a fully connected layer, a residual layer, and a LayerNorm layer; the second sub-layer includes two layers of fully connected layers, a residual layer, and a LayerNorm layer.

[0046] In this embodiment, the Transformer model is beneficial for capturing the interactions between amino acid residues that are far apart, avoiding ignoring the associations between amino acid residues that are far apart due to the overly long sequence, thereby improving the CD4 + accuracy of the correlation between TCR and polypeptide.

[0047] Exemplarily, the cross-attention network consists of two parts. The first sub-layer includes cross multi-head attention, a fully connected layer, a residual layer, and a LayerNorm layer; the second sub-layer includes two layers of fully connected layers, a residual layer, and a LayerNorm layer.

[0048] In this embodiment, cross-attention allows the TCR features to focus on the key parts of the polypeptide features. For example, moreover, the cross-attention model can simulate the CD4 + synergistic complementarity of the binding between TCR and polypeptide, capturing the features of multi-region synergistic effects.

[0049] In this way, by using the Transformer model to capture long-range dependencies, the accuracy of the correlation between CD4 + TCR and polypeptide is improved, and the cross-attention model is used to achieve adaptive focusing on the key interaction regions, which is beneficial to improving the prediction accuracy.

[0050] Returning to the embodiment of the present application, in step S102, based on the CD4 +The TCR sequence and the polypeptide sequence are respectively used to obtain TCR descriptor features and polypeptide descriptor features, and the TCR descriptor features and the polypeptide descriptor features are fused to obtain fused descriptor features.

[0051] The descriptor features refer to numerical features that can quantitatively describe the physicochemical properties or structural characteristics of biomolecules (such as proteins, polypeptides) obtained through calculation or prediction. These features do not directly depend on the amino acid arrangement order of the sequence itself, but reflect the overall or local properties of the molecule.

[0052] Exemplarily, for CD4 + The TCR descriptor features obtained from the TCR sequence may include physicochemical properties (such as hydrophobicity, molecular weight), structural properties (such as solvent accessibility, secondary structure content), and functional properties (such as conservation score).

[0053] The polypeptide descriptor features obtained from the polypeptide sequence may include MHC binding characteristics (such as MHC affinity prediction value), immunogenicity index (protease cleavage site prediction), and physicochemical properties (such as hydrophilicity, charge distribution).

[0054] Specifically, descriptors can be calculated through amino acid composition and sequence information. For example, the average hydrophobicity of the sequence can be calculated based on the Kyte-Doolittle scale, and the net charge and charge distribution of the sequence at physiological pH can be calculated. For example, machine learning or deep learning models can be used to predict descriptors related to structure or function. For example, for secondary structure prediction, tools such as PSIPRED and JPred can be used to predict the proportions of alpha helix, beta sheet, and random coil. Pre-computed descriptors can also be extracted from bioinformatics databases, such as the AAindex database.

[0055] This is only an example and does not constitute a limiting description of the specific solution.

[0056] In this embodiment, the descriptor features can capture the intermolecular interaction mechanism that cannot be directly reflected by sequence features, and moreover, the descriptor features (such as physicochemical properties) have higher conservation, which helps the generalization ability of the prediction model in different scenarios. By extracting and fusing the descriptor features of CD4 + TCR and polypeptide, the physicochemical properties and structural characteristics of biomolecules can be integrated into the prediction model, so as to more comprehensively capture CD4 + TCR-polypeptide interaction mechanism, which helps to improve the accuracy of CD4 + TCR and polypeptide binding force prediction.

[0057] As a preferred embodiment, the descriptors are amino acid composition (AAC), dipeptide composition (DPC), composition of k-spaced amino acid groups pairs (CKSAAGP), pseudo amino acid composition (PAAC), and physicochemical properties (PHYC).

[0058] Among them, the amino acid composition (AAC) refers to the statistical occurrence frequency (percentage) of 20 standard amino acids in the sequence, ignoring the amino acid sequence information, which is conducive to reflecting the overall amino acid preference.

[0059] The dipeptide composition (DPC) refers to the statistical occurrence frequency of all possible dipeptides in the sequence, which is conducive to capturing local structural motifs (such as the anchor residue pairs of MHC-binding peptides).

[0060] The composition of k-spaced amino acid groups pairs (CKSAAGP) refers to calculating the occurrence frequency of groups pairs spaced k amino acids apart, and the group types are usually classified based on the properties of amino acid side chains (such as hydrophobicity, charge), which is conducive to revealing long-range interaction patterns.

[0061] The pseudo amino acid composition (PAAC) refers to a mixed feature combining amino acid composition and sequence order information. The first part includes the composition frequencies of 20 amino acids (same as AAC), and the second part includes amino acid correlations (such as hydrophobic correlation, charge correlation) based on different intervals (such as 1-20 residues), which is conducive to balancing global and local information and is applicable to distinguishing proteins with similar functions but large sequence differences.

[0062] The physicochemical properties (PHYC) refer to calculating various physicochemical attributes of the sequence, which is conducive to directly correlating the physical behavior of the molecule (such as hydrophobicity affecting membrane-binding ability, charge affecting receptor-binding affinity).

[0063] By comprehensively considering these five types of descriptors, namely AAC, DPC, CKSAAGP, PAAC, and PHYC, the descriptor features obtained based on these five types of descriptors can comprehensively characterize the CD4 + TCR and the sequence features of polypeptides from at least five dimensions such as composition frequency, local order, long-range interaction, mixed characteristics, and physicochemical properties. This multi-dimensional feature fusion strategy significantly improves the prediction model's ability to predict the CD4 + TCR-polypeptide binding force.

[0064] In some embodiments, based on the CD4 + TCR sequence and polypeptide sequence, TCR sequence descriptor information and polypeptide descriptor information are respectively obtained; based on the TCR sequence descriptor information and polypeptide descriptor information, the TCR descriptor features and polypeptide descriptor features are respectively obtained using the Transformer model.

[0065] As a preferred embodiment, both the TCR sequence descriptor information and the polypeptide descriptor information can be amino acid composition (AAC), dipeptide composition (DPC), composition of k-spaced amino acid groups pairs (CKSAAGP), pseudo amino acid composition (PAAC), and physicochemical properties (PHYC).

[0066] Specifically, as Figure 2 , the TCR sequence descriptor information and the polypeptide descriptor information are each transformed into a numerical matrix acceptable to the prediction model. The Transformer model calculates the correlation weights between different descriptor dimensions through the self-attention mechanism. For example, for AAC, it captures the synergistic effects of different amino acid frequencies (such as the impact on sequence function when a certain type of amino acid appears frequently); for DPC, it combines the combination patterns of adjacent amino acids with the global composition features to enhance the representation of local structures; for PHYC, it combines numerical values such as hydrophilicity and charge with sequence position information to generate features containing "position-property" correlations. Through multiple layers of Transformer encoders, the original descriptor information is mapped into high-dimensional feature vectors (i.e., TCR descriptor features and polypeptide descriptor features), which integrate the statistical laws, structural information, and physicochemical properties of the sequences.

[0067] Specifically, Figure 6 shows a polypeptide sequence descriptor vector, which is composed of calculating the overall composition of the polypeptide: amino acid composition (AAC), dipeptide composition (DPC), composition of k-spaced amino acid groups pairs (CKSAAGP), pseudo amino acid composition (PAAC), and physicochemical properties (PHYC). The polypeptide sequence descriptor vector is of floating-point data type. Among them, Figure 6 This is only an exemplary illustration and does not constitute a limitation on the specific solution.

[0068] In this embodiment, obtaining TCR descriptor features and polypeptide descriptor features based on the Transformer model is beneficial to enhancing the global correlation of features, improving the abstraction level and discriminative ability of features. For example, the descriptors are statistical values based on fixed rules, while the Transformer model can transform them into more abstract function-related features through multiple layers of neural networks. Moreover, the descriptor information usually contains multiple types (statistical, structural, physicochemical), and the parallel computing ability of the Transformer can process information in different dimensions simultaneously. During the prediction process, using the Transformer model to extract TCR descriptor features and polypeptide descriptor features can directly reflect the association between the sequence and function and the potential mechanism of molecular binding.

[0069] In this embodiment, as Figure 2, based on the TCR descriptor features and polypeptide descriptor features, a feed-forward neural network is used to obtain fused descriptor features. The TCR descriptor features and polypeptide descriptor features are usually high-dimensional vectors, which respectively characterize the abstract characteristics of CD4 + TCR and polypeptide sequences. The feed-forward neural network performs non-linear transformation and fusion on the two sets of features through multiple fully-connected layers, and can capture the non-linear interaction relationship between CD4 + TCR and polypeptide, and output the fused descriptor features (new high-dimensional vectors).

[0070] Feature fusion based on a feed-forward neural network can be flexibly applied to different feature dimensions. For example, if the TCR feature is 128-dimensional and the polypeptide feature is 256-dimensional, the feed-forward neural network can directly splice them into a 384-dimensional input and reduce the dimension to a suitable dimension (such as 256-dimensional) through the hidden layer. Moreover, the weights of the fully-connected layers of the feed-forward neural network can intuitively reflect the importance of features (for example, the larger the absolute value of the weight corresponding to a certain input feature, the greater the impact on the fused feature), which helps to understand the key interaction sites between CD4 + TCR and polypeptide.

[0071] In this embodiment, the fused features of the feed-forward neural network and the fused features of the cross-attention model complement each other. For example, the cross-attention model focuses on capturing the position-dependent relationship between features (such as the corresponding association between a certain region feature of CD4 + TCR and a certain region feature of polypeptide), and is suitable for modeling the spatial interaction pattern of sequences. The feed-forward neural network focuses on the non-linear combination of global features, is suitable for compressing features from different sources into a unified representation, and is more computationally efficient. In this way, local-global multi-level features can be formed, which are suitable for scenarios that require efficient integration of multi-source features and capture complex associations, and are beneficial to improving the prediction accuracy of the binding force between CD4 + TCR and polypeptide.

[0072] Returning to the embodiment of the present application, in step S103, based on the CD4 + TCR sequence and polypeptide sequence, TCR spatial features and polypeptide spatial features are respectively obtained, and the TCR spatial features and polypeptide spatial features are fused to obtain fused spatial features.

[0073] Among them, the binding of CD4 + TCR and polypeptide is a conformational match in three-dimensional space, not a simple match of one-dimensional sequences. By fusing the TCR spatial features and polypeptide spatial features to obtain fused spatial features, it can be closer to the essence of the interaction between biomolecules, can make up for the limitations of sequence features, and can also improve the prediction ability in complex interaction scenarios.

[0074] It is beneficial for the prediction model to understand CD4 from a multi-dimensional perspective +The interaction mechanism between TCR and peptide can significantly improve the ability to model biomolecular interactions, especially for CD4 + Scenarios that are highly dependent on spatial conformation, such as TCR recognition.

[0075] In some embodiments, based on the CD4 + The three-dimensional coordinates of each amino acid in each sequence can be obtained by using the TCR sequence and the peptide sequence. Specifically, the HelixFoldSingle protein structure prediction model can be used to obtain the three-dimensional coordinates of the peptide and CD4 + The atomic three-dimensional space coordinates of TCR are calculated by taking the coordinates of each 'CA' atom in the three-dimensional space of the predicted sequence as the three-dimensional coordinates x, y, z of each amino acid, and at the same time taking the peptide and CD4 + The sequence length of TCR is padded with 0 to a certain length.

[0076] In some embodiments, according to the CD4 + The three-dimensional coordinates of each amino acid in the TCR sequence are used to obtain a TCR amino acid distance matrix; the polypeptide amino acid distance matrix is obtained based on the three-dimensional coordinates of each amino acid in the polypeptide sequence; the TCR spatial characteristics are obtained based on the TCR amino acid distance matrix using a convolutional neural network model, and the polypeptide spatial characteristics are obtained based on the polypeptide amino acid distance matrix.

[0077] Specifically, the distance matrix is an n×n symmetric matrix (where n is the length of the amino acid sequence), and the element Dij represents the Euclidean distance between the i-th and j-th amino acid residues in the sequence in three-dimensional space. + The TCR amino acid distance matrix for the sequence is obtained by calculating the Euclidean distance between each atom (i.e., amino acid) in the TCR sequence based on the three-dimensional coordinates of each atom (i.e., amino acid). The peptide amino acid distance matrix for the sequence is obtained by calculating the Euclidean distance between each atom (i.e., amino acid) in the peptide sequence based on the three-dimensional coordinates of each atom (i.e., amino acid). For example, the TCR amino acid distance matrix is 30×30, and the peptide amino acid distance matrix is 22×22.

[0078] The three-dimensional coordinates of the sequence are Figure 7 As shown, the position coordinates corresponding to 'CA' are extracted as the three-dimensional coordinates of the amino acid, and the Euclidean distances between the two amino acid atoms are calculated based on the three-dimensional coordinates to generate a distance matrix. The distance matrix between amino acids is as follows: Figure 8 As shown, the peptide amino acid distance matrix is 22×22, and the TCR amino acid distance matrix is 30×30. They are floating point types, and the label data is 0, 1, where 0 represents that the peptide does not bind to the TCR, and 1 represents that the peptide binds to the TCR.

[0079] This embodiment is merely an illustrative description and does not constitute a limitation on the specific solution.

[0080] The convolutional neural network model of this embodiment consists of a multi-scale convolutional neural network and a feed-forward fully-connected network. As Figure 4 shown, the convolutional neural network is composed of multiple two-dimensional convolutional layers, normalization layers, Dropout layers, and max pooling layers. In the fully-connected network, there are Dropout layers and normalization layers, and the activation function of the intermediate layer uses LeakyReLU.

[0081] As Figure 2 , the TCR amino acid distance matrix and the polypeptide amino acid distance matrix are respectively input into two convolutional neural network models to extract the TCR spatial features and polypeptide spatial features, and the outputs of these two convolutional neural network models are concatenated and then input into the feed-forward neural network. That is to say, based on the TCR spatial features and polypeptide spatial features, the fused spatial features are obtained by using the feed-forward neural network.

[0082] Specifically, the TCR amino acid distance matrix and the polypeptide amino acid distance matrix are two-dimensional matrices. Using the convolutional neural network model for feature extraction can capture the distance patterns of local amino acid clusters, such as the spatial triangular structure formed by three consecutive amino acids, can also automatically capture long-distance spatial dependencies, and can also retain the relative distance features between amino acids in three-dimensional space.

[0083] Using the feed-forward neural network to fuse the TCR spatial features and polypeptide spatial features to obtain the fused spatial features can flexibly integrate spatial feature vectors of different dimensions and can simulate the global conformational matching when CD4 + TCR binds to the polypeptide in three dimensions, and can better capture the synergistic effects between features.

[0084] Returning to the embodiment of the present application, in step S104, based on the fused sequence features, fused descriptor features, and fused spatial features, the binding force between CD4 + TCR and the polypeptide is predicted, which can significantly improve the prediction accuracy.

[0085] Specifically, as Figure 2 , the fused sequence features, fused descriptor features, and fused spatial features can be concatenated and then input into the feed-forward neural network to obtain the prediction result of the binding force between CD4 + TCR and the polypeptide. For example, the prediction model is trained multiple times using the training set, and the model parameters with the best prediction results are saved as the final prediction model.

[0086] Exemplarily, nn.init.xavier_uniform_ can be used to initialize the network parameters during the training process, the optimizer uses Adam, and a weight decay of 1e-6 is added at the same time.

[0087] Multimodal polypeptide and CD4+ The prediction model for T cell receptor binding is a binary classification model. The network output is processed by softmax. If the predicted value is greater than 0.5, it indicates that the polypeptide binds to CD4 + TCR, otherwise it does not bind.

[0088] In some embodiments, a validation set is used to verify the prediction effect of the optimal prediction model. The AUC value of the prediction model on the validation set is 0.85, and the prediction accuracy is relatively high.

[0089] Sequence features cannot directly reflect the spatial conformation, spatial features are difficult to explain some phenomena that depend on sequence context, and descriptor features lack dynamic interaction information. By fusing fused sequence features, fused descriptor features, and fused spatial features to obtain multimodal fusion features, the limitations of single features can be compensated for, and the three work together to point to strong binding force. Through the non-linear integration of the feed-forward neural network, not only can the prediction accuracy be improved, but also the biological understanding can be fed back through feature importance analysis, providing more comprehensive computational support for immunotherapy and vaccine development.

[0090] In some embodiments, a CD4-based deep learning + prediction system for T cell receptor binding to polypeptides, such as Figure 9 The prediction system 900 includes a sequence processing unit 901, a descriptor processing unit 902, a spatial processing unit 903, and a fusion unit 904. Among them, the sequence processing unit 901 is configured to: respectively obtain TCR sequence features and polypeptide sequence features based on the CD4+ TCR sequence and the polypeptide sequence, and fuse the TCR sequence features and the polypeptide sequence features to obtain fused sequence features. The descriptor processing unit 902 is configured to: respectively obtain TCR descriptor features and polypeptide descriptor features based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR descriptor features and the polypeptide descriptor features to obtain fused descriptor features. The spatial processing unit 903 is configured to: respectively obtain TCR spatial features and polypeptide spatial features based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR spatial features and the polypeptide spatial features to obtain fused spatial features. The fusion unit 904 is configured to predict the binding force between CD4 + TCR and the polypeptide based on the fused sequence features, fused descriptor features, and fused spatial features, and the prediction accuracy is relatively high.

[0091] In some embodiments, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the deep learning-based CD4 of each embodiment of the present application +Steps in a method for predicting the binding of a T cell receptor to a polypeptide.

[0092] The above computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a phase change random access memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), a flash drive or other forms of flash memory, a cache, a register, a static memory, a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD) or other optical memory, a cassette tape or other magnetic storage device, or any other possible non-transitory medium used to store information or instructions that can be accessed by a computer device, etc.

[0093] According to an embodiment of the present application, an electronic device is further provided. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the CD4-based deep learning described in various embodiments of the present application are implemented. + Steps in a method for predicting the binding of a T cell receptor to a polypeptide.

[0094] As Figure 10 shown, the electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.

[0095] The processor may be a processing device including more than one general-purpose processing device, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be more than one dedicated processing device, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), etc.

[0096] Those skilled in the art can understand that Figure 10 the structure shown in is only a structural diagram of a part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have a different component arrangement.

[0097] This application describes various operations or functions, which may be implemented as software code or instructions or defined as software code or instructions. Such content may be source code that can be directly executed or differential code (“incremental” or “patch” code) (“object” or “executable” form). The software code or instructions may be stored in a computer-readable storage medium, and when executed, may cause a machine to perform the described functions or operations, and include any mechanism for storing information in a form accessible by a machine (e.g., a computing device, an electronic system, etc.), such as a recordable or non-recordable medium (e.g., read only memory (ROM), random access memory (RAM), magnetic disk storage medium, optical storage medium, flash device, etc.).

[0098] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application having equivalent elements, modifications, omissions, combinations (e.g., solutions that cross various embodiments), adaptations, or alterations. The elements in the claims will be broadly interpreted based on the language employed in the claims and not limited to the examples described in the specification or during the implementation of the present application, and the examples will be interpreted as non-exclusive. Thus, the specification and examples are intended to be considered only as examples, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.

[0099] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description. Additionally, in the above detailed description, various features may be grouped together to simplify the present application. This should not be construed as an intention that features disclosed without claim are necessary for any claim. On the contrary, the subject matter of the present application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by way of example or embodiment into the detailed description, where each claim stands on its own as a separate embodiment, and these embodiments may be combined with each other in various combinations or permutations. The scope of the present application should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.

[0100] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A prediction method for the binding of CD4 + T cell receptors to polypeptides, characterized in that The prediction method includes: Based on CD4 + The TCR sequence features and the polypeptide sequence features are obtained from the TCR sequence and the polypeptide sequence respectively, and the TCR sequence features and the polypeptide sequence features are fused to obtain the fused sequence features; Based on the CD4 + TCR sequence and the polypeptide sequence respectively obtain the TCR descriptor feature and the polypeptide descriptor feature, and fuse the TCR descriptor feature and the polypeptide descriptor feature to obtain a fused descriptor feature; Based on the CD4 + TCR sequence and polypeptide sequence, the TCR spatial feature and polypeptide spatial feature are obtained respectively, and the TCR spatial feature and polypeptide spatial feature are fused to obtain a fused spatial feature; Predict the binding affinity of CD4 + TCR and polypeptides based on the fused sequence features, fused descriptor features, and fused spatial features.

2. The prediction method according to claim 1, wherein The descriptors are amino acid composition, dipeptide composition, composition of k-spaced amino acid groups, pseudo amino acid composition, and physicochemical properties.

3. The prediction method according to claim 1, characterized in that, Based on the CD4 + TCR sequence and the polypeptide sequence, the TCR spatial features and the polypeptide spatial features are obtained respectively, specifically including: Based on the CD4 + The three-dimensional coordinates of each amino acid of each sequence are obtained based on the TCR sequence and the polypeptide sequence; According to the three-dimensional coordinates of each amino acid of the CD4 + TCR sequence, a TCR amino acid distance matrix is obtained; Obtain the distance matrix between amino acids of the polypeptide according to the three-dimensional coordinates of each amino acid of the polypeptide sequence; Use a convolutional neural network model to obtain the TCR spatial features based on the TCR amino acid distance matrix, and obtain the polypeptide spatial features based on the polypeptide amino acid distance matrix.

4. The prediction method according to claim 1, characterized in that Based on CD4 + The TCR sequence features and the polypeptide sequence features are obtained from the TCR sequence and the polypeptide sequence respectively, and the TCR sequence features and the polypeptide sequence features are fused to obtain the fused sequence features, specifically including: Digitize the amino acid sequences of the CD4 + TCR sequence and the polypeptide sequence respectively to obtain their respective digitized sequences; Based on CD4 + The digitized sequence obtained from the TCR sequence, and the Transformer model is used to obtain the TCR sequence features; Based on the digital sequence obtained from the polypeptide sequence, use the Transformer model to obtain the polypeptide sequence features; Based on the TCR sequence features and polypeptide sequence features, use the cross-attention model to obtain the fused sequence features.

5. The prediction method according to claim 4, wherein During the digital processing, it includes: Preset CD4 + The numerical length of the amino acid sequence of the TCR sequence is the first preset length. If the CD4 + actual numerical length of the amino acid sequence of the TCR sequence is less than the first preset length, it is padded with a preset number; Preset the digital length of the amino acid sequence of the polypeptide sequence as the second preset length. If the actual digital length of the amino acid sequence of the polypeptide sequence is less than the second preset length, it is padded with a preset number.

6. The prediction method according to claim 1, wherein Based on the above-mentioned CD4 + The TCR descriptor features and polypeptide descriptor features are obtained respectively from the TCR sequence and the polypeptide sequence, and the TCR descriptor features and the polypeptide descriptor features are fused to obtain the fused descriptor features, which specifically include: Based on the CD4 + Obtain TCR sequence descriptor information and polypeptide descriptor information from the TCR sequence and the polypeptide sequence respectively; Based on the TCR sequence descriptor information and polypeptide descriptor information, use the Transformer model to obtain the TCR descriptor features and polypeptide descriptor features respectively; Based on the TCR descriptor features and polypeptide descriptor features, use a feed-forward neural network to obtain the fused descriptor features.

7. The prediction method according to claim 2, characterized in that Fuse the TCR spatial features and polypeptide spatial features to obtain the fused spatial features, specifically including using a feed-forward neural network to obtain the fused spatial features based on the TCR spatial features and polypeptide spatial features.

8. A prediction system for the binding of CD4 + T cell receptor to polypeptide based on deep learning, characterized in that The prediction system includes: A sequence processing unit, configured to: based on the CD4 + TCR sequence and the polypeptide sequence respectively obtain the TCR sequence feature and the polypeptide sequence feature, and fuse the TCR sequence feature and the polypeptide sequence feature to obtain a fused sequence feature; Descriptor processing unit, configured to: obtain TCR descriptor features and polypeptide descriptor features respectively based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR descriptor features and the polypeptide descriptor features to obtain fused descriptor features; A spatial processing unit, configured to: obtain TCR spatial features and polypeptide spatial features respectively based on the CD4 + TCR sequence and the polypeptide sequence, and fuse the TCR spatial features and the polypeptide spatial features to obtain fused spatial features; A fusion unit configured to predict the CD4 + TCR and polypeptide binding affinity based on the fusion sequence feature, the fusion descriptor feature, and the fusion spatial feature.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the prediction method of binding of a CD4 + T cell receptor to a polypeptide according to any one of claims 1-7 are implemented. + ​ 10. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in the prediction method of the CD4 + T cell receptor binding to a polypeptide according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Method and system for predicting combination of T cell receptor and epitope

    CN114360644A

  • Method and system for predicting interaction between T cell receptor and peptide

    CN116665769A

  • Method for characterizing interaction conformation between T cell receptor and epitope

    CN117347613A

  • Prediction method and system for combination of polypeptide and TCR based on deep learning

    CN120108510A

  • Method for predicting protein-protein interaction

    US20230011678A1

Cited By

  • 6-HB targeted membrane fusion inhibitory peptide prediction method, device, equipment and medium

    CN120913657A

  • 6-hb targeting membrane fusion inhibiting peptide prediction method, apparatus, device, and medium

    CN120913657B