Intelligent prediction and antibody optimization design method for addiction-related epitopes

Through intelligent prediction and optimization design methods, low immunogenic regions and hidden epitopes are identified and antibody design is optimized, which solves the problems of low efficiency and poor accuracy of antibody development in the existing technology, achieves high efficiency and safety optimization of antibodies, and extends the drug effectiveness cycle.

CN120496624AInactive Publication Date: 2025-08-15张晨
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510434441.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot achieve dual optimization of antibody affinity and safety, resulting in low efficiency and poor accuracy of antibody development, and the problems of blind screening and inefficient verification.

Method used

The intelligent prediction and antibody optimization design method of addiction-related antigen epitopes was adopted. By integrating multi-source heterogeneous biological information data, low immunogenic regions were identified, high adaptability candidate epitope was generated, and antibody design was optimized using real-time molecular dynamics simulation and evolutionary data, state space and action space were constructed, and antibody sequence generation was performed based on the bidirectional GRU and softmax strategies, and cryptic epitopes with evolutionary conservative and structurally stable were screened out.

Benefits of technology

The dual optimization of antibody affinity and safety is achieved, the design efficiency and accuracy are improved, blind screening and inefficient verification in traditional methods are avoided, the drug effectiveness cycle is extended, and the risk of off-target is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496624A_ABST
    Figure CN120496624A_ABST
Patent Text Reader

Abstract

The invention discloses an addiction-related epitope intelligent prediction and antibody optimization design method, and belongs to the field of medicine.The design method specifically comprises the following steps that firstly, multi-source heterogeneous biological information data are integrated, and the potential space structure, dynamic stability and regulation and control characteristics of an epitope are extracted; iI, identifying a low-immunogenicity region in the addiction target, and generating a high-adaptability candidate epitope according to the identified low-immunogenicity region; iII, optimizing antibody design according to the high-adaptability candidate epitopes, and dynamically adjusting the antibody design according to real-time molecular dynamics simulation feedback; according to the method, double optimization of antibody affinity and safety can be realized, the design efficiency and accuracy are improved, the problems of blind screening and low-efficiency verification in a traditional method are avoided, and the recognition redundancy of the antibody is improved, so that the effective period of a drug is prolonged, and the off-target risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medicine, and in particular to a method for intelligent prediction of addiction-related antigen epitopes and optimized antibody design. Background Art

[0002] Addictive diseases have become a serious public health problem worldwide. Their long-term mechanisms involve a series of complex neurotransmitter pathways, immune inflammatory responses, and changes in specific antigenic structures. In recent years, antibody drugs have shown great potential in blocking the effects of addictive substances, clearing toxic metabolites from the blood, and preventing relapse. In particular, they have been initially clinically applied in the intervention of opioid overdose. However, current antibody development still faces many challenges: addiction-related antigen epitopes are highly dynamic, prone to mutation, and difficult to accurately locate; cross-species immunogenicity differences make it difficult to smoothly transfer antibodies that have been proven effective in early animal models to humans; antibody affinity, safety, drug resistance and other properties are difficult to optimize; there is a lack of interpretable modeling methods for antibody-antigen interaction mechanisms; therefore, it is particularly important to invent intelligent prediction methods for addiction-related antigen epitopes and antibody optimization design.

[0003] Existing methods for intelligent prediction of addiction-related antigen epitopes and optimized antibody design cannot achieve dual optimization of antibody affinity and safety, reducing design efficiency and accuracy, and are prone to blind screening and inefficient verification, which reduces the recognition redundancy of antibodies. To this end, we propose a method for intelligent prediction of addiction-related antigen epitopes and optimized antibody design. Summary of the Invention

[0004] The purpose of the present invention is to address the deficiencies in the prior art and to propose a method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] Intelligent prediction of addiction-related antigen epitopes and antibody optimization design method. The specific steps of this design method are as follows:

[0007] Ⅰ. Integrate multi-source heterogeneous bioinformatics data and extract the potential spatial structure, dynamic stability and regulatory characteristics of antigen epitopes;

[0008] II. Identify low immunogenic regions within addiction targets and generate highly adaptable candidate epitopes based on these identified low immunogenic regions;

[0009] III. Optimize antibody design based on highly adaptable candidate epitopes and dynamically adjust antibody design based on real-time molecular dynamics simulation feedback;

[0010] IV. Simulate the possible mutation paths of antigens during evolution, identify hidden epitopes, and then analyze the residue interaction paths during the binding process between antibodies and antigen epitopes;

[0011] V. Integrate the results of epitope prediction, antibody design and optimization analysis, and use biological experiments to evaluate the affinity, specificity and safety performance of antibodies.

[0012] As a further embodiment of the present invention, the specific steps of extracting the potential spatial structure, kinetic stability and regulatory characteristics of the antigen epitope in step I are as follows:

[0013] S1.1: Based on the amino acid sequence of the antigen, obtain its known or predicted three-dimensional structure model from the PDB and AlphaFold DB databases. Then, use homology modeling and the AlphaFold2 algorithm to complete the missing structure regions to generate the corresponding antigen structure. Then, perform energy minimization to generate the antigen structure.

[0014] S1.2: Use GROMACS simulation software to simulate the temporal conformational evolution of the processed antigen structure while sampling at the microsecond level. Based on the multiple timeframes obtained from the simulated sampling, the true dynamic changes of the protein under physiological conditions are captured.

[0015] S1.3: Perform cluster analysis on the evolutionary trajectories and divide each antigen conformation into multiple stable state clusters based on similarity to generate a conformational space corresponding to the antigen structure. Then, use free energy surface analysis to identify energy valleys in the conformational space, i.e., structural states exposed to the immune system.

[0016] S1.4: Use RMSD, RMSF, B-factor, and H-bond number to evaluate the stability of exposed regions. Then, use the rolling ball method to calculate the accessible surface area of each exposed region. Identify highly exposed or periodically exposed regions as candidate epitopes. Perform high-level feature encoding on the selected candidate epitopes to generate multiple sets of feature data, and normalize each feature data.

[0017] As a further embodiment of the present invention, the specific steps of identifying low immunogenicity regions in addiction-related targets in step II are as follows:

[0018] S2.1: Using the amino acid residues in the antigen as nodes, identify the proximity relationships between two residues in the three-dimensional structure based on the extracted feature data sets, and use these as edges between the two sets of nodes. Construct a protein structure graph based on the identified nodes and edges.

[0019] S2.2: Perform graph convolution on the protein structure graph to propagate the neighborhood information of each node in the protein structure graph to extract local conformational context features. Use a bidirectional recurrent neural network to model the conformational features of the antigen at different time frames to obtain the dynamic features of each antigen residue at each time frame.

[0020] S2.3: Based on the dynamic characteristics of each antigen residue, the attention weight of each antigen residue at the time step is calculated, and based on the weight of each antigen residue, the dynamic features of different time steps are integrated to generate the corresponding key residue representation. Then, the epitope prediction is performed on the attention-weighted key residue representation through the fully connected layer, and the residues are binary classified according to the preset threshold to output the final dynamic epitope prediction result, and the low immunogenicity region in the addictive target is identified based on the prediction result.

[0021] As a further embodiment of the present invention, the specific steps for generating highly adaptable candidate epitopes in step II are as follows:

[0022] S3.1: Collect highly immunogenic epitope data under various human HLA restrictions and use them as a real sample set, including high-affinity HLA-binding peptides, HLA allele labels corresponding to epitope sequences, and structural stability scores, and encode the identified low immunogenic regions as input latent variables;

[0023] S3.2: Build and initialize the candidate construction model, including the input layer, generator, discriminator, and output layer. Input the latent vector, the context encoding of the low immunogenicity region, and the real sample set into the candidate construction model. The input layer inputs the latent vector and the context encoding of the low immunogenicity region into the generator, and the real sample set into the discriminator.

[0024] S3.3: The generator processes each set of data transmitted by the input layer based on the initial weight matrix and bias term, and performs nonlinear activation processing through the softmax function, outputting a new candidate epitope sequence to the discriminator. The discriminator then receives the new candidate epitope sequence and the real sample set. The discriminator calculates the probability that each candidate epitope sequence is real human high adaptability data, and outputs the new candidate epitope sequence generated by the generator and the discriminator's discrimination result through the output layer;

[0025] S3.4: Based on the new candidate epitope sequence output by the generator and the discrimination result of the discriminator, the loss value of the candidate construction model is calculated using the adversarial loss function. The loss value is then input from the output layer and backpropagated based on the chain rule. The gradient of the loss value for each network layer of the candidate construction model is calculated layer by layer, and the parameters of each layer are adjusted using the Adam optimizer.

[0026] S3.5: After parameter adjustment is completed, the candidate construction model is retrained until the loss value of the candidate construction model converges to the preset threshold. Training is then stopped, and the low immunogenicity region of the newly identified addictive target is input into the candidate construction model. The corresponding candidate epitope is output through the forward propagation algorithm;

[0027] S3.6: Calculate the HLA binding affinity of candidate epitopes using NetMHCpan software, and screen out candidate epitopes with HLA binding affinity below a preset threshold. Then, embed the candidate epitopes into the original antigen structure and evaluate their impact on the overall conformational stability. If the energy change is higher than the preset threshold, screen out the corresponding candidate epitope and record the remaining candidate epitopes as highly adaptable candidate epitopes.

[0028] As a further embodiment of the present invention, the specific steps for optimizing antibody design based on the highly adaptable candidate epitopes in step III are as follows:

[0029] S4.1: Establish a state space based on the generated antibody sequence fragments and their binding simulation results to the highly adaptable candidate epitopes, and establish an action space based on different operations of selecting amino acid residues or structural fragments to be added to the antibody sequence under different states;

[0030] S4.2: A dual-objective reward function is established based on the affinity of the antibody-epitope binding and the risk of antibody-antibody cross-reactivity. At each time step, the current state is converted into a vector embedding representation. The current state includes the generated antibody residue sequence fragment, the current antibody-epitope docking structure information, and the structural fingerprint of the epitope region. A bidirectional GRU is then used to encode the antibody sequence context.

[0031] S4.3: Based on the vector embedding representation corresponding to the state, output the selection probability of each action in the action space for the current state through a fully connected layer and a softmax function to obtain the corresponding policy distribution. Then, sample from the policy distribution and select the next action.

[0032] S4.4: After executing an action, the corresponding residues are spliced into the current antibody sequence to construct a new state. Action selection and state update are then repeated until the length of the antibody sequence reaches a set threshold. The expected total reward of the generated antibody sequence is calculated based on the dual-objective reward function.

[0033] S4.5: Screen out antibody sequences whose expected total reward does not reach the preset threshold, then model the optimized antibody sequence as a complete structure, evaluate the folding quality of the optimized antibody sequence structure, and simulate the binding conformation and affinity again to evaluate the accuracy and stability of its binding position on the target antigen, and record each optimized antibody and its test results.

[0034] As a further embodiment of the present invention, the specific steps of simulating the possible mutation paths of the antigen during the evolution process and identifying the cryptic epitope in step IV are as follows:

[0035] S5.1: Based on the mutation data of multiple species or individuals of the existing antigen protein, construct an antigen residue mutation graph. In the graph, each node represents each amino acid residue in the protein sequence, and an edge represents the mutation path or co-mutation relationship between two residues. Based on the mutation frequency in the evolution database, set the initial pheromone concentration for each edge, and then initialize a set of populations;

[0036] S5.2: Initialize the starting nodes of each group of individuals in the population in the antigen residue mutation graph by random selection. Calculate the selection probability of each node based on the pheromone concentration of each edge and the structural perturbation cost of each residue after the change. Each individual at the current position selects the next residue using the roulette wheel method based on the selection probability of each node.

[0037] S5.3: Repeat node selection until a complete residue path is constructed. After each iteration, all individuals complete path sampling and score the conservation and structural stability of each residue path constructed. The pheromone concentration is updated based on the scoring results. Pathway construction and pheromone updating are repeated until the pheromone concentration change value of each residue path converges to the preset range.

[0038] S5.4: Count the residue paths with the highest frequency among all individuals, and select the set of residues with the least structural perturbation and that appear stably in multiple rounds of iterations as cryptic epitopes. Then, based on the identified candidate cryptic epitopes, increase the redundancy of antibody recognition.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] This method for intelligent prediction of addiction-related antigen epitopes and optimal design of antibodies constructs state space and action space, and based on the antibody-epitope binding structure, uses bidirectional GRU and softmax strategy sampling to gradually generate antibody sequences, and constructs a dual-objective reward function based on affinity and cross-reaction risk for optimization. Subsequently, an antigen mutation map is constructed through evolutionary data to generate a population containing multiple groups of individuals, and each individual is used to simulate the mutation path. The path search is guided by pheromone concentration and structural perturbation cost, and hidden epitopes with evolutionary conserved and stable structure are screened out. It can achieve dual optimization of antibody affinity and safety, improve design efficiency and accuracy, avoid the problems of blind screening and inefficient verification in traditional methods, and improve the recognition redundancy of antibodies, thereby extending the effective period of the drug and reducing the risk of off-target. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0042] Figure 1 This is a flowchart of the method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design proposed in the present invention. DETAILED DESCRIPTION

[0043] Example 1, with reference to Figure 1 , a method for intelligent prediction of addiction-related antigen epitopes and optimized antibody design, the specific steps of the design method are as follows:

[0044] Integrate multi-source heterogeneous bioinformatics data and extract the potential spatial structure, dynamic stability and regulatory characteristics of antigen epitopes.

[0045] Specifically, according to the amino acid sequence of the antigen, from PDB and AlphaFold The known or predicted three-dimensional structural model is obtained from the DB database, and then the regions with missing structures are completed using homology modeling and the AlphaFold2 algorithm to generate the corresponding antigen structure. The antigen structure is then generated by energy minimization. The conformational evolution behavior of the processed antigen structure in the time dimension is simulated by using GROMACS simulation software, and microsecond sampling is performed at the same time. Based on the multiple sets of time frame structures obtained by simulated sampling, the real dynamic change process of the protein under physiological conditions is obtained. The evolutionary trajectory is clustered and analyzed. Each antigen conformation is divided into multiple stable state clusters according to similarity to generate the conformational space of the corresponding antigen structure. The free energy surface analysis method is then used to identify the low energy region in the conformational space, that is, the structural state exposed to the immune system. The stability of the exposed region is evaluated using RMSD, RMSF, B-factor and H-bond number. The rolling ball method is then used to calculate the accessible surface area of each exposed region, and highly exposed or periodically exposed regions are identified as candidate epitopes. The selected candidate epitopes are encoded with advanced features to generate multiple sets of feature data, and each feature data is normalized.

[0046] Identify low-immunogenic regions in addiction targets and generate highly adaptable candidate epitopes based on the identified low-immunogenic regions.

[0047] Specifically, the amino acid residues in the antigen are used as nodes. According to the extracted sets of feature data, the proximity relationship between two residues in the three-dimensional structure is identified, and it is used as the edge between the two sets of nodes. According to the identified nodes and edges, a protein structure graph is constructed, and the protein structure graph is subjected to graph convolution processing. The neighborhood information of each node in the protein structure graph is propagated to extract local conformational context features. A bidirectional recurrent neural network is used to model the conformational features of the antigen in different time frames to obtain the dynamic features of each antigen residue in each time frame. According to the dynamic features of each antigen residue, the attention weight of each antigen residue in the time step is calculated, and based on the weight of each antigen residue, the dynamic features of different time steps are integrated to generate the corresponding key residue representation. Then, the epitope prediction of the attention-weighted key residue representation is performed through the fully connected layer, and the residues are binary classified according to the preset threshold to output the final dynamic epitope prediction result, and the low immunogenicity region in the addictive target is identified based on the prediction result.

[0048] Specifically, highly immunogenic epitope data under the restrictions of various human HLAs are collected and used as real sample sets, including high-affinity HLA binding peptides, HLA allele markers corresponding to epitope sequences, and structural stability scores. The identified low immunogenicity regions are encoded as input latent variables, and a candidate construction model is constructed and initialized, including an input layer, a generator, a discriminator, and an output layer. The latent vector, the context encoding of the low immunogenicity region, and the real sample set are input into the candidate construction model. The input layer inputs the latent vector and the context encoding of the low immunogenicity region into the generator, and the real sample set is input into the discriminator. The generator processes each set of data transmitted by the input layer based on the initial weight matrix and the bias term, and performs nonlinear activation processing through the softmax function, and outputs the new candidate epitope sequence to the discriminator. The discriminator then receives the new candidate epitope sequence and the real sample set. The discriminator calculates the probability that each candidate epitope sequence is real human high adaptability data, and outputs the new candidate epitope sequence generated by the generator through the output layer. The loss value of the candidate construction model is calculated by the adversarial loss function based on the new candidate epitope sequence output by the generator and the discrimination result of the discriminator. The loss value is then input from the output layer. Based on the chain rule, backpropagation is performed, and the gradient of the loss value for each network layer of the candidate construction model is calculated layer by layer. The parameters of each layer are then adjusted using the Adam optimizer. After the parameter adjustment is completed, the candidate construction model is retrained until the loss value of the candidate construction model converges to the preset threshold. Training is stopped, and the low immunogenicity region in the newly identified addictive target is input into the candidate construction model. The corresponding candidate epitope is output through the forward propagation algorithm, and the HLA binding affinity of the candidate epitope is calculated using the NetMHCpan software. The candidate epitope with HLA binding affinity below the preset threshold is screened out. The candidate epitope is then embedded into the original antigen structure to evaluate its impact on the overall conformational stability. If the energy change is higher than the preset threshold, the corresponding candidate epitope is screened out, and the remaining candidate epitopes are regarded as high-adaptability candidate epitopes and recorded.

[0049] Example 2, reference Figure 1 , a method for intelligent prediction of addiction-related antigen epitopes and optimized antibody design, the specific steps of the design method are as follows:

[0050] Optimize antibody design based on highly adaptable candidate epitopes and dynamically adjust antibody design based on real-time molecular dynamics simulation feedback.

[0051] Specifically, based on the generated antibody sequence fragments and their binding simulation results to highly adaptable candidate epitopes, a state space is established. Based on different operations of selecting amino acid residues or structural fragments to be added to the antibody sequence under different states, an action space is established. The affinity of the antibody and epitope binding, as well as the risk of antibody and cross-reaction, are comprehensively considered to establish a dual-objective reward function. Then, at each time step, the current state is converted into a vector embedding representation. The current state includes the generated antibody residue sequence fragments, the current antibody-epitope docking structure information, and the structural fingerprint of the epitope region. The bidirectional GRU is then used to encode the antibody sequence context. According to the vector embedding representation corresponding to the state, the current state is output through the fully connected layer and the softmax function. The probability of selecting each action in the state in the action space is used to obtain the corresponding strategy distribution, and then sampling is performed from the strategy distribution to select the next action. After the action is executed, the corresponding residues are spliced into the current antibody sequence to construct a new state. The action selection and state update are then repeated until the length of the antibody sequence reaches the set threshold. The expected total reward of the generated antibody sequence is calculated based on the dual-objective reward function, and the antibody sequences whose expected total reward does not reach the preset threshold are screened out. The optimized antibody sequence is then modeled as a complete structure, and the folding quality of the optimized antibody sequence structure is evaluated. The binding conformation and affinity are simulated again to evaluate the accuracy and stability of its binding position on the target antigen, and each optimized antibody and its test results are recorded.

[0052] Simulate the possible mutation paths of antigens during evolution, identify hidden epitopes, and then analyze the residue interaction paths during the binding process between antibodies and antigen epitopes.

[0053] Specifically, based on the variation data of multiple species or individuals of existing antigen proteins, an antigen residue mutation graph is constructed. In the graph, each node represents each amino acid residue in the protein sequence, and an edge represents a mutation path or a co-mutation relationship between two residues. The initial pheromone concentration is set for each edge according to the mutation frequency of the evolutionary database, and then a group of populations is initialized. By random selection, the starting nodes of each group of individuals in the population in the antigen residue mutation graph are initialized. The selection probability of each node is calculated based on the pheromone concentration of each edge and the structural perturbation cost after the change of each residue. Each individual is at the current position, and the selection probability of each node is calculated based on the selection probability of each node. The next residue is selected by the roulette wheel method, and the node selection is repeated until a complete residue path is constructed. After each round of iteration, all individuals complete the path sampling, and the conservation and structural stability of each group of residue paths constructed are scored. The pheromone concentration is updated according to the scoring results, and the path construction and pheromone update are repeated until the pheromone concentration change value of each residue path converges to a preset range. The residue paths with the highest frequency of all individuals are counted, and the residue set with the smallest structural perturbation and stable appearance in multiple rounds of iterations is selected as the cryptic epitope. Then, based on the identified candidate cryptic epitopes, the antibody recognition redundancy is increased.

[0054] Integrate epitope prediction, antibody design and optimization analysis results, and use biological experiments to evaluate the affinity, specificity and safety performance of antibodies.

Claims

1. A method for intelligent prediction of addiction-related antigen epitopes and optimized antibody design, characterized in that: The specific steps of this design method are as follows: Ⅰ. Integrate multi-source heterogeneous bioinformatics data and extract the potential spatial structure, dynamic stability and regulatory characteristics of antigen epitopes; II. Identify low immunogenic regions within addiction targets and generate highly adaptable candidate epitopes based on these identified low immunogenic regions; III. Optimize antibody design based on highly adaptable candidate epitopes and dynamically adjust antibody design based on real-time molecular dynamics simulation feedback; IV. Simulate the possible mutation paths of antigens during evolution, identify hidden epitopes, and then analyze the residue interaction paths during the binding process between antibodies and antigen epitopes; V. Integrate the results of epitope prediction, antibody design and optimization analysis, and use biological experiments to evaluate the affinity, specificity and safety performance of antibodies.

2. The method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design according to claim 1, characterized in that: The specific steps for extracting the potential spatial structure, kinetic stability and regulatory characteristics of the antigen epitope described in step I are as follows: S1.1: Based on the amino acid sequence of the antigen, obtain its known or predicted three-dimensional structure model from the PDB and AlphaFold DB databases. Then, use homology modeling and the AlphaFold2 algorithm to complete the missing structure regions to generate the corresponding antigen structure. Then, perform energy minimization to generate the antigen structure. S1.2: Use GROMACS simulation software to simulate the temporal conformational evolution of the processed antigen structure while sampling at the microsecond level. Based on the multiple timeframes obtained from the simulated sampling, the true dynamic changes of the protein under physiological conditions are captured. S1.3: Perform cluster analysis on the evolutionary trajectories and divide each antigen conformation into multiple stable state clusters based on similarity to generate a conformational space corresponding to the antigen structure. Then, use free energy surface analysis to identify energy valleys in the conformational space, i.e., structural states exposed to the immune system. S1.4: Use RMSD, RMSF, B-factor, and H-bond number to evaluate the stability of exposed regions. Then, use the rolling ball method to calculate the accessible surface area of each exposed region. Identify highly exposed or periodically exposed regions as candidate epitopes. Perform high-level feature encoding on the selected candidate epitopes to generate multiple sets of feature data, and normalize each feature data.

3. The method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design according to claim 2, characterized in that: The specific steps for identifying low immunogenic regions in addiction-related targets in Step II are as follows: S2.1: Using the amino acid residues in the antigen as nodes, identify the proximity relationships between two residues in the three-dimensional structure based on the extracted feature data sets, and use these as edges between the two sets of nodes. Construct a protein structure graph based on the identified nodes and edges. S2.2: Perform graph convolution on the protein structure graph to propagate the neighborhood information of each node in the protein structure graph to extract local conformational context features. Use a bidirectional recurrent neural network to model the conformational features of the antigen at different time frames to obtain the dynamic features of each antigen residue at each time frame. S2.3: Based on the dynamic characteristics of each antigen residue, the attention weight of each antigen residue at the time step is calculated, and based on the weight of each antigen residue, the dynamic features of different time steps are integrated to generate the corresponding key residue representation. Then, the epitope prediction is performed on the attention-weighted key residue representation through the fully connected layer, and the residues are binary classified according to the preset threshold to output the final dynamic epitope prediction result, and the low immunogenicity region in the addictive target is identified based on the prediction result.

4. The method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design according to claim 3, characterized in that: The specific steps for generating highly adaptable candidate epitopes in step II are as follows: S3.1: Collect highly immunogenic epitope data under various human HLA restrictions and use them as a real sample set, including high-affinity HLA-binding peptides, HLA allele labels corresponding to epitope sequences, and structural stability scores, and encode the identified low immunogenic regions as input latent variables; S3.2: Build and initialize the candidate construction model, including the input layer, generator, discriminator, and output layer. Input the latent vector, the context encoding of the low immunogenicity region, and the real sample set into the candidate construction model. The input layer inputs the latent vector and the context encoding of the low immunogenicity region into the generator, and the real sample set into the discriminator. S3.3: The generator processes each set of data transmitted by the input layer based on the initial weight matrix and bias term, and performs nonlinear activation processing through the softmax function, outputting a new candidate epitope sequence to the discriminator. The discriminator then receives the new candidate epitope sequence and the real sample set. The discriminator calculates the probability that each candidate epitope sequence is real human high adaptability data, and outputs the new candidate epitope sequence generated by the generator and the discriminator's discrimination result through the output layer; S3.4: Based on the new candidate epitope sequence output by the generator and the discrimination result of the discriminator, the loss value of the candidate construction model is calculated using the adversarial loss function. The loss value is then input from the output layer and backpropagated based on the chain rule. The gradient of the loss value for each network layer of the candidate construction model is calculated layer by layer, and the parameters of each layer are adjusted using the Adam optimizer. S3.5: After parameter adjustment is completed, the candidate construction model is retrained until the loss value of the candidate construction model converges to the preset threshold. Training is then stopped, and the low immunogenicity region of the newly identified addictive target is input into the candidate construction model. The corresponding candidate epitope is output through the forward propagation algorithm; S3.6: Calculate the HLA binding affinity of candidate epitopes using NetMHCpan software, and screen out candidate epitopes with HLA binding affinity below a preset threshold. Then, embed the candidate epitopes into the original antigen structure and evaluate their impact on the overall conformational stability. If the energy change is higher than the preset threshold, screen out the corresponding candidate epitope and record the remaining candidate epitopes as highly adaptable candidate epitopes.

5. The method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design according to claim 4, characterized in that: The specific steps for optimizing antibody design based on the highly adaptable candidate epitopes described in Step III are as follows: S4.1: Establish a state space based on the generated antibody sequence fragments and their binding simulation results to the highly adaptable candidate epitopes, and establish an action space based on different operations of selecting amino acid residues or structural fragments to be added to the antibody sequence under different states; S4.2: A dual-objective reward function is established based on the affinity of the antibody-epitope binding and the risk of antibody-antibody cross-reactivity. At each time step, the current state is converted into a vector embedding representation. The current state includes the generated antibody residue sequence fragment, the current antibody-epitope docking structure information, and the structural fingerprint of the epitope region. A bidirectional GRU is then used to encode the antibody sequence context. S4.3: Based on the vector embedding representation corresponding to the state, output the selection probability of each action in the action space for the current state through a fully connected layer and a softmax function to obtain the corresponding policy distribution. Then, sample from the policy distribution and select the next action. S4.4: After executing an action, the corresponding residues are spliced into the current antibody sequence to construct a new state. Action selection and state update are then repeated until the length of the antibody sequence reaches a set threshold. The expected total reward of the generated antibody sequence is calculated based on the dual-objective reward function. S4.5: Screen out antibody sequences whose expected total reward does not reach the preset threshold, then model the optimized antibody sequence as a complete structure, evaluate the folding quality of the optimized antibody sequence structure, and simulate the binding conformation and affinity again to evaluate the accuracy and stability of its binding position on the target antigen, and record each optimized antibody and its test results.

6. The method for intelligent prediction of addiction-related antigen epitopes and optimal antibody design according to claim 5, characterized in that: The specific steps for simulating the possible mutation paths of antigens during evolution and identifying cryptic epitopes in step IV are as follows: S5.1: Based on the mutation data of multiple species or individuals of the existing antigen protein, construct an antigen residue mutation graph. In the graph, each node represents each amino acid residue in the protein sequence, and an edge represents the mutation path or co-mutation relationship between two residues. Based on the mutation frequency in the evolution database, set the initial pheromone concentration for each edge, and then initialize a set of populations; S5.2: Initialize the starting nodes of each group of individuals in the population in the antigen residue mutation graph by random selection. Calculate the selection probability of each node based on the pheromone concentration of each edge and the structural perturbation cost of each residue after the change. Each individual at the current position selects the next residue using the roulette wheel method based on the selection probability of each node. S5.3: Repeat node selection until a complete residue path is constructed. After each iteration, all individuals complete path sampling and score the conservation and structural stability of each residue path constructed. The pheromone concentration is updated based on the scoring results. Pathway construction and pheromone updating are repeated until the pheromone concentration change value of each residue path converges to the preset range. S5.4: Count the residue paths with the highest frequency among all individuals, and select the set of residues with the least structural perturbation and that appear stably in multiple rounds of iterations as cryptic epitopes. Then, based on the identified candidate cryptic epitopes, increase the redundancy of antibody recognition.

Citation Information

Cited By

  • Automatic screening and optimizing method and system for bispecific immune receptors

    CN121191579A