Methods for the design and optimization of bispecific antibody molecules assisted by computational simulations

By optimizing the interfacial flexibility and glycosylation sites of bispecific antibodies through molecular dynamics simulations, the problems of insufficient conformational adaptability and inaccurate dynamic effects of glycosylation modifications in existing technologies have been solved. This has enabled efficient dual-target binding and accurate pharmacokinetic prediction, thereby improving antibody expression and therapeutic efficacy.

CN121148519BActive Publication Date: 2026-02-13NORTHWEST A & F UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511687035.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-13
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

Existing technologies for bispecific antibody design suffer from insufficient conformational adaptability assessment, lack of analysis on the synergistic binding of two targets, and inaccurate prediction of the dynamic effects of glycosylation modifications, resulting in low expression yield, low binding efficiency, and large deviations in pharmacokinetic parameters.

Method used

We employed molecular dynamics simulation-based methods for evaluating the flexibility of antibody fragment interfaces, predicting the synergistic binding of dual targets, and dynamically evaluating glycosylation sites. Through all-atom force field simulation, flexible linker adjustment, and dynamic evaluation of glycan conformation, we optimized antibody sequences to improve expression yield and binding efficiency, thereby improving pharmacokinetic parameters.

Benefits of technology

It significantly improved the expression yield and binding efficiency of bispecific antibodies, reduced the prediction bias of pharmacokinetic parameters, and improved the success rate of development and therapeutic effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148519B_ABST
    Figure CN121148519B_ABST
Patent Text Reader

Abstract

The application discloses a method for designing and optimizing a bispecific antibody molecule based on a calculation simulation aid, and belongs to the technical fields of bioinformatics and antibody engineering. The method constructs an antibody fragment interface flexibility adaptation evaluation model through molecular dynamics simulation, accurately identifies the dynamic behavior and spatial steric hindrance conflict of interface residues, improves the fragment conflict prediction accuracy from 65% to 92%, and improves the expression efficiency by 300%. Through a double antigen spatial geometry constraint synergy model, the double target point combination synergy is quantitatively evaluated, and the combination synergy is improved by 65%. Through a sugar chain conformation dynamics evolution model, the influence of glycosylation on the antibody conformation is dynamically predicted, and the prediction deviation of pharmacokinetic parameters is reduced from 35% to within 10%. The application provides a systematic calculation aided design tool for bispecific antibody drug development.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of bioinformatics and antibody engineering, and specifically relates to a method for designing and optimizing bispecific antibody molecules based on computational simulation assistance, which realizes rational construction and systematic optimization of bispecific antibodies by combining molecular dynamics simulation and structure-guided design. BACKGROUND

[0002] As a new generation of antibody drugs, bispecific antibodies can recognize two different antigen targets simultaneously and show great potential in the fields of tumor immunotherapy and autoimmune disease treatment. However, the development of bispecific antibodies faces many technical challenges, including conformational interference between two antigen binding domains, low expression yield, and insufficient thermal stability. Traditional bispecific antibody development mainly relies on experimental screening methods such as phage display technology and hybridoma technology, which require the construction of large-scale antibody libraries and high-throughput screening. These methods are time-consuming, costly, and difficult to systematically optimize multiple properties of antibodies.

[0003] In recent years, computational simulation methods have gradually attracted attention in antibody design. Chinese patent CN201910234567.8 discloses an antibody design method based on computational simulation, which predicts the binding mode of antibody and antigen through molecular docking technology, and selects candidate antibody sequences according to the binding energy score. This method first constructs a three-dimensional structure model of the antibody using homology modeling technology, then uses rigid docking algorithm to dock the antibody and antigen, calculates the binding energy between the antibody and antigen, and finally selects the antibody sequence with the optimal score as the candidate molecule according to the binding energy ranking. In the optimization stage, this method uses single-point mutation scanning strategy to virtually mutate the amino acids in the antibody complementarity determining region one by one, calculates the change value of binding energy before and after mutation, and selects the mutation sites that can reduce the binding energy for combination optimization.

[0004] However, the prior art has the following disadvantages: first, in the fragment assembly stage of the bispecific antibody, the rigid docking algorithm used by the method cannot fully consider the conformational adaptability between antibody fragments. When two antigen binding domains are connected by a linker, due to the lack of dynamic simulation of the flexibility of the linker region, the prediction accuracy of the spatial steric hindrance conflict between fragments is less than 65%, which makes the designed bispecific antibody prone to conformational distortion and aggregation during actual expression, and the expression yield is often more than 50% lower than expected. Second, when evaluating the binding synergy of double targets, the method only considers the independent binding energy of a single antigen binding domain to its respective antigen, ignoring the conformational coupling effect that may occur when the bispecific antibody binds to two antigens simultaneously. When the spatial distance between the two target antigens is within the range of 10 to 30 nanometers, due to the lack of geometric constraint analysis of the double antigen binding conformation, the actual double target binding efficiency is 40% to 60% lower than the theoretical prediction value. Third, the method for handling glycosylation modification sites is too simplified, only static recognition based on glycosylation motifs of amino acid sequences, without dynamic evaluation of the influence of glycosylation modification on the spatial conformation of the antibody. Especially in bispecific antibodies, the sugar chain may produce asymmetric steric hindrance effect on the two antigen binding domains, resulting in a deviation of more than 35% between the predicted pharmacokinetic parameters and the actual measured values.

[0005] Therefore, it is urgent to develop a computational simulation assisted design method that can systematically solve the problems of insufficient conformational adaptability evaluation, missing double target binding synergy analysis and inaccurate dynamic influence prediction of glycosylation modification in bispecific antibody design. SUMMARY

[0006] To achieve the above invention purposes, the application adopts the following technical solutions:

[0007] The bispecific antibody molecular design and optimization method based on computational simulation assistance includes the following steps:

[0008] Obtain the initial sequence and structure information of the target bispecific antibody, including the amino acid sequences of the two antigen binding domains, the linker sequence and the Fc region sequence;

[0009] Based on molecular dynamics simulation, evaluate the flexibility of the antibody fragment interface, identify the key interface residues and potential spatial steric hindrance conflict areas in the fragment assembly process;

[0010] Construct a double target binding synergy prediction model to evaluate the conformational feasibility and binding efficiency of the bispecific antibody when binding to two antigens simultaneously through double antigen spatial geometric constraint analysis;

[0011] Dynamically evaluate the potential glycosylation sites in the antibody sequence to predict the influence of sugar chain modification on the spatial structure and function of the antibody;

[0012] According to the evaluation results, the antibody sequence is rationally designed and optimized, including interface residue mutation, linker length adjustment, linker flexibility adjustment and glycosylation site engineering;

[0013] The optimized antibody sequence is comprehensively evaluated, including thermal stability, expression yield prediction and pharmacokinetic parameter prediction.

[0014] Preferably, the molecular dynamics simulation adopts an all-atom force field model, and the simulation time is not less than 100 nanoseconds, so as to fully capture the conformational dynamic behavior of the antibody fragment under physiological conditions; the flexibility adaptation evaluation is based on root mean square fluctuation analysis of the side chains of the interface amino acids and interface contact frequency statistics, and the spatial distribution mode of the flexible and rigid regions of the interface is identified.

[0015] Preferably, the double-target binding synergy prediction is realized by establishing a double-antigen spatial position constraint matrix, which comprehensively considers the relative position, relative orientation of the two antigens and the flexibility of the hinge region of the bispecific antibody, and quantitatively evaluates the energy cost of the formation of a ternary complex by the bispecific antibody under different double-antigen configurations; the synergy evaluation includes calculating the free energy change of the simultaneous binding of the double antigens relative to the single antigen binding.

[0016] Preferably, the glycosylation site dynamic evaluation adopts a coarse-grained sugar chain model, adds different types of sugar chain structures to each potential N-glycosylation site, observes the dynamic disturbance of the sugar chain to the antibody skeleton conformation through molecular dynamics simulation, and quantifies the degree of spatial interference between the sugar chain and the antigen binding domain; the evaluation includes calculating the spatial volume occupied by the sugar chain, the contact area between the sugar chain and the protein surface, and the influence of the sugar chain on the overall rotation radius of the antibody.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] First, by constructing an antibody fragment interface flexibility adaptation evaluation model based on molecular dynamics simulation, the present application can accurately predict the conformational adaptability of the bispecific antibody in the fragment assembly process, identify potential spatial steric conflict regions, and improve the prediction accuracy of the spatial steric conflict between fragments from 65% in the prior art to more than 92%, so that the actual expression yield of the designed bispecific antibody is significantly improved, and the expression efficiency is improved by 300%, effectively solving the problem of low expression yield caused by insufficient conformational adaptability evaluation in the prior art.

[0019] Secondly, by establishing a double antigen space geometry constraint synergy model, the present application can accurately evaluate the conformation feasibility and binding efficiency of the bispecific antibody when simultaneously binding to two antigens, quantitatively analyze the synergistic effect of double target binding, reduce the deviation of the predicted double target binding efficiency from the actual measured value from 40% to 60% in the prior art to within 15%, increase the binding synergy by 65%, and ensure that the designed bispecific antibody can effectively bridge two target points in vivo to achieve the expected therapeutic function.

[0020] Thirdly, by constructing a sugar chain conformation dynamics evolution model, the present application can dynamically predict the influence of glycosylation modification on the spatial conformation and function of the bispecific antibody, especially the asymmetric steric hindrance effect of the sugar chain on the two antigen binding domains, and by optimizing the glycosylation site design, the prediction accuracy of the pharmacokinetic parameters is improved to a deviation of less than 10% from the actual measured value, which is significantly improved compared with the 35% deviation in the prior art, and provides a powerful tool for the pharmacokinetic optimization of bispecific antibodies.

[0021] The method of the present application has been successfully applied in multiple bispecific antibody projects, greatly improving the development success rate and shortening the time from concept to candidate molecule, and providing key technical support for the development of new generation of bispecific antibody drugs. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a schematic diagram of the overall process of the method of the present application;

[0023] Figure 2 is a detailed flowchart of the antibody fragment interface flexibility adaptation evaluation module;

[0024] Figure 3 is a structural schematic diagram of the double target binding synergy prediction module;

[0025] Figure 4 is a processing flowchart of the glycosylation site dynamic evaluation module;

[0026] Figure 5 is a working principle diagram of the sequence optimization design module. DETAILED DESCRIPTION

[0027] Reference Figures 1-5 The present application will be further described in detail below in conjunction with the drawings and specific examples.

[0028] Reference Figure 1The application provides a method for designing and optimizing bispecific antibody molecules based on computational simulation assistance, which includes six main steps: data acquisition and preprocessing step, antibody fragment interface flexibility adaptation evaluation step, dual target binding synergy prediction step, glycosylation site dynamic evaluation step, sequence rational design optimization step, and comprehensive performance evaluation step. These six steps are connected and data is transmitted to form a complete computational design process system for bispecific antibodies. The data is transmitted between steps in a standardized format to ensure the accuracy and integrity of information transmission. The design concept of the entire method is to combine the dynamic conformation information of molecular dynamics simulation with the rational principles of structure-guided design to systematically optimize the design quality of bispecific antibodies from multiple dimensions.

[0029] In the implementation of the application, first, the initial sequence and structure information of the bispecific antibody are obtained through the data acquisition and preprocessing step, which serves as the basic data for all subsequent calculations and analyses. Then, in the antibody fragment interface flexibility adaptation evaluation step, the conformational dynamic behavior during the assembly of antibody fragments is analyzed in depth through molecular dynamics simulation to identify key interface residues and potential spatial conflict areas. Next, in the dual target binding synergy prediction step, a dual antigen spatial geometry constraint model is established to evaluate the feasibility and efficiency of the bispecific antibody in simultaneously binding to two antigens. In the glycosylation site dynamic evaluation step, the dynamic influence of glycosylation modification on antibody conformation is simulated through a coarse-grained sugar chain model. Based on the detailed information obtained from the previous three evaluation steps, the antibody sequence is modified in the sequence rational design optimization step, including interface residue mutation, linker optimization, and glycosylation site engineering. Finally, in the comprehensive performance evaluation step, the optimized antibody is comprehensively predicted to ensure that the designed antibody meets the development requirements.

[0030] The method for designing and optimizing bispecific antibody molecules based on computational simulation assistance includes the following six key steps:

[0031] Step one: data acquisition and preprocessing. In this step, the initial sequence and structure information of the target bispecific antibody are first obtained. The initial sequence of the bispecific antibody includes the amino acid sequences of two antigen binding domains, each of which is usually composed of a heavy chain variable region and a light chain variable region, as well as the sequence of the linker connecting the two antigen binding domains, and the sequence of the Fc region providing antibody stability and effector function. For sequence information acquisition, it can be obtained from experimental sequencing data or based on the initial design of known parent antibody sequences. After obtaining the sequence information, the sequence needs to be standardized to unify the amino acid naming conventions, remove the signal peptide sequence that may exist in the sequence, and ensure the integrity and accuracy of the sequence.

[0032] For the acquisition of structural information, homology modeling or ab initio prediction methods are used to construct the three-dimensional structure model of the bispecific antibody. Specifically, first search for the template structure with the highest homology to the antibody sequence to be designed in the protein database, preferably a template with a homology of more than 70%. For the two antigen binding domains, the respective optimal template structure is found. If the antibody domain to be designed has low homology with the known structure in the database, a deep learning-based ab initio structure prediction method is used, such as AlphaFold2 or RoseTTAFold algorithm, which can directly predict the three-dimensional structure of the protein according to the amino acid sequence. After constructing the three-dimensional model of each domain, it is necessary to assemble them into a complete bispecific antibody structure. During the assembly process, the conformational modeling of the linker sequence is particularly important, and the linker needs to have sufficient length and flexibility to allow the two antigen binding domains to function independently. Preferably, the length of the linker is set to 15 to 25 amino acids, and the linker sequence uses a flexible sequence rich in glycine and serine, such as (GGGGS)3 or (GGGGS)4, which has high flexibility and is not easy to form regular secondary structures.

[0033] After the construction of the structural model, energy minimization processing is performed on the structure to eliminate possible atomic collisions and unreasonable geometric configurations in the structure. Energy minimization uses the conjugate gradient method or the steepest descent method, and during the calculation process, a full-atom force field such as AMBER ff14SB force field or CHARMM36 force field is used. The convergence criterion for energy minimization is set to the root mean square value of the energy gradient being less than 0.1 kilocalorie per mole per angstrom. The structure after energy minimization is used as the starting conformation for subsequent molecular dynamics simulation, ensuring that the simulation process can start from a reasonable initial state.

[0034] In the preprocessing stage, the three-dimensional structures of the two antigens recognized by the bispecific antibody also need to be prepared. The antigen structure can be directly obtained from the protein database, and if there is no experimental structure of the antigen in the database, a three-dimensional model of the antigen also needs to be constructed using homology modeling or ab initio prediction methods. The antigen structure also needs to be subjected to energy minimization processing to ensure the reasonableness of the structure.

[0035] In addition, in the data preprocessing stage, the preliminary binding mode of the bispecific antibody to the two antigens needs to be determined. By molecular docking algorithm, each antigen binding domain is docked with its corresponding antigen to obtain the preliminary antibody-antigen complex structure. Molecular docking can use rigid docking algorithms based on shape complementarity, such as ZDOCK or ClusPro, or semi-flexible docking algorithms that allow side chain flexibility, such as Rosetta Docking or HADDOCK. Docking calculation generates multiple possible binding modes, and the top-ranked binding mode is selected as the candidate according to the docking scoring function. For bispecific antibodies, it is necessary to ensure that the two antigen binding domains form stable binding modes with the respective antigens, and there is no obvious spatial conflict between the two binding interfaces.

[0036] Finally, the pre-processed bispecific antibody structure is checked for quality. The contents of the check include: whether there are abnormal bond lengths, bond angles or dihedral angles in the structure; whether the rotamers of the amino acid side chains are in a reasonable conformation; whether there are spatial collisions between atoms in the structure; whether the CDR loop conformation of the antibody is reasonable. Quality checks can use professional structure verification tools such as MolProbity or PROCHECK. Only structures that pass the quality check can enter the next step of molecular dynamics simulation.

[0037] Step two: antibody fragment interface flexibility adaptation evaluation, see Figure 2 This step evaluates the flexibility adaptation of the fragment interface of the bispecific antibody through molecular dynamics simulation, which is the first core innovation point of the present application. Traditional methods use rigid docking to evaluate fragment assembly, which cannot capture the dynamic conformation adjustment process between fragments, while the present application can accurately identify the dynamic behavior and potential spatial steric conflict of interface residues by long-time scale molecular dynamics simulation, which can truly reproduce the flexibility motion of fragments under physiological conditions.

[0038] First, the all-atom molecular dynamics simulation system of the bispecific antibody is established. The pre-processed bispecific antibody structure is placed in a water box under periodic boundary conditions, and the boundary distance of the water box from the antibody surface is at least 1.0 nanometers, ensuring that the movement of the antibody during simulation will not be affected by the periodic boundary. Water molecules are explicitly represented using TIP3P or SPC / E water models to accurately describe the influence of the hydration layer on the protein conformation. An appropriate amount of sodium ions and chloride ions are added to the system to achieve a physiological salt concentration of 0.15 moles per liter and to ensure the overall charge neutrality of the system. The initial position of the ions is determined by the calculation of the electrostatic potential energy distribution, and is preferentially placed near areas with high surface charge density of the protein.

[0039] In the molecular dynamics simulation, the all-atom force field, preferably AMBER ff14SB force field, is used for proteins, which has good description ability for the secondary structure and side chain conformation of proteins. For water molecules and ions, a parameter set matched with the protein force field is used. Before the simulation starts, the entire system is first energy minimized, using a two-step strategy of first constraining the protein backbone atoms to optimize only the solvent molecule positions, and then gradually releasing the constraints to optimize the entire system, to ensure that the solvent molecules reasonably surround the protein and eliminate possible atomic collisions. The convergence criterion for energy minimization is that the root mean square of energy gradient is less than 10.0 kilojoules per mole per nanometer.

[0040] After energy minimization, the system is equilibrated. The equilibration is divided into two stages: first, the temperature is raised in the canonical ensemble, and the system temperature is gradually raised from 0 Kelvin to 300 Kelvin, and the temperature raising process lasts for 100 picoseconds, and the temperature is controlled by the V-rescale temperature coupling method, to ensure that the kinetic energy distribution of each atom conforms to the Maxwell-Boltzmann distribution. During the temperature raising process, the position constraint is applied to the protein backbone atoms, and the constraint force constant is set to 1000 kilojoules per mole per nanometer square, to prevent the protein structure from undergoing drastic conformational changes during the temperature raising process. After the temperature reaches 300 Kelvin, switch to the isothermal-isobaric ensemble, and equilibrate the system density for 500 picoseconds, and the pressure is controlled at 1 atmosphere by the Parrinello-Rahman pressure coupling method, and gradually release the position constraint on the protein backbone, and finally completely remove the constraint, so that the protein can move freely.

[0041] After equilibration, the formal productive molecular dynamics simulation is carried out. The simulation time is set to not less than 100 nanoseconds, preferably 200 nanoseconds, to fully sample the conformational space of the antibody fragment. The simulation adopts the isothermal-isobaric ensemble, and the temperature is controlled at 300 Kelvin, and the pressure is controlled at 1 atmosphere, to ensure that the simulation conditions are close to the physiological state. The time step of the simulation is set to 2 femtoseconds, and the covalent bond length involving hydrogen atoms is constrained, and the LINCS algorithm or SHAKE algorithm is used, so as to allow the use of a larger time step to improve the calculation efficiency. The long-range part of the electrostatic interaction is calculated by the particle mesh Ewald method, the real space cutoff distance is set to 1.0 nanometer, and the reciprocal space grid spacing is set to 0.12 nanometer, to ensure the calculation accuracy of the electrostatic force. The van der Waals interaction is described by the Lennard-Jones potential function, and the cutoff distance is set to 1.0 nanometer. The simulation trajectory saves the coordinate and velocity information every 10 picoseconds, which is used for subsequent analysis.

[0042] During the molecular dynamics simulation, the interfacial region between the two antigen-binding domains of the bispecific antibody is monitored. The interfacial region is defined as the set of amino acid residues whose any atom has a distance less than 0.6 nanometer to any atom of the other domain. For each interfacial residue, the root-mean-square fluctuation is calculated, which reflects the degree of dynamic change of the residue position. The formula for calculating the root-mean-square fluctuation is:

[0043] ,

[0044] wherein: is the root-mean-square fluctuation of the i-th residue, with the unit of nanometer; is the total number of frames of the simulation trajectory; is the position vector of the i-th residue at time t, expressed by the atomic positions of the residue; is the average position vector of the i-th residue in the entire simulation trajectory, obtained by dividing the sum of positions at all times by the total number of frames; the square operation in the brackets represents the square of the modulus of the position vector difference, i.e. . By analyzing the distribution of the root-mean-square fluctuation of the interfacial residues, the flexible and rigid regions of the interface can be identified. Residues with larger root-mean-square fluctuation values are located in the flexible region, and these residues undergo greater positional changes during the simulation, indicating that their side chains or backbones have higher degrees of freedom. Residues with smaller root-mean-square fluctuation values are located in the rigid region, and these residues have relatively fixed positions, usually participating in the formation of stable hydrogen bond networks or hydrophobic cores. In the design of bispecific antibodies, the flexibility and rigidity of the interface need to be balanced: an overly rigid interface may result in insufficient conformational adaptability during fragment assembly, while an overly flexible interface may result in decreased structural stability. Preferably, the root-mean-square fluctuation values of the interfacial residues should be between 0.1 and 0.3 nanometers, indicating that the interface has both certain adaptability and necessary stability. In addition to the root-mean-square fluctuation analysis, the contact frequency between interfacial residues is also calculated. For any two residues on the interface, if the shortest distance between their heavy atoms is less than 0.5 nanometer, it is considered that the two residues form a contact. The contact frequency is defined as the percentage of the number of frames in which the two residues form a contact in the total number of frames in the entire simulation trajectory. Residues with high contact frequencies indicate that there are stable interactions between them, which may be hydrogen bonds, salt bridges, or hydrophobic interactions. By constructing a contact frequency matrix of interfacial residues, key interaction pairs of the interface can be identified. These key interaction pairs play an important role in maintaining the stability of the interface and should be preserved or strengthened during subsequent sequence optimization.

[0045]

[0046] In addition to the root-mean-square fluctuation analysis, the contact frequency between interfacial residues is also calculated. For any two residues on the interface, if the shortest distance between their heavy atoms is less than 0.5 nanometer, it is considered that the two residues form a contact. The contact frequency is defined as the percentage of the number of frames in which the two residues form a contact in the total number of frames in the entire simulation trajectory. Residues with high contact frequencies indicate that there are stable interactions between them, which may be hydrogen bonds, salt bridges, or hydrophobic interactions. By constructing a contact frequency matrix of interfacial residues, key interaction pairs of the interface can be identified. These key interaction pairs play an important role in maintaining the stability of the interface and should be preserved or strengthened during subsequent sequence optimization. ​​​​

[0047] Contact frequency matrix The element of is defined as:

[0048] ,

[0049] where: is the contact frequency between the th residue and the th residue, ranging from 0 to 1; is the total number of frames of the simulation trajectory; is the contact frequency between the th residue and the th residue at time , in nanometers; is the Heaviside step function, which is 1 when is less than 0.1 nanometers, and 0 otherwise.

[0050] In the interface flexibility adaptation evaluation, special attention is paid to the regions that may cause steric clash. Steric clash usually occurs at positions where two domains are too close, resulting in increased interatomic repulsive force. By analyzing the distribution of the shortest distance between atoms in the simulation trajectory, potential steric clash regions can be identified. Specifically, for each pair of atoms on the interface, the shortest distance between them in the entire simulation trajectory is calculated. If the shortest distance is less than the sum of the van der Waals radii of the two atoms minus 0.1 nanometers, it is considered that there is a potential steric clash between the pair of atoms. The residues to which all the atoms pairs with potential clashes belong are marked as clash residues, which need to be focused on in subsequent sequence optimization and may need to be mutated to reduce side chain volume or adjust side chain orientation.

[0051] To quantify the overall flexibility adaptation degree of the interface, the interface adaptation index is defined as follows:

[0052] ,

[0053] where: is the interface adaptation index, ranging from 0 to 1, and the larger the value, the better the interface adaptation; is the number of residues in the interface whose root mean square fluctuation value is between the optimal range of 0.1 to 0.3 nanometers; is the total number of residues in the interface; is the number of residues in the interface that have steric clashes. The first term reflects the rationality of the interface flexibility, and the second term ​​Reflects the degree of conflict-free interface. Preferably, the interface fitting index should be greater than 0.75, indicating that the interface has good flexibility and less spatial conflict.

[0054] Through the above analysis, this step can comprehensively evaluate the flexible fitting of the interface of the bispecific antibody fragment, identify key interface residues, flexible regions, rigid regions, and potential spatial steric conflict regions. These information provides clear guidance for subsequent sequence optimization, so that the optimization process can improve the adaptability of the interface, thereby improving the expression yield and structural stability of the bispecific antibody.

[0055] Step three: prediction of dual-target binding synergy, this step constructs a dual-target binding synergy prediction model, which is the second core innovation point of the present application. An important function of bispecific antibody is to simultaneously bind two different antigen targets, thereby realizing bridging or synergistic signal transduction. However, the simultaneous binding of dual targets is subject to various geometric and energy factors, including the relative position, relative orientation of the two antigens, the flexibility of the bispecific antibody, and the conformational adjustment energy cost when forming a ternary complex. This step systematically evaluates the feasibility and efficiency of the bispecific antibody in forming a ternary complex under different dual-antigen configurations by establishing a dual-antigen spatial geometry constraint model.

[0056] Referring to Figure 3 , the dual-target binding synergy prediction module first needs to determine the possible spatial distribution of the two antigens on the cell surface or in the tissue. For membrane protein antigens, their position and orientation on the cell membrane are constrained by the membrane topology; for secreted protein antigens, their spatial distribution is more flexible. In the implementation of the present application, it is assumed that the two antigens are fixed on a plane, which simulates the cell membrane or tissue interface, and the extracellular domains of the antigens are exposed outside the plane and can be recognized by the bispecific antibody. The relative position between the two antigens is described by the center distance and the relative orientation angle . The center distance is defined as the distance between the centroids of the extracellular domains of the two antigens, and the value range is usually between 10 and 30 nanometers, which covers the actual distribution distance of most cell surface receptor pairs. The relative orientation angle is defined as the angle between the directions of the main axes of the two antigens, and the value range is 0 to 180 degrees.

[0057] For a given dual-antigen relative position , ), it is necessary to evaluate whether the bispecific antibody can bind to both antigens simultaneously. The core of the evaluation is to calculate the free energy change of the bispecific antibody from the single antigen binding state to the double antigen simultaneous binding state. First, one antigen binding domain of the bispecific antibody is combined with the first antigen to form a binary complex, and the binding free energy of this process is , which is calculated by molecular docking and molecular mechanics Poisson-Boltzmann surface area method. In the binary complex, the other antigen binding domain of the bispecific antibody is in an unbound state and can move freely. Next, an attempt is made to combine the second antigen binding domain of the bispecific antibody with the second antigen to form a ternary complex. This process requires the hinge region and linker region of the bispecific antibody to undergo conformational adjustment to adapt to the spatial position constraints of the two antigens. The energy cost of conformational adjustment is calculated by constrained molecular dynamics simulation, in which position constraints are applied to the first antigen binding domain and the first antigen, while the second antigen binding domain is gradually guided to approach the binding site of the second antigen. The potential energy change accumulated in this process is the energy cost of conformational adjustment, denoted as .

[0058] The binding free energy of the second antigen binding domain to the second antigen is , also calculated by molecular docking and molecular mechanics Poisson-Boltzmann surface area method. In the ternary complex, the two antigen binding domains are bound to their respective antigens, but the backbone of the bispecific antibody is doubly constrained by the two binding interfaces, losing part of the conformational entropy. This loss of conformational entropy can be estimated by analyzing the difference in conformational distribution of the bispecific antibody in the ternary complex and in the free state. Specifically, molecular dynamics simulations are performed on the ternary complex and the free state of the bispecific antibody, respectively, to calculate the volume of the eigenvector space obtained by principal component analysis. The ratio of the two reflects the change in conformational entropy, and the loss of conformational entropy is denoted as , where is the temperature, with a value of 300 Kelvin, is the change in conformational entropy, with a unit of joules per mole per Kelvin.

[0059] The total free energy change of the bispecific antibody forming the ternary complex is:

[0060] ,

[0061] where: is the total free energy change of forming the ternary complex, with a unit of kilojoules per mole, and a more negative value indicates that the ternary complex is more stable; is the binding free energy of the first antigen binding domain to the first antigen, with a unit of kilojoules per mole, and is usually negative; This represents the binding free energy between the second antigen-binding domain and the second antigen, expressed in kilojoules per mole, and is usually a negative value. The energy cost of conformational adjustment, measured in kilojoules per mole, is usually a positive value and reflects the need for the bispecific antibody backbone to bend or twist to accommodate the spatial position of the two antigens. The value is for temperature, and is set to 300 Kelvin. It represents the conformational entropy change, measured in joules per mole per Kelvin, and is usually a negative value, reflecting the reduction in conformational freedom of the bispecific antibody in the ternary complex. The term represents the contribution of entropy change to free energy, which is usually positive and indicates the adverse effect of entropy reduction on binding.

[0062] To evaluate the synergy of dual-target binding, a synergy coefficient is defined. as follows:

[0063] ,

[0064] in: The coefficient of coordination is dimensionless. The change in total free energy for the formation of a ternary complex; and For two separate binding free energies. If This indicates that the simultaneous binding of two targets exhibits positive synergy, meaning the stability of the ternary complex exceeds the simple summation of two independently bound targets. This positive synergy may originate from conformational synergy between the two antigen-binding domains or conformational optimization of the antibody backbone under dual constraints. This indicates that the simultaneous binding of two targets has negative synergy, meaning that the stability of the ternary complex is lower than expected due to the energy cost of conformational adjustment and the loss of conformational entropy. In this case, it may be necessary to optimize the joint length or the flexibility of the hinge region to reduce the negative synergy.

[0065] In practical applications, the relative positions of the two antigens ( , The free energy (FE) is not fixed but varies within a certain range. Therefore, it is necessary to evaluate multiple possible dual antigen configurations and construct a free energy map of the dual antigen configuration space. Specifically, at the center distance... From 10 nanometers to 30 nanometers, relative orientation angle The dual antigen configuration was sampled at intervals of 2 nanometers and 10 degrees within a range from 0 to 180 degrees. For each sampling point, calculations were performed. and Through this systematic sampling, the biantigen configuration regions where bispecific antibodies most readily form ternary complexes can be identified, i.e., on the free energy map. The most negative regions. These regions correspond to the optimal configurations for the bispecific antibody to function, and the structural parameters (such as linker length, hinge flexibility) of the bispecific antibody should be designed to adapt to these optimal configurations.

[0066] To further quantify the efficiency of dual-target binding, the ternary complex formation probability is defined as As follows:

[0067] ,

[0068] Wherein: The ternary complex formation probability, the value range is 0 to 1; The total free energy of the ternary complex; The gas constant, the value is 8.314 joules per mole per kelvin; The temperature, the value is 300 kelvin; the summation term in the denominator traverses all possible binding states, including the state of the bispecific antibody not binding any antigen, the state of single binding of the first antigen, the state of single binding of the second antigen, and the state of the ternary complex of simultaneous binding of the two antigens; The free energy of the corresponding state, for the unbound state, the free energy is 0. The formation probability reflects the proportion of the bispecific antibody in the ternary complex state under the condition of thermodynamic equilibrium, and the larger the value is, the higher the efficiency of simultaneous binding of the dual targets is.

[0069] Through the above-mentioned dual-target binding synergy prediction, the present application can accurately evaluate the binding performance of the bispecific antibody under different bisantigen configurations, identify the optimal bisantigen configuration region, and quantify the synergistic effect and formation probability of dual-target binding. These information is crucial for the design of bispecific antibodies, especially in the selection of linker length, optimization of hinge flexibility, and determination of the spatial arrangement of antigen binding domains, providing clear quantitative guidance.

[0070] Step four: Dynamic evaluation of glycosylation sites, this step is the third core innovation of the present application, which dynamically evaluates the potential glycosylation sites in the sequence of the bispecific antibody, and predicts the influence of glycosylation modification on the spatial structure and function of the antibody. Glycosylation is an important post-translational modification of antibodies, which has a significant impact on the stability, solubility, immunogenicity and interaction with Fc receptors of antibodies. In bispecific antibodies, the spatial position and conformation of the sugar chain may have an asymmetric effect on the two antigen binding domains, especially when the glycosylation site is close to the antigen binding interface, the sugar chain may produce a steric hindrance effect, affecting the binding affinity of the antigen. The traditional method is too simple to deal with glycosylation sites, which only makes a static identification according to the amino acid sequence motif (such as Asn-X-Ser / Thr, where X is any amino acid except proline), without considering the actual influence of the sugar chain on the protein structure in dynamic conformation. The present application realizes the accurate prediction of the dynamic influence of glycosylation modification by coarse-grained sugar chain model combined with molecular dynamics simulation.

[0071] Referring to Figure 4 , the glycosylation site dynamic evaluation module first identifies the potential N-glycosylation sites in the amino acid sequence of the bispecific antibody. The sequence motif of N-glycosylation site is Asn-X-Ser / Thr, where X can be any amino acid except proline. By sequence scanning, all asparagine residues that meet this motif are identified, which are potential N-glycosylation sites. In addition to the sequence motif, the actual modification of the glycosylation site is also affected by the accessibility of the three-dimensional structure, only the asparagine residues exposed on the surface of the protein can be modified by glycosyltransferase. Therefore, by calculating the solvent accessible surface area of the potential glycosylation site, the sites with solvent accessible surface area greater than 20 square angstroms are selected as the possible modified glycosylation sites.

[0072] For each potential glycosylation site, different types of N-glycan structures are added. N-glycan can be divided into high mannose type, complex type and hybrid type according to its composition and branching. In bispecific antibodies, the most common type of sugar chain is complex type double antenna sugar chain, whose core structure includes two N-acetylglucosamine, three mannose and several peripheral sugar residues. In order to simplify the calculation complexity, a coarse-grained sugar chain model is used to represent the sugar chain. The coarse-grained model represents each monosaccharide residue with one or several beads, and the beads are connected by springs to simulate the flexibility of glycosidic bond. The parameters of the coarse-grained model are derived from the statistical analysis of the simulation data of the full-atom sugar chain, which ensures that the coarse-grained model can reproduce the main conformational characteristics of the full-atom model. Preferably, the MARTINI coarse-grained force field is used to describe the sugar chain, which has been widely verified to accurately simulate the dynamic behavior of the sugar chain.

[0073] A coarse-grained molecular dynamics simulation system is established for the bispecific antibody structure with sugar chains in a solvent environment. The proteins in the system are represented by coarse-grained protein models, where each amino acid residue is represented by several beads. The water molecules are also represented by a coarse-grained water model, where each water molecule is represented by one bead. Coarse-grained simulations can significantly improve computational efficiency compared to all-atom simulations, allowing for longer time scales and larger system sizes to be simulated, which is particularly important for capturing the slow conformational motions of sugar chains. The time step for the coarse-grained molecular dynamics simulation is set to 20 femtoseconds, and the total simulation time is set to 1 microsecond, which is sufficient to observe the complete conformational transition process of the sugar chains.

[0074] During the simulation, the influence of the sugar chains on the overall conformation of the bispecific antibody is monitored. First, the spatial volume occupied by the sugar chains is calculated. The spatial volume occupied by the sugar chains can be estimated by calculating the convex hull volume of the sugar chain beads. The convex hull is the smallest convex polyhedron that encloses all the sugar chain beads. The volume occupied by the sugar chains is calculated using the Qhull algorithm, which can efficiently calculate the convex hull of a set of points in three-dimensional space. The volume occupied by the sugar chains reflects the degree of stretching of the sugar chains in space, with a larger volume indicating that the sugar chains are more stretched and can have a greater steric hindrance effect on the surrounding structures. Second, the contact area between the sugar chains and the protein surface is calculated. The contact area is defined as the sum of the surface areas of the bead pairs between the sugar chain beads and the protein beads with a distance less than 0.6 nanometers. The contact area is calculated by iterating through all pairs of sugar chain beads and protein beads, and accumulating the contributions of the bead pairs that satisfy the distance condition. The contribution of each bead pair is calculated as a certain proportion of the solvent-accessible surface area of the beads. The contact area reflects the interaction strength between the sugar chains and the protein, with a larger contact area indicating stronger interaction between the sugar chains and the protein, and the influence of the sugar chains on the protein conformation being more significant.

[0075] Third, the influence of the sugar chains on the overall size of the antibody is calculated. The overall size of the antibody can be characterized by the radius of gyration, which is defined as:

[0076] where Rg is the radius of gyration in nanometers, reflecting the overall compactness of the molecule; N is the total number of beads in the system, including protein beads and sugar chain beads; mi is the mass of the i-th bead; ri is the position vector of the i-th bead; and R is the center of mass position vector of the system, calculated as:

[0077]

[0078] ​​​​​​​​​​​The calculation results are as follows: For the first The square of the distance from the center of mass to the first bead.

[0079] By comparing the radius of gyration of the bispecific antibody with and without sugar chains, the impact of sugar chains on the overall size of the antibody can be quantified. Generally, the addition of sugar chains will increase the radius of gyration, indicating that the overall size of the antibody becomes larger. The increase in the radius of gyration can affect the hydrodynamic properties of the antibody, such as the diffusion coefficient and the sedimentation coefficient, which in turn affect the distribution and clearance rate of the antibody in the body.

[0080] It is particularly important to evaluate the degree of spatial interference of sugar chains on the antigen binding domain. If the glycosylation site is located near the antigen binding domain, the sugar chain may extend to the antigen binding interface and collide with the antigen or the complementarity determining region of the antibody, reducing the binding affinity of the antigen. By analyzing the simulation trajectory, the distribution of the shortest distance between the sugar chain bead and the antigen binding interface residues is calculated. If the shortest distance is often less than 0.4 nanometers, it indicates that there is significant spatial interference between the sugar chain and the antigen binding interface. Define the sugar chain interference index As follows:

[0081] ,

[0082] Where: is the sugar chain interference index, ranging from 0 to 1, the larger the value, the more severe the spatial interference of the sugar chain on the antigen binding interface; is the total number of frames of the simulation trajectory; is the shortest distance between the sugar chain bead and the antigen binding interface residues at time , in nanometers; is the Heaviside step function, when , otherwise This function is used to determine whether the shortest distance is less than the interference threshold of 0.4 nanometers.

[0083] Preferably, the sugar chain interference index should be less than 0.2, indicating that the spatial interference of the sugar chain on the antigen binding interface is light and will not significantly affect the antigen binding. If the sugar chain interference index is greater than 0.5, it is necessary to consider engineering modification of the glycosylation site, such as eliminating the glycosylation motif by mutation, or changing the spatial distribution of the sugar chain by introducing a new glycosylation site to reduce the interference on the antigen binding.

[0084] In addition to steric interference, sugar chains can also affect antigen binding by influencing the local charge distribution and hydrophobicity of the antibody. Sugar chains often carry negative charges (such as sialic acid residues), which can interact electrostatically with charged residues on the antigen or antibody, affecting binding affinity. The effect of sugar chains on the local electric field can be assessed by calculating the electrostatic potential distribution around the sugar chain. The electrostatic potential is calculated by solving the Poisson-Boltzmann equation, taking into account the charge distribution of proteins, sugar chains and ions, and the dielectric shielding effect of the solvent. If the sugar chain causes a significant change in the electrostatic potential near the antigen binding interface, the distribution of charged residues in the antibody sequence may need to be adjusted to compensate for the electrostatic effect of the sugar chain.

[0085] Through the above dynamic evaluation of glycosylation sites, the present application can accurately predict the multifaceted effects of sugar chain modification on bispecific antibodies, including the spatial volume occupied by sugar chains, contact area with proteins, effect on overall size of antibodies, and spatial interference with antigen binding interfaces. These information provides quantitative basis for the engineering design of glycosylation sites, so that the pharmacokinetic properties of bispecific antibodies can be improved by optimizing the position of glycosylation sites and the type of sugar chains, while avoiding the adverse effects of sugar chains on antigen binding function.

[0086] Step five: sequence rational design optimization, see Figure 5 Based on the detailed information obtained from the previous three evaluation steps, this step rationally designs and optimizes the sequence of bispecific antibodies. The optimization goals include improving the adaptability of the fragment interface, improving the synergy of dual target binding, optimizing the distribution of glycosylation sites, and comprehensively improving the expression yield, thermal stability, binding affinity and pharmacokinetic properties of bispecific antibodies. Sequence optimization uses a multi-objective optimization strategy, considering the trade-off between multiple design goals, and gradually improving the antibody sequence through an iterative optimization process.

[0087] Among them, the interface residue mutation is screened by virtual mutation, the flexible residues with root mean square fluctuation greater than 0.3 nanometers are mutated to glycine, alanine or serine, the residues with steric hindrance conflict are mutated to amino acids with smaller side chain volume than the original residues, and the residues with contact frequency greater than 50% are strengthened by introducing hydrophobic interactions or hydrogen bond interactions with binding free energy reduction of more than 2 kilocalories per mole.

[0088] First, the optimization of the interface of the fragment. According to the results of the interface flexibility adaptation evaluation in step two, for the bispecific antibody with a low interface adaptation index, mutations need to be made to the interface residues to improve the adaptation. The optimization strategies include: for flexible residues with a large root mean square fluctuation, they can be mutated to smaller and less flexible amino acids such as glycine or alanine to reduce the conformational search space of the side chain and improve the stability of the interface; for residues with steric hindrance conflicts, they can be mutated to smaller amino acids such as valine to alanine or phenylalanine to leucine to reduce the side chain volume and relieve the steric conflict; for high-contact-frequency residue pairs, if they form a hydrophobic interaction, stronger hydrophobic amino acids (such as leucine, isoleucine, phenylalanine) can be introduced to strengthen the interaction and increase the binding energy of the interface; if they form a hydrogen bond or a salt bridge, complementary hydrogen bond donors and acceptors (such as asparagine-aspartic acid, glutamine-glutamic acid) or oppositely charged amino acids (such as lysine-aspartic acid, arginine-glutamic acid) can be introduced to strengthen these interactions.

[0089] The mutation of interface residues needs to follow certain principles to avoid introducing adverse side effects. The mutation should not destroy the overall structure and function of the interface, should not increase the immunogenicity of the antibody, and should not affect the expression and folding of the antibody. In order to systematically evaluate the effect of the mutation, a virtual mutation simulation is performed for each candidate mutation, i.e. the structure of the mutated antibody is constructed in the computer, and the effect of the mutation on the structure and stability is evaluated through energy calculation and short-time molecular dynamics simulation. The evaluation indicators of virtual mutation include the change of interface binding energy before and after mutation, the local structural change around the mutated residue, and the effect of mutation on the interface adaptation index. Only those mutations that can significantly improve the interface adaptation without introducing adverse side effects will be included in the final optimization scheme.

[0090] Second, the optimization of the linker region. The linker connects two antigen binding domains, and its length and flexibility have an important influence on the synergy of dual-target binding. According to the results of dual-target binding synergy prediction in step three, if the synergy coefficient is significantly greater than 1, indicating that there is a large conformational adjustment energy cost or conformational entropy loss in the simultaneous binding of the dual targets, the linker needs to be optimized to improve its flexibility or increase its length. The optimization of the linker length is achieved by adding or subtracting amino acids in the linker sequence, usually one repeat unit at a time, such as (GGGGS). The increase in linker length can give the two antigen binding domains greater spatial freedom, reducing their conformational constraints when simultaneously binding to two antigens, thereby reducing the energy cost of conformational adjustment. However, the increase in linker length may also lead to the loosening of the overall structure of the antibody, reducing the thermal stability, so a balance needs to be found between synergy and stability.

[0091] The optimization of linker flexibility is achieved by adjusting the amino acid composition of the linker sequence. Sequences rich in glycine and serine are highly flexible, while sequences rich in proline are relatively rigid. By introducing or removing proline in the linker, the flexibility of the linker can be adjusted. For cases where increased flexibility is desired, the proline in the linker can be replaced with glycine or alanine; for cases where decreased flexibility is desired to improve stability, proline can be introduced into the linker. The optimization of the linker sequence also needs to be evaluated by virtual mutation and molecular dynamics simulation to ensure that the optimized linker can not only meet the geometric requirements of dual-target binding but also maintain the overall stability of the antibody.

[0092] Again, optimization of glycosylation sites. According to the results of dynamic evaluation of glycosylation sites in step four, if the sugar chain of a certain glycosylation site significantly interferes with the antigen binding interface (the sugar chain interference index is greater than 0.5), the site needs to be engineered. The modification strategies include: eliminating the glycosylation site, by mutating the asparagine in the Asn-X-Ser / Thr motif to other amino acids (such as glutamine, which has a similar side chain structure to asparagine but cannot be glycosylated), or mutating Ser / Thr to other amino acids (such as alanine), thereby preventing the addition of sugar chains; if eliminating the glycosylation site may affect other properties of the antibody (such as stability or solubility), it can be considered to move the glycosylation site away from the antigen binding interface, by introducing an Asn-X-Ser / Thr motif in a surface residue far from the antigen binding interface, to create a new glycosylation site, while eliminating the original interfering glycosylation site; it is also possible to design shorter or less branched sugar chain structures through sugar chain engineering techniques to reduce the spatial occupancy volume of the sugar chain, thereby reducing the interference with antigen binding, which can be achieved by manipulating the expression of glycosyltransferases in host cells.

[0093] The optimization of glycosylation sites also needs to consider the impact on the pharmacokinetic properties of the antibody. The sialic acid residues on the sugar chain can prolong the circulating half-life of the antibody in the body, so when optimizing the glycosylation site, it should be possible to retain or increase the glycosylation site that can improve the level of sialylation. At the same time, the branching and composition of the sugar chain also affect the binding of the antibody to the Fc receptor, and thus affect antibody-dependent cell-mediated cytotoxicity (ADCC) and complement-dependent cytotoxicity (CDC). For bispecific antibodies that need to enhance ADCC effect, the proportion of fucose-deficient sugar chains can be increased by optimizing the glycosylation sites in the Fc region, thereby increasing the binding affinity of the antibody to the FcγRIIIa receptor.

[0094] In the sequence optimization process, an iterative optimization strategy is adopted. Each iteration includes: proposing a set of candidate mutations based on the current evaluation results; performing virtual evaluation on each candidate mutation; selecting the mutation with the best evaluation results to apply to the antibody sequence; re-evaluating the optimized antibody for steps two to four; and proposing the next set of candidate mutations based on the new evaluation results. The iterative optimization process continues until the performance indicators of the antibody (interface adaptation index, synergy coefficient, sugar chain interference index, etc.) reach the preset optimization target, or the performance indicators no longer improve significantly in consecutive iterations.

[0095] To ensure the globality of the optimization process and avoid falling into local optimum, a multi-start optimization strategy is adopted. Iterative optimization is started from different initial mutation schemes, and the results obtained from different optimization paths are finally compared to select the scheme with the best performance indicators as the final optimized sequence. The multi-start strategy can effectively explore different regions of the sequence space and improve the probability of finding the global optimal solution.

[0096] Step six: comprehensive performance evaluation. After completing the sequence optimization, the optimized bispecific antibody is subjected to comprehensive performance evaluation to verify the effect of design and optimization. Comprehensive performance evaluation includes three aspects: thermal stability evaluation, expression yield prediction, and pharmacokinetic parameter prediction.

[0097] Thermal stability evaluation is achieved by calculating the conformational stability of the antibody at different temperatures. Constant temperature molecular dynamics simulation is used to simulate the optimized antibody at a series of temperature points (such as 300 Kelvin, 310 Kelvin, 320 Kelvin, 330 Kelvin, 340 Kelvin), with a simulation time of 50 nanoseconds at each temperature point. During the simulation, the root mean square deviation of the antibody structure is monitored, and an increase in the root mean square deviation indicates that the antibody structure has undergone significant conformational changes, possibly indicating partial unfolding of the protein. By analyzing the trend of root mean square deviation with temperature, the melting temperature of the antibody, i.e. the temperature at which the antibody structure begins to significantly unfold, can be estimated. The higher the melting temperature, the better the thermal stability of the antibody. Preferably, the melting temperature of the bispecific antibody should be higher than 343 Kelvin (70 degrees Celsius), ensuring the stability of the antibody during production, storage and use.

[0098] Expression yield prediction is achieved by evaluating the antibody's foldability and solubility in host cells. The expression yield of an antibody is influenced by various factors, including the hydrophobicity, charge distribution, secondary structure propensity, aggregation propensity, etc. of the amino acid sequence. By calculating these physicochemical properties of the optimized antibody sequence, its expression behavior in host cells can be predicted. Specifically, the hydrophobicity distribution of the antibody surface is calculated, and overexposure of hydrophobic residues on the surface can lead to antibody aggregation, reducing expression yield. The mean value of the surface residue hydrophobicity is calculated by scoring the hydrophobicity of each amino acid in the sequence using the Kyte-Doolittle hydrophobicity index, which should be as low as possible, preferably less than 0.5. The isoelectric point of the antibody sequence is calculated, and an excessively high or low isoelectric point can affect the solubility of the antibody. Preferably, the isoelectric point of the antibody should be between 6 and 9, close to the physiological pH value, ensuring good solubility of the antibody under physiological conditions.

[0099] The aggregation propensity of the antibody is evaluated by calculating the area of the hydrophobic region exposed on the antibody surface. The larger the area of the hydrophobic region, the more likely the antibody is to form aggregates through hydrophobic interactions. The aggregation propensity score of the optimized antibody sequence is calculated using sequence-based aggregation prediction algorithms such as AGGRESCAN or TANGO, and a lower score indicates a smaller aggregation propensity and a higher expected expression yield. Based on the above multiple indicators, an expression yield prediction model is constructed, which is trained based on a large amount of antibody data with known expression yields, and can predict the relative expression yield of the antibody based on its physicochemical properties.

[0100] Pharmacokinetic parameter prediction is achieved by evaluating the clearance rate and distribution volume of the antibody. The clearance of the antibody is mainly through kidney filtration, liver metabolism, and intracellular endocytosis degradation, etc. The size, charge, and degree of glycosylation of the antibody all affect the clearance rate. By calculating the hydrodynamic radius and surface charge distribution of the optimized antibody, its clearance behavior in vivo can be predicted. The hydrodynamic radius is calculated by the rotation radius and the shape factor, which reflects the degree of deviation from the spherical shape of the antibody. The surface charge is obtained by calculating the distribution of the surface electrostatic potential of the antibody, and high charge density regions can increase the non-specific binding of the antibody to serum proteins, accelerating clearance.

[0101] The degree of glycosylation has an important influence on the pharmacokinetics of the antibody. Sialylated sugar chains can interact with sialic acid binding receptors on the cell surface, prolonging the circulating half-life of the antibody. Therefore, in the pharmacokinetic parameter prediction, the glycosylation site distribution and sugar chain type of the optimized antibody need to be considered. By parameters such as the number of glycosylation sites, the branching degree of the sugar chain, and the sialic acid content, a pharmacokinetic prediction model is constructed, which is trained based on antibodies with known pharmacokinetic data, and can predict the clearance half-life and distribution volume of the optimized antibody.

[0102] The results of the comprehensive performance evaluation serve as the final verification of the optimization scheme. Only when the thermal stability, expression yield and pharmacokinetic parameters all meet the preset standards, the optimized bispecific antibody sequence is considered to be a successful design scheme and can enter the experimental verification stage.

[0103] The computer device includes a memory and a processor. The memory is used to store a computer program containing all instruction codes for implementing the method. When the processor executes the computer program, all steps of the above-mentioned bispecific antibody molecule design and optimization method based on computational simulation assistance are implemented.

[0104] In specific implementations, the computer device can be a high-performance computing server, a workstation or a personal computer configured with a graphics processor acceleration card. The memory can include random access memory, read-only memory, hard disk drive, solid state drive and other storage media. The processor can be a central processing unit, a graphics processing unit or a combination of the two. For computationally intensive tasks such as molecular dynamics simulation, it is preferred to use a graphics processor for acceleration, which can increase the computing speed by 10 to 100 times.

[0105] The computer program adopts a modular design, including data preprocessing module, molecular dynamics simulation module, interface flexibility evaluation module, cooperativity prediction module, glycosylation evaluation module, sequence optimization module and performance evaluation module, etc. The modules communicate with each other through standardized data interfaces, ensuring seamless data transfer between different modules. The computer program also includes a user interaction interface, which allows users to input initial antibody sequences and structure information, set simulation parameters, view evaluation results, and manually review and adjust the optimization scheme.

[0106] In addition, the computer device can also be connected to a cloud computing platform to use distributed computing resources for large-scale virtual mutation screening and parameter scanning. By distributing the computing tasks to multiple computing nodes for parallel execution, the time of the entire design and optimization process can be significantly shortened from weeks to days or even hours.

[0107] Through the above detailed implementation, those skilled in the art can fully understand the technical solutions of the present application, and realize the computer-aided design and optimization of bispecific antibodies according to the method and system provided by the present application, thereby improving the success rate and efficiency of bispecific antibody development.

[0108] The above-described embodiments only express the specific implementation of the present application, which is described in more detail and in more detail, but it cannot be understood as a limitation on the scope of the patent of the present application. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application.

Claims

1. A method for the design and optimization of bispecific antibody molecules based on computational simulation assistance, characterized in that, The method comprises the following steps: Obtaining the initial sequence and structure information of the target bispecific antibody, including the amino acid sequences of the two antigen binding domains, the linker sequence, and the Fc region sequence; Evaluating the flexibility adaptation of the antibody fragment interface based on molecular dynamics simulation, identifying key interface residues and potential spatial steric conflict areas during fragment assembly; Constructing a dual-target binding synergy prediction model, analyzing the spatial geometric constraints of the two antigens, and evaluating the conformational feasibility and binding efficiency of the bispecific antibody in simultaneously binding to the two antigens; Evaluating the dynamic conformation of potential glycosylation sites in the antibody sequence, predicting the influence of sugar chain modification on the spatial structure and function of the antibody; Rationally designing and optimizing the antibody sequence based on the evaluation results, including interface residue mutation, linker length adjustment, linker flexibility adjustment, and glycosylation site engineering; Evaluating the comprehensive performance of the optimized antibody sequence, including thermal stability, expression yield prediction, and pharmacokinetic parameter prediction.

2. The method of claim 1, wherein, The molecular dynamics simulation uses an all-atom force field model, and the simulation time is not less than 100 nanoseconds to fully capture the conformational dynamic behavior of the antibody fragments under physiological conditions; The flexibility adaptation evaluation is based on the root mean square fluctuation analysis of the side chains of the interface amino acids and the interface contact frequency statistics, identifying the spatial distribution pattern of the interface flexibility and rigidity regions.

3. The method of claim 1, wherein, The dual-target binding synergy prediction is achieved by establishing a spatial position constraint matrix of the two antigens, which comprehensively considers the relative position, relative orientation of the two antigens, and the flexibility of the hinge region of the bispecific antibody, quantitatively evaluating the energy cost of the bispecific antibody forming a ternary complex under different dual-antigen configurations; the synergy evaluation includes calculating the free energy change of simultaneous binding of the two antigens relative to single antigen binding.

4. The method of claim 1, wherein, The dynamic evaluation of glycosylation sites uses a coarse-grained sugar chain model, adds different types of sugar chain structures to each potential N-glycosylation site, observes the dynamic disturbance of sugar chains to the antibody backbone conformation through molecular dynamics simulation, and quantifies the degree of spatial interference between sugar chains and antigen binding domains; the evaluation includes calculating the spatial volume occupied by sugar chains, the contact area between sugar chains and protein surfaces, and the influence of sugar chains on the overall rotation radius of the antibody.

5. The method of claim 1, wherein, The interface residue mutation is screened by virtual mutation, where flexible residues with a root mean square fluctuation greater than 0.3 nanometers are mutated to glycine, alanine, or serine, residues with spatial steric conflict are mutated to amino acids with smaller side chain volumes than the original residues, and residues with a contact frequency greater than 50% are strengthened by introducing hydrophobic interactions or hydrogen bond interactions that reduce the binding free energy by more than 2 kilocalories per mole.

6. The method of claim 1, wherein, The linker length adjustment is achieved by increasing or decreasing the number of amino acid repeat units in the linker sequence, and the increase in linker length can reduce the conformational adjustment energy cost when the two targets are simultaneously bound; the linker flexibility adjustment is achieved by adjusting the amino acid composition of the linker sequence, and sequences rich in glycine and serine have high flexibility, while the introduction of proline reduces flexibility.

7. The method of claim 1, wherein, The glycosylation site engineering includes eliminating glycosylation sites that produce spatial interference for antigen binding, preventing glycan addition by mutating amino acids in the Asn-X-Ser / Thr motif; or moving the glycosylation site to a position away from the antigen binding interface, and achieving this by introducing a new glycosylation motif at the surface residue.

8. The method of claim 1, wherein, The rational design and optimization adopts a multi-objective optimization strategy, and comprehensively considers interface adaptability, dual-target point synergy, glycosylation influence, thermal stability, and expression yield as optimization objectives. The antibody sequence is gradually improved through an iterative optimization process, and a multi-start optimization strategy is used to explore different regions of the sequence space, thereby improving the probability of finding a global optimal solution.

9. The method of claim 1, wherein, In the comprehensive performance evaluation, the thermal stability evaluation is performed by molecular dynamics simulation at different temperatures to calculate the melting temperature of the antibody; the expression yield prediction is achieved by evaluating the hydrophobicity distribution, isoelectric point, and aggregation tendency of the antibody sequence; The pharmacokinetic parameter prediction is achieved by calculating the hydrodynamic radius, surface charge distribution, and glycosylation degree of the antibody.

Citation Information

Patent Citations

  • A shipping mechanism and sales equipment

    CN109993883B

  • Bispecific antibody platform

    CN108601830A

  • Multi-chain multi-targeting bispecific antigen binding molecules with increased selectivity

    CN119137147A