Model training method, nano antibody inverse folding method and electronic equipment

By combining molecular dynamics simulations and machine learning models, the problem of insufficient nanobody structure data was solved, the predictive performance of the sequence prediction model was improved, and efficient prediction of nanobody structures was achieved.

CN121483385APending Publication Date: 2026-02-06HUASHEN INTELLIGENT MEDICINE BIOTECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511648761.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Due to insufficient antibody structure data for nanobodies, existing machine learning models perform poorly in nanobodies refolding tasks, and the complex dynamics of the complementarity determination region of nanobodies affects the training process of sequence prediction models.

Method used

By performing molecular dynamics simulations on the original antibody structure samples, enhanced antibody structure samples are extracted, an enhanced dataset is constructed, and these samples are processed using machine learning models to generate predicted amino acid sequences. The sequence prediction model is then trained using a loss function, including specific processing of the CDR region and data augmentation techniques.

Benefits of technology

The predictive performance of the sequence prediction model has been improved, enhancing the accuracy and stability of nanobody structure prediction and overcoming the challenges of insufficient data and kinetic complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483385A_ABST
    Figure CN121483385A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a nanometer antibody inverse folding method and electronic equipment. The model training method comprises the following steps: performing molecular dynamics simulation on an original antibody structure sample to obtain a molecular dynamics track of the original antibody structure sample; extracting a plurality of enhanced antibody structure samples from the molecular dynamics trajectory so as to form an enhanced data set; processing each enhanced antibody structure sample in the enhanced data set by using a machine learning model to obtain a predicted amino acid sequence of each enhanced antibody structure sample; determining a loss function according to the predicted amino acid sequence of each enhanced antibody structure sample; and training the machine learning model by using the loss function to obtain a sequence prediction model. The trained sequence prediction model is used for carrying out nano antibody inverse folding treatment, and a more accurate prediction result can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computational biology, and in particular, to a model training method, a nanobody inverse folding method, and an electronic device. BACKGROUND

[0002] Nanobody is the smallest antibody known to bind to target antigens. Compared with traditional antibodies, nanobody has the advantages of small size and strong penetration, and can be applied to the fields of tumor targeted therapy and diagnostic reagent development.

[0003] Nanobody inverse folding aims to predict the amino acid sequence of nanobody according to the structure of nanobody. At present, machine learning models can predict the amino acid sequence of proteins according to protein structures, and such machine learning models are also applied to sequence prediction tasks of traditional antibodies. SUMMARY

[0004] The inventors have noticed that when a sequence prediction model is used to realize nanobody inverse folding, the antibody structure data of nanobody is very limited, which cannot meet the data volume requirement of machine learning model training, resulting in poor training effect of the sequence prediction model.

[0005] Accordingly, the present disclosure provides a model training method capable of improving the prediction performance of a sequence prediction model.

[0006] According to a first aspect of the embodiments of the present disclosure, a model training method is provided, including: performing molecular dynamics simulation on an original antibody structure sample to obtain a molecular dynamics trajectory of the original antibody structure sample; extracting a plurality of enhanced antibody structure samples from the molecular dynamics trajectory to constitute an enhanced data set; processing each enhanced antibody structure sample in the enhanced data set by using a machine learning model to obtain a predicted amino acid sequence of the each enhanced antibody structure sample; determining a loss function according to the predicted amino acid sequence of the each enhanced antibody structure sample; and training the machine learning model by using the loss function to obtain a sequence prediction model.

[0007] In some embodiments, the determining of the loss function includes: determining a first loss function according to the predicted amino acid sequence of the each enhanced antibody structure sample and an amino acid sequence sample corresponding to the each enhanced antibody structure sample; processing the predicted amino acid sequence of the each enhanced antibody structure sample by using a structure prediction model to obtain a predicted antibody structure corresponding to the each enhanced antibody structure sample; determining a second loss function according to the each enhanced antibody structure sample and the predicted antibody structure corresponding to the each enhanced antibody structure sample; and determining the loss function according to the first loss function and the second loss function.

[0008] In some embodiments, the determining the first loss function comprises: obtaining the first loss function according to a probability that a k-th amino acid in the predicted amino acid sequence is the k-th amino acid in the amino acid sequence sample, and a weight of the k-th amino acid, wherein K is a total number of amino acid sites of the enhanced antibody structure sample.

[0009] In some embodiments, the obtaining the first loss function comprises: calculating a product of a logarithm value of the probability and the weight to obtain a probability weighted value of the k-th amino acid; and obtaining the first loss function according to a sum of the probability weighted values of the K amino acid sites.

[0010] In some embodiments, if the k-th amino acid belongs to a complementarity determining region, the weight of the k-th amino acid is a first weight value; if the k-th amino acid belongs to a framework region, the weight of the k-th amino acid is a second weight value, wherein the first weight value is greater than the second weight value.

[0011] In some embodiments, the first weight value is 0.8; and the second weight value is 0.2.

[0012] In some embodiments, the determining the second loss function comprises: obtaining the second loss function according to a coordinate of a k-th amino acid in each of the enhanced antibody structure samples, and a coordinate of the k-th amino acid in the predicted antibody structure corresponding to each of the enhanced antibody structure samples, wherein K is a total number of amino acid sites of the enhanced antibody structure sample.

[0013] In some embodiments, the obtaining the second loss function comprises: calculating a deviation between the coordinate of the k-th amino acid in each of the enhanced antibody structure samples and the coordinate of the k-th amino acid in the predicted antibody structure corresponding to each of the enhanced antibody structure samples to obtain a coordinate deviation of the k-th amino acid; and obtaining the second loss function according to a sum of the coordinate deviations of the K amino acid sites.

[0014] In some embodiments, the determining the loss function comprises: calculating a weighted sum of the first loss function and the second loss function to obtain the loss function.

[0015] In some embodiments, the extracting the plurality of enhanced antibody structure samples from the molecular dynamics trajectory comprises: extracting an initial conformation of the original antibody structure sample from the molecular dynamics trajectory; extracting, from the molecular dynamics trajectory, conformations having a similarity greater than a perturbation threshold to the initial conformation as preselected conformations; clustering the preselected conformations to obtain a plurality of conformation clusters; and extracting one or more representative conformations from each of the plurality of conformation clusters as the enhanced antibody structure samples.

[0016] In some embodiments, the extracting one or more representative conformations from each of the plurality of conformation clusters comprises: determining, according to a sampling weight of the each of the plurality of conformation clusters, a number of representative conformations to be extracted from the each of the plurality of conformation clusters; and extracting the number of representative conformations from the each of the plurality of conformation clusters.

[0017] In some embodiments, the processing each of the enhanced antibody structure samples in the enhanced dataset using the machine learning model comprises: extracting node features and edge features of the each of the enhanced antibody structure samples; extracting an entropy feature of the each of the enhanced antibody structure samples; and processing the node features, the edge features, and the entropy feature to obtain a predicted amino acid sequence of the each of the enhanced antibody structure samples.

[0018] In some embodiments, the extracting the entropy feature of the each of the enhanced antibody structure samples comprises: performing principal component analysis on the molecular dynamics trajectory to obtain a plurality of principal components and a variance of each of the plurality of principal components; selecting, in order from large to small variance, a predetermined number of principal components from the plurality of principal components as target principal components; and determining the entropy feature of the each of the enhanced antibody structure samples according to the variance of the target principal components.

[0019] In some embodiments, the processing each of the enhanced antibody structure samples in the enhanced dataset using the machine learning model further comprises: masking the each of the enhanced antibody structure samples using the machine learning model according to a preset masking rule to obtain a masked antibody structure sample; and processing the masked antibody structure sample using the machine learning model to obtain the predicted amino acid sequence of the each of the enhanced antibody structure samples.

[0020] In some embodiments, the preset masking rule comprises: a masking rate of a complementarity determining region of the enhanced antibody structure sample is a first masking rate, and a masking rate of a framework region of the enhanced antibody structure sample is a second masking rate, wherein the first masking rate is greater than the second masking rate.

[0021] In some embodiments, the first masking rate is 40%, and the second masking rate is 15%.

[0022] In some embodiments, the preset mask rule further comprises: masking a plurality of consecutive amino acids on the enhanced antibody structure sample.

[0023] In some embodiments, the processing each enhanced antibody structure sample in the enhanced dataset by using the machine learning model further comprises: dividing the plurality of enhanced antibody structure samples into a plurality of homologous clusters according to the similarity between the amino acid sequences of the plurality of enhanced antibody structure samples, wherein if the similarity between the amino acid sequences of the complementarity determining regions of two enhanced antibody structure samples is greater than a homologous threshold, the two enhanced antibody structure samples are divided into a homologous cluster; dividing the enhanced dataset into a training set for training the model, a validation set for adjusting the parameters of the model, and a test set for evaluating the performance of the model, wherein the enhanced antibody structure samples belonging to the same homologous cluster are placed in the same set of the training set, the validation set, and the test set; processing each enhanced antibody structure sample in the training set by using the machine learning model.

[0024] In some embodiments, the machine learning model comprises M layers of encoders and M layers of decoders, the learning rate of an i-th layer of encoder is less than the learning rate of a j-th layer of encoder, and the learning rate of an i-th layer of decoder is less than the learning rate of a j-th layer of decoder, .

[0025] According to a second aspect of the embodiments of the present disclosure, a nanobody reverse folding method is provided, comprising: obtaining a target antibody structure; processing the target antibody structure by using a sequence prediction model to obtain a predicted amino acid sequence of the target antibody structure, wherein the sequence prediction model is obtained according to the model training method of any of the above embodiments.

[0026] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a memory; a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to implement the method according to any of the above embodiments.

[0027] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, wherein the computer readable storage medium stores computer instructions, and the instructions are executed by a processor to implement the method according to any of the above embodiments.

[0028] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, comprising computer instructions, and the computer instructions are executed by a processor to implement the method according to any of the above embodiments.

[0029] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments thereof, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those of ordinary skill in the art without creative labor under the premise of not paying creative labor.

[0031] Figure 1 A flowchart of a model training method according to an embodiment of the present disclosure;

[0032] Figure 2 A flowchart of a model training method according to another embodiment of the present disclosure;

[0033] Figure 3 A flowchart of a nanobody reverse folding method according to an embodiment of the present disclosure;

[0034] Figure 4 A structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the at least one exemplary embodiment is actually only illustrative, but not as any limitation on the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present disclosure.

[0036] Unless otherwise specified, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0037] Meanwhile, it should be understood that, for the convenience of description, the sizes of the various parts shown in the drawings are not drawn in accordance with the actual proportional relationship.

[0038] The technology, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered as part of the authorized description under appropriate circumstances.

[0039] In all the examples shown and discussed here, any specific value should be interpreted as merely exemplary, and not as a limitation. Therefore, other examples of the exemplary embodiments can have different values.

[0040] It should be noted that similar reference numbers and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0041] The inventors noticed that the sequence prediction task of a protein is to predict the amino acid sequence of a protein according to the structure of the protein, and the ESM-IF model of Facebook and the MPNN (Message Passing Neural Network) model of Baker have made some progress in the sequence prediction task of a protein. At the same time, the AbMPNN model and the AntiFold model based on the protein language model (PLM) are also used in the sequence prediction task of an antibody in order to realize antibody design.

[0042] The defects of existing machine learning models in the nanobody unfolding task are shown in Table 1.

[0043] Table 1

[0044]

[0045] In combination with the above analysis, when a sequence prediction model is used to realize nanobody unfolding, the following problems exist:

[0046] 1) Due to the small amount of antibody structure data of nanobody, the data amount requirement for training of a machine learning model cannot be met, resulting in poor training effect of the sequence prediction model.

[0047] 2) The nanobody is a single-chain structure, and the kinetic process of the complementarity determining region (CDR) of the nanobody is more complex than that of the traditional antibody. At the same time, the region used by the nanobody to bind to the antigen can also be a beta-sheet segment, and the instability of such natural structure also affects the training process of the sequence prediction model.

[0048] Accordingly, the present disclosure provides a model training method capable of improving the prediction performance of a sequence prediction model.

[0049] Figure 1 A flowchart of the model training method of one embodiment of the present disclosure includes steps 11-15.

[0050] In step 11, molecular dynamics (MD) simulation is performed on the original antibody structure sample to obtain the molecular dynamics trajectory of the original antibody structure sample. ​

[0051] In some embodiments, raw antibody structure samples are collected from a database.

[0052] For example, filtering from a PDB (Protein Data Bank) for samples with a resolution smaller than [missing information]. Nanobody structures, such as the nanobody structure with PDB ID 7QN4 and the nanobody structure with PDB ID 1ZVH, were used as original antibody structure samples, resulting in 1200 original antibody structure samples.

[0053] It should be noted that, considering that CDR-H3 in nanobodies is the key region for antigen binding, nanobodies containing the complete CDR region, especially those containing the complete CDR-H3 region, are preferred.

[0054] In some embodiments, the method for performing molecular dynamics simulations on the original antibody structure sample includes the following steps S11-S15.

[0055] In step S11, the original antibody structure sample is preprocessed.

[0056] For example, the original antibody structure sample can be pretreated using molecular modeling tools such as Chimera or PyMOL. Pretreatment includes at least one of the following: (1) repairing missing residues, especially CDR regions; (2) adding hydrogen atoms and optimizing the protonation state to achieve a pH of 7.0; (3) removing water of crystallization molecules and non-protein ligands.

[0057] In step S12, a solvation model is constructed, and the ion balance of the solvation model is adjusted.

[0058] For example, the original antibody structure sample is placed in a cubic water box, using the TIP3P (Transferable Intermolecular Potential 3-Point) water model, with a boundary distance greater than or equal to... Ensure adequate solubilization of the protein surface. Add a physiological concentration (e.g., 150 mmol / L). Ions, in order to neutralize the charge, mimic the physiological environment.

[0059] In step S13, select the force field.

[0060] For example, using AMBER ff19SB (Assisted Model Building with Energy Refinement - force field 2019 Side-chain & Backbone) or CHARMM36m (Chemistry at Harvard Macromolecular Mechanics - version 36 modified) force field, the accuracy of the simulation of the nanobody conformation (especially the flexible loop region) is higher.

[0061] In step S14, the CDR region of the enhanced antibody structure sample is processed.

[0062] For example, the length of the CDR-H3 loop is greater than 12 residues, and the enhanced dihedral potential is applied to the CDR-H3 loop to increase its conformational sampling efficiency.

[0063] In step S15, a thermal perturbation heating simulation is performed.

[0064] In some embodiments, the temperature range of the thermal perturbation heating simulation is 300-500K, and the time length is greater than or equal to 100ns.

[0065] For example, at a temperature of 400K, a 150ns thermal perturbation heating simulation is performed.

[0066] It should be noted here that lowering the energy barrier by high temperature can promote the CDR loop conformational flip.

[0067] For example, in the thermal perturbation heating simulation, the temperature coupling uses the Berendsen or Nose-Hoover algorithm, and the step size is 2fs. The hydrogen bond constraint uses the LINCS or SHAKE algorithm.

[0068] It should be noted here that the serial numbers of steps S11-S15 do not represent the execution order of the above steps.

[0069] In some embodiments, the method of performing molecular dynamics simulation on the original antibody structure sample further comprises performing molecular dynamics simulation on the original antibody structure sample using an enhanced sampling technique.

[0070] For example, the enhanced sampling techniques include Gaussian Accelerated Molecular Dynamics (GaMD) and Replica-Exchange Molecular Dynamics (REMD). Among them, GaMD accelerates the conformational change of the CDR region by applying a Gaussian enhanced potential to the CDR region, while maintaining the energy landscape authenticity. REMD enhances the CDR conformational space sampling by setting 8-12 temperature replicas in the temperature range of 300-500K, with an exchange frequency of 1-2 ps.

[0071] By the method involved in the above embodiment, the molecular dynamics simulation is performed on the original antibody structure sample to obtain a molecular dynamics trajectory of the original antibody structure sample.

[0072] In step 12, a plurality of enhanced antibody structure samples are extracted from the molecular dynamics trajectory to constitute an enhanced data set.

[0073] In some embodiments, the method of extracting a plurality of enhanced antibody structure samples from the molecular dynamics trajectory includes steps S21-S24.

[0074] In step S21, the initial conformation of the original antibody structure sample is extracted from the molecular dynamics trajectory.

[0075] In step S22, a conformation with a similarity greater than a perturbation threshold to the initial conformation is extracted from the molecular dynamics trajectory as a preselected conformation.

[0076] In some embodiments, the RMSD (Root Mean Square Deviation) between each conformation in the molecular dynamics trajectory and the initial conformation is calculated as the similarity of each conformation to the initial conformation. The perturbation threshold can be set to The conformation with an RMSD greater than to the initial conformation is selected as the preselected conformation.

[0077] For example, assuming that the initial conformation in the molecular dynamics trajectory is , the fth conformation is , then the fth conformation has an to the initial conformation , as shown in equation (1).

[0078] (1)

[0079] wherein is the total number of atoms involved in the calculation, for the initial conformation the number of atoms participating in the calculation of the the number of atoms participating in the calculation of the for the f-th conformation the number of atoms participating in the calculation of the the number of atoms participating in the calculation of the

[0080] It should be noted that the RMSD is an important indicator for measuring the difference between two sets of data (for example, predicted values and true values, or two molecular structures), which can be calculated for any two structures. For example, the difference between two minimum point configurations, the overall change in geometry before and after electronic excitation / ionization / external electric field, or the change in monomer structure after forming a complex can be measured. The value range of the RMSD is The smaller the RMSD value, the higher the consistency of the two sets of data; the larger the RMSD value, the greater the difference between the two sets of data. The conformational difference of the CDR region is focused on.

[0081] In step S23, the preselected conformations are clustered to obtain a plurality of conformation clusters.

[0082] In some embodiments, the preselected conformations are clustered based on the atoms of the CDR region of the preselected conformation, to obtain a plurality of conformation clusters.

[0083] For example, using the cluster command of GROMACS (Groningen Machine for Chemical Simulations) or AMBER (Assisted Model Building with Energy Refinement), the preselected conformations are clustered based on the atom positions of the CDR region of the preselected conformation, to obtain a plurality of conformation clusters. The clustering algorithm can use the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm.

[0084] In step S24, one or more representative conformations are extracted from each of the plurality of conformation clusters as enhanced antibody structure samples.

[0085] It should be noted that one or more conformations can be initially extracted from each conformation cluster, for example, more than or equal to 50 conformations are extracted from each 100 ns trajectory, and the extracted conformations are further screened to obtain representative conformations.

[0086] It should be noted that the conformation clusters can also be screened. For example, conformation clusters containing only one conformation are removed. Alternatively, multiple conformation clusters are sorted according to the number of conformations contained in the conformation clusters in descending order, and only the top ten conformation clusters are retained.

[0087] In some embodiments, the number of representative conformations extracted from each conformation cluster is determined according to the sampling weight of each conformation cluster. The number of representative conformations is extracted from each conformation cluster.

[0088] It should be noted that the sampling weight of each conformation cluster can be determined according to the number of conformations contained in each conformation cluster.

[0089] For example, the first conformation cluster contains 4000 conformations, and the number of representative conformations extracted therefrom is 20.

[0090] For another example, the second conformation cluster contains 2000 conformations, and the number of representative conformations extracted therefrom is 10.

[0091] For another example, the fifth conformation cluster contains 600 conformations, and the number of representative conformations extracted therefrom is 3.

[0092] It should be noted that the sampling weight of the conformation cluster of interest can also be increased.

[0093] It should be noted that the number of enhanced antibody structure samples in the enhanced data set can reach 10,000.

[0094] In some embodiments, Gaussian noise is added to the enhanced antibody structure samples to form the enhanced data set.

[0095] It should be noted that adding Gaussian noise to the enhanced antibody structure samples can avoid model overfitting and improve the robustness of the model.

[0096] Through the method involved in the above embodiments, multiple enhanced antibody structure samples are extracted from the molecular dynamics trajectory of an original antibody structure sample to form an enhanced data set, so as to realize data enhancement of the antibody structure sample.

[0097] In step 13, each enhanced antibody structure sample in the enhanced data set is processed using a machine learning model to obtain a predicted amino acid sequence of each enhanced antibody structure sample.

[0098] In some embodiments, the enhanced antibody structure samples in the enhanced data set are subjected to structure standardization processing.

[0099] For example, the coordinates of the backbone atoms in the enhanced antibody structure sample are extracted, including , and atoms; according to IMGT (International ImMunoGeneTics Information System) numbering, the complementarity determining regions and framework regions (FR) between different enhanced antibody structure samples are aligned so as to eliminate structural bias.

[0100] In some embodiments, the plurality of enhanced antibody structure samples are divided into a plurality of homology clusters according to the similarity between the amino acid sequences of the plurality of enhanced antibody structure samples, wherein if the similarity between the amino acid sequences of the complementarity determining regions of two enhanced antibody structure samples is greater than a homology threshold, the two enhanced antibody structure samples are divided into a homology cluster. The enhanced dataset is divided into a training set for training the model, a validation set for adjusting the model parameters, and a test set for evaluating the performance of the model, wherein the enhanced antibody structure samples belonging to the same homology cluster are placed in the same set in the training set, the validation set and the test set. Each enhanced antibody structure sample in the training set is processed by using a machine learning model.

[0101] For example, the amino acid sequences of the CDR regions of the enhanced antibody structure samples are spliced together to form the joint CDR sequences of the enhanced antibody structure samples. The homology threshold can be set to 90%. If the similarity between the joint CDR sequences of two enhanced antibody structure samples is greater than 90%, the two enhanced antibody structure samples are considered to be homologous and are divided into the same homology cluster.

[0102] It should be noted here that the enhanced antibody structure samples belonging to the same homology cluster are all placed in the training set, or all placed in the validation set, or all placed in the test set, ensuring that there is no sequence homology overlap between any two of the training set, the validation set and the test set, thereby avoiding overfitting.

[0103] For example, the enhanced dataset is divided into the training set, the validation set and the test set in a ratio of 8:1:1.

[0104] In some embodiments, each enhanced antibody structure sample in the enhanced dataset is processed by using a machine learning model, including the following steps S31-S33.

[0105] In step S31, the node features and edge features of each enhanced antibody structure sample are extracted.

[0106] For example, the node features include residue type, secondary structure type, and solvent accessible surface area (SASA). The edge features include atomic distance, dihedral angle difference, and hydrogen bond energy.

[0107] In step S32, entropy features of each enhanced antibody structure sample are extracted.

[0108] In some embodiments, the entropy features of each enhanced antibody structure sample are extracted by a conformation entropy encoding module in the machine learning model.

[0109] In some embodiments, principal component analysis (PCA) is performed on the molecular dynamics trajectory to obtain a plurality of principal components and a variance of each principal component in the plurality of principal components. According to the order of variance from large to small, a predetermined number of principal components are selected from the plurality of principal components as target principal components. According to the variance of the target principal components, the entropy features of each enhanced antibody structure sample are determined.

[0110] For example, the predetermined number is 3.

[0111] For example, principal component analysis is performed on the molecular dynamics trajectory to obtain a plurality of principal components and a variance of each principal component in the plurality of principal components. According to the order of variance from large to small, 3 principal components are selected from the plurality of principal components as target principal components. The variance of the target principal components is taken as the entropy features of the enhanced antibody structure samples extracted from the above molecular dynamics trajectory.

[0112] It should be noted that by extracting the entropy features of the enhanced antibody structure samples for training the machine learning model, the flexibility changes of the CDR regions can be dynamically captured, the modeling of the CDR regions can be enhanced, and thus the training effect of the machine learning model can be improved.

[0113] In some embodiments, a temporal convolutional network (TCN) in the machine learning model is used to extract the time domain features of the molecular dynamics trajectory. The machine learning model is used to process the time domain features so as to obtain the predicted amino acid sequence of each enhanced antibody structure sample.

[0114] For example, principal component analysis is performed on the molecular dynamics trajectory to obtain a plurality of principal components, which constitute a principal component sequence. The principal component sequence is processed by the temporal convolutional network to obtain the time domain features of the molecular dynamics trajectory.

[0115] It should be noted that temporal convolutional networks are used to extract the temporal features of molecular dynamics trajectories, learn the temporal dynamics of conformational transitions in molecular dynamics trajectories, and predict state evolution or rare events.

[0116] In step S33, the node features, edge features, and entropy features are processed to obtain the predicted amino acid sequence of each enhanced antibody structure sample.

[0117] In some embodiments, the machine learning model employs a Transformer architecture, including a multi-layer GVP-GNN (Geometric Vector Perceptron – Graph Neural Network), a multi-layer encoder, and a multi-layer decoder. The multi-layer GVP-GNN is used to process geometric vector features. The multi-layer encoder is used to learn sequence-structure dependencies. The multi-layer decoder is used to generate conditional probability distributions.

[0118] In some embodiments, the machine learning model includes a 4-layer GVP-GNN. The 4-layer GVP-GNN processes geometric vector features including atomic coordinates and bond angle orientations. The 4-layer GVP-GNN can update node features through a message passing mechanism.

[0119] For example, suppose the p-th node and the q-th node are adjacent, the p-th node... The node features of the p-th node output by the layer are: , No. The node features of the q-th node output by the layer are: The edge characteristics between the p-th node and the q-th node are: Then the first Node features of the p-th node output by the layer As shown in formula (2).

[0120] (2)

[0121] in, For message functions, Let be the set of all neighboring nodes of the p-th node. It is a geometric vector perceptron. It includes scalar features (e.g., residue type) and vector features (e.g., spatial geometry information).

[0122] For example, a machine learning model includes an 8-layer Transformer encoder and an 8-layer Transformer decoder. The 8-layer decoder is used to generate conditional probability distributions. where Y is the predicted amino acid sequence and X is the antibody structure coordinates. For each amino acid site, a 20-dimensional probability distribution is obtained. The machine learning model can generate the predicted amino acid sequence according to the conditional probability distribution.

[0123] In some embodiments, the machine learning model comprises M layers of encoders and M layers of decoders, a learning rate of an i-th layer of encoder is less than a learning rate of a j-th layer of encoder, a learning rate of an i-th layer of decoder is less than a learning rate of a j-th layer of decoder, .

[0124] For example, the machine learning model determines the learning rate of each layer of encoder and each layer of decoder by adopting a hierarchical learning rate decay strategy. Assuming that the initial learning rate is LR and the hierarchical decay factor is , the learning rate of the i-th layer of encoder and the i-th layer of decoder is as shown in formula (3).

[0125] (3)

[0126] wherein the initial learning rate LR can be set to 5e-4.

[0127] It should be noted that setting a lower learning rate for the shallow layer of encoder and the shallow layer of decoder in the machine learning model and setting a higher learning rate for the deep layer of encoder and the deep layer of decoder can avoid loss of pre-training knowledge or catastrophic forgetting.

[0128] In some embodiments, according to a preset masking rule, the machine learning model is used to mask each enhanced antibody structure sample to obtain a masked antibody structure sample. The machine learning model is used to process the masked antibody structure sample to obtain the predicted amino acid sequence of each enhanced antibody structure sample.

[0129] For example, the machine learning model is a Masked Language Model (MLM).

[0130] In some embodiments, the preset masking rule comprises: a masking rate of a complementarity determining region of the enhanced antibody structure sample is a first masking rate, and a masking rate of a framework region of the enhanced antibody structure sample is a second masking rate, wherein the first masking rate is greater than the second masking rate.

[0131] In some embodiments, the first masking rate is 40%, and the second masking rate is 15%.

[0132] It should be noted that by setting a higher masking rate for the complementarity determining region of the enhanced antibody structure sample, the learning of the model on the hypervariable region is strengthened.

[0133] In some embodiments, the preset mask rule further comprises: masking a plurality of amino acids that are continuous on the enhanced antibody structure sample.

[0134] It should be noted here that the plurality of amino acids that are continuous on the enhanced antibody structure sample refers to a plurality of amino acid residues that are continuous on the enhanced antibody structure sample. At most, 6 continuous amino acid residues can be masked. By span masking, the machine learning model is forced to learn long-range dependencies.

[0135] In step 14, a loss function is determined according to the predicted amino acid sequence of each enhanced antibody structure sample.

[0136] In some embodiments, the method of determining the loss function comprises steps S41-S44.

[0137] In step S41, a first loss function is determined according to the predicted amino acid sequence of each enhanced antibody structure sample, and the amino acid sequence sample corresponding to each enhanced antibody structure sample.

[0138] In some embodiments, the first loss function is obtained according to the probability that the kth amino acid in the predicted amino acid sequence is the kth amino acid in the amino acid sequence sample, and the weight of the kth amino acid, where K is the total number of amino acid positions of the enhanced antibody structure sample.

[0139] In some embodiments, the product of the logarithmic value of the probability and the weight is calculated to obtain the probability weighting value of the kth amino acid. The sum of the probability weighting values of the K amino acid positions is obtained to obtain the first loss function.

[0140] For example, assuming that the kth amino acid in the amino acid sequence sample is , and the predicted amino acid sequence before the kth amino acid of the enhanced antibody structure sample is , the probability that the kth amino acid in the predicted amino acid sequence is is , and the weight of the kth amino acid is , then the first loss function is as shown in formula (4).

[0141] (4)

[0142] In some embodiments, if the kth amino acid belongs to a complementarity determining region, the weight of the kth amino acid is a first weight value. If the kth amino acid belongs to a framework region, the weight of the kth amino acid is a second weight value, where the first weight value is greater than the second weight value.

[0143] In some embodiments, the first weight value is 0.8. The second weight value is 0.2.

[0144] For example, the weight of the amino acid (e.g., Gly112-Ser124) in the CDR-H3 loop of the enhanced antibody structure sample is set to 0.8.

[0145] In step S42, the predicted amino acid sequence of each enhanced antibody structure sample is processed by using the structure prediction model to obtain a predicted antibody structure corresponding to each enhanced antibody structure sample.

[0146] For example, the structure prediction model is an Alphafold model, which can predict the antibody structure according to the amino acid sequence.

[0147] In step S43, a second loss function is determined according to each enhanced antibody structure sample and the predicted antibody structure corresponding to each enhanced antibody structure sample.

[0148] In some embodiments, the second loss function is obtained according to the coordinates of the kth amino acid in each enhanced antibody structure sample and the coordinates of the kth amino acid in the predicted antibody structure corresponding to each enhanced antibody structure sample, wherein K is the total number of amino acid sites of the enhanced antibody structure sample.

[0149] In some embodiments, the deviation between the coordinates of the kth amino acid in each enhanced antibody structure sample and the coordinates of the kth amino acid in the predicted antibody structure corresponding to each enhanced antibody structure sample is calculated to obtain the coordinate deviation of the kth amino acid. The sum of the coordinate deviations of the K amino acid sites is obtained to obtain the second loss function.

[0150] For example, assuming that the coordinates of the kth amino acid in the enhanced antibody structure sample are and the coordinates of the kth amino acid in the predicted antibody structure corresponding to the enhanced antibody structure sample are then the second loss function is as shown in formula (5).

[0151] (5)

[0152] In step S44, a loss function is determined according to the first loss function and the second loss function.

[0153] In some embodiments, the weighted sum of the first loss function and the second loss function is calculated to obtain the loss function.

[0154] It should be noted here that the loss function adds a structure consistency constraint on the basis of dynamic weighted cross-entropy.

[0155] In step 15, the sequence prediction model is obtained by training the machine learning model by using the loss function.

[0156] It should be noted that the machine learning model can update the model parameters using the AdamW optimizer. The training period is greater than 200 epochs. The GPU hardware used for training is NVIDIA Tesla V100 / A100, and a single card can process 300 antibody structure samples per minute. CUDA (Compute Unified Device Architecture) unified memory management is enabled to support long sequence training. The sequence prediction model obtained by training is used to obtain the predicted amino acid sequence of the nanobody structure according to the nanobody structure.

[0157] In some embodiments, a pre-training model is obtained by training a machine learning model on a general dataset. The sequence prediction model is obtained by training the pre-training model on an enhanced dataset.

[0158] For example, the general dataset is constructed using the predicted structure of the amino acid sequence in the OAS (Observed Antibody Space) database.

[0159] It should be noted that training on the massive data of the general dataset enables the pre-training model to learn general protein knowledge, and training on the enhanced dataset constructed for nanobodies enables the sequence prediction model to improve antibody-specific modeling capability and achieve pre-training knowledge transfer.

[0160] Through the model training method involved in the above embodiments, the prediction performance of the sequence prediction model can be improved.

[0161] Figure 2 The flowchart of the model training method of another embodiment of the present disclosure includes steps 21-26.

[0162] In step 21, the original antibody structure sample is obtained.

[0163] In step 22, molecular dynamics simulation is performed on the original antibody structure sample to obtain an enhanced dataset. The enhanced dataset includes a plurality of enhanced antibody structure samples.

[0164] In step 23, the enhanced antibody structure samples in the enhanced dataset are subjected to feature extraction to obtain node features, edge features, and entropy features of the enhanced antibody structure samples.

[0165] In step 24, the extracted node features, edge features, and entropy features are subjected to feature encoding.

[0166] In step 25, the machine learning model is subjected to mask training to obtain a sequence prediction model.

[0167] At step 26, a predicted amino acid sequence of the enhanced antibody structure sample is generated by using the sequence prediction model.

[0168] By the model training method involved in the above embodiments, the original antibody structure sample is subjected to molecular dynamics simulation to obtain an enhanced data set, the machine learning model is subjected to mask training by using the enhanced data set, and a predicted amino acid sequence of the enhanced antibody structure sample is generated, so as to realize data enhancement and improve the prediction performance of the sequence prediction model. Meanwhile, by extracting the entropy feature of the enhanced antibody structure sample, the machine learning model is used to process the entropy feature of the enhanced antibody structure sample, which is conducive to the model to capture the flexible change of the CDR region of the sample and enhance the modeling of the CDR region.

[0169] Figure 3 A flowchart of the nanobody reverse folding method of one embodiment of the present disclosure is shown, which includes steps 31-32.

[0170] At step 31, a target antibody structure is obtained.

[0171] At step 32, the target antibody structure is processed by using a sequence prediction model to obtain a predicted amino acid sequence of the target antibody structure. The sequence prediction model is obtained according to the model training method involved in any of the above embodiments.

[0172] By the nanobody reverse folding method involved in the above embodiments, the target antibody structure is processed by using a sequence prediction model to obtain a predicted amino acid sequence of the target antibody structure, which can improve the accuracy of the nanobody reverse folding result, thereby helping to design an antibody with a specific function.

[0173] Figure 4 A structural diagram of an electronic device of one embodiment of the present disclosure is shown. As shown in Figure 4 The electronic device 40 includes a memory 41, a processor 42, and a bus 43 connecting different system components.

[0174] The memory 41 may, for example, include a system memory, a non-volatile storage medium, etc. The system memory, for example, stores an operating system, an application program, a Boot Loader, and other programs, etc. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium, for example, stores instructions of at least one embodiment of the method being executed. The non-volatile storage medium includes, but is not limited to, a disk storage, an optical storage, a flash memory, etc.

[0175] The processor 42 can be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete hardware component such as a discrete gate or transistor, and the like. Accordingly, the method in any of the above embodiments can be implemented by a central processing unit (CPU) running instructions stored in a memory to perform corresponding steps, or by a special-purpose circuit performing corresponding steps.

[0176] For example, the processor 42 is configured to perform the sequence prediction model training method as described in any of Figure 1 and Figure 2 to implement the model training method and the nanobody unfolding method as described in any of Figure 3 and

[0177] The bus 43 can use any of a variety of bus structures. For example, the bus structure includes, but is not limited to, an industry standard architecture (ISA) bus, a micro channel architecture (MCA) bus, a peripheral component interconnect (PCI) bus.

[0178] The interfaces 44, 45, 46 of the electronic device 40, and the memory 41 and the processor 42 can be connected through the bus 43. The input / output interface 44 can provide a connection interface for input / output devices such as a display, a mouse, a keyboard, and the like. The network interface 45 provides a connection interface for various networking devices. The storage interface 46 provides a connection interface for external storage devices such as a floppy disk, a U disk, an SD card, and the like.

[0179] Through the implementation of the above embodiments of the present disclosure, the beneficial effects obtained are shown in Table 2.

[0180] Table 2

[0181] It should be noted here that in Table 2, the first row is the RMSD index calculated on the nanobody single domain test set, with the unit of . The second row is the RMSD index calculated on the CDR-H3 loop test set, with the unit of . The third row is the thermodynamic stability (ΔG) index, with the unit of kcal / mol.

[0182] As shown in Table 2, the sequence prediction model obtained by the model training method proposed in the present disclosure has smaller RMSD on the nanobody single domain test set and the CDR-H3 loop test set, and the predicted amino acid sequence has better thermodynamic stability.

[0183] ​Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations thereof, can be implemented by computer-readable program instructions.

[0184] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, such that execution of the instructions by the processor produces means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.

[0185] These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions cause a computer to work in a particular manner to produce an article of manufacture, including instructions that implement the functions specified in one or more boxes in a flowchart and / or block diagram.

[0186] This disclosure may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.

[0187] This disclosure also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement... Figure 1 and Figure 2 The model training method involved in any embodiment and such Figure 3 The nanobody defolding method involved in any of the embodiments.

[0188] This disclosure also provides a computer program product, including computer instructions, wherein the computer instructions, when executed by a processor, implement as follows: Figure 1 and Figure 2 The model training method involved in any embodiment and such Figure 3 The nanobody defolding method involved in any of the embodiments.

[0189] In some embodiments, the functional units described above may be implemented as general-purpose processors, programmable logic controllers (PLCs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any suitable combination thereof for performing the functions described herein.

[0190] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be instructed by programs to complete the related hardware, and the programs can be stored in a computer readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0191] The description of the present disclosure is given for the purpose of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are chosen and described in order to better illustrate the principles and practical application of the present disclosure, and to enable others skilled in the art to understand the present disclosure in order to design various embodiments with various modifications for specific use.

Claims

1. A method for training a model, comprising: performing molecular dynamics simulation on an original antibody structure sample to obtain a molecular dynamics trajectory of the original antibody structure sample; extracting a plurality of enhanced antibody structure samples from the molecular dynamics trajectory to form an enhanced dataset; processing each enhanced antibody structure sample in the enhanced dataset by using a machine learning model to obtain a predicted amino acid sequence of the each enhanced antibody structure sample; determining a loss function according to the predicted amino acid sequence of the each enhanced antibody structure sample; training the machine learning model by using the loss function to obtain a sequence prediction model. 2.The model training method of claim 1, wherein, The determining of the loss function comprises: determining a first loss function according to the predicted amino acid sequence of the each enhanced antibody structure sample and an amino acid sequence sample corresponding to the each enhanced antibody structure sample; processing the predicted amino acid sequence of the each enhanced antibody structure sample by using a structure prediction model to obtain a predicted antibody structure corresponding to the each enhanced antibody structure sample; determining a second loss function according to the each enhanced antibody structure sample and the predicted antibody structure corresponding to the each enhanced antibody structure sample; determining the loss function according to the first loss function and the second loss function. 3.The model training method of claim 2, wherein, The determining of the first loss function comprises: According to the probability that the k-th amino acid in the predicted amino acid sequence is the k-th amino acid in the amino acid sequence sample, and the weight of the k-th amino acid, the first loss function is obtained, wherein K is the total number of amino acid sites of the enhanced antibody structure sample. 4.The model training method of claim 3, wherein, The obtaining of the first loss function comprises: calculating a product of a logarithm value of the probability and the weight to obtain a probability weighted value of the k th amino acid; obtaining the first loss function according to a sum of the probability weighted values of the K amino acid sites. 5.The method of claim 3, wherein, if the k th amino acid belongs to a complementarity determining region, the weight of the k th amino acid is a first weight value; if the k th amino acid belongs to a framework region, the weight of the k th amino acid is a second weight value, wherein the first weight value is greater than the second weight value. 6.The method of claim 5, wherein, the first weight value is 0.8; the second weight value is 0.

2. 7.The method of Claim 2, wherein, The determining of the second loss function comprises: According to the coordinates of the k-th amino acid in each enhanced antibody structure sample and the coordinates of the k-th amino acid in the predicted antibody structure corresponding to each enhanced antibody structure sample, the second loss function is obtained, wherein K is the total number of amino acid sites of the enhanced antibody structure sample. 8.The method of claim 7, the obtaining of the second loss function comprises: calculating a deviation between a coordinate of a k th amino acid in the each enhanced antibody structure sample and a coordinate of the k th amino acid in the predicted antibody structure corresponding to the each enhanced antibody structure sample to obtain a coordinate deviation of the k th amino acid; obtaining the second loss function according to a sum of the coordinate deviations of the K amino acid sites. 9.The method of Claim 2, wherein, The determining of the loss function comprises: calculating a weighted sum of the first loss function and the second loss function to obtain the loss function. 10.The method of Claim 1, wherein The extracting of the plurality of enhanced antibody structure samples from the molecular dynamics trajectory comprises: extracting an initial conformation of the original antibody structure sample from the molecular dynamics trajectory; extracting a conformation with a similarity greater than a perturbation threshold to the initial conformation from the molecular dynamics trajectory as a preselected conformation; clustering the preselected conformation to obtain a plurality of conformation clusters; extract one or more representative conformations from each of the plurality of conformational clusters as the augmented antibody structure samples.

11. The model training method according to claim 10, wherein, The extracting one or more representative conformations from each of the plurality of conformational clusters comprises: determining a number of extractions of representative conformations extracted from each of the conformational clusters according to a sampling weight of the each of the conformational clusters; extracting the number of representative conformations from the each of the conformational clusters. 12.The method of Claim 1, wherein The processing each of the augmented antibody structure samples in the augmented dataset by using the machine learning model comprises: extracting node features and edge features of the each of the augmented antibody structure samples; extracting entropy features of the each of the augmented antibody structure samples; processing the node features, the edge features and the entropy features to obtain a predicted amino acid sequence of the each of the augmented antibody structure samples. 13.The model training method of claim 12, wherein, The extracting the entropy features of the each of the augmented antibody structure samples comprises: performing principal component analysis on the molecular dynamics trajectory to obtain a plurality of principal components and a variance of each of the plurality of principal components; selecting a predetermined number of principal components from the plurality of principal components as target principal components in a descending order of the variances; determining the entropy features of the each of the augmented antibody structure samples according to the variances of the target principal components. 14.The method of Claim 1, wherein The processing each of the augmented antibody structure samples in the augmented dataset by using the machine learning model further comprises: masking the each of the augmented antibody structure samples by using the machine learning model according to a preset masking rule to obtain a masked antibody structure sample; processing the masked antibody structure sample by using the machine learning model to obtain the predicted amino acid sequence of the each of the augmented antibody structure samples. 15.The model training method of claim 14, wherein, The preset masking rule comprises: a masking rate of a complementarity determining region of the augmented antibody structure sample is a first masking rate, and a masking rate of a framework region of the augmented antibody structure sample is a second masking rate, wherein the first masking rate is greater than the second masking rate.

16. The model training method of claim 15, wherein, the first masking rate is 40%; the second masking rate is 15%.

17. The model training method of claim 14, wherein, The preset masking rule further comprises: masking a plurality of amino acids that are continuous on the augmented antibody structure sample. 18.The method of Claim 1, wherein, The processing each of the augmented antibody structure samples in the augmented dataset by using the machine learning model further comprises: dividing the plurality of augmented antibody structure samples into a plurality of homologous clusters according to similarities between amino acid sequences of the plurality of augmented antibody structure samples, wherein if a similarity between amino acid sequences of complementarity determining regions of two augmented antibody structure samples is greater than a homology threshold, the two augmented antibody structure samples are divided into a homologous cluster; dividing the augmented dataset into a training set for training a model, a validation set for adjusting model parameters, and a test set for evaluating model performance, wherein augmented antibody structure samples belonging to the same homologous cluster are placed in the same set among the training set, the validation set and the test set; processing each of the augmented antibody structure samples in the training set by using the machine learning model.

19. The model training method of claim 1, wherein, The machine learning model comprises M layers of encoders and M layers of decoders, a learning rate of an i-th layer of encoders is less than a learning rate of a j-th layer of encoders, and a learning rate of an i-th layer of decoders is less than a learning rate of a j-th layer of decoders, .

20. A nanobody reverse folding method, comprising: obtaining a target antibody structure; processing the target antibody structure using a sequence prediction model to obtain a predicted amino acid sequence of the target antibody structure, wherein the sequence prediction model is obtained according to the model training method of any one of claims 1-19.

21. An electronic device, comprising: a memory; a processor coupled to the memory, the processor configured to perform a method as claimed in any one of claims 1-20 based on instructions stored in the memory.

22. A computer readable storage medium, wherein, a computer readable storage medium storing computer instructions, the instructions being executed by a processor to implement a method as claimed in any one of claims 1-20.

23. A computer program product, comprising computer instructions, wherein the computer instructions, when executed by a processor, implement a method as claimed in any one of claims 1-20.