Protein structure prediction

By generating single and paired representations of antibody structure prediction using antibody language model (ALM) and making predictions without relying on multi-sequence alignment (MSA), the problems of inefficiency and loss of accuracy in antibody structure prediction are solved, and efficient and accurate antibody structure prediction is achieved.

CN120092293APending Publication Date: 2025-06-03BIOMAP (BEIJING) INTELLIGENCE TECH LTD

Patent Information

Application Number
CN202380068654.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-27
Filing Date
2023-09-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency and loss of accuracy in antibody structure prediction, especially when relying on multi-sequence alignment (MSA), it is difficult to accurately predict proteins with less homologous information or rapidly evolving antibodies.

Method used

Using a computer-implemented method based on the antibody language model (ALM), residue encoding and attention weight encoding are generated through ALM, and structure prediction is performed without MSA. The method includes inputting the target antibody sequence into the ALM, generating a single representation and paired representation, and inputting it into a structural prediction model for prediction.

Benefits of technology

The efficiency and accuracy of antibody structure prediction are improved, especially in the complementary determination region (CDR) of the antibodies, and a significant performance improvement compared with traditional methods is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092293A_ABST
    Figure CN120092293A_ABST
Patent Text Reader

Abstract

Disclosed herein are methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for antibody structure prediction. In one example method, a target antibody sequence for a target antibody comprising an amino acid sequence is received. The target antibody sequence is processed by an antibody language model (ALM) to obtain residue coding and attention weight coding without multiple sequence alignment (MSA), where the ALM is a protein language model trained according to the antibody sequence, and the ALM comprises a plurality of self-attention layers. The residue coding and attention weight coding are converted into a single representation and paired representations and input into a structure prediction model. And determining a prediction structure of the target antibody by using the structure prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to protein structure prediction, particularly antibody structure prediction based on machine learning techniques. Background Art

[0002] Protein structure prediction is to infer the three-dimensional (3D) structure of a protein from its amino acid sequence. Machine learning methods, such as deep learning methods, can be used for protein structure prediction. Deep learning methods combine the evolutionary and geometric information of protein structures and deep neural networks. Among these deep learning methods, progress has been made in using the co-evolution information of multiple sequence alignments (MSAs), such as AlphaFold, AlphaFold2, OpenFold, and RoseTTAFold. For example, AlphaFold2 provides an architecture to jointly model MSAs and pairwise information and predict protein structures based on protein sequences and MSAs. However, these methods are time-consuming and rely on MSAs. This remains a challenge for the structure prediction of orphan proteins with less homologous information or antibodies for which MSAs are not always useful due to rapid evolution.

[0003] Recently, protein structure prediction has been carried out on large protein language models (PLMs), which no longer rely on MSAs. In particular, models such as DeepAb, ABlooper, and IgFold have been developed for antibody structure prediction. These models can reduce the calculation time but result in a certain loss of prediction accuracy.

[0004] There is a desire to obtain efficient and accurate antibody structure prediction techniques. Summary of the Invention

[0005] Embodiments of the described subject matter may include one or more features, either alone or in combination.

[0006] For example, in one embodiment, a computer-implemented method for antibody structure prediction includes: a data processing device receiving a target antibody sequence of a target antibody, the sequence comprising an amino acid sequence; the data processing device inputting the target antibody sequence into an Antibody Language Model (ALM), where the ALM is a protein language model trained based on antibody sequences and the ALM includes a plurality of self-attention layers; the data processing device using the ALM without performing a Multiple Sequence Alignment (MSA) to obtain residue encoding and attention weight encoding, where the residue encoding includes corresponding first embeddings output by the ALM for each amino acid in the target antibody sequence; the attention weight encoding includes corresponding second embeddings calculated from the attention weights of the self-attention layers of the ALM for a pair of amino acids in the target antibody sequence; the data processing device converting the residue encoding and the attention weight encoding into a single representation and a pairwise representation; the data processing device inputting the single representation and the pairwise representation into a structure prediction model, where the parameters of the structure prediction model are trained based on a loss function reflecting the difference between the predicted structure and the actual structure of the antibody; the data processing device using the structure prediction model to determine the predicted structure of the target antibody based on the single representation and the pairwise representation; and the data processing device outputting the predicted structure of the target antibody.

[0007] In some embodiments, these general and specific aspects can be implemented using a system, a method, or a computer program, or any combination of a system, a method, and a computer program. Each of the above-described and other embodiments may optionally include one or more of the following aspects:

[0008] In some embodiments, the ALM is pre-trained using an antibody database according to the Bidirectional Encoder Representations from Transformers (BERT) architecture, and the antibody database consists of antibody sequences.

[0009] In some embodiments, the ALM includes L self-attention layers, each self-attention layer includes H attention heads, and where the second embedding qij corresponds to a pair of amino acids i and amino acid j in the target antibody sequence, and where the data processing device obtaining the attention weight encoding using the ALM without performing MSA includes: obtaining the attention weights of the H attention heads of each of the L self-attention layers when amino acid i is used as a query and amino acid j is used as a key; and concatenating these attention weights to obtain the second embedding qij.

[0010] In some embodiments, the data processing device converting the residue encoding and the attention weight encoding into a single representation and a pairwise representation includes: converting the residue encoding into a single representation through a first linear neural network layer; and converting the attention weight encoding into a pairwise representation through a second linear neural network layer; wherein the parameters of the first linear neural network layer and the second linear neural network layer are updated based on the gradient of the loss function.

[0011] In some embodiments, the loss function does not include the loss due to the MSA.

[0012] In some embodiments, the loss function includes a Framed Aligned Point Error (FAPE) loss, a torsion angle loss, and a loss for the Complementarity Determining Region (CDR).

[0013] In some embodiments, the loss function includes a differential Root-Mean-Squared-Eeviation (RMSD) in addition to the Framed Aligned Point Error (FAPE) loss.

[0014] In some embodiments, before the single representation and the pairwise representation are input into the structure prediction model, the single representation and the pairwise representation do not contain template features.

[0015] In some embodiments, before the data processing device inputs the single representation and the pairwise representation into the structure prediction model, the computer-implemented method further includes: the data processing device performing a template search based on the target antibody sequence without performing a multiple sequence alignment (MSA) to find one or more template candidates similar to the target antibody structure; and the data processing device obtaining template features based on the one or more template candidates; and wherein the data processing device converting the residue encoding and the attention weight encoding into a single representation and a pairwise representation includes: the data processing device converting the residue encoding and the attention weight encoding into a preliminary single representation and a preliminary pairwise representation; the data processing device merging the template features into the preliminary single representation and the preliminary pairwise representation to obtain the single representation and the pairwise representation.

[0016] In some embodiments, where the data processing device performs a template search to find one or more template candidates includes: the data processing device performs a sequential modal search in a first structure database to find a first structure template, where the antibody sequence corresponding to the first structure template is similar to the target antibody sequence; and the data processing device performs a structural modal search in a second structure database to find a second structure template, where the structure of the second structure template is similar to the coarse-grained structure of the target antibody sequence, and where one or more template candidates include one or more of the first structure template or the second structure template, and where the coarse-grained structure is the default structure, or a structure predicted by another structure prediction algorithm or another structure prediction model.

[0017] It should be understood that the methods according to this specification may include any combination of the aspects and features described herein. That is, the methods according to this specification are not limited to the combinations of aspects and features specifically described herein, but also include any combination of the provided aspects and features.

[0018] Details of one or more embodiments of this specification are set forth in the accompanying drawings and the description below. Other features and advantages of this specification will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic diagram showing an example computer-implemented system configured for protein structure prediction according to an embodiment of this specification.

[0020] Figure 2 is a schematic diagram showing an example input and output of an ALM according to an embodiment of this specification.

[0021] Figure 3 is a schematic diagram showing an example residue pair communication in an example computer-implemented system configured for protein structure prediction according to an embodiment of this specification.

[0022] Figure 4 is a table showing example dataset statistics for protein structure prediction according to an embodiment of this specification.

[0023] Figure 5 is a table showing the accuracy performance of different example protein structure prediction models in antibody structure prediction according to an embodiment of this specification.

[0024] Figure 6 are two tables showing the accuracy performance of different example protein structure prediction models in complementarity determining region (CDR) loop structure prediction according to an embodiment of this specification.

[0025] Figure 7A drawing showing an example of a protein structure predicted by an example computer-implemented system configured for protein structure prediction and other baseline predictions according to an embodiment of this specification.

[0026] Figure 8 A drawing showing an example of a protein structure predicted by xTrimoABFold and other baseline predictions according to an embodiment of this specification.

[0027] Figure 9 A chart showing example experimental results of the antibody structure prediction performance of an example computer-implemented system configured for protein structure prediction with and without focal loss according to an embodiment of this specification.

[0028] Figure 10 A schematic diagram showing another example computer-implemented system configured for protein structure prediction according to an embodiment of this specification.

[0029] Figure 11 Another table showing the accuracy performance of different example protein structure prediction models on antibody structure prediction according to an embodiment of this specification.

[0030] Figure 12 A drawing showing an example of a protein structure predicted by xTrimoABFold++ and other baseline predictions according to an embodiment of this specification.

[0031] Figure 13 A flowchart showing an example of a protein structure prediction process according to an embodiment of this specification.

[0032] Figure 14 A block diagram showing an example computer-implemented system for providing computing functions related to the algorithms, methods, functions, processes, flows, and programs described according to an embodiment of this disclosure.

[0033] Figure 15 Depicts example modules of a device according to an embodiment of this specification.

[0034] The same reference numerals and marks in the various drawings represent the same elements. Detailed Description

[0035] This specification describes protein structure prediction techniques based on machine learning or artificial intelligence (AI) technologies, such as antibody structure prediction. The described techniques can be applied to fields such as antibody engineering, drug design, and / or discovery.

[0036] In some embodiments, techniques are described for predicting, perturbing, or otherwise identifying protein structures, particularly antibody structures. A protein can be defined or specified by one or more amino acid chains or sequences in two dimensions (2D), three dimensions (3D), or higher dimensions. The amino acid sequence can include, for example, a long polypeptide, a short polypeptide, or a peptide. When amino acids are linked by peptide bonds in a sequence, these amino acids can be referred to as amino acid residues or simply residues. Thus, an amino acid sequence or amino acid chain is also referred to as an amino acid sequence or residue sequence.

[0037] The structure of a protein defines the three-dimensional (3D) configuration of the atoms in the protein's amino acid sequence. In some embodiments, the structure of a protein can be defined or represented by the values of structural parameters, such as the positions and angles of the atoms in the protein's amino acid sequence. For example, the structural parameters of a protein can include the 3D coordinates of the atoms and / or the relative translation and rotation between the atoms in the protein.

[0038] An antibody can include, for example, a protein used by the immune system to identify and neutralize foreign objects such as pathogenic bacteria and viruses. An antibody recognizes or otherwise corresponds to an antigen. For example, an antibody can include one or more complementarity-determining regions (CDRs), where each CDR is specific for a particular epitope on the antigen, such that the two structures can bind precisely. In this application, the terms "antigen" or "antibody" are broad enough to encompass one or more of a protein, a peptide, or another type of amino acid sequence.

[0039] Antibodies are an important type of protein for disease diagnosis and treatment. The structure of an antibody is closely related to its function. Therefore, antibody structure prediction, which aims to predict the 3D coordinates of the atoms in an antibody, is crucial in biological and medical applications such as protein engineering, modifying antigen-binding affinity, and identifying the epitopes of specific antibodies. However, manual experimental methods such as X-ray crystallography are both time-consuming and expensive.

[0040] The techniques described provide a computer-implemented solution based on machine learning or artificial intelligence (AI) techniques for predicting protein structures, particularly antibody structures. The techniques described include example models, architectures, or systems (collectively referred to as "systems") configured to predict antibody structures from antibody sequences using an antibody language model (ALM). An example system is referred to as "xTrimoABFold", which will be described in more detail below in conjunction with Figure 1 more detail. Different variants or extensions of xTrimoABFold are also described. For example, one variant is referred to as "xTrimoABFold++", which will be described in more detail below in conjunction with Figure 10 more detail.

[0041] Traditional protein structure prediction techniques typically rely on MSA to predict the structure of a target protein sequence. MSA refers to the process or result of sequence alignment of three or more biological sequences. The MSA of amino acid sequences can include using computational sequence alignment techniques (such as progressive alignment construction) to align an amino acid sequence (such as a target antibody sequence) with multiple other amino acid sequences (such as sequences from other homologous proteins). MSA involves computationally expensive MSA searches.

[0042] The described techniques are based on non-MSA or MSA-free protein structure prediction techniques. The described techniques use ALM via a Transformer model, for example, to learn an information representation of an antibody. ALM can mine homologous sequence information without complex manual preparation of MSA. In some embodiments, the described techniques use ALM to generate single and pairwise representations instead of MSA.

[0043] Compared with MSA-based protein structure prediction techniques, the described techniques can also improve prediction accuracy. Different from general proteins, antibodies are not naturally evolved but bind to specific antigens and evolve specifically (rapidly and unidirectionally). The MSA of antibodies, especially in the complementarity-determining regions (CDRs), is not always available or reliable, which may impair the accuracy of the model on antibody data.

[0044] In addition, the described techniques employ a pre-trained ALM to extract information from a single sequence, which performs better than protein structure prediction techniques using general protein language models (PLMs) trained on protein databases. In some embodiments, the described techniques specifically train an ALM based on antibody sequences for antibody applications. For example, the ALM is trained or fine-tuned on a large-scale observed antibody space (OAS) database. ALM can learn more specific language information and can perform more powerful representations for downstream tasks related to antibodies than general PLMs.

[0045] In some embodiments, for protein structure prediction, a template structure can be an auxiliary information to improve the quality of the structure model. The described techniques also include a computationally efficient template search algorithm designed based on sequence modality and / or structure modality. For example, a cross-modal homologous structure search algorithm is designed to search for templates and provide a good starting point for antibody structure prediction.

[0046] In some embodiments, the described techniques can train an overall model by solving an optimization problem to minimize a loss function to predict an antibody structure in an end-to-end manner. For example, the described techniques can use a structure prediction model including an evoformer and a structure module (e.g., those similar to those in AlphaFold2) to learn an antibody structure in an end-to-end manner. In some embodiments, the described techniques introduce several forms of loss functions that can provide more accurate prediction results. For example, the described techniques also introduce a Domain Specific Focal Loss on the complementarity-determining regions (CDRs) of an antibody, and / or a differentiable root mean square deviation (RMSD) loss to supplement or replace the frame alignment point loss, thereby better modeling the difference between the predicted structure and the accurate structure of the antibody. In some embodiments, one or more losses (e.g., the Domain Specific Focal Loss or RMSD loss on the CDR) can be used during the training and / or fine-tuning of the model. In some embodiments, one or more losses (e.g., the Domain Specific Focal Loss or RMSD loss on the CDR) are used only during fine-tuning, rather than during model training. Compared with the prior art, the described techniques can achieve better prediction performance.

[0047] In some embodiments, the described techniques can improve computational efficiency and achieve higher antibody prediction accuracy, especially on the CDRs of antibodies. The described techniques can be applied to scenarios such as industrial high-throughput drug design, which is impractical or infeasible for the prior art. Although some examples are described with respect to antibody structure prediction, which is important in drug discovery, the described techniques can be applied to general protein prediction and complex prediction. In some embodiments, compared with the prior art, the described techniques can improve the accuracy and efficiency of antibody structure prediction, making it a valuable tool for de novo antibody design and enabling further improvements in immunological theory.

[0048] In some embodiments, the described techniques can help better understand antibody structure and its paratopes to facilitate a mechanistic understanding of its function. The described techniques can facilitate the design of a novel antibody whose paratope can bind to a specific antigen with the correct epitope. In some embodiments, the described techniques can facilitate the generation, synthesis, screening, modification, or otherwise design of proteins and more accurately and efficiently predict the structure of proteins.

[0049] The technologies described in this disclosure may produce additional or different technical effects. In some embodiments, the described technologies may be implemented as software-implemented applications or software packages that can effectively predict the structure of target proteins. Compared with other computer-aided protein structure prediction technologies, the described technologies can reduce the computational load and improve computational efficiency. Experiments that have been conducted show that the described technologies are significantly superior (more than 30% improvement in RMSD) to AlphaFold2 and other state-of-the-art PLM-based technologies, such as OmegaFold, HelixFold-Single, and IgFold, while being 151 times faster than AlphaFold2.

[0050] Figure 1 FIG. 4 is a schematic diagram showing an example computer-implemented system 100 configured for protein structure prediction according to an embodiment of this specification. In some embodiments, the example computer-implemented system 100 provides an antibody structure prediction process based on the AlphaFold2 architecture without performing computationally expensive MSA searches. The example computer-implemented system 100 provides non-MSA or MSA-free protein structure prediction. In this specification, the example computer-implemented system 100 is referred to as "xTrimoABFold".

[0051] In some embodiments, xTrimoABFold 100 takes an amino acid sequence (also referred to as a residue sequence) 110 as input and generates a fine-grained antibody structure prediction 160 as output.

[0052] In some embodiments, xTrimoABFold 100 uses a pre-trained ALM 130 to generate a residue encoding 125 and an attention weight encoding 135, and initializes a single representation 175 and a pairwise representation 185 with the transformed results of the residue encoding 125 and the attention weight encoding 135 respectively, which can compensate for the loss of MSA homology information.

[0053] In some embodiments, a structural template that simulates the homologous structure of the target antibody can provide good prior information for structure prediction. In some embodiments, xTrimoABFold 100 may additionally use a template search algorithm to find a structural template 140 based on the sequence of the target antibody and / or the coarse-grained predicted structure of the target antibody. xTrimoABFold with template search may be referred to as xTrimoABFold+Tmpl. Features extracted from the structural template (referred to as template features) 165 can be merged into the transformed result of the residue encoding 125 (preliminary single representation 145) and the transformed result of the attention weight encoding (preliminary pairwise representation 155) to obtain a single representation 175 and a pairwise representation 185 respectively.

[0054] The single representation 175 and the paired representation 185 are input into the structure prediction model 150 to predict the fine-grained 3D prediction structure 160. In some embodiments, the structure prediction model 150 includes a combination of an encoder and a decoder. As Figure 1 shown in the example, the encoder can be a Transformer-based encoder that mixes information between the single representation and the paired representation to obtain updated single and paired representations. An example of the encoder is the evoformer 152 similar to that used in AlphaFold2. In some embodiments, the decoder can be a structure module that converts the abstract representation into specific 3D atomic coordinates. As shown in the example architecture 100, the decoder can be a structure module 154 similar to that used in AlphaFold2. In some embodiments, for further refinement, the structure prediction model 150 can iteratively update the input of the encoder by recycling the output of the encoder and the output of the decoder.

[0055] For the single representation, the pre-trained ALM (e.g., ALM 130) takes a single sequence (e.g., the residue sequence 110) as input and generates a residue (token) - level representation (e.g., the residue encoding 125). The residue - level representation can be used, through appropriate transformation, as the initial value of the single representation 175 of the subsequent encoder (e.g., evoformer 152).

[0056] Figure 2 FIG. 200 is a schematic diagram of an example input 210 and output 250 of the ALM 230 in an example computer - implemented system (e.g., xTrimoABFold 100) configured for antibody structure prediction according to an embodiment of the present specification. In some embodiments, the ALM 230 can be an example implementation of the ALM 130, or another computer - implemented system configured for antibody structure prediction. In some embodiments, the ALM 230 can be a deep machine - learning model including multiple neural network blocks, such as blocks 232, 234, and 236. In some embodiments, each block of the ALM 230 can be a self - attention network including one or more self - attention layers.

[0057] For an input x, the output z of the ALM can be represented as follows:

[0058] z = ALM(x), z ∈ R N×d lm (1 - 1)

[0059] where x = {x 1 , x 2 , ···, x N} represents the residue sequence (e.g., the residue sequence 110), N refers to the number of residues in the given protein, dlm is the hidden layer size of the ALM, where ALM represents the pre-trained ALM.

[0060] In some embodiments, the residue sequence can be a sequence of amino acid type identifiers (IDs) (e.g., represented by letters A, R, M, F, G, etc.). Each amino acid can correspond to a d lm -dimensional embedding, e.g., based on one-hot encoding. Thus, N amino acids correspond to an N×d lm embedding. In this case, before the ALM, there can be an embedding layer that maps the amino acid type ID to a d lm -dimensional embedding (e.g., a 1×d lm vector), and the x input to the ALM in formula (1-1) can be an embedding of size N×d lm .

[0061] In some other embodiments, the x input to the ALM can be a sequence of amino acid type IDs of size N×1. The ALM can include an embedding layer as its first layer that maps the amino acid type ID to a d lm -dimensional embedding. The ALM can also include other layers, such as self-attention layers, to update the embedding output by the first layer.

[0062] Taking the residue sequence 110 as an example of the residue sequence of protein x, the output z of the ALM can be an example of residue encoding 125.

[0063] The output z of the ALM can be used for the following calculation of the preliminary single representation (e.g., preliminary single representation 145):

[0064] s 0 = Linear(z), s 0 ∈ R N×ds (1-2)

[0065] where s 0 is the preliminary single representation, d s is the hidden layer size of the corresponding single representation in the subsequent encoder (e.g., evoformer 152), and Linear refers to the linear layer of a neural network (e.g., a fully convolutional neural network (FCNN)) used to convert the output z into the preliminary single representation. In some embodiments, without using a structure template in structure prediction, s 0 can be directly used as the initial single representation of the subsequent encoder; in some embodiments, when using a structure template in structure prediction, s 0 can be combined with the template features to obtain the initial single representation.

[0066] In some embodiments, the input 210 of the ALM 230 can be a token sequence. In some embodiments, the input 210 can be an amino acid sequence or a residue sequence, comprising a plurality of amino acids or residues, such as residue sequence 110. As Figure 2 shown in the example, the input 210 contains N = 5 residues, i.e., x = {A, R, M, F, G}. Each residue can be regarded as a token, and the ALM 230 can generate an embedding for each residue in the residue sequence 210. In Figure 2 the example shown, the output 250z of the ALM 230 contains 5 embeddings 252, 254, 256, 258, and 260, corresponding to the 5 residues A, R, M, F, and G respectively. In this example, the dimension of each embedding can be 1×d lm , and the dimension of the output z 250 is 5×d lm .

[0067] In some embodiments, the ALM 230 employs a multi - head self - attention mechanism, where each token can obtain information from other tokens, which can be regarded as residue - pair communication. For pairwise representation, the attention weights of the multi - head self - attention mechanism in the ALM are rich in prior knowledge about the relationships between residues, such as positional information, which can be combined into a preliminary single representation 155 through an adaptive transformation.

[0068] For example, the ALM can have a multi - head self - attention structure (e.g., an ALM with L attention layers and H attention heads in each layer). The h - th attention head of the l - th layer has learnable parameters These parameters represent the learnable parameters corresponding to queries, keys, and values in the self - attention neural network (i.e., the ALM in this example). In some embodiments, each residue can be represented by its respective embedding. For each attention head of each layer, the embedding corresponding to a residue in the input residue sequence 110 can serve as at least two roles, namely as a query and a key, for updating its own embedding and helping to update the embedding of another residue. For example, the input to the multi - head attention layer of the l - th layer of the ALM can be an embedding x l (including ), where corresponds to the embedding of residue i in the residue sequence corresponding to N residues. The multi - head attention layer of the l - th layer of the ALM with H attention heads can process x l and obtain x out , x out can be directly used as or transformed into x l+1 (including ), which can be input into the multi - head attention layer of the (l + 1) - th layer of the ALM.

[0069] In some embodiments, the ALM is used to generate a preliminary pairwise representation p0 It can be formalized as follows: A h,l = softmax(B h,l ),(2 - 4) p 0 = Linear(q),(2 - 6)

[0070] where Q i h,l and K j h,l respectively represent the query and key vectors / embeddings of residues i and j in the k-th attention head of the l-th layer, a ij represents the relative position encoding between residue i and residue j (e.g., a ij can represent the relative position of residue i and residue j in the residue sequence, which can be a learnable embedding), A h,l represents the attention weight matrix obtained by the h-th attention head of the l-th layer, represents the (i, j) element of matrix A h,l , represents the (i, j) element of matrix B h,l ,q ij represents the (i, j) element of matrix q ∈ R N ×N×HL ,p 0 ∈ R N×N×dp ,d p is the hidden layer size corresponding to a single representation in the encoder.

[0071] In addition, there is another learnable parameter in the l-th layer that can be used to generate x out ,for example, as follows:

[0072] where x out’ can be obtained through V 1,l A 1,l ,V 2,l A 2,l ,…,V H,l A H,l ,for example, by concatenation, where:

[0073] x out can be directly used as or converted to x l+1 。In some embodiments, the conversion includes, for example, normalization and / or feed-forward.

[0074] The above calculation can be regarded as an example of residue pair communication, as this step involves the multi-head query-key product of residue pairs. For example, given a pair of amino acid residues i and j in the input residue sequence 110, the multi-head query-key product Q i k,l (K j k,l ) T is calculated. For example, if the ALM has L = 10 layers and H = 3 attention heads per layer, q ij can be a vector of size HL = 30, where the first 3 elements of q ij (e.g., elements 0 - 2) correspond to the attention weights of the 3 attention heads in the first layer, the second 3 elements of q ij (e.g., elements 3 - 5) correspond to the attention weights of the 3 attention heads in the second layer, and so on. In some embodiments, q ij can include the attention weights of the H attention heads in the L layers concatenated, collected, or otherwise combined.

[0075] Figure 3 is an example schematic diagram 300 of residue pair communication in an example computer-implemented system (e.g., xTrimoABFold 100) configured for protein structure prediction according to an embodiment of the present specification. In this example schematic diagram 300, the attention weight encoding 335 (e.g., attention weight encoding 135) of the multi-head self-attention mechanism in the ALM can include a second embedding (e.g., q ij ) obtained when the amino acid residue A (e.g., residue i) is used as a query and the amino acid residue V (e.g., residue j) is used as a key in the multi-head self-attention mechanism.

[0076] In some embodiments, the structural template can provide good prior information for structure prediction. Different from previous works such as AlphaFold2 that searched for templates through MSA-based algorithms (e.g., HHSearch that detected templates through Hidden Markov Model (HMM)-HMM alignment between the query and the target database), the present disclosure introduces a template search algorithm without MSA. This template search algorithm does not rely on MSA and is more efficient in terms of memory and computation. In some embodiments, the template search algorithm can be a cross-modal homology search algorithm that searches for templates from both the sequence and structure perspectives without MSA.

[0077] For example, xTrimoABFold+Tmpl employs a cross-modal template search algorithm to search for homologous structures in both the sequential and structural modalities. The cross-modal template search algorithm includes a sequence modality search (also known as a sequential modality search) 122 and a structural modality search 124. The sequence modality search 122 searches for the structures of one or more sequences similar to the input amino acid sequence 110 in a template database. For example, when using the structural modality search 124, the coarse-grained structure 120 can be part of the input. The structural modality search 124 searches for one or more structures similar to the input coarse-grained structure 120 in a template database. The template databases used in the sequence modality search 122 and the structural modality search 124 can be the same database or different databases. In some embodiments, xTrimoABFold+Tmpl can use a single-modal template search.

[0078] In some embodiments, the template search algorithm can be performed in a protein structure database or an antibody database. In some embodiments, before performing the template search, a protein structure database and / or an antibody database can be constructed and used as a structural template database.

[0079] For the sequence modality search 122, considering the idea that similar antibody sequences may have similar 3D structures, a similarity score or an alignment score (e.g., a similarity score based on sequence alignment) can be used to search for the structures of sequences similar to the target antibody sequence in the template database as templates. An example similarity score function is formalized as follows:

[0080] Sim(x 1 ,x 2 ) = Align(x 1 ,x 2 )) / max(len(x 1 ),len(x 2 )), (3)

[0081] where x 1 and x 2 are residue sequences, Align(.,.) is sequence alignment, representing the maximum number of matching residues between two amino acid sequences (e.g., Align('GVI', 'GIV') = 2). Various existing algorithms (e.g., the Needleman-Wunsch algorithm) can be used for sequence alignment calculation. Additional or different formulas or algorithms can also be used as the similarity score or for calculating the similarity score.

[0082] In some embodiments, sequential modality search first filters out all sequences with similarity scores within a certain range (e.g., within the range (0.4, 0.95)), and restricts the available templates to a certain number Tse (e.g., Tse = 10) of sequences with the highest similarity scores to the target antibody sequence. Subsequently, the structures corresponding to these top Tse sequences will be regarded as part of the template candidates for subsequent training or inference.

[0083] In some embodiments, in terms of the efficiency of the search algorithm, sequential modality search is more efficient than MSA-based algorithms. Sequential modality search can provide real-time search and batch search. In some embodiments, real-time search can search for templates of the target sequence within 1 second through a parallel search algorithm. In some embodiments, real-time search divides the template database into N workers parts and conducts parallel search to select N workers *T se candidates, and then sorts the searched candidates according to the similarity scores through merge sort. Since merge sort is a stable algorithm, the results of each real-time search can be guaranteed to be the same. Finally, the top Tse structures are selected from the sorted homologous structures as templates. In some embodiments, batch search can compress the time cost of template search for a single sequence to the millisecond level by parallel searching and storing a large number of sequences.

[0084] Structure modality search 124 focuses on finding similar structures in the database based on the coarse-grained structure 120 of the target antibody, even if the sequences of these structures may not match the target antibody. The coarse-grained structure 120 can be an estimated, predicted, or otherwise obtained structure, which is used as an initial or baseline structure template to search for similar structures. In some embodiments, the coarse-grained structure 120 can be configured as a default structure (e.g., based on knowledge of structures similar to the target antibody structure, or knowledge of structures that provide a good starting point for the target antibody). In some embodiments, the coarse-grained structure 120 can be the structure prediction result obtained based on the target antibody sequence through another structure prediction algorithm or model.

[0085] Compared with sequential modality search 122, structure modality search 124 can use the same or different similarity scores. In some embodiments, similar to sequential modality search 122, the similarity scores between the coarse-grained structure of the target antibody and the structures in the template database (e.g., template database 115) are calculated. Various existing algorithms or tools applicable to structure alignment comparison (such as the FoldSeek tool) can be used to calculate the alignment scores. Structure modality search 124 can determine a maximum number T st , (e.g., T st= 10). In some embodiments, structures with too high similarity (e.g., greater than 0.95 or other thresholds) are removed to exclude the target antibody itself. The finally obtained top-T st structures can be added to the template candidate set.

[0086] After cross-modal template search, a total of T template candidates can be obtained. In some embodiments, since there may be duplicates in the search results of the two modalities, T is less than or equal to T se + T st . T, T se and T st values can be configured. For example, when T = 4, T se = 2 and T st = 2, during inference, 4 templates can be selected from the candidate sets of the first 2 sequence modality templates and the first 2 structure modality templates. In some embodiments, during the training step, a certain number (e.g., min(Uniform[0, T], S)) of templates can be randomly selected from this restricted set of T templates, where S can also be configured. For example, S = 4. In some embodiments, the structures selected by the two search algorithms contain more homologous structure information, so higher sampling probabilities can be assigned to these structures.

[0087] In some embodiments, the features extracted from the structure template (referred to as template feature 165) can be merged into the conversion result of the residue encoding 125 (preliminary single representation 145) and the conversion result of the attention weight encoding 135 (preliminary pairwise representation 155) to obtain a single representation 175 and a pairwise representation 185 respectively. For example, a template encoder (e.g., the template encoder of AlphaFold2) can be used to encode the template structure into two types of template features, namely template angle features and template pair features. Then the template angle features and template pair features are respectively merged into the preliminary single representation and the preliminary pairwise representation, and the formalization is as follows: s^ 0 = Concat(s 0 , f ta ), (4 - 1) p^ 0 = p 0 + f tp ; (4 - 2)

[0088] where f ta ∈R T×N×ds , s^ 0 ∈R (T+1)×N×ds , p^ 0 , f tp ∈R N×N×dp , f ta and f tpThey are the template angles and pair features, s^ 0 and p^ 0 are the single and paired representations with template features, and T is the number of templates. In some embodiments, methods similar to AlphaFold2 can be used to extract f ta and f tp . For example, f ta can be constructed by concatenating: template_aatype, template_torsion_angles, template_alt_torsion_angles, and template_torsion_angles_mask. f tp can include the concatenation of the residue features template_distogram, template_unit_vector, and some residue features converted to pair features.

[0089] s^ 0 and p^ 0 can be used as the input to the encoder of the structure prediction model 150. In some embodiments, the evoformer 152 of AlphaFold2 can be used as the encoder to model the complex information in the initial single and paired representations. Note that the column-wise gated self-attention of the evoformer 152 can exchange the sequence information modeled by the ALM 130 with the structural information of the template 140. The structure module 154 can employ some geometric transformation operators, such as Invariant Point Attention (IPA), to predict the 3D structure of the protein end-to-end. In this example, the evoformer 152 includes 48 blocks, and the structure module 154 includes 8 blocks. In some other embodiments, the evoformer and the structure module can include different numbers of blocks. For example, when the embedding predicted by the ALM is good, the number of blocks in the evoformer can be less, such as 1 block. In addition, a recycling mechanism 170 is adopted to iteratively refine the predicted structure 160.

[0090] In some embodiments, the xTrimoABFold 100 is trained end-to-end to optimize the objective function or minimize the loss function. Compared with the loss functions used by AlphaFold2 (including the Frame Aligned Point Error (FAPE) and some auxiliary losses), the loss function of the xTrimoABFold 100 (a non-MSA or MSA-free structure prediction system) removes the loss regarding the masked MSA.

[0091] In some embodiments, the loss function used by the xTrimoABFold 100 can be formalized as follows:

[0092] L train = 0.5L FAPE + 0.5L aux + 0.3L dist + 0.01L conf , (5)

[0093] where L FAPE denotes the FAPE of all atoms in the amino acid sequence, and L aux is the average FAPE and torsional loss of the intermediate structure only on C α , and L dist is the average cross-entropy loss predicted by the distance histogram, and L conf is the model confidence loss. These losses can be calculated according to existing methods (such as the methods disclosed in AlphaFold2).

[0094] In some embodiments, the loss function of xTrimoABFold 100 may include other loss / error / distance metrics. For example, since the structure of the complementarity-determining regions (CDRs) in antibodies is usually more difficult to predict than other framework regions (FRs), the loss function may further include a CDR focal loss. In some embodiments, the CDR focal loss can be used both during training and fine-tuning of xTrimoABFold. In some embodiments, the CDR focal loss can be used only for fine-tuning xTrimoABFold after it has been trained using a loss function that does not include the CDR focal loss. In some embodiments, this variant of xTrimoABFold that uses the CDR focal loss during fine-tuning but not during training is referred to as xTrimoABFold-FL (Focal Loss). In one example, the CDR focal loss is expressed as:

[0095] where x i and x true i are the predicted and true 3D coordinates of atom i in the CDR region, respectively, and T j and T j true represent the SE(3) transformations calculated based on x i and x true , respectively, including rotation and translation denotes the Hadamard product, N atoms CDR denotes the number of atoms in the antibody CDR region, and N frames is the number of local frames. Using Lfine-tune Performing fine - tuning helps xTrimoABFold pay more attention to difficult CDR regions. In this example, d clamp and Z are both set to This means that if d ij is greater than then d ij is set to because any greater distance is considered not beneficial for prediction. In some embodiments, d clamp and Z can be set to other values to improve prediction performance.

[0096] In some embodiments, the loss function can further include an RMSD loss to supplement or replace the FAPE loss (and / or other losses). The RMSD loss can be a more accurate metric because the FAPE loss is an upper bound of the RMSD. In some embodiments, a differentiable RMSD loss is developed to improve prediction accuracy:

[0097] where N atom is the number of atoms, x i pred , x i gt are the predicted and true 3D coordinates respectively, and T align is their SE(3) transformation. Different from the FAPE loss that performs transformation at the frame level, here T align can be performed at the global level of the entire amino acid sequence.

[0098] In some embodiments, one or more protein structure databases can be collected, created, downloaded, received, or otherwise obtained, for example, for template search, and / or for training the ALM, and / or configuring other components of a computer - implemented system for protein structure prediction (such as xTrimoABFold 100 or xTrimoABFold++1000). In an experiment, two large - scale datasets were created. The first is the 19K antibody structure dataset 105 as Figure 1 shown. A total of 18937 antibody data were obtained, which include amino acid sequences and structures selected from the RCSB Protein Data Bank (PDB) released before April 13, 2022. The specific selection for structures and sequences is as follows. First, each PDB file is split into single chains and then selected. On the one hand, among the 19736 BCR chains in the PDB, those without a structure resolution value or with a structure resolution greater than Samples are taken to maintain the quality of the structural data. On the other hand, for sequences, samples with an empty sequence or a repetition rate of a certain amino acid in the sequence exceeding 90% are filtered out. In addition, the sequences are also de-duplicated, and samples with a lower structural resolution are retained. After these filtering processes, 18,937 antibody data are obtained as the antibody structure dataset 105. In one example implementation, among these data, the data containing 18,470 samples released before January 17, 2022 are used as the training set, while the other 470 samples are used as the test set.

[0099] In some embodiments, the antibody structure dataset 105 is used as the training dataset for xTrimoABFold (and its variants). In the training phase, antibody data (including antibody sequences and corresponding actual structures) of training antibodies can be selected from the antibody structure dataset 105 to obtain their coarse-grained structures, and template candidates can be determined by the above-mentioned sequence search based on the antibody sequence and / or structure mode search based on the coarse-grained structure. After the template search, T templates can be selected from the template candidates. Based on the antibody sequence and the templates of the training antibodies, the initial xTrimoABFold (for example, an untrained model with initial model parameters, or a model whose parameters have been updated through several training iterations but have not been fully trained) can be used to predict the structures of the training antibodies. For example, based on the techniques described in the present disclosure, the loss between the predicted structure and the actual structure of the training antibody can be calculated. Then, the model parameters of xTrimoABFold are updated according to the loss. The above process can be repeated for the antibody data of other training antibodies in the training database.

[0100] The second dataset is the 501K Protein Structure Database. The entire protein database can be downloaded from the RCSB PDB. After filtering out the missing structure files, a total of 593,491 protein chains can be obtained. Subsequently, as described above, parts that do not meet the specifications of structural resolution and sequence similarity are removed, and duplicate examples are also removed. Finally, the 501K Protein Structure Database is obtained, which includes a total of 501,533 protein chains. This protein structure database can be used as a template database, such as the template database 115, for template search.

[0101] Figure 4 Table 1 in [the relevant part] shows the statistical data of the example datasets of the 19K antibody structure dataset 105 and the template database 115 containing 501K protein structures according to the embodiments of the present specification.

[0102] The xTrimoABFold method was compared with several of the latest state-of-the-art protein structure prediction methods: AlphaFold2, OmegaFold, PLM-based HelixFold-Single, ESMFold, ALM-based IgFold, and DeepAb, which were used as the baselines for comparison. For AlphaFold2, five different models were used for inference, and the structure with the highest predicted local distance difference test (pLDDT) confidence was selected for benchmarking. In some experiments, a variant of the xTrimoABFold model, called xTrimoABFold-ESM, was trained. xTrimoABFold-ESM replaced the ALM with the general protein language model ESM2. The performance of xTrimoABFold-ESM was worse than that of xTrimoABFold, indicating that the ALM is a better choice than the general protein language model.

[0103] To evaluate the quality of antibody structure prediction, root mean square deviation (RMSD), TM-Score, GDT_TS, and GDT_HA can be used as evaluation metrics. These two values can be calculated on the backbone heavy atoms after aligning their respective framework residues through DeepAlign. To evaluate the performance of the model in predicting the complementarity-determining region (CDR) loops, which are more difficult to predict, three CDR regions of the antibody structure were extracted and evaluated for these regions based on local and global alignments respectively. In the local alignment scheme, two local CDR regions were aligned, and the RMSD was calculated on the local alignment matrix. In the global alignment scheme, an alignment matrix was generated using two complete antibody structures, and the RMSD was calculated based on this alignment matrix.

[0104] In some embodiments, the TM-score is calculated as follows:

[0105] where L target is the sequence length of the target protein, and L common is the number of residues that appear in both the template and the target structure.

[0106] In an example experiment, for the ALM 130, AntiBERTy (version 0.0.5, installed from PyPI) was used to generate residue-level representations. AntiBERTy is a BERT-based pre-trained protein language model trained on the Observed Antibody Space (OAS) with 5.58M natural antibody sequences. The hidden dimension of the ALM is 512 and the feed-forward dimension is 2048. AntiBERTy contains 8 layers with 8 attention heads per layer. Overall, AntiBERTy contains approximately 26M trainable parameters. In some embodiments, during the training phase, the gradient backpropagation of the ALM can be blocked, and only the evoformer 152 and the structure module 154 are trained. In some embodiments, the Adam optimizer with a learning rate of 1e-3, β 1 = 0.9, β 2 = 0.999, ∈ = 8 and a weight decay of 0 is used. In some embodiments, the gradient is clipped using a threshold of 10e9. In this example experiment, the model was trained for 25 epochs with a fixed batch size of 8 on 8 NVIDIA A100 GPUs, taking 46 hours. Similar to AlphaFold2, the cropped size of the sequence is set to 256. Since the single-sequence representation of the ALM replaces the MSA representation, compared to AlphaFold2, the InputEmbedder, ExtraMSAEmbedder, and ExtraMSAStack as well as the masked MSA loss are removed. When performing the structure modal search, Foldseek was used, which is capable of making fast and sensitive comparisons of large structure sets. The 3Di Gotoh-Smith-Waterman was selected as the alignment type, and max-seq was set to 2000.

[0107] The main experimental results of xTrimoABFold compared with the baseline include two parts: one is the performance of the model on the evaluation metrics, and the other is the time efficiency. Figure 5 and Figure 6 Tables 2, 3, and 4 in respectively show the accuracy performance of the model on antibody structure prediction and CDR loop structure prediction. For simplicity, only the RMSD and TM-score of three CDR loops are shown. Specifically, Table 2 shows the experimental results of antibody structure prediction on the test dataset with a 95% confidence interval. xTrimoABFold-ESM refers to a method similar to xTrimoABFold, except that the pre-trained PLM ESM2 with 15 billion parameters (the current largest PLM) is used to replace the pre-trained ALM. The results show that the ALM is more suitable for antibody structure prediction.

[0108] For the protein structure prediction of the CDR loop, which is known to be a region difficult for models to accurately predict, xTrimoABFold also performs excellently. Figure 6 Tables 3 and 4 in Figure 6 show the RMSD of all models based on local alignment and global alignment respectively. Specifically, Table 3 shows the experimental results of antibody CDR loop structure prediction based on local alignment on the test dataset with a 95% confidence interval. Table 4 shows the experimental results of antibody CDR loop structure prediction based on global alignment on the test dataset with a 95% confidence interval. As shown in the figure, xTrimoABFold has improvements in CDR1 and CDR2 loops compared to HelixFold-Single and IgFold trained based on large-scale protein language models and ALM. xTrimoABFold performs best in the CDR3 loop, which has been proven to be a difficult-to-predict region due to its high variability and conformational diversity.

[0109] Figure 7 is Chart 700, showing an example of experimental results on the antibody structure prediction time of different methods for different lengths of amino acid sequences in the test dataset. Specifically, Figure 7 shows the median time of MSA search, AlphaFold2, and xTrimoABFold. AlphaFold2 predicts protein structures based on MSA, resulting in a large amount of time consumption. Compared with AlphaFold2, xTiomoABFold is a MSA-free model that predicts protein structures based on a single amino acid sequence through ALM. As Figure 7 shown, xTrimoABFold is 151 times faster than AlphaFold2, indicating that xTrimoABFold can overcome the bottleneck of time efficiency in protein structure prediction and is capable of quickly performing large-scale antibody structure prediction. Compared with the baseline, xTrimoABFold achieves better time efficiency in structure prediction and can quickly perform antibody structure prediction.

[0110] In terms of antibody structure prediction performance, xTrimoABFold significantly outperforms all baselines on the test dataset. In terms of RMSD, xTrimoABFold has improvements of 37.20%, 40.06%, 34.08%, 38.05%, 86.28%, and 93.52% compared to AlphaFold2, OmegaFold, HelixFold-Single, ESMFold, IgFold, and DeepAb respectively, as shown in Table 2. At the same time, this trend also appears in other evaluation metrics. Compared with protein structure prediction methods based on PLM and MSA, xTrimoABFold achieves state-of-the-art performance in antibody structure prediction.

[0111] Figure 8 It is Plot 800, which shows an example of the protein structures predicted by xTrimoABFold and other baselines according to the embodiments of this specification. As shown, in terms of prediction accuracy, xTrimoABFold is superior to other baselines including AlphaFold2, OmegaFold, and ESMFold.

[0112] In the experiment, an ablation study was conducted to evaluate the performance improvement brought by introducing pre-trained ALM (such as the model based on AntiBERTy) and adding CDR focal loss when fine-tuning the xTrimoABFold model.

[0113] xTrimoABFold uses pre-trained ALM (such as the AntiBERTy-based model) to generate residue-level representations, which contain more specific antibody information compared with general protein language models such as OmegaPLM and ESM-2. In the example ablation study, the variant xTrimoABFold-ESM of xTrimoABFold was used to verify the selection of ALM instead of the general protein language model. xTrimoABFold-ESM replaced ALM with the large-scale protein language model ESM-2 trained on 250 million protein sequences while keeping other parts of xTrimoABFold unchanged. In the experiment, xTrimoABFold-ESM was trained on the same dataset as xTrimoABFold, and its prediction performance was worse than that of xTrimoABFold, as shown in Table 2, which shows the performance improvement brought by the pre-trained ALM in xTrimoABFold.

[0114] To prove the effectiveness of the focal loss, an ablation study was conducted on another variant xTrimoABFold+FL of xTrimoABFold. As mentioned above, xTrimoABFold+FL added the focal loss to the loss function of xTrimoABFold for fine-tuning. The performance of xTrimoABFold+FL is also shown in Table 2. The experiment found that the designed focal loss can effectively improve the performance and reduce the variance.

[0115] In addition, in another experiment, ten samples were randomly selected from the test dataset, and the performance of xTrimoABFold before and after adding the CDR focal loss was compared. Figure 9 It is Chart 900, which shows the example experimental results on the antibody structure prediction performance of xTrimoABFold with and without the focal loss. Figure 9In these examples shown, compared with xTrimoABFold without CDR focal loss, xTrimoABFold with CDR focal loss (such as xTrimoABFold+FL) has varying degrees of reduction in the RMSD value between the predicted structure and the true structure. The performance improvement brought by CDR focal loss indicates that focal loss is effective in antibody structure prediction, especially for CDR loops that are difficult to predict by conventional models.

[0116] Another ablation experiment was also conducted to demonstrate the effectiveness of the templates found through cross-modal homologous structure search. Another variant of the xTrimoABFold model, xTrimoABFold+Tmpl, was used. xTrimoABFold+Tmpl integrates cross-modal homologous structure search into xTrimoABFold and adds template feature 140 to the single representation 175 and the pairwise representation 185. Table 2 shows the performance of xTrimoABFold+Tmpl, and its prediction accuracy has improved compared with xTrimoABFold. The experimental results of xTrimoABFold+Tmpl show that the templates found through cross-modal homologous structure search can effectively reduce variance and improve prediction accuracy.

[0117] Figure 10 is a schematic diagram showing another example computer-implemented system 1000 configured for protein structure prediction according to an embodiment of the present specification. The example computer-implemented system 1000 provides protein structure prediction based on non-MSA or without MSA. The example computer-implemented system 1000 can be regarded as Figure 1 another variant of xTrimoABFold 100 in. In this specification, the example computer-implemented system 1000 is referred to as "xTrimoABFold++". Compared with Figure 1 xTrimoABFold 100 in, xTrimoABFold++1000 does not need to perform template search, which further reduces the computational complexity.

[0118] In some embodiments, xTrimoABFold++1000 takes an amino acid sequence (also called residue sequence) 1010 as input and generates a fine-grained structure prediction 1060 as output. xTrimoABFold++1000 can include two subsystems, the ALM subsystem 1005 and the structure prediction model 1050.

[0119] The ALM subsystem 1005 uses the pre-trained ALM 1030 to model homologous antibody sequences and learn the representation of antibodies, such as a single representation, without performing an expensive MSA search. ALM 1030 can be associated with Figure 1 orFigure 2 similar to the ALM130 or 230 described in. The ALM 1030 receives the input amino acid sequence 1010 and outputs the final hidden state 1025 of the ALM 1030. In some embodiments, the final hidden state 1025 can be represented as a vector, matrix, tensor, or other embedding form. The final hidden state 1025 can be converted to a single representation 1175, for example, by a fully convolutional neural network (FCNN) 1045 or other methods, such that the single representation 1175 has dimensions suitable for input to the subsequent structure prediction model 1050 (e.g., the input to the encoder 1052 of the structure prediction model 1050). Using the examples described with respect to formulas (1-1) and (1-2) and Figure 2 the example described in, if the hidden layer size of the encoder 1052 is d s , the dimension of the final hidden state 1025 can be N×d lm , and the FCNN 1045 is used to convert the final hidden state 1025 to a single representation with dimensions of N×d s .

[0120] The ALM 1030 can also be used to obtain a pairwise representation 1185 to be input into the subsequent structure prediction model 1050. In some embodiments, residue pair communication 1015 can be used to obtain multi-head attention weights 1035, for example, according to the example techniques described above with respect to formulas (2-1)-(2-8) and Figure 3 the example described in or other techniques. The multi-head attention weights 1035 can be converted to a pairwise representation 1185, for example, by another fully convolutional neural network (FCNN) 1055 or other methods, such that the pairwise representation 1185 has dimensions suitable for input to the subsequent structure prediction model 1050 (e.g., the input to the encoder 1052 of the structure prediction model 1050). Using the examples described with respect to formulas (2-1)-(2-8) and Figure 3 the example described in, the dimension of the multi-head attention weights 1035 can be N×N*HL, and the FCNN 1045 is used to convert the multi-head attention weights 1035 to a pairwise representation with dimensions of N×N×dp.

[0121] The structure prediction model 1050 can be the same as or different from the structure prediction model 150 in Figure 1 . In some embodiments, the structure prediction model 1050 has a deep learning architecture. In some embodiments, the structure prediction model 1050 includes a combination of an encoder 1052 (e.g., the evoformer in AlphaFold2) and a decoder 1054 (e.g., the structure module in AlphaFold2). As Figure 10In the example shown, the encoder 1052 can use row-wise gated self-attention 3, triangular updates, and triangular self-attention, and the decoder 1054 uses invariant point attention to learn amino acid interactions and geometric representations. In this example, the encoder 1052 includes 48 blocks, while the decoder 1054 includes 8 blocks.

[0122] Similar to xTrimoABFold 100, xTrimoABFold++ 1000 can be end-to-end trained using the various loss functions described above. For example, the loss function of xTrimoABFold++ 1000 can include a CDR focus loss and an RMSD loss, as described in formulas (9) and (10), to supplement or replace some of the losses used in existing protein structure prediction models.

[0123] Figure 11 Table 5 in shows the accuracy performance of different example protein structure prediction models (including xTrimoABFold++ 1000) according to the embodiments of the present specification in antibody structure prediction. As shown, xTrimoABFold++ is superior to all baselines in antibody structure prediction, especially in the CDR-H3 region of the antibody dataset consisting of 68 antibody complexes.

[0124] Figure 12 is Plot 1200, which shows an example of a protein structure predicted by xTrimoABFold++ and other baselines according to the embodiments of the present specification. Plot 1200 shows an example of a target protein, PDB 7WVM_B, which is the light chain of cemiplimab for PD-1. As shown, xTrimoABFold++ is superior to other baselines in terms of RMSD.

[0125] Figure 13 is a flowchart of an example process 1300 for protein structure prediction according to the embodiments of the present specification. Process 1300 can be an example of a protein structure prediction algorithm without MSA executed by a data processing device, such as Figure 1 the computer-implemented system 100 in Figure 10 or the computer-implemented system 1000 in Figure 14 In some embodiments, the data processing device can be a system of one or more computers located at one or more locations and appropriately programmed according to the present specification. For example, Figure 14 the computer-implemented system 1400 in

[0126] In some embodiments, Figure 13The example process 1300 shown can be modified or reconfigured to include more, fewer, or different operations, which can be performed in the order shown or in a different order. In some cases, one or more operations can be repeated or iterated, for example, until a termination condition is reached. In some implementations, Figure 13 one or more of the individual operations shown can be performed as multiple individual operations, or Figure 13 one or more subsets of the operations shown can be combined and performed as a single operation.

[0127] Although Figure 13 it is described with reference to antibodies and antibody sequences (e.g., a target antibody sequence), the example process 1300 can be more broadly applied to protein structure prediction, e.g., based on a target protein sequence.

[0128] In step 1310, the data processing device inputs, configures, identifies, obtains, or otherwise receives a target antibody sequence that includes an amino acid (or amino acid residue) sequence. The target antibody sequence can represent an antibody specified by the amino acid sequence. The example process 1300 can be used to predict the structure of an antibody specified by an amino acid sequence. The target antibody sequence can be the example amino acid sequence or residue sequence 110 or 1010.

[0129] In some embodiments, receiving the target antibody sequence includes receiving data representing the target antibody sequence. For example, the data representing the target antibody sequence can include an embedding representing the amino acids in the target antibody sequence. An "embedding" can be an ordered set of numerical values, such as a vector, matrix, or tensor of numerical values. Thus, the target antibody sequence can be represented as a vector, matrix, tensor, or other form or data structure. In some embodiments, the target antibody sequence includes additional data related to the target antibody sequence, such as embedding data (e.g., one-hot encoded data). For example, different amino acids can be represented by different letters, such as A to Z. For each amino acid, the corresponding embedding data can be a word2vec vector or other type of embedding code. Thus, an antibody composed of amino acids can be represented by the respective letters of the amino acids and / or the embedding data. In some embodiments, amino acids and antibodies can be represented in another way or data structure for computer processing.

[0130] In step 1320, the target antibody sequence is input into the ALM. The ALM can be a protein language model trained according to the antibody sequence. The ALM can be the example ALM 130, 230, or 1030.

[0131] For example, the ALM can be trained using an antibody database that includes antibody sequences or consists only of antibody sequences. In some embodiments, the ALM can be pre-trained, e.g., independently or separately from an overall model configured for protein structure prediction. In some embodiments, the ALM can be part of an overall model configured for protein structure prediction (e.g., xTrimoABFold 100 or xTrimoABFold++1000), and be trained or fine-tuned using the loss function of the overall model (e.g., one or more of the loss functions in formulas (5), (8), (9), or (10)). In the latter case, the parameters of the first machine learning model and the second machine learning model can be trained or updated based on the gradient of the loss function of the overall model configured for protein structure prediction.

[0132] In some embodiments, the ALM can be a neural network, such as a self-attention model that includes multiple self-attention neural network layers (also referred to as self-attention layers). Various types of self-attention models or architectures can be used as the basis for training the ALM. In some embodiments, the ALM is pre-trained according to the BERT architecture (e.g., the AntiBERTy architecture), using an antibody database.

[0133] In step 1330, residue encodings and attention weight encodings are obtained using the ALM without performing a multiple sequence alignment (MSA). The residue encodings are used to generate a single representation to be input into a structure prediction model (e.g., structure prediction model 150 or 1050). The attention weight encodings are used to generate a pairwise representation to be input into a structure prediction model (e.g., structure prediction model 150 or 1050).

[0134] The residue encoding can be a residue-level data representation that includes corresponding first embeddings for each amino acid in the target antibody sequence. For example, according to Figure 1 , Figure 2 and Figure 10 the example techniques described therein, the ALM outputs the corresponding first embeddings by taking the target antibody sequence as the input to the ALM. For example, the residue encoding can be example residue encoding 125, output 250, or last hidden state 1025. The residue encoding can be represented by a numerical vector, matrix, tensor, or other data structure. Different from traditional protein structure prediction methods that generate a single representation based on MSA embeddings, the residue encoding is output by the ALM without performing an MSA, thereby improving the computational efficiency of process 1300.

[0135] The attention weight encoding can be a pairwise data representation that includes corresponding second embeddings corresponding to a pair of amino acids in the target antibody sequence. If the number of residues in the sequence is N, the number of pairs and the size of the attention weight encoding is N*N. The corresponding second embeddings are calculated from the attention weights of the self-attention layer of the ALM. For example, the attention weight encoding can include the exemplary attention weight encoding 135 or the attention weight 1035, for example, according to Figure 1 , Figure 3 and Figure 10 the exemplary techniques described in.

[0136] The attention weight encoding can be represented by a numerical vector, matrix, tensor, or other data structure. Different from traditional protein structure prediction methods that generate pairwise representations based on MSA embeddings, the attention weight encoding is generated based on the attention weights of the ALM without using MSA embeddings, thereby improving the computational efficiency of process 1300.

[0137] In some embodiments, if the ALM includes L self-attention layers and each self-attention layer includes H attention heads, the attention weight encoding can include a second embedding corresponding to a pair of amino acids i and amino acid j in the target antibody sequence (e.g., q in formula (2-4)) ij ). Without performing MSA, obtaining the second embedding q ij using the ALM includes: obtaining the attention weights of the H attention heads of each of the L self-attention layers when amino acid i is used as the query and amino acid j is used as the key; and connecting these attention weights to obtain the second embedding q ij , for example, according to formula (2-4). In some embodiments, the embedding q ij can include the attention weights of the L layers of H attention heads connected, collected, or otherwise combined. When amino acid i is used as the query and amino acid j is used as the key in the ALM, the attention weights can be calculated based on the query-key product (e.g., Q i k,l (K j k,l ) T ). The attention weights can be A h,l , which is calculated, for example, according to the softmax operation shown in formula (2-3), another normalization operation B h,l , or another variant of B h,l , or B h,l itself.

[0138] In step 1340, the residue encoding and the attention weight encoding are converted into a single representation and a pairwise representation. The single representation may include data representing the features of individual residues in the amino acid sequence of the target antibody sequence. The pairwise representation may include data representing the features of residue pairs in the amino acid sequence of the target antibody sequence. The single representation and the pairwise representation may be represented in the form of vectors, matrices, tensors, or other data structures. The single representation and the pairwise representation may be the initial single representation (e.g., initial single representation 175 or 1175) and the initial pairwise representation (e.g., initial pairwise representation 185 or 1185) to be input into the structure prediction model. In some embodiments, the data processing device converting the residue encoding and the attention weight encoding into a single representation and a pairwise representation includes: converting the residue encoding into a single representation through a first machine learning model (e.g., a first linear neural network layer (e.g., FCNN 1045)); and converting the attention weight encoding into a pairwise representation through a first machine learning model (e.g., a second linear neural network layer (e.g., FCNN 1055)). The first machine learning model and the second machine learning model may be trained separately or as part of an overall model configured for protein structure prediction (e.g., xTrimoABFold 100 or xTrimoABFold++ 1000), and are trained using the loss function of the overall model (e.g., one or more of the loss functions in formulas (5), (8), (9), or (10)). In the latter case, the first machine learning model and the second machine learning model may be trained, for example, by updating the parameters based on the gradient of the loss function of the overall model configured for protein structure prediction.

[0139] In some embodiments, example process 1300 further includes a template search to identify one or more template candidates that are similar to the target antibody structure. The one or more template candidates may be used to initialize the single representation and the pairwise representation before they are input into the structure prediction model. In some embodiments, steps 1325, 1335, and 1345 related to the template search may be performed.

[0140] In step 1325, a template search is performed based on the target antibody sequence without performing a multiple sequence alignment (MSA) to find one or more template candidates that are similar to the target antibody structure. The template search may use Figure 1The example cross-modal template search algorithm described in or other template search algorithms. For example, performing template search to find one or more template candidates includes: performing a sequential modality search in a first structure database to find a first structure template, where the antibody sequence corresponding to the first structure template is similar to the target antibody sequence; and performing a structural modality search in a second structure database to find a second structure template, where the structure of the second structure template is similar to the coarse-grained structure of the target antibody sequence. One or more template candidates include the first structure template and / or the second structure template. The first structure database and the second structure database can be the same or different.

[0141] In step 1335, template features (such as template feature 165) are obtained based on one or more template candidates. For example, the template features can be obtained by extracting matching features from one or more template candidates and adding them or otherwise incorporating them into the corresponding features of the single representation and the pairwise representation.

[0142] In step 1345, the template features are incorporated into the single representation and the pairwise representation generated in step 1340. For example, the single representation and the pairwise representation generated in step 1340 can be regarded as the generated preliminary single representation and preliminary pairwise representation, and the template features are added to the preliminary single representation and preliminary pairwise representation.

[0143] In some embodiments, process 1300 does not include any template search (such as any one of steps 1325, 1335, and 1345). In this case, before the single representation and the pairwise representation are input into the structure prediction model, the single representation and the pairwise representation do not contain template features.

[0144] In step 1350, the single representation and the pairwise representation are input into a structure prediction model (such as structure prediction model 150 or 1050). The parameters of the structure prediction model are trained or otherwise obtained based on a loss function that reflects the difference between the predicted structure of the antibody and the actual structure. For example, the parameters of the structure prediction model are trained by solving an optimization problem to minimize the loss function, for example, by updating the parameters based on the gradient of the loss function. The loss function can be one or more of the loss functions in formulas (5), (8), (9), or (10), or can include other or different losses. However, the loss function does not include the loss due to MSA. For example, the loss function includes the Frame Aligned Point Error (FAPE) loss and the torsion angle loss, as well as the loss for the Complementary Determining Region (CDR). This loss represents the difference between the predicted structure of the target antibody and the actual structure. As another example, the loss function also includes the differentiable Root Mean Square Deviation (RMSD) to supplement or replace the Frame Aligned Point Error (FAPE) loss between the predicted structure and the actual structure of the target antibody sequence.

[0145] In step 1350, a predicted structure of the target antibody is determined using a structure prediction model based on the single representation and the pairwise representation. For example, after training an overall model (such as xTrimoABFold 100 or xTrimoABFold++1000) including an ALM and a structure prediction model for protein structure prediction, the structure prediction model is used in the inference phase to determine the predicted structure of the target antibody. In some embodiments, the structure prediction model is used to iteratively determine the predicted structure of the target antibody sequence until a convergence or other termination condition (such as the number of iterations) is met.

[0146] In step 1360, the predicted structure of the target antibody is output. The predicted structure of the target antibody can be defined by values of a plurality of structural parameters, such as atomic positions and angles, to represent the 3D structure of the target antibody specified by the target antibody sequence. In some embodiments, experiments, tests, and further processing, such as drug discovery and design, can be performed based on the predicted structure of the target antibody.

[0147] Figure 14 FIG. 14 is a block diagram of an example computer-implemented system 1400 for providing computing functionality related to the algorithms, methods, functions, processes, flows, and programs described herein. For example, system 1400 can be an example of a data processing apparatus configured to perform protein structure prediction according to an embodiment of this specification. In the illustrated embodiment, system 1400 includes a computer 1402 and a network 1430.

[0148] The illustrated computer 1402 is intended to encompass any computing device, such as a server, a desktop computer, a notebook / laptop, a wireless data port, a smartphone, a personal digital assistant (PDA), a tablet, one or more processors in these devices, another computing device, or a combination of computing devices, including physical or virtual instances of computing devices, or a combination of physical or virtual instances of computing devices. Additionally, computer 1402 can include an input device, such as a keypad, a keyboard, a touch screen, another input device, or a combination of input devices that can accept user information, and an output device that conveys information related to the operation of computer 1402 in a graphical user interface (UI) (or GUI) or other UI, including digital data, visual, audio, another type of information, or a combination of information types.

[0149] The computer 1402 can act as a client, a network component, a server, a database, or another persistent storage, another role, or a combination of multiple roles in a distributed computing system to perform the subject matter described in this disclosure. The illustrated computer 1402 is communicatively coupled to the network 1430. In some embodiments, one or more components of the computer 1402 can be configured to operate in an environment, including a cloud-computing-based, on-premises, global, another environment, or a combination of multiple environments.

[0150] At a high level, the computer 1402 is an electronic computing device capable of receiving, transmitting, processing, storing, or managing data and information related to the subject matter described. According to some embodiments, the computer 1402 can also include or be communicatively coupled to a server, including an application server, an email server, a web server, a caching server, a streaming data server, another server, or a combination of servers.

[0151] The computer 1402 can receive requests (e.g., from a client software application executing on another computer 1402) via the network 1430 and respond to the received requests by processing the received requests using a software application or a combination of software applications. Additionally, requests can also be sent to the computer 1402 from an internal user (e.g., from a command console or via another internal access method), an external or third party, or other entities, individuals, systems, or computers.

[0152] Each component of computer 1402 can communicate using system bus 1403. In some embodiments, any or all components of computer 1402, including hardware, software, or a combination of hardware and software, can interface on system bus 1403 through application programming interface (API) 1412, service layer 1413, or a combination of API 1412 and service layer 1413. API 1412 can include specifications of routines, data structures, and object classes. API 1412 can be language-independent or language-dependent, and can refer to a complete interface, a single function, or even a set of APIs. Service layer 1413 provides software services for computer 1402 or other components (whether shown or not) communicatively coupled to computer 1402. All service consumers using service layer 1413 can access the functions of computer 1402. Software services, such as those provided by service layer 1413, provide reusable defined functions through defined interfaces. For example, the interface can be software written in JAVA, C++, another computing language, or a combination of computing languages, and provide data in extensible markup language (XML) format, another format, or a combination of formats. Although shown as an integrated component of computer 1402, alternative embodiments can show API 1412 or service layer 1413 as independent components relative to other components of computer 1402 or other components (whether shown or not) communicatively coupled to computer 1402. Additionally, any or all parts of API 1412 or service layer 1413 can be implemented as sub-modules or sub-components of another software module, enterprise application, or hardware module without departing from the scope of the present disclosure.

[0153] Computer 1402 includes interface 1404. Although shown as a single interface 1404, two or more interfaces 1404 can be used depending on the specific requirements, expectations, or specific embodiments of computer 1402. Interface 1404 is used for computer 1402 to communicate with another computing system (whether shown or not) communicatively linked to network 1430 in a distributed environment. Generally, interface 1404 is operable to communicate with network 1430 and includes logic encoded in software, hardware, or a combination of software and hardware. More specifically, interface 1404 can include software that supports one or more communication protocols related to communication, such that network 1430 or the hardware of interface 1404 is operable to transmit physical signals inside and outside the shown computer 1402.

[0154] The computer 1402 includes a processor 1405. Although shown as a single processor 1405, two or more processors 1405 may be used depending on the specific requirements, expectations, or particular embodiments of the computer 1402. Generally, the processor 1405 executes instructions and manipulates data to perform the operations of the computer 1402 and any algorithms, methods, functions, processes, flows, and programs described in this disclosure.

[0155] The computer 1402 also includes a database 1406, which may store data for the computer 1402, another component communicatively linked to the network 1430 (whether shown or not), or a combination of the computer 1402 and another component. For example, the database 1406 may be an in-memory database, a conventional database, or another type of database that stores data consistent with this disclosure. In some embodiments, depending on the specific requirements, expectations, or particular embodiments of the computer 1402 and the described functionality, the database 1406 may be a combination of two or more different database types (e.g., a hybrid in-memory database and a conventional database). Although shown as a single database 1406, two or more similar or different types of databases may be used depending on the specific requirements, expectations, or particular embodiments of the computer 1402 and the described functionality. Although the database 1406 is shown as a component of the computer 1402, in alternative embodiments, the database 1406 may be external to the computer 1402.

[0156] For example, the database 1406 may store data related to the embodiments of this specification. For example, the database 1406 may store one or more databases (e.g., the antibody structure dataset 105 and the template database 115), training data 1416 for training the ALM and / or configuring the overall model for protein structure prediction (e.g., xTrimoABFold 100 or xTrimoABFold++1000), a pre-trained ALM 1418 (e.g., ALM 130, 230, or 1030), a structure prediction model 1422 (e.g., the structure prediction model 150 or 150), or another component or sub-model of the overall model configured for protein structure prediction (e.g., FCNN 1045 or 1055), a target protein 1423 (e.g., the target protein sequence 110, 210, or 1010), a predicted protein structure 1428, or other test / experimental results 1432.

[0157] The computer 1402 also includes a memory 1407, which can store data for the computer 1402, another component or multiple components communicatively linked to the network 1430 (whether shown or not), or a combination of the computer 1402 and another component. The memory 1407 can store any data consistent with the present disclosure. In some embodiments, the memory 1407 can be composed of two or more different types of memory (such as a combination of semiconductor and magnetic storage) according to the specific requirements, expectations or specific embodiments of the computer 1402 and the functions described. Although shown as a single memory 1407, two or more similar or different types of memory can be used according to the specific requirements, expectations or specific embodiments of the computer 1402 and the functions described. Although the memory 1407 is shown as a component of the computer 1402, in alternative embodiments, the memory 1407 can be external to the computer 1402.

[0158] The application 1408 is an algorithmic software engine that provides corresponding functions according to the specific requirements, expectations or specific embodiments of the computer 1402, particularly with respect to the functions described in the present disclosure. For example, the application 1408 can be implemented as one or more components, modules or applications. Additionally, although shown as a single application 1408, the application 1408 can be implemented as multiple applications 1408 on the computer 1402. Further, although shown as a part of the computer 1402, in alternative embodiments, the application 1408 can be external to the computer 1402.

[0159] The computer 1402 can also include a power supply 1414. The power supply 1414 can include a rechargeable or non-rechargeable battery, which can be configured to be user-replaceable or non-user-replaceable. In some embodiments, the power supply 1414 can include a power conversion or management circuit (including charging, standby or other power management functions). In some embodiments, the power supply 1414 can include a power plug that allows the computer 1402 to be plugged into a wall socket or other power source, such as to power the computer 1402 or charge a rechargeable battery.

[0160] Any number of computers 1402 can be associated with or external to the computer system including the computer 1402, and each computer 1402 communicates through the network 1430. Additionally, the terms "client", "user" or other suitable terms can be used interchangeably as needed without departing from the scope of the present disclosure. Moreover, the present disclosure contemplates that many users can use one computer 1402, or one user can use multiple computers 1402.

[0161] Figure 15FIG. 0 is a schematic diagram of an example module of apparatus 1500 according to an embodiment of the present specification. Apparatus 1500 may be an example embodiment of a data processing apparatus for protein structure prediction according to an embodiment of the present specification. Apparatus 1500 may correspond to the above embodiments, and apparatus 1500 includes the following parts: a receiving module 1501 for receiving a target antibody sequence of a target antibody containing an amino acid sequence; a first input module 1502 for inputting the target antibody sequence into an antibody language model (ALM), where ALM is a protein language model trained according to antibody sequences, and ALM includes a plurality of self-attention layers; an obtaining module 1503 for obtaining residue encoding and attention weight encoding using ALM without performing multiple sequence alignment (MSA); a conversion module 1505 for converting the residue encoding and attention weight encoding into a single representation and a pairwise representation; a second input module 1506 for inputting the single representation and the pairwise representation into a structure prediction model, where the parameters of the structure prediction model are trained based on a loss function reflecting the difference between the predicted structure and the actual structure of the antibody; a determining module 1507 for determining the predicted structure of the target antibody using the structure prediction model based on the single representation and the pairwise representation; and an output module 1508 for outputting the predicted structure of the target antibody.

[0162] In some embodiments, apparatus 1500 further includes the following parts: a search module 1504 for performing a template search based on the target antibody sequence without performing multiple sequence alignment (MSA) to find one or more template candidates similar to the target antibody structure before inputting the single representation and the pairwise representation into the structure prediction model; and a second obtaining module 1509 for obtaining template features based on the one or more template candidates; and wherein converting the residue encoding and the attention weight encoding into a single representation and a pairwise representation includes: converting the residue encoding and the attention weight encoding into a preliminary single representation and a preliminary pairwise representation; and merging the template features into the preliminary single representation and the preliminary pairwise representation to obtain the single representation and the pairwise representation.

[0163] In some embodiments, performing a template search to find one or more template candidates includes: performing a sequential mode search in a first structure database to find a first structure template, where the antibody sequence corresponding to the first structure template is similar to the target antibody sequence; and performing a structure mode search in a second structure database to find a second structure template, where the structure of the second structure template is similar to the coarse-grained structure of the target antibody sequence, and where one or more template candidates include one or more of the first structure template or the second structure template, and where the coarse-grained structure is a default structure or a structure predicted by another structure prediction algorithm or another structure prediction model.

[0164] In some embodiments, the ALM is pre-trained according to the BERT architecture using an antibody database, and the antibody database consists of antibody sequences.

[0165] In some embodiments, the ALM includes L self-attention layers, each self-attention layer includes H attention heads, and wherein the second embedding qij corresponds to a pair of amino acids i and amino acids j in the target antibody sequence, and wherein the data processing device obtains the attention weight encoding using the ALM without performing MSA, including: obtaining the attention weights of the H attention heads of each of the L self-attention layers when amino acid i is used as a query and amino acid j is used as a key; and connecting these attention weights to obtain the second embedding qij.

[0166] In some embodiments, converting the residue encoding and the attention weight encoding into single representations and pairwise representations includes: converting the residue encoding into single representations through a first linear neural network layer; and converting the attention weight encoding into pairwise representations through a second linear neural network layer; wherein the parameters of the first linear neural network layer and the second linear neural network layer are updated based on the gradient of the loss function.

[0167] In some embodiments, the loss function does not include the loss due to MSA.

[0168] In some embodiments, the loss function includes a frame alignment point error (FAPE) loss, a torsion angle loss, and a loss for the complementarity-determining region (CDR).

[0169] In some embodiments, the loss function includes a differentiable root mean square deviation (RMSD) in addition to the frame alignment point error (FAPE) loss.

[0170] In some embodiments, before inputting the single representation and the pairwise representation into the structure prediction model, the single representation and the pairwise representation do not contain template features.

[0171] Embodiments of the described subject matter may include one or more features, either alone or in combination. For example, in a first embodiment, a computer-implemented method for antibody structure prediction includes one or more of the following steps: receiving a target antibody sequence of a target antibody containing an amino acid sequence; inputting the target antibody sequence into an Antibody Language Model (ALM), where the ALM is a protein language model trained based on antibody sequences and the ALM includes multiple self-attention layers; obtaining residue encoding and attention weight encoding using the ALM without performing a Multiple Sequence Alignment (MSA), where the residue encoding includes corresponding first embeddings output by the ALM for each amino acid in the target antibody sequence, and the attention weight encoding includes corresponding second embeddings corresponding to a pair of amino acids in the target antibody sequence calculated from the attention weights of the self-attention layers of the ALM; converting the residue encoding and attention weight encoding into a single representation and a pairwise representation; inputting the single representation and the pairwise representation into a structure prediction model, where the parameters of the structure prediction model are trained based on a loss function reflecting the difference between the predicted structure and the actual structure of the antibody; using the structure prediction model to determine the predicted structure of the target antibody based on the single representation and the pairwise representation; and outputting the predicted structure of the target antibody.

[0172] The above and other described embodiments may each optionally include one or more of the following features:

[0173] A first feature, combinable with any of the following features, stipulating that the ALM is pre-trained according to the BERT architecture using an antibody database consisting of antibody sequences.

[0174] A second feature, combinable with any of the following features, stipulating that the ALM includes L self-attention layers, each self-attention layer includes H attention heads, and where the second embedding qij corresponds to a pair of amino acids i and j in the target antibody sequence, and where obtaining the attention weight encoding using the ALM without performing an MSA includes: obtaining the attention weights of the H attention heads of each of the L self-attention layers when amino acid i is used as the query and amino acid j is used as the key; and concatenating these attention weights to obtain the second embedding qij.

[0175] A third feature, combinable with any of the following features, stipulating that converting the residue encoding and attention weight encoding into a single representation and a pairwise representation includes: converting the residue encoding into a single representation through a first linear neural network layer; and converting the attention weight encoding into a pairwise representation through a second linear neural network layer; where the parameters of the first linear neural network layer and the second linear neural network layer are updated based on the gradient of the loss function.

[0176] A fourth feature, combinable with any of the following features, stipulating that the loss function does not include a loss due to MSA.

[0177] The fifth feature, which can be combined with any of the following features, stipulates that the loss function includes a Frame Aligned Point Error (FAPE) loss, a twist angle loss, and a loss for Complementary Determining Regions (CDRs).

[0178] The sixth feature, which can be combined with any of the following features, stipulates that the loss function includes a differentiable Root Mean Square Deviation (RMSD) in addition to the Frame Aligned Point Error (FAPE) loss.

[0179] The seventh feature, which can be combined with any of the following features, stipulates that before inputting the single representation and the pairwise representation into the structure prediction model, the single representation and the pairwise representation do not contain template features.

[0180] The eighth feature, which can be combined with any of the following features, stipulates that before inputting the single representation and the pairwise representation into the structure prediction model, the computer-implemented method further includes: performing a template search based on the target antibody sequence to find one or more template candidates similar to the target antibody structure without performing a Multiple Sequence Alignment (MSA); and obtaining template features based on the one or more template candidates; and wherein converting the residue encoding and the attention weight encoding into the single representation and the pairwise representation includes: converting the residue encoding and the attention weight encoding into a preliminary single representation and a preliminary pairwise representation; and merging the template features into the preliminary single representation and the preliminary pairwise representation to obtain the single representation and the pairwise representation.

[0181] The ninth feature, which can be combined with any of the following features, stipulates that performing a template search to find one or more template candidates includes: performing a sequential mode search in a first structure database to find a first structure template, wherein the antibody sequence corresponding to the first structure template is similar to the target antibody sequence; and performing a structural mode search in a second structure database to find a second structure template, wherein the structure of the second structure template is similar to the coarse-grained structure of the target antibody sequence, and wherein one or more template candidates include one or more of the first structure template or the second structure template, and wherein the coarse-grained structure is a default structure or a structure predicted by another structure prediction algorithm or another structure prediction model.

[0182] In a second embodiment, a system includes: one or more processors; and one or more computer-readable memories coupled to the one or more processors, having instructions stored thereon that are executable by the one or more processors to perform the method of the first embodiment and an optional combination of one or more of the above features.

[0183] In a third embodiment, a device for identifying a target protein corresponding to a target protein. The device includes one or more modules (e.g., as Figure 15The described module) for performing the method in the first embodiment and optional combinations of one or more of the above features.

[0184] The systems, devices, modules or units shown in the foregoing embodiments can be implemented by using computer chips or entities, or by using products with specific functions. Typical example devices are computers (a computer can be a personal computer), laptop computers, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email receiving and sending devices, game consoles, tablet computers, wearable devices, or any combination of these devices.

[0185] Regarding the implementation process of the functions and roles of each module in the device, reference can be made to the implementation process of the corresponding steps in the foregoing method. For the sake of brevity, the detailed content is omitted here.

[0186] Since the device embodiments and the method embodiments are basically corresponding, for the relevant parts, reference can be made to the relevant descriptions in the method embodiments. The foregoing device embodiments are merely examples. The modules described as separate parts may or may not be physically separated, and the parts shown as modules may or may not be physical modules, may be located in one place, or may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions in this specification. Those of ordinary skill in the art can understand and implement the embodiments of the present application without creative efforts.

[0187] Referring again to Figure 15 , which can be interpreted as showing the internal functional modules and structure of a computing implementation device. This computing implementation device can be an example of a computing system configured to identify a target protein corresponding to a target protein. In essence, the execution subject can be an electronic device, which includes: one or more processors; and one or more computer-readable memories configured to store executable instructions of one or more processors. In some embodiments, one or more computer-readable memories are coupled to one or more processors, and programming instructions are stored thereon, which can be executed by one or more processors to execute the algorithms, methods, functions, processes, flows and programs described in this specification. This specification also provides one or more non-transitory computer-readable storage media, which are coupled to one or more processors, and instructions are stored thereon, and when these instructions are executed by one or more processors, cause one or more processors to perform the operations according to the method embodiments provided herein.

[0188] This specification further provides a system for implementing the methods provided herein. The system includes one or more processors, and a computer-readable storage medium coupled to the one or more processors, having instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations in accordance with the method embodiments provided herein.

[0189] Embodiments of the subject matter described in this specification, as well as the acts and operations, can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in one or more combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions encoded on a computer program carrier for execution by, or to control the operation of, a data processing apparatus. For example, the computer program carrier can include one or more computer-readable storage media having instructions encoded or stored thereon. The carrier can be a tangible non-transitory computer-readable medium, such as a magnetic, magneto-optical, or optical disk, a solid state drive, a random access memory (RAM), a read only memory (ROM), or other type of medium. Alternatively, or additionally, the carrier can be an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof, or can be a part of them. A computer storage medium is not a propagated signal.

[0190] A computer program, which may also be referred to as or described as a program, software, a software application, an application, a module, a software module, an engine, a script, or code, can be written in any form of programming language, including a compiled or interpreted language, or a declarative or procedural language; and it can be deployed in any form, including as a stand-alone program or as a module, a component, an engine, a subroutine, or other unit suitable for execution in a computing environment that can include one or more computers interconnected at one or more locations by a data communication network.

[0191] A computer program can (but need not) correspond to a file in a file system. A computer program can be stored in a part of a file that holds other programs or data, such as one or more scripts in a markup language document, in a single file dedicated to the program, or in multiple coordinated files, such as files that store one or more modules, subroutines, or portions of code.

[0192] A processor for executing a computer program includes, for example, general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive computer program instructions and data for execution from a non-transitory computer-readable medium coupled to the processor.

[0193] The term "data processing apparatus" encompasses all types of apparatus, devices, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The data processing apparatus may include dedicated logic circuitry, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or a graphics processing unit (GPU). In addition to hardware, the apparatus may also include code that creates an execution environment for computer programs, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.

[0194] The processes and logical flows described in this specification may be performed by a data processing apparatus as software, hardware, firmware, or a hybrid implementation thereof. For example, the processes and logical flows described in this specification may be performed by one or more computers or processors executing one or more computer programs, by operating on input data and generating output. These processes and logical flows may also be performed by dedicated logic circuitry, such as an FPGA, an ASIC, or a GPU, or by a combination of dedicated logic circuitry and one or more programmed computers.

[0195] A computer suitable for executing a computer program may be based on a general or special purpose microprocessor or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read only memory or a random access memory or both. The elements of a computer may include a central processing unit for executing instructions and one or more storage devices for storing instructions and data. The central processing unit and the storage devices may be supplemented by, or incorporated in, dedicated logic circuitry.

[0196] Typically, a computer will also include, or be operatively coupled to receive data from and transfer data to one or more storage devices. The storage devices can be, for example, magnetic, magneto-optical or optical disks, solid state drives, or any other type of non-transitory computer-readable medium. However, a computer does not necessarily have such devices. Thus, a computer can be coupled to one or more storage devices, such as one or more local and / or remote memories. For example, a computer can include one or more local memories as an integral component of the computer, or a computer can be coupled to one or more remote memories in a cloud network. Additionally, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a gaming console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name just a few.

[0197] Components can be "coupled" to each other in an exchangeable manner (e.g., electrically or optically), whether directly or through one or more intermediate components. If one component is integrated into another component, then those components can also be "coupled" to each other. For example, a storage component (e.g., an L2 cache component) integrated into a processor is "coupled" to the processor.

[0198] To enable interaction with a user, embodiments of the subject matter described in this specification can be implemented on or configured to communicate with a computer that has a display device (e.g., a liquid crystal display (LCD) monitor) for displaying information to the user, and an input device through which the user can provide input to the computer, such as a keyboard and a pointing device (e.g., a mouse, a trackball, or a touchpad). Other types of devices can also be used to enable interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on the user device, or by interacting with an application running on the user device (e.g., a smartphone or an electronic tablet). Further, a computer can interact with a user by sending a text message or other form of message to a personal device (e.g., a smartphone running a messaging application) and receiving a response message from the user.

[0199] This specification associates the term "configured to" with systems, devices, and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system is installed with software, firmware, hardware, or a combination thereof that, when operating, cause the system to perform those operations or actions. For one or more computer programs configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations or actions. For a special-purpose logic circuit configured to perform particular operations or actions means that the circuit has electronic logic that performs those operations or actions.

[0200] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of the claims, which is determined by the claims themselves, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be removed, and the claim may be directed to a sub-combination or a variant of the sub-combination.

[0201] Similarly, although the operations depicted in the drawings and recited in the claims are in a particular order, this should not be understood to require that the operations must be performed in the particular order shown or in a sequential order, or that all of the operations shown must be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be understood to be required in all embodiments, and it should be understood that the described program components and systems may generally be integrated in a single software product or packaged into multiple software products.

[0202] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes depicted in the drawings do not necessarily need to be performed in the particular order shown, or in a sequential order, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A computer-implemented method for antibody structure prediction, wherein the predicted structure of a given antibody is defined by values of a plurality of structural parameters, the method comprises: A data processing device receives a target antibody sequence of a target antibody, the sequence comprising an amino acid sequence; The data processing device inputs the target antibody sequence into an Antibody Language Model (ALM), wherein the ALM is a protein language model trained based on antibody sequences, and the ALM comprises a plurality of self-attention layers; The data processing device uses the ALM without performing a Multiple Sequence Alignment (MSA) to obtain residue encoding and attention weight encoding, wherein: The residue encoding comprises corresponding first embeddings output by the ALM for each amino acid in the target antibody sequence; The attention weight encoding comprises corresponding second embeddings calculated from the attention weights of the self-attention layers of the ALM for a pair of amino acids in the target antibody sequence; The data processing device converts the residue encoding and the attention weight encoding into a single representation and a pairwise representation; The data processing device inputs the single representation and the pairwise representation into a structure prediction model, wherein the parameters of the structure prediction model are trained based on a loss function reflecting the difference between the predicted structure and the actual structure of the antibody; The data processing device uses the structure prediction model to determine the predicted structure of the target antibody based on the single representation and the pairwise representation; and The data processing device outputs the predicted structure of the target antibody.

2. The computer-implemented method according to claim 1, wherein, The ALM is pre-trained according to the Bidirectional Encoder Representations from Transformers (BERT) architecture using an antibody database, and the antibody database consists of antibody sequences.

3. The computer-implemented method according to claim 1, wherein, The ALM comprises L self-attention layers, each self-attention layer comprises H attention heads, and wherein, the second embedding q ij corresponds to a pair of amino acids i and amino acid j in the target antibody sequence, and wherein, the data processing device uses the ALM without performing an MSA to obtain attention weight encoding, comprising: When the amino acid i is used as a query and the amino acid j is used as a key, obtaining the attention weights of the H attention heads of each of the L self-attention layers; and Connect the attention weights to obtain the second embedding q ij .

4. The computer-implemented method according to claim 1, wherein, The data processing device converts the residue encoding and the attention weight encoding into a single representation and a pairwise representation, comprising: Converting the residue encoding into the single representation through a first linear neural network layer; and Converting the attention weight encoding into a pairwise representation through a second linear neural network layer; wherein the parameters of the first linear neural network layer and the second linear neural network layer are updated based on the gradients of the loss function.

5. The computer-implemented method according to claim 1, wherein, The loss function does not include the loss due to MSA.

6. The computer-implemented method according to claim 1, wherein, The loss function includes a Frame Alignment Point Error (FAPE) loss, a torsion angle loss, and a loss for the Complementary Determining Region (CDR).

7. The computer-implemented method according to claim 1, wherein, in addition to the Frame Alignment Point Error (FAPE) loss, the loss function further includes a differentiable Root Mean Square Deviation (RMSD).

8. The computer-implemented method according to claim 1, wherein, before inputting the single representation and the pairwise representation into the structure prediction model, the single representation and the pairwise representation do not contain template features.

9. The computer-implemented method according to claim 1, wherein, before the data processing device inputs the single representation and the pairwise representation into the structure prediction model, the computer-implemented method further includes: the data processing device performs a template search based on the target antibody sequence without performing a Multiple Sequence Alignment (MSA) to find one or more template candidates that are structurally similar to the target antibody; and the data processing device obtains template features based on the one or more template candidates; and wherein the data processing device converts the residue encoding and the attention weight encoding into the single representation and the pairwise representation, including: the data processing device converts the residue encoding and the attention weight encoding into a preliminary single representation and a preliminary pairwise representation; the data processing device merges the template features into the preliminary single representation and the preliminary pairwise representation to obtain the single representation and the pairwise representation.

10. The computer-implemented method according to claim 9, wherein, the data processing device performs a template search to find one or more template candidates, including: the data processing device performs a sequential modal search in a first structure database to find a first structure template, wherein the antibody sequence corresponding to the first structure template is similar to the target antibody sequence; and the data processing device performs a structural modal search in a second structure database to find a second structure template, wherein the structure of the second structure template is similar to the coarse-grained structure of the target antibody sequence, and wherein the one or more template candidates include one or more of the first structure template or the second structure template, and wherein the coarse-grained structure is a default structure or a structure predicted by another structure prediction algorithm or another structure prediction model.

11. A system for a software-implemented application for performing antibody structure prediction, the system comprises: one or more processors; and one or more computer-readable memories coupled to the one or more processors, storing instructions executable by the one or more processors to perform the method according to any one of claims 1 to 10.

12. A device for antibody structure prediction, the device comprising a plurality of modules for performing the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Protein structure prediction method, protein structure prediction device and medium

    CN114220479A

  • Protein structure reasoning method based on distributed technology

    CN115034393A

  • Protein Structure Prediction from Amino Acid Sequences Using Self-Attention Neural Networks

    US20210166779A1

Cited By

  • Targeted target epitope antibody design method

    CN120833853A

  • Method for predicting protein structure, medium and electronic equipment

    CN120853661A

  • Antibacterial peptide prediction method based on dual-channel sparse attention

    CN121354654A