In silico antibody library construction method and antigen-antibody binding affinity prediction model

WO2026168918A1PCT designated stage Publication Date: 2026-08-13AINB INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-03
Publication Date
2026-08-13

Smart Images

  • Figure KR2026001956_13082026_PF_FP_ABST
    Figure KR2026001956_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The present specification provides a method for constructing an in-silico antibody library, the method comprising the steps of: acquiring two or more pieces of antibody sequence information; and generating new antibody sequences by combining a variable heavy chain (VH) and a variable light chain (VL) of each antibody. The in-silico antibody library has the characteristics of: 1) not being so large that in-silico calculation costs are unmanageable; 2) not being so small that in-silico prediction becomes meaningless; and 3) including sufficiently various antibodies such that the possibility of discovering a good antibody is high. In addition, the present specification provides a model for predicting antigen-antibody binding affinity, wherein the model is formed to receive an antigen sequence and an antibody sequence and output antigen-antibody binding affinity, and integrates and utilizes chemical characteristics of the antigen sequence and the antibody sequence in a cross-attention mechanism. The antigen-antibody binding affinity prediction model integrates chemical characteristic information of each antigen sequence and antibody sequence into a process of modeling the interaction of each sequence, and thus can be effectively trained even in a situation where there is not much data, and can be applied in various ways.
Need to check novelty before this filing date? Find Prior Art

Description

Method for constructing an in silico antibody library and antigen-antibody binding affinity prediction model

[0001] This specification discloses a method for constructing an in-silico antibody library and a technology related to a model that takes an antigen sequence and an antibody sequence as input and predicts antigen-antibody binding affinity.

[0002] [Prior Art Literature]

[0003] Prior Literature 1: Ruofan Jin, Qing Ye, Jike Wang, Zheng Cao, Dejun Jiang, Tianyue Wang, Yu Kang, Wanting Xu, Chang-Yu Hsieh, Tingjun Hou, AttABseq: an attention-based deep learning prediction method for antigen-antibody binding affinity changes based on protein sequences, Briefings in Bioinformatics, Volume 25, Issue 4, July 2024, bbae304, https: / / doi.org / 10.1093 / bib / bbae304

[0004] Prior Literature 2: Yuan, Y., Chen, Q., Mao, J. et al. DG-Affinity: predicting antigen-antibody affinity with language models from sequences. BMC Bioinformatics 24, 430 (2023). https: / / doi.org / 10.1186 / s12859-023-05562-z

[0005] Conventional antibody library construction method

[0006] Traditionally, antibodies that specifically bind to antigens of interest have been identified using phage display methods or hybridoma production methods. This approach involves constructing a library by synthesizing large quantities of random antibodies, antibody variable regions, or other equivalents, and then experimentally selecting antibodies from this library that bind well and specifically to the antigen of interest. Traditional methods are time-consuming and labor-intensive because they require 1) the actual synthesis of numerous antibodies or their equivalents to build a library, and 2) the identification of antibodies that specifically bind to the antigen of interest through repetitive experiments. Even if antibodies that specifically bind to the antigen of interest are selected through this process, a problem remains: the selected antibodies may trigger an immune response in the human body or possess properties unfavorable for mass production, making them difficult to use as therapeutic agents. This is because the characteristic of specifically binding to the antigen of interest does not guarantee low immunogenicity or high productivity. Therefore, it is necessary to determine through further experiments whether the properties are suitable for therapeutic use. Determining this requires additional effort and cost. Consequently, it is very difficult to discover antibodies of interest through experiments using traditional methods without prior information, and the probability of success is not high.

[0007] Conventional antigen-antibody binding affinity prediction model

[0008] The following prior art documents disclose a model that predicts antigen-antibody binding affinity by taking antigen and antibody sequences as input:

[0009] Prior Literature 1: Ruofan Jin, Qing Ye, Jike Wang, Zheng Cao, Dejun Jiang, Tianyue Wang, Yu Kang, Wanting Xu, Chang-Yu Hsieh, Tingjun Hou, AttABseq: an attention-based deep learning prediction method for antigen-antibody binding affinity changes based on protein sequences, Briefings in Bioinformatics, Volume 25, Issue 4, July 2024, bbae304, https: / / doi.org / 10.1093 / bib / bbae304

[0010] Prior Literature 2: Yuan, Y., Chen, Q., Mao, J. et al. DG-Affinity: predicting antigen-antibody affinity with language models from sequences. BMC Bioinformatics 24, 430 (2023). https: / / doi.org / 10.1186 / s12859-023-05562-z

[0011] Despite the aforementioned prior research, the performance of antigen-antibody binding affinity prediction models published to date is not particularly outstanding. This is because there is a lack of training data.

[0012] The process of actually synthesizing antibodies, reacting them with antigens, and measuring antigen-antibody binding affinity is costly. Generally, rather than being performed on a large scale, this experiment is conducted on a limited basis, specifically for antibodies of interest with therapeutic potential. Therefore, data on antigen-antibody binding affinity is very limited. Approximately 10 per antigen 2 to 10 3 It is understood that antibody data of a certain scale has been released. Considering the amount of data generally required to train a generative model, this is woefully insufficient.

[0013] Therefore, there is a need for a new model capable of effective learning even with limited data.

[0014] This specification addresses the technical task of providing a method for constructing an in-silico antibody library that can reasonably reduce the search space, and a model that can predict antigen-antibody binding affinity using antibody sequence information and antigen sequence information.

[0015] To solve the above technical problem, this specification provides the following:

[0016] A method for constructing an in-silico antibody library, comprising the steps of acquiring two or more antibody sequence information; and generating new antibody sequences by combining the Variable Heavy Chain (VH) and Variable Light Chain (VL) of each antibody.

[0017] A method for deriving antibody candidates of interest by predicting the characteristics of each antibody included in the above-constructed in-silico antibody library;

[0018] An antigen-antibody binding affinity prediction model configured to receive antigen sequences and antibody sequences as input and output antigen-antibody binding affinity, and a model that utilizes the chemical characteristics of antigen sequences and antibody sequences by integrating them into a cross-attention mechanism;

[0019] A method for training the above antigen-antibody binding affinity prediction model; and

[0020] A method for predicting antigen-antibody binding affinity using the above antigen-antibody binding affinity prediction model.

[0021] The in-silico antibody library disclosed in this specification has the characteristics that 1) it is not so large that the in-silico computational cost is manageable, 2) it is not so small that the in-silico prediction is meaningless, and 3) it contains sufficiently diverse antibodies so that there is a high probability of discovering good antibodies.

[0022] The antigen-antibody binding affinity prediction model disclosed in this specification incorporates chemical characteristic information of each antigen sequence and antibody sequence into the process of modeling the interaction between each sequence, and has a structure that can be effectively learned even in situations where there is not much data, and has the characteristic of being applicable in various ways.

[0023] Figure 1 schematically illustrates a method for constructing an in-silico antibody library by randomly combining the sequence of a light chain variable region and the sequence of a heavy chain variable region from antibody information disclosed according to Experimental Example 1.1.1.

[0024] Figure 2 shows the results of calculating the human germline similarity and CNN-P scores of antibodies included in an in-silico antibody library according to Experimental Example 1.2.1 and plotting them on a two-dimensional plane. Each point on the graph represents a single antibody, with the horizontal axis representing the CNN-P score and the vertical axis representing the human germline similarity.

[0025] Figure 3 illustrates a graph showing the antigen-antibody binding affinity for antigens of Target 1, Target 2, and Target 3 for each antibody derived according to Experimental Example 1.2.2, predicted using the antigen-antibody binding affinity prediction model in Experimental Example 1.1.5. The top graph represents the predicted binding affinity for Target 1 for each antibody derived according to Experimental Example 1.2.2, the middle graph represents the predicted binding affinity for Target 2 for each antibody derived according to Experimental Example 1.2.2, and the bottom graph represents the predicted binding affinity for Target 3 for each antibody derived according to Experimental Example 1.2.2.

[0026] Figure 4 shows the results of verifying the binding affinity of the antibody candidates of Experimental Example 1.3.1 to Target 1 using an immunoassay according to Experimental Example 1.3.2. Each bar graph represents the results for one antibody.

[0027] Figure 5 shows the results of verifying the binding affinity of the antibody candidates of Experimental Example 1.3.1 to Target 2 or Target 3 using an immunoassay according to Experimental Example 1.3.2. The area in front of the dotted line represents the results for Target 2, and the area behind the dotted line represents the results for Target 3. Each bar graph represents the results for a single antibody.

[0028] Figure 6 is a graph showing the results of confirming whether the antibody candidate group for Target 1 specifically binds to Target 1 by measuring the binding affinity to antigens other than the target of interest (Target 1) according to Experimental Example 1.3.3. Each label on the horizontal axis of the graph represents one antibody, and the four bar graphs for each label represent, from left to right, the binding affinity to Target 1 (target of interest), the binding affinity to Target A, the binding affinity to Target B, and the binding affinity to Target C.

[0029] FIG. 7 is a flowchart illustrating one embodiment of the chemical cross-attention mechanism of this specification.

[0030] FIG. 8 is a flowchart illustrating one implementation example of multi-head attention calculation of this specification.

[0031] FIG. 9 is a diagram illustrating the structure of an embodiment model of this specification.

[0032] FIG. 10 is a flowchart illustrating one embodiment of a method for training an antigen-antibody binding affinity prediction model of this specification.

[0033] FIG. 11 is a flowchart illustrating a detailed implementation example of the process of training a model (s430) during the process of FIG. 10.

[0034] FIG. 12 is a schematic diagram of a computer device (600) implementing the method of this specification.

[0035] Figure 13 is a graph showing the change in the loss function of the training process step by step when training the model according to Experimental Example 2.2.

[0036] Figure 14 is a graph showing the change in the loss function of the verification process step by step when verifying the implemented model according to Experimental Example 2.2.

[0037] Figure 15 is a graph showing the change in F1 score by step in the implementation model according to Experimental Example 2.2.

[0038] Figure 16 is a graph showing the learning results of the antibody generation model according to Experimental Example 2.3. The X-axis represents the number of backpropagations of the generation model using reinforcement learning, and the Y-axis represents the results of evaluating whether specific antigens are bound using an antigen-antibody binding prediction model using CCA for antibody sequences generated by reinforcement learning.

[0039] The best modes for carrying out the invention are disclosed below by way of example. These include some embodiments of the invention disclosed herein, but not all embodiments. The embodiments described in this paragraph are merely illustrative and should not be understood as the only "best modes of the invention." A person skilled in the art would be able to conceive of many variations and more preferred embodiments of the examples described in this paragraph, and such should also be considered to be included in the best modes for carrying out the invention.

[0040] This specification discloses a method for training a model for predicting antigen-antibody binding affinity (hereinafter referred to as the training method), wherein the model is configured to receive an antigen sequence and an antibody sequence as input and output an antigen-antibody binding affinity, and the model includes a Chemical Cross Attention (CCA) mechanism for modeling the interaction between the antigen sequence and the antibody sequence, and the method comprises the following:

[0041] (a) The process of obtaining training data,

[0042] Here, the above training data includes multiple antigen sequence-antibody sequence pairs and antigen-antibody binding affinities corresponding to each pair;

[0043] (b) A process of identifying chemical characteristics for each antigen sequence and antibody sequence of the above training data,

[0044] Here, the chemical properties are information related to the intermolecular interactions between each amino acid of the antigen sequence and each amino acid of the antibody sequence; and

[0045] (c) A process of training the model using the training data obtained in (a) above and the chemical characteristics identified in (b) above,

[0046] Here, during the training process of the above model, a chemical cross-attention operation is performed between the characteristic representation of the antigen sequence and the characteristic representation of the antibody sequence during each input data forward propagation, and

[0047] The above chemical cross-attention operation performs a cross-attention operation between the information derived from the antigen sequence and the information derived from the antibody sequence, wherein

[0048] By applying an attention guide based on the above chemical properties to the calculated cross-attention scores, the cross-attention scores at positions related to intermolecular interactions are relatively amplified, and

[0049] A predicted value is derived based on the chemical attention value calculated as a result of performing the above chemical cross-attention operation, and

[0050] Calculate the loss function based on the above predicted values ​​and the actual antigen-antibody binding affinity, and

[0051] Adjust the model parameters in a way that minimizes the loss function.

[0052] In one embodiment, the chemical properties may include information selected from the following:

[0053] Information on amino acid pairs capable of forming hydrogen bonds with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence;

[0054] Information on amino acid pairs that can interact with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence;

[0055] Information on amino acid pairs that can interact hydrophobically with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence;

[0056] Information regarding amino acid pairs that can electrostatically interact with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence; or

[0057] Any combination of the above information.

[0058] In one embodiment, the information regarding the amino acid pair may include the location of the amino acid within the antigen sequence capable of interacting with the corresponding molecules and the location of the amino acid within the antibody sequence.

[0059] In one embodiment, the information regarding the amino acid pair further includes information on the direction and strength of the interaction between the corresponding molecules, and the strength information may be an absolute value or a relative value.

[0060] In one embodiment, the information derived from the antigen sequence is a value embedding the antigen sequence or a feature representation converted from the value embedding the antigen sequence, and the information derived from the antibody sequence is a value embedding the antibody sequence or a feature representation converted from the value embedding the antibody sequence, wherein the feature representation may mean an artificially extracted feature or a hidden feature.

[0061] In one embodiment, the chemical cross attention determines a query from information derived from the antigen sequence, and

[0062] Keys and values ​​are determined from the information derived from the above antibody sequence,

[0063] The above attention scores are represented as an attention score matrix or its equivalent, and

[0064] The element in the i-th row and j-th column of the above attention score matrix represents the attention score between the i-th amino acid in the antigen sequence and the j-th amino acid in the antibody sequence, and

[0065] The attention guide based on the above chemical properties is represented as an attention guide matrix or its equivalent, and

[0066] The element in the i-th row and j-th column of the above attention guide matrix indicates whether there is an intermolecular interaction or the degree of interaction between the i-th amino acid in the antigen sequence and the j-th amino acid in the antibody sequence, and

[0067] The chemical cross-attention operation described above may include a process of performing a Hadamard product operation between the attention score matrix and the attention guide matrix.

[0068] In one embodiment, the chemical cross attention determines a query from information derived from the antibody sequence, and

[0069] Keys and values ​​are determined from the information derived from the above antigen sequence,

[0070] The above attention scores are represented as an attention score matrix or its equivalent, and

[0071] When i and j are integers,

[0072] The element in the i-th row and j-th column of the above attention score matrix represents the attention score between the i-th amino acid in the antibody sequence and the j-th amino acid in the antigen sequence, and

[0073] The attention guide based on the above chemical properties is represented as an attention guide matrix or its equivalent, and

[0074] The element in the i-th row and j-th column of the above attention guide matrix indicates whether there is an intermolecular interaction or the degree of interaction between the i-th amino acid in the antibody sequence and the j-th amino acid in the antigen sequence, and

[0075] The chemical cross-attention operation described above may include a process of performing a Hadamard product operation between the attention score matrix and the attention guide matrix.

[0076] In one embodiment, the antigen-antibody binding affinity prediction model comprises a first chemical cross-attention layer and a second chemical cross-attention layer, and

[0077] During the training process of the above model, at the forward propagation of each input data,

[0078] In the first chemical cross-attention layer above, a chemical cross-attention operation according to claim 6 is performed, and

[0079] In the second chemical cross-attention layer above, a chemical cross-attention operation according to claim 7 can be performed.

[0080] In one embodiment, the element of the i-row j-column of the attention guide matrix may be 1 if there is an interaction between the antigen sequence and antibody sequence amino acids at the corresponding position, and 0 if there is no interaction.

[0081] In one embodiment, the chemical properties include first intermolecular interaction information and second intermolecular interaction information, and

[0082] The above chemical cross-attention operation is a multi-head attention including a first head operation and a second head operation, and

[0083] During the first head operation, an attention guide based on the first molecular interaction information is applied, and

[0084] During the second head operation, an attention guide based on the second molecular interaction information may be applied.

[0085] In one embodiment, the model for predicting antigen-antibody binding affinity further comprises a first embedding layer, a second embedding layer, a first self-attention layer, and a second self-attention layer, and

[0086] During the training process of the above model, the following calculation is performed during the forward propagation of each input data:

[0087] 1) The above antigen sequence passes through a first embedding layer and is converted into an antigen embedding expression;

[0088] 2) The antigen embedding expression passes through the first self-attention layer and is converted into an antigen characteristic expression;

[0089] 3) The above antibody sequence passes through a second embedding layer and is converted into an antibody embedding expression;

[0090] 4) The antibody embedding expression passes through the second self-attention layer and is converted into an antibody characteristic expression;

[0091] 5) Chemical cross-attention operations can be performed using the above antigen characteristic expression and the above antibody characteristic expression.

[0092] In one embodiment, the model for predicting antigen-antibody binding affinity outputs a real value between 0 and 1, and

[0093] The above real value may represent the probability that the input antigen sequence and the input antibody sequence will bind.

[0094] In one embodiment, the antibody sequence may include a sequence of a variable heavy chain (VH) and a variable light chain (VL), and optionally a distinguishing token.

[0095] This specification provides a method for predicting antigen-antibody binding affinity (hereinafter, prediction method), comprising the following:

[0096] (a) The process of obtaining antigen sequences and antibody sequences;

[0097] (b) The process of identifying chemical characteristics of antigen sequences and antibody sequences,

[0098] Here, the chemical properties are information related to the intermolecular interactions between each amino acid of the antigen sequence and each amino acid of the antibody sequence; and

[0099] (c) A process of deriving an antigen-antibody binding affinity prediction value by inputting the antigen sequence and antibody sequence of (a) and the chemical characteristics of (b) into an antigen-antibody binding affinity prediction model trained by any one of the training methods selected from claims 1 to 13,

[0100] Here, the chemical properties used in the above training method and the chemical properties identified in (b) above are information regarding the same type of intermolecular interactions.

[0101] This specification provides a method for providing an antibody of interest that binds to an antigen of interest, comprising the following:

[0102] (a) A process of obtaining sequence information of the first antibody and the second antibody;

[0103] (b) Process of generating candidate antibodies of interest,

[0104] Here, the provisional candidate for the antibody of interest is:

[0105] Having a variable heavy chain having the same sequence as the variable heavy chain of the first antibody above, and

[0106] Having a variable light chain with the same sequence as the variable light chain of the second antibody above;

[0107] (c) A process for predicting the characteristics of the above-mentioned candidate antibody of interest,

[0108] Here, the above characteristic includes the binding affinity with the antigen of interest;

[0109] (d) a process of synthesizing the antibody of interest when the characteristics of the antibody candidate of interest satisfy predetermined criteria; and

[0110] (e) A process of determining whether the candidate antibody of interest synthesized in the above process (d) is the antibody of interest by confirming the binding affinity of the candidate antibody of interest to the antigen of interest.

[0111] In one embodiment, the characteristics of the antibody candidate of interest predicted in the process (c) above may further include physicochemical stability, immunogenicity, or a combination thereof.

[0112] In one embodiment, the binding affinity of the antibody candidate of interest and the antigen of interest predicted in process (c) is predicted by the method of claim 14, and

[0113] Here:

[0114] The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and

[0115] The sequence of the antibody candidate of interest can be used as the antibody sequence of the above prediction method.

[0116] This specification discloses a method for providing an antibody of interest that binds to an antigen of interest, comprising the following:

[0117] (a) The process of obtaining sequence information of a candidate antibody of interest,

[0118] Here:

[0119] The above candidate antibody of interest is generated by combining the sequence information of the first antibody and the second antibody;

[0120] The variable heavy chain sequence of the above-mentioned candidate antibody of interest is obtained from the variable heavy chain sequence of the above-mentioned first antibody;

[0121] The variable light chain sequence of the above-mentioned candidate antibody of interest is obtained from the variable light chain sequence of the above-mentioned first antibody;

[0122] The predicted values ​​for the characteristics of the above-mentioned antibody candidate of interest satisfy predetermined criteria; and

[0123] The characteristics of the candidate antibody of interest to be predicted above include binding affinity for the antigen of interest;

[0124] (b) a process for synthesizing the above-mentioned candidate antibody of interest; and

[0125] (c) A process of determining whether the candidate antibody of interest synthesized in the above process (b) is the antibody of interest by confirming the binding affinity of the candidate antibody of interest to the antigen of interest.

[0126] In one embodiment, the characteristics of the candidate antibody of interest to be predicted in the above (a) process may further include physicochemical stability, immunogenicity, or a combination thereof.

[0127] In one embodiment, the binding affinity of the antibody candidate of interest and the antigen of interest predicted in process (a) is predicted by the method of claim 14, and

[0128] Here:

[0129] The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and

[0130] The sequence of the antibody candidate of interest can be used as the antibody sequence of the above prediction method.

[0131] This specification provides a method for providing a candidate antibody of interest expected to bind to an antigen of interest, comprising the following:

[0132] (a) A process of obtaining sequence information of the first antibody and the second antibody;

[0133] (b) Process of generating temporary candidates for the antibody of interest,

[0134] Here, the provisional candidate for the antibody of interest is:

[0135] Having a variable heavy chain having the same sequence as the variable heavy chain of the first antibody above, and

[0136] Having a variable light chain with the same sequence as the variable light chain of the second antibody above;

[0137] (c) A process for predicting the characteristics of a provisional candidate for the antibody of interest,

[0138] Here, the above characteristic includes the binding affinity with the antigen of interest;

[0139] (d) If the characteristics of the provisional candidate for the antibody of interest satisfy predetermined criteria, the provisional candidate for the antibody of interest is provided as the antibody of interest candidate.

[0140] In one embodiment, the characteristics of the provisional candidate for the antibody of interest predicted in the process (c) above may further include physicochemical stability, immunogenicity, or a combination thereof.

[0141] In one embodiment, the binding affinity of the temporary candidate for the antibody of interest and the antigen of interest predicted in the process (c) is predicted by the method of claim 14, and

[0142] Here:

[0143] The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and

[0144] The sequence of the antibody candidate of interest can be used as the antibody sequence of the above prediction method.

[0145] common parts

[0146] Definition of Terms

[0147] nearby, approximately, about

[0148] As used in this specification, the terms “near,” “approximately,” or “about” mean a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by about 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1%, or 0% with respect to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. Additionally, the terms are used to mean a reasonable range that takes into account measurement errors, experimental conditions, manufacturing deviations, etc.

[0149] Singular and plural expressions

[0150] In this specification, unless otherwise specified in the context, terms expressed in the singular form are used to include the plural form, and terms expressed in the plural form are also used to include the singular form. For example, the term "bilirubin-amino sugar complex" may include not only a single bilirubin-amino sugar complex but also multiple bilirubin-amino sugar complexes. As another example, the expression "multiple bilirubin-amino sugar complexes" may be used to include individual bilirubin-amino sugar complexes.

[0151] Includes

[0152] Expressions such as "includes" used in this specification are used to mean that they do not exclude other components not enumerated, unless otherwise specified.

[0153] and / or

[0154] As used in this specification, "and / or" is used to mean including all of one or more combinations of the listed items.

[0155] It could be

[0156] As used in this specification, expressions such as "may be" or "may be" indicate optional configurations or optional features. They do not imply that they are mandatory components.

[0157] procession

[0158] In this specification, the term "matrix" encompasses all collections of numerical values ​​that can be expressed as a matrix or their equivalents. That is, it refers not only to numerical values ​​expressed as a matrix, but also to collections of numerical values ​​that can be converted into a matrix even if those values ​​are not explicitly expressed as a matrix. The equivalents of the matrix include not only cases where a matrix can be constructed directly from the numerical values ​​alone, but also cases where a matrix can be constructed by appropriately adding arbitrary values ​​while preserving the meaning of the numerical values. The term "matrix" encompasses all other meanings that a person skilled in the art would recognize and should be interpreted appropriately according to the context.

[0159] Chapter 1. Methods for Constructing an In-Silico Antibody Library

[0160] Chapter 1: Overview of the Invention

[0161] Difficulty in selecting antibodies of interest

[0162] Traditionally, antibodies that specifically bind to antigens of interest have been identified using phage display methods or hybridoma production methods. This approach involves constructing a library by synthesizing large quantities of random antibodies, antibody variable regions, or other equivalents, and then experimentally selecting antibodies from this library that bind well and specifically to the antigen of interest. Traditional methods are time-consuming and labor-intensive because they require 1) the actual synthesis of numerous antibodies or their equivalents to build a library, and 2) the identification of antibodies that specifically bind to the antigen of interest through repetitive experiments. Even if antibodies that specifically bind to the antigen of interest are selected through this process, a problem remains: the selected antibodies may trigger an immune response in the human body or possess properties unfavorable for mass production, making them difficult to use as therapeutic agents. This is because the characteristic of specifically binding to the antigen of interest does not guarantee low immunogenicity or high productivity. Therefore, it is necessary to determine through further experiments whether the properties are suitable for therapeutic use. Determining this requires additional effort and cost. Consequently, it is very difficult to discover antibodies of interest through experiments using traditional methods without prior information, and the probability of success is not high.

[0163] Concept of antibody property prediction and search space

[0164] Significant time and cost can be saved if, before undergoing the series of processes of synthesizing and evaluating numerous antibodies in large quantities, selecting them, and re-determining their sequences, the properties of the antibodies can be predicted to pre-select candidate antibodies of interest that are "highly likely to produce the intended effect," and only those candidates can be synthesized and evaluated. As technology advances and large amounts of data accumulate, in silico prediction models have been developed to predict various properties based on antibody sequence information. Using these models, it is possible to predict what properties an antibody will possess using only basic information, such as antibody sequence data. However, even in this case, basic antibody information is required to apply the in silico prediction model. The set of basic antibody information to be searched can be conceptualized as the search space.

[0165] It is necessary to set the search space reasonably.

[0166] As discussed above, in order to derive antibody candidates of interest by predicting antibody properties in silico, a search space must be established to apply the aforementioned models. If the search space is excessively wide, the computational cost of applying in silico models increases significantly. This contradicts the purpose of applying models, which is to save time and cost. If the search space is excessively narrow or lacks diversity, the benefit of applying in silico models is minimal because the probability of discovering good antibodies is low to begin with. Therefore, the search space must be established reasonably.

[0167] For example, let's assume we define a search space by generating a completely random sequence. Even if we randomly construct only the variable sequence of the antibody, we must generate a sequence of approximately 250 amino acids. The search space is 20 250 , that is, 10 325 More than one sequence must be searched. In this case, the search space is too wide, making efficient computation difficult.

[0168] As another example, let's consider a scenario where only the CDR sequence of a specific, already known antibody is mutated. In this case, the limited length of the CDR sequence makes it difficult to secure a sufficient number of search spaces. Furthermore, since this involves only a partial modification of a specific antibody, it is difficult to consider such a search space as encompassing a sufficiently diverse range of antibody sequences. Only antibody sequences with properties similar to the existing antibody will be generated, and the probability of deriving antibodies with improved properties, such as physicochemical stability or immunogenicity, will be low. In other words, the search space is too narrow and lacks diversity, making it difficult to discover effective antibodies.

[0169] It is necessary to establish a search space that is diverse and wide enough to discover new antibodies with good properties, while having a reasonable size to avoid incurring excessive computational costs.

[0170] The technical problem to be solved by this specification

[0171] The inventors of this application have as their technical objective the establishment of a search space of a reasonable level. Furthermore, they have as their technical objective the development of a method to derive a novel antibody candidate (hereinafter referred to as "antibody candidate of interest") that binds to an antigen of interest by predicting the properties of antibodies included in the said search space. Additionally, they have as their technical objective the development of a method to discover an actual antibody of interest based on the information of the said antibody candidate of interest. Furthermore, they have as their technical objective the resolution of technical problems necessary to achieve the above goals.

[0172] Methods for solving technical problems

[0173] The inventors of this application separated the information on the Variable Light Chain (VL) and Variable Heavy Chain (VH) of antibodies, from which sequence information of the Variable Region can be obtained, and established a reasonable search space by recombining the VL and VH between different antibodies. Specifically, they generated a new antibody by combining the VL of one antibody with the VH of another antibody, and constructed an in silico antibody library by generating all possible combinations from the given antibody information. Furthermore, they developed a method to derive antibody candidates of interest by predicting the properties of antibodies included in the above search space in silico. In addition, they developed a method to identify the antibody of interest by actually synthesizing and evaluating the above antibody candidates of interest. The inventors actually performed the above series of methods and demonstrated that the antibody of interest can be identified.

[0174] Method for selecting antibodies of interest

[0175] Overview of Antibody Selection Methods

[0176] This specification discloses a method for selecting antibodies of interest. The above method for selecting antibodies of interest can be broadly divided into 1) a process of deriving antibody candidates of interest in silico, and 2) a process of selecting antibodies of interest from among the antibody candidates of interest. The above process of deriving antibody candidates of interest in silico includes a process of constructing an in silico antibody library and a process of deriving antibody candidates of interest by predicting the properties of each antibody in silico. The above process of constructing an in silico antibody library includes acquiring information on two or more antibody sequences; and a process of generating new antibody sequences by combining the Variable Heavy Chain (VH) and Variable Light Chain (VL) of each antibody. For example, in the above process, information on a first antibody and a second antibody may be acquired; a new antibody sequence having a VH identical to the VH of the first antibody and a VL identical to the VL of the second antibody may be generated; and this may be added to the in silico antibody library. The above process may be repeated as many times as necessary. Meanwhile, the above process can be performed by deriving new antibody sequences all at once based on multiple antibody information. Furthermore, it includes a process of independently acquiring sequence information for each heavy chain variable region or light chain variable region, and then combining them to generate new antibody sequences.

[0177] Antibody of interest, candidate antibody of interest, provisional candidate

[0178] In this specification, "antibody of interest" refers to an antibody suitable for production and use as a therapeutic agent, such as by specifically binding to an antigen of interest, having low immunogenicity, and possessing physicochemical properties favorable for production.

[0179] In this specification, "candidate antibody of interest" means a candidate that has not been confirmed as an antibody of interest but has the potential to become one. This primarily refers to antibody sequence information that satisfies certain criteria based on the prediction of in silico properties.

[0180] In this specification, "provisional candidate" refers to a sequence that is evaluated during the process of selecting candidates for the antibody of interest. This generally refers to individual sequence information contained within the search space itself.

[0181] In this specification, candidates and provisional candidates for the antibody of interest primarily refer to antibodies whose sequence information is known only before actual synthesis. Based on the sequence information of the provisional candidates, various properties of the corresponding antibody are predicted, and sequences that satisfy certain criteria become candidates for the antibody of interest. The above candidates for the antibody of interest are actually synthesized and evaluated to identify whether they are the antibody of interest.

[0182] Even if the above terms are described in the singular in this specification, they should be interpreted as including multiple sequences depending on the context.

[0183] The above terms include all meanings that a person skilled in the art would recognize and may be interpreted appropriately depending on the context.

[0184] Process of Deriving Antibody Candidates of Interest #1 - In Silico Antibody Library

[0185] The antibody screening method disclosed in this specification includes a process for deriving antibody candidates of interest. This specification discloses an in silico antibody library that serves as a search space in the process of deriving antibody candidates of interest, and a method for constructing the same. The above in silico antibody library is constructed based on antibody information from which sequence information of variable regions can be obtained. Specifically, the light chain variable region (VL) and heavy chain variable region (VH) of each antibody are separated, and the library is constructed by recombining each separated variable region. Alternatively, it may be constructed by combining the sequences of the light chain variable region and the heavy chain variable region obtained independently. Here, the above light chain variable region may be a kappa light chain sequence (Vκ), a lambda light chain sequence (Vλ), or both. That is, the library is constructed by generating all possible combinations of the light chain variable region and the heavy chain variable region. Each of the above combinations of the light chain variable region and the heavy chain variable region is considered as a variable region of a new antibody. The library constructed in this way is called the Base Library. When a specific combination is selected by applying one or more criteria to the above basic library, it is referred to as a Selected Library. For example, a Selected Library composed of antibodies satisfying specific criteria can be constructed by applying antibodies included in the above basic library to one or more in silico property prediction models. Both the above basic library and the above Selected Library may be referred to as in silico antibody libraries.

[0186] Process of Deriving Antibody Candidates of Interest #2 - Predicting In Silico Properties

[0187] The process of deriving the antibody candidates of interest involves predicting their properties in silico for each antibody sequence (i.e., a combination of light chain variable regions and heavy chain variable regions) included in the above in silico antibody library. Since the objective is to discover antibodies of interest that possess characteristics suitable for actual antibody therapeutics, properties that can determine suitability as an antibody therapeutic are predicted. Specifically, 1) binding affinity to the antigen of interest, 2) selective binding affinity to the antigen of interest, 3) immunogenicity, and 4) physicochemical properties suitable for therapeutic use are predicted. Here, physicochemical properties suitable for therapeutic use include i) purity (also expressed as homogeneity) during antibody production, ii) structural stability, and iii) solubility. The prediction of these properties utilizes information on each antibody sequence, and additional information if necessary. If the properties satisfy specific criteria, the antibody is selected as a candidate of interest. For example, antibody sequences predicted to specifically bind to the antigen of interest above a certain level, have very low immunogenicity, high purity during production, be structurally stable in the body environment, and have high solubility are selected as candidates for the antibody of interest. All of the above processes can be performed in silico.

[0188] The process of selecting antibodies of interest from among candidate antibodies of interest

[0189] The antibody selection method disclosed in this specification includes a process of selecting an antibody of interest from among antibody candidates of interest. The process of deriving antibody candidates of interest has been described above. The process of selecting an antibody of interest involves actually synthesizing the antibody candidates of interest derived above and evaluating their properties to select the antibody of interest. Proteins are synthesized based on sequence information, various properties of the antibodies are evaluated, such as binding affinity to the antigen of interest or physicochemical characteristics suitable for therapeutic use, and antibody candidates of interest that satisfy the criteria are selected. The above series of processes can be performed using known methods.

[0190] Characteristics of the Antibody of Interest Screening Method #1 - The search space is limited to a reasonable range

[0191] The antibody screening method disclosed in this specification uses the aforementioned in silico antibody library as the search space. As described above, it is most desirable to set the search space such that 1) the in silico computational cost is not excessively large to the extent that it is manageable, 2) the in silico prediction is not excessively small to the extent that it is meaningless, and 3) it includes a sufficiently diverse range of antibodies to ensure a high probability of discovering a good antibody.

[0192] The size of the in silico antibody library disclosed in this specification is determined by (number of sequences of known light chain variable regions) × (number of sequences of known heavy chain variable regions). For example, calculated based on the number of disclosed antibody sequences, (10 6 (10 known kappa light chain sequences) Х (10 8 (Known variable heavy chain sequences) = 10 14 Search spaces can be secured. This means that when searching a randomly constructed 250-mer amino acid sequence, 10 325 Compared to the need to search through sequences, it is a sufficiently manageable size. Furthermore, it is larger than the square of the number of known antibodies, so it is not excessively small. Moreover, the sequences included in the above in silico antibody library can be seen as containing a sufficiently diverse range of antibodies, as 1) they include conventionally existing light chain variable regions or heavy chain variable regions, thus having a sufficiently high probability of forming antibodies themselves, and 2) they are expected to have completely new structures and functions through combinations of new light chain variable regions and heavy chain variable regions.

[0193] Therefore, the in silico antibody library disclosed in this specification is a reasonable range of search space.

[0194] Characteristics of the Antibody of Interest Screening Method #2 - Suitable antibodies of interest can be derived

[0195] The antibody screening method disclosed in this specification derives antibody candidates of interest by predicting not only whether they specifically bind to the antigen of interest, but also whether they possess characteristics suitable for use as a therapeutic agent (e.g., low immunogenicity, physicochemical properties suitable for therapeutic use, etc.). Therefore, there is a high probability of deriving antibodies of interest that not only specifically bind to the antigen of interest and exhibit the desired effect, but are also advantageous for production and use.

[0196] Feature of the Antibody of Interest Selection Method #3 - No Loss of Selected Sequence Information

[0197] According to conventional technology, multiple antibodies are synthesized and evaluated simultaneously using a high-throughput method for efficient experiments. In the high-throughput method, a library is constructed by synthesizing antibodies with known sequences and antibodies with random sequences all at once. Therefore, even if an antibody of interest is selected through an experiment, there is a problem in that the experiment becomes complicated because it must go through the step of sequencing it to determine its sequence again.

[0198] On the other hand, the antibody screening method disclosed in this specification has the advantage of being able to screen antibodies of interest while maintaining sequence information, as both the library construction and the selection of antibody candidates of interest are performed in silico. In other words, the sequence information generated and screened by the above method is not lost and can be appropriately reused later.

[0199] in silico antibody library

[0200] Overview of the in silico antibody library

[0201] This specification discloses an in silico antibody library and a method for constructing said library. The said in silico antibody library is constructed by generating a new antibody sequence by combining the light chain variable region of one antibody and the heavy chain variable region of another antibody, based on antibodies for which variable region information is known. A library containing all possible combinations of light chain variable regions and heavy chain variable regions is referred to as the Base Library. A library composed of combinations included in the Base Library that satisfy one or more criteria is referred to as the Selected Library. Each combination included in the said Base Library or Selected Library has the potential to become a new antibody. For convenience, these may be referred to as "combinations," "provisional candidates," "hypothetical antibodies," or simply "antibodies," and should be interpreted appropriately according to the context.

[0202] The above in silico antibody library serves as a search space for the antibody screening method of interest disclosed in this specification.

[0203] Base Library

[0204] The basic library contains information on virtual antibodies formed by combining the sequences of all light chain variable regions of available antibodies and the sequences of all available heavy chain variable regions. The above basic library contains one or more combinations of light chain variable region sequences and heavy chain variable region sequences. The light chain variable region sequences and heavy chain variable region sequences included in each combination are sequences derived from different antibodies. That is, each combination represents a virtual antibody generated in silico.

[0205] Characteristics of antibody information included in the basic library

[0206] As previously mentioned, the above basic library contains one or more combinations consisting of sequences of light chain variable regions and heavy chain variable regions. New antibodies can be constructed by adding known invariant regions to the light chain variable regions and heavy chain variable regions of each combination. In other words, each combination represents a new antibody possessing the light chain and heavy chain variable regions of the corresponding sequences. Since this information is generated by randomly combining known antibody sequence data, it is impossible to know whether each antibody actually forms an antibody structure or specifically binds to a specific antigen until it is actually verified. However, because the composition is formed by recombining light chain variable regions and heavy chain variable regions that function as actual antibodies, it is judged to have a higher probability of functioning as an antibody than completely randomly generated amino acid sequences.

[0207] Selected Library

[0208] The above basic library refers to the maximum search space constructed by utilizing all available antibody information. Among the virtual antibodies included in the above basic library, a separate library can be constructed by collecting only those antibodies that satisfy specific criteria. This is referred to as the Selected Library. The Selected Library corresponds to a part of the above basic library. By applying appropriate criteria to the above basic library, a high-quality Selected Library can be constructed and utilized for selecting antibodies of interest. For example, among the virtual antibodies included in the above basic library, a Selected Library can be constructed by separately selecting antibodies whose immunogenicity is below a certain level and whose physicochemical characteristics suitable for therapeutic use are above a certain level. Here, immunogenicity and physicochemical characteristics can be determined based on values ​​predicted in silico.

[0209] Predicting properties in silico

[0210] Overview of In Silico Property Prediction

[0211] The method for selecting antibodies of interest disclosed in this specification includes a process of deriving antibody candidates of interest in silico. Specifically, for each antibody included in an in silico antibody library, properties are predicted in silico, and antibodies satisfying certain criteria are derived as antibody candidates of interest. This process reduces the time and cost required for the process of synthesizing and verifying the actual antibody of interest by preferentially selecting in silico antibodies among the new antibodies included in the in silico antibody library that have a high probability of possessing the desired properties. The properties predicted in silico may be, for example, binding affinity to the antigen of interest, binding specificity to the antigen of interest, immunogenicity, or physicochemical characteristics suitable for therapeutic use. To predict these properties, predicted values ​​for various factors can be derived. Various methods may be used to derive predicted values ​​for properties and are not limited to a specific method. For example, various methods may be utilized, such as sequence identity comparison, calculations through simulation models, calculations through structure prediction, and derivation of predicted values ​​through machine learning models. A person skilled in the art can predict the properties of antibodies included in a basic library in silico using a known method and derive candidate antibodies of interest using this.

[0212] For convenience, this specification distinguishes between properties and factors, and specific details are explained in the paragraph below.

[0213] Properties and Factors

[0214] In this specification, the term "Property" refers to a qualitative characteristic conceptualized from a specific perspective. For example, "property of binding to an antigen of interest," "property of binding specifically only to an antigen of interest," "immunogenicity," or "physicochemical properties suitable for use as a therapeutic agent" are considered properties.

[0215] In this specification, the term "Factor" refers to any value that can be expressed quantitatively, such as a measurable physical quantity, a relative value according to a specific criterion, or a score that can be calculated. For example, "antigen-antibody binding strength," "similarity to human antibodies," and "light-heavy chain binding potential" are considered factors. The above factors may be derived in various ways and may be predicted by specific methods. For example, the above factors may be measured through experiments, calculated statistically, calculated based on relevant information, derived using specific models or methodologies, or determined by other known methods.

[0216] To select antibody candidates of interest, it is necessary to predict whether antibodies within the search space satisfy specific properties. Whether these specific properties are satisfied is determined based on predicted values ​​for specific factors. For example, broadly speaking, these specific properties refer to "whether the antibody produces sufficient therapeutic effects, has low immunogenicity, and possesses physicochemical characteristics favorable for use as a therapeutic agent," and this is determined by predicting or calculating factors such as "antigen-antibody binding affinity," "similarity to human antibodies," and "light-heavy chain binding potential."

[0217] The above distinction between properties and factors is merely a rough distinction. In other words, they are not always strictly and clearly distinguished, and a specific factor may directly represent a specific property. For example, "solubility under specific conditions" is a measurable physical quantity and thus corresponds to a factor, but it is also treated as a physicochemical characteristic (property) suitable for therapeutic use.

[0218] Properties can be determined by a single factor or a combination of various factors

[0219] A property may be a characteristic resulting from the contribution of a single factor or various factors. Furthermore, a single factor can serve as the basis for judging two or more properties. Properties are judged based on measured or predicted factors, and the relationship between these factors and properties can be one-to-one, one-to-many, many-to-one, or many-to-many. These relationships can be expressed and interpreted in various ways.

[0220] Predicted Property #1 - Binding Ability to Antigen of Interest

[0221] For an antibody to produce the intended effect, it must be able to bind to the antigen of interest. In other words, the antibody of interest exhibits a binding affinity to the antigen of interest above a certain level. This binding affinity to the antigen of interest can be determined by the binding affinity of the antibody to the antigen of interest.

[0222] Predicted Property #2 - Binding Specificity to Antigen of Interest

[0223] For antibodies to avoid unwanted side effects, they must be able to bind selectively only to the antigen of interest. In other words, the antibody of interest exhibits a binding affinity below a certain level for antigens other than the antigen of interest. This binding specificity to the antigen of interest can be assessed by the binding affinity of the antibody to antigens other than the antigen of interest.

[0224] Predicted Property #3 - Immunogenicity

[0225] In order for an antibody to be used as a therapeutic agent for humans, it must not trigger an immune response when delivered into the body. In other words, the antibody of interest exhibits an immune response below a certain level when delivered to the human body. There may be various criteria for determining this. For example, immunogenicity can be determined based on the similarity measured by how similar the sequence of the antibody above is to the sequence of a human antibody.

[0226] Predicted Property #4 - Physicochemical Properties Suitable for Therapeutic Use

[0227] To commercialize an antibody as a therapeutic agent, it is desirable for it to possess physicochemical properties suitable for therapeutic use. In other words, the antibody of interest possesses physicochemical properties suitable for therapeutic use. Physicochemical properties suitable for therapeutic use encompass various attributes. For example, they may include high purity during production, structural stability in the in vivo environment, and high solubility in the in vivo environment. Below, several physicochemical properties suitable for therapeutic use are described as examples. However, physicochemical properties suitable for therapeutic use are not limited to the examples described below. If necessary, physicochemical properties suitable for therapeutic use can be determined through other appropriate attributes.

[0228] Physicochemical properties suitable for therapeutic use #1 - Purity (or homogeneity) at production

[0229] Generally, antibody drugs are produced by delivering the amino acid sequence information of an antibody to specific cells and having those cells express the antibody sequence. If the cells producing the antibody express proteins or polypeptides other than the antibody, these are all treated as impurities. There are various causes for the formation of impurities. For example, proteins or polypeptides other than the intended antibody may be expressed due to various reasons, such as: 1) only a partial fragment is expressed instead of the full sequence; 2) the full sequence is expressed but does not fold properly; 3) disulfide bonds between the antibody light and heavy chains are not properly formed; or 4) additional modifications (e.g., enzymatic cleavage) occur after the full sequence is expressed. Since all of these impurities are derived from the antibody expression process, the characteristics of the antibody's amino acid sequence itself contribute to a certain extent to whether the impurities are more or less formed. Therefore, the ability to produce with high purity is a property advantageous for use as a therapeutic agent.

[0230] Purity during production can be determined by one or more factors. For example, it can be assessed by confirming molecular weight via SDS-PAGE, or by performing antibody purity and SE-HPLC analysis to determine the presence of aggregation.

[0231] Physicochemical properties suitable for therapeutic use #2 - Structural stability

[0232] Since antibody drugs must function within the human body, structural stability in the in vivo environment is desirable. Structural stability means maintaining an intact antibody structure within the body without clumping, cleavage, or other deformations. Because the antibody must maintain its three-dimensional structure to bind to a target antigen, structural stability is crucial for producing a stable therapeutic effect.

[0233] Structural stability can be assessed by one or more factors. For example, it can be determined by observing whether the structure is maintained when heat is applied. As another example, it can be predicted using a score derived by inputting antibody sequence information into a model trained to predict the degree of protein aggregation. The above model may be, but is not limited to, the prediction model disclosed in Sun et al., *Enhancing protein aggregation prediction: a unified analysis leveraging graph convolutional networks and active learning*, RSC Adv., 2024, 14, 31439-31450. Another example is solubility in the in vivo environment. Generally, proteins with low solubility in the in vivo environment tend to aggregate and be structurally unstable.

[0234] Physicochemical Properties Suitable for Therapeutic Use #3 - Solubility

[0235] Antibody drugs typically move to a target site via body fluids after being injected into the human body to perform their function. Therefore, to move to the target site through body fluids (e.g., blood, lymph, etc.) within the body, it is advantageous to have high solubility in the body fluids. Since body fluids consist mostly of water, gastric solubility generally refers to water solubility.

[0236] Since solubility can be directly measured through experiments, it is also treated as a factor. Solubility can be determined by measuring or predicting how much dissolves in water, blood, or other body fluids at a specific temperature and pH.

[0237] In Silico Property Prediction Method #1 - Method for Predicting Binding Strength and Binding Specificity for Antigens of Interest

[0238] The binding strength and binding specificity of an antibody to an antigen of interest can be predicted in various ways. For example, the binding strength to an antigen of interest can be predicted using a predicted value obtained by applying the sequence of the antibody above to an antigen-antibody binding prediction model for the antigen of interest. As another example, the binding specificity to an antigen of interest can be predicted by comparing the predicted values ​​obtained by inputting the sequence of the antibody above into an antigen-antibody binding prediction model for the antigen of interest and an antigen-antibody binding prediction model for other antigens. Here, the above antigen-antibody binding prediction model may be a model trained using data in which antibody sequences are labeled as binding strengths with specific antigens. In one embodiment, the model disclosed in the paragraph [Chapter 2. Antigen-Antibody Binding Affinity Prediction Model] may be used to predict the antigen-antibody binding affinity.

[0239] Method for Predicting In Silico Properties #2 - Method for Predicting Immunogenicity

[0240] Whether an antibody will exhibit immunogenicity in the body can be predicted by various methods. For example, immunogenicity can be predicted using a predicted value obtained by applying the sequence of the antibody above to a human antibody prediction model. Here, the human antibody prediction model above may be an antibody model. As another example, the immunogenicity of an antibody can be predicted by determining how similar the antibody sequence is to a human antibody by appropriately utilizing known methods. As an example, the following method can be followed: 1) input the sequences of the antibody's light chain variable region and heavy chain variable region into the ANARCI program (James Dunbar, Charlotte M. Deane, ANARCI: antigen receptor numbering and receptor classification, Bioinformatics, Volume 32, Issue 2, January 2016, Pages 298-300, https: / doi.org / 10.1093 / bioinformatics / btv552); 2) Each sequence and the IMGT GENE database (Vιronique Giudicelli, Denys Chaume, Marie-Paule Lefranc, IMGT / GENE-DB: a comprehensive database for human and mouse immunoglobulin and T cell receptor genes, Nucleic Acids Research, Volume 33, Issue suppl_1, 1 January 2005, Pages D256-D261, https: / doi.org / 10.1093 / nar / gki010) compares the similarity between species-specific antibody sequences to find the closest species; 3) output the closest antibody species found in ANARCI to determine immunogenicity.

[0241] In Silico Property Prediction Method #3 - Method for Predicting Physicochemical Properties Suitable for Therapeutic Use

[0242] Whether an antibody possesses physicochemical properties suitable for therapeutic use can be predicted by various methods. For example, the above properties can be predicted by applying a machine learning model trained on data labeled with a composite score for "properties suitable for therapeutic use" to the antibody sequence. Here, the machine learning model may be, but is not limited to, the CNN-P model disclosed in Chinery, L., Jeliazkov, JR, & Deane, CM (2024). Humatch - fast, gene-specific joint humanisation of antibody heavy and light chains. mAbs, 16(1). https: / doi.org / 10.1080 / 19420862.2024.2434121. As another example, the above properties can be predicted by applying a model that predicts the degree of protein aggregation to the antibody sequence. Here, the model that predicts the degree of protein aggregation is Sun et. The prediction model disclosed in al., Enhancing protein aggregation prediction: a unified analysis leveraging graph convolutional networks and active learning, RSC Adv., 2024, 14, 31439-31450 may be, but is not limited to, the model disclosed therein.

[0243] Can be used for curated library configuration

[0244] The above method for predicting in silico properties can be applied to a basic library to construct a selected library. For example, the immunogenicity and physicochemical properties suitable for therapeutic use of the antibodies included in the basic library can be predicted, and a selected library can be constructed by gathering antibodies that meet certain criteria. Although it is not yet identified which antigen the antibodies included in the selected library will bind to, it can be expected that they at least have low immunogenicity and physicochemical properties suitable for therapeutic use. If the antigen of interest is determined later, candidate antibodies of interest can be derived more efficiently based on the selected library.

[0245] The process of selecting antibodies of interest from among candidate antibodies of interest

[0246] Overview of the process of selecting the antibody of interest from among candidate antibodies of interest

[0247] The method for selecting an antibody of interest in this specification includes a process of selecting an antibody of interest from among antibody candidates of interest. Here, the above antibody candidates of interest refer to antibodies derived from the aforementioned in silico library through in silico property prediction. All of the above antibody candidates of interest are antibodies generated and selected in silico, and have not been actually produced or evaluated. Therefore, in order to identify which of the above antibody candidates of interest is the antibody of interest, they must be actually synthesized and their properties evaluated. Conceptually, the process of selecting an antibody of interest from among the above antibody candidates of interest includes 1) synthesizing antibody candidates of interest, 2) evaluating the properties of the synthesized antibody candidates of interest, and 3) selecting the antibody of interest. As long as the respective objectives are achieved, the means for these processes are not limited and can be performed using known techniques.

[0248] Synthesize candidate antibodies of interest

[0249] The process of selecting the antibody of interest from the above candidate antibodies includes the process of synthesizing the candidate antibodies of interest. Here, the sequences of the light chain variable region and the heavy chain variable region of each candidate antibody of interest are specified. Based on these variable regions, an antibody of an appropriate form is synthesized according to the purpose.

[0250] For example, a full-length antibody containing a constant region can be synthesized. Here, the isotype of the full-length antibody can be appropriately selected depending on the purpose. As another example, a single-chain variable fragment (scFv) containing a light-chain variable region and a heavy-chain variable region can be synthesized. It can be synthesized in various other forms, examples of which are disclosed in the "Possible Examples of the Invention" paragraph.

[0251] Evaluate the properties of the synthesized candidate antibody of interest

[0252] The process of selecting the antibody of interest from the above candidate antibodies of interest includes evaluating the properties of the synthesized candidate antibodies of interest. This is a process of actually verifying whether the synthesized candidate antibodies of interest possess properties suitable for use as the antibody of interest. The properties to be verified may be appropriately selected depending on the purpose. The method for verifying such properties may be performed by a person skilled in the art using known methods. The properties evaluated through this process may be identical to, but are not limited to, the properties predicted in the in silico property prediction process. For example, with the synthesized candidate antibodies of interest, the following properties may be verified: binding affinity to the antigen of interest; specific binding affinity to the antigen of interest; immunogenicity; physicochemical properties favorable for production; physicochemical properties suitable for use as a therapeutic agent; or any combination of the above properties.

[0253] Select antibodies of interest

[0254] The process of selecting the antibody of interest from the above candidate antibodies of interest includes the selection process. Whether each candidate antibody of interest evaluated above satisfies predetermined criteria is assessed to identify whether it is an antibody of interest. The predetermined criteria can be appropriately set according to the purpose.

[0255]

[0256] Chapter 2. Antigen-Antibody Binding Affinity Prediction Models

[0257] Chapter 2 Overview of the Invention

[0258] The usefulness of predicting antigen-antibody binding affinity

[0259] When determining whether an antibody is commercially useful, its binding ability to a target antigen is assessed. This is because antibodies function by binding to the target antigen. Traditionally, screening methods have been used in which a library containing numerous antibodies is exposed to a specific antigen to identify antibodies with high binding affinity, which are then isolated, analyzed, and sequenced. Since this method is based on the premise that an existing antibody library exists and that the binding ability of such antibodies to the antigen is tested and selected, there has been relatively little need to predict antigen-antibody binding affinity.

[0260] On the other hand, recent advancements in computer-based modeling technology have made it possible to generate entirely new antibody sequences. In such cases, actual antibodies must be synthesized based on the antibody sequence information, and their properties, such as binding affinity to antigens, must be analyzed. The problem is that the actual production and property analysis of antibodies are costly, and the cost increases in proportion to the number of candidate antibody sequences.

[0261] In this context, if antigen-antibody binding affinity can be predicted using computer-based models, significant costs can be saved. Therefore, there is an increasing need for technologies that can predict antigen-antibody binding affinity with high accuracy.

[0262] A model has been developed to predict antigen-antibody binding affinity using only antigen and antibody sequences.

[0263] Specifically, the following prior art discloses a model that predicts antigen-antibody binding affinity by taking antigen and antibody sequences as input:

[0264] Prior Literature 1: Ruofan Jin, Qing Ye, Jike Wang, Zheng Cao, Dejun Jiang, Tianyue Wang, Yu Kang, Wanting Xu, Chang-Yu Hsieh, Tingjun Hou, AttABseq: an attention-based deep learning prediction method for antigen-antibody binding affinity changes based on protein sequences, Briefings in Bioinformatics, Volume 25, Issue 4, July 2024, bbae304, https: / / doi.org / 10.1093 / bib / bbae304

[0265] Prior Literature 2: Yuan, Y., Chen, Q., Mao, J. et al. DG-Affinity: predicting antigen-antibody affinity with language models from sequences. BMC Bioinformatics 24, 430 (2023). https: / / doi.org / 10.1186 / s12859-023-05562-z

[0266] The results are unsatisfactory due to insufficient training data.

[0267] Despite the aforementioned prior research, the performance of antigen-antibody binding affinity prediction models published to date is not particularly outstanding. This is because there is a lack of training data.

[0268] The process of actually synthesizing antibodies, reacting them with antigens, and measuring antigen-antibody binding affinity is costly. Generally, rather than being performed on a large scale, this experiment is conducted on a limited basis, specifically for antibodies of interest with therapeutic potential. Therefore, data on antigen-antibody binding affinity is very limited. Approximately 10 per antigen 2 to 10 3 It is understood that antibody data of a certain scale has been released. Considering the amount of data generally required to train a generative model, this is woefully insufficient.

[0269] Therefore, there is a need for a new model capable of effective learning even with limited data.

[0270] Technical problems of this specification

[0271] This specification sets forth the technical objective of developing an antigen-antibody binding affinity prediction model that takes an antigen sequence and an antibody sequence as input and outputs an antigen-antibody binding affinity. In particular, the objective is to develop a model that can be efficiently trained with relatively small amounts of data. Additionally, the technical objective is to implement a method for training the above antigen-antibody binding affinity prediction model and a method for predicting actual antigen-antibody binding affinities using the trained model. Furthermore, the technical objective is to disclose specific examples of utilizing the above antigen-antibody binding affinity prediction model.

[0272] Problem Solving Method #1 - Additionally utilizing the chemical properties of antigen and antibody sequences

[0273] The inventors of this application sought to improve learning efficiency by additionally utilizing the chemical characteristics of antigen and antibody sequences when training an antigen-antibody binding affinity prediction model. Specifically, they additionally utilized information on amino acids that contribute to antigen-antibody binding. For example, information on amino acids expected to form hydrogen bonds when an antigen and antibody bind, information on amino acids expected to form strong van der Waals interactions, and information on hydrophobic amino acids may be utilized.

[0274] Problem-Solving Method #2 - Integrating and Utilizing Chemical Properties into the Model's Information Processing Structure

[0275] Instead of simply using the chemical properties as additional input variables, the inventors integrated them into a structure that models the interaction between antigen and antibody sequences. Specifically, they introduced a cross-attention mechanism to learn the interaction between antigen and antibody sequences, and configured the model to enable more efficient learning by focusing on the locations contributing to antigen-antibody binding during cross-attention to learn the context. Since the locations of amino acids contributing to antigen-antibody binding differ depending on the respective antigen and antibody sequences, this aspect was dynamically reflected in the learning process.

[0276] Problem Solving Method #3 - Specifically implement the above idea and demonstrate performance improvement

[0277] The inventors constructed a model that specifically implements the above idea and demonstrated that efficient learning is achieved by training with available data.

[0278] Antigen-antibody binding affinity prediction model

[0279] Overview of Antigen-Antibody Binding Affinity Prediction Models

[0280] This specification discloses an antigen-antibody binding affinity prediction model. The above model is configured to take an antigen sequence and an antibody sequence as input and output an antigen-antibody binding affinity. In the process of predicting binding affinity, the above model utilizes the chemical characteristics of the antigen sequence and the antibody sequence as additional information. These chemical characteristics refer to information regarding amino acids that interact during antigen-antibody binding, such as the positions of amino acids involved in hydrogen bonding. Instead of simply using these chemical characteristics as additional input variables, the above model integrates them into a cross-attention mechanism. The cross-attention mechanism employed in the above model is referred to as the chemical cross-attention mechanism. The above chemical cross-attention is designed to determine which positions on the sequence to focus on learning by considering these chemical characteristics during the process of learning how the antigen sequence and the antibody sequence will interact. The antigen-antibody binding affinity prediction model of this specification is characterized by learning antigen-antibody interactions by considering chemical characteristics. Accordingly, it can learn effectively even with limited data. Since the above model can predict antigen-antibody binding affinity with high accuracy, it can be applied in various ways to technologies for generating novel antibodies. For example, if the above antigen-antibody binding affinity prediction model is used as a reward for the antibody generation model, the antibody generation model can be improved to generate antibody sequences that bind better to the antigen.

[0281] Model Input #1 - Antigen Sequence

[0282] The antigen-antibody binding affinity prediction model of this specification receives an antigen sequence as input. The antigen sequence is a sequence of amino acids of a fixed length. The antigens that can be input into the model are not otherwise limited, as long as they are targets to which antibodies can bind. The antigen sequence may be the full-length sequence or a partial sequence of a known antigen.

[0283] Model Input #2 - Antibody Sequence

[0284] The antigen-antibody binding affinity prediction model of this specification receives an antibody sequence as input. The antibody sequence is an amino acid sequence of a fixed length. The antibodies that can be input into the model are not otherwise limited. However, considering the performance and efficiency of the model, only the sequence of an important region of the antibody may be received as input. In this case, the important region does not necessarily refer only to a continuous portion of the antibody's full-length sequence, but may include multiple regions within the full-length sequence that are not contiguous with one another. For example, the antibody sequence may be the full-length amino acid sequence of the antibody. As another example, the antibody sequence may be the amino acid sequence of the antibody's variable region. As yet another example, the antibody sequence may be the amino acid sequence of the antibody's heavy chain variable region and the amino acid sequence of the antibody's light chain variable region. As yet another example, the antibody sequence may be the sequences of the antibody's complementary determining region (CDR).

[0285] Model Output - Antigen-Antibody Binding Affinity

[0286] The antigen-antibody binding affinity prediction model of this specification outputs an antigen-antibody binding affinity. The above antigen-antibody binding affinity is represented by at least one real number. Depending on the purpose and configuration of the model, the scale on which the antigen-antibody binding affinity is calculated may be determined and is not otherwise limited as long as it can be represented as a number. For example, the above antigen-antibody binding affinity may be 0 or 1. Here, 0 means that the antigen and antibody do not bind, and 1 means that the antigen and antibody bind. As another example, the above antigen-antibody binding affinity may be a real number between 0 and 1. Here, the output real number represents the probability that the antigen and antibody will bind. That is, 0 means that the antigen and antibody do not bind, and 1 means that the antigen and antibody bind with a 100% probability.

[0287] Additional utilization of the chemical properties of antigen and antibody sequences

[0288] The antigen-antibody binding affinity prediction model of this specification utilizes the chemical features of the antigen sequence and antibody sequence when predicting antigen-antibody binding affinity. The chemical features of the antigen sequence and antibody sequence include information regarding amino acids that are highly likely to interact during antigen-antibody binding. For example, the chemical features may include information regarding amino acids within the antigen sequence and antibody sequence where hydrogen bonding may occur. As another example, the chemical features may include information regarding amino acids within the antigen sequence and antibody sequence where strong van der Waals forces act between them. Among the chemical features, the positional information of the amino acids—specifically, the positional information of the corresponding amino acids within the antigen sequence and antibody sequence—is important. This is because the model utilizes this positional information to learn the interaction between the antigen and the antibody. Specifically, the model includes a chemical cross-attention mechanism to learn the interaction between the antigen and the antibody, and the chemical features are utilized in this process.

[0289] A method for modeling the interaction between antigen and antibody sequences - Chemical Cross Attention

[0290] The antigen-antibody binding affinity prediction model of this specification employs a chemical cross-attention mechanism to model the interaction between antigen sequences and antibody sequences. Cross-attention mechanisms are commonly used in machine learning models to model interactions between two objects. The chemical cross-attention described above is a modified version of the general cross-attention mechanism designed to account for the chemical properties of antigen and antibody sequences. Specifically, the chemical cross-attention described above is configured so that characteristic values ​​at locations associated with the chemical properties are given greater weight during cross-attention calculations. For example, characteristic values ​​at locations associated with the chemical properties may be amplified and reflected. As another example, characteristic values ​​at locations not associated with the chemical properties may all be filtered out and removed. The chemical cross-attention mechanism described above enables the model to focus on locations where significant interactions occur when learning the interaction between antigen and antibody sequences.

[0291] Feature of the Antigen-Antibody Binding Affinity Prediction Model #1 - Utilizes the chemical properties of each sequence in interaction modeling

[0292] The antigen-antibody binding affinity prediction model of this specification is characterized by utilizing chemical property information of each antigen sequence and antibody sequence. In particular, it is characterized by integrating this chemical property information into the process of modeling the interactions between the sequences, rather than using it as a new input variable or integrating it into existing input variables. When deriving interaction information between antigen and antibody sequences, the model is configured to learn antigen-antibody binding affinities more efficiently by focusing on regions where actual amino acid interactions are highly likely to occur. In the model, the interaction modeling using the chemical properties is implemented as chemical cross-attention.

[0293] Feature of the Antigen-Antibody Binding Affinity Prediction Model #2 - Can train effectively even with small amounts of data

[0294] The antigen-antibody binding affinity prediction model of this specification is characterized by the ability to be effectively trained even in situations where there is limited data. This is because chemical property information is effectively utilized during the training process of the above model, thereby properly guiding which information to focus on. The inventors of this application have demonstrated that the idea applied to the above model actually enables efficient training (see experimental example).

[0295] Features of the Antigen-Antibody Binding Affinity Prediction Model #3 - Versatile Applications

[0296] With the advancement of computer-based technology, unlike traditional antibody discovery processes, it has become possible to directly generate and use "antibody sequence information" rather than actual antibodies. In this context, the antigen-antibody binding affinity prediction model of this specification can be applied in various ways. For example, if an in silico antibody sequence library is available, the above model can be used to screen antibody sequences with a high probability of binding to a specific antigen. As another example, the above model can be utilized as a compensation model for an antibody generation model to optimize the model and generate high-quality antibody sequences. Furthermore, the above model can be applied in various ways to computer-based antibody research.

[0297] Chemical properties of antigen sequences and antibody sequences

[0298] Overview of the Chemical Characteristics of Antigen and Antibody Sequences

[0299] The antigen-antibody binding affinity prediction model of this specification utilizes the chemical features of the antigen and antibody sequences. Since the purpose of the above model is to predict antigen-antibody binding affinity, the above chemical features include information related to intermolecular interactions that contribute to antigen-antibody binding. The above chemical features may also be simply referred to as intermolecular interaction features. The above intermolecular interactions may be, but are not limited to, hydrogen bonds, van der Waals interactions, hydrophobic interactions, or electrostatic interactions. The information related to the above intermolecular interactions refers to information regarding amino acid pairs capable of interacting within the antigen and antibody sequences. For example, when the above chemical features include information related to hydrogen bonds, information regarding amino acid pairs capable of hydrogen bonding with each other within the antigen and antibody sequences is utilized in the training of the above model. Here, the above amino acid pair refers to a pair consisting of one amino acid in the antigen sequence and one amino acid in the antibody sequence. The above chemical features may be derived by considering only the types of amino acids in each sequence. Depending on the implementation of the model, chemical features may also be derived by considering the stereochemical structures of the antigen and antibody, but this is not a mandatory requirement.

[0300] Examples of important chemical properties are described in more detail below.

[0301] Chemical Property Example #1 - Hydrogen Bond

[0302] The above chemical properties may include information related to hydrogen bonding. Specifically, the antigen-antibody binding affinity prediction model of this specification may utilize information on amino acid pairs that can hydrogen bond with each other in the antigen sequence and the antibody sequence.

[0303] Hydrogen bonding occurs when a hydrogen atom (or group) with a polar covalent bond to a highly electronegative atom (or group) within a molecule interacts with a highly electronegative atom (or group) possessing an electron pair within another molecule. Here, the atom (or group) with a polar covalent bond to the hydrogen is referred to as the donor atom (Dn), and the other atom (or group) interacting with the hydrogen is referred to as the hydrogen bond acceptor (Ac). Accordingly, an amino acid containing the donor atom in the antigen sequence can form hydrogen bonds with an amino acid containing the hydrogen bond acceptor in the antibody sequence. The same applies to an amino acid containing the donor atom in the antibody sequence and an amino acid containing the hydrogen bond acceptor in the antigen sequence.

[0304] The above information regarding hydrogen bonds may include: 1) information about a pair of amino acids containing a donor atom in the antigen sequence and an amino acid containing a hydrogen bond acceptor in the antibody sequence, 2) information about a pair of amino acids containing a donor atom in the antibody sequence and an amino acid containing a hydrogen bond acceptor in the antigen sequence, or 3) both 1) and 2). Herein, the information about the amino acid pairs includes the position of each amino acid within the antigen sequence or antibody sequence and may further include the strength of the hydrogen bonding of the said amino acid pair.

[0305] Chemical Property Example #2 - Van Der Waals Interaction

[0306] The above chemical properties may include information related to van der Waals interactions. Specifically, the antigen-antibody binding affinity prediction model of this specification may utilize information on amino acid pairs that interact with each other van der Waals in the antigen sequence and the antibody sequence.

[0307] Van der Waals interactions refer to weak electrical interactions arising from instantaneous changes in charge distribution within molecules. As long as the distance between molecules is not excessively close, this interaction acts primarily as an attractive force. Furthermore, because it does not require special chemical conditions, it is observed between all types of molecules.

[0308] Information related to the above van der Waals interactions may include information on all possible amino acid pairs within the antigen and antibody sequences. This is because van der Waals interactions occur between all amino acids. Here, the information on amino acid pairs includes the position of each amino acid within the antigen or antibody sequence and may further include the strength of the van der Waals interaction of the corresponding amino acid pair.

[0309] Chemical Property Example #3 - Hydrophobic Interaction

[0310] The above chemical properties may include information related to hydrophobic interactions. Specifically, the antigen-antibody binding affinity prediction model of this specification may utilize information on amino acid pairs that hydrophobically interact with each other in the antigen sequence and the antibody sequence.

[0311] Hydrophobic interactions refer to the attractive forces that occur between residues of hydrophobic amino acids. When hydrophobic residues are exposed to an aqueous environment, water molecules form a clathrate structure, which is entropically disadvantageous. Consequently, proteins tend to clump together their hydrophobic parts to minimize the surface area of ​​contact with water. This interaction also plays a role during antigen-antibody binding. When a hydrophobic amino acid in the antigen sequence comes into close proximity to a hydrophobic amino acid in the antibody sequence, water molecules are released, increasing entropy and consequently lowering the free energy of the binding. This contributes to the greater stabilization of the antigen-antibody binding.

[0312] The information related to the above hydrophobic interactions may include information regarding hydrophobic amino acids in the antigen sequence and hydrophobic amino acid pairs in the antibody sequence. Here, the information regarding amino acid pairs includes the position of each amino acid within the antigen sequence or antibody sequence, and may further include the strength of the hydrophobic interaction of the corresponding amino acid pair.

[0313] Chemical Property Example #4 - Electrostatic Interaction

[0314] The above chemical properties may include information related to electrostatic interactions. Specifically, the antigen-antibody binding affinity prediction model of this specification may utilize information on amino acid pairs that electrostatically interact with each other in the antigen sequence and the antibody sequence.

[0315] Electrostatic interactions refer to the attractive or repulsive forces that occur between charged amino acid residues. In the body's environment, there are positively charged and negatively charged amino acids. For example, lysine, arginine, and histidine carry a positive charge, while aspartic acid and glutamic acid carry a negative charge. Electrostatic interactions occur between charged objects. Specifically, forces of attraction arise between opposite charges, while forces of repulsion arise between like charges. Consequently, during antigen-antibody binding, amino acid pairs with opposite charges cause the antigen and antibody to bind more strongly, while amino acid pairs with like charges hinder the binding of the antigen and antibody.

[0316] The information related to the above electrostatic interactions may include: 1) information about a pair of positively charged amino acids in the antigen sequence and a negatively charged amino acid in the antibody sequence, 2) information about a pair of positively charged amino acids in the antibody sequence and a negatively charged amino acid in the antigen sequence, or 3) both 1) and 2). Here, the information about the amino acid pairs includes the position of each amino acid within the antigen sequence or antibody sequence and may further include the strength of the electrostatic interaction of the said amino acid pair.

[0317] Examples of expressions of information regarding chemical properties

[0318] The information regarding the above chemical properties can be represented as a matrix or its equivalent. This is referred to as a chemical property matrix. The above chemical property matrix is ​​a two-dimensional matrix with dimensions of (length of antigen sequence) × (length of antibody sequence), or (length of antibody sequence) × (length of antigen sequence). For example, let the above chemical property matrix be C, and when it has dimensions of (length of antigen sequence) × (length of antibody sequence), each element of C represents the presence or strength of interaction between the i-th amino acid of the antigen sequence and the j-th amino acid of the antibody sequence. As another example, let C be the chemical property matrix above, and when it has dimensions (length of antibody sequence) × (length of antigen sequence), each element of C represents whether there is an interaction or the strength of the interaction between the i-th amino acid of the antibody sequence and the j-th amino acid of the antigen sequence. Here, the above interaction may be the aforementioned hydrogen bond, van der Waals interaction, hydrophobic interaction, or electrostatic interaction.

[0319] Method of utilizing chemical properties in antigen-antibody binding affinity prediction models

[0320] The antigen-antibody binding affinity prediction model of this specification is characterized by learning antigen-antibody interactions by considering chemical properties. Specifically, these chemical properties are integrated into the process in which the model performs cross-attention calculations. During cross-attention, contextual information is generated by deriving interaction information between all amino acids within the antibody sequence and the antigen sequence. In this process, locations that interact significantly are designated based on the chemical properties, and high weights are assigned to the corresponding property values. This is why the positional information of each amino acid within the antigen or antibody sequence is important.

[0321] The cross-attention calculation characteristically used in this model is referred to as "chemical cross-attention" and is described in detail in the following paragraphs.

[0322] Chemical Cross Attention

[0323] The meaning of chemical cross-attention

[0324] The antigen-antibody binding affinity prediction model of this specification includes a chemical cross-attention mechanism.

[0325] The chemical cross-attention described above is a mechanism introduced to learn information regarding interactions between antigen sequences or information derived therefrom and antibody sequences or information derived therefrom, and to reflect this in the results. Through chemical cross-attention, the system intensively learns about interactions at chemically significant positions among the interactions between antigen and antibody sequences. This mechanism differs from standard cross-attention, which treats all elemental interactions equally for learning.

[0326] Calculation Structure of Chemical Cross Attention #1 - Overview (s100)

[0327] The computational structure of Chemical Cross Attention (CCA) can be expressed by the following equation:

[0328]

[0329] Here, Q represents the query, K represents the key, and V represents the value matrix. The query originates from antigen sequence information, and the key and value originate from antibody sequence information, or the query originates from antibody sequence information, and the key and value originate from antigen sequence information. This is a variable also used in standard cross-attention. G c represents the Chemical Attention Guide. This is a characteristic element of Chemical Cross-Attention and is derived from chemical properties. Here, the query, key, and value are determined based on information derived from the antigen sequence or antibody sequence. The Chemical Attention Guide is determined from the aforementioned chemical properties. The CCA Value is the result of Chemical Cross-Attention and includes interaction information between the antigen sequence information and the antibody sequence information. In the context in which this model is used, the above "interaction information" refers to the interactions that contribute to antigen-antibody binding.

[0330] A method for implementing the chemical cross-attention mechanism above is explained with reference to the flowchart of FIG. 7. The chemical cross-attention of FIG. 7 includes the following calculations:

[0331] 1) Receive input information (s110). Here, input information refers to the input for chemical cross attention, and is referred to as "block input" to distinguish it from the input of the model below.

[0332] 2) Obtain a query (s120) and key and value (s130) from the input information.

[0333] 3) Calculate the similarity between the query and the key, and derive attention scores based on the similarity (s140).

[0334] 4) Adjust the importance of attention scores calculated using a chemical attention guide derived from the chemical characteristics between the query sequence and the key sequence (s150).

[0335] 5) A CCA value reflecting context information is derived using the adjusted attention scores and values ​​(s160). The above CCA value is referred to as "block output".

[0336] The meaning and specific implementation of each variable and calculation are explained below.

[0337] Calculation Structure of Chemical Cross Attention #2 - Block Input (s110) and Block Output (s160)

[0338] Chemical cross-attention receives information derived from an antigen sequence and information derived from an antibody sequence as block inputs (s110). Here, "information derived from the sequence" encompasses both the input embedding information for the corresponding sequence and intermediate representations obtained by transforming that input embedding information in various ways. Generally, the above chemical cross-attention is implemented as a single module and can receive information that has undergone computations by other modules (e.g., embedding, self-attention, neural network computation, etc.) as input. In this case, it can be described that intermediate representations for the antigen sequence and intermediate representations for the antibody sequence are received as block inputs.

[0339] The Chemical Cross Attention above calculates and outputs a CCA value from the above block input (s160). This is referred to as the Block Output. From the perspective of the overall model, the above Block Output can be viewed as an Intermediate Representation, as it represents a value that appears during the intermediate process where the information input to the model is transformed in various ways. The above Block Output can be used as an input for another Chemical Cross Attention or other modules.

[0340] Calculation Structure of Chemical Cross Attention #3 - Query Derivation (s120)

[0341] Chemical cross-attention includes the process of calculating a query from information derived from an antigen sequence or information derived from an antibody sequence (s120). A unit of chemical cross-attention uses only one of the information derived from the antigen sequence or the information derived from the antibody sequence as a query. For convenience of description, the sequence determined as the query of the cross-attention is referred to as the "query sequence." Generally, the query is derived by computing the information derived from the query sequence with query weights. Here, the above query weights are learnable. The above query is represented as a query matrix or its equivalent.

[0342] Calculation Structure of Chemical Cross Attention #4 - Derivation of Keys and Values ​​(s130)

[0343] Chemical Cross Attention involves the process of calculating keys and values ​​from information derived from sequences other than the query sequence. Since Chemical Cross Attention belongs to the category of Cross Attention that models the interaction between two objects, the query information and the key and value information must originate from different objects. The sequence from which the key and value are derived is referred to as the "key sequence." Generally, the key and value are derived by computing the information derived from the key sequence with its corresponding weights, respectively. That is, the key is the result of computing the information derived from the key sequence with its key weights, and the value is the result of computing the information derived from the value sequence with its value weights. Here, the key weights and value weights are learnable. The above key and value are represented by their respective matrices or equivalents. For example, if the query sequence is an antigen sequence and the key sequence is an antibody sequence, Chemical Cross Attention calculates the query from the information derived from the antigen sequence and calculates the key and value from the information derived from the antibody sequence. In this case, due to the Cross Attention calculation structure, the key and value must have the same sequence length.

[0344] Chemical Cross Attention Calculation Structure #5 - Calculation of Query and Key Similarity and Derivation of Attention Score (s140)

[0345] Chemical cross attention includes a process of determining the similarity between a query and a key and deriving an attention score based thereon (s140). The similarity between a query and a key can be determined based on various calculated values. For example, the similarity between a query and a key can be determined based on the dot product, weighted dot product, cosine similarity, Euclidean distance, or Manhattan distance between the query matrix and the key matrix. The attention score refers to values ​​obtained by appropriately transforming the similarity between a query and a key into a probability distribution. For example, the attention score may be a value obtained by applying a softmax, entropymax, sigmoid, or Gaussian kernel function to the similarity values ​​between the query and the key. In the context of determining antigen-antibody binding affinity, the attention score contains information about the degree of interaction at each position of the query sequence and the key sequence.

[0346] Calculation Structure of Chemical Cross Attention #6 - Chemical Attention Guide Calculation (s150)

[0347] Chemical cross-attention includes a process of converting the attention scores derived above according to a chemical attention guide (s150). The chemical attention guide is derived from the chemical characteristics between the antigen sequence and the antibody sequence. These chemical characteristics are used in a way that utilizes the chemical attention guide based thereon to calculate the CCA value.

[0348] The attention score value calculated above represents the degree of interaction at each position between the query sequence and the key sequence. Since this value is based on the similarity between the query and the key, it is a result in which information from all positions within the sequence is considered equally. However, when actual antigen-antibody binding occurs, positions involved in intermolecular binding will contribute significantly to the binding, while other positions will contribute less. Therefore, if greater weight is given to interactions between chemically significant positions during model training, it can be expected that the model will not only learn more efficiently but also predict antigen-antibody binding affinity with higher accuracy. To implement this process using chemical cross-attention, a chemical attention guide was introduced.

[0349] The chemical attention guide operation is designed to highlight attention scores corresponding to the positions of amino acid pairs included in the aforementioned chemical characteristics. For example, if the chemical characteristics include information related to hydrogen bonding, the chemical attention guide is designed to increase the weight of attention scores corresponding to the positions of amino acid pairs involved in hydrogen bonding. The above chemical attention guide operation can be implemented in various ways. For example, the chemical attention guide operation can be designed to retain only the attention scores corresponding to the positions of amino acid pairs related to the chemical characteristics and remove (mask) all attention scores for the remaining positions. As another example, if the interaction strength for each amino acid pair related to the chemical characteristics is available, the chemical attention guide operation can be designed to amplify or attenuate each attention score in proportion to the respective interaction strength. As a more specific example, the chemical attention guide operation may be a process of performing operations on the chemical attention score matrix and the aforementioned chemical characteristic matrix or its equivalent. Here, let S be the chemical attention score matrix, and each element of S is the cross-attention score between the i-th amino acid of the query matrix and the j-th amino acid of the key matrix. Also, when the chemical property matrix is ​​denoted as C, each element of C represents whether there is an interaction or the degree of interaction between the i-th amino acid of the query matrix and the j-th amino acid of the key matrix. Here, the above operation can be matrix addition or Hadamad product.

[0350] If the above chemical properties contain information regarding various types of intermolecular interactions, chemical attention guides can be designed differently for each type of interaction. For example, chemical cross-attention may include a first chemical guide computation process reflecting van der Waals interactions and a second chemical guide computation process reflecting electrostatic interactions. Multiple chemical attention guide computations can be applied in parallel for each head when the chemical cross-attention is a multi-head attention structure. This is explained in more detail in the section "Implementation of Chemical Attention Guides #3 - Case of Multi-Head Attention".

[0351] Chemical Cross Attention Calculation Structure #7 - Derivation of CCA Value (s160)

[0352] Chemical Cross Attention calculates the CCA value by weighting the value sum according to the attention score converted according to the chemical attention guide and converting it into a learnable weight (s160). The CCA value is information about the query sequence and reflects the interaction information between the query sequence and the key sequence. By utilizing this, information about the chemical interaction between the antigen sequence and the antibody sequence can be strongly reflected when predicting the final output value (antigen-antibody binding affinity).

[0353] Calculation Structure of Chemical Cross Attention #8 - Case of Multi-Head Attention (s200)

[0354] Chemical Cross Attention may include multi-head attention calculation. A method for implementing the above multi-head attention calculation is described with reference to the flowchart in FIG. 8. This is a method for deriving information by considering the interaction between the query and the key in parallel from various perspectives. In the multi-head attention calculation, the block input (s210) is split into multiple heads to derive the query (s220) and the key and value (s230) respectively, and the aforementioned calculation process is performed for each head (s240). Then, the CCA values ​​of each head are connected (s250) and appropriately transformed to derive a single CCA value (s260). Here, the appropriate transformation may be a value obtained by converting the connected CCA value into a learnable weight.

[0355] Implementation of Chemical Attention Guide #1 - When Sequence Position Information Is Preserved

[0356] When designing a model, it can be configured so that sequence position information is preserved during the process of converting input antigen and antibody sequence information. In such a model, even if the information value of each position changes, the specific amino acid position within the sequence that each position represents remains unchanged. When chemical cross-attention is performed using information with preserved sequence position data as block input, each attention score can be viewed as a characteristic value representing the interaction between each amino acid pair. Therefore, when applying a chemical attention guide to such a model, the information regarding each amino acid pair included in the chemical characteristics can be utilized as is. For example, if the chemical characteristic utilizes information that the 1st amino acid of the antigen and the 3rd amino acid of the antibody form a hydrogen bond, the chemical attention guide can be designed to emphasize the attention score between the 1st position derived from the antigen sequence and the 3rd position derived from the antibody sequence.

[0357] Implementation of Chemical Attention Guide #2 - When Sequence Position Information Is Not Preserved

[0358] Sequence position information may not be preserved during the conversion process between the antigen and antibody sequences input into the model. For example, this can occur when long sequence information is restructured and compressed, or when additional information is added to extend the sequence. If an intermediate representation in which sequence position information is not preserved is used as a block input for a chemical attention guide, the calculated attention score can no longer be regarded as a characteristic value representing the interaction of each amino acid pair. This is because the meaning of each position has changed during the information conversion process. In this case, chemical characteristic information must be used after being appropriately converted based on the logic used to transform each sequence information.

[0359] Implementation of Chemical Attention Guides #3 - For Multi-Head Attention

[0360] As described above, multiple chemical attention guides can be configured according to the types of intermolecular interactions. When the chemical cross-attention is a multi-head attention structure, the multiple chemical attention guide operations can be applied in parallel for each head. For example, when the above multi-head attention includes a first head, a second head, a third head, and a fourth head, the chemical attention guide operation for hydrogen bonding can be performed on the first head, the chemical attention guide operation for van der Waals interactions can be performed on the second head, the chemical attention guide operation for hydrophobic interactions can be performed on the third head, and the chemical attention guide operation considering electrostatic interactions can be performed on the fourth head. In this case, since the types of intermolecular interactions considered for each head are clearly separated and the results are integrated into the block output value, the interactions of each amino acid pair during antigen-antibody binding are learned more efficiently.

[0361] Chemical Cross Attention

[0362] The meaning of chemical cross-attention

[0363] The antigen-antibody binding affinity prediction model of this specification includes a chemical cross-attention mechanism.

[0364] The chemical cross-attention described above is a mechanism introduced to learn information regarding interactions between antigen sequences or information derived therefrom and antibody sequences or information derived therefrom, and to reflect this in the results. Through chemical cross-attention, the system intensively learns about interactions at chemically significant positions among the interactions between antigen and antibody sequences. This mechanism differs from standard cross-attention, which treats all elemental interactions equally for learning.

[0365] Calculation Structure of Chemical Cross Attention #1 - Overview (s100)

[0366] The computational structure of Chemical Cross Attention (CCA) can be expressed by the following equation:

[0367]

[0368] Here, Q represents the query, K represents the key, and V represents the value matrix. The query originates from antigen sequence information, and the key and value originate from antibody sequence information, or the query originates from antibody sequence information, and the key and value originate from antigen sequence information. This is a variable also used in standard cross-attention. G c represents the Chemical Attention Guide. This is a characteristic element of Chemical Cross-Attention and is derived from chemical properties. Here, the query, key, and value are determined based on information derived from the antigen sequence or antibody sequence. The Chemical Attention Guide is determined from the aforementioned chemical properties. The CCA Value is the result of Chemical Cross-Attention and includes interaction information between the antigen sequence information and the antibody sequence information. In the context in which this model is used, the above "interaction information" refers to the interactions that contribute to antigen-antibody binding.

[0369] A method for implementing the chemical cross-attention mechanism above is explained with reference to the flowchart of FIG. 7. The chemical cross-attention of FIG. 7 includes the following calculations:

[0370] 1) Receive input information (s110). Here, input information refers to the input for chemical cross attention, and is referred to as "block input" to distinguish it from the input of the model below.

[0371] 2) Obtain a query (s120) and key and value (s130) from the input information.

[0372] 3) Calculate the similarity between the query and the key, and derive attention scores based on the similarity (s140).

[0373] 4) Adjust the importance of attention scores calculated using a chemical attention guide derived from the chemical characteristics between the query sequence and the key sequence (s150).

[0374] 5) A CCA value reflecting context information is derived using the adjusted attention scores and values ​​(s160). The above CCA value is referred to as "block output".

[0375] The meaning and specific implementation of each variable and calculation are explained below.

[0376] Calculation Structure of Chemical Cross Attention #2 - Block Input (s110) and Block Output (s160)

[0377] Chemical cross-attention receives information derived from an antigen sequence and information derived from an antibody sequence as block inputs (s110). Here, "information derived from the sequence" encompasses both the input embedding information for the corresponding sequence and intermediate representations obtained by transforming that input embedding information in various ways. Generally, the above chemical cross-attention is implemented as a single module and can receive information that has undergone computations by other modules (e.g., embedding, self-attention, neural network computation, etc.) as input. In this case, it can be described that intermediate representations for the antigen sequence and intermediate representations for the antibody sequence are received as block inputs.

[0378] The Chemical Cross Attention above calculates and outputs a CCA value from the above block input (s160). This is referred to as the Block Output. From the perspective of the overall model, the above Block Output can be viewed as an Intermediate Representation, as it represents a value that appears during the intermediate process where the information input to the model is transformed in various ways. The above Block Output can be used as an input for another Chemical Cross Attention or other modules.

[0379] Calculation Structure of Chemical Cross Attention #3 - Query Derivation (s120)

[0380] Chemical cross-attention includes the process of calculating a query from information derived from an antigen sequence or information derived from an antibody sequence (s120). A unit of chemical cross-attention uses only one of the information derived from the antigen sequence or the information derived from the antibody sequence as a query. For convenience of description, the sequence determined as the query of the cross-attention is referred to as the "query sequence." Generally, the query is derived by computing the information derived from the query sequence with query weights. Here, the above query weights are learnable. The above query is represented as a query matrix or its equivalent.

[0381] Calculation Structure of Chemical Cross Attention #4 - Derivation of Keys and Values ​​(s130)

[0382] Chemical Cross Attention involves the process of calculating keys and values ​​from information derived from sequences other than the query sequence. Since Chemical Cross Attention belongs to the category of Cross Attention that models the interaction between two objects, the query information and the key and value information must originate from different objects. The sequence from which the key and value are derived is referred to as the "key sequence." Generally, the key and value are derived by computing the information derived from the key sequence with its corresponding weights, respectively. That is, the key is the result of computing the information derived from the key sequence with its key weights, and the value is the result of computing the information derived from the value sequence with its value weights. Here, the key weights and value weights are learnable. The above key and value are represented by their respective matrices or equivalents. For example, if the query sequence is an antigen sequence and the key sequence is an antibody sequence, Chemical Cross Attention calculates the query from the information derived from the antigen sequence and calculates the key and value from the information derived from the antibody sequence. In this case, due to the Cross Attention calculation structure, the key and value must have the same sequence length.

[0383] Chemical Cross Attention Calculation Structure #5 - Calculation of Query and Key Similarity and Derivation of Attention Score (s140)

[0384] Chemical cross attention includes a process of determining the similarity between a query and a key and deriving an attention score based thereon (s140). The similarity between a query and a key can be determined based on various calculated values. For example, the similarity between a query and a key can be determined based on the dot product, weighted dot product, cosine similarity, Euclidean distance, or Manhattan distance between the query matrix and the key matrix. The attention score refers to values ​​obtained by appropriately transforming the similarity between a query and a key into a probability distribution. For example, the attention score may be a value obtained by applying a softmax, entropymax, sigmoid, or Gaussian kernel function to the similarity values ​​between the query and the key. In the context of determining antigen-antibody binding affinity, the attention score contains information about the degree of interaction at each position of the query sequence and the key sequence.

[0385] Calculation Structure of Chemical Cross Attention #6 - Chemical Attention Guide Calculation (s150)

[0386] Chemical cross-attention includes a process of converting the attention scores derived above according to a chemical attention guide (s150). The chemical attention guide is derived from the chemical characteristics between the antigen sequence and the antibody sequence. These chemical characteristics are used in a way that utilizes the chemical attention guide based thereon to calculate the CCA value.

[0387] The attention score value calculated above represents the degree of interaction at each position between the query sequence and the key sequence. Since this value is based on the similarity between the query and the key, it is a result in which information from all positions within the sequence is considered equally. However, when actual antigen-antibody binding occurs, positions involved in intermolecular binding will contribute significantly to the binding, while other positions will contribute less. Therefore, if greater weight is given to interactions between chemically significant positions during model training, it can be expected that the model will not only learn more efficiently but also predict antigen-antibody binding affinity with higher accuracy. To implement this process using chemical cross-attention, a chemical attention guide was introduced.

[0388] The chemical attention guide operation is designed to highlight attention scores corresponding to the positions of amino acid pairs included in the aforementioned chemical characteristics. For example, if the chemical characteristics include information related to hydrogen bonding, the chemical attention guide is designed to increase the weight of attention scores corresponding to the positions of amino acid pairs involved in hydrogen bonding. The above chemical attention guide operation can be implemented in various ways. For example, the chemical attention guide operation can be designed to retain only the attention scores corresponding to the positions of amino acid pairs related to the chemical characteristics and remove (mask) all attention scores for the remaining positions. As another example, if the interaction strength for each amino acid pair related to the chemical characteristics is available, the chemical attention guide operation can be designed to amplify or attenuate each attention score in proportion to the respective interaction strength. As a more specific example, the chemical attention guide operation may be a process of performing operations on the chemical attention score matrix and the aforementioned chemical characteristic matrix or its equivalent. Here, let S be the chemical attention score matrix, and each element of S is the cross-attention score between the i-th amino acid of the query matrix and the j-th amino acid of the key matrix. Also, when the chemical property matrix is ​​denoted as C, each element of C represents whether there is an interaction or the degree of interaction between the i-th amino acid of the query matrix and the j-th amino acid of the key matrix. Here, the above operation can be matrix addition or Hadamad product.

[0389] If the above chemical properties contain information regarding various types of intermolecular interactions, chemical attention guides can be designed differently for each type of interaction. For example, chemical cross-attention may include a first chemical guide computation process reflecting van der Waals interactions and a second chemical guide computation process reflecting electrostatic interactions. Multiple chemical attention guide computations can be applied in parallel for each head when the chemical cross-attention is a multi-head attention structure. This is explained in more detail in the section "Implementation of Chemical Attention Guides #3 - Case of Multi-Head Attention".

[0390] Chemical Cross Attention Calculation Structure #7 - Derivation of CCA Value (s160)

[0391] Chemical Cross Attention calculates the CCA value by weighting the value sum according to the attention score converted according to the chemical attention guide and converting it into a learnable weight (s160). The CCA value is information about the query sequence and reflects the interaction information between the query sequence and the key sequence. By utilizing this, information about the chemical interaction between the antigen sequence and the antibody sequence can be strongly reflected when predicting the final output value (antigen-antibody binding affinity).

[0392] Calculation Structure of Chemical Cross Attention #8 - Case of Multi-Head Attention (s200)

[0393] Chemical Cross Attention may include multi-head attention calculation. A method for implementing the above multi-head attention calculation is described with reference to the flowchart in FIG. 8. This is a method for deriving information by considering the interaction between the query and the key in parallel from various perspectives. In the multi-head attention calculation, the block input (s210) is split into multiple heads to derive the query (s220) and the key and value (s230) respectively, and the aforementioned calculation process is performed for each head (s240). Then, the CCA values ​​of each head are connected (s250) and appropriately transformed to derive a single CCA value (s260). Here, the appropriate transformation may be a value obtained by converting the connected CCA value into a learnable weight.

[0394] Implementation of Chemical Attention Guide #1 - When Sequence Position Information Is Preserved

[0395] When designing a model, it can be configured so that sequence position information is preserved during the process of converting input antigen and antibody sequence information. In such a model, even if the information value of each position changes, the specific amino acid position within the sequence that each position represents remains unchanged. When chemical cross-attention is performed using information with preserved sequence position data as block input, each attention score can be viewed as a characteristic value representing the interaction between each amino acid pair. Therefore, when applying a chemical attention guide to such a model, the information regarding each amino acid pair included in the chemical characteristics can be utilized as is. For example, if the chemical characteristic utilizes information that the 1st amino acid of the antigen and the 3rd amino acid of the antibody form a hydrogen bond, the chemical attention guide can be designed to emphasize the attention score between the 1st position derived from the antigen sequence and the 3rd position derived from the antibody sequence.

[0396] Implementation of Chemical Attention Guide #2 - When Sequence Position Information Is Not Preserved

[0397] Sequence position information may not be preserved during the conversion process between the antigen and antibody sequences input into the model. For example, this can occur when long sequence information is restructured and compressed, or when additional information is added to extend the sequence. If an intermediate representation in which sequence position information is not preserved is used as a block input for a chemical attention guide, the calculated attention score can no longer be regarded as a characteristic value representing the interaction of each amino acid pair. This is because the meaning of each position has changed during the information conversion process. In this case, chemical characteristic information must be used after being appropriately converted based on the logic used to transform each sequence information.

[0398] Implementation of Chemical Attention Guides #3 - For Multi-Head Attention

[0399] As described above, multiple chemical attention guides can be configured according to the types of intermolecular interactions. When the chemical cross-attention is a multi-head attention structure, the multiple chemical attention guide operations can be applied in parallel for each head. For example, when the above multi-head attention includes a first head, a second head, a third head, and a fourth head, the chemical attention guide operation for hydrogen bonding can be performed on the first head, the chemical attention guide operation for van der Waals interactions can be performed on the second head, the chemical attention guide operation for hydrophobic interactions can be performed on the third head, and the chemical attention guide operation considering electrostatic interactions can be performed on the fourth head. In this case, since the types of intermolecular interactions considered for each head are clearly separated and the results are integrated into the block output value, the interactions of each amino acid pair during antigen-antibody binding are learned more efficiently.

[0400] Training method of antigen-antibody binding affinity prediction model

[0401] Overview of the Training Method for Antigen-Antibody Binding Affinity Prediction Models (s400)

[0402] This specification discloses a method for training an antigen-antibody binding affinity prediction model. The antigen-antibody binding affinity prediction model described above has been explained above. A method for training the antigen-antibody binding affinity prediction model is illustrated in FIGS. 10 and FIGS. 11. The following description refers to FIGS. 10 and FIGS. 11. The training method includes a process of acquiring training data, a process of identifying chemical properties, and a process of training a model. The training data includes pairs of antigen sequences and antibody sequences labeled with antigen-antibody binding affinities. The process of identifying chemical properties is a process of identifying the aforementioned chemical properties for each pair of antigen sequences and antibody sequences in the training data. For example, for each antigen sequence and antibody sequence, it may be a process of identifying information on amino acid pairs capable of forming hydrogen bonds with each other. The method includes a process of training an antigen-antibody binding affinity prediction model using the training data and chemical properties. The training process follows a standard training method.

[0403] Structure Summary of the Antigen-Antibody Binding Affinity Prediction Model

[0404] The antigen-antibody binding affinity prediction model trained in the above method is as described above. The key points regarding the above model are summarized as follows:

[0405] 1) Receive antigen sequences and antibody sequences as input, and output antigen-antibody binding affinity.

[0406] 2) Adopts a chemical cross-attention mechanism.

[0407] 3) The above chemical cross-attention mechanism receives information derived from the antigen sequence and information derived from the antibody sequence as block inputs, and outputs a CCA value.

[0408] 4) The above chemical cross-attention mechanism is performed using a chemical attention guide based on the chemical properties between the antigen sequence and the antibody sequence. Here, the above chemical attention guide is performed by assigning weights to the attention scores of important positions among the cross-attention scores.

[0409] Acquired training data (s410)

[0410] The above method for training the antigen-antibody binding affinity prediction model includes the process of acquiring training data. The above training data includes pairs of antigen sequences and antibody sequences and antigen-antibody binding affinity values ​​between the two sequences. These correspond to the input and output of the model, respectively. As long as the above training data contains the above information, the source thereof is not otherwise limited. For example, the above training data may be obtained from a known antigen-antibody database.

[0411] Identifying chemical properties (s420)

[0412] The above method for training the antigen-antibody binding affinity prediction model includes a process of identifying chemical characteristics. Here, chemical characteristics are identified for each pair of antigen and antibody sequences, and the specific details are as described above. The timing at which chemical characteristics are identified may vary depending on the specific model. For example, the above chemical characteristics may be identified immediately after acquiring the above training data. As another example, the above chemical characteristics may be identified when calculating chemical cross-attention during the process of training the model.

[0413] Train model(s430)

[0414] This will be explained with reference to Figures 10 and 11. The method for training the above antigen-antibody binding affinity prediction model includes a process of training the model. The process of training the above model is performed using a standard method.

[0415] For example, the process of training the above model may include the following:

[0416] 1) Initialize the parameters of the model (s431). Here, the parameters include all learnable weights and are initialized in an appropriate way depending on the specific implementation of the model.

[0417] 2) Input one unit of data (s432). Here, the one unit of data may be one data (antigen sequence-antibody sequence pair labeled with antigen-antibody binding affinity), or one batch of data.

[0418] 3) A predicted value of antigen-antibody binding affinity is calculated from the above data (s433). Here, the process of calculating the predicted value of antigen-antibody binding affinity includes the process of performing chemical cross-attention calculation. Here, information derived from the antigen sequence and information derived from the antibody sequence are used as block inputs for the chemical cross-attention. In addition, when calculating the chemical cross-attention, a chemical attention guide operation based on the chemical characteristics identified in the previous process is performed. In addition, the antigen-antibody binding affinity is calculated using the result of the chemical cross-attention calculation, i.e., the aforementioned CCA value.

[0419] 4) Calculate the loss function using the above predicted values ​​and the actual antigen-antibody binding affinities (s434). The above loss function can be appropriately selected depending on the configuration of the model and is not otherwise limited. For example, it may be a Mean Squared Error function, Cross-Entropy Loss, or Huber Loss function.

[0420] 5) Adjust the above parameters in a direction that minimizes the above loss function (s435).

[0421] 6) Repeat steps 2) through 5) with new data until a predetermined criterion is satisfied (s436) (s437).

[0422] Antigen-antibody binding affinity prediction method

[0423] This specification discloses a method for predicting antigen-antibody binding affinity. The above method uses the aforementioned antigen-antibody binding affinity prediction model. A method for training the above antigen-antibody binding affinity prediction model has also been described above. The above method includes a process of obtaining an antigen sequence and an antibody sequence, a process of identifying the chemical characteristics of the above antigen sequence and antibody sequence, and a process of predicting antigen-antibody binding affinity. Here, the chemical characteristics of the antigen sequence and antibody sequence refer to information necessary to derive a chemical attention guide applied to the chemical cross-attention calculation of the antigen-antibody binding affinity prediction model. Once only the antigen sequence and antibody sequence are identified, the above method can be performed to predict antigen-antibody binding affinity.

[0424] Examples of applications for antigen-antibody binding affinity prediction models

[0425] Application Example #1 - Reinforcement Learning for Antibody Generation Models

[0426] This specification discloses a method for reinforcement learning an antibody generation model utilizing an antigen-antibody binding affinity prediction model as a reward model. The method for reinforcement learning the above antibody generation model includes the following:

[0427] 1) A process of pre-training an antibody generation model, wherein the antibody generation model is configured to output antibodies that bind to the antigen of interest, and through pre-training, learns the general patterns and characteristics of antibodies that bind to the antigen of interest.

[0428] 2) A reinforcement learning process is performed by combining a pre-trained antibody generation model with a completed antigen-antibody binding affinity prediction model, wherein the antigen-antibody binding affinity prediction model is the aforementioned model and is trained using the method described above, and the antibody sequence output from the antibody generation model and the sequence of the antigen of interest are input into the antigen-antibody binding affinity prediction model to obtain a prediction value, and the parameters of the antibody generation model are adjusted in a direction that maximizes the prediction value.

[0429] Through the above method, the antibody generation model can be optimized to produce antibodies that bind better to the antigen of interest. This enables the construction of an effective antibody generation model even in environments with limited training data.

[0430] Usage Example #2 - In Silico Library Screening

[0431] This specification discloses an in silico library screening method utilizing an antigen-antibody binding affinity prediction model. The above in silico library screening method includes the following:

[0432] 1) A process of obtaining an in silico library, wherein the above in silico library contains multiple antibody sequence information.

[0433] 2) A process for selecting antibody candidates that bind to the antigen of interest, wherein each antibody sequence and the antigen of interest included in the above in silico library are input into the aforementioned antigen-antibody binding affinity prediction model to obtain a prediction value, and if the prediction value satisfies a predetermined criterion, it is selected as an antibody candidate.

[0434] Through the above method, antibody candidates with a high probability of binding to the antigen of interest can be selected from among multiple antibody sequences included in an in silico library. This allows for more efficient screening by reducing the cost of verification through actual experiments.

[0435] Device implementing an antigen-antibody binding affinity prediction model

[0436] Overview of a device implementing an antigen-antibody binding affinity prediction model

[0437] Referring to FIG. 12, the following is explained. Information for implementing the antigen-antibody binding affinity prediction model and the antibody generation model disclosed herein may be stored in a computer-readable storage medium (500 or 630). The antigen-antibody binding affinity prediction model, the learning method of the said model, the antigen-antibody binding affinity prediction method, and examples of their application disclosed herein may all be implemented in a computer device (600). A computer device (600) implementing the said methods is illustrated in the figure below:

[0438] Computer-readable storage media (500 or 630)

[0439] Referring to FIG. 12, the description is as follows: This specification discloses a computer-readable storage medium (500 or 630) in which information related to the aforementioned model is stored. The computer-readable storage medium may exist independently of a computer device (600) (500) and may be part of a computer device (630). The computer-readable storage medium retains information about the model even when the model is not in use and allows it to be loaded into memory (610) when needed. The computer-readable storage medium is not otherwise limited as long as it is a medium capable of storing information. For example, the computer-readable storage medium may be a hard disk drive (HDD), a solid-state drive (SSD), flash memory (FLASH), or cloud storage.

[0440] One or more processors (610)

[0441] Referring to FIG. 12, this specification discloses a computer device (600) that implements the aforementioned model. The computer device includes one or more processors (610). The processors (610) perform calculations essential to the implementation of the model, such as matrix multiplication, convolution, and various function calculations. The processors (610) process input data using model computation structures and parameters loaded in memory and perform prediction or generation through forward propagation. Additionally, when training the model, they process calculations to update parameters through back propagation or gradient descent. The processors (610) are not otherwise limited as long as they are processors capable of performing the aforementioned operations and operating or training the model. For example, the processors may be a CPU, GPU, TPU, or other special-purpose processors.

[0442] One or more memories (620)

[0443] Referring to FIG. 12, this specification discloses a computer device (600) that implements the aforementioned model. The computer device includes one or more memories (620). The memory (620) is a temporary workspace that stores information that the processor needs to access immediately while running or training the model. When the model is executed, the computational structure and parameters of the model, and other necessary information, are loaded into the memory (620) from a computer-readable storage medium (630). Additionally, values ​​necessary for model execution, such as input data, intermediate calculation results, activation values, and gradients, are also stored. The memory (620) is not otherwise limited as long as it is a memory capable of achieving the above purpose. For example, the memory may be RAM, HBM, or a distributed memory system.

[0444] Model Implementation Device #1 - Device storing the antigen-antibody binding affinity prediction model

[0445] Referring to FIG. 12, this specification discloses a computer-readable storage medium (500 or 630) in which an antigen-antibody binding affinity prediction model is stored. The computer-readable storage medium may be an independent medium (500) or part of a computer device (630). The computer-readable storage medium stores the computational structure (architecture) of the antigen-antibody binding affinity prediction model disclosed in this specification, learned parameters (weights and biases), hyperparameters, and other metadata necessary for model implementation.

[0446] Model Implementation Device #2 - Device storing the antibody generation model

[0447] Referring to FIG. 12, this specification discloses a computer-readable storage medium (500) in which an antibody generation model is stored. The computer-readable storage medium may be an independent medium (500) or part of a computer device (630). The computer-readable storage medium stores the computational structure (architecture) of the antigen generation model disclosed in this specification, learned parameters (weights and biases), hyperparameters, and other metadata necessary for model implementation.

[0448] Model Implementation Device #3 - Device implementing a method for training an antigen-antibody binding affinity prediction model

[0449] Referring to FIG. 12, the following description is provided. This specification discloses a computer device (600) for implementing a method for learning an antigen-antibody binding affinity prediction model. The computer device (600) comprises one or more processors (610), one or more memories (620), and one or more computer-readable storage media (630). Herein, the processor (610), the memory (620), and the computer-readable storage media (630) are capable of communicating with each other. The computer-readable storage media (630) stores training data, a model computational structure, and a computer-implementable method for learning the antigen-antibody binding affinity prediction model (hereinafter referred to as a learning algorithm). Herein, the learning algorithm is as described above. When the above computer device (600) performs the antigen-antibody binding affinity prediction model learning method, the learning data, model computation structure, and learning algorithm are loaded from the above computer-readable storage medium (630) into memory (620), the above processor (610) performs the above learning algorithm using the learning data and model computation structure loaded into the above memory (620), and the learned parameters as a result of performing the above learning algorithm are stored in the above computer-readable storage medium (630).

[0450] Model Implementation Device #4 - Device implementing a method to predict antigen-antibody binding affinity

[0451] Referring to FIG. 12, the description is as follows. This specification discloses a computer device (600) for implementing a method for learning an antigen-antibody binding affinity prediction model. The computer device (600) comprises one or more processors (610), one or more memories (620), and one or more computer-readable storage media (630). Here, the processors (610), the memories (620), and the computer-readable storage media (630) are capable of communicating with each other. The computer-readable storage media (630) stores the computational structure (architecture) of the antigen-antibody binding affinity prediction model disclosed in this specification, learned parameters (weights and biases), hyperparameters, and other metadata necessary for model implementation. When the above computer device (600) performs an antigen-antibody binding affinity prediction method, the computational structure (architecture) of the antigen-antibody binding affinity prediction model, learned parameters (weights and biases), hyperparameters, and other metadata required for model implementation are loaded from the above computer-readable storage medium (630) to the memory (620), input data is loaded into the above memory (620), and the above processor (610) forward propagates the input data loaded into the above memory (620) to the antigen-antibody binding affinity prediction model to calculate a prediction value.

[0452] Model Implementation Device #5 - Device implementing a method for training an antibody sequence generation model

[0453] Referring to FIG. 12, the description is as follows. This specification discloses a computer device (600) for implementing a method for learning an antigen-antibody binding affinity prediction model. The computer device (600) comprises one or more processors (610), one or more memories (620), and one or more computer-readable storage media (630). Herein, the processor (610), the memory (620), and the computer-readable storage media (630) are capable of communicating with each other. The computer-readable storage media (630) stores learning data, a model computational structure, and a computer-implementable method for learning the antibody sequence generation model (hereinafter, a learning algorithm). Herein, the learning algorithm is as described above. When the above computer device (600) performs the antibody sequence generation model learning method, the learning data, model computation structure, and learning algorithm are loaded from the above computer-readable storage medium (630) into memory (620), and the above processor (610) performs the above learning algorithm using the learning data and model computation structure loaded into the above memory (620), and the learned parameters as a result of performing the above learning algorithm are stored in the above computer-readable storage medium (630).

[0454] Model Implementation Device #6 - Device for implementing an in silico antibody candidate screening method

[0455] Referring to FIG. 12, the description is as follows. This specification discloses a computer device (600) for implementing an in silico antibody candidate screening method. The computer device (600) includes one or more processors (610), one or more memories (620), and one or more computer-readable storage media (630). Here, the processors (610), the memories (620), and the computer-readable storage media (630) are capable of communicating with each other. The computer-readable storage media (630) stores an in silico antibody library, an algorithm for computer-executing the in silico antibody screening method disclosed in this specification (hereinafter referred to as the screening algorithm), and a computational structure (architecture) of an antigen-antibody binding affinity prediction model disclosed in this specification, learned parameters (weights and biases), hyperparameters, and other metadata necessary for model implementation. When the computer device (600) performs an in silico antibody screening method, the in silico antibody library, the screening algorithm, and the computational structure (architecture) of the antigen-antibody binding affinity prediction model disclosed in this specification, learned parameters (weights and biases), hyperparameters, and other metadata required for model implementation are loaded from the computer-readable storage medium (630) to memory (620), and the processor (610) utilizes the in silico antibody library and the antigen-antibody binding affinity model loaded in memory (620) to screen antibody sequence candidates according to the screening algorithm, and the antibody sequence candidates screened as a result of performing the screening algorithm are stored in the computer-readable storage medium (630).

[0456]

[0457] [Numbered Examples (Enumerated Embodiments)]

[0458] Hereinafter, various numbered embodiments (or modes) are described sequentially to aid in understanding the invention. However, the present invention should not be interpreted as being limited to the embodiments listed below, but should be understood to include all variations, equivalents, and additional combinations that a person skilled in the art can easily derive from this specification, the accompanying drawings, and the claims.

[0459] For the numbers attached to each process

[0460] Where each example is an example of a method, numbers such as (a), (b), (a.1), (a.2), (1), and (i) were used. These numbers are assigned for convenience based solely on the order of description and do not restrict the order of each process. That is, the process numbered (b) does not necessarily mean that it is performed after the process numbered (a). The order of each process should be interpreted appropriately according to the meaning and context of each process, and may be performed sequentially or in parallel.

[0461]

[0462] Chapter 1. Examples of Methods for Constructing an In-Silico Antibody Library

[0463] Method for generating temporary antibody information

[0464] Example 1, Method for generating temporary antibody information

[0465] A method for generating temporary antibody information based on available antibody information, comprising:

[0466] (a.1) The process of obtaining available antibody information; and

[0467] (a.2) A process of generating provisional antibody information based on available antibody information obtained in (a.1).

[0468] Example 2, Condition of Available Antibody Information #1 - Including Different Antigen Information

[0469] A method for generating temporary antibody information based on available antibody information of Example 1,

[0470] The above available antibody information includes two or more types of antibody information, and

[0471] At least two types of antibodies included in the above available antibody information each bind to different antigens.

[0472] Example 3, Condition of Available Antibody Information #2 - Antigen Information Not Considered

[0473] A method for generating provisional antibody information based on available antibody information of Examples 1 and 2,

[0474] The above available antibody information includes two or more types of antibody information, and

[0475] The type of antigen to which each antibody binds is not considered.

[0476] Example 4, including sequence information of the first antibody and the second antibody

[0477] A method for generating temporary antibody information based on available antibody information selected from any one of Examples 1 to 3, wherein

[0478] The above available antibody information includes sequence information of a first antibody and sequence information of a second antibody, and

[0479] The above provisional antibody information includes the first provisional antibody information.

[0480] Example 5, specification of sequence information of the first antibody and the second antibody

[0481] In a method for generating temporary antibody information based on available antibody information of Example 4,

[0482] The sequence information of the first antibody includes sequence information of the variable light chain (VL) of the first antibody, and

[0483] The sequence information of the second antibody above includes the sequence information of the variable heavy chain (VH) of the second antibody.

[0484] Example 6, Light chain variable region segmentation

[0485] In a method for generating temporary antibody information based on available antibody information of Example 5,

[0486] The light chain variable region of the first antibody is the kappa light chain (Vκ) or lambda light chain (Vλ).

[0487] Example 7, Content of the first novel antibody sequence

[0488] A method for generating temporary antibody information based on available antibody information selected from any one of Examples 1 to 6,

[0489] The sequence of the light chain variable region of the first temporary antibody is the same as the sequence of the light chain variable region of the first antibody, and

[0490] The sequence of the heavy chain variable region of the first temporary antibody is the same as the sequence of the heavy chain variable region of the second antibody.

[0491] Example 8, including multiple antibody sequence information

[0492] A method for generating temporary antibody information based on available antibody information selected from any one of Examples 1 to 7, wherein

[0493] The above available antibody information includes the following:

[0494] (1) Sequence information of the variable region of the light chain of at least one antibody; and

[0495] (2) Sequence information of the variable region of the heavy chain of at least one antibody.

[0496] Example 9, Conditions of Novel Antibody Information

[0497] In a method for generating temporary antibody information based on available antibody information of Example 8,

[0498] Each temporary antibody information generated in the above (a.2) process satisfies the following conditions:

[0499] (i) includes sequence information of the light chain variable region and sequence information of the heavy chain variable region;

[0500] (ii) The sequence of the light chain variable region is the same as the sequence of one light chain variable region included in the available antibody information; and

[0501] (iii) The sequence of the heavy chain variable region is the same as the sequence of one heavy chain variable region included in the available antibody information;

[0502] Here, the antibody of (ii) and the antibody of (iii) are different from each other.

[0503] Example 10, generation of all combinations

[0504] In a method for generating temporary antibody information based on available antibody information of Example 9,

[0505] One or more temporary antibody information generated in the above (a.2) process is,

[0506] Includes all possible provisional antibody information that can be generated from the above available antibody information to satisfy the conditions described in Example 9.

[0507] Method for selecting candidate antibodies of interest

[0508] Example 11, Method for selecting antibody candidates of interest

[0509] A method for screening candidate antibodies of interest, comprising the following:

[0510] (b.1) A process of predicting one or more properties for each antibody based on provisional antibody information; and

[0511] (b.2) A process of selecting antibodies that satisfy predetermined criteria for properties predicted in the above (b.1) process as candidates for antibodies of interest.

[0512] Example 12, properties to be predicted

[0513] In the method for selecting antibody candidates of interest of Example 11,

[0514] One or more properties predicted in (b.1) above are properties for determining whether each antibody is suitable for use as a therapeutic agent.

[0515] Example 13, Example of properties

[0516] A method for selecting any one of the antibody candidates of interest selected from Examples 11 to 12, wherein

[0517] The above properties are selected from the following:

[0518] (i) binding affinity to the antigen of interest; (ii) whether it specifically binds to the antigen of interest; (iii) immunogenicity; (iv) physicochemical properties favorable for therapeutic use; or (v) any combination of (i) to (iv) above.

[0519] Example 14, limiting physicochemical properties favorable for therapeutic use

[0520] In the method for selecting antibody candidates of interest of Example 13,

[0521] The physicochemical properties advantageous for the use of the above therapeutic agent are selected from the following:

[0522] (1) Purity during antibody production; (2) Homogineity of the produced antibody; (3) Stability; (4) Solubility; or (5) Any combination of (1) to (4) above.

[0523] Example 15, prediction of binding strength with antigen of interest

[0524] A method for selecting any one of the antibody candidates of interest selected from Examples 11 to 14, wherein

[0525] The above process (b.1) includes predicting the binding affinity of each antibody included in the above temporary antibody information (if the temporary antibody information includes only one type of temporary antibody information, the corresponding antibody) with the antigen of interest in silico, and

[0526] The above process (b.2) includes determining whether the binding affinity with the antigen of interest satisfies a predetermined criterion.

[0527] Example 16, use of antigen-antibody binding strength prediction model

[0528] In the method for selecting antibody candidates of interest of Example 15,

[0529] A method for predicting the binding affinity of each antibody included in the above provisional antibody information with the antigen of interest in silico is,

[0530] The method includes a process of predicting the binding strength with the antigen of interest by applying the sequence of each antibody included in the above-mentioned temporary antibody information to an antigen-antibody binding strength prediction model of interest, and

[0531] Here, the above-mentioned antigen-antibody binding strength prediction model is a model trained with data in which the antibody sequence and optionally additional information are labeled as the binding strength with the antigen of interest.

[0532] Example 17, Antigen-Antibody Binding Affinity Prediction Model - Transformer Model

[0533] In the method for selecting antibody candidates of interest of Example 16,

[0534] The antigen-antibody binding strength prediction model used to predict binding strength with the antigen of interest is a transformer model, and

[0535] Antigen-antibody binding affinity was derived through the following method:

[0536] 1) Embed the above antigen sequence of interest and perform self-attention using a multi-head attention algorithm;

[0537] 2) Embed a temporary antibody sequence and perform self-attention using a multi-head attention algorithm;

[0538] 3) Using the results of 1) and 2), perform antigen-antibody cross-attention; and

[0539] 4) Input the attention result of 3) into a neural network to derive antigen-antibody binding strength.

[0540] Example 18, CCA added

[0541] In the method for selecting antibody candidates of interest of Example 16,

[0542] The antigen-antibody binding strength prediction model used to predict binding strength with the antigen of interest is a transformer model, and

[0543] Antigen-antibody binding affinity was derived through the following method:

[0544] 1) Embed the antigen sequence of interest and perform self-attention using a multi-head attention algorithm;

[0545] 2) Embed a temporary antibody sequence and perform self-attention using a multi-head attention algorithm;

[0546] 3) Using the results of 1) and 2), perform antigen-antibody cross-attention;

[0547] 4) Chemical Cross Attention (CCA) is performed on the results of 1) and 2), wherein the CCA is performed using a multi-head attention algorithm, but with specific positions masked for each head, and the masking includes the following: masking all positions other than the amino acids at the electron donor and electron acceptor sites; masking all positions other than the amino acids at hydrogen bonding sites; or masking all positions other than the amino acids at van der Waals bonding sites; and

[0548] 5) The attention results of 3) and 4) are input into a neural network to derive antigen-antibody binding strength.

[0549] Example 19, Prediction of specific binding affinity with antigen of interest

[0550] A method for selecting any one of the antibody candidates of interest selected from Examples 11 to 18, wherein

[0551] The above process (b.1) includes predicting in silico the binding affinity of each antibody included in the above provisional antibody information (the antibody if the provisional antibody information includes only one type of provisional antibody information) with the antigen of interest and the binding affinity with one or more non-target antigens, and

[0552] The above process (b.2) includes determining whether the binding affinity with the antigen of interest and the binding affinity with one or more non-target antigens satisfy predetermined criteria.

[0553] Example 20, method for predicting specific binding force

[0554] In the method for selecting antibody candidates of interest of Example 19,

[0555] A method for predicting in silico the binding affinity of each antibody included in the above-mentioned provisional antibody information with the antigen of interest and with one or more non-target antigens is,

[0556] The method includes a process of applying the sequence of each antibody included in the above-mentioned temporary antibody information to an antigen-antibody binding affinity prediction model to predict the binding affinity with the antigen of interest or each non-target antigen, and

[0557] Here, the antigen-antibody binding strength prediction model is a model trained with data in which the antibody sequence and optionally additional information are labeled as the binding strength with the corresponding antigen.

[0558] Example 21, limited to an antigen-antibody model

[0559] In the method for selecting antibody candidates of interest of Example 20,

[0560] The above antigen-antibody binding strength prediction model is the antigen-antibody binding strength prediction model described in Example 17 or Example 18.

[0561] Example 22, Prediction of physicochemical properties favorable for therapeutic use

[0562] A method for selecting any one of the antibody candidates of interest selected from Examples 11 to 21, wherein

[0563] The above (b.1) process includes calculating the light chain-heavy chain binding potential of each antibody included in the above temporary antibody information (if the temporary antibody information includes only one type of temporary antibody information, the corresponding antibody), and

[0564] Here, the light chain-heavy chain binding potential is obtained by inputting the sequences of the light chain variable region and the heavy chain variable region of each antibody into the CNN-P model disclosed in Chinery, L., Jeliazkov, JR, & Deane, CM (2024). Humatch - fast, gene-specific joint humanization of antibody heavy and light chains. mAbs, 16(1). https: / doi.org / 10.1080 / 19420862.2024.2434121, and

[0565] The above process (b.2) includes determining whether the possibility of light chain-heavy chain bonding satisfies a predetermined criterion.

[0566] Example 23, Immunogenicity Prediction

[0567] A method for selecting any one of the antibody candidates of interest selected from Examples 11 to 22,

[0568] The above (b.1) process includes predicting the similarity between the sequence of each antibody included in the above temporary antibody information (the antibody in cases where the temporary antibody information includes only one type of temporary antibody information) and a human antibody, and

[0569] The above process (b.2) includes determining whether the similarity satisfies a predetermined criterion.

[0570] Example 24, similarity determination method

[0571] In the method for selecting antibody candidates of interest according to Examples 22 to 23,

[0572] The prediction of similarity between the sequences of each of the above antibodies and human antibodies was performed by the following method:

[0573] 1) The light chain variable region sequences and heavy chain variable region sequences of the antibody were entered into the ANARCI program (James Dunbar, Charlotte M. Deane, ANARCI: antigen receptor numbering and receptor classification, Bioinformatics, Volume 32, Issue 2, January 2016, Pages 298-300, https: / doi.org / 10.1093 / bioinformatics / btv552);

[0574] 2) The closest species was identified by comparing the similarity between each sequence and species-specific antibody sequences in the IMGT GENE database (Vιronique Giudicelli, Denys Chaume, Marie-Paule Lefranc, IMGT / GENE-DB: a comprehensive database for human and mouse immunoglobulin and T cell receptor genes, Nucleic Acids Research, Volume 33, Issue suppl_1, 1 January 2005, Pages D256-D261, https: / / doi.org / 10.1093 / nar / gki010) using the HMMER program (A NEW GENERATION OF HOMOLOGY SEARCH TOOLS BASED ON PROBABILISTIC INFERENCE SEAN R. EDDY (USA) Genome Informatics 2009. October 2009, 205-211); and

[0575] 3) Predict human germline similarity by outputting the closest antibody species found in ANARCI.

[0576] Method to verify whether it is an antibody of interest

[0577] Example 25, method for verifying whether it is an antibody of interest

[0578] A method for verifying whether a candidate antibody of interest is the antibody of interest, comprising:

[0579] (c.1) Synthesize candidate antibody of interest; and

[0580] (c.2) Verify one or more properties for each candidate antibody of interest synthesized in (c.1) above.

[0581] Example 26, Example of a property to be verified

[0582] In a method for verifying whether the antibody candidate of Example 25 is the antibody of interest,

[0583] The property to be verified in the above (c.2) process is selected from the following:

[0584] (i) binding affinity to the antigen of interest; (ii) whether it specifically binds to the antigen of interest; (iii) immunogenicity; (iv) physicochemical properties favorable for therapeutic use; or (v) any combination of (i) to (iv) above.

[0585] Example 27, limiting physicochemical properties favorable for therapeutic use

[0586] In a method for verifying whether the antibody candidate of Example 26 is the antibody of interest,

[0587] The physicochemical properties advantageous for the use of the above therapeutic agent are selected from the following:

[0588] (1) Purity during antibody production; (2) Homogineity of the produced antibody; (3) Stability; (4) Solubility; or (5) Any combination of (1) to (4) above.

[0589] Example 28, Method for verifying binding affinity with antigen of interest

[0590] In a method for verifying whether the antibody candidate of Example 26 is the antibody of interest,

[0591] The binding affinity with the antigen of interest is verified by an immunoassay using the antigen of interest and the candidate antibody of interest.

[0592] Example 29, form of the synthesized antibody

[0593] A method for verifying whether any one of the selected antibody candidate of Examples 25 to 28 is the antibody of interest,

[0594] The antibody candidate of interest synthesized in the above (c.1) process is a full-length antibody or an antibody fragment.

[0595] Example 30, Example of a full-length antibody isomorphism

[0596] In a method for verifying whether the antibody candidate of Example 29 is the antibody of interest,

[0597] The antibody candidate of interest synthesized in the above (c.1) process is a full-length antibody, and

[0598] The above full-length antibody is an isotype selected from the following:

[0599] IgA; IgD; IgE; IgG; IgM; or IgY.

[0600] Example 31, Example of antibody fragment

[0601] In a method for verifying whether the antibody candidate of Example 29 is the antibody of interest,

[0602] The antibody candidate of interest synthesized in the above (c.1) process is an antibody fragment, and

[0603] The above antibody fragment is a structure selected from the following:

[0604] F(ab); F(ab'); F(ab')2; Monospecific F(ab')2; Bispecific F(ab')2; single-chain variable fragment (scFv); scFv-Fc; or single-domain antibody (sdAb).

[0605] In Silico Antibody Library #1 - Basic Library

[0606] Example 32, Basic Library

[0607] in silico antibody library,

[0608] The above in silico antibody library contains information on one or more temporary antibodies, and

[0609] The above temporary antibody information is generated by a method of generating temporary antibody information selected from any one of Examples 1 to 10.

[0610] In silico antibody library screening method

[0611] Example 33, in silico antibody library screening method #1

[0612] A method for constructing an antibody library in silico, comprising:

[0613] (a) The process of generating temporary antibody information,

[0614] Here, the process of generating the above-mentioned temporary antibody information is performed by a method of generating temporary antibody information selected from any one of Examples 1 to 10; and

[0615] (b) A process of selecting candidate antibodies of interest based on provisional antibody information generated in the above (a) process,

[0616] Here, the above process is performed by a method of selecting any one of the antibody candidates of interest selected from Examples 11 to 24, and

[0617] In the above method for selecting candidate antibodies of interest, the provisional antibody information of process (b.1) is the provisional antibody information generated in process (a), and

[0618] The above-mentioned in silico library is constructed using the selected antibody candidates of interest as a result of performing the above-mentioned method for selecting antibody candidates of interest.

[0619] Example 34, Selection of immunogenicity and physicochemical properties favorable for therapeutic use

[0620] In the method for constructing an in silico antibody of Example 33,

[0621] The method for screening the candidate antibody of interest performed in the above (b) process comprises the following:

[0622] (b.1) A process of predicting one or more properties for each antibody based on provisional antibody information,

[0623] Here, the above temporary antibody information is the temporary antibody information generated in the above (a) process, and

[0624] The above properties are selected from the following:

[0625] (i) immunogenicity; (ii) physicochemical properties favorable for therapeutic use; or (iii) any combination of (i) to (ii) above; and

[0626] (b.2) Antibodies whose properties predicted in the above (b.1) process satisfy predetermined criteria are selected as candidates for antibodies of interest.

[0627] Example 35, limiting physicochemical properties favorable for therapeutic use

[0628] In the method for constructing an in silico antibody of Example 34,

[0629] The physicochemical properties of the above (ii) favorable for the use of the therapeutic agent in the above (b.1) process of the method for screening the antibody candidate of interest performed in the above (b) process include the following:

[0630] (1) Purity during antibody production; (2) Homogineity of the produced antibody; (3) Stability; (4) Solubility; or (5) Any combination of (1) to (4) above.

[0631] Example 36, in silico antibody library screening method #2

[0632] A method for constructing an antibody library in silico, comprising:

[0633] (b) A process of selecting candidate antibodies of interest based on provisional antibody information,

[0634] Here, the above process is performed by a method of selecting any one of the antibody candidates of interest selected from Examples 11 to 24, and

[0635] In the method for selecting the antibody candidate of interest above, the provisional antibody information of process (b.1) is provisional antibody information generated by performing a method for generating provisional antibody information selected from any one of Examples 1 to 10, and

[0636] The above-mentioned in silico library is constructed using the selected antibody candidates of interest as a result of performing the above-mentioned method for selecting antibody candidates of interest.

[0637] Example 37, Selection of immunogenicity and physicochemical properties favorable for therapeutic use

[0638] In the method for constructing an in silico antibody of Example 36,

[0639] The method for screening the candidate antibody of interest performed in the above (b) process comprises the following:

[0640] (b.1) A process of predicting one or more properties for each antibody based on provisional antibody information,

[0641] Here, the above temporary antibody information is temporary antibody information generated by performing a method for generating temporary antibody information selected from any one of Examples 1 to 10, and

[0642] The above properties are selected from the following:

[0643] (i) immunogenicity; (ii) physicochemical properties favorable for therapeutic use; or (iii) any combination of (i) to (ii) above; and

[0644] (b.2) Antibodies whose properties predicted in the above (b.1) process satisfy predetermined criteria are selected as candidates for antibodies of interest.

[0645] Example 38, limitation of physicochemical properties favorable for therapeutic use

[0646] In the method for constructing an in silico antibody of Example 37,

[0647] The physicochemical properties of the above (ii) favorable for the use of the therapeutic agent in the above (b.1) process of the method for screening the antibody candidate of interest performed in the above (b) process include the following:

[0648] (1) Purity during antibody production; (2) Homogineity of the produced antibody; (3) Stability; (4) Solubility; or (5) Any combination of (1) to (4) above.

[0649] in silico antibody library #2 - Selected library

[0650] Example 39, selected library

[0651] An in silico antibody library composed of any one of the methods of Examples 33 to 38.

[0652] Method for Selecting Antibodies of Interest #1

[0653] Example 40, Method for Selecting Antibody of Interest #1

[0654] A method for screening antibodies of interest, comprising:

[0655] (a) The process of generating temporary antibody information,

[0656] Here, the process of generating the above temporary antibody information is performed by a method of generating temporary antibody information selected from any one of Examples 1 to 10;

[0657] (b) A process of selecting candidate antibodies of interest based on provisional antibody information generated in the above (a) process,

[0658] Here, the above process is performed by a method of selecting any one of the antibody candidates of interest selected from Examples 11 to 24, and

[0659] In the above method for screening candidate antibodies of interest, the provisional antibody information of process (b.1) is the provisional antibody information generated in process (a); and

[0660] (c) A process for verifying whether the antibody candidate selected in the above (b) process is the antibody of interest,

[0661] Here, the above process is performed as a method to verify whether any one of the selected antibody candidate among Examples 25 to 31 is the antibody of interest, and

[0662] In a method for verifying whether the above-mentioned candidate antibody of interest is the antibody of interest, the candidate antibody of interest is the candidate antibody of interest selected in the above-mentioned (b) process.

[0663] Method for Selecting Antibodies of Interest #2

[0664] Example 41, Method for screening antibodies of interest #2

[0665] A method for screening antibodies of interest, comprising:

[0666] (b) Process of selecting candidate antibodies of interest,

[0667] Here, the above process is performed by a method of selecting any one of the antibody candidates of interest selected from Examples 11 to 24, and

[0668] In the method for screening the antibody candidate of interest above, the provisional antibody information of process (b.1) is generated by a method for generating provisional antibody information selected from any one of Examples 1 to 10; and

[0669] (c) A process for verifying whether the antibody candidate selected in the above (b) process is the antibody of interest,

[0670] Here, the above process is performed as a method to verify whether any one of the selected antibody candidate among Examples 25 to 31 is the antibody of interest, and

[0671] In a method for verifying whether the above-mentioned candidate antibody of interest is the antibody of interest, the candidate antibody of interest is the candidate antibody of interest selected in the above-mentioned (b) process.

[0672] How to Select Antibodies of Interest #3

[0673] Example 42, Method for Selecting Antibody of Interest #3

[0674] A method for screening antibodies of interest, comprising:

[0675] (c) A process of verifying whether a candidate antibody of interest is the antibody of interest,

[0676] Here, the above process is performed as a method to verify whether any one of the selected antibody candidate among Examples 25 to 31 is the antibody of interest, and

[0677] In a method for verifying whether the above-mentioned candidate antibody of interest is an antibody of interest, the candidate antibody of interest in step (c.1) is a candidate antibody of interest selected by a method of selecting any one of the candidate antibody of interest selected from Examples 11 to 24, and

[0678] In the above method for selecting candidate antibodies of interest, the provisional antibody information of process (b.1) is generated by a method for generating provisional antibody information selected from any one of Examples 1 to 10.

[0679]

[0680] Chapter 2. Examples of Antigen-Antibody Binding Affinity Prediction Models

[0681] Layers that can be used in the model

[0682] Example 43, Chemical Attention Guide

[0683] Chemical attention guide derived based on the chemical properties of antigen and antibody sequences.

[0684] Example 44, limiting intermolecular interactions

[0685] In the chemical attention guide of Example 43,

[0686] The chemical properties of the above antigen sequence and antibody sequence are the intermolecular interactions selected from the following:

[0687] Hydrogen bond between each amino acid of the antigen sequence and each amino acid of the antibody sequence;

[0688] Van der Waals interaction between each amino acid of the antigen sequence and each amino acid of the antibody sequence;

[0689] Hydrophobic interaction between each amino acid of the antigen sequence and each amino acid of the antibody sequence;

[0690] Electrostatic interaction between each amino acid of the antigen sequence and each amino acid of the antibody sequence; or

[0691] Any combination of the above interactions.

[0692] Example 45, chemical attention guide element limitation

[0693] In the chemical attention guide of Example 44,

[0694] The length of the above antigen sequence is L ag , the length of the above antibody sequence is L ab When saying,

[0695] The above chemical attention guide is as follows Represented as a matrix or its equivalent:

[0696]

[0697] Each element of the matrix above represents whether there is an intermolecular interaction between the i-th amino acid of the antigen and the j-th amino acid of the antibody, or the degree of intermolecular interaction.

[0698] Example 46, presence of intermolecular interactions

[0699] In the chemical attention guide of Example 45,

[0700] Each element included in the above chemical attention guide If means whether there is an intermolecular interaction between the i-th amino acid of the antigen and the j-th amino acid of the antibody,

[0701] If the element is 0, it means there is no interaction, and if the element is 1, it means there is interaction.

[0702] Example 47, degree of intermolecular interaction

[0703] In the chemical attention guide of Example 45,

[0704] Each element included in the above chemical attention guide If represents the degree of intermolecular interaction between the i-th amino acid of the antigen and the j-th amino acid of the antibody,

[0705] The degree of the above interaction is determined based on the type of the i-th amino acid of the antigen and the type of the j-th amino acid of the antibody.

[0706] Example 48, embedding layer

[0707] Embedding layer,

[0708] Here, the layer input of the above embedding layer includes an amino acid sequence, and

[0709] The layer output of the above embedding layer contains an embedding representation matrix or its equivalent.

[0710] Example 49, Dimensional Limitation

[0711] In the embedding layer of Example 48,

[0712] When the length of the amino acid sequence of the upper layer input is L,

[0713] The embedding representation matrix of the upper layer output contains (L + p) vectors of E dimension, and

[0714] The vectors included in the above embedding representation matrix include L E-dimensional vectors and p padding vectors corresponding to each amino acid of the above amino acid sequence, and

[0715] L and E above are integers greater than or equal to 1, and p above is an integer greater than or equal to 0.

[0716] Example 50, order preservation

[0717] In the embedding layer of Example 49,

[0718] Among the vectors included in the above embedding representation matrix, the order of the vectors corresponding to each amino acid of the above amino acid sequence preserves the order of the amino acid sequence.

[0719] Example 51, positional encoding

[0720] In the embedding layer of Example 50,

[0721] Among the vectors included in the above embedding representation matrix, the vector corresponding to each amino acid of the above amino acid sequence contains the following information:

[0722] The type of amino acid embedded as an E-dimensional vector; and the position of the corresponding amino acid within the above amino acid sequence.

[0723] Example 52, learnable weights

[0724] In any one of the selected embedding layers among Examples 48 to 51,

[0725] The above embedding layer includes an embedding matrix or its equivalent, and

[0726] The above embedding matrix contains an E-dimensional vector corresponding to each of all possible types of amino acids, and

[0727] When the above embedding layer generates the above embedding representation matrix, it references the above embedding matrix, and

[0728] The E-dimensional vectors corresponding to each of the above amino acid types are initialized with random values ​​and adjusted during the training process of the model containing the above embedding layer.

[0729] Example 53, self-attention layer

[0730] Self-attention layer,

[0731] Here, the layer inputs of the above self-attention layer include the following:

[0732] As a first input, a characteristic expression derived from an antigen sequence or a characteristic expression derived from an antibody sequence,

[0733] When H1 is an integer greater than or equal to 1, the above self-attention layer includes H1 attention heads, and

[0734] The operations performed when the input from the upper layer is forward propagated to the upper self-attention layer include the following:

[0735] For each case where n is an integer between 1 and H1 inclusive, the operations 1) through 6) below:

[0736] 1) An operation to derive the nth query from the first input above,

[0737] The first query above contains one or more vectors;

[0738] 2) An operation to derive the nth key from the first input above,

[0739] The first key above contains one or more vectors;

[0740] 3) An operation to derive the nth value from the first input above,

[0741] The above n-th value contains one or more vectors, and the number of vectors included in the above n-th key is the same as the number of vectors included in the above n-th value;

[0742] 4) An operation to derive the similarity between each vector included in the above n-th query and each vector included in the above n-th key,

[0743] As a result of the above operation, the similarity for all possible pairs between the vector included in the above nth query and the vector included in the above nth key is derived;

[0744] 5) An operation to derive the nth attention score corresponding to each similarity of 4) above; and

[0745] 6) An operation to derive the nth attention value using the nth attention score of 5) above and the nth value above;

[0746] Here, the layer output of the above self-attention layer is calculated based on the above first to nth attention values.

[0747] Example 54, multi-head attention, each head dimension

[0748] In the self-attention layer of Example 53,

[0749] The dimension of each vector included in the first input is E t and,

[0750] For each case where i is an integer greater than or equal to H1,

[0751] The dimension of each vector included in the i-th attention value is E i If we say,

[0752] lim.

[0753] Example 55, each head dimension is the same

[0754] In the self-attention layer of Example 54,

[0755] For each case where i is an integer greater than or equal to H1,

[0756] The dimension of each vector included in the i-th attention value is E t / H1

[0757] Example 56, Multi-head Attention, Output Value

[0758] In the self-attention layer of Example 54 or Example 55,

[0759] The layer output of the above self-attention layer is selected from the following when H1 is 1:

[0760] A value obtained by linearly or non-linearly transforming the first attention value into a learnable weight matrix;

[0761] A first attention value linearly or non-linearly transformed into a learnable weight matrix, and a value connected to the first input and residuals above; or

[0762] Linearly or non-linearly transforming the first attention value into a learnable weight matrix, connecting the residuals with the first input above, and normalizing the value;

[0763] The layer output of the above self-attention layer is selected from the following if H1 is an integer greater than or equal to 2:

[0764] By connecting the 1st attention value, ..., and the H1th attention value, each vector is E t Values ​​that are linearly or non-linearly transformed into a learnable weight matrix, making the dimensionality such that it becomes a value;

[0765] By connecting the 1st attention value, 쪋, and the H1th attention value, each vector is E t To make it dimensional, transform it linearly or non-linearly into a learnable weight matrix, and the value connected to the first input and the residual above; or

[0766] By connecting the 1st attention value, ..., and the H1th attention value, each vector is E t Make the dimension, linearly or non-linearly transform into a learnable weight matrix, connect the residuals with the first input above, and normalize the value.

[0767] Example 57, Normalization Type

[0768] In the self-attention layer of Example 56,

[0769] The above normalization is Layer Normalization, Batch Normalization, or Group Normalization.

[0770] Example 58, Input Specification

[0771] In any one of the selected self-attention layers from Examples 53 to 57,

[0772] The characteristic expression derived from the above antigen sequence or the characteristic expression derived from the antibody sequence is selected from the following:

[0773] i) an embedding representation matrix or its equivalent output by inputting an antigen sequence into any one of the selected embedding layers of Examples 48 to 52;

[0774] ii) A matrix or equivalent generated by transforming the embedding representation matrix or its equivalent of i) above;

[0775] iii) an embedding representation matrix or its equivalent output by inputting an antibody sequence into any one of the selected embedding layers of Examples 48 to 52;

[0776] iv) A matrix or equivalent generated by transforming the embedding representation matrix of iii) above or its equivalent.

[0777] Example 59, Query, Key, Value

[0778] In any one of Examples 53 to 58, a self-attention layer selected

[0779] For each case where n is an integer between 1 and H1 inclusive, the operations in 1) through 3) above are as follows:

[0780] 1) The above first input An operation that derives the nth query by performing a linear or non-linear transformation on a matrix;

[0781] 2) The above first input An operation to derive the nth key by performing a linear or non-linear transformation on a matrix; and

[0782] 3) The above first input An operation that derives the nth value by performing a linear or non-linear transformation on a matrix,

[0783] The above n-th value contains one or more vectors, and the number of vectors included in the above n-th key is the same as the number of vectors included in the above n-th value;

[0784] Here, above , , and is a learnable weight matrix.

[0785] Example 60, Similarity Calculation

[0786] In any one of the selected self-attention layers from Examples 53 to 59,

[0787] For each case where n is an integer between 1 and H1 inclusive, the operation in 4) above is as follows:

[0788] 4) An operation to derive the similarity between each vector included in the above n-th query and each vector included in the above n-th key,

[0789] As a result of the above operation, the similarity for all possible pairs between the vector included in the above nth query and the vector included in the above nth key is derived, and

[0790] The above similarity derivation method is selected from the following: Dot Product; Weighted Dot Product, where the weight matrix of the dot product values ​​is learnable; Scaled Dot Product, where the scale value is the square root of the dimension of the key vector; Cosine Similarity; Euclidean Distance; or Manhattan Distance.

[0791] Example 61, probability distribution transformation function

[0792] In any one of the selected self-attention layers among Examples 53 to 60,

[0793] For each case where n is an integer between 1 and H1 inclusive, the operation in 5) above is as follows:

[0794] 5) An operation to derive the nth attention score corresponding to each similarity of 4) above,

[0795] Here, the operation for deriving the nth attention score is derived by applying one of the following functions to each of the above similarity values: Softmax; a variation of the above Softmax function; Entmax; Sigmoid; or a Gaussian Kernel function.

[0796] Example 62, padding mask added

[0797] In any one of the selected self-attention layers from Examples 53 to 61,

[0798] The first input above includes a padding vector, and

[0799] For each case where n is an integer between 1 and H1 inclusive, the operation in 6) above is as follows:

[0800] 6) An operation to mask the attention score associated with the padding vector among the nth attention scores of 5) above, and to derive the nth attention value using the masked nth attention score and the nth value above.

[0801] Example 63, Chemical Cross Attention Layer #1 - Using antigen sequence as query

[0802] Chemical Cross Attention Layer,

[0803] Here, the layer input of the chemical cross-attention layer above includes the following:

[0804] As a first input, a characteristic expression derived from an antigen sequence; and

[0805] As a second input, it is a characteristic expression derived from the antibody sequence;

[0806] When H2 is an integer greater than or equal to 1, the above chemical cross-attention layer includes H2 attention heads, and

[0807] The operations performed when the input from the upper layer is forward propagated to the upper chemical cross-attention layer include the following:

[0808] For each case where m is an integer between 1 and H2 inclusive, the operations 1) through 7) below:

[0809] 1) An operation to derive the m-th query from the first input above,

[0810] The above query contains one or more vectors;

[0811] 2) An operation to derive the m-th key from the second input above,

[0812] The above key m contains one or more vectors;

[0813] 3) An operation to derive the m-th value from the second input above,

[0814] The above m-th value contains one or more vectors, and the number of vectors included in the above m-th key is the same as the number of vectors included in the above m-th value;

[0815] 4) An operation to derive the similarity between each vector included in the above m-th query and each vector included in the above m-th key,

[0816] As a result of the above operation, the similarity for all possible pairs between the vector included in the above m-th query and the vector included in the above m-th key is derived;

[0817] 5) An operation to derive the m-th attention score corresponding to each similarity of 4) above;

[0818] 6) An operation to derive the m-th chemical attention score by adjusting the weights of the m-th attention score above in 5),

[0819] Here, the above operation is performed using any one selected from Examples 43 to 47 as the m-th chemical attention guide; and

[0820] 7) An operation to derive the m-th attention value using the m-th attention score of 6) above and the m-th value above;

[0821] Here, the layer output of the chemical cross-attention layer above is calculated based on the first to mth attention values ​​above.

[0822] Example 64, Chemical Cross-Attention Layer #2 - Using antibody sequence as query

[0823] Chemical Cross Attention Layer,

[0824] Here, the layer input of the chemical cross-attention layer above includes the following:

[0825] As a first input, a characteristic expression derived from an antibody sequence; and

[0826] As a second input, it is a characteristic expression derived from the antigen sequence;

[0827] When H2 is an integer greater than or equal to 1, the above chemical cross-attention layer includes H2 attention heads, and

[0828] The operations performed when the input from the upper layer is forward propagated to the upper chemical cross-attention layer include the following:

[0829] For each case where m is an integer between 1 and H2 inclusive, the operations 1) through 7) below:

[0830] 1) An operation to derive the m-th query from the first input above,

[0831] The above query contains one or more vectors;

[0832] 2) An operation to derive the m-th key from the second input above,

[0833] The above key m contains one or more vectors;

[0834] 3) An operation to derive the m-th value from the second input above,

[0835] The above m-th value contains one or more vectors, and the number of vectors included in the above m-th key is the same as the number of vectors included in the above m-th value;

[0836] 4) An operation to derive the similarity between each vector included in the above m-th query and each vector included in the above m-th key,

[0837] As a result of the above operation, the similarity for all possible pairs between the vector included in the above m-th query and the vector included in the above m-th key is derived;

[0838] 5) An operation to derive the m-th attention score corresponding to each similarity of 4) above;

[0839] 6) An operation to derive the m-th chemical attention score by adjusting the weights of the m-th attention score above in 5),

[0840] Here, the above operation is performed using any one selected from Examples 43 to 47 as the m-th chemical attention guide; and

[0841] 7) An operation to derive the m-th attention value using the m-th attention score of 6) above and the m-th value above;

[0842] Here, the layer output of the chemical cross-attention layer above is calculated based on the first to mth attention values ​​above.

[0843] Example 65, multi-head attention, each head dimension

[0844] In any one of Examples 63 to 64, a chemical cross-attention layer,

[0845] The dimension of each vector included in the first input is E t and,

[0846] For each case where i is an integer greater than or equal to H2,

[0847] The dimension of each vector included in the i-th attention value is E i If we say,

[0848] lim.

[0849] Example 66, each head dimension is the same

[0850] In the chemical cross-attention layer of Example 65,

[0851] For each case where i is an integer greater than or equal to H2,

[0852] The dimension of each vector included in the i-th attention value is E t / H2

[0853] Example 67, Multi-head Attention, Output Value

[0854] In the chemical cross-attention layer of Example 65 or Example 66,

[0855] The layer output of the above chemical cross-attention layer is selected from the following when H2 is 1:

[0856] A value obtained by linearly or non-linearly transforming the first attention value into a learnable weight matrix;

[0857] A first attention value linearly or non-linearly transformed into a learnable weight matrix, and a value connected to the first input and residuals above; or

[0858] Linearly or non-linearly transforming the first attention value into a learnable weight matrix, connecting the residuals with the first input above, and normalizing the value;

[0859] The layer output of the above chemical cross-attention layer is selected from the following if H2 is an integer greater than or equal to 2:

[0860] By connecting the 1st attention value, ..., the H2nd attention value, each vector is E t Values ​​that are linearly or non-linearly transformed into a learnable weight matrix, making the dimensionality such that it becomes a value;

[0861] By connecting the 1st attention value, ..., the H2nd attention value, each vector is E t To make it dimensional, transform it linearly or non-linearly into a learnable weight matrix, and the value connected to the first input and the residual above; or

[0862] By connecting the 1st attention value, ..., the H2nd attention value, each vector is E t Make the dimension, linearly or non-linearly transform into a learnable weight matrix, connect the residuals with the first input above, and normalize the value.

[0863] Example 68, Input Specification

[0864] In any one of the selected chemical cross-attention layers of Examples 63 to 67,

[0865] The characteristic expression derived from the above antigen sequence is selected from the following:

[0866] i) an embedding representation matrix or its equivalent output by inputting an antigen sequence into any one of the selected embedding layers of Examples 48 to 52;

[0867] ii) A matrix or equivalent generated by transforming the embedding representation matrix or its equivalent of i) above;

[0868] iii) a feature representation or its equivalent output by passing the matrix of i) or ii) above through one or more layers; or

[0869] iv) a layer output or an equivalent thereof obtained by passing the matrix of i), ii), or iii) above through any one of the self-attention layers selected from Examples 53 to 62;

[0870] The characteristic expression derived from the above antibody sequence is selected from the following:

[0871] v) an embedding representation matrix or its equivalent output by inputting an antibody sequence into any one of the selected embedding layers of Examples 48 to 52;

[0872] vi) A matrix or equivalent generated by transforming the embedding representation matrix of v) above or its equivalent;

[0873] vii) A feature representation or its equivalent output by passing the matrix of v) or vi) above through one or more layers; or

[0874] viii) a layer output or an equivalent thereof obtained by passing the matrix of v), vi), or vii) above through any one of the self-attention layers selected from Examples 53 to 62.

[0875] Example 69, Query, Key, Value

[0876] In any one of Examples 63 to 68, a chemical cross-attention layer selected

[0877] For each case where m is an integer between 1 and H2 inclusive, the operations in 1) through 3) above are as follows:

[0878] 1) The above first input An operation to derive the m-th query by performing a linear or non-linear transformation on a matrix;

[0879] 2) The above second input An operation to derive the m-th key by performing a linear or non-linear transformation on a matrix; and

[0880] 3) The above second input An operation that derives the m-th value by performing a linear or non-linear transformation on a matrix,

[0881] The above m-th value contains one or more vectors, and the number of vectors included in the above m-th key is the same as the number of vectors included in the above m-th value;

[0882] Here, above , , and is a learnable weight matrix.

[0883] Example 70, Similarity Calculation

[0884] In any one of Examples 63 to 69, a chemical cross-attention layer selected

[0885] For each case where m is an integer between 1 and H2 inclusive, the operation in 4) above is as follows:

[0886] 4) An operation to derive the similarity between each vector included in the above m-th query and each vector included in the above m-th key,

[0887] As a result of the above operation, the similarity for all possible pairs between the vector included in the above m-th query and the vector included in the above m-th key is derived, and

[0888] The above similarity derivation method is selected from the following: Dot Product; Weighted Dot Product, where the weight matrix of the dot product values ​​is learnable; Scaled Dot Product, where the scale value is the square root of the dimension of the key vector; Cosine Similarity; Euclidean Distance; or Manhattan Distance.

[0889] Example 71, probability distribution transformation function

[0890] In any one of Examples 63 to 70, a chemical cross-attention layer selected

[0891] For each case where m is an integer between 1 and H2 inclusive, the operation in 5) above is as follows:

[0892] 5) An operation to derive the m-th attention score corresponding to each similarity of 4) above,

[0893] Here, the operation for deriving the nth attention score is derived by applying one of the following functions to each of the above similarity values: Softmax; a variation of the above Softmax function; Entmax; Sigmoid; or a Gaussian Kernel function.

[0894] Example 72, Addition of padding mask to chemical attention guide

[0895] In any one of the selected chemical cross-attention layers from Examples 63 to 71,

[0896] The first input above, the second input above, or both include a padding vector, and

[0897] For each case where m is an integer between 1 and H2 inclusive, the operation in 6) above is as follows:

[0898] 6) An operation to derive the m-th chemical attention score by adjusting the weights of the m-th attention score above in 5),

[0899] Here, the above operation is performed using any one of the selected chemical attention guides from Examples 43 to 47, and

[0900] Additionally, among the above m attention scores, the attention scores associated with the padding vector are masked.

[0901] Example 73, Forward Neural Network

[0902] Forward neural network layer,

[0903] Here, the layer input of the above forward neural network layer includes the first input, and

[0904] Here, the first input above is a characteristic expression derived from an antigen sequence or a characteristic expression derived from an antibody sequence, and

[0905] When T is an integer greater than or equal to 1, the above forward neural network layer includes T hidden layers, and

[0906] The layer output of the above forward neural network layer has the same dimension as the first input.

[0907] Example 74, hidden layer operation specific

[0908] In the forward neural network layer of Example 73,

[0909] The operations performed when the input to the upper layer is forward propagated to the upper forward neural network layer include the following:

[0910] For each case where t is an integer greater than or equal to 1 and less than or equal to T,

[0911] 1) An operation that receives the operation value of the t-1th hidden layer as input, performs a linear transformation, and optionally adds a bias;

[0912] 2) An operation that performs a non-linear transformation on the result of 1) above;

[0913] 3) an operation to linearly transform the result of 2) above and optionally add a bias; and

[0914] 4) An operation to derive the operation value of the t-th hidden layer using the result of 3) above;

[0915] Here, the operation value of the above t-th hidden layer is transferred to the t+1-th hidden layer, and

[0916] Here, the operation value of the 0th hidden layer corresponds to the above 1st input, and

[0917] The T+1 hidden layer refers to the output layer, and passing to the T+1 hidden layer means setting the corresponding value as the layer output.

[0918] Example 75, Dimensional expansion and reduction

[0919] In the forward neural network layer of Example 74,

[0920] The operation 1) above, which takes the operation value of the t-1 hidden layer as input and performs a linear transformation, is an operation that expands the dimension of the operation value of the t-1 hidden layer, and

[0921] The operation that linearly transforms the result of 2) above is a dimensionality reduction operation such that the dimension of the result of 2) above becomes the same as the dimension of the operation value of the t-1 hidden layer above.

[0922] Example 76, Dimensional expansion and Dimensional reduction specific

[0923] In the forward neural network layer of Example 75,

[0924] Let d be the dimension of the operation value of the t-1 hidden layer above, and the dimension d is obtained by the linear transformation operation of 1) above. ff When it is said to have been expanded as,

[0925] And, above, f is 2, 3, 4, 5 or an integer greater than or equal to 6.

[0926] Example 77, residual linkage operation

[0927] In any one of the forward neural network layers selected from Examples 73 to 76,

[0928] The operation that derives the operation value of the t-th hidden layer using the result of 4) above 3) above is one of the following selected operations:

[0929] The result of 3) above is used as is as the operation value of the t-th hidden layer;

[0930] The result of 3) above is used as the operation value of the t-1 hidden layer by concatenating the residuals with the operation value of the t-1 hidden layer;

[0931] The result of 3) above is combined with the residual of the operation value of the t-1 hidden layer above, normalized, and used as the operation value of the t hidden layer.

[0932] Example 78, specific trainable weights

[0933] In any one of the forward neural network layers selected from Examples 73 to 77,

[0934] For each case where t is an integer between 1 and T, the operations in 1) through 3) above are as follows:

[0935] 1) The operation value of the t-1 hidden layer Linearly transform into a matrix, and optionally, Operation to add bias;

[0936] 2) An operation to perform a non-linear transformation by applying a non-linear activation function to the result of 1) above;

[0937] 3) The result of 2) above Linearly transform into a matrix, and optionally Operation to add bias;

[0938] Here, above , , , and is a learnable weight matrix.

[0939] Example 79, non-linear activation function

[0940] In Example 78, the above non-linear activation function is ReLU (Rectified Linear Unit), GELU (Gaussian Error Linear Unit), or SiLU (Sigmoid Linear Unit).

[0941] Example 80, layer for deriving antigen-antibody binding affinity

[0942] Antigen-antibody binding affinity derivation layer,

[0943] Here, the layer input of the antigen-antibody binding affinity derivation layer above includes the following:

[0944] As a first input, a characteristic expression derived from an antigen sequence; and

[0945] As a second input, it is a characteristic expression derived from the antibody sequence;

[0946] The antigen-antibody binding affinity derivation layer above outputs a predicted antigen-antibody binding affinity value based on the first input and the second input above.

[0947] Example 81, combination, linear transformation

[0948] In the antigen-antibody binding affinity derivation layer of Example 80,

[0949] The operations performed when the input from the upper layer is forward propagated to the upper antigen-antibody binding affinity derivation layer include the following:

[0950] 1) An operation connecting the first input and the second input; and

[0951] 2) An operation to pool the result of 2) above into a scalar dimension.

[0952] Example 82, Scalar Derivation

[0953] In the antigen-antibody binding affinity derivation layer of Example 81,

[0954] The above predicted antigen-antibody binding affinity value represents the probability that the antigen and antibody will bind, and

[0955] The operations performed when the input from the upper layer is forward propagated to the upper antigen-antibody binding affinity derivation layer further include the following:

[0956] 3) An operation that derives a real number between 0 and 1 by introducing the result of 2) above into a specific function.

[0957] Example 83, specific function limitation

[0958] In Example 82, the specific function is a softmax, a modified softmax function, an entropymax, a sigmoid, or a Gaussian kernel function.

[0959] Antigen-antibody binding affinity prediction model

[0960] Example 84, Antigen-Antibody Binding Affinity Prediction Model

[0961] Antigen-antibody binding affinity prediction model (hereinafter referred to as the prediction model),

[0962] Here, the above prediction model takes antigen sequences and antibody sequences as input and outputs predicted antigen-antibody binding affinity values, and

[0963] The above prediction model includes a chemical cross-attention layer selected from any one of Examples 63 to 72, and

[0964] The layer output of the chemical cross-attention layer above is reflected in the predicted antigen-antibody binding affinity value above.

[0965] Example 85, including two independent chemical cross-attention layers

[0966] In the prediction model of Example 84,

[0967] Here, the above prediction model includes a first chemical cross-attention layer and a second chemical cross-attention layer, and

[0968] The above first chemical cross-attention layer and second chemical cross-attention layer are each independently selected from any one of Examples 63 to 72, and

[0969] The first input of the above first chemical cross-attention layer is a characteristic expression derived from an antigen sequence, and the second input is a characteristic expression derived from an antibody sequence, and

[0970] The first input of the second chemical cross-attention layer above is a characteristic expression derived from an antibody sequence, and the second input is a characteristic expression derived from an antigen sequence, and

[0971] The predicted antigen-antibody binding affinity values ​​above reflect the layer output of the first chemical cross-attention layer above and the layer output of the second chemical cross-attention layer above.

[0972] Example 86, including an embedding layer

[0973] In any one of the prediction models of Examples 84 to 85,

[0974] The above prediction model further includes a first embedding layer and a second embedding layer, and

[0975] The above first embedding layer and second embedding layer are each independently selected from any one of Examples 48 to 52, and

[0976] The input to the first embedding layer above is an antigen sequence, and the input to the second embedding layer above is an antibody sequence.

[0977] Example 87, including a self-attention layer

[0978] In any one of the prediction models of Examples 84 to 86,

[0979] The above prediction model further includes a first self-attention layer and a second self-attention layer, and

[0980] The above first self-attention layer and second self-attention layer are each independently selected from any one of Examples 53 to 62, and

[0981] The first input of the first self-attention layer above is a feature expression derived from an antigen sequence, and

[0982] The second input of the second self-attention layer above is a characteristic expression derived from the antibody sequence.

[0983] Example 88, including a forward neural network layer

[0984] In any one of the prediction models of Examples 84 to 87,

[0985] The above prediction model further includes a first forward neural network layer and a second forward neural network layer, and

[0986] The first forward neural network layer and the second forward neural network layer are each independently selected from any one of Examples 73 to 79, and

[0987] The first input of the first self-attention layer above is a feature expression derived from an antigen sequence, and

[0988] The second input of the second self-attention layer above is a characteristic expression derived from the antibody sequence.

[0989] Example 89, model structure specification, 1 block

[0990] In any one of the prediction models of Examples 84 to 88,

[0991] The above prediction model includes the following transformation process during forward propagation of the input data:

[0992] 1) The input antigen sequence passes through the first embedding layer above and is converted into an antigen embedding representation;

[0993] 2) The antigen embedding expression of 1) above passes through the first self-attention layer above and is converted into a first antigen sequence characteristic expression;

[0994] 3) The input antibody sequence passes through the second embedding layer and is converted into an antibody embedding expression;

[0995] 4) The antibody embedding expression of 3) above passes through the second self-attention layer above and is converted into a first antibody sequence characteristic expression;

[0996] 5) The first antigen sequence characteristic expression of 2) above and the first antibody sequence characteristic expression of 4) above pass through the first chemical cross-attention layer above and are converted into a second antigen sequence characteristic expression,

[0997] Here, the first input of the above first chemical cross-attention layer is a first antigen sequence characteristic expression, and the second input is a first antibody sequence characteristic expression;

[0998] 6) The first antibody sequence characteristic expression of 4) above and the first antigen sequence characteristic expression of 2) above pass through the second chemical cross-attention layer above and are converted into a second antibody sequence characteristic expression,

[0999] Here, the first input of the second chemical cross-attention layer above is a first antibody sequence characteristic expression, and the second input is a first antigen sequence characteristic expression;

[1000] 7) The second antigen sequence characteristic representation of 5) above passes through the first forward neural network and is converted into a third antigen sequence characteristic representation;

[1001] 8) The second antibody sequence characteristic expression of 6) above passes through the second forward neural network and is converted into a third antibody sequence characteristic expression;

[1002] 9) Based on the third antigen characteristic expression of 7) above and the third antibody characteristic expression of 8) above, the predicted antigen-antibody binding affinity value is derived.

[1003] Example 90, output operation specific

[1004] In the prediction model of Example 89,

[1005] The conversion process of 9) above is a predicted value derived by inputting the third antigen characteristic expression of 7) above and the third antibody characteristic expression of 8) above into an antigen-antibody binding affinity derivation layer selected from any one of Examples 80 to 83.

[1006] Example 91, model structure specific, Nx block

[1007] In any one of the prediction models of Examples 84 to 90,

[1008] The above prediction model is each independently an embedding layer selected from Examples 48 to 52 class Includes,

[1009] The above prediction model includes Nx modules when Nx is an integer greater than or equal to 2, and

[1010] Each module is called the nth module (where n is an integer between 1 and Nx inclusive) in the order in which operations occur, and

[1011] For each case where n is an integer between 1 and Nx inclusive, the above nth module includes the following:

[1012] , and , stomach , and Each is independently a self-attention layer selected from any one of Examples 53 to 62;

[1013] , and , stomach , and Each is independently a chemical cross-attention layer selected from any one of Examples 63 to 72; and

[1014] , and , stomach , and Each is independently a forward neural network layer selected from any one of Examples 73 to 79;

[1015] The above prediction model includes the following transformation process during forward propagation of the input data:

[1016] 1) The input antigen sequence is the stomach Passes through and is converted into a zero antigen sequence characteristic representation;

[1017] 2) The entered antibody sequence is above Passes through and is converted into a characteristic expression of the 0th antibody sequence;

[1018] 3) For each case where n is an integer between 1 and Nx, the following a) through f) are performed:

[1019] a) The representation of the n-1 antigen sequence characteristics is above Passes through and is converted into the (2n-1) intermediate antigen sequence characteristic representation;

[1020] b) The representation of the n-1 antibody sequence characteristics is above Passes through and is converted into the characteristic expression of the (2n-1) intermediate antibody sequence;

[1021] c) The representation of the intermediate antigen sequence characteristics of a) above and the representation of the intermediate antibody sequence characteristics of b) above It passes through the chemical cross-attention layer and is converted into the (2n) intermediate antigen sequence characteristic expression,

[1022] Here, above The first input is a (2n-1) intermediate antigen sequence characteristic expression, and the second input is a (2n-1) intermediate antibody sequence characteristic expression;

[1023] d) The representation of the intermediate antigen sequence characteristics of a) above and the representation of the intermediate antibody sequence characteristics of b) above It passes through the chemical cross-attention layer and is converted into the (2n) intermediate antibody sequence characteristic expression,

[1024] Here, above The first input is a (2n-1) intermediate antibody sequence characteristic expression, and the second input is a (2n-1) intermediate antigen sequence characteristic expression;

[1025] e) The (2n) intermediate antigen sequence characteristic expression of c) above It passes through the forward neural network layer and is converted into a characteristic representation of the nth antigen sequence;

[1026] f) The expression of the (2n) intermediate antibody sequence characteristics of d) above It passes through the forward neural network layer and is converted into a characteristic representation of the nth antibody sequence;

[1027] g) The n-th antigen sequence characteristic representation of e) above and the n-th antibody sequence characteristic representation of f) above are transferred to the n+1 module, where, if n is Nx, this process is not performed;

[1028] 4) Predict antigen-antibody binding affinity values ​​based on the Nx antigen sequence characteristic representation and the Nx antibody sequence characteristic representation.

[1029] Example 92, output operation specific

[1030] In the prediction model of Example 91,

[1031] The conversion process of 4) above is a predicted value derived by inputting the above Nx antigen characteristic expression and the above Nx antibody characteristic expression into an antigen-antibody binding affinity derivation layer selected from Examples 80 to 83.

[1032] Example 93, Meaning of predicted values

[1033] In any one of the prediction models selected from Examples 84 to 92,

[1034] The antigen-antibody binding affinity output by the above prediction model is a value between 0 and 1 indicating the probability that the input antigen sequence and the input antibody sequence will bind.

[1035] Training method for antigen-antibody binding affinity prediction model

[1036] Example 94, antigen-antibody binding affinity

[1037] A method for training an antigen-antibody binding affinity prediction model (hereinafter referred to as the training method),

[1038] The above antigen-antibody binding affinity prediction model is any one of the prediction models selected from Examples 84 to 93, and

[1039] The above learning method includes the following:

[1040] a) The process of acquiring training data,

[1041] Here, the above training data includes antigen sequence-antibody sequence pairs labeled with antigen-antibody binding affinities;

[1042] b) The process of identifying chemical characteristics between antigen sequences and antibody sequences,

[1043] Here, the chemical properties are chemical properties capable of deriving one or more chemical attention guides selected from Examples 43 to 47; and

[1044] c) The process of training the above prediction model using the training data of a) above and the chemical characteristics of b) above.

[1045] Example 95, specification of the learning process

[1046] In the learning method of Example 94,

[1047] The learning process of c) above includes the following:

[1048] c-1) Divide the training data of a) above into two or more sets including a training set and a validation set;

[1049] c-2) Initialize the parameters of the above prediction model,

[1050] Here, the above parameter includes all learnable weights;

[1051] c-3) Input one unit of data included in the above training set,

[1052] Here, the data of the above unit may be data of a batch comprising a pair of antigen sequences-antibody sequences labeled with antigen-antibody binding affinities, or a plurality of pairs of antigen sequences-antibody sequences labeled with antigen-antibody binding affinities;

[1053] c-4) Input the data from the above one unit into the above prediction model and perform forward propagation to derive the predicted antigen-antibody binding affinity value;

[1054] c-5) Calculate the loss function using the above predicted values ​​and labeled antigen-antibody binding affinities,

[1055] c-6) Adjust the above parameters in a direction that minimizes the above loss function;

[1056] c-7) Repeat the above c-1) through c-5) by inputting new data until the predetermined criteria are satisfied.

[1057] Example 96, learning termination condition

[1058] In the learning method of Example 95,

[1059] The predetermined criteria in c-7) above include the following:

[1060] 1) Whenever the number of iterations of c-1) through c-5) above reaches a specific epoch or step, one unit of data included in the above validation set is forward propagated to the model, the validation loss is calculated using the derived antigen-antibody binding affinity predictions and the labeled antigen-antibody binding affinities, and the iteration is stopped if the above validation loss does not improve for a specific number of iterations;

[1061] 2) Whenever the number of iterations in c-1) through c-5) above reaches a specific epoch or step, one unit of data included in the validation set above is propagated to the model, and the iteration is stopped if the model's Precision, Recall, or F1 Score satisfies certain criteria when comparing the derived antigen-antibody binding affinity predictions with the labeled antigen-antibody binding affinities;

[1062] 3) Stop the iteration when the number of iterations of c-1) through c-5) above reaches a specific epoch or a specific step;

[1063] 4) Stop repetition after a certain period of time;

[1064] 5) Stop the iteration when the slope of the learning curve drops below a specific value;

[1065] 6) Stop the iteration if the rate of change of loss between consecutive epochs falls below a specific threshold;

[1066] 7) Stop the iteration when the learning rate drops below a specific value; or

[1067] 8) Any combination of the above criteria.

[1068] Example 97, Loss Function Example

[1069] In the learning method of Example 96,

[1070] The above loss function is the Mean Squared Error function, Cross-Entropy Loss, or Huber Loss function.

[1071] Example 98, Hyperparameter Optimization

[1072] In any one of the learning methods selected from Examples 94 to 97,

[1073] The above prediction model includes the following hyperparameters:

[1074] 1) The dimension E of the vector embedding each amino acid, as the size of the embedding dimension;

[1075] 2) Number of attention heads of the self-attention layer H1;

[1076] 3) Number of attention heads H2 of the chemical cross-attention layer;

[1077] 4) Number of prediction modules Nx;

[1078] 5) When using a batch with one unit of data, batch size B; and

[1079] 6) Number of hidden layers T of the forward neural network layers;

[1080] The above learning method further includes the process of optimizing the above hyperparameters.

[1081] Example 99, Hyperparameter Optimization Method

[1082] In Example 98, the method for optimizing the above hyperparameters may be Grid Search, Random Search, Bayesian Optimization, Genetic Algorithm, Hyperband, or any combination of the above methods.

[1083] Antigen-antibody binding affinity prediction method

[1084] Example 100, Method for Predicting Antigen-Antibody Binding Affinity

[1085] A method for predicting antigen-antibody binding affinity (hereinafter referred to as the prediction method) comprises the following:

[1086] The antigen sequence and antibody sequence are input into any one of the selected prediction models from Examples 84 to 93, and a result value is output.

[1087] Example 101, trained model

[1088] In the prediction method of Example 100,

[1089] The above prediction model is trained using any one of the learning methods selected from Examples 94 to 99.

[1090] Reinforcement learning method for antibody generation models

[1091] Example 102, Antibody generation model

[1092] As a method for reinforcement learning an antibody generation model (hereinafter referred to as the generative model learning method),

[1093] The above antibody generation model is configured to generate antibodies against the antigen of interest, and

[1094] The above method includes the following:

[1095] a) The process of pre-training the above antibody generation model,

[1096] Accordingly, the above antibody generation model learns the general patterns and characteristics of antibody sequences;

[1097] b) The process of fine-tuning the above antibody generation model using a compensation signal,

[1098] Here, the above compensation signal includes the predicted antigen-antibody binding affinity value output from a prediction model trained by any one of Examples 84 to 93 and any one of Examples 94 to 99.

[1099] Example 103, specification of the fine-tuning process

[1100] As a method for training a generative model of Example 102,

[1101] The fine-tuning process of b) above includes the following:

[1102] b-1) A process of obtaining predicted antigen-antibody binding affinity values ​​by inputting the antibody sequence and the antigen sequence of interest into the prediction model;

[1103] b-2) A process of determining a compensation signal using predicted antigen-antibody binding affinity values;

[1104] b-3) The process of adjusting the parameters of the above pre-trained antibody sequence in a direction that maximizes the above reward signal;

[1105] Here, the above fine-tuning process repeats steps b-1) through b-3) above by inputting a new antibody sequence until a preset standard is reached.

[1106] in silico antibody screening method

[1107] Example 104, in silico library screening method

[1108] An in silico antibody screening method (hereinafter referred to as the screening method) comprises the following:

[1109] a) The process of obtaining an in silico library,

[1110] Here, the above in silico library contains multiple antibody sequence information; and

[1111] b) Process of selecting candidate antibodies of interest,

[1112] Here, characteristic information is determined based on the sequence information of each antibody included in the above in silico library, and

[1113] If the above characteristic information satisfies predetermined criteria, it is selected as a candidate antibody of interest, and

[1114] The above characteristic information includes the antigen-antibody binding affinity with the antigen of interest, and

[1115] The above predetermined criterion is whether the output antigen-antibody binding affinity falls within a predetermined range by inputting the above antigen sequence of interest and each antibody sequence into any one of the prediction models selected from Examples 84 to 93, and

[1116] The above prediction model was trained using any one of Examples 94 to 99.

[1117] Example 105, addition of synthesis and evaluation process

[1118] In the screening method of Example 104,

[1119] The above screening method includes the following additional steps:

[1120] c) the process of synthesizing the antibody candidates of interest in b) above; and

[1121] d) A process of identifying the antibody of interest by evaluating the properties of each candidate antibody of interest synthesized in c) above.

[1122]

[1123] [Experimental Examples]

[1124] The invention provided by this specification will be described in more detail below through experimental examples and embodiments. These embodiments are intended solely to illustrate the contents disclosed by this specification, and it will be obvious to those skilled in the art that the scope of the contents disclosed by this specification is not to be interpreted as being limited by these embodiments.

[1125]

[1126] Chapter 1. Experimental Examples for the Method of Constructing an In-Silico Antibody Library

[1127] Experimental Example 1.1. Experimental Method and Materials

[1128] Experimental Example 1.1.1. Construction of an in silico antibody library

[1129] Antibody information was obtained from available sources. The above antibody information included the sequence of the light chain variable region, the sequence of the heavy chain variable region, or a pair of the two. An in silico antibody library was constructed using all possible combinations of the light chain variable region and the heavy chain variable region included in the above antibody information. Each combination resulted in a new antibody sequence. This is schematically illustrated in Figure 1.

[1130] Experimental Example 1.1.2. Prediction of Human Germline Similarity

[1131] For the antibody sequences included in the in silico antibody library according to Experimental Example 1.1.1, human germline similarity was predicted by the following method:

[1132] 1) The sequences of the light chain variable region and the heavy chain variable region of the antibody were entered into the ANARCI program (James Dunbar, Charlotte M. Deane, ANARCI: antigen receptor numbering and receptor classification, Bioinformatics, Volume 32, Issue 2, January 2016, Pages 298-300, https: / doi.org / 10.1093 / bioinformatics / btv552).

[1133] 2) Using the HMMER program (A NEW GENERATION OF HOMOLOGY SEARCH TOOLS BASED ON PROBABILISTIC INFERENCE SEAN R. EDDY (USA) Genome Informatics 2009. October 2009, 205-211), the similarity between each sequence and the species-specific antibody sequences in the IMGT GENE database (Vιronique Giudicelli, Denys Chaume, Marie-Paule Lefranc, IMGT / GENE-DB: a comprehensive database for human and mouse immunoglobulin and T cell receptor genes, Nucleic Acids Research, Volume 33, Issue suppl_1, 1 January 2005, Pages D256-D261, https: / doi.org / 10.1093 / nar / gki010) was compared to find the closest species.

[1134] 3) The closest antibody species found in ANARCI was output to predict human germline similarity.

[1135] Experimental Example 1.1.3. Prediction of Light-Heavy Chain Binding Possibility

[1136] For the antibody sequences included in the in silico antibody library according to Experimental Example 1.1.1, the possibility of light chain-heavy chain binding was scored using the following method:

[1137] 1) The light chain variable sequence and heavy chain variable sequence of each antibody were entered into the CNN-P model (Chinery, L., Jeliazkov, JR, & Deane, CM (2024). Humatch - fast, gene-specific joint humanisation of antibody heavy and light chains. mAbs, 16(1). https: / doi.org / 10.1080 / 19420862.2024.2434121).

[1138] 2) The above CNN-P model infers and outputs the probability that the input antibody sequence is a human sequence, and this was obtained and used for productivity prediction.

[1139] Experimental Example 1.1.4. Prediction of Protein Aggregation Degree

[1140] For the antibody sequences included in the in silico antibody library according to Experimental Example 1.1.1, the degree of protein aggregation was predicted by the following method:

[1141] 1) Using the sequences of each antibody, the protein structure was predicted by applying them to Alphafold2 or other 3D structure prediction methods.

[1142] 2) Based on the structure predicted in 1), graph data containing information on adjacent amino acids was extracted.

[1143] 3) The above graph data was applied to a model for predicting the degree of protein aggregation (Sun et. al., Enhancing protein aggregation prediction: a unified analysis leveraging graph convolutional networks and active learning, RSC Adv., 2024,14, 31439-31450) to predict the degree of antibody sequence aggregation.

[1144] Experimental Example 1.1.5. Prediction of Antigen-Antibody Binding Affinity

[1145] Antibody sequences included in the in silico antibody library according to Experimental Example 1.1.1 were input into an antigen-antibody binding strength prediction model to predict antigen-antibody binding strength.

[1146] The antigen-antibody binding strength prediction model used in this experiment is an embodiment of the antigen-antibody binding affinity prediction model disclosed in Chapter 2, and its structure is as follows:

[1147] 1) Embed the antigen sequence and perform self-attention using the multi-head attention algorithm.

[1148] 2) Embed the antibody sequence and perform self-attention using a multi-head attention algorithm.

[1149] 3) Using the results of 1) and 2), perform cross-attention between antigen and antibody.

[1150] 4) Chemical Cross Attention (CCA) is performed on the results of 1) and 2). The above CCA is performed using a multi-head attention algorithm, but with specific locations masked for each head. Masking includes, for example, the following: masking all locations other than the amino acids at the electron donor and electron acceptor sites of the antigen; masking all locations other than the amino acids at the electron donor and electron acceptor sites of the antibody; masking all locations other than the amino acids at hydrogen bonding sites; or masking all locations other than the amino acids at van der Waals bonding sites.

[1151] 5) The attention results of 3) and 4) are input into a neural network to derive antigen-antibody binding strength.

[1152] The above antigen-antibody binding strength prediction model was trained using data in which the antigen-antibody binding strength was labeled to the antibody sequence for a specific antigen sequence.

[1153] Experimental Example 1.1.6. Preparation of a vector for producing candidate antibodies of interest

[1154] To produce selected antibody candidates of interest by applying the prediction method of Experimental Examples 1.2 to 1.5 to each antibody sequence of the in silico antibody library constructed according to Experimental Example 1.1.1, a production vector was constructed as follows:

[1155] 1) The sequences of selected antibody candidates of interest were configured in the scFv form.

[1156] 2) The scFv sequence above was converted into a DNA sequence using DNA codons suitable for E. coli.

[1157] 3) The converted antibody DNA sequence was cloned into a pET-based vector to enable expression in E. coli.

[1158] Experimental Example 1.1.7. Antibody expression in Escherichia coli periplasm

[1159] Antibodies were expressed in E. coli periplasm using the production vector prepared according to Experimental Example 1.1.6. The specific procedure is as follows:

[1160] 1) Transformation: Escherichia coli BL21 strain was transformed with the scFv gene cloned into an AinB production vector. An overnight pre-culture of Escherichia coli was prepared.

[1161] 2) Deep Well Plate Preparation: LB medium containing antibiotics (containing 50 μg / mL kanamycin) was dispensed into deep well plates at a rate of 1300 μl per well.

[1162] 3) Addition of pre-culture: 50 μl of the E. coli pre-culture from 1) was added to each well of the deep-well plate above, and pipettes were made 3 to 4 times.

[1163] 4) Culture #1: Place the above deep-well plate in a shaking incubator and incubate overnight at 37°C.

[1164] 5) Culture #2: The above deep-well plates were cultured at 37 °C and 220 rpm until the OD (600 nm) of the cell solution reached 0.6 to 0.8. The above OD was measured using a 100 μl sample of the cell solution. After confirming the OD value, the medium was removed until the volume of medium remaining in each well was 1200 μl.

[1165] 6) IPTG induction: 120 μl of 10 mM IPTG was dispensed into the deep well plates above and pipetted 3-4 times per well.

[1166] 7) Culture #3: The above deep-well plates were incubated overnight in a shaking incubator at 30 °C and 220 rpm.

[1167] Experimental Example 1.1.8. Antibody Extraction

[1168] 1) Centrifugation #1: After antibody expression according to Experimental Example 1.7, the deep well plate used was centrifuged at 3,000g for 15 minutes at 4 °C to pellet the cells.

[1169] 2) Resuspension: The supernatant was removed from the product of 1), and the cell pellet was resuspended in 200 μl of 1X STE buffer.

[1170] 3) Culture #4: The mixture of 2) was cultured at 4 °C for 10 minutes.

[1171] 4) Addition of buffer and mixing: 300 μl of 0.2X STE buffer was added to the mixture from 3), and mixed by pipetting 2 to 3 times.

[1172] 5) Culture #5: The mixture of 4) was cultured at 4 °C for 30 minutes.

[1173] 6) Centrifugation #2: The mixture of 5) was centrifuged at 3,000g for 15 minutes at 4 °C.

[1174] 7) Transfer of supernatant: 250 μl of supernatant from a deep-well plate was transferred to a 96-well non-binding plate and stored at 4 °C. The supernatant obtained here is the periplasmic portion containing scFv.

[1175] Experimental Example 1.1.9. Dot Blot

[1176] 1) Membrane preparation: The NC membrane was cut to an appropriate size.

[1177] 2) Sample application: 5 μL of the periplasm portion obtained in Experimental Example 1.8 was applied to a designated location on the NC membrane. The sample order was maintained.

[1178] 3) Drying: The membrane was completely dried at 37°C for 10 minutes.

[1179] 4) Blocking: 20 mL of 5% skim milk PBST was added to the dot blot case along with the membrane. Blocked in a shaker at room temperature for 1 hour.

[1180] 5) Wash #1: 20 mL of 0.1% PBST was added to the upper dot blot case. Washed in a shaker at room temperature for 5 minutes. The wash solution was carefully discarded.

[1181] 6) Repeat #1: The process in 5) above was repeated three more times.

[1182] 7) Antibody culture: 20 mL of anti-HA-tag HRP conjugated antibody was added to the dot blot case. It was incubated in a shaker at room temperature for 1 hour.

[1183] 8) Wash #2: 20 mL of 0.1% PBST was added to the dot blot case. Washed in a shaker at room temperature for 5 minutes. The wash solution was carefully discarded.

[1184] 9) Repeat #2: The process in 8) above was repeated three more times.

[1185] 10) Preparation of Chemiluminescent Substrate: A chemiluminescent substrate was prepared by adding 1 mL of luminol and 1 mL of peroxide to a 5 mL tube (stored in foil).

[1186] 11) Substrate application: The NC membrane was placed in a transparent plastic sleeve. 1 mL of chemiluminescent substrate was applied.

[1187] 12) Substrate dispersion: The substrate was evenly dispersed throughout the membrane.

[1188] 13) Using detection tweezers, the membrane was carefully transferred to ChemiDoc and detection was performed.

[1189] Experimental Example 1.1.10. Immunoassay

[1190] 1) Plate preparation: The plates were washed three times with 300 μl / well of PBST (0.05% Tween).

[1191] 2) Antigen coating: PBS-based antigen (target antigen) was coated at 200 ng / well for 1 hour at 37°C.

[1192] 3) Wash #3: The plates were washed three times with 300 μl / well of 0.05% PBST.

[1193] 4) Periplasmic Fraction binding: 100 μl of the periplasmic fraction according to Experimental Example 1.7 was added and incubated in a 37°C incubator for 1 hour.

[1194] 5) Wash #4: The plates were washed three times with 300 μl / well of 0.05% PBST.

[1195] 6) Secondary antibody treatment: 100 μl of 1:5000 diluted 6X-HA tag, HRP conjugated antibody was added to each well.

[1196] 7) Wash #5: The plates were washed three times with 300 μl / well of 0.05% PBST.

[1197] 8) TMB treatment: 100 μl of TMB was added to each well.

[1198] 9) Stop TMB: 100 μl of 1N H2SO4 solution was added to each well.

[1199] 10) Detection: Absorbance was measured at 450 nm.

[1200] Experimental Example 1.2. Derivation of in-silico antibody candidates

[1201] Experimental Example 1.2.1. Construction of a basic library using available antibody data

[1202] Using antibody data extracted from public data (OAS) and patents, 2.07e+09 (2,071,324,363) VHs and 3.57e+08 (356,953,703) VLs were collected. An in silico antibody library was constructed according to Experimental Example 1.1.1. The antibodies contained in the above in silico antibody library totaled 7.4 x 10 17 It is a dog.

[1203] Experimental Example 1.2.2. Prediction of Immunogenicity and Productivity

[1204] For the antibodies in the basic library according to Experimental Example 1.2.1, antibodies were primarily selected by predicting immunogenicity and productivity. Specifically, for each antibody sequence included in the basic library, human germline similarity was calculated according to Experimental Example 1.1.2, and the CNN-P score was calculated according to Experimental Example 1.1.3. The results plotted on a two-dimensional plane are shown in Fig. 2.

[1205] Among these, subsequent exploration was conducted using antibody sequences that were predicted to have good productivity due to high light-heavy chain binding potential scores according to the CNN-P model, and were expected to have low immunogenicity due to high human germline similarity.

[1206] Experimental Example 1.2.3. Prediction of binding strength and binding specificity for antigen of interest

[1207] For the antibody sequences derived according to Experimental Example 1.2.2, binding strength and binding specificity for the antigen of interest were predicted. The antigens of interest that were searched were designated as Target 1, Target 2, and Target 3, respectively.

[1208] The above antibody sequences were input into the antibody side of the antigen-antibody binding strength prediction model. As the prediction model, a model trained on data in which antigen sequences were labeled as binding strengths with specific antigens was used. In order to predict not only the binding strength for the antigen of interest but also whether it binds specifically to the antigen of interest, the antigen-antibody binding strengths for other antigens to check for cross-reactivity were also predicted. For all antigen-antibody pairs used, the predictions of the binding prediction model were organized by antigen and are shown in Figure 3.

[1209] In the results of Figure 3, if the probability of binding to a different antigen is higher than the probability of binding to the antigen of interest, the antibody is considered to have low binding specificity to the antigen of interest and is excluded from experimental verification.

[1210] Experimental Example 1.2.4. Derivation of Candidate Antibodies of Interest

[1211] Based on the predicted values ​​of Experimental Example 1.2.2 and Experimental Example 1.2.3, antibodies with low immunogenicity (high human sequence similarity), good productivity (high light-heavy chain binding potential score according to the CNN-P model), and high binding affinity and binding specificity for the antigen of interest were derived as candidate antibodies of interest.

[1212] The above antibody candidate of interest was experimentally verified through Experimental Example 1.3 below.

[1213] Experimental Example 1.3. Screening of Antibodies of Interest

[1214] Experimental Example 1.3.1. Synthesis of Candidate Antibodies of Interest

[1215] According to Experimental Examples 1.1.6 to 1.1.8, the antibody candidate of interest derived in Experimental Example 1.2 was synthesized.

[1216] Experimental Example 1.3.2. Verification of binding affinity to the antigen of interest

[1217] For the candidate antibodies of interest according to Experimental Example 1.3.1, the binding affinity to the target antigen was verified by immunoassay according to Experimental Example 1.1.10. The experimental results for each target are shown in Figures 4 and 5.

[1218] As a result of the experiment, among the antibody candidates derived through Experimental Example 1.2, the proportion of antibodies that actually bind well to the antigen of interest was high. This demonstrates that the method for selecting antibodies of interest in this specification, particularly the method for deriving in silico antibody candidates, is a method for efficiently selecting antibodies of interest within a reasonable search space.

[1219] Experimental Example 1.3.3. Verification of specific binding affinity to the antigen of interest

[1220] The antibody candidate group for Target 1 was tested for specific binding to the antigen of interest by measuring the binding affinity to three additional antigens in addition to the antigen of interest. For some antibodies that bound to Target 1, the binding of the antigen of interest to three other non-target antigens (hereinafter referred to as Targets A, B, and C) was measured. The results are shown in Figure 6.

[1221] As a result of the experiment, it was confirmed that the candidate antibody selected according to Experimental Example 2 specifically binds only to the antigen of interest when its binding affinity to the antigen of interest and three non-target antigens was measured.

[1222]

[1223] confirmed.

[1224] Chapter 2. Experimental Examples for Antigen-Antibody Binding Affinity Prediction Models

[1225] Experimental Example 2.1. Experimental Method and Materials

[1226] Experimental Example 2.1.1. Acquisition of Antigen-Antibody Binding Affinity Training Data

[1227] Antibody patent sequences corresponding to each antigen are obtained. Then, the data is clustered based on 90% antibody sequence similarity to divide it into training, validation, and test data. In each dataset, antibody sequences corresponding to the antigen are labeled as 1, and non-corresponding antibody sequences are labeled as 0 to finally obtain training, validation, and test data.

[1228] Experimental Example 2.1.2. Acquisition of Antibody Generation Training Data

[1229] Antibody patent sequences are collected, and the corresponding antibody sequences are used as training data when generating new antibody sequences.

[1230] Experimental Example 2.1.3. Implementation of Computational Structures for Each Model

[1231] Antigen-antibody binding affinity prediction and generative models were implemented using Python 3.10, Flax, and the JAX library.

[1232] The antigen-antibody binding affinity prediction model consists of an embedding layer and 12 attention modules. The embedding layer has a dimension of 512, and the self-attention layer and chemical cross-attention layer consist of 8 heads, with the query, key, and value of each head having a dimension of 64. The final MLP layer consists of binary classification.

[1233] Experimental Example 2.1.4. Model Training

[1234] Chemical features were obtained using a CPU, and the antigen, antibody sequences, and chemical features were input into the model. The model was trained on a TPU v3-8 environment using a 5-fold cross-validation method with 100 epochs per fold and a batch size of 256. After training, the model with the highest F1 score was finally selected using validation data.

[1235] Experimental Example 2.2. Verification of the effect of introducing chemical cross-attention to an antigen-antibody binding affinity prediction model

[1236] According to Experimental Example 2.1.3, an antigen-antibody binding affinity prediction model (implementation model) described in the paragraph "Example of implementation of antigen-antibody binding affinity prediction model" was implemented. To demonstrate that learning is possible more efficiently when chemical cross-attention is introduced, an antigen-antibody binding affinity prediction model (comparison model) was also implemented, which has the same structure but without chemical attention guidance. Training data was acquired according to Experimental Example 2.1.1, and training was performed according to Experimental Example 2.1.4 using a method suitable for the model structure.

[1237] The learning results of the implemented model are shown in Figures 13 to 15.

[1238] The results of comparing the precision, recall, and F1 score of the implementation model and the comparison model above are shown in the following table:

[1239] LabelPrecisionRecallF1-ScoreF1-SupportsVanilla AttentionCCAVanilla AttentionCCAVanilla AttentionCCAAntigen 10.3331.0000.1110.2220.1670.3649Antigen 21.0001.0001.0001.0000.9990.9992Antigen 30.6900.8400.5000.5250.5800.64640

[1240] In the table above, Vanilla Attention refers to the comparison model, and CCA refers to the implementation model. Antigen 1, 2, and 3 each refer to different antigens.

[1241] The F1 score represents the harmonic mean of the model's precision and recall. When comparing the implementation model with the comparison model, it was confirmed that the implementation model had a higher MACRO F1 score (0.669). In conclusion, the implementation model showed improved performance compared to the comparison model (0.087 performance improvement), which demonstrates that the chemical attention guide applied to the implementation model contributed to efficient learning.

[1242] Experimental Example 2.3. Implementation of an antibody generation model using an antigen-antibody binding affinity prediction model as a compensation model

[1243] According to Experimental Example 2.1.3, the antibody generation model described in the paragraph "Reinforcement Learning of Antibody Generation Model" was implemented. Training data was obtained according to Experimental Example 2.1.2, and training was performed using a method suitable for the model structure according to Experimental Example 2.1.4.

[1244] The training results of the above model are shown in Figure 16.

[1245] In Figure 16, the X-axis represents the number of backpropagations of the generative model using reinforcement learning, and the Y-axis represents the result of evaluating whether the antibody sequences generated by reinforcement learning bind to a specific antigen using an antigen-antibody binding prediction model with CCA. As reinforcement learning progresses, the generative model modifies the generated results so that the antigen-antibody binding prediction model evaluates the binding status highly, and consequently, the probability of the generated antibody sequences binding to a specific antigen increases when verified experimentally.

[1246] According to the method for constructing an in-silico antibody library provided in this specification, 1) the in-silico computational cost is not excessively large to the point of being manageable, 2) the in-silico prediction is not excessively small to the point of being meaningless, and 3) an in-silico antibody library is constructed that contains a sufficiently diverse range of antibodies and has a high probability of discovering good antibodies, and using this, it is possible to screen for antibodies of interest that not only specifically bind to the antigen of interest and exhibit the desired effect, but are also advantageous for production and use.

[1247] According to the antigen-antibody binding affinity prediction model provided in this specification, antigen-antibody binding affinity can be effectively predicted even in situations where there is not much data, and by applying it to the in-silico antibody library, antibody sequences with a high probability of binding to a specific antigen can be screened, and by utilizing it as a compensation model for the antibody generation model, the antibody generation model can be optimized to generate high-quality antibody sequences.

Claims

1. A method for training a model to predict antigen-antibody binding affinity (hereinafter referred to as the training method), The above model is configured to receive an antigen sequence and an antibody sequence as input and output the antigen-antibody binding affinity, and The above model includes a Chemical Cross Attention (CCA) mechanism for modeling the interaction between the antigen sequence and the antibody sequence, and The above method includes the following: (a) The process of obtaining training data, Here, the above training data includes multiple antigen sequence-antibody sequence pairs and antigen-antibody binding affinities corresponding to each pair; (b) A process of identifying chemical characteristics for each antigen sequence and antibody sequence of the above training data, Here, the chemical properties are information related to the intermolecular interactions between each amino acid of the antigen sequence and each amino acid of the antibody sequence; and (c) A process of training the model using the training data obtained in (a) above and the chemical characteristics identified in (b) above, Here, during the training process of the above model, a chemical cross-attention operation is performed between the characteristic representation of the antigen sequence and the characteristic representation of the antibody sequence during each input data forward propagation, and The above chemical cross-attention operation performs a cross-attention operation between the information derived from the antigen sequence and the information derived from the antibody sequence, wherein By applying an attention guide based on the above chemical properties to the calculated cross-attention scores, the cross-attention scores at positions related to intermolecular interactions are relatively amplified, and A predicted value is derived based on the chemical attention value calculated as a result of performing the above chemical cross-attention operation, and Calculate the loss function based on the above predicted values ​​and the actual antigen-antibody binding affinity, and Adjust the model parameters in a way that minimizes the loss function.

2. In the training method of Paragraph 1, The above chemical properties include information selected from the following: Information on amino acid pairs capable of forming hydrogen bonds with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence; Information on amino acid pairs that can interact with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence; Information on amino acid pairs that can interact hydrophobically with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence; Information regarding amino acid pairs that can electrostatically interact with each other among amino acid pairs in the antigen sequence and amino acid pairs in the antibody sequence; or Any combination of the above information.

3. In the training method of Paragraph 2, The information regarding the above amino acid pairs includes the position of the amino acid within the antigen sequence capable of interacting with the corresponding molecules and the position of the amino acid within the antibody sequence.

4. In the training method of Paragraph 3, The information regarding the above amino acid pairs further includes information on the direction and strength of the interaction between the corresponding molecules, and The above intensity information may be an absolute or relative value.

5. In any one of the training methods selected from paragraphs 1 to 4, The information derived from the above antigen sequence is a value embedding the above antigen sequence, or a feature representation converted from a value embedding the above antigen sequence, and The information derived from the above antibody sequence is a value embedding the above antibody sequence, or a feature representation converted from a value embedding the above antibody sequence, and Here, the above feature expression refers to an artificially extracted feature (Extracted Feature) or a hidden feature (Hidden Feature).

6. In any one of the training methods selected from paragraphs 1 to 5, The above chemical cross-attention determines a query from information derived from the above antigen sequence, and Keys and values ​​are determined from the information derived from the above antibody sequence, The above attention scores are represented as an attention score matrix or its equivalent, and The element in the i-th row and j-th column of the above attention score matrix represents the attention score between the i-th amino acid in the antigen sequence and the j-th amino acid in the antibody sequence, and The attention guide based on the above chemical properties is represented as an attention guide matrix or its equivalent, and The element in the i-th row and j-th column of the above attention guide matrix indicates whether there is an intermolecular interaction or the degree of interaction between the i-th amino acid in the antigen sequence and the j-th amino acid in the antibody sequence, and The above chemical cross-attention operation includes a process of performing a Hadamard product operation between the above attention score matrix and the above attention guide matrix.

7. In any one of the training methods selected from paragraphs 1 through 6, The above chemical cross-attention determines a query from information derived from the above antibody sequence, and Keys and values ​​are determined from the information derived from the above antigen sequence, The above attention scores are represented as an attention score matrix or its equivalent, and When i and j are integers, The element in the i-th row and j-th column of the above attention score matrix represents the attention score between the i-th amino acid in the antibody sequence and the j-th amino acid in the antigen sequence, and The attention guide based on the above chemical properties is represented as an attention guide matrix or its equivalent, and The element in the i-th row and j-th column of the above attention guide matrix indicates whether there is an intermolecular interaction or the degree of interaction between the i-th amino acid in the antibody sequence and the j-th amino acid in the antigen sequence, and The above chemical cross-attention operation includes a process of performing a Hadamard product operation between the above attention score matrix and the above attention guide matrix.

8. In any one of the training methods selected from paragraphs 1 through 7, The above antigen-antibody binding affinity prediction model includes a first chemical cross-attention layer and a second chemical cross-attention layer, and During the training process of the above model, at the forward propagation of each input data, In the first chemical cross-attention layer above, a chemical cross-attention operation according to claim 6 is performed, and In the above second chemical cross-attention layer, a chemical cross-attention operation according to claim 7 is performed.

9. In the training method of either Paragraph 6 or Paragraph 8, The element in row i-j-column of the above attention guide matrix is ​​1 if there is an interaction between the amino acids of the antigen sequence and antibody sequence at the corresponding position, and 0 if there is no interaction.

10. In any one of the methods selected from paragraphs 1 to 9, The above chemical characteristics include first molecular interaction information and second molecular interaction information, and The above chemical cross-attention operation is a multi-head attention including a first head operation and a second head operation, and During the first head operation, an attention guide based on the first molecular interaction information is applied, and During the second head operation, an attention guide based on the second molecular interaction information is applied.

11. In any one of the training methods selected from paragraphs 1 to 10, The above model for predicting antigen-antibody binding affinity further includes a first embedding layer, a second embedding layer, a first self-attention layer, and a second self-attention layer, and During the training process of the above model, the following calculation is performed during the forward propagation of each input data: 1) The above antigen sequence passes through a first embedding layer and is converted into an antigen embedding expression; 2) The antigen embedding expression passes through the first self-attention layer and is converted into an antigen characteristic expression; 3) The above antibody sequence passes through a second embedding layer and is converted into an antibody embedding expression; 4) The antibody embedding expression passes through the second self-attention layer and is converted into an antibody characteristic expression; 5) Chemical cross-attention operations are performed using the above antigen characteristic expression and the above antibody characteristic expression.

12. In any one of the training methods selected from paragraphs 1 to 11, The above model for predicting antigen-antibody binding affinity outputs a real value between 0 and 1, and The above real value represents the probability that the input antigen sequence and the input antibody sequence will bind.

13. In any one of the training methods selected from paragraphs 1 to 12, The above antibody sequence comprises a sequence of a variable heavy chain (VH) and a variable light chain (VL), and optionally a distinguishing token.

14. A method for predicting antigen-antibody binding affinity (hereinafter referred to as the prediction method) comprises the following: (a) The process of obtaining antigen sequences and antibody sequences; (b) The process of identifying chemical characteristics of antigen sequences and antibody sequences, Here, the chemical properties are information related to the intermolecular interactions between each amino acid of the antigen sequence and each amino acid of the antibody sequence; and (c) A process of deriving an antigen-antibody binding affinity prediction value by inputting the antigen sequence and antibody sequence of (a) and the chemical characteristics of (b) into an antigen-antibody binding affinity prediction model trained by any one of the training methods selected from claims 1 to 13, Here, the chemical properties used in the above training method and the chemical properties identified in (b) above are information regarding the same type of intermolecular interactions.

15. A method for providing an antibody of interest that binds to an antigen of interest, comprising: (a) A process of obtaining sequence information of the first antibody and the second antibody; (b) Process of generating candidate antibodies of interest, Here, the provisional candidate for the antibody of interest is: Having a variable heavy chain having the same sequence as the variable heavy chain of the first antibody above, and Having a variable light chain with the same sequence as the variable light chain of the second antibody above; (c) A process for predicting the characteristics of the above-mentioned candidate antibody of interest, Here, the above characteristic includes the binding affinity with the antigen of interest; (d) a process of synthesizing the antibody of interest when the characteristics of the antibody candidate of interest satisfy predetermined criteria; and (e) A process of determining whether the candidate antibody of interest synthesized in the above process (d) is the antibody of interest by confirming the binding affinity of the candidate antibody of interest to the antigen of interest.

16. In the method of paragraph 15, The characteristics of the antibody candidate of interest predicted in the above (c) process further include physicochemical stability, immunogenicity, or a combination thereof.

17. In any one of the methods selected from paragraphs 15 to 16, The binding affinity of the candidate antibody of interest and the antigen of interest predicted in the above (c) process is predicted by the method of claim 14, and Here: The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and The sequence of the antibody candidate of interest is used as the antibody sequence of the above prediction method.

18. A method for providing an antibody of interest that binds to an antigen of interest, comprising: (a) The process of obtaining sequence information of a candidate antibody of interest, Here: The above candidate antibody of interest is generated by combining the sequence information of the first antibody and the second antibody; The variable heavy chain sequence of the above-mentioned candidate antibody of interest is obtained from the variable heavy chain sequence of the above-mentioned first antibody; The variable light chain sequence of the above-mentioned candidate antibody of interest is obtained from the variable light chain sequence of the above-mentioned first antibody; The predicted values ​​for the characteristics of the above-mentioned antibody candidate of interest satisfy predetermined criteria; and The characteristics of the candidate antibody of interest to be predicted above include binding affinity for the antigen of interest; (b) a process for synthesizing the above-mentioned candidate antibody of interest; and (c) A process of determining whether the candidate antibody of interest synthesized in the above process (b) is the antibody of interest by confirming the binding affinity of the candidate antibody of interest to the antigen of interest.

19. In the method of paragraph 18, The characteristics of the candidate antibody of interest to be predicted in the above (a) process further include physicochemical stability, immunogenicity, or a combination thereof.

20. In any one of the methods selected from paragraphs 18 to 19, The binding affinity of the antibody candidate of interest and the antigen of interest predicted in the above process (a) is predicted by the method of claim 14, and Here: The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and The sequence of the antibody candidate of interest is used as the antibody sequence of the above prediction method.

21. A method for providing a candidate antibody of interest expected to bind to an antigen of interest, comprising: (a) A process of obtaining sequence information of the first antibody and the second antibody; (b) Process of generating temporary candidates for the antibody of interest, Here, the provisional candidate for the antibody of interest is: Having a variable heavy chain having the same sequence as the variable heavy chain of the first antibody above, and Having a variable light chain with the same sequence as the variable light chain of the second antibody above; (c) A process for predicting the characteristics of a provisional candidate for the antibody of interest, Here, the above characteristic includes the binding affinity with the antigen of interest; (d) If the characteristics of the provisional candidate for the antibody of interest satisfy predetermined criteria, the provisional candidate for the antibody of interest is provided as the antibody of interest candidate.

22. In the method of paragraph 21, The characteristics of the provisional candidate for the antibody of interest predicted in the above process (c) further include physicochemical stability, immunogenicity, or a combination thereof.

23. In any one of the methods selected from paragraphs 21 to 22, The binding affinity of the temporary candidate for the antibody of interest and the antigen of interest predicted in the above process (c) is predicted by the method of claim 14, and Here: The sequence of the antigen of interest is used as the antigen sequence of the above prediction method; and The sequence of the antibody candidate of interest is used as the antibody sequence of the above prediction method.