Humanized antibody generation method and device, database construction method and device and storage medium
By evaluating the immunogenicity and humanization of antibody protein sequences, a variable region fragment encoded by the V(D)J gene with low immunogenicity and high humanization was screened out, generating humanized antibody sequences. This solved the problems of low humanization and high immunogenicity in existing technologies, and achieved efficient and low-cost antibody development.
Patent Information
- Application Number
- CN202410826883.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2025-12-26
AI Technical Summary
Existing technologies have low levels of humanization when generating human antibodies, which can easily lead to immunogenicity, causing the antibodies to develop drug resistance in the body, increasing toxic side effects and reducing efficacy. Furthermore, the antibody development process lacks diversity, making it difficult to efficiently screen for highly exploitable antibody sequences.
By evaluating the immunogenicity and humanization of antibody protein sequences, V(D)J gene variable region fragments with low immunogenicity and high humanization were screened out. Antibody sequences were generated by simulating V(D)J rearrangement, and a humanized antibody database was constructed to ensure the humanization and development potential of antibodies.
It reduces the immunogenicity of antibodies, improves the efficiency and developability of antibody development, reduces the cost of experimental validation, and generates antibody libraries with high humanization and low immunogenicity, as well as high developability.
Smart Images

Figure CN121215033A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of bioinformatics, and more particularly, relates to a method for generating humanized antibodies, a database construction method, a device, and a storage medium. BACKGROUND
[0002] An antibody is a Y-shaped structure composed of two heavy chains and two light chains, the heavy chain contains one N-terminal variable region and three C-terminal constant regions, and the light chain contains one N-terminal variable region and one C-terminal constant region. The variable region of the heavy chain and the variable region of the light chain are combined, the first constant region of the heavy chain is combined with the constant region of the light chain, and the two light chains and the variable region and the first constant region of the heavy chain together form the arms of the Y-shaped structure of the antibody, while the other constant regions of the two heavy chains are combined to form the Fc region of the antibody. Generally, the variable region of the antibody binds to the antigen, and the constant region does not participate in the binding of the antigen. The variable region is composed of a β-sheet conformation formed by four relatively conserved framework regions and three complementarity determining regions (CDR regions) with relatively high diversity, which form loops connecting the β-sheet, and the CDR region is the key part of the antibody variable region that binds to the antigen. Antibody chains are encoded by genes from three independent loci on different chromosomes, one of which encodes the heavy chain, and the light chain has two types, lambda (λ) light chain and kappa (κ) light chain, which are encoded by two loci. For each heavy chain of the antibody, lambda (λ) light chain and kappa (κ) light chain are composed of different gene segments. The heavy chain variable region of the antibody is encoded by three gene segments, V (Variable), D (Diversity), and J (Joining), while the lambda (λ) light chain and kappa (κ) light chain of the antibody are encoded by two gene segments, V and J, and the constant region is encoded by C (Constant) gene Figure 1 ) gene. Through V(D)J gene rearrangement, mutation, and heavy chain and light chain recombination, almost infinite antibodies can be generated, so for almost all antigens, corresponding antibodies can be generated in vivo.
[0003] At present, the method for generating human antibody sequence is mainly through extracting lymphocytes of healthy people or people with related diseases, and combining human antibody library by PCR technology, such as CN101892525A and CN101294308A. Other antibody libraries generated by random mutation and the like often have low humanization degree of antibodies, are easy to produce immunogenicity, produce drug resistance antibodies in the body of patients, and cause side effects and reduced efficacy. Generally, a large number of experiments are carried out to verify immunogenicity after the antibodies are extracted, and then further drug development and the like are applied, which consumes a large amount of manpower and material resources. Patent US10774138B2 provides a method for constructing a human antibody DNA library by rearranging V(D)J fragments of human germ lines. However, this method does not consider the diversity in the expression process of the generated antibody sequence, the V(D)J fragments of human germ lines are maternal fragments in the human body without any in vivo affinity maturation, the sequence of the CDR region fragment, especially the CDRH3, is relatively single, the antigen binding diversity is not sufficient, and in the process of antibody development, the developability of the generated fragments is not high, such as low expression, aggregation and precipitation, and insufficient stability. SUMMARY
[0004] In view of the above defects or improvement needs of the prior art, the present application provides a humanized antibody generation method, a database construction method, equipment and a storage medium, which aims to evaluate the immunogenicity and / or humanization degree of antibody protein sequences, screen out antibody V(D)J gene coded variable region fragments with low immunogenicity and high humanization degree, simulate V(D)J rearrangement to generate antibody protein sequences, ensure the humanization degree of antibodies, thereby reducing the immunogenicity of antibody proteins and improving the developability of the database, thereby solving the technical problems that the existing antibody acquisition method needs experimental evaluation of immunogenicity or the obtained human antibody library has low developability.
[0005] To achieve the above-mentioned purpose, according to one aspect of the present application, a humanized antibody generation method is provided, comprising the following steps:
[0006] (1) obtaining sequence data set of antibody protein and evaluating its immunogenicity and / or humanization degree; wherein the antibody protein immunogenicity is the ability of the antibody protein to induce immune response of the body to itself or related proteins or cause immune related events, and the humanization degree refers to the similarity between the antibody and human antibody;
[0007] (2) obtaining the variable region sequence fragments encoded by the V(D)J genes of the antibody sequence obtained in step (1), deleting the variable region sequence fragments encoded by the V genes and J genes with immunogenicity higher than a preset threshold or humanization degree lower than a preset threshold according to their immunogenicity and humanization degree, as the variable region sequence fragments encoded by the V(D)J genes of the humanized antibody, and constructing a data set of the variable region sequence fragments encoded by the V(D)J genes of the humanized antibody based on the variable region sequence fragments encoded by the V(D)J genes of the humanized antibody;
[0008] The fragments of the variable region encoded by the V(D)J genes include:
[0009] The fragments of the variable region encoded by the V genes: the sequence fragments encoded by the heavy chain V genes, the sequence fragments encoded by the Kappa light chain V genes, and the sequence fragments encoded by the Lambda light chain V genes;
[0010] The sequence fragments encoded by the heavy chain D genes,
[0011] The fragments of the variable region encoded by the J genes: the sequence fragments encoded by the heavy chain J genes, the sequence fragments encoded by the Kappa light chain J genes, and the sequence fragments encoded by the Lambda light chain J genes;
[0012] (3) according to the principles of V(D)J rearrangement and antibody light / heavy chain recombination, selecting the variable region sequence fragments encoded by the V(D)J genes of the humanized antibody from the data set of the variable region sequence fragments encoded by the V(D)J genes of the humanized antibody obtained in step (2), and combining them into antibody sequences as humanized antibody sequences.
[0013] Preferably, the method for generating the humanized antibody, wherein the immunogenicity of the antibody protein in step (1) is evaluated according to the frequency of drug resistance antibody and / or the frequency of neutralizing antibody, according to the principle that the higher the frequency of drug resistance antibody and / or the frequency of neutralizing antibody, the stronger the immunogenicity;
[0014] The humanization degree of the antibody protein is evaluated according to its sequence similarity with known human antibodies, according to the principle that the higher the sequence similarity with human antibodies, the higher the humanization degree;
[0015] The sequence data of the antibody protein is the sequence of the experimentally verified therapeutic antibody protein.
[0016] Preferably, the method for generating the humanized antibody, wherein step (2) retains all the variable region sequence fragments encoded by the heavy chain D genes.
[0017] Preferably, the method for generating humanized antibodies, wherein the sequence fragment encoded by the light chain V gene in step (3) is selected from the sequence fragment encoded by the Kappa light chain V gene or the sequence fragment encoded by the Lambda light chain V gene; the sequence fragment encoded by the light chain J gene is selected from the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the light chain V gene; the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the Kappa light chain V gene is the sequence fragment encoded by the Kappa light chain J gene, and the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the Lambda light chain V gene is the sequence fragment encoded by the Lambda light chain J gene.
[0018] According to another aspect of the present application, a method for constructing a virtual database of humanized antibodies is provided, characterized by comprising the following steps: collecting sequence data of humanized antibodies generated by the method for generating humanized antibodies according to the present application.
[0019] Preferably, the method for constructing a virtual database of humanized antibodies, when the CDR region of the sequence fragment of the variable region encoded by the V(D)J gene of the humanized antibody selected in step (3) is the same as the CDR region of the existing humanized antibody sequence data, the sequence fragment of the variable region encoded by the V(D)J gene encoding a different sequence from the existing humanized antibody sequence data is selected for the framework region, and the combination of the antibody sequence as the humanized antibody sequence.
[0020] According to another aspect of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method for generating humanized antibodies according to the present application when executing the program.
[0021] According to another aspect of the present application, a non-transitory computer readable storage medium is provided, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the method for generating humanized antibodies according to the present application.
[0022] According to another aspect of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method for constructing a virtual database of humanized antibodies according to the present application when executing the program.
[0023] According to another aspect of the present application, a non-transitory computer readable storage medium is provided, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the method for constructing a virtual database of humanized antibodies according to the present application.
[0024] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0025] The method for generating humanized antibodies provided by the present application uses verified low-immunogenicity and high-humanization-degree variable region sequence fragments encoded by V(D)J genes to obtain new antibody sequences according to V(D)J rearrangement, simulates the normal biological process in the human body, guarantees the low immunogenicity of the generated antibodies from the source, reduces the difficulty of antibody development, and saves the cost generated due to a large number of immunogenicity verification experiments.
[0026] The database obtained according to the method for constructing a humanized antibody virtual database provided by the present application has the characteristics of high humanization and low immunogenicity, and the preferred scheme uses the sequences of experimentally verified therapeutic antibody proteins as source data, so that the humanized antibody virtual database constructed has the characteristics of high developability and sequence diversity.
[0027] In the preferred scheme, the CDR region of the selected variable region sequence fragment data of the humanized antibody V(D)J gene is the same as the CDR region of the existing humanized antibody sequence data, and the framework region is selected from the variable region sequence fragment data of the V(D)J gene encoding a different sequence from the existing humanized antibody sequence data, which not only maintains the diversity of the framework region but also reduces the redundancy of the CDR region. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a schematic diagram of the sequence structure encoded by the V(D)J gene of the variable region of the antibody protein; wherein the distribution of the four framework regions (black) and the three CDR regions (red) of the heavy chain (top) and the light chain (bottom) of the antibody, wherein the V gene encodes most of the framework region and CDR1, CDR2 and part of CDR3, the D gene encodes part of CDR3, the J gene encodes part of CDR3 and the last framework region.
[0029] Figure 2 is a schematic diagram of the combination of the variable region sequence fragment data of the humanized antibody heavy chain V(D)J gene selected by the embodiment of the present application into an antibody sequence as a humanized antibody sequence;
[0030] Figure 3 is a schematic diagram of the combination of the variable region sequence fragment data of the humanized antibody light chain V(D)J gene selected by the embodiment of the present application into an antibody sequence as a humanized antibody sequence;
[0031] Figure 4 is a diagram of the evaluation results of the humanization degree of the humanized antibody virtual database generated by the embodiment of the present application;
[0032] Figure 5is a score result plot of the aggregation and precipitation developed by the virtual database of humanized antibodies generated by the embodiment of the present application for evaluation;
[0033] Figure 6 is a score result plot of the viscosity developed by the virtual database of humanized antibodies generated by the embodiment of the present application for evaluation;
[0034] Figure 7 is a score result plot of the non-specific binding developed by the virtual database of humanized antibodies generated by the embodiment of the present application for evaluation. DETAILED DESCRIPTION
[0035] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0036] The present application provides a method for generating humanized antibodies, comprising the following steps:
[0037] (1) Obtain the sequence data set of the antibody protein and evaluate its immunogenicity and / or humanization degree; wherein the immunogenicity of the antibody protein is the ability of the antibody protein to induce the immune response of the body to the self or related protein or to cause the immune related event, the stronger the immunogenicity of the antibody protein, the higher the probability of inducing the immune response of the body to the self or related protein or causing the immune related event, and the frequency of inducing the immune response of the body to the self or related protein or causing the immune related event can be evaluated, such as the frequency of forming Anti-Drug Antibody (ADA) and / or Neutralizing Antibody (NAb) of the antibody protein; the humanization degree refers to the similarity between the antibody and the human antibody, which can be measured by the sequence similarity of the antibody, and the higher the humanization degree, the more likely the antibody is suitable for drug development; specifically:
[0038] The immunogenicity of the antibody protein is evaluated according to the frequency of the anti-drug antibody and / or the frequency of the neutralizing antibody, according to the principle that the higher the frequency of the anti-drug antibody and / or the frequency of the neutralizing antibody, the stronger the immunogenicity;
[0039] The humanization degree of the antibody protein is evaluated according to the sequence similarity with the known human antibody, according to the principle that the higher the sequence similarity with the human antibody, the higher the humanization degree.
[0040] Preferably, the sequence data of the antibody protein is the sequence of a therapeutic antibody protein verified by experiment, ensuring that the virtual antibody database constructed using the antibody generation method has high developability.
[0041] (2) For the antibody sequence obtained in step (1), obtain the fragment of the variable region encoded by the V(D)J gene, delete the variable region sequence fragment encoded by the V gene and J gene with immunogenicity higher than a preset threshold or humanization degree lower than a preset threshold according to its immunogenicity and humanization degree, as the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody, and construct a data set of the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody based on the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody.
[0042] The fragment of the variable region encoded by the V(D)J gene includes:
[0043] The fragment of the variable region encoded by the V gene: the sequence fragment encoded by the heavy chain V gene (HV), the sequence fragment encoded by the Kappa light chain V gene (KV), and the sequence fragment encoded by the Lambda light chain V gene (LV);
[0044] The sequence fragment encoded by the heavy chain D gene (HD),
[0045] The fragment of the variable region encoded by the J gene: the sequence fragment encoded by the heavy chain J gene (HJ), the sequence fragment encoded by the Kappa light chain J gene (KJ), and the sequence fragment encoded by the Lambda light chain J gene (LJ).
[0046] Since the sequence fragment encoded by the heavy chain D gene (HD) has high variability, it can be directly retained regardless of its immunogenicity and humanization degree.
[0047] The data of the data set of the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody include: the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody, and / or a virtual humanized antibody V(D)J gene encoded variable region sequence fragment with immunogenicity lower than a preset threshold or humanization degree higher than a preset threshold constructed using a machine learning algorithm; Preferably, according to the sequence data set of the antibody protein, construct an antibody sequence with immunogenicity lower than a preset threshold or humanization degree higher than a preset threshold using a protein language model, and extract the fragment of the variable region encoded by the V(D)J gene as the virtual humanized antibody V(D)J gene encoded variable region sequence fragment.
[0048] (3) according to the principle of V(D)J rearrangement and antibody light heavy chain recombination, variable region sequence fragment data encoded by humanized antibody V(D)J gene from the data set of variable region sequence fragment data of humanized antibody V(D)J gene obtained in step (2) are selected respectively, and combined into an antibody sequence as a humanized antibody sequence;
[0049] Wherein, the sequence fragment encoded by the light chain V gene is selected from the sequence fragment encoded by the Kappa light chain V gene (KV) or the sequence fragment encoded by the Lambda light chain V gene (LV); the sequence fragment encoded by the light chain J gene is selected from the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the light chain V; the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the Kappa light chain V gene (KV) is the sequence fragment encoded by the Kappa light chain J gene (KJ), and the sequence fragment encoded by the J gene of the same gene as the sequence fragment encoded by the Lambda light chain V gene (LV) is the sequence fragment encoded by the Lambda light chain J gene (LJ).
[0050] The application provides a humanized antibody virtual database construction method, including the following steps: collecting humanized antibody sequence data generated according to the humanized antibody generation method provided by the application.
[0051] Preferably, when the CDR region of the variable region sequence fragment data of the humanized antibody V(D)J gene selected in step (3) is the same as the CDR region of the existing humanized antibody sequence data, the variable region sequence fragment data of the V(D)J gene encoding a different sequence from the existing humanized antibody sequence data is selected, and the antibody sequence is combined as a humanized antibody sequence.
[0052] The following is an example:
[0053] Example 1
[0054] The humanized antibody generation method provided in this embodiment includes the following steps:
[0055] (1) Obtain the sequence data set of the antibody protein and evaluate its immunogenicity and humanization degree; specifically: in this embodiment, 940 WHO registered therapeutic antibody data (Thera-SAbDab database) are downloaded https: / / opig.stats.ox.ac.uk / webapps / sabdab-sabpred / therasabdab / search / ?all=true ).
[0056] The immunogenicity of the antibody protein is evaluated by obtaining the frequency data of the antibody drug resistance antibody (ADA) of the antibody on the market in FDA or searching the corresponding therapeutic antibody ADA data in PubMed https: / / www.accessdata.fda.gov / scripts / cder / daf / Antibody proteins with an ADA occurrence frequency of ≥10% are labeled as high immunogenic antibody proteins, and antibody proteins with an ADA occurrence frequency of <10% are labeled as low immunogenic antibody proteins.
[0057] The degree of humanization of antibody proteins was evaluated using BioPhi to calculate the degree of humanization of the heavy or light chain of the antibody. https: / / biophi.dichlab.org / humanization / humanness / If the OASis Identity of the heavy or light chain of an antibody exceeds 80%, the antibody is considered to have a good degree of humanization and is marked as a humanized antibody chain; if the OASis Identity of the heavy or light chain of an antibody is less than 80%, the antibody is considered to have a poor degree of humanization and is marked as a non-humanized antibody chain.
[0058] (2) For the antibody sequence obtained in step (1), obtain the variable region segment encoded by its V(D)J gene. According to its immunogenicity and humanization degree, delete the variable region sequence segments encoded by the V and J genes whose immunogenicity is higher than the preset threshold or whose humanization degree is lower than the preset threshold. These segments are used as the variable region sequence segments encoded by the humanized antibody V(D)J gene. Based on the variable region sequence segments encoded by the humanized antibody V(D)J gene, construct a data set of variable region sequence segments encoded by the humanized antibody V(D)J gene.
[0059] The segments of the variable region encoded by the V(D)J gene include:
[0060] Segments of the variable regions encoded by the V gene: the sequence segment encoded by the heavy chain V gene (HV), the sequence segment encoded by the Kappa light chain V gene (KV), and the sequence segment encoded by the Lambda light chain V gene (LV);
[0061] The sequence segment encoded by the heavy chain D gene (HD);
[0062] Fragments of the variable regions encoded by the J gene: the sequence fragment encoded by the heavy chain J gene (HJ), the sequence fragment encoded by the Kappa light chain J gene (KJ), and the sequence fragment encoded by the Lambda light chain J gene (LJ).
[0063] Because the sequence fragment encoded by the heavy chain D gene (HD) is highly variable, it is directly retained regardless of its immunogenicity and degree of humanization.
[0064] The specific process of this embodiment is as follows:
[0065] The sequence of antibody heavy chain is divided into V, D, J three gene coding fragments, respectively denoted as HV, HD and HJ; the light chain of antibody is first divided into Kappa light chain and Lambda light chain, the Kappa light chain is divided into V gene coding and J gene coding fragments, respectively denoted as KV and KJ; the Lambda light chain is divided into V gene coding and J gene coding fragments, respectively denoted as LV and LJ.
[0066] The HV, HJ, KV, KJ, LV and LJ fragments marked as high immunogenicity or non-human source are removed, and the HD fragment is not screened. The repeated HV, HD and HJ, KV, KJ, LV and LJ fragments are removed.
[0067] The HD fragments are coded as HD1, HD2, HD3, …, HDn (n = 1, 2, 3, …). n The CDR1 and CDR2 sequences of HV fragments are extracted and spliced into a CDR sequence, and coded as CDR1, CDR2, …, CDRn (n = 1, 2, 3, …) according to the uniqueness of the CDR sequence. For different HV fragments of the same CDR sequence, code from 1 to m (m = 1, 2, 3, …), and the code of each HV fragment is HV1, HV2, HV3, …, HVn (n = 1, 2, 3, …). n-m The CDR1 and CDR2 and part of the CDR3 sequences of KV and LV fragments are extracted and spliced into a CDR sequence, and coded as CDR1, CDR2, …, CDRn (n = 1, 2, 3, …) according to the uniqueness of the CDR sequence. For different KV and LV fragments of the same CDR sequence, code from 1 to m (m = 1, 2, 3, …), and the code of each KV and LV fragment is KV1, KV2, KV3, …, KVn (n = 1, 2, 3, …), LV1, LV2, LV3, …, LVn (n = 1, 2, 3, …). n-m n-m The part of the CDR3 sequences of HJ, KV and LJ fragments are extracted and coded as CDR1, CDR2, …, CDRn (n = 1, 2, 3, …) according to the uniqueness of the CDR sequence. For different HJ, KV and LJ fragments of the same CDR sequence, code from 1 to m (m = 1, 2, 3, …), and the code of each HJ, KJ and LJ fragment is HJ1, HJ2, HJ3, …, HJn (n = 1, 2, 3, …), KJ1, KJ2, KJ3, …, KJn (n = 1, 2, 3, …), LJ1, LJ2, LJ3, …, LJn (n = 1, 2, 3, …). n-m n-m n-m .
[0068] The data of the variable region sequence fragment data set of the humanized antibody V(D)J gene coding in the embodiment includes: the variable region sequence fragment of the humanized antibody V(D)J gene coding, which is directly collected as the variable region sequence fragment data set of the humanized antibody V(D)J gene coding.
[0069] (3) According to the principle of V(D)J rearrangement and antibody light and heavy chain recombination, variable region sequence fragment data encoded by humanized antibody V(D)J gene from the data set obtained in step (2) is selected respectively, and combined into an antibody sequence as a humanized antibody sequence.
[0070] The humanized antibody virtual database construction method provided in the embodiment comprises the following steps: humanized antibody sequence data generated according to the humanized antibody generation method provided in the embodiment is generated and collected.
[0071] When the CDR region of the variable region sequence fragment data of the humanized antibody V(D)J gene selected in step (3) is the same as the CDR region of the existing humanized antibody sequence data, the variable region sequence fragment data of the V(D)J gene encoding a different sequence from the existing humanized antibody sequence data is selected, and combined into an antibody sequence as a humanized antibody sequence. Specifically as follows:
[0072] S1, the HV n-m , HD n , HJ n-m fragments of the heavy chain are arranged and combined according to the uniqueness of the CDR sequence of each sequence, that is, the first subscript number n is arranged and combined, and for each combination, for the fragments with the same CDR sequence, that is, the fragments with the same first subscript n, the second subscript m is randomly extracted, that is, the fragments with different framework regions, to ensure the diversity of the framework region. The randomly extracted HV, HD and HJ fragments are arranged into an antibody heavy chain ( Figure 2 ). In the embodiment, there are 243 HV fragments with different CDR sequences, 817 HD fragments with different CDR sequences, and 158 HJ fragments with different CDR sequences. A maximum of 243x817x158=31,367,898 CDR sequence unique heavy chain sequences can be generated.
[0073] S2, the KV n-m , KJ n-m and Lambda light chain LV n-m , LJ n-m fragments of the Kappa light chain are arranged and combined according to the uniqueness of the CDR sequence of each sequence, that is, the first subscript number n is arranged and combined, and for each combination, for the fragments with the same CDR sequence, that is, the fragments with the same first subscript n, the second subscript m is randomly extracted, that is, the fragments with different framework regions, to ensure the diversity of the framework region. The randomly extracted KV, KJ fragments are combined into an antibody Kappa light chain, and the LV, LJ fragments are combined into an antibody Lambda light chain ( Figure 3). In this embodiment, there are 269 CDR sequence different KV fragments, 30 CDR sequence different KJ fragments, 64 CDR sequence different LV fragments, and 18 CDR sequence different LJ fragments, and a maximum of 9,222 CDR sequence unique light chain sequences can be generated.
[0074] S3, the recombination of the heavy chain and the two light chains, a maximum of about 2.9 trillion CDR region sequence different antibody sequences can be generated, which maintains the diversity of the framework region while reducing the redundancy of the CDR region, and the antibody library has the characteristics of high humanization, low immunogenicity, high developability, and sequence diversity.
[0075] According to the above method, a humanized antibody virtual database is generated.
[0076] The humanization degree and developability of the antibody library provided in this embodiment are evaluated as follows:
[0077] Evaluate the degree of humanization: 100 antibody sequences are randomly generated, and the ratio of Kappa chain to Lambda chain in the antibody light chain is 4:1, i.e. 80 Kappa light chains and 20 Lambda light chains. The humanization degree of 100 randomly generated antibody sequences and 940 WHO registered therapeutic antibody sequences is calculated by BioPhi (https: / / biophi.dichlab.org / humanization / humanness / ), and the OASis percentile is used as the humanization score, and a box plot comparison is made, as shown in Figure 4 , it can be seen that the humanization degree of the antibody library has been significantly improved compared with the 940 therapeutic antibodies.
[0078] Evaluate the developability: the developability of 100 randomly generated antibody sequences and 940 WHO registered therapeutic antibody sequences is evaluated by the developability evaluation tool of Great Bay Bio https: / / alfadax.greatbay-bio.com / , and a box plot comparison is made for the aggregation precipitation score, the viscosity score, and the non-specific binding score. As shown in Figure 5 , 6, 7, the aggregation precipitation score of Example 1 is lower than that of the 940 therapeutic antibodies, and the viscosity score and the non-specific binding score are comparable to those of the 940 therapeutic antibodies, indicating that the developability of the antibody library at least reaches the level of the 940 therapeutic antibodies, and in terms of aggregation precipitation, it has been significantly improved compared with the 940 therapeutic antibodies.
[0079] The humanized antibody virtual database provided in this embodiment is applied to antibody development, and the following scheme can be used:
[0080] 1. Downstream users can directly build phage antibody libraries, yeast antibody libraries, etc. based on this antibody library, and use phage display technology, yeast display technology, etc. to screen antibody sequences.
[0081] 2. Downstream users can first use a computer to perform virtual screening of this antibody library. The screening can be carried out using the above methods, such as phage display and yeast display.
[0082] 3. Molecules that have undergone further screening can be used for final affinity verification using affinity assays such as surface plasmon resonance (SRR), biomembrane interference (BLI), and enzyme-linked immunosorbent assay (ELISA).
[0083] Example 2
[0084] The method for generating humanized antibodies provided in this embodiment includes the following steps:
[0085] (1) Obtain the sequence dataset of antibody proteins and evaluate their immunogenicity and humanization; specifically: in this embodiment, data on 940 WHO-registered therapeutic antibodies were downloaded from the Thera-SAbDab database. https: / / opig.stats.ox.ac.uk / webapps / sabdab-sabpred / therasabdab / search / ?all=true ).
[0086] Immunogenicity evaluation of antibody proteins involves obtaining data on the frequency of anti-drug antibodies (ADAs) against FDA-approved antibodies or searching for ADA data for corresponding therapeutic antibodies in PubMed. https: / / www.accessdata.fda.gov / scripts / cder / daf / Antibody proteins with an ADA occurrence frequency of ≥10% are labeled as high immunogenic antibody proteins, and antibody proteins with an ADA occurrence frequency of <10% are labeled as low immunogenic antibody proteins.
[0087] The degree of humanization of antibody proteins was evaluated using BioPhi to calculate the degree of humanization of the heavy or light chain of the antibody. https: / / biophi.dichlab.org / humanization / humanness / If the OASis Identity of the heavy or light chain of an antibody exceeds 80%, the antibody is considered to have a good degree of humanization and is marked as a humanized antibody chain; if the OASis Identity of the heavy or light chain of an antibody is less than 80%, the antibody is considered to have a poor degree of humanization and is marked as a non-humanized antibody chain.
[0088] (2) For the antibody sequence obtained in step (1), obtain the fragment of the variable region encoded by the V(D)J gene, delete the variable region sequence fragment encoded by the V gene and the J gene with immunogenicity higher than a preset threshold or humanization degree lower than a preset threshold according to the immunogenicity and humanization degree thereof, as the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody, and construct a data set of the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody based on the variable region sequence fragment encoded by the V(D)J gene of the humanized antibody;
[0089] The fragment of the variable region encoded by the V(D)J gene includes:
[0090] The fragment of the variable region encoded by the V gene: the sequence fragment encoded by the heavy chain V gene (HV), the sequence fragment encoded by the Kappa light chain V gene (KV), and the sequence fragment encoded by the Lambda light chain V gene (LV);
[0091] The sequence fragment encoded by the heavy chain D gene (HD);
[0092] The fragment of the variable region encoded by the J gene: the sequence fragment encoded by the heavy chain J gene (HJ), the sequence fragment encoded by the Kappa light chain J gene (KJ), and the sequence fragment encoded by the Lambda light chain J gene (LJ).
[0093] Since the sequence fragment encoded by the heavy chain D gene (HD) has high variability, it is directly retained regardless of its immunogenicity and humanization degree.
[0094] The specific process of the embodiment is as follows:
[0095] The sequence of the antibody heavy chain is divided into three gene coding fragments V, D, and J, which are denoted as HV, HD, and HJ respectively; the light chain of the antibody is first divided into Kappa light chain and Lambda light chain, and the Kappa light chain is divided into V gene coding and J gene coding fragments, which are denoted as KV and KJ respectively; and the Lambda light chain is divided into V gene coding and J gene coding fragments, which are denoted as LV and LJ respectively.
[0096] Remove the HV, HJ, KV, KJ, LV, and LJ fragments marked as high immunogenicity or non-humanization, and do not screen the HD fragment. Remove the repeated HV, HD, and HJ, KV, KJ, LV, and LJ fragments.
[0097] The HD fragments are coded according to 1 to n (n = 1, 2, 3, …), and the code of each HD fragment is HD nCDR1 and CDR2 sequences of the HV fragments are extracted and spliced into a CDR sequence, and coded from 1 to n (n = 1, 2, 3,...) according to the uniqueness of the CDR sequence, and coded from 1 to m (m = 1, 2, 3,...) for different HV fragments of the same CDR sequence, and the coding of each HV fragment is HV n-m CDR1 and CDR2 sequences of the KV, LV fragments are extracted and spliced into a CDR sequence, and coded from 1 to n (n = 1, 2, 3,...) according to the uniqueness of the CDR sequence, and coded from 1 to m (m = 1, 2, 3,...) for different KV, LV fragments of the same CDR sequence, and the coding of each KV, LV fragment is KV n-m , LV n-m Part of the CDR3 sequences of the HJ, KV, LJ fragments are extracted, and coded from 1 to n (n = 1, 2, 3,...) according to the uniqueness of the CDR sequence, and coded from 1 to m (m = 1, 2, 3,...) for different HJ, KV, LJ fragments of the same CDR sequence, and the coding of each HJ, KJ, LJ fragment is HJ n-m , KJ n-m , LJ n-m .
[0098] The data of the humanized antibody variable region sequence fragment data set in the embodiment includes: humanized antibody variable region sequence fragments encoded by the humanized antibody variable region gene fragments and variable region sequence fragments encoded by virtual humanized antibody V(D)J genes with immunogenicity lower than a preset threshold or humanization degree higher than a preset threshold constructed by a machine learning algorithm.
[0099] In the embodiment, the protein language model is used to construct antibody sequences with immunogenicity lower than a preset threshold or humanization degree higher than a preset threshold according to the sequence data set of the antibody protein, and the coding of the variable region of the V(D)J gene is extracted as the variable region sequence fragment encoded by the virtual humanized antibody V(D)J gene. The specific steps are as follows:
[0100] a) The heavy chain or light chain of 940 antibody sequences is respectively input into 4 protein language models (esm1v_t33_650M_UR90S_1, esm1v_t33_650M_UR90S_4, esm2_t33_650M_UR50D, esm2_t36_3B_UR50D), and the probability of each model output is calculated;
[0101] b) record the highest probability amino acid output by each protein language model, calculate the frequency of the highest probability amino acid, and if the frequency of the highest probability amino acid is greater than or equal to 3 (i.e. at least 3 of the 4 models recommend the same mutation, there is more than two protein language models that agree on the highest probability amino acid, and this amino acid is different from the original amino acid), mutate the amino acid at this position to the amino acid with a frequency of the highest probability amino acid greater than or equal to 3.
[0102] c) apply all the mutations recommended by the language model to generate two sequences, one being the original sequence and the other being the mutated sequence with all the recommended positions mutated, and keep both sequences.
[0103] d) for the mutated sequence generated in step 4c), apply the immunogenicity and humanization degree tags, the immunogenicity tag is the same as the original sequence, and the humanization degree tag is re-labeled according to the method of step 3, and screen out antibody protein sequences with high immunogenicity or non-humanization.
[0104] (3) According to the principle of V(D)J rearrangement and antibody light and heavy chain recombination, select the variable region sequence fragment data encoded by the humanized antibody V(D)J gene from the humanized antibody V(D)J gene encoded variable region sequence fragment data set obtained in step (2), and combine them into an antibody sequence as a humanized antibody sequence.
[0105] The humanized antibody virtual database construction method provided in this embodiment includes the following steps: collecting the humanized antibody sequence data generated according to the humanized antibody generation method provided in this embodiment.
[0106] When the CDR region of the humanized antibody V(D)J gene encoded variable region sequence fragment data selected in step (3) is the same as the CDR region of the existing humanized antibody sequence data, select the variable region sequence fragment data encoded by the V(D)J gene with a different sequence from the existing humanized antibody sequence data, and combine them into an antibody sequence as a humanized antibody sequence. The specific steps are as follows:
[0107] S1, arrange and combine the HV n-m , HD n , HJ n-m fragments according to the uniqueness of the CDR sequence of each sequence, i.e. arrange and combine the first subscript number n, and for each combination, randomly extract fragments with different second subscripts m, i.e. fragments with different framework regions, to ensure the diversity of the framework region. Arrange the randomly extracted HV, HD and HJ fragments into an antibody heavy chain ( Figure 2). In this embodiment, there are 316 HV fragments with different CDR sequences, 910 HD fragments, and 176 HJ fragments with different CDR sequences. A maximum of 316 x 910 x 176 = 50,610,560 CDR sequence unique heavy chain sequences can be generated.
[0108] S2, arrange and combine KVn-m, KJn-m of Kappa light chain and LVn-m, LJn-m of Lambda light chain according to the uniqueness of CDR sequence of each sequence, i.e. arrange and combine the first subscript number n, and in each combination, for fragments with the same CDR sequence, i.e. fragments with the same first subscript n, randomly extract fragments with different second subscripts m, i.e. fragments with different framework regions, to ensure the diversity of the framework region. The randomly extracted KV and KJ fragments form the antibody Kappa light chain, and the LV and LJ fragments form the antibody Lambda light chain. Figure 3 ) In this embodiment, there are 357 KV fragments with different CDR sequences, 31 KJ fragments with different CDR sequences, 83 LV fragments with different CDR sequences, and 19 LJ fragments with different CDR sequences. A maximum of 357 x 31 + 83 x 19 = 12,644 CDR sequence unique light chain sequences can be generated.
[0109] S3, recombination of heavy chain and two light chains, a maximum of about 640 billion CDR region sequence different antibody sequences can be generated, about 2 times more than example 1, increasing the sequences consistent with evolutionary information, maintaining the diversity of the framework region while reducing the redundancy of the CDR region, and the antibody library has the characteristics of high humanization, low immunogenicity, high developability, and sequence diversity.
[0110] According to the above method, a humanized antibody virtual database is generated.
[0111] The humanization degree and developability of the antibody library provided in this embodiment are evaluated as follows:
[0112] Randomly generate 100 antibody sequences, in which the ratio of Kappa chain to Lambda chain of antibody light chain is 4:1, i.e. 80 Kappa light chains and 20 Lambda light chains. Calculate the humanization degree of 100 randomly generated antibody sequences and 940 WHO registered therapeutic antibody sequences by BioPhi (https: / / biophi.dichlab.org / humanization / humanness / ), and take OASis percentile as the humanization score to make a box plot comparison, as shown in Figure 4As shown, the humanization degree of Example 2 is not only significantly improved compared with the 940 therapeutic antibodies, but also significantly improved compared with the humanization degree of Example 1, indicating that the humanization degree of the antibody library evolved by the language model has been further improved.
[0113] The developability of 100 randomly generated antibody sequences and 940 WHO registered therapeutic antibody sequences was evaluated by the developability evaluation tool of Great Bay Bio https: / / alfadax.greatbay-bio.com / , and the aggregation precipitation score, viscosity score, and non-specific binding score were compared by box plot. Figure 5 As shown in Figs. 6 and 7, Example 2 has significantly improved aggregation precipitation score, viscosity score, and non-specific binding score compared with the 940 therapeutic antibodies and Example 1, indicating that the developability of the antibody library evolved by the language model has been significantly improved.
[0114] The virtual database of humanized antibodies provided in this embodiment has the following applications:
[0115] 1. Downstream can directly establish phage antibody library, yeast antibody library, etc. on this antibody library, and screen the antibody sequences by phage display technology, yeast display technology, etc.
[0116] 2. Downstream can first screen this antibody library by computer, and screen by the above-mentioned methods by phage, yeast display, etc.
[0117] 3. The molecules screened further can be verified for affinity by surface plasmon resonance (SRR), biological membrane interference technology (BLI), enzyme-linked immunosorbent assay (ELISA), etc.
[0118] Those skilled in the art will readily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, and improvement within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for generating a humanized antibody, characterized by, The method comprises the following steps: (1) obtaining a sequence dataset of an antibody protein and evaluating its immunogenicity and / or humanization degree; wherein the immunogenicity of the antibody protein is the ability of the antibody protein to induce an immune response of the body to self or related proteins or to cause an immune-related event, and the humanization degree refers to the similarity between the antibody and a human antibody; (2) obtaining a fragment of a variable region encoded by a V(D)J gene of the antibody sequence obtained in step (1), deleting a variable region sequence fragment encoded by a V gene and a J gene with a high immunogenicity higher than a preset threshold or a low humanization degree lower than a preset threshold according to the immunogenicity and the humanization degree, as a variable region sequence fragment encoded by a humanized antibody V(D)J gene, and constructing a humanized antibody V(D)J gene variable region sequence fragment dataset based on the variable region sequence fragment encoded by the humanized antibody V(D)J gene; The fragment of the variable region encoded by the V(D)J gene comprises: The fragment of the variable region encoded by the V gene: a sequence fragment encoded by a heavy chain V gene, a sequence fragment encoded by a Kappa light chain V gene, and a sequence fragment encoded by a Lambda light chain V gene; a sequence fragment encoded by a heavy chain D gene, The fragment of the variable region encoded by the J gene: a sequence fragment encoded by a heavy chain J gene, a sequence fragment encoded by a Kappa light chain J gene, and a sequence fragment encoded by a Lambda light chain J gene; (3) selecting a variable region sequence fragment data of a humanized antibody V(D)J gene from the humanized antibody V(D)J gene variable region sequence fragment dataset obtained in step (2) according to the principle of V(D)J rearrangement and antibody light / heavy chain recombination, and combining the variable region sequence fragment data into an antibody sequence as a humanized antibody sequence.
2. The method of generating a humanized antibody according to claim 1, wherein The immunogenicity of the antibody protein in step (1) is evaluated according to the frequency of drug resistance antibodies and / or the frequency of neutralizing antibodies, and the principle is that the higher the frequency of drug resistance antibodies and / or the frequency of neutralizing antibodies, the stronger the immunogenicity; The humanization degree of the antibody protein is evaluated according to the sequence similarity with a known human antibody, and the principle is that the higher the sequence similarity with the human antibody, the higher the humanization degree; The sequence data of the antibody protein is the sequence of a therapeutic antibody protein verified by experiments.
3. The method of generating a humanized antibody according to claim 1, wherein All the sequence fragments of the variable regions encoded by the heavy chain D genes are retained in step (2).
4. The method of generating a humanized antibody according to claim 1, wherein The sequence fragment of the light chain V gene is selected from a sequence fragment of a Kappa light chain V gene or a sequence fragment of a Lambda light chain V gene; the sequence fragment of the light chain J gene is selected from a sequence fragment of a J gene of the same gene as the sequence fragment of the light chain V gene; the sequence fragment of the J gene of the same gene as the sequence fragment of the Kappa light chain V gene is a sequence fragment of a Kappa light chain J gene, and the sequence fragment of the J gene of the same gene as the sequence fragment of the Lambda light chain V gene is a sequence fragment of a Lambda light chain J gene.
5. A method for constructing a virtual database of humanized antibodies, characterized by, The method comprises the following steps: The humanized antibody sequence data generated according to the method for generating a humanized antibody according to any one of claims 1 to 4 is recorded.
6. The method of constructing a humanized antibody virtual database according to claim 5, wherein, When the CDR region of the variable region sequence fragment data encoded by the humanized antibody V(D)J gene selected in step (3) is the same as the CDR region of the existing humanized antibody sequence data, the variable region sequence fragment data encoded by the V(D)J gene encoding a different sequence from the existing humanized antibody sequence data is selected, and combined into an antibody sequence as a humanized antibody sequence.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method for generating a humanized antibody according to any one of claims 1 to 4 when executing the program.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method for generating a humanized antibody according to any one of claims 1 to 4.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method for constructing a virtual database of humanized antibodies according to claim 5 or 6 when executing the program. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method for constructing a virtual database of humanized antibodies according to claim 5 or 6.
Citation Information
Patent Citations
Method for humanizing recombinant phages antibody library
CN101294308A
Method for constructing natural humanized IgG Fab phage antibody library
CN101892525A
Combinatorial antibody libraries and uses thereof
US10774138B2