antibody library
Synthetic antibody libraries are designed to mimic the natural human repertoire by generating theoretical segment pools and matching them to a reference set, addressing sequence diversity limitations and immunogenicity issues, resulting in a diverse and effective antibody library for therapeutic applications.
Patent Information
- Application Number
- JP2020193227
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2010-07-16
- Filing Date
- 2020-11-20
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2031-07-14
AI Technical Summary
Existing antibody libraries are limited by sequence diversity, disproportionately enriching certain sequences and overrepresenting others, often requiring redesign or 'humanization' for therapeutic use, and may contain non-natural amines that are immunogenic.
Development of synthetic antibody libraries that mimic the natural human repertoire by generating theoretical segment pools and matching them to a reference set to ensure diversity and human character, using methods to select and include segments based on frequency and physicochemical properties.
The solution provides a diverse antibody library with improved human sequence representation, reducing immunogenicity and enhancing the ability to recognize a wide variety of antigens, while maintaining desirable properties.
Smart Images

Figure 0007788794000407 
Figure 0007788794000408 
Figure 0007788794000409
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application is a continuation of U.S. Provisional Patent Application No. 61 / 365,194, filed July 16, 2010. No. 60 / 699,994, filed on Dec. 1, 2003, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Antibodies have profound implications as research tools and in diagnostic and therapeutic applications. However, identifying useful antibodies can be difficult, and once identified, the antibodies They require considerable redesign or "humanization" before they are suitable for therapeutic application in humans. This often requires "processing."
[0003] Many methods for identifying antibodies are derived from amplification of nucleic acid from B cells or tissues. Some of these methods involve the display of libraries of antibodies that are synthesized. However, many of these approaches have limitations. For example, most human antibody libraries known in the art are can be obtained experimentally or cloned from biological sources (e.g., B cells) Such libraries therefore contain only the antibody sequence diversity that can be matched. This may disproportionately enrich some sequences while overrepresenting other sequences, particularly those that bind to human antigens. Completely lacking or disproportionately reducing the amount of Most synthetic libraries contain non-natural (i.e., non-human) amines that have the potential to be immunogenic. It has other limitations, such as the occurrence of amino acid sequence motifs.
[0004] Thus, it is non-immunogenic (i.e., it is a human antibody) and has desirable properties (e.g., A diverse antibody library containing candidate antibodies with the ability to recognize a wide variety of antigens However, achieving such a library requires a Competitive strategies to generate diverse libraries while still maintaining the human character of the sequences within The present invention provides these and other desirable features. Antibody libraries having the same structure and methods for making and using such libraries This provides a method for Summary of the Invention [Means for solving the problem]
[0005] The present invention relates to, inter alia, CDRH3, CDRL3, heavy chain, light chain, and / or full-length ( Design and development of synthetic libraries that mimic the diversity of the natural human repertoire of (complete) antibody sequences. In some embodiments, the present invention provides improvements in the production and maintenance of CDRH3 sequences. A library containing or encoding (e.g., polynucleotides or polypeptides) ( TN1, DH, for consideration of inclusion in the physical expression of (e.g., antibody libraries) We defined a method for generating theoretical segment pools of N2 and H3-JH segments. In one embodiment, the present invention provides a method for determining the number of segments in a theoretical segment pool. Determine the frequency of occurrence (or segment usage weight) in the reference set for each of the segments To achieve this, we compared the individual members of these theoretical segment pools to a reference set of CDRH3 sequences. It defines and provides a method for matching members of a given set of CDRH3 sequences. Although the present invention may be used as a subject-specific reference set or subset, The present invention also defines and provides methods for generating subsets. Filter the original reference set to obtain the provided reference set with pre-epidemic features. It provides a method for applying the method to the reference set rather than the theoretical segment pool. Methods for defining and / or identifying segments present within a CDRH3 sequence are also provided. Such segments may be used, for example, for inclusion in a physical library. can be added to the theoretical segment pool for consideration. The frequency of occurrence of a particular segment may be used to identify the segment for inclusion in a physical library. While useful for selecting for nucleotide sequences, the present invention also provides for the selection of nucleotide sequences for inclusion in a physical library. can be used to select segments (alone or in addition to any other criteria) The standard also provides a number of physicochemical and biological properties.
[0006] In some embodiments, the present invention provides methods for detecting stochastically distributed sequences of molecules that are site-by-site stochastic in composition or sequence. (sitewise-stochastic), and therefore, It is well known in the art in that it is not as inherently random as certain other libraries. provides a library that is different from certain other libraries known in the , for a discussion of information content and randomness, see U.S. Pat. No. 6,119,493, which is incorporated by reference in its entirety. (See Example 14 of US Patent Application Publication No. 2009 / 0181855). In this embodiment, the degenerate oligonucleotides have sequences (e.g., CDRH3, CDRL3, Matching to a reference set of heavy chain, light chain, and / or full-length (complete) antibody sequences This will be used to further increase the diversity of library members, with further improvements. That's fine.
[0007] The present invention also provides a method for detecting a patient's condition by performing the analysis described herein. For example, by generating a CDRH3 reference set as in Example 3; By generating theoretical segment pools as in Examples 5-7; Example 4 and 8, matching members of the theoretical segment pool to the reference set. and by using envelopes in physical libraries, as in Examples 8-9. By selecting members of the theoretical segment pool for inclusion, the physical library sequences whose members are related to one another in that they are selected for inclusion in a library. As in Examples 12-16, a library having degenerate oligonucleotides is provided. Methods for further increasing the sequence diversity of a species by utilizing tides are also , is provided.
[0008] In some embodiments, the present invention provides a method for determining the activity of a CDRH3, a CDRL3, a heavy chain, a light chain, and / or a nucleotide sequence. or polynucleotide libraries and polypeptides containing full-length (complete) antibody sequences Libraries and methods for making and using such libraries are provided. Provide.
[0009] In some embodiments, the present invention provides a library or Comprising, consisting essentially of, or consisting of any of the theoretical segment pools Provide a library.
[0010] In some embodiments, the present invention provides a computational method for determining the in vivo activity of the enzyme TdT. By mimicking the activity, a theoretical pool of segments is generated, followed by the library Matching to a large reference dataset of CDR sequences to select for inclusion in the These theoretical segments can be optimized to match the CDR sequences in the reference dataset. Recognize that it will also be reproduced.
[0011] In one embodiment, the present invention provides a compound having the structure: [TN1]-[DH]-[N2]-[H3-JH at least about 10 encoding a CDRH3 polypeptide having 4 Polynucleotides of TN1 is a library of polynucleotides comprising the sequences of Tables 9-10 and 18-19. A polypeptide corresponding to any of the TN1 polypeptides of Tables 25-26 or a polypeptide corresponding to any of the TN1 polypeptides of Tables 25-26 DH is a polypeptide produced by translation of any one of the nucleotides of A polypeptide corresponding to any of the DH polypeptides in Tables 9, 11, 17-25, and 28 or the translation of any of the DH-encoding polynucleotides in Tables 16, 25, and 27. N2 is a polypeptide produced by the method of any one of Tables 9, 12, 18-25, and 30. A polypeptide corresponding to any of the N2 polypeptides of Tables 25 and 29 a polypeptide produced by translation of either the H3-J or H3-J polynucleotide H is any of the H3-JH polypeptides of Tables 9, 13, 15, 18-25, and 32. or the H3-JH-encoding polypeptides of Tables 14, 25, and 31. A polynucleotide is a polypeptide produced by translation of any of the nucleotides. Provides a library of.
[0012] In some embodiments, the present invention provides a method for identifying at least about 1% of the sequences in a library, 5%, or 10% of the structure provided above or the live structure provided herein A library having any of the following structures is provided:
[0013] In one embodiment, the present invention provides a TN1 CDRH3 polypeptides produced by a set of DH, N2, and H3-JH polypeptides A library containing polynucleotides encoding polypeptides is provided.
[0014] In some embodiments, the present invention provides a set of TN1 polypeptides provided in Table 26. , a set of DH polypeptides provided in Table 28, an N2 polypeptide provided in Table 30 and the set of H3-JH polypeptides provided in Table 32. and providing a library containing polynucleotides encoding CDRH3 polypeptides. .
[0015] In certain embodiments, the invention relates to a polypeptide whose members have at least one of the polypeptides described above. exhibiting (or encoding a polypeptide exhibiting) at least one certain percent identity A library, for example, having the structure: [TN1]-[DH]-[N2]-[H3-JH] at least about 10 4 containing a polynucleotide of a library, wherein TN1 is a TN1 polypeptide of Tables 9-10 and 18-26; a polypeptide that is at least about 80%, 90%, or 95% identical to any of Polypeptides produced by translation of any of the TN1 polynucleotides of Tables 25-26 a polypeptide that is at least about 80%, 90%, or 95% identical to DH is at least about 100% identical to any of the DH polypeptides of Tables 9, 11, 17-25, and 28. A polypeptide that is 80%, 90%, or 95% identical to a polypeptide of Tables 16, 25, and 2 Polypeptides and at least one other polypeptide produced by translation of any of the seven DH-encoding polynucleotides N2 is a polypeptide that is at least about 80%, 90%, or 95% identical to a polypeptide of the present invention. At least about 80% of any of N2 polypeptides 9, 12, 18-25, and 30 , 90%, or 95% identical polypeptides or N2 codes of Tables 25 and 29 and a polypeptide produced by translation of any of the polynucleotides, %, 90%, or 95% identical polypeptides, and H3-JH is a polypeptide that is identical to any of the polypeptides in Tables 9, 13 1, 15, 18-25, and 32 H3-JH polypeptides and at least about 8 0%, 90%, or 95% identical to a polypeptide of Tables 14, 25, and 31 Polypeptides produced by translation of any of the H3-JH encoding polynucleotides a library of polypeptides that are at least about 80%, 90%, or 95% identical to Offer Lee.
[0016] In some embodiments, the invention includes a polynucleotide encoding a light chain variable region. a library comprising: a light chain variable region comprising: (a) one or more of positions 4, 49, and (b) VK1-05 sequences differing at one or more positions 4, 49, 46, and (c) VK1-12 sequences differing at positions 4, 49, and 6 (d) VK1-33 sequences differing at one or more positions 4, 49, and 46; (e) VK1-39 sequences differing at one or more positions 2, 4, 46, and 49; (f) a VK2-28 sequence that differs at one or more positions 2, 4, 36, and 49; (g) VK3-11 sequences; (g) VK3-11 sequences differing at one or more positions 2, 4, 48, and 49 (h) VKs differing at one or more positions 2, 4, 48, and 49; 3-20 sequence; and / or (i) one or more of positions 4, 46, 49, and A library is provided in which the VK4-1 sequences are selected from the group consisting of 66 different VK4-1 sequences.
[0017] In one embodiment, the present invention provides a method for producing a light chain polypeptide comprising administering to a subject a subject a light chain polypeptide sequence comprising two or more of the light chain polypeptide sequences provided in Table 3. a light chain comprising a polypeptide sequence at least about 80%, 90%, or 95% identical to a sequence A library containing polynucleotides encoding the variable regions is provided.
[0018] In some embodiments, the present invention provides polypeptides in which the light chain variable region is a polypeptide provided in Table 3. A library containing the sequence of the target gene is provided.
[0019] In one embodiment, the present invention provides a light chain comprising a polynucleotide encoding a light chain variable region. and the L3-VL polypeptide sequence of the light chain variable region is an L3-VL germline Provide libraries that differ by two or three residues in positions 89-94 compared to the sequence In some embodiments, a light chain containing a single light chain germline sequence and variants thereof is used. In one embodiment, a library is provided, which is produced from different light chain germline sequences. The variants produce a library encoding multiple light chain germline sequences and their variants. The light chain L3-VL antibodies provided herein can be combined to produce the same. Any of the germline sequences may differ at two or three residues between positions 89 and 94; Those skilled in the art will appreciate that any other L3-VL sequences can also be used in the libraries provided by the present invention. that can be diversified according to the principles described herein to produce In some embodiments, the present invention provides a method for producing a polypeptide encoding an antibody light chain variable region. The library contains polynucleotides, and the antibody light chain variable region is one or more (i) an amino acid sequence identical to the L3-VL germline sequence ( See, e.g., Table 1); (ii) residue 89 compared to the L3-VL germline sequence (iii) an amino acid sequence containing two substitutions at ~94; and (iii) an L3-VL germline sequence In some embodiments, the amino acid sequence contains three substitutions at residues 89-94 compared to: In the library, each antibody light chain variable region is one or more of the above L3-VL In some embodiments, such libraries comprise antibody light chain variable regions. It may or may not encode such an L3-VL sequence. It may or may not be combined with one or more sets of other nucleic acids. In some embodiments, the present invention provides antibodies having an amino acid sequence set forth in Table 4. A polynucleotide encoding a light chain variable region or one or more of the polynucleotides in Tables 5 to 7 A library containing the polynucleotide sequence described herein, comprising two sequences from positions 89 to 94 Or three residues are different.
[0020] In some embodiments, the present invention provides a polynucleotide encoding an antibody light chain variable region. and a library containing all the coding antibody light chain variable The regions are identical to each other except for the substitution of residues at positions 89 to 94. Across the library, the sequences of any two coded antibody light chain variable regions are aligned at no more than three positions. are different from each other.
[0021] In some embodiments, the present invention provides a method for the preparation of a nucleic acid sequence comprising administering to a subject a nucleic acid sequence selected from the group consisting of two or more polynucleotides provided in Tables 5-7. at least about 80%, 90%, or or 95% identical polypeptide sequence. In one embodiment, all members of the library are provided. is produced by translation of two or more polynucleotide sequences provided in Tables 5-7. The polypeptide is at least about 80%, 90%, or 95% identical to the polypeptide of interest.
[0022] In one embodiment, the present invention provides translations of the polynucleotide sequences provided in Tables 5-7. A library containing light chain variable regions containing polypeptides produced by translation is provided. In one embodiment, all members of the library have the sequences provided in Tables 5-7. It includes polypeptides produced by translation of a nucleotide sequence.
[0023] In some embodiments, a nucleic acid encoding a nucleic acid sequence containing or comprising CDRL3 and / or a light chain variable region. Any of the libraries described herein that load the complete light chain associated Further, some embodiments contain a CDRL3 and / or light chain variable region such as So, such a library (and / or a complete light chain library) can be one or Further containing or encoding multiple heavy chain CDRH3s, variable domains, or complete heavy chains. In some embodiments, the libraries provided are comprised of antibodies, such as whole IgG. The antibody may comprise or encode a complete antibody such as
[0024] In some embodiments, the libraries provided comprise human antibodies or antibody fragments. or encodes, and in some such embodiments, the provided library is The antibody comprises or encodes a human antibody.
[0025] In one embodiment, the present invention provides a method for the preparation of a nucleic acid library comprising the library nucleic acids described herein. In many embodiments, such libraries include nucleic acid vectors. Each rally member contains the same vector.
[0026] In some embodiments, the present invention provides one or more provided, including, for example, vectors. In some embodiments, the host cell contains the provided library. The cell is a yeast, and in one embodiment, the yeast is Saccharomyces cerevisiae )
[0027] In some embodiments, the present invention provides a method for producing a medicament solely from the libraries described herein. The antibody to be isolated is provided.
[0028] In one embodiment, the present invention provides a method for producing any of the libraries described herein. A kit containing the same is provided.
[0029] In some embodiments, the present invention provides a library in computer readable form. Li and / or theoretical segment pools, e.g., Tables 10, 23-25, and 26 TN1 polypeptides of Tables 11, 23-25, and 28; DH polypeptides of Table 12, 23-25, and 30 N2 polypeptides; Tables 13, 15, 17, 23-25, and 32 H3-JH polypeptides; TN1 polynucleotides of Tables 25-26; Table 25 and 27 DH polynucleotides; N2 polynucleotides of Tables 25 and 29; and / or or a representative of the H3-JH polynucleotides of Tables 25 and 31 ion).
[0030] In one embodiment, the present invention provides a surrogate (or a complement) of polynucleotide sequences of a human pre-immune set. and providing a computer readable form of the polypeptide expression product thereof. do.
[0031] In some embodiments, the present invention provides a synthetic polynucleotide encoding CDRH3 library. 1. A method for producing a nucleotide comprising: (a) selecting a TN1, a DH, an N2, and an H3-J (b) providing a theoretical segment pool containing H segments; (c) providing a reference set of sequences; and (b) each C Theoretical segment analysis of (a) to identify the closest match(s) to the DRH3 sequence. (d) utilizing the theoretical segmentation rules for inclusion in a synthetic library; (e) selecting segments from the top pool; and (f) a synthetic CDRH3 library. In one embodiment, the present invention provides a method for synthesizing a compound comprising: In some embodiments, a synthetic library is provided. The segments selected for inclusion in the reference set of CDRH3 sequences are They are selected according to their segment usage weights.
[0032] In one embodiment, the present invention provides a synthetic polynucleotide encoding a CDRL3 library. A method for generating a peptide comprising the steps of: (i) obtaining a reference set of light chain sequences; Therefore, the reference set is composed of individuals with the same IGVL germline gene and / or its allelic variants. (ii) containing a light chain sequence having a VL segment derived from the IGVL gene; Which amino acids are present at each of the CDRL3 positions in the reference set encoded by (iii) determining whether a light chain variable domain coding sequence is present; The two positions 89 to 94 correspond to the corresponding positions in the reference set. These steps contain degenerate codons that encode the five most frequently occurring amino acid residues. and (iv) a step of synthesizing polynucleotides encoding the CDRL3 library. In one embodiment, the present invention provides a method for producing a medicament for the treatment of a cancer, comprising the steps of: Provide a library where
[0033] In some embodiments, the present invention provides a method for isolating antibodies that bind to an antigen using a laser of the present invention. a method for using any of the libraries, comprising: contacting the expression product with an antigen and isolating a polypeptide expression product that binds to the antigen. The present invention provides a method comprising the steps of:
[0034] In one embodiment, the N-linked glycosylation sites in the libraries of the invention, deamidation The number of cysteine motifs and / or Cys residues in the repertoire derived from biological sources Decreased or reduced compared to libraries produced by amplification.
[0035] The present invention relates to many polynucleotide and polypeptide sequences, as well as to larger polynucleotides. Segments that can be used to construct nucleotide and polypeptide sequences ( For example, TN1, DH, N2, and and H3-JH segments). Those skilled in the art will recognize that in some instances, these sequences , providing a consensus sequence after alignment of the sequences provided by the present invention. These consensus sequences are within the scope of the present invention. and any of the sequences provided herein may be used for more concise presentation. You will easily recognize that. In certain embodiments, for example, the following are provided: (Item 1) Structure: Contains a CDRH3 sequence with [TN1]-[DH]-[N2]-[H3-JH]. At least about 10 encoding a polypeptide containing 4 Synthetic polynucleotides containing the polynucleotide A library of leotids, TN1 is a polypeptide corresponding to any of the TN1 polypeptides in Tables 9-10 and 18-26. peptide, or produced by translation of any of the TN1 polynucleotides in Tables 25-26. is a polypeptide that is DH corresponds to any of the DH polypeptides in Tables 9, 11, 17-25, and 28. polypeptide, or any of the DH-encoding polynucleotides of Tables 16, 25, and 27 a polypeptide produced by translation of N2 corresponds to any of the N2 polypeptides in Tables 9, 12, 18-25, and 30. A translation of any of the polypeptides or N2-encoding polynucleotides of Tables 25 and 29. and H3-JH is the H3-JH polypeptide of Tables 9, 13, 15, 18-25, and 32. A polypeptide corresponding to any of the H3-JH coding polypeptides in Tables 14, 25, and 31 is also included. a library of polypeptides produced by translation of any of the oligonucleotides . (Item 2) At least about 1%, 5%, or 10% of the sequences in the library are provided 2. The library according to item 1, having a structure (Item 3) The polynucleotide is selected from the group consisting of TN1, D, and TN2, D, and D, provided in any one of Tables 23 to 25. CDRH3 polypeptides produced by a set of H, N2, and H3-JH polypeptides 2. The library according to item 1, wherein the library encodes a peptide. (Item 4) The polynucleotide is a set of TN1 polypeptides provided in Table 26, Table 27 the set of DH polypeptides provided in Table 28, the set of N2 polypeptides provided in Table 30 The set of peptides and the set of H3-JH polypeptides provided in Table 32 2. The library of item 1, encoding a CDRH3 polypeptide produced thereby. (Item 5) A method of using the library according to item 1 to isolate antibodies that bind to an antigen. contacting the polypeptide expression products of the library with an antigen; isolating a polypeptide expression product that binds to the antigen. (Item 6) The number of N-linked glycosylation sites, deamidation motifs, and / or Cys residues Low compared to libraries produced by amplification of repertoires from biological sources 2. The library according to item 1, wherein the library is reduced or reduced in size. (Item 7) and a polynucleotide encoding one or more light chain variable domain polypeptides. Item 1. The library according to item 1, comprising: (Item 8) 8. The library of item 7, wherein the polypeptides are expressed as full-length IgG. (Item 9) 9. Polypeptide expression products of the library according to item 8. (Item 10) 10. An antibody isolated from the polypeptide expression product of item 9. (Item 11) A vector containing the library described in Item 1. (Item 12) A host cell containing the vector according to item 11. (Item 13) 13. The host cell according to item 12, wherein the host cell is a yeast cell. (Item 14) Item 14. The yeast cell according to item 13, wherein the yeast is a budding yeast. (Item 15) A kit containing the library according to item 1. (Item 16) A representation of the library of item 1 in computer readable form. (Item 17) Structure: Contains a CDRH3 sequence with [TN1]-[DH]-[N2]-[H3-JH]. At least about 10 encoding a polypeptide containing 4 Synthetic polynucleotides containing the polynucleotide A library of leotids, TN1 may be at least about 5'-6' with any of the TN1 polypeptides of Tables 9-10 and 18-26. A polypeptide that is 80%, 90%, or 95% identical to a TN1 polypeptide of Tables 25-26 at least about 80% of the polypeptide produced by translation of any of the nucleotides , 90%, or 95% identical polypeptides to the sequence DH is at least one of the DH polypeptides in Tables 9, 11, 17-25, and 28. or a polypeptide that is about 80%, 90%, or 95% identical to a polypeptide of Tables 16, 25, and and a polypeptide produced by translation of any of 27 DH-encoding polynucleotides. a polypeptide that is at least about 80%, 90%, or 95% identical to N2 may be at least one of the N2 polypeptides in Tables 9, 12, 18-25, and 30. or a polypeptide of Tables 25 and 29 that is about 80%, 90%, or 95% identical and at least one polypeptide produced by translation of any of the N2-encoding polynucleotides. and H3-JH is the H3-JH polypeptide of Tables 9, 13, 15, 18-25, and 32. a polypeptide that is at least about 80%, 90%, or 95% identical to any of by translation of any of the H3-JH-encoding polynucleotides in Tables 14, 25, and 31 A polypeptide that is at least about 80%, 90%, or 95% identical to a polypeptide produced by A library of peptides. (Item 18) A library of synthetic polynucleotides encoding light chain variable regions, comprising: The light chain variable region (a) VK1-05 sequences differing at one or more positions 4, 49, and 46; (b) VK1-12 sequences that differ at one or more positions 4, 49, 46, and 66; (c) VK1-33 sequences that differ at one or more positions 4, 49, and 66; (d) VK1-39 sequences that differ at one or more positions 4, 49, and 46; (e) VK2-28 sequences that differ at one or more positions 2, 4, 46, and 49; (f) VK3-11 sequences that differ at one or more positions 2, 4, 36, and 49; (g) VK3-15 sequences that differ at one or more positions 2, 4, 48, and 49; (h) VK3-20 sequences that differ at one or more positions 2, 4, 48, and 49; Ravini (i) from VK4-1 sequences that differ at one or more of positions 4, 46, 49, and 66; A library selected from the group consisting of: (Item 19) The library comprises two or more light chain polypeptide sequences provided in Table 3 and at least one light chain polypeptide sequence. light chain variable regions comprising polypeptide sequences that are at least about 80%, 90%, or 95% identical to each other; 19. The library according to item 18, comprising polynucleotides encoding: (Item 20) Item 18, wherein the light chain variable region comprises a polypeptide sequence provided in Table 3. Library of content. (Item 21) A library of synthetic polynucleotides encoding light chain variable regions, The polypeptide sequence of the variable region is either one of positions 89 to 94 of the variable light chain polypeptide sequence or Libraries differing in three residues. (Item 22) The library comprises two or more polynucleotide sequences provided in Tables 5-7. at least about 80%, 90%, or 95% identical to the polypeptide produced by translation a polynucleotide encoding a light chain variable region comprising a polypeptide sequence 21. The library described in 21. (Item 23) The light chain variable region is obtained by translation of the polynucleotide sequences provided in Tables 5-7. 22. The library according to Item 21, comprising polypeptides produced by (Item 24) A machine-readable representation of any of the following: TN1 polypeptides of Tables 10, 23-25, and 26; DH polypeptides of Tables 11, 23-25, and 28; N2 polypeptides of Tables 12, 23-25, and 30; H3-JH polypeptides of Tables 13, 15, 17, 23-25, and 32; TN1 polynucleotides in Tables 25-26; DH polynucleotides of Tables 25 and 27; N2 polynucleotides of Tables 25 and 29; and H3-JH polynucleotides of Tables 25 and 31. (Item 25) Polynucleotide sequence surrogates of the human pre-immune set (Appendix A) or computer readable and the polypeptide expression product in a digestible form. (Item 26) A library of synthetic polynucleotides encoding polypeptides containing CDRH3 sequences. 1. A method for making a (a) Theoretical segment profiles containing TN1, DH, N2, and H3-JH segments. providing a rule; (b) providing a reference set of CDRH3 sequences; (c) The closest match(s) to each CDRH3 sequence in the reference set in (b). utilizing the theoretical segment pool of (a) to identify (d) selecting segments from the theoretical segment pool for inclusion in a synthetic library; selecting step; and (e) synthesizing a synthetic CDRH3 library. (Item 27) A library of polynucleotides produced according to the method of item 26. (Item 28) The segments selected for inclusion in the synthetic library are Item 26, which is selected according to their segment usage weights in the reference set of The method described. (Item 29) The segments selected for inclusion in the synthetic library may comprise one or more 27. The method according to item 26, wherein the compound is selected according to its physicochemical properties. (Item 30) wherein the reference set of CDRH3 sequences is a reference set of pre-immune CDRH3 sequences. Item 27. The method according to item 26. (Item 31) Additional TN1 and TN2 fragments present in the reference set but not in the theoretical segment pool 27. The method of claim 26, further comprising the step of selecting N2 segments. (Item 32) 26. The method according to claim 26, wherein stop codons are reduced or eliminated from the library. The method described. (Item 33) Unpaired Cys residues, N-linked glycosylation motifs, and deamidation motifs Item 27. The method according to item 26, wherein the translation products of the library are reduced or eliminated. How to post. (Item 34) The D H segment and the H3-J H segment are CDRH 27. The method of item 26, wherein the fragment is sequentially cleaved before matching with three sequences. (Item 35) DH and N2 or a combination thereof 27. The method of item 26, further comprising the step of introducing one or two degenerate codons. (Item 36) Item 2, further comprising the step of introducing one degenerate codon into the H3-JH segment. 6. The method according to claim 6. (Item 37) 27. Polypeptide expression products of the library according to item 26. (Item 38) 38. An antibody isolated from the polypeptide expression product of item 37. (Item 39) A method for generating synthetic polynucleotides encoding a CDRL3 library. So, (i) obtaining a reference set of light chain sequences, the reference set comprising sequences from the same IGVL genome; Light chains having VL segments derived from lineage genes and / or allelic variants thereof containing the sequence; (ii) that of the CDRL3 position in the reference set encoded by the IGVL gene determining which amino acids are present in each; (iii) synthesizing a light chain variable domain coding sequence, Two or three positions have two or more of the five most common positions in the corresponding position in the reference set. including degenerate codons encoding frequently occurring amino acid residues; and (iv) synthesizing polynucleotides encoding the CDRL3 library. A method comprising: (Item 40) A library of polynucleotides produced according to the method of item 39. (Item 41) A library of synthetic polynucleotides encoding polypeptides, comprising: one or more VH chassis containing Kabat residues 1-94 of the human IGHV sequence; one or more TN1 segments selected from a reference set of human CDRH3 sequences; Selected from a theoretical pool of segments matched to a reference set of human CDRH3 sequences one or more DH segments; one or more N2 segments selected from a reference set of human CDRH3 sequences; and Beauty Selected from a theoretical pool of segments matched to a reference set of human CDRH3 sequences A library comprising one or more H3-JH segments. [Brief explanation of the drawings]
[0036] [Figure 1] Vernier residues 4 and 49 (asterisks) in VK1-39 are shown to have a diversity index that is equal to or exceeds the diversity index of the CDR positions (ie, 0.07 or greater in this example). [Figure 2] Clinically validated CDRL3 sequences show little deviation from the germline-like sequence (n=35). [Figure 3] Figure 1 shows the percent of sequences in the jumping dimer CDRL3 library of the present invention and the previous CDRL3 library, VK-v1.0, that have X or fewer mutations from germline, where FX is the fraction of sequences in the library that have X or fewer mutations from germline. [Figure 4] 1 shows the application of the provided methods used to generate the nucleotide sequences encoding the parent H3-JH segments (SEQ ID NOS: 8748-8759, respectively, in order of appearance). [Figure 5] FIG. 1 shows a schematic representation of the approach used to select segments from a theoretical segment pool for inclusion in theoretical and / or synthetic libraries. [Figure 6]1 shows the frequency of "good" and "poor" expressing CDRH3 sequences isolated from the yeast-based library described in U.S. Patent Application Publication No. 2009 / 0181855 as a function of DH segment hydrophobicity (increasing to the right) and their comparison to sequences contained in the library design described therein ("Design"). [Figure 7] The percentage of CDRH3 sequences in the LUA-141 library and Exemplary Library Design 3 (ELD-3) that match CDRH3 sequences from Lee-666 and Boyd-3000 with 0, 1, 2, 3, or more than 3 amino acid mismatches is shown. [Figure 8] Both Exemplary Library Design 3 (ELD-3) and Extended Diversity Library Design are shown to yield better matches to clinically relevant CDRH3 sequences than the LUA-141 library. [Figure 9] We demonstrate that the combinatorial efficiency of Exemplary Library Design 3 (ELD-3) exceeds that of the LUA-141 library. In particular, ELD-3 segments are more likely to yield unique CDRH3s than LUA-141 library segments. [Figure 10] The amino acid compositions of Kabat-CDRH3 of LUA-141, Exemplary Library Design 3 (ELD-3), and the human CDRH3 sequence (Human H3) derived from HPS are shown. [Figure 11] The Kabat-CDRH3 length distributions of human CDRH3 sequences (Human H3) derived from LUA-141, Exemplary Library Design 3 (ELD-3), and HPS are shown. [Figure 12] The percentage of CDRH3 sequences in the Extended Diversity library that match the CDRH3 sequences from Boyd et al. with 0 to 32 amino acid mismatches is shown. [Figure 13]1 shows the Kabat-CDRH3 length distributions for Exemplary Library Design 3 ("ELD-3"), Extended Diversity Library Design ("Extended Diversity"), and human CDRH3 sequences from the Boyd et al. dataset ("Boyd 2009"). [Figure 14] 1 shows the amino acid composition of Kabat-CDRH3 for the human CDRH3 sequence from the Extended Diversity Library Design ("Extended Diversity") and the Boyd et al. dataset ("Boyd 2009"). [Figure 15] The combinatorial efficiency of the Extended Diversity Library Design is shown by matching 20,000 randomly selected sequences from the same design. Approximately 65% of the sequences appear only once in the design, and approximately 17% appear twice. DETAILED DESCRIPTION OF THE INVENTION
[0037] The present invention relates to, inter alia, polynucleotide and polypeptide libraries. , methods for producing and using libraries, kits containing libraries, and the libraries and / or theoretical segment pools disclosed herein. This application provides a computer-readable representation of the method taught in this application. Libraries are constructed from components (e.g., polynucleotides or polypeptides) from which they are constructed. It can be described, at least in part, in terms of peptide "segments." The present invention is particularly directed to these polynucleotide or polypeptide segments, such Methods for producing and using such library segments, and and providing a computer readable form of the kit and surrogate containing the segment. do.
[0038] In one embodiment, the present invention provides a method for identifying sequences and sequences in the naturally occurring human antibody repertoire. The present invention provides an antibody library specifically designed based on the distribution of antigenic stimuli and CDR lengths. Even in the absence of stimulation, individual humans have at least about 10 7 Create different antibody molecules It is estimated that the enzyme is produced by the tional Medicine, 2009, 1: 1). The antigen binding properties of many antibodies The site can cross-react with a variety of related but distinct epitopes. The antibody repertoire contains almost every possible antibody, albeit potentially low affinity. enough to ensure that there is an antigen-binding site that matches the desired epitope. big.
[0039] The mammalian immune system controls chromosomally separated gene segments prior to transcription. By linking it to the binatorial system, it can be used to generate almost infinite amounts of They have evolved unique genetic mechanisms that allow them to generate distinct light and heavy chains. Each type of immunoglobulin (Ig) chain (i.e., kappa light chain, lambda light chain) , and heavy chains) are composed of two or more gene segments to produce a single polypeptide chain. The DNA sequences are synthesized by combinatorial assembly of DNA sequences selected from a family of In addition, the heavy and light chains each consist of a variable region and a constant (C) region. The region is encoded by DNA sequences that are constructed from three families of gene sequences: Variable (IGHV), Divergent (IGHD), and Joining (IGHJ). The variable region of the light chain is Constructed from two families of gene sequences for the α and λ light chains, respectively. They are encoded by DNA sequences: variable (IGLV) and joining (IGLJ), respectively. These variable regions (heavy and light) also contain constant regions to produce full-length immunoglobulin chains. It can be recombined with.
[0040] Combinatorial assembly of V, D, and J gene segments contributes to antibody variable region diversity Although these gene segments make substantial contributions to the genome, further diversity is due to the imprecise sequencing of these gene segments. Pre-ligation occurs through the introduction of non-templated nucleotides at the junctions between ligated and gene segments. It is introduced in vivo at the B cell stage (more specifically, for example, in its entirety). See U.S. Patent Application Publication No. 2009 / 0181855, which is incorporated by reference. sea bream).
[0041] After a B cell recognizes an antigen, it is induced to proliferate. During proliferation, B cells The receptor locus undergoes an extremely high rate of somatic mutation, far exceeding the normal rate of genomic mutation. The resulting mutations are primarily localized in the Ig variable region and include substitutions, insertions, and This somatic hypermutation results in the development of B cells that express antibodies with increased affinity for antigens. Such antigen-driven somatic hypermutation allows the production of antibody responses to a given antigen. Fine-tune the.
[0042] The synthetic antibody library of the present invention recognizes any antigen, including antigens of human origin. It is possible that autoreactive antibodies are generated by the donor's immune system through negative selection. Therefore, the ability to recognize antigens of human origin is lost from human biological origin (even Other antibody libraries, such as those prepared from human cDNA, It cannot exist in this world.
[0043] Additionally, the present invention simplifies certain aspects of library development and / or screening. For example, in some embodiments, The present invention utilizes cell sorting techniques (e.g., fluorescent activity) to identify positive clones. This allows the use of a fast cell sorter (FACS) to sort the hybridoma library. This avoids or eliminates the need for standard, lengthy methods of generating and screening supernatants.
[0044] Additionally, in some embodiments, the present invention provides a laser that provides multiple screening passes. For example, some embodiments provide libraries and / or sub-libraries. In this case, the provided libraries and / or sub-libraries may be screened multiple times. In some such embodiments, each provided library Libraries and / or sub-libraries are used to discover additional antibodies against many targets. It can be used for
[0045] Before further describing the present invention, some terms will be defined.
[0046] definition Unless otherwise defined, all technical and scientific terms used herein are All terms and phrases have the meaning commonly understood by one of ordinary skill in the art. The numbering system of 1 to 10 is used throughout the application. The following definitions are within the skill in the art. This supplements the terminology in the specification and relates to the embodiments described in this application. be.
[0047] The term "amino acid" or "amino acid residue" refers to a typical amino acid, as understood by those skilled in the art. Specifically, alanine (Ala or A); arginine (Arg or R); asparagine (A sn or N); aspartic acid (Asp or D); cysteine (Cys or C); Glutamine (Gln or Q); glutamic acid (Glu or E); glycine (Gly or or G); histidine (His or H); isoleucine (Ile or I): leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine ( Tyr or Y); and valine (Val or V). and the like, but also refers to amino acids having their art-recognized definition, such as modified amino acids, Synthetic or rare amino acids may be used as desired. Amino acids with nonpolar side chains (e.g., Ala, Cys, Ile, Leu, Met, Phe, Pro, Val); negatively charged side chains (e.g., Asp, Glu); positively charged side chains (e.g., Ar g, His, Lys); or uncharged polar side chains (e.g., Asn, Cys, Gln, Gl y, His, Met, Phe, Ser, Thr, Trp, and Tyr) They can be classified as follows.
[0048] As will be understood by those skilled in the art, the term "antibody" is used herein in its broadest sense. In particular, at least one monoclonal antibody, polyclonal antibody, polyspecific antibody, antibodies (e.g., bispecific antibodies), chimeric antibodies, humanized antibodies, human antibodies, and antibody fragments. Antibodies are expressed by immunoglobulin genes or fragments of immunoglobulin genes. A protein containing one or more polypeptides that are qualitatively or partially encoded by The recognized immunoglobulin genes are kappa, lambda, alpha, gamma, and delta , epsilon, and mu constant region genes and numerous immunoglobulin variable region genes Includes genes.
[0049] The term "antibody binding region" refers to an immunoglobulin or antibody variable region capable of binding to an antigen. Typically, an antibody binding region refers to one or more portions of an antibody binding region, such as an antibody light chain ( or variable region or one or more CDRs thereof), antibody heavy chain (or variable region or or one or more CDRs thereof), heavy chain Fd region, Fab, F(ab')2, single domain or combined antibody light and heavy chains, such as single chain antibodies (scFv) (or its variable region), or a full-length antibody that recognizes the antigen, such as an IgG (e.g. IgG1, IgG2, IgG3, or IgG4 subtype), IgA1, IgA2, Any region of an IgD, IgE, or IgM antibody.
[0050] An "antibody fragment" is a portion of an intact antibody, e.g., one or more of its antigen-binding regions. Examples of antibody fragments include intact antibodies and Fab, F, and F-antibodies formed from antibody fragments. ab', F(ab')2, and Fv fragments, bispecific antibodies, linear antibodies, single-chain antibodies, etc. and polyspecific antibodies.
[0051] The term "antibody of interest" refers to an antibody identified and / or isolated from the library of the present invention. refers to an antibody having a property of interest. Exemplary properties of interest include the ability to bind to a particular antigen or epitope. binding to, binding with a certain affinity, cross-reactivity, blocking of the binding interaction between two molecules These include, but are not limited to, the induction of a biological effect, and / or the elicitation of a biological effect.
[0052] The term "canonical structure" refers to a structure that includes antigen-binding (CDR) sequences, as understood by those skilled in the art. This refers to the main chain conformation adopted by the loops. Comparative structural studies have identified six antigen-binding loops. Five of the groups were found to have only a limited repertoire of available conformations. Each canonical structure can be characterized by the torsion angles of the polypeptide backbone. The corresponding loops between antibodies are therefore highly amino- Page 11 ... Despite their amino acid sequence diversity, they can have very similar three-dimensional structures (respectively, their Chothia and Lesk, J. Mol., incorporated by reference in its entirety. Biol., 1987, 196: 901; Chothia et al., Nature, 1989, 342: 877; Martin and Thorn ton, J. Mol. Biol., 1996, 263: 800). There is a relationship between the loop structure adopted and the amino acid sequence surrounding it. As is known, the conformation of a particular canonical class varies depending on the length of the loop as well as the loop size. and exist in key positions within the loop and within the storage framework (i.e., outside the loop). The assignment to a particular canonical class is therefore determined by the amino acid sequence and amino acid residue. The prediction can be made based on the presence of these important amino acid residues. The term "canonical structure" also refers to the linear sequence of an antibody, as classified by Kabat. (Kabat et al., in "Sequence" ces of Proteins of Immunological Interes t,” 5 th Edition, US Department of Heat (Human Services, 1992). Kabat numbering The method is widely used to number the amino acid residues of antibody variable domains in a consistent manner. These are widely accepted standards and will be used herein unless otherwise indicated. Consideration of other structures can also be used to determine the canonical structure of an antibody. For example, differences not fully reflected by Kabat numbering are and / or can be explained by the numbering system of ia et al. revealed by other techniques, such as crystallography and two- or three-dimensional computational modeling Thus, a given antibody sequence can be, among other things, tailored to a suitable chassis sequence ( canonical classes that allow identification of the chassis sequence. canonical structures in a library. (Based on the request) Kabat numbering and Chothia e of antibody amino acid sequences Structural considerations as described by et al. as well as canonical aspects of antibody structure. Their relevance for interpretation is described in the literature.
[0053] The term "CDR" and its plural "CDRs" refer to the three CDRs of a light chain variable region (CDR The heavy chain variable region (CDRs) comprise the binding characteristics of the heavy chain variable region (CDRs 1, 2, and 3). Complementarity-determining regions (CDRs) that make up the binding characteristics of H1, CDRH2, and CDRH3 CDRs refer to amino acids that contribute to the functional activity of an antibody molecule and include the framework regions. The exact boundaries and lengths of CDRs are defined in various classifications and Therefore, the CDRs may be numbered, for example, according to the CDRH numbering system described below. 3 numbering systems, including Kabat, Chothia, contact, or other boundary definitions. Despite the different boundaries, Each of the formulas has some degree of variation in what constitutes the so-called "hypervariable regions" within the variable domain. Therefore, the CDR definitions in these methods are based on adjacent frameworks. The regions may have different boundary areas of different lengths. Kabat et al., in “Sequen ces of Proteins of Immunological Interes t,” 5 th Edition, US Department of Heal th and Human Services, 1992; al., J. Mol. Biol., 1987, 196: 901; and Ma cCallum et al., J. Mol. Biol., 1996, 262 : 732.
[0054] As used herein, the "CDRH3 numbering system" refers to the CDRH3 numbering system at position 95. Define the first amino acid of CDRH3 and the last amino acid of CDRH3 as position 102. It should be noted that this is a custom numbering system that does not follow Kabat. The amino acid segment beginning at position 95 is designated "TN1" and, if present, is numbered 95, 96, 96A, 96B, etc. The nomenclature used in this application is , U.S. Patent Application Publication Nos. 2009 / 0181855 and 2010 / 0056386 and slightly different from that used in International Publication No. WO / 2009 / 036379 Note that in those applications, position 95 is referred to as the "Tail" residue. Here, Tail(T) is combined with N1 segment and becomes "TN1". The TN1 segment is followed by a "DH" segment. Successively, this is assigned the numbers 97, 97A, 97B, 97C, etc. followed by an "N2" segment, which, if present, is 98, 98A, 98B, etc. Finally, the most C-terminal amino acid residue of the "H3-JH" segment is , designated as number 102. The residue immediately preceding it (at the N-terminus), if present, is designated as 101. The previous one (if present) is 100. The remaining H3-JH amino acids are in the reverse order. , starting with 99 for the immediately N-terminal amino acid to 100, The N-terminal residues of 99 are 99A, 99B, 99C, etc., and the same applies to others. Thus, examples of CDRH3 sequence residue numbers may include: [ka]
[0055] The "chassis" of the present invention is the a portion of an antibody heavy chain variable (IGHV) or light chain variable (IGLV) domain, but not The chassis of the present invention begins with the first amino acid of FRM1 and ends with the last amino acid of FRM3. It is defined as the portion of the variable region of an antibody that ends with the amino acid contains amino acids from position 1 to position 94 inclusive. For light chains (kappa and lambda), The chassis is defined as including positions 1 to 88. The chassis of the present invention is They may contain certain modifications compared to the germline variable domain sequence. may be engineered (e.g., to remove N-linked glycosylation sites) or may be naturally occurring may occur (e.g., to give rise to naturally occurring allelic variations), e.g. It is known in the art that the immunoglobulin gene repertoire is polymorphic. (Wang et al., Imm, each of which is incorporated by reference in its entirety). unol. Cell. Biol., 2008, 86: 111; Collin s et al., Immunogenetics, 2008, 60: 669) the chassis, CDRs, and constant regions representing these allelic variants are also included in the present invention. In some embodiments, the invention provides a method for the preparation of a medicament for use in a particular embodiment of the invention. The allelic variant(s) to be tested may be, for example, non-immunogenic in these patient populations. To identify antibodies, selection was based on allelic variations present in various patient populations. In certain embodiments, the immunogenicity of the antibodies of the invention may be selected based on the major tissues of a patient population. It may depend on allelic variation in the MHC genes. Allelic variation may also be considered in the design of the libraries of the present invention. In one embodiment, the chassis and constant region are contained in a vector, and the CDR3 region is introduced between them via homologous recombination.
[0056] As used herein, "directed divers" refers to a group of diversified organisms. The sequences designed by "ity" contain both sequence diversity and length diversity. The specified diversity is not probabilistic.
[0057] As used herein, the term "diversity" refers to a wide variety or significant heterogeneity. The term "sequence diversity" refers to the diversity of sequences found in, for example, naturally occurring human antibodies. For example, the CDRH3 sequence Diversity was achieved by combining known human TN1, DH, N2, and The CDRL3 sequence may refer to various possibilities for combining the H3-JH segment and the H3-JH segment. Diversity (kappa or lambda) is determined by the naturally occurring light chains to form the CDRL3 sequence. The CDRL3 (i.e., "L3-VL") and linking (i.e., "L3- The term "combination" may refer to various possibilities for combining "combination" ("combination") segments. As used herein, "H3-JH" refers to the portion of the IGHJ gene that contributes to CDRH3. As used herein, "L3-VL" and "L3-JL" refer to , the portions of the IGLV and IGLJ genes (kappa or lambda) that contribute to CDRL3 Point.
[0058] As used herein, the term "expression" refers to any combination of transcription, post-transcriptional modification, translation, and transcription. The steps involved in the production of a polypeptide, including, but not limited to, post-modification, and secretion Point to Tep.
[0059] The term "framework region" refers to regions that lie between the more divergent (i.e., hypervariable) CDRs. refers to the art-recognized portions of antibody variable regions. The region typically comprises frameworks 1-4 (FRM1, FRM2, FRM3, and FRM4). 4), and the six CDRs (heavy chain Scaffolds for presentation of the IgG1-derived IgG1 and IgG2 ... provide.
[0060] The term "full-length heavy chain" refers to a chain that includes four framework regions, three CDRs, and a constant region. immunoglobulins containing each of the canonical structural domains of immunoglobulin heavy chains, This refers to the heavy chain.
[0061] The term "full-length light chain" refers to a light chain that comprises four framework regions, three CDRs, and a constant region. immunoglobulins containing each of the canonical structural domains of immunoglobulin light chains, Refers to the phosphodiesterase light chain.
[0062] The term "germline-like" when used in reference to the CDRL3 sequences of the light chains of the present invention means that the CDRL3 sequences of the light chains of the present invention are similar to those of the CDRL3 sequences of the light chains of the present invention. The first six wild-type residues (i.e., positions 89-94 in the Kabat numbering system; "L" stands for kappa or lam and (ii) long chains derived mostly, but not exclusively, from the JL segment. One of several amino acid sequences of length 2, 1 to 4 amino acids ("L" also refers to It means a sequence of the most common lengths (i.e., For kappa CDRL3 sequences (8, 9, and 10 residues), the sequence in (ii) is 0, FT, LT, IT, RT, WT, YT, [X]T, [X]PT, [X]FT, [X]LT, [X]IT, [X]RT, [X]WT, [X]YT, [X]PFT, [X] PLT, [X]PIT, [X]PRT, [X]PWT, and [X]PYT, and [X ] is the amino acid sequence found at position 95 (Kabat) in each VK germline sequence. X corresponds to a carboxylic acid residue. X is most commonly P, but can also be S or VK germline sequences. It may also be any other amino acid residue found at position 95 of the sequence. For the eight exemplified VK chassis, the corresponding 160 germline-like sequences are sequence (i.e., combined with positions 89-94 of each of the eight VK germline sequences, 20 sequences of 2 to 4 amino acids in length are provided in Table 1. Similar approaches is applied to define germline-like CDRL3 sequences for lambda light chains. The kappa sequences described are encoded by the IGVL gene (in this case IGVλ). The complete non-mutated portion of CDRL3 that is observed is predominantly, but not exclusively, the Jλ segment. The latter sequence (see (ii) above) will be combined with a sequence derived from the original The corresponding values are YV, VV, WV, AV, or V. As described in Example 7 of Japanese Patent Application Publication No. 2009 / 0818155, The resulting "germline-like" sequences are still considered, while partial codons are considered. and mutations at the last position of the Vλ-encoded portion of CDRL3. Further considerations can be taken into account. More specifically, U.S. Patent Application Publication No. 2009 / 08 Example 7 of the "Minimalist Library" in 18155 The entire "germline-like" sequence would be defined as "germline-like." Those skilled in the art will recognize that other VK and Vλ sequences It will be readily apparent that these methods can be extended.
[0063] The term "genotype-phenotype linkage" is understood by those skilled in the art to refer to a specific expression. The nucleic acid (genotype) encoding the protein having the genotype (e.g., binds to an antigen) is labeled. For illustrative purposes, the term "phage" refers to the fact that it can be isolated from a library. Antibody fragments expressed on the surface of the antibody can be isolated based on their binding to the antigen ( (See, e.g., U.S. Patent No. 5,837,500.) Binding of an antibody to an antigen simultaneously results in the formation of an antibody fragment. This allows the isolation of phages containing nucleic acids encoding the antibody. The antigen-binding characteristics of the antibody fragment are "linked" to the genotype (the nucleic acid encoding the antibody fragment). Another method for maintaining genotype-phenotype linkage is described by Wittrup et al. (U.S. Patent No. 6,300,065, each of which is incorporated by reference in its entirety. No. 6,331,391, No. 6,423,538, No. 6,696,251, No. 6,6 99,658, and U.S. Patent Application Publication No. 20040146976), Milten yi (U.S. Patent No. 7,166,423, which is incorporated by reference in its entirety), Fan dl (U.S. Patent No. 6,919,183, each of which is incorporated by reference in its entirety) , U.S. Patent Application Publication No. 20060234311), Clausell-Tormos et al. (Chem. Biol., 200 8, 15: 427), Love et al. (incorporated by reference in its entirety) Nat. Biotechnol., 2006, 24: 703), and Ke lly et al. (Chem. Commun., incorporated by reference in its entirety) ., 2007, 14: 1773). The terms are used in a way that maintains a connection between them. The antibody proteins are encoded in a manner that allows both to be recovered while the antibody is being transferred. It can be used to refer to any method of locating an antibody protein along with its associated gene. This can be done.
[0064] The term "heterologous moiety" is used herein to refer to the addition of a moiety to an antibody, The heterologous moiety is not a part of a naturally occurring antibody. Exemplary heterologous moieties include drugs, toxins, imaging agents, and any other composition that provides activity not essential to the antibody itself.
[0065] As used herein, the term "host cell" refers to a cell in which a polynucleotide of the invention is expressed. Such terms are intended to refer to cells containing a particular target cell, as well as It should be understood that the term also refers to the progeny or potential progeny of such cells. , may arise in subsequent generations due to mutations or environmental influences, so that such offspring In fact, the cell may not be identical to the parent cell, but is within the scope of the term as used herein. Included.
[0066] As used herein, the term "human antibody CDR library" refers to a library of human antibodies. The sequences are designed to represent the sequence and length diversity of naturally occurring CDRs in the human genome. at least one polynucleotide or polypeptide library (For example, the term "CDR" in "human antibody CDR library" includes "CDRL1", "CDRL2", "CDRL3", "CDRH1", "CDRH2", and / or may be substituted with "CDRH3"). Known human CDR sequences include: Jackson et al., J. Immunol Methods, 2007, 324: 26;Martin, Proteins, 1996, 25: 130;Lee et al., Immu nogenetics, 2006, 57: 917, Boyd et al., S Science Translational Medicine, 2009, 1: 1, and International Publication No. WO / 2009 / 036379. and in the HPS provided in Appendix A.
[0067] The term "human preimmune set" or "HPS" refers to the GI numbers provided in Appendix A. A reference set of 3,571 corresponding curated human pre-immune heavy chain sequences was Point.
[0068] A "complete antibody" is an antibody that contains full-length heavy and light chains (i.e., heavy and light chains, respectively). The complete antibody is composed of four framework regions, three CDRs, and a constant region. These antibodies are also called "full-length" antibodies.
[0069] The term "length diversity" refers to the variation in the length of a family of nucleotide or amino acid sequences. For example, in naturally occurring human antibodies, the heavy chain CDR3 sequences are For example, the length can vary from about 2 amino acids to about 35 amino acids or more. For example, they vary in length from about 5 to about 16 amino acids.
[0070] The term "library" refers to a library having the diversity and / or It refers to a set of elements comprising two or more elements designed according to the methods of the present invention, for example: A "library of polynucleotides" is a polynucleotide library having the diversity described herein. and / or a polynucleotide comprising two or more polynucleotides designed according to the methods of the present invention. A "library of polypeptides" refers to a set of nucleotides as described herein. Two or more polypeptides with the desired diversity and / or designed according to the methods of the present invention A "synthetic polynucleotide library" refers to a set of polypeptides containing a plurality of polypeptides. , having the variability described herein and / or designed according to the methods of the present invention refers to a set of polynucleotides that includes two or more synthetic polynucleotides. Libraries in which the members are synthetic are also encompassed by the present invention. A "library" is a library that has the diversity described herein and / or is a library that is based on the method of the present invention. A set of polypeptides comprising two or more polypeptides designed according to the method, e.g. Lipoproteins designed to represent the sequence and length diversity of naturally occurring human antibodies. In some embodiments, the term "library" refers to a library of similar structural features. or a set of elements sharing sequence characteristics, e.g., a "heavy chain library," a "light chain library," "CDRH3 Library," "Antibody Library," and / or "CDRH3 Library" That's fine.
[0071] The term "physical realization" refers to the actual physical representation of a display, e.g., by any display method. theoretical (e.g., computer-based) or synthetic (e.g., refers to a portion of the diversity (e.g., oligonucleotide-based). These include phage display, ribosome display, and yeast display. For synthetic sequences, the size of the physical realization of the library depends on (1) the number of sequences that can be actually synthesized; and (2) the theoretical diversity that can be detected, depending on the limitations of the particular screening method. An exemplary limitation of the screening method is the lack of specific assays (e.g., ribosomal dissociation). Prey, phage display, yeast display) The number of mutants that can be produced and the host cells (e.g., For illustrative purposes, the transformation efficiency of 10 12 of Considering the theoretical diversity of the library members, up to 10 11 Members of Exemplary physical realizations of libraries that can be included (e.g., yeast, bacterial cells, or or ribosome display) therefore limits the theoretical diversity of the library. Approximately 10% will be sampled. 12 Theoretical diversity 10 of the libraries 11 If fewer than 10 members are synthesized, the physical implementation of the library is Currently, up to 10 11 members, giving the library a theoretical diversity of 10 Less than 10% are sampled in the physical realization of the library. 12 The physical realization of a library that can contain more than 100 members would allow the theoretical diversity to be "on- "over-sampling" means that each member is repeated. U(10 12 (This assumes that the entire theoretical diversity of
[0072] The term "polynucleotide(s)" refers to nucleic acids such as DNA molecules and RNA molecules. Acids and their analogs (e.g., nucleotide analogs or nucleic acid chemistry) As desired, polynucleotides may be used to refer to DNA or RNA produced using a variety of methods. The nucleic acids can be synthesized synthetically, for example, using art-recognized nucleic acid chemistry. Alternatively, it may be made enzymatically, for example using a polymerase, if desired. Exemplary modifications include methylation, biotinylation, and other modifications known in the art. It includes known modifications. Furthermore, the nucleic acid molecule can be single-stranded or double-stranded, If desired, it can be linked to a detectable moiety. The expression for base is used by the International Union of Pure and Applied Chemistry. Follow the nomenclature of the International Union of Chemical Applied Chemistry (IUPAC) (see reference in its entirety) See U.S. Patent Application Publication No. 2009 / 0181855, which is hereby incorporated by reference.
[0073] "Pre-immune" antibody libraries are libraries in which naturally occurring human antibody sequences are negatively selected. sequences similar to naturally occurring human antibody sequences prior to undergoing mutation and / or somatic hypermutation diversity of size and length. See, for example, Lee et al. (see the entire Immunogenetics, 2006, 57: 917 ) and the human pre-immune set described herein (HPS) (see Appendix A) is considered to represent sequences derived from the pre-immune repertoire. In certain embodiments of the invention, the sequences of the invention will be similar to these sequences. (e.g. in terms of composition and length).
[0074] As used herein, the term "site-by-site probabilistic" refers to a method for generating a sequence of amino acids. The paper describes a process for constructing a nucleotide sequence in which only the occurrence of an amino acid at each position is considered, and a higher-order moiety is created. The chief (e.g., pair-wise correlation) Not described (e.g., Knappik, each of which is incorporated by reference in its entirety). , et al., J Mol Biol, 2000, 296: 57 and the United States See the analysis provided in Patent Application Publication No. 2009 / 0181855).
[0075] The term "split-pool synthesis" refers to the synthesis of multiple individual first reactions. The products are combined (pooled) and then separated before participating in multiple second reactions. For example, U.S. Patent Application Publication No. 2009 / 01 No. 81855 (incorporated by reference in its entirety) discloses 27 methods for preparing 27 hydroxybenzoates in separate reactions. The synthesis of 8 DH segments (products) is described. After synthesis, these 278 segments are combined (pooled), and the segments are then used for the synthesis of N2. This is distributed (split) into columns. This allows pairing of each of the 278 DH segments with the corresponding nucleotide.
[0076] As used herein, "probabilistic" refers to the selection of an element from a probability distribution. A process that generates random sequences of nucleotides or amino acids that are considered samples (See, for example, U.S. Pat. No. 5,723,323).
[0077] As used herein, the term "synthetic polynucleotide" refers to a polynucleotide derived from a molecule of natural origin. In contrast to a molecule, a molecule formed through a chemical process or a template-based molecule of natural origin refers to molecules derived through amplification of a gene (e.g., from a population of B cells via PCR amplification). Immunoglobulin chains cloned from the host are referred to as "synthetic" immunoglobulin chains as used herein. In some instances, for example, multiple segments (e.g., TN1, D When referring to a library of the invention comprising a plurality of nucleotides (e.g., H, N2, and / or H3-JH), the library of the invention includes a plurality of nucleotides (e.g., H, N2, and / or H3-JH). The invention provides a library in which at least one, two, three, or four of the aforementioned components are synthetic. By way of example, certain components may be synthetic, while others may be of natural origin. Libraries derived through template-based amplification of molecules of natural or natural origin are Libraries that are entirely synthetic will also, of course, be encompassed by the present invention. would be encompassed by the present invention.
[0078] The term "theoretical diversity" refers to the maximum number of variants in a library design. For example, if we consider the amino acid sequence of three residues, residues 1 and 3 each contain five amino acids. Residue 2 may be any one of the 20 amino acid types If only one is allowed, the theoretical diversity is 5 x 20 x 5 = 500 possible sequences. Similarly, if sequence X is constructed by combining four amino acid segments, Segment 1 has 100 possible sequences and segment 2 has 75 possible sequences. Segment 3 has 250 possible sequences and segment 4 has 30 possible sequences. Then the theoretical diversity of fragment X is 100 × 75 × 200 × 30 or 5.6 × 10 5 would be a possible sequence of
[0079] The term "theoretical segment pool" refers to a larger polynucleotide or polypeptide Polynucleotides or nucleotides that can be used as building blocks to construct or a set of polypeptide segments. For example, TN1, DH, N2, and H The theoretical segment pool containing the 3-JH segment is [TN1]-[DH]-[N 2]-[H3-JH] to form a sequence represented by The CDRH3 sequence was derived by ligating the fragments to the corresponding oligonucleotides and synthesizing the corresponding oligonucleotides. The term "theoretical segment pool" refers to can be applied to any set of polynucleotide or polypeptide segments Thus, the set of TN1, DH, N2, and H3-JH segments collectively Although this can be thought of as a theoretical segment pool, each individual set of segments is also In addition, the theoretical segment pool, especially the TN1 theoretical segment pool and the DH theoretical segment pool, The theoretical segment pool includes the top pool, the N2 theoretical segment pool, and the H3-JH theoretical segment pool. Any subset of the theoretical segment pool that contains more than one sequence can also be used. It can be thought of as a logical segment pool.
[0080] As used herein, the term "unique" refers to a designed set (e.g., It is different (e.g., has a different chemical structure) from all other sequences in the theoretical diversity. refers to an arrangement of many unique properties from the theoretical diversity in a particular physical realization. It is understood that there may be more than one copy of a sequence. For example, If a sequence occurs three times in the physical realization of the library, then at the theoretical level there are three A library containing unique sequences may contain a total of nine members. In some embodiments, each unique sequence is present only once, less than once, or more than once. You may do so.
[0081] The term "variable" refers to the variability in their sequence and the specificity and binding of a particular antibody. The part of an immunoglobulin domain (i.e., the "variable domain") that is involved in determining affinity ( The diversity is not evenly distributed throughout the variable domains of antibodies, It is concentrated in the subdomains of each of the heavy and light chain variable regions. The main regions are called "hypervariable" or "complementarity determining regions" (CDRs). The more conserved (i.e., non-hypervariable) parts of the gene are called "framework" regions (FRMs). Naturally occurring heavy and light chain variable domains each contain four FRM regions. Most of them bind and in some cases form part of the β-sheet structure. It adopts a β-sheet configuration held together by three hypervariable regions that form loops that The hypervariable regions in each chain are linked to the hypervariable regions from the other chain by FRMs. are held together in close proximity and contribute to the formation of the antigen-binding site (all of which The body is incorporated by reference. Proteins of Immunological Interest, 5th Ed. Public Health Service, National Institute See Statutes of Health, Bethesda, Md., 1991 The constant domains are not directly involved in antigen binding, but are involved in the regulation of, for example, antibody-dependent cellular mediators. They exhibit a variety of effector functions, such as endothelial cell cytotoxicity and complement activation.
[0082] The libraries of the present invention containing "VKCDR3" and "VλCDR3" sequences are , refer to the kappa and lambda subsets of light chain CDR3 (CDRL3) sequences, respectively. Such libraries represent a comprehensive representation of the length and sequence of the human antibody CDRL3 repertoire. These libraries may be designed with a specified diversity to represent the diversity of the sequences. The "pre-immune" version is a negative selection of naturally occurring human antibody CDRL3 sequences. Similar to naturally occurring human antibody CDRL3 sequences prior to undergoing mutation and / or somatic hypermutation. The known human CDRL3 sequences are: NCBI database, International Publication No. W O / 2009 / 036379, and Martin, Proteins, 1996, 25: 130.
[0083] General Library Design The antibody library provided by the present invention is a library of pre-immune antibodies generated by the human immune system. Certain libraries of the present invention may be designed to reflect certain aspects of the library's repertoire. The library is a rationally designed library characterized by a collection of human V, D, and J genes. and human heavy and light chain sequences (e.g., each of which is incorporated by reference in its entirety). Jackson et al., J. Immunol Methods, 2 007, 324: 26;Lee et al., Immunogenetics, 2006, 57: 917;Boyd et al., Science Tran Publicly known from Relational Medicine, 2009, 1: 1-8 The sequences were compiled from germline and rearranged VK and Vλ sequences. The sequence (International Publication No. WO / 2009 / 0 For further information, see e.g. ,Scaviner et al., each of which is incorporated by reference in its entirety. Exp. Clin. Immunogenet., 1999, 16: 234;T omlinson et al., J. Mol. Biol., 1992, 22 7: 799; and Matsuda et al., J. Exp. Med., 1998, 188:2151.
[0084] In one embodiment of the invention, the possible V, D, and segments showing J diversity and junctional diversity (i.e., TN1 and N2) The DNA fragments are synthesized de novo as single-stranded or double-stranded DNA oligonucleotides. In embodiments, the oligonucleotides encoding the CDR sequences are in yeast together with one or more acceptor vectors containing the sequences and constant domains. Primer-based PCR amplification from mammalian cDNA or mRNA Alternatively, the template-directed cloning step may be Through standard homologous recombination, the recipient yeast receives the chassis sequence and The CDR segments are recombined with an acceptor vector containing the constant region and genetically amplified. A properly ordered synthesis that can be synthesized, expressed, displayed and screened. A full-length human heavy and / or light chain immunoglobulin library is generated. Acceptor vectors may be designed to yield constructs other than full-length human heavy and / or light chains. For example, one embodiment of the present invention may be designed as follows: In some embodiments, the chassis may comprise a polypeptide encoding an antibody fragment or a subunit of an antibody fragment. The oligonucleotide sequences containing the CDRs may be designed to encode portions of the CDRs. When the set is recombined with the acceptor vector, the antibody fragment or subunit thereof A sequence encoding the
[0085] Thus, in one embodiment, the present invention provides a synthetic pre-immune human antibody repertoire comprising: (a) one or more selected human antibody heavy chain chassis (i.e., according to the Kabat definition) When used, amino acids 1 to 94 of the heavy chain variable region); (b) Human IGHD and IGHJ germline sequences from the reference set of human CDRH3 sequences. and a CDRH3 repertoire designed based on the extraction of TN1 and N2 sequences (see below). (i) a TN1 segment; (ii) a DH segment; (iii) the N2 segment; (iv) the CDRH3 replicator containing the H3-JH segment. Tory, (c) one or more selected human antibody kappa and / or lambda light chain chassis; and (d) CDRL3 replicators designed based on human IGLV and IGLJ germline sequences. and a CDRL3 repertoire in which "L" may be a kappa or lambda light chain. We offer a repertoire that includes:
[0086] The present invention also provides methods for producing such libraries and one or more immunoglobulin domains. The present invention provides methods for producing and using libraries containing antibodies or antibody fragments. The design and synthesis of each component of the antibody library is provided in more detail below. can be.
[0087] Antibody library chassis sequence design In one embodiment, the libraries provided are composed of naturally occurring variable domain sequences (e.g., constructed from selected chassis sequences based on the IGHV and IGLV genes The selection of such chassis arrangements may be arbitrary or based on some predetermined criteria. For example, electronic data containing non-overlapping rearranged antibody sequences can be generated. The Kabat database, a database of the most frequently represented heavy and light chain germline sequences, Sequences can be queried using algorithms such as BLAST or S More specialized tools such as oDA (Vol. 10, incorporated by reference in its entirety) PE et al., Bioinformatics, 2006, 22: 438 -44) identifies the germline families most frequently used to generate functional antibodies. To identify the gene, germline sequences (e.g., V) are identified using the BASE2 database; e.g., , Retter et al., Nucleic Acids Res., 2005, 33: D671-D674) or used to compare rearranged antibody sequences with similar sets of human V, D, and J genes. It can be used.
[0088] Several criteria are used for the selection of chassis for inclusion in the library of the present invention. For example, yeast or other organisms (e.g. poor expression in bacteria, mammalian cells, fungi, or plants Sequences for which the identity is known (or determined) can be removed from the library. The chassis are also based on the expression of their corresponding germline genes in human peripheral blood. In one embodiment of the present invention, a protein highly expressed in human peripheral blood may be selected. It may be desirable to select a chassis that corresponds to a germline sequence. In some cases, for example, less frequent sequencing is required to increase the canonical diversity of the library. It may be desirable to select a chassis that corresponds to a germline sequence not shown in Therefore, the chassis represents the largest and most structurally diverse group of functional human antibodies. may be selected to produce
[0089] In some embodiments of the present invention, for example, chassis diversity is less prevalent and more Producing smaller, more focused libraries with more CDR diversity If desired, a less diverse chassis may be utilized. In some embodiments, the chassis are designed to inhibit their expression in cells of the invention (e.g., yeast cells). Selection based on both the diversity of canonical structures exhibited by the present and selected sequences Therefore, the diversity of canonical structures that are fully expressed in the cells of the present invention can be A library may be produced having:
[0090] Heavy chain chassis sequence design The design and selection of heavy chain chassis sequences that can be used in the present invention are described in U.S. Pat. Publication No. 2009 / 0181855 and U.S. Patent Application Publication No. 2010 / 00563 86 and International Publication No. WO / 2009 / 036379. , each of which is incorporated by reference in its entirety and is therefore briefly described herein. All that is needed is to
[0091] Generally, the VH domain of the library contains three components: (1) amino acids 1 to 5; 94 (using Kabat numbering) VH "chassis", (2) appropriate Kab CDRs defined herein to include at CDRH3 (positions 95-102) H3, and (3) FRM containing amino acids 103–113 (Kabat numbering). 4 regions. The overall VH domain structure may therefore be represented schematically as follows: (Not to scale). [ka]
[0092] In one embodiment of the invention, the VH chassis of the library may comprise one or more of the following: It may include about Kabat residue 1 to about Kabat residue 94 of the IGHV germline sequence. :IGHV1-2, IGHV1-3, IGHV1-8, IGHV1-18, IGHV1- 24, IGHV1-45, IGHV1-46, IGHV1-58, IGHV1-69, I GH8, IGH56, IGH100, IGHV3-7, IGHV3-9, IGHV3-1 1, IGHV3-13, IGHV3-15, IGHV3-20, IGHV3-21, IG HV3-23, IGHV3-30, IGHV3-33, IGHV3-43, IGHV3- 48, IGHV3-49, IGHV3-53, IGHV3-64, IGHV3-66, I GHV3-72, IGHV3-73, IGHV3-74, IGHV4-4, IGHV4- 28, IGHV4-31, IGHV4-34, IGHV4-39, IGHV4-59, I GHV4-61, IGHV4-B, IGHV5-51, IGHV6-1, and / or In some embodiments of the invention, the library comprises one or It may contain multiple of these sequences, one or more allelic variants of these sequences. or one or more of these sequences. 9%, 98.5%, 98%, 97.5%, 97%, 96.5%, 96%, 95.5%, 9 5%, 94.5%, 94%, 93.5%, 93%, 92.5%, 92%, 91.5%, 9 1%, 90.5%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83% ,82%,81%,80%,77.5%,75%,73.5%,70%,65%,60% , 55%, or 50% identical amino acid sequences. In view of the definition of chassis provided in, any IGHV coding sequence may be used as a chassis of the present invention. It will be appreciated that the present invention can be adapted for use as a computer. 2009 / 0181855 and 2010 / 0056386 and International Publication Nos. WO / 2009 / 036379 (each of which is incorporated by reference in its entirety). As exemplified in the following example, these chassis also have specific functions in the CDRH1 and CDRH2 regions. The diversity of the library was further increased by modifying the amino acid residues in the It can also be increased.
[0093] Light chain chassis sequence design The design and selection of light chain chassis sequences that can be used in the present invention are described in U.S. Pat. Publication No. 2009 / 0181855 and U.S. Patent Application Publication No. 2010 / 00563 86 and International Publication No. WO / 2009 / 036379. , each of which is incorporated by reference in its entirety and is therefore briefly described herein. The light chain chassis of the present invention are based on kappa and / or lambda light chain sequences. It may also be something.
[0094] The VL domain of the library contains three major components: (1) the VL "chassis" (2) an appropriate Kabat numbering system containing amino acids 1–88 (using Kabat numbering). CDRs defined herein to include bat CDRL3 (positions 89-97) L3, and (3) FRM4, including amino acids 98–107 (Kabat numbering). Therefore, the overall VL domain structure may be represented schematically as follows: (Not to scale). [ka]
[0095] In one embodiment of the invention, the VL chassis of the library is based on the IGKV germline sequence. In one embodiment of the present invention, the library comprises one or more chassis based on The L chassis may be formed by ligating one or more of the following Kabat residues from about 1 to about 1K of the IGKV germline sequence: and may contain about abat residue 88: IGKV1-05, IGKV1-06, IGK V1-08, IGKV1-09, IGKV1-12, IGKV1-13, IGKV1-1 6, IGKV1-17, IGKV1-27, IGKV1-33, IGKV1-37, IG KV1-39, IGKV1D-16, IGKV1D-17, IGKV1D-43, IGK V1D-8, IGK54, IGK58, IGK59, IGK60, IGK70, IGKV 2D-26, IGKV2D-29, IGKV2D-30, IGKV3-11, IGKV3 -15, IGKV3-20, IGKV3D-07, IGKV3D-11, IGKV3D- 20, IGKV4-1, IGKV5-2, IGKV6-21, and / or IGKV6 D-41. In some embodiments of the invention, the library comprises one or more of the following: These sequences may contain one or more allelic variants of these sequences, or One or more of these sequences have at least about 99.9%, 99.5%, 99%, 98%, or 0.5%, 98%, 97.5%, 97%, 96.5%, 96%, 95.5%, 95%, 94 0.5%, 94%, 93.5%, 93%, 92.5%, 92%, 91.5%, 91%, 90 .5%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 77.5%, 75%, 73.5%, 70%, 65%, 60%, 55%, Alternatively, it may encode an amino acid sequence that is 50% identical.
[0096] In one embodiment of the invention, the VL chassis of the library is based on an IGλV germline sequence. In one embodiment of the present invention, the library comprises one or more chassis based on The L chassis may be one or more of the following: May contain approximately 88 abat residues: IGλV3-1, IGλV3-21, IGλ4 4, IGλV1-40, IGλV3-19, IGλV1-51, IGλV1-44, IG λV6-57, IGλ11, IGλV3-25, IGλ53, IGλV3-10, IGλ V4-69, IGλV1-47, IGλ41, IGλV7-43, IGλV7-46, I GλV5-45, IGλV4-60, IGλV10-54, IGλV8-61, IGλV 3-9, IGλV1-36, IGλ48, IGλV3-16, IGλV3-27, IGλ V4-3, IGλV5-39, IGλV9-49, and / or IGλV3-12. In some embodiments of the invention, the library contains one or more of these sequences, may contain one or more allelic variants of these sequences or Multiple of these sequences have at least about 99.9%, 99.5%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%, 75%, Encoding amino acid sequences that are 70%, 65%, 60%, 55%, or 50% identical Good too.
[0097] Those skilled in the art will recognize any IGKV or IGλ chassis in light of the chassis definitions provided above. It will be appreciated that V coding sequences can be adapted for use as chassis in the present invention. It will be.
[0098] Design and selection of TN1, DH, N2, and H3-JH segments The human germline repertoire contains at least six IGHJ genes (IGHJ1, IGH J2, IGHJ3, IGHJ4, IGHJ5, and IGHJ6; the main The essential allele is designated "01" and the selected allelic variant is designated "02" or " 03) and at least 27 IGHD genes (Table 16, allelic variants In some embodiments, the present invention provides a CDRH3 polypeptide sequence. Libraries of polynucleotide sequences encoding the sequences or CDRH3 sequences and the A library containing any member of the theoretical segment pool disclosed in the document. Includes.
[0099] Those skilled in the art will recognize all of the segments in the theoretical segment pool provided herein. The addition of a CDRH3 fragment is not necessarily required to generate a functional CDRH3 library of the present invention. Therefore, in one embodiment, the CDRH3 live antibody of the present invention may be used in combination with a CDRH3 live antibody. The pool may be any segment of the theoretical segment pool described herein. For example, in one embodiment of the invention, Theoretical segmentation methods provided in or generated by the methods described herein At least about 15, 30, 45, 60, 75, 90, 100, 100% of any of the 05, 120, 135, 150, 165, 180, 195, 200, 210, 225, 2 40, 255, 270, 285, 300, 320, 340, 360, 380, 400, 4 20, 440, 460, 480, 500, 520, 540, 560, 580, 600, 6 20, 640, or 643 H3-JH segments are included in the library. In some embodiments of the present invention, the at least about 15 of any of the theoretical segment pools generated by the method of 30, 45, 60, 75, 90, 100, 105, 120, 135, 150, 165, 1 80, 195, 200, 250, 300, 350, 400, 450, 500, 550, 6 00, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 , 1100, 1111, 2000, 3000, 4000, 5000, 6000, 7000 , 14000, 21000, 28000, 35000, 42000, 49000, 560 DH segments of 00, 63000, or 68374 are included in the library. In some embodiments of the invention, the compounds provided or described herein are At least about 10 of any of the theoretical segment pools generated by the methods described , 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130 , 140, 141, 150, 160, 170, 180, 190, or 200, 220 , 240, 260, 280, 300, 320, 340, 360, 380, 400, 420 , 424, 440, 460, 480, 500, 550, 600, 650, 700, 727 , 750, 800, 850, 900, 950, or 1000 TN1 and / or N In one embodiment, the library of the invention comprises two segments. It may contain less than a certain number of polynucleotide or polypeptide segments, The number of segments is one of the integers provided above for each segment. In some embodiments of the present invention, the specified numerical ranges are inclusive or or using any two of the integers provided above as the lower and upper bounds of an exclusive range. All combinations of integers provided that define upper and lower limits are contemplated. will be done.
[0100] In one embodiment, the present invention provides a method for the preparation of a theoretical segment pool as provided herein. At least about 1%, 2.5%, 5%, 10%, 15%, 20% of the segments from either %, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75 CDRH3 library offerings include 80%, 85%, 90%, 95%, or 99% For example, the present invention provides a method for producing a nucleotide sequence comprising the steps of: at least about one TN1, DH, N2, and / or H3-JH segment derived from %, 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50 %, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or In some embodiments of the present invention, a library containing 99% of the specific parts is provided. The percentage ranges provided above are inclusive or exclusive of the lower and upper limits of the ranges. The percentages are set using any two of the following: All combinations of the percentages provided are contemplated.
[0101] In some embodiments of the present invention, H3-JH, DH in the CDRH3 library , TN1, and / or N2 segments at least approximately 1%, 2.5%, 5%, or 10% , 15%, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65% , 70%, 75%, 80%, 85%, 90%, 95%, or 99% are used herein to mean Theoretical segments provided in or generated by the methods described herein any of the H3-JH, DH, TN1, and / or N2 segments of the top pool In some embodiments of the invention, the CDRH3 fragment is isolated from a library (e.g., , which bind to specific antigens and / or generic ligands through one or more rounds of selection. (by using a small amount of the H3-JH, DH, TN1, and / or N2 segments of the antibody) At least about 1%, 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the compounds provided or described herein H3-JH, DH, TN from any of the theoretical segment pools generated by the method In one embodiment, the CDRH3 library of the present invention is a CDRH3 segment. A brary is less than a certain percentage as provided herein or as defined herein. Any of the theoretical segment pools H3- generated by the method described in may contain JH, DH, TN1, and / or N2 segments, and the number of segments Use one of the percentages provided above for each segment. In some embodiments of the present invention, a particular percentage range is determined using: the percentages provided above as the lower and upper limits of inclusive or exclusive ranges; The percentages provided are used to define upper and lower limits. All page combinations are contemplated.
[0102] Those skilled in the art, upon appreciating the disclosure herein, will recognize that the present invention is not limited to the above-described embodiments. Any of the theoretical segment pools generated by the methods described herein Consider the TN1, DH, N2, and / or H3-JH segments of It may not be 100% identical to the provided one in terms of functionality, but may be very similar in function. , analogous TN1, DH, N2, and / or H3-JH segments and corresponding It will be appreciated that CDRH3 libraries can be generated. Such theoretical segment pools and CDRH3 libraries are also within the scope of the present invention. Mutagenesis techniques well known in the art, including those provided herein, are also known. A variety of techniques known in the art can be used to obtain these additional sequences. Therefore, each explicitly recited embodiment of the present invention is also provided herein. Theoretical segment pools obtained by or generated by the methods described herein Use segments that share a certain percent identity with any of the segments in For example, each of the above-described embodiments of the present invention can be implemented using , provided herein or produced by the methods described herein TN1, DH, N2, and / or H3-JH in any of the theoretical segment pools Segment at least approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, TN1, DH, N2, and / or TN3 that are 99%, 99.5%, or 99.9% identical This can be done using the H3-JH segment.
[0103] In some embodiments, the present invention provides one or more TN1 segments, one or Multiple DH segments, one or more N2 segments, and one or more H3 segments - produced from one or more VH chassis sequences combined with a JH segment In one embodiment, a library is provided. , 75, or 100 chassis, TN1, DH, N2, or H3-JH The fragments are included in the libraries of the present invention.
[0104] In some embodiments, the present invention provides a method for the detection of CDRH3 variants of GABA-1, GABA-2, GABA-3, GABA-4, GABA-5, GABA-6, GABA-7, GABA-8, GABA-9, GABA-10, GABA-11, GABA-12, GABA-13, GABA-14, GABA-15, GABA-16, GABA-17, GABA-18, GABA-19, GABA-20, GABA-21, GABA-22, GABA-23, GABA-24, GABA-25, GABA-26, GABA-27, GABA-28, GABA-29 ... Select the TN1, DH, N2, and H3-JH segments from the theoretical segment pool to 1. A method for: (i) a nucleotide sequence containing one or more TN1, DH, N2, and H3-JH segments; providing a logical segment pool; (ii) providing a reference set of CDRH3 sequences; (iii) The closest match to each CDRH3 sequence in the reference set in (ii) ( utilizing the theoretical segment pool of (i) to identify (a) or (b) segments; to (iv) selecting a segment from the theoretical segment pool for inclusion in a synthetic library; The method includes selecting:
[0105] In some embodiments, the selection process of (iv) optionally includes selecting a reference cell of (ii). The frequency of occurrence of the segment (i) in the set; the corresponding segment usage weight; and Any physicochemical properties of the segments (www.genome.jp / aaindex / See all metrics in (e.g., hydrophobicity, alpha-helical propensity, and Optionally, (i) (ii) are not present in the theoretical segment pool but are found in the reference set The TN1 and / or N2 segments involved may be identified and potential theoretical In the segment pool and / or synthetic library of the invention, TN1 and / or To generate a theoretical segment pool with increased N2 or N3 diversity, may be added to the target segment pool.
[0106] For example, one or more biological properties (e.g., immunogenicity, stability, half-life) and and / or one or more physicochemical properties, such as the numerical indices provided above. Any feature or set of features of a segment, including its sex, may be considered for inclusion in a library. In some embodiments, at least one 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or These and other such characteristics may be used to select segments for inclusion in the libraries of the present invention. The physicochemical properties included in the indicators provided above are used to select , ANDN920101 alpha-CH chemical shift (Andersen et al. , 1992);ARGP820101 hydrophobicity index (Argos et al., 1 982);ARGP820102 signal sequence helix potential (helica l potential)(Argos et al., 1982);ARGP820 103 Membrane-embedded selection parameters (Argos et al., 1982); BEGF 750101 Conformational parameters of the internal helix (Beghin-Dirkx, 1975);BEGF750102 3D structural parameters of beta structure (Beghin -Dirkx, 1975);BEGF750103 Beta turn conformation parameters Tar (Beghin-Dirkx, 1975); BHAR880101 Average mobility of fingers Mark (Bhaskaran-Ponnuswamy, 1988);BIGC670101 Residue volume (Bigelow, 1967); BIOV880101 accessibility Information value: average fraction 35% (Biou et al., 1988); BIOV8801 02 Information value of reachability; average fraction 23% (Biou et al., 198 8);BROC820101 Retention factor in TFA (Browne et al., 1982);BROC820102 Retention factor in HFBA (Browne et al., 1982);BULH740101 Transfer free energy to the surface (Bul l-Breese, 1974);BULH740102 apparent partial specific volume (Bul l-Breese, 1974);BUNA790101 alpha-NH chemical shift ( Bundi-Wuthrich, 1979);BUNA790102 Alpha-CH Chemical shifts (Bundi-Wuthrich, 1979); BUNA790103 The pin-spin coupling constant 3JHalpha-NH (Bundi-Wuthrich, 19 79);BURA740101 α-helix normalized frequency (Burgess et al. l., 1974);BURA740102 Standardized frequency of extended structures (Burgess et al., 1974); CHAM810101 Stereoscopic parameters (Chart on, 1981); CHAM820101 polarizability parameters (Charton-C harton, 1982);CHAM820102 Free energy of aqueous solution, kca l / mol (Charton-Charton, 1982); CHAM830101 Chou-Fasman parameters (Charton-Charton) of the 3D structure of the silane , 1983);CHAM830102 Chou-Fasman parameters of beta sheet The parameter (Charton- Charton, 1983); CHAM830103 in the side chain labeled 1+1 Number of atoms (Charton-Charton, 1983); CHAM830104 The number of atoms in the side chain labeled 2+1 (Charton-Charton, 198 3); CHAM830105 The number of atoms in the side chain labeled 3+1 (Chart n-Charton, 1983); CHAM830106, the bond in the longest chain Number (Charton-Charton, 1983);CHAM830107 Charge transfer Ability parameters (Charton-Charton, 1983); CHAM830 108 Charge transfer donor ability parameters (Charton-Charton, 19 83);CHOC750101 average volume of buried residues (Chothia, 1975); CHOC760101 Residue nearest neighbor area in tripeptides (Chothia, 1976);CHOC760102 Residue nearest neighbor area in folded proteins (Chothia, 1976); CHOC760103 95% buried residue percentage (Ch othia, 1976); CHOC760104 100% buried residue ratio (Chot hia, 1976);CHOP780101 Normalized frequency of beta turns (Chou- Fasman, 1978a); CHOP780201 Normalized frequency of alpha helices degree (Chou-Fasman, 1978b); CHOP780202 beta sheet Standardized frequency (Chou-Fasman, 1978b); CHOP780203 Beta Normalized frequency of turns (Chou-Fasman, 1978b); CHOP780204 Normalized frequency of N-terminal helices (Chou-Fasman, 1978b); CHOP 780205 Normalized frequency of C-terminal helix (Chou-Fasman, 1978b );CHOP780206 Normalized frequency of N-terminal non-helical region (Chou-Fasm an, 1978b);CHOP780207 Normalized frequency of the C-terminal non-helical region ( Chou-Fasman, 1978b);CHOP780208 N-terminal beta sheet Normalized frequency of (Chou-Fasman, 1978b); CHOP780209 C-terminus Normalized frequency of edge beta sheets (Chou-Fasman, 1978b); CHOP78 0210 Normalized frequency of N-terminal non-beta region (Chou-Fasman, 1978b) ;CHOP780211 Normalized frequency of C-terminal non-beta region;CHOP780212 Frequency of the first residue in the sequence (Chou-Fasman, 1978b); CHOP7 Frequency of the second residue in a turn (Chou-Fasman, 1978b );CHOP780214 Frequency of the third residue in a turn (Chou-Fasman , 1978b);CHOP780215 Frequency of the fourth residue in a turn (Chou -Fasman, 1978b);CHOP780216 Second and Third Turns Normalized frequencies of residues in 3 (Chou-Fasman, 1978b); CIDH92010 1 Standardized hydrophobicity scale for alpha proteins (Cid et al., 1 992);CIDH920102 Standardized Hydrophobicity Scale for Beta Proteins ( Cid et al., 1992);CIDH920103 alpha and beta proteins Standardized Hydrophobicity Scale for Proteins (Cid et al., 1992); CIDH 920104 Standardized Hydrophobicity Scale for Alpha / Beta Proteins (Cid et al., 1992);CIDH920105 Normalized Average Hydrophobicity Scale (Ci d et al., 1992);COHE430101 partial specific volume (Cohn-Eds all, 1943); CRAJ730101 Normalized frequency of the central helix (Cra wford et al., 1973);CRAJ730102 Beta sheet standard Frequency of mutation (Crawford et al., 1973); CRAJ730103 Standardized frequency of DAWD72010 (Crawford et al., 1973) 1 Size (Dawson, 1972); DAYM780101 Amino acid composition (Da yhoff et al., 1978a);DAYM780201 relative mutagenicity (Dayhoff et al., 1978b);DESM900101 cytochrome Membrane selection for b: MPH89 (Degli Esposti et al., 1 990);DESM900102 Average membrane selection:AMP07 (Degli Espos ti et al., 1990);EISD840101 Consensus Normalized Hydrophobicity Scale (Eisenberg, 1984); EISD860101 Solvation Free Energy Lugie (Eisenberg-McLachlan, 1986); EISD86010 2 Atom-based hydrophobic moment (Eisenberg-McLachlan, 19 86);EISD860103 Hydrophobic moment direction;FASG760101 Molecule Amount (Fasman, 1976); FASG760102 Melting point (Fasman, 19 76);FASG760103 Optical rotation (Fasman, 1976);FASG760 104 pK-N (Fasman, 1976);FASG760105 pK-C (Fasman, 1976);FAUJ830101 Hydrophobicity parameter pi(Fa uchere-Pliska, 1983);FAUJ880101 Graph Shape Index (Fauchere et al., 1988); FAUJ880102 Smoothed Upsilon steric parameters (Fauchere et al., 1988) ;FAUJ880103 Normalized van der Waals volume (Fauchere et al. l., 1988);FAUJ880104 Side chain STERIMOL length (Fauche re et al., 1988);FAUJ880105 Side chain STERIMOL Small width (Fauchere et al., 1988); FAUJ880106 side chain STERIMOL maximum width; FAUJ880107 NMR chemical shift of alpha carbon (Fauchere et al., 1988); FAUJ880108 Local Electrical effects (Fauchere et al., 1988); FAUJ880109 Number of hydrogen bond donors (Fauchere et al., 1988); FAUJ880 110 Number of total nonbonding orbitals (Faucher et al., 1988); FAUJ 880111 Positive charge (Faucher et al., 1988); FAUJ88 0112 Negative charge (Faucher et al., 1988); FAUJ8801 13 pK-a(RCOOH) (Faucher et al., 1988);F INA770101 Helix-coil equilibrium constant (Finkelstein-Pitts yn, 1977);FINA910101 Helix initiation parameter at position i-1 -(Finkelstein et al., 1991);FINA910102nd place Helix initiation parameters at positions i, i+1, and i+2 (Finkelstein et al., 1991);FINA910103 Helices at positions j-2, j-1, and j FINA 910104 Helix end parameter at position j+1 (Finkelstein et al., 1991);GARJ730101 partition coefficient (Garel et a l., 1973);GEIM800101 Alpha Helix Index (Geisow- Roberts, 1980);GEIM800102 for alpha protein Alpha Helix Index (Geisow-Roberts, 1980); GEIM80 Alpha helix index for beta proteins (Geisow-Rob erts, 1980);GEIM800104 Alpha / beta proteins Alpha helix index (Geisow-Roberts, 1980); GEIM8 00105 Beta strand index (Geisow-Roberts, 1980); G EIM800106 Beta Strand Index for Beta Proteins (Geisow -Roberts, 1980);GEIM800107 alpha / beta protein Beta strand index, GEIM800108 non-periodic index (Geisow- R oberts, 1980);GEI M800109 for alpha protein Aperiodic indicators (Geisow-Roberts, 1980); GEIM800110 Aperiodic index for beta proteins (Geisow-Roberts, 1980 );GEIM800111 Aperiodic Index for Alpha / Beta Proteins (Ge isow-Roberts, 1980); GOLD730101 Hydrophobicity factor (Gol dsack-Chalifoux, 1973);GOLD730102 Residue volume (G oldsack-Chalifoux, 1973);GRAR740101 Composition (G Grantham, 1974);GRAR740102 Polarity(Grantham, 1 974), GRA740103 volume (Grantham, 1974); GUYH8 50101 Distributed Energy (Guy, 1985);Charton-Charton (1982) cited the HOPA770101 hydration number (Hopfinger, 1971), HOPT810101 hydrophilicity value (Hopp-Woods, 1981) ;HUTJ700101 Heat capacity (Hutchens, 1970);HUTJ7001 02 Absolute entropy (Hutchens, 1970); HUTJ700103 Raw Entropy of formation (Hutchens, 1970); ISOY800101 Alpha Relative standardized frequency of lichens (Isogai et al., 1980);ISOY 800102 Relative normalized frequency of extended structures (Isogai et al., 1980 );ISOY800103 relative normalized frequency of bending (Isogai et al., 1980);ISOY800104 Relative standardized frequency of bending radius (Isogai et al., 1980);ISOY800105 Relative standardized frequency of bending S (Isoga i et al., 1980);ISOY800106 relative standard for helix ends frequency (Isogai et al., 1980); ISOY800107 double bending Relative standardized frequency (Isogai et al., 1980); ISOY80010 Relative normalized frequency of 8 coils (Isogai et al., 1980); JANJ 780101 Average nearest enveloping area (Janin et al., 1978); JAN J780102 Percentage of buried residues (Janin et al., 1978) ;JANJ780103 Percentage of exposed residues (Janin et al., 1 978);JANJ790101 Ratio of buried and accessible mole fractions (Janin, 1 979);JANJ790102 Transfer free energy (Janin, 1979);J OND750101 hydrophobic (Jones, 1975); JOND750102 pK (-COOH) (Jones, 1975);JOND920101 Relative frequency degree (Jones et al., 1992); JOND920102 relative mutation Sex (Jones et al., 1992), JUKT750101 amino acid distribution ( Jukes et al., 1975);JUNJ780101 sequence frequency (Jung ck, 1978);KANM800101 Average relative likelihood of helix (Kane hisa-Tsong, 1980);KANM800102 beta sheet average relative possibility (Kanehisa-Tsong, 1980); KANM800103 internal Average relative likelihood of helix (Kanehisa-Tsong, 1980); KAN M800104 Average relative probability of inner beta sheets (Kanehisa-Tsong , 1980);KARP850101 Mobility parameters for the absence of strong neighbors One meter (Karplus-Schulz, 1985); KARP850102 Mobility parameters for the rigid neighbors of (Karplus-Schulz, 19 85);KARP850103 Mobility parameter (K arplus-Schulz, 1985);KHAG800101 Kerr constant increment ment (Khanarian-Moore, 1980);KLEP840101 Net charge (Klein et al., 1984); KRIW710101 side chain interactions Parameters for (Krigbaum-Rubin, 1971); KRIW790101 Side chain interaction parameters (Krigbaum-Komoriya, 1979); K RIW790102 Fraction of the area occupied by water (Krigbaum-Komor iya, 1979);KRIW790103 Side chain volume (Krigbaum-Komo riya, 1979);KYTJ820101 Hydropathy Index (Kyte-Doo little, 1982);LAWE840101 Transfer free energy, CHP / water (Lawson et al., 1984); LEVM760101 Hydrophobicity parameter (Levitt, 1976); LEVM760102 Side chain C-alpha and Distance between centers of mass (Levitt, 1976); LEVM760103 Side chain angle theta (AAR)(Levitt, 1976);LEVM760104 Side chain torsion angle phi (AAAR)(Levitt, 1976);LEVM760105 Radius of gyration of side chain ( Levitt, 1976);LEVM760106 van der Waals parameters R0 (Levitt, 1976), LEVM760107 van der Waals parameters ter epsilon (Levitt, 1976); LEVM780101 important Normalized frequency of alpha helices (Levitt, 1978); LEVM78010 2 Normalized frequency of beta sheets with significance (Levitt, 1978); LEVM 780103 Standardized frequency of reverse turns with significance (Levitt, 1978);L EVM780104 Normalized frequency of alpha helices not considered significant (Levitt t, 1978);LEVM780105 Standardization frequency of beta sheets not considered important degree (Levitt, 1978);LEVM780106 Reverse turn not considered important Normalized frequency of (Levitt, 1978); LEWP710101 in beta bending Frequency of occurrence (Lewis et al., 1971); LIFS790101 Conformational selection for all beta strands (Lifson-Sander, 197 9);LIFS790102 Conformational selection for parallel beta strands (Li fson-Sander, 1979);LIFS790103 Anti-parallel beta Conformational selection for strands (Lifson-Sander, 1979); MA NP780101 Average Peripheral Hydrophobicity (Manavalan-Ponnuswamy, 1 978);MAXF760101 alpha helix normalized frequency (Maxfield -Scheraga, 1976);MAXF760102 Standardized Frequency of Extended Structure (M axfield-Scheraga, 1976);MAXF760103 Zeta R Standardized frequency (Maxfield-Scheraga, 1976); MAXF76010 4 Normalized frequency of left-handed alpha helices (Maxfield-Scheraga, 1976);MAXF760105 Zeta L Standardized Frequency (Maxfield-Sc heraga, 1976);MAXF760106 alpha region normalized frequency (Ma xfield-Scheraga, 1976);MCMT640101 Jones (1975) cited refractive power (McMeekin et al., 1964 ); MEEJ800101 HPLC, retention factor at pH 7.4 (Meek, 19 80); MEEJ800102 HPLC, retention factor at pH 2.1 (Meek, 1980); MEEJ810101 Retention factor in NaClO4 (Meek-Ros setti, 1981); MEEJ810102 Retention factor in NaH2PO4 ( Meek-Rossetti, 1981); MEIH800101 C-alpha Average equivalent distance (Meirovitch et al., 1980); MEIH8 Average reduced distance for side chains (Meirovitch et al., 1 980);MEIH800103 Average side chain orientation angle (Meirovitch et al ., 1980);MIYS850101 Effective distribution energy (Miyazawa- Jernigan, 1985); NAGK730101 Alpha Helix Standardization Frequency (Nagano, 1973); NAGK730102 Standardized frequency of beta structure ( Nagano, 1973), NAGK730103 coil standardized frequency (Nagan o, 1973);NAKH900101 Amino acid composition of total protein (Nakash ima et al., 1990);NAKH900102 Amino acids of total protein Composition SD (Nakashima et al., 1990); NAKH900103 Amino acid composition of mt-protein (Nakashima et al., 1990) ; NAKH900104 Standardized composition of mt-protein (Nakashima et al., 1990);NAKH900105 Amino acids of animal-derived mt-proteins Composition (Nakashima et al., 1990);NAKH900106 Animal Standardized composition derived from (Nakashima et al., 1990); NAKH900 107 Amino acid composition of mt-proteins from fungi and plants (Nakashima et al., 1990);NAKH900108 Standardized composition of fungi and plant origin (Nakashima et al., 1990); NAKH900109 membrane protein Amino acid composition of protein (Nakashima et al., 1990); NAKH90 Normalized composition of membrane proteins (Nakashima et al., 1990 );NAKH900111 Non-mt-protein transmembrane domain (Nakashima e t al., 1990);NAKH900112 transmembrane domain of mt-protein (N Akashima et al., 1990);NAKH900113 Mean and arithmetic Ratio of the composition of the ash (Nakashima et al., 1990); NAKH920101 Amino acid composition of single-pass protein CYT (Nakashima-Nishikaw a, 1992);NAKH920102 Amino acid sequence of single-pass protein CYT2 (Nakashima-Nishikawa, 1992); NAKH920103 Amino acid composition of the EXT of single-pass proteins (Nakashima-Nishikawa , 1992);NAKH920104 Amino acid composition of the single-pass protein EXT2 (Nakashima-Nishikawa, 1992);NAKH920105 1 Amino acid composition of MEM of transmembrane proteins (Nakashima-Nishikawa, 1992);NAKH920106 Amino acid composition of multi-transmembrane protein CYT ( Nakashima-Nishikawa, 1992);NAKH920107 Multiple Amino acid composition of the EXT of transmembrane proteins (Nakashima-Nishikawa, 1992);NAKH920108 Amino acid composition of MEM of multi-transmembrane proteins ( Nakashima-Nishikawa, 1992);NISK800101 8 Number of contacts (Nishikawa-Ooi, 1980); NISK860101 14 contacts Touch count (Nishikawa-Ooi, 1986); NOZY710101 Transfer energy Ghee, organic solvent / water (Nozaki-Tanford, 1971); OOBM7701 01 average non-bonded energy per atom (Oobatake-Ooi, 1977); OOBM770102 Short and medium range non-bonding energies per atom (Oobatak e -Ooi, 1977);OOBM770103 Long-range non-bonding energy per atom (Oobatake-Ooi, 1977), OOBM770104 average per residue Non-bonded energy (Oobatake-Ooi, 1977); OOBM770105 Short- and medium-range non-bonding energies per residue (Oobatake-Ooi, 1977 );OOBM850101 Optimized beta-structure coil equilibrium constant (Oobatake et al., 1985);OOBM850102 Optimal tendency to form reverse turns (Oo batake et al., 1985);OOBM850103 Optimized Transfer Energy Ghee parameters (Oobatake et al., 1985); OOBM8501 04 Optimized average non-bond energy per atom (Oobatake et al., 1985);OOBM850105 Optimized side chain interaction parameters (Oobatak e et al., 1985);PALJ810101 LG-derived alpha helices Standardized frequency of serotonin (Palau et al., 1981); PALJ810102C normalized frequency of alpha helices derived from F ( Palau et al., 1981 ); Normalized frequency of beta sheets derived from PALJ810103LG (Palau et al. , 1981); PALJ810104 CF-derived beta sheet normalized frequency (Pal au et al., 1981);PALJ810105 Standardization of LG-derived turns Frequency (Palau et al., 1981); PALJ810106 CF-derived Normalized frequency of the serotonin (Palau et al., 1981); PALJ810107 Normalized frequency of alpha helices across all alpha classes (Palau et al. , 1981);PALJ810108 Alpha helicopter in Alpha + Beta class Standardized frequency of serotonin-dependent serotonin-dependent mutations (Palau et al., 1981); PALJ810109 Normalized frequency of alpha helices in the alpha / beta class (Palau et al., 1981);PALJ810110 Beta sheets in all beta classes Standardized frequency of (Palau et al., 1981); PALJ810111 Al Normalized frequency of beta sheets in the pha + beta class (Palau et al., 1981);PALJ810112 Beta sheet standard in the alpha / beta class Normalized frequency (Palau et al., 1981); PALJ810113 Total Alf Normalized frequency of turns in the class (Palau et al., 1981); PA LJ810114 Normalized frequency of turns across all beta classes (Palau et al. l., 1981);PALJ810115 Turns in the Alpha + Beta Class Standardized frequency (Palau et al., 1981); PALJ810116 Alf Standardized frequency of turns in the a / beta class (Palau et al., 1981 );PARJ860101 HPLC parameters (Parker et al., 1 986);PLIV810101 partition coefficient (Pliska et al., 1981 );PONP800101 Surrounding hydrophobicity in the folded form (Ponnuswamy et al., 1980);PONP800102 Average increase in ambient hydrophobicity (Ponnuswamy et al., 1980);PONP800103 Average increase ratio in aqueous solution (Ponnuswamy et al., 1980); PO NP800104 Surrounding hydrophobicity in alpha helix (Ponnuswamy e t al., 1980);PONP800105 Surrounding hydrophobicity in the beta sheet ( Ponnuswamy et al., 1980);PONP800106 turn Ambient hydrophobicity in (Ponnuswamy et al., 1980); PONP80 0107 Reachability reduction ratio (Ponnuswamy et al., 1980); PON P800108 Average number of surrounding residues (Ponnuswamy et al., 1980 );PRAM820101 Intercept in regression analysis (Prabhakaran-Ponn uswamy, 1982);PRAM820102 Regression analysis x 1.0E1 Slope (Prabhakaran-Ponnuswamy, 1982); PRAM820 103 Correlation coefficient in regression analysis (Prabhakaran-Ponnuswamy, 1982);PRAM900101 Hydrophobic (Prabhakaran, 1990) ;PRAM900102 Relative frequency of in alpha helices (Prabhaka ran, 1990); PRAM900103 relative frequency in beta sheet (Pr abhakaran, 1990);PRAM900104 Relative frequency of reverse turns degree (Prabhakaran, 1990);PTIO830101 Helix coil Equilibrium constant (Ptitsyn-Finkelstein, 1983); PTIO8301 02 Beta-coil equilibrium constant (Ptitsyn-Finkelstein, 1983) ;QIAN880101 -6 window position for alpha helix QIAN880102 -5 Weights for alpha helices at window positions (Qian-Sejnowsk i, 1988);QIAN880103 Alpha helices at window positions of -4 Weights on the basis of (Qian-Sejnowski, 1988); QIAN880 The weights for the alpha helix at the window position of 104-3 (Qian- Sejnowski, 1988);QIAN880105 -2 window position Weights for alpha helices (Qian-Sejnowski, 1988) ;QIAN880106 -1 window position for alpha helix Qian-Sejnowski, 1988; QIAN880107 0 Weights for alpha helices at the wind position (Qian-Sejnowski , 1988);QIAN880108 1 window position in the alpha helix Weights (Qian-Sejnowski, 1988);QIAN8801 09 2 window position weights for alpha helices (Qian-Se jnowski, 1988);QIAN880110 3 window positions Alf Weights for helices (Qian-Sejnowski, 1988); QI AN880111 Weight (Q) for alpha helix at window position 4 ian-Sejnowski, 1988);QIAN880112 5 window positions The weight for the alpha helix in the position (Qian-Sejnowski, 19 88);QIAN880113 for alpha helix at window position 6 Weight (Qian-Sejnowski, 1988); QIAN880114 -6 The weights for the beta sheet at the window position (Qian-Sejnowski , 1988); QIAN880115 -5 window position for beta sheet Hand weight (Qian-Sejnowski, 1988); QIAN880116 -4 window position weight for beta sheet (Qian-Sejnows ki, 1988);QIAN880117 -3 window position in the beta sheet Weights (Qian-Sejnowski, 1988); QIAN88011 8 -2 window position weights for beta sheets (Qian-Sejno wski, 1988);QIAN880119 Beta at window position -1 Weights for the metric (Qian-Sejnowski, 1988); QIAN880 Weights for beta sheets at window positions of 120 0 (Qian-Sejn owski, 1988);QIAN880121 Beta at 1 window position Weights for the metric (Qian-Sejnowski, 1988); QIAN880 122 Weights for beta sheets at window positions of 2 (Qian-Sejn owski, 1988);QIAN880123 Beta at 3 window positions Weights for the metric (Qian-Sejnowski, 1988); QIAN880 124 Weights for beta sheets at 4 window positions (Qian-Sejn owski, 1988);QIAN880125 Beta at 5 window positions Weights for the metric (Qian-Sejnowski, 1988); QIAN880 126 Weights for beta sheets at 6 window positions (Qian-Sejn owski, 1988);QIAN880127 -6 window position coil Weights (Qian-Sejnowski, 1988); QIAN88012 8 -5 window position coil weights (Qian-Sejnowsk i, 1988);QIAN880129 -4 window position coil Weight (Qian-Sejnowski, 1988); QIAN880130 -3 The weights for the coils at the window positions (Qian-Sejnowski, 1 988);QIAN880131 -2 Weight for coil at window position (Qian-Sejnowski, 1988);QIAN880132 -1 Win Weights for coils in dough position (Qian-Sejnowski, 1988) ;QIAN880133 0 window position coil weight (Qian -Sejnowski, 1988);QIAN880134 1 window position Weights on coils (Qian-Sejnowski, 1988); QIAN8 80135 Weights for coils at window positions 2 (Qian-Sejno wski, 1988);QIAN880136 3 window position coil Hand weight (Qian-Sejnowski, 1988); QIAN880137 Weights for coils at window positions of 4 (Qian-Sejnowski, 1988);QIAN880138 Weight for coil at window position 5 (Qian-Sejnowski, 1988);QIAN880139 6 Wind Weights for coils in the U position (Qian-Sejnowski, 1988); RACS770101 Average conversion distance for C-alpha (Rackovsky- Scheraga, 1977);RACS770102 Average reduced distance for side chains (Rackovsky-Scheraga, 1977);RACS770103 side chain Orientational selection (Rackovsky-Scheraga, 1977); RACS82 Average relative fragment abundance in A0(i) (Rackovsky-Sc heraga, 1982);RACS820102 Mean relative value in AR(i) the presence of fragments (Rackovsky-Scheraga, 1982); RACS82 Average relative fragment abundance (Rackovsky-Sc) in AL(i) heraga, 1982);RACS820104 EL(i) average relative the presence of fragments (Rackovsky-Scheraga, 1982); RACS82 Average relative fragment abundance in E0(i) (Rackovsky-Sc heraga, 1982);RACS820106 ER(i) average relative the presence of fragments (Rackovsky-Scheraga, 1982); RACS82 Average relative fragment abundance in A0(i-1) (Rackovsky- S cheraga, 1982);RACS820108 Average of AR(i-1) the relative presence of fragments (Rackovsky-Scheraga, 1982); RAC S820109 Average relative fragment abundance in AL(i-1) (Rackovs ky-Scheraga, 1982);RACS820110 EL(i-1) The presence of average relative fractions (Rackovsky-Scheraga, 1982) Average relative fragment abundance in RACS820111 E0(i-1) (Rac kovsky-Scheraga, 1982); RACS820112 ER(i-1 ) the average relative fractional presence (Rackovsky-Scheraga, 1 982);RACS820113 Theta(i) value (Rackovsky-Scher aga, 1982);RACS820114 Theta(i-1) value (Rackovs ky-Scheraga, 1982);RADA880101 chx~wat movement Free energy (Radzicka-Wolfenden, 1988); RADA88 0102 oct~wat transfer free energy (Radzicka-Wolfende n, 1988);RADA880103 vap~chx transfer free energy (Ra dzicka-Wolfenden, 1988);RADA880104 chx~o ct transfer free energy (Radzicka-Wolfenden, 1988); R ADA880105 Transfer free energy of vap~oct (Radzicka-Wol fenden, 1988); RADA880106 Nearest Envelope Area (Radzick a-Wolfenden, 1988);RADA880107 Energy from outside to inside -Transfer (95% embedded) (Radzicka-Wolfenden, 1988); RAD A880108 Average polarity (Radzicka-Wolfenden, 1988);R Relative selectivity in ICJ880101 N” (Richardson-Richard son, 1988);RICJ880102 Relative selection value at N' (Richar dson-Richardson, 1988);RICJ880103 N-cap Relative selection value of (Richardson-Richardson, 1988); RI CJ880104 Relative selectivity at N1 (Richardson-Richards on, 1988); RICJ880105 Relative selection value in N2 (Richard son-Richardson, 1988);RICJ880106 Relative to N3 Selection Value (Richardson-Richardson, 1988);RICJ88 Relative selectivity at N4 (Richardson-Richardson, 1988);RICJ880108 Relative selection value in N5 (Richardson- Richardson, 1988);RICJ880109 Relative selection in Mid Value (Richardson-Richardson, 1988);RICJ88011 0 Relative selection value at C5 (Richardson-Richardson, 198 8);RICJ880111 Relative selectivity in C4 (Richardson-Ric hardson, 1988); RICJ880112 Relative selection value (Ri chardson-Richardson, 1988);RICJ880113 C2 Relative selection value at (Richardson-Richardson, 1988); R Relative selectivity in ICJ880114 C1 (Richardson-Richard son, 1988);RICJ880115 Relative selectivity in C-cap (Ric hardson-Richardson, 1988);RICJ880116 C' Relative selection value of (Richardson-Richardson, 1988); RI Relative selectivity at C'' (Richardson-Richard son, 1988);ROBB760101 Information measurement of alpha helices value (Robson-Suzuki, 1976); ROBB760102 N-terminal helix Information Measures for Cox (Robson-Suzuki, 1976); ROBB76 Information measures for the central helix (Robson-Suzuki, 1 976); ROBB760104 Information measurement for the C-terminal helix (Robso n-Suzuki, 1976);ROBB760105 Information measurement of expansion ( Robson-Suzuki, 1976);ROBB760106 Pleated sheet Information measures about (Robson-Suzuki, 1976); ROBB76010 7 Information measures for extensions without H-bonds (Robson-Suzuki, 19 76);ROBB760108 Information measurements about turns (Robson-Suzu ki, 1976);ROBB760109 Information measurement of N-terminal turn (Ro bson-Suzuki, 1976);ROBB760110 About the central turn Information measurement (Robson-Suzuki, 1976);ROBB760111 C Information measure for terminal turns (Robson-Suzuki, 1976); ROB B760112 Coil Information Measurements (Robson-Suzuki, 197 6);ROBB760113 Information Measures for Loops (Robson-Suzuk i, 1976);ROBB790101 Hydration free energy (Robson-Osg uthorpe, 1979);ROSG850101 Mean variance embedded by migration Area (Rose et al., 1985); ROSG850102 Average fragment area Loss of solvation (Rose et al., 1985); ROSM880101 Unmodified side chain hydropathic (Roseman, 1988); ROSM880102 Side-chain hydropathy corrected for solvation (Roseman, 1988); ROSM 880103 Loss of side chain hydropathy due to helix formation (Roseman, 19 88);SIMZ760101 by Charton-Charton (1982) Transfer free energy cited (Simon, 1976); SNEP660101 main Component I (Sneath, 1966); SNEP660102 Principal component II (Sneath h, 1966);SNEP660103 Principal component III (Sneath, 1966) ;SNEP660104 Principal component IV (Sneath, 1966);SUEM8401 01 Zimm-Bragg parameters at 20C (Sueki et al., 1 984);SUEM840102 Zimm-Bragg parameter sigma x 1. 0E4(Sueki et al., 1984);SWER830101 best match Hydrophobicity (Sweet-Eisenberg, 1983); TANS770101 Normalized frequency of alpha helices (Tanaka-Scheraga, 1977); T ANS770102 Normalized frequency of isolated helices (Tanaka-Scheraga, 1977);TANS770103 Standardized frequency of extended structure (Tanaka-Sche raga, 1977);TANS770104 Normalized frequency of strand inversion R (Tanaka -Scheraga, 1977);TANS770105 Normalized frequency of strand inversion S (T anaka-Scheraga, 1977);TANS770106 Standard for chain inversion D Frequency of neurogenesis (Tanaka-Scheraga, 1977); TANS770107 left gyrus Normalized frequency of helices (Tanaka-Scheraga, 1977); TAN S770108 Zeta R normalized frequency (Tanaka-Scheraga, 1977 );TANS770109 Standardized Frequency of Coil (Tanaka-Scheraga, 1977), TANS770110 normalized frequency of strand inversion (Tanaka-Schera ga, 1977);VASM830101 Relative occupancy of conformational state A (Vas quez et al., 1983);VASM830102 relative conformational state C occupancy rate (Vasquez et al., 1983); VASM830103 Relative occupancy of body structural state E (Vasquez et al., 1983); VEL V850101 electron-ion interaction potential (Veljkovic et al. , 1985);VENT840101 bitterness(Venanzi, 1984);VHE G790101 Transfer free energy to lipophilic phase (von Heijne-Blomb erg, 1979);WARP780101 Average interactions per side chain atom (War me-Morgan, 1978);WEBA780101 for high salt chromatography RF value in (Weber-Lacey, 1978); WERD780101 inside Tendency to be embedded (Wertz-Scheraga, 1978); WERD780102 Free energy change from epsilon(i) to epsilon(ex) (Wertz-Sche raga, 1978);WERD780103 Alpha (Ri) ~ Alpha (Rh) Free energy change (Wertz-Scheraga, 1978); WERD780 104 Free energy change from epsilon(i) to alpha(Rh) (Wertz-Sc heraga, 1978);WOEC730101 Polarity Requirement (Woese, 19 73);WOLR810101 hydration potential (Wolfenden et al. , 1981);WOLS870101 Primary attribute value z1(Wold et al., 1987);WOLS870102 Primary attribute value z2 (Wold et al., 1987);WOLS870103 Primary attribute value z3 (Wold et al., 1 987);YUTK870101 Gibbs energy of denaturation in water, pH 7.0 (Yu Tani et al., 1987); YUTK870102 in water, pH 9.0 The modified Gibbs energy (Yutani et al., 1987); YUTK870 103 Denaturation, Gibbs energy of activation at pH 7.0 (Yutani et al., 1987);YUTK870104 Denaturation, Gibbs energy of activation at pH 9.0 (Yu Tani et al., 1987);ZASB820101 Partition coefficient for ionic strength Number dependence (Zaslavsky et al., 1982); ZIMJ680101 Hydrophobic (Zimmerman et al., 1968); ZIMJ680102 or Sabari (Zimmerman et al., 1968); ZIMJ680103 Polar (Zimmerman et al., 1968); ZIMJ680104 Isoelectric point (Zimmerman et al., 1968); ZIMJ680105 RF Run (Zimmerman et al., 1968);AURR980101 Helic Normalized positional residue frequencies at the N4' end of the base (Aurora-Rose, 1998); A URR980102 Normalized positional residue frequency at helix terminal N'' (Aurora- Rose, 1998);AURR980103 normalized positional alignment at the helix terminal N Residue frequency (Aurora-Rose, 1998);AURR980104 helix Normalized positional residue frequency at terminal N' (Aurora-Rose, 1998); AURR 980105 Normalized positional residue frequency at helix terminal Nc (Aurora-Rose , 1998);AURR980106 Normalized positional residue frequencies at helix terminal N1 (Aurora-Rose, 1998);AURR980107 at helix terminal N2 Normalized positional residue frequencies (Aurora-Rose, 1998); AURR98010 8 Normalized positional residue frequency at helix terminal N3 (Aurora-Rose, 199 8 );AURR980109 Normalized positional residue frequency at helix terminal N4 (Auror a-Rose, 1998);AURR980110 normalized position at helix terminal N5 positional residue frequency (Aurora-Rose, 1998); AURR980111 Normalized positional residue frequencies at the terminal C5 of the auxin (Aurora-Rose, 1998); AU RR980112 Normalized positional residue frequency at helix terminal C4 (Aurora-Ro se, 1998);AURR980113 Standard positional residues at helix terminal C3 Frequency (Aurora-Rose, 1998);AURR980114 Helix end Normalized positional residue frequency at C2 (Aurora-Rose, 1998); AURR98 Normalized positional residue frequencies at helix terminal C1 (Aurora-Rose, 1998);AURR980116 Normalized positional residue frequency (A urora-Rose, 1998);AURR980117 at the C' end of the helix Standardized positional residue frequencies (Aurora-Rose, 1998); AURR980118 Normalized positional residue frequencies at helix-terminal C's (Aurora-Rose, 1998 );AURR980119 Normalized positional residue frequency at helix terminal C'' (Auro ra-Rose, 1998);AURR980120 Standard at C4' of helix terminal positional residue frequency (Aurora-Rose, 1998);ONEK900101 0 Delta G values for predicted peptides adjusted for M urea (O'Neil-DeGra do, 1990);ONEK900102 Helix formation parameters (Delta de (O'Neil-DeGrado, 1990); VINM940101 Standard Mobility parameter (B-value), mean (Vihinen et al., 1994) ;VINM940102 Target for each residue not surrounded by strong neighbors Standardized mobility parameter (B-value) (Vihinen et al., 1994); V INM940103 Target for each residue surrounded by one strong neighbor Standardized mobility parameter (B-value) (Vihinen et al., 1994); V INM940104 Target for each residue surrounded by two strong neighbors Standardized mobility parameter (B-value) (Vihinen et al., 1994); M UNV940101 Free energy in alpha helix conformation (Munoz -Serrano, 1994);MUNV940102 in the alpha helix region Free energy in (Munoz-Serrano, 1994); MUNV94010 3 Free energy in beta-strand conformation (Munoz-Serrano, 1994);MUNV940104 Free energy in the beta-strand region ( Munoz-Serrano, 1994);MUNV940105 beta strand Free energy in the region (Munoz-Serrano, 1994), WIMW9 60101 Free energy of transfer of AcWl-X-LL peptide from the bilayer interface to water (Wimley-White, 1996);KIMC930101 Thermodynamic Betacy (Kim-Berg, 1993); MONM990101 transmembrane helix Turning Tendency Scale (Monne et al., 1999); BLAM93 Alpha-helical propensity of position 44 in T4 lysozyme (Blaber et al., 1993);PARS000101 Mesophilic tannins based on the distribution of B values p-value of the quality (Parthasarathy-Murthy, 2000); PARS0 p-values of thermophilic proteins based on the distribution of B-values (Parthasarath y-Murthy, 2000);KUMS000101 18 thermophilic proteins Distribution of amino acid residues in non-redundant families ( Kumar et al., 2000 );KUMS000102 Amino acids in 18 non-redundant families of mesophilic proteins Distribution of acid residues (Kumar et al., 2000); KUMS000103 Distribution of amino acid residues in alpha helices in thermogenic proteins (Kumar et al., 2000);KUMS000104 A study of mesophilic proteins Distribution of amino acid residues in the ruffa helix ( Kumar et al., 2000 );TAKK010101 Side chain contribution to protein stability (kJ / mol) (Taka no-Yutani, 2001);FODM020101 Amino acid in the pi helix Acid tendency (Fodje-Al-Karadaghi, 2002);NADH01010 1. Hydropathy scale based on self-information value in the two-state model (5% attainability) Naderi-Manesh et al., 2001);NADH010102 II Hydropathy scale based on self-information value in the state model (9% attainability) (Nad eri-Manesh et al., 2001);NADH010103 Two-state model Self-Information Value-Based Hydropathy Scale (16% attainable) (Nader i-Manesh et al., 2001);NADH010104 two-state model Self-information-based hydropathy scale (20% attainable) (Naderi- Manesh et al., 2001);NADH010105 In the two-state model Self-information-based hydropathy scale (25% attainable) (Naderi-Ma nesh et al., 2001);NADH010106 in the two-state model Self-Information Value-Based Hydropathy Scale (36% attainable) (Naderi-Mane sh et al., 2001);NADH010107 Self in a two-state model Information-based hydropathy scale (50% attainable) (Naderi-Manesh et al., 2001);MONM990201 average in transmembrane helices Turn tendency (Monne et al., 1999); KOEP990101 designed Alpha helix propensity derived from the designed sequence; KOEP990102 Derived beta-sheet propensity; CEDJ970101 Amino acids in extracellular proteins Composition (percent) (Cedano et al., 1997); CEDJ9701 02 Amino acid composition (percentage) of immobilized proteins (Cedano e t al., 1997);CEDJ970103 Amino acid combinations in membrane proteins Composition (percent) (Cedano et al., 1997); CEDJ970104 Percent amino acid composition of intracellular proteins (Cedano et al. ., 1997);CEDJ970105 Amino acid composition of nuclear proteins (part cent) (Cedano et al., 1997);FUKS010101 Thermophilic bacteria Surface composition (percentage) of amino acids in intracellular proteins (Fukuchi-Ni shikawa, 2001);FUKS010102 Intracellular proteins of mesophilic bacteria Surface composition (percentage) of amino acids in the 01);FUKS010103 Amino acid surface composition of extracellular proteins of mesophilic bacteria (Percent)(Fukuchi-Nishikawa, 2001);FUKS010 104 Surface composition of amino acids in nuclear proteins (percent) (Fukuchi-N ishikawa, 2001);FUKS010105 Intracellular proteins of thermophilic bacteria Internal composition (percentage) of amino acids in the cereals (Fukuchi-Nishikawa, 2 001);FUKS010106 Internal amino acid composition of intracellular proteins of mesophilic bacteria Composition (percent) (Fukuchi-Nishikawa, 2001);FUKS01 Internal composition (percentage) of amino acids in extracellular proteins of mesophilic bacteria (F ukuchi-Nishikawa, 2001);FUKS010108 Nuclear protein Internal composition (percentage) of amino acids in proteins (Fukuchi-Nishikawa, 2001);FUKS010109 Total amino acids in intracellular proteins of thermophilic bacteria Chain composition (percent) (Fukuchi-Nishikawa, 2001); FUKS Total amino acid chain composition (percentage) of intracellular proteins of mesophilic bacteria (Fukuchi-Nishikawa, 2001);FUKS010111 Mesophilic bacteria Total amino acid chain composition (percentage) of extracellular proteins of shikawa, 2001);FUKS010112 Amino acids in nuclear proteins Total chain composition (percent) of (Fukuchi-Nishikawa, 2001); AV BF000101 Screening factor gamma, local (Avbelj, 2000); AVBF000102 Screening coefficient gamma, non-local (Avbelj, 200 0);AVBF000103 Slope Tripeptide, FDPB VFF Neutral (Avbel j, 2000);AVBF000104 Slope Tripeptide, LD VFF Neutral (A vbelj, 2000);AVBF000105 Slope tripeptide, FDPB VF F Noside (Avbelj, 2000);AVBF000106 Slope Tripeptide Do FDPB VFF All (Avbelj, 2000);AVBF000107 Slope tripeptide FDPB PARSE neutral (Avbelj, 2000); AVB F000108 Sloped decapeptide FDPB VFF Neutral (Avbelj, 200 0);AVBF000109 Slope protein, FDPB VFF Neutral (Avbelj , 2000);YANJ020101 Side chain stereochemistry by Gaussian evolution Structure (Yang et al., 2002); MITS020101 Amphiphilicity Index (M itaku et al., 2002);TSAJ990101 ProtOr used The volume including the water of crystallization (Tsai et al., 1999); TSAJ99010 2 Using ProtOr, volume without water of crystallization; COSI940101 electron ion Interaction potential value (Cosic, 1994); PONP930101 Hydrophobicity Scale (Ponnuswamy, 1993); WILM950101 0.1%TF RP-HPLC with A / MeCN / HO, hydrophobicity index (Wilce et al. 1995);WILM950102 0.1%TFA / MeCN / H2 RP-HPLC using O, hydrophobicity index in C8 (Wilce et al. 19 95);WILM950103 RP-HP using 0.1% TFA / MeCN / H2O LC, hydrophobicity coefficient at C4 (Wilce et al. 1995); WILM95 RP-HPLC with 0.1% TFA / 2-PrOH / MeCN / HO; Hydrophobicity coefficient for C18 (Wilce et al. 1995); KUHL9501 01 Hydrophilic scale (Kuhn et al., 1995); GUOD860101 Retention factor at pH 2 (Guo et al., 1986); JURD980101 Modified Kyte-Doolittle hydrophobicity scale (Juretic et al., 1998);BASU050101 Interaction scale obtained from contact matrix ( Bastolla et al., 2005);BASU050102 single domain The interaction sequence obtained by maximizing the average correlation coefficient for globular proteins is Kale (Bastolla et al., 2005); BASU050103 TI Maximizing the average correlation coefficient for pairs of sequences that share an M barrel fold The interaction scale obtained by (Bastolla et al., 2005); S U YM030101 Linker tendency index (Suyama-Ohara, 2003);PU NT030101 Information base from 1d_Helix in the MPtopo database Film Tendency Scale (Punta-Maritan, 2003); PUNT03010 2. Information-based membrane tendency scheme derived from 3D_Helix in the MPtopo database (Punta-Maritan, 2003);GEOR030101 All data Linker trends from datasets (George-Heringa, 2003); GEO R030102 Linker trends from the 1-linker dataset (George-Herman inga, 2003); GEOR030103 2-Linker dataset Carr tendency (George-Heringa, 2003);GEOR030104 3- From the Linker dataset (George-Heringa, 2003); GEOR 030105 Linker propensity (G) from a small dataset (linker length less than 6 residues) eorge-Heringa, 2003);GEOR030106 Medium data set Linker tendencies (George-Heringa et al., 2001) from the set (linker length 6–14 residues) , 2003);GEOR030107 long dataset (linker length > 14 residues) Linker tendency from (George-Heringa, 2003); GEOR0 Linker propensity (G) from the 30108 helix (annotated by DSSP) dataset eorge-Heringa, 2003);GEOR030109 Non-helix (D Linker trends from the SSP (annotated by George-Heringa) dataset , 2003);ZHOH040101 Stability of information-based atom-atom potentials Scale (Zhou-Zhou, 2004); ZHOH040102 Mutation experiment? The relative stability scale extracted from (Zhou-Zhou, 2004); ZHOH0 40103 Buriability (Zhou-Zhou, 2004); BAEK050101 Linker Index (Bae et al., 2005);HARY 940101 Average volume of buried residues inside proteins (Harpaz et al. ., 1994);PONJ960101 average volume of residues (Pontius et al. l., 1996);DIGM050101 Hydrostatic pressure asymmetry index, PAI(Di Giu lio, 2005);WOLR790101 Hydrophobicity Index (Wolfenden et al., 1979);OLSK800101 Mean internal selection (Olsen, 198 0);KIDA850101 Hydrophobicity-related index (Kidera et al., 198 5);GUYH850102 Apparent calculated from the Wertz-Scheraga index Above distribution energy (Guy, 1985);GUYH850103 Robson-O Apparent partitioning energy calculated from the Guthorpe index (Guy, 1985 );GUYH850104 Apparent partition energy calculated from Janin index ( Guy, 1985);GUYH850105 The appearance calculated from the Chothia index Partition energy on the chain (Guy, 1985); ROSM880104 amino acid side chain, Neutral morphological hydropathic (Roseman, 1988); ROSM880105 Amis Hydropathy of the acid side chain, pi value at pH 7.0 (Roseman, 1988); JACR890101 Weights from the IFH scale (Jacobs-White, 1989); COWR900101 Hydrophobicity Index, pH 3.0 (Cowan-Whitt aker, 1990), BLAS910101 side chain hydrophobicity value (Black-Moul d, 1991);CASG920101 Hydrophobicity scale from native protein structures (Casari-Sippl, 1992);CORJ870101 NNEIG indicator( Cornette et al., 1987);CORJ870102 SWEIG finger Mark (Cornette et al., 1987);CORJ870103 PRIF T index (Cornette et al., 1987); CORJ870104 PR ILS index (Cornette et al., 1987); CORJ870105 ALTFT index (Cornette et al., 1987), CORJ87010 6 ALTLS index (Cornette et al., 1987); CORJ870 107 TOTFT index (Cornette et al., 1987); CORJ8 70108 TOTLS indicator (Cornette et al., 1987); MIY S990101 Relative partition energy derived by the Bethe approximation (Miyaz awa-Jernigan, 1999);MIYS990102 Optimized relative distribution equation Energy - Method A (Miyazawa-Jernigan, 1999); MIYS99 Optimized Relative Partition Energy - Method B (Miyazawa-Jernigan , 1999);MIYS990104 Optimized relative distribution energy - Method C(Miy azawa-Jernigan, 1999);MIYS990105 Optimized relative fraction Energy Distribution Method D (Miyazawa-Jernigan, 1999); ENGD 860101 Hydrophobicity Index (Engelman et al., 1986); and FASG890101 may include the hydrophobicity index (Fasman, 1989).
[0107] In some embodiments of the invention, the degenerate oligonucleotide is one or more of the oligonucleotides of the invention. used to synthesize multiple TN1, DH, N2, and / or H3-JH segments In one embodiment of the present invention, an oligonucleotide encoding an H3-JH segment is A codon at or near the 5' end of a given sequence is a degenerate codon. , the first codon from the 5' end, the second codon from the 5' end, the third codon from the 5' end the fourth codon from the 5' end, the fifth codon from the 5' end, and / or the above In some embodiments of the invention, the DH segment may be any combination of One or more codons at or near the 5' and / or 3' end of Such degenerate codons are the first codons from the 5' and / or 3' ends. , the second codon from the 5' and / or 3' end, the third codon from the 5' and / or the fourth codon from the 3' end, or the fifth codon from the 3' end, and / or any combination of the above The degeneracy used in each of the oligonucleotides encoding the segments may be Codons are identified by the sequences in the theoretical segment pool and / or CDRH3 reference set. They may be selected for their ability to reproduce optimally.
[0108] In some embodiments, the present invention provides a method for detecting H3-JH receptors, as described in the Examples. In Example 5, a method for generating a theoretical segment pool of fragments is provided. Utilizing NNN triplets instead of or in addition to the NN doublets described The theoretical segment pools generated by the Some implementations are within the scope of the present invention, such as synthetic libraries incorporating fragments. In one embodiment, the present invention provides a method for generating a theoretical segment of a DH segment, as described in the Examples. In particular, for example, the present invention provides a method for producing a soluble fraction pool by the method of Example 6. The theoretical segment pool of DH segments described by the YTHON program is generated. Example 6 provides a method for generating a 68K theoretical segment pool. The application of this program is described below. (Minimum length of DNA sequence after stepwise deletion = 4 bases) and the minimum length of a peptide sequence for inclusion in the theoretical segment pool = 2). The minimum length of the DNA sequence after stepwise deletion is one base, and the minimum length of the peptide sequence is one amino acid. An alternative embodiment is provided in which other values can be used for these parameters. For example, the minimum length of the DNA sequence after stepwise deletion is 1, Set to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. The minimum length of the peptide sequence in the theoretical segment pool can be 1, 2, 3, It can be set to 4 or 5.
[0109] of the CDRH3 library using the TN1, DH, N2, and H3-JH segments. design The CDRH3 library of the present invention contains TN1, DH, N2, and H3-JH segments. Thus, in one embodiment of the present invention, the entire CDRH3 library The design can be expressed by the following formula: [TN1]-[DH]-[N2]-[H3-JH].
[0110] In one embodiment of the invention, the synthetic CDRH3 repertoire comprises a VH chassis of choice. The sequences and heavy chain constant regions are combined via homologous recombination. In one embodiment, a synthetic CDRH3 library and selected chassis and constant regions are used. To promote homologous recombination between vectors containing the CDRH3 region, a synthetic CDRH3 library It may be desirable to include DNA sequences flanking the 5' and 3' ends of the In embodiments, the vector also contains the uncleaved region of the IGHJ gene (i.e., FRM4-JH). The N-terminal sequence contains a sequence encoding at least a portion of the region. The loading polynucleotide (e.g., CA(K / R / T)) is added to the synthetic CDRH3 sequence. and the N-terminal polynucleotide is homologous to FRM3 of the chassis, while A polynucleotide encoding a C-terminal sequence (e.g., WG(Q / R / K)G) is synthesized using synthetic C The C-terminal polynucleotide may be added to DRH3 and be homologous to FRM4-JH. The sequence WG(Q / R)G is shown in this exemplary embodiment, but FRM4-J Additional amino acids C-terminal to this sequence in H also encode the C-terminal sequence The polynucleotide may contain N-terminal and C-terminal sequences. The purpose of the nucleotides in this case is to facilitate homologous recombination, and those skilled in the art will appreciate that these It will be appreciated that the sequences may be longer or shorter than those shown below. Therefore, in one embodiment of the present invention, a method for promoting homologous recombination with a selected chassis is provided. The overall design of the CDRH3 repertoire, including the sequences required for transcription, is given by the following formula: (regions of homology with the vector are underlined): CA[R / K / T] -[TN1]-[DH]-[N2]-[H3-JH]- [WG(Q / R / K)G] .
[0111] In some embodiments of the invention, the CDRH3 repertoire is represented by the formula: which, excluding the T residues shown in the blueprint above, can be: CA[R / K] -[TN1]-[DH]-[N2]-[H3-JH]- [WG(Q / R / K)G] .
[0112] Each reference describing a collection of V, D, and J genes is incorporated herein by reference in its entirety. Incorporated by Scaviner et al., Exp. Clin. mmunogenet., 1999, 16: 243 and Ruiz et al. , Exp. Clin. Immunogenet, 1999, 16: 173 Homologous recombination is one method for generating the libraries of the present invention, but those skilled in the art The present inventors have also performed DNA constructions and / or Other methods of DNA synthesis can be used to generate the libraries of the present invention. You will easily recognize that you can.
[0113] CDRH3 length The length of the segments can also be varied, for example, to provide a library with a particular distribution of CDRH3 lengths. In one embodiment of the present invention, the H3-JH segment may be diversified to produce is about 0 to about 10 amino acids in length, and the DH segment is about 0 to about 12 amino acids in length. The TN1 segment is about 0 to about 4 amino acids in length, and the N2 segment is about 0 to about 4 amino acids in length. In one embodiment, the H3-JH segment is at least about 0, 1, 2 amino acids in length. , 3, 4, 5, 6, 7, 8, 9, and / or 10 amino acids in length. In the form, the DH segment has at least about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, In one embodiment, the TN1 segment is 10, 11, and / or 12 amino acids in length. In some embodiments, the fragment is at least about 0, 1, 2, 3, or 4 amino acids in length. In some embodiments, the N2 amino acid is at least about 0, 1, 2, 3, or 4 amino acids in length. In certain embodiments, CDRH3 has a molecular weight of about 2 to about 35, about 2 to about 28, or about 5 to about 26 In some embodiments, the CDRH3 is at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 3 In some embodiments, the nucleic acids of the present invention are 3, 34, and / or 35 amino acids in length. The length of either the CDRH3 or CDRH4 fragment may be less than a certain number of amino acids. The number of amino acids is the integer provided above for each segment or CDRH3. In some embodiments of the present invention, the range of specific numbers is Any of the integers provided above as the lower and upper limits of an inclusive or exclusive range The upper and lower bounds are determined using the combination of integers provided. Everything is planned.
[0114] CDRL3 library design The design of the CDRL3 library and the light chain sequences are described in U.S. Patent Application Publication No. 2009 / 014494. 81855 and 2010 / 0056386 and International Publication No. WO / 2009 / 036379, each of which is incorporated by reference in its entirety. The invention is incorporated herein and is therefore only briefly described herein. The resulting libraries are designed according to similar principles, but with three important differences: That is, the library of the present invention has (1) diversity in CDRL1 and CDRL2, (2) diversity in CDRL3 and CDRL4, ) diversity in the framework regions, and / or (3) as defined above , a light chain library with a CDRL3 closely resembling a human germline-like CDRL3 sequence The nucleotide sequences contain diversity in CDRL3 (Table 1) designed to generate:
[0115] The CDRL3 library of the present invention may be a VKCDR3 library and / or a VλC In one embodiment of the present invention, a defined VL sequence may be a DR3 library. The pattern of occurrence of specific amino acids at positions is based on publicly available data or other data. It is determined by analyzing a database, e.g., the NCBI database (e.g., (See International Publication No. WO / 2009 / 036379). In one embodiment of the present invention These sequences are compared based on identity and matched to the germline genes from which they were derived. The sequence is then assigned to a family. The amino acid composition at each position may be determined. This process is referred to herein as This is illustrated in the examples provided below.
[0116] Light chains with framework diversity In some embodiments, the present invention provides a method for the preparation of a light chain variable domain comprising administering to a mammalian subject the invention in one or more frame of the light chain variable domain, which differ at work positions 2, 4, 36, 46, 48, 49, and 66 In some embodiments, the present invention provides libraries comprising one or more amino acid sequences. are identical to each other except for substitutions at positions 2, 4, 36, 46, 48, 49, and 66 a library of light chain variable domains comprising at least a plurality of light chain variable domains, In one embodiment, the present invention provides a method for the preparation of a medicament for the treatment ... At least about 70%, 75%, 80%, 85%, 90%, or 100% of any of the light chain variable domain sequences %, 95%, 96%, 97%, 98%, 99%, and / or 99.5%, and one or further having substitutions at positions 2, 4, 36, 46, 48, 49, and 66; A library of light chain variable domains is provided, comprising at least a plurality of light chain variable domains. In some embodiments, the amino acids selected for inclusion at these positions are at about 2, 3, 4, 5, 6, 7, 8, 9 at the corresponding positions in the reference set of chain variable domains , and / or selected from among the 10 most frequently occurring amino acids.
[0117] In some embodiments, the present invention provides a method for diversifying the light chain variable domain. A system and method for selecting framework positions comprising: (i) obtaining a reference set of light chain sequences, the reference set comprising a single IGVL gene Sequences found in or encoded by germline genes and / or is found in or among allelic variants of a single IGVL germline gene a light chain sequence having a VL segment selected from the group consisting of sequences encoded by containing, steps, (ii) which framework positions in the reference set are identical to one or more of the sequences in the reference set; have a degree of diversity similar to the degree of diversity present at multiple CDR positions. (e.g., diversity at framework positions is determined by comparing the sequences in the reference set) At least about 70%, 80%, 90%, or 90% of the diversity found in the CDR positions of 5%, 100%, or more), (iii) an amino acid sequence for each of the framework positions identified in (ii) determining the frequency of occurrence of acid residues; (iv) synthesizing a light chain variable domain coding sequence, The framework positions defined are 2, 3, 4, 5, 6, 7, 8, 9, 1 at the corresponding positions. 0, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 most frequently The system is diversified to include the present amino acid residues (identified in (iii)). and a method.
[0118] Those skilled in the art will appreciate that this disclosure provides a method for developing framework variants of heavy chain sequences. It will be appreciated that the present invention provides a similar method for
[0119] Light chains with CDR1 and / or CDR2 diversity In some embodiments, the present invention provides a method for the preparation of a light chain variable domain comprising administering to a mammalian subject the method of the present invention ... 1. Light chains differing at positions 28, 29, 30, 30A, 30B, 30E, 31, and 32 Providing a library of variable domains (Chothia-Lesk numbering scheme; Chothia and Lesk, J. Mol. Biol., 1987, 1 96:901). In some embodiments, the invention provides a method for treating a leukemia comprising administering to a mammalian subject the method ... The light chain variable domains differ at multiple CDRL2 positions 50, 51, 53, and 55. In some embodiments, these CDRL1 and / or C The amino acids selected for inclusion at the DRL2 position are selected from a reference set of light chain variable domains. At the corresponding positions in the , 14, 15, 16, 17, 18, 19, and / or 20 most frequently occurring Amis In some embodiments, the present invention provides a light chain variable domain comprising a nucleotide sequence selected from the group consisting of nucleotides 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, A system for selecting the CDRL1 and / or CDRL2 positions to be diversified and a method comprising: (i) obtaining a reference set of light chain sequences, the reference set comprising a single IGVL gene Sequences found in or encoded by germline genes and single Found in or due to an allelic variant of the IGVL germline gene a light chain sequence having a VL segment encoded by a sequence selected from the group consisting of , step, (ii) determining which CDRL1 and / or CDRL2 positions are diverse within the reference set; determining, (iii) synthesizing a light chain variable domain coding sequence, wherein in (ii), The identified CDRL1 and / or CDRL2 positions may be substituted with 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or diversified to include the 20 most frequently occurring amino acid residues. Those skilled in the art will appreciate that this disclosure provides CDRH2 and / or It will be appreciated that the present invention provides similar methods for developing CDRH2 variants. cormorant.
[0120] Light chain sequence In some embodiments, the present invention provides one or more of the methods provided herein. Light chain sequences, e.g., the polypeptide sequences of Table 3 and / or Table 4 and / or Table 5 a light chain library comprising any of the polynucleotide sequences of Table 1, Table 2, and / or Table 3; Those skilled in the art will appreciate that all of the light chain sequences provided herein are compatible with the functions of the present invention. It will be appreciated that a single light chain library may not necessarily be required to generate a functional light chain library. Therefore, in one embodiment, the light chain library of the present invention comprises the sequences described above. For example, in certain embodiments of the present invention, at least about 10 of the light chain polynucleotide and / or polypeptide sequences provided herein , 100, 200, 300, 400, 500, 600, 700, 800, 900, 10 3 , 10 4 , and / or 10 5 is included in the library. The libraries of the invention may contain less than a certain number of polynucleotide or polypeptide segments. The number of segments may be as provided above for each segment. In some embodiments of the present invention, the arithmetic unit is defined using any one of the integers Ranges may be expressed as any of the integers provided above as the lower and upper limits of the inclusive or exclusive range. The upper and lower bounds are determined by using any two of the provided integer combinations. All combinations are contemplated.
[0121] In certain embodiments, the present invention provides a method for the production of any of the sets of light chain sequences provided herein. at least about 1%, 2.5%, 5%, 10%, 15%, 20%, 25% of the sequences derived from 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, Light chain libraries containing 85%, 90%, 95%, or 99% of the total DNA fragments are provided. The invention provides at least one of the light chain sequences provided in Table 3, Table 4, Table 5, Table 6, and / or Table 7. At least about 1%, 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, In some embodiments of the present invention, libraries containing 95%, or 99% of the Certain percentage ranges may be used as lower and upper limits of inclusive or exclusive ranges. Defined using any two of the percentages provided above. Upper and lower bounds All combinations of the provided percentages are contemplated.
[0122] In some embodiments of the invention, at least about 1% of the light chain sequences in the library , 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50% , 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% are the light chain sequences provided herein. isolated from a library of antibodies (e.g., specific antigens and / or general ligands) by binding to at least about 1%, 2.5%, 5%, 10%, 15%, or %, 20%, 25%, 30%, 40%, 45%, 50%, 55%, 60%, 65%, 70 %, 75%, 80%, 85%, 90%, 95%, or 99% are as provided herein In some embodiments, the light chain library of the present invention is a light chain library of a particular The present invention also provides a method for the preparation of a light chain comprising administering to a mammalian subject the invention, ... a method for the preparation of a light chain comprising administering to a mammalian subject the invention, The percentage is defined using one of the percentages provided above. In some embodiments of the invention, the specific percentage ranges may be inclusive or exclusive ranges. Defined using any two of the percentages provided above as the lower and upper bounds of All combinations of percentages provided that define upper and lower limits are It is planned.
[0123] Those skilled in the art will appreciate that the overall level of the light chain sequences specified may be readily apparent in light of the light chain sequences provided herein. sequence identity and / or one or more of the characteristic sequences described herein Similar light chain sequences can be produced that share sequence elements, and the sequence identity can be determined by the Global extent and / or characteristic sequence elements may confer general functional attributes Those skilled in the art will further recognize that the mutagenesis methods provided herein are thoroughly familiar with the various techniques for preparing such related sequences, including techniques Therefore, each explicitly recited embodiment of the present invention will also be Light chain sequences that share a specified percent identity with any of the light chain sequences provided in the publication. For example, each of the above-described embodiments of the present invention can be implemented using the Embodiments include those having at least about 70%, 75%, 80%, 85%, 90%, 95%, 100%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 5%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 9 This can be done using light chain sequences that are 9%, 99.5%, or 99.9% identical. For example, in some embodiments, the light chain libraries provided by the present invention include , at least about 70%, 75%, 80%, 85% to the light chain sequences provided herein , 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% , 99.5%, or 99.9% identical light chain variable domains, Framework positions 2, 4, 36, 46, 48, 49, and 66, CDRL1 position 28 , 29, 30, 30A, 30B, 30E, 31, and 32 (Chothia-Lesk numbering scheme), and / or at CDRL2 positions 50, 51, 53, and 55 It has a substitution in
[0124] In some embodiments, the present invention provides a method for detecting a IgG1-associated IgG1 gene, which is encoded by a specific IGVL germline gene. 1. A system and method for diversifying positions within a portion of a CDRL3 comprising: (i) obtaining a reference set of light chain sequences, the reference set comprising sequences from the same IGVL genome; Light chains having VL segments derived from lineage genes and / or allelic variants thereof containing the sequence; (ii) that of the CDRL3 position in the reference set encoded by the IGVL gene Determining which amino acids are present at each position (i.e., positions 89-94, inclusive) ); (iii) synthesizing a light chain variable domain coding sequence, Two positions in each light chain variable domain are aligned to the corresponding positions in the reference set. , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 , degenerate codons encoding the 18, 19, or 20 most frequently occurring amino acid residues The present invention provides a system and method comprising the steps of:
[0125] As described in the Examples, the degenerate codons in (iii) are The amino acid diversity contained in the reference set for each of the two positions that differ is calculated as follows: Finally, the methods and systems described above can be used to Although described with respect to RL3, those skilled in the art will recognize that the IGHV gene is entirely encoded by the RL3 gene. It is understood that the same principle can be applied to CDRH1 and / or CDRH2 of the heavy chain. You will easily recognize it.
[0126] CDRL3 length In some embodiments, in another method or other manner described herein, In addition to the above features, the present invention provides libraries in which the CDRL3 length may vary. Therefore, the present invention provides, inter alia, a library having a specific distribution of CDRL3 lengths. CDRL3 libraries of lengths 8, 9, and 10 are exemplified, but those skilled in the art will appreciate that One may also apply the methods described herein to different lengths of time, which are also within the scope of the present invention. (e.g., about 5, 6, 7, 11, 12, 13, 14, 15, and / or 16) C It will be readily apparent that this can be adapted to produce light chains with DRL3. In some embodiments, the length of any of the CDRL3s of the invention may be any amino acid sequence. The number of amino acids may be less than a certain number, and the number of amino acids may be any one of the integers provided above. In some embodiments of the invention, the specified numerical ranges are inclusive or Use any two of the integers provided above as the lower and upper bounds of an exclusive range. All combinations of the integers provided that define the upper and lower limits are contemplated. can be.
[0127] Synthetic Antibody Library In some embodiments of the invention, the libraries provided comprise one or more synthetic In some embodiments, the provided library comprises: (a) a polynucleotide; (b) heavy chain chassis polynucleotide; (b) light chain chassis polynucleotide; (c) CDR3 (d) constant domain polynucleotides; and (e) combinations thereof. Those skilled in the art will appreciate that such synthetic polynucleotides may be selected from the group consisting of: The polynucleotide may be amplified by other synthetic or non-synthetic polynucleotides in the provided library. It will be appreciated that the present invention may be linked to a nucleotide. Synthetic polynucleotides may be prepared by any available method.
[0128] For example, in some embodiments, synthetic polynucleotides are prepared using methods such as Feldhausen t al., Nucleic Acids Research, 2000, 28: 534;Omstein et al., Biopolymers, 1978, 17: 2341;Brenner and Lerner, PNAS, 1992, 87: 6378, U.S. Patent Application Publication Nos. 2009 / 0181855 and 2010 / 0056386, and International Publication No. WO / 2009 / 036379 (respectively Split Pool DN, as described in It can be synthesized by A synthesis.
[0129] In some embodiments of the invention, The segments showing TN1, DH, N2, and JH diversity were identified by double-stranded DNA oligonucleotides. oligonucleotide, a single-stranded DNA oligonucleotide representing the coding strand, or a single-stranded DNA oligonucleotide representing the non-coding strand The DNA is then newly synthesized as a single-stranded DNA oligonucleotide representing the target gene. The array contains chassis sequences and in some cases portions of FRM4 and constant regions. The vector can be introduced into a host cell along with an acceptor vector that supports the vector. Primer-based PCR amplification from NA or mRNA or mammalian cDNA or No template-specific cloning step from mRNA needs to be used.
[0130] Library construction by yeast homologous recombination In one embodiment, the present invention provides a method for the intrinsic ability of yeast cells to promote homologous recombination with high efficiency. The mechanism of homologous recombination in yeast and its applications are briefly described below. (See, e.g., U.S. Patent Nos. 6, 409, 410, 420, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 44 No. 6,863; No. 6,410,246; No. 6,410,271; No. 6,610,47 2; and see also 7,700,302).
[0131] As an illustrative embodiment, homologous recombination can be performed, for example, by This can be carried out in budding yeast with genetic machinery designed to The budding yeast strains used were EM93, CEN.PK2, RM11-1a, YJM789, and B This mechanism is thought to have evolved for the purpose of chromosome repair and is called "Gallial "Gap repair" or "gap filling" By utilizing this mechanism, mutations can be generated in specific regions of the yeast genome. For example, a vector carrying a mutant gene can be introduced into a The 5' and 3' open ends of the gene intended to be disrupted or mutated It can contain two sequence segments that are homologous to the original reading frame (ORF) sequence. The vector also contains a nutrient allele flanked by two homologous DNA segments. nucleotides (e.g., URA3) and / or antibiotic resistance markers (e.g., Genetic It may also encode a positive selection marker such as cin / G418. Genetic markers and antibiotic resistance markers are known to those skilled in the art.
[0132] In some embodiments of the invention, the vector (e.g., a plasmid) is linearized. The plasmid and yeast genome are transfected with the vector at two homologous recombination sites. Through homologous recombination between two homologous sequence segments, the reciprocal exchange of DNA content occurs. The wild-type and mutant genes (selectable marker genes) in the flanking yeast genome Select one or more selectable markers between the gene(s) By this, the surviving yeast cells are those in which the wild-type gene has been replaced with the mutant gene. (Pearson et al., Ye, incorporated by reference in its entirety) Ast, 1998, 14: 391). This mechanism is important for functional genomics research. All 6,000 yeast genes or open reading frames (ORFs) in It has been used to generate systematic mutations in A similar approach has also been used to clone yeast genomic DNA fragments into plasmid vectors. (Iwas incorporated by reference in its entirety.) aki et al., Gene, 1991, 109: 81).
[0133] By utilizing the endogenous homologous recombination machinery present in yeast, gene fragments or syntheses can be synthesized. The synthetic oligonucleotides can also be used to construct plasmid vectors without a ligation step. In this application of homologous recombination, the target gene fragment can be cloned into (i.e., a fragment to be inserted into a plasmid vector, e.g., CDR3) (e.g., oligonucleotide synthesis, PCR amplification, restriction digestion from other vectors, etc.) DNA sequences homologous to the selected region of the plasmid vector are used to induce target gene fragmentation. These homologous regions may be entirely synthetic. Tracking is achieved through PCR amplification of target gene fragments with primers incorporating identical or homologous sequences. The plasmid vector may contain a nutrient enzyme allele (e.g., URA3) or an antibiotic Positives such as substance resistance markers (e.g., Geneticin / G418) The plasmid vector may then be transfected with the target gene fragment. A unique restriction site located in the middle of a region of shared sequence homology tion cut), thereby creating an artificial gap at the cut site. The linearized plasmid vector and the plasmid vector are flanked by sequences homologous to the plasmid vector. The target gene fragments located therein are co-transformed into a yeast host strain. The yeast then Recognizes two stretches of sequence homology between the target and target gene fragments and compares the differences at the gaps This facilitates the reciprocal exchange of DNA content through recombination. The gene fragment is inserted into the vector without ligation.
[0134] The methods described above also include those in which the target gene fragment is derived from, for example, a circular M13 phage. In the form of single-stranded DNA or as a single-stranded oligonucleotide (each of which is incorporated by reference in its entirety). mon and Moore, Mol. Cell Biol., 1987, 7: 2329; Ivanov et al., Genetics, 1996, 14 2: 693; and DeMarini et al., 2001, 30: 5 20.) Therefore, recombination into a gapped vector The target form that can be obtained can be double-stranded or single-stranded and can be synthesized chemically, PCR , restriction digestion, or other methods.
[0135] Several factors can affect the efficiency of homologous recombination in yeast. For example, gap The efficiency of repair depends on the homologous sequences flanking both the linearized vector and the target gene. In one embodiment, a length of about 20 or more base pairs is used for the length of the homologous sequence. may be used, and about 80 base pairs may give near-optimal results (respectively, Hua et al., Plasmid, 19, which is incorporated by reference in its entirety. 97, 38: 91; Raymond et al., Genome Res., 2002, 12: 190). In some embodiments of the invention, the 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 3 2, 33, 34 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65 ,70,75,80,85,90,95,100,110,120,130,140,1 50, 160, 170, 180, 187, 190, or 200 homologous base pairs are In one embodiment, about 20 to about 40 base pairs may be used to facilitate the synthesis of a nucleic acid sequence. Furthermore, the mutual exchange between vectors and gene fragments is strictly sequence-dependent. that is, it does not cause a frameshift. Scanning ensures the insertion of gene fragments with both high efficiency and high precision. , two, three, or more target genes in the same vector in a single transformation attempt. allowing for simultaneous cloning of the offspring fragments (incorporated by reference in its entirety). Raymond et al., Biotechniques, 1999, 26: 134) Furthermore, the nature of high-precision sequence conservation through homologous recombination provides direct functional For testing, the selected gene or gene fragment is inserted into an expression vector or fusion vector. (each of which is incorporated by reference in its entirety) El-Deiry et al., Nature Genetics, 1992, 1: 4549; Ishioka et al., PNAS, 1997, 94: 2449).
[0136] Libraries of gene fragments have also been constructed in yeast using homologous recombination. For example, a human brain cDNA library was constructed in the vector pJG4-5. The hybrid fusion library was constructed (Gu, which is incorporated by reference in its entirety). idotti and Zervos, Yeast, 1999, 15: 715) A total of 6,000 pairs of PCR primers were developed for studying yeast genome-protein interactions. It has also been reported that it has been used to amplify 6,000 known yeast ORFs for (Hudson et al., Geno, which is incorporated by reference in its entirety) me Res., 1997, 7: 1169). In 2000, Uetz et al. performed a comprehensive analysis of protein-protein interactions in budding yeast (see references in their entirety). Incorporated by Uetz et al., Nature, 2000, 40 3: 623). The protein-protein interaction map of budding yeast shows all the interactions between yeast proteins. Using a comprehensive system to examine two-hybrid interactions in various combinations (Ito et al., 2014, incorporated by reference in its entirety). PNAS, 2000, 97: 1143), proteins of the vaccinia virus genome A genetic linkage map was studied using this system (Mc Craith et al., PNAS, 2000, 97: 4879).
[0137] In one embodiment of the invention, the synthetic CDR3 (heavy or light chain) is a full-length heavy or light chain CDR3. The heavy chain chassis or light chain chassis, FRM4, are synthesized by homologous recombination to form the The vector may be linked to a vector encoding the portion, and the constant region. In some embodiments, homologous recombination is carried out directly in yeast cells. A method like (a) In yeast cells, (i) a fragment encoding a heavy or light chain chassis, a portion of FRM4, and a constant region; a linearized vector, wherein the linearization site is between the end of FRM3 of the chassis and The linearized vector between the beginning of the constant region and (ii) a library of linear, double-stranded CDR3 insert nucleotide sequences, Each of the R3 inserts contains a nucleotide sequence encoding a CDR3 and a homologous recombination between the vector and the library of CDR3 insert sequences, (i) 5' and 3' flanking sequences sufficiently homologous to the ends of the vector at the site of linearization a library of CDR3 insert nucleotide sequences, including the sequence Ravini (b) a CDR3 insert sequence to generate a vector encoding a full-length heavy or light chain; In the transformed yeast cells, the vector and CD The step of allowing homologous recombination to occur between the R3 insertion sequences.
[0138] As specified above, the CDR3 insert is inserted into the 5′ end of the linearized vector, which is homologous to the end of the linearized vector. and 3' flanking sequences. When the linearized vector is introduced into a host cell, such as a yeast cell, the vector The "gap" (linearization site) created by linearization is the gap between these two linear duplexes. Recombination of homologous sequences at the 5' and 3' ends of the DNA (i.e., vector and insert) Through this event of homologous recombination, the variable C A library of circular vectors encoding full-length heavy or light chains containing DR3 inserts was generated. Specific examples of these methods are provided in the Examples.
[0139] Subsequent analysis can identify homologous sequences that result in the correct insertion of, for example, CDR3 sequences into a vector. For example, a recombinant vector may be prepared from a selected yeast clone. PCR amplification of the direct CDR3 insert revealed how many clones were recombinant. In one embodiment, a library having at least about 90% recombinant clones may be prepared. In some embodiments, a minimum of about 1%, 5%, 10%, 15%, 20%) is utilized. 5%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 7 5%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 9 3%, 94%, 95%, 96%, 97%, 98% or 99% recombinant clones A library containing the selected clones is utilized. Similar PCR amplification of the selected clones is also performed to identify the inserts. Size may be specified.
[0140] To verify the sequence diversity of the inserts in the selected clones, inserts of the correct size were PCR amplification products containing the target gene are known to cut or not cut within the amplified region. The DNA may be "fingerprinted" using restriction enzymes known to From the analysis, the clones analyzed may be of the same identity or of different or various identities. The PCR products may also be used to determine the identity and cloning of the insert. To clarify the fitness of the screening procedure and to evaluate the independence and diversity of clones To verify, it may be sequenced directly.
[0141] Expression and screening systems Generated by any of the techniques described herein or other suitable techniques A library of polynucleotides is used to identify antibodies with desired structure and / or activity. The expression of antibodies can be performed, for example, in cells. Cell extracts (and e.g., ribosome display), phage display, prokaryotic cells (e.g., bacterial display) or eukaryotic cells (e.g., yeast display) In one embodiment of the present invention, the antibody library can be cultured in yeast. It is expressed as
[0142] In some embodiments, the polynucleotide is expressed in a cell-free extract. See, e.g., U.S. Pat. No. 5,324,637. Nos. 5,492,817; 5,665,563 (each incorporated by reference in its entirety) The vectors and extracts described in the above (incorporated herein) can be used, or Many are commercially available. Ribosome display and other cell-free technologies for linking (i.e., phenotype) For example, Profusion™ can be used (e.g., each of its entirety). The bodies of which are incorporated by reference are U.S. Patent Nos. 6,348,315; 6,261,804 (See Nos. 6,258,558; and 6,214,553).
[0143] Alternatively or additionally, the polynucleotides of the invention may be d Skerra (Meth. Enz., each of which is incorporated by reference in its entirety). ymol., 1989, 178: 476; Biotechnology, 19 In E. coli expression systems such as those described by (91, 9: 273), The mutein can be expressed in any of the following forms, all of which are incorporated by reference in their entirety: Better and Horwitz, Meth. Enzymol., 19 89, 178: 476, In some embodiments, the VH and VH can be expressed for secretion into the cytoplasm. The single domain encoding L is the ompA, phoA, or pelB signal, respectively. The signal sequence is added to the 3' end of the sequence encoding the signal sequence, such as the Incorporated by illumination Lei et al., J. Bacteriol., 19 87, 169: 4379). These gene fusions are derived from a single vector. These are constructed in a dicistronic construct so that they can be expressed. It is secreted into the periplasmic space of E. coli and can be recovered in an active form. (Skerra et al., Biot (Echnology, 1991, 9: 273). For example, antibody heavy chain genes are The antibody can be co-expressed with an antibody light chain gene to produce an antibody or antibody fragment.
[0144] In some embodiments of the invention, the antibody sequences are those described in U.S. Patent Application Publication No. 2004 / 007 2740; U.S. Patent Application Publication No. 2003 / 0100023; and U.S. Patent Application Publication No. No. 2003 / 0036092 (each of which is incorporated by reference in its entirety). For example, secretion signals and lipidation components may be used to produce prokaryotic vectors, as described in , for example, expressed on the membrane surface of E. coli.
[0145] Mammalian cells, such as myeloma cells (e.g., NS / 0 cells), hybridoma cells, cells, Chinese hamster ovary (CHO) cells, human embryonic kidney (HEK) cells, etc. Such higher eukaryotic cells can also be used for expressing the antibodies of the present invention. Antibodies expressed in mammalian cells may be secreted into the medium or expressed on the surface of the cells. The antibody or antibody fragment may be, for example, an intact antibody molecule. or as individual VH and VL fragments, Fab fragments, single domains or as single chain ( scFv) (Huston et al., which is incorporated by reference in its entirety). on et al., PNAS, 1988, 85: 5879).
[0146] Alternatively or additionally, antibodies may be used as described, for example, in Jeong et al., PNAS , 2007, 104: 8247 (incorporated by reference in its entirety). Also, by fixed cell periphery expression (APEx2 hybrid surface display) as described See, for example, Mazor et al., Nature Biotechnology, 2007, 25:563 (incorporated by reference in its entirety). can be expressed and screened by other immobilization methods
[0147] In some embodiments of the invention, antibodies are selected using mammalian cell display. (Ho et al., PNA, incorporated by reference in its entirety) S, 2006, 103: 9637).
[0148] Screening of antibodies from the libraries of the present invention can be carried out by any suitable means. For example, binding activity can be measured by standard immunoassays and / or Catalytic function, e.g., can be assessed by affinity chromatography. Screening of the antibodies of the invention for proteolytic function can be performed using standard assays, such as For example, in U.S. Pat. No. 5,798,208 (incorporated by reference in its entirety), This can be accomplished using the hemoglobin plaque assay described in Determining the ability of a body to bind to a therapeutic target can be accomplished using, for example, surface plasmon resonance based assays. Using a BIACORE™ instrument, which measures the binding rate of an antibody to a given target or antigen, In vitro assays can be used. In vivo assays are often used. These can be performed using any of the animal models, and then subsequently in humans, as needed. Cell-based biological assays are also contemplated.
[0149] One feature of the invention is that the antibodies in the library can be expressed and screened. In one embodiment of the invention, the antibody library is about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 2 can be expressed in yeast with a doubling time of less than 1, 22, 23, or 24 hours In some embodiments, the doubling time is about 1 to about 3 hours, about 2 to about 4 hours, or about 3 to about 8 hours, about 3 to about 24, about 5 to about 24, about 4 to about 6, about 5 to about 22, about 6 to about 8, about 7 to Approximately 22, approximately 8 to approximately 10 hours, approximately 7 to approximately 20, approximately 9 to approximately 20, approximately 9 to approximately 18, approximately 11 to approximately 1 8, about 11 to about 16, about 13 to about 16, about 16 to about 20, or about 20 to about 30 hours. In one embodiment of the present invention, the antibody library is incubated for about 16 to about 20 hours, about 8 to about 1 hour, or The doubling time is about 6 hours, or about 4 to about 8 hours, expressed in yeast. The antibody library of the present invention is easy to express and screen. Expression and screening can be performed in a few hours, compared to previously known techniques that require several days. The throughput of such a screening process in mammalian cells can be The restriction step in the method typically involves repeatedly regrowing the isolated cell population. These are the times required for the The yeast has a doubling time that exceeds that of the yeast to which it is subjected.
[0150] In certain embodiments of the invention, the compositions of the library are enriched in one or more enrichment steps. (e.g., screening for antigen binding, binding to a general ligand, or other properties) For example, about x% of the sequences or libraries of the present invention may be determined after screening. The library having compositions containing the library is subjected to one or more screening steps. After that, about 2x%, 3x%, 4x%, 5x%, 6x%, 7x%, 8x%, 9x% of the present invention , 10x%, 20x%, 25x%, 40x%, 50x%, 60x% 75x%, 80x% Enriched to contain 90x%, 95x%, or 99x% of the sequence or library In some embodiments of the invention, the sequences or libraries of the invention may be Approximately 2-, 3-, 4-, or 5-fold higher than their occurrence before one or more enrichment steps , 6x, 7x, 8x, 9x, 10x, 100x, 1,000x, or more, concentrated In one embodiment of the invention, the library may comprise CDRH3, CDRL3, At least one of a number of specific types of sequences, such as a chain, a light chain, or an entire antibody may contain (e.g., at least about 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 1 0 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 101 6 , 10 17 , 10 18 , 10 19 , or 10 20 In some embodiments, these sequences The columns should be at least about 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , or 10 19 To provide a library containing each of the sequences Alternatively, the concentration may be carried out between multiple concentration steps.
[0151] Mutagenesis approaches for affinity maturation As described above, antibody leads may be assayed for binding to one or more antigens. and screening antibodies from the library of the invention for a specific biological activity. The coding sequences of these antibody leads can be initially identified through a selection process. To generate a secondary library with diversity introduced in relation to the antibody leads, Further mutagenesis was performed in vitro or in vivo to Such mutagenized antibody leads may then be further purified by the addition of the mutagenized antibody leads from the primary library. In vitro antibody selection was performed following a procedure similar to that used for initial antibody lead selection. for binding to target antigens or biological activity in vitro or in vivo Such mutagenesis and further screening of primary antibody leads can be performed. and selection is performed to identify mammalian cells that produce antibodies with stepwise increases in affinity for the antigen. This effectively mimics the affinity maturation process that occurs naturally in
[0152] In some embodiments of the invention, only the CDRH3 region is mutagenized. In some embodiments of the present invention, the entire variable region is mutagenized. In embodiments, one or more of CDRH1, CDRH2, CDRH3, CDRL1, CDRL2 and / or CDRL3 may be mutagenized. In an embodiment, "light chain shuffling" is used as part of an affinity maturation protocol. In certain embodiments, this may be done to enhance the affinity and / or biological activity of the antibody. This involves pairing many light chains with one or more heavy chains to select the light chain that best suits the needs of the individual. In certain embodiments of the invention, one or more heavy chains may be capable of pairing. The number of light chains in the antibody may be at least about 2, 5, 10, 100, 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , or 10 10 In one embodiment of the present invention, these The light chain is encoded by a plasmid. In some embodiments of the invention, the light chain is It may be integrated into the genome of the host cell.
[0153] The coding sequence of the antibody lead is mutagenized using one of a wide variety of methods. Examples of methods for mutagenesis include site-directed mutagenesis and error-prone PCR. Mutagenesis, cassette mutagenesis, and random PCR mutagenesis are included. Alternatively or additionally, the region encoding the desired mutation may be Oligonucleotides that bind to the nucleotides can be synthesized and then, for example, via recombination or ligation, It can be introduced into the sequence to be mutagenized.
[0154] Site-directed or point mutagenesis involves gradually altering the CDR sequences in specific regions. For example, this may be used to vary the length of an oligonucleotide-directed collision. This may be achieved by using natural mutagenesis or PCR. Short sequences of the code are synthetically mutagenized in the heavy and / or light chain regions. Such a method may be used to exchange the CDR sequences with oligonucleotides derived from the CDR sequences. Although it may not be as efficient at mutagenizing specific target proteins, it may be more efficient at mutagenizing specific target proteins. It may be used for fine tuning of specific leads to achieve high affinity.
[0155] Cassette mutagenesis is used to mutagenize CDR sequences in specific regions. Alternatively or additionally, a single The sequence blocks or regions of the template are replaced with fully or partially randomized sequences. However, the maximum information content that can be obtained is determined by the run length of the oligonucleotide. Similar to point mutagenesis, this method may be statistically limited by the number of dumb sequences. In addition, specific lead-binding proteins can be used to achieve higher affinity for specific target proteins. This may be used for fine tuning the
[0156] Error-prone PCR or "poison" PCR is described, for example, in U.S. Pat. No. 6,153, No. 745; Caldwell and Joyce, PCR Methods and Applications, 1992, 2: 28;Leung et al., Technique, 1989, 1: 11;Shafikhani et al. ., Biotechniques, 1997, 23: 304; and Stemm er et al., PNAS, 1994, 91: 10747 (respectively, CDR sequences were prepared according to the protocol described in the paper "Generic Isolation of CDRs in a Novel Cell Line," which is incorporated by reference in its entirety. It may be used to mutagenize the sequence.
[0157] Conditions for error-prone PCR include, for example, (a) Taq DNA polymerase High concentrations of Mn efficiently induce the dysfunction of 2+ (For example, about 0.4 to about 0.6 mM) and / or (b) causing incorrect incorporation of a disproportionately high concentration of substrate into the template. This results in a disproportionately high concentration of one nucleotide in the PCR reaction, leading to mutations. Alternatively or additionally, the number of PCR cycles, Other factors, such as the type of DNA polymerase used and the length of the template, affect PC This can affect the rate of misincorporation of the "wrong" nucleotide into the R product. Commercially available products such as the "Clontech™ PCR Random Mutagenesis Kit" (CLONTECH™) Kits available at are utilized for mutagenesis of selected antibody libraries. Good too.
[0158] The primer pairs used in PCR-based mutagenesis, in one embodiment, are: It may also contain a region that matches the homologous recombination site in the expression vector. The design involves the insertion of heavy or light chain chassis vectors via homologous recombination followed by mutagenesis. This allows for easy reintroduction of PCR products back into the
[0159] Other PCR-based mutagenesis methods may also be used, either alone or in combination with the error promoters described above. For example, PCR-amplified CDR sequences can be used in conjunction with PCR. The fragment is digested with DNase to generate nicks in the double-stranded DNA. These nicks may be cleaved by other exonucleases such as Bal 31. The gap can then be expanded by the disproportionately high concentration of Along with one substrate (e.g., dGTP), the usual substrates dGTP, dATP, dTTP, and Randomization was achieved by using DNA Klenow polymerase with low concentrations of dCTP and The filling reaction may be highly efficient in the filled gap region. Such a method of DNase digestion should result in a high frequency of mutations. together with error-prone PCR to generate high frequency mutations in the R segment. may be used.
[0160] CDRs or antibody segments amplified from primary antibody reads are used to identify targeting genes in pre-B cells. Mutagenesis is achieved in vivo by exploiting the intrinsic ability of spontaneous mutation. Ig genes in pre-B cells are particularly susceptible to rapid mutation. As B cells proliferate, Ig promoters and enhancers are released into the pre-B cell environment. Therefore, the CDR gene segments are Cloned into a mammalian expression vector containing a human Ig enhancer and promoter. Such constructs may be introduced into pre-B cell lines such as 38B9. The VH and VL gene segments may be introduced into pre-B cells, which allows for the rapid translation of VH and VL gene segments. (Liu and Van, incorporated by reference in their entirety) Ness, Mol. Immunol., 1999, 36: 461). mutation The induced CDR segments can be amplified from cultured pre-B cell lines and subsequently purified by, for example, homologous recombination. The vector can be reintroduced into the chassis-containing vector(s) via a transfection.
[0161] In some embodiments, CDR "hits" isolated from library screening are The "gap repair" can be performed by resynthesizing the gene using, for example, degenerate codons or trinucleotides. The fragments can be recloned into heavy or light chain vectors using the same techniques.
[0162] Other variants of the polynucleotide sequences of the present invention In one embodiment, the present invention provides a method for hybridizing the polynucleotides taught herein. Recombinase or hybridize with the complement of the polynucleotides taught herein. For example, the polynucleotides taught herein are provided. The protease was subjected to hybridization and cleavage under low, medium, or high stringency conditions. The isolated polynucleotides or polynucleotides taught herein that remain hybridized after washing and cleavage are Complements of the depicted polynucleotides are encompassed by the invention.
[0163] Exemplary low stringency conditions include about 30% to about 35% formamide at about 37°C. A buffer solution of about 1M NaCl and about 1% SDS (sodium dodecyl sulfate) was used. Hybridization and elution at about 50°C to about 55°C in about 1× to about 2× SSC (20× SSC) C=3.0M NaCl / 0.3M trisodium citrate).
[0164] Exemplary medium stringency conditions include about 40% to about 45% formamide at about 37°C. , about 1 M NaCl, about 1% SDS, and at about 55°C to about 6 This includes washing in about 0.5x to about 1x SSC at 0°C.
[0165] Exemplary high stringency conditions are about 50% formamide, about 1 M at about 37°C. Hybridization in NaCl, about 1% SDS and at about 60°C to about 65°C Includes a wash in 0.1x SSC.
[0166] Optionally, the wash buffer may contain about 0.1% to about 1% SDS.
[0167] The hybridization period is generally less than about 24 hours, usually about 4 to about 12 hours. be.
[0168] Sub-libraries of the invention and larger libraries or sub-libraries Library The libraries described herein (e.g., CDRH3 and CDRL3 libraries) Libraries containing combinations of the above-mentioned sequences are encompassed by the present invention. Sub-libraries containing portions of the libraries described in are also provided by the present invention. Included are, for example, CDRH3 libraries of particular heavy chain chassis or, for example, lengths (a subset of the CDRH3 library based on
[0169] Additionally, a library containing one of the libraries or sub-libraries of the present invention may be used. For example, in one embodiment of the present invention, One or more libraries or sub-libraries are part of a larger library (logical The sequences may be contained within a molecule (logical or physical), which may be sequences derived by other means, For example, non-human or human sequences derived by stochastic or site-by-site stochastic synthesis. In one embodiment of the present invention, the polynucleotide library At least about 1% of the sequences in the present invention are those of the present invention, regardless of the composition of the other 99% of the sequences. (e.g., CDRH3 sequence, CDRL3 sequence, VH sequence, VL sequence). For purposes of illustration only, those skilled in the art will 7 The members of the library of the present invention are (i.e. 1%), totaling 10 9 A library containing members of the present invention has utility. It has been shown that clear library members can be isolated from such libraries. It will be readily appreciated that in some embodiments of the present invention, any polynucleotide label At least about 0.001%, 0.01%, 0.1%, 2%, 5%, 1% 0%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 81%, 82%, 8 3%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 9 3%, 94%, 95%, 96%, 97%, 98%, or 99% of the sequences are from a set of other sequences. In some embodiments, the sequences of the present invention may be of any composition. is approximately 0.05 in any polynucleotide library, regardless of the composition of other sequences. 0.01% to approximately 1%, approximately 1% to approximately 2%, approximately 2% to approximately 5%, approximately 5% to approximately 10%, approximately 10% to approximately 15%, about 15% to about 20%, about 20% to about 25%, about 25% to about 30%, about 30% to about 35%, approx. 35% to approx. 40%, approx. 40% to approx. 45%, approx. 45% to approx. 50%, approx. 50% to approx. 55%, about 55% to about 60%, about 60% to about 65%, about 65% to about 70%, about 70% to about 75%, about 75% to about 80%, about 80% to about 85%, about 85% to about 90%, about 90% to about 95%, or about 95% to about 99% of the sequence of one or more of the present invention. library or sub-library of the present invention, but A quantity of books that can be used to efficiently screen a library or sub-library Still comprising one or more libraries or sub-libraries of the invention, or isolating sequences encoded by multiple libraries or sub-libraries. Libraries that can be used to generate such a library are also within the scope of the present invention.
[0170] Alternative scaffolds As will be apparent to one skilled in the art, the CDRH3 and / or C provided by the present invention The DRL3 polypeptide may also be displayed on alternative scaffolds. Scaffolds such as these have been shown to yield molecules with specificity and affinity comparable to antibodies. Exemplary alternative scaffolds include fibronectin (e.g., AdNectin), β -Sandwich (e.g., iMab), lipocalin (e.g., Anticalin), E ETI-II / AGRP, BPTI / LACI-D1 / ITI-D2 (e.g. Kuni tz domain), thioredoxin (e.g., peptide aptamers), protein A (e.g., Affibody), ankyrin repeat (DARPin), γB crystal Phospho / ubiquitin (e.g., Affilin), CTLD3 (e.g., tetranectin) , and (LDLR-A module) 3 (e.g., Avimer) For example, further information on alternative scaffolds may be found, e.g., in the respective references in their entirety. Incorporated by Binz et al., Nat. Biotechnol. , 2005 23: 1257 and Skerra, Current Opin. in Biotech., 2007 18: 295-304.
[0171] Further embodiments of the present invention Library Size In some embodiments of the invention, the library comprises about 10 1 ~about 10 20 Different positions including oligonucleotide or polypeptide sequences (e.g., antibodies, heavy chain, CDRH3, light chain and / or encodes or comprises CDRL3). The library of the present invention comprises at least about 10 1 , 10 2 , 10 3 , 10 4 , 10 5 , 1 0 6 , 10 7 , 10 8 , 109 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 16 , 10 17 , 10 18 , 10 19 , or 10 20 or more different antibodies, heavy chain, CDRH3, light chain, and / or CDRL3 polynucleotides or In some embodiments, the libraries of the invention are designed to contain a polypeptide sequence. A library may contain less than a specified number of polynucleotide or polypeptide sequences. The number of sequences is defined using any one of the integers provided above. In certain embodiments, the specified numerical ranges may be used as lower and upper limits of inclusive or exclusive ranges. The upper and lower bounds are defined using any two of the integers provided above. All combinations of the provided integers are contemplated.
[0172] In some embodiments, the present invention provides a method for detecting a small proportion of the members of a library. The members of the present invention are lyophilized and / or lyophilized, and are produced according to the methods, systems, and compositions provided herein. One important property of the libraries of the present invention is that they contain a variety of lengths. advantageously mimic certain aspects of the human preimmune repertoire, including chromosome diversity and sequence diversity Those skilled in the art will understand that the libraries provided by the present invention are A subset of library members may be isolated according to the methods, systems, and compositions provided herein. It will be readily appreciated that the present invention includes a library of members that are produced. Ba, 10 6Members of the group are produced according to the methods, systems, and compositions provided herein. 10 born 8 The library containing members of the method provided herein is The sequences produced in accordance with the present invention will comprise 1% of the sequences produced in accordance with the present invention, systems, and compositions. or multiple 10 6 members using screening techniques known in the art. It will be appreciated that the library can be easily isolated using More specifically, at least about 1%, 2.5%, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 45%, 50%, 55% 60% 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the present CDRH3, CDRL3, light chain, or heavy chain, and / or full-length Libraries containing antibody sequences are within the scope of the present invention. 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 The CDRH3, CDRL3, light chain, heavy chain Libraries containing nucleotides, and / or full-length antibody sequences are also within the scope of the present invention.
[0173] Human Pre-immune Set In some embodiments, the present invention provides a method for producing 3,571 controlled human ovarian tumors contained within an HPS. A set of pre-immune antibody sequences, their corresponding CDRH3 sequences (Appendix A), and / or or computer readable form of these CDRH3 sequences (and / or TN1, DH, N2, and / or H3-JH segments of the target gene. In embodiments, the present invention provides a method for generating a CDRH3 library, the method comprising: Candidate segments from the segment pool (i.e., TN1, DH, N2, and H3-JH) ) to the CDRH3 sequence in the HPS and / or any other replica of the CDRH3 sequence. In some embodiments, the present invention provides a method for matching a tree with a candidate segments and / or fragments derived from the theoretical segment pool disclosed herein or segments selected for inclusion in a physical library.
[0174] Embodiment The methods described herein use a limited number of allelic variants, H3 Although the production of a theoretical segment pool of -JH and -DH segments is shown, those skilled in the art will be able to The methods taught herein are also applicable to any other allelic variants as well as all non-human applied to any IGHJ and IGHD gene, including the IGHJ and IGHD genes It will be appreciated that alternatively or additionally, the The method may, for example, include: Alternatively or additionally, the present invention may be applied to any reference set of CDRH3 sequences. The supplier expressly acknowledges that each of the described embodiments of the present invention may be used in combination with a vector, virus, or microorganism. (e.g., yeast or bacteria) in the form of a polynucleotide or polypeptide It will be appreciated that the present invention also provides a fully enumerated synthetic library. Therefore, an embodiment of the present invention provides a computer readable version of the method described above. Any of the described embodiments and uses thereof.
[0175] Non-human antibody libraries are also within the scope of the present invention.
[0176] The present disclosure provides a method for identifying Cys residues, N-linked glycosylation motifs, from the libraries of the present invention. This paper describes the removal of sequences containing deamidation motifs and highly hydrophobic sequences. The trader may determine that one or more (i.e., not necessarily all) of these criteria are met in the present invention. that can be applied to remove undesired sequences from any library However, it will be appreciated that a label containing one or more of these types of sequences may be used. Libraries are also within the scope of the present invention. Other criteria may also be used; What is described in the specification is not limiting.
[0177] In one embodiment, the present invention provides a library (theoretical, synthetic, or physical realization) of To provide a library in which the number of times a particular sequence of In some embodiments, the present invention provides libraries (e.g., CDRH3, CDRL3, The frequency of occurrence of any of the sequences in the heavy chain, light chain, and full-length antibody is about 2, 5, 10, 15, , 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800 In some embodiments, a library is provided in which the number of nucleotides is less than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, , the frequency of occurrence of any of the sequences in the library is less than a multiple of the occurrence frequency of the sequence, e.g., less than the occurrence frequency of any other sequence in the library Current frequency of approximately 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 6 0, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 5 00, 600, 700, 800, 900, or 1000 times less.
[0178] In some embodiments, the library is used to generate CDRH3 sequences. Combinatorial diversity of the segments, particularly those that generate specific CDRH3 sequences, is desired. It is defined by the number of non-degenerate segment combinations that can be used for In some embodiments, this metric is, for example, approximately 200 0, 5000, 10000, 20000, 50000, 100000, or more The sequence samples and the sequences used to generate the CDRH3 sequences of the library In some embodiments, the metric may be calculated using "self-matching" using the segment. The present invention provides a method for detecting at least about 95%, 90%, 85%, 80%, 70%, 85%, 90 ...0%, 90%, 90%, 90%, 90%, 90%, 90%, 90%, 90%, 5%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 2 5%, 20%, 15%, 10%, or 5% of the CDRH3 sequences were from a single set of segments. The present invention provides a library that may be formed by combining
[0179] In one embodiment of the invention, statistical bootstrap analysis is performed to generate the CDRH3 reference set. Although it may be convenient to use this method, it is This is not required for all embodiments of the present invention.
[0180] In some embodiments, the present invention provides a method for the production of a polynucleotide that encodes a polypeptide of the present invention. A method and system for selecting nucleotides, individually and / or in other segments. After combinatorial ligation with
[0013] Methods and systems are provided (e.g., performing) a method for selecting a leotide segment. See Example 9.3.7).
[0181] The exemplary libraries provided herein are for illustrative purposes only and are not limiting. Provided for your convenience. [Example]
[0182] (Example) This invention is further illustrated by the following examples which should not be construed as limiting. The contents of all references, patents, and published patent applications cited throughout this application is hereby incorporated by reference.
[0183] In general, the practice of the present invention will involve, unless otherwise indicated, techniques in chemistry, molecular biology, recombinant DNA technology, and the like. techniques, PCR techniques, immunology (e.g., antibody techniques, among others), expression systems (e.g., yeast expression, Cell-free expression, phage display, ribosome display, and profusion N™), as well as any known compounds within the skill of the art and described in the literature. Use conventional techniques for cell culture, as required. See, e.g., Sambrook, Fritsch and Maniatis, Molecular Cloning: Cold S pring Harbor Laboratory Press (1989);DNA Cloning, Vols. 1 and 2, (DN Glover, E d. 1985);Oligonucleotide Synthesis (M.J. Gait, Ed. 1984);PCR Handbook Current Pr otocols in Nucleic Acid Chemistry, Beauc age, Ed. John Wiley & Sons (1999) (Edito r);Oxford Handbook of Nucleic Acid Struc ture, Neidle, Ed., Oxford Univ Press (19 99);PCR Protocols: A Guide to Methods an d Applications, Innis et al., Academic P ress (1990);PCR Essential Techniques: Es sential Techniques, Burke, Ed., John Wil ey & Son Ltd (1996);The PCR Technique: R T-PCR, Siebert, Ed., Eaton Pub. Co. (199 8);Antibody Engineering Protocols (Metho ds in Molecular Biology), 510, Paul, S., Humana Pr (1996);Antibody Engineering: A Practical Approach (Practical Approach Series, 169), McCafferty, Ed., Irl Pr ( 1996);Antibodies: A Laboratory Manual, H arlow et al., C.S.H.L. Press, Pub. (1999 );Current Protocols in Molecular Biology , eds. Ausubel et al., John Wiley & Sons (1992);Large-Scale Mammalian Cell Cultu re Technology, Lubiniecki, A., Ed., Marc. el Dekker, Pub., (1990);Phage Display: A Laboratory Manual, C. Barbas (Ed.), CSH L Press, (2001);Antibody Phage Display, P O'Brien (Ed.), Humana Press (2001);Bor der et al., Nature Biotechnology, 1997, 15: 553;Border et al., Methods Enzymol. 2000, 328: 430; Pluck in U.S. Pat. No. 6,348,315 Ribosome display as described by Thun et al. and U.S. Pat. Nos. 6,258,558; 6,261,804; and 6,214,553 Profusion™, as described by Szostak et al.; and the bacterial cell periphery described in U.S. Patent Application Publication No. 20040058403 A1. See, e.g., 1999, pp. 111-114, 2000. Each of the references cited in this paragraph is incorporated by reference in its entirety. It is incorporated by lighting.
[0184] Antibody sequence analysis using Kabat rules and aligned nucleotide sequences Further details regarding sequences and programs for analyzing amino acid sequences can be found, for example, in J. ohnson et al., Methods Mol. Biol., 2004, 248: 11;Johnson et al., Int. Immunol. 1998, 10: 1801;Johnson et al., Methods M ol. Biol., 1995, 51: 1;Wu et al., Protei ns, 1993, 16: 1; and Martin, Proteins, 199 6, 25:130. References cited in this paragraph Each of the references is incorporated by reference in its entirety.
[0185] Further details regarding antibody sequence analysis using Chothia's rules can be found, for example, in Ch othia et al., J. Mol. Biol., 1998, 278: 457;Morea et al., Biophys. Chem., 1997, 68: 9;Morea et al., J. Mol. Biol., 1998, 275: 269;Al-Lazikani et al., J. Mol. Bi ol., 1997, 273: 927.Barre et al., Nat. S truct. Biol., 1994, 1: 915;Chothia et al. ., J. Mol. Biol., 1992, 227: 799;Chothia et al., Nature, 1989, 342: 877; and Choth ia et al., J. Mol. Biol., 1987, 196: 901 Further analysis of the CDRH3 conformation may be found in Shirai et al. t al., FEBS Lett., 1999, 455: 188 and Shir Ai et al., FEBS Lett., 1996, 399: 1 Further details on Chothia's analysis can be found, for example, in Choth ia et al., Cold Spring Harb. Symp. Quant Biol., 1987, 52: 399. In this paragraph Each of the references cited herein is incorporated by reference in its entirety.
[0186] Further details regarding CDR contact issues can be found, for example, in MacCallum et al., J. Mol. Biol., 199 6, 262:732.
[0187] Further details regarding the antibody sequences and databases referred to herein can be found in For example, Tomlinson et al., J. Mol. Biol., 1992 , 227: 776, VBASE2 (Retter et al., Nuclei c Acids Res., 2005, 33: D671);BLAST (www .ncbi.nlm.nih.gov / BLAST / );CDHIT (bioinfo rmatics.ljcrf.edu / cd-hi / );EMBOSS (www.hg mp.mrc.ac.uk / Software / EMBOSS / );PHYLIP (e evolution.genetics.washington.edu / phylip. html); and FASTA (fasta.bioch.virginia.edu ) Each of the references cited in this paragraph is incorporated by reference in its entirety. is incorporated by reference.
[0188] Light Chain Library Example 1. Framework and / or CDRL1 and / or CDRL2 Diversity Light chain library containing Diversity in antibody sequences is concentrated in the CDRs, but some diversity in the framework regions Residues in may also affect antigen recognition and / or alter affinity (respectively , Queen et al., Proc. Na, which is incorporated by reference in its entirety. tl. Acad. Sci. USA, 1989, 86: 10029; Car ter et al., Proc. Natl. Acad. Sci. USA, 1992, 89: 4285). These residues have been classified and are used in, for example, antibody human have been used to make framework substitutions that improve antibody affinity during the synthesis process. (See, e.g., Foote and Winter, incorporated by reference in its entirety.) "Vernie" in J. Mol. Biol., 1992, 224: 487 In the heavy chain, the Vernier residues are Barring residues 2, 27-30, 47-49, 67, 69, 71, 73, 78, 93-94 In the light chain, the Vernier residues include Kabat residues 2, 4, 35-36. , 46-49, 64, 66, 68-69, 71, and 98. The Vernier residue numbers are The sequences for the kappa and lambda light chains are the same (incorporated by reference in their entirety). Chothia et al., J. Mol. Biol., 1985, 18 6: 651). Furthermore, the framework of the VL-VH interface Position can also affect affinity. In the heavy chain, interface residues include Kabat residue 35, 37, 39, 45, 47, 91, 93, 95, 100, and 103 (in their entirety) Chothia et al., J. Mol. Biol., incorporated by reference. ., 1985, 186: 651). In the light chain, the interface residues are Kabat residue 34, Includes 36, 38 44, 46, 87, 89, 91, 96, and 98.
[0189] The following procedure identifies the framework residues to be diversified and the amino acid sequence to be diversified. was used to select amino acids.
[0190] a. A set of human VK light chain DNA sequences was obtained from NCBI (International Publication No. (See Appendix A of publication number WO / 2009 / 036379.) These sequences are were classified according to the germline origin of the VK germline segments. b. The patterns of mutations at each of the Vernier and interface positions were examined as follows: i. Formula 1 (Makowski & Soares, 2002, incorporated by reference in its entirety) Bioinformatics, 2003, 19: 483) is a To calculate diversity indices for positions, interface positions, CDRL1, and CDRL2 Used.
number
[0191] Those skilled in the art will appreciate that the procedures outlined above also work well for diversifying Vγ germline sequences. Libraries containing Vγ chains can also be used to select positions. will readily recognize that such modifications are within the scope of the invention.
[0192] In addition to framework mutations, diversity also exists within CDRL1 and CDRL2. This is because the residues in CDRL1 and CDRL2 are different from those used above. The present invention also provides a synthetic library of genetic information, which is used to determine the diversity within a particular germline in a VK dataset. The two to four most frequent mutations in CDRL1 and CDRL2 in the library This was achieved by integrating the VK1-5 germline clone into the CDRL2 region of the VK1-5 germline. Thus, these alternatives did not arise from allelic variation. Table 3 shows the presently exemplified variants of the present invention. Nine light chain chassis and their frameworks and C The polypeptide sequences of the DR L1 / L2 mutants are shown. The amino acid residues at the / L2 position are diversified as follows (Chothia-Le sk numbering system; Chothia and Lesk, J. Mol. Biol., 1987, 196: 901). 1. Position 28: Germline S or G was optionally changed to G, A, or D. 2. Position 29: germline V was optionally changed to I. 3. Position 30: Germline S was optionally changed to N, D, G, T, A, or R. 4. Position 30A: Germline H was optionally changed to Y. 5. Position 30B: Germline S was optionally changed to R or T. 6. Position 30E: Germline Y was optionally changed to N. 7. Position 31: Germline S was optionally changed to D, R, I, N, or T. 8.Position 32: Germline Y or N was optionally changed to F, S, or D. 9.Position 50: Germline A, D, or G were optionally changed to G, S, E, K, or D. 10.Position 51: Germline G or A was optionally changed to A, S, or T. 11.Position 53: Germline S or N were optionally changed to N, H, S, K, or R. 12.Position 55: The germline E was optionally changed to A or Q.
[0193] Example 2. Light chain library with increased diversity in CDRL3 Various methods for generating light chain libraries are known in the art ( See, e.g., U.S. Patent Application Publication Nos. 2009 / 0181855 and 2010 / 0056386 (See, for example, International Publication No. WO / 2009 / 036379). Analysis of the antibody sequences revealed that these sequences were germline-like VL-JL (herein referred to as VL-JL) sequences prior to somatic mutation. (where "L" can be a kappa or lambda germline sequence) We showed that germline-like rearrangements are rarely found in the V and J regions (Fig. 2). For purposes of this specific example, the CDRL 3 is a rearrangement limited to 8, 9, or 10 amino acids in length (U.S. Patent Application Publication No. 2009 / 0181855, 2010 / 0056386, and International Publication No. WO However, the IGHJK1 gene The WT (Trp-Thr) and RT (Arg-Thr) sequences (the first two N-terminal Both of these sequences (residues) are considered to be "germline-like" and therefore contain such sequences. All L3 rearrangements are considered to be "germline-like." Therefore, the new light chain library At the same time, (1) the sequence minimizes deviation from the germline-like sequence, as defined above. (2) It was designed and constructed with the goal of generating maximum diversity. Maximize the types of diversity shown to be most favorable by previously validated antibody sequences. In particular, the designed library consisted of length-matched germline sequences. We sought to maximize the diversity of CDRL3 sequences that differed by no more than two amino acids.
[0194] This is because the light chain oligonucleotide design does not include "jumping dimers" or "jumping This was achieved by utilizing the "jumping dimer" approach. , each of the six positions of CDRL3 (L3-VL) encoded by the VL segment At most two positions are required for each individual L3-VL Although the sequence differs from the germline, the two positions do not have to be adjacent to each other. The total number of designed degenerate oligonucleotides synthesized per VL chassis was 6! / (4!2!) or 15 (VL and VL for each kappa germline chassis) Six of the most commonly occurring amino acids at the junction between the α and β positions (position 96) are described. (i.e., F, L, I, R, W, Y, and P; for the amino acid at the junction at position 96 For more details regarding the present invention, see U.S. Pat. Nos. 5,629,293, 5,629,550 ... Publication Nos. 2009 / 0181855 and 2010 / 0056386 and International See Publication No. WO / 2009 / 036379. This approach is similar to the jumping dimer approach, but the two Instead of one, three positions differ from the germline in each L3-VL sequence. The degenerate code selected for each position in the jumping dimer and trimer approach The sequences are (1) a large number of sequences contained in the known repertoire of publicly available human VK sequences; (See Appendix A of International Publication No. WO / 2009 / 036379) (2) N-linked glycosylation motifs (NXS / NXT), Cys residues, Resulting codons, such as stop codons and deamidation-prone NG motifs, Selected to minimize or eliminate undesirable sequences within CDRL3 of the synthetic light chain Table 4 shows the VK1-39 nucleotide sequence with a length of 9 amino acids and an F or Y junction amino acid. The 15 degenerate oligonucleotides encoding the CDRL3 sequences and the corresponding degenerate polynucleotides were Tables 5, 6, and 7 show the peptide sequences of CD8, CD9, and CD10, respectively. VK sequences of exemplary jumping dimer and trimer libraries for RL3 length Oligonucleotide sequences for each and the corresponding sequence for CDRL3 to provide.
[0195] The number of unique CDRL3 sequences in each germline library is then counted, A different light chain library, designated "VK-v1.0", for each of the three lengths The number of unique CDRL3 sequences was compared to that in U.S. Patent Application Publication No. 2009 / 01 (See Example 6.2 in U.S. Pat. No. 81855.) Table 8 shows the respective germline libraries. The number of unique CDRL3 sequences in each of the libraries is provided.
[0196] FIG. 3 shows the results of the comparison of the nucleotide sequences containing no mutations from the germline-like sequence (Table 1) or the germline-like sequence. 9 amino acid CDs containing 1, 2, 3, or 4 or fewer mutations from the sequence Jumping dimers with RL3 length and sequences in the VK-v1.0 library The naturally occurring VK1-05 sequence has a nucleotide sequence at Kabat position 95. , which has almost as many Ser (germline amino acid type) as Pro, and therefore Both residues (S and P) were incorporated into the synthetic library representing the VK1-05 repertoire. However, as shown in Table 1, the VK gene was VK1-05. In some cases, only Ser is considered to be the germline-like residue at position 95 for the purposes of this analysis. The plot for VK3-20 shows that for length 9, The sequences in the VK1-05 library are representative of the remaining chassis. Within 3 amino acids of the germline sequence, approximately 63% of the sequence is human germline-like. The remainder of the library was designed as follows: 100% of the sequences are within 2 amino acids of the human germline-like sequence, and therefore More than 95% of the sequences of length 9 in the jumping dimer library considered were It was within 2 amino acids of the germline-like sequence. In comparison, the 9 amino acid long VK-v1 Only 16% of the members of the .0 library matched the 2 amino acid sequence of the corresponding human germline-like sequence. For length 8, the sequence in the jumping dimer library is in the range of 98% are within 2 amino acids of the germline-like sequence compared to approximately 19% for VK-v1.0 For length 10, more than 95% of the sequences in the jumping dimer library were , which was within 2 amino acids of the germline-like sequence compared to approximately 8% for VK-v1.0.
[0197] In some embodiments, the most likely solvent-exposed portion of the folded antibody is To focus on the diversity at high positions, positions 89 and 90 (Kabat numbering) The sequences are not modified from the germline - they are most often QQ, but the sequence is VK. For the 2-28 chassis, it is MQ. Other VK germline genes are located at positions 88-89. The use of these genes as chassis with different sequences is also within the scope of the present invention. For example, VK1-27 has QK, and VK1-17 and VK1-6 both have LQ, etc. The sequences at these positions are known in the art. See, for example, Scaviner et al., which is incorporated by reference in its entirety. , Exp. Clin. Immunogenet., 1999, 16: 234 (See Figure 2).
[0198] CDRH3 Library The following examples demonstrate improved libraries compared to those known in the art. Methods and compositions useful for the design and synthesis of antibody libraries containing CDRH3 sequences are provided. The CDRH3 sequences of the present invention are described in detail below. Synthetic CDRH3 segments that have increased diversity compared to human sequences but retain the characteristics of the human sequence. and / or improve the combinatorial efficiency of the synthetic CDRH3 sequence and one Alternatively, improving the match between human CDRH3 sequences in multiple reference sets.
[0199] Example 3. Generation of a curated reference set of human pre-immune CDRH3 sequences A file containing approximately 84,000 human and mouse heavy chain DNA sequences is available from BLA Downloaded from the ST public resource (ftp.ncbi.nih.gov / bl ast / db / FASTA / ; File name: igSeqNt.gz; Download date: 2 (August 29, 2008). Of these approximately 84,000 sequences, approximately 34,000 The sequence header annotation These sequences were then identified as human heavy chain sequences based on the analysis of the The sequences were filtered as follows: first, all sequences were filtered to their nearest neighbors. The VH regions were grouped according to the VH germline (incorrect or poorly matched). sequences that could not be matched due to insufficient length or extensive mutations. Second, the DNA sequences were significantly different when compared to their corresponding germline VH sequences. Any sequences containing more than five mutations at the same level were also discarded. nd Milstein, EMBO J., 2001, 20: 4570 Thus, mutations (or lack thereof) in the N-terminal portion of the variable region may be present in the C-terminal portion of the variable region. Conservative surrogates for mutations (or lack thereof) in the may be used as a conservative surrogate The CDRH3 variants were designed to have 5 or fewer nucleotide mutations in the N-terminal VH region. Selection of only sequences that are slightly or not mutated (i.e. It is also likely that CDRH3 sequences with higher pre-immune characteristics will be selected.
[0200] After translating the remaining DNA sequence into their amino acid counterparts, the heavy chain germline amino acid sequence Identify the appropriate reading frame containing the CDRs, including the sequence of CDRH3. The list of CDRH3 sequences obtained at this point was at least At least three amino acids were different from any other sequence in the set (for length After matching, we further filtered the members to reduce the number of members. This resulted in 11,411 CDRH3 sequences, of which 3,571 sequences were identified in healthy adults. Annotate as derived from "healthy pre-immune set" or "HPS"; for GI numbers (See Appendix A for details), and the other 7,840 sequences were derived from individuals with the disease. were annotated as being of fetal origin or of antigen-specific origin The method described below then separates the four segments that make up CDRH3: 1) TN1, (2) DH, (3) N2, and (4) H3-JH sequences in HPS Each was used to deconvolute.
[0201] Example 4. Segments from a theoretical segment pool are matched to CDRH3 in the reference set How to match This example shows the TN1, DH, N2, and H3-JH segments of CDRH3 in HPS. The methods used to identify markers are described. Design and synthesis of human CDRH3 sequences. Currently exemplified approaches to this synthesis involve the human immune system generating a pre-immune CDRH3 repertoire. This mimics the segmented VDJ recombination process that produces the VDJ fragments described herein. The matching method is to match a specific CD across a reference set (e.g., HPS) of CDRH3. Which TN1, DH, N2, and H3-JH segments are used to produce RH3? This information is then used to determine which TNs from the theoretical segment pool are being used. 1, DH, N2, and H3-JH segments (or TN1 and N2 in the case of reference The CDRH3 library (segments extracted from the CDRH3 sequences in the set) Optionally, other information described below (such as It is used together with the physical and chemical properties.
[0202] The inputs to the matching method are: (1) a reference set of CDRH3 sequences (e.g., HPS (2) multiple TN1, DH, N2, and / or A theoretical segment pool containing the H3-JH segment. The manner in which members of the set are generated is described more fully below. For each CDRH3, the matching method produces two outputs: (i) the theoretical The closest matching C can be generated using segments from the segment pool. Generate a list of DRH3 sequences and (ii) their closest matching CDRH3 sequences. One or more segments from a theoretical segment pool that can be used to A combination of
[0203] The matching method was performed as follows: TN1 segments in the theoretical segment pool. The segments correspond to the first amino acid (position 95) and The alignment was performed at the first amino acid of each segment. All segments that contribute to the The retained TN1 segment is then truncated to the [TN1]-[DH] segment. All DH segments from the theoretical segment pool are concatenated to generate a DH segment. These segments were then aligned as above to form the [TN1]-[DH] segments. The best match for each segment is kept. The procedure is The length of the RH3 sequence is determined by combining segments from a theoretical segment pool. Until the same is reproduced, [TN1]-[DH]-[N2] and [TN1]-[DH]-[ Repeated by [N2]-[H3-JH] segments. All combinations of segments that result in the best match are considered the output of the matching method. and retain it.
[0204] Table 9 shows a theoretical segment pool called "Theoretical Segment Pool 1" or "TSP1." The output of the matching method using the top pool, specifically the four individual sequences from HPS, We provide an example of the output for TSP1. TSP1 uses several theoretical segment pools, i.e. , 212 TN1 segments (Table 10), 1,111 DH segments (Table 11), 14 Contains 1 N2 segment (Table 12) and 285 H3-JH segments (Table 13) The CDRH3 sequence in test case 1 is a unique combination of four segments. Test Case 2.1 and 2.2 and 2.3 are two distinct combinations of different TN1 and DH segments, respectively. Therefore, the same match is obtained in test cases 3.1, 4.1, and 4.2. The closest matches all deviate by just one amino acid from the reference CDRH3, and T One (3.1) or two (4.1 and 4.2) combinations of segments from SP1 This approach can be achieved by All closest matches to any reference CDRH3 sequence as well as the reference CDRH3 All of the segments in the theoretical segment pool that can accurately produce a sequence Can be generalized to find combinations and / or their closest matches .
[0205] Example 5. Derivation of theoretical segment pools of H3-JH segments Theoretical sequences of the H3-JH segment considering inclusion in the synthetic CDRH3 library To generate a fragment pool, the following method involves subcloning the 12 germline IGHJ sequences of Table 14. Seven of them (IGHJ1-01, IGHJ2-01, IGHJ3-02, IGHJ4-0 2, IGHJ5-02, IGHJ6-02, and IGHJ6-03) These seven alleles were selected because they were the most identical in the human sequence. These were chosen because they were among the commonly occurring alleles. All sequences in Table 14 (some with Libraries using sequences that differ only in FRM4 are also within the scope of the present invention. Method for generating junctional diversity during the VDJ recombination process in vivo This is intended to simulate the nucleotide sequence for germline gene segments. This is done by enzyme-mediated addition and deletion of nucleotides. The method proceeds as follows: This results in a completely enumerated theoretical segment pool of H3-JH segments. 1. Pretreatment yields well-known JH framework regions, leading to translation of the JH segment. two nucleotide bases at their 5' end before the first nucleotide that encodes the translation IGHJ genes containing partial codons consisting of IGHJ3-02, IGHJ4-0 2, IGHJ5-02, IGHJ6-02, and IGHJ6-03). For example, the IGHJ3-02 gene encodes the JH segment that produces the JH framework region. It contains an AT dinucleotide sequence before the first nucleotide encoding the translation (Figure 4, (Top). All partial codons consisting of two nucleotide bases are The ' position was completed using all possible nucleotide doublets (i.e., NN). (Figure 4, top, second row for IGHJ3-02). More specifically, in the germline sequence The 5'-most nucleotide in the It was added 5' to the do. 2. IGHJ genes IGHJ1-01 (Figure 4, center) and IGHJ2-01 (Figure 4, bottom) ) is the first nucleotide that encodes the translation of the JH segment that produces the JH framework region. They contain zero and one nucleotide bases at their 5' ends before the nucleotides. For these genes, the pretreatment described in step 1 was not performed. The 5' doublet was mutated to NN (Fig. 4, middle and bottom, second row of each). Therefore, after performing this step, the seven IGHJ genes listed above Each was converted to a mutant with NN doublets as its first two 5' positions. . 3. The 5' codon of the sequence produced by steps 1 and 2 is then deleted, resulting in The first two bases of the resulting DNA sequence are then mutated to an NN doublet. (Figure 4, columns 3-4 for all). 4. The 5' codon of the sequence produced in step 3 is then deleted, resulting in The first two bases of the resulting DNA sequence were subsequently mutated to an NN doublet (Figure 4 , all columns 5-6). 5. Next, each of the polynucleotide sequences generated by steps (1) to (4) is This involves the reading frame for each sequence that produced the JH framework region. A theoretical segment consisting of 248 parent H3-JH polypeptide segments (Table 15) was obtained from the Translated to get the top. 6. The parent H3-JH polypeptide segment contains only the portion of the JH segment that includes FW4. one amino acid at a time until only one H3-JH segment remains (i.e., a zero amino acid long H3-JH segment). The method described above was carried out by truncating the N-terminus of the nucleotides by removing the amino acid. This resulted in the production of a theoretical segment pool of 285 H3-JH segments (Table 13). .
[0206] Example 6. Derivation of theoretical segment pools of DH segments The two theoretical pools of DH segments are "Translation Method 0" (TM0) or "Translation Method 1" (TM2). and (ii) the translation is generated using one or more of two translation methods, referred to as "Translation Method 1" ("TM1"), Each of the 27 human germline IGHD DNA sequences or segments derived therefrom was The sequences were run in the forward reading frame (Table 16).
[0207] 1K DH Theoretical Segment Pool (1K DH) TM1 is the "1K DH theoretical segment pool" ("1K DH"; 1,1 in Table 11). 11 DH segments). It contains two untranslated bases after translation in either of the forward reading frames. IGHD sequences with partial codons containing two bases can be completed with only one Only when the amino acid can be coded for is the complete codon completed to produce a complete codon. For example, a DNA sequence such as TTA-GCT-CG has two sequences that are translated into LA. The complete codon for R and any of CGA, CGC, CGG, or CGT encodes R. Therefore, there is a remaining partial codon (CG) that can code for only R. Therefore, applying TM1 to this sequence will result in LAR. Partial codons (e.g., GA or AG) that can code for more than one amino acid For sequences with T, partial codons were ignored. By applying M1, the theoretical segments containing the 73 DH parent segments of Table 17 were obtained. A pool of nucleotides was generated (some containing a stop codon ("Z") and an unpaired Cys residue). These sequences are then split into their N and C positions until only two amino acids remain. The truncated segments were deleted sequentially at the amino acid level at the termini. containing an unpaired Cys residue, an N-linked glycosylation motif, or a deamidation motif. If it contained any fragments, they were truncated. This process resulted in the 1,111 DH segments in Table 11 was brought about.
[0208] 68K DH Theoretical Segment Pool (68K DH) The 27 IGHD genes and alleles in Table 16 are separated by 4 bases until four bases remain. Consecutive deletions at either or both the 5' and 3' ends, resulting in four or more nucleotides 5,076 unique polynucleotide sequences of these genes were obtained. The sequences of 6 may have 0, 1, and / or 2N nucleotides at their 5' and / or 3' ends. The resulting sequences were translated using TM0. In TM0, only complete codons are translated, and partial codons (i.e., 1 or 2 bases) are translated. This method ignores stop codons, unpaired Cys residues, and N-linked glycosylation moieties. Asn at the last or next to the last position that can lead to a deamidation motif After elimination of the corresponding segments, 68,374 unique DH polypeptide segments were obtained. ("68K DH Theoretical Segment Pool") as input to PYTHON. Using the IGHD genes in Table 16, the computer code provided below This would reproduce the exact theoretical segment pool of 374 DH segments. There are two free parameters in the program: (1) the remaining DNA after stepwise deletion; (1) the minimum length of the A sequence (four bases in this example) and (2) the number of bases in the theoretical segment pool. The minimum length of a peptide sequence that is acceptable for inclusion (in this example, 2 amino acids). These parameters can be changed to modify the output of the program. For example, changing the first parameter to one base and the second parameter to one amino acid. Furthermore, it has 68,396 unique sequences, each containing 18 single amino acid segments. This would lead to a larger theoretical segment pool. H segment, e.g., 1, 2, 3, or 4 or more amino acids or Cleavage to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides before Also within the scope of the present invention.
[0209] A Python computer program for generating 68,374 DH segments
number
number
number
number
number
number
[0210] Example 7. Derivation of theoretical segment pools of TN1 and N2 segments The libraries of this example are, in some instances, libraries known in the art. Compared to other libraries, they have larger TN1 and N2 segments. The diversity of the TN1 and N2 segments is designed to be sufficient for HPS. The CDRH3 sequences in each of these segments (i.e., TN1, DH, N2, and and H3-JH), and then To extract the novel TN1 and N2 segments, the nucleotide sequences described in Example 4 were used. For the purposes of this invention, the "novel" T The N1 and N2 segments are theoretical segments that match a reference set of CDRH3 sequences. The following are the TN1 and N2 segments that do not appear in the top pool: This is an example of the method used to extract the relevant TN1 and N2 segments. Any theoretical segment containing a TN1, DH, N2, and / or H3-JH segment. The fragment pools were used to identify novel TN1 and TN2 sequences from any reference set of CDRH3 sequences. It can be generalized to extract N2 segments.
[0211] Table 9 shows the HPS-derived reference C sequences using theoretical segment pool 1 ("TSP1"). DRH3 sequence ERTINWGWGVYAFDI (SEQ ID NO: 8760) Test Case 5. The matching results for the nucleotides 1 to 5.4 are provided. The best match against the reference CDRH3 is The four CDRH3 sequences are each a 3-amino acid range of the reference CDRH3 sequence. In each of these matches, TN1, DH, N2, and H3-JH The segments are 4, 3, 3, and 5 amino acids in length, respectively. CDRH3 can be deconvoluted into the following segments: ERTI- NWG-WGW-YAFDI (SEQ ID NO: 8761) (That is, [TN1]-[D H]-[N2]-[H3-JH]). DH and H3-JH segments from reference CDRH3. NWG and YAFDI are the partners (SEQ ID NO: 4540) are the same in TSP1, respectively. However, TN1(ERTI) derived from reference CDRH3 (SEQ ID NO: 87 18) and N2(WGW) segments are absent in TSP1, and one or more These "novel" TN1 segments match the TSP1 segment with several amino acid mismatches. and N2 segments were extracted from reference CDRH3 and analyzed to determine possible theoretical inclusions. It is believed that the segment pool and / or synthetic library. and N2 segment by applying this analysis to all members of the HPS. To ensure the identification of TN1 and N2 sequences, the extracts were cloned using the reference CDRH3 and and the DH and H3-JH segments in TSP1 cumulatively contain at most three amino acids. Only CDRH3 sequences that result in mismatches are run, which is the reference CDRH This implied that the DH and H3-JH segments of 3 were reliably assigned.
[0212] Example 8. Calculation of segment usage weight The segment usage weight is calculated based on the theoretical segment pool (e.g., TSP1 and TS P1 and the novel TN1 and N2 segments identified as described in Example 7. Their usefulness in determining which segments from the nucleotide sequence should be included in the synthetic library The segment usage weights were calculated based on the matching method and and by using Eq. 2 we obtain:
number
[0213] The procedure for calculating segment usage weights is further illustrated below. In each example, the best match combination from TSP1 is a single CD For the RH3 sequence, Sm=1, the degeneracy (k) and fragment mismatch (f) are given. Dependency weight calculation is explained.
[0214] Example 8.1. Segment Usage Weights for Test Case 1 in Table 9 See Test Case 1 in Table 9. CDRH3 sequence RTAHHFDY (array Number 3660) is the unique segment combination (g = 1) of TSP1(f = 1, subscripts omitted for brevity). Table 18 shows the test case The segment corresponding to the best match from TSP1 for CDRH3 of SEQ ID NO:1 Provides the weights used.
[0215] Example 8.2. Segment Usage for Test Cases 2.1 and 2.2 in Table 9 Weight See test cases 2.1 and 2.2 in Table 9. CDRH3 sequence VGI VGAASY (SEQ ID NO: 3661) is a combination of two separate segments (g=2) Table 19 shows the results of Test Case 2.1. and the segment corresponding to the best match from TSP1 for CDRH3 of 2.2 Provides the weights used for.
[0216] Example 8.3. Segment Usage Weights for Test Case 3.1 in Table 9 See test case 3.1 in Table 9. CDRH3 sequence DRYSGHDLG Y (SEQ ID NO: 3662) is a unique segment combination with only one amino acid difference. It can be similarly located in TSP1 by combining (g=1). As provided below, The TN1, N2, and H3-JH segments match similarly to the corresponding reference sequence fragments. However, four of the five DH amino acids are matched as well. HPS-derived sequence: DR-YSGHD-LG-Y (SEQ ID NO: 3662) Nearest neighbor in TSP1: DR-YSGYD-LG-Y (SEQ ID NO: 8719) Therefore, here, f = 4 / 5 for the DH segment and = TN1, N2, and H3-JH segments The result is 1 for (Table 20).
[0217] Example 8.4. Matching Test Cases 4.1 and 4.2 in Table 9 See test cases 4.1 and 4.2 in Table 9. CDRH3 sequence GIA AADSNWLDP (SEQ ID NO: 3663) There are two amino acid differences, each with only one amino acid difference. The following distinct segments can be combined (g=2) to form TSP1: The TN1, DH, and N2 segments are shown in Table 1 as corresponding reference sequence fragments. It matches as well, but five of the six H3-JH amino acids match as well. HPS-derived sequence: (-)-GIAAA-D-SNWLDP (SEQ ID NO: 3663) Nearest neighbor in TSP1: (-)-GIAAA-D-SNWFDP (SEQ ID NO:8 720) HPS-derived sequence: G-IAAA-D-SNWLDP (SEQ ID NO: 3663) Nearest neighbor in TSP1: G-IAAA-D-SNWFDP (SEQ ID NO: 8720 ) Here, (-) indicates an "empty" TN1 segment. Application of Equation 2 results in the segment usage weights provided in Table 21.
[0218] Example 8.5. Calculation of segment usage weights for test cases 1 to 4.2 in Table 9 The individual calculations described above include all test cases 1 to 4.2 in Table 9 simultaneously. Expanding this yields the segment usage weights in Table 22.
[0219] Example 8.6. Segment usage weights for test cases 5.1 to 5.4 in Table 9 calculation CDRH3 sequence ERTINWGWGVYAFDI (SEQ ID NO: 8760) and examples 7, which referenced novel TN1 and N2 segments extracted from the CDRH3 sequence. In this case, the novel TN1 and N2 segments (ERITs, respectively) (SEQ ID NO: 871 8) and WGV) and the DH and H3-JH segments from TSP1 (respectively NWG and YAFDI (SEQ ID NO: 4540) ) are each assigned a single usage weight. It can be guessed.
[0220] Example 9. TN1, DH, N2, and JH Segments for Inclusion in Synthetic Libraries Selection of ment Figure 5 provides the general methodology used for the design of the synthetic CDRH3 library. The method involves, as input, (1) a sequence containing the TN1, DH, N2, and H3-JH segments; The theoretical segment pool (e.g., TSP and novel TN1 and N2 segments) ) and (2) a set of reference CDRH3 sequences (e.g., HPS). From this, a specific subset of segments from the theoretical segment pool can be identified as a physical CDR. Selected for inclusion in the H3 library.
[0221] First, the best match of HPS to CDRH3 is determined using the matching method described above. Using the method, TSP1 with or without the novel TN1 and N2 segments was This data was then used to calculate the segment usage weights using Equation 2. The segments were used to identify their relative positions in the CDRH3 sequence of HPS. Frequency of occurrence (indicated by segment usage weights) and hydrophobicity, alpha helicity and expressibility in yeast. of inclusion in the physical library based on other factors (described more fully below). Therefore, we prioritized them.
[0222] Example 9.1. Exemplary Library Design (ELD-1) ELD-1 contains as input the HPS and TSP1-derived segments (9.5 × 1 0 9 members) and ordered by their usage weight in the HPS. From TSP1, 100 TN1, 200 DH, 141 N2, and 1 00, resulting in an output of the H3-JH segment of 2.82 x 10 8 Theoretical complexity The segments corresponding to ELD-1 are provided in Table 23. Here, all combinations of segments (i.e., TN1, DH, N2, and H3- JH) as well as individual sets of segments (i.e., TN1 only, DH only, N2 only, and It was noted that each of the H3-JH and H3-JH segments constituted a theoretical segment pool. stomach.
[0223] Example 9.2. Exemplary Library Design 2 (ELD-2) The inputs for this design were HPS and a segment from TSP1 and HPS. The extracted novel TN1 and N2 segments (Example 7). The output is: (1) TS 200 DH and 100 H3-JH segments from P1 and (2) 100 TN1 and 200 N2, which originally contained the TN1 and N2 segments of TSP1 The sequences were extracted from the segments and HPS. In Example 7, to extract the N2 segment (i.e., not included in TSP1), Application of the described method resulted in the identification of 1,710 novel TN1 segments and 1,024 novel This resulted in the identification of the N2 segment, which corresponds to ELD-2, as presented in Table 24. where all combinations of segments (i.e. TN1, DH, N2, and and H3-JH) as well as individual sets of segments (i.e., TN1 only, DH only, N2 Note that the H3-JH and H3-JH segments (only H3-JH and only H3-JH) constitute the theoretical segment pools, respectively. Please note that, as in ELD-1, all segments in ELD-2 are Para-HPS was selected for inclusion based on their usage weights.
[0224] Example 9.3. Exemplary Library Design 3 (ELD-3) The inputs for this design are the same as for ELD-2. As in ELD-2, the output The forces were (1) 200 DH and 100 H3-JH segments from TSP1, respectively; and (2) a set of 100 Ts originally containing the TN1 and N2 segments of TSP1. Extracted from the set of N1 and 200 N2 segments and sequences in HPS However, due to the selection of segments for ELD-3, The approach used differs in two respects: the first selected physicochemical The physical properties (hydrophobicity, isoelectric point, and alpha-helical propensity) are In addition to the segment usage weight to prioritize segments for inclusion in Hydrophobicity was experimentally demonstrated in poorly expressed antibodies isolated from yeast-based De-prioritize disproportionately abundant hydrophobic DH segments The isoelectric point and tendency to form alpha helices are well known in the art. The physicochemical properties of the CDRH3 library, which is known in the literature, were relatively unexplored. The method was used to identify segments located in regions of spatial space (see, e.g., U.S. Pat. App. Pub. No. 2005 / 0100222). Publication Nos. 2009 / 0181855 and 2010 / 0056386 and International Publication Second, the segment usage weights are calculated based on the HPS data. These methods are described more fully below. The segments corresponding to ELD-3 are provided in Table 25. All combinations of the elements (i.e., TN1, DH, N2, and H3-JH) and Individual sets of segments (i.e., TN1 only, DH only, N2 only, and H3-JH Note that each of the segments (i.e., ∇ ...
[0225] Example 9.3.1. Generating Segment Usage Weights with Bootstrap Analysis Bootstrap analysis (Efron & Tibshirani, An Intro duction to the Bootstrap, 1994 Chapman H (Sill, New York) is a widely used method for estimating the variability of statistics in a given sample. This estimate is based on a sample with a size equal to the original sample and with The statistical values calculated on several subsamples derived from it by sampling Members of the original sample are randomly selected to form subsamples. and typically included multiple times in each subsample (hence the term "restored sample"). ring").
[0226] Here, the original sample is the HPS dataset with n=3,571 members. The statistics are segment usage weights. Each has 3,571 members. The 1000 subsamples were generated by randomly selecting sequences from the HPS dataset. generated (allowing at most 10 repeats of a given sequence in each subsample) The matching method described above was then applied to each subsample, The final segment usage weight is calculated as the average of the values obtained for the individual subsamples. The mean values derived by this bootstrap procedure were calculated using the parent HPS dataset. Unless otherwise indicated, the values are calculated based on 1,000 subsamples. These average values of the samples were used in the selection of segments for ELD-3.
[0227] Example 9.3.2. Amino Acid Characteristic Index AAind available online at www.genome.jp / aaindex / The ex database provides various physicochemical and biochemical information for amino acids and amino acid pairs. It provides over 500 numerical indices that indicate the characteristics of a given property. In many cases, several useful indices are available, including hydrophobicity, electrostatic behavior, secondary structure propensity, and and other characteristics. The following three indices are based on the well-understood Kyte-Dooli Starting with the ttle hydropathic index (KYJT820101), Selected by adding the most numerically de-correlated indices They therefore potentially account for non-overlapping regions of amino acid property space and EL The DH and H3-JH segments for D-3 were used for analysis and selection. 1.KYTJ820101 (Hydrophilic Index) 2.LEVM780101 (normalized frequency of alpha helix) 3.ZIMJ680104 (isoelectric point)
[0228] Example 9.3.3. Hydrophobic DH Segments Are Isolated from a Yeast-Based Library , disproportionately overrepresented in poorly expressed antibodies Based on the protein expression levels of approximately 1200 antibodies expressed in Saccharomyces cerevisiae Antibodies were classified as "good" or "poor" expressors. The CDRH3 sequences of each antibody in each class were compared to the sequences that correlated with expression levels. Sequence features were examined to identify the sequence features. One such sequence feature is KYTJ820101 The hydrophobicity of the DH segment is calculated using the index. Provides the frequency of "good" and "bad" expressions according to gender (increasing to the right). The distribution is also consistent with that expected from the synthetic library used to isolate these antibodies. and serves as a reference ("design"). The DH segment was disproportionately overrepresented among the "poor" expressors (compared to design expectations). ), was disproportionately underrepresented among the "good" expressors. (far left) are disproportionately overrepresented among "good" expressors and disproportionately overrepresented among "poor" expressors. From this data, it can be seen that the overall expression capacity of the library antibodies was due to the hydrophobic DH It is speculated that this could be improved by synthesizing fewer CDRH3 sequences. was done.
[0229] Example 9.3.4. Selection of 200 DH Segments for Inclusion in ELD-3 A set of 71 DH segments from TSP1 was selected for automatic inclusion in ELD-3. These segments are referred to as the "core" DH segments for the following desirable It had the following characteristics: 53 of the 1.71 are based on segment usage weights derived from bootstrap analysis. It was in the top 7% of ordered DH segments. 18 of the 2.71 antibodies were isolated from libraries expressed in Saccharomyces cerevisiae It is in the top 7% of DH segments ordered by usage weight derived from Ta.
[0230] The remaining 1,040 segments were designated as "non-core." To complete the set of 00 segments, add the 129 segments in the following way: were drawn from the segment's "non-core" pool. The 1.65 segments are N-linked via combination with (a) N2 amino acids. As at or next to the last position that has the potential to form a type glycosylation motif n residue or (b) contained the amino acid sequence NG involved in deamidation, and therefore was reduced. 2. KYTJ820101 Median value for hydropathic index (median = 1K DH) Segments higher than 2.9 were eliminated from further consideration. Given the known importance of Tyr (each incorporated by reference in its entirety), Fellouse et al., PNAS, 2004, 101: 12467; and Hofstadter et al., J. Mol. Biol., 199 9, 285: 805), segments containing at least one Tyr residue are Unless the KYTJ820101 value is in the highest hydrophobicity quartile (KYTJ820101 value higher than 9.4), This reduced the number of segments by 443. The final set of 3.129 segments is the following variables: (1) for nearest neighbors (2) multidimensional structures defined by the values of three physicochemical property indices Euclidean distance between the "core" and the remaining 443 "non-core" segments in the original space was obtained by using an objective function that aimed to maximize the distance between
[0231] Example 9.3.5. Selection of 100 H3-JH Segments for Inclusion in ELD-3 Choice The 100 H3-JH segments were selected for inclusion in ELD-3 in the following manner: I chose it. The H3-JH segments of 1.28 are similar to other 1.28 genes containing only these H3-JH segments. The library was selected after experimental validation (U.S. Patent Application Publication No. 2009 / 000266). 181855 and 2010 / 0056386 and International Publication No. WO / 200 (See 9 / 036379). The 2.57 segments were used with weights derived from the bootstrap analysis described above. Therefore, based on their presence within the top 25% of the ordered H3-JH segments (1) Of these, 57 H3-JH segments and 28 H3-JH segments were selected. These segments (i.e., a total of 85 segments) are referred to as the "core" H3-JH segments. These, like the core DH segment, were automatically included in ELD-3. Further segments of 4.15 are based on the following variables: (1) amino acid sequence relative to the nearest neighbor mismatch and (2) in a multidimensional space defined by the values of three physicochemical property indices. The Euclidean distance between the "core" and the remaining 200 "non-core" segments is The selection was made by using an objective function that aimed to maximize
[0232] Example 9.3.6. 100 TN1 and 200 N2 for inclusion in ELD-3 Selecting a Segment The TN1 and N2 segments were analyzed in the respective subsamples of the bootstrap procedure. The 100 TN1 and TN2 sequences with the highest average segment usage weights were extracted from the sequences. The N2 segments of 200 and 200 contain undesirable motifs, namely Cys and Asn residues. After elimination of sequences bearing the group, the nucleotide sequence was selected for inclusion in the library.
[0233] Example 9.3.7. Nucleotides encoding segments selected for inclusion in ELD-3 Nucleotide sequence selection Each of the polypeptide segments selected for inclusion in the library is The polypeptide must be translated back into the corresponding oligonucleotide sequence (DNA Many oligonucleotides are required for each polypeptide due to the degeneracy of the genetic code. Although it is possible to encode a gene segment, a more desirable oligonucleotide is We imposed certain constraints on the selection. First, we investigated whether ELD-3 is a nucleotide sequence that can be expressed in yeast (Saccharomyces cerevisiae). Since the gene was expressed, codons rarely used in yeast were avoided. For example, Ar Of the six possible codons for g, three, CGA, CGC, and CGG, are Less than 0% of genes are used to encode yeast proteins (e.g., Nakamur a et al., Nucleic Acids Res., 2000, 28:2 92), therefore, those three codons were avoided to the extent possible. Second, many antibodies are produced in Chinese hamster ovary (CHO) cells. (e.g., after its discovery in yeast), the CCG codon (encoding Pro) also Because it is rarely used by hamsters (Nakamura et al. ), avoided.
[0234] A number of restriction enzymes were used during the actual construction of the CDRH3 oligonucleotide library. (See Example 10 of US Patent Application Publication No. 2009 / 0181855). Therefore, recognition motifs for these restriction enzymes within the CDRH3 polynucleotide sequence It is desirable to avoid the presence of codons for restriction enzymes that may be used downstream. Selection is made at the individual segment level to avoid introducing recognition motifs for the Motifs such as ,The combination of segments is also checked and codons are ,considered whenever possible. In particular, the three restriction enzymes currently exemplified were altered to reduce the presence of specific motifs. The following primers were used during the construction of the CDRH3 library: BsrDI, BbsI, and Avr II. The first two are type II enzymes with non-palindromic recognition sites. The opposite strand of the oligonucleotide encoding the nucleotide sequence is attached to the recognition sites for these two enzymes. In particular, the opposite strand contains the motifs GCAATG and CATTGC ( for BsrDI) and GAAGAC and GTCTTC (for BbsI) The recognition motif for AvrII is a palindrome. Therefore, the oligonucleotides were checked only for the sequence CCTAGG. While AvrII is only used to process the TN1 segment, There is no need to evaluate its presence in other segments or combinations thereof.
[0235] Tasked with improving the engineering of polypeptides to polynucleotides The additional constraints introduced are likely to increase errors during solid-phase oligonucleotide synthesis. Therefore, the avoidance of consecutive sequences of 6 or more bases of the same type was the aim. The DNA sequences for the three segments were chosen to avoid such motifs. The DNA sequences for the ELD-3 segments are shown in Table 25. Those skilled in the art will appreciate that these methods may also be used with any other library, any Restriction sites, any number of nucleotide repeats, and / or any desired sequence in any organism. It should be noted that this can also be applied to avoid the presence of any codons that are considered to be incorrect. will be easily recognized.
[0236] Example 10. Human CDRH3 dataset and ELD-3 for clinically relevant antibodies Matching Among the objectives of the present invention are the generation of human CDRH3 repertoires in vivo. mimicking the underlying VDJ recombination process and thereby the human characteristics of CDRH3 CDRH compared to other libraries known in the art while maintaining 3) increase the diversity of the library. One measure of success is the number of human The set of reference CDRH3 sequences may be similar or closely related in any library of the invention. The degree of difference is indicated by a difference (e.g., less than about 5, 4, 3, or 2 amino acids). The inventors have developed two human CDRH3 sequence reference data sets, both of which are non-redundant with each other. and HPS: (1) a collection of 666 human CDRH3 sequences (Lee et al., Immunogenetics, 2006, 57: 917; ”Lee-666 ”); and (2) Boyd et al., Science Translation nal Medicine, 2009, 1: 1-8 (“Boyd-3000”) 3,000 human sequences randomly selected from over 200,000 sequences disclosed in A collection of CDRH3 sequences was used to evaluate this metric. The results of a recent random sample of 3,000 human CDRH3 sequences are presented in Boyd et al. was applied to all members (>200,000 CDRH3 sequences) of the set of l. Results were representative of the same analysis.
[0237] Figure 7 shows the results of the analysis of the nucleotide sequences of Lee-666 with zero, one, two, three or more amino acid mismatches. or two synthetic libraries matching sequences from the Boyd-3000 set, LUA-141” and the percentage of CDRH3 sequences in ELD-3. Here, "LUA-141" is TN1 of 212, DH of 278, N2 of 141, and and 28 H3-JH containing libraries (see U.S. Patent Application Publication No. (See 2009 / 0181855). In particular, ELD-3 is more efficient than LUA-141. The CDRH3 sequences matched similarly to the reference CDRH3 sequence (Lee-666 and Boyd-667, respectively). 8.4% and 6.3% for the 3000 set) It is important to show (for the Lee-666 and Boyd-3000 sets, respectively) 12.9% and 12.1% for ELD-3 and LUA-141 (respectively Lee- compared with 41.2% and 43.7%) for the 666 and Boyd-3000 sets. Therefore, a higher degree of overlap of human CDRH3 sequences was found with at most two amino acid mismatches. It is also important to show the product percentages (Lee-666 and B, respectively). 54.1% and 52.5%) for the oyd-3000 set.
[0238] Other metrics by which antibody libraries can be evaluated include the ability to measure the performance of "clinically relevant" reference C Figure 8 shows that ELD-3 is a LUA-141 rRNA gene and its ability to match the DRH3 sequence. yielding better matches to clinically relevant CDRH3 sequences than the library In particular, ELD-3 demonstrates the ability to replicate 55 clinically validated antibodies within a single amino acid sequence. The LUA-141 library matches 34 of the 55 (62%). It only matches 0 (37%).
[0239] Example 11. Comparison of ELD-3 and LUA-141 ELD-3, like LUA-141, has 73 TN1, 92 DH, 119 N2, and 28 H3-JH. Therefore, ELD-3 (4.0 × 10 8 member) 94.5% of the sequences in the LUA-141 library (2.3 × 10 8 member) Figure 9 shows that the combinatorial efficiency of segments in ELD-3 is different from that of LUA- 141. In particular, the ELD-3 segment is There is a higher chance of obtaining a unique CDRH3 than the -141 library segment. However, fewer segments were used to synthesize libraries with increased CDRH3 diversity. This is advantageous because it allows
[0240] Figure 10 shows the K values of human CDRH3 sequences from LUA-141, ELD-3, and HPS. The amino acid composition of abat-CDRH3 is provided.
[0241] Figure 11 shows the K values of human CDRH3 sequences from LUA-141, ELD-3, and HPS. The length distribution of abat-CDRH3 is provided. CDRH3 library synthesized with degenerate oligonucleotides
[0242] Example 12. Further CDRH3 diversity by utilizing degenerate oligonucleotides increase The method described in this example allows for the production of more members than the libraries described above. Extending the methods taught above to generate a CDRH3 library with In particular, one or two degenerate codons are present in the DH and / or N2 polynucleotide sequences. Degenerate codons were introduced into the H3-JH segment, and (generally) degenerate codons were not introduced into the H3-JH segment. Segments with different numbers of degenerate codons, or one degenerate codon, were introduced. For example, a DH having 0, 1, 2, 3, 4, 5, 6, 7, 8, or more degenerate codons H3- with segments and 0, 1, 2, 3, 4, 5, or more degenerate codons JH segments are also contemplated, which are particularly relevant to the reference set of human CDRH3 sequences. Approximately 10 11 (about 2×10 11 ) This results in a CDRH3 library containing distinct CDRH3 amino acid sequences. As described in Considering the sequence, each is usually, but not always, very N- and / or C-terminal End positions or 5' and 3' terminal codons (i.e., not necessarily the first or last Degenerate codons were also used to synthesize the N2 segment. 200 of the TN1 segments were as described in ELD-3, but Ligands with degenerate TN1 segments or alternative choices of TN1 segment sequences A further 100 TN1 segments are included in this live A set of 300 TN1 segments for the library is completed. The nucleotide sequences are listed in Table 26. Alternatively, degenerate oligonucleotides may be used. In addition, the use of mixtures of trinucleotides can also be used to identify "bases" or " at one or more selected positions within the "seed" segment sequence (defined below). It is possible to allow for variations in amino acid type.
[0243] Example 13. Selection of DH segments for synthesis with degenerate oligonucleotides Segment usage weights are based on the sequence contained in Boyd et al. The theoretical segment pool of 68K DH was calculated by comparison. DH segments with acid lengths are used in conjunction with their segment usage weights (as described above). The top 201 were designated as "seed" sequences. The code sequence is then diversified by selecting certain positions for incorporating degenerate codons. The positions chosen for mutation, the amino acid types they diversify, are 68K A reference set of 9,171 DH segments, a subset of the DH theoretical segment pool, was created. The DH segments were determined by comparing the seed sequence with the DH segments. The weights of those segments used in Boyd et al. are significant, This meant that the cumulative segment usage weight (Example 8) was at least 1.0. So I chose it.
[0244] Each of the 201 seed sequences was compared to a reference set of 9,171 DH segments. The sequences are compared with each other, and those that are the same length but differ at one position are further identified as seeds. Thus, the most abundant mutants for each seed were characterized. Various positions are identified, and a set of candidate amino acid types is also identified for each position. Finally, the set of degenerate codons provides a set of candidate amino acid types for each specific position. The set of codons considered was to identify those that most faithfully represented the set of codons. N-linked glycosylation motifs (i.e., NXS, where X is any amino acid type) Degenerate codons encoding the nucleotide sequence (e.g., NXT) or the deamidation motif (NG) are This process generated 149 unique degenerate oligonucleotide sequences, In total, this encodes 3,566 unique polypeptide sequences. The resulting alternative designs also take into account greater diversity (unique polypeptides). (in terms of number of code sequences) and those with smaller RMAX values (see below) Priority was given for inclusion in the library of inventions. However, 68K DH Using different criteria to select DH segments from the theoretical segment pool and then create a library containing DH segments selected by these different criteria. are also contemplated to be within the scope of the present invention.
[0245] Not all degenerate oligonucleotides encode the same number of polypeptides Thus, the latter is a given theoretical segment contained within the CDRH3 library of the present invention. The same virus was generated uniformly across the pools (i.e., TN1, DH, N2, and H3-JH). For example, the total degeneracy is encoded by four oligonucleotides. Each amino acid sequence X has a "weight" of 1 / 4, but is a degenerate oligonucleotide of 6. The other individual amino acid sequence encoded by the octide, Y, has a weight of 1 / 6. Furthermore, certain amino acid sequences may be amplified by more than one degenerate oligonucleotide. and therefore their weights can be coded by Within a given theoretical segment pool, The polypeptide with the highest weighting against the polypeptide with the lowest weighting. The ratio of petite weight, RMAX, is an important design criterion that ideally should be minimized. The RMAX value may be calculated by length or globally for all of the segments of a given type. may be defined (i.e., all DH segments or all H3-JH segments) (e.g., for TN1 and / or N2 segments). Table 27 shows the degenerate oligonucleotides Table 28 lists the oligonucleotide sequences resulting from these oligonucleotides. These two tables list unique polypeptide sequences, the designs of which are detailed below. Contains a DH dimer segment.
[0246] Example 13.1. Selection of DH dimer segments A different method involves designing a set of degenerate oligonucleotides that encode the DH dimer sequence. The method was used to isolate Asn (N) residues and excessively hydrophobic dimers (i.e., F, I). Segments containing only L, M, and / or V residues) All 45 dimer sequences in ELD-3, minus ment, and 400 other theories The top possible dimer sequences (i.e., 20 possible residues at each of the two positions = 20* The goal of this design process was to generate 213 unique peptide dimers. This ultimately resulted in 35 degenerate oligonucleotide sequences encoding the target sequences. Similar to the selection process used for all of the other segments, one skilled in the art will Other criteria can be used to select body segments, and these segments can be included It will be readily apparent that libraries containing such sequences are also within the scope of the present invention.
[0247] By combining the longer DH segment of Example 13 with the DH dimer segment, A total of 184 oligonucleotides were used compared to the 200 oligonucleotides in ELD-3. (35 coding dimers and 149 coding segments with three or more amino acids) The final set of DH segments of the presently exemplified library was encoded by The 184 oligonucleotides yielded a total of 3,779 unique polypeptides. Sequences: 213 dimers and 3,566 longer segments of three or more amino acids Code.
[0248] Example 14. Generation of expanded N2 diversity As described above, ELD-3 contains 200 N2 segments. In the library shown, there is an empty N2 segment (i.e., there is no N2, and consequently, segment is directly linked to the H3-JH segment) and the monomeric N2 segment is However, the degenerate oligonucleotides were similar to ELD-3. Not only does it reproduce all of the corresponding sequences in the genome, but it also provides additional diversity. This was used to generate sets of 3 and 4 base lengths. These degenerate oligonucleotides contain Asn (at the wrong position) and Cys residues and More specifically, the Asn residue was designed to eliminate the codons at the following amino acid Whenever the first amino acid is not Gly and the next amino acid is not Ser or Thr, position and the first or second position of the tetramer, and therefore the candidate N2 segment The presently exemplified laminin sequences were designed to avoid deamidation or N-linked glycosylation motifs within the laminin. The N2 theoretical segment pool for the library contains a total of one zeromer (zeromer). o-mer) (i.e., no N2 segment), 18 monomers, 279 dimers, 33 Contains 9 trimer and 90 tetramer N2 amino acid sequences or 727 segments. These amino acid sequences were identified as 1, 18, 8, and 146 oligonucleotides in total. The first 19 are encoded by 1, 36, and 10 oligonucleotides, respectively. Except for oligonucleotides encoding zero and monomers, these are degenerate. The 146 oligonucleotide sequences are listed, while Table 30 lists the resulting 727 unique sequences. Suitable polypeptide sequences are listed below.
[0249] Example 15. Generation of expanded H3-JH diversity until only the DNA sequence corresponding to FRM4 remains (i.e., no H3-JH remains). , a nucleotide level sequence relative to the 5' end of the human IGHJ polynucleotide segment Stepwise deletion followed by systematic completion of 1 or 2 bp to the same 5' end The application resulted in 643 unique H3-JH peptide segments after translation (" 643 H3-JH set). Like the DH segment, Boyd et al. Their usage weights obtained after comparison to approximately 237,000 previous human sequences It is possible to order each of the 643 segments by The top 200 individual sequences from those free of unwanted motifs are currently exemplified. The H3-JH segments were chosen to provide a set of H3-JH segments for the library.
[0250] Instead, in the illustrated embodiment, 46 of the 200 H3-JH segments Overall, 200 oligonucleotides encode 246 unique peptide sequences. The peptides were designed with a two-fold degenerate codon at the first position to allow for a at the oligonucleotide level, N- or 5'-end, respectively).
[0251] In another alternative exemplary embodiment, the additional use of degenerate codons is 90, 100, 200 or more oligonucleotides representing distinct polypeptide sequences of It may be envisaged to generate libraries encoded by nucleotides. These up to 500 unique sequences may, but are not necessarily, the same as those described above. A subset of the sequences in the 43 H3-JH reference set or variants of these sequences As exemplified above, undesired polypeptide motifs can be The H3-JH segment containing the fragment may be omitted from the design. The oligonucleotide sequences for the The polypeptide sequences are provided in Table 32. In Table 31, the nucleotides corresponding to the FRM4 region are The peptide sequence is also provided, but the "peptide length" value refers to only the H3-JH portion. For simplicity, only the H3-JH peptide sequences are included in Table 32.
[0252] Example 16. Extended Diversity Library Design ( EDLD) The above selected TN1, DH, H3-JH, and The N2 and N3 segments are approximately 2 × 10 11 (300 TN1 x 3,779 DH x 727 N Extended Diversity with the theoretical diversity of 2×246 H3-JH The selected sequences were combined to generate the E-Library Design (EDLD). Oligonucleotides encoding the selected segments were selected according to the principles of Example 9.3.7. That's right.
[0253] Figures 12-15 illustrate certain features of this design. For example, Boyd et al. Approximately 50% of the approximately 237,000 CDRH3 sequences in This indicates that the sequence can be reproduced by library sequences that have the same or no mismatches ( That is, by summing the "0" and "1" bins in Figure 12. The theoretical length distribution (Figure 13) and amino acid composition (Figure 14) of the library from [the authors] are also closely matches each feature observed in the same set of human CDRH3 sequences Figure 15 shows the Extended Diversity Library Design Approximately 65% of the sequences appear only once in the design (Fig. (i.e., generated by one non-degenerate combination of segments). Figure 8 shown above is , in terms of matching to clinically relevant human antibody sequences, Extended Div ersity Library Design for both LUA-141 and ELD-3 It shows that one method is superior to the other. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 2] [Table 3-1] Table 3-2 Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 4-1 Table 4-2 Table 4-3 Table 5-1 Table 5-2 Table 5-3 Table 5-4 Table 5-5 Table 5-6 Table 5-7 Table 5-8 Table 5-9 Table 5-10 Table 5-11 Table 5-12 Table 5-13 Table 5-14 Table 5-15 Table 5-16 Table 5-17 Table 5-18 Table 5-19 Table 5-20 Table 5-21 Table 5-22 Table 5-23 Table 5-24 Table 5-25 Table 5-26 Table 5-27 Table 5-28 Table 5-29 Table 5-30 Table 5-31 Table 5-32 Table 5-33 Table 5-34 Table 5-35 Table 5-36 Table 5-37 Table 5-38 Table 5-39 Table 5-40 Table 6-1 Table 6-2 Table 6-3 Table 6-4 Table 6-5 Table 6-6 Table 6-7 Table 6-8 Table 6-9 Table 6-10 Table 6-11 Table 6-12 Table 6-13 Table 6-14 Table 6-15 Table 6-16 Table 6-17 Table 6-18 Table 6-19 Table 6-20 Table 6-21 Table 6-22 Table 6-23 Table 6-24 Table 6-25 Table 6-26 Table 6-27 Table 6-28 Table 6-29 Table 6-30 Table 6-31 Table 6-32 Table 6-33 Table 6-34 Table 6-35 Table 6-36 Table 6-37 Table 6-38 Table 6-39 Table 7-1 Table 7-2 Table 7-3 Table 7-4 Table 7-5 Table 7-6 Table 7-7 Table 7-8 Table 7-9 Table 7-10 Table 7-11 Table 7-12 Table 7-13 Table 7-14 Table 7-15 Table 7-16 Table 7-17 Table 7-18 Table 7-19 Table 7-20 Table 7-21 Table 7-22 Table 7-23 Table 7-24 Table 7-25 Table 7-26 Table 7-27 Table 8 Table 9-1 Table 9-2 Table 10-1 Table 10-2 Table 10-3 Table 10-4 Table 11-1 Table 11-2 Table 11-3 Table 11-4 Table 11-5 Table 11-6 Table 11-7 Table 11-8 Table 11-9 Table 11-10 Table 11-11 Table 11-12 Table 11-13 Table 11-14 Table 11-15 Table 11-16 Table 11-17 Table 11-18 Table 11-19 Table 12 Table 13-1 Table 13-2 Table 13-3 Table 13-4 Table 13-5 Table 13-6 Table 14 Table 15-1 Table 15-2 Table 15-3 Table 15-4 Table 15-5 Table 16 Table 17-1 Table 17-2 Table 18 Table 19 Table 20 Table 21 Table 22 Table 23-1 Table 23-2 Table 23-3 Table 23-4 Table 23-5 Table 23-6 Table 23-7 Table 24-1 Table 24-2 Table 24-3 Table 24-4 Table 24-5 Table 24-6 Table 24-7 Table 25-1 Table 25-2 Table 25-3 Table 25-4 Table 26-1 Table 26-2 Table 26-3 Table 26-4 Table 26-5 Table 26-6 Table 26-7 Table 26-8 Table 26-9 Table 26-10 Table 26-11 Table 27-1 Table 27-2 Table 27-3 Table 27-4 Table 27-5 Table 27-6 Table 27-7 Table 28-1 Table 28-2 Table 28-3 Table 28-4 Table 28-5 Table 28-6 Table 28-7 Table 28-8 Table 28-9 Table 28-10 Table 28-11 Table 28-12 Table 28-13 Table 28-14 Table 28-15 Table 28-16 Table 28-17 Table 28-18 Table 28-19 Table 28-20 Table 28-21 Table 28-22 Table 28-23 Table 28-24 Table 28-25 Table 28-26 Table 28-27 Table 28-28 Tables 28-29 Table 28-30 Table 28-31 Table 28-32 Table 28-33 Table 28-34 Table 28-35 Table 28-36 Table 28-37 Tables 28-38 Tables 28-39 Table 28-40 Table 28-41 Table 28-42 Table 28-43 Table 28-44 Table 28-45 Table 28-46 Table 28-47 Table 28-48 Table 28-49 Table 28-50 Table 28-51 Table 28-52 Table 28-53 Table 28-54 Table 28-55 Table 28-56 Table 28-57 Table 28-58 Table 28-59 Table 28-60 Table 28-61 Table 28-62 Table 28-63 Table 28-64 Table 28-65 Table 28-66 Table 28-67 Table 28-68 Table 28-69 Table 28-70 Table 28-71 Table 28-72 Table 28-73 Table 28-74 Table 28-75 Table 28-76 Table 28-77 Table 28-78 Table 28-79 Table 28-80 Table 28-81 Table 28-82 Table 28-83 Table 28-84 Table 28-85 Table 28-86 Table 28-87 Table 28-88 Table 28-89 Table 28-90 Table 28-91 Table 28-92 Table 28-93 Table 28-94 Table 28-95 Table 28-96 Table 28-97 Table 28-98 Table 28-99 Table 28-100 Table 28-101 Table 28-102 Table 28-103 Table 28-104 Table 28-105 Table 28-106 Table 28-107 Table 28-108 Table 28-109 Table 28-110 Table 28-111 Table 28-112 Table 28-113 Table 28-114 Table 28-115 Table 28-116 Table 28-117 Table 28-118 Table 28-119 Table 28-120 Table 28-121 Table 28-122 Table 28-123 Table 28-124 Table 28-125 Table 28-126 Table 28-127 Table 29-1 Table 29-2 Table 29-3 Table 29-4 Table 29-5 Table 30-1 Table 30-2 Table 30-3 Table 30-4 Table 30-5 Table 30-6 Table 30-7 Table 30-8 Table 30-9 Table 30-10 Table 30-11 Table 30-12 Table 30-13 Table 30-14 Table 30-15 Table 30-16 Table 30-17 Table 30-18 Table 30-19 Table 30-20 Table 30-21 Table 30-22 Table 30-23 Table 30-24 Table 30-25 Table 31-1 Table 31-2 Table 31-3 Table 31-4 Table 31-5 Table 31-6 Table 31-7 Table 32-1 Table 32-2 Table 32-3 Table 32-4 Table 32-5 Table 32-6 Table 32-7 Table 32-8 [Table 32-9]
[0254] equivalent Those skilled in the art will be able to easily implement the methods described herein using no more than routine experimentation. You will recognize or identify many equivalents to the specific embodiments and methods of the invention. Such equivalents are intended to be encompassed by the scope of the following claims. It is intended.
[0255] Appendix A GI numbers of 3,571 sequences in the healthy pre-immune set (HPS) [Table 33-1] [Table 33-2] [Table 33-3] [Table 33-4] [Table 33-5] [Table 33-6] [Table 33-7] [Table 33-8] [Table 33-9] Table 33-10 Table 33-11 Table 33-12 Table 33-13 Table 33-14 Table 33-15
Claims
1. 1. A method for generating a library of variant CDRL1 or CDRL2 or synthetic light chain variable domain polypeptides comprising CDRL1 and CDRL2 sequences, the method comprising: (a) providing a reference set of light chain sequences that have sequence diversity and length diversity similar to naturally occurring antibody sequences, including CDRL1, or CDRL2, or CDRL1 and CDRL2 sequences, before the naturally occurring antibody sequences have been subjected to negative selection and / or hypermutation; (b) Formula [Equation 10] calculating a diversity index for a position in said CDRL1, or CDRL2, or CDRL1 and CDRL2 sequence using (c) selecting positions for variation in the CDRL1, or CDRL2, or CDRL1 and CDRL2 sequences based on the diversity index of (b); (d) altering the amino acid residues at said positions in (c) to include the two to four most frequently occurring amino acid residues at corresponding positions in the germline of said reference set of light chain sequences; (e) synthesizing a polypeptide comprising the variant CDRL1 or CDRL2, or CDRL1 and CDRL2 sequences of (d); A method comprising:
2. The location selected in (c) is: at least one of positions 28, 29, 30B, and 30E of CDRL1 (Chothia-Lesk numbering scheme); and below: (a) CDRL1 positions 30, 30A, 31, and 32 (Chothia-Lesk numbering scheme); and (b) CDRL2 positions 50, 51, 53 and 55 (Chothia-Lesk numbering scheme) The method of claim 1 , comprising one or more of:
3. The positions selected for variation in the variant CDRL1 or CDRL2, or CDRL1 and CDRL2 sequences, are: (a) Position 28: germline S or G is changed to G, A, or D; (b) position 29: germline V is changed to I; (c) Position 30A: germline H is changed to Y; (d) position 30B: germline S is changed to R or T; and (e) position 30E: germline Y is changed to N; at least one of: If necessary, the following: (f) position 30: germline S is changed to N, D, G, T, A, or R; (g) Position 31: germline S is changed to D, R, I, N, or T; (h) position 32: germline Y or N is changed to F, S, or D; (i) Position 50: germline A, D, or G is changed to G, S, E, K, or D; (j) position 51: germline G or A is changed to A, S, or T; (k) position 53: germline S or N is changed to N, H, S, K, or R; or (l) position 55: germline E is changed to A or Q; The method of claim 1 , wherein the ion exchange rate is selected from one or more of:
4. The method of claim 1 , wherein the library comprises kappa or lambda light chain sequences.
5. 2. The method of claim 1, wherein the positions selected for variation in step (c) have a diversity index of (i) greater than 0.05 or (ii) greater than or equal to 0.07.
Citation Information
Patent Citations
Superhumanization of Antibodies by Blasting Putative Mature CDRs and Creating and Screening Cohort Libraries
JP2009515516A
Rationally designed, synthetic antibody libraries and uses therefor
WO2009036379A2