Variant nucleic acid library of GLP-1 receptor
By designing GLP1R antibodies and nucleic acid libraries with specific amino acid sequences, protein libraries are generated, and the problem of targeting GLP1R is solved, efficient GLP1R binding and regulation is achieved, and applied to the treatment of metabolic disorders.
Patent Information
- Application Number
- CN202080031649.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-14
- Filing Date
- 2020-02-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-02-26
AI Technical Summary
The prior art is difficult to effectively target the glucagon-like peptide-1 receptor (GLP1R) and prepare stable antibodies, resulting in difficulty in developing therapeutic intervention agents.
An antibody or antibody fragment thereof that binds GLP1R is provided, containing immunoglobulin heavy and light chains of specific amino acid sequences, and a nucleic acid library is designed to encode these antibodies, and a protein library is generated through a mottled nucleic acid library for targeting GLP1R.
High affinity and stable GLP1R binding is achieved, which can effectively inhibit or activate GLP1R, and is used to treat metabolic disorders such as type II diabetes and obesity.
Smart Images

Figure CN113766930B_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 810,377, filed Feb. 26, 2019; U.S. Provisional Patent Application No. 62 / 830,316, filed Apr. 5, 2019; U.S. Provisional Patent Application No. 62 / 855,836, filed May 31, 2019; U.S. Provisional Patent Application No. 62 / 904,563, filed Sep. 23, 2019; U.S. Provisional Patent Application No. 62 / 945,049, filed Dec. 6, 2019; and U.S. Provisional Patent Application No. 62 / 961,104, filed Jan. 14, 2020, each of which is incorporated herein by reference in its entirety.
[0003] Sequence Listing [0001.1] This application contains a Sequence Listing that has been submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy, created on Mar. 31, 2020, is named 44854-787_601_SL.txt and is 1,080,880 bytes in size. BACKGROUND OF THE INVENTION
[0004] G protein-coupled receptors (GPCRs) are associated with a variety of diseases. Because GPCRs are typically expressed at low levels in cells and are very unstable upon purification, there are problems in obtaining suitable antigens, and thus it is difficult to generate antibodies against GPCRs. Accordingly, there is a need for improved agents for therapeutic intervention that target GPCRs.
[0005] Incorporation by Reference
[0006] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference to the extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. SUMMARY OF THE INVENTION
[0007] The present disclosure provides antibodies or antibody fragments that bind to GLP1R, which comprise an immunoglobulin heavy chain and an immunoglobulin light chain: (a) wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320 or 2321; and (b) wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315 or 2316. The present disclosure further provides antibodies or antibody fragments that bind to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2303; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2310. The present disclosure further provides antibodies or antibody fragments that bind to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2304; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2311. The present disclosure further provides antibodies or antibody fragments that bind to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2305; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2312. The present disclosure further provides antibodies or antibody fragments that bind to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2306; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2313.The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2307; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2314. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2308; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2315. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequences shown in SEQ ID NO: 2309, 2317, 2318, 2319; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2316. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody is a monoclonal antibody, polyclonal antibody, bispecific antibody, multispecific antibody, transplantation antibody, human antibody, humanized antibody, synthetic antibody, chimeric antibody, camelized antibody, single-chain Fv (scFv), single-chain antibody, Fab fragment, F(ab')2 fragment, Fd fragment, Fv fragment, single-domain antibody, isolated complementarity-determining region (CDR), diabody, fragment consisting of only a single monomer variable domain, disulfide-linked Fv (sdFv), intracellular antibody, anti-idiotype (anti-Id) antibody, or an antigen-binding fragment thereof. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody or the antibody fragment is chimeric or humanized. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the EC50 of the antibody in the cAMP assay is less than about 25 nanomolar. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the EC50 of the antibody in the cAMP assay is less than about 20 nanomolar. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the EC50 of the antibody in the cAMP assay is less than about 10 nanomolar. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody is an agonist of GLP1R.The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody is an antagonist of GLP1R. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody is an allosteric modulator of GLP1R. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the allosteric modulator of GLP1R is a negative allosteric modulator. The present invention further provides an antibody or an antibody fragment that binds to GLP1R, wherein the antibody or antibody fragment comprises CDR-H3, and the CDR-H3 comprises the sequence of any one of SEQ ID NO: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294, 2295, 2296, 2297, 2298, 2299, 2300, 2301 or 2302.
[0008] The present disclosure provides a nucleic acid library comprising multiple nucleic acids, wherein each nucleic acid encodes a sequence that encodes an immunoglobulin scaffold when translated, wherein the immunoglobulin scaffold comprises a CDR-H3 loop that comprises a GLP1R binding domain, and wherein each nucleic acid comprises a sequence encoding a sequence variant of the GLP1R binding domain. The present disclosure further provides a nucleic acid library, wherein the CDR-H3 loop has a length of about 20 to about 80 amino acids. The present disclosure further provides a nucleic acid library, wherein the CDR-H3 loop has a length of about 80 to about 230 base pairs. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin scaffold further comprises one or more domains selected from a light chain variable domain (VL), a heavy chain variable domain (VH), a light chain constant domain (CL), and a heavy chain constant domain (CH). The present disclosure further provides a nucleic acid library, wherein the VH domain is IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV3-74, IGHV4-39, or IGHV4-59 / 61. The present disclosure further provides a nucleic acid library, wherein the VH domain is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. The present disclosure further provides a nucleic acid library, wherein the VH domain is IGHV1-69 and IGHV3-30. The present disclosure further provides a nucleic acid library, wherein the VL domain is IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, or IGLV3-1. The present disclosure further provides a nucleic acid library, wherein the VH domain has a length of about 90 to about 100 amino acids. The present disclosure further provides a nucleic acid library, wherein the VL domain has a length of about 90 to about 120 amino acids. The present disclosure further provides a nucleic acid library, wherein the VH domain has a length of about 280 to about 300 base pairs. The present disclosure further provides a nucleic acid library, wherein the VL domain has a length of about 300 to about 350 base pairs. The present disclosure further provides a nucleic acid library, wherein the library comprises at least 10 5 different nucleic acids. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin scaffold comprises a single immunoglobulin domain. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin scaffold comprises a peptide of at most 100 amino acids.
[0009] The present disclosure provides a protein library comprising a plurality of proteins, wherein each protein of the plurality of proteins comprises an immunoglobulin scaffold, and wherein the immunoglobulin scaffold comprises a CDR-H3 loop that comprises a sequence variant of a GLP1R-binding domain. The present disclosure further provides a protein library, wherein the CDR-H3 loop has a length of from about 20 to about 80 amino acids. The present disclosure further provides a protein library, wherein the immunoglobulin scaffold further comprises one or more domains selected from the group consisting of a light chain variable domain (VL), a heavy chain variable domain (VH), a light chain constant domain (CL), and a heavy chain constant domain (CH). The present disclosure further provides a protein library, wherein the VH domain is IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV3-74, IGHV4-39, or IGHV4-59 / 61. The present disclosure further provides a protein library, wherein the VH domain is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. The present disclosure further provides a protein library, wherein the VH domain is IGHV1-69 and IGHV3-30. The present disclosure further provides a protein library, wherein the VL domain is IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, or IGLV3-1. The present disclosure further provides a protein library, wherein the VH domain has a length of from about 90 to about 100 amino acids. The present disclosure further provides a protein library, wherein the VL domain has a length of from about 90 to about 120 amino acids. The present disclosure further provides a protein library, wherein the plurality of proteins is used to generate a peptidomimetic library. The present disclosure further provides a protein library, wherein the protein library comprises antibodies.
[0010] The present disclosure provides a protein library comprising a plurality of proteins, wherein the plurality of proteins comprises sequences encoding different GPCR-binding domains, and wherein each GPCR-binding domain has a length of from about 20 to about 80 amino acids. The present disclosure further provides a protein library, wherein the protein library comprises peptides. The present disclosure further provides a protein library, wherein the protein library comprises immunoglobulins. The present disclosure further provides a protein library, wherein the protein library comprises antibodies. The present disclosure further provides a protein library, wherein the plurality of proteins is used to generate a peptidomimetic library.
[0011] The present disclosure provides a library of vectors comprising a nucleic acid library as described herein.
[0012] The present disclosure provides a cell library comprising a nucleic acid library as described herein.
[0013] The present disclosure provides a cell library comprising a protein library as described herein.
[0014] The present disclosure provides an antibody, wherein the antibody comprises a CDR-H3 that comprises the sequence of any one of SEQ ID NO: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294, 2295, 2296, 2297, 2298, 2299, 2300, 2301, or 2302.
[0015] The present disclosure provides an antibody, wherein the antibody comprises a CDR-H3 that comprises the sequence of any one of SEQ ID NO: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294, 2295, 2296, 2297, 2298, 2299, 2300, 2301, or 2302; and wherein the antibody is a monoclonal antibody, polyclonal antibody, bispecific antibody, multispecific antibody, transplantation antibody, human antibody, humanized antibody, synthetic antibody, chimeric antibody, camelized antibody, single-chain Fv (scFv), single-chain antibody, Fab fragment, F(ab')2 fragment, Fd fragment, Fv fragment, single-domain antibody, isolated complementarity-determining region (CDR), diabody, fragment consisting of only a single monomeric variable domain, disulfide-linked Fv (sdFv), intracellular antibody, anti-idiotypic (anti-Id) antibody, or an antigen-binding fragment thereof.
[0016] The present disclosure provides methods of inhibiting GLP1R activity, which comprise administering an antibody or antibody fragment as described herein. The present disclosure further provides methods of inhibiting GLP1R activity, wherein the antibody or antibody fragment is an allosteric modulator. The present disclosure further provides methods of inhibiting GLP1R activity, wherein the antibody or antibody fragment is a negative allosteric modulator. The present disclosure further provides methods of treating a metabolic disorder, which comprise administering an antibody or antibody fragment as described herein to a subject in need thereof. The present disclosure further provides methods of treating a metabolic disorder, wherein the metabolic disorder is type II diabetes or obesity.
[0017] The present disclosure provides nucleic acid libraries comprising: a plurality of nucleic acids, each of which encodes a sequence that encodes a GLP1R-binding immunoglobulin when translated, wherein the GLP1R-binding immunoglobulin comprises a variant of a GLP1R-binding domain, wherein the GLP1R-binding domain is a ligand of GLP1R, and wherein the nucleic acid library comprises at least 10,000 variant immunoglobulin heavy chains and at least 10,000 variant immunoglobulin light chains. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 50,000 variant immunoglobulin heavy chains and at least 50,000 variant immunoglobulin light chains. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 100,000 variant immunoglobulin heavy chains and at least 100,000 variant immunoglobulin light chains. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 10 5 distinct nucleic acids. The present disclosure further provides nucleic acid libraries, wherein the immunoglobulin heavy chain has a length of from about 90 to about 100 amino acids when translated. The present disclosure further provides nucleic acid libraries, wherein the immunoglobulin heavy chain has a length of from about 100 to about 400 amino acids when translated. The present disclosure further provides nucleic acid libraries, wherein the variant immunoglobulin heavy chain has at least 80% sequence identity to SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320, or 2321 when translated. The present disclosure further provides nucleic acid libraries, wherein the variant immunoglobulin light chain has at least 80% sequence identity to SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315, or 2316 when translated.
[0018] The present disclosure provides nucleic acid libraries comprising: a plurality of nucleic acids, each of which encodes a sequence that encodes a GLP1R single-domain antibody when translated, wherein each of the plurality of sequences comprises a variant sequence encoding at least one of CDR1, CDR2, and CDR3 on the heavy chain; wherein the library comprises at least 30,000 variant sequences; and wherein the antibody or antibody fragment binds to its antigen with a K D less than 100 nM. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 50,000 variant immunoglobulin heavy chains and at least 50,000 variant immunoglobulin light chains. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 100,000 variant immunoglobulin heavy chains and at least 100,000 variant immunoglobulin light chains. The present disclosure further provides nucleic acid libraries, wherein the nucleic acid library comprises at least 10 5A plurality of distinct nucleic acids. The present invention further provides a nucleic acid library, wherein the length of the immunoglobulin heavy chain when translated is about 90 to about 100 amino acids. The present invention further provides a nucleic acid library, wherein the length of the immunoglobulin heavy chain when translated is about 100 to about 400 amino acids. The present invention further provides a nucleic acid library, wherein the variant immunoglobulin heavy chain has at least 80% sequence identity with SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320 or 2321 when translated. The present invention further provides a nucleic acid library, wherein the variant immunoglobulin light chain has at least 80% sequence identity with SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315 or 2316 when translated.
[0019] The present invention provides an antagonist of GLP1R, which comprises SEQ ID NO: 2279 or 2320. The present invention further provides an antagonist, wherein the EC50 of the antagonist does not exceed 1.5 nM. The present invention further provides an antagonist, wherein the EC50 of the antagonist does not exceed 1.0 nM. The present invention further provides an antagonist, wherein the EC50 of the antagonist does not exceed 0.5 nM. The present invention further provides an antagonist, wherein the antagonist is an antibody or an antibody fragment thereof.
[0020] The present invention provides a nucleic acid library, which comprises: a plurality of nucleic acids, wherein each nucleic acid encodes a sequence encoding a GLP1R-binding immunoglobulin when translated, wherein the GLP1R-binding immunoglobulin comprises a variant of the GLP1R-binding domain, wherein the GLP1R-binding domain is a ligand of GLP1R, and wherein the nucleic acid library comprises at least 10,000 variant immunoglobulin heavy chains and at least 10,000 variant immunoglobulin light chains. The present invention further provides a nucleic acid library, wherein the nucleic acid library comprises at least 50,000 variant immunoglobulin heavy chains and at least 50,000 variant immunoglobulin light chains. The present invention further provides a nucleic acid library, wherein the nucleic acid library comprises at least 100,000 variant immunoglobulin heavy chains and at least 100,000 variant immunoglobulin light chains. The present invention further provides a nucleic acid library, wherein the nucleic acid library comprises at least 10 5A plurality of distinct nucleic acids. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin heavy chain has a length of about 90 to about 100 amino acids when translated. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin heavy chain has a length of about 100 to about 400 amino acids when translated. The present disclosure further provides a nucleic acid library, wherein the variant immunoglobulin heavy chain has at least 90% sequence identity with SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320 or 2321 when translated. The present disclosure further provides a nucleic acid library, wherein the variant immunoglobulin light chain has at least 90% sequence identity with SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315 or 2316 when translated.
[0021] The present disclosure provides a nucleic acid library comprising: a plurality of nucleic acids, wherein each nucleic acid encodes a sequence that encodes a GLP1R single domain antibody when translated, wherein each sequence in the plurality of sequences comprises a variant sequence encoding at least one of CDR1, CDR2, and CDR3 on the heavy chain; wherein the library comprises at least 30,000 variant sequences; and wherein the antibody or antibody fragment binds to its antigen with a K D less than 100 nM. The present disclosure further provides a nucleic acid library, wherein the nucleic acid library comprises at least 50,000 variant immunoglobulin heavy chains and at least 50,000 variant immunoglobulin light chains. The present disclosure further provides a nucleic acid library, wherein the nucleic acid library comprises at least 100,000 variant immunoglobulin heavy chains and at least 100,000 variant immunoglobulin light chains. The present disclosure further provides a nucleic acid library, wherein the nucleic acid library comprises at least 10 5 A plurality of distinct nucleic acids. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin heavy chain has a length of about 90 to about 100 amino acids when translated. The present disclosure further provides a nucleic acid library, wherein the immunoglobulin heavy chain has a length of about 100 to about 400 amino acids when translated. The present disclosure further provides a nucleic acid library, wherein the variant immunoglobulin heavy chain has at least 90% sequence identity with SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320 or 2321 when translated. The present disclosure further provides a nucleic acid library, wherein the variant immunoglobulin light chain has at least 90% sequence identity with SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315 or 2316 when translated.
[0022] The present disclosure provides antibodies or antibody fragments that bind to GLP1R and comprise an immunoglobulin heavy chain and an immunoglobulin light chain: (a) wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320, or 2321; and (b) wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315, or 2316. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2303; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2310. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2304; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2311. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2305; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2312. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2306; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2313. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2307; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2314. The present disclosure further provides antibodies or antibody fragments wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2308; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence set forth in SEQ ID NO: 2315.The present disclosure further provides an antibody or an antibody fragment, wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence shown in SEQ ID NO: 2309; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90% identical to the amino acid sequence shown in SEQ ID NO: 2316. The present disclosure further provides an antibody or an antibody fragment thereof, wherein the antibody is a monoclonal antibody, a polyclonal antibody, a bispecific antibody, a multispecific antibody, a transplantation antibody, a human antibody, a humanized antibody, a synthetic antibody, a chimeric antibody, a camelized antibody, a single-chain Fv (scFv), a single-chain antibody, a Fab fragment, an F(ab')2 fragment, an Fd fragment, an Fv fragment, a single-domain antibody, an isolated complementarity-determining region (CDR), a diabody, a fragment consisting of only a single monomer variable domain, a disulfide-linked Fv (sdFv), an intrabody, an anti-idiotypic (anti-Id) antibody, or an antigen-binding fragment thereof. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody or the antibody fragment thereof is chimeric or humanized. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody has an EC50 in a cAMP assay of less than about 25 nanomolar. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody has an EC50 in a cAMP assay of less than about 20 nanomolar. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody has an EC50 in a cAMP assay of less than about 10 nanomolar. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody is an agonist of GLP1R. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody is an antagonist of GLP1R. The present disclosure further provides an antibody or an antibody fragment, wherein the antibody is an allosteric modulator of GLP1R. The present disclosure further provides an antibody or an antibody fragment, wherein the allosteric modulator of GLP1R is a negative allosteric modulator.
[0023] The present disclosure provides an antibody or an antibody fragment, wherein the antibody or the antibody fragment comprises the sequence of any one of SEQ ID NO: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294, 2295, 2296, 2297, 2298, 2299, 2300, 2301 or 2302 or the sequence shown in Table 27.
[0024] The present invention provides an antibody or antibody fragment, wherein the antibody or antibody fragment comprises the sequence of any one of SEQ ID NO: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294, 2295, 2296, 2297, 2298, 2299, 2300, 2301 or 2302, or the sequence shown in Table 27; and wherein the antibody is a monoclonal antibody, polyclonal antibody, bispecific antibody, multispecific antibody, transplantation antibody, human antibody, humanized antibody, synthetic antibody, chimeric antibody, camelized antibody, single-chain Fv (scFv), single-chain antibody, Fab fragment, F(ab')2 fragment, Fd fragment, Fv fragment, single-domain antibody, isolated complementarity-determining region (CDR), diabody, fragment consisting of only a single monomer variable domain, disulfide-linked Fv (sdFv), intracellular antibody, anti-idiotype (anti-Id) antibody, or an antigen-binding fragment thereof.
[0025] The present invention provides an antagonist of GLP1R, which comprises SEQ ID NO: 2279 or 2320. The present invention further provides an antagonist of GLP1R, wherein the EC50 of the antagonist does not exceed 1.5 nM. The present invention further provides an antagonist of GLP1R, wherein the EC50 of the antagonist does not exceed 1.0 nM. The present invention further provides an antagonist of GLP1R, wherein the EC50 of the antagonist does not exceed 0.5 nM. The present invention further provides an antagonist of GLP1R, wherein the antagonist is an antibody or antibody fragment.
[0026] The present invention provides an agonist of GLP1R, which comprises SEQ ID NO: 2317. The present invention further provides an agonist of GLP1R, wherein the EC50 of the agonist does not exceed 1.5 nM. The present invention further provides an agonist of GLP1R, wherein the EC50 of the agonist does not exceed 1.0 nM. The present invention further provides an agonist of GLP1R, wherein the EC50 of the agonist does not exceed 0.5 nM. The present invention further provides an agonist of GLP1R, wherein the agonist is an antibody or antibody fragment.
[0027] The present invention provides a method for inhibiting GLP1R activity, which comprises administering an antibody or antibody fragment as described herein. The present invention further provides a method for inhibiting GLP1R activity, wherein the antibody or antibody fragment is an allosteric modulator. The present invention further provides a method for inhibiting GLP1R activity, wherein the antibody or antibody fragment is a negative allosteric modulator.
[0028] The present invention provides methods for treating metabolic disorders, which include administering to a subject in need an antibody as described herein. The present invention provides methods for treating metabolic disorders, wherein the metabolic disorder is type II diabetes or obesity.
[0029] The present invention provides a protein library encoded by a nucleic acid library as described herein, wherein the protein library contains peptides. The present invention further provides a protein library, wherein the protein library contains immunoglobulins. The present invention further provides a protein library, wherein the protein library contains antibodies. The present invention further provides a protein library, wherein the protein library is a peptidomimetic library.
[0030] The present invention provides a library of vectors comprising a nucleic acid library as described herein. The present invention provides a library of cells comprising a nucleic acid library as described herein. The present invention provides a library of cells comprising a protein library as described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1A A first schematic diagram depicting an immunoglobulin scaffold is provided.
[0032] Figure 1B A second schematic diagram depicting an immunoglobulin scaffold is provided.
[0033] Figure 2 A schematic diagram depicting a motif for placement within a scaffold is provided.
[0034] Figure 3 A step diagram presenting an exemplary processing workflow illustrating gene synthesis as disclosed herein is provided.
[0035] Figure 4 An example of a computer system is shown.
[0036] Figure 5 A block diagram showing the architecture of a computer system is provided.
[0037] Figure 6 A diagram illustrating a network configured to incorporate multiple computer systems, multiple cellular telephones and personal data assistants, and network attached storage (NAS) is provided.
[0038] Figure 7 A block diagram of a multiprocessor computer system using a shared virtual address storage space is provided.
[0039] Figure 8A A schematic diagram depicting an immunoglobulin scaffold comprising a VH domain linked to a VL domain using a linker is provided.
[0040] Figure 8BDepicts a schematic diagram of the full - structure of an immunoglobulin scaffold, which contains a VH domain, a leader sequence, and a pIII sequence linked to the VL domain using a linker.
[0041] Figure 8C Depicts a schematic diagram of the four framework elements (FW1, FW2, FW3, FW4) and three variable CDR (L1, L2, L3) elements of the VL or VH domain.
[0042] Figures 9A - 9O Depicts GLP1R - 2( Figure 9A )、GLP1R - 3( Figure 9B )、GLP1R - 8( Figure 9C )、GLP1R - 26( Figure 9D )、GLP1R - 30( Figure 9E )、GLP1R - 56( Figure 9F )、GLP1R - 58( Figure 9G )、GLP1R - 10( Figure 9H )、GLP1R - 25( Figure 9I )、GLP1R - 60( Figure 9J )、GLP1R - 70( Figure 9K )、GLP1R - 72( Figure 9L )、GLP1R - 83( Figure 9M )、GLP1R - 93( Figure 9N ) and GLP1R - 98( Figure 9O ) cell - binding data.
[0043] Figures 10A - 10O Depicts GLP1R - 2( Figure 10A )、GLP1R - 3( Figure 10B )、GLP1R - 8( Figure 10C )、GLP1R - 26( Figure 10D )、GLP1R - 30( Figure 10E )、GLP1R - 56( Figure 10F )、GLP1R - 58( Figure 10G )、GLP1R - 10( Figure 10H )、GLP1R - 25( Figure 10I )、GLP1R - 60( Figure 10J )、GLP1R - 70( Figure 10K )、GLP1R - 72( Figure 10L )、GLP1R - 83( Figure 10M )、GLP1R - 93( Figure 10N ) and GLP1R - 98( Figure 10O ) variant's inhibition of GLP1 - 7 - 36 peptide - induced cAMP activity diagram.
[0044] Figures 11A - 11G Depicts the cellular function data of GLP1R-2( Figure 11A ), GLP1R-3( Figure 11B ), GLP1R-8( Figure 11C ), GLP1R-26( Figure 11D ), GLP1R-30( Figure 11E ), GLP1R-56( Figure 11F ) and GLP1R-58( Figure 11G ).
[0045] Figures 12A - 12G Depicts the illustration of the inhibition of exendin-4 peptide-induced cAMP activity by GLP1R-2( Figure 12A ), GLP1R-3( Figure 12B ), GLP1R-8( Figure 12C ), GLP1R-26( Figure 12D ), GLP1R-30( Figure 12E ), GLP1R-56( Figure 12F ) and GLP1R-58( Figure 12G ) variants.
[0046] Figure 13 Depicts a schematic diagram of glucagon (SEQ ID NO:2740), GLP1-1 (SEQ ID NO:6) and GLP-2 (SEQ ID NO:2741).
[0047] Figures 14A - 14C Depicts the cellular binding affinity of purified immunoglobulin.
[0048] Figure 14D Depicts the cAMP activity of purified immunoglobulin.
[0049] Figures 15A - 15H Depicts the binding curves plotted for GLP1R-238( Figure 15A ), GLP1R-240( Figure 15B ), GLP1R-241( Figure 15C ), GLP1R-242( Figure 15D ), GLP1R-243( Figure 15E ), GLP1R-244( Figure 15F ), pGPCR-GLP1R-43( Figure 15G ) and pGPCR-GLP1R-44( Figure 15H ), with IgG concentration in nanomolar (nM) plotted against MFI (mean fluorescence intensity).
[0050] Figures 16A - 16IDepicts the flow cytometry data of the binding assay presented in dot plot format for 100 nM IgG of GLP1R-238( Figure 16A ), GLP1R-240( Figure 16B ), GLP1R-241( Figure 16C ), GLP1R-242( Figure 16D ), GLP1R-243( Figure 16E ), GLP1R-244( Figure 16F ), pGPCR-GLP1R-43( Figure 16G ), pGPCR-GLP1R-44( Figure 16H ) and GLP1R-239( Figure 16I ).
[0051] Figures 17A - 17B Depicts the data from the cAMP assay, with the y-axis being relative light units (RLU) and the x-axis being the concentration in nanomoles (nM). cAMP was measured in response to GLP1(7-36), GLP1R-238, GLP1R-239, GLP1R-240, GLP1R-241, GLP1R-242, GLP1R-243, GLP1R-244, pGPCR-GLP1R-43, pGPCR-GLP1R-44 and buffer.
[0052] Figure 17C Depicts the allosteric effect plot of GLP1R-241.
[0053] Figure 17D Depicts the β-arrestin recruitment plot of GLP1R-241.
[0054] Figure 17E Depicts the internalization plot of GLP1R-241.
[0055] Figures 18A - 18B Depicts the data from the cAMP assay, with the y-axis being relative light units (RLU) and the x-axis being the concentration of GLP1(7-36) in nanomoles (nM). The allosteric effects of GLP1R-238, GLP1R-239, GLP1R-240, GLP1R-241, GLP1R-242, GLP1R-243, GLP1R-244, pGPCR-GLP1R-43, pGPCR-GLP1R-44 and no antibody were tested.
[0056] Figures 19A - 19F Depicts for GLP1R-59-2( Figure 19A ), GLP1R-59-241( Figure 19B ), GLP1R-59-243( Figure 19C)、GLP1R-3( Figure 19D )、GLP1R-241( Figure 19E ) and GLP1R-2( Figure 19F ), Flow cytometry data of binding assays presented as dot plots and histograms. Figures 19A - 19F Also depicted is the GLP1R-59-2 ( Figure 19A )、GLP1R-59-241( Figure 19B )、GLP1R-59-243( Figure 19C )、GLP1R-3( Figure 19D )、GLP1R-241( Figure 19E ) and GLP1R-2( Figure 19F ), titration curves plotted with IgG concentration in nanomolar (nM) versus MFI (mean fluorescence intensity).
[0057] Figures 20A - 20F Depicted is the GLP1R-59-2 ( Figure 20A )、GLP1R-59-241( Figure 20B )、GLP1R-59-243( Figure 20C )、GLP1R-3( Figure 20D )、GLP1R-241( Figure 20E ) and GLP1R-2( Figure 20F ), data from a cAMP assay, with relative luminescence units (RLU) on the y-axis and GLP1(7-36) concentration in nanomolar (nM) on the x-axis, and β-arrestin recruitment and receptor internalization.
[0058] Figures 21A - 21B The TIGIT affinity distribution of the VHH library is depicted, depicting affinity thresholds from 20 to 4000 ( Figure 21A ) or an affinity threshold between 20 and 1000 ( Figure 21B ). Among the 140 VHH binders, 51 variants were <100 nM and 90 variants were <200 nM.
[0059] Figures 22A - 22B Depicted is the FAC analysis of GLP1R-43-77 ( Figure 22A ) and dose curves and specificity plots ( Figure 22B ).
[0060] Figure 23A A schematic diagram of the heavy chain IGHV3-23 design is depicted. Figure 23A SEQ ID NOs 2742-2747 are disclosed respectively in the order of appearance.
[0061] Figure 23B A schematic diagram of the heavy chain IGHV1-69 design is depicted.Figure 23B SEQ ID NOs 2748 - 2753 are disclosed in the order of their appearance.
[0062] Figure 23C A schematic diagram depicting the design of light chains IGKV 2 - 28 and IGLV 1 - 51. Figure 23C SEQ ID NOs 2754 - 2759 are disclosed in the order of their appearance.
[0063] Figure 23D A schematic diagram depicting the theoretical and final diversity of the GLP1R library.
[0064] Figures 23E - 23F A FACS binding plot of GLP1R IgG.
[0065] Figures 23G - 23H A diagram depicting the cAMP assay using purified GLP1R IgG.
[0066] Figure 24A A GLP1R - 3 inhibition plot compared to no antibody. The y - axis depicts relative light units (RLU), and the x - axis depicts the concentration of GLP1(7 - 36) in nanomoles (nM).
[0067] Figure 24B A diagram depicting GLP1R - 3 inhibition at high concentrations after stimulation with 0.05 nM GLP1(7 - 36). The y - axis depicts relative light units (RLU), and the x - axis depicts the concentration of GLP1R - 3 in nanomoles (nM).
[0068] Figure 24C A plot depicting glucose levels after administration of glucose in a mouse model of diet - induced obesity when treated with vehicle (triangle), liraglutide (square), and GLP1R - 3 (circle).
[0069] Figure 24D A plot depicting glucose levels after administration of glucose in a mouse model of diet - induced obesity when treated with vehicle (open triangle), liraglutide (square), and GLP1R - 59 - 2 (solid triangle).
[0070] Figure 25A A plot depicting the blood glucose levels (mg / dL; y - axis) of mice treated with GLP1R - 59 - 2 (agonist), GLP1R - 3 (antagonist), and control over time (in minutes, x - axis).
[0071] Figure 25BGraph showing blood glucose levels (mg / dL; y-axis) of mice treated with GLP1R-59-2 (agonist), GLP1R-3 (antagonist), and control.
[0072] Figure 25C Graph showing blood glucose levels (mg / dL; y-axis) of mice treated with GLP1R-59-2 (agonist) compared to control in fasted (p = 0.0008) and non-fasted (p < 0.0001) mice.
[0073] Figure 25D Graph showing blood glucose levels (mg / dL / min; y-axis) of pre-administered GLP1R-59-2 (agonist), GLP1R-3 (antagonist), and control mice. Detailed Description
[0074] Unless otherwise indicated, the present disclosure employs conventional molecular biology techniques within the skill of the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0075] Definitions
[0076] Throughout this disclosure, various embodiments are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of any embodiment. Thus, unless the context clearly dictates otherwise, the description of a range should be considered to have expressly disclosed all possible sub-ranges as well as the individual values within the range to the tenth of the lower limit unit. For example, a description of a range such as from 1 to 6 should be considered to have expressly disclosed sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and the individual values within the range, e.g., 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the width of the range. The upper and lower limits of these intermediate ranges may be independently included in a smaller range and are also covered by the present disclosure, subject to any specific excluded limits of the recited range. Unless the context clearly dictates otherwise, ranges excluding either or both of these included limits are also included in the present disclosure when the recited range includes one or both of the limits.
[0077] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiments. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein are also intended to include the plural forms. Further, it should be understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0078] Unless specifically stated or apparent from the context, as used herein, the term "about" with respect to a number or numerical range shall be understood to mean the stated number and numbers within + / −10% thereof, or for values listed for a range, from 10% below the listed lower limit to 10% above the listed upper limit.
[0079] Unless otherwise specified, as used herein, the term "nucleic acid" encompasses double-stranded or triple-stranded nucleic acids as well as single-stranded molecules. In double-stranded or triple-stranded nucleic acids, the nucleic acid strands need not co-extend (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). When provided, nucleic acid sequences are listed in the 5' to 3' direction, unless otherwise indicated. The methods described herein provide for the generation of isolated nucleic acids. The methods described herein additionally provide for the generation of isolated and purified nucleic acids. The "nucleic acids" referred to herein can be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more bases in length. Moreover, the methods provided herein provide for the synthesis of nucleotide sequences encoding any number of polypeptide segments, including sequences encoding non-ribosomal peptides (NRPs), sequences encoding non-ribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments from other protein families, including non-coding DNA or RNA such as regulatory sequences, e.g., promoters, transcription factors, enhancers, siRNA, shRNA, RNAi, miRNA, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, intergenic DNA, loci (loci) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), small nucleolar RNA, ribozymes, complementary DNA (cDNA) (which is the DNA representation of mRNA, typically obtained by reverse transcription of messenger RNA (mRNA) or by amplification); DNA molecules generated synthetically or by amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. The cDNA encoding the genes or gene fragments referred to herein can contain at least one region encoding an exon sequence without the intervening intron sequences present in the genomic equivalent sequence.
[0080] GPCR library of GLP1 receptor
[0081] The present disclosure provides methods and compositions related to a G-protein coupled receptor (GPCR) binding library of glucagon-like peptide-1 receptor (GLP1R), the library comprising nucleic acids encoding scaffolds containing a GPCR binding domain. The scaffolds as described herein can stably support the GPCR binding domain. The GPCR binding domain can be designed based on the surface interaction of a GLP1R ligand with GLP1R. The library as described herein can be further variegated to provide a variant library of nucleic acids comprising respective variants encoding at least one predetermined reference nucleic acid sequence. Protein libraries that can be generated upon translation of the nucleic acid library are further described herein. In some instances, the nucleic acid library as described herein is transferred into cells to generate a cell library. Downstream applications of libraries synthesized using the methods described herein are also provided. Downstream applications include identification of variant nucleic acid or protein sequences having enhanced biologically relevant functions (e.g., improved stability, affinity, binding, functional activity) and for treating or preventing disease states associated with GPCR signaling.
[0082] Scaffold library
[0083] The present disclosure provides a library comprising nucleic acids encoding a scaffold, wherein a sequence encoding a GPCR binding domain is positioned within the scaffold. The scaffolds as described herein allow for a range of GPCR binding domain encoding sequences to have improved stability when inserted into the scaffold as compared to when inserted into an unmodified scaffold. Exemplary scaffolds include, but are not limited to, proteins, peptides, immunoglobulins, derivatives thereof, or combinations thereof. In some instances, the scaffold is an immunoglobulin. The scaffolds as described herein have improved functional activity, structural stability, expression, specificity, or combinations thereof. In some instances, the scaffold comprises a long region for supporting the GPCR binding domain.
[0084] The present disclosure provides libraries of nucleic acids encoding scaffolds, where the scaffold is an immunoglobulin. In some cases, the immunoglobulin is an antibody. As used herein, the term antibody will be understood to include proteins having the characteristic two-armed, Y-shaped structure of a typical antibody molecule, as well as one or more fragments of an antibody that retain the ability to specifically bind an antigen. Exemplary antibodies include, but are not limited to, monoclonal antibodies, polyclonal antibodies, bispecific antibodies, multispecific antibodies, transplantation antibodies, human antibodies, humanized antibodies, synthetic antibodies, chimeric antibodies, camelized antibodies, single-chain Fv (scFv) (including fragments in which VL and VH are joined by a synthetic or natural linker using recombinant methods such that they can be a single protein chain, where the VL and VH regions pair to form a monovalent molecule, including single-chain Fab and scFab), single-chain antibodies, Fab fragments (including monovalent fragments containing the VL, VH, CL, and CH1 domains), F(ab′)2 fragments (including divalent fragments containing two Fab fragments linked by a disulfide bond in the hinge region), Fd fragments (including fragments containing the VH and CH1 fragments), Fv fragments (including fragments containing the VL and VH domains of a single arm of an antibody), single-domain antibodies (dAb or sdAb) (including fragments containing the VH domain), isolated complementarity-determining regions (CDR), diabodies (including fragments containing a divalent dimer, such as two VL and VH domains that bind to each other and recognize two different antigens), fragments consisting of only a single monomeric variable domain, disulfide-linked Fv (sdFv), intracellular antibodies, anti-idiotypic (anti-Id) antibodies, or antigen-binding fragments thereof. In some cases, the libraries disclosed herein contain nucleic acids encoding scaffolds, where the scaffold is an Fv antibody, including an Fv antibody consisting of the minimal antibody fragment containing the complete antigen recognition and antigen-binding site. In some embodiments, the Fv antibody consists of a dimer of one heavy-chain variable domain and one light-chain variable domain that are tightly, non-covalently associated, and the three hypervariable regions of each variable domain interact to define an antigen-binding site on the surface of the VH-VL dimer. In some embodiments, these six hypervariable regions confer antigen-binding specificity to the antibody. In some embodiments, a single variable domain (or half of an Fv containing only the three hypervariable regions specific for an antigen, including single-domain antibodies isolated from camelids, such as VHH antibodies or nanobodies, containing one heavy-chain variable domain) has the ability to recognize and bind an antigen. In some cases, the libraries disclosed herein contain nucleic acids encoding scaffolds, where the scaffold is a single-chain Fv or scFv, including antibody fragments containing the VH, VL, or both the VH and VL domains, where both domains are present in a single polypeptide chain. In some embodiments, the Fv polypeptide further contains a polypeptide linker between the VH and VL domains, thereby allowing the scFv to form the required structure for antigen binding.In some cases, the scFv is linked to an Fc fragment, or the VHH is linked to an Fc fragment (including nanobodies). In some cases, the antibody comprises an immunoglobulin molecule and an immunologically active fragment of the immunoglobulin molecule, e.g., a molecule containing an antigen-binding site. The immunoglobulin molecule can be of any type (e.g., IgG, IgE, IgM, IgD, IgA, and IgY), class (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2), or subclass.
[0085] In some embodiments, the library comprises immunoglobulins of a class suitable for the intended therapeutic target. Generally, these methods involve "mammalianization" and include methods for transferring donor antigen-binding information to a less immunogenic mammalian antibody acceptor to generate a useful therapeutic agent. In some cases, the mammal is a mouse, rat, horse, sheep, cow, primate (e.g., chimpanzee, baboon, gorilla, orangutan, monkey), dog, cat, pig, donkey, rabbit, and human. In some cases, libraries and methods for feline and canine antibody engineering are provided herein.
[0086] A "humanized" form of a non-human antibody can be a chimeric antibody containing minimal sequences derived from the non-human antibody. A humanized antibody is generally a human antibody (acceptor antibody) in which residues from one or more CDRs are replaced with residues from one or more CDRs of a non-human antibody (donor antibody). The donor antibody can be any suitable non-human antibody, such as a mouse, rat, rabbit, chicken, or non-human primate antibody having the desired specificity, affinity, or biological effect. In some cases, selected framework region residues of the acceptor antibody are replaced with the corresponding framework region residues from the donor antibody. A humanized antibody can also contain residues not found in either the acceptor antibody or the donor antibody. In some cases, these modifications are made to further improve antibody properties.
[0087] "Caninization" can include methods of treatment that transfer non-canine antigen-binding information from a donor antibody to a less immunogenic canine antibody acceptor to generate a therapeutic agent that can be used in dogs. In some cases, the caninized form of the non-canine antibody provided herein is a chimeric antibody containing a minimal sequence derived from the non-canine antibody. In some cases, the caninized antibody is a canine antibody sequence ("acceptor" or "recipient" antibody), where the hypervariable region residues of the acceptor are replaced with hypervariable region residues from a non-canine species ("donor" antibody) such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant sequence, or an engineered sequence with desired properties. In some cases, the framework region (FR) residues of the canine antibody are replaced with the corresponding non-canine FR residues. In some cases, the caninized antibody includes residues not found in the acceptor antibody or the donor antibody. In some cases, these modifications are made to further improve antibody performance. The caninized antibody may also contain at least a portion of the immunoglobulin constant region (Fc) of the canine antibody.
[0088] "Felinization" can include methods of treatment that transfer non-feline antigen-binding information from a donor antibody to a less immunogenic feline antibody acceptor to generate a therapeutic agent that can be used in cats. In some cases, the felinized form of the non-feline antibody provided herein is a chimeric antibody containing a minimal sequence derived from the non-feline antibody. In some cases, the felinized antibody is a feline antibody sequence ("acceptor" or "recipient" antibody), where the hypervariable region residues of the acceptor are replaced with hypervariable region residues from a non-feline species ("donor" antibody) such as mouse, rat, rabbit, cat, dog, goat, chicken, cow, horse, llama, camel, dromedary, shark, non-human primate, human, humanized, recombinant sequence, or an engineered sequence with desired properties. In some cases, the framework region (FR) residues of the feline antibody are replaced with the corresponding non-feline FR residues. In some cases, the felinized antibody includes residues not found in the acceptor antibody or the donor antibody. In some cases, these modifications are made to further improve antibody performance. The felinized antibody may also contain at least a portion of the immunoglobulin constant region (Fc) of the feline antibody.
[0089] The present disclosure provides libraries of nucleic acids encoding scaffolds, where the scaffold is a non-immunoglobulin. In some cases, the scaffold is a non-immunoglobulin binding domain. For example, the scaffold is an antibody mimetic. Exemplary antibody mimetics include, but are not limited to, anticalins, affilins, affibody molecules, affimers, affitins, alphabodies, avimers, atrimers, DARPins, fynomers, Kunitz domain-based proteins, monobodies, anticalins, knottins, armadillo repeat protein-based proteins, and bicyclic peptides.
[0090] The libraries of nucleic acids encoding scaffolds (where the scaffold is an immunoglobulin) described herein contain variations in at least one region of the immunoglobulin. Exemplary regions of the antibody for variation include, but are not limited to, complementarity determining regions (CDRs), variable domains, or constant domains. In some cases, the CDR is CDR1, CDR2, or CDR3. In some cases, the CDR is a heavy chain domain, including but not limited to CDR-H1, CDR-H2, and CDR-H3. In some cases, the CDR is a light chain domain, including but not limited to CDR-L1, CDR-L2, and CDR-L3. In some cases, the variable domain is a light chain variable domain (VL) or a heavy chain variable domain (VH). In some cases, the VL domain contains a κ or λ chain. In some cases, the constant domain is a light chain constant domain (CL) or a heavy chain constant domain (CH).
[0091] The methods described herein provide for the synthesis of libraries of nucleic acids comprising nucleic acids encoding a scaffolding, where each nucleic acid encodes a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library comprises sequences encoding variations of at least a single codon such that multiple different variants of a single residue in the subsequent protein encoded by the synthetic nucleic acid are generated by the standard translation process. In some cases, the scaffolding library comprises nucleic acids encoding variations that co-encode variations at multiple positions. In some cases, the variant library comprises sequences encoding variations of at least a single codon of a CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, or VH domain. In some cases, the variant library comprises sequences encoding variations of multiple codons of a CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, or VH domain. In some cases, the variant library comprises sequences encoding variations of multiple codons of a framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). Exemplary numbers of codons for variation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0092] In some cases, at least one region of the immunoglobulin for variation is from a heavy chain V gene family, a heavy chain D gene family, a heavy chain J gene family, a light chain V gene family, or a light chain J gene family. See Figures 1A - 1B. In some cases, the light chain V gene family comprises immunoglobulin kappa (IGK) genes or immunoglobulin lambda (IGL). Exemplary genes include, but are not limited to, IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV1-69, IGHV3-74, IGHV4-39, IGHV4-59 / 61, IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, and IGLV3-1. In some cases, the gene is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, or IGHV1-8. In some cases, the gene is IGHV1-69 and IGHV3-30. In some cases, the gene is IGHJ3, IGHJ6, IGHJ, IGHJ4, IGHJ5, IGHJ2, or IGH1. In some cases, the gene is IGHJ3, IGHJ6, IGHJ, or IGHJ4.
[0093] The present disclosure provides libraries comprising nucleic acids encoding immunoglobulin scaffolds, wherein the libraries are synthesized with various numbers of segments. In some cases, the segments comprise CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, or VH domains. In some cases, the segments comprise framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some cases, the scaffold library is synthesized with at least or about 2 segments, 3 segments, 4 segments, 5 segments, or more than 5 segments. The length of each nucleic acid segment or the average length of the synthesized nucleic acids can be at least or about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some cases, the length is about 50 to 600, 75 to 575, 100 to 550, 125 to 525, 150 to 500, 175 to 475, 200 to 450, 225 to 425, 250 to 400, 275 to 375, or 300 to 350 base pairs.
[0094] When translating, a library of nucleic acids comprising immunoglobulin scaffolds encoded as described herein comprises amino acids of various lengths. In some cases, the length of each amino acid fragment or the average length of the synthesized amino acids can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more than 150 amino acids. In some cases, the length of the amino acids is about 15 to 150, 20 to 145, 25 to 140, 30 to 135, 35 to 130, 40 to 125, 45 to 120, 50 to 115, 55 to 110, 60 to 110, 65 to 105, 70 to 100 or 75 to 95 amino acids. In some cases, the length of the amino acids is about 22 amino acids to about 75 amino acids. In some cases, the immunoglobulin scaffold comprises at least or about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more than 5000 amino acids.
[0095] Using the methods described herein, a large number of variant sequences are de novo synthesized for at least one region of an immunoglobulin for mutagenesis. In some cases, a large number of variant sequences are de novo synthesized for CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH or a combination thereof. In some cases, a large number of variant sequences are de novo synthesized for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3) or framework element 4 (FW4). The number of variant sequences can be at least or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more than 500 sequences. In some cases, the number of variant sequences is at least or about 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000 or more than 8000 sequences. In some cases, the number of variant sequences is about 10 to 500, 25 to 475, 50 to 450, 75 to 425, 100 to 400, 125 to 375, 150 to 350, 175 to 325, 200 to 300, 225 to 375, 250 to 350 or 275 to 325 sequences.
[0096] In some cases, the variant sequences for at least one region of the immunoglobulin vary in length or sequence. In some cases, at least one region synthesized de novo is for CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH, or a combination thereof. In some cases, at least one region synthesized de novo is for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3), or framework element 4 (FW4). In some cases, the variant sequence contains at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more than 50 variant nucleotides or amino acids compared to the wild type. In some cases, the variant sequence contains at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 additional nucleotides or amino acids compared to the wild type. In some cases, the variant sequence contains at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 fewer nucleotides or amino acids than the wild type. In some cases, the library contains at least or about 10 1 、10 2 、10 3 、10 4 、10 5 、10 6 、10 7 、10 8 、10 9 、10 10 or more than 10 10 variants.
[0097] After synthesizing the scaffold library, the scaffold library can be used for screening and analysis. For example, analyzing the library displayability and panning of the scaffold library. In some cases, selective tag analysis is used to analyze displayability. Exemplary tags include, but are not limited to, radiolabels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some cases, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG.
[0098] In some cases, the scaffold library is analyzed by sequencing using various methods, including but not limited to single molecule real-time (SMRT) sequencing, polymerase cloning (Polony) sequencing, ligation sequencing, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or synthesis sequencing.
[0099] In some cases, the functional activity, structural stability (e.g., thermostable or pH-stable), expression, specificity, or a combination thereof of a scaffold library is analyzed. In some cases, scaffolds capable of folding in the scaffold library are analyzed. In some cases, the functional activity, structural stability, expression, specificity, folding, or a combination thereof of an antibody region is analyzed. For example, the functional activity, structural stability, expression, specificity, folding, or a combination thereof of the VH region or the VL region is analyzed.
[0100] GLP1R library
[0101] Provided herein is a GLP1R-binding library comprising nucleic acids encoding scaffolds that comprise sequences of GLP1R-binding domains. In some cases, the scaffold is an immunoglobulin. In some cases, scaffolds comprising sequences of GLP1R-binding domains are identified by the interaction between the GLP1R-binding domain and GLP1R.
[0102] Provided herein is a library comprising nucleic acids encoding scaffolds that comprise GLP1R-binding domains, wherein the GLP1R-binding domains are designed based on surface interactions on GLP1R. In some cases, the GLP1R comprises the sequence defined in SEQ ID NO:1. In some cases, the GLP1R-binding domain interacts with the amino or carboxyl terminus of GLP1R. In some cases, the GLP1R-binding domain interacts with at least one transmembrane domain, which includes but is not limited to transmembrane domain 1 (TM1), transmembrane domain 2 (TM2), transmembrane domain 3 (TM3), transmembrane domain 4 (TM4), transmembrane domain 5 (TM5), transmembrane domain 6 (TM6), and transmembrane domain 7 (TM7). In some cases, the GLP1R-binding domain interacts with the intracellular surface of GLP1R. For example, the GLP1R-binding domain interacts with at least one intracellular loop, which includes but is not limited to intracellular loop 1 (ICL1), intracellular loop 2 (ICL2), and intracellular loop 3 (ICL3). In some cases, the GLP1R-binding domain interacts with the extracellular surface of GLP1R. For example, the GLP1R-binding domain interacts with at least one extracellular domain (ECD) or extracellular loop (ECL) of GLP1R. The extracellular loop includes but is not limited to extracellular loop 1 (ECL1), extracellular loop 2 (ECL2), and extracellular loop 3 (ECL3).
[0103] This text describes a GLP1R binding domain, which is designed based on the surface interaction between a GLP1R ligand and GLP1R. In some cases, the ligand is a peptide. In some cases, the ligand is glucagon, glucagon-like peptide 1-(7-36) amide, glucagon-like peptide 1-(7-37), liraglutide, exendin-4, lixisenatide, T-0632, GLP1R0017, or BETP. In some cases, the ligand is a GLP1R agonist. In some cases, the ligand is a GLP1R antagonist. In some cases, the ligand is a GLP1R allosteric modulator. In some cases, the allosteric modulator is a negative allosteric modulator. In some cases, the allosteric modulator is a positive allosteric modulator.
[0104] The sequence of the GLP1R binding domain based on the surface interaction between a GLP1R ligand and GLP1R is analyzed using various methods. For example, a multi-species computational analysis is performed. In some cases, a structural analysis is performed. In some cases, a sequence analysis is performed. Sequence analysis can be carried out using databases known in the art. Non-limiting examples of databases include but are not limited to NCBI BLAST (blast.ncbi.nlm.nih.gov / Blast.cgi), UCSC Genome Browser (genome.ucsc.edu / ), UniProt (www.uniprot.org / ), and IUPHAR / BPS Guide to PHARMACOLOGY (guidetopharmacology.org / ).
[0105] This text describes a GLP1R binding domain designed based on sequence analysis among various organisms. For example, sequence analysis is performed to identify homologous sequences in different organisms. Exemplary organisms include but are not limited to mice, rats, horses, sheep, cows, primates (such as chimpanzees, baboons, gorillas, orangutans, monkeys), dogs, cats, pigs, donkeys, rabbits, fish, flies, and humans.
[0106] After identifying the GLP1R binding domain, a library comprising nucleic acids encoding the GLP1R binding domain can be generated. In some cases, the GLP1R binding domain library comprises GLP1R binding domain sequences designed based on conformational ligand interactions, peptide ligand interactions, small molecule ligand interactions, the extracellular domain of GLP1R, or antibodies targeting GLP1R. In some cases, the GLP1R binding domain library comprises sequences of GLP1R binding domains designed based on peptide ligand interactions. The library of GLP1R binding domains can be translated to generate a protein library. In some cases, the library of GLP1R binding domains is translated to generate a peptide library, an immunoglobulin library, derivatives thereof, or combinations thereof. In some cases, the library of GLP1R binding domains is translated to generate a protein library, and the protein library is further modified to generate a peptidomimetic library. In some cases, the library of GLP1R binding domains is translated to generate a protein library for generating small molecules.
[0107] The methods described herein provide for the synthesis of a GLP1R binding domain library that comprises nucleic acids each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library comprises sequences encoding variations of at least a single codon such that multiple different variants of a single residue in the subsequent protein encoded by the synthetic nucleic acids are generated by standard translation processes. In some cases, the GLP1R binding domain library comprises nucleic acids encoding variations at multiple positions. In some cases, the variant library comprises sequences encoding variations of at least a single codon in the GLP1R binding domain. In some cases, the variant library comprises sequences encoding variations of multiple codons in the GLP1R binding domain. Exemplary numbers of codons for variation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0108] The methods described herein provide for the synthesis of a library of nucleic acids encoding a GLP1R binding domain, wherein the library comprises sequences encoding length variants of the GLP1R binding domain. In some cases, the library comprises sequences encoding length variants that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300 or more codons shorter than a predetermined reference sequence. In some cases, the library comprises sequences encoding length variants that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300 or more codons longer than a predetermined reference sequence.
[0109] After identifying the GLP1R binding domain, the GLP1R binding domain can be placed in a scaffold as described herein. In some cases, the scaffold is an immunoglobulin. In some cases, the GLP1R binding domain is placed in the CDR-H3 region. The GPCR binding domain that can be placed in the scaffold can also be referred to as a motif. Scaffolds containing the GLP1R binding domain can be designed based on binding, specificity, stability, expression, folding, or downstream activity. In some cases, scaffolds containing the GLP1R binding domain enable contact with GLP1R. In some cases, scaffolds containing the GLP1R binding domain enable high-affinity binding to GLP1R. Exemplary amino acid sequences of the GLP1R binding domain are described in Table 1.
[0110] Table 1. GLP1R Amino Acid Sequences
[0111]
[0112] The present disclosure provides scaffolds comprising GLP1R binding domains, wherein the sequence of the GLP1R binding domain supports interaction with GLP1R. The sequence can be homologous or identical to the sequence of a GLP1R ligand. In some cases, the GLP1R binding domain sequence has at least or about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO:1. In some cases, the GLP1R binding domain sequence has at least or about 95% homology with SEQ ID NO:1. In some cases, the GLP1R binding domain sequence has at least or about 97% homology with SEQ ID NO:1. In some cases, the GLP1R binding domain sequence has at least or about 99% homology with SEQ ID NO:1. In some cases, the GLP1R binding domain sequence has at least or about 100% homology with SEQ ID NO:1. In some cases, the GLP1R binding domain sequence comprises at least a portion having at least or about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400 or more than 400 amino acids of SEQ ID NO:1.
[0113] The term "sequence identity" means that two polynucleotide sequences are identical (i.e., on a nucleotide-by-nucleotide basis) within a comparison window. The term "percent sequence identity" is calculated as follows: compare two optimally aligned sequences within the comparison window, determine the number of positions at which the identical nucleic acid bases (e.g., A, T, C, G, U or I) occur in both sequences to yield the number of matching positions, divide the number of matching positions by the total number of positions within the comparison window (i.e., window size), and multiply the result by 100 to obtain the percent sequence identity.
[0114] The "homology" or "similarity" between two proteins is determined by comparing the amino acid sequence of one protein sequence and its conservative amino acid substitutions with a second protein sequence. Similarity can be determined by procedures well known in the art, such as the BLAST program (Basic Local Alignment Search Tool of the National Center for Biotechnology Information).
[0115] The present disclosure provides a GLP1R-binding library that includes nucleic acids encoding scaffolds that include GLP1R-binding domains, where the scaffolds include variations in domain type, domain length, or residue variation. In some cases, the domain is a region in a scaffold that includes a GLP1R-binding domain. For example, the region is a VH, CDR-H3, or VL domain. In some cases, the domain is the GLP1R-binding domain.
[0116] The methods described herein provide a GLP1R-binding library of nucleic acids that each encode a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding variations of at least a single codon such that multiple different variants of a single residue in the subsequent protein encoded by the synthetic nucleic acid are generated by standard translation processes. In some cases, the GLP1R-binding library includes nucleic acids encoding variations at multiple positions. In some cases, the variant library includes sequences encoding variations of at least a single codon of a VH, CDR-H3, or VL domain. In some cases, the variant library includes sequences encoding variations of at least a single codon in a GLP1R-binding domain. For example, at least one single codon of the GLP1R-binding domain listed in Table 1 is different. In some cases, the variant library includes sequences encoding variations of multiple codons of a VH, CDR-H3, or VL domain. In some cases, the variant library includes sequences encoding variations of multiple codons in a GLP1R-binding domain. Exemplary numbers of codons for variation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons.
[0117] The methods described herein provide a GLP1R-binding library of nucleic acids that synthesize respective variants encoding at least one predetermined reference nucleic acid sequence, wherein the GLP1R-binding library comprises sequences with length variations in the encoding domain. In some cases, the domain is a VH, CDR-H3, or VL domain. In some cases, the domain is a GLP1R-binding domain. In some cases, the library comprises sequences encoding length variations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons less than the predetermined reference sequence. In some cases, the library comprises sequences encoding length variations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 codons more than the predetermined reference sequence.
[0118] The present invention provides a GLP1R-binding library comprising nucleic acids encoding a scaffold containing a GLP1R-binding domain, wherein the GLP1R-binding library is synthesized with various numbers of fragments. In some cases, the fragments comprise VH, CDR-H3, or VL domains. In some cases, the GLP1R-binding library is synthesized with at least or about 2, 3, 4, 5, or more than 5 fragments. The length of each nucleic acid fragment or the average length of the synthesized nucleic acids can be at least or about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, or more than 600 base pairs. In some cases, the length is about 50 to 600, 75 to 575, 100 to 550, 125 to 525, 150 to 500, 175 to 475, 200 to 450, 225 to 425, 250 to 400, 275 to 375, or 300 to 350 base pairs.
[0119] When translated, a GLP1R binding domain library comprising nucleic acids encoding scaffolds containing a GLP1R binding domain as described herein contains amino acids of various lengths. In some cases, the length of each amino acid fragment or the average length of the synthesized amino acids can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more than 150 amino acids. In some cases, the amino acid length is about 15 to 150, 20 to 145, 25 to 140, 30 to 135, 35 to 130, 40 to 125, 45 to 120, 50 to 115, 55 to 110, 60 to 110, 65 to 105, 70 to 100 or 75 to 95 amino acids. In some cases, the amino acid length is about 22 to about 75 amino acids.
[0120] A GLP1R binding library comprising de novo synthesized variant sequences encoding scaffolds containing a GLP1R binding domain contains a large number of variant sequences. In some cases, a large number of variant sequences are de novo synthesized for CDR-H1, CDR-H2, CDR-H3, CDR-L1, CDR-L2, CDR-L3, VL, VH or combinations thereof. In some cases, a large number of variant sequences are de novo synthesized for framework element 1 (FW1), framework element 2 (FW2), framework element 3 (FW3) or framework element 4 (FW4). In some cases, a large number of variant sequences are de novo synthesized for the GPCR binding domain. For example, the number of variant sequences is about 1 to about 10 sequences for the VH domain, 10 8 sequences for the GLP1R binding domain, and about 1 to about 44 sequences for the VK domain. The number of variant sequences can be at least or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more than 500 sequences. In some cases, the number of variant sequences is about 10 to 300, 25 to 275, 50 to 250, 75 to 225, 100 to 200 or 125 to 150 sequences.
[0121] A GLP1R-binding library comprising variant sequences encoding scaffolds containing GLP1R-binding domains synthesized de novo has improved diversity. For example, variants are generated by placing GLP1R-binding domain variants in immunoglobulin scaffold variants containing N-terminal CDR-H3 variants and C-terminal CDR-H3 variants. In some cases, the variants include affinity matured variants. Alternatively or in combination, the variants include variants in other regions of the immunoglobulin, including but not limited to CDR-H1, CDR-H2, CDR-L1, CDR-L2, and CDR-L3. In some cases, the number of variants in the GLP1R-binding library is at least or about 10 4 、10 5 、10 6 、10 7 、10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16 、10 17 、10 18 、10 19 、10 20 or more than 10 20 distinct sequences. For example, a library containing approximately 10 VH region variant sequences, approximately 237 CDR-H3 region variant sequences, and approximately 43 VL and CDR-L3 region variant sequences contains 10 5 distinct sequences (10 x 237 x 43).
[0122] The present disclosure provides a library comprising nucleic acids encoding GLP1R antibodies that contain variants in at least one region of the antibody, wherein the region is a CDR region. In some cases, the GLP1R antibody is a single domain antibody comprising a heavy chain variable domain, such as a VHH antibody. In some cases, the VHH antibody contains variants in one or more CDR regions. In some cases, the libraries described herein contain at least or about 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000 or more than 3000 sequences of CDR1, CDR2 or CDR3. In some cases, the libraries described herein contain at least or about 10 4 、10 5 、106 and 10 7 and 10 8 and 10 9 and 10 10 and 10 11 and 10 12 and 10 13 and 10 14 and 10 15 and 10 16 and 10 17 and 10 18 and 10 19 and 10 20 or more than 10 20 sequences. For example, the library contains at least 2000 CDR1 sequences, at least 1200 CDR2 sequences, and at least 1600 CDR3 sequences. In some cases, each sequence is different.
[0123] In some cases, the CDR1, CDR2, or CDR3 is of the variable light chain domain (VL). The CDR1, CDR2, or CDR3 of the variable light chain domain (VL) may be referred to as CDR-L1, CDR-L2, or CDR-L3, respectively. In some cases, the library described herein contains at least or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000 or more than 3000 sequences of CDR1, CDR2, or CDR3 of VL. In some cases, the library described herein contains at least or approximately 10 4 and 10 5 and 10 6 and 10 7 and 10 8 and 10 9 and 10 10 and 10 11 and 10 12 and 10 13 and 10 14 and 10 15 and 10 16 and 10 17 and 10 18 and 10 19 and 10 20 or more than 10 20sequences. For example, the library comprises at least 20 CDR1 sequences of VL, at least 4 CDR2 sequences of VL, and at least 140 CDR3 sequences of VL. In some cases, the library comprises at least 2 CDR1 sequences of VL, at least 1 CDR2 sequence of VL, and at least 3000 CDR3 sequences of VL. In some cases, the VL is IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, IGLV2-14, IGLV1-40, or IGLV3-1. In some cases, the VL is IGKV2-28. In some cases, the VL is IGLV1-51.
[0124] In some cases, the CDR1, CDR2, or CDR3 is of the variable heavy domain (VH). The CDR1, CDR2, or CDR3 of the variable heavy domain (VH) may be referred to as CDR-H1, CDR-H2, or CDR-H3, respectively. In some cases, the library described herein comprises at least or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2400, 2600, 2800, 3000, or more than 3000 sequences of CDR1, CDR2, or CDR3 of VH. In some cases, the library described herein comprises at least or about 10 4 、10 5 、10 6 、10 7 、10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16 、10 17 、10 18 、10 19 、10 20 or more than 10 20 sequences. For example, the library comprises at least 30 CDR1 sequences of VH, at least 570 CDR2 sequences of VH, and at least 10 8CDR3 sequences of a number of VHs. In some cases, the library comprises at least 30 CDR1 sequences of VHs, at least 860 CDR2 sequences of VHs and at least 10 7 CDR3 sequences of VHs. In some cases, the VH is IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV3-74, IGHV4-39 or IGHV4-59 / 61. In some cases, the VH is IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1 or IGHV1-8. In some cases, the VH is IGHV1-69 and IGHV3-30. In some cases, the VH is IGHV3-23.
[0125] In some embodiments, the libraries as described herein comprise CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2 or CDR-H3 of different lengths. In some cases, the lengths of the CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2 or CDR-H3 comprise at least or about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90 or more than 90 amino acids. For example, the CDR-H3 comprises at least or about 12, 15, 16, 17, 20, 21 or 23 amino acids in length. In some cases, the CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2 or CDR-H3 comprise amino acids in the range of about 1 to about 10, about 5 to about 15, about 10 to about 20 or about 15 to about 30 in length.
[0126] When translating, a library of nucleic acids encoding antibodies with variant CDR sequences as described herein contains amino acids of various lengths. In some cases, the length of each amino acid fragment or the average length of the synthesized amino acids can be at least or about 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150 or more than 150 amino acids. In some cases, the amino acid length is about 15 to 150, 20 to 145, 25 to 140, 30 to 135, 35 to 130, 40 to 125, 45 to 120, 50 to 115, 55 to 110, 60 to 110, 65 to 105, 70 to 100 or 75 to 95 amino acids. In some cases, the amino acid length is about 22 amino acids to about 75 amino acids. In some cases, the antibody contains at least or about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more than 5000 amino acids.
[0127] The length ratios of CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2 or CDR-H3 can vary in the libraries described herein. In some cases, CDR-L1, CDR-L2, CDR-L3, CDR-H1, CDR-H2 or CDR-H3 containing at least or about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90 or more than 90 amino acids in length account for about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more than 90% of the library. For example, CDR-H3 containing about 23 amino acids in length is present in the library at 40%, CDR-H3 containing about 21 amino acids in length is present in the library at 30%, CDR-H3 containing about 17 amino acids in length is present in the library at 20%, and CDR-H3 containing about 12 amino acids in length is present in the library at 10%. In some cases, CDR-H3 containing about 20 amino acids in length is present in the library at 40%, CDR-H3 containing about 16 amino acids in length is present in the library at 30%, CDR-H3 containing about 15 amino acids in length is present in the library at 20%, and CDR-H3 containing about 12 amino acids in length is present in the library at 10%.
[0128] The library of encoded VHH antibodies as described herein contains diversified CDR sequences of a library of at least or about 10 7 、10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16 、10 17 、10 18 、10 19 、10 20 or more than 10 20 sequences. In some cases, the final library diversity of the library is at least or about 10 7 、10 8 、10 9 、10 10 、10 11 、10 12 、10 13 、10 14 、10 15 、10 16 、10 17 、10 18 、10 19 、10 20 or more than 10 20 sequences.
[0129] The present disclosure provides a GLP1R-binding library encoding immunoglobulins. In some cases, the GLP1R immunoglobulin is an antibody. In some cases, the GLP1R immunoglobulin is a VHH antibody. In some cases, the binding affinity (e.g., kD) of the GLP1R immunoglobulin for GLP1R is less than 1 nM, less than 1.2 nM, less than 2 nM, less than 5 nM, less than 10 nM, less than 11 nM, less than 13.5 nM, less than 15 nM, less than 20 nM, less than 25 nM, or less than 30 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 1 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 1.2 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 2 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 5 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 10 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 13.5 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 15 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 20 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 25 nM. In some cases, the GLP1R immunoglobulin has a kD of less than 30 nM.
[0130] In some cases, the GLP1R immunoglobulin is a GLP1R agonist. In some cases, the GLP1R immunoglobulin is a GLP1R antagonist. In some cases, the GLP1R immunoglobulin is a GLP1R allosteric modulator. In some cases, the allosteric modulator is a negative allosteric modulator. In some cases, the allosteric modulator is a positive allosteric modulator. In some cases, the GLP1R immunoglobulin causes an agonistic, antagonistic or allosteric effect at a concentration of at least or about 1 nM, 2 nM, 4 nM, 6 nM, 8 nM, 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM, 120 nM, 140 nM, 160 nM, 180 nM, 200 nM, 300 nM, 400 nM, 500 nM, 600 nM, 700 nM, 800 nM, 900 nM, 1000 nM or more than 1000 nM. In some cases, the GLP1R immunoglobulin is a negative allosteric modulator. In some cases, the GLP1R immunoglobulin is a negative allosteric modulator at a concentration of at least or about 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1 nM, 2 nM, 4 nM, 6 nM, 8 nM, 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM or more than 100 nM. In some cases, the GLP1R immunoglobulin is a negative allosteric modulator at a concentration in the range of about 0.001 to about 100, 0.01 to about 90, about 0.1 to about 80, 1 to about 50, about 10 to about 40 nM or about 1 to about 10 nM. In some cases, the EC50 or IC50 of the GLP1R immunoglobulin is at least or about 0.001, 0.0025, 0.005, 0.01, 0.025, 0.05, 0.06, 0.07, 0.08, 0.9, 0.1, 0.5, 1, 2, 3, 4, 5, 6 nM or more than 6 nM. In some cases, the EC50 or IC50 of the GLP1R immunoglobulin is at least or about 1 nM, 2 nM, 4 nM, 6 nM, 8 nM, 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM or more than 100 nM.
[0131] The present disclosure provides a GLP1R-binding library encoding an immunoglobulin having a long half-life. In some cases, the half-life of the GLP1R immunoglobulin is at least or about 12 hours, 24 hours, 36 hours, 48 hours, 60 hours, 72 hours, 84 hours, 96 hours, 108 hours, 120 hours, 140 hours, 160 hours, 180 hours, 200 hours, or more than 200 hours. In some cases, the half-life of the GLP1R immunoglobulin ranges from about 12 hours to about 300 hours, from about 20 hours to about 280 hours, from about 40 hours to about 240 hours, or from about 60 hours to about 200 hours.
[0132] The GLP1R immunoglobulins as described herein may have improved properties. In some cases, the GLP1R immunoglobulin is monomeric. In some cases, the GLP1R immunoglobulin is less prone to aggregation. In some cases, at least or about 70%, 75%, 80%, 85%, 90%, 95%, or 99% of the GLP1R immunoglobulins are monomeric. In some cases, the GLP1R immunoglobulin is thermostable. In some cases, the GLP1R immunoglobulin results in reduced non-specific binding.
[0133] After synthesizing a GLP1R-binding library comprising a nucleic acid encoding a scaffold containing a GLP1R-binding domain, the library can be used for screening and analysis. For example, library display and panning of the library can be analyzed. In some cases, selective tags are used to analyze library display. Exemplary tags include, but are not limited to, radiolabels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some cases, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. In some cases, the GLP1R-binding library comprises a nucleic acid encoding a scaffold comprising a GPCR-binding domain having multiple tags such as GFP, FLAG, and Lucy, as well as a DNA barcode. In some cases, the library is characterized by sequencing using various methods, including but not limited to single molecule real-time sequencing (SMRT), Polony sequencing, ligation sequencing, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis.
[0134] Expression system
[0135] The present disclosure provides a library comprising a nucleic acid encoding a scaffold containing a GLP1R-binding domain, wherein the library has improved specificity, stability, expression, folding, or downstream activity. In some cases, the libraries as described herein are used for screening and analysis.
[0136] The present disclosure provides a library of nucleic acids comprising nucleic acids encoding scaffolds containing a GLP1R binding domain, wherein the nucleic acid library is for screening and analysis. In some cases, the screening and analysis include in vitro, in vivo, or ex vivo assays. Cells for screening include primary cells or cell lines taken from a living subject. The cells can be from prokaryotes (such as bacteria and fungi) or eukaryotes (such as animals and plants). Exemplary animal cells include, but are not limited to, animal cells from mice, rabbits, primates, and insects. In some cases, the cells for screening include cell lines, including but not limited to, Chinese hamster ovary (CHO) cell line, human embryonic kidney (HEK) cell line, or baby hamster kidney (BHK) cell line. In some cases, the nucleic acid library described herein can also be delivered to a multicellular organism. Exemplary multicellular organisms include, but are not limited to, plants, mice, rabbits, primates, and insects.
[0137] The nucleic acid library described herein or a library of proteins encoded thereby can be screened for various pharmacological or pharmacokinetic properties. In some cases, in vitro assays, in vivo assays, or ex vivo assays are used to screen the library. For example, the in vitro pharmacological or pharmacokinetic properties screened include, but are not limited to, binding affinity, binding specificity, and binding avidity. Exemplary in vivo pharmacological or pharmacokinetic properties of the library described herein include, but are not limited to, therapeutic efficacy, activity, preclinical toxicity properties, clinical efficacy properties, clinical toxicity properties, immunogenicity, potency, and clinical safety properties.
[0138] Pharmacological or pharmacokinetic properties that can be screened include, but are not limited to, cell binding affinity and cell activity. For example, cell binding affinity assays or cell activity assays are performed to determine the agonistic, antagonistic, or allosteric effects of the library described herein. In some cases, the cell activity assay is a cAMP assay. In some cases, the library described herein is compared with the cell binding or cell activity of a GLP1R ligand.
[0139] The library described herein can be screened in cell-based assays or in cell-free assays. Examples of cell-free assays include, but are not limited to, using viral particles, using in vitro translated proteins, and using liposomes having GLP1R.
[0140] Nucleic acid libraries as described herein can be screened by sequencing. In some cases, next-generation sequencing is used to determine sequence enrichment of GLP1R-binding variants. In some cases, V gene distribution, J gene distribution, V gene family, CDR3 count per length, or a combination thereof is determined. In some cases, clonal frequency, clonal accumulation, lineage accumulation, or a combination thereof is determined. In some cases, sequence number, sequences with VH clones, clones, clones greater than 1, clonotypes, clonotypes greater than 1, lineages, simpsons, or a combination thereof is determined. In some cases, the percentage of distinct CDRs3 is determined. For example, the percentage of distinct CDR3s is calculated as the number of distinct CDR3s in the sample divided by the total number of sequences with CDR3s in the sample.
[0141] Nucleic acid libraries are provided herein, where the nucleic acid libraries can be expressed in a vector. Expression vectors for insertion of the nucleic acid libraries disclosed herein can include eukaryotic or prokaryotic expression vectors. Exemplary expression vectors include, but are not limited to, mammalian expression vectors: pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His ("6His” is disclosed as SEQ ID NO: 2410), pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), and pSF-CMV-PURO-NH2-CMYC; bacterial expression vectors: pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, and pSF-Tac; plant expression vectors: pRI 101-AN DNA and pCambia2301; and yeast expression vectors: pTYB21 and pKLAC2, and insect vectors: pAc5.1 / V5-His A and pDEST8. In some cases, the vector is pcDNA3 or pcDNA3.1.
[0142] This text describes a nucleic acid library expressed in a vector to generate constructs containing a scaffold, which scaffold contains the sequence of a GLP1R binding domain. In some cases, the constructs vary in size. In some cases, the constructs contain at least or about 500, 600, 700, 800, 900, 1000, 1100, 1300, 1400, 1500, 1600, 1700, 1800, 2000, 2400, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, 5000, 6000, 7000, 8000, 9000, 10000 or more than 10000 bases. In some cases, the constructs contain in the range of about 300 to 1,000, 300 to 2,000, 300 to 3,000, 300 to 4,000, 300 to 5,000, 300 to 6,000, 300 to 7,000, 300 to 8,000, 300 to 9,000, 300 to 10,000, 1,000 to 2,000, 1,000 to 3,000, 1,000 to 4,000, 1,000 to 5,000, 1,000 to 6,000, 1,000 to 7,000, 1,000 to 8,000, 1,000 to 9,000, 1,000 to 10,000, 2,000 to 3,000, 2,000 to 4,000, 2,000 to 5,000, 2,000 to 6,000, 2,000 to 7,000, 2,000 to 8,000, 2,000 to 9,000, 2,000 to 10,000, 3,000 to 4,000, 3,000 to 5,000, 3,000 to 6,000, 3,000 to 7,000, 3,000 to 8,000, 3,000 to 9,000, 3,000 to 10,000, 4,000 to 5,000, 4,000 to 6,000, 4,000 to 7,000, 4,000 to 8,000, 4,000 to 9,000, 4,000 to 10,000, 5,000 to 6,000, 5,000 to 7,000, 5,000 to 8,000, 5,000 to 9,000, 5,000 to 10,000, 6,000 to 7,000, 6,000 to 8,000, 6,000 to 9,000, 6,000 to 10,000, 7,000 to 8,000, 7,000 to 9,000, 7,000 to 10,000, 8,000 to 9,000, 8,000 to 10,000 or 9,000 to 10,000 bases.
[0143] The present disclosure provides libraries of nucleic acids comprising nucleic acids encoding scaffolds containing GPCR-binding domains, wherein the nucleic acid libraries are expressed in cells. In some cases, the libraries are synthesized to express a reporter gene. Exemplary reporter genes include, but are not limited to, acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), β-galactosidase (LacZ), β-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), cerulean fluorescent protein, citrine fluorescent protein, orange fluorescent protein, cherry fluorescent protein, turquoise fluorescent protein, blue fluorescent protein, horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase and derivatives thereof. Methods for determining the regulation of reporter genes are well known in the art and include, but are not limited to, fluorometry (e.g., fluorescence spectroscopy, fluorescence-activated cell sorting (FACS), fluorescence microscopy) and antibiotic resistance determination.
[0144] Diseases and disorders
[0145] The present disclosure provides GLP1R-binding libraries comprising nucleic acids encoding scaffolds containing GLP1R-binding domains, which may have therapeutic effects. In some cases, the GLP1R-binding libraries produce a protein upon translation, which is used to treat a disease or disorder. In some cases, the protein is an immunoglobulin. In some cases, the protein is a peptidomimetic.
[0146] The GLP1R libraries as described herein may comprise modulators of GLP1R. In some cases, the modulator of GLP1R is an inhibitor. In some cases, the modulator of GLP1R is an activator. In some cases, the GLP1R inhibitor is a GLP1R antagonist. In some cases, the GLP1R antagonist is GLP1R-3. In some cases, GLP1R-3 comprises SEQ ID NO:2279. In some cases, GLP1R-3 comprises SEQ ID NO:2320. In some cases, various diseases or disorders are treated using a modulator of GLP1R.
[0147] Exemplary diseases include, but are not limited to, cancer, inflammatory diseases or disorders, metabolic diseases or disorders, cardiovascular diseases or disorders, respiratory diseases or disorders, pain, digestive diseases or disorders, reproductive diseases or disorders, endocrine diseases or disorders, or nervous system diseases or disorders. In some cases, the cancer is a solid cancer or a hematological cancer. In some cases, the modulators of GLP1R described herein are used to treat weight gain (or to cause weight loss), treat obesity, or treat type II diabetes. In some cases, the GLP1R modulator is used to treat hypoglycemia. In some cases, the GLP1R modulator is used to treat post-weight loss hypoglycemia. In some cases, the GLP1R modulator is used to treat severe hypoglycemia. In some cases, the GLP1R modulator is used to treat hyperinsulinemia. In some cases, the GLP1R modulator is used to treat congenital hyperinsulinemia.
[0148] In some cases, the subject is a mammal. In some cases, the subject is a mouse, rabbit, dog, or human. The subject treated by the methods described herein can be an infant, adult, or child. The pharmaceutical composition comprising the antibody or antibody fragment described herein can be administered intravenously or subcutaneously.
[0149] The present disclosure describes a pharmaceutical composition comprising an antibody or an antibody fragment that binds to GLP1R. In some embodiments, the antibody or antibody fragment comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320, or 2321; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315, or 2316. In some embodiments, the antibody or antibody fragment comprises an immunoglobulin heavy chain and an immunoglobulin light chain. Wherein the immunoglobulin heavy chain comprises the amino acid sequence set forth in SEQ ID NO: 2303, 2304, 2305, 2306, 2307, 2308, 2309, 2317, 2318, 2319, 2320, or 2321; and wherein the immunoglobulin light chain comprises the amino acid sequence set forth in SEQ ID NO: 2310, 2311, 2312, 2313, 2314, 2315, or 2316.
[0150] In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2303; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2310. In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2304; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2311. In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2305; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2312. In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2306; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2313. In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2307; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99%, or 100% identical to the amino acid sequence shown in SEQ ID NO: 2314.In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2308; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2315. In some embodiments, the antibody or antibody fragment thereof comprises an immunoglobulin heavy chain and an immunoglobulin light chain: wherein the immunoglobulin heavy chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2309; and wherein the immunoglobulin light chain comprises an amino acid sequence that is at least about 90%, 95%, 97%, 99% or 100% identical to the amino acid sequence shown in SEQ ID NO: 2316.
[0151] In some cases, the pharmaceutical composition comprises an antibody or antibody fragment comprising CDR-H3 described herein, and the CDR-H3 comprises a sequence of any one of SEQ ID NOs: 2260-2276. In some cases, the pharmaceutical composition comprises an antibody or antibody fragment described herein, and the antibody or antibody fragment comprises a sequence of any one of SEQ ID NOs: 2277-2295. In some cases, the pharmaceutical composition comprises an antibody or antibody fragment described herein, and the antibody or antibody fragment comprises a sequence of any one of SEQ ID NOs: 2277, 2278, 2281, 2282, 2283, 2284, 2285, 2286, 2289, 2290, 2291, 2292, 2294 or 2295. In a further case, the pharmaceutical composition is used for treating metabolic disorders.
[0152] Variant library
[0153] Codon variation
[0154] The mutant nucleic acid libraries described herein can contain multiple nucleic acids, where each nucleic acid encodes a mutant codon sequence compared to a reference nucleic acid sequence. In some cases, each nucleic acid in a first nucleic acid population contains a variant at a single mutation site. In some cases, the first nucleic acid population contains multiple variants at a single mutation site, such that the first nucleic acid population contains more than one variant at the same mutation site. The first nucleic acid population can contain nucleic acids that together encode multiple codon variants at the same mutation site. The first nucleic acid population can contain nucleic acids that together encode up to 19 or more codons at the same position. The first nucleic acid population can contain nucleic acids that together encode up to 60 mutant triplets at the same position, or the first nucleic acid population can contain nucleic acids that together encode up to 61 different codon triplets at the same position. Each variant can encode a codon that produces a different amino acid during translation. Table 3 provides a list of each possible codon (and representative amino acid) for the mutation site.
[0155] Table 2. List of Codons and Amino Acids
[0156]
[0157] A nucleic acid population can contain altered nucleic acids that together encode up to 20 codon mutations at multiple positions. In such cases, each nucleic acid in the population contains codon mutations at more than one position in the same nucleic acid. In some cases, each nucleic acid in the population contains codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more codons in a single nucleic acid. In some cases, each mutant long nucleic acid contains codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more codons in a single long nucleic acid. In some cases, the mutant nucleic acid population contains codon mutations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more codons in a single nucleic acid. In some cases, the mutant nucleic acid population contains codon mutations at at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more codons in a single long nucleic acid.
[0158] Highly parallel nucleic acid synthesis
[0159] The present disclosure provides a platform approach that utilizes miniaturization, parallelization, and vertical integration of an end-to-end process from polynucleotide synthesis to gene assembly within silicon nanopores to create a revolutionary synthesis platform. The devices described herein provide such a silicon synthesis platform with the same footprint as a 96-well plate, which is capable of increasing throughput by up to 1,000-fold or more compared to traditional synthesis methods, where up to about 1,000,000 or more polynucleotides or 10,000 or more genes are generated in a single highly parallelized run.
[0160] With the advent of next-generation sequencing, high-resolution genomic data has become an important factor in studies that delve into the biological roles of various genes in normal biology and disease pathogenesis. Central to this research is the central dogma of molecular biology and the concept of "residue-by-residue transfer of sequential information." The genomic information encoded in DNA is transcribed into messenger RNA, which is subsequently translated into a protein, the active product within a given biological pathway.
[0161] Another exciting area of research is the discovery, development, and preparation of therapeutic molecules that target highly specific cellular targets. Highly diverse DNA sequence libraries are central to the development pipeline of targeted therapeutics. Mutant genes are used to express proteins in the design, construction, and testing protein engineering cycles, which, ideally, are optimized for the high expression of proteins with high affinity for their therapeutic targets. As an example, consider the binding pocket of a receptor. The ability to test all sequence permutations of all residues within the binding pocket would allow for a thorough exploration, thereby increasing the likelihood of success. Saturation mutagenesis, where researchers attempt to generate all possible mutations at specific sites within a receptor, represents one approach to this development challenge. While it is costly, time-consuming, and laborious, it enables the introduction of each variant into each position. In contrast, combinatorial mutagenesis, where several selected positions or short DNA segments are extensively modified, generates incomplete libraries of variants with a biased representation.
[0162] To accelerate the drug development process, libraries with the desired variants available at the correct positions and at the expected frequencies for testing (in other words, precise libraries) enable cost reduction and screening turnaround time. The present disclosure provides methods for synthesizing nucleic acid synthesis variant libraries that are capable of precisely introducing each desired variant at the desired frequency. For the end user, this means not only being able to thoroughly sample sequence space but also being able to query these hypotheses in an efficient manner, thereby reducing costs and screening times. Genome-wide editing can elucidate important pathways, can detect each variant and sequence permutation for a library with optimal functionality, and can use thousands of genes to reconstruct entire pathways and genomes to re-engineer biological systems for drug discovery.
[0163] In the first instance, the drug itself can be optimized using the methods described herein. For example, to improve a specified function of an antibody, a variant polynucleotide library encoding a portion of the antibody is designed and synthesized. A variant nucleic acid library of the antibody can then be generated by the processes described herein (e.g., insertion into a vector following PCR mutagenesis). The antibody is then expressed in a production cell line and screened for enhanced activity. Exemplary screens include examining the binding affinity, stability, or modulation of effector function (e.g., ADCC, complement, or apoptosis) for an antigen. Exemplary regions used to optimize an antibody include, but are not limited to, the Fc region, the Fab region, the variable region of the Fab region, the constant region of the Fab region, the variable domain of the heavy or light chain (V H or V L ) and the specific complementarity determining regions (CDRs) of V H or V L .
[0164] The nucleic acid libraries synthesized by the methods described herein can be expressed in a variety of cells associated with a disease state. Cells associated with a disease state include cell lines, tissue samples, primary cells from a subject, cultured cells expanded from a subject, or cells in a model system. Exemplary model systems include, but are not limited to, plant and animal models of a disease state.
[0165] To identify variant molecules associated with the prevention, alleviation, or treatment of a disease state, the variant nucleic acid libraries described herein are expressed in cells associated with a disease state, or in cells that can induce Cell a disease state. In some cases, an agent is used to induce a disease state in the cells. Exemplary tools for disease state induction include, but are not limited to, the Cre / Lox recombination system, LPS inflammation induction, and streptozotocin used to induce hypoglycemia. Cells associated with a disease state can be cells from a model system or cultured cells, as well as cells from a subject with a specific disease condition. Exemplary disease conditions include bacterial, fungal, viral, autoimmune, or proliferative disorders (e.g., cancer). In some cases, the variant nucleic acid libraries are expressed in a model system, cell line, or primary cells derived from a subject and screened for an alteration in at least one cellular activity. Exemplary cellular activities include, but are not limited to, proliferation, cell cycle progression, cell death, adhesion, migration, replication, cell signaling, energy production, oxygen utilization, metabolic activity, and aging, response to free radical damage, or any combination thereof.
[0166] Substrate
[0167] A device for use as a polynucleotide synthesis surface can be in the form of a substrate, including but not limited to a homogeneous array surface, a patterned array surface, channels, beads, gels, etc. Substrates containing a plurality of clusters are provided herein, where each cluster contains a plurality of sites that support polynucleotide attachment and synthesis. In some cases, the substrate comprises a homogeneous array surface. For example, the homogeneous array surface is a homogeneous plate. As used herein, the term "site" refers to a discrete region of structure that provides support for the extension of a polynucleotide encoding a single predetermined sequence from the surface. In some cases, the sites are on a two-dimensional surface (e.g., a substantially planar surface). In some cases, the sites are on a three-dimensional surface (e.g., a pore, a micro-pore, a channel, or a post). In some cases, the surface of the site contains a material that is activated and functionalized to attach at least one nucleotide for polynucleotide synthesis, or preferably, a population of the same nucleotide for polynucleotide population synthesis. In some cases, the polynucleotide refers to a population of polynucleotides encoding the same nucleic acid sequence. In some cases, the surface of the substrate includes one or more surfaces of the substrate. The average error rate of polynucleotides synthesized within the libraries described herein using the provided systems and methods is typically less than 1 / 1000, less than about 1 / 2000, less than about 1 / 3000, or lower, typically without error correction.
[0168] The present disclosure provides a surface that supports the parallel synthesis of multiple polynucleotides having different predetermined sequences at addressable locations on a common support. In some cases, the substrate supports the synthesis of more than 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more different polynucleotides. In some cases, the surface supports the synthesis of more than 50, 100, 200, 400, 600, 800, 1000, 1200, 1400, 1600, 1800, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more polynucleotides encoding different sequences. In some cases, at least a portion of the polynucleotides have the same sequence or are configured to be synthesized with the same sequence. In some cases, the substrate provides a surface environment for growing polynucleotides having at least 80, 90, 100, 120, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more bases.
[0169] The present disclosure provides methods for synthesizing polynucleotides at different sites of a substrate, where each site supports the synthesis of a population of polynucleotides. In some cases, each site supports the synthesis of a population of polynucleotides having a different sequence from the population of polynucleotides growing at another site. In some cases, each polynucleotide sequence is synthesized to have a redundancy of 1, 2, 3, 4, 5, 6, 7, 8, 9, or more at different sites within the same site cluster on the surface used for polynucleotide synthesis. In some cases, the sites of the substrate are located within multiple clusters. In some cases, the substrate comprises at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000, or more clusters. In some cases, the substrate comprises more than 2,000, 5,000, 10,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,100,000, 1,200,000, 1,300,000, 1,400,000, 1,500,000, 1,600,000, 1,700,000, 1,800,000, 1,900,000, 2,000,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, or 10,000,000, or more different sites. In some cases, the substrate comprises approximately 10,000 different sites. The amount of sites within a single cluster varies in different cases. In some cases, each cluster comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 300, 400, 500, or more sites. In some cases, each cluster comprises approximately 50 - 500 sites. In some cases, each cluster comprises approximately 100 - 200 sites. In some cases, each cluster comprises approximately 100 - 150 sites. In some cases, each cluster comprises approximately 109, 121, 130, or 137 sites. In some cases, each cluster comprises approximately 19, 20, 61, 64, or more sites.Alternatively or in combination, polynucleotide synthesis is performed on a uniform array surface.
[0170] In some cases, the number of different polynucleotides synthesized on a substrate depends on the number of different sites available in the substrate. In some cases, the density of sites within a cluster or on a surface of the substrate is at least or about 1, 10, 25, 50, 65, 75, 100, 130, 150, 175, 200, 300, 400, 500, 1,000 or more sites / mm 2 . In some cases, the substrate comprises 10 - 500, 25 - 400, 50 - 500, 100 - 500, 150 - 500, 10 - 250, 50 - 250, 10 - 200 or 50 - 200 mm 2 . In some cases, the distance between the centers of two adjacent sites within a cluster or on a surface is about 10 - 500, about 10 - 200 or about 10 - 100 µm. In some cases, the distance between the centers of two adjacent sites is greater than about 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 µm. In some cases, the distance between the centers of two adjacent sites is less than about 200, 150, 100, 80, 70, 60, 40, 30, 20 or 10 µm. In some cases, each site has a width of about 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 µm. In some cases, each site has a width of about 0.5 - 100, 0.5 - 50, 10 - 75 or 0.5 - 50 µm.
[0171] In some cases, the density of clusters within the substrate is at least or about 1 cluster / 100 mm 2 , 1 cluster / 10 mm 2 , 1 cluster / 5 mm 2 , 1 cluster / 4 mm 2 , 1 cluster / 3 mm 2 , 1 cluster / 2 mm 2 , 1 cluster / 1 mm 2 , 2 clusters / 1 mm 2 , 3 clusters / 1 mm 2 , 4 clusters / 1 mm 2 , 5 clusters / 1 mm 2 , 10 clusters / 1 mm 2 , 50 clusters / 1 mm 2 or higher. In some cases, the substrate comprises about 1 cluster / 10 mm 2 to about 10 clusters / 1 mm 2. In some cases, the distance between the centers of two adjacent clusters is at least or about 50, 100, 200, 500, 1000, 2000 or 5000 um. In some cases, the distance between the centers of two adjacent clusters is about 50 - 100, 50 - 200, 50 - 300, 50 - 500 and 100 - 2000 um. In some cases, the distance between the centers of two adjacent clusters is about 0.05 - 50, 0.05 - 10, 0.05 - 5, 0.05 - 4, 0.05 - 3, 0.05 - 2, 0.1 - 10, 0.2 - 10, 0.3 - 10, 0.4 - 10, 0.5 - 10, 0.5 - 5 or 0.5 - 2 mm. In some cases, each cluster has a cross - section of about 0.5 to about 2, about 0.5 to about 1 or about 1 to about 2 mm. In some cases, each cluster has a cross - section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2 mm. In some cases, each cluster has an internal cross - section of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.15, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2 mm.
[0172] In some cases, the substrate is about the size of a standard 96 - well plate, e.g., about 100 to about 200 mm by about 50 to about 150 mm. In some cases, the substrate has a diameter less than or equal to about 1000, 500, 450, 400, 300, 250, 200, 150, 100 or 50 mm. In some cases, the diameter of the substrate is about 25 - 1000, 25 - 800, 25 - 600, 25 - 500, 25 - 400, 25 - 300 or 25 - 200 mm. In some cases, the substrate has a planar surface area of at least about 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 12,000, 15,000, 20,000, 30,000, 40,000, 50,000 mm 2 or greater. In some cases, the thickness of the substrate is about 50 - 2000, 50 - 1000, 100 - 1000, 200 - 1000 or 250 - 1000 mm.
[0173] Surface material
[0174] The substrates, devices, and reactors provided herein are made of any kind of material suitable for the methods, compositions, and systems described herein. In some cases, the substrate material is fabricated to exhibit a low level of nucleotide binding. In some cases, the substrate material is modified to generate a different surface that exhibits a high level of nucleotide binding. In some cases, the substrate material is transparent to visible light and / or ultraviolet light. In some cases, the substrate material has sufficient conductivity, e.g., capable of forming a uniform electric field across the entire substrate or a portion thereof. In some cases, the conductive material is electrically grounded. In some cases, the substrate is thermally conductive or thermally insulating. In some cases, the material is chemically resistant and heat resistant to support chemical or biochemical reactions, such as during polynucleotide synthesis reactions. In some cases, the substrate comprises a flexible material. For flexible materials, the material can include, but is not limited to: modified and unmodified nylon, nitrocellulose, polypropylene, etc. In some cases, the substrate comprises a rigid material. For rigid materials, the material can include, but is not limited to: glass; fused silica; silicon, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof, etc.); metals (e.g., gold, platinum, etc.). The substrate, solid support, or reactor can be made of a material selected from silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), and glass. The substrate / solid support or the microstructures therein, the reactor can be made using a combination of the materials listed herein or any other suitable materials known in the art.
[0175] Surface architecture
[0176] The present disclosure provides substrates for the methods, compositions, and systems described herein, wherein the substrates have a surface architecture suitable for the methods, compositions, and systems described herein. In some cases, the substrate comprises raised and / or recessed features. One benefit of having such features is an increased surface area for supporting polynucleotide synthesis. In some cases, a substrate having raised and / or recessed features is referred to as a three-dimensional substrate. In some cases, the three-dimensional substrate comprises one or more channels. In some cases, one or more seats comprise channels. In some cases, the channels can be used for reagent deposition by a deposition device such as a material deposition device. In some cases, reagents and / or fluids are collected in larger pores that are in fluid communication with one or more channels. For example, the substrate comprises a plurality of channels corresponding to a plurality of seats having clusters, and the plurality of channels are in fluid communication with a pore of the cluster. In some methods, a polynucleotide library is synthesized in the plurality of seats of the cluster.
[0177] The present disclosure provides a substrate for the methods, compositions, and systems described herein, wherein the substrate is configured for polynucleotide synthesis. In some cases, the structure is formulated to allow controlled flow and mass transfer pathways for polynucleotide synthesis on the surface. In some cases, the construction of the substrate allows for a controlled and uniform distribution of mass transfer pathways, number of chemical exposures, and / or washing efficacy during polynucleotide synthesis. In some cases, the construction of the substrate allows for increased scanning efficiency, e.g., by providing a volume sufficient for growing polynucleotides such that the volume excluded by the growing polynucleotides is no more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the initial available volume that is available or suitable for growing polynucleotides. In some cases, the three-dimensional structure allows for controlled flow of fluids, thereby allowing for rapid exchange of chemical exposures.
[0178] The present disclosure provides a substrate for the methods, compositions, and systems described herein, wherein the substrate comprises a structure suitable for the methods, compositions, and systems described herein. In some cases, isolation is achieved through a physical structure. In some cases, isolation is achieved by differential functionalization of the surface to generate activated and passivated regions for polynucleotide synthesis. In some cases, differential functionalization is achieved by alternating the presentation of hydrophobicity across the surface of the substrate, thereby creating a beading of the reagent that can cause deposition or a water contact angle effect that can cause wetting. Larger structures can reduce splashing and cross-contamination of different polynucleotide synthesis sites by reagents from adjacent spots. In some cases, a device such as a material deposition device is used to deposit reagents onto different polynucleotide synthesis sites. The substrate having three-dimensional features is configured in a manner that allows for the synthesis of a large number of polynucleotides (e.g., more than about 10,000) at a low error rate (e.g., less than about 1:500, 1:1000, 1:1500, 1:2,000; 1:3,000; 1:5,000; or 1:10,000). In some cases, the substrate comprises features having a density of about or greater than about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500 features / mm 2 of features.
[0179] The pores of the substrate may have the same or different widths, heights, and / or volumes as another pore of the substrate. The channels of the substrate may have the same or different widths, heights, and / or volumes as another channel of the substrate. In some cases, the diameter of the cluster or the diameter of the pore containing the cluster or both are about 0.05 - 50, 0.05 - 10, 0.05 - 5, 0.05 - 4, 0.05 - 3, 0.05 - 2, 0.05 - 1, 0.05 - 0.5, 0.05 - 0.1, 0.1 - 10, 0.2 - 10, 0.3 - 10, 0.4 - 10, 0.5 - 10, 0.5 - 5, or 0.5 - 2 mm. In some cases, the diameter of the cluster or the pore or both are less than or about 5, 4, 3, 2, 1, 0.5, 0.1, 0.09, 0.08, 0.07, 0.06, or 0.05 mm. In some cases, the diameter of the cluster or the pore or both are about 1.0 mm to 1.3 mm. In some cases, the diameter of the cluster or the pore or both are about 1.150 mm. In some cases, the diameter of the cluster or the pore or both are about 0.08 mm. The diameter of the cluster refers to the cluster within the two - dimensional or three - dimensional substrate.
[0180] In some cases, the height of the pore is about 20 - 1000, 50 - 1000, 100 - 1000, 200 - 1000, 300 - 1000, 400 - 1000, or 500 - 1000 μm. In some cases, the height of the pore is less than about 1000, 900, 800, 700, or 600 μm.
[0181] In some cases, the substrate contains a plurality of channels corresponding to a plurality of seats within the cluster, wherein the height or depth of the channel is 5 - 500, 5 - 400, 5 - 300, 5 - 200, 5 - 100, 5 - 50, or 10 - 50 μm. In some cases, the height of the channel is less than 100, 80, 60, 40, or 20 μm.
[0182] In some cases, the diameter of the channel, the seat (e.g., in a substantially flat substrate), or both the channel and the seat (e.g., in a three - dimensional substrate where the seat corresponds to the channel) is about 1 - 1000, 1 - 500, 1 - 200, 1 - 100, 5 - 100, or 10 - 100 μm, such as about 90, 80, 70, 60, 50, 40, 30, 20, or 10 μm. In some cases, the diameter of the channel, the seat, or both the channel and the seat is less than about 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 μm. In some cases, the distance between the centers of two adjacent channels, seats, or both the channel and the seat is about 1 - 500, 1 - 200, 1 - 100, 5 - 200, 5 - 100, 5 - 50, or 5 - 30, such as about 20 μm.
[0183] Surface modification
[0184] The present disclosure provides methods for synthesizing polynucleotides on a surface, where the surface comprises various surface modifications. In some cases, surface modifications are employed to chemically and / or physically alter the surface by an addition process or a subtraction process to change one or more chemical and / or physical properties of the substrate surface or selected sites or regions of the substrate surface. For example, surface modifications include, but are not limited to: (1) changing the wetting properties of the surface; (2) functionalizing the surface, i.e., providing, modifying, or replacing surface functional groups; (3) de-functionalizing the surface, i.e., removing surface functional groups; (4) otherwise changing the chemical composition of the surface, e.g., by etching; (5) increasing or decreasing surface roughness; (6) providing a coating on the surface, e.g., a coating exhibiting wetting properties different from those of the surface; and / or (7) depositing microparticles on the surface.
[0185] In some cases, adding a chemical layer (referred to as an adhesion promoter) on top of the surface facilitates the patterned structuring of sites on the substrate surface. Exemplary surfaces for applying the adhesion promoter include, but are not limited to, glass, silicon, silica, and silicon nitride. In some cases, the adhesion promoter is a chemical with a high surface energy. In some cases, a second chemical layer is deposited on the surface of the substrate. In some cases, the second chemical layer has a low surface energy. In some cases, the surface energy of the chemical layer coated on the surface supports the positioning of small droplets on the surface. Depending on the selected patterning arrangement, the proximity of the sites and / or the fluid contact area at the sites can be varied.
[0186] In some cases, the substrate surface or resolved sites onto which nucleic acids or other moieties are deposited (e.g., for polynucleotide synthesis) are smooth or substantially planar (e.g., two-dimensional), or have irregularities such as raised or recessed features (e.g., three-dimensional features). In some cases, the substrate surface is modified with one or more different compound layers. Such modification layers of interest include, but are not limited to, inorganic layers and organic layers such as metals, metal oxides, polymers, organic small molecules, and the like.
[0187] In some cases, one or more moieties that increase and / or decrease surface energy are used to functionalize the resolution sites of a substrate. In some cases, the moieties are chemically inert. In some cases, the moieties are configured to support desired chemical reactions, such as one or more processes in a polynucleotide synthesis reaction. The surface energy or hydrophobicity of the surface is a factor that determines the affinity of nucleotides to attach to the surface. In some cases, the substrate functionalization method includes: (a) providing a substrate having a surface comprising silica; and (b) silanizing the surface using a suitable silanizing agent described herein or known in the art (e.g., an organofunctional alkoxysilane molecule). The method and the silanizing agent are described in U.S. Patent 5,474,796, which is incorporated herein by reference in its entirety.
[0188] In some cases, the substrate surface is generally functionalized by contacting the substrate surface with a derivatization composition containing a silane mixture under reaction conditions effective to couple the silane to the substrate surface via reactive hydrophilic moieties present on the substrate surface. Silanization generally involves self-assembly using organofunctional alkoxysilane molecules to cover the surface. A variety of siloxane functionalization reagents currently known in the art can also be used, e.g., to reduce or increase surface energy. Organofunctional alkoxysilanes are classified according to their organofunctionality.
[0189] Polynucleotide synthesis
[0190] The methods of the present disclosure for polynucleotide synthesis can include processes involving phosphoramidite chemistry. In some cases, polynucleotide synthesis includes coupling a base with a phosphoramidite. Polynucleotide synthesis can include coupling a base by depositing a phosphoramidite under coupling conditions, wherein the same base is optionally deposited with the phosphoramidite more than once, i.e., double coupling. Polynucleotide synthesis can include capping of unreacted sites. In some cases, capping is optional. Polynucleotide synthesis can also include oxidation or an oxidation step or multiple oxidation steps. Polynucleotide synthesis can include deprotection, detritylation, and sulfurization. In some cases, polynucleotide synthesis includes oxidation or sulfurization. In some cases, during one step or between each step in the polynucleotide synthesis reaction, the device is washed, e.g., using tetrazole or acetonitrile. The time range for any step in the phosphoramidite synthesis method can be less than about 2 min, 1 min, 50 sec, 40 sec, 30 sec, 20 sec, and 10 sec.
[0191] Polynucleotide synthesis using phosphoramidite methods can include subsequently adding phosphoramidite building blocks (e.g., nucleoside phosphoramidites) to a growing polynucleotide chain to form phosphite triesters. Phosphoramidite polynucleotide synthesis proceeds in the 3' to 5' direction. Phosphoramidite polynucleotide synthesis allows for the controlled addition of one nucleotide per synthesis cycle to a growing nucleic acid chain. In some cases, each synthesis cycle includes a coupling step. Phosphoramidite coupling involves forming a phosphite triester bond between an activated nucleoside phosphoramidite and a nucleoside attached to a substrate (e.g., via a linker). In some cases, the nucleoside phosphoramidite is provided to an activation device. In some cases, the nucleoside phosphoramidite is provided to a device with an activator. In some cases, the nucleoside phosphoramidite is provided to the device in an excess of 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100-fold or more relative to the nucleoside attached to the substrate. In some cases, the addition of the nucleoside phosphoramidite is carried out in an anhydrous environment (e.g., in anhydrous acetonitrile). After adding the nucleoside phosphoramidite, the device is optionally washed. In some cases, the coupling step is repeated one or more additional times, optionally with a washing step between additions of the nucleoside phosphoramidite to the substrate. In some cases, the polynucleotide synthesis methods used herein include 1, 2, 3 or more consecutive coupling steps. In many cases, prior to coupling, the nucleoside attached to the device is deprotected by removing a protecting group that serves to prevent polymerization. A common protecting group is 4,4'-dimethoxytrityl (DMT).
[0192] After coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step. In the capping step, the growing polynucleotide is treated with a capping agent. The capping step can be used to block unreacted 5'-OH groups attached to the substrate after coupling to prevent further chain extension and thus prevent the formation of polynucleotides with internal base deletions. Additionally, phosphoramidites activated with 1H-tetrazole can react to a small extent with the O6 position of guanosine. Without being bound by theory, after oxidation with I2 / water, this byproduct (possibly via O6-N7 migration) can undergo depurination. The apurinic site can terminate cleavage during the final deprotection of the polynucleotide, thereby reducing the yield of the full-length product. The O6 modification can be removed by treating with a capping reagent prior to oxidation with I2 / water. In some cases, including a capping step during polynucleotide synthesis reduces the error rate compared to synthesis without capping. As an example, the capping step includes treating the polynucleotide attached to the substrate with a mixture of acetic anhydride and 1-methylimidazole. After the capping step, the device is optionally washed.
[0193] In some cases, after addition of the nucleoside phosphoramidite and optionally after capping and one or more wash steps, the growing nucleic acid attached to the device is oxidized. The oxidation step involves oxidizing the phosphite triester to a tetracoordinated phosphate triester - a protected precursor of the naturally occurring phosphodiester internucleoside linkage. In some cases, oxidation of the growing polynucleotide is achieved by treatment with iodine and water, optionally in the presence of a weak base (e.g., pyridine, lutidine, collidine). Oxidation can be carried out under anhydrous conditions using, for example, tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, a capping step is carried out after oxidation. The second capping step allows the device to dry, as residual water that may persist from the oxidation can inhibit subsequent coupling. After oxidation, the device and the growing polynucleotide are optionally washed. In some cases, the oxidation step is replaced by a sulfurization step to obtain a polynucleotide phosphorothioate, and any capping step can be carried out after sulfurization. Many reagents are capable of effecting efficient sulfur transfer, including but not limited to 3-(dimethylaminomethylene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide (also known as the Beaucage reagent), and N,N,N'N'-tetraethylthiuram disulfide (TETD).
[0194] In order for subsequent nucleoside incorporation cycles to occur by coupling, the protected 5' end of the growing polynucleotide attached to the device is removed such that the primary hydroxyl group can react with the next nucleoside phosphoramidite. In some cases, the protecting group is DMT and deblocking is carried out with trichloroacetic acid in dichloromethane. Extended detritylation times or the use of a stronger acid solution than the recommended acid solution for detritylation can result in increased depurination of the polynucleotide attached to the solid support and thus reduced yield of the desired full-length product. The methods and compositions of the present disclosure described herein provide controlled deblocking conditions that limit unwanted depurination reactions. In some cases, the polynucleotide attached to the device is washed after deblocking. In some cases, effective washing after deblocking contributes to the synthesis of polynucleotides with a low error rate.
[0195] Polynucleotide synthesis methods generally include a series of iterative steps of: applying a protected monomer to an activated functionalized surface (e.g., a seat) for attachment to the activated surface, linker, or to a pre-deprotected monomer; deprotecting the applied monomer such that it can react with a subsequently applied protected monomer; and applying another protected monomer for attachment. One or more intermediate steps include oxidation or sulfurization. In some cases, there are one or more wash steps before or after one or all of the steps.
[0196] Phosphoramidite-based polynucleotide synthesis methods include a series of chemical steps. In some cases, one or more steps of the synthesis method involve reagent cycling, where one or more steps of the method include applying a reagent useful for the step to the device. For example, the reagent is cycled through a series of liquid phase deposition and vacuum drying steps. For substrates containing three-dimensional features such as pores, micropores, channels, etc., the reagent optionally passes through one or more regions of the device via the pores and / or channels.
[0197] The methods and systems described herein relate to a polynucleotide synthesis device for synthesizing polynucleotides. The synthesis can be parallel. For example, at least or about at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 1000, 10000, 50000, 75000, 100000 or more polynucleotides can be synthesized in parallel. The total number of polynucleotides that can be synthesized in parallel can be 2 - 100000, 3 - 50000, 4 - 10000, 5 - 1000, 6 - 900, 7 - 850, 8 - 800, 9 - 750, 10 - 700, 11 - 650, 12 - 600, 13 - 550, 14 - 500, 15 - 450, 16 - 400, 17 - 350, 18 - 300, 19 - 250, 20 - 200, 21 - 150, 22 - 100, 23 - 50, 24 - 45, 25 - 40, 30 - 35. Those skilled in the art will appreciate that the total number of polynucleotides synthesized in parallel can be within any range defined by any of these values, such as 25 - 100. The total number of polynucleotides synthesized in parallel can be within any range defined by any value serving as the range endpoints. The total molar mass of the polynucleotides synthesized in the device or the molar mass of each polynucleotide can be at least or at least about 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles or greater. The length of each polynucleotide or the average length of the polynucleotides in the device can be at least or about at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500 or more nucleotides. The length of each polynucleotide or the average length of the polynucleotides in the device can be at most or about at most 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 or fewer nucleotides. The length of each polynucleotide or the average length of the polynucleotides in the device can be between 10 - 500, 9 - 400, 11 - 300, 12 - 200, 13 - 150, 14 - 100, 15 - 50, 16 - 45, 17 - 40, 18 - 35, 19 - 25.Those skilled in the art are aware that the length of each polynucleotide or the average length of polynucleotides within a device can be within any range defined by any of these values, such as 100 - 300. The length of each polynucleotide or the average length of polynucleotides within a device can be within any range defined by any value serving as a range endpoint.
[0198] The methods for synthesizing polynucleotides on a surface provided herein allow for synthesis at a relatively fast rate. As an example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 125, 150, 175, 200 or more nucleotides are synthesized per hour. Nucleotides include adenine, guanine, thymine, cytosine, uridine building blocks, or analogs / modified forms thereof. In some cases, polynucleotide libraries are synthesized in parallel on a substrate. For example, a device containing approximately or at least about 100, 1,000, 10,000, 30,000, 75,000, 100,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000 or 5,000,000 resolved sites is capable of supporting the synthesis of at least the same number of different polynucleotides, where polynucleotides encoding different sequences are synthesized at the resolved sites. In some cases, a polynucleotide library is synthesized on the device with a low error rate as described herein in less than about three months, two months, one month, three weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 24 hours or less. In some cases, larger nucleic acids assembled from polynucleotide libraries synthesized with a low error rate using the substrates and methods described herein are prepared in less than about three months, two months, one month, three weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 24 hours or less.
[0199] In some cases, the methods described herein provide for generating nucleic acid libraries containing variant nucleic acids that differ at multiple codon positions. In some cases, the nucleic acid can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50 or more variant codon positions.
[0200] In some cases, one or more of the variant codon sites can be adjacent. In some cases, one or more of the variant codon sites can be non - adjacent and separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more codons.
[0201] In some cases, a nucleic acid can contain multiple sites of variant codon sites, where all the variant codon sites are adjacent to each other, forming a stretch of variant codon sites. In some cases, a nucleic acid can contain multiple sites of variant codon sites, where the variant codon sites are not adjacent to each other. In some cases, a nucleic acid can contain multiple sites of variant codon sites, where some of the variant codon sites are adjacent to each other, forming a stretch of variant codon sites, while some of the variant codon sites are not adjacent to each other.
[0202] See the accompanying drawings, Figure 3 which shows an exemplary processing workflow for synthesizing nucleic acids (e.g., genes) from shorter nucleic acids. The workflow is generally divided into the following stages: (1) de novo synthesis of a single - stranded nucleic acid library, (2) ligation of nucleic acids to form larger fragments, (3) error correction, (4) quality control, and (5) transportation. Before de novo synthesis, an expected nucleic acid sequence or a set of nucleic acid sequences is pre - selected. For example, a set of genes is pre - selected for generation.
[0203] Once a large nucleic acid for generation is selected, a predefined nucleic acid library is designed for de novo synthesis. Various suitable methods for generating high - density polynucleotide arrays are known. In this workflow example, a surface layer of the device is provided. In this example, the chemical properties of the surface are altered to improve the polynucleotide synthesis process. Low - surface - energy regions are generated to repel liquids, while high - surface - energy regions are generated to attract liquids. The surface itself can be in the form of a planar surface or contain shape variations, such as protrusions or micropores that increase the surface area. In this workflow example, as disclosed in International Patent Application Publication WO / 2015 / 021080, which is incorporated herein by reference in its entirety, the selected high - surface - energy molecules perform a dual function of supporting DNA chemistry processes.
[0204] In - situ preparation of the polynucleotide array is carried out on a solid support and utilizes a single - nucleotide extension process to extend multiple oligomers in parallel. A deposition device, such as a material deposition device, is designed to release reagents in a step - by - step manner such that multiple polynucleotides are extended one residue at a time in parallel to generate oligomers 302 with a predefined nucleic acid sequence. In some cases, the polynucleotides are cleaved from the surface at this stage. Cleavage includes, for example, gas cleavage using ammonia or methylamine.
[0205] The generated polynucleotide library is placed in a reaction chamber. In this exemplary workflow, the reaction chamber (also referred to as a "nanoreactor") is a silicon-coated well that contains PCR reagents and is lowered onto the polynucleotide library 303. Before or after polynucleotide sealing 304, reagents are added to release the polynucleotides from the substrate. In this exemplary workflow, the polynucleotides are released after nanoreactor sealing 305. Once released, the fragments of the single-stranded polynucleotides hybridize to span the entire long-range DNA sequence. Partial hybridization 305 is possible because each synthesized polynucleotide is designed to have a small portion that overlaps with at least one other polynucleotide in the pool.
[0206] After hybridization, the PCA reaction is initiated. During the polymerase cycling process, the polynucleotides anneal to complementary fragments and the gaps are filled in with polymerase. Depending on which polynucleotides find each other, the length of each fragment is randomly increased in each cycle. Complementarity between the fragments allows for the formation of a complete large-span double-stranded DNA 306.
[0207] After PCA is completed, the nanoreactor is separated from the device 307 and positioned to interact with a device having PCR primers 308. After sealing, the nanoreactor undergoes PCR 309 and amplifies the larger nucleic acids. After PCR 310, the nanochamber is opened 311, error correction reagents are added 312, the chamber is sealed 313 and an error correction reaction is performed to remove mismatched base pairs and / or strands with poor complementarity from the double-stranded PCR amplification product 314. The nanoreactor is opened and separated 315. The error correction product then undergoes additional processing steps such as PCR and molecular barcoding, followed by packaging 322 for shipping 323.
[0208] In some cases, quality control measures are taken. After error correction, the quality control steps include, for example, interacting with a wafer having sequencing primers for amplifying the error correction product 316, sealing the wafer into a chamber containing the error correction amplification product 317, and performing another round of amplification 318. The nanoreactor is opened 319, the products are combined 320 and sequenced 321. After acceptable quality control results are obtained, the packaged product 322 is approved for shipping 323.
[0209] In some cases, the nucleic acids generated by workflows such as Figure 3 are mutagenized using the overlapping primers disclosed herein. In some cases, a primer library is generated by in situ preparation on a solid support and multiple oligomers are extended in parallel using a single nucleotide extension process. A deposition device such as a material deposition device is designed to release reagents in a stepwise manner such that multiple polynucleotides are extended one residue at a time in parallel to generate oligomers having a predetermined nucleic acid sequence 302.
[0210] Computer system
[0211] Any system described herein can be operably connected to a computer and can be automated locally or remotely via the computer. In various cases, the methods and systems of the present disclosure can further include software programs on a computer system and their use. Thus, computerized control for the synchronization of dispense / evacuate / refill functions, such as choreographing and synchronizing the movement of a material deposition device, dispense actions, and vacuum actuations, is within the scope of the present disclosure. The computer system can be programmed to interface between a user-specified base sequence and the position of the material deposition device to deliver the correct reagent to a specified area of the substrate.
[0212] Figure 4 The computer system 400 shown therein can be understood as a logical device capable of reading instructions from a medium 411 and / or a network port 405, which can optionally be connected to a server 409 having a fixed medium 412. Such as Figure 4 The system shown can include a CPU 401, a disk drive 403, optional input devices such as a keyboard 415 and / or a mouse 416, and an optional monitor 407. Data communication with a server at a local or remote location can be achieved through the shown communication medium. The communication medium can include any means for transmitting and / or receiving data. For example, the communication medium can be a network connection, a wireless connection, or an Internet connection. Such a connection can provide communication via the World Wide Web. It is contemplated that data related to the present disclosure can be transmitted over such a network or connection for reception and / or review by Figure 4 the user side 422 shown.
[0213] Figure 5 is a block diagram showing a first example architecture of a computer system 500 that can be used in conjunction with an example instance of the present disclosure. As Figure 5 shown, the example computer system can include a processor 502 for processing instructions. Non-limiting examples of processors include: Intel Xeon TM processors, AMD Opteron TM processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0 TM processors, ARM Cortex-A8 Samsung S5PC100TM processors, ARM Cortex-A8 Apple A4 TM processors, Marvell PXA930 TMA processor or a functionally equivalent processor. Multiple execution threads can be used for parallel processing. In some cases, multiple processors or processors with multiple cores can also be used, whether in a single computer system, in a cluster, or distributed across systems through a network including multiple computers, cellular phones, and / or personal data assistant devices.
[0214] As Figure 5 shown, the cache memory 504 can be connected to or incorporated into the processor 502 to provide a high-speed memory for instructions or data newly or frequently used by the processor 502. The processor 502 is connected to the north bridge 506 through the processor bus 508. The north bridge 506 is connected to the random access memory (RAM) 510 through the memory bus 512 and manages the access of the processor 502 to the RAM 510. The north bridge 506 is also connected to the south bridge 514 through the chipset bus 516. The south bridge 514 is in turn connected to the peripheral bus 518. The peripheral bus can be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The north bridge and the south bridge are generally referred to as the processor chipset and manage the data transfer between the processor, the RAM, and the peripheral components on the peripheral bus 518. In some alternative architectures, the functionality of the north bridge can be incorporated into the processor instead of using a separate north bridge chip. In some cases, the system 500 can include an accelerator card 522 attached to the peripheral bus 518. The accelerator can include a field programmable gate array (FPGA) or other hardware for accelerating a certain process. For example, the accelerator can be used for adaptive data reconstruction or for evaluating algebraic expressions used in extended set processing.
[0215] The software and data are stored in the external memory 524 and can be loaded into the RAM 510 and / or the cache memory 504 for use by the processor. The system 500 includes an operating system for managing system resources; non-limiting examples of the operating system include: Linux, Windows TM , MACOS TM , BlackBerry OS TM , iOS TM and other functionally equivalent operating systems, as well as application software running on top of the operating system for managing data storage and optimization according to the example scenarios of the present disclosure. In this instance, the system 500 also includes network interface cards (NICs) 520 and 521 connected to the peripheral bus to provide a network interface to external storage such as network attached storage (NAS) and other computer systems available for distributed parallel processing.
[0216] Figure 6FIG. 0 is a diagram showing a network 600 with multiple computer systems 602a and 602b, multiple cellular phones and personal data assistants 602c, and network attached storage (NAS) 604a and 604b. In an example instance, systems 602a, 602b, and 602c may manage data storage and optimize data access to data stored in network attached storage (NAS) 604a and 604b. Mathematical models may be used for the data and evaluated using distributed parallel processing across computer systems 602a and 602b, cellular phones, and personal data assistant systems 602c. Computer systems 602a and 602b, cellular phones, and personal data assistant systems 602c may also provide parallel processing for adaptive data reconstruction of data stored in network attached storage (NAS) 604a and 604b. Figure 6 Only one instance is shown, and a wide variety of other computer architectures and systems may be used with multiple instances of the present disclosure. For example, blade servers may be used to provide parallel processing. Processor blades may be connected via a backplane to provide parallel processing. Storage may also be connected to the backplane via a separate network interface or as network attached storage (NAS). In some example instances, processors may maintain separate storage spaces and transfer data via network interfaces, backplanes, or other connectors for parallel processing by other processors. In other cases, some or all of the processors may use a shared virtual address storage space.
[0217] Figure 7 FIG. 6 is a block diagram of a multiprocessor computer system 700 using a shared virtual address storage space according to an example scenario. The system includes multiple processors 702a-f that can access a shared memory subsystem 704. Multiple programmable hardware storage algorithm processors (MAP) 706a-f are incorporated in the memory subsystem 704 in the system. Each MAP 706a-f may include a memory 708a-f and one or more field programmable gate arrays (FPGA) 710a-f. The MAP provides configurable functional units and may provide specific algorithms or portions of algorithms to the FPGA 710a-f for processing in close cooperation with the corresponding processors. For example, in an example scenario, the MAP may be used to evaluate algebraic expressions related to data models and for adaptive data reconstruction. In this example, each MAP may be globally accessible to all processors for these purposes. In one configuration, each MAP may use direct memory access (DMA) to access the associated memory 708a-f, enabling it to perform tasks independently and asynchronously of their respective microprocessors 702a-f. In this configuration, the MAP may feed the results directly to another MAP for pipelining and parallel execution of algorithms.
[0218] The above computer architectures and systems are merely examples, and a wide variety of other computer, cellular phone, and personal data assistant architectures and systems can be used in conjunction with the example instances, including systems using any combination of general-purpose processors, coprocessors, FPGAs, and other programmable logic devices, system-on-chips (SOCs), application-specific integrated circuits (ASICs), and other processing and logic elements. In some cases, all or part of the computer system can be implemented in software or hardware. Any kind of data storage medium can be used in conjunction with the example instances, including random access memory, hard disk drives, flash memory, tape drives, disk arrays, network-attached storage (NAS), and other local or distributed data storage devices and systems.
[0219] In the example instances, the computer system can be implemented using software modules executed on any of the above or other computer architectures and systems. In other instances, the functionality of the system can be implemented in part or in whole in firmware, programmable logic devices such as Figure 5 the field programmable gate arrays (FPGAs), system-on-chips (SOCs), application-specific integrated circuits (ASICs), or other processing and logic elements mentioned. For example, the Set Processor and optimizer can be implemented in a hardware-accelerated manner by using a hardware accelerator card such as Figure 5 the accelerator card 522 shown.
[0220] The following embodiments are presented to more clearly illustrate the principles and practices of the disclosed embodiments to those skilled in the art and should not be construed as limiting the scope of any claimed embodiment. Unless otherwise indicated, all parts and percentages are by weight.
[0221] Embodiments
[0222] The following embodiments are given for the purpose of illustrating various embodiments of the present disclosure and are not intended to limit the invention in any way. These embodiments and the methods described herein, which currently represent the preferred embodiments, are exemplary and are not intended to limit the scope of the present disclosure. Variations and other uses within the spirit of the present disclosure as defined by the scope of the claims will occur to those skilled in the art.
[0223] Example 1: Functionalization of the Device Surface
[0224] Functionalize the device to support the attachment and synthesis of polynucleotide libraries. First, wet clean the device surface with a piranha solution containing 90% H2SO4 and 10% H2O2 for 20 minutes. Rinse the device in several beakers containing deionized water, hold it under a deionized water gooseneck stopcock for 5 min, and dry it with N2. Subsequently, soak the device in NH4OH (1:100; 3 mL:300 mL) for 5 min, rinse it with deionized water using a handgun, soak it in each of three consecutive beakers containing deionized water for 1 min, and then rinse it again with deionized water using a handgun. Then, plasma clean the device by exposing the device surface to O2. Perform O2 plasma etching at 250 watts for 1 min in downstream mode using a SAMCO PC-300 instrument.
[0225] Use a YES-1224P chemical vapor deposition oven system with the following parameters to activate-functionalize the clean device surface with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutyramide: 0.5 to 1 torr, 60 min, 70 °C, 135 °C vaporizer. Coat the device surface with a resist using a BrewerScience 200X spin coater. Spin coat SPR TM 3612 photoresist on the device at 2500 rpm for 40 sec. Pre-bake the device on a Brewer hot plate at 90 °C for 30 min. Perform photolithography on the device using a Karl Suss MA6 mask aligner. Expose the device for 2.2 sec and develop it in MSF 26A for 1 min. Rinse the remaining developer with a handgun and soak the device in water for 5 min. Bake the device in an oven at 100 °C for 30 min, and then visually inspect for photolithography defects using a Nikon L200. Use a descum process to remove the residual resist by performing O2 plasma etching at 250 watts for 1 min using a SAMCO PC-300 instrument.
[0226] Passivate-functionalize the device surface with a 100 μL solution of perfluorooctyltrichlorosilane mixed with 10 μL light mineral oil. Place the device in a chamber, pump for 10 min, then close the valve leading to the pump and let it stand for 10 min. Vent the chamber. Strip the resist from the device by soaking it twice for 5 min in 500 mL NMP at 70 °C while sonicating at maximum power (9 on a Crest system). Then soak the device in 500 mL isopropyl alcohol at room temperature for 5 min while sonicating at maximum power. Immerse the device in 300 mL of 200 proof ethanol and dry it with N2. Activate the functionalized surface to serve as a support for polynucleotide synthesis.
[0227] Example 2: Synthesis of a 50-mer sequence on an oligonucleotide synthesizer
[0228] A two-dimensional oligonucleotide synthesizer was assembled into a flow cell and connected to the flow cell (Applied Biosystems (ABI 394 DNA synthesizer")). The two-dimensional oligonucleotide synthesizer was uniformly functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide (Gelest) and used to synthesize an exemplary polynucleotide of 50 bp ("50-mer polynucleotide") using the polynucleotide synthesis method described herein.
[0229] The sequence of the 50-mer is as set forth in SEQ ID NO.:2. 5'AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3' (SEQ ID NO.:2), where # represents thymidine-succinylhexylamide CED phosphoramidite (CLP-2244 from ChemGenes), which is a cleavable linker that allows the release of the polynucleotide from the surface during deprotection.
[0230] Synthesis was completed using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking) according to the protocol in Table 3 and the ABI synthesizer.
[0231] Table 3: Synthesis protocol
[0232]
[0233]
[0234]
[0235] The phosphoramidite / activator combination was delivered in a manner similar to the delivery of bulk reagents through the flow cell. When the environment was kept "wet" with reagent at all times, no drying step was performed.
[0236] Remove the flow restrictor from the ABI 394 synthesizer to enable faster flow. In the absence of the flow restrictor, the flow rates of amidites (0.1 M in ACN), activator (0.25 M benzoylthiotetrazole (“BTT”; 30-3070-xx from GlenResearch) in ACN), and Ox (0.02 M I2 in 20% pyridine, 10% water, and 70% THF) are approximately about 100 uL / sec, the flow rates of acetonitrile (“ACN”) and capping reagents (a 1:1 mixture of cap A and cap B, where cap A is acetic anhydride in THF / pyridine and cap B is 16% 1-methylimidizole in THF) are approximately about 200 uL / sec, and the flow rate of the deblocking agent (3% dichloroacetic acid in toluene) is approximately about 300 uL / sec (in contrast, with the flow restrictor, the flow rate of all reagents is about 50 uL / sec). Observe the time for complete expulsion of the oxidizer, adjust the timing of the chemical flow times accordingly, and introduce additional ACN washes between different chemicals. After polynucleotide synthesis, deprotect the chip overnight in gaseous ammonia at 75 psi. Apply five drops of water to the surface to recover the polynucleotide. Then analyze the recovered polynucleotide on a BioAnalyzer small RNA chip.
[0237] Example 3: Synthesis of a 100-mer sequence on an oligonucleotide synthesizer
[0238] Using the same procedure described in Example 2 for the synthesis of the 50-mer sequence, synthesize a 100-mer polynucleotide (“100-mer polynucleotide”; 5'CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3', where # represents thymidine-succinylhexylamide CED phosphoramidite (CLP-2244 from ChemGenes); SEQ ID NO.:3) on two different silicon chips, the first functionalized uniformly with N-(3-triethoxysilylpropyl)-4-hydroxybutyramide, and the second functionalized with a 5 / 95 mixture of 11-acetoxydodecyltriethoxysilane and decyltriethoxysilane, and analyze the polynucleotide extracted from the surface on a BioAnalyzer instrument.
[0239] Using the following thermal cycling program, further PCR amplify all ten samples from the two chips in a 50 μL PCR mixture (25 μL NEB Q5 Master Mix, 2.5 μL 10 μM forward primer, 2.5 μL 10 μM reverse primer, 1 μL polynucleotide extracted from the surface, made up to 50 μL with water) using the forward primer (5'ATGCGGGGTTCTCATCATC3'; SEQ ID NO.: 4) and the reverse primer (5'CGGGATCCTTATCGTCATCG3'; SEQ ID NO.: 5):
[0240] 98 °C, 30 sec
[0241] 98 °C, 10 sec; 63 °C, 10 sec; 72 °C, 10 sec; repeat 12 cycles
[0242] 72 °C, 2 min
[0243] The PCR products were also run on a BioAnalyzer, showing sharp peaks at the 100-mer position. Then, the PCR-amplified samples were cloned and subjected to Sanger sequencing. Table 4 summarizes the Sanger sequencing results of the samples collected from spots 1-5 from chip 1 and spots 6-10 from chip 2.
[0244] Table 4: Sequencing results
[0245] Spot Error Rate Cycle Efficiency 1 1 / 763 bp 99.87% 2 1 / 824 bp 99.88% 3 1 / 780 bp 99.87% 4 1 / 429 bp 99.77% 5 1 / 1525 bp 99.93% 6 1 / 1615 bp 99.94% 7 1 / 531 bp 99.81% 8 1 / 1769 bp 99.94% 9 1 / 854 bp 99.88% 10 1 / 1451 bp 99.93%
[0246] Thus, the high quality and uniformity of the synthesized polynucleotides were reproduced on two chips with different surface chemistries. Overall, 89% of the 100-mers sequenced were perfect error-free sequences, corresponding to 233 out of 262.
[0247] Table 5 summarizes the error characteristics of the sequences obtained from the polynucleotide samples from spots 1-10.
[0248] Table 5: Error characteristics
[0249]
[0250]
[0251] Example 4: Design of the GLP1R binding domain based on peptide-ligand interactions
[0252] A GLP1R binding domain was designed based on the interaction surface between peptide ligands that interact with GLP1R. Motif variants were generated based on the interaction surface of the peptide with the ECD and the interaction surface with the N-terminal GLP1R ligand. This was accomplished using structural modeling. As shown in Table 6, exemplary motif variants were created based on the interaction of glucagon-like peptide with GLP1R. Motif variant sequences were generated using the following sequence from glucagon-like peptide: HAEGTFTSDVSSYLEGQAAKEFIAWLVKGRG (SEQ ID NO:6).
[0253] Table 6. Variant Amino Acid Sequences of Glucagon-Like Peptide
[0254]
[0255]
[0256] Example 5: Design of Antibody Scaffolds
[0257] To generate the scaffolds, structural analysis, heavy chain repertoire sequencing analysis, and specificity analysis of the heterodimer high-throughput sequencing dataset were performed. Each heavy chain was associated with each light chain scaffold. Five different long CDR-H3 loop options were assigned to each heavy chain scaffold. Five different L3 scaffolds were assigned to each light chain scaffold. The heavy chain CDR-H3 stem was selected from the frequently observed long H3 loop stem (10 amino acids at the N-terminus and C-terminus) found between the individual and V gene segments. The light chain scaffold L3 was selected from heterodimers containing long H3. Direct heterodimers based on information from the Protein Data Bank (PDB) and deep sequencing datasets were used, where CDR H1, H2, L1, L2, L3, and the CDR-H3 stem were fixed. Then the various scaffolds were formatted for display on phage to evaluate expression.
[0258] Structural Analysis
[0259] Approximately 2,017 antibody structures were analyzed, and 22 structures with CDR-H3 of at least 25 amino acid lengths were observed. The heavy chains included: IGHV1-69, IGHV3-30, IGHV4-49, and IGHV3-21. The identified light chains included: IGLV3-21, IGKV3-11, IGKV2-28, IGKV1-5, IGLV1-51, IGLV1-44, and IGKV1-13. In the analysis, four heterodimer combinations were observed multiple times, including: IGHV4-59 / 61-IGLV3-21, IGHV3-21-IGKV2-28, IGHV1-69-IGKV3-11, and IGHV1-69-IGKV1-5. Sequence and structural analysis determined internal disulfide bonds in CDR-H3 in some structures, and bulky side chains such as tyrosine were stacked in the stem, thus providing support for long-term H3 stability. Secondary structures including β-turn-β sheet and "hammerhead" subdomains were also observed.
[0260] Repertoire analysis
[0261] Repertoire analysis was performed on 1,083,875 IgM+ / CD27- naive B cell receptor (BCR) sequences and 1,433,011 CD27+ sequences obtained from 12 healthy controls by unbiased 5′ RACE. These 12 healthy controls included equal numbers of males and females and consisted of 4 white, 4 Asian, and 4 Hispanic individuals. Repertoire analysis showed that less than 1% of the human repertoire contains BCRs with CDR-H3 longer than 21 amino acids. V gene bias was observed in the long CDR3 subgroup repertoire, where IGHV1-69, IGHV4-34, IGHV1-18, and IGHV1-8 showed preferential enrichment in BCRs with long H3 loops. Bias away from long loops was observed for IGHV3-23, IGHV4-59 / 61, IGHV5-51, IGHV3-48, IGHV3-53 / 66, IGHV3-15, IGHV3-74, IGHV3-73, IGHV3-72, and IGHV2-70. The IGHV4-34 scaffold was shown to be autoreactive and have a short half-life.
[0262] Viable N-terminal and C-terminal CDR-H3 scaffold variants for long loops were also designed based on the 5′ RACE reference repertoire. Approximately 81,065 CDR-H3s with an amino acid length of 22 amino acids or longer were observed. By comparing in V gene scaffolds, scaffold-specific H3 stem variants were avoided, thus allowing scaffold diversity to be cloned into multiple scaffold references.
[0263] Heterodimer analysis
[0264] The scaffold was analyzed for heterodimers, and the variant sequences and lengths of the scaffold were determined.
[0265] Structural analysis
[0266] Structural analysis was performed on the GPCR scaffold using the variant sequences, and the lengths were determined.
[0267] Example 6: Generation of a GPCR antibody library
[0268] Based on the GPCR-ligand interaction surface and scaffold arrangement, the library was designed and de novo synthesized. See Example 4. Ten variant sequences were designed for the heavy chain variable domain, 237 variant sequences were designed for the heavy chain complementarity determining region 3, and 44 variant sequences were designed for the light chain variable domain. These fragments were synthesized into three fragments according to a method similar to that described in Examples 1-3.
[0269] After de novo synthesis, ten variant sequences were generated for the heavy chain variable domain, 236 variant sequences were generated for the heavy chain complementarity determining region 3, and 43 variant sequences were designed for the region containing the light chain variable domain and CDR-L3, of which nine variants were designed for the light chain variable domain. This produced a library with a diversity of approximately 10 5 (10x236x43). It was confirmed using next-generation sequencing (NGS) with 16 million reads. The normalized sequencing reads for each of the ten variants of the heavy chain variable domain were approximately 1 (data not shown). The normalized sequencing reads for each of the 43 variants of the light chain variable domain were approximately 1 (data not shown). The normalized sequencing reads for each of the 236 variant sequences of the heavy chain complementarity determining region 3 were approximately 1 (data not shown).
[0270] Then the expression and protein folding of various light and heavy chains were tested. The ten variant sequences for the heavy chain variable domain included: IGHV1-18, IGHV1-69, IGHV1-8, IGHV3-21, IGHV3-23, IGHV3-30 / 33rn, IGHV3-28, IGHV3-74, IGHV4-39, and IGHV4-59 / 61. Among these ten variant sequences, IGHV1-18, IGHV1-69, and IGHV3-30 / 33rn exhibited improved properties such as improved thermal stability. The nine variant sequences for the light chain variable domain included: IGKV1-39, IGKV1-9, IGKV2-28, IGKV3-11, IGKV3-15, IGKV3-20, IGKV4-1, IGLV1-51, and IGLV2-14. Among these nine variant sequences, IGKV1-39, IGKV3-15, IGLV1-51, and IGLV2-14 exhibited improved properties such as improved thermal stability.
[0271] Example 7: Expression of GPCR Antibody Library in HEK293 Cells
[0272] After generating the GPCR antibody library, approximately 47 GPCRs were selected for screening. GPCR constructs sized from approximately 1.8 kb to approximately 4.5 kb were designed in the pCDNA3.1 vector. Then, GPCR constructs were synthesized following a method similar to that described in Examples 2 - 4, including hierarchical assembly. Among these 47 GPCR constructs, 46 GPCR constructs were synthesized.
[0273] The synthesized GPCR constructs were transfected into HEK293, and their expression was detected using immunofluorescence. HEK293 cells were transfected with a GPCR construct containing an N - terminal hemagglutinin (HA) - tagged human Y1 receptor. After 24 - 48 hours of transfection, the cells were washed with phosphate - buffered saline (PBS) and fixed with 4% paraformaldehyde. The cells were stained with a fluorescent primary antibody against the HA tag or a secondary antibody containing a fluorophore and DAPI to render the cell nuclei blue. The human Y1 receptor was presented on the cell surface of non - permeabilized cells and on the cell surface and intracellularly in permeabilized cells.
[0274] The GPCR constructs were also visualized by designing GPCR constructs containing autofluorescent proteins. The human Y1 receptor contains EYFP fused to its C - terminus, while the human Y5 receptor contains ECFP fused to its C - terminus. HEK293 cells were transfected with the human Y1 receptor or co - transfected with the human Y1 receptor and the human Y5 receptor. After transfection, the cells were washed and fixed with 4% paraformaldehyde. The cells were stained with DAPI. The localization of the human Y1 receptor and the human Y5 receptor was visualized by fluorescence microscopy.
[0275] Example 8: Design of Immunoglobulin Library
[0276] An immunoglobulin scaffold library was designed for the placement of GPCR - binding domains and to enhance the stability of a series of GPCR - binding domain - encoding sequences. The immunoglobulin scaffold consists of a VH domain attached to a VL domain with a linker. Variant nucleic acid sequences were generated for the framework and CDR elements of the VH and VL domains. The designed structure is shown in Figure 8A and the full - domain architecture is shown in Figure 8B The sequences of the leader sequence, linker, and pIII are listed in Table 7.
[0277] Table 7. Nucleotide Sequences
[0278]
[0279] The designed VL domains include IGKV1-39, IGKV3-15, IGLV1-51 and IGLV2-14. Each of the four VL domains is assembled with its respective invariant four framework elements (FW1, FW2, FW3, FW4) and variable three CDR (L1, L2, L3) elements. For IGKV1-39, there are 490 variants designed for L1, 420 variants designed for L2, and 824 variants designed for L3, resulting in a diversity of 1.7x10 8 (490*420*824). For IGKV3-15, there are 490 variants designed for L1, 265 variants designed for L2, and 907 variants designed for L3, resulting in a diversity of 1.2x10 8 (490*265*907). For IGLV1-51, there are 184 variants designed for L1, 151 variants designed for L2, and 824 variants designed for L3, resulting in a diversity of 2.3x10 7 (184*151*824). For IGLV2-14, there are 967 variants designed for L1, 535 variants designed for L2, and 922 variants designed for L3, resulting in a diversity of 4.8x10 8 (967*535*922). Table 8 lists the amino acid sequences and nucleotide sequences of the four framework elements (FW1, FW2, FW3, FW4) of IGLV1-51. Table 9 lists the variable three CDR (L1, L2, L3) elements of IGLV1-51. Variant amino acid sequences and nucleotide sequences of the four framework elements (FW1, FW2, FW3, FW4) and variable three CDR (L1, L2, L3) elements are also designed for IGKV1-39, IGKV3-15 and IGLV2-14.
[0280] Table 8. Sequences of IGLV1-51 Framework Elements
[0281]
[0282] Table 9. Sequences of IGLV1-51 CDR Elements
[0283]
[0284]
[0285]
[0286]
[0287]
[0288]
[0289]
[0290]
[0291]
[0292]
[0293]
[0294]
[0295]
[0296]
[0297]
[0298]
[0299]
[0300]
[0301]
[0302]
[0303]
[0304]
[0305]
[0306]
[0307]
[0308]
[0309]
[0310]
[0311]
[0312]
[0313]
[0314]
[0315]
[0316] Pre-screen the CDRs to be free of amino acid liabilities, cryptic splice sites or nucleotide restriction sites. CDR variants are observed in at least two individuals and encompass a near-germline space of single, double, and triple mutations. The order of assembly is shown in Figure 8C as follows.
[0317] The designed VH domains include IGHV1-69 and IGHV3-30. Each of the two heavy-chain VH domains is assembled with its respective invariant four framework elements (FW1, FW2, FW3, FW4) and variable three CDR (H1, H2, H3) elements. For IGHV1-69, 417 variants were designed for H1 and 258 variants for H2. For IGHV3-30, 535 variants were designed for H1 and 165 variants for H2. For CDR H3, the same cassette is used in both IGHV1-69 and IGHV-30 because both are designed to use the same FW4, and the edges of FW3 are also the same for IGHV1-69 and IGHV3-30. CDR H3 contains N-terminal and C-terminal elements that combinatorially engage to a central intermediate element to generate 1×10 10 diversity. The N-terminal and intermediate elements overlap with the "GGG" glycine codon. The intermediate and C-terminal elements overlap with the "GGT" glycine codon. CDR H3 contains 5 separately assembled sub-pools. The individual N-terminal and C-terminal elements contain the sequences shown in Table 10.
[0318] Table 10. Sequences of N-terminal and C-terminal elements
[0319]
[0320]
[0321] Example 9. Enrichment of GPCR GLP1R-binding proteins
[0322] Antibodies having a CDR-H3 region comprising a GPCR-binding protein variant fragment are generated by the methods described herein and panned by a cell-based method as described in Example 4 to identify variants enriched for binding to a specific GPCR.
[0323] Variants of the GLP C-terminal peptide (listed in Table 11) were identified that were repeatedly and selectively enriched for binding to the GPCR GLP1R when embedded in the CDR-H3 region of an antibody.
[0324] Table 11. Sequences of GLP1 embedded in CDR-H3
[0325]
[0326] Example 10. Analysis of GLP1R-binding protein variants
[0327] Antibodies with a CDR-H3 region containing a variant fragment of a GLP1R-binding protein were generated by the methods described herein and panned by a cell-based method as described in Example 4 to identify variants enriched for binding to GLP1R.
[0328] Next-generation sequencing (NGS) enrichment of variants of the GLP1R peptide was performed (data not shown). Briefly, the phage population was deep-sequenced after each round of selection to track enrichment and identify cross-sample clones. After filtering out CHO background clones from the NGS data, target-specific clones were selected. For the GLP1R peptide, approximately 2000 VH and VL pairs were directly barcoded and sequenced from bacterial colonies to identify distinct clones.
[0329] The V gene distribution, J gene distribution, V gene families, and CDR3 counts for each length of the GLP1R-1 variant were analyzed. The frequencies of the V genes IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV3-53, IGHV3-NL1, IGHV3-d, IGHV1-46, IGHV3-h, IGHV1, IGHV3-38, IGHV3-48, IGHV1-18, IGHV1-3, and IGHV3-15 were determined (data not shown). High frequencies of IGHV1-69 and IGHV3-30 were observed. The frequencies of the J genes IGHJ3, IGHJ6, IGHJ, IGHJ4, IGHJ5, mIGHJ, IGHJ2, and IGH1 were also determined (data not shown). High frequencies of IGHJ3 and IGHJ6 were observed, while lower frequencies of IGHJ and IGHJ4 were observed. The frequencies of the V genes IGHV1-69, IGHV3-30, IGHV3-23, IGHV3, IGHV1-46, IGHV3-7, IGHV1, and IGHV1-8 were determined (data not shown). High frequencies of IGHV1-69 and IGHV3-30 were observed. The frequencies of the J genes IGHJ3, IGHJ6, IGHJ, IGHJ4, IGHJ5, IGHJ2, and IGH1 were determined (data not shown). High frequencies of IGHJ3 and IGHJ6 were observed, while lower frequencies of IGHJ and IGHJ4 were observed.
[0330] The H accumulation and frequencies of GLP1R-1, GLP1R-2, GLP1R-3, GLP1R-4, and GLP1R-5 were determined (data not shown).
[0331] Sequence analysis was performed on the GLP1R-1, GLP1R-2, GLP1R-3, GLP1R-4, and GLP1R-5 variants (data not shown).
[0332] Cell binding was determined for the GLP1R variants. Figures 9A - 9O GLP1R-2 ( Figure 9A )、GLP1R-3 ( Figure 9B )、GLP1R-8 ( Figure 9C )、GLP1R-26 ( Figure 9D )、GLP1R-30 ( Figure 9E )、GLP1R-56 ( Figure 9F )、GLP1R-58 ( Figure 9G )、GLP1R-10 ( Figure 9H )、GLP1R-25 ( Figure 9I )、GLP1R-60 ( Figure 9J )、GLP1R-70 ( Figure 9K)、GLP1R-72( Figure 9L )、GLP1R-83( Figure 9M )、GLP1R-93( Figure 9N ) and GLP1R-98( Figure 9O ) cell binding data.
[0333] Then, the allosteric effects of GLP1R-3, GLP1R-8, GLP1R-26, GLP1R-56, GLP1R-58, and GLP1R-10 on the GLP1-7-36 peptide were analyzed in a cAMP assay. Figures 10A - 10O Graphs showing the inhibition of GLP1-7-36 peptide-induced cAMP activity by GLP1R variants are presented. GLP1R-3( Figure 10B )、GLP1R-8( Figure 10C )、GLP1R-26( Figure 10D )、GLP1R-30( Figure 10E )、GLP1R-56( Figure 10F )、GLP1R-58( Figure 10G )、GLP1R-10( Figure 10H , right panel), GLP1R-25( Figure 10I ) and GLP1R-60( Figure 10J ) show allosteric inhibition of GLP1-7-36 peptide-induced cAMP activity. Figure 10H Further shown is the effect of GLPR-10 on cAMP signaling compared to exendin-4( Figure 10H , left panel).
[0334] GLP1R variants were tested in a cAMP assay to determine if these variants are antagonists that block exendin-4-induced cAMP activity. Figures 11A - 11G Cell function data for GLP1R-2( Figure 11A )、GLP1R-3( Figure 11B )、GLP1R-8( Figure 11C )、GLP1R-26( Figure 11D )、GLP1R-30( Figure 11E )、GLP1R-56( Figure 11F ) and GLP1R-58( Figure 11G ) are depicted.
[0335] Then, the allosteric effects of GLP1R-2, GLP1R-3, GLP1R-8, GLP1R-26, GLP1R-30, GLP1R-56, and GLP1R-58 on exendin-4 were analyzed in a cAMP assay. Figures 12A - 12G Cell function data for GLP1R-2( Figure 12A )、GLP1R-3(Figure 12B )、GLP1R-8( Figure 12C )、GLP1R-26( Figure 12D )、GLP1R-30( Figure 12E )、GLP1R-56( Figure 12F ) and GLP1R-58( Figure 12G ) Variant's illustration of inhibiting the cAMP activity induced by exendin-4 peptide. Table 12 shows the EC50 (nM) data of exendin-4 alone or together with GLP1R-2, GLP1R-3, GLP1R-8, GLP1R-26, GLP1R-30, GLP1R-56 and GLP1R-58.
[0336] Table 12. EC50 (nM) data
[0337]
[0338] FACS screening was performed on GLP1R variants. GLP1R-2, GLP1R-3, GLP1R-8, GLP1R-10, GLP1R-25, GLP1R-26, GLP1R-30, GLP1R-56, GLP1R-58, GLP1R-60, GLP1R-70, GLP1R-72, GLP1R-83, GLP1R-93 and GLP1R-98 were identified, as shown in Table 13. GLP1R-3, GLP1R-8, GLP1R-56, GLP1R-58, GLP1R-60, GLP1R-72 and GLP1R-83 contain the GLP1 motif. See Figure 13 . GLP1R-25, GLP1R-30, GLP1R-70, GLP1R-93 and GLP1R-98 contain the GLP2 motif. See Figure 13 . GLP1R-50 and GLP1R-71 contain the CC chemokine 28 motif.
[0339] Table 13. GLP1R variants
[0340]
[0341]
[0342] * Bold corresponds to the GLP1 or GLP2 motif.
[0343] The aggregation of GLP1R variants was evaluated. Size exclusion chromatography (SEC) was performed on GLP1R-30 and GLP1R-56 variants. 82.64% of GLP1R-30 was monomeric (about 150 Kd). 97.4% of GLP1R-56 was monomeric (about 150 Kd).
[0344] Example 11. Functionality of GPCR-Binding Proteins
[0345] For GPCR-binding proteins, the top 100 - 200 scFvs from phage selection were converted to full-length immunoglobulins. After immunoglobulin conversion, the clones were transiently transfected in ExpiCHO to produce immunoglobulins. Kingfisher and Hamilton were used for bulk IgG purification, followed by lab-on-a-chip to collect purity data for all purified immunoglobulins. As shown in Table 14, high yields and purities were obtained from 10 mL cultures.
[0346] Table 14. Immunoglobulin Purity Percentages
[0347] Name IgG% Purity mAb1 100 mAb2 100 mAb3 100 mAb4 100 mAb5 98 mAb6 100 mAb7 97 mAb8 100 100 100 100 100 100 100 100
[0348] Then, stable cell lines expressing the GPCR target were generated and confirmed by FACS (data not shown). Then, cells expressing >80% of the target were directly used for cell-based selection. Five rounds of selection were performed against cells overexpressing the target of interest. 10 8 cells were used per round of selection. Before selection of cells expressing the target, phage from each round was first depleted on 10 8 CHO background cells. The stringency of selection was increased by increasing the number of washes in subsequent selection rounds. The enrichment ratio for each round of selection was monitored.
[0349] The cell-binding affinity of the purified IgG was tested using FACS ( ) and cAMP activity ( ). Allosteric inhibition was observed.
[0350] The purified IgG was tested using BVP ELISA. BVP ELISA showed that some clones had BVP scores comparable to the comparator antibody (data not shown).
[0351] Example 12. GLP1R scFv Modulators
[0352] This example illustrates the identification of GLP1R modulators.
[0353] Library Panning
[0354] The GPCR1.0 / 2.0 scFv-phage library was incubated with CHO cells at room temperature (RT) for 1 h to deplete CHO cell binders. After incubation, the cells were pelleted by centrifugation at 1,200 rpm for 10 min to remove non-specific CHO cell binders. The phage supernatant depleted of CHO cell binders was then transferred to CHO cells expressing GLP1R. The phage supernatant and CHO cells expressing GLP1R were incubated at room temperature for 1 h to select for GLP1R binders. After incubation, non-binding clones were washed away by washing several times with PBS. Finally, to selectively elute agonist clones, the phage bound to GLP1R cells was competed off with 1 μM GLP1 peptide (residues 7 to 36), the native ligand of GLP1R. Clones eluted from the cells were likely to bind to the GLP1 ligand-binding epitope on GLP1R. The cells were pelleted by centrifugation at 1,200 rpm for 10 min to remove clones that still bound to GLP1R on the cells and did not bind to the endogenous GLP1 ligand-binding site (orthosteric site). The supernatant was amplified in TG1 E. coli cells for the next round of selection. This selection strategy was repeated for five rounds. The phage amplified from one round was used as the input phage for the next round, and the stringency of washing was increased in each subsequent round of selection. After five rounds of selection, 500 clones each from the 4th and 5th rounds were subjected to Sanger sequencing to identify clones of GLP1R modulators. Seven unique clones were reformatted into IgG2, purified, and tested for binding by FACS and functional assays.
[0355] Binding assay
[0356] Seven GLP1R scFv clones (GLP1R-238, GLP1R-239, GLP1R-240, GLP1R-241, GLP1R-242, GLP1R-243, and GLP1R-244) and two GLP1R IgGs used as controls (pGPCR-GLP1R-43 and pGPCR-GLP1R-44, Janssen Biotech, J&J) were tested in a binding assay coupled with flow cytometry analysis. The CDR3 sequences (Table 15), heavy chain sequences (Table 16), and light chain sequences (Table 17) of GLP1R-238, GLP1R-239, GLP1R-240, GLP1R-241, GLP1R-242, GLP1R-243, and GLP1R-244 are shown below.
[0357] Table 15. CDR3 sequences
[0358]
[0359] *Bold corresponds to GLP-1 or GLP-2 motifs.
[0360] Table 16. Variable heavy chain sequences
[0361]
[0362]
[0363]
[0364] Table 17. Variable light chain sequences
[0365]
[0366]
[0367] Briefly, CHO cells expressing flag-GLP1R-GFP (CHO-GLP1R) and CHO parental cells were incubated with 100 nM IgG on ice for 1 hour, washed three times, and then incubated with Alexa 647-conjugated goat-anti-human antibody (1:200) on ice for 30 minutes, followed by three washes. All incubations and washes were performed in a buffer containing PBS and 0.5% BSA. For titration, IgG was serially diluted 1:3 starting from 100 nM. Cells were analyzed by flow cytometry, and hits where IgG bound to CHO-GLP1R were identified by measuring GFP signal against Alexa647 signal. GLP1R-238, GLP1R-240, GLP1R-241, GLP1R-242, GLP1R-243, and GLP1R-244 were found to bind to CHO-GLP1R. GLP1R-238 bound equally to CHO-GLP1R and CHO parental cells and thus appeared to be a non-specific binder. Binding assay analysis using IgG titration presented a binding curve plotted as IgG concentration against MFI (mean fluorescence intensity), see .. Flow cytometry data from the binding assay using 100 nM IgG were presented as dot plots, see .
[0368] Functional assays
[0369] All GLP1R scFv clones, as well as pGPCR-GLP1R-43 and pGPCR-GLP1R-44, were analyzed for their potential impact on GLP1R signaling by performing cAMP assays (Eurofins DiscoverX Corporation). These assays involved CHO cells that were engineered to overexpress native Gα s-Coupled wild-type GLP1R and designed to detect changes in intracellular cAMP levels in response to agonist stimulation of the receptor. The technique involved in detecting cAMP levels is a wash-free signal amplification competitive immunoassay based on enzyme fragment complementation technology and generates a luminescence signal proportional to the amount of cAMP in the cells. The experiment aimed to determine the agonist or allosteric activity of IgG. To test the agonist activity of IgG, cells were stimulated with IgG (titrated 1:3 starting from 100 nM) or with the known agonist GLP1(7-36) peptide (titrated 1:6 starting from 12.5 nM) at 37 °C for 30 minutes. To test the allosteric activity of IgG, cells were incubated with a fixed concentration of 100 nM IgG at room temperature for 1 hour to allow binding, and then stimulated with GLP1(7-36) peptide (titrated 1:6 starting from 12.5 nM) at 37 °C for 30 minutes. Intracellular cAMP levels were detected according to the instructions of the detection kit.
[0370] As shown, none of the IgGs elicited an agonist signal. The allosteric cAMP effects of GLP1R-241 ( ), β-arrestin recruitment ( ), and internalization ( ) were also tested. By altering the signal response of these cells to GLP1(7-36) in an inhibitory manner, several IgGs acted as negative allosteric modulators, as shown. Table 18 shows the EC50 (nM) values corresponding to Figure 18A , and Table 19 shows the EC50 corresponding to Figure 18B .
[0371] Table 18. EC50 (nM) values
[0372]
[0373] Table 19. EC50 (nM) values
[0374]
[0375] The data showed the pharmacological and functional effects of GLP1R modulators.
[0376] Example 13: GLP1R Agonists and Antagonists
[0377] This example illustrates the identification of GLP1R agonists and antagonists.
[0378] Experiments were conducted similarly to Example 12. Binding and functional assays of six GLP1R immunoglobulins (IgG) were analyzed to determine which clones were agonists or antagonists. The GLP1R IgGs tested included GLP1R-59-2, GLP1R-59-241, GLP1R-59-243, GLP1R-3, GLP1R-241, and GLP1R-2. GLP1R-241, GLP1R-3, and GLP1R-2 were previously described in Examples 10 and 12. The heavy chain sequences of GLP1R-59-2, GLP1R-59-241, GLP1R-59-243, GLP1R-43-8, and GLP1R-3 are shown in Table 20.
[0379] Table 20. Variable heavy chain sequences
[0380]
[0381]
[0382] The thermal ramp stability (T m and T agg ) of GLP1R IgG was characterized. The IgG was characterized using the UNcle platform, and the data are shown in Table 21.
[0383] Table 21. Thermal ramp stability measurements
[0384]
[0385] Then, GLP1R IgG was assayed in a binding assay coupled with flow cytometry analysis using a similar method as described in Example 12. Briefly, CHO cells stably expressing Flag-GLP1R-GFP or CHO parental cells were incubated with the first IgG (100 nM or 1:3 titration). The secondary antibody incubation involved Alexa 647-conjugated goat anti-human IgG. Flow cytometry measured the GFP signal against the Alexa 647 signal to identify the IgG that specifically binds to the target (GLP1R). The ligand competition assay involved co-incubating the first IgG with 1 μM GLP1(7-36). The data for GLP1R-59-2, GLP1R-59-241, GLP1R-59-243, GLP1R-3, GLP1R-241, and GLP1R-2 are shown in Figures 19A - 19F .
[0386] Functional assays were also performed using GLP1R IgG using a similar method as described in Example 12. Briefly, cAMP, β-arrestin recruitment, and agonist internalization assays were obtained from Eurofins DiscoverX and utilized CHO-K1 or U2OS cells overexpressing unlabeled GLP-1R. These were used to test the agonist activity of the IgG compared to GLP1(7-36) or the antagonistic activity of the IgG by pre-incubating the cells with the IgG and examining their effect on GLP1(7-36)-induced signaling. For the cAMP assay, after stimulation with GLP1(7-36) or IgG, a homogeneous, wash-free, signal-amplified competitive immunoassay based on enzyme fragment complementation (EFC) technology was used to measure cellular cAMP levels. Data from the functional assays of GLP1R-59-2, GLP1R-59-241, GLP1R-59-243, GLP1R-3, GLP1R-241, and GLP1R-2 are shown in Figures 20A - 20F . EC50 (nM) data for GLP1R-59-2, GLP1R-59-241, GLP1R-59-243, GLP1R-3, GLP1R-241, and GLP1R-2 are shown in Tables 22-23. As seen in Table 23, the EC50 data for GLP1R-3 show a 2.2-fold difference. The EC50 data for GLP1-241 show a 1.7-fold difference. The EC50 data for GLP1R-2 show a 0.8-fold difference.
[0387] Table 22. EC50 (nM) for GLP1R-59-2, GLP1R-59-241, and GLP1R-59-243
[0388] GLP1R IgG EC50 GLP1(7 - 36) EC50 GLP1R - 59 - 2 0.842 0.4503 GLP1R - 59 - 241 0.7223 0.4731 GLP1R - 59 - 243 0.8209 0.4731
[0389] Table 23. EC50 (nM) for GLP1R-3, GLP1R-241, and GLP1R-2
[0390]
[0391] GLP1R-3 was also assayed to determine the specificity of GLP1R binding compared to GL1P2R and to determine that it is specific for GLP1R over GLP2R (data not shown). Binding of GLP1R-3, GLP1R-59-242, and GLP1R-43-8 to mouse, cynomolgus monkey, and human GLP1R was also assayed. It was found that 100 nM of GLP1R-3, 100 nM of GLP1R-59-242, and 100 nM of GLP1R-43-8 bind to mouse, cynomolgus monkey, and human GLP1R (data not shown). It was also found that GLP1R-3 binds to human pancreatic progenitor cells expressing endogenous GLP1R.
[0392] Binding of GLP1R-59-2, GLP1R-59-241, and GLP1R-59-243 to mouse, cynomolgus monkey, and human GLP1R was determined. It was found that 100 nM of GLP1R-59-2, 100 nM of GLP1R-59-241, and 50 nM of GLP1R-59-243 bound to mouse, cynomolgus monkey, and human GLP1R (data not shown). It was also found that GLP1R-59-2 bound to human pancreatic progenitor cells expressing endogenous GLP1R.
[0393] This example shows GLP1R IgGs with agonistic and antagonistic properties. Several IgGs induced cAMP signaling, β-arrestin recruitment, and receptor internalization, similar to GLP1(7-36).
[0394] Example 14: VHH Library
[0395] A synthetic VHH library was developed. For the "VHHRatio" library with customized CDR diversity, 2391 VHH sequences (iCAN database) were aligned using ClustalOmega to determine the consensus sequence at each position, and the framework was derived from the consensus sequence at each position. Position-specific variability analysis was performed on the CDRs of all 2391 sequences, and this diversity was introduced into the library design. For the "VHH Shuffle" library with shuffled CDR diversity, the iCAN database was scanned for unique CDRs in the nanobody sequences. 1239 unique CDR1s, 1600 unique CDR2s, and 1608 unique CDR3s were identified, and the framework was derived from the consensus sequence at each framework position in the 2391 sequences in the iCAN database. Each unique CDR was synthesized and shuffled individually in the consensus framework to generate a library with a theoretical diversity of 3.2x10^9. Then, the library was cloned into a phagemid vector using restriction enzyme digestion. For the 'VHHhShuffle' library (synthetic "human" VHH library with shuffled CDR diversity), the iCAN database was scanned for unique CDRs in the nanobody sequences. 1239 unique CDR1s, 1600 unique CDR2s, and 1608 unique CDR3s were identified, and frameworks 1, 3, and 4 were derived from the human germline DP-47 framework. Framework 2 was derived from the consensus sequence at each framework position in the 2391 sequences in the iCAN database. Each unique CDR was synthesized and shuffled individually in the partially humanized framework using the NUGE tool to generate a library with a theoretical diversity of 3.2x10^9. Then, the library was cloned into a phagemid vector using the NUGE tool.
[0396] The binding affinity and affinity distribution of VHH-Fc variants were evaluated using the Carterra SPR system. The VHH-Fc showed a range of affinities for TIGIT, with a lower limit of 12 nM K D , and an upper limit of 1685 nM K D (data not shown). Table 23A provides the specific values of VHH-Fc clones for ELISA, Protein A (mg / ml), and K D (nM). Figure 21A and Figure 21B The TIGIT affinity distribution of the VHH library was depicted at 20 - 4000 affinity thresholds ( Figure 21A ; monovalent K D ) and 20 - 1000 affinity thresholds ( Figure 21B ; monovalent K D ). Among 140 tested VHH binders, 51 variants had an affinity < 100 nM and 90 variants had an affinity < 200 nM.
[0397] Table 23A. ELISA, Protein A, and K of VHH-Fc clones D
[0398]
[0399]
[0400] Example 15: VHH Library of GLP1R
[0401] A VHH library of GLP1R was developed by a method similar to that described in Example 14. Briefly, a stable cell line expressing GLP1R was generated, and target expression was confirmed by FACS. Then, cells expressing > 80% of the target were used for cell-based selection. Five rounds of cell-based selection were performed against cells stably overexpressing the target of interest. 10 8 cells were used in each round of selection. Before selection of cells expressing the target, phage from each round was depleted on 10 8 CHO background cells. The stringency of selection was increased by increasing the number of washes in subsequent rounds of selection. Then, cells were eluted from the phage using trypsin, and the phage was amplified for the next round of panning. A total of 1000 clones from rounds 4 and 5 were sequenced by NGS to identify unique clones for reformulation into VHH-Fc.
[0402] Among 156 unique GLP1R VHH Fc binders, 53 had a mean fluorescence intensity (MFI) value of the target cells that was 2-fold that of the parental cells. Data for variant GLP1R-43-77 are shown in Figures 22A - 22B and Table 23B - 24.
[0403] Table 23B. Overview of panning
[0404]
[0405] Table 24. GLP1R-43-77 data
[0406] Subset name with gated path Count Median: RL1 - A Sample E10.fcs / CHO - parental 11261 237 Sample E10.fcs / CHO - GLP1R 13684 23439
[0407] Example 16. GLP1R library with different CDRs
[0408] Create a GLP1R library using a CDR randomization scheme.
[0409] Briefly, design a GLP1R library based on GPCR antibody sequences. Over 60 different GPCR antibodies were analyzed, and the sequences from these GPCRs were modified using a CDR randomization scheme.
[0410] The heavy chain IGHV3-23 design is shown in Figure 23A . As Figure 23A shown, IGHV3-23 CDRH3 has four different lengths: 23 amino acids, 21 amino acids, 17 amino acids, and 12 amino acids, each with its residue diversity. The proportions of the four lengths are as follows: 40% for the 23-amino-acid length of CDRH3, 30% for the 21-amino-acid length of CDRH3, 20% for the 17-amino-acid length of CDRH3, and 10% for the 12-amino-acid length of CDRH3. The CDRH3 diversity was determined to be 9.3×10 8 , and the total heavy chain IGHV3-23 diversity was 1.9x10 13 .
[0411] The heavy chain IGHV1-69 design is shown in Figure 23B . As Figure 23B shown, IGHV1-69 CDRH3 has four different lengths: 20 amino acids, 16 amino acids, 15 amino acids, and 12 amino acids, each with its residue diversity. The proportions of the four lengths are as follows: 40% for the 20-amino-acid length of CDRH3, 30% for the 16-amino-acid length of CDRH3, 20% for the 15-amino-acid length of CDRH3, and 10% for the 12-amino-acid length of CDRH3. The CDRH3 diversity was determined to be 9×10 7 , and the total heavy chain IGHV-69 diversity was 4.1x10 12 .
[0412] The light chain IGKV 2-28 and IGLV 1-51 designs are shown in Figure 23CAnalyze the position-specific variations in the CDR sequences of antibody light chains. Two light chain frameworks with fixed CDR lengths were selected. The theoretical diversities of the κ chain and the light chain were determined to be 13,800 and 5,180, respectively.
[0413] The final theoretical diversity was determined to be 4.7x10 17 , and the diversity of the finally generated Fab library was 6x10 9 . See Figure 23D .
[0414] Purified GLP1R IgG was assayed to determine cell-based affinity measurements and functional analysis. FACS binding was measured using purified GLP1R IgG. As Figure 23E shown, GLP1R IgG selectively binds to cells expressing GLP1R with an affinity in the low nanomolar range, demonstrating that IgG selectively binds to cells expressing the target with an affinity of 1.1 nM. FACS binding was also measured in GLP1R IgG generated using the methods described in Examples 4-10. As Figure 23F shown, GLP1R IgG selectively binds to cells expressing GLP1R with an affinity in the low nanomolar range.
[0415] The cAMP assay using purified GLP1R IgG demonstrated that the presence of GLP1R IgG caused a left shift in the dose-response curve of GLP1 agonist-induced cAMP response in CHO cells expressing GLP1R, as Figure 23G shown. GLP1R IgG generated using the methods described in Examples 4-10 also caused a left shift in the dose-response curve of receptor agonist-induced cAMP response in CHO cells expressing GLP1R ( Figure 23H ).
[0416] Data showed the design and generation of GLP1R IgG with improved potency and function.
[0417] Example 17. Oral Glucose Tolerance Mouse Model
[0418] The aim of this study was to evaluate the acute effects of chimeric antibody GLP1R agonists and antagonists on blood glucose control in a diet-induced obesity mouse model in C57BL / 6J DIO mice. The test articles are shown in Table 25 below.
[0419] Table 25. Test Article Identity
[0420]
[0421] - = Not applicable.
[0422] For each test article, seven different groups of test articles were generated, as summarized in Table 26 below, with eight animals in each group.
[0423] Table 26. Experimental Design
[0424]
[0425] No. = number; HFD = high-fat diet; QD = once daily; SC = subcutaneous injection
[0426] On Day 3 (for all animals) and Day 1 (for Groups 1 - 7), non-fasting blood glucose was measured by tail snip. Approximately 5 - 10 μL of blood was collected. The second drop of blood from the animal was placed on a blood glucose test strip and analyzed using a handheld glucometer (Abbott Alpha Trak).
[0427] After non-fasting blood glucose measurement on the day of the procedure, the animals were weighed, their tails were marked, and the animals were placed in clean cages without food. The animals were fasted for 4 hours and fasting blood glucose was measured. Then the animals were treated with the designated test article as shown in Table 26.
[0428] Sixty minutes later, an oral glucose tolerance test (OGTT) was performed on each animal. The animals were administered 2 g / kg glucose (10 mL / kg) by oral gavage. At the following times relative to glucose administration, the second drop of blood from the animal was placed on a handheld glucometer (Abbott Alpha Trak) and blood glucose was measured by tail snip: 0 (just before glucose administration), 15, 30, 60, 90, and 120 minutes. Additional blood samples were obtained at the 15 - minute and 60 - minute time points of the OGTT for estimation of serum insulin.
[0429] Figures 24A - 24B Shown is that GLP1R - 3 inhibits GLP1:GLP1R signaling ( Figure 24A ), and completely inhibits at higher concentrations ( Figure 24B ). As Figure 24C shown, animals administered GLP1R - 3 maintained a sustained high glucose level after glucose administration, indicating GLP1:GLP1R signaling blockade. As Figure 24D shown, a dose of 10 mg / kg of GLP1R - 59 - 2 exhibited a sustained low glucose level similar to the liraglutide control.
[0430] Data show that the generated GLP1R antibodies have a functional impact on glucose tolerance in a mouse model.
[0431] Example 18. Effects of GLP1R Agonists and Antagonists in Wild - type Mice
[0432] In this example, the effects of GLP1R-59-2 (agonist) and GLP1R-3 (antagonist) in wild-type mice were determined.
[0433] Fifteen C57BL / 6NHsd mice were used and a glucose tolerance test (GTT) was performed. An in vivo GTT was performed on three groups of mice, with 5 mice in each group. All three groups were fasted for 13.5 hours, then weighed, and their blood glucose at time zero was measured. Then, a 30% glucose solution was injected intraperitoneally at a dose of 10 μL / g body weight. Blood glucose measurements were recorded for each mouse at 15, 30, 60, 120, and 180 minutes after glucose injection. The first group of mice was treated with two doses of GLP1R-59-2: 10 mg / kg of GLP1R-59-2 at fasting (13.5 hours before GTT), and again 10 mg / kg of GLP1R-59-2 two hours before the start of GTT. The second group of mice was treated with two doses of GLP1R-3: 10 mg / kg of GLP1R-3 at fasting (13.5 hours before GTT), and again 10 mg / kg of GLP1R-3 two hours before the start of GTT. The third group of mice was the control group and was not treated. Data for GLP1R-59-2 (agonist), GLP1R-3 (antagonist), and the control are shown in Figures 25A - 25D . Figure 25A Shows the blood glucose levels (y-axis) of mice treated with GLP1R-59-2 (agonist), GLP1R-3 (antagonist), and the control over time (in minutes, x-axis). Figure 25B Shows the blood glucose levels (y-axis) of mice treated with GLP1R-59-2 (agonist), GLP1R-3 (antagonist), and the control. As Figure 25C shown, a significant decrease in blood glucose was observed in mice treated with GLP1R-59-2 (agonist) compared to the control, in both fasting (p = 0.0008) and non-fasting (p < 0.0001) mice. As Figure 25D shown, animals pre-administered with GLP1R-3 (antagonist) did not show a decrease in blood glucose within 6 hours of fasting, while the control mice showed a decrease.
[0434] Example 19. Exemplary Sequences
[0435] Exemplary sequences of GLP1R are shown in Table 27.
[0436] Table 27. GLP1R Sequences
[0437]
[0438]
[0439]
[0440]
[0441]
[0442]
[0443]
[0444]
[0445]
[0446]
[0447]
[0448]
[0449]
[0450]
[0451]
[0452]
[0453]
[0454] Although the preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Many variations, changes, and substitutions will occur to those skilled in the art without departing from the content of the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. It is intended that the scope of the present disclosure be defined by the appended claims and that the methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. An antibody or antibody fragment that binds to GLP1R, wherein the antibody or antibody fragment comprises a heavy chain and a light chain, wherein: (a) the heavy chain consists of the amino acid sequence of residues 20 to 504 of SEQ ID NO: 2305, and the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2312; (b) the heavy chain consists of the amino acid sequence of residues 20 to 507 of SEQ ID NO: 2306, and the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2313; (c) the heavy chain consists of the amino acid sequence of residues 20 to 504 of SEQ ID NO: 2307, and the light chain consists of the amino acid sequence of residues 21 to 235 of SEQ ID NO: 2314; or (d) the heavy chain consists of the amino acid sequence of residues 20 to 494 of SEQ ID NO: 2309, and the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2316.
2. The antibody or antibody fragment according to claim 1, wherein the heavy chain consists of the amino acid sequence of residues 20 to 504 of SEQ ID NO: 2305; and wherein the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2312.
3. The antibody or antibody fragment according to claim 1, wherein the heavy chain consists of the amino acid sequence of residues 20 to 507 of SEQ ID NO: 2306; and wherein the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2313.
4. The antibody or antibody fragment according to claim 1, wherein the heavy chain consists of the amino acid sequence of residues 20 to 504 of SEQ ID NO: 2307; and wherein the light chain consists of the amino acid sequence of residues 21 to 235 of SEQ ID NO: 2314.
5. The antibody or antibody fragment according to claim 1, wherein the heavy chain consists of the amino acid sequence of residues 20 to 494 of SEQ ID NO: 2309; and wherein the light chain consists of the amino acid sequence of residues 21 to 236 of SEQ ID NO: 2316.
6. The antibody or antibody fragment according to any one of claims 1-5, wherein the antibody is a monoclonal antibody, bispecific antibody, multispecific antibody, transplant antibody, human antibody, humanized antibody, synthetic antibody or chimeric antibody.
7. The antibody or antibody fragment according to any one of claims 1-5, wherein the antibody or its antibody fragment is chimeric or humanized.
8. The antibody or antibody fragment according to any one of claims 1-5, wherein the antibody has an EC50 in a cAMP assay of less than 25 nanomoles.
9. The antibody or antibody fragment according to any one of claims 1-5, wherein the antibody has an EC50 in a cAMP assay of less than 20 nanomoles.
10. The antibody or antibody fragment according to any one of claims 1-5, wherein the EC50 of the antibody in the cAMP assay is less than 10 nanomoles.
11. Use of the antibody or antibody fragment according to any one of claims 1-10 in the preparation of a pharmaceutical composition for treating hypoglycemia or hyperinsulinemia.
Citation Information
Patent Citations
Method and apparatus for conducting an array of chemical reactions on a support surface
US5474796A
De novo synthesized gene libraries
WO2015021080A2
Method for engineering immunoglobulins
WO2008003116A2
Anti-GLP-1r antibodies and their uses
WO2011056644A2